An optimization method and system based on spectral data

CN122778298APending Publication Date: 2026-09-18SUZHOU PLATING LIANGHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610938959.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]然而,在实际测量环境中,光谱仪的单次采集极易受到仪器噪声、环境干扰或样品偶然因素的影响,从而产生基线偏移、特征峰畸变或异常尖峰等严重偏离样品真实光谱的异常数据,也就是说,由于算术平均是赋予所有光谱完全相同的权重,当采集的多次光谱中混入一条上述异常数据,那么该异常数据就会与正常数据一同被纳入平均计算,从而污染最终的标准光谱,造成特征峰变形、信噪比提升受限甚至产生系统性预测偏差,进而导致最后生成的光谱失真且鲁棒性差

Benefits of technology

1.本发明通过对原始光谱数据进行校正、特征提取及异常评分,并依据评分结果为各校正光谱数据分配差异化权重后加权融合,能够有效剔除或抑制异常光谱的干扰,提升目标优化光谱的信噪比与保真度,解决了现有算术平均法易被异常数据污染导致光谱失真的问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778298A_ABST
    Figure CN122778298A_ABST
Patent Text Reader

Abstract

The application provides an optimization method and system based on spectral data, and relates to the technical field of spectral data processing, and comprises the following steps: collecting a plurality of original spectral data of a to-be-tested sample; performing data correction on the plurality of original spectral data to obtain a plurality of corrected spectral data; performing spectral feature extraction on each of the corrected spectral data respectively, and calculating an abnormal score based on the extracted spectral features to obtain an abnormal score result corresponding to each of the corrected spectral data; determining a weight coefficient corresponding to each of the corrected spectral data based on the abnormal score result; performing spectral optimization processing on each of the corrected spectral data based on the weight coefficient to obtain a target optimized spectrum of the to-be-tested sample; and solving the problem that the prior art cannot eliminate significant abnormal spectra, resulting in distortion of the final average spectrum and poor robustness, and having the effect of improving the accuracy and robustness of spectral reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral data processing technology, and in particular to an optimization method and system based on spectral data. Background Technology

[0002] Spectroscopic analysis technology, due to its advantages such as speed, non-destructive nature, and rich information content, has been widely used in fields such as industrial process control, environmental monitoring, biomedical diagnostics, and food safety testing. In these applications, the accuracy and reliability of the analytical results directly determine the correctness of subsequent decisions. Therefore, spectral data is required to have a high signal-to-noise ratio, high fidelity, and good repeatability to ensure the credibility of qualitative and quantitative analysis results.

[0003] To suppress random noise and improve measurement reliability, the current mainstream approach is to repeatedly acquire spectra of the same sample, directly perform an arithmetic average of the acquired spectra, and use the averaged spectrum as a standard spectrum representing the true physicochemical information of the sample for subsequent modeling or discriminant analysis.

[0004] However, in actual measurement environments, a single acquisition by a spectrometer is highly susceptible to instrument noise, environmental interference, or accidental factors related to the sample, resulting in abnormal data that deviates significantly from the true spectrum of the sample, such as baseline shift, characteristic peak distortion, or unusual spikes. In other words, since the arithmetic mean assigns the same weight to all spectra, when one such abnormal data point is mixed into multiple acquired spectra, it will be included in the averaging calculation along with the normal data, thus contaminating the final standard spectrum. This can cause characteristic peak distortion, limited signal-to-noise ratio improvement, or even systematic prediction bias, ultimately leading to distorted and poorly robust spectra. Summary of the Invention

[0005] In view of the above-mentioned shortcomings in the existing technology, the purpose of this invention is to provide an optimization method based on spectral data, which has the characteristics of improving the accuracy and robustness of spectral reconstruction.

[0006] The above-mentioned objective of this invention is achieved through the following technical solution: An optimization method based on spectral data includes: Collect multiple raw spectral data of the sample to be tested; Data correction is performed on the multiple original spectral data to obtain multiple corrected spectral data; Spectral features are extracted from each of the corrected spectral data, and anomaly scores are calculated based on the extracted spectral features to obtain the anomaly score results corresponding to each of the corrected spectral data. Based on the anomaly scoring results, determine the weighting coefficients corresponding to each of the corrected spectral data. Based on the weighting coefficients, the calibration spectral data are subjected to spectral optimization processing to obtain the target optimized spectrum of the sample to be tested.

[0007] By adopting the above technical solution, the original spectral data is first corrected to eliminate systematic bias, and each spectrum is quantitatively scored based on the extracted spectral features to accurately identify its degree of abnormality. Then, a reasonable weight coefficient is assigned to each corrected spectrum according to the abnormality score, so that the high-quality spectrum dominates the fusion process, while the contribution of abnormal spectra is effectively suppressed. The final generated target optimized spectrum not only improves the signal-to-noise ratio, but also effectively avoids spectral distortion caused by the mixing of abnormal data, which greatly improves the accuracy and robustness of spectral reconstruction.

[0008] Preferably, the step of performing data correction on the plurality of original spectral data to obtain plurality of corrected spectral data includes: The original spectral data is smoothed and denoised to obtain denoised spectral data; Identify the characteristic peaks in the denoised spectral data and match them with the standard characteristic peaks in the preset standard spectrum to determine the correspondence; Based on the correspondence, the wavelength shift of the characteristic peak relative to the standard characteristic peak is calculated; Based on the wavelength offset, wavelength axis correction is performed on the denoised spectral data to obtain corrected spectral data.

[0009] By adopting the above technical solution, the random noise and spur interference in the original spectrum are effectively suppressed by smoothing and denoising, which improves the identification accuracy of characteristic peaks. Then, the wavelength offset is calculated based on the characteristic peak matching and axis correction is performed, which can systematically eliminate the wavelength drift error caused by instrument fluctuations or environmental changes. The corrected spectral data obtained thus have good consistency at different wavelength positions.

[0010] Preferably, the spectral features include global baseline deviation and local noise level; the step of extracting spectral features from each of the corrected spectral data and calculating anomaly scores based on the extracted spectral features to obtain anomaly score results corresponding to each of the corrected spectral data includes: Calculate the average intensity value of each of the corrected spectral data in the preset background band, and use it as the individual background intensity of each of the corrected spectral data. Calculate the median of the average intensity values ​​of all corrected spectral data within the background band, and use it as the overall background intensity reference. The difference between the individual background intensity of each of the corrected spectral data and the overall background intensity benchmark is calculated to obtain the global baseline deviation of each of the corrected spectral data. Within a preset noise assessment band, the intensity difference between adjacent wavelength points in each of the corrected spectral data is calculated sequentially to form a difference sequence, and the standard deviation of the difference sequence is calculated to obtain the local noise level of each of the corrected spectral data. The global baseline deviation and the local noise level are fused together to obtain the anomaly score results corresponding to each of the corrected spectral data.

[0011] By adopting the above technical solution, the data quality of each spectrum is comprehensively evaluated from two dimensions: global baseline deviation and local noise level. The global baseline deviation reflects the overall baseline drift of the spectrum, while the local noise level reflects the detail fluctuations and random interference of the spectrum. The anomaly score obtained by fusing the two can comprehensively and quantitatively characterize the degree of anomaly of each corrected spectrum relative to the normal spectrum.

[0012] Preferably, determining the weighting coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results includes: The anomaly score results corresponding to each of the corrected spectral data are compared with the preset anomaly score threshold. If the anomaly score result is greater than the anomaly score threshold, the weight coefficient of the corresponding corrected spectral data is set to zero. If the abnormal score result is less than or equal to the abnormal score threshold, then the weight coefficient corresponding to the corrected spectral data is determined based on the degree of deviation between the abnormal score result and the preset benchmark score, wherein the smaller the abnormal score result, the larger the weight coefficient corresponding to the corrected spectral data.

[0013] By adopting the above technical solution, the corrected spectral data is screened using anomaly scoring thresholds. The weight of spectra with excessively high anomaly scores is directly set to zero, thus completely eliminating severely anomalous spectra and avoiding their contamination of subsequent spectral fusion. At the same time, for the selected spectra, differentiated weights are assigned according to the degree of deviation between their anomaly scores and benchmark scores, so that the higher the quality of the spectra, the greater their contribution to the fusion. This achieves a reasonable balance between eliminating extreme anomalies and retaining valid data.

[0014] Preferably, determining the weighting coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results further includes: The first intermediate value is obtained by calculating the ratio of the square of the abnormal score result to the square of twice the preset attenuation adjustment parameter. Calculate the negative exponent of the first intermediate value to obtain the weighting coefficient of the corrected spectral data, wherein the attenuation adjustment parameter is used to control the attenuation rate of the weighting coefficient as the abnormal scoring result changes.

[0015] By adopting the above technical solution, the weight coefficient is calculated using a continuous function based on negative exponents, which makes the weight decrease smoothly as the abnormal score increases. This avoids the weight jump problem caused by hard threshold division and ensures the continuity and stability of weight allocation. At the same time, by introducing an attenuation adjustment parameter, the attenuation rate of the weight coefficient can be flexibly controlled according to the actual application scenario, which enhances the adaptability of the method to different data types and noise levels.

[0016] Preferably, the step of performing spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested includes: The corrected spectral data are multiplied by the corresponding weighting coefficients to obtain the weighted spectral data. All the weighted spectral data are summed to obtain the spectral weighted sum; The weight normalization coefficients are obtained by summing all the weight coefficients. Calculate the ratio of the spectral weighted sum to the weighted normalization coefficient, and determine the resulting ratio as the target optimized spectrum of the sample to be tested.

[0017] By adopting the above technical solution, each calibration spectral data and its corresponding weight coefficient are weighted, summed, and normalized, so that the higher quality calibration spectral data accounts for a larger proportion in the target optimized spectrum, thereby effectively improving the signal-to-noise ratio and fidelity of the optimized spectrum.

[0018] Preferably, after performing spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested, the process includes: The target optimized spectrum is used as a dynamic reference spectrum; The similarity between each of the corrected spectral data and the dynamic reference spectrum is calculated to obtain multiple similarity values; Calculate the average similarity of all similarity values ​​and determine whether the average similarity meets the preset average similarity threshold; If the condition is not met, the abnormal score result of each of the corrected spectral data is updated based on the corresponding similarity value, and the process jumps to the step of determining the weight coefficient corresponding to each of the corrected spectral data based on the abnormal score result, until the average similarity meets the average similarity threshold, and the current target optimized spectrum is taken as the final optimized spectrum.

[0019] By adopting the above technical solution, the initially obtained target optimized spectrum is used as a dynamic reference spectrum. The anomaly scoring results are iteratively updated by using the similarity between each corrected spectral data and the dynamic reference spectrum. The iteration convergence is controlled by judging whether the average similarity meets the preset threshold, so that the final optimized spectrum is closer to the real spectral characteristics, further improving the accuracy and robustness of spectral reconstruction.

[0020] The second objective of this invention is to provide an optimization system based on spectral data, which improves the accuracy and robustness of spectral reconstruction.

[0021] The second objective of this invention is achieved through the following technical solution: An optimization system based on spectral data includes: The data acquisition module is used to acquire multiple raw spectral data of the sample to be tested; The data correction module is used to perform data correction on the multiple original spectral data to obtain multiple corrected spectral data; An anomaly scoring module is used to extract spectral features from each of the corrected spectral data, and calculate anomaly scores based on the extracted spectral features to obtain anomaly score results corresponding to each of the corrected spectral data. The weight determination module is used to determine the weight coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results. The spectral optimization module is used to perform spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested.

[0022] By adopting the above technical solution, the data acquisition module, data correction module, anomaly scoring module, weight determination module and spectral optimization module work together to achieve highly robust fusion and high-fidelity reconstruction of spectral data.

[0023] The third objective of this invention is to provide an electronic device that improves the accuracy and robustness of spectral reconstruction.

[0024] The above-mentioned objective three of this invention is achieved through the following technical solution: An electronic device includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the optimization method based on spectral data described in any of the preceding claims.

[0025] The fourth objective of this invention is to provide a computer-readable storage medium capable of storing corresponding programs, which facilitates the improvement of the accuracy and robustness of spectral reconstruction.

[0026] The fourth objective of this invention is achieved through the following technical solution: A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the optimization method based on spectral data described in any of the preceding claims.

[0027] In summary, the present invention has at least one of the following beneficial technical effects: 1. This invention corrects, extracts features and scores anomalies in the original spectral data, and then weights and fuses the corrected spectral data according to the scoring results by assigning differentiated weights to each data point. This effectively eliminates or suppresses the interference of abnormal spectra, improves the signal-to-noise ratio and fidelity of the target optimized spectrum, and solves the problem that the existing arithmetic mean method is easily contaminated by abnormal data, leading to spectral distortion. 2. This invention uses the initially obtained target optimized spectrum as a dynamic reference spectrum, calculates the similarity between each corrected spectral data and the reference spectrum, and iteratively updates the anomaly scoring results and weight coefficients until the average similarity meets a preset threshold, which can further improve the consistency between the optimized spectrum and the sample's true spectrum. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating the steps of an optimization method based on spectral data provided in Embodiment 1 of the present invention.

[0029] Figure 2 This is a structural block diagram of an optimization system based on spectral data provided in Embodiment 2 of the present invention. Detailed Implementation

[0030] This invention provides an optimization method and system based on spectral data to address the technical problem that existing technologies cannot remove significantly anomalous spectra, leading to distortion of the final average spectrum and poor robustness. It effectively improves the accuracy and robustness of spectral reconstruction.

[0031] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0032] It should be noted that in this embodiment of the invention, all content involving object data must be obtained with the object's authorization and consent, and must comply with current laws and standards. If the embodiment involves personal information, it must ensure that the individual's consent has been obtained; if it involves sensitive information, the separate consent of the information subject must be obtained. The implementation of the entire embodiment should also be based on the object's authorization and consent.

[0033] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this disclosure.

[0034] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Example 1

[0035] Please see Figure 1 The present invention provides an optimization method based on spectral data, comprising: Step 101: Collect multiple raw spectral data of the sample to be tested.

[0036] The sample to be tested refers to a chemical substance, biological sample, or industrial material in solid, liquid, or gaseous form.

[0037] Raw spectral data refers to measurement results directly output by a spectrometer without any preprocessing or correction. Spectrometers include, but are not limited to, ultraviolet-visible-near-infrared spectrometers, Raman spectrometers, or fluorescence spectrometers.

[0038] Understandably, during data acquisition, the sample to be tested is placed in the measurement optical path of the spectrometer, and under strictly identical measurement conditions (such as the spectrometer's light source intensity, integration time, detector gain, ambient temperature, sample position and orientation, and the starting wavelength and resolution of data acquisition remaining constant), the sample is continuously or intermittently acquired multiple times, resulting in a total of N raw spectral data points, denoted as N0. ,in , These are wavelength or wavenumber coordinates.

[0039] It is worth mentioning that, in order to balance data redundancy and processing efficiency, N is a positive integer greater than or equal to 3, and the preferred value range is 5 to 50.

[0040] In this embodiment of the invention, under the same measurement conditions, raw spectral data of the sample to be tested are collected continuously or at intervals.

[0041] Step 102: Perform data correction on multiple original spectral data to obtain multiple corrected spectral data.

[0042] Preferably, step 102 may include the following sub-steps: S11. Perform smoothing and denoising processing on the original spectral data to obtain the denoised spectral data.

[0043] It should be noted that, since the original spectral data contains high-frequency random noise, which mainly comes from the dark current of the spectrometer detector, readout noise, and intensity fluctuations of the light source, the moving window averaging method or the Savitzky-Golay convolution smoothing method is used to smooth and denoise each original spectral data separately.

[0044] Taking the Savitzky-Golay convolution smoothing method as an example, the smoothing window width is set to 2m+1 wavelength points. A k-order polynomial is used to perform least-squares fitting on the data points within the window. The center value of the fitted result is taken as the smoothing result for that wavelength point. The window is traversed sequentially to obtain the denoised spectral data. The smoothing window width and polynomial order are preset according to the sampling interval and noise level of the original spectrum. Typical values ​​are, for example, a window width of 11 to 25 wavelength points and a polynomial order of 2 to 4.

[0045] In this embodiment of the invention, the Savitzky-Golay convolution smoothing method is used to smooth and denoise the original spectral data, so that the high-frequency spikes caused by random noise in the original spectrum are effectively suppressed, thereby obtaining the denoised spectral data.

[0046] S12. Identify the characteristic peaks in the denoised spectral data and match them with the standard characteristic peaks in the preset standard spectrum to determine the correspondence.

[0047] Characteristic peaks refer to the convex regions on the spectral curve where the local intensity is significantly higher than the surrounding background. They usually correspond to the characteristic absorption, scattering, or emission responses of specific chemical functional groups or components in the sample being tested.

[0048] A preset standard spectrum refers to a reference spectrum of a known component or standard sample stored in advance under the same measurement conditions. This spectrum contains several standard characteristic peaks at known wavelength positions and is used as a reference standard for wavelength correction.

[0049] It is worth mentioning that if a sufficient number of characteristic peaks cannot be identified in the denoised spectrum (e.g., less than 2), an alternative solution is adopted, such as directly using the preset offset for correction, or prompting the user to check the measurement conditions.

[0050] In this embodiment of the invention, the first or second derivative of the spectrum is first calculated, and extreme points where the derivative is zero and the second derivative is negative are marked as candidate peak positions. Then, the peak height and peak width are evaluated near each candidate peak, and spurious peaks with a signal-to-noise ratio lower than a preset threshold (e.g., peak height less than 3 times the standard deviation of the neighborhood background noise) are eliminated. Finally, a set of identified characteristic peaks is obtained, where each characteristic peak includes peak position and peak height parameters. Then, the identified characteristic peaks are matched with standard characteristic peaks in a preset standard spectrum using a nearest neighbor matching strategy: for each standard characteristic peak in the preset standard spectrum, the characteristic peak with the smallest deviation from its preset wavelength position is searched in the characteristic peak set of the denoised spectrum. If the absolute value of this deviation is less than a preset matching tolerance window (e.g., ...), the peak is considered a candidate peak. If the deviation exceeds the tolerance window, the two are considered to be matched successfully, and a correspondence between the characteristic peak and the standard characteristic peak is established; if the deviation exceeds the tolerance window, the standard characteristic peak is considered to be missing in the current spectrum, and no correspondence is established.

[0051] S13. Based on the correspondence, calculate the wavelength offset of the characteristic peak relative to the standard characteristic peak.

[0052] Wavelength offset refers to the difference between the actual measured position of the characteristic peak and the corresponding position of the standard characteristic peak in the preset standard spectrum.

[0053] It is understandable that, for the k-th pair of successfully matched characteristic peaks, the standard wavelength of the standard characteristic peak is denoted as . The actual measured wavelength of the identified characteristic peak is Then the wavelength offset Calculate using the following formula:

[0054] in, It can be a positive or negative value. A positive value indicates that the actual measured wavelength is greater than the standard wavelength (i.e., redshift), and a negative value indicates that the actual measured wavelength is less than the standard wavelength (i.e., blueshift).

[0055] When multiple successfully matched feature peak pairs exist, the wavelength offset corresponding to each pair is calculated to obtain the offset set. , where m is the total number of successfully matched feature peak pairs.

[0056] It is worth mentioning that by calculating the wavelength shift, the degree of deviation of the denoised spectrum from the standard spectrum on the wavelength axis can be quantitatively described. If only one pair of characteristic peaks is successfully matched, the shift can be directly used as the basis for subsequent wavelength axis correction. If multiple pairs of characteristic peaks are successfully matched, the mean or median of multiple shifts can be further calculated to suppress the influence of single-peak matching error and obtain a more robust shift estimate.

[0057] In this embodiment of the invention, based on the established correspondence between the characteristic peaks and the standard characteristic peaks, the wavelength offset corresponding to each successfully matched characteristic peak is calculated one by one.

[0058] S14. Based on the wavelength offset, perform wavelength axis correction on the denoised spectral data to obtain corrected spectral data.

[0059] Wavelength axis correction refers to the remapping or translation adjustment of wavelength coordinates on the denoised spectral data based on the calculated wavelength offset, in order to eliminate systematic wavelength drift errors and ensure that the position of the spectral data on the wavelength axis is consistent with the preset standard spectrum, thereby improving the comparability between multiple spectra and between the spectrum and the standard reference.

[0060] Corrected spectral data refers to the spectral data obtained after wavelength axis correction of denoised spectral data. Its wavelength coordinates are consistent with the wavelength coordinate system of a preset standard spectrum, eliminating systematic wavelength deviations introduced by factors such as instrument drift and environmental changes. This ensures direct comparability in wavelength position between different batches of acquisitions and different spectra. It is denoted as […]. , where i is the spectral index and λ is the standard wavelength axis coordinate.

[0061] When there is only one wavelength offset (e.g., only one pair of characteristic peaks are successfully matched, or the average of multiple offsets is taken as the global offset), a global translation correction method is adopted: each wavelength point in the denoised spectral data is shifted to the global offset. Intensity value Remapped to new wavelength position The corrected spectral data is then generated.

[0062] It is worth mentioning that, since the actual sampling wavelength points are discretely and equally spaced, the new wavelength positions after translation... The new wavelength may not coincide with the original sampling point. In this case, interpolation methods (such as linear interpolation or cubic spline interpolation) are used to calculate the intensity value corresponding to the new wavelength position, thus obtaining the corrected spectral data. ,in, Standard wavelength axis coordinates are used uniformly.

[0063] When multiple characteristic peak matching pairs exist and the wavelength shifts are not completely consistent (i.e., different bands exhibit different degrees of drift, such as nonlinear wavelength distortion), piecewise correction or polynomial fitting correction methods are used: control point pairs are constructed using the standard wavelength of each matching characteristic peak and the actual measured wavelength. , This represents the actual measured wavelength position of the k-th characteristic peak identified in the denoised spectral data. To determine the standard wavelength position of the k-th standard characteristic peak corresponding to the predefined characteristic peak in the standard spectrum, a first-order or second-order polynomial is used to fit the wavelength mapping function. ,in, This refers to the actual measured wavelength value at a specific wavelength point in the denoised spectral data. This is the standard wavelength value corresponding to this wavelength point after correction. The mapping function is then applied to each measurement wavelength point of the denoised spectrum. The mapping function is used to calculate the corresponding standard wavelength position, and then the corrected intensity value is obtained by interpolation, thereby realizing nonlinear correction of the full spectrum wavelength.

[0064] For example, the measurement wavelength at each control point pair As input, standard wavelength As the output, a first-order linear function is used. or second-order polynomial function Perform fitting to determine function parameters Thus, the complete mapping expression is obtained.

[0065] In this embodiment of the invention, wavelength offset is used to perform wavelength correction processing on the denoised spectral data, so that each corrected spectral data has good consistency on the wavelength axis, eliminating wavelength deviation caused by instrument drift or environmental changes.

[0066] Based on the above feasible embodiments, the dark background subtraction method can also be used to correct the original spectral data. The dark background subtraction method refers to acquiring the dark background spectral signal output by the spectrometer itself under the condition of no light input or the optical path being blocked, using the same integration time and gain parameters as the sample acquisition. This signal mainly comes from the dark current noise of the detector, the background bias of the electronic readout, and non-signal components such as environmental thermal radiation.

[0067] Specifically, when performing dark background subtraction, the raw spectral data collected will be subtracted from the original data. Subtract the pre-collected and stored dark background spectrum ,Right now In the formula, This represents the i-th spectral data after dark background subtraction correction, which effectively suppresses fixed-mode noise introduced by the instrument background and improves the signal-to-noise ratio and baseline flatness of the spectral data.

[0068] It should be noted that the dark background subtraction method can be used in conjunction with the aforementioned smoothing and denoising methods, wavelength axis correction, etc. For example, the original spectral data can be subjected to dark background subtraction first, and then smoothing and denoising can be performed. Alternatively, dark background subtraction can be used as a preprocessing step before smoothing and denoising to eliminate different types of noise interference step by step and further improve the quality of the corrected spectral data.

[0069] Step 103: Extract spectral features from each corrected spectral data, and calculate anomaly scores based on the extracted spectral features to obtain the anomaly score results corresponding to each corrected spectral data.

[0070] Preferably, step 103 may include the following sub-steps: Spectral characteristics include global baseline deviation and local noise level.

[0071] Spectral features refer to quantitative indicators extracted from calibrated spectral data that can quantitatively describe the quality or degree of anomalies of the spectral data.

[0072] Global baseline deviation refers to the degree of difference between the average intensity of a single corrected spectral data point in a preset background band and the median of the average intensity of all corrected spectral data points in that band.

[0073] Local noise level refers to the standard deviation of the intensity difference between adjacent wavelengths in a single calibrated spectral data point within a preset noise assessment band.

[0074] S21. Calculate the average intensity value of each corrected spectral data in the preset background band, and use it as the individual background intensity of each corrected spectral data.

[0075] The preset background band refers to one or more continuous intervals selected in advance on the wavelength or wavenumber coordinate axis. The selection of the background band is determined based on the spectral characteristics of the sample to be tested, and it is usually located at the two ends of the spectrum or in a flat region with no known characteristic response.

[0076] The average intensity value refers to the arithmetic mean obtained by summing the spectral intensity at each wavelength point within a preset background band and dividing it by the total number of wavelength points within that band.

[0077] Individual background intensity refers to the average intensity value calculated for a specific calibrated spectral data point within its preset background band.

[0078] Understandably, for each corrected spectral data Within the preset background band, the spectral intensity values ​​corresponding to all wavelength points within the band are accumulated, and then divided by the total number of wavelength points within the band to obtain the average intensity value within the band. This value is used as the individual background intensity of the i-th corrected spectral data, denoted as . The specific calculation formula is as follows:

[0079] In the formula, L is the set of wavelength point indices included in the preset background band, where L is the total number of wavelength points within that band. For the i-th corrected spectral data at wavelength point The intensity value at that location.

[0080] In this embodiment of the invention, the average intensity value of the corrected spectral data is calculated within a preset background band and used as the individual background intensity.

[0081] S22. Calculate the median of the average intensity values ​​of all corrected spectral data in the background band, and use it as the overall background intensity benchmark.

[0082] The overall background intensity reference refers to the median of the average intensity values ​​of all calibrated spectral data within the background band, denoted as . .

[0083] In this embodiment of the invention, the individual background intensity is based on each corrected spectral data. ,in Collect the individual background intensities of all N calibrated spectral data points to form a numerical sequence of length N. Then, the sequence is sorted, and the value in the middle position after sorting is taken as the median. Specifically, when N is odd, the median is the (N+1) / 2th value after sorting; when N is even, the median is the arithmetic mean of the N / 2th and N / 2+1th values ​​after sorting. This median is used as the benchmark for the overall background intensity.

[0084] S23. Calculate the difference between the individual background intensity of each corrected spectral data and the overall background intensity benchmark to obtain the global baseline deviation of each corrected spectral data.

[0085] For each corrected spectral data point, the absolute value of the difference between its individual background intensity and the overall background intensity benchmark is calculated. This difference is used as the global baseline deviation of that corrected spectral data point. The specific calculation formula is as follows:

[0086] In the formula, This represents the global baseline deviation.

[0087] It is worth mentioning that the magnitude of the global baseline deviation directly reflects the baseline shift of the spectrum relative to the median level of all spectra: the larger the value, the more serious the background intensity of the corrected spectrum deviates from the overall reference level, that is, the more significant the baseline drift and the worse the spectral quality; the smaller the value, the closer the baseline of the spectrum is to the typical normal level.

[0088] To further eliminate the influence of intensity dimensions between different measurement batches or different samples, the global baseline deviation can be normalized: for example, using the overall background intensity benchmark as the denominator, the relative deviation can be calculated. The deviation is expressed as a percentage. The normalized global baseline deviation is highly comparable across different measurement batches, making it easy to set a uniform anomaly detection threshold.

[0089] In this embodiment of the invention, the absolute value of the difference between the individual background intensity of each corrected spectral data and the overall background intensity benchmark is calculated to obtain the global baseline deviation of each corrected spectral data.

[0090] S24. Within the preset noise assessment band, calculate the intensity difference between adjacent wavelength points in each corrected spectral data in sequence to form a difference sequence, and calculate the standard deviation of the difference sequence to obtain the local noise level of each corrected spectral data.

[0091] A preset noise assessment band refers to one or more continuous intervals pre-selected on the wavelength or wavenumber coordinate axis. The spectral signal within this interval does not contain characteristic peaks, and the actual signal changes relatively smoothly with wavelength (i.e., the theoretical slope is close to zero). The noise assessment band can be the same as, partially overlap with, or be completely independent of the preset background band; the specific selection depends on the spectral characteristics of the sample under test. The purpose of setting a noise assessment band is to provide a stable assessment interval with a flat signal, unaffected by characteristic peaks, for calculating the local noise level.

[0092] The difference sequence refers to the numerical sequence formed by calculating the spectral intensity difference between two adjacent wavelength points in ascending order within a preset noise assessment band, and arranging these differences in wavelength order.

[0093] Understandably, for each corrected spectral data point, within the preset noise assessment band, adjacent wavelength points are traversed sequentially in ascending order, and the spectral intensity difference between the subsequent wavelength point and the preceding wavelength point is calculated. Let there be M wavelength points in the noise assessment band, denoted in wavelength order as follows: The corresponding spectral intensity value is The intensity difference between adjacent wavelength points Calculate using the following formula:

[0094] The above M−1 differences constitute a difference sequence. The difference sequence reflects the intensity fluctuation of the spectrum in a local range: when the spectral signal is smooth and the noise is small, the absolute value of the difference between adjacent points is small and the difference sequence is relatively concentrated; when the spectrum has high-frequency spikes or large random noise, the absolute value of the difference between adjacent points will fluctuate significantly and the dispersion of the difference sequence will increase accordingly.

[0095] Subsequently, the standard deviation of this difference sequence is calculated, which is used as the local noise level of this corrected spectral data, denoted as . The formula for calculating the standard deviation is as follows:

[0096] in, It is the arithmetic mean of the difference sequence, i.e. M is the total number of wavelength points included in the preset noise evaluation band, and J is the index number of the difference element in the difference sequence.

[0097] In this embodiment of the invention, within a preset noise assessment band, the intensity difference between adjacent wavelength points in each corrected spectral data is calculated sequentially to form a difference sequence, and the standard deviation of the difference sequence is calculated to obtain the local noise level of each corrected spectral data, providing an important quantitative basis for the diagnosis of abnormal spectra from the perspective of local noise.

[0098] S25. The global baseline deviation and local noise level are fused and calculated to obtain the anomaly score results corresponding to each corrected spectral data.

[0099] Anomaly score results refer to the quantitative values ​​obtained through fusion calculations used to quantitatively characterize the overall anomaly degree of a single corrected spectral data point, denoted as... , where the subscript i is the spectral number.

[0100] The anomaly score results can be determined using a weighted fusion method or anomaly scoring model: When using the weighted summation method

[0101] In the formula, and These are the dimensionless values ​​of the global baseline deviation and the local noise level after normalization, respectively. and The preset weighting coefficients satisfy... .

[0102] Normalization can be achieved using Z-score standardization or min-max normalization. Specifically, let's take min-max normalization as an example: , Similarly, to ensure that the normalized values ​​all fall within the [0,1] interval, the weighting coefficients α and β are set according to the relative impact of baseline drift and local noise on data quality in the actual application scenario. For example, in measurement environments where baseline drift is significant, the weighting coefficients α and β can be adjusted accordingly. Set it to a high level (e.g.) =0.7, =0.3); In cases of severe noise interference, β can be set higher (e.g., =0.3, =0.7); if the two are of equal importance, then they can be set to 0.7. = =0.5.

[0103] When using an anomaly scoring model, multiple spectral features can be extracted from each corrected spectral data, and these features can be combined into a feature vector. , where p is the total number of feature dimensions, and then the feature vector is input into the anomaly scoring model, and the model output is the anomaly score result of the corrected spectral data.

[0104] It is worth mentioning that the anomaly scoring model can be an anomaly detection model based on isolated forest or single-class support vector machine. That is, the feature vectors of all corrected spectral data are used as training samples to build an isolated forest or single-class support vector machine model. This model can learn the distribution boundary of normal spectra in high-dimensional feature space. For each spectrum, the model outputs its anomaly score (such as the inverse of the path length in isolated forest or the distance to the decision boundary in single-class support vector machine), and this score is used as the anomaly scoring result.

[0105] Understandably, multi-dimensional feature fusion diagnosis, compared to weighted summation of single or a few features, can capture spectral abnormal patterns from a more comprehensive perspective. Especially when the abnormal manifestations are complex and a single indicator is difficult to effectively distinguish, the abnormality scoring model can utilize the nonlinear relationships and interactions between features to improve the accuracy and robustness of abnormality diagnosis. The selection of the above-mentioned abnormality scoring models and the specific parameter settings can be flexibly determined according to the availability of training data and the requirements for model interpretability in the actual application scenario.

[0106] In this embodiment of the invention, the global baseline deviation and local noise level are fused and calculated using a weighted fusion method or by constructing an anomaly scoring model to obtain the anomaly scoring results corresponding to each corrected spectral data.

[0107] Based on the above-described feasible embodiments, the spectral features may also include the overall signal-to-noise ratio and / or the consistency of characteristic peak parameters.

[0108] The overall signal-to-noise ratio (SNR) refers to the ratio of the mean to the standard deviation of the spectral signal within a preset stationary band. It characterizes the overall signal quality of the corrected spectral data. The stationary band should be selected as a region in the spectrum without characteristic peaks and with a relatively flat signal to avoid interference from fluctuations in the real signal on noise estimation. Specifically, for the i-th corrected spectral data, the arithmetic mean of the spectral intensities at all wavelengths within the preset stationary band is calculated. and standard deviation The overall signal-to-noise ratio A higher overall signal-to-noise ratio (SNR) indicates a greater signal intensity relative to the noise level in that spectrum, resulting in better data quality. Conversely, a lower overall SNR indicates more severe noise contamination and poorer data quality. In subsequent anomaly scoring, the reciprocal or negative value of the overall SNR can be used as a quantitative indicator of the degree of anomaly; that is, a spectrum with a higher SNR should receive a lower anomaly score.

[0109] Characteristic peak parameter consistency refers to the consistency of key characteristic peak parameters (including but not limited to peak positions) identified in the corrected spectral data. Peak height Peak area The deviation is calculated by comparing the corresponding characteristic peak parameters with the median or mean of all calibrated spectral data. This deviation characterizes the consistency level of the spectrum with most spectra in terms of characteristic peak morphology. Specifically, for a preset key characteristic peak, the peak height values ​​at that characteristic peak are collected from all N calibrated spectral data. Calculate the median or mean Then calculate the peak height deviation of the i-th spectrum: (Expressed as a percentage of relative deviation), similarly, the peak position deviation can be calculated. and peak area deviation The larger the deviation value, the more abnormal the morphology of the spectrum at that characteristic peak, which may be due to factors such as sample bubbles, particulate matter obstruction, or slight changes in the measurement position. In multi-feature fusion, the parameter consistency indices of multiple characteristic peaks can be weighted and combined to obtain an overall characteristic peak parameter consistency score.

[0110] Understandably, the aforementioned overall signal-to-noise ratio and characteristic peak parameter consistency indicators can be used as a supplement to the global baseline deviation and local noise level, and incorporated into the fusion calculation to form a more comprehensive and multi-dimensional spectral quality assessment system, thereby further improving the accuracy and robustness of abnormal spectral identification. The specific features selected can be flexibly configured according to the main manifestations of anomalies in the actual application scenario.

[0111] Step 104: Based on the anomaly scoring results, determine the weighting coefficients corresponding to each corrected spectral data.

[0112] Preferably, step 104 may include the following sub-steps: S31. Compare the anomaly score results corresponding to each corrected spectral data with the preset anomaly score threshold.

[0113] The preset anomaly scoring threshold refers to a pre-set numerical limit used to determine whether the degree of anomaly in the corrected spectral data has reached the level that requires complete removal. This threshold can be set according to the data quality requirements in the actual application scenario, and is usually between 0.5 and 0.9.

[0114] In this embodiment of the invention, the calculated anomaly score results of each corrected spectral data are compared one by one with a preset anomaly score threshold.

[0115] S32. If the abnormal score result is greater than the abnormal score threshold, the weight coefficient of the corresponding corrected spectral data is set to zero.

[0116] A spectral spectrum with a weighting coefficient of zero means that the spectrum makes no contribution whatsoever in the subsequent weighted average calculation, which is equivalent to being completely removed from the valid spectral dataset.

[0117] In this embodiment of the invention, when the abnormal score of a certain calibration spectral data exceeds the preset abnormal score threshold, the spectrum is determined to be a severely abnormal spectrum. Its data quality can no longer be effectively utilized by weighting down. If it is included in the subsequent spectral fusion, it may still cause non-negligible pollution to the final optimized spectrum. Therefore, the weight coefficient of the calibration spectral data is directly set to zero, which can effectively block the interference of severely abnormal spectra on the target optimized spectrum and ensure that the final fusion result comes only from spectral data of acceptable quality. At the same time, the spectrum number and its abnormal score result can be recorded and output for the operator to trace the cause of the abnormality or determine whether the sample measurement needs to be repeated.

[0118] S33. If the abnormal score result is less than or equal to the abnormal score threshold, the weight coefficient corresponding to the corrected spectral data is determined based on the degree of deviation between the abnormal score result and the preset benchmark score. The smaller the abnormal score result, the larger the weight coefficient corresponding to the corrected spectral data.

[0119] The preset baseline score refers to the reference value used to measure the relative size of abnormal scores.

[0120] In this embodiment of the invention, when the abnormal score result of a certain corrected spectral data is less than or equal to a preset abnormal score threshold, the quality of the spectrum is determined to be within an acceptable range, and it is retained and participates in subsequent spectral fusion. Subsequently, based on the degree of deviation between the abnormal score result and the preset benchmark score (usually set to 0), its weight coefficient is determined. Specifically, a linear weight allocation method is adopted: ,in The result is the abnormal rating, and T is the preset abnormal rating threshold.

[0121] Preferably, step 104 may further include the following sub-steps: S41. Calculate the ratio of the square of the abnormal score result to the square of twice the preset attenuation adjustment parameter to obtain the first intermediate value.

[0122] The preset decay adjustment parameter refers to the adjustment factor used to control how quickly the weight coefficient decays as the abnormal score increases, denoted as... The value of the attenuation adjustment parameter is greater than 0. The larger the attenuation adjustment parameter is, the slower the weight coefficient decays as the abnormal score results increase, that is, the higher the tolerance for abnormal scores. The smaller the attenuation adjustment parameter is, the faster the decay rate is, that is, the higher the sensitivity to abnormal scores. This parameter can be flexibly set according to the data quality requirements in the actual application scenario. The typical value range is 0.1 to 0.5.

[0123] It should be noted that the formula for determining the first intermediate value can be encapsulated as follows:

[0124] In the formula, This is the first intermediate value.

[0125] In this embodiment of the invention, the average value of the preset attenuation adjustment parameter is calculated and then multiplied by 2 to obtain... Then the squared value of the abnormal score results and Divide them and calculate the ratio, then use this ratio as the first intermediate value.

[0126] S42. Calculate the negative exponent of the first intermediate value to obtain the weighting coefficient of the corrected spectral data. The attenuation adjustment parameter is used to control the attenuation rate of the weighting coefficient as the abnormal scoring results change.

[0127] It should be noted that the formula for determining the weighting coefficients of the corrected spectral data can be encapsulated as follows:

[0128] In the formula, is the weighting coefficient, and exp(·) is an exponential function with the natural constant ee as the base.

[0129] As shown in the above formula, the weighting coefficient ranges from (0,1], when the abnormal scoring result... hour, When the weighting coefficient reaches its maximum value; when abnormal scoring results hour, ,and The larger, The smaller the value, the more the weighting coefficient monotonically decreases as the abnormal rating results increase.

[0130] In this embodiment of the invention, the weighting coefficients of the corrected spectral data are determined by the soft weighting method of Gaussian function, which can assign continuously changing weighting coefficients to all spectral data, avoiding the instability of the spectrum being completely retained or completely removed due to small fluctuations in the threshold boundary.

[0131] Step 105: Based on the weighting coefficients, perform spectral optimization processing on each corrected spectral data to obtain the target optimized spectrum of the sample to be tested.

[0132] Preferably, step 105 may include the following sub-steps: S51. Multiply each corrected spectral data with its corresponding weighting coefficient to obtain each weighted spectral data.

[0133] Weighted spectral data refers to the spectral data obtained by multiplying the intensity value of each wavelength point in the corrected spectral data by the weighting coefficient corresponding to that spectrum. It is denoted as... , where the subscript i is the spectral number and λ is the wavelength or wavenumber coordinate.

[0134] For each corrected spectral data point, the spectral intensity values ​​at all wavelength points are multiplied by the corresponding weighting coefficients to obtain the weighted spectral data for that spectrum. The specific calculation formula is as follows:

[0135] In the formula, For weighted spectral data, Let be the intensity value of the i-th corrected spectral data at wavelength λ. This represents the weighting coefficient corresponding to this spectrum.

[0136] In this embodiment of the invention, by performing wavelength-by-wavelength product operation on the corrected spectral data and the corresponding weighting coefficients, the enhancement and preservation of high-quality spectra and the effective suppression of low-quality spectra are achieved.

[0137] S52. Accumulate all weighted spectral data to obtain the spectral weighted sum.

[0138] Spectral weighted sum refers to the spectral data obtained by summing the weighted spectral data corresponding to all corrected spectral data point by wavelength, denoted as . It reflects the total contribution of each spectrum under the influence of its respective weighting coefficient.

[0139] For each wavelength point, the intensity values ​​of all N weighted spectral data at that wavelength point are summed to obtain the cumulative result for that wavelength point. After traversing all wavelength points, a complete weighted spectral sum is formed. The calculation formula is as follows:

[0140] In this embodiment of the invention, all weighted spectral data are summed to obtain a spectral weighted sum.

[0141] S53. Sum all the weight coefficients to obtain the weight normalization coefficient.

[0142] The weighting normalization coefficient is a scalar used to normalize the weighted sum of spectra, denoted as Its value is the sum of the weighting coefficients of all the fused corrected spectral data. Its purpose is to eliminate the influence of the magnitude of the weighting coefficients on the amplitude of the weighted average result, so that the final output target optimized spectral intensity range is consistent with the original spectrum.

[0143] The weighting coefficients of all determined corrected spectral data are summed to obtain the weighted normalization coefficients, calculated using the following formula:

[0144] In the formula, N represents the total number of times the original spectral data was collected.

[0145] In this embodiment of the invention, all weight coefficients are summed to obtain the weight normalization coefficient.

[0146] S54. Calculate the ratio of the spectral weighted sum to the weighted normalization coefficient, and determine the obtained ratio as the target optimized spectrum of the sample to be tested.

[0147] Target optimized spectrum refers to the final spectral data obtained after a series of optimization processes, including anomaly diagnosis, weight allocation, and weighted fusion, denoted as .

[0148] For each wavelength point, calculate the ratio of the intensity value of the weighted spectral sum at that wavelength point to the weighted normalization coefficient. Use this ratio as the intensity value of the target optimized spectrum at that wavelength point. After traversing all wavelength points, the complete target optimized spectrum is obtained. The calculation formula is as follows:

[0149] It should be noted that when all weight coefficients are zero (i.e. all spectra are judged to be severely abnormal), the denominator is zero and cannot be directly calculated. In this case, abnormal handling measures can be taken, such as prompting the user to re-acquire the spectrum, or using the arithmetic mean of all corrected spectral data as an alternative output. However, in practical applications, a reasonably set abnormal scoring threshold will usually not cause all weight coefficients to be zero.

[0150] In this embodiment of the invention, based on the weighting coefficient, each corrected spectral data is weighted, summed, and normalized to obtain the target optimized spectrum of the sample to be tested. This results in the target optimized spectrum improving the signal-to-noise ratio while avoiding spectral distortion caused by abnormal data contamination.

[0151] To further improve the signal-to-noise ratio and robustness of the final optimized spectrum, based on the above-described feasible embodiments, the following sub-steps may be included after step 105: A1. Use the target optimized spectrum as the dynamic reference spectrum.

[0152] Dynamic reference spectrum refers to using the target optimized spectrum obtained in the previous iteration as a reference benchmark for evaluating the similarity of each corrected spectral data in the current iteration. , where the superscript k is the iteration number.

[0153] In this embodiment of the invention, the calculated target optimized spectrum is used as the dynamic reference spectrum for the first iteration. In subsequent iterations, the newly calculated target optimized spectrum for each iteration will be used as the dynamic reference spectrum for the next iteration, thereby realizing the dynamic updating of the reference spectrum and making the reference standard approach the true spectral characteristics of the sample round by round.

[0154] A2. Calculate the similarity between each corrected spectral data and the dynamic reference spectrum to obtain multiple similarity values.

[0155] Similarity is a quantitative indicator used to quantitatively measure the degree of waveform consistency between calibrated spectral data and dynamic reference spectra, denoted as . Its value range is usually [−1,1] or [0,1]. The larger the value, the more similar the waveforms of the two spectra are.

[0156] It should be noted that the similarity can be calculated using the Pearson correlation coefficient, and the formula is as follows:

[0157] In the formula, M represents the total number of wavelength points. Let be the intensity mean of the i-th corrected spectral data. This represents the average intensity of the dynamic reference spectrum.

[0158] The Pearson correlation coefficient ranges from [−1, 1]. The closer the value is to 1, the stronger the positive correlation between the two spectra and the more similar the waveforms. The closer the value is to 0, the weaker the correlation between the two spectra. The closer the value is to -1, the negative correlation between the two spectra. In actual spectral analysis, the similarity between the normal spectrum and the dynamic reference spectrum is usually close to 1. As an alternative, cosine similarity or Euclidean distance after conversion can also be used as a similarity measure.

[0159] Through the above calculations, a total of N similarity values ​​were obtained. , which correspond to the similarity between N corrected spectral data and the current dynamic reference spectrum.

[0160] In this embodiment of the invention, the similarity between each corrected spectral data and the dynamic reference spectrum is calculated to obtain multiple similarity values.

[0161] A3. Calculate the average similarity of all similarity values ​​and determine whether the average similarity meets the preset average similarity threshold.

[0162] Average similarity refers to the arithmetic mean of the similarity values ​​between all corrected spectral data and the current dynamic reference spectrum, denoted as . It is used to characterize the overall consistency between the current batch of calibrated spectral data and the dynamic reference spectrum.

[0163] The preset average similarity threshold refers to a pre-defined numerical limit used to determine whether the iterative process has reached convergence. It is denoted as... The typical value range is 0.95 to 0.99, which can be flexibly set according to the requirements of spectral consistency and noise level in actual applications.

[0164] In this embodiment of the invention, for all the calculated similarity values, their arithmetic mean is calculated, and the calculated average similarity is used. With the preset average similarity threshold Compare: If This indicates that the overall consistency between the current dynamic reference spectrum and each corrected spectral data has met the preset requirements, and the iteration process can be terminated; if This indicates that there is still a significant difference between the current dynamic reference spectrum and the corrected spectrum data, requiring further iterative optimization.

[0165] A4. If not satisfied, update the anomaly score results of each corrected spectral data based on each similarity value, and jump to the step of determining the weight coefficients corresponding to each corrected spectral data based on the anomaly score results, until the average similarity meets the average similarity threshold, and take the current target optimized spectrum as the final optimized spectrum.

[0166] It should be noted that the specific steps for updating the anomaly score results of each corrected spectral data based on each similarity value are as follows: First, the temporary anomaly score derived from similarity in the current round is weighted and fused with the historical anomaly score from the previous round to preserve historical information and improve the stability of the score. The fusion and update formula is as follows:

[0167] In the formula, This is the abnormal score result before the update (i.e., the abnormal score of the corrected spectral data in the previous iteration). This is a temporary anomaly score calculated based on similarity. This is the updated anomaly scoring result used for the next iteration. The preset forgetting factor has a value range of 100%. Preferred It is 0.7.

[0168] Temporary anomaly score This is obtained by converting similarity into anomaly scores, specifically: Among them, similarity The range of values ​​is .

[0169] This conversion formula makes: when When there is a perfect positive correlation, i.e., the spectra are completely similar, =0 (The anomaly score is the lowest); when (Perfectly negative correlation) , =1 (The highest score is for anomalies); when (When unrelated) =0.5 (Medium anomaly), thus ensuring that the lower the similarity, the higher the temporary anomaly score.

[0170] In this embodiment of the invention, when the average similarity is less than a preset average similarity threshold, it is determined that the iteration has not yet converged and further optimization is needed. After optimization, the process jumps back to step 104 (determining the weight coefficients corresponding to each corrected spectral data based on the anomaly scoring results), uses the updated anomaly scoring results to redetermine the weight coefficients of each corrected spectral data, and sequentially executes step 105 (performing spectral optimization processing based on the weight coefficients) to calculate a new round of target optimized spectrum. The new round of target optimized spectrum is used as the updated dynamic reference spectrum, and steps A1 to A4 are repeated to enter the next round of iteration. The above iteration process continues until the average similarity calculated in a certain round of iteration meets the preset average similarity threshold. At this point, it is determined that the iteration has converged, the iteration process is terminated, and the target optimized spectrum obtained in the current round is output as the final optimized spectrum. Example 2

[0171] Please see Figure 2 The present invention provides an optimization system based on spectral data, comprising: The data acquisition module 201 is used to acquire multiple raw spectral data of the sample to be tested.

[0172] The data correction module 202 is used to perform data correction on multiple raw spectral data to obtain multiple corrected spectral data.

[0173] The anomaly scoring module 203 is used to extract spectral features from each corrected spectral data and calculate anomaly scores based on the extracted spectral features to obtain the anomaly score results corresponding to each corrected spectral data.

[0174] The weight determination module 204 is used to determine the weight coefficients corresponding to each corrected spectral data based on the anomaly scoring results.

[0175] The spectral optimization module 205 is used to perform spectral optimization processing on each corrected spectral data based on weighting coefficients to obtain the target optimized spectrum of the sample to be tested.

[0176] Since the above is a system corresponding to an optimization method based on spectral data, its implementation principle is consistent with that of an optimization method based on spectral data. For the sake of convenience and brevity, those skilled in the art can clearly understand that the specific working process of the system and modules described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here. Example 3

[0177] An electronic device according to an embodiment of the present invention includes: a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs an optimization method based on spectral data as described in any of the above embodiments.

[0178] The memory can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory has storage space for program code used to perform any of the method steps described above. For example, the storage space for program code may include individual program codes for implementing the various steps in the methods described above. This program code can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above. Example 4

[0179] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed, implements the optimization method based on spectral data of any of the above embodiments.

[0180] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0183] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0184] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0185] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An optimization method based on spectral data, characterized in that, include: Collect multiple raw spectral data of the sample to be tested; Data correction is performed on the multiple original spectral data to obtain multiple corrected spectral data; Spectral features are extracted from each of the corrected spectral data, and anomaly scores are calculated based on the extracted spectral features to obtain the anomaly score results corresponding to each of the corrected spectral data. Based on the anomaly scoring results, determine the weighting coefficients corresponding to each of the corrected spectral data. Based on the weighting coefficients, the calibration spectral data are subjected to spectral optimization processing to obtain the target optimized spectrum of the sample to be tested.

2. The optimization method based on spectral data according to claim 1, characterized in that, The process of performing data correction on the multiple original spectral data to obtain multiple corrected spectral data includes: The original spectral data is smoothed and denoised to obtain denoised spectral data; Identify the characteristic peaks in the denoised spectral data and match them with the standard characteristic peaks in the preset standard spectrum to determine the correspondence; Based on the correspondence, the wavelength shift of the characteristic peak relative to the standard characteristic peak is calculated; Based on the wavelength offset, wavelength axis correction is performed on the denoised spectral data to obtain corrected spectral data.

3. The optimization method based on spectral data according to claim 1, characterized in that, The spectral features include global baseline deviation and local noise level; the process of extracting spectral features from each of the corrected spectral data and calculating anomaly scores based on the extracted spectral features to obtain anomaly score results corresponding to each of the corrected spectral data includes: Calculate the average intensity value of each of the corrected spectral data in the preset background band, and use it as the individual background intensity of each of the corrected spectral data. Calculate the median of the average intensity values ​​of all corrected spectral data within the background band, and use it as the overall background intensity reference. The difference between the individual background intensity of each of the corrected spectral data and the overall background intensity benchmark is calculated to obtain the global baseline deviation of each of the corrected spectral data. Within a preset noise assessment band, the intensity difference between adjacent wavelength points in each of the corrected spectral data is calculated sequentially to form a difference sequence, and the standard deviation of the difference sequence is calculated to obtain the local noise level of each of the corrected spectral data. The global baseline deviation and the local noise level are fused together to obtain the anomaly score results corresponding to each of the corrected spectral data.

4. The optimization method based on spectral data according to claim 1, characterized in that, The step of determining the weighting coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results includes: The anomaly score results corresponding to each of the corrected spectral data are compared with the preset anomaly score threshold. If the anomaly score result is greater than the anomaly score threshold, the weight coefficient of the corresponding corrected spectral data is set to zero. If the abnormal score result is less than or equal to the abnormal score threshold, then the weight coefficient corresponding to the corrected spectral data is determined based on the degree of deviation between the abnormal score result and the preset benchmark score, wherein the smaller the abnormal score result, the larger the weight coefficient corresponding to the corrected spectral data.

5. The optimization method based on spectral data according to claim 1, characterized in that, The step of determining the weighting coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results further includes: The first intermediate value is obtained by calculating the ratio of the square of the abnormal score result to the square of twice the preset attenuation adjustment parameter. Calculate the negative exponent of the first intermediate value to obtain the weighting coefficient of the corrected spectral data, wherein the attenuation adjustment parameter is used to control the attenuation rate of the weighting coefficient as the abnormal scoring result changes.

6. The optimization method based on spectral data according to claim 1, characterized in that, The step of performing spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested includes: The corrected spectral data are multiplied by the corresponding weighting coefficients to obtain the weighted spectral data. All the weighted spectral data are summed to obtain the spectral weighted sum; The weight normalization coefficients are obtained by summing all the weight coefficients. Calculate the ratio of the spectral weighted sum to the weighted normalization coefficient, and determine the resulting ratio as the target optimized spectrum of the sample to be tested.

7. The optimization method based on spectral data according to claim 1, characterized in that, After performing spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested, the process includes: The target optimized spectrum is used as a dynamic reference spectrum; The similarity between each of the corrected spectral data and the dynamic reference spectrum is calculated to obtain multiple similarity values; Calculate the average similarity of all similarity values ​​and determine whether the average similarity meets the preset average similarity threshold; If the condition is not met, the abnormal score result of each of the corrected spectral data is updated based on the corresponding similarity value, and the process jumps to the step of determining the weight coefficient corresponding to each of the corrected spectral data based on the abnormal score result, until the average similarity meets the average similarity threshold, and the current target optimized spectrum is taken as the final optimized spectrum.

8. An optimization system based on spectral data, characterized in that, include: The data acquisition module is used to acquire multiple raw spectral data of the sample to be tested; The data correction module is used to perform data correction on the multiple original spectral data to obtain multiple corrected spectral data; An anomaly scoring module is used to extract spectral features from each of the corrected spectral data, and calculate anomaly scores based on the extracted spectral features to obtain anomaly score results corresponding to each of the corrected spectral data. The weight determination module is used to determine the weight coefficients corresponding to each of the corrected spectral data based on the anomaly scoring results. The spectral optimization module is used to perform spectral optimization processing on each of the corrected spectral data based on the weighting coefficients to obtain the target optimized spectrum of the sample to be tested.

9. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described in any one of claims 1 to 7, based on the optimization method for spectral data.

10. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and execute the optimization method based on spectral data as described in any one of claims 1 to 7.