A self-supervised one-dimensional spectrum denoising method based on redundant observation construction
By constructing redundant observation pairs and feature fidelity rules, a self-supervised one-dimensional spectral denoising method is developed, which solves the problem of distinguishing noise from weak peaks in existing technologies and achieves accurate denoising and efficient analysis in spectral analysis.
Patent Information
- Application Number
- CN202511768968.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-11-28
AI Technical Summary
Existing spectral analysis methods struggle to effectively distinguish noise from weak peaks in unlabeled data, leading to excessive suppression of weak characteristic peaks. Furthermore, uniform denoising across the entire spectrum results in peak distortion or residual noise, and the lack of a structured strategy for redundant observations leads to unstable denoising performance.
By acquiring two-dimensional spectral data of the target substance, defining characteristic peak regions and constructing redundant sub-observation pairs, and performing calibration processing in conjunction with characteristic fidelity rules, a one-dimensional spectrum is generated. The validity of the results is verified through qualitative and quantitative analysis, and the redundant construction strategy library is dynamically optimized to adapt to different sample characteristics.
It achieves accurate denoising in unlabeled data, retains strong peaks, enhances weak peaks, and suppresses noise, thereby improving the accuracy and adaptability of spectral analysis and ensuring stable denoising effect and analysis accuracy in multi-sample scenarios.
Smart Images

Figure CN121233906B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spectral analysis technology, and more specifically, to a self-supervised one-dimensional spectral denoising method based on redundant observation construction. Background Technology
[0002] Spectral analysis is a key tool for material analysis and is widely used in many fields. However, spectral signals are easily affected by noise and interference, resulting in the distortion of strong peaks and the submergence of weak peaks. Traditional denoising methods are difficult to balance noise suppression and feature preservation, and efficient denoising under unlabeled data has become a bottleneck.
[0003] Various spectral denoising methods have emerged in the prior art, but these methods have some limitations. Supervised denoising methods rely on a large number of clean-noise spectral pairs as labeled data, making it difficult to adapt to multi-sample scenarios. Unsupervised denoising methods are based on a single observation signal and cannot effectively distinguish weak peaks from noise, which can easily lead to excessive suppression of weak feature peaks. At the same time, most existing methods perform uniform denoising on the entire spectrum without distinguishing regional characteristics, resulting in peak distortion or noise residue. In addition, although some methods introduce redundant observations, they do not construct a structured redundancy strategy, which cannot guarantee the signal consistency and noise independence of redundant sub-observation pairs, resulting in unstable denoising effects and no dynamic optimization, and poor adaptability. Summary of the Invention
[0004] To overcome the aforementioned shortcomings of existing technologies and achieve the above objectives, this invention provides the following technical solution: a self-supervised one-dimensional spectral denoising method based on redundant observation construction, comprising:
[0005] S1: Acquire the original spectral signal of the target substance, collect two-dimensional spectral data containing characteristic peaks, and simultaneously record metadata of the substance sample and the environment to form a two-dimensional spectral dataset containing redundant observations;
[0006] S2: Based on a two-dimensional spectral dataset, characteristic peak regions are divided according to signal repeatability. The corresponding characteristic peak region construction strategy is matched with the redundancy construction strategy library to generate qualified redundant sub-observation pairs.
[0007] S3: Based on the characteristic peak region division results and redundant sub-observation pairs, merge them to generate one-dimensional spectral pairs, and obtain the denoised one-dimensional spectrum through calibration processing; combine with the feature fidelity rules for verification to obtain the feature fidelity verification report;
[0008] S4: Based on the denoised one-dimensional spectrum and feature fidelity verification report, qualitative analysis is completed by comparison with the standard spectral library, and quantitative analysis is completed by obtaining the substance concentration according to the peak intensity, forming qualitative and quantitative analysis results, and marking the validity of the analysis results;
[0009] S5: Based on the analysis results and validity indicators, a comprehensive evaluation result is obtained. When the expected results are not met, targeted strategy optimization is performed, the redundancy construction strategy library is updated and the adjustment process is recorded, and an optimized redundancy construction strategy library and closed-loop iteration record are generated.
[0010] Furthermore, the two-dimensional spectral dataset is formed in the following ways:
[0011] By using a suitable spectroscopic device, a horizontal channel covering the wavelength range of the characteristic peak is set for the target substance sample, and a vertical redundant observation parameter is set to make the same characteristic peak appear repeatedly in multiple vertical rows.
[0012] According to the set horizontal and vertical dimensions, collect two-dimensional spectral data of the target substance containing characteristic peaks; simultaneously monitor data quality in real time, and record substance sample information, environmental parameters and equipment parameters as metadata;
[0013] The two-dimensional spectral data is correlated and verified with the metadata to form a two-dimensional spectral dataset containing redundant observations.
[0014] Furthermore, the method of dividing the characteristic peak region according to signal repeatability includes:
[0015] By calling a two-dimensional spectral dataset, signal quality is enhanced through baseline correction and smoothing / denoising processes.
[0016] Furthermore, based on the repeatability of the signal in the vertical rows, for each horizontal channel, the signal mean and baseline in all vertical rows are obtained, and channels with signal mean significantly stronger than the baseline are selected as potential strong feature peak regions.
[0017] In the potential characteristic peak region, strong characteristic peak regions are first defined by high frequency of occurrence and stable peak position; then, signals with medium frequency of occurrence and weaker intensity than strong characteristic peaks are selected around the strong characteristic peak regions to define characteristic peak regions; the remaining regions without obvious peak shapes are defined as background regions.
[0018] Furthermore, the generation method of the redundant sub-observation pairs includes:
[0019] Based on the feature region segmentation results, a pre-set redundant construction strategy library is invoked, and the optimal construction strategy is matched between strong feature peak regions and weak feature peak regions through the corresponding matching rules.
[0020] Sub-observation pairs are generated according to a preset number of sub-observation pairs. The original two-dimensional spectral data is cut according to the matching construction strategy, and then sub-observation fragments of each region are extracted and combined into complete redundant sub-observation pairs.
[0021] The generated redundant sub-observation pairs are checked for signal consistency and noise independence. If the check fails, a retry mechanism is triggered to regenerate and check again until a qualified redundant sub-observation pair is obtained.
[0022] Furthermore, the method for obtaining the denoised one-dimensional spectrum includes:
[0023] Based on the feature region division results and qualified redundant sub-observation pairs, each group of redundant sub-observation pairs is first merged vertically along the horizontal channel to generate a one-dimensional spectral pair.
[0024] Calibration processing is performed on different regions of one-dimensional spectral alignment. Specifically, peak position alignment and peak height calibration are performed on strong characteristic peak regions, and the blurred signal is enhanced by cross-spectral comparison on weak characteristic peak regions. Background regions are smoothed and denoised.
[0025] Then, following the principle of selecting the best, the different regions after calibration are merged to obtain the denoised one-dimensional spectrum.
[0026] Furthermore, the method for generating the feature fidelity verification report includes:
[0027] The verification logic of the feature fidelity rule is defined as follows: for strong feature peak areas, verify the peak position shift, peak shape distortion and peak height rationality; for weak feature peak areas, verify whether weak peaks are lost; and for background areas, verify whether noise is effectively suppressed and clearly isolated from feature peaks.
[0028] Based on the denoised one-dimensional spectrum, each region is verified according to the verification logic of the feature fidelity rule. If all verifications pass, the feature fidelity verification is deemed to have passed.
[0029] If any verification fails, the feature fidelity verification is deemed to have failed. For the spectrum that failed verification, the calibration process for the one-dimensional spectrum pair is triggered in reverse and reprocessed until the feature fidelity verification passes.
[0030] The verification results from each region and the final feature fidelity determination results are integrated to generate feature fidelity verification results.
[0031] Furthermore, the qualitative and quantitative analysis results are formed in the following ways:
[0032] The qualitative analysis logic is defined as follows: Based on the denoised one-dimensional spectrum that has passed the feature fidelity verification, combined with the standard spectral library of the substance, all peaks in the strong feature peak region and the weak peaks that must be retained in the weak feature peak region are extracted to form a list of peaks to be compared.
[0033] Based on the list of peaks to be compared, the denoised one-dimensional spectrum is compared with the standard spectral library in multiple dimensions to obtain the characteristic peak matching degree. The corresponding substance name is matched by the characteristic peak matching degree as the qualitative analysis result.
[0034] The quantitative analysis logic is defined as follows: based on the denoised one-dimensional spectrum, select highly applicable quantitative characteristic peaks in the strong characteristic peak region, extract the peak intensity, and then combine the peak intensity-concentration standard curve of the target substance to obtain the concentration of the target substance, which is used as the quantitative analysis result.
[0035] Furthermore, the methods for identifying the validity of the labeled analysis results include:
[0036] The validity of the qualitative and quantitative analysis results is determined separately. If the results pass the test, they are marked as valid candidates; otherwise, they are marked as invalid candidates.
[0037] The method for determining the validity of qualitative analysis results is as follows:
[0038] If the qualitative matching degree is greater than or equal to the qualitative validity threshold and there are no missing key peaks, the qualitative analysis result is initially determined to be a valid candidate; if the matching degree is less than the qualitative validity threshold, or there are missing key peaks, the qualitative analysis result is marked as an invalid candidate.
[0039] The method for determining the validity of quantitative analysis results is as follows:
[0040] If the concentration calculation error is less than or equal to the quantitative validity threshold, and the peak intensity of the quantitative characteristic peak is within the linear range of the standard curve, the quantitative result is initially determined to be a valid candidate; if the error is greater than the quantitative validity threshold, or the peak intensity of the quantitative characteristic peak exceeds the linear range, it is marked as an invalid candidate.
[0041] If both the qualitative and quantitative analysis results are marked as valid candidates, then the analysis results are marked as valid by adding a comprehensive validity label. If any invalid candidate label exists, then the analysis results are marked as invalid by adding a comprehensive validity label.
[0042] Furthermore, the comprehensive evaluation results are obtained in the following ways:
[0043] The denoising effectiveness is evaluated based on the one-dimensional spectra before and after denoising, through the signal-to-noise ratio improvement, weak peak retention, and peak shape fidelity preservation effects.
[0044] Based on the matching construction strategy, the utility of the redundancy strategy is evaluated through the adaptability of the redundancy construction strategy and the redundancy gain.
[0045] The analytical results that are identified as valid based on the overall validity are recorded as valid analytical results. The qualitative accuracy rate of the qualitative analysis and the quantitative average error of the quantitative analysis are statistically analyzed based on the valid analytical results to evaluate the analytical utility.
[0046] The denoising utility, redundancy strategy utility, and analysis utility are combined to generate a comprehensive evaluation result.
[0047] Furthermore, the optimized redundant construction strategy library is generated in the following ways:
[0048] If the overall evaluation results do not meet expectations, the problem category will be identified through the problem classification and positioning mechanism, and targeted strategy optimization will be carried out.
[0049] The optimized parameters and construction strategies are written into the redundant construction strategy library and dynamic adaptation rules are generated. The optimized redundant construction strategy library is generated, and closed-loop iteration records are generated simultaneously.
[0050] The technical effects and advantages of the self-supervised one-dimensional spectral denoising method based on redundant observation construction proposed in this invention are as follows:
[0051] This invention acquires the original spectral signal of the target substance, collects two-dimensional spectral data and records metadata simultaneously, forming a two-dimensional spectral dataset with redundant observations. This provides reliable basic data for subsequent processes and ensures the data reliability of the entire technical process.
[0052] Secondly, based on the two-dimensional spectral dataset, feature regions are divided according to signal repeatability and targeted redundancy strategies are matched to generate qualified sub-observation pairs. This regional processing method accurately identifies the characteristics of different feature regions and solves the problem that single observations are difficult to distinguish between noise and weak peaks.
[0053] Next, one-dimensional spectral pairs are generated by merging redundant sub-observations. After calibration, a denoised one-dimensional spectrum is obtained. At the same time, a report is generated by verifying the feature fidelity rule. This step achieves accurate denoising in different regions, retains strong peaks, enhances weak peaks, and suppresses noise, overcoming the limitations of uniform processing of the entire spectrum.
[0054] Then, based on the denoised one-dimensional spectrum and feature fidelity verification report, qualitative and quantitative analysis was completed, and validity labels were added, which improved the accuracy of the substance analysis. The validity labels provided a clear basis for judging the reliability of the analysis results and enhanced the credibility of the results.
[0055] Finally, by evaluating the overall effectiveness, the solution is dynamically optimized through evaluation and strategy iteration when expectations are not met. This dynamic optimization mechanism can continuously adapt to different sample characteristics and analysis scenarios, significantly improving the adaptability and effectiveness of the overall solution and ensuring stable denoising effect and analysis accuracy in multi-sample analysis. Attached Figure Description
[0056] Figure 1 This is a flowchart illustrating a self-supervised one-dimensional spectral denoising method based on redundant observation construction according to the present invention.
[0057] Figure 2 This is a schematic diagram of the process for generating a denoised one-dimensional spectrum in a self-supervised one-dimensional spectral denoising method based on redundant observation construction according to the present invention.
[0058] Figure 3 This is a schematic diagram of the process of a self-supervised one-dimensional spectral denoising system based on redundant observation construction according to the present invention.
[0059] Figure 4 This invention relates to a self-supervised one-dimensional spectral denoising method based on redundant observation construction, specifically an image of the Raman signal of *E. coli*.
[0060] Figure 5 This is a Raman spectral image of Escherichia coli, based on a self-supervised one-dimensional spectral denoising method constructed using redundant observations, according to the present invention.
[0061] Figure 6 This is an odd-row projection image of the Escherichia coli Raman signal, based on a self-supervised one-dimensional spectral denoising method constructed with redundant observations according to the present invention.
[0062] Figure 7 This is an even-row projection image of the Escherichia coli Raman signal, based on a self-supervised one-dimensional spectral denoising method constructed with redundant observations according to the present invention.
[0063] Figure 8 This image shows the denoising effect of Raman spectroscopy on Escherichia coli using a self-supervised one-dimensional spectral denoising method based on redundant observation construction according to the present invention. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0065] Please see Figure 1 and Figure 2 As shown in this embodiment, a self-supervised one-dimensional spectral denoising method based on redundant observation construction includes:
[0066] S1: Acquire the original spectral signal of the target substance, collect two-dimensional spectral data containing characteristic peaks, and simultaneously record metadata of the substance sample and the environment to form a two-dimensional spectral dataset containing redundant observations;
[0067] S2: Based on a two-dimensional spectral dataset, characteristic peak regions are divided according to signal repeatability. The corresponding characteristic peak region construction strategy is matched with the redundancy construction strategy library to generate qualified redundant sub-observation pairs.
[0068] S3: Based on the characteristic peak region division results and redundant sub-observation pairs, merge them to generate one-dimensional spectral pairs, and obtain the denoised one-dimensional spectrum through calibration processing; combine with the feature fidelity rules for verification to obtain the feature fidelity verification report;
[0069] S4: Based on the denoised one-dimensional spectrum and feature fidelity verification report, qualitative analysis is completed by comparison with the standard spectral library, and quantitative analysis is completed by obtaining the substance concentration according to the peak intensity, forming qualitative and quantitative analysis results, and marking the validity of the analysis results;
[0070] S5: Based on the analysis results and validity indicators, evaluate the denoising utility, redundancy strategy utility and analysis utility to obtain a comprehensive evaluation result. When the expected results are not met, perform targeted strategy optimization, update the redundancy construction strategy library and record the adjustment process, and generate an optimized redundancy construction strategy library and closed-loop iteration record.
[0071] The methods for forming two-dimensional spectral datasets include:
[0072] Select the appropriate spectroscopic equipment for the target substance analysis based on the type of substance (e.g., micro Raman spectrometer for analyzing organic pollutants, fluorescence spectrophotometer for analyzing heavy metal fluorescence), and ensure that the equipment supports area array sensors.
[0073] Adjust the grating and detector of the device according to the characteristic peak wavelength range of the target substance sample to ensure that the lateral channel can cover the entire peak region (for example, assuming the Raman peak of the antibiotic is 1600 cm⁻¹, set the lateral channel to cover 1500-1700 cm⁻¹, the number of channels ≥200, and the wavelength interval of a single channel ≤1 cm⁻¹), and then set the device to continuous scanning mode to ensure that there are no breaks in the lateral channel;
[0074] In setting the longitudinal redundancy observation parameters, a large number of rows needs to be set in the longitudinal dimension to ensure that the same characteristic peak appears repeatedly in multiple rows. Therefore, the minimum number of rows in the longitudinal dimension should be ≥30, and the default can be set to 50 rows to ensure redundancy. Each row corresponds to one repeated observation, and the row interval should be as small as possible (e.g., ≤10μm) to avoid signal jumps caused by sample inhomogeneity. The exposure time should be adjusted according to the sample signal intensity (e.g., 50-100ms for strong signal samples and 500-1000ms for weak signal samples) to ensure that the signal intensity of a single row is ≥10 times that of the baseline (to avoid excessively low signal-to-noise ratio). The exposure time of the same batch of samples should be kept consistent. The laser power should be set according to the sample type. For photosensitive samples (e.g., biological tissue), the setting should be as small as possible (e.g., ≤5mW), and for stable samples, the setting should be slightly larger (e.g., 10-20mW) to avoid sample burning.
[0075] Two-dimensional spectral data containing characteristic peaks of the target substance are collected according to the set horizontal and vertical dimensions; the data quality is monitored in real time simultaneously, and the substance sample information (including sample identification (e.g., drug-batch number-concentration), pretreatment method (e.g., grinding + quartz slide, solid phase extraction enrichment), sample status (e.g., no precipitation + no bubbles, uniform particles)), environmental parameters (temperature, humidity and air pressure at the time of collection) and equipment parameters (including laser wavelength, power, exposure time, horizontal channel range and vertical row number) are recorded as metadata;
[0076] Two-dimensional spectral data is saved in TIFF format (preserving the original pixel values), and metadata is saved in JSON format. The two are linked by a unique number.
[0077] Then, the integrity of the TIFF file is verified (checking whether the TIFF file size conforms to the vertical number of rows (e.g., 50 rows) × horizontal number of channels (e.g., 200 columns), ensuring no missing pixels), and the consistency of the metadata (whether the acquisition parameters in the metadata are consistent with the device log records, and whether the sample identifier matches the actual data); thus forming a verified two-dimensional spectral dataset with redundant observations (containing TIFF format spectral images and associated JSON metadata).
[0078] Methods for dividing characteristic peak regions based on signal repeatability include:
[0079] Call the two-dimensional spectral dataset, load the spectral images and metadata through spectral analysis software, and ensure that the horizontal channels (spectral wavelengths) correspond correctly to the vertical rows (redundant observations);
[0080] Then, signal quality is enhanced through baseline correction and smoothing / denoising processes.
[0081] The baseline correction involves automatically correcting the original two-dimensional spectrum (which can be achieved using 3rd-5th order polynomial fitting) to eliminate continuous noise interference such as fluorescence background. For each horizontal channel (single wavelength), the mean signal value of the vertical row is calculated as the baseline. Then, the baseline is subtracted from the original signal to obtain the corrected data matrix (the matrix dimension is: vertical row × horizontal channel).
[0082] The smoothing and denoising process involves: lightly smoothing the corrected data matrix (which can be achieved using Savitzky-Golay filtering, with the window size set to 5×5 and the polynomial order set to 2), thereby preserving the characteristic peak shape while suppressing high-frequency random noise and preventing subsequent peak identification from being falsely triggered by noise;
[0083] For the two-dimensional spectral data after signal quality enhancement, based on the repeatability of the signal in the vertical rows, for each horizontal channel, the signal mean and baseline in all vertical rows are calculated (the baseline is the signal mean in the background area, which can be obtained from the peakless areas marked in the metadata).
[0084] Channels with a signal mean significantly stronger than the baseline are selected from all horizontal channels (this can be achieved through the 3σ principle, i.e., channels with a signal mean > 3 times the baseline are marked as channels significantly stronger than the baseline) as potential strong characteristic peak regions;
[0085] In the region of potential strong characteristic peaks, areas with high frequency and stable peak positions are defined as strong characteristic peak regions. The purpose is to identify the core characteristic peaks for qualitative analysis of substances, ensuring their stable appearance and significant signal in redundant observations. Specifically:
[0086] For signals within the potential characteristic peak region, count their frequency of occurrence in longitudinal redundant observation (i.e., count the number of times the peak appears in the longitudinal row and record the proportion of the number of occurrence rows to the total number of longitudinal rows as the frequency of occurrence). If a certain peak signal can be stably identified in most longitudinal rows (one way to judge is to repeatedly identify it as a peak by visual inspection or analysis software, or to directly identify peaks with an occurrence frequency ≥ the occurrence frequency threshold (e.g., 80%) as stably identified), and the lateral position (wavelength) deviation of the peak is extremely small (i.e., the peak position deviation is within the acceptable range of instrument error, for example, the peak position deviation ≤ 1), then it is marked as a candidate strong peak.
[0087] Based on the lateral center position of the candidate strong peak, extend to both sides to the baseline where the peak shape disappears (including the rising edge, peak value, and falling edge of the peak), and vertically cover all redundant observation rows where the peak appears to form a continuous strong characteristic peak region, which is marked in the spectrum.
[0088] Based on the defined strong feature peak regions, the weak feature peak regions and background regions are defined as follows:
[0089] Within the horizontal region surrounding the strong characteristic peak region (one way to define the horizontal region is based on the range of weak peaks that may appear as predicted by the material properties, and other ways, such as the wavelength range of low concentration characteristic peaks marked in the metadata), channels with signal values greater than the baseline and less than the average value of the strong characteristic peak signal are selected (to avoid confusion with strong peak signals).
[0090] For the selected signals, their frequency of occurrence in the vertical rows is counted. If a signal is identifiable in some vertical rows (distinguished from baseline fluctuations), but its frequency of occurrence is lower than that of strong peaks (not reaching the proportion of strong peak occurrences), it is marked as a candidate weak peak.
[0091] Based on the lateral center position of the candidate weak peak, extend to both sides to the channel where the signal falls back to the baseline (the range is smaller than the strong characteristic peak region to avoid excessive noise inclusion), and vertically cover the redundant observation line where the weak peak appears to form the weak characteristic peak region, which is marked on the spectrum with a specific marker.
[0092] From the complete spectral region, remove the defined strong and weak characteristic peak regions, and use the remaining region as the candidate range for the background region;
[0093] If the signal in the vertical row of the candidate range of the background region is observed, and the signal has no obvious peak shape and exhibits random fluctuations in the vertical row (without repeated peak structures), it is determined to be the background region. The signal fluctuation in this region can represent the noise level in the spectrum.
[0094] Continuous peakless regions are merged into background regions and also marked in the spectrum to ensure that they cover a sufficient horizontal and vertical range (including at least one complete horizontal channel segment and ≥10 rows of vertical observations) to meet the needs of subsequent noise assessment.
[0095] After verifying and manually reviewing each of the divided regions, a structured feature region division result is generated, specifically:
[0096] The verification process involves using software to calculate whether the boundaries of each region overlap (e.g., when the horizontal ranges of strong and weak feature peak regions overlap, the strong feature peak region is retained first) to ensure that the region division is unique; the proportion of strong feature peak regions in the total horizontal channels is statistically analyzed, and it is best to be ≤30% to avoid excessive expansion of strong feature peak regions, which would lead to redundant construction and waste of resources.
[0097] Manual verification involves visualizing the classification results (drawing a spectral heatmap and marking strong characteristic peak areas, weak characteristic peak areas, and background areas with different colors). Spectroscopic analysts confirm whether the strong characteristic peak areas contain known core characteristic peaks of the substance (e.g., checking the position of strong peaks in the standard spectral library based on the substance type in the metadata), whether the weak characteristic peak areas cover potential trace characteristic peak areas, and whether the background areas have no obvious peak shape interference. If necessary, the region boundaries can be manually fine-tuned (e.g., including missed weak characteristic peak areas in the classification).
[0098] Redundant sub-observation pairs can be generated in the following ways:
[0099] Based on the feature region segmentation results and associated metadata, key parameters are extracted, including the number of horizontally covered channels and the percentage of vertically occurring rows in the strong feature peak region, the horizontal position of the weak feature peak region (distance relative to the strong feature peak region), the ratio of signal intensity to strong peak, and sample characteristics (high concentration / low concentration, organic / inorganic labels, etc.).
[0100] The pre-defined redundancy construction strategy library includes basic strategies (including strong feature peak area strategy set, weak feature peak area strategy set, and low redundancy remediation strategy set) and adaptation rules for different regions.
[0101] The strong feature peak region strategy set includes odd-even row partitioning, upper and lower half region partitioning, and backup strategy (interval row sampling). The core requirement for strong feature peak regions is large granularity and no disruption to the peak structure. The weak feature peak region strategy set includes local window partitioning, continuous multi-row merging, and backup strategy (random row sampling). The core requirement for weak feature peak regions is small granularity and focusing on weak peak regions. The low redundancy remediation strategy set includes cross-sample redundancy supplementation and time-dimensional observation overlay (for scenarios where the number of vertical rows may be insufficient).
[0102] The pre-built redundant construction strategy library is invoked, and the optimal construction strategy is matched between strong feature peak regions and weak feature peak regions according to the corresponding matching rules. Specifically:
[0103] For strong characteristic peak regions: if the number of horizontal channels in a strong characteristic peak region is less than or equal to the preset channel number threshold (e.g., 50), it indicates that the peak shape is simple. Odd-even row division is preferred (the vertical row is divided into two groups of sub-observation segments according to the odd and even numbers, so as to ensure that the strong peak signal is consistent in the two groups); if the number of horizontal channels in a strong characteristic peak region is greater than the channel number threshold (it indicates that the peak shape is complex), upper and lower half region division is preferred (divided according to the midpoint of the vertical row, so as to avoid the peak structure being fragmented).
[0104] The matching rule for strong characteristic peak regions is defined as follows: ensure that the peak position deviation of the strong characteristic peak regions of the two groups of sub-observation segments after division is ≤1 horizontal channel, and the peak height ratio is within the expected range (e.g., within 0.9-1.1).
[0105] For weak feature peak regions: Set a channel distance threshold (i.e., the number of channels between them, such as 20). If the horizontal channel between a weak feature peak region and a strong feature peak region is less than or equal to the channel distance threshold (e.g., the weak feature peak region is located within ±20 horizontal channels of the strong feature peak region, indicating that the weak feature peak is adjacent to the strong feature peak), select local window division (window size = number of horizontal channels in the weak feature peak region × 2, only splitting the vertical row where the weak feature peak is located); if the horizontal distance between a weak feature peak region and a strong feature peak region is greater than the channel distance threshold (e.g., horizontal > 20 channels), select random row sampling (randomly select 2 groups from the vertical rows, the number of rows in each group is greater than or equal to a certain proportion (e.g., 80%) of the number of rows where the weak feature peak region appears, to improve the diversity of weak peak observations).
[0106] The matching rule for weak feature peak regions is defined as follows: ensure that the signal intensity of the weak feature peak is ≥ 1.5 times that of the baseline in at least one set of sub-observations (i.e., ensure that it is identifiable).
[0107] If the original number of vertical rows is less than a certain number (e.g., <30), it is judged to be insufficient in redundancy, and the metadata is labeled as a low-concentration sample (i.e., weak feature peak signal). At this time, it is judged to be a low-redundancy scenario, and the cross-sample redundancy supplementation mechanism is automatically triggered. That is, the redundant rows of high-concentration material samples of the same batch and type are called through the low-redundancy remediation strategy set (ensuring that strong feature peak areas and weak feature peak areas have been extracted and labeled), and merged with the vertical rows of the current material sample before being divided to supplement the observation data of weak feature peak areas.
[0108] Sub-observation pairs are generated according to the preset number of sub-observation pairs. By default, 2 sets of sub-observation pairs (P1, P2) are generated. If the signal in the weak characteristic peak region is extremely weak (i.e., the signal strength is less than 5% of the signal strength of the strong characteristic peak), then 3 sets of sub-observation pairs (P1, P2, P3) are generated to increase redundancy.
[0109] The logic for handling boundaries is defined as follows: when dividing, it is necessary to ensure that the vertical rows of strong feature peak areas and weak feature peak areas are not split (for example, if a strong feature peak appears in vertical row 10-20, then the two sets of sub-observations after the division must completely include the row of that interval).
[0110] The original two-dimensional spectral data is cut according to the matching construction strategy. For example, when the strong characteristic peak region is divided into odd and even rows, P1 contains vertical rows 1, 3, 5... and P2 contains vertical rows 2, 4, 6...; when the weak characteristic peak region is divided into local windows, the rows where the weak characteristic peak appears (such as rows 15-25) are extracted from the vertical rows and split into P1_weak and P2_weak.
[0111] Then, the sub-observation segments of the strong feature peak region, the weak feature peak region and the background region are combined to form a complete redundant sub-observation pair (P1, P2). It is necessary to ensure that each redundant sub-observation pair contains all feature regions.
[0112] The generated redundant sub-observation pairs are checked for signal consistency and noise independence.
[0113] One method for performing signal consistency verification includes:
[0114] For strong feature peak regions: compare the peak positions (lateral channel difference ≤ 1) and peak shapes of strong feature peaks of P1 and P2 (through visual comparison or simple correlation calculation, the correlation coefficient needs to be ≥ 0.9).
[0115] For weak feature peak regions: check whether the number of rows in P1 and P2 where weak feature peaks appear is greater than or equal to a certain percentage (e.g., 70%) of the number of rows in the original weak feature peak region, to ensure that weak feature peaks are not lost.
[0116] One method for performing noise independence verification includes:
[0117] For the background region: Calculate the signal correlation between P1 and P2 in the background region (this can be done through a simple correlation calculation; the correlation coefficient needs to be less than the correlation threshold, such as ≤0.3), to ensure that there is no significant correlation between noise; if the correlation is too large (such as >0.3), then re-perform random row sampling (adjust the sampling seed) until the correlation coefficient meets the requirements;
[0118] If the verification fails, a retry mechanism is triggered. That is, a backup strategy is called from the redundancy construction strategy library (such as switching from odd-even row partitioning to interval row sampling in strong feature peak regions), and the redundancy sub-observation pairs are regenerated and verified until a qualified redundant sub-observation pair is obtained.
[0119] Methods for obtaining the denoised one-dimensional spectrum include:
[0120] Based on the feature region segmentation results and qualified redundant sub-observation pairs, data is loaded using spectral processing tools to ensure that the horizontal channels (wavelengths) and vertical rows (redundant observations) of the redundant sub-observation pairs are aligned with the original data. Then, the vertical row signals in the redundant sub-observation pairs are checked. If the signal of a certain row is completely offset (e.g., the peak height of the strong feature peak region is less than one-third of that of other rows), it is judged as an abnormal row and removed (retaining ≥80% of the original rows) to avoid abnormal data interfering with the merging results.
[0121] For each redundant sub-observation pair (P1, P2), vertical merging is performed according to the horizontal channel (wavelength point). For each horizontal channel c, the signal mean at c of all vertical rows in P1 is calculated to obtain the one-dimensional spectrum s1 of P1. Similarly, the one-dimensional spectrum s2 of P2 is calculated to form a one-dimensional spectrum pair (s1, s2).
[0122] Each set of redundant sub-observations is vertically merged along the horizontal channel to generate a one-dimensional spectral pair. Specifically:
[0123] Regarding the parameter settings during merging: For strong characteristic peak regions, a weighted average method is used during merging (the higher the peak height of a row, the greater the weight, and the weight value is equal to the ratio of the peak height of that row to the sum of the peak heights of all rows), thereby enhancing the stability of strong characteristic peak signals; for weak characteristic peak regions, an arithmetic average method is used to avoid weight bias from masking weak peak signals; for background regions, a median average method is used to reduce the impact of extreme noise points.
[0124] Then, calibration is performed on different regions of the one-dimensional spectrum. Specifically:
[0125] For strong characteristic peak regions, peak position alignment and peak height calibration are performed. Peak position alignment is as follows: compare the peak positions (lateral channels) of strong characteristic peak regions in s1 and s2. If the deviation is ≤1 channel (i.e., it meets the signal consistency requirements), proceed directly to peak height calibration; if the deviation is >1 channel, use the standard peak position marked in the metadata as the reference and fine-tune the spectrum with larger deviation (e.g., if the peak position of s1 is too far to the left, shift the whole spectrum to the right by 1 channel).
[0126] Peak height calibration is performed as follows: Calculate the peak height ratio of the same strong peak in s1 and s2, that is, the ratio of the peak height of s1 to the peak height of s2; preset the peak height ratio threshold range (here, it is set to 0.8-1.2 for example). If 0.8 ≤ peak height ratio ≤ 1.2 (indicating that the difference is small), keep the peak heights of both unchanged; if the peak height ratio < 0.8 (indicating that the peak height of s1 is too low), correct s1 using the peak height trend of s2 (e.g., peak height of s1 = peak height of s1 × k, and through the coefficient k, make the new peak height ratio ≥ 0.8).
[0127] The coefficient k is the minimum adjustment coefficient (not a fixed value). The specific logic is that when the original peak height ratio is <0.8 (e.g., peak height ratio = 0.7, meaning the peak height of s1 is only 70% of that of s2), the peak height of s1 needs to be amplified by the coefficient k so that the new peak height ratio = (s1 peak height × k) / s2 peak height ≥ 0.8. Substituting the peak height ratio = 0.7, we get: 0.7 × k ≥ 0.8 × s2 peak height → k ≥ (0.8 × s2 peak height) / 0.7 ≈ 1.14. Therefore, the coefficient k here is 1.14 (in practical applications, k needs to be calculated based on the specific peak height ratio to ensure that the peak height difference after correction is ≤20%).
[0128] If the peak height ratio is >1.2 (s2 peak height is too low), s1 is used to correct s2 in the same way, so that the difference in peak height of strong peaks is controlled within a certain proportion (such as 20%).
[0129] For weak characteristic peak regions, the blurred signal is enhanced by cross-spectral comparison. Specifically, in the weak characteristic peak regions of s1 and s2, weak peaks are identified by the baseline lifting method. That is, for each horizontal channel c, if the signal value is greater than the baseline mean + 1.5 times the baseline standard deviation, and it is peak-shaped (rising first and then falling) within the c±2 channels, it is marked as a weak peak.
[0130] Then, the weak peaks are strengthened and calibrated: if a weak peak signal in s1 is blurred (i.e., peak height < baseline mean + 2 times standard deviation), but a clear weak peak exists at the corresponding position in s2 (i.e., peak height ≥ baseline mean + 2.5 times standard deviation), then the shape of the weak peak in s2 is used to correct s1 (i.e., the baseline of s1 is retained, and the peak height is referenced to 70% of s2).
[0131] If the same weak peak in s1 and s2 is relatively blurry, but appears in ≥3 sets of redundant sub-observation pairs, then the average enhancement of the two is taken (i.e., the peak height of s1 after correction = (peak height of s1 + peak height of s2) / 2).
[0132] The background region is smoothed and denoised. Specifically, noise consistency is checked first by calculating the signal standard deviations σ1 and σ2 of s1 and s2 in the background region. If the ratio of σ1 to σ2 is within the range of 0.8-1.2 (i.e., the peak height ratio threshold range), it indicates that the noise level is consistent, and smoothing and denoising are performed. Otherwise, the noise independence of the sub-observation pair is rechecked (i.e., noise independence is checked again).
[0133] The smoothing and denoising process involves using a sliding window mean smoothing method on the background area. The window size should not be too small or too large. It can be set to 5 horizontal channels to avoid affecting the characteristic peak area. The background area signals of s1 and s2 are smoothed separately to suppress high-frequency random noise.
[0134] After calibration, s1 and s2 are merged according to the principle of selecting the best to obtain the denoised one-dimensional spectrum. In the strong characteristic peak region, the signal with the sharper peak shape (i.e., the one with a smaller full width at half maximum) between s1 and s2 is selected. In the weak characteristic peak region, the part of the signal with the stronger signal between the two is selected. In the background region, the average of the two is selected.
[0135] The methods for generating feature fidelity verification reports include:
[0136] The verification logic of the feature fidelity rule is defined as follows: for strong feature peak areas, verify the peak position shift, peak shape distortion and peak height rationality; for weak feature peak areas, verify whether weak peaks are lost; and for background areas, verify whether noise is effectively suppressed and whether it is clearly isolated from feature peaks.
[0137] The verification of peak position shift is defined as follows: locate all characteristic peaks in the strong characteristic peak region in the denoised one-dimensional spectrum and record their lateral channel positions; compare the peak positions with the reference strong peaks and calculate the shift (i.e., the difference between the denoised peak position and the reference peak position). If the absolute value of the shift is within the error range of the wavelength resolution of the spectrometer (e.g., the absolute value of the shift is ≤ 1 channel), the peak position is deemed qualified; otherwise, it is marked as peak position shift exceeding the standard.
[0138] The verification of peak shape distortion is defined as follows: Calculate the full width at half maximum (FWHM) of the strong peak in the denoised one-dimensional spectrum and compare it with the FWHM of the reference strong peak. If the deviation is ≤10% (e.g., reference FWHM = 5 channels, denoised FWHM ≤ 5.5 channels), the peak shape is determined to be undistorted; compare the rising edge, peak value, and falling edge trends of the peak using visual recognition tools. If they are consistent with the reference peak shape (no obvious flattening or sharpening), the peak shape verification is passed; otherwise, the peak shape is marked as distorted.
[0139] The verification definition of peak height rationality is as follows: if the peak height of the strong peak after denoising is less than a certain percentage (e.g., 70%) of the reference peak height, it indicates that denoising may be excessive; or if it is greater than a certain percentage (e.g., 130%) of the reference peak height, it indicates that noise may be retained, and the peak height is marked as abnormal. If the peak height is within the acceptable range of the reference (e.g., 70%-130%), and the deviation of the average peak height of the strong peak from the redundant sub-observation pair is ≤ a certain percentage (usually 20%), the peak height is deemed acceptable.
[0140] The logic for checking for weak peaks that must be retained is defined as follows: Traverse the weak characteristic peak region of the denoised one-dimensional spectrum and check whether all weak peaks that must be retained (i.e., weak peaks that appear in ≥3 groups in the original sub-observation pair) exist. If the signal value of a weak peak in the denoised one-dimensional spectrum is < the baseline mean + 1.5 times the standard deviation (indicating that it cannot be identified), it is marked as a lost weak peak that must be retained.
[0141] For any weak peaks that exist, record their peak positions and compare them with the benchmark weak peak positions. If the weak offset is ≤1 channel, the peaks are considered qualified.
[0142] Then, the integrity of the weak peak signal is verified by calculating the peak height ratio of the weak peak to the strong peak in the denoised one-dimensional spectrum. If the deviation of this ratio from the average ratio in the original sub-observation pair is ≤ a certain proportion (default is 30%, such as 10% in the original ratio and ≥7% after denoising), it is determined that the weak peak signal has not been excessively suppressed.
[0143] If a new peak (a peak not present in the original sub-observation pair) appears in the weak peak region, and the signal value is less than the baseline mean plus 2 times the standard deviation, it is determined to be noise introduced and marked as a false weak peak.
[0144] The verification method for whether the background noise is effective is as follows: calculate the standard deviation of the signal in the one-dimensional spectral background region after denoising, and compare it with the average standard deviation of the background region of the original sub-observation. If the standard deviation of the signal in the background region is less than or equal to a certain percentage of the average standard deviation (usually 50%), it indicates that the noise is effectively suppressed and the background region is deemed qualified. If the signal in the background region shows periodic fluctuations (indicating that noise from the denoising algorithm may have been introduced), it is marked as abnormal fluctuation in the background region.
[0145] The verification method for whether the characteristic peak is clearly isolated is as follows: check the difference between the peak bottom of a strong or weak peak and the background signal. If the difference is greater than or equal to a certain percentage of the peak height (usually 50%), it indicates that the peak and background boundary is clear and the isolation is qualified; otherwise, it is marked as peak and background confusion.
[0146] Based on the denoised one-dimensional spectrum, each region is verified according to the verification logic of the feature fidelity rule. If all verifications pass, the feature fidelity verification is deemed to have passed.
[0147] If any verification fails, the feature fidelity verification is deemed to have failed. For the spectrum that failed verification, the calibration process for the one-dimensional spectrum pair is triggered in reverse and reprocessed until the feature fidelity verification passes.
[0148] The verification results from each region (including strong peak verification results, weak peak verification results, and background region verification results) and the final feature fidelity determination results are integrated to generate feature fidelity verification results.
[0149] The methods for forming qualitative and quantitative analytical results include:
[0150] The qualitative analysis logic is defined as follows: Based on the denoised one-dimensional spectrum that has passed the feature fidelity verification, combined with the standard spectral library of the substance (i.e., the standard spectral library obtained through laboratory experiments, classified according to the type of substance, such as organic compounds, inorganic ions, etc., each standard spectrum contains parameters such as characteristic peak position, peak height ratio, and full width at half maximum), all peaks in the strong characteristic peak region and the weak peaks that must be retained in the weak characteristic peak region are extracted from the denoised one-dimensional spectrum, and the wavelength (wavelength value corresponding to the horizontal channel), peak height, and full width at half maximum of each peak are recorded to form a list of peaks to be compared.
[0151] Based on the list of peaks to be compared, the denoised one-dimensional spectrum is compared with the standard spectral library in multiple dimensions to obtain the characteristic peak matching degree. The corresponding substance name is then matched using the characteristic peak matching degree as the qualitative analysis result. Specifically:
[0152] Candidate spectra are screened from standard spectral libraries according to the type of substance sample (e.g., when analyzing water samples, common water pollutant spectral libraries are searched first). Peak position, peak height ratio, and peak shape are compared using a multi-peak joint matching algorithm. Based on the matching results of peak position, peak height ratio, and peak shape, a weighted fusion is performed to calculate the total matching degree (the corresponding weights can be set according to experience, such as 40% for peak position, 30% for peak height ratio, and 30% for peak shape). If the total matching degree meets the expectation (e.g., ≥90%), and the characteristic peaks of the corresponding candidate substance are all reflected in the denoised one-dimensional spectrum (no key peaks are missing), it is judged as a successful qualitative match, and the name of the substance and the corresponding matching degree are recorded. If the total matching degree does not meet the expectation (e.g., <90%), or there are two or more missing key peaks, it is marked as a failed qualitative match, and the three most likely candidate substances and their matching degrees are recorded.
[0153] Peak matching is defined as follows: the wavelength deviation between each peak in the denoised one-dimensional spectrum and the corresponding peak in the candidate material spectrum is calculated, and if the expected value is met (e.g., ≤1cm⁻¹), it is considered a match.
[0154] Peak height ratio matching is: calculate the peak height ratio of any two strong peaks in the denoised one-dimensional spectrum and compare it with the ratio deviation of the corresponding peaks in the candidate material spectrum. If it meets the expectation (e.g., ≤15%), it is considered a match.
[0155] Peak shape matching is defined as follows: the overall peak shape trend is compared using the Dynamic Time Warping (DTW) algorithm, and a similarity greater than a certain proportion (e.g., ≥85%) is considered a match.
[0156] The quantitative analysis logic is defined as follows: Based on the denoised one-dimensional spectrum, select highly applicable quantitative characteristic peaks in the strong characteristic peak region, extract the peak intensity, and then combine it with the peak intensity-concentration standard curve of the target substance to obtain the concentration of the target substance, which is used as the quantitative analysis result. Specifically:
[0157] From the mass spectra that were successfully matched by qualitative analysis, candidate quantitative characteristic peaks were screened. Specifically, peaks with sharp peak shapes (small half-maximum and width at half-maximum), stable peak heights (coefficient of variation <10% in the original redundant observations), and no interference from other mass peaks were selected from the strong characteristic peak regions (e.g., a characteristic peak of a certain substance at 1600 cm⁻¹ with no other peaks overlapping around it). If multiple candidate quantitative characteristic peaks exist, the quantitative applicability index of each peak (i.e., the linear correlation coefficient between peak height and known concentration) was calculated, and the peak with the highest index was selected as the final quantitative characteristic peak (i.e., the quantitative benchmark).
[0158] Call the peak intensity-concentration standard curve of the target substance (pre-plotted by the laboratory, containing peak intensity values corresponding to concentration gradients of 0.1-100 mg / L); measure the peak intensity of the quantitative characteristic peak in the denoised one-dimensional spectrum (take the peak height or integral area, which needs to be consistent with the standard curve), and substitute it into the standard curve equation (e.g., concentration = a × peak intensity + b) to calculate the concentration of the substance sample;
[0159] If the material sample has undergone enrichment treatment, the calculation results need to be corrected according to the enrichment factor (actual concentration = calculated concentration / enrichment factor).
[0160] Calculate the deviation between the concentration value and the estimated concentration range in the metadata (e.g., about 5 mg / L). If the deviation is less than a certain percentage requirement (e.g., ≤5%), the quantitative result is considered reliable. If the deviation is greater than the percentage requirement (e.g., >5%), check whether it is caused by peak intensity saturation (exceeding the linear range of the standard curve) or baseline drift. If necessary, recalculate the quantitative peak until the error meets the percentage requirement.
[0161] Methods for indicating the validity of analysis results include:
[0162] The qualitative and quantitative analysis results are evaluated for validity separately. If they pass the evaluation, they are marked as valid candidates; otherwise, they are marked as invalid candidates. Specifically:
[0163] The system includes a pre-defined validity judgment rule base, which includes qualitative validity thresholds (e.g., total matching degree ≥ 90%, and no missing key peaks (the core characteristic peaks of the substance in the standard spectral library are present in the spectrum to be analyzed)) and quantitative validity thresholds (e.g., concentration calculation deviation ≤ 5% (compared with the standard reference value or the estimated range), and the peak intensity of the quantitative characteristic peak is within the linear range of the standard curve (unsaturated)).
[0164] The method for determining the validity of qualitative analysis results is as follows:
[0165] If the qualitative matching degree is greater than or equal to the qualitative validity threshold and there is no missing key peak, the qualitative analysis result is initially determined to be a valid candidate; if the matching degree is less than the qualitative validity threshold, or there is a missing key peak, the qualitative analysis result is marked as an invalid candidate and the specific reason is recorded.
[0166] The method for determining the validity of quantitative analysis results is as follows:
[0167] If the concentration calculation error is less than or equal to the quantitative validity threshold, and the peak intensity of the quantitative characteristic peak is within the linear range of the standard curve, the quantitative result is initially determined to be a valid candidate; if the error is greater than the quantitative validity threshold, or the peak intensity of the quantitative characteristic peak exceeds the linear range, it is marked as an invalid candidate, and the reason is recorded (e.g., peak intensity saturation leads to excessive error).
[0168] If both the qualitative and quantitative analysis results are marked as valid candidates, then the analysis results are marked as valid by adding a comprehensive validity label. If any invalid candidate label exists, then the analysis results are marked as invalid by adding a comprehensive validity label.
[0169] The methods for obtaining the comprehensive evaluation results include:
[0170] Based on the one-dimensional spectra before and after denoising, obtain the signal-to-noise ratio (SNR) of the strong characteristic peak region of the one-dimensional spectra before and after denoising (denoised as SNR1 before denoising and SNR2 after denoising, i.e., the ratio of the peak height of the strong peak to the standard deviation of the background noise).
[0171] The effectiveness of the noise reduction process is evaluated by combining the signal-to-noise ratio improvement, weak peak retention, and peak shape fidelity preservation effects.
[0172] The evaluation method for the signal-to-noise ratio improvement effect is as follows:
[0173] First, calculate the signal-to-noise ratio (SNR) improvement rate. The formula is: SNR improvement rate = (SNR2 - SNR1) / SNR1. If the SNR improvement rate meets the expectation (e.g., ≥30%), the SNR improvement effect is recorded as effective denoising; otherwise, it is recorded as inefficient denoising.
[0174] The evaluation method for the weak peak retention effect is as follows:
[0175] The number of identifiable weak peaks that must be retained in the denoised one-dimensional spectrum is counted, and then the weak peak retention rate (i.e., the ratio of the number of identifiable weak peaks that must be retained after denoising to the number of identifiable weak peaks in the original spectrum) is calculated.
[0176] If the retention rate of weak peaks meets expectations (e.g., ≥90%) and no necessary weak peaks are lost, the weak peak preservation is deemed effective; if the retention rate does not meet expectations, it is marked as excessive suppression of weak peaks.
[0177] The evaluation method for peak shape fidelity is as follows:
[0178] The distortion rate is calculated by comparing the full width at half maximum (FWHM) of the strong characteristic peaks in the one-dimensional spectrum before and after denoising. The formula is: Distortion rate = |FWHM after denoising - FWHM before denoising| / FWHM before denoising.
[0179] If the distortion rate meets the expectations (e.g., ≤5%), it indicates that the peak shape remains basically unchanged and is judged as having no distortion of the characteristic peak; if it does not meet the expectations, it is marked as peak shape distortion impact analysis.
[0180] Based on the matching construction strategy, the adaptability and redundancy gain of the redundancy construction strategy are evaluated and combined as the redundancy strategy utility.
[0181] The evaluation method for the adaptability of the redundancy construction strategy is as follows: For the construction strategy used, a strategy-region matching score is calculated. An exemplary calculation logic is as follows: For the strong feature peak region strategy, a signal consistency of ≥90% for sub-observation pairs is awarded 3 points, 80%-90% is awarded 2 points, and <80% is awarded 1 point; For the weak feature peak region strategy, a weak peak detection rate (i.e., the ratio of the number of identifiable weak peaks to the total number of weak peaks that must be retained) ≥80% is awarded 3 points, 60%-80% is awarded 2 points, and <60% is awarded 1 point; A total score ≥5 indicates strategy adaptability, and <5 indicates strategy mismatch.
[0182] The evaluation method for redundancy gain is as follows: calculate the redundancy gain using the formula: Redundancy gain = (Number of valid analysis results - Number of analysis results without redundancy) / Number of analysis results without redundancy (data without redundancy refers to analysis results using only a single set of observations).
[0183] If the redundancy gain meets expectations (e.g., ≥20%), it indicates that redundancy significantly increases the number of valid results, and is marked as having significant redundancy construction gain.
[0184] The analytical results that are deemed valid based on overall validity are recorded as valid analytical results. The qualitative accuracy of the qualitative analysis is calculated based on the valid analytical results (i.e., the ratio of the number of samples with a matching degree that meets expectations (≥90%) to the total number of valid samples). The denoising utility of the qualitative analysis is then evaluated. Specifically, if the qualitative accuracy meets expectations (e.g., ≥95%) and is higher than the results of the analysis without redundancy (e.g., 85% accuracy without redundancy), the qualitative analysis is considered to have excellent utility; if it does not meet expectations, it is marked as having insufficient qualitative accuracy.
[0185] Based on the valid results, the quantitative average error of the quantitative analysis is obtained (i.e., the mean of the absolute values of the quantitative errors of all samples, where the quantitative error is the deviation between the concentration of the substance sample and the predicted concentration range). Then, the denoising effect of the quantitative analysis is evaluated. That is, if the quantitative average error meets the expectations (e.g., ≤3%) and is lower than the quantitative error of the one-dimensional spectrum before denoising, it is judged as a significant improvement in quantitative accuracy. If the quantitative average error does not meet the expectations, it is marked as a failure to meet the quantitative accuracy standard.
[0186] The denoising utility of qualitative and quantitative analysis is integrated together as the analytical utility;
[0187] The denoising utility, redundancy strategy utility, and analytical utility are integrated. Specifically, each indicator is assigned a weight. An example of the indicator weight assignment logic is as follows: Denoising utility (30%), including signal-to-noise ratio improvement rate (15%), weak peak retention rate (10%), and characteristic peak distortion rate (5%); Redundancy strategy utility (20%), including strategy-region matching score (12%) and redundancy gain (8%); Analytical utility (50%), including qualitative accuracy rate (25%) and quantitative average error (25%).
[0188] The indicators are weighted and integrated according to the assigned weights, and the results are mapped to an evaluation score (0-100 points). The final comprehensive evaluation result is determined based on the evaluation score. An example of the determination method is: an evaluation score ≥80 points indicates excellent utility, 60-79 points indicates adequate utility, and <60 points indicates insufficient utility.
[0189] The optimized methods for generating the redundancy construction strategy library include:
[0190] If the overall evaluation results do not meet expectations, that is, the overall evaluation results are judged to be insufficient in utility, the problem category is located through the problem classification and positioning mechanism. Specifically, the problem classification and positioning mechanism includes problem classification and positioning in the noise reduction stage, the redundancy strategy stage, and the analysis stage.
[0191] The problem classification and localization logic in the denoising process is as follows: if the weak peak retention rate is low and the signal-to-noise ratio improvement rate meets the standard, it is determined that the denoising parameters excessively suppress the weak signal; if the peak distortion rate is greater than the normal distortion threshold (such as 10%), it is located as the smoothing window being too large, resulting in blurred peak shape.
[0192] The problem classification and location logic of the redundancy strategy is as follows: if the strategy adaptability score is low and the signal consistency in the strong peak area does not meet expectations (e.g., <80%), it is determined that the strong peak area division strategy does not match the peak shape complexity (e.g., complex peaks are divided using simple odd and even rows); if the redundancy gain is small (e.g., <10%), it is located as insufficient number of sub-observation pairs.
[0193] The problem classification and localization logic in the analysis stage is as follows: if the qualitative accuracy is low (e.g., <85%) and most of the samples are low-concentration samples, it is determined that the weak peak features have not been effectively matched; if the quantitative average error does not meet expectations, it is determined that the quantitative peak selection is affected by background noise.
[0194] One way to perform targeted strategy optimization is:
[0195] For adjusting denoising parameters: if weak peaks are not retained sufficiently, reduce the smoothing window in the weak peak region (e.g., reduce from 5 channels to 3 channels) and increase the threshold of the weak peak signal (e.g., adjust from baseline +1.5σ to baseline +1.2σ) to ensure that weak peaks are not over-smoothed; if peak shape distortion occurs, reduce the weighted average weight deviation in the strong peak region (e.g., narrow the peak height weight coefficient range from 0.8-1.2 to 0.9-1.1) to avoid extreme values distorting the peak shape;
[0196] For redundancy strategy iteration: if the strong peak region has poor adaptability, switch complex peaks (i.e., full width at half maximum (FWHM) > 8 channels) to upper and lower half region division (replacing odd and even row division) to ensure the peak structure is complete; simple peaks (FWHM ≤ 5 channels) retain interval row sampling to improve efficiency; if the redundancy gain is insufficient, increase the number of low-concentration sample sub-observation pairs from 2 to 3, and simultaneously enable cross-sample redundancy supplementation (calling weak peak region data of high-concentration samples in the same batch).
[0197] For optimizing the analysis process: if the qualitative accuracy is low, force multi-peak joint verification (≥2 weak peaks must be matched) for low-concentration samples; if the quantitative error is large, add a quantitative peak anti-interference screening step (prioritize strong peaks with low background noise).
[0198] The optimized parameters and construction strategies are written into the redundant construction strategy library to obtain the optimized redundant construction strategy library. A sample feature-optimization strategy mapping table (e.g., low concentration + complex peak → 3 sets of sub-observation pairs + weak peak weight enhancement) is established as a dynamic adaptation rule.
[0199] Simultaneously, the entire process optimization trajectory is recorded to generate a closed-loop iteration record, enabling continuous adaptive closed-loop iterative improvement. Example 2
[0200] Please see Figure 3 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A self-supervised one-dimensional spectral denoising system based on redundant observation construction is provided, including:
[0201] Data acquisition module: acquires the original spectral signal of the target substance, collects two-dimensional spectral data containing characteristic peaks, and synchronously records metadata of the substance sample and the environment to form a two-dimensional spectral dataset containing redundant observations;
[0202] Redundancy construction module: Based on the two-dimensional spectral dataset, the characteristic peak regions are divided according to the signal repeatability. The construction strategy of the corresponding characteristic peak region is matched with the redundancy construction strategy library to generate qualified redundant sub-observation pairs.
[0203] Spectral denoising module: Based on the characteristic peak region division results and redundant sub-observation pairs, it merges them to generate one-dimensional spectral pairs, and obtains the denoised one-dimensional spectrum through calibration processing; it is then verified in combination with feature fidelity rules to obtain a feature fidelity verification report;
[0204] Material Analysis Module: Based on the denoised one-dimensional spectrum and feature fidelity verification report, qualitative analysis is completed by comparison with the standard spectral library, and quantitative analysis is completed by obtaining the substance concentration according to the peak intensity, forming qualitative and quantitative analysis results, and marking the validity of the analysis results;
[0205] Utility evaluation and iteration module: Based on the analysis results and validity indicators, a comprehensive evaluation result is obtained. When the expected results are not met, targeted strategy optimization is performed, the redundancy construction strategy library is updated and the adjustment process is recorded, and an optimized redundancy construction strategy library and closed-loop iteration record are generated. Example 3
[0206] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the self-supervised one-dimensional spectral denoising system based on redundant observation construction described above.
[0207] Since the electronic device described in this embodiment is the one used to implement the self-supervised one-dimensional spectral denoising method based on redundant observation construction described in this application, those skilled in the art can understand the specific implementation and various variations of the electronic device in this embodiment based on the self-supervised one-dimensional spectral denoising method based on redundant observation construction described in this application. Therefore, how the electronic device implements the method in this application will not be described in detail here. Any electronic device used by those skilled in the art to implement the self-supervised one-dimensional spectral denoising method based on redundant observation construction described in this application falls within the scope of protection of this application.
[0208] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0209] The above description is merely a preferred embodiment of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for users of ordinary technical skills, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A self-supervised one-dimensional spectral denoising method based on redundant observation construction, characterized in that, include: S1: Acquire the original spectral signal of the target substance, collect two-dimensional spectral data containing characteristic peaks, and simultaneously record metadata of the substance sample and the environment to form a two-dimensional spectral dataset containing redundant observations; S2: Based on a two-dimensional spectral dataset, characteristic peak regions are divided according to signal repeatability. The corresponding characteristic peak region construction strategy is matched with the redundancy construction strategy library to generate qualified redundant sub-observation pairs. S3: Based on the characteristic peak region division results and redundant sub-observation pairs, one-dimensional spectral pairs are generated by merging them, and the denoised one-dimensional spectrum is obtained through calibration. The feature fidelity is verified using the feature fidelity rules, resulting in a feature fidelity verification report. S4: Based on the denoised one-dimensional spectrum and feature fidelity verification report, qualitative analysis is completed by comparison with the standard spectral library, and quantitative analysis is completed by obtaining the substance concentration according to the peak intensity, forming qualitative and quantitative analysis results, and marking the validity of the analysis results; S5: Based on the analysis results and validity indicators, a comprehensive evaluation result is obtained. When the expected results are not met, targeted strategy optimization is performed, the redundancy construction strategy library is updated and the adjustment process is recorded, and an optimized redundancy construction strategy library and closed-loop iteration record are generated.
2. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 1, characterized in that, The two-dimensional spectral dataset is formed in the following ways: By using a suitable spectroscopic device, a horizontal channel covering the wavelength range of the characteristic peak is set for the target substance sample, and a vertical redundant observation parameter is set to make the same characteristic peak appear repeatedly in multiple vertical rows. According to the set horizontal and vertical dimensions, collect two-dimensional spectral data of the target substance containing characteristic peaks; simultaneously monitor data quality in real time, and record substance sample information, environmental parameters and equipment parameters as metadata; The two-dimensional spectral data is correlated and verified with the metadata to form a two-dimensional spectral dataset containing redundant observations.
3. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 2, characterized in that, The method of dividing characteristic peak regions according to signal repeatability includes: By calling a two-dimensional spectral dataset, signal quality is enhanced through baseline correction and smoothing / denoising processes. Furthermore, based on the repeatability of the signal in the vertical rows, for each horizontal channel, the signal mean and baseline in all vertical rows are obtained, and channels with signal mean significantly stronger than the baseline are selected as potential strong feature peak regions. In the potential characteristic peak region, strong characteristic peak regions are first defined by high frequency of occurrence and stable peak position; then, signals with medium frequency of occurrence and weaker intensity than strong characteristic peaks are selected around the strong characteristic peak regions to define characteristic peak regions; the remaining regions without obvious peak shapes are defined as background regions.
4. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 3, characterized in that, The methods for generating the redundant sub-observation pairs include: Based on the feature region segmentation results, a pre-set redundant construction strategy library is invoked, and the optimal construction strategy is matched between strong feature peak regions and weak feature peak regions through the corresponding matching rules. Sub-observation pairs are generated according to a preset number of sub-observation pairs. The original two-dimensional spectral data is cut according to the matching construction strategy, and then sub-observation fragments of each region are extracted and combined into complete redundant sub-observation pairs. The generated redundant sub-observation pairs are checked for signal consistency and noise independence. If the check fails, a retry mechanism is triggered to regenerate and check again until a qualified redundant sub-observation pair is obtained.
5. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 4, characterized in that, The methods for obtaining the denoised one-dimensional spectrum include: Based on the feature region division results and qualified redundant sub-observation pairs, each group of redundant sub-observation pairs is first merged vertically along the horizontal channel to generate a one-dimensional spectral pair. Calibration processing is performed on different regions of one-dimensional spectral alignment. Specifically, peak position alignment and peak height calibration are performed on strong characteristic peak regions, and the blurred signal is enhanced by cross-spectral comparison on weak characteristic peak regions. Background regions are smoothed and denoised. Then, following the principle of selecting the best, the different regions after calibration are merged to obtain the denoised one-dimensional spectrum.
6. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 5, characterized in that, The generation methods for the feature fidelity verification report include: The verification logic of the feature fidelity rule is defined as follows: for strong feature peak areas, verify the peak position shift, peak shape distortion and peak height rationality; for weak feature peak areas, verify whether weak peaks are lost; and for background areas, verify whether noise is effectively suppressed and clearly isolated from feature peaks. Based on the denoised one-dimensional spectrum, each region is verified according to the verification logic of the feature fidelity rule. If all verifications pass, the feature fidelity verification is deemed to have passed. If any verification fails, the feature fidelity verification is deemed to have failed. For the spectrum that failed verification, the calibration process for the one-dimensional spectrum pair is triggered in reverse and reprocessed until the feature fidelity verification passes. The verification results from each region and the final feature fidelity determination results are integrated to generate feature fidelity verification results.
7. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 6, characterized in that, The qualitative and quantitative analysis results are formed in the following ways: The qualitative analysis logic is defined as follows: Based on the denoised one-dimensional spectrum that has passed the feature fidelity verification, combined with the standard spectral library of the substance, all peaks in the strong feature peak region and the weak peaks that must be retained in the weak feature peak region are extracted to form a list of peaks to be compared. Based on the list of peaks to be compared, the denoised one-dimensional spectrum is compared with the standard spectral library in multiple dimensions to obtain the characteristic peak matching degree. The corresponding substance name is matched by the characteristic peak matching degree as the qualitative analysis result. The quantitative analysis logic is defined as follows: based on the denoised one-dimensional spectrum, select highly applicable quantitative characteristic peaks in the strong characteristic peak region, extract the peak intensity, and then combine the peak intensity-concentration standard curve of the target substance to obtain the concentration of the target substance, which is used as the quantitative analysis result.
8. The self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 7, characterized in that, The methods for identifying the validity of the labeled analysis results include: The validity of the qualitative and quantitative analysis results is determined separately. If the results pass the test, they are marked as valid candidates; otherwise, they are marked as invalid candidates. The method for determining the validity of qualitative analysis results is as follows: If the qualitative matching degree is greater than or equal to the qualitative validity threshold and there are no missing key peaks, the qualitative analysis result is initially determined to be a valid candidate; if the matching degree is less than the qualitative validity threshold, or there are missing key peaks, the qualitative analysis result is marked as an invalid candidate. The method for determining the validity of quantitative analysis results is as follows: If the concentration calculation error is less than or equal to the quantitative validity threshold, and the peak intensity of the quantitative characteristic peak is within the linear range of the standard curve, the quantitative result is initially determined to be a valid candidate; if the error is greater than the quantitative validity threshold, or the peak intensity of the quantitative characteristic peak exceeds the linear range, it is marked as an invalid candidate. If both the qualitative and quantitative analysis results are marked as valid candidates, then the analysis results are marked as valid by adding a comprehensive validity label. If any invalid candidate label exists, then the analysis results are marked as invalid by adding a comprehensive validity label.
9. A self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 8, characterized in that, The comprehensive evaluation results are obtained in the following ways: The denoising effectiveness is evaluated based on the one-dimensional spectra before and after denoising, through the signal-to-noise ratio improvement, weak peak retention, and peak shape fidelity preservation effects. Based on the matching construction strategy, the utility of the redundancy strategy is evaluated through the adaptability of the redundancy construction strategy and the redundancy gain. The analytical results that are identified as valid based on the overall validity are recorded as valid analytical results. The qualitative accuracy rate of the qualitative analysis and the quantitative average error of the quantitative analysis are statistically analyzed based on the valid analytical results to evaluate the analytical utility. The denoising utility, redundancy strategy utility, and analysis utility are combined to generate a comprehensive evaluation result.
10. A self-supervised one-dimensional spectral denoising method based on redundant observation construction according to claim 9, characterized in that, The optimized redundant construction strategy library is generated in the following ways: If the overall evaluation results do not meet expectations, the problem category will be identified through the problem classification and positioning mechanism, and targeted strategy optimization will be carried out. The optimized parameters and construction strategies are written into the redundant construction strategy library and dynamic adaptation rules are generated. The optimized redundant construction strategy library is generated, and closed-loop iteration records are generated simultaneously.
Citation Information
Patent Citations
Raman spectrum denoising method based on adaptive sparse decomposition
CN116380869A
Raman spectrum nonlinear generation method driven by mixed machine learning
CN120470287A