In-situ multi-modal spectral identification and drift adaptive method for mine water contaminants

By employing a multimodal spectral recognition method, absorption, fluorescence, and scattering signals are simultaneously acquired. Sparse coding decomposition is performed using an end-member dictionary, which solves the problem of real-time monitoring and identification accuracy of mine water pollutants, and achieves high-precision, low-maintenance on-site monitoring of mine water pollutants.

CN120992533BActive Publication Date: 2026-02-10CHINA INST OF WATER RESOURCES & HYDROPOWER RES +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511508507.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-10
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing methods for detecting pollutants in mine water cannot achieve in-situ real-time monitoring. Single-channel spectral detection cannot fully reflect the coupling effect of multiple pollutants in the water. Multi-channel fusion methods fail to effectively consider the inconsistencies between channels, resulting in low identification accuracy. Furthermore, equipment drift leads to identification errors.

Method used

A multimodal spectral identification method is adopted, which simultaneously acquires absorption, fluorescence and scattering channel signals, performs sparse coding endmember decomposition using an endmember dictionary, and combines point-by-point subtraction and residual auditing to achieve unified analysis of multi-channel spectra. Wavelength and intensity drift are corrected by incremental domain adaptation and reference calibration loop.

Benefits of technology

It achieves high-resolution identification of complex mixed pollutants in mine water, avoids identification errors caused by equipment aging and environmental changes, has self-calibration characteristics, can operate stably for a long time, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120992533B_ABST
    Figure CN120992533B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of optical detection, and particularly relates to a multi-modal spectrum in-situ recognition and drift self-adaptive method for mine water pollutants, which comprises the following steps: step 1: synchronously acquiring signals of three channels including absorption, fluorescence and scattering, sequentially performing dark signal removal, baseline flattening, time alignment and amplitude normalization pretreatment, and forming standardized multi-modal spectrum joint observation data; step 2: calling a preset endmember dictionary library, and performing sparse coding endmember decomposition on the standardized multi-modal spectrum joint observation data; and step 3: generating a pollutant category and a relative grade based on an endmember abundance table, and performing an incremental domain self-adaptive and calibration loop. The present application significantly improves the resolution capability of complex mixed pollutants in mine water; avoids recognition errors caused by equipment aging, light source attenuation and environmental changes; has self-calibration and self-adaptive characteristics without manual intervention, can be operated stably for a long time, and reduces maintenance cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of optical detection, and particularly relates to a multi-modal spectrum in-situ recognition and drift self-adaptive method for mine water pollutants. BACKGROUND

[0002] Mine water is a by-product fluid commonly existing in mine production activities, and its composition is complex, usually containing suspended particles, metal ions, organic solutes and various inorganic salts. Real-time monitoring and pollutant identification of mine water are important foundations for protecting ecological safety and groundwater resource utilization in mining areas. At present, the detection methods of mine water pollutants mainly include three types: laboratory offline chemical analysis, single-channel spectral detection and multi-spectral remote sensing inference. Traditional laboratory chemical analysis methods, such as atomic absorption spectrometry, ion chromatography and ICP-OES, have high accuracy, but the sampling, transportation and sample preparation process are complex, which cannot meet the needs of in-situ real-time monitoring. For dynamic water flow in deep wells or closed stope, this method has obvious limitations in sampling representativeness and time response.

[0003] Single-channel spectral detection method is earlier used in online monitoring of mine water, in which absorption spectrum is used to determine the concentration of metal ions and inorganic salts, fluorescence spectrum is used to characterize organic pollutants, and scattering spectrum is used to estimate turbidity and particle content. However, a single channel cannot fully reflect the coupling effect of multiple pollutants in water. For example, in mine water containing iron and manganese, the absorption signal is easily covered by the colloidal scattering background; the fluorescence channel is greatly disturbed by turbidity, and the signal often drifts and quenches; although the scattering signal can reflect the particle concentration, it lacks material characteristics. In view of these problems, some studies propose a multi-channel spectral fusion idea, which combines absorption, fluorescence and scattering information to improve the identification accuracy. However, the existing multi-channel fusion methods are mostly based on simple linear combination or principal component transformation, without considering the non-uniformity of different channels in response dynamics and noise structure, resulting in limited fusion effect. In terms of spectral data analysis algorithm, existing studies often use full-spectrum regression, non-negative matrix decomposition, independent component analysis or sparse representation. Full-spectrum regression relies on a large number of calibration samples, and has poor generalization ability for new scenes or drifting samples; non-negative matrix decomposition and independent component analysis often assume a fixed number of end members, while the pollution components of mine water change with time and water flow conditions; traditional sparse representation methods can recover end members to some extent, but cannot effectively distinguish the mixing characteristics between similar spectral shapes, and the decomposition results are easy to lose stability when there is drift. In addition, most algorithms only analyze single channel, without considering the consistency constraint between cross-channel end members, so it is difficult to achieve accurate synchronous identification of pollutants. SUMMARY

[0004] In view of this, the main purpose of the present application is to provide a mine water pollutant multi-modal spectral in-situ identification and drift adaptive method, which comprehensively utilizes the spectral information of absorption, fluorescence and scattering three channels, realizes the unified analysis of multi-channel spectrum and pollutant class identification through synchronous acquisition, standardization processing and sparse coding end member decomposition. The method uses the channel corresponding structure of the end member dictionary library to maintain the spectral consistency in the end member selection, amplitude alignment and peak position fine tuning process, and realizes high resolution decomposition through point-by-point subtraction and residual audit. In the long-term operation, the system automatically corrects the wavelength and intensity drift through the incremental domain adaptive and reference calibration loop, ensuring the stability and reliability of the results. The beneficial effects of the present application are: realizing the multi-channel collaborative pollutant identification, significantly improving the resolution capability of complex mixed pollutants in mine water; avoiding the identification error caused by equipment aging, light source attenuation and environmental change; having the self-calibration and adaptive characteristics without manual intervention, which can be operated stably for a long time, reduces the maintenance cost, and provides a high-precision and sustainable technical solution for the on-site real-time monitoring of mine water pollutants.

[0005] The technical solutions adopted by the present application are as follows:

[0006] The mine water pollutant multi-modal spectral in-situ identification and drift adaptive method comprises:

[0007] Step 1: synchronously acquiring signals including absorption, fluorescence and scattering three channels, sequentially executing dark signal removal, baseline flattening, time alignment and amplitude normalization preprocessing to form standardized multi-modal spectral joint observation data;

[0008] Step 2: calling a preset end member dictionary library, performing sparse coding end member decomposition on the standardized multi-modal spectral joint observation data: according to the main peak and auxiliary peak marks in each end member item in the end member dictionary library, synchronously dividing the absorption, fluorescence and scattering three channels into several aligned segments; for each aligned segment, multiplying and summing the aligned segment and its corresponding end member item point by point to obtain an overlap amount, taking the square root of the point-by-point square sum of the aligned segment and the end member item as a scale, dividing the overlap amount by the product of the two scales to obtain a segment similarity score, and taking the minimum value of the segment similarity scores of the three channels as the channel consistency score of the aligned segment; selecting the end member with the highest score, performing amplitude alignment at the peak height ratio, and performing peak position fine tuning on the segments marked as deformable by piecewise linear interpolation; subtracting the end member item after amplitude alignment and peak position fine tuning from the standardized multi-modal spectral joint observation data point by point, setting the negative value to zero and merging small peaks, and repeating this process until the residual meets the stopping criterion to record the alignment factor at each subtraction and generate an end member abundance table;

[0009] Step 3: generating the pollutant class and relative grade based on the endmember abundance table, and performing incremental domain adaptation and calibration loop: taking reference measurements at the beginning and end of sample processing, comparing with the reference standard stored in the system to obtain sampling point offset and amplitude offset; using the obtained sampling point offset and amplitude offset to implement wavelength and intensity alignment on the three channels, and synchronously correcting the entries in the endmember dictionary; when the difference between the reference measurement and the reference standard exceeds the preset threshold, triggering a complete sparse coding endmember decomposition recalculation, and expanding the segment fine-tuning range according to the piecewise linear interpolation rule of deformable segments; after recalculation, consistency check is performed on the generated class and grade, if inconsistent, use the recalculation result, if the difference exceeds the discrimination threshold, take the adjacent grade, the endmember that does not pass the consistency check is marked as to be confirmed and the final result is output.

[0010] Further, step 1 specifically includes: using a coaxial probe to simultaneously obtain signals of the absorption channel, the fluorescence channel and the scattering channel at the same measurement point to form multi-modal spectral joint observation data; sequentially performing dark signal removal, baseline flattening, channel time alignment and amplitude normalization on the multi-modal spectral joint observation data; performing Raman and Rayleigh band masking on the fluorescence channel and interpolating with adjacent effective bands, and performing angular consistency setting and turbidity reference binding on the scattering channel to obtain standardized multi-modal spectral joint observation data.

[0011] Further, in step 2, the endmember dictionary library includes an absorption endmember sub-library, a fluorescence endmember sub-library and a scattering endmember sub-library; each endmember in the dictionary library has a corresponding entry in the three sub-libraries, and the three entries are stored according to the same sampling point sequence; each entry includes four types of markers: main peak sampling point, auxiliary peak group, background segment and deformable segment; the standardized multi-modal spectral joint observation data is unified to the sampling point sequence used by the endmember dictionary library on the three channels; using the main peak and auxiliary peak markers in each endmember entry of the endmember dictionary library as anchor points, the three channels are simultaneously divided into a plurality of aligned segments, and each aligned segment corresponds to a main peak and its adjacent background segment.

[0012] Further, in step 2, the calculation of the consistency score is as follows: first, multiply each aligned segment with its corresponding endmember entry point by point and sum up to get the overlap amount, then use the point-by-point square sum of the aligned segment and the endmember entry respectively, and take the square root of the square sum as the normalization scale, divide the overlap amount by the product of the two normalization scales to get the segment similarity score of the aligned segment in each channel; take the minimum value between the segment similarity scores of the three channels as the segment channel consistency score of the aligned segment; take the median of the segment channel consistency scores of all segments corresponding to any endmember as the endmember similarity score of the endmember; select the endmember with the highest score among all endmembers as the current target endmember; locate the main peak sampling point of the current target endmember in the three channels respectively, and calculate the ratio of the peak height of the standard multi-modal spectral joint observation data at the target main peak sampling point to the peak height of the current target endmember entry to obtain the amplitude alignment factor; use the amplitude alignment factor to adjust the amplitude of the current target endmember entry in the three channels simultaneously.

[0013] Further, in step 2, for the section of the current target endmember entry marked as a deformable segment, a three-segment polyline connection section mapping is established around the sampling points of the main peak and the auxiliary peak, and the control points of the section mapping are taken from the peak and valley positions of the standard multi-modal spectral joint observation data in the same section; based on the established section mapping, perform point-by-point linear interpolation deformation on the corresponding section of the current target endmember entry, so that the peak position of the endmember entry is consistent with the peak position of the standard multi-modal spectral joint observation data.

[0014] Further, in step 2, the loop process is as follows: subtract the current target endmember entry after amplitude alignment and deformation from the standard multi-modal spectral joint observation data in the three channels point by point, and record the subtraction result as the residual sequence; set the negative values in the residual sequence to zero point by point, and merge and smooth adjacent small peaks to obtain updated standard multi-modal spectral joint observation data; if the total energy of the residual sequence in the three channels is lower than the preset threshold, or the number of stripped endmembers reaches the preset upper limit, stop the loop; otherwise, return to the calculation rule of the endmember similarity score and continue to perform endmember selection, fine tuning and stripping on the updated standard multi-modal spectral joint observation data; record the amplitude alignment factor at each stripping time according to the endmember name to form an endmember abundance table, and record the corresponding segment similarity score and residual sequence energy as the matching credentials of the corresponding endmember; eliminate or set to zero the negative items and abnormal items in the endmember abundance table; according to the endmember abundance table and the calibration table stored in the system in the endmember dictionary library, find the pollutant category and relative grade in the calibration interval of each endmember; when the same pollutant is hit in multiple endmember entries, select the first listed entry as the final source of the pollutant in the priority order recorded in the calibration table, and mark the entries that are not hit as not detected.

[0015] Further, in step 3, at the beginning and end of each sample processing, trigger the built-in reference component to perform a reference measurement, respectively, to obtain the reference absorption signal, the reference fluorescence signal, and the reference scattering signal; compare the main peak position and amplitude of the reference measurement with the reference benchmark stored with the system point by point to obtain the sampling point offset and the amplitude offset of the three channels; use the sampling point offset to perform sampling point translation and endpoint extrapolation interpolation on the three channels to complete the alignment in the wavelength direction; use the amplitude offset to perform amplitude stretching or compression on the three channels to complete the alignment in the intensity direction; correct only for the current sample processing process.

[0016] Further, in step 3, apply the sampling point offset and the amplitude offset obtained by channel correction to the corresponding channels of all endmember entries in the endmember dictionary library simultaneously; for the section marked as a deformable segment, use the peak position and valley position of the reference measurement as new control points, and deform the endmember entry in a piecewise linear interpolation manner to make the endmember entry consistent with the three corrected channels; when the difference between the reference measurement and the reference benchmark exceeds the preset threshold, trigger an endmember entry synchronous correction and a complete sparse coded endmember decomposition recalculation immediately; follow the endmember similarity score calculation rules, endmember selection and stripping process of step 2 during the recalculation process; after the recalculation is completed, compare the endmember abundance table and the residual sequence energy before and after the recalculation, if the difference still exceeds the preset threshold, then expand the control point range of the deformation section according to the deformable segment fine-tuning rules and recalculate again until the difference falls within the preset threshold or reaches the preset number of upper limit.

[0017] Further, in step 3, check the consistency of the pollutant category and the relative grade after recalculation with the results before recalculation: when the category is inconsistent, use the result after recalculation as the standard; when the relative grade difference exceeds the preset grade step, take the adjacent grade between the results before and after recalculation as the final grade; endmembers that do not pass the consistency check are excluded from the endmember abundance table and are marked as to be confirmed in the output; at the end of the current sample processing, write the sampling point offset, the amplitude offset, and the deformation control points of the deformable segment obtained by the current reference measurement into the device operating parameters for direct use in the channel correction and endmember entry synchronous correction of the next sample, without retaining any original observation data.

[0018] With the above technical solutions, the present application has the following beneficial effects: the present application realizes synchronous acquisition, standardization processing and sparse coding end member decomposition in three channels of absorption, fluorescence and scattering, and establishes a unified identification framework of multi-modal spectrum. Through the multi-channel corresponding structure of the end member dictionary library, the present application makes each end member maintain the consistency of peak position, shape and response characteristics in different spectral channels, so that it can identify multiple pollutant components at the same time under the complex mine water background. Compared with the existing single channel or linear fusion method, the present application introduces a point-by-point deduction and segment consistency comparison mechanism in the end member selection and spectral shape stripping process, effectively avoiding the misidentification of interference and overlapping spectral peaks between channels. Through segmented linear interpolation fine tuning, the peak position can be adaptively corrected without changing the overall characteristics of the spectral shape, so that the identification process has higher stability to environmental disturbance and sample fluctuation. In terms of device and environmental drift, the present application proposes an incremental domain adaptive and reference calibration loop mechanism, which performs reference measurement at the beginning and end of sample processing and dynamically corrects the wavelength and intensity coordinates, so that the system can automatically maintain accuracy under long-term unattended conditions. Through the cooperation of trigger condition recalculation and end member dictionary library synchronous correction, the drift error caused by factors such as light source attenuation, probe fouling and temperature change can be eliminated in real time, so as to maintain the repeatability and traceability of the identification results. Overall, the present application realizes multi-channel collaborative detection of mine water pollutants, self-consistent decomposition of spectral shape level and continuous drift self-adaptation under field conditions, and provides a high-robustness, high-resolution and low-maintenance-cost solution for mine water online monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A flowchart of a multi-modal spectrum in-situ identification and drift self-adaptation method of mine water pollutants is provided for the embodiments of the present application.

[0020] Figure 2 A schematic diagram of three-channel spectrum signal synchronous acquisition is provided for the embodiments of the present application.

[0021] Figure 3 A schematic diagram of aligned segment similarity score calculation is provided for the embodiments of the present application.

[0022] Figure 4 A principle schematic diagram of end member stripping and residual calculation process is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0023] All features disclosed in this specification, or all steps of any methods or processes disclosed, may be combined in any combination, except combinations where at least some features and / or steps are mutually exclusive.

[0024] Any feature in the disclosure (including any accompanying claims, abstract) can be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise. That is, unless expressly stated, each feature is one of a range of equivalent or similar features.

[0025] Reference Figure 1 : In-situ multi-modal spectral identification and drift adaptive method of mine water pollutants, comprising:

[0026] Step 1: Synchronously acquire signals of three channels including absorption, fluorescence and scattering, and sequentially perform preprocessing of dark signal removal, baseline flattening, time alignment and amplitude normalization to form standardized multi-modal spectral joint observation data.

[0027] Specifically, at the same measurement point, the absorption channel, the fluorescence channel and the scattering channel are placed on the same coaxial observation path, and a single system clock is used to issue a sampling start instruction, so that the first sampling point time of the three channels is consistent. To avoid the slight phase difference caused by mechanical vibration and water flow pulsation, the sampling duration is set to an integer multiple of the single flow-through time, for example, the sampling duration is 3 seconds, and the light source is kept stable during this duration. The absorption channel and the fluorescence channel are sequentially scanned with uniform wavelength steps, for example, from 230 nanometers to 900 nanometers with a step of 1 nanometer, obtaining 671 sampling points for each channel; the scattering channel records the intensity sequence within the same duration, and resamples to the same number of sampling points as the absorption channel and the fluorescence channel at the end of sampling by linear interpolation. In this way, the three channels observe the same water body in space and time, and have consistent length in data structure, which facilitates subsequent point-by-point processing.

[0028] Before and after sampling, each performs a light blocking measurement, and the light blocking measurement keeps the light path and probe position unchanged, and only closes the light source. Two sets of dark signals are obtained, and the point-by-point average value is used as the dark signal of the sample. If the maximum difference of the two sets of dark signals exceeds 5 count values, it is judged that the environmental temperature or electronic noise has changed rapidly, and the light blocking measurement is immediately repeated until the maximum difference does not exceed 5 count values. The dark signal is subtracted from the observation sequence of the three channels point by point to obtain the multi-modal spectral joint observation data after dark signal removal. By performing light blocking measurement before and after sampling, the fixed offset caused by temperature drift and readout bias can be offset, so that the weak signal region is not lifted by the noise floor, and misleading to subsequent baseline flattening is avoided. In the common scene of 25 degrees Celsius ground temperature, the average value of the dark signal is usually between 12 and 18 count values, and the random fluctuation range near zero intensity is maintained at 3 to 5 count values after removal.

[0029] After dark signal removal, estimate and subtract the slowly varying background from each channel. For absorption and fluorescence channels, use moving window lower envelope estimation. Specifically, use a moving window of length 101 samples on each channel, calculate the 20th percentile of each window as the local background of the window, and then linearly stitch all local backgrounds to form the full spectrum background curve. Subtract the background curve from the observed sequence point by point, and set the negative samples to zero. For the scattering channel, first find a flat section as reference, for example, take 200 consecutive samples around the median of the channel sequence, and subtract the average of the section from the full sequence point by point. The reason for using percentile background and flat section background is that suspended particles and colored dissolved substances in mine water will cause the background to rise slowly, while the true absorption and fluorescence peaks are narrow or locally protruding. By fitting the slowly varying lower envelope first, the peak shape can be preserved and the background can be eliminated, so that the subsequent time alignment and amplitude normalization are not affected by the background fluctuations. After baseline flattening, the common background fluctuation amplitude can be compressed from the original tens of percent to the order of one percent.

[0030] Although the three channels start sampling at the same time, due to the difference in internal integration time and readout method, there may be a slight shift in the time axis of the peak value. For this reason, a rectangular step of light intensity is arranged within the sampling duration, the step occurs at a position about 1 second after the start of sampling, and the step width is set to 100 milliseconds. The three channels will all respond to this step. Extract the sampling point index of the rising edge of the step for each channel, and align the indexes of the three channels to the index of the absorption channel. If there is a non-integer sampling point difference, linearly interpolate and shift the fluorescence and scattering channels so that the rising edge occurs at the same sampling point. This way, without introducing complex calculations, we can use the same physical event as a common marker to eliminate the time shift caused by different channel sampling links, ensuring that the peak and valley of the three channels are matched one by one, facilitating consistent segment division and point-by-point operations on the three channels. In practical applications, the maximum offset before alignment is commonly 20 to 120 milliseconds, and the residual offset after alignment is less than 2 milliseconds.

[0031] To eliminate the amplitude inconsistency caused by different exposure time, incident light fluctuation and fiber coupling efficiency difference, two-stage amplitude normalization is adopted. The first stage is zero-point normalization: in the absorption channel and the fluorescence channel, select the low signal area far from the known main peak position, for example, in the 400-420 nm and 820-840 nm segments, calculate the average value of each segment, and subtract the average value from the full sequence point by point to make the low signal area close to zero. The scattering channel uses the average value of the aforementioned flat segment to perform the same processing. The second stage is unit amplitude normalization: find the main peak of this sample on each channel, which is defined as a peak that is more than three times higher than the average value of the neighborhood and has a duration of more than 5 sampling points. The peak height of this main peak is taken as the unit amplitude of this channel, and the full sequence is divided by the unit amplitude point by point to make the peak height of the main peak approximately 1. The basis for selecting the main peak as the unit amplitude is that the main peak has the highest signal-to-noise ratio and can stably reflect the optical path and light source state of this sample, thereby pulling the intensity scale between channels to a comparable interval. If this sample does not contain a significant main peak, use the standard deviation of the aforementioned flat segment as the unit amplitude to avoid amplifying pure noise samples. After completing the two-stage amplitude normalization, the numerical range of the three channels is unified between 0 and 1, which facilitates subsequent point-by-point subtraction and similarity calculation on the three channels.

[0032] After completing dark signal removal, baseline flattening, time alignment and amplitude normalization, the data of the three channels are combined according to the unified sampling point sequence to form the standardized multi-modal spectral joint observation data. Each channel contains 671 sampling points, the data saves the original count value as a sixteen-bit unsigned integer, while the normalized value is saved as a floating-point type, with a value range of 0 to 1, and is accompanied by the index value of the sampling start time and the time alignment marker. The standardized multi-modal spectral joint observation data is directly input to the subsequent steps and does not depend on historical samples.

[0033] In an optional embodiment, in the fluorescence channel, the Raman bands and Rayleigh bands can be masked and interpolated before baseline flattening. Specifically, a range of 5 to 15 sampling points is taken on both sides of the center of each band, and the sampling points in the range are replaced by linear interpolation of the adjacent valid bands. In this way, the bias caused by strong bands to the background estimation can be avoided. In the scattering channel, angular consistency adjustment and turbidity reference binding can be performed before amplitude normalization. Specifically, when the scattering channel is derived from multiple observation angles, the intensities of each angle are first aligned according to the maximum intensity, and then the average value of the multi-angle sequence is taken as the intensity sequence of the channel; subsequently, the intensity obtained from a reference water sample with known turbidity in the same sampling period is used as the unit intensity, and the intensity of the sample is divided by the unit intensity, so that the scattering intensities in different sampling periods have comparability. In baseline flattening, the background can be estimated by replacing the morphological opening operation. Specifically, a sliding window with a structural element length of 51 sampling points is used to perform a sequence operation of erosion followed by dilation on each channel to obtain the background, and then subtract each point from the observation sequence. This method can better handle the background with slow fluctuations and isolated peaks. In time alignment, if it is not convenient to set the light intensity rectangular step, the wavelength of the light source can be switched twice quickly before sampling starts, and the common transition edge generated in the three channels at the switching moment is used as the alignment marker. This method does not change the chemical state of the measured water body, while providing a clear synchronization event.

[0034] Step 2: Call the preset endmember dictionary library to perform sparse coding endmember decomposition on the normalized multi-modal spectral joint observation data: according to the main peak and auxiliary peak markers of each endmember item in the endmember dictionary library, the absorption, fluorescence and scattering three channels are synchronously divided into a plurality of aligned segments; for each aligned segment, the overlap quantity is obtained by multiplying and summing each point of the aligned segment and its corresponding endmember item, and the channel consistency score of the aligned segment is obtained by taking the minimum value of the segment similarity scores of the three channels as the square root of the sum of the squares of each point of the aligned segment and the endmember item respectively, and dividing the overlap quantity by the product of the two scales; select the endmember with the highest score, and perform amplitude alignment at the peak height ratio at the main peak, and perform peak position fine tuning according to the segmented linear interpolation for the segments marked as deformable; the endmember item after amplitude alignment and peak position fine tuning is subtracted from the normalized multi-modal spectral joint observation data point by point, the negative values after subtraction are set to zero and small peaks are combined, and the process is repeated until the residual meets the stopping criterion, and the endmember abundance table is recorded.

[0035] Specifically, under the premise that the standardized multi-modal spectral joint observation data has been arranged according to a uniform sampling point sequence, a library of end member dictionary is called. The library of end member dictionary includes an absorption end member sub-library, a fluorescence end member sub-library and a scattering end member sub-library. Each end member has a same-named end member entry in the three sub-libraries. The end member entry includes labels of a main peak, an auxiliary peak, a background segment and a deformable segment, and is stored in the same sampling point sequence as the standardized multi-modal spectral joint observation data. According to the main peak and auxiliary peak labels of each end member entry in the library of end member dictionary, the absorption channel, the fluorescence channel and the scattering channel are synchronously divided into a plurality of aligned segments in the same sampling point sequence. The boundary of the aligned segment is centered on the main peak sampling point, and the same number of sampling points are taken on both sides; when the end member entry contains multiple auxiliary peaks that are adjacent to each other, the background segment adjacent to the main peak and the auxiliary peak is included in the same aligned segment. In actual operation, the typical width of a single-peak aligned segment is 31 to 61 sampling points, and the typical width of a multi-peak aligned segment is 81 to 121 sampling points. Synchronous division enables point-by-point calculation of the three channels under the same segment boundary, avoiding misalignment of peak-valley correspondence between channels.

[0036] For each aligned segment, the same length sequence of the aligned segment and the corresponding end member entry is taken in the absorption channel, the fluorescence channel and the scattering channel respectively, and point-by-point multiplication and summation of all sampling points are performed to obtain an overlap quantity; point-by-point squaring and summation of all sampling points are performed on the same length sequence of the aligned segment and the end member entry respectively, and the square roots of the summation results are obtained; the overlap quantity is divided by the product of the two scales to obtain the segment similarity score of the aligned segment in the channel. The minimum value of the segment similarity scores of the three channels of the same aligned segment is taken as the channel consistency score of the aligned segment. The reason for using the minimum value is that there is instantaneous turbidity disturbance or local occlusion of a channel in the mine water scene, and using the minimum value can preferentially exclude accidental high scores of a single channel, retain the segments that coincide with the three channels at the same time, reduce the number of attempts on the wrong end member in subsequent amplitude alignment and peak position fine tuning, thereby improving the stripping efficiency and stability.

[0037] For all aligned segments of an end member, the corresponding channel consistency scores are collected, aggregated by median, and the end member similarity score of the end member is obtained. The median is not sensitive to a small number of abnormal segments and can stably reflect the coincidence degree of the whole end member entry and the standardized multi-modal spectral joint observation data. The end member similarity scores of all end members are sorted, and the end member with the highest score is selected as the current target end member. If the difference between the first and second ranked end member similarity scores is less than 0.02, parallel review can be enabled in optional implementation, see the end of the specification.

[0038] After the current target endmember is selected, the main peak sampling point of the endmember signature is located in the absorption channel, the fluorescence channel and the scattering channel respectively. The peak height of the normalized multi-modal spectral joint observation data at the sampling point is read, the peak height of the endmember signature at the sampling point is read, the ratio of the two is calculated, and the amplitude alignment factor of the channel is obtained. The full segment of the current target endmember signature is adjusted in amplitude by the amplitude alignment factor. The advantage of selecting the main peak for amplitude alignment is that the main peak has the highest signal-to-noise ratio and is least affected by background uplift and random noise, so that a stable scale relationship can be established without adding additional references. For the section of the current target endmember signature marked as a deformable segment, a piecewise linear interpolation three-segment mapping is established around the sampling points of the main peak and the auxiliary peak. The control points of the three-segment mapping come from the peak position and the valley position of the normalized multi-modal spectral joint observation data in the same section. Apply the piecewise linear interpolation to the deformable segment of the endmember signature to make the peak position of the endmember signature consistent with the peak position of the normalized multi-modal spectral joint observation data. Piecewise linear interpolation is relatively direct and can achieve small-range peak position adjustment without introducing waveform distortion, avoiding ringing and false peaks caused by high-order interpolation.

[0039] The current target endmember signature after amplitude alignment and peak position fine-tuning is subtracted from the normalized multi-modal spectral joint observation data point by point in the three channels to obtain the residual sequence of the three channels. The negative values in the residual sequence are set to zero point by point to avoid reverse compensation in subsequent segment similarity score calculation. Adjacent small peaks are merged, specifically: when the distance between the two adjacent local peaks is less than 7 sampling points and the peak height of the smaller peak is lower than 20% of the peak height of the larger peak, the two are merged into one peak, and the sampling point of the larger peak is taken as the new peak position. The merging rule can remove the fragmented small peaks caused by subtraction errors, making the residual closer to the true unexplained components, and avoiding misjudgment as a new main peak in the next round of endmember selection.

[0040] After completing a stripping, the aforementioned alignment segment division, segment similarity score calculation, channel consistency score calculation, endmember similarity score aggregation and endmember selection are performed again on the updated normalized multi-modal spectral joint observation data. The cycle is repeated until any of the following stopping criteria is met: the residual sequence of the three channels no longer contains consecutive main peaks, specifically, the number of peaks in any channel that are continuously higher than twice the average value of the neighborhood and have a duration of more than 5 sampling points is 0; the cumulative number of stripped endmembers reaches 20; after two consecutive stripping, the highest value of the segment similarity score of the three channels is less than 0.15. The reason for using multiple stopping criteria is that in the mine water scene, there may be a small number of strong endmembers or a large number of weak endmembers, and the three criteria correspond to saturated interpretation, time limit and low separability protection three cases, which can end the cycle stably under different sample structures.

[0041] In each stripping, the amplitude alignment factors of the three channels are recorded by endmember name to form an endmember abundance table. Each row of the endmember abundance table corresponds to an endmember name, containing the amplitude alignment factors of the absorption channel, the amplitude alignment factors of the fluorescence channel, and the amplitude alignment factors of the scattering channel, as well as the median of the fragment similarity scores of the endmember in the three channels and the residual sequence energy at the time of stripping. Negative and abnormal items in the endmember abundance table are removed or set to zero. The endmember abundance table will be the direct basis for generating the subsequent steps of generating the pollutant category and relative level. The width of the alignment fragment can be automatically selected according to the distance between the main peak and the auxiliary peak of the endmember strip record. If the distance between the main peak and the nearest auxiliary peak is less than 15 sampling points, the typical width of the multi-peak alignment fragment is used; otherwise, the typical width of the single-peak alignment fragment is used. The calculation of the fragment similarity score only uses basic operations such as point-by-point multiplication, summation, squaring, square root, and division, and does not rely on any probability model or training data, which is convenient for fixed-point implementation on edge devices. The calculation of the amplitude alignment factor uses the peak height ratio of the main peak sampling point. When the main peak sampling point is blocked or saturated in a certain channel, the second-ranked peak in that channel can be used instead, provided that the peak height of the peak reaches more than 30% of the unit amplitude of that channel. The number of three-section mapping control points of piecewise linear interpolation is 3, which are located at the nearest valley on the left side of the main peak, the main peak sampling point, and the nearest valley on the right side of the main peak. The section outside the control points is not stretched or compressed, avoiding affecting the spectral shape that is not marked as a deformable fragment. The smoothing operation of merging small peaks can be performed once in a moving window of 5 to 9 sampling points for mean smoothing to suppress the sharp noise remaining after subtraction. The evaluation of the residual sequence energy is represented by the square root value of the point-by-point square sum of the non-negative residual, which is convenient for comparison with the numerical range of different channels. The evaluation only uses the current sample and does not involve any historical data.

[0042] In an optional implementation, when the difference between the first-ranked and second-ranked endmember similarity scores is less than 0.02, amplitude alignment and peak position fine-tuning can be performed on both endmembers simultaneously, and each is subtracted once. Compare the residual sequence energies after the two subtractions, and select the one with smaller residual as the effective stripping. This method can reduce the misselection when adjacent endmember entries are highly similar. If after the deformable fragment completes peak position fine-tuning, the residual still appears as a continuous main peak at the fragment boundary, the alignment fragment can be expanded by 5 to 10 sampling points on both sides and the fragment similarity score and channel consistency score are recalculated, and then the endmember is selected. This can cover the missing information caused by the initial boundary being too narrow. After completing 5 strippings, consistency checks are performed on different endmember entries from the same pollutant in the endmember abundance table. If the median of the fragment similarity scores of multiple endmember entries of the same pollutant in the three channels is higher than 0.5, the endmember entry with the main peak position closer to the global main peak of the standardized multi-modal spectral joint observation data is preferred to reduce redundant interpretation.

[0043] Step 3: Generating pollutant classes and relative grades based on endmember abundance table, and performing incremental domain adaptation and calibration loop: Reference measurements are taken at the beginning and end of sample processing, and compared with reference benchmarks stored in the system to obtain sample point offset and amplitude offset; the obtained sample point offset and amplitude offset are used to align the wavelengths and intensities of the three channels, and the entries in the endmember dictionary library are corrected synchronously; when the difference between the reference measurement and the reference benchmark exceeds the preset threshold, a complete sparse coding endmember decomposition recalculation is triggered, and the segment fine-tuning range is expanded according to the segmented linear interpolation rule of deformable segments; after recalculation, consistency check is performed on the generated classes and grades, if not consistent, the recalculation result is used as the reference, if the difference exceeds the discrimination threshold, the adjacent grade is taken, and the endmember that does not pass the consistency check is marked as to be confirmed and the final result is output.

[0044] Specifically, within the same sample processing, reference measurements are performed at the beginning and end respectively, and the reference measurements include water reference for the absorption channel, fixed fluorescent standard sheet reference for the fluorescence channel, and dark reference and water reference for the scattering channel. In this way, short-term drift within the same period can be captured without relying on historical data, such as slow decline of light source output, weak obstruction caused by probe window attachments, or slight shift of spectral axis. By comparing the two reference measurements with the reference benchmarks stored in the system item by item, the numerical range of the sample point offset and the amplitude offset can be obtained, thereby guiding the alignment of the wavelengths and intensities, and determining whether to trigger the recalculation of sparse coding endmember decomposition. In the absorption channel, the main peak and the secondary peak in the reference benchmark are located, and the common positions are near 275 nanometers and 730 nanometers. In the reference measurement at the beginning, the peak position index of the two peaks is searched point by point, and the difference from the reference benchmark peak position index is calculated to obtain the sample point offset. The same steps are repeated at the end, and if the absolute difference of the two sample point offsets exceeds 2 sample points, it is determined that there is a spectral axis change during this sample processing, and the sample point offset at the end is used for alignment. The amplitude offset is calculated using the peak height ratio: the peak height of the main peak of the reference measurement is divided by the peak height of the main peak of the reference benchmark to obtain the amplitude offset. For example, in the absorption channel, if the ratio is 0.92, it means that the intensity is reduced by about 8%. In the fluorescence channel, the strongest emission peak of the fixed fluorescent standard sheet is used as the reference point of the peak position and the peak height. In the scattering channel, the average value of the dark reference and the average value of the flat section of the water reference are used as the zero point and the unit amplitude reference, and the sample point offset and the amplitude offset are obtained by subtracting and dividing the corresponding values of the reference benchmark.

[0045] According to the aforementioned sampling point offset, the absorption channel, fluorescence channel and scattering channel are respectively shifted by an integer on the sampling point sequence; when the sampling point offset contains a non-integer part, a transition point is generated between the original sampling points using linear interpolation, so that the reference peak position coincides with the reference benchmark peak position. Then, according to the aforementioned amplitude offset, the intensity of the three channels is scaled, so that the reference peak height or the average value of the flat section is consistent with the reference benchmark. The reason for this processing is that the device gain change and the optical path change in the same period often show overall scaling of the intensity, and the slight drift of the spectral axis shows overall translation of the peak position. By aligning these two dimensions, the equivalent observation conditions under the reference benchmark can be restored without changing the peak shape, avoiding misjudgment of device changes as material changes. In order to maintain the consistency of the endmember dictionary library and the current channel state, the same sampling point offset and amplitude offset are applied to the corresponding channels of all endmember entries in the endmember dictionary library. Specifically, the same sampling point shift and linear interpolation are performed on the endmember entries of the three channels as the observation data, and the same scaling is performed on the intensity sequence as the observation data. For the sections marked as deformable segments, the deformable mark is not changed, and only the coordinates and amplitudes of the control points are re-labeled according to the coordinates and scales after this alignment. Synchronous correction can avoid the situation that the segment boundary and the peak position label are out of sync in subsequent recalculation, ensuring that the segment similarity score and the channel consistency score are still calculated in the same coordinate system.

[0046] The reference measurements at the start time and the end time are compared with the reference benchmark respectively, and if the sampling point offset of any channel exceeds 2 sampling points, or the amplitude offset deviates by more than 10%, a complete sparse coding endmember decomposition recalculation is triggered. Before recalculation, the segment fine-tuning range is expanded according to the piecewise linear interpolation rule of the deformable segment: 5 to 10 sampling points are added on the left and right sides of each deformable segment, and peak and valley control points are added in the new range, which are taken from the most significant peak and valley positions of the current observation data in the adjacent area. This expansion ensures that the endmember entries can still be consistent with the peak positions of the observation data in the case of slight stretching or compression of the peak position caused by device alignment, avoiding residual accumulation due to the narrowness of the deformable segment. Then, according to the endmember similarity score calculation rule, endmember selection and stripping process described in step 2, based on the aligned three channels and the synchronously corrected endmember dictionary library, recalculation is performed until the stop criterion is met.

[0047] Without triggering recalculation, the current endmember abundance table is directly used to look up the pollution class and relative rank from the scaling table stored in the system. With triggering recalculation, the same look-up process is performed using the endmember abundance table generated by recalculation. The rank interval can be set to 5 grades from 0 to 4 or 4 grades from 0 to 3, depending on the grade definition of the scaling table. By using look-up instead of fitting, the mapping can be completed with fixed computation on the edge device, and the adjustment can be made on demand by updating the scaling table without changing the algorithm flow.

[0048] The class and rank generated without recalculation are compared with those generated after recalculation. When the classes are inconsistent, the result after recalculation is used. When the difference between the relative ranks exceeds the discrimination threshold, the adjacent rank between the two is taken as the final rank; the discrimination threshold can be set to 1 rank step. For abnormal endmembers appearing in the comparison process, if the median of the fragment similarity scores of the endmember in the three channels is lower than 0.15, or the residual sequence energy of the endmember when stripped is higher than 3 times the standard deviation of the reference in the same channel, the endmember is marked as to be confirmed in the output and temporarily excluded from the statistics of the endmember abundance table. The final result includes the pollution class, relative rank, and alignment state prompt, such as whether the sampling point alignment occurs, whether the intensity alignment occurs, whether the recalculation is performed, and whether the fine-tuning range of the deformable fragment is expanded.

[0049] Within the same sample processing, the reference alignment at the start time is first used to generate the initial class and rank; if the reference alignment at the end time shows that the offset is aggravated and triggers recalculation, the class and rank after recalculation are used to overwrite the initial result. In this way, a closed loop of start alignment, processing, and end review is formed, so that short-term changes in device state are immediately corrected in the same sample processing, avoiding cross-period dependence on historical data. For the case where recalculation is not triggered, the reference alignment at the end time is used to confirm that the current alignment is still valid, and is output as the alignment record of the current sample together with the result.

[0050] In an alternative embodiment, the average shift of the two reference peaks is used as the final sampling point shift in the calculation of the sampling point shift. When the shift directions of the two reference peaks are opposite and the absolute value difference exceeds 1 sampling point, the reference peak closer to 500 nanometers is preferred, because this band is less affected by strong absorption edges in mine water and the peak position is more stable. For the absorption channel, 2 sampling points and a threshold of 8% are used, for the fluorescence channel, 1 sampling point and a threshold of 10% are used, and for the scattering channel, 3 sampling points and a threshold of 12% are used to adapt to the noise and stability characteristics of different channels. To control the calculation amount, the upper limit of the recalculation times is set to 2 times. When the relative grade difference of the end member abundance table after the first recalculation and the end member abundance table without recalculation does not exceed 1 grade for the main pollutants, the second recalculation is not performed. After expanding the segment fine-tuning range, if the new control points cause the local peak shape of the end member entry to appear obvious sharp peaks, the three-segment mapping can be changed to five-segment mapping, and the two new control points are located at 10 sampling points on the left and right of the main peak respectively, to more smoothly match the observed peak shape.

[0051] Reference Figure 2The present application synchronously obtains the spectral signals of three channels at the same measuring point by using a coaxial probe, the horizontal axis represents the wavelength range of 200nm to 700nm, and the vertical axis represents the intensity after normalization processing, the value range is 0 to 1.00. The solid line curve in the figure represents the spectral signal of the absorption channel, which presents two main absorption peaks at wavelengths of about 300nm and 550nm, the normalized intensities of the peaks are 0.72 and 0.58 respectively, and the baseline intensity is about 0.08. The absorption channel mainly reflects the absorption characteristics of the pollutant molecules in the water sample to the light of a specific wavelength, and the position and intensity of the absorption peak are directly related to the type and concentration of the pollutant. The dashed line curve represents the spectral signal of the fluorescence channel, which shows two fluorescence emission peaks at wavelengths of about 380nm and 680nm, the normalized intensities of the peaks are 0.55 and 0.48 respectively, and the baseline intensity is about 0.05. The fluorescence channel captures the fluorescence signal emitted by the pollutant molecules after excitation, and the peak shape is relatively wider than the absorption peak, which is due to the spectral broadening phenomenon caused by energy relaxation in the fluorescence emission process. The dotted line curve represents the spectral signal of the scattering channel, which shows scattering peaks at wavelengths of about 220nm and 750nm, the normalized intensities of the peaks are 0.40 and 0.35 respectively, and the overall baseline intensity is relatively high about 0.15. The scattering channel mainly reflects the scattering effect of suspended particulate matter in the water sample to the incident light, and its spectral characteristics are related to the particle size distribution and concentration of the particulate matter. The spectral curves of the three channels are distributed in the wavelength dimension, avoiding the complete overlap of the peak positions, which reflects the complementary advantage of multi-modal spectral joint observation. By synchronously collecting the data of the three channels, the present application can comprehensively characterize the optical characteristics of pollutants in the water sample from the three dimensions of absorption, fluorescence and scattering, and provide rich multi-dimensional information source for subsequent end member decomposition. The three curves all contain a certain degree of noise fluctuation, which is inevitable in actual measurement, and the subsequent end member decomposition algorithm has the robustness to noise.

[0052] As Figure 3As shown, the present application realizes the accurate identification and matching of endmembers by calculating the similarity score between the aligned segment and the endmember entry. The horizontal axis represents the wavelength range of 200-700 nm, and the vertical axis represents the intensity value range of 0-1.00. The thick solid curve in the figure represents the aligned segment extracted from the normalized multi-modal spectral joint observation data, which is centered on a certain main peak and reaches a peak value of about 0.78 at a wavelength of about 500 nm, with a peak width of about 65 nm. This aligned segment is divided according to the main peak and auxiliary peak markers in the endmember dictionary library and represents the characteristic section in the observation data that needs to be matched with the endmember library. The dashed curve represents the spectral characteristics of a certain candidate endmember entry in the endmember dictionary library in the corresponding wavelength band, and the main peak of the endmember entry is located at a wavelength of about 480 nm, with a peak height of about 0.72 and a peak width of about 60 nm. The endmember entry and the observation segment have a shift of about 20 nm in the peak position, a difference of about 0.06 in the peak height, and a slight difference in the peak width, which are caused by factors such as instrument drift and environmental changes in actual measurement, and are the problems that need to be solved by peak position fine-tuning. The gray filled area in the figure represents the overlap between the observation segment and the endmember entry in the intensity dimension, i.e., the area formed by taking the smaller value of the two curves at each sampling point. This overlap is the core basis for calculating the similarity score. The specific calculation process is shown in the formula on the right: first, multiply and sum the observation segment and the endmember entry at all sampling points to get the overlap value; then calculate the square sum of each sampling point of the observation segment and the endmember entry respectively, and take the square root as the normalization scale; finally, divide the overlap by the product of the two normalization scales to get the normalized similarity score. In this example, the calculated segment similarity score S = 0.876, which is a relatively high similarity value, indicating that the endmember entry has a good matching relationship with the observation segment. In the actual endmember decomposition process, the system will calculate the similarity score for all candidate endmembers in the endmember dictionary library, and select the endmember with the highest score as the current target endmember for amplitude alignment and peak position fine-tuning. It is worth noting that the present application calculates the segment similarity score on three channels respectively, and takes the minimum value of the three channel scores as the channel consistency score. This minimum value strategy can ensure that the selected endmember has good matching quality in all channels, avoiding the problem of strong correlation in a certain channel masking the weak correlation in other channels, thereby improving the reliability and accuracy of endmember identification.

[0053] As Figure 4As shown, the present application adopts the sparse coding endmember decomposition method to extract each endmember component from the mixed spectrum in turn through iterative stripping. The horizontal axis represents the wavelength range of 200-700 nm, and the vertical axis represents the intensity value range of 0-1.00. The thick solid line curve in the figure represents the original mixed spectrum, which is the spectral signal of a certain channel extracted from the standardized multi-modal spectral joint observation data. The mixed spectrum shows three obvious peaks at wavelengths of about 300 nm, 520 nm and 720 nm, with peak heights of about 0.48, 0.85 and 0.40, respectively, and a baseline intensity of about 0.03. The three peaks correspond to the three different pollutant components present in the water sample, and their spectral characteristics superimpose on each other to form the observed mixed spectrum. The complexity of the mixed spectrum is the main difficulty in pollutant identification, and it needs to be decomposed into a linear combination of each pure endmember component through endmember decomposition. The dashed curve represents the first endmember identified by similarity calculation, which corresponds to the highest peak in the mixed spectrum at 520 nm. According to the entries in the endmember dictionary library, after amplitude alignment, the peak shape of this endmember matches the corresponding peak in the mixed spectrum. The annotations in the figure show that the amplitude alignment factor a = 0.85, indicating that the standard intensity of this endmember in the dictionary library needs to be multiplied by a coefficient of 0.85 to match the amplitude of the observation data. This coefficient also represents the relative abundance or concentration of the endmember in the mixed spectrum. The dotted curve represents the residual sequence obtained after stripping the first endmember from the original mixed spectrum. The calculation method of the residual is to subtract the intensity value of the endmember after amplitude alignment from the intensity value of the mixed spectrum at each sampling point, and set the negative values to zero point by point. As can be seen from the figure, the peak at 520 nm in the residual curve has been completely stripped, and the remaining peaks are mainly at 300 nm and 720 nm. The shape and position of these two peaks remain basically the same as the corresponding peaks in the original mixed spectrum, but the amplitude is slightly reduced, because the endmember entries may contain certain background fragments or auxiliary peaks that contribute slightly to these positions. The residual energy E = 18.5% is also marked in the figure, which is obtained by squaring and summing the intensity values of the residual sequence at all sampling points and normalizing it. It reflects the energy proportion of the remaining signal after stripping an endmember. The residual energy is an important basis for determining whether to continue endmember stripping. When the residual energy is lower than a pre-set threshold (such as 3%), it means that the mixed spectrum has been sufficiently decomposed, and the iteration process can be stopped. In this example, the residual energy is 18.5%, which is significantly higher than the stop threshold, so the next round of endmember identification and stripping needs to be performed on the residual sequence, until the peaks at 300 nm and 720 nm are also matched to the corresponding endmember entries and stripped, and finally the abundance table of each endmember and the low-energy residual that meets the stopping criterion are obtained. This iterative strategy of successive stripping can effectively handle multi-component mixed samples and achieve accurate identification and quantitative analysis of complex pollutant systems.

[0054] While specific embodiments of the application have been described above, it will be appreciated that those skilled in the art within the scope of the application can make modifications, substitutions and changes in the form and details of the method and system described above without departing from the spirit and scope of the application. For example, it is well within the scope and spirit of the application to combine various of the illustrative steps of the above methods to carry out essentially the same function to achieve essentially the same result in substantially the same way. Accordingly, the scope of the application should be determined sole by the appended claims.

Claims

1. A method for in-situ identification and drift adaptation of multimodal spectra of mine water pollutants, characterized in that, The method includes: Step 1: Simultaneously acquire signals from three channels: absorption, fluorescence, and scattering. Perform preprocessing steps in sequence, including dark signal removal, baseline flattening, time alignment, and amplitude normalization, to form standardized multimodal spectral joint observation data. Step 2: Call the preset endmember dictionary library and perform sparse coding endmember decomposition on the standardized multimodal spectral joint observation data: According to the main peak and auxiliary peak labels of each endmember entry in the endmember dictionary library, the absorption, fluorescence, and scattering channels are simultaneously divided into several aligned segments; for each aligned segment, the overlap is obtained by multiplying the aligned segment and its corresponding endmember entry point by point; the overlap is divided by the product of the two scales to obtain the segment similarity score, and the minimum value among the segment similarity scores of the three channels is taken as the channel consistency score of this aligned segment; the endmember with the highest score is selected, and amplitude alignment is performed at the main peak by the peak height ratio; the peak position is fine-tuned by piecewise linear interpolation for segments marked as deformable; the endmember entries after amplitude alignment and peak position fine-tuning are subtracted point by point from the standardized multimodal spectral joint observation data, the negative values ​​after subtraction are set to zero and small peaks are merged, and this process is repeated until the residual meets the stopping criterion; the alignment factor at each subtraction is recorded to generate an endmember abundance table; Step 3: Generate pollutant categories and relative levels based on the endmember abundance table, and execute an incremental domain adaptation and calibration loop: Reference measurements are performed at the beginning and end of sample processing, and the sampling point offset and amplitude offset are obtained by comparing them with the reference standard stored in the system; the obtained sampling point offset and amplitude offset are used to align the wavelength and intensity of the three channels, and the entries in the endmember dictionary are corrected simultaneously; when the difference between the reference measurement and the reference standard exceeds a preset threshold, a complete sparse coding endmember decomposition recalculation is triggered, and the segment fine-tuning range is expanded according to the piecewise linear interpolation rules of deformable segments; after recalculation, the generated categories and levels are checked for consistency. If there is any inconsistency, the recalculated result shall prevail. If the difference exceeds the discrimination threshold, the adjacent level shall be taken. Endmembers that fail the consistency check shall be marked as pending confirmation and the final result shall be output. Step 1 specifically includes: using a coaxial probe to synchronously acquire signals from the absorption channel, fluorescence channel and scattering channel at the same measurement point to form multimodal spectral joint observation data; performing dark signal removal, baseline flattening, channel time alignment and amplitude normalization on the multimodal spectral joint observation data in sequence; performing Raman and Rayleigh strip masking on the fluorescence channel and interpolating with adjacent effective bands; performing angular consistency tuning and turbidity reference binding on the scattering channel to obtain standardized multimodal spectral joint observation data.

2. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 1, characterized in that, In step 2, the endmember dictionary includes an absorption endmember sub-library, a fluorescence endmember sub-library, and a scattering endmember sub-library. Each endmember in the dictionary has a corresponding entry with the same name in the three sub-libraries, and the three entries are stored according to the same sampling point sequence. Each entry contains four types of labels: the sampling point where the main peak is located, the auxiliary peak group, the background segment, and the deformable segment. The standardized multimodal spectral joint observation data are unified across the three channels to the sampling point sequence used by the endmember dictionary. Using the main peak and auxiliary peak labels in each endmember entry of the endmember dictionary as anchor points, the three channels are simultaneously divided into several aligned segments, and each aligned segment corresponds to a main peak and its adjacent background segment.

3. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 2, characterized in that, In step 2, the consistency score is calculated as follows: First, multiply each aligned segment by its corresponding endmember entry point by point and sum the results to obtain the overlap. Then, sum the squares of each aligned segment and endmember entry point by point and take the square root of the sum as a normalization scale. Divide the overlap by the product of the two normalization scales to obtain the segment similarity score of the aligned segment in each channel. Take the minimum value among the segment similarity scores of the three channels as the segment channel consistency score of the aligned segment. Take the median of the segment channel consistency scores of all segments corresponding to any endmember as the endmember similarity score of the endmember. Select the endmember with the highest score among all endmembers as the current target endmember. Locate the main peak sampling point of the current target endmember in each of the three channels. Calculate the ratio of the peak height of the standardized multimodal spectral joint observation data at the target main peak sampling point to the peak height of the current target endmember entry. Use the calculated ratio as the amplitude alignment factor. The amplitude alignment factor is used to simultaneously adjust the amplitude of the current target end-member entry in all three channels.

4. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 3, characterized in that, In step 2, for the segments marked as deformable fragments in the current target endmember entry, a three-segment broken line mapping is established around the sampling points of the main peak and auxiliary peaks. The control points of the segment mapping are taken from the peak and valley positions of the standardized multimodal spectral joint observation data in the same segment. Based on the established segment mapping, point-by-point linear interpolation deformation is performed on the current target endmember entry in the corresponding segment to make the peak position of the endmember entry consistent with the peak position of the standardized multimodal spectral joint observation data.

5. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 4, characterized in that, In step 2, the loop process is as follows: the current target endmember entry, which has been aligned and deformed, is subtracted point by point from the standardized multimodal spectral joint observation data of the three channels, and the subtraction result is recorded as the residual sequence; the negative values ​​in the residual sequence are set to zero point by point, and adjacent small peaks are merged and smoothed to obtain the updated standardized multimodal spectral joint observation data; if the total energy of the residual sequence in the three channels is lower than the preset threshold, or the number of stripped endmembers reaches the preset upper limit, the loop stops; otherwise, the calculation rules of the endmember similarity score are returned, and endmember selection, fine-tuning and stripping are performed on the updated standardized multimodal spectral joint observation data. The amplitude alignment factor at each stripping step is recorded by endmember name to form an endmember abundance table, and the corresponding fragment similarity score and residual sequence energy are recorded as matching criteria for the corresponding endmembers. Negative and outlier items in the endmember abundance table are removed or set to zero. Based on the endmember abundance table and the calibration table stored with the system in the endmember dictionary, the pollutant category and relative level are obtained by looking up the table in the calibration interval of each endmember. When the same pollutant is hit in multiple endmember entries, the first entry listed is selected as the final source of the pollutant according to the priority order recorded in the calibration table, and the entries that are not hit are marked as undetected.

6. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 5, characterized in that, In step 3, at the beginning and end of each sample processing, the built-in reference component is triggered to perform a reference measurement to obtain the reference absorption signal, reference fluorescence signal and reference scattering signal; the position and amplitude of the main peak of the reference measurement are compared point by point with the reference reference stored with the system to obtain the sampling point offset and amplitude offset of the three channels. Using the sampling point offset, the sampling points are translated and the endpoints are extrapolated to align the wavelength direction of the three channels; using the amplitude offset, the amplitude is stretched or compressed to align the intensity direction of the three channels; the correction only applies to the current sample processing flow.

7. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 6, characterized in that, In step 3, the sampling point offset and amplitude offset obtained from channel correction are simultaneously applied to the corresponding channels of all endmember entries in the endmember dictionary. For segments marked as deformable fragments, the peak and valley positions of the reference measurement are used as new control points, and the endmember entries are deformed using piecewise linear interpolation to ensure that the endmember entries are consistent with the three corrected channels. When the difference between the reference measurement and the reference benchmark exceeds a preset threshold, a synchronous correction of the endmember entries and a complete sparse coding endmember decomposition recalculation are immediately triggered. During the recalculation, the endmember similarity score calculation rules, endmember selection, and stripping process of step 2 are followed. After the recalculation is completed, the endmember abundance table and residual sequence energy before and after the recalculation are compared. If the difference still exceeds the preset threshold, the control point range of the deformable segment is expanded according to the deformable fragment fine-tuning rules and the recalculation is repeated until the difference falls back to within the preset threshold or reaches the preset number of times limit.

8. The multimodal spectral in-situ identification and drift adaptive method for mine water pollutants as described in claim 7, characterized in that, In step 3, a consistency check is performed between the recalculated pollutant categories and relative levels and the results before recalculation: when the categories are inconsistent, the recalculated results shall prevail; when the difference in relative levels exceeds the preset level step size, the adjacent level between the results before and after recalculation shall be taken as the final level. Endmembers that fail the consistency check are removed from the endmember abundance table and marked as pending confirmation in the output. At the end of the current sample processing, the sampling point offset, amplitude offset and deformation control points of deformable segments obtained from the current reference measurement are written into the equipment operating parameters for direct use in the channel correction and endmember entry synchronous correction of the next sample, without retaining any original observation data.

Citation Information

Patent Citations

  • Water pollutant detection method, system and equipment based on spectral analysis

    CN120558868A

  • Water quality pollution multi-dimensional analysis system

    CN120778663A