Multi-mode spectrum in-situ identification and drift self-adaption method for mine water pollutants

By employing a multimodal spectral recognition method, combined with an end-member dictionary and adaptive calibration technology, the problem of real-time monitoring and identification accuracy of mine water pollutants was solved, achieving high-precision and low-maintenance online monitoring of mine water.

CN120992533AActive Publication Date: 2025-11-21CHINA INST OF WATER RESOURCES & HYDROPOWER RES +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511508507.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing methods for detecting pollutants in mine water cannot achieve real-time monitoring. Single-channel spectral detection cannot fully reflect the coupling effect of multiple pollutants in the water. Multi-channel fusion methods do not consider the inconsistency between channels, resulting in low identification accuracy. Furthermore, equipment drift leads to large identification errors.

Method used

A multimodal spectral identification method is adopted. By simultaneously acquiring absorption, fluorescence and scattering channel signals, sparse coding endmember decomposition is performed using an endmember dictionary, and an incremental domain adaptive and reference calibration loop is combined to maintain spectral consistency and automatically correct drift, thereby achieving multi-channel collaborative identification.

Benefits of technology

It achieves high-resolution identification of complex mixed pollutants in mine water, avoids identification errors caused by equipment drift and environmental changes, has self-calibration characteristics, reduces maintenance costs, and provides a high-precision solution for real-time on-site monitoring of pollutants in mine water.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120992533A_ABST
    Figure CN120992533A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of optical detection, and particularly relates to a mine water pollutant multi-mode spectrum in-situ recognition and drifting self-adaption method which comprises the following steps: step 1, synchronously acquiring signals comprising an absorption channel, a fluorescence channel and a scattering channel, and sequentially performing pretreatment of dark signal removal, baseline leveling, time alignment and amplitude normalization to obtain a signal sequence; standardized multi-mode spectrum combined observation data are formed; 2, calling a preset end member dictionary database, and executing sparse coding end member decomposition on the standardized multi-mode spectrum joint observation data; and 3, generating a pollutant category and a relative grade based on an end member abundance table, and executing an incremental domain self-adaption and calibration loop. According to the invention, the distinguishing capability of complex mixed pollutants in mine water is obviously improved; identification errors caused by equipment aging, light source attenuation and environment change are avoided; the system has self-calibration and self-adaption characteristics without manual intervention, can stably operate for a long time, and reduces the maintenance cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of optical detection technology, specifically relating to a multimodal spectral in-situ identification and drift adaptive method for mine water pollutants. Background Technology

[0002] Mine water is a common byproduct fluid in mining operations, with a complex composition typically including suspended particles, metal ions, organic solutes, and various inorganic salts. Real-time monitoring and pollutant identification of mine water are crucial for ensuring ecological safety and groundwater resource utilization in mining areas. Currently, methods for detecting pollutants in mine water mainly fall into three categories: offline laboratory chemical analysis, single-channel spectroscopic detection, and multispectral remote sensing inference. Traditional laboratory chemical analysis methods, such as atomic absorption spectrometry, ion chromatography, and ICP-OES, while highly accurate, suffer from complex sampling, transportation, and sample preparation processes, failing to meet the needs of in-situ real-time monitoring. For dynamic water flows in deep wells or enclosed mining areas, these methods have significant limitations in terms of sample representativeness and time response.

[0003] Single-channel spectral detection methods were first adopted in online monitoring of mine water. Absorption spectroscopy is used to determine the concentration of metal ions and inorganic salts, fluorescence spectroscopy to characterize organic pollutants, and scattering spectroscopy to estimate turbidity and particulate content. However, a single channel cannot comprehensively reflect the coupling effects of multiple pollutants in the water. For example, in iron- and manganese-containing mine water, absorption signals are easily masked by colloidal scattering background; fluorescence channels are significantly affected by turbidity, often exhibiting signal drift and quenching; and while scattering signals can reflect particle concentration, they lack specific material characteristics. To address these issues, some studies have proposed multi-channel spectral fusion approaches, improving identification accuracy by jointly analyzing absorption, fluorescence, and scattering information. However, most existing multi-channel fusion methods are based on simple linear combinations or principal component transformations, failing to consider the inconsistencies in response dynamics and noise structure between different channels, resulting in limited fusion effectiveness. Regarding spectral data analysis algorithms, existing research often employs methods such as full-spectrum regression, non-negative matrix factorization, independent component analysis, or sparse representation. Full-spectrum regression relies on a large number of calibrated samples and has poor generalization ability to new scenarios or drifting samples. Nonnegative matrix factorization and independent component analysis often assume a fixed number of endmembers, while the pollutant composition of mine water changes with time and flow conditions. Although traditional sparse representation methods can recover endmembers to some extent, they cannot effectively distinguish the mixed features between similar spectral shapes, and the decomposition results are prone to instability when drift exists. In addition, most algorithms only perform decomposition analysis on a single channel and do not consider the consistency constraints between endmembers across channels, making it difficult to achieve accurate and synchronous identification of pollutants. Summary of the Invention

[0004] Therefore, the main objective of this invention is to provide a multi-modal spectral in-situ identification and drift-adaptive method for mine water pollutants. This method comprehensively utilizes spectral information from absorption, fluorescence, and scattering channels, achieving unified analysis of multi-channel spectra and pollutant category identification through synchronous acquisition, standardization processing, and sparse coding endmember decomposition. The method leverages the channel correspondence structure of the endmember dictionary to maintain spectral consistency during endmember selection, amplitude alignment, and peak fine-tuning, and achieves high-resolution decomposition through point-by-point subtraction and residual auditing. During long-term operation, the system automatically corrects wavelength and intensity drift through incremental domain adaptation and a reference calibration loop, ensuring stable and reliable results. The beneficial effects of this invention are: it achieves multi-channel collaborative pollutant identification, significantly improving the resolution of complex mixed pollutants in mine water; it avoids identification errors caused by equipment aging, light source attenuation, and environmental changes; it possesses self-calibration and adaptive characteristics without manual intervention, enabling long-term stable operation and reducing maintenance costs, providing a high-precision and sustainable technical solution for real-time on-site monitoring of mine water pollutants.

[0005] The technical solution adopted in this invention is as follows: Multimodal spectral in-situ identification and drift adaptive method for mine water pollutants, including: Step 1: Simultaneously acquire signals from three channels: absorption, fluorescence, and scattering. Perform preprocessing steps in sequence, including dark signal removal, baseline flattening, time alignment, and amplitude normalization, to form standardized multimodal spectral joint observation data. Step 2: Call the preset endmember dictionary library and perform sparse coding endmember decomposition on the standardized multimodal spectral joint observation data: According to the main peak and auxiliary peak labels of each endmember entry in the endmember dictionary library, the absorption, fluorescence, and scattering channels are simultaneously divided into several aligned segments; for each aligned segment, the overlap is obtained by multiplying the aligned segment and its corresponding endmember entry point by point; the overlap is divided by the product of the two scales to obtain the segment similarity score, and the minimum value among the segment similarity scores of the three channels is taken as the channel consistency score of this aligned segment; the endmember with the highest score is selected, and amplitude alignment is performed at the main peak by the peak height ratio; the peak position is fine-tuned by piecewise linear interpolation for segments marked as deformable; the endmember entries after amplitude alignment and peak position fine-tuning are subtracted point by point from the standardized multimodal spectral joint observation data, the negative values ​​after subtraction are set to zero and small peaks are merged, and this process is repeated until the residual meets the stopping criterion; the alignment factor at each subtraction is recorded to generate an endmember abundance table; Step 3: Generate pollutant categories and relative levels based on the endmember abundance table, and execute an incremental domain adaptation and calibration loop: Perform reference measurements at the beginning and end of sample processing, and compare them with the reference benchmark stored with the system to obtain the sampling point offset and amplitude offset; use the obtained sampling point offset and amplitude offset to perform wavelength and intensity alignment on the three channels, and simultaneously correct the entries in the endmember dictionary; when the difference between the reference measurement and the reference benchmark exceeds a preset threshold, trigger a complete sparse coding endmember decomposition recalculation, and expand the fragment fine-tuning range according to the piecewise linear interpolation rules of deformable fragments; after recalculation, perform a consistency check on the generated categories and levels. If they are inconsistent, the recalculation result shall prevail. If the difference exceeds the discrimination threshold, the adjacent level shall be taken. Endmembers that fail the consistency check are marked as pending confirmation and the final result is output.

[0006] Furthermore, step 1 specifically includes: using a coaxial probe to simultaneously acquire signals from the absorption channel, fluorescence channel, and scattering channel at the same measurement point to form multimodal spectral joint observation data; sequentially performing dark signal removal, baseline flattening, channel time alignment, and amplitude normalization on the multimodal spectral joint observation data; performing Raman and Rayleigh strip masking on the fluorescence channel and interpolating with adjacent effective bands; and performing angular consistency tuning and turbidity reference binding on the scattering channel to obtain standardized multimodal spectral joint observation data.

[0007] Furthermore, in step 2, the endmember dictionary includes an absorption endmember sub-library, a fluorescence endmember sub-library, and a scattering endmember sub-library; each endmember in the dictionary has a corresponding entry with the same name in the three sub-libraries, and the three entries are stored according to the same sampling point sequence; each entry contains four types of labels: the sampling point where the main peak is located, the auxiliary peak group, the background segment, and the deformable segment; the standardized multimodal spectral joint observation data are unified on the three channels to the sampling point sequence used by the endmember dictionary; using the main peak and auxiliary peak labels in each endmember entry of the endmember dictionary as anchor points, the three channels are simultaneously divided into several aligned segments, and each aligned segment corresponds to a main peak and its adjacent background segment.

[0008] Further, in step 2, the consistency score is calculated as follows: First, multiply each aligned segment by its corresponding endmember entry point by point and sum the results to obtain the overlap. Then, sum the squares of each aligned segment and endmember entry point by point and take the square root of the sum as a normalization scale. Divide the overlap by the product of the two normalization scales to obtain the segment similarity score of the aligned segment in each channel. Take the minimum value among the segment similarity scores of the three channels as the segment channel consistency score of the aligned segment. Take the median of the segment channel consistency scores of all segments corresponding to any endmember as the endmember similarity score of the endmember. Select the endmember with the highest score among all endmembers as the current target endmember. Locate the main peak sampling point of the current target endmember in each of the three channels. Calculate the ratio of the peak height of the standardized multimodal spectral joint observation data at the target main peak sampling point to the peak height of the current target endmember entry. Use the calculated ratio as the amplitude alignment factor. Adjust the amplitude of the current target endmember entry simultaneously in the three channels using the amplitude alignment factor.

[0009] Furthermore, in step 2, for the segments marked as deformable fragments in the current target endmember entry, a segmented mapping is established around the sampling points of the main peak and auxiliary peaks, with three segments connected by broken lines. The control points of the segmented mapping are taken from the peak and valley positions of the standardized multimodal spectral joint observation data in the same segment. Based on the established segmented mapping, point-by-point linear interpolation deformation is performed on the current target endmember entry in the corresponding segment to make the peak position of the endmember entry consistent with the peak position of the standardized multimodal spectral joint observation data.

[0010] Furthermore, in step 2, the loop process is as follows: the current target endmember entry, after amplitude alignment and deformation, is subtracted point by point from the standardized multimodal spectral joint observation data of the three channels, and the subtraction result is recorded as the residual sequence; negative values ​​in the residual sequence are set to zero point by point, and adjacent small peaks are merged and smoothed to obtain the updated standardized multimodal spectral joint observation data; if the total energy of the residual sequence in the three channels is lower than the preset threshold, or the number of stripped endmembers reaches the preset upper limit, the loop stops; otherwise, the calculation rules for endmember similarity scores are returned, and endmember selection is performed on the updated standardized multimodal spectral joint observation data. Selection, fine-tuning, and stripping; recording the amplitude alignment factor by endmember name at each stripping step to form an endmember abundance table, and recording the corresponding segment similarity score and residual sequence energy as matching criteria for the corresponding endmembers; removing or zeroing negative and outlier entries in the endmember abundance table; looking up the pollutant category and relative level in the calibration interval of each endmember based on the endmember abundance table and the calibration table stored with the system in the endmember dictionary; when the same pollutant is hit in multiple endmember entries, selecting the first listed entry as the final source of the pollutant according to the priority order recorded in the calibration table, and marking the undetected entries as not detected.

[0011] Furthermore, in step 3, at the beginning and end of each sample processing step, the built-in reference component is triggered to perform a reference measurement to obtain the reference absorption signal, reference fluorescence signal, and reference scattering signal. The position and amplitude of the main peak of the reference measurement are compared point by point with the reference reference stored in the system to obtain the sampling point offset and amplitude offset of the three channels. Using the sampling point offset, the sampling point is translated and the endpoints are extrapolated to complete the alignment of the wavelength direction. Using the amplitude offset, the amplitude is stretched or compressed to complete the alignment of the intensity direction. The correction only applies to the current sample processing flow.

[0012] Furthermore, in step 3, the sampling point offset and amplitude offset obtained from channel correction are simultaneously applied to the corresponding channels of all endmember entries in the endmember dictionary. For segments marked as deformable fragments, the peak and valley positions of the reference measurement are used as new control points, and the endmember entries are deformed using piecewise linear interpolation to ensure that the endmember entries are consistent with the three corrected channels. When the difference between the reference measurement and the reference benchmark exceeds a preset threshold, a synchronous correction of the endmember entries and a complete sparse coding endmember decomposition recalculation are immediately triggered. During the recalculation, the endmember similarity score calculation rules, endmember selection, and stripping process of step 2 are followed. After the recalculation is completed, the endmember abundance table and residual sequence energy before and after the recalculation are compared. If the difference still exceeds the preset threshold, the control point range of the deformable segment is expanded according to the deformable fragment fine-tuning rules and the recalculation is repeated until the difference falls back to within the preset threshold or reaches the preset number of times limit.

[0013] Furthermore, in step 3, a consistency check is performed on the recalculated pollutant categories and relative levels with the results before recalculation: when the categories are inconsistent, the recalculated results shall prevail; when the difference in relative levels exceeds the preset level step size, the adjacent level between the results before and after recalculation shall be taken as the final level; endmembers that fail the consistency check shall be removed from the endmember abundance table and marked as pending confirmation in the output; at the end of the current sample processing, the sampling point offset, amplitude offset and deformation control points of deformable segments obtained from this reference measurement shall be written into the equipment operating parameters for direct use in the channel correction and endmember entry synchronous correction of the next sample, without retaining any original observation data.

[0014] By employing the above technical solutions, this invention achieves the following beneficial effects: It establishes a unified identification framework for multimodal spectroscopy by simultaneously acquiring, standardizing, and decomposing sparsely encoded endmembers across three channels: absorption, fluorescence, and scattering. This method, through a multi-channel correspondence structure of the endmember dictionary, ensures that each endmember maintains consistency in peak position, shape, and response characteristics across different spectral channels, thereby enabling the simultaneous identification of multiple pollutant components against complex mine water backgrounds. Compared to existing single-channel or linear fusion methods, this invention introduces a point-by-point subtraction and fragment consistency comparison mechanism during endmember selection and spectral shape stripping, effectively avoiding inter-channel interference and misidentification of overlapping peaks. Through piecewise linear interpolation fine-tuning, peak position adaptive correction can be achieved without altering the overall spectral characteristics, resulting in higher stability of the identification process against environmental disturbances and sample fluctuations. Regarding equipment and environmental drift, this invention proposes an incremental domain adaptive and reference calibration loop mechanism, performing reference measurements and dynamically correcting wavelength and intensity coordinates at the beginning and end of sample processing, enabling the system to automatically maintain accuracy under long-term unattended conditions. By combining trigger condition recalculation with synchronous correction of the end-member dictionary, drift errors caused by factors such as light source attenuation, probe scaling, and temperature changes can be eliminated in real time, thereby maintaining the repeatability and traceability of the identification results. Overall, this invention achieves multi-channel collaborative detection of mine water pollutants, spectral-level self-consistent decomposition, and continuous drift adaptation under field conditions, providing a highly robust, high-resolution, and low-maintenance-cost solution for online monitoring of mine water. Attached Figure Description

[0015] Figure 1 A schematic flowchart illustrating the multimodal spectral in-situ identification and drift adaptive method for mine water pollutants provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of synchronous acquisition of three-channel spectral signals provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the calculation of the similarity score of aligned segments provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the principle of the endmember stripping and residual calculation process provided in an embodiment of the present invention. Detailed Implementation

[0016] All features disclosed in this specification, or steps in all methods or processes disclosed herein, may be combined in any way, except for mutually exclusive features and / or steps.

[0017] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0018] refer to Figure 1 A multimodal spectral in-situ identification and drift adaptive method for mine water pollutants, including: Step 1: Simultaneously acquire signals from three channels: absorption, fluorescence, and scattering. Perform preprocessing steps in sequence, including dark signal removal, baseline flattening, time alignment, and amplitude normalization, to form standardized multimodal spectral joint observation data.

[0019] Specifically, at the same measurement point, the absorption, fluorescence, and scattering channels are placed on a coaxial observation path. A single system clock is used to issue the sampling start command, ensuring that the first sampling point of each of the three channels is synchronized. To avoid minor phase differences caused by mechanical vibration and water flow pulsation, the sampling duration is set to an integer multiple of the single flow duration, for example, a sampling duration of 3 seconds, and the light source is kept stable within this duration. The absorption and fluorescence channels are scanned sequentially with a uniform wavelength step, for example, from 230 nm to 900 nm in 1 nm steps, resulting in 671 sampling points for each channel. The scattering channel records the intensity sequence within the same duration and resamples to the same number of sampling points as the absorption and fluorescence channels using linear interpolation at the end of the sampling period. This ensures that the three channels view the same water body both spatially and temporally, and form a consistent data structure, facilitating subsequent point-by-point processing.

[0020] A shading measurement was performed before and after sampling, keeping the optical path and probe position unchanged during the shading measurements, with only the light source turned off. Two sets of dark signals were obtained, and the point-by-point average was used as the dark signal for that sample. If the maximum difference between the two sets of dark signals exceeded 5 counts, it was determined that the ambient temperature or electronic noise had changed rapidly, and the shading measurement was immediately repeated until the maximum difference did not exceed 5 counts. The dark signals were subtracted point-by-point from the observation sequences of the three channels to obtain the multimodal spectral joint observation data after dark signal removal. By performing shading measurements before and after sampling, the fixed offset caused by temperature drift and readout bias can be offset, preventing weak signal areas from being raised by the background noise and avoiding misleading subsequent baseline flattening. In a typical scenario with a ground temperature of 25 degrees Celsius, the average dark signal value is usually between 12 and 18 counts, and the random fluctuation range near zero intensity after removal is maintained at 3 to 5 counts.

[0021] After removing the dark signal, the slowly changing background is estimated and subtracted channel by channel. For absorption and fluorescence channels, a moving window lower envelope estimation is used. Specifically, a moving window of 101 sampling points is used for each channel, and the 20th percentile within each window is calculated as the local background for that window. The local backgrounds of all windows are then linearly stitched together to form a full-spectrum background curve. The background curve is subtracted point by point from the observation sequence of the corresponding channel, and sampling points with negative results are set to zero. For scattering channels, a flat, unstructured region is first identified as a reference. For example, 200 consecutive sampling points are taken near the median value of the channel sequence, and the average value of this region is used as the scattering background, which is then subtracted point by point from the entire sequence. The percentile background and flat region background are used because suspended particles and colored dissolved substances in mine water cause slowly fluctuating background elevation, while the true absorption and fluorescence peaks exhibit narrow shapes or local protrusions. By prioritizing the fitting of the slowly changing lower envelope, peak shape can be preserved and background eliminated, ensuring that subsequent time alignment and amplitude normalization are not affected by background fluctuations. After the baseline is flattened, the amplitude of common background fluctuations can be compressed from the original ten percent to the one percent level.

[0022] Although the three channels begin sampling at the same time, the peak values ​​may exhibit slight shifts on the time axis due to differences in integration time and readout methods within each channel. To address this, a rectangular intensity step is introduced within the sampling duration, occurring approximately 1 second after the start of sampling, with a step width of 100 milliseconds. All three channels respond to this step. The sampling point index where the rising edge of the step occurs is extracted for each channel, and the indices of the three channels are aligned with the index of the absorption channel. If a difference in non-integer sampling points occurs, linear interpolation shifts are performed on the fluorescence and scattering channels to ensure the rising edge occurs at the same sampling point. This approach eliminates time shifts caused by different channel sampling links without introducing complex calculations, using the same physical event as a common marker. This ensures a one-to-one match between the peaks and valleys of the three channels, facilitating consistent segmentation and point-by-point calculations across the three channels. In practical applications, the maximum shift when unaligned is typically 20 to 120 milliseconds, while the residual shift after alignment is less than 2 milliseconds.

[0023] To eliminate amplitude inconsistencies caused by differences in exposure time, incident light fluctuations, and fiber coupling efficiency, a two-stage amplitude normalization is employed. The first stage is zero-point normalization: in the absorption and fluorescence channels, low-signal regions far from the known main peak position are selected, such as the 400nm-420nm and 820nm-840nm segments. The average value for each segment is calculated, and this average value is subtracted point-by-point from the entire sequence to bring the low-signal region close to zero. The scattering channel undergoes the same processing using the average value of the aforementioned flat segments. The second stage is unit amplitude normalization: in each channel, the main peak of the current sample is identified, defined as a peak more than three times higher than the neighborhood average and with a duration exceeding five sampling points. The peak height of this main peak is used as the unit amplitude for that channel. The entire sequence is divided point-by-point by this unit amplitude until the peak height of the main peak is approximately 1. The main peak is chosen as the unit amplitude because it has the highest signal-to-noise ratio, stably reflecting the optical path length and light source state of the current sample, thus bringing the intensity scale between channels into a comparable range. If the current sample does not contain a significant main peak, the standard deviation of the aforementioned flat section is used as the unit amplitude to avoid amplifying pure noise samples. After completing the two-level amplitude normalization, the numerical range of the three channels is unified between 0 and 1, which facilitates subsequent point-by-point subtraction and similarity calculation on the three channels.

[0024] After completing dark signal removal, baseline flattening, time alignment, and amplitude normalization, the data from the three channels are combined according to a unified sampling point sequence to form standardized multimodal spectral joint observation data. Each channel contains 671 sampling points. The data stores the original count values ​​as 16-bit unsigned integers and the normalized values ​​as floating-point numbers, ranging from 0 to 1, along with index values ​​indicating the sampling start time and time alignment. The standardized multimodal spectral joint observation data serves as direct input for subsequent steps and does not depend on historical samples.

[0025] In one alternative implementation, in the fluorescence channel, Raman and Rayleigh bands can be masked and interpolated before baseline flattening. Specifically, 5 to 15 sampling points are taken on each side of the center of each band, and the sampling points within this range are replaced with linear interpolation of adjacent effective bands. This avoids strong bands biasing the background estimation. In the scattering channel, angular consistency tuning and turbidity reference binding can be performed before amplitude normalization. Specifically, when the scattering channel originates from multiple observation angles, the intensities of each angle are first aligned to the maximum intensity, and then the average of the multi-angle sequences is calculated as the intensity sequence of the channel. Subsequently, the intensity obtained from a reference water sample with known turbidity within the same sampling period is used as the unit intensity, and the intensity of the current sample is divided by this unit intensity to make the scattering intensity comparable under different sampling periods. In baseline flattening, morphological opening operations can be used to estimate the background. Specifically, a sliding window with a structuring element length of 51 sampling points is used to perform a sequence operation of erosion followed by dilation on each channel to obtain the background, and then the background is subtracted point by point from the observation sequence. This method handles backgrounds with slow fluctuations and isolated spikes better. For time alignment, if setting a rectangular step for light intensity is inconvenient, the light source wavelength can be quickly switched twice before sampling begins, using the common transition edge generated in the three channels at the moment of switching as an alignment marker. This method does not change the chemical state of the water being measured and provides clear synchronization events.

[0026] Step 2: Call the preset endmember dictionary library and perform sparse coding endmember decomposition on the standardized multimodal spectral joint observation data: According to the main peak and auxiliary peak labels of each endmember entry in the endmember dictionary library, the absorption, fluorescence and scattering channels are simultaneously divided into several aligned segments; for each aligned segment, the overlap is obtained by multiplying the aligned segment and its corresponding endmember entry point by point; the overlap is divided by the product of the two scales to obtain the segment similarity score, and the minimum value among the segment similarity scores of the three channels is taken as the channel consistency score of this aligned segment; the endmember with the highest score is selected, and amplitude alignment is performed at the main peak by the peak height ratio; the peak position is fine-tuned by piecewise linear interpolation for segments marked as deformable; the endmember entries after amplitude alignment and peak position fine-tuning are subtracted point by point from the standardized multimodal spectral joint observation data, the negative values ​​after subtraction are set to zero and small peaks are merged, and this process is repeated until the residual meets the stopping criterion; the alignment factor at each subtraction is recorded to generate an endmember abundance table.

[0027] Specifically, assuming the standardized multimodal spectral joint observation data has been arranged according to a unified sampling point sequence, an endmember dictionary is invoked. The endmember dictionary contains absorption, fluorescence, and scattering endmember sub-libraries. Each endmember has a corresponding endmember entry in each of the three sub-libraries. Endmember entries include labels for the main peak, auxiliary peaks, background segments, and deformable segments, and are stored using a sampling point sequence consistent with the standardized multimodal spectral joint observation data. Based on the main peak and auxiliary peak labels of each endmember entry in the endmember dictionary, the absorption, fluorescence, and scattering channels are synchronously divided into several aligned segments on the same sampling point sequence. The boundary of the aligned segment is centered on the main peak sampling point, with the same number of sampling points taken on both sides. When an endmember entry contains multiple adjacent auxiliary peaks, the background segments adjacent to the main peak and auxiliary peaks are included in the same aligned segment. In practice, the typical width of a single-peak aligned segment is 31 to 61 sampling points, and the typical width of a multi-peak aligned segment is 81 to 121 sampling points. Synchronous partitioning allows the three channels to be calculated point by point under the same segment boundary, avoiding misalignment of peak-valley correspondence between channels.

[0028] For each aligned segment, in the absorption, fluorescence, and scattering channels, the same-length sequences of the aligned segment and its corresponding endmember entries are taken, multiplied point-by-point, and summed over all sampling points to obtain the overlap. The same-length sequences of the aligned segment and its endmember entries are then squared point-by-point and summed over all sampling points. The square roots of the sums are then taken to obtain two scales. The overlap is divided by the product of the two scales to obtain the segment similarity score of the aligned segment in that channel. The minimum segment similarity score across the three channels of the same aligned segment is taken as the channel consistency score of that aligned segment. The reason for using the minimum channel consistency score is that there are instantaneous turbidity disturbances or local occlusion of a channel in the mine water scene. Using the minimum value can preferentially exclude occasional high scores in a single channel, retain segments that match simultaneously across all three channels, reduce the number of attempts at erroneous endmembers for subsequent amplitude alignment and peak fine-tuning, thereby improving stripping efficiency and stability.

[0029] For all aligned segments of a given endmember, the corresponding channel consistency scores are collected and aggregated according to the median to obtain the endmember similarity score. The median is not sensitive to a small number of outlier segments and can stably reflect the overall agreement between the endmember entry and the standardized multimodal spectral joint observation data. The endmember similarity scores of all endmembers are sorted, and the endmember with the highest score is selected as the current target endmember. If the difference between the similarity scores of the first and second ranked endmembers is less than 0.02, parallel review can be enabled in an optional implementation, as explained at the end of the document.

[0030] After the current target endmember is selected, the main peak sampling point marked for that endmember entry is located in the absorption, fluorescence, and scattering channels, respectively. The peak height of the normalized multimodal spectral joint observation data at that sampling point is read, and the peak height of the endmember entry at that sampling point is read as well. The ratio of the two is calculated to obtain the amplitude alignment factor for that channel. The amplitude alignment factor is used to adjust the amplitude of the entire segment of the current target endmember entry. The advantage of selecting the main peak for amplitude alignment is that the main peak has the highest signal-to-noise ratio, is least affected by background uplift and random noise, and can establish a stable scale relationship without adding additional references. For the segments marked as deformable segments in the current target endmember entry, a three-segment mapping of piecewise linear interpolation is established around the sampling points of the main peak and auxiliary peaks. The control points of the three-segment mapping are derived from the peak and valley positions of the normalized multimodal spectral joint observation data in the same segment. This piecewise mapping is applied to the deformable segments of the endmember entry to ensure that the peak position of the endmember entry is consistent with the peak position of the normalized multimodal spectral joint observation data. Piecewise linear interpolation is more direct and can achieve small-range peak position adjustment without introducing waveform distortion, avoiding ringing and spurious peaks caused by higher-order interpolation.

[0031] The current target endmember entry, after amplitude alignment and peak position fine-tuning, is subtracted point-by-point from the standardized multimodal spectral joint observation data in each of the three channels to obtain a residual sequence for the three channels. Negative values ​​in the residual sequence are set to zero point-by-point to avoid reverse compensation in subsequent segment similarity score calculations. Adjacent small peaks are merged: when the distance between two adjacent local peaks is less than 7 sampling points and the peak height of the smaller peak is less than 20% of the peak height of the larger peak, the two are merged into one peak, and the sampling point of the larger peak is used as the new peak position. This merging rule removes fragmented small peaks caused by subtraction errors, making the residuals closer to the true unexplained components and preventing them from being misjudged as new main peaks in the next round of endmember selection.

[0032] After one stripping operation, the aforementioned alignment segmentation, segment similarity score calculation, channel consistency score calculation, endmember similarity score aggregation, and endmember selection are re-executed on the updated standardized multimodal spectral joint observation data. This process is repeated until any of the following stopping criteria are met: the residual sequences of the three channels no longer contain continuous main peaks, specifically, the number of peaks in any channel that are continuously more than twice the neighborhood average and have a duration of more than 5 sampling points is 0; the cumulative number of stripped endmembers reaches 20; and after two consecutive rounds of stripping, the highest value of the segment similarity score for the three channels is below 0.15. Multiple stopping criteria are used because the mine water scenario may contain both a small number of strong endmembers and a large number of weak endmembers. The three criteria correspond to three different scenarios: saturation interpretation, number of attempts, and low separability protection, respectively, ensuring stable termination of the cycle under different sample structures.

[0033] During each stripping, the amplitude alignment factors of each of the three channels are recorded by endmember name, forming an endmember abundance table. Each row of the endmember abundance table corresponds to an endmember name, including the amplitude alignment factors of the absorption channel, fluorescence channel, and scattering channel, as well as the median of the fragment similarity scores of that endmember in the three channels and the residual sequence energy at the time of stripping. Negative and outlier entries in the endmember abundance table are removed or set to zero. The endmember abundance table will serve as the direct basis for generating pollutant categories and relative grades in subsequent steps. The width of the aligned fragment can be automatically selected based on the distance between the main peak and the auxiliary peak marked by the endmember entry. If the distance between the main peak and the nearest auxiliary peak is less than 15 sampling points, the typical width of the multi-peak aligned fragment is used; otherwise, the typical width of the single-peak aligned fragment is used. The fragment similarity score is calculated using only basic operations such as pointwise multiplication, summation, squaring, square root, and division, without relying on any probability model or training data, making it easy to implement in a fixed-point manner on edge devices. The amplitude alignment factor is calculated using the peak height ratio of the main peak sampling point. When the main peak sampling point is obscured or saturated in a channel, the second-ranked peak in that channel can be used instead, provided that the peak height of that peak reaches more than 30% of the unit amplitude of that channel. The piecewise linear interpolation uses three control points: the nearest valley to the left of the main peak, the sampling point of the main peak, and the nearest valley to the right of the main peak. No stretching or compression is applied to segments outside the control points to avoid affecting the spectral shape of segments not marked as deformable. Smoothing of merging small peaks can be performed once within a moving window of 5 to 9 sampling points to suppress sharp noise remaining after subtraction. The residual sequence energy is evaluated as the square root of the sum of the squares of the non-negative residuals, facilitating comparison with the numerical ranges of different channels. This evaluation is only used within the current sample and does not involve any historical data.

[0034] In one alternative implementation, when the difference in similarity scores between the first and second ranked endmembers is less than 0.02, amplitude alignment and peak position fine-tuning can be performed on both endmembers simultaneously, with point-by-point subtraction performed on each. The residual sequence energy after the two subtractions is compared, and the one with the smaller residual is selected as the effective stripping. This method can reduce misselection when adjacent endmember entries are highly similar. If, after peak position fine-tuning is completed on a deformable segment, a continuous main peak still appears at the boundary of the segment, the aligned segment can be extended to both sides by 5 to 10 sampling points, and the segment similarity score and channel consistency score can be recalculated before endmember selection. This can cover information missed due to an initially narrow boundary. After every 5 stripping operations, consistency checks are performed on different endmember entries from the same pollutant in the endmember abundance table. If the median segment similarity score of multiple endmember entries for the same pollutant is higher than 0.5 in all three channels, the endmember entry whose main peak position is closer to the global main peak of the normalized multimodal spectral joint observation data is preferentially retained to reduce redundant interpretation.

[0035] Step 3: Generate pollutant categories and relative levels based on the endmember abundance table, and execute an incremental domain adaptation and calibration loop: Perform reference measurements at the beginning and end of sample processing, and compare them with the reference benchmark stored with the system to obtain the sampling point offset and amplitude offset; use the obtained sampling point offset and amplitude offset to perform wavelength and intensity alignment on the three channels, and simultaneously correct the entries in the endmember dictionary; when the difference between the reference measurement and the reference benchmark exceeds a preset threshold, trigger a complete sparse coding endmember decomposition recalculation, and expand the fragment fine-tuning range according to the piecewise linear interpolation rules of deformable fragments; after recalculation, perform a consistency check on the generated categories and levels. If they are inconsistent, the recalculation result shall prevail. If the difference exceeds the discrimination threshold, the adjacent level shall be taken. Endmembers that fail the consistency check are marked as pending confirmation and the final result is output.

[0036] Specifically, within the same sample processing cycle, two reference measurements are performed at the start and end times. These measurements include a clear water reference for the absorption channel, a fixed fluorescence standard reference for the fluorescence channel, and dark and clear water references for the scattering channel. This allows for the capture of short-term drifts within the same period without relying on historical data, such as slow decreases in light source output, slight obstruction caused by deposits on the probe window, or minor shifts in the spectral axis. By comparing the two reference measurements with the reference standards stored in the system, the numerical ranges of the sampling point offset and amplitude offset can be obtained, guiding wavelength and intensity alignment and determining whether to trigger a recalculation of sparse coding endmember decomposition. In the absorption channel, the known main and secondary peaks in the reference standards are located, commonly around 275 nm and 730 nm. In the reference measurement at the start time, the peak position indices of these two peaks are searched point by point, and the difference between these and the reference peak position indices is calculated to obtain the sampling point offset. Repeat the same steps at the end. If the absolute difference between the two sampling point offsets exceeds two sampling points, it is determined that there is a spectral axis change during this sample processing, and the sampling point offset at the end needs to be used for alignment. The amplitude offset is calculated using the peak height ratio: the ratio of the peak height of the main peak measured in the reference measurement to the peak height of the main peak of the reference standard is used to calculate the amplitude offset. For example, in the absorption channel, if the ratio is 0.92, it indicates that the overall intensity has decreased by about 8%. In the fluorescence channel, the strongest emission peak of the fixed fluorescence standard is used as the reference point for peak position and peak height. In the scattering channel, the average value of the dark reference and the average value of the flat section of the clear water reference are used as the zero point and unit amplitude reference. The sampling point offset and amplitude offset are obtained by subtracting and dividing the corresponding values ​​from the reference standard.

[0037] Based on the aforementioned sampling point offsets, integer shifts are performed on the sampling point sequences for the absorption, fluorescence, and scattering channels. When the sampling point offsets contain non-integer parts, linear interpolation is used to generate transition points between the original sampling points, ensuring that the reference peak position coincides with the reference baseline peak position. Subsequently, the intensity of the three channels is scaled according to the aforementioned amplitude offsets, so that the average value of the reference peak height or flat section is consistent with the reference baseline. The reason for this approach is that changes in device gain and optical path within the same period often manifest as an overall intensity scaling, while slight shifts in the spectral axis manifest as an overall peak position shift. By aligning these two dimensions, equivalent observation conditions under the reference baseline can be restored without changing the peak shape, avoiding misinterpreting device changes as material changes. To maintain consistency between the endmember dictionary and the current channel state, the same sampling point offsets and amplitude offsets are applied to the corresponding channels of all endmember entries in the endmember dictionary. Specifically, the same sampling point shifts and linear interpolations as the observed data are performed on the endmember entries of the three channels, and the same scaling as the observed data is performed on the intensity sequence. For segments marked as deformable fragments, the deformability markings are not changed; only the positions and amplitudes of the control points are re-labeled according to the new coordinates and scale. This synchronous correction avoids the situation where the fragment boundaries and peak positions become disconnected during subsequent recalculations, ensuring that the fragment similarity score and channel consistency score are still calculated within the same coordinate system.

[0038] The reference measurements at the start and end times are compared with the reference baseline. If the sampling point offset of any channel exceeds two sampling points, or the amplitude offset deviates by more than 10%, a complete sparse coded endmember decomposition recalculation is triggered. Before recalculation, the segment fine-tuning range is expanded according to the piecewise linear interpolation rules of deformable segments: 5 to 10 sampling points are added to the left and right sides of each deformable segment, and peak and valley control points are added within the newly added range. The control points are taken from the most significant peak and valley of the current observation data in the adjacent region. This expansion ensures that even with slight peak stretching or compression caused by device alignment, the endmember entries can still be completely consistent with the peaks of the observation data, avoiding residual accumulation caused by excessively narrow deformable segments. Subsequently, recalculation is performed based on the aligned three channels and the synchronously corrected endmember dictionary, according to the endmember similarity score calculation rules, endmember selection and stripping process described in step 2, until the stopping criteria are met.

[0039] Without triggering a recalculation, the pollutant category and relative level are obtained by directly looking up each item in the current endmember abundance table and the calibration table stored with the system in the endmember dictionary. When a recalculation is triggered, the same lookup process is performed using the recalculated endmember abundance table. The level range can be set to 5 levels from 0 to 4 or 4 levels from 0 to 3, depending on the level definition in the calibration table. Using a lookup-based approach instead of fitting allows for mapping on edge devices with a fixed computational cost, and adjustments can be made as needed by updating the calibration table without changing the algorithm flow.

[0040] The categories and grades generated before recalculation are compared item by item with those generated after recalculation. When categories are inconsistent, the recalculated result prevails. When the difference between relative grades exceeds the discrimination threshold, the adjacent grade between the two is taken as the final grade; the discrimination threshold can be set to a grade step size. For abnormal endmembers appearing during the comparison process, if the median segment similarity score of the endmember in the three channels is less than 0.15, or if the residual sequence energy of the endmember during stripping is higher than three times the standard deviation of the reference in the same channel, the endmember is marked as pending confirmation in the output and temporarily excluded from the endmember abundance table statistics. The final result includes pollutant category, relative grade, and alignment status prompts, such as whether sampling point alignment occurred, whether intensity alignment occurred, whether recalculation was performed, and whether the fine-tuning range of deformable segments was expanded.

[0041] Within the same sample processing cycle, the reference alignment at the start is first used to generate the initial category and grade. If the reference alignment at the end shows an increased offset and triggers a recalculation, the recalculated category and grade overwrite the initial result. This forms a closed loop of start alignment, processing, and end verification, ensuring that short-term changes in device status are corrected instantly within the same sample processing cycle, avoiding cross-cycle reliance on historical data. For cases where a recalculation is not triggered, the reference alignment at the end is used to confirm that the current alignment is still valid and is output as the alignment record for this sample along with the results.

[0042] In one optional implementation, the average offset of two reference peaks is used as the final sampling point offset when calculating the sampling point offset. When the offset directions of the two reference peaks are opposite and the absolute difference exceeds one sampling point, the reference peak closer to the 500 nm side is preferred because this band is less affected by strong absorption edges in mine water, resulting in a more stable peak position. Two sampling points and an 8% threshold are used for the absorption channel, one sampling point and a 10% threshold are used for the fluorescence channel, and three sampling points and a 12% threshold are used for the scattering channel to adapt to the noise and stability characteristics of different channels. To control the computational load, the maximum number of recalculations is set to two. If the difference in the relative levels of the major pollutants between the recalculated endmember abundance table and the unrecalculated endmember abundance table does not exceed one level, a second recalculation is not performed. If adding control points causes obvious spikes in the local peak shape of endmember entries after expanding the range of segment fine-tuning, the three-segment mapping can be changed to a five-segment mapping. The two newly added control points are located 10 sampling points to the left and right of the main peak, respectively, to more smoothly match the observed peak shape.

[0043] refer to Figure 2This invention employs a coaxial probe to simultaneously acquire spectral signals from three channels at the same measurement point. The horizontal axis represents the wavelength range of 200 nm to 700 nm, and the vertical axis represents the normalized intensity, ranging from 0 to 1.00. The solid curve in the figure represents the spectral signal of the absorption channel, which exhibits two main absorption peaks at approximately 300 nm and 550 nm, with normalized peak intensities of 0.72 and 0.58, respectively, and a baseline intensity of approximately 0.08. The absorption channel primarily reflects the absorption characteristics of pollutant molecules in the water sample to specific wavelengths of light; the position and intensity of the absorption peaks are directly related to the type and concentration of the pollutants. The dashed curve represents the spectral signal of the fluorescence channel, which displays two fluorescence emission peaks at approximately 380 nm and 680 nm, with normalized peak intensities of 0.55 and 0.48, respectively, and a baseline intensity of approximately 0.05. The fluorescence channel captures the fluorescence signal emitted by pollutant molecules after excitation; its peak shape is broader than the absorption peaks due to spectral broadening caused by energy relaxation during fluorescence emission. The dotted-line curve represents the spectral signal of the scattering channel, which shows scattering peaks at approximately 220 nm and 750 nm, with peak normalized intensities of 0.40 and 0.35, respectively, and a relatively high overall baseline intensity of approximately 0.15. The scattering channel primarily reflects the scattering effect of suspended particulate matter in the water sample on incident light, and its spectral characteristics are related to the particle size distribution and concentration. The spectral curves of the three channels are staggered in wavelength dimension, avoiding complete peak overlap, which demonstrates the complementary advantages of multimodal spectral joint observation. By simultaneously acquiring data from the three channels, this invention can comprehensively characterize the optical properties of pollutants in the water sample from three dimensions: absorption, fluorescence, and scattering, providing a rich multidimensional information source for subsequent endmember decomposition. All three curves contain a certain degree of noise fluctuation, which is unavoidable in actual measurements; the subsequent endmember decomposition algorithm is robust to noise.

[0044] like Figure 3As shown, this invention achieves accurate endmember identification and matching by calculating the similarity score between aligned segments and endmember entries. The horizontal axis represents the wavelength range of 200 nm to 700 nm, and the vertical axis represents the intensity value range of 0 to 1.00. The thick solid line curve in the figure represents the aligned segment extracted from the standardized multimodal spectral joint observation data. This segment is centered on a main peak, reaches its peak at a wavelength of approximately 500 nm, has a peak height of approximately 0.78, and a peak width of approximately 65 nm. This aligned segment is obtained by dividing the data according to the main peak and auxiliary peak labels in the endmember dictionary, representing the feature segment in the observation data that needs to be matched with the endmember library. The dashed line curve represents the spectral characteristics of a candidate endmember entry in the endmember dictionary in the corresponding band. The main peak of this endmember entry is located at a wavelength of approximately 480 nm, with a peak height of approximately 0.72 and a peak width of approximately 60 nm. There is an approximately 20 nm offset in peak position between the endmember entry and the observed segment, a difference of about 0.06 in peak height, and slight differences in peak width. These differences are caused by factors such as instrument drift and environmental changes during actual measurements, which is precisely the problem that this invention aims to solve through peak position fine-tuning. The gray-filled area in the figure represents the overlap between the observed segment and the endmember entry in the intensity dimension, i.e., the area formed by taking the smaller value of the two curves at each sampling point. This overlap is the core basis for calculating the similarity score. The specific calculation process is shown in the formula on the right: First, multiply the observed segment and the endmember entry point-by-point at all sampling points and sum them to obtain the overlap value; then, calculate the sum of squares point-by-point for each of the observed segment and the endmember entry, and take the square root as the normalization scale; finally, divide the overlap by the product of the two normalization scales to obtain the normalized similarity score. In this example, the calculated segment similarity score S = 0.876, which is a high similarity value, indicating that the endmember entry and the observed segment have a good matching relationship. In the actual endmember decomposition process, the system calculates similarity scores for all candidate endmembers in the endmember dictionary and selects the endmember with the highest score as the current target endmember for amplitude alignment and peak fine-tuning. It is worth noting that this invention calculates segment similarity scores on each of the three channels separately and takes the minimum score among the three channels as the channel consistency score. This minimum-value strategy ensures that the selected endmembers have good matching quality across all channels, avoiding the problem of strong correlation in one channel masking weak correlation in other channels, thereby improving the reliability and accuracy of endmember identification.

[0045] like Figure 4As shown, this invention employs a sparse coding endmember decomposition method, iteratively extracting each endmember component from the mixed spectrum. The horizontal axis represents the wavelength range of 200 nm to 700 nm, and the vertical axis represents the intensity value range of 0 to 1.00. The thick solid line curve in the figure represents the original mixed spectrum, which is the spectral signal of a certain channel extracted from the standardized multimodal spectral joint observation data. The mixed spectrum shows three distinct peaks at wavelengths of approximately 300 nm, 520 nm, and 720 nm, with peak heights of approximately 0.48, 0.85, and 0.40, respectively, and a baseline intensity of approximately 0.03. These three peaks correspond to three different pollutant components present in the water sample, and their spectral characteristics are superimposed to form the observed mixed spectrum. The complexity of the mixed spectrum is the main challenge in pollutant identification, requiring endmember decomposition to break it down into linear combinations of individual pure endmember components. The dashed curve represents the first endmember identified through similarity calculation, which corresponds to the highest peak at 520 nm in the mixed spectrum. According to the entry in the endmember dictionary, after amplitude alignment, the peak shape of this endmember closely matches the corresponding peak height in the mixed spectrum. The figure shows an amplitude alignment factor α = 0.85, indicating that the standard intensity of this endmember in the dictionary needs to be multiplied by a coefficient of 0.85 to match the amplitude of the observed data. This coefficient also represents the relative abundance or concentration of this endmember in the mixed spectrum. The dotted line curve represents the residual sequence obtained after removing the first endmember from the original mixed spectrum. The residual is calculated by subtracting the amplitude-aligned endmember intensity value from the intensity value of the mixed spectrum at each sampling point, setting negative values ​​to zero point by point. As can be seen from the figure, the peak at 520 nm in the residual curve has been completely removed, leaving mainly two peaks at 300 nm and 720 nm. The shape and position of these two peaks are basically consistent with the corresponding peaks in the original mixed spectrum, but the amplitude is slightly reduced. This is because the endmember entry may contain some background fragments or auxiliary peaks that slightly contribute to these positions. The figure also indicates a residual energy of E=18.5%. This parameter is obtained by summing the squares of the intensity values ​​of the residual sequence at all sampling points and normalizing them. It reflects the energy percentage of the remaining signal after removing an endmember. The residual energy is a crucial criterion for determining whether to continue endmember stripping. When the residual energy is below a preset threshold (e.g., 3%), it indicates that the mixed spectrum has been sufficiently decomposed, and the iterative process can be stopped. In this example, the residual energy is 18.5%, significantly higher than the stopping threshold. Therefore, it is necessary to continue performing the next round of endmember identification and stripping on the residual sequence until the peaks at 300nm and 720nm are also matched to their corresponding endmember entries and stripped. This ultimately yields the abundance table of each endmember and the low-energy residuals that meet the stopping criteria. This iterative strategy of successive stripping can effectively handle multi-component mixed samples, enabling accurate identification and quantitative analysis of complex pollutant systems.

[0046] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. A multimodal spectral in-situ identification and drift adaptive method for mine water contaminants, characterized in that, The method comprises: Step 1: synchronously acquiring signals of three channels including absorption, fluorescence and scattering, sequentially performing preprocessing of dark signal removal, baseline flattening, time alignment and amplitude normalization to form standardized multi-modal spectral joint observation data; Step 2: calling a preset endmember dictionary library, performing sparse coding endmember decomposition on the standardized multi-modal spectral joint observation data: synchronously dividing the three channels of absorption, fluorescence and scattering into a plurality of aligned segments according to the main peak and auxiliary peak markers of each endmember entry in the endmember dictionary library; for each aligned segment, multiplying and summing the aligned segment and its corresponding endmember entry point by point to obtain an overlap amount, taking the square root of the point-by-point square sum of the aligned segment and the endmember entry as a scale, dividing the overlap amount by the product of the two scales to obtain a segment similarity score, and taking the minimum value of the segment similarity scores of the three channels as the channel consistency score of the aligned segment; selecting the endmember with the highest score, performing amplitude alignment at the peak height ratio at the main peak, and performing peak position fine tuning according to the segmented linear interpolation for the segments marked as deformable; subtracting the endmember entry after amplitude alignment and peak position fine tuning from the standardized multi-modal spectral joint observation data point by point, setting the negative value after subtraction to zero and merging small peaks, and repeating the process until the residual meets the stopping criterion to record the alignment factor at each subtraction to generate an endmember abundance table; Step 3: generating a pollutant class and a relative grade based on the endmember abundance table, and performing an incremental domain adaptive and calibration loop: performing reference measurement at the beginning and end of sample processing, and comparing with the reference standard stored in the system to obtain sampling point offset and amplitude offset; aligning the wavelengths and intensities of the three channels using the obtained sampling point offset and amplitude offset, and simultaneously correcting the entries in the endmember dictionary library; when the difference between the reference measurement and the reference standard exceeds a preset threshold, trigger a complete sparse coding endmember decomposition recalculation, and expand the segment fine tuning range according to the segmented linear interpolation rule of the deformable segment; after recalculation, check the consistency of the generated class and grade, if not consistent, take the recalculation result as the standard, if the difference exceeds the discrimination threshold, take the adjacent grade, and mark the endmember that does not pass the consistency check as to be confirmed and output the final result.

2. The mine water contaminant multi-modal spectral in-situ identification and drift adaptation method of claim 1, wherein, Step 1 specifically comprises: using a coaxial probe to synchronously acquire signals of absorption, fluorescence and scattering channels at the same measurement point to form multi-modal spectral joint observation data; sequentially performing dark signal removal, baseline flattening, channel time alignment and amplitude normalization on the multi-modal spectral joint observation data; performing Raman and Rayleigh band masking on the fluorescence channel and interpolating adjacent effective bands, performing angular consistency setting and turbidity reference binding on the scattering channel to obtain standardized multi-modal spectral joint observation data.

3. The mine water contaminant multi-modal spectral in-situ identification and drift adaptation method of claim 2, wherein, In step 2, the endmember dictionary library comprises an absorption endmember sub-library, a fluorescence endmember sub-library and a scattering endmember sub-library; each endmember in the dictionary library has a corresponding entry in the three sub-libraries, and the three entries are stored in the same sequence of sampling points; each entry comprises four types of markers, i.e., a sampling point where a main peak is located, an auxiliary peak group, a background segment and a deformable segment; the normalized multi-modal spectral joint observation data is unified to the sequence of sampling points used by the endmember dictionary library on the three channels; the main peak and the auxiliary peak markers in each endmember entry of the endmember dictionary library are used as anchor points to simultaneously divide the three channels into a plurality of aligned segments, and each aligned segment corresponds to a main peak and a background segment adjacent to the main peak.

4. The multi-modal in-situ identification and drift adaptive method of mine water contaminants of claim 3, wherein, In step 2, the calculation of the consistency score is as follows: first, the overlap is obtained by multiplying and summing the points of any aligned segment and the corresponding endmember entry; then, the sum of the squares of the points of the aligned segment and the endmember entry is taken as the square root of the square of the sum of the squares, as a normalization scale; the overlap is divided by the product of the two normalization scales to obtain the segment similarity score of the aligned segment on each channel; the minimum value of the segment similarity scores of the three channels is taken as the segment channel consistency score of the aligned segment; the median of the segment channel consistency scores of all segments corresponding to any endmember is taken as the endmember similarity score of the endmember; the endmember with the highest score is selected from all endmembers as the current target endmember; the main peak sampling point of the current target endmember is located on the three channels; the peak height of the normalized multi-modal spectral joint observation data at the target main peak sampling point is calculated by ratio calculation with the peak height of the current target endmember entry; and the calculated ratio is taken as the amplitude alignment factor; The amplitude alignment factor is used to simultaneously adjust the amplitude of the current target endmember entry on the three channels.

5. The mine water contaminant multi-modal spectral in-situ identification and drift adaptation method of claim 4, wherein, In step 2, for the section of the current target endmember entry marked as a deformable segment, a piecewise mapping of three connected broken lines is established around the sampling points of the main peak and the auxiliary peak, and the control points of the piecewise mapping are taken from the peak positions and valley positions of the normalized multi-modal spectral joint observation data in the same section; based on the established piecewise mapping, a point-by-point linear interpolation deformation is performed on the corresponding section of the current target endmember entry, so that the peak positions of the endmember entry are consistent with the peak positions of the normalized multi-modal spectral joint observation data.

6. The multi-modal in-situ identification and drift adaptive method of mine water contaminants of claim 5, wherein, In step 2, the loop process is as follows: the current target endmember entry after amplitude alignment and deformation is subtracted from the normalized multi-modal spectral joint observation data on the three channels point by point, and the result is recorded as a residual sequence; the negative values in the residual sequence are set to zero point by point, and adjacent small peaks are merged and smoothed to obtain updated normalized multi-modal spectral joint observation data; if the total energy of the residual sequence on the three channels is lower than a preset threshold, or the number of stripped endmembers reaches a preset upper limit, the loop is stopped; otherwise, the calculation rule of the endmember similarity score is returned to, and the endmember selection, fine tuning and stripping are continued on the updated normalized multi-modal spectral joint observation data. The amplitude alignment factor at each peeling is recorded by end member name to form an end member abundance table, and the corresponding fragment similarity score and residual sequence energy are recorded as the matching credentials of the corresponding end member; negative and abnormal items in the end member abundance table are deleted or set to zero; according to the end member abundance table and the scaling table stored in the system in the end member dictionary library, the pollutant category and relative grade are obtained within the scaling interval of each end member; when the same pollutant is hit in multiple end member entries, the first listed entry is selected as the final source of the pollutant in the order of priority recorded in the scaling table, and the unmatched entry is marked as not detected.

7. The mine water pollutant multi-modal spectral in-situ identification and drift adaptation method of claim 6, wherein, In step 3, at the beginning and end of each sample processing, the built-in reference component is triggered to perform a reference measurement to obtain the reference absorption signal, the reference fluorescence signal and the reference scattering signal; the main peak position and amplitude of the reference measurement are compared with the reference benchmark stored in the system point by point to obtain the sampling point offset and amplitude offset of the three channels; Using the sampling point offset, the sampling point translation and endpoint extrapolation interpolation are performed on the three channels to complete the alignment in the wavelength direction; using the amplitude offset, the amplitude stretching or compression is performed on the three channels to complete the alignment in the intensity direction; the correction only acts on the current sample processing process.

8. The mine water contaminant multi-modal spectral in-situ identification and drift adaptation method of claim 7, wherein, In step 3, the sampling point offset and amplitude offset obtained by channel correction are synchronously applied to the corresponding channels of all end member entries in the end member dictionary library; for the section marked as a deformable fragment, the peak position and valley position of the reference measurement are used as new control points, and the end member entry is deformed by piecewise linear interpolation to make the end member entry consistent with the corrected three channels; when the difference between the reference measurement and the reference benchmark exceeds the preset threshold, an end member entry synchronization correction and a complete sparse coded end member decomposition recalculation are triggered immediately; the end member similarity score calculation rule, end member selection and peeling process in step 2 are followed during the recalculation; after the recalculation, the end member abundance table and the residual sequence energy before and after the recalculation are compared, and if the difference still exceeds the preset threshold, the control point range of the deformation section is expanded according to the deformable fragment fine tuning rule and recalculated again until the difference falls within the preset threshold or reaches the preset upper limit of times.

9. The mine water pollutant multi-modal spectral in-situ identification and drift adaptation method of claim 8, wherein, In step 3, the consistency of the pollutant category and relative grade before and after recalculation is checked: when the categories are inconsistent, the result after recalculation is used as the final result; when the relative grade difference exceeds the preset grade step, the adjacent grade between the results before and after recalculation is taken as the final grade; The end member that does not pass the consistency check is deleted from the end member abundance table and marked as to be confirmed in the output; at the end of the current sample processing, the sampling point offset, the amplitude offset and the deformation control point of the deformable fragment obtained by the current reference measurement are written into the device operating parameters for direct use in the channel correction and end member entry synchronization correction of the next sample, without retaining any original observation data.

Citation Information

Patent Citations

  • Microscopic multimodal fusion spectrum detection system

    CN107044959A

  • Discrete three-dimensional fluorescence / visible light absorption spectrum detection device for judging water quality pollution

    CN114166747A

  • Water pollutant detection method, system and equipment based on spectral analysis

    CN120558868A

  • Marine plankton hyperspectral imaging detection system

    CN120635685A

  • Water quality pollution multi-dimensional analysis system

    CN120778663A