Quantitative method for oil pollutants in water based on ultraviolet fluorescence characteristic spectrum
By acquiring multi-angle polarization spectra, peak correction, and multi-dimensional feature modulation, combined with a spectral correlation model and concentration inversion network, the problem of difficult separation of overlapping spectra in multi-component oil mixtures by ultraviolet fluorescence spectroscopy was solved, achieving high-precision quantitative analysis of oil components.
Patent Information
- Application Number
- CN202611118084.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-08-25
AI Technical Summary
Existing ultraviolet fluorescence spectroscopy methods are difficult to accurately separate and quantify multi-component oil mixtures, especially in complex mixed systems where overlapping spectra are difficult to separate precisely, leading to distorted quantitative results.
Multi-angle polarization-induced spectra of the water sample to be tested at multiple preset wavelengths are collected. Through peak position correction and multi-dimensional feature modulation, composite structure feature spectrum is generated. Group matching analysis is performed using a pre-constructed spectrum association model, and quantitative inversion is performed by combining it with a concentration inversion network.
It enables precise localization and high-precision quantification of different oil components in complex mixed systems, effectively overcoming spectral morphology variations and matrix interference, and improving the robustness of concentration prediction.
Smart Images

Figure CN122631610A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water quality testing technology, specifically to a quantitative method for oil pollutants in water based on ultraviolet fluorescence characteristic spectra. Background Technology
[0002] Ultraviolet fluorescence detection of oil pollutants in water typically relies on fluorescence emission spectra obtained at a fixed excitation wavelength and a single acquisition angle, using the linear relationship between fluorescence characteristic peak intensity and concentration for quantitative analysis. When the water sample contains only a single or simple mixture of oil components, this method can provide relatively accurate concentration information. However, actual polluted water bodies often contain multiple petroleum hydrocarbons and aromatic compounds simultaneously, resulting in severe overlap of fluorescence emission bands among the components. Furthermore, differences in fluorescence quantum efficiency, molecular quenching effects, and polarization characteristics among different components make it difficult for conventional single-angle, fixed-wavelength fluorescence spectra to reflect complete component information. The envelope peaks formed by the superposition of multi-component signals are difficult to separate, and direct quantification based on mixed peak height or peak area will produce significant allocation errors, especially for oil components with similar chemical structures, where the misjudgment rate and concentration deviation are even more pronounced.
[0003] Existing multi-component spectral analysis methods mostly employ linear unmixing or principal component regression based on known pure component spectra. These methods are highly dependent on the completeness of the prior spectral library and cannot adequately address variations in fluorescence peak positions caused by matrix effects, such as peak broadening and asymmetric tailing. If the mixed system contains components whose spectral characteristics are not accurately characterized, the unmixing results are highly susceptible to failure. For some applications, attempts have been made to introduce synchronous fluorescence or three-dimensional fluorescence matrix techniques to increase dimensionality. However, this increase in data dimensionality has not fundamentally solved the problems of adaptive modeling of overlapping peaks and refined reconstruction of component spatial distribution. Quantitative results still fall short of the growing demand for simultaneous detection of low-concentration multi-component systems.
[0004] The problem this application aims to solve is how to accurately separate and characterize the feature information of different oil components from highly overlapping ultraviolet fluorescence spectra in complex mixed systems, and achieve simultaneous high-precision inversion of multi-component concentrations. Another problem to be solved is how to effectively overcome the influence of spectral morphology variations and matrix interference on quantitative accuracy, achieve multi-peak adaptive nonlinear fitting correction, and improve the robustness of concentration prediction. Summary of the Invention
[0005] The present invention aims to provide a quantitative method for oil pollutants in water based on ultraviolet fluorescence characteristic spectra, in order to solve the problems of existing fluorescence spectroscopy methods which make it difficult to accurately separate overlapping spectra when dealing with multi-component oil mixtures, and the distortion of quantitative results caused by spectral morphology variations.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra. This method includes the following steps: acquiring multi-angle polarization-induced spectra of the water sample at multiple preset wavelengths as an initial set of fluorescence spectra; utilizing multi-dimensional modulation of excitation wavelength and polarization angle to obtain anisotropic information of oil fluorophores, enriching the feature dimensions; correcting the peak positions of each spectrum in the initial set of fluorescence spectra to obtain a corrected set of characteristic spectra, thereby eliminating peak shifts caused by instrument drift and sample matrix, ensuring the accuracy of the spectral data. Figure 1 The corrected feature spectrum set is subjected to multi-dimensional feature modulation to generate a composite structure feature spectrum. By fusing polarization dimension, wavelength dimension, and time-frequency features of independent components, the spectral differences between components are amplified and background interference is suppressed. Based on a pre-built spectral association model, the composite structure feature spectrum is subjected to group matching analysis to generate spatial distribution data of oil components, realizing the precise positioning of the contribution of different oil groups in the spectral space. According to the spatial distribution data of oil components, the composite structure feature spectrum is subjected to feature morphology dynamic fitting to obtain feature peak correction spectrum, so as to adaptively separate overlapping fluorescence peaks and restore the true peak shape of each component. Based on the feature peak correction spectrum, a pre-trained concentration inversion network is called to perform quantitative inversion to obtain the concentration value of each oil component.
[0007] As a preferred embodiment of the present invention, the process of acquiring multi-angle polarization-induced spectra of the water sample under multiple preset wavelengths specifically involves: selecting a first excitation wavelength from the excitation wavelength set to irradiate the water sample, while simultaneously controlling the rotation angle of the polarization filter assembly, and synchronously recording fluorescence intensity at multiple angles to form a first angular polarization spectral line; traversing the remaining excitation wavelengths in the excitation wavelength set, sequentially executing the rotation control of the polarization filter assembly and fluorescence intensity recording to obtain multiple angular polarization spectral lines corresponding to each excitation wavelength; and splicing the first angular polarization spectral line and the multiple angular polarization spectral lines corresponding to the remaining excitation wavelengths according to the excitation wavelength order to generate an initial fluorescence spectrum set. This method constructs a spectral data volume covering three-dimensional information of excitation-emission-polarization in one go, significantly improving the resolution capability for complex oil mixtures.
[0008] As another preferred technical solution of the present invention, the process of correcting the peak positions of the initial fluorescence spectrum set specifically involves: extracting the set of peak position coordinates for each spectrum in the initial fluorescence spectrum set; obtaining the standard fluorescence characteristic peak positions of the internal standard substance in the water sample to be tested, and calculating the offset matrix between the set of peak position coordinates and the standard fluorescence characteristic peak positions; performing a translation correction on the wavelength axis of each spectrum based on the offset matrix, and using a prediction algorithm to perform peak shape matching iteration on the translated spectrum until all peak position deviations are less than a preset threshold, thereby obtaining the corrected characteristic spectrum set. This correction process can eliminate peak drift caused by environmental fluctuations and improve the reproducibility of the spectrum.
[0009] Furthermore, the multi-dimensional feature modulation of the corrected feature spectrum set includes: performing independent component analysis on each corrected feature spectrum in the corrected feature spectrum set to extract multiple independent spectral line components; performing time-frequency transformation on each independent spectral line component to generate a corresponding time-frequency energy distribution matrix; superimposing all time-frequency energy distribution matrices corresponding to the same corrected feature spectrum according to preset spectral line intensity weights to generate a single-spectrum modulation feature map; calculating the interaction mapping matrix of the single-spectrum modulation feature maps corresponding to all corrected feature spectra, and using the interaction mapping matrix to perform weighted fusion of each single-spectrum modulation feature map to generate a composite structure feature spectrum. This modulation method mines potential structural information from multi-spectral correlation, effectively enhancing the fluorescence characteristics of trace oil components.
[0010] As a preferred technical solution of the present invention, the pre-constructed spectral association model is constructed in the following manner: Pure fluorescence spectral data of multiple known oil components are acquired, and molecular structure feature labels are labeled for each pure fluorescence spectral data; multi-dimensional feature modulation is performed on the pure fluorescence spectral data, including joint modulation of the polarization angle dimension and wavelength dimension, to obtain a pure composite structure feature spectrum; the pure composite structure feature spectrum is clustered according to the molecular structure feature labels to generate multiple group category clusters; group feature basis functions are extracted from the pure composite structure feature spectrum within each group category cluster, and the group feature basis functions are associated and stored with the corresponding molecular structure feature labels to construct the spectral association model. This model directly associates spectral features with molecular groups, providing a basis for component spatial analysis.
[0011] Based on the aforementioned spectral association model, the preferred process for group matching analysis includes: slicing the composite structure feature spectrum according to wavelength dimension to obtain multiple wavelength slice sub-images; performing template matching between each wavelength slice sub-image and the feature basis function corresponding to each group category cluster in the spectral association model, and calculating the sub-image matching similarity; generating a similarity space matrix based on the sub-image matching similarity of all wavelength slice sub-images, and dynamically reconstructing the similarity space matrix using a fractal mapping algorithm to obtain the spatial distribution data of oil components. Specifically, during dynamic reconstruction using the fractal mapping algorithm, the fractal dimension of the similarity space matrix is calculated to determine the self-similarity scale parameter of the reconstructed image; an iterative shrinking transformation is performed on the similarity space matrix based on the self-similarity scale parameter to generate a reconstructed spectrum; the residual space between the reconstructed spectrum and the composite structure feature spectrum is calculated, and the difference feature vectors in the residual space are extracted; the difference feature vectors are mapped to the corresponding positions in the reconstructed spectrum using a pre-constructed trend association matrix to generate the spatial distribution data of oil components. This analysis method can automatically adapt to the nonlinear superposition effect of the spectrum, improving the accuracy of component localization.
[0012] As a preferred technical solution of the present invention, the process of performing dynamic fitting of characteristic morphology based on the spatial distribution data of oil components specifically includes: labeling the characteristic peak regions corresponding to each oil component in the composite structure characteristic spectrum based on the spatial distribution data of oil components; extracting the contour lines of each characteristic peak region to form a set of local peak curves; performing point-by-point nonlinear fitting on each local peak curve in the set of local peak curves by calling a predefined mathematical function to generate a single characteristic peak fitting curve; superimposing all the single characteristic peak fitting curves and performing residual analysis with the composite structure characteristic spectrum, iteratively adjusting the fitting parameters, and outputting the corrected spectrum of the characteristic peak when the sum of squared residuals converges to a stable value. This fitting strategy can effectively remove peak distortion caused by fluorescence quenching and matrix effects, and obtain high-fidelity component characteristic peaks.
[0013] The quantitative inversion process of the concentration inversion network further includes: analyzing the morphological features of the characteristic peak correction spectrum to extract the kurtosis, skewness, and half-width at half-maximum (WHM) parameters of the characteristic peak; concatenating the kurtosis, skewness, and WHM parameters with the peak height data of the characteristic peak correction spectrum to generate a concentration inversion input vector; and inputting the concentration inversion input vector into a pre-trained concentration inversion network, which is composed of multiple fully connected layers and nonlinear activation layers stacked alternately, and calculates and outputs the concentration values of each oil component through layer-by-layer mapping. The network integrates peak morphology and intensity information, maintaining excellent quantitative accuracy and anti-interference ability even in complex backgrounds.
[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: Multi-angle polarization fluorescence acquisition was performed on the water sample under multiple preset excitation wavelengths to obtain an initial set of fluorescence spectra covering both excitation wavelength and polarization angle. After peak position correction, multi-dimensional feature modulation was further performed. This modulation process involved independent component analysis of the corrected spectra, extracting independent spectral line components reflecting the intrinsic characteristics of different oil components. Then, time-frequency transformation was used to map each independent spectral line component to the time-frequency domain, generating a time-frequency energy distribution matrix. This matrix was then superimposed according to spectral line intensity weights to form a single-spectrum modulation feature map. Subsequently, through the calculation of the interaction mapping matrix between different spectra, all single-spectrum modulation feature maps were weighted and fused to obtain a composite structure feature spectrum. The composite structure feature spectrum integrates the joint modulation information of polarization angle and excitation wavelength dimension. Compared with the traditional single-angle or single-wavelength acquisition method, the fluorescence feature dimension is greatly extended. The characteristic peaks of different oil components that originally overlapped severely in a single dimension are presented in the composite structure spectrum with differentiated time-frequency textures and energy distribution patterns. This effectively alleviates the problem of spectral feature stacking and confusion in the mixed system and provides a more distinctive spectral characterization form for distinguishing oil components with similar chemical structures.
[0015] A pre-constructed spectral correlation model was used to perform group matching analysis on the feature spectra of composite structures. During the model establishment process, the same multi-dimensional feature modulation was applied to the pure fluorescence spectra of known oil components. The acquired pure composite structure feature spectra were clustered into group category clusters based on molecular structure feature labels, and group feature basis functions were extracted and associated within each cluster. During analysis, wavelength slices were performed on the composite structure feature spectra, and the template matching similarity between each slice and the feature basis functions of each group was calculated, forming a similarity space matrix. The self-similarity scaling parameter of this matrix was obtained using a fractal mapping algorithm, and an iterative shrinking transformation was performed to generate a reconstructed spectrum. The difference feature vector of the residual space between the reconstructed spectrum and the original composite structure feature spectrum was then extracted. The spatial distribution data of the oil components was obtained by mapping with a trend correlation matrix. This data precisely marks the spectral contribution positions and relative distribution patterns of each oil component. When performing dynamic fitting of characteristic morphology based on this distribution data, contour lines are extracted for the characteristic peak regions corresponding to each oil component. A predefined mathematical function is used for point-by-point nonlinear fitting, and the fitting parameters are iteratively adjusted through residual analysis with the composite structure characteristic spectrum until the sum of squared residuals converges and stabilizes, outputting the corrected characteristic peak spectrum. This process utilizes template matching and fractal reconstruction at the group dimension to accurately locate the hidden component characteristic regions under overlapping bands. Simultaneously, dynamic fitting can adaptively correct peak asymmetry, broadening, and shift, eliminating the limitations of fixed function fitting methods for fitting biases in complex deformed peaks. This ensures that the obtained corrected characteristic peak spectrum retains the true peak shape information of each oil component. Based on the highly accurate corrected characteristic peak spectrum, morphological parameters such as kurtosis, skewness, and half-width are extracted and concatenated with peak height to form an input vector. Quantitative mapping is then completed through a concentration inversion network composed of alternating fully connected layers and nonlinear activation layers. The network can learn the complex nonlinear relationship between component concentration and multidimensional morphological features from the finely peak-corrected input, significantly improving the accuracy and robustness of the output oil component concentration values. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0017] Figure 1 This is a flowchart of a method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra; Figure 2 This is a schematic diagram of the process for obtaining the initial set of fluorescence spectra; Figure 3 It is the emission wavelength distribution curve of the characteristic spectrum of the complex structure of the functional group category; Figure 4This is a schematic diagram of the characteristic peak correction spectrum and the characteristic peak regions of oil components. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1 This invention provides a method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra. The method includes: collecting multi-angle polarization-induced spectra of the water sample at multiple preset wavelengths as an initial set of fluorescence spectra; correcting the peak positions of each spectrum in the initial set of fluorescence spectra to obtain a corrected set of characteristic spectra; performing multi-dimensional feature modulation on the corrected set of characteristic spectra to generate a composite structure characteristic spectrum; performing group matching analysis on the composite structure characteristic spectrum based on a pre-constructed spectral association model to generate spatial distribution data of oil components; performing dynamic fitting of the characteristic morphology of the composite structure characteristic spectrum based on the spatial distribution data of oil components to obtain a characteristic peak corrected spectrum; and calling a pre-trained concentration inversion network based on the characteristic peak corrected spectrum to perform quantitative inversion and obtain the concentration values of each oil component.
[0020] Example 1: In specific implementation, please refer to Figure 2 The excitation wavelength set consists of nineteen wavelengths: 220 nm, 230 nm, 240 nm, 250 nm, 260 nm, 270 nm, 280 nm, 290 nm, 300 nm, 310 nm, 320 nm, 330 nm, 340 nm, 350 nm, 360 nm, 370 nm, 380 nm, 390 nm, and 400 nm. One of these excitation wavelengths is selected as the first excitation wavelength, which is the wavelength with the smallest numerical value in the set.
[0021] The excitation source is controlled to output monochromatic ultraviolet light with a wavelength of the first excitation wavelength. This monochromatic ultraviolet light is irradiated into a sample cell containing the water sample to be tested. Oil pollutant molecules in the water sample are excited and produce fluorescence. A polarization filter assembly is placed in the fluorescence emission path. The polarization filter assembly includes a polarizer rotatable around the optical axis and an angle controller. The angle controller drives the polarizer to rotate within a preset set of polarization angles, which consists of thirteen polarization angles: 0°, 15°, 30°, 45°, 60°, 75°, 90°, 105°, 120°, 135°, 150°, 165°, and 180°. Whenever the polarization filter assembly rotates to a polarization angle within the preset set, a fluorescence detector located at the end of the emission path simultaneously records a fluorescence spectrum. The fluorescence spectrum represents a curve showing the change in fluorescence intensity with the emission wavelength. After recording the fluorescence spectra corresponding to all polarization angles at the first excitation wavelength, the first angle polarization spectrum is obtained, which contains fluorescence spectra at thirteen polarization angles.
[0022] After acquiring the first angular polarization spectral line, an excitation wavelength other than the first excitation wavelength is selected from the excitation wavelength set as the current remaining excitation wavelength. The order of selection of the current remaining excitation wavelength is based on the ascending order of wavelengths in the excitation wavelength set. The excitation source is controlled to output monochromatic ultraviolet light at the current remaining excitation wavelength, and the process of the angle controller driving the polarizer to rotate within the preset polarization angle set and the fluorescence detector synchronously recording the fluorescence spectrum is repeated to obtain multiple angular polarization spectral lines corresponding to the current remaining excitation wavelength. By traversing all remaining excitation wavelengths in the excitation wavelength set, a total of eighteen sets of multiple angular polarization spectral lines corresponding to the remaining excitation wavelengths are obtained. Each set of multiple angular polarization spectral lines corresponding to the remaining excitation wavelengths contains fluorescence spectra at thirteen polarization angles.
[0023] The first angular polarization spectral line and the multiple angular polarization spectral lines corresponding to the eighteen remaining excitation wavelengths are stitched together in ascending order of excitation wavelength value. This stitching operation treats the fluorescence spectrum under the same excitation wavelength and polarization angle as an independent spectral element. All independent spectral elements are combined to form the initial fluorescence spectrum set. Each spectrum in the initial fluorescence spectrum set is described by four dimensions: excitation wavelength, polarization angle, emission wavelength, and fluorescence intensity. The emission wavelength is used as the horizontal axis of the spectrum, and the fluorescence intensity is used as the vertical axis.
[0024] Example 2: In practice, before collecting the water sample, an internal standard is added. Quinine sulfate is chosen as the internal standard, as it exhibits a single, sharp fluorescence emission peak around 450 nm under UV excitation. The position of the standard fluorescence characteristic peak of the internal standard is obtained through pre-determination. The determination process involves preparing a solution with a concentration of 1.0 × 10⁻⁶.-6 A standard aqueous solution of quinine sulfate (mol / L) was used. Under the same excitation wavelength and polarization angle as the initial fluorescence spectrum, the fluorescence spectrum of the quinine sulfate standard aqueous solution was measured. The emission wavelength corresponding to the maximum fluorescence intensity in the fluorescence spectrum was extracted, and the extracted emission wavelength was used as the position of the standard fluorescence characteristic peak of the internal standard and stored in the system database. The position of the standard fluorescence characteristic peak of the internal standard corresponds one-to-one with the excitation wavelength and polarization angle.
[0025] Extract the peak position coordinates of each spectrum in the initial fluorescence spectrum set. For a spectrum in the initial fluorescence spectrum set, use the first derivative peak finding method to process the fluorescence intensity versus emission wavelength curve. Calculate the first derivative of the curve and search for zero-crossing points where the first derivative changes from positive to negative. Record the emission wavelength value corresponding to the zero-crossing point as a peak position. Traverse the entire emission wavelength range to obtain all peak positions of the spectrum, and combine all peak positions to form the peak position coordinates of the spectrum. Perform the above first derivative peak finding operation on each spectrum in the initial fluorescence spectrum set to obtain the peak position coordinates of each spectrum.
[0026] The positions of the standard fluorescence characteristic peaks of the internal standard in the water sample to be tested are obtained, and the offset matrix between the set of peak position coordinates and the standard fluorescence characteristic peak positions is calculated. For a spectrum currently being calibrated, the corresponding standard fluorescence characteristic peak positions of the internal standard are retrieved from the system database based on the excitation wavelength and polarization angle of the spectrum. The peak positions belonging to the internal standard in the set of peak position coordinates are detected by searching for the peak position with the smallest wavelength difference from the standard fluorescence characteristic peak position. The detected peak positions are denoted as... The position of the standard fluorescence characteristic peak is denoted as The offset matrix is a one-dimensional column vector containing a single element. , Calculated by the following formula:
[0027] in, This indicates the offset of the wave crest position, in nanometers. This indicates the peak position corresponding to the internal standard substance in the set of peak position coordinates. Indicates the position of the standard fluorescence characteristic peak of the internal standard substance.
[0028] The wavelength axis of each spectrum is shifted and corrected based on the offset matrix. The offset is then subtracted from all emission wavelength coordinates on the wavelength axis of the current spectrum. The translated and corrected spectrum is obtained, and the translated and corrected spectrum replaces the original spectrum in subsequent processing.
[0029] The peak shape matching iteration of the shifted spectrum is performed using a prediction algorithm. The prediction algorithm employs a gradient descent-based wavelength fine-tuning network architecture, which includes an input module, a wavelength shift parameter module, a loss function module, and a parameter update module. The input module receives the shift-corrected spectrum data and the standard fluorescence characteristic peak positions of the internal standard. The wavelength shift parameter module maintains a trainable wavelength fine-tuning parameter. Wavelength fine-tuning parameters The initial value is set to 0 nanometers. The loss function module constructs the loss function. loss function Defined as the square of the difference between the position of the internal standard peak and the position of the standard fluorescence characteristic peak in the spectrum after translation correction. The parameter update module is based on the loss function. Wavelength fine-tuning parameters The gradient, according to the learning rate Update wavelength fine-tuning parameters Learning rate Set the learning rate to 0.0005 nanometers per iteration. The setting is based on avoiding iterative oscillations while ensuring convergence speed. Through multiple experiments, 0.0005 was selected as a fixed value in the range of 0.0001 to 0.001.
[0030] A complete peak matching iteration includes the following steps: adding the current wavelength fine-tuning parameter to all wavelength axis coordinates of the translated and corrected spectrum. A temporary spectrum is generated; first-order derivative peak finding is performed on the temporary spectrum to obtain the peak position of the internal standard in the temporary spectrum; the peak positions of the internal standard in the temporary spectrum and the standard fluorescence characteristic peak positions are substituted into the loss function. Calculate the loss value; calculate the loss value relative to the wavelength fine-tuning parameters. The gradient, the parameter update module will fine-tune the wavelength parameters. Updated to Determine if the loss value is less than a preset convergence threshold. The preset convergence threshold is set to... The nanometer squared value, with its preset convergence threshold set to meet the wavelength accuracy requirements of quantitative analysis, terminates the iteration when the loss value falls below the preset convergence threshold, and the current wavelength fine-tuning parameters are adjusted. This serves as the final offset correction. The final offset correction is then superimposed onto the wavelength axis of the translated spectrum to obtain the corrected characteristic spectrum.
[0031] The peak matching iteration process is performed on all translated spectra one by one until the peak position deviation of all spectra meets the preset convergence condition. All corrected feature spectra are collected to form a set of corrected feature spectra.
[0032] Example 3: In practice, the process of performing multi-dimensional feature modulation on the corrected feature spectrum set to generate composite structure feature spectra involves joint processing of the polarization angle dimension and the emission wavelength dimension. Each corrected feature spectrum in the corrected feature spectrum set contains an excitation wavelength and a polarization angle identifier. The horizontal axis of the corrected feature spectrum is the emission wavelength, and the vertical axis is the fluorescence intensity.
[0033] Independent component analysis (ICA) is performed on each corrected feature spectrum in the set of corrected feature spectra to extract multiple independent spectral components. ICA is implemented using a fast fixed-point iterative algorithm based on maximizing negative entropy. The fluorescence intensity variation of a corrected feature spectrum with emission wavelength is considered as a one-dimensional observation signal vector, with the vector dimension equal to the number of emission wavelength sampling points. The number of independent spectral components to be extracted is set to 5. Centering is performed, setting the mean of the observation signal vector to zero; whitening is performed, calculating the covariance matrix of the observation signal vector, performing eigenvalue decomposition on the covariance matrix, and constructing a whitening transformation matrix using eigenvalues and eigenvectors to transform the centered observation signal vector into a whitened vector. A 5-row unmixing matrix is initialized, with each row vector in the unmixing matrix corresponding to the projection direction of an independent component. The row vectors of the unmixing matrix are estimated row by row iteratively, using the negative entropy approximation as a measure of non-Gaussianity during the iteration process. The negative entropy approximation is calculated using the following formula:
[0034] in, Represents random variables The approximate value of negative entropy; This represents the projection value obtained by the inner product of a row vector in the unmixing matrix and the whitening vector; Operator for mathematical expectation; Represents the cube of the projected value; This represents the fourth power of the projection value. The iterative update rule is as follows: the current row vector is iteratively updated to the expectation of the nonlinear function of the whitening vector and the projection value minus the projection of the row vector onto the previous row vector, and then normalized until convergence. After traversing the 5 independent components, the unmixing matrix is obtained. Multiplying the unmixing matrix with the whitening vector yields 5 independent spectral components, each of which is represented as an emission wavelength-fluorescence intensity curve.
[0035] A time-frequency transform is performed on each independent spectral line component to generate the corresponding time-frequency energy distribution matrix. The time-frequency transform uses continuous wavelet transform, with Morse wavelet as the wavelet basis function, a symmetry parameter of 3, and a time-bandwidth product of 60. Using the independent spectral line components as input signals and the emitted wavelength as the time variable, a continuous wavelet transform is performed, with the scale parameter ranging from 1 to 64 at exponential intervals of 64. The continuous wavelet transform outputs a complex matrix, where the rows correspond to the 64 scales, and the columns correspond to the emitted wavelength sampling points. The square of the modulus of each element in the complex matrix is calculated to obtain the time-frequency energy distribution matrix. The rows of the time-frequency energy distribution matrix represent the frequency scale dimension, the columns represent the emitted wavelength dimension, and the matrix elements represent the energy intensity at the corresponding scale and emitted wavelength position.
[0036] Five time-frequency energy distribution matrices corresponding to the same corrected characteristic spectrum are superimposed according to preset spectral intensity weights to generate a single-spectrum modulation feature map. The preset spectral intensity weights are determined based on the variance contribution rate of each independent spectral line component, calculated as the variance of the independent spectral line component divided by the sum of the variances of the five independent spectral line components. The five time-frequency energy distribution matrices are multiplied by their corresponding spectral intensity weights and then summed element-wise to obtain a superimposed time-frequency energy distribution matrix. This superimposed time-frequency energy distribution matrix is the single-spectrum modulation feature map. The single-spectrum modulation feature map has 64 rows and the number of columns corresponds to the number of sampling points at the transmission wavelength.
[0037] The interaction mapping matrix is calculated for the single-spectrum modulation feature maps corresponding to all corrected feature spectra. This interaction mapping matrix is then used to weight and fuse the single-spectrum modulation feature maps to generate a composite structure feature spectrum. The interaction mapping matrix is calculated as follows: all single-spectrum modulation feature maps are flattened into one-dimensional vectors along the frequency scale dimension. These one-dimensional vectors are then arranged column-wise to form a data matrix. The number of rows in the data matrix equals the total number of elements in the single-spectrum modulation feature maps, and the number of columns equals the number of single-spectrum modulation feature maps. Singular value decomposition is performed on the data matrix to obtain a left singular vector matrix, a singular value diagonal matrix, and a right singular vector matrix. The first two column vectors of the right singular vector matrix are taken as the mapping directions to construct the interaction mapping matrix. The dimensions of the interaction mapping matrix are the number of rows and two columns of the single-spectrum modulation feature maps. Each row vector of the interaction mapping matrix corresponds to the coordinates of a single-spectrum modulation feature map in the two-dimensional interaction space. The weighting coefficients are determined based on the distances between coordinates in the two-dimensional interaction space: For each single-spectrum modulation feature map, the negative exponent sum of its Euclidean distances to all other single-spectrum modulation feature maps in the two-dimensional interaction space is calculated, and this sum is used as the fusion weight for that single-spectrum modulation feature map. The fusion weight is negatively correlated with the distance. The obtained fusion weights are then used to perform a weighted average on all single-spectrum modulation feature maps. The weighted averaging operation is performed independently at each element position of the time-frequency energy distribution matrix, resulting in a composite structure feature spectrum map. The composite structure feature spectrum map is also a time-frequency energy distribution matrix with 64 rows and the number of columns equal to the number of sampling points for the transmitted wavelength.
[0038] Example 4: In practice, the process of pre-constructing the spectral correlation model begins with acquiring the pure fluorescence spectral data of known oil components. The pure fluorescence spectral data of the known oil components are obtained using a fluorescence spectral acquisition device with the same configuration as in Example 1, which includes an excitation light source, a polarization filter assembly, a sample cell, and a fluorescence detector. Twenty pure standard substances of known oil components are selected, covering alkanes, cycloalkanes, aromatic hydrocarbons, and nitrogen-containing heterocyclic oils. The pure standard substance for each known oil component is measured separately. For each known oil component, a pure standard solution with a concentration of 100 mg / L was prepared. The single-component standard solution was injected into the sample cell. Using all nineteen excitation wavelengths in the excitation wavelength set and all thirteen polarization angles in the preset polarization angle set, the fluorescence spectrum under each combination condition was recorded according to the procedure for collecting multi-angle polarization induced spectra in Example 1. This formed the pure fluorescence spectrum data of the known oil component. The pure fluorescence spectrum data contained information in four dimensions: excitation wavelength, polarization angle, emission wavelength, and fluorescence intensity.
[0039] For each known oil component, the pure fluorescence spectral data are labeled with corresponding molecular structure feature tags. The molecular structure feature tags consist of a group category code and structural feature parameters. The group category code uses a three-digit integer, where 100-199 represents alkyl groups, 200-299 represents cycloalkyl groups, 300-399 represents monocyclic aromatic groups, 400-499 represents polycyclic aromatic groups, and 500-599 represents nitrogen-containing heterocyclic groups. The structural feature parameters include the number of carbon atoms, the number of rings, the number of aromatic rings, and the number of nitrogen atoms. Taking toluene as an example, the group category code for toluene's molecular structure feature tag is 301, the number of carbon atoms is 7, the number of rings is 0, the number of aromatic rings is 1, and the number of nitrogen atoms is 0. The molecular structure feature tags are stored as bond-value pairs associated with the corresponding pure fluorescence spectral data.
[0040] Multidimensional feature modulation was performed on the fluorescence spectral data of the pure product to obtain the characteristic spectrum of the pure product's composite structure. Multidimensional feature modulation included joint modulation of the polarization angle dimension and the emission wavelength dimension of the pure product's fluorescence spectral data. The same multi-dimensional feature modulation process as in Example 3 was performed on the pure product fluorescence spectral data: the pure product fluorescence spectral data were organized into multiple pure product spectra according to the excitation wavelength and polarization angle, and independent component analysis was performed on each pure product spectrum. The number of independent spectral line components in the independent component analysis was set to 5, resulting in 5 independent spectral line components for each pure product spectrum; continuous wavelet transform was used for time-frequency transformation of each independent spectral line component, with Morse wavelet as the wavelet basis function, symmetry parameter set to 3, time-bandwidth product set to 60, and scale parameter from 1 to 64 taking 64 scale values at exponential intervals to generate the corresponding pure product time-frequency energy distribution matrix; the 5 pure product time-frequency energy distribution matrices corresponding to the same pure product spectrum were superimposed according to the spectral line intensity weight, which was determined according to the variance contribution rate of each independent spectral line component; the single spectrum modulation feature maps corresponding to all pure product spectra were interactively mapped and weighted to obtain the pure product composite structure feature spectrum of the known oil component. The above operation was performed on pure standard substances of twenty known oil components one by one, and a total of twenty pure composite structure characteristic spectra were obtained.
[0041] The characteristic spectra of pure compound structures were clustered according to molecular structure feature labels, generating multiple group category clusters. A hierarchical clustering method based on group category coding was used. Pure compound structure characteristic spectra with the same group category code were grouped into the same group category cluster. If a group category code contains only a single pure compound structure characteristic spectrum, then that group category cluster contains only one element. After clustering, five group category clusters were obtained: alkyl chain cluster, cycloalkyl chain cluster, monocyclic aromatic hydrocarbon cluster, polycyclic aromatic hydrocarbon cluster, and nitrogen-containing heterocyclic hydrocarbon cluster.
[0042] Group feature basis functions were extracted from the feature spectra of pure compound structures within each group category cluster. Principal component analysis was used for group feature basis function extraction. Taking a group category cluster as the processing object, all pure compound structure feature spectra within that group category cluster were flattened into row vectors, with each pure compound structure feature spectrum having a dimension of [missing value]. 64 is the frequency scale dimension. The number of sampling points is the emission wavelength. All flattened row vectors are stacked to form a sample matrix. The covariance matrix of the sample matrix is calculated, and eigenvalue decomposition is performed on the covariance matrix. The eigenvector corresponding to the largest eigenvalue is extracted as the group feature basis function for that group cluster. The dimension of the group feature basis function is consistent with the dimension of the flattened pure product composite structure feature spectrum. The group feature basis function is then reshaped into a 64-row matrix. The matrix form of columns. Group feature basis functions are associated and stored with the group category codes in the corresponding molecular structure feature labels to construct a spectral association model. The spectral association model contains five entries, each storing a group category code and a group feature basis function.
[0043] Based on a pre-constructed spectral association model, group matching analysis is performed on the composite structure feature spectrum to generate spatial distribution data of oil components. The composite structure feature spectrum refers to the composite structure feature spectrum generated from the water sample under test through multi-dimensional feature modulation as described in Example 3. The composite structure feature spectrum has a size of 64 rows. List.
[0044] The composite structure feature spectrum is sliced along the wavelength dimension to obtain multiple wavelength slice sub-images. The wavelength dimension corresponds to the column direction of the composite structure feature spectrum. The slicing operation is performed along the column direction using a sliding window with a step size of 1 and a window width of 8 columns, resulting in a total of [number missing] wavelength slice sub-images. Each wavelength slice has a size of 64 rows and 8 columns.
[0045] For each wavelength slice sub-image, template matching is performed between the sub-image and the feature basis function corresponding to each group category cluster in the spectral association model, and the sub-image matching similarity is calculated. The mapping relationship is: the first sub-image in the spectral association model is used to match the feature basis function of each group category cluster in the spectral association model. The characteristic basis functions of each group are sliced along the wavelength dimension with the same step size of 1 and a window width of 8 columns to obtain the first group. Multiple template slices corresponding to the characteristic basis functions of each group. For the first characteristic spectrum of the composite structure... The wavelength slice sub-image, compared with the first wavelength slice sub-image. The first characteristic basis function corresponding to the group For each template slice, calculate the subgraph matching similarity between the two. :
[0046] in, Indicates the first The wavelength slice sub-image and the first Subgraph matching similarity between characteristic basis functions of each group. The value range is between -1 and 1; Indicates the index of the wavelength slice subplot. The value range is from 1 to Integers; Indicates the index of the characteristic basis function of the group. The value of is an integer from 1 to 5; An index representing the frequency scale dimension. The value of is an integer from 1 to 64; This indicates the column index within the wavelength slice window. The value of is an integer from 1 to 8; The first characteristic spectrum of the composite structure The wavelength slice subplot is located at the frequency scale index. Column index The element value at that position; The first characteristic spectrum of the composite structure The arithmetic mean of all element values within a wavelength slice subplot; Indicates the first The first characteristic basis function of the group The template slice is located at the frequency scale index Column index The element value at that position; Indicates the first The first characteristic basis function of the group The arithmetic mean of all element values within a template slice.
[0047] A similarity space matrix is generated based on the subgraph matching similarity of all wavelength slice subgraphs. The dimension of the similarity space matrix is... Row 5, 5th column, the similarity space matrix Line 1 Column elements are .
[0048] A fractal mapping algorithm is used to dynamically reconstruct the similarity space matrix to obtain spatial distribution data of oil components. The fractal dimension of the similarity space matrix is calculated to determine the self-similarity scale parameter of the reconstructed image. The fractal dimension calculation uses box counting: the similarity space matrix is treated as a two-dimensional grayscale image, and the grayscale values are mapped by linearly scaling the elements of the similarity space matrix to the integer range of 0 to 255, resulting in a scaled matrix. The scaled matrix is then mapped using a side length of... Covered by a square box, The value sequence is 2, 4, 8, 16, 32. Count the dimensions of each side. Minimum number of boxes required to completely cover all non-zero grayscale values .calculate and The slope of the linear regression between the two values, and the absolute value of the linear regression slope as the fractal dimension. Self-similarity scale parameter Depend on The self-similarity scale parameter was calculated. The value range is between 0 and 1.
[0049] An iterative shrinkage transformation is performed on the similarity space matrix based on the self-similarity scaling parameter to generate a reconstructed spectral map. The iterative shrinkage transformation is performed in 10 iterations. In the... In the next iteration, the matrix to be transformed is multiplied by the shrinkage factor. , ,in Indicates the iteration number. The value of is an integer from 1 to 10. Represents the self-similarity scale parameter of The power of 1. In the first iteration, the matrix to be transformed is the similarity space matrix. In each iteration, the matrix's row and column dimensions are shrunk to a fraction of the original matrix's dimensions. The rounding and shrinking operation employs bilinear interpolation resampling. The matrix obtained after each iteration of shrinkage is upsampled back to the original size of the similarity space matrix using bilinear interpolation, resulting in 10 intermediate shrunken matrices. The arithmetic mean of these 10 intermediate shrunken matrices is calculated based on their corresponding element positions to obtain the reconstructed spectrogram. The reconstructed spectrogram has the same dimension as the similarity space matrix. Rows, 5 columns.
[0050] Calculate the residual space between the reconstructed spectrum and the composite structure feature spectrum, and extract the difference feature vectors from the residual space. Expand the reconstructed spectrum column-wise into a length of... A one-dimensional reconstruction vector. The composite structure feature spectrum is slide-sampled along the wavelength dimension with a step size of 1 and a window width of 8 columns. The arithmetic mean of all elements in each sample is calculated, resulting in a vector of length [missing information]. The one-dimensional spectral summary vector is copied 5 times and horizontally concatenated to obtain a length of [length missing]. The one-dimensional expanded spectral vector. The residual space is obtained by element-wise subtraction of the one-dimensional expanded spectral vector and the one-dimensional reconstructed vector. The residual space is of length . A one-dimensional vector. The difference feature vector is a vector directly formed by all elements of the residual space, and its dimension is... .
[0051] A pre-constructed trend correlation matrix is used to map the difference feature vectors to the corresponding positions in the reconstructed spectra, generating spatial distribution data of oil components. The pre-constructed trend correlation matrix is constructed as follows: pure fluorescence spectra of pure standard substances of twenty known oil components are collected; the residual vector between the reconstructed spectrum and the one-dimensional expanded spectral vector of each known oil component is calculated according to the aforementioned group matching analysis procedure; and the residual vector of each known oil component is concatenated with its corresponding one-dimensional reconstructed vector to form a matrix of size [missing information]. A 20-dimensional correlation sample vector was obtained for twenty known oil components. A correlation sample matrix was constructed with these vectors, consisting of 20 rows and columns. The trend correlation matrix is obtained by solving the correlation sample matrix using ridge regression, with the regularization coefficient of the ridge regression set to 0.01. The dimension of the trend correlation matrix is... List, Okay. Multiply the difference feature vectors by the trend correlation matrix on the right to obtain the mapping vector, the length of which is... Reshape the mapping vector to The matrix is in the form of 5 rows and 5 columns. The reshaped matrix is the spatial distribution data of oil components. Each row of the spatial distribution data of oil components corresponds to a wavelength position, and each column corresponds to a group category. The element value represents the distribution intensity of the corresponding group category at the corresponding wavelength position.
[0052] See Figure 3 The figure shows the intensity distribution curves of five types of oil components within the emission wavelength range of 270 nm to 490 nm. The horizontal axis represents the emission wavelength (unit: nanometers), and the vertical axis represents the normalized intensity distribution. The five types of oil components are: alkyl groups (solid blue line), cycloalkyl groups (dashed orange line), monocyclic aromatics (dotted green line), polycyclic aromatics (dotted red line), and nitrogen-containing heterocyclic groups (dotted purple line). The intensity distribution curves of each category all exhibit clear peak characteristics, and the peak positions are correlated with the molecular structure characteristics of the oil components.
[0053] Specifically, the fluorescence intensity of alkyl groups peaks at approximately 320 nm, with a peak intensity close to 0.95, before rapidly decreasing to near zero, indicating that the fluorescence response of this type of component is mainly concentrated in the emission wavelength range around 320 nm. Cycloalkyl groups peak at approximately 360 nm, with a peak intensity of approximately 0.82. The peak shape is relatively symmetrical and the coverage is slightly wider than that of alkyl groups, indicating a clear wavelength distinction in their fluorescence characteristic spectrum compared to alkyl groups. Monocyclic aromatic hydrocarbons peak at approximately 400 nm, with a peak intensity of approximately 0.92. The peak shape is relatively sharp, showing that the fluorescence emission of this type of component has strong wavelength selectivity. Polycyclic aromatic hydrocarbons peak at approximately 445 nm, with a peak intensity of approximately 0.70. The peak shape is relatively broad, suggesting that their fluorescence response is distributed over a wide wavelength range. Nitrogen-containing heterocyclic hydrocarbons peak at approximately 475 nm, with a peak intensity of approximately 0.60. The peak shape exhibits a slight right-biased characteristic, demonstrating the unique fluorescence emission features of this type of component.
[0054] Example 5: In practice, dynamic fitting of characteristic morphology is performed on the composite structure characteristic spectrum based on the spatial distribution data of oil components to obtain the characteristic peak corrected spectrum. The spatial distribution data of oil components was generated in Example 4. A matrix of 5 rows and 5 columns, the composite structure feature spectrum is generated in Example 3 with 64 rows and 5 columns. The time-frequency energy distribution matrix of the column. The purpose of characteristic morphology dynamic fitting is to separate and accurately reconstruct the fluorescence response of each oil component in the characteristic spectrum of the composite structure along the emission wavelength dimension.
[0055] Characteristic peak regions corresponding to each oil component in the composite structure characteristic spectrum are labeled based on the spatial distribution data of oil components. Each column of the spatial distribution data of oil components corresponds to a group category, including alkyl groups, cycloalkyl groups, monocyclic aromatic groups, polycyclic aromatic groups, and nitrogen-containing heterocyclic groups. List, The value of is an integer from 1 to 5. Calculate the th... The arithmetic mean and standard deviation of all element values in the column will be used to calculate the first element. Wavelength slice locations whose element values are greater than the arithmetic mean plus 1.5 standard deviations are recorded as significant response locations. These significant response locations are then mapped back to the emission wavelength coordinates of the composite structure feature spectrum along the emission wavelength dimension. The mapping relationship is as follows: The emission wavelength of the column corresponds to the first The central emission wavelength of each wavelength slice sub-image. On the emission wavelength axis of the composite structure characteristic spectrum, a wavelength interval is defined by extending 2 nm to the left and right as the center of the mapped emission wavelength corresponding to each significant response position. This defined wavelength interval is labeled as the characteristic peak region of the oil component of the corresponding functional group category. If the wavelength intervals of two adjacent characteristic peak regions overlap, the overlapping wavelength intervals are merged into a continuous characteristic peak region, and the merged characteristic peak region is simultaneously labeled as the coexistence region of the two functional group categories involved.
[0056] For each characteristic peak region, its contour line is extracted to form a set of local peak curves. The contour line extraction operation is performed along the emission wavelength dimension on the composite structure characteristic spectrum. For the emission wavelength range corresponding to the characteristic peak region in the composite structure characteristic spectrum, the energy values at 64 frequency scales are calculated column-wise to obtain a one-dimensional average energy curve that varies with the emission wavelength. This one-dimensional average energy curve is the contour line of the characteristic peak region. The corresponding contour line is extracted for each labeled characteristic peak region using the above method, and all extracted contour lines are compiled to form a set of local peak curves.
[0057] For each local peak curve in the set of local peak curves, a predefined mathematical function is used to perform point-by-point nonlinear fitting, generating a fitting curve for a single characteristic peak. The predefined mathematical function is an asymmetric Gaussian function, which is expressed by the following formula:
[0058] in, Indicates the emission wavelength The asymmetric Gaussian function value at that location; Indicates the peak amplitude parameter. The initial value is taken as the maximum value of fluorescence intensity in the local peak curve; Indicates the peak center wavelength parameter. The initial value is taken as the emission wavelength corresponding to the maximum fluorescence intensity in the local peak curve; This represents the left half width parameter. The initial value is taken from the emission wavelength. Extend to the left until the fluorescence intensity decreases. When and The wavelength interval between them It is a natural constant; This represents the right half-width parameter. The initial value is taken from the emission wavelength. Extend to the right until the fluorescence intensity drops to When and The wavelength interval between them; Indicates the asymmetric adjustment factor. The value ranges from 0.1 to 0.8. The initial value is set to 0.3. The 0.3 value was chosen based on statistical analysis of the fluorescence peak shapes of pure standard substances of twenty known oil components. Most oil fluorescence peaks exhibited a slight right-biased characteristic. A value of 0.3, slightly lower than the middle range of 0.1 to 0.8, was chosen to balance the fitting convergence speed and fitting accuracy. For each emission wavelength sampling point on the local peak shape curve... Substitute the initial parameters into the formula for the asymmetric Gaussian function and calculate the fitted value. The Levenberg-Marquardt algorithm was used to optimize the parameters. , , , and Iterative optimization is performed, with the objective function being the sum of the actual value of the local peak curve and... The sum of squares of the residuals between the fitted values. The iteration terminates when the parameter update amount is less than... Or the number of iterations reaches 100. The asymmetric Gaussian function curve at the end of the iteration is the curve fitted by a single characteristic peak.
[0059] All individual characteristic peak fitting curves are superimposed, and residual analysis is performed with the composite structure characteristic spectrum to iteratively adjust the fitting parameters. The superposition operation is performed along the emission wavelength dimension, adding the fitting curves of all individual characteristic peaks belonging to the same characteristic peak region point by point at the emission wavelength sampling point to obtain the superimposed fitting curve. For regions belonging to the same characteristic peak region but labeled as coexisting regions of different group categories, the fitting curves of individual characteristic peaks within the coexisting region also participate in the superposition. Using the one-dimensional average energy curve corresponding to the superimposed fitting curve as a reference, the residual between the superimposed fitting curve and the reference reference is calculated. The residual is calculated as the difference between the superimposed fitting curve value and the reference reference value at each emission wavelength sampling point. The residuals are summed at squares over all emission wavelength sampling points to obtain the residual sum of squares. The residual sum of squares is compared with the residual sum of squares from the previous iteration. If the decrease in the residual sum of squares is less than... If the sum of squared residuals converges to a stable value, then the parameters of the fitting curve for each individual characteristic peak are adjusted according to the distribution of the residuals at each emission wavelength sampling point. The adjustment method is to subtract the partial derivative of the residuals with respect to the parameters and multiply it by the step size factor, which is set to 0.05, from the original parameters, and then proceed to the next iteration. When the sum of squared residuals converges to a stable value, all the fitting curves of each individual characteristic peak are re-summed point by point along the emission wavelength dimension to obtain the final characteristic peak corrected spectrum, which is a one-dimensional fluorescence intensity curve with respect to the emission wavelength.
[0060] Based on characteristic peak correction spectroscopy, a pre-trained concentration inversion network is used for quantitative inversion to obtain the concentration values of each oil component. The pre-trained concentration inversion network consists of an input layer, three fully connected layers, two nonlinear activation layers, and an output layer connected sequentially. The input layer contains 4 neurons to receive the concentration inversion input vector. The first fully connected layer is fully connected to the input layer and contains 32 neurons. Each neuron in the first fully connected layer applies a bias to the weighted sum of the output values of the four neurons in the input layer. The first nonlinear activation layer is then connected to the first fully connected layer, using the hyperbolic tangent function as its activation function. The second fully connected layer is fully connected to the first nonlinear activation layer and contains 16 neurons. The second nonlinear activation layer is then connected to the second fully connected layer, also using the hyperbolic tangent function as its activation function. The third fully connected layer is fully connected to the second nonlinear activation layer and contains 8 neurons. The output layer is fully connected to the third fully connected layer. The output layer contains five neurons, each corresponding to a predicted concentration of one of the following oil components: alkyl, cycloalkyl, monocyclic aromatic, polycyclic aromatic, and nitrogen-containing heterocyclic. The activation function for the output layer is a linear function, and the output values of the neurons are directly used as the initial predicted concentration values for the corresponding oil components.
[0061] Morphological analysis was performed on the corrected spectrum of the characteristic peaks to extract kurtosis, skewness, and half-maximum width (HWHM) parameters. The corrected spectrum was treated as a probability density curve, and the first, second, third, and fourth central moments were calculated. Kurtosis was obtained by dividing the fourth central moment by the square of the second central moment and then subtracting 3. Skewness was obtained by dividing the third central moment by the second central moment raised to the power of 1.5. The HWHM parameter was the average of the HWHMs corresponding to all local maxima in the corrected spectrum. The HWHM of a single local maximum was calculated as the difference between the two emission wavelengths to the left and right of the HWHM at that local maximum. The peak height data of the corrected spectrum was the difference between the average fluorescence intensity of all local maxima and the average fluorescence intensity of all local minima.
[0062] The kurtosis, skewness, and half-width at half-maximum (HWHM) parameters are concatenated with the peak height data of the characteristic peak correction spectrum to generate a concentration inversion input vector. The concentration inversion input vector is a column vector of length 4. The first element of the concentration inversion input vector is the kurtosis, the second element is the skewness, the third element is the HWHM parameter, and the fourth element is the peak height data.
[0063] The concentration inversion input vector is fed into the pre-trained concentration inversion network. After weighting and biasing operations in the first fully connected layer, the input vector outputs a 32-dimensional intermediate vector. This 32-dimensional intermediate vector is then transformed element-wise by the hyperbolic tangent function in the first nonlinear activation layer and fed into the second fully connected layer. After weighting and biasing operations, it outputs a 16-dimensional intermediate vector. This 16-dimensional intermediate vector is then transformed element-wise by the hyperbolic tangent function in the second nonlinear activation layer and fed into the third fully connected layer. After weighting and biasing operations, it outputs an 8-dimensional intermediate vector. This 8-dimensional intermediate vector is fed into the output layer and, after linear weighting and biasing operations, outputs a 5-dimensional column vector. Each element of this 5-dimensional column vector represents the concentration value of the corresponding oil component.
[0064] The training process of the pre-trained concentration inversion network is as follows: 100 sets of mixed oily water samples with known concentrations were collected. These mixed oily water samples contained five types of oil components: alkyl groups, cycloalkyl groups, monocyclic aromatics, polycyclic aromatics, and nitrogen-containing heterocyclic compounds. The concentration range of each oil component was from 0.1 mg / L to 100 mg / L. For each set of mixed oily water samples, all steps prior to the characteristic peak correction spectrum in Examples 1 to 5 were performed to obtain the characteristic peak correction spectrum for each set of mixed oily water samples. Kuroism, skewness, half-width at half-maximum (HWHM), and peak height data were extracted from each set of characteristic peak correction spectra to generate 100 concentration inversion input vectors. The 100 concentration inversion input vectors and the corresponding true concentration values of the five oil components constituted the training sample set. The network training used the mean squared error loss function, which calculated the average of the squares of the differences between each component of the 5-dimensional concentration prediction vector output by the network and the true concentration vector. The Adam optimizer was selected, with a learning rate of 0.001, a first-order momentum decay coefficient of 0.9, and a second-order momentum decay coefficient of 0.999. The training batch size was set to 16 samples, and the training epochs were set to 200. In each epoch, the training sample set was randomly sorted, and samples were extracted according to the batch size and input into the network. The loss function value was calculated, and the gradients of all weight and bias parameters were calculated through backpropagation. The Adam optimizer updated all weight and bias parameters based on the gradients. After 200 epochs of training, the network weight and bias parameters were saved, resulting in the pre-trained concentration inversion network.
[0065] See Figure 4 In the figure, the horizontal axis represents the emission wavelength, ranging from 250 nm to 500 nm, and the vertical axis represents the fluorescence intensity, in au (any unit). The black curve is the characteristic peak corrected spectrum obtained by dynamic fitting of characteristic morphology in Example 5 of this invention, which clearly shows the fluorescence response characteristics of oil pollutants in the water sample to be tested.
[0066] The figure shows the characteristic peak regions of five types of oil groups, marked according to the spatial distribution data of oil components, and distinguished by different colored shading: alkyl group characteristic peak region (light blue, around 325 nm), cycloalkyl group characteristic peak region (light orange, around 365 nm), monocyclic aromatic hydrocarbon characteristic peak region (light green, around 404 nm), polycyclic aromatic hydrocarbon characteristic peak region (light red, around 440 nm), and nitrogen-containing heterocyclic hydrocarbon characteristic peak region (light purple, around 479 nm). Each characteristic peak region is marked with its corresponding dashed line indicating its center wavelength. The indicated wavelengths closely match the peak positions of the corrected spectra of the black characteristic peaks, demonstrating the accurate separation and localization of the fluorescence peaks of each oil component using this method.
[0067] The fluorescence intensity curves show a peak value of approximately 1.0 in the alkyl characteristic region, followed by a second strong peak in the cycloalkyl characteristic region, slightly above 1.0. In the monocyclic aromatic hydrocarbon characteristic region, the fluorescence intensity significantly increases to a peak value of approximately 1.5, indicating that this component has the most significant fluorescence response. The fluorescence intensities in the polycyclic aromatic hydrocarbon and nitrogen-containing heterocyclic characteristic regions are both above 1.2 and exhibit distinct peak shapes, suggesting that these two types of oil components also contribute strongly to the fluorescence in the sample.
[0068] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra, characterized in that, The method includes: Collect multi-angle polarization-induced spectra of the water sample to be tested at multiple preset wavelengths to serve as the initial set of fluorescence spectra; Each spectrum in the initial set of fluorescence spectra is corrected for its peak position to obtain a set of corrected characteristic spectra. Multi-dimensional feature modulation is performed on the corrected feature spectrum set to generate a composite structure feature spectrum; Based on a pre-constructed spectral association model, the feature spectrum of the composite structure is analyzed by group matching to generate spatial distribution data of oil components; Based on the spatial distribution data of the oil components, dynamic fitting of the characteristic morphology of the composite structure is performed on the characteristic morphology of the composite structure to obtain the characteristic peak corrected spectrum; Based on the characteristic peak corrected spectrum, a pre-trained concentration inversion network is invoked to perform quantitative inversion and obtain the concentration values of each oil component.
2. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The multi-angle polarization-induced spectra of the water sample to be tested at multiple preset wavelengths specifically include: The first excitation wavelength is selected from the set of excitation wavelengths to irradiate the water sample to be tested. At the same time, the rotation angle of the polarization filter assembly is controlled to record the fluorescence intensity at multiple angles simultaneously, forming the first angle polarization spectrum. By traversing the remaining excitation wavelengths in the set of excitation wavelengths, the rotation control and fluorescence intensity recording of the polarization filter assembly are executed sequentially to obtain multiple angular polarization spectral lines corresponding to each excitation wavelength; The first angular polarization spectral line and multiple angular polarization spectral lines corresponding to the remaining excitation wavelengths are spliced together in order of excitation wavelength to generate the initial fluorescence spectrum set.
3. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The step of correcting the peak positions of each spectrum in the initial set of fluorescence spectra specifically includes: Extract the set of peak position coordinates for each spectrum in the initial set of fluorescence spectra; Obtain the position of the standard fluorescence characteristic peak of the internal standard substance in the water sample to be tested, and calculate the offset matrix between the set of peak position coordinates and the position of the standard fluorescence characteristic peak. Based on the offset matrix, the wavelength axis of each spectrum is shifted and corrected, and the shifted spectrum is iterated by peak matching using the prediction algorithm until all peak position deviations are less than a preset threshold, thereby obtaining the corrected feature spectrum set.
4. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The step of performing multi-dimensional feature modulation on the corrected feature spectrum set specifically includes: Perform independent component analysis on each corrected feature spectrum in the set of corrected feature spectra to extract multiple independent spectral components; Perform time-frequency transformation on each independent spectral line component to generate the corresponding time-frequency energy distribution matrix; All time-frequency energy distribution matrices corresponding to the same corrected feature spectrum are superimposed according to the preset spectral intensity weights to generate a single-spectrum modulation feature map; The interaction mapping matrix is calculated for the single-spectrum modulation feature maps corresponding to all the corrected feature spectra, and the single-spectrum modulation feature maps are weighted and fused using the interaction mapping matrix to generate the composite structure feature spectra.
5. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The pre-built graph association model is constructed in the following way: Obtain pure fluorescence spectral data of multiple known oil components and label the molecular structure feature tags corresponding to each pure fluorescence spectral data; Multidimensional feature modulation is performed on the fluorescence spectrum data of the pure product to obtain the composite structure feature spectrum of the pure product; The composite structure feature spectrum of the pure product is clustered according to the molecular structure feature tags to generate multiple functional group category clusters; Group feature basis functions are extracted from the pure composite structure feature spectra within each group category cluster, and the group feature basis functions are associated and stored with the corresponding molecular structure feature labels to construct the spectrum association model.
6. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 5, characterized in that, The group matching analysis of the composite structure feature spectrum based on the pre-constructed spectral association model specifically includes: The composite structure feature spectrum is sliced according to the wavelength dimension to obtain multiple wavelength slice sub-images; For each wavelength slice sub-image, template matching is performed with the feature basis function corresponding to each group category cluster in the spectrum association model, and the sub-image matching similarity is calculated. A similarity space matrix is generated based on the subgraph matching similarity of all wavelength slice subgraphs, and the similarity space matrix is dynamically reconstructed using a fractal mapping algorithm to obtain the spatial distribution data of the oil components.
7. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 6, characterized in that, The dynamic reconstruction of the similarity space matrix using a fractal mapping algorithm specifically includes: The fractal dimension of the similarity space matrix is calculated to determine the self-similarity scale parameter of the reconstructed graphic; Based on the self-similarity scale parameter, the similarity space matrix is iteratively shrunk to generate a reconstructed spectrogram; Calculate the residual space between the reconstructed spectrum and the composite structure feature spectrum, and extract the difference feature vector in the residual space; The differential feature vectors are mapped to the corresponding positions in the reconstructed spectrum using a pre-constructed trend correlation matrix to generate the spatial distribution data of the oil components.
8. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The step of performing dynamic fitting of the feature morphology of the composite structure feature spectrum based on the spatial distribution data of the oil components specifically includes: Based on the spatial distribution data of the oil components, the characteristic peak regions corresponding to each oil component in the characteristic spectrum of the composite structure are marked; Extract the contour line of each characteristic peak region to form a set of local peak curves; For each local peak curve in the set of local peak curves, a predefined mathematical function is called to perform point-by-point nonlinear fitting to generate a single characteristic peak fitting curve; The fitted curves of all individual characteristic peaks are superimposed and residual analysis is performed with the characteristic spectrum of the composite structure. The fitting parameters are iteratively adjusted, and when the sum of squared residuals converges to a stable value, the corrected spectrum of the characteristic peaks is output.
9. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 1, characterized in that, The quantitative inversion based on the characteristic peak-corrected spectrum, using a pre-trained concentration inversion network, specifically includes: The morphological characteristics of the corrected spectrum of the characteristic peak are analyzed to extract the kurtosis, skewness and full width at half maximum (FWHM) parameters of the characteristic peak. The kurtosis, skewness, and half-width at half-maximum parameters are concatenated with the peak height data of the characteristic peak corrected spectrum to generate a concentration inversion input vector. The concentration inversion input vector is input into the pre-trained concentration inversion network, which is composed of multiple fully connected layers and nonlinear activation layers stacked alternately. The network calculates and outputs the initial concentration prediction values of each oil component through layer-by-layer mapping.
10. The method for quantifying oil pollutants in water based on ultraviolet fluorescence characteristic spectra as described in claim 5, characterized in that, The multidimensional feature modulation includes joint modulation of the polarization angle dimension and wavelength dimension of the pure fluorescence spectral data.