Cross-scale spectrum and chromatographic data associated pollutant conversion evaluation method and system
Through the combination of multi-point sensing acquisition and deep self-attention neural network, spectral and chromatographic data are integrated, and the problem of difficulty in integrating multi-source data in the existing technology is solved, and the accurate prediction and optimization of organic pollutants in the coal gangue/persulfate system is achieved, which improves the scientificity and efficiency of the treatment plan.
Patent Information
- Application Number
- CN202510539921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology lacks methods to effectively integrate multi-source data, and cannot fully utilize cross-scale analysis methods to reveal the organic pollutant conversion process in the coal gangue/persulfate system, resulting in inefficient formulation of treatment plans and difficulty in coping with the diversity changes in complex environmental samples.
Reaction parameters are collected through multi-point sensing, and spectral and chromatographic data are normalized. Multi-layer correlation tensor decomposition and fused data at different scales are used to build a deep self-attention neural network to generate a coal gangue degradation performance predictor to form an optimal processing solution report card.
It realizes accurate prediction and optimization of the degradation process of organic pollutants in the coal gangue/persulfate system, reduces the number of experiments, and improves the scientificity and efficiency of the treatment plan design, especially when dealing with complex environmental samples.
Smart Images

Figure CN120452576A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a pollutant conversion assessment method and system that correlates cross-scale spectral and chromatographic data. Background Art
[0002] The treatment of organic pollutants has always been an important research direction in the field of environmental science. Traditional methods for treating organic pollutants mainly include activated carbon adsorption, ozone oxidation, Fenton oxidation, biodegradation and other technologies. These methods each have their own advantages and disadvantages in practical applications: activated carbon adsorption easily reaches saturation and requires regeneration; ozone oxidation equipment is expensive and energy-intensive; Fenton oxidation has strict requirements on pH value and produces a large amount of iron sludge; and biodegradation has limited effect on difficult-to-degrade organic matter. In recent years, gangue, as a by-product of coal mining, has been studied for its rich mineral composition and large specific surface area for the synergistic treatment of organic pollutants with persulfate, showing good application prospects. However, the optimization of gangue / persulfate combined treatment technology still relies mainly on a large number of experimental studies, including trial and error methods to screen the optimal treatment conditions and explore the mechanism.
[0003] The main shortcoming of existing technologies is the lack of methods to effectively integrate multi-source data (such as spectra, chromatography, etc.), and the inability to fully utilize cross-scale analytical methods to reveal the transformation process of pollutants. Traditional evaluation methods often independently examine the data of a single measurement technology, such as analyzing the structural changes of pollutants only by a single method such as FTIR or UV-Vis, or monitoring concentration changes only by chromatographic methods such as GC-MS. This fragmented analysis method makes it difficult to establish a comprehensive structure-activity relationship. At the same time, the optimization of treatment conditions mainly relies on a large number of repeated experiments, which is time-consuming, labor-intensive and resource-intensive. Especially for complex environmental samples, the lack of systematic predictive evaluation tools leads to inefficient formulation of treatment plans and makes it difficult to cope with the diversity of pollutant types and environmental conditions. Summary of the Invention
[0004] This application provides a pollutant conversion assessment method and system that associates cross-scale spectral and chromatographic data, which is used to accurately predict and optimize the degradation process of organic pollutants in coal gangue / persulfate systems, reduce the number of experiments, and improve the scientificity and efficiency of treatment scheme design.
[0005] In the first aspect, the present application provides a pollutant transformation assessment method that associates cross-scale spectral and chromatographic data, and the pollutant transformation assessment method that associates cross-scale spectral and chromatographic data includes: collecting reaction parameters in the coal gangue degradation system through multi-point sensing, normalizing the acquired spectral maps and chromatographic peaks, and obtaining a pollutant characteristic data matrix; according to the pollutant characteristic data matrix, using multi-layer correlation tensor decomposition to fuse and map data of different scales to obtain a pollutant structure-activity relationship map; based on the pollutant structure-activity relationship map, constructing the correspondence rules between molecular changes and treatment conditions through a deep self-attention neural network to generate a coal gangue degradation efficiency predictor; based on the coal gangue degradation efficiency predictor, simulating and evaluating multiple groups of treatment parameters to form an optimal treatment solution report card.
[0006] In a second aspect, the present application provides a pollutant conversion assessment system that associates cross-scale spectral and chromatographic data, the pollutant conversion assessment system that associates cross-scale spectral and chromatographic data comprising:
[0007] The acquisition module is used to collect reaction parameters in the gangue degradation system through multi-point sensing, normalize the acquired spectral patterns and chromatographic peaks, and obtain a pollutant characteristic data matrix;
[0008] A mapping module is used to fuse and map data of different scales using multi-layer correlation tensor decomposition based on the pollutant characteristic data matrix to obtain a pollutant structure-activity relationship map;
[0009] A construction module is used to construct a correspondence rule between molecular changes and treatment conditions based on the pollutant structure-activity relationship map through a deep self-attention neural network to generate a coal gangue degradation efficiency predictor;
[0010] The simulation module is used to simulate and evaluate multiple groups of treatment parameters based on the coal gangue degradation efficiency predictor to form an optimal treatment solution report card.
[0011] In a third aspect, a pollutant conversion assessment device associated with cross-scale spectral and chromatographic data is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the pollutant conversion assessment device associated with cross-scale spectral and chromatographic data to execute the above-mentioned pollutant conversion assessment method associated with cross-scale spectral and chromatographic data.
[0012] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned pollutant conversion assessment method associated with cross-scale spectral and chromatographic data.
[0013] In the technical solution provided in this application, reaction parameters in the gangue degradation system are collected through multi-point sensing and spectral maps and chromatographic peaks are normalized, achieving unified representation of multi-source data, effectively solving the problem of difficulty in directly comparing data obtained by different measurement methods, and laying a standardized foundation for subsequent analysis; multi-layer correlation tensor decomposition is used to fuse and map data of different scales, overcoming the limitations of traditional methods of independent data analysis, establishing intrinsic connections between cross-scale data, and forming a pollutant structure-activity relationship map containing rich structural information; a deep self-attention neural network is used to construct correspondence rules between molecular changes and treatment conditions, innovatively introducing graph-structured data processing and attention mechanisms, enabling the model to automatically identify key active sites in molecular structures and reaction patterns under different treatment conditions, thereby improving the accuracy and interpretability of predictions; multiple sets of treatment parameters are simulated and evaluated based on the gangue degradation efficiency predictor, significantly reducing the number of experiments and resource consumption. At the same time, a multi-objective comprehensive evaluation more comprehensively considers multi-dimensional indicators such as degradation efficiency, economic cost, and treatment time. The resulting optimal treatment solution report card provides intuitive and clear guidance for practical applications. This method excels in analyzing degradation mechanisms and optimizing treatment conditions, particularly for processing complex environmental samples and rapidly screening treatment conditions. From the perspective of artificial intelligence algorithm application, this approach innovatively combines tensor decomposition technology with deep neural networks, leveraging the advantages of tensor decomposition for high-dimensional data reduction and feature extraction, as well as the ability of deep self-attention neural networks to capture complex nonlinear relationships. The multi-layer correlation tensor decomposition algorithm offers unique advantages for revealing potential correlations between heterogeneous data from multiple sources, while the self-attention mechanism accurately captures structure-activity relationships by dynamically weighting different molecular structural sites. These algorithmic features contribute to this approach by: first, enabling intelligent processing throughout the entire process from data acquisition to output; second, providing data-driven mechanism analysis capabilities; and third, possessing predictive generalization capabilities for emerging pollutants. Overall, this approach effectively integrates multi-source data through algorithmic innovation, significantly improving the evaluation efficiency and optimization accuracy of the gangue / persulfate treatment system, providing a more scientific and efficient technical solution for environmental pollutant remediation. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1This is a schematic diagram of an embodiment of a pollutant conversion assessment method based on correlation of cross-scale spectral and chromatographic data in an embodiment of the present application;
[0016] Figure 2 This is a schematic diagram of an embodiment of a pollutant conversion assessment system for correlating cross-scale spectral and chromatographic data in an embodiment of the present application;
[0017] Figure 3 It is a schematic block diagram of the structure of a pollutant conversion assessment device that associates cross-scale spectral and chromatographic data in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present application embodiment provides a pollutant conversion assessment method and system associated with cross-scale spectral and chromatographic data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in a sequence other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0019] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation includes:
[0020] Step S101: collecting reaction parameters in the gangue degradation system through multi-point sensing, normalizing the acquired spectrum and chromatographic peaks, and obtaining a pollutant characteristic data matrix;
[0021] Step S102: Based on the pollutant characteristic data matrix, multi-layer correlation tensor decomposition is used to fuse and map data of different scales to obtain a pollutant structure-activity relationship map;
[0022] Step S103: Based on the pollutant structure-activity relationship map, a deep self-attention neural network is used to construct a correspondence rule between molecular changes and treatment conditions to generate a coal gangue degradation efficiency predictor;
[0023] Step S104: Based on the gangue degradation efficiency predictor, multiple sets of treatment parameters are simulated and evaluated to form an optimal treatment solution report card.
[0024] It is understood that the execution subject of this application can be a pollutant conversion assessment system that associates cross-scale spectral and chromatographic data, or a terminal or server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0025] In an embodiment of the present application, multiple sensor nodes are arranged in the gangue degradation system to collect key parameters in the reaction process. In specific operations, the reaction parameters collected by the sensor include adsorption temperature (range 20-40°C), adsorption time (0-120 minutes), waste concentration (10-1000 mg / L), pH value (2-10) and gangue particle size (0.1-2.0 mm). At the same time, the gangue sample is subjected to spectral analysis to obtain FTIR (Fourier transform infrared spectroscopy) data to record molecular vibration information in the range of 4000-400 cm-1; UV-Vis (ultraviolet-visible spectroscopy) data to obtain electronic transition characteristics in the wavelength range of 200-800 nm; XRD (X-ray diffraction) data to determine the crystalline phase composition of the gangue; XPS (X-ray photoelectron spectroscopy) data to characterize the chemical state of surface elements. For chromatographic analysis, UPLC and UPLC-MS / MS techniques were used to separate and detect pollutants and intermediates during the degradation process. Chromatographic peak retention times, peak areas, and mass spectral characteristics were recorded. Spectral and chromatographic data were then normalized to eliminate scale differences between different measurement techniques. Normalization employed a min-max normalization method, mapping all data types to the [0, 1] interval. The processed multimodal data were organized according to the three-dimensional relationship of time, feature, and sample, forming a pollutant feature data matrix. Based on the pollutant feature data matrix obtained in the previous step, multi-layer correlation tensor decomposition was used to fuse and map the data at different scales. First, the feature data matrix was preprocessed using tensor nuclear norm minimization to reduce data redundancy and noise. The CANDECOMP / PARAFAC decomposition algorithm was then used to decompose the three-dimensional data tensor into a series of factor matrices to extract underlying data patterns. Non-negativity constraints and sparsity regularization were applied to these factor matrices to ensure physical interpretability of the decomposition results. Kendall correlation coefficients were then calculated between the features of different modalities to construct a cross-scale feature correlation network. Hierarchical clustering and edge-weight threshold filtering are performed on the network to identify highly correlated feature clusters. These feature clusters are cross-mapped with molecular structure information to establish a correspondence between spectral features and molecular structural units. Finally, a weighted bidirectional correlation graph is constructed to form a pollutant structure-activity relationship map, which clearly demonstrates the quantitative relationship between molecular structural features and degradation activity. Using the generated pollutant structure-activity relationship map, a deep self-attention neural network is used to construct correspondence rules between molecular changes and treatment conditions. The relationship map is first converted into a standardized graph data input, consisting of a node feature matrix and a topological adjacency matrix. Spatial features are then extracted from the data using a multi-layer graph convolutional network to capture the reactivity characteristics of local regions within the molecular structure. The extracted structural features are processed using a multi-head self-attention mechanism, with each attention head learning the interdependencies between features from a different perspective.A dot-product attention method is used to assign weights to features at different locations. Reaction condition parameters (temperature, pH, particle size, and persulfate dosage) are fused and encoded with structural features to generate a condition-aware attention feature map. This feature map is then subjected to deep feature extraction using a multi-layer residual connection module. This feature map is then subjected to dimensionality reduction mapping using a feedforward neural network with regularization to generate a prediction vector for pollutant degradation efficiency. Finally, through error calculation with experimental data and parameter optimization, a coal gangue degradation efficiency predictor with inference capabilities is developed. Based on the degradation efficiency predictor constructed in the previous step, simulations and evaluations of multiple treatment parameter combinations are conducted. First, a parameter search space is constructed, covering four key dimensions: temperature, pH, coal gangue particle size, and persulfate dosage. Latin hypercube sampling is used to generate uniformly distributed test samples for each parameter combination within the search space. These samples are standardized and fed into the degradation efficiency predictor to obtain the corresponding degradation efficiency prediction results. The prediction results are then subjected to a multi-objective evaluation, comprehensively considering three key indicators: degradation efficiency, economic cost, and treatment time. The analytic hierarchy process is used to assign appropriate weights to each evaluation metric, generating a comprehensive performance score that balances all factors. Treatment options are ranked based on their comprehensive scores to identify the optimal candidate. Sensitivity analysis and stability assessments are conducted on these candidate options to ensure their reliability in practical applications. Finally, detailed information on the optimal option is compiled into a treatment plan report card to provide specific guidance for project implementation.
[0026] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0027] The adsorption temperature, adsorption time, waste concentration, pH value and gangue particle size in the gangue degradation system were monitored in real time through multiple channels to obtain the original environmental parameter set;
[0028] The FTIR spectrum data was subjected to wavelet transform denoising and baseline correction to obtain the infrared characteristic fingerprint region signal;
[0029] The UV-Vis and fluorescence three-dimensional excitation-emission matrix data were subjected to scattered light elimination and inner filter effect correction to obtain purified spectral signals.
[0030] Perform peak identification and peak area integration calculation on UPLC-MS / MS chromatograms to obtain quantitative retention time-response intensity pairs;
[0031] Based on synchronized timestamps, the multimodal monitoring data are aligned in time dimension and missing values are interpolated to obtain a complete time series dataset.
[0032] Apply the minimum-maximum normalization algorithm to the complete time series data set to obtain a multi-source data matrix with unified dimensions;
[0033] The multi-source data matrix is reorganized according to the time-feature-sample three-dimensional structure to obtain the pollutant feature data matrix.
[0034] Specifically, multi-channel real-time monitoring involves deploying multiple parameter sensors simultaneously within the gangue / persulfate treatment system to measure key parameters such as adsorption temperature, adsorption time, waste concentration, pH, and gangue particle size. In practice, a temperature sensor (PT100) monitors the reaction temperature with an accuracy of ±0.1°C; a photoelectric timer records adsorption time; waste concentration is measured using an online conductivity meter or colorimetric method; a pH electrode monitors changes in solution pH; and gangue particle size is controlled through pre-screening. These sensors convert analog signals into digital signals via a data acquisition module, generating a raw environmental parameter set containing the timestamp and corresponding parameter values for each measurement point, recording the changes in external conditions throughout the degradation process. For FTIR spectral data, wavelet transform denoising involves using wavelet functions to perform multi-scale decomposition of the spectral signal to distinguish between signal and noise components. In practice, an appropriate wavelet basis function (such as the Daubechies wavelet) is selected to decompose the raw FTIR signal, followed by thresholding the high-frequency noise coefficient, and finally reconstructing the denoised signal. Baseline correction is a common drift issue in FTIR spectral data processing. Polynomial fitting or adaptive iterative algorithms are used to remove the effects of baseline drift. The processed spectral data focuses on the infrared fingerprint region (1800-600 cm⁻¹), which contains key vibrational information of the pollutant's molecular structure, such as the characteristic absorption peaks of functional groups such as benzene rings, carboxyl groups, and hydroxyl groups.
[0035] Processing UV-Vis and fluorescence three-dimensional excitation-emission matrix data is more complex. Scattered light removal involves removing Rayleigh and Raman scattering interference from the spectrum. These scattering signals appear as diagonal bands in the fluorescence matrix. These scattering signals are eliminated by setting an appropriate wavelength range mask or interpolation algorithm. Inner filter correction addresses signal distortion caused by sample self-absorption and reabsorption, using dilution methods or mathematical models (such as the internal standard method). The resulting purified spectral signal accurately reflects the electronic structure characteristics of the pollutant and its degradation products, providing a basis for subsequent structural elucidation. Processing UPLC-MS / MS chromatograms is key to monitoring changes in pollutant concentration during degradation. Peak identification begins by determining the background noise level. A signal-to-noise ratio threshold (typically 3) is then set, and peak detection is performed for signals exceeding the threshold. Detected chromatographic peaks are fitted with a Gaussian or exponentially modified Gaussian fit to determine the peak start, end, and apex positions. Peak area integral calculation calculates the quantitative value of the component content by summing the product of signal intensity and time within the peak range. Each chromatographic peak is characterized by two parameters: retention time and response intensity, forming a retention time-response intensity pair for qualitative and quantitative analysis.
[0036] Multimodal data alignment based on synchronized timestamps is a key step in processing heterogeneous data. Due to varying sampling frequencies and trigger times across analytical instruments, raw data can be misaligned in the temporal dimension. By adding precise timestamps to each data point, a unified temporal reference system is established. Linear or spline interpolation is then used to interpolate sparsely sampled regions and address missing values during the measurement process. This results in a complete time series dataset that is consistent in the temporal dimension, facilitating subsequent correlation analysis. The min-max normalization algorithm unifies data of varying dimensions and numerical ranges to the same scale. Specifically, the minimum value of each feature dimension is subtracted and then divided by the difference between the maximum and minimum values to map the data to the [0, 1] interval. This normalization eliminates dimensional differences, preventing large numerical features from dominating the analysis results, and resulting in a dimensionally unified multi-source data matrix. The processed multi-source data is then reorganized into a three-dimensional structure: time-feature-sample. The time dimension reflects the dynamics of the degradation process, the feature dimension contains molecular features acquired by different measurement methods, and the sample dimension distinguishes between different experimental conditions or replicate measurements. This three-dimensional tensor structure can simultaneously capture the temporal dynamics, feature diversity, and sample variability of the data, forming a characteristic data matrix that comprehensively describes the pollutant degradation process.
[0037] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0038] Applying tensor low-rank representation and nuclear norm minimization to the pollutant characteristic data matrix, a refined data tensor is obtained that eliminates redundancy and noise.
[0039] The refined data tensor is input into the improved CANDECOMP / PARAFAC decomposition algorithm for three-dimensional decomposition operation to obtain a core factor matrix group containing time dimension, feature dimension and sample dimension;
[0040] Apply L1 / L2 mixed norm regularization and non-negative constraint optimization to the core factor matrix group to obtain a set of characteristic component expressions with physical meaning;
[0041] Based on the characteristic component expression set, the Kendall correlation coefficient matrix of cross-spectral modalities and cross-chromatographic modalities was calculated to obtain the characteristic dependency network reflecting the correlation between multiple monitoring methods;
[0042] Applying hierarchical clustering and edge weight threshold filtering to the feature dependency network, we obtain highly correlated feature clusters that reveal the characteristics of pollutants.
[0043] Cross-mapping the highly correlated feature clusters with the molecular vibrational mode database calculated based on density functional theory to obtain a spectral feature-molecular structural unit correspondence table;
[0044] Based on the spectral feature-molecular structure unit correspondence table, a weighted bidirectional correlation graph structure containing active site markers was constructed to obtain a pollutant structure-activity relationship map describing the degradation mechanism.
[0045] Specifically, tensor low-rank representation is a technique for reducing the complexity of high-dimensional data, which assumes that the real data lies on a low-dimensional manifold. Nuclear norm minimization is an implementation method of tensor low-rank representation, which eliminates redundant information and random noise in the data by minimizing the nuclear norm of the tensor. In specific operations, the pollutant characteristic data matrix is regarded as a third-order tensor X, and each mode is expanded to obtain matrices X1, X2, and X3. Then, the optimization problem is solved to minimize the sum of the nuclear norms of these matrices while maintaining closeness to the original data. This process filters out small disturbances and measurement errors in the data to obtain a refined data tensor with a clearer structure.
[0046] Feeding the refined data tensor into the improved CANDECOMP / PARAFAC decomposition algorithm for three-dimensional decomposition is central to multidimensional data analysis. The CANDECOMP / PARAFAC algorithm decomposes high-order tensors into the sum of a series of rank-1 tensors. The improved algorithm enhances convergence and stability by adding adaptive step sizes and alternating direction multiplication to the traditional algorithm. Specifically, the third-order refined data tensor is decomposed into the outer product of three low-dimensional factor matrices, corresponding to the time, feature, and sample dimensions, respectively. Each column in the factor matrix represents a latent component, each capturing an underlying pattern in the data. A mixed L1 / L2 norm regularization is applied to the core factor matrix group to promote both sparsity and structured sparsity. The L1 norm forces most parameters to zero, preserving a small number of important features; the L2 norm prevents excessive sparsity and maintains model stability. Furthermore, non-negativity constraint optimization ensures that all factor matrix elements are greater than or equal to zero, which is physically meaningful, as physical quantities such as spectral intensity and concentration are typically non-negative. This processing simplifies complex data patterns into a small number of characteristic components with clear physical meaning, such as the characteristic peaks of specific functional groups or the concentration trend of a certain type of compound.
[0047] Calculating a correlation coefficient matrix based on the characteristic component expression set is key to revealing the intrinsic connections between multi-source data. The Kendall correlation coefficient is a nonparametric statistic that measures the consistency of the ordinal relationship between two variables and is not strongly affected by the data distribution form or outliers. The calculation process calculates the Kendall correlation coefficient τ for each pair of spectral and chromatographic features to form a correlation coefficient matrix. This matrix reflects the degree of correlation between data obtained by different monitoring methods (such as FTIR, UV-Vis, UPLC-MS / MS), reveals the intrinsic connections between cross-scale data, and forms a feature dependency network.
[0048] Applying hierarchical clustering to feature dependency networks is an effective method for identifying clusters of highly correlated features. Hierarchical clustering begins with each feature as an independent cluster and gradually merges the most similar clusters until a predetermined number of clusters or similarity threshold is reached. Simultaneously, edge weight threshold filtering removes connections with correlations below a certain threshold, retaining strong correlations. The resulting highly correlated feature clusters contain features that exhibit consistent behavior during degradation, often corresponding to the same chemical processes or structural changes. Cross-mapping these highly correlated feature clusters with a database of molecular vibrational modes is a crucial step in establishing a correlation between spectral features and molecular structure. Density functional theory is a quantum chemical method for calculating molecular vibrational modes. By solving the Schrödinger equation, it derives molecular energies and wave functions, and then calculates vibrational frequencies and intensities. The cross-mapping process compares experimentally observed spectral features with theoretically calculated vibrational modes to determine which molecular building blocks correspond to peaks at specific wavelengths or frequencies. This process establishes a spectral feature-building block correspondence table, clarifying which spectral changes reflect changes in the molecular structure.
[0049] The final step in representing the degradation mechanism is to construct a weighted bidirectional association graph based on the correspondence table. The bidirectional association graph uses nodes to represent molecular building blocks and treatment conditions, and edges to represent the relationships between them. Weighting means that the thickness or color of the edges indicates the strength of the relationship. Active site markers indicate locations in the molecule that are susceptible to oxidation or reduction, which are often the starting points of degradation reactions. The entire construction process forms a structure-activity relationship map for the pollutant, visually demonstrating the quantitative relationship between molecular structural characteristics and degradation activity, revealing the degradation mechanism.
[0050] Taking the case of treating nitrophenol-containing contaminants as an example, low-rank representation and nuclear norm minimization were first applied to the data matrices obtained from various spectroscopic and chromatographic methods to reduce background noise in the raw data. The processed data were then decomposed into three factor matrices using CP decomposition, extracting the underlying patterns corresponding to the time, feature, and sample dimensions. After applying mixed-norm regularization to the factor matrices, the time dimension showed a clear degradation trend that started fast and then slowed, while the feature dimension highlighted characteristic peaks associated with nitro groups and phenolic hydroxyl groups. Calculation of the Kendall correlation coefficient revealed a high correlation between the nitro group's characteristic peak at 1520 cm-1 in FTIR, the absorption peak at 380 nm in UV-Vis, and the peak at a specific retention time in UPLC, indicating that they reflect changes in the same molecular structure. Hierarchical clustering was used to organize these highly correlated features into clusters, and comparison with a database of vibrational modes calculated from quantum chemistry confirmed the correspondence between the asymmetric stretching vibration of the nitro group and specific electronic transitions. The final structure-activity relationship map clearly shows the degradation pathway of nitrophenol under different pH and temperature conditions, revealing a step-by-step degradation mechanism in which the nitro group is first reduced and then the phenol ring is opened.
[0051] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0052] The pollutant structure-activity relationship map is converted into a node feature matrix and a topological structure adjacency matrix to obtain standardized input graph data;
[0053] The standardized input graph data is subjected to spatial feature extraction through a multi-layer graph convolutional network to obtain a structure-aware node representation vector;
[0054] The node representation vector is input into the self-attention calculation unit composed of multiple attention heads to obtain the weighted context feature representation;
[0055] The features are weighted based on the dot-product attention calculation method and fused with the reaction condition parameters for conditional encoding to obtain a conditionally perceived attention feature map. The reaction condition parameters include temperature, pH value, particle size, and dosage.
[0056] The conditionally perceived attention feature map is passed through a multi-layer residual connection module to obtain a deep feature representation;
[0057] The deep feature representation is mapped to dimension reduction using a multi-layer feedforward neural network with a regularization mechanism to obtain a pollutant degradation efficiency prediction vector.
[0058] The pollutant degradation efficiency prediction vector and experimental label data are subjected to error calculation and parameter optimization adjustment to obtain a coal gangue degradation efficiency predictor with reasoning ability.
[0059] Specifically, converting pollutant structure-activity relationship maps into standardized input graph data requires two key matrices: a node feature matrix and a topological adjacency matrix. The node feature matrix X records the characteristic information of each node in the graph. Each row corresponds to a node (such as a molecular building block, functional group, or atom), and each column represents a characteristic (such as spectral peak intensity, chemical shift, or bond energy). The topological adjacency matrix A represents the connectivity between nodes, with each element Aij representing the strength of the connection between node i and node j. For undirected graphs, A is symmetric; for directed graphs, A may be asymmetric. To ensure numerical stability, the adjacency matrix is typically normalized so that the sum of each row or column is 1. Applying spatial feature extraction to the standardized input graph data through a multi-layer graph convolutional network is an effective means of capturing the inter-node relationships. Graph convolutional networks extend the concept of traditional convolutional neural networks and are applicable to graph-structured data in non-Euclidean spaces. In each layer of graph convolution, the feature representation of a node is updated by aggregating information from the node itself and its neighbors. The specific operation is to multiply the node feature matrix X by the normalized adjacency matrix A, and then process it through the weight matrix W and a nonlinear activation function (such as ReLU). By stacking multiple layers of graph convolution, the network can capture the relationships between nodes at longer distances, forming node representation vectors that contain rich structural information. These vectors not only contain the node's own characteristics, but also implicitly contain information about its position and connection pattern within the entire graph structure.
[0060] Inputting node representation vectors into the self-attention computation unit is a key step in enhancing the model's ability to identify important features. The self-attention mechanism consists of multiple attention heads, each of which independently learns feature associations from different perspectives. Each attention head contains three transformation matrices: a query matrix Q, a key matrix K, and a value matrix V. The input node representation vector is multiplied by these three matrices to generate a query vector, a key vector, and a value vector. Attention scores are calculated based on the interaction between the query and key vectors, reflecting the importance of different nodes to the current node. The multi-head design allows the model to simultaneously focus on different aspects of feature relationships, enhancing representational capabilities. The outputs of each attention head are concatenated or averaged to form the final weighted contextual feature representation. The dot product attention calculation is the core of the multi-head self-attention mechanism and is used to determine the attention weights between features. The calculation process first calculates the dot product between each query vector and all key vectors to obtain the raw attention scores. These scores are then divided by the square root of the key vector's dimension to avoid the vanishing gradient problem caused by high dimensionality. A softmax function is then used to convert the scores into weighted distributions that sum to 1. Finally, these weights are weighted summed over the value vector to obtain the output features. In this method, dot-product attention not only processes the node representation vectors but also incorporates reaction condition parameters (temperature, pH, particle size, and dosage) into the calculation through conditional encoding. Conditional encoding converts the reaction parameters into embedding vectors of the same dimension as the node representation vectors, which are then concatenated or additively fused. This processing approach enables the model to learn the complex interactions between molecular structural features and reaction conditions, forming a condition-aware attention feature map. Passing this condition-aware attention feature map through multi-layer residual connection modules is an effective strategy for training deep networks. Residual connections alleviate the vanishing gradient problem in deep network training by adding short-circuit connections between network layers. Each residual module consists of two fully connected layers and a normalization layer. The input features are transformed by these layers and then directly passed to the module output, where they are summed with the transformed features. This structure makes it easier for the network to learn the identity mapping, preserving the original information while learning the residual component. Through processing by multi-layer residual connection modules, the network extracts a deep feature representation containing rich high-order features.
[0061] Dimensionality reduction of deep feature representations using a multi-layer feedforward neural network is a key step in generating final predictions. A feedforward neural network consists of fully connected layers, each of which processes input data through linear transformations and nonlinear activation functions. To prevent overfitting, the network incorporates regularization mechanisms, including weight decay (L2 regularization) and dropout. Weight decay constrains the weights by adding a sum-of-squares term to the loss function; dropout randomly disables some neurons during training to enhance network robustness. Through these processes, the multi-layer feedforward network maps high-dimensional deep features into a low-dimensional vector of pollutant degradation efficiency predictions, with each element corresponding to the predicted degradation rate under specific conditions. The final step in the training process is comparing the pollutant degradation efficiency prediction vector with experimental labeled data to calculate the prediction error. The network parameters are then adjusted using a backpropagation algorithm. Error calculation typically uses mean squared error or cross-entropy loss functions to measure the difference between the predicted value and the true value. Parameter optimization utilizes gradient descent algorithms such as the Adam optimizer, with adaptive learning rate adjustments to accelerate convergence. Through the iterative training process, the network parameters are continuously optimized, and finally a coal gangue degradation efficiency predictor with reasoning ability is formed.
[0062] Taking the case study of treating chlorinated organic pollutants as an example, the relationship graph containing molecular structure information and activity data is first converted into a matrix representation. The node feature matrix records the 128-dimensional features of 35 nodes, and the adjacency matrix represents the chemical bonds and interaction strengths between nodes. A three-layer graph convolutional network processes this data, with each layer extracting progressively higher-level structural features. For example, the first layer identifies local functional groups, the second layer captures molecular fragments, and the third layer extracts global structural features. An eight-head self-attention mechanism processes the node representations output by the graph convolution, with each head focusing on different feature associations. For example, some heads focus on the relationship between chlorine atoms and aromatic rings, while others focus on the interaction between carboxyl groups and neighboring groups. The dot product attention calculation integrates reaction conditions (temperature 25-35°C, pH 3-9, coal gangue particle size 0.1-1.0 mm, persulfate dosage 1-5 mmol / L) with structural features to generate a feature map that comprehensively considers both molecular properties and environmental conditions. Three layers of residual connections further extract deep features, and finally a prediction vector is generated through a three-layer feedforward network (with a 50% dropout rate and a 0.001 weight decay). After 500 rounds of training, the predictor accurately captures the degradation behavior of chlorinated organic compounds under different conditions, specifically identifying the relationship between the number and location of chlorine atoms and pH value on degradation rate.
[0063] In a specific embodiment, the process of performing the step of performing dimensionality reduction mapping on the deep feature representation through a multi-layer feedforward neural network with a regularization mechanism may specifically include the following steps:
[0064] Perform batch normalization on the deep feature representation to obtain standardized features with a mean of zero and a variance of one;
[0065] The standardized features are input into the first fully connected layer for linear transformation to obtain the intermediate hidden layer representation;
[0066] Apply a nonlinear activation function to the intermediate hidden layer representation to obtain a hidden layer activation value with nonlinear characteristics;
[0067] Apply a random dropout operation with dropout rate control to the hidden layer activation values to obtain a sparse representation to prevent overfitting;
[0068] The sparse representation is passed through the second fully connected layer with L2 norm penalty to extract features and obtain a reduced-dimensional feature vector;
[0069] The reduced dimension feature vector is correlated and matched with historical experimental data to obtain a set of candidate degradation efficiency values;
[0070] Probability density estimation is performed based on the set of candidate degradation efficiency values to obtain a pollutant degradation efficiency prediction vector containing the predicted value and its uncertainty quantification.
[0071] Specifically, in the pollutant conversion assessment method that associates cross-scale spectral and chromatographic data, converting deep feature representations into pollutant degradation efficiency prediction vectors is a key link in the entire assessment process. First, batch normalization of deep feature representations is a common technique for neural network training. Batch normalization standardizes each feature in the batch dimension, calculates the mean and variance of each feature in the current batch, and then subtracts the mean from the feature value of each sample and divides it by the variance, so that the mean of the data distribution is 0 and the variance is 1. Batch normalization helps alleviate the problem of internal covariate shift and accelerates the convergence of network training. At the same time, it standardizes the input of each neuron, making the subsequent layers less sensitive to changes in the input distribution. In the coal gangue degradation system, this step converts deep features with different dimensions and numerical ranges (such as spectral peak intensity, molecular structure parameters, etc.) into standardized features of a unified scale.
[0072] Inputting the standardized features into the first fully connected layer for linear transformation is the basic step of feature mapping. The fully connected layer is the basic structure in the neural network, and each neuron is connected to all neurons in the previous layer. The specific process is to multiply the input feature vector with the weight matrix and then add the bias vector to achieve a linear mapping from the input space to the output space. Each element in the weight matrix represents the contribution of the corresponding input feature to the output, and is continuously adjusted through the training process. In the pollutant degradation assessment system, the first fully connected layer maps standardized multidimensional features (such as comprehensive features from FTIR, UV-Vis and UPLC-MS / MS) to an intermediate hidden layer representation space to capture the linear relationship between features.
[0073] Applying a nonlinear activation function to the intermediate hidden layer representation is key to introducing nonlinear capabilities into the model. Common activation functions include ReLU (rectified linear unit), Sigmoid, or Tanh (hyperbolic tangent function). Taking ReLU as an example, it leaves positive inputs unchanged while mapping negative inputs to zero, offering the advantages of simple computation and alleviating the vanishing gradient problem. The introduction of nonlinear activation functions enables neural networks to learn the complex nonlinear relationship between input and output, which is crucial for accurately predicting the complex pollutant degradation process in the gangue / persulfate system. The hidden layer representation processed by the activation function contains information after the original features have been nonlinearly transformed, and has stronger expressive power.
[0074] Applying a random dropout operation to the hidden layer activation values, which controls the dropout rate, is an important technique for preventing neural networks from overfitting. The core idea of the dropout method is to randomly set the output of a portion of neurons to zero during training, simulating the effect of integrating multiple different networks. The dropout rate parameter controls the probability of a neuron being dropped and is usually set between 0.2 and 0.5. This random dropout forces the network to learn more robust feature representations, avoiding overfitting to noise or irrelevant patterns in the training data. In the prediction of pollutant degradation, the dropout method effectively reduces the model's dependence on specific experimental conditions, improves the generalization ability to new samples, and obtains a more sparse feature representation.
[0075] Passing the sparse representation through a second fully connected layer with an L2-norm penalty for feature extraction is an important means of controlling model complexity. The L2-norm penalty (also known as weight decay) penalizes large weights by adding a sum-of-squares term to the loss function, resulting in a smoother distribution of weight values. This approach encourages the model to learn smoother decision boundaries and improves generalization performance. The second fully connected layer maps the sparse representation from the previous step to a lower-dimensional feature space, achieving feature dimensionality reduction while retaining key information. The resulting reduced-dimensional feature vector is more compact and contains key information about pollutant degradation.
[0076] The association matching of reduced-dimensionality feature vectors with historical experimental data is an important step in case-based reasoning. The association matching process first establishes a distance metric in the feature space (such as Euclidean distance or cosine similarity), and then calculates the similarity between the current reduced-dimensionality feature vector and each sample in the historical database. Based on the similarity sorting, the K most similar historical samples are selected, and their actual degradation efficiency values constitute the candidate degradation efficiency value set. This similarity-based reasoning method combines historical empirical data to make the prediction results more reliable, especially for new pollutants or non-standard treatment conditions, and can use existing knowledge to make reasonable extrapolations.
[0077] Performing probability density estimation based on a set of candidate degradation efficiency values is an innovative step in achieving uncertainty quantification. Probability density estimation no longer outputs a single predicted value, but instead provides a complete probability distribution. Specific methods include kernel density estimation or Bayesian inference, which estimate the probability distribution of degradation efficiency based on the candidate set. This method outputs not only the predicted degradation efficiency value, but also the uncertainty of the prediction (such as variance or confidence interval), forming a complete pollutant degradation efficiency prediction vector. Uncertainty quantification provides a risk assessment basis for decision-making and is of great value in practical applications.
[0078] Taking the treatment of wastewater containing polycyclic aromatic hydrocarbons (PAHs) as an example, the deep feature representation (including features from FTIR, GC-MS, and fluorescence spectroscopy) was first batch normalized to adjust the data distribution to a mean of 0 and a variance of 1. The standardized feature input size was 256×128 for the first fully connected layer, resulting in a 128-dimensional intermediate hidden layer representation. Nonlinearity was then introduced through the Reluctant Unit (ReLU) activation function. A dropout rate of 0.3 was applied to the activation values to enhance model robustness. Dimensionality reduction was then performed using a second fully connected layer (size 128×64) with a weight decay of 0.001, resulting in a 64-dimensional feature vector. These feature vectors were then matched with historical PAH treatment records in a database, and the 10 most similar cases and their corresponding actual degradation rates were selected based on cosine similarity. Finally, the probability density of these 10 candidate values was estimated using a Gaussian mixture model, and the degradation efficiency prediction and 95% confidence interval of polycyclic aromatic hydrocarbons such as naphthalene, phenanthrene, and pyrene under specific coal gangue conditions (temperature 30°C, pH 4, particle size 0.3mm, and persulfate dosage 2.5mmol / L) were obtained, providing a scientific basis for the selection of treatment options.
[0079] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0080] Based on historical experimental data, a parameter search space was constructed, which included four treatment parameter dimensions: temperature, pH value, gangue particle size, and persulfate dosage, and a parameter combination matrix was obtained.
[0081] Input each set of parameters in the parameter combination matrix into the coal gangue degradation efficiency predictor to obtain the corresponding degradation efficiency prediction result set;
[0082] Based on the degradation efficiency prediction result set, a multi-objective evaluation function is calculated, and the three indicators of degradation efficiency, economic cost and processing time are comprehensively considered to obtain a comprehensive score table of the scheme;
[0083] Applying the analytic hierarchy process to optimize the weights of the comprehensive score table of the schemes to obtain the weighted performance scores of the balanced multi-dimensional evaluation indicators;
[0084] The parameter combinations are sorted and screened based on the weighted performance scores to obtain the top three candidate optimal treatment solutions;
[0085] Conduct sensitivity analysis and stability testing on the candidate optimal treatment solutions to obtain a solution detail card containing performance boundaries and reliability assessment;
[0086] The scheme detail cards are integrated into an optimal treatment scheme report card including treatment parameter configuration, expected degradation effect, economic evaluation and engineering implementation suggestions.
[0087] Specifically, constructing a parameter search space based on historical experimental data is a key step in determining the range of feasible treatment conditions. Historical experimental data includes records of previous gangue / persulfate treatment experiments, documenting the treatment effects under different parameter conditions. The parameter search space primarily considers four key treatment parameter dimensions: temperature (typically within the range of 15-45°C), pH (typically within the range of 2-10), gangue particle size (typically within the range of 0.1-2.0 mm), and persulfate dosage (typically within the range of 0.5-5.0 mmol / L). Each parameter dimension is discretized, and several representative points are selected within its reasonable range. A grid search or Latin hypercube sampling method is then used to generate all possible combinations of these parameters, forming a parameter combination matrix. This matrix is an n×4 table, where n is the total number of parameter combinations and each row represents a set of treatment conditions to be evaluated. Inputting each parameter set in the parameter combination matrix into the gangue degradation efficiency predictor is the core process for obtaining the expected treatment effect. The gangue degradation efficiency predictor, a deep self-attention neural network model trained in the previous steps, predicts pollutant degradation efficiency based on input treatment conditions. Each row in the parameter combination matrix (i.e., a set of treatment parameters) is standardized and input into the predictor to obtain the corresponding degradation efficiency prediction value. These predictions form the degradation efficiency prediction result set, which intuitively reflects the expected pollutant removal performance under different treatment conditions.
[0088] Calculating a multi-objective evaluation function based on the degradation efficiency prediction results is a key step in the comprehensive evaluation of treatment options. This multi-objective evaluation function quantitatively assesses three conflicting objectives: degradation efficiency, economic cost, and treatment time. Degradation efficiency is directly derived from the prediction results; economic cost is calculated based on the energy and reagent consumption corresponding to each parameter. For example, higher temperatures increase energy consumption, and higher persulfate dosages increase reagent costs. Treatment time is estimated using a degradation kinetic model to achieve a specific removal rate. These three indicators are calculated for each parameter set to form a comprehensive scheme score table, which contains all parameter combinations and their corresponding multi-dimensional evaluation indicators. Applying the Analytic Hierarchy Process (AHP) to optimize weights in this comprehensive scheme score table is an effective means of balancing multi-dimensional evaluation indicators. The AHP establishes a hierarchical model, compares evaluation indicators pairwise to determine their relative importance, generates a judgment matrix, and then calculates eigenvalues and eigenvectors to determine the weight coefficients for each indicator. In the field of pollutant treatment, weights for the three primary indicators (degradation efficiency, economic cost, and treatment time) are typically determined first, followed by weights for each sub-indicator. For example, different application scenarios may prioritize different approaches: emergency response scenarios prioritize processing time, while routine treatment prioritizes a balance between economic cost and degradation efficiency. Based on the determined weights, a weighted performance score is calculated for each parameter group, comprehensively reflecting the overall merits of the solution.
[0089] Sorting and screening parameter combinations based on weighted performance scores is a direct means of identifying the optimal solution. All parameter combinations are sorted from high to low according to the weighted performance scores, and the top three with the highest scores are selected as candidate optimal treatment solutions. This multi-candidate solution strategy takes into account the parameter adjustment limitations that may exist in actual applications and provides decision makers with multiple near-optimal options. Sensitivity analysis identifies the key parameters with the greatest impact by observing the degree of change in the output results with small fluctuations around the baseline parameters. The sensitivity coefficient is calculated to quantify the degree of impact of parameter changes on the results. Stability testing evaluates the performance of the solution under non-ideal conditions by simulating interference factors that may occur in actual applications, including fluctuations in influent water quality, temperature fluctuations, pH buffering capacity, and other effects. These analysis results are integrated into the solution details card, which contains the performance boundaries and reliability assessment information of each candidate solution.
[0090] Integrating the solution detail cards into the optimal treatment solution report card is the final step in translating technical achievements into practical application guidance. The optimal treatment solution report card contains four main sections: treatment parameter configuration (detailed process parameter settings), expected degradation results (removal rate and dynamic changes of target pollutants and intermediates), economic evaluation (energy consumption, reagent costs, and total treatment cost analysis), and project implementation recommendations (equipment selection, process design, and operating instructions). This report card serves as a bridge between laboratory research results and engineering applications, providing a scientific basis for actual pollution treatment.
[0091] In a specific embodiment, the process of inputting each set of parameters in the parameter combination matrix into the gangue degradation efficiency predictor may specifically include the following steps:
[0092] Perform Latin hypercube sampling on the parameter combination matrix to obtain a uniformly distributed parameter combination subset;
[0093] Normalizing each set of parameter vectors in the parameter combination subset to obtain a normalized parameter vector;
[0094] The normalized parameter vector is fused with the pollutant characteristic information to obtain the model input data packet;
[0095] Send the model input data packets to the gangue degradation efficiency predictor in batches to obtain the initial prediction results;
[0096] Calculate confidence intervals and screen outliers for the initial prediction results to obtain prediction data with credibility marks;
[0097] Monte Carlo sampling and multiple predictions are performed based on the prediction data marked with credibility to obtain the prediction distribution characteristics;
[0098] The predicted distribution features are reorganized and arranged according to the parameter combination index to obtain a complete set of degradation efficiency prediction results.
[0099] Specifically, Latin hypercube sampling is performed on the parameter combination matrix to efficiently obtain representative samples in a large parameter space. Latin hypercube sampling is a stratified sampling technique that divides each parameter dimension into N equal intervals and ensures that there is exactly one sample point in each interval for each dimension. In practice, the four parameter dimensions (temperature, pH, gangue particle size, and persulfate dosage) are divided into intervals. Then, a point is randomly selected within each interval and these points are randomly combined to form N parameter vectors. Compared to simple random sampling or grid sampling, this sampling method achieves better coverage of the parameter space with a smaller sample size, reducing the computational burden while maintaining a comprehensive exploration of the parameter space. Each parameter vector in the parameter combination subset is normalized to eliminate the influence of different parameter dimensions. Normalization typically uses min-max normalization or Z-score normalization. Min-max normalization maps parameter values to the [0, 1] interval by subtracting the minimum value from the original value and dividing the result by the difference between the maximum and minimum values. Z-score normalization transforms parameters into a distribution with a mean of 0 and a standard deviation of 1. This is calculated by subtracting the mean from the original value and dividing it by the standard deviation. For gangue degradation systems, parameters such as temperature (e.g., 20-40°C), pH (e.g., 3-9), gangue particle size (e.g., 0.1-2.0 mm), and persulfate dosage (e.g., 0.5-5.0 mmol / L) are typically standardized to the same numerical range to prevent certain parameters from dominating the model's predictions due to their large values.
[0100] Fusion of the normalized parameter vector with pollutant characteristic information is a necessary step in constructing a complete model input. Pollutant characteristic information includes molecular structure and chemical property data obtained from spectral analysis (such as FTIR, UV-Vis) and chromatographic analysis (such as UPLC-MS / MS). Fusion typically involves vector concatenation or embedding. Vector concatenation directly concatenates the normalized parameter vector and the pollutant characteristic vector into a longer vector; embedding transforms the two types of information into the same feature space using a specific mapping function. This fusion process generates a model input data packet that contains both treatment condition information and the characteristics of the pollutant itself, making the prediction results more accurate and targeted. Feeding the model input data packets into the gangue degradation efficiency predictor in batches is a core step in the prediction process. Batch processing aims to balance computational efficiency and memory usage, typically dividing the data packets into batches of 16, 32, or 64. Data from each batch is simultaneously fed into the predictor, where the corresponding degradation efficiency prediction value is obtained through forward propagation. The neural network structure inside the predictor (including graph convolution layer, self-attention mechanism and feedforward network layer) performs feature extraction and nonlinear transformation on the input data, and finally outputs the initial prediction result representing the percentage of pollutant degradation.
[0101] Calculating confidence intervals and screening for outliers on initial predictions are important means of ensuring prediction reliability. Confidence interval calculations are typically based on the statistical properties of the prediction model, such as using bootstrap or Bayesian inference methods to estimate the uncertainty of the predicted value. Bootstrap methods use multiple resampling and predictions to obtain a predicted distribution, then calculate the interval limits at a specific confidence level (e.g., 95%). Screening for outliers uses statistical tests to identify predictions that significantly deviate from a reasonable range. Commonly used methods include the Z-score method (which considers values that deviate from the mean by more than 3 standard deviations to be outliers) or the quartile-based boxplot method (which labels values that exceed the upper quartile plus 1.5 times the interquartile range or fall below the lower quartile minus 1.5 times the interquartile range as outliers). After this process, each prediction result is assigned a confidence level, indicating its reliability level.
[0102] Monte Carlo sampling and multiple predictions based on confidence-labeled prediction data are effective methods for quantifying forecast uncertainty. Monte Carlo methods repeatedly simulate the forecasting process by introducing randomness to generate a probability distribution of the forecast results. In practice, for each parameter set, a perturbation range is determined based on its confidence label. Multiple sets of perturbation parameter values are then generated within this range and fed back into the predictor to obtain multiple forecast results. This method is particularly well-suited for handling uncertainties in model parameters or inputs, enabling a more comprehensive assessment of the stability and reliability range of forecast results. From these results, statistical features such as the expected value, variance, and quantiles of the forecast can be extracted to form a forecast distribution.
[0103] The final step in forming a complete set of prediction results is to reorganize the predicted distribution characteristics according to the parameter combination index. The reorganization process first establishes a unique index for each set of original parameters and then associates the corresponding predicted distribution characteristics (such as mean, variance, 95% confidence interval, etc.) with this index. This organization allows users to easily query the prediction results and their reliability assessment for specific parameter combinations, facilitating subsequent scenario comparison and decision-making. The complete degradation efficiency prediction result set includes not only the point prediction value but also the quantification of the prediction uncertainty, providing a scientific basis for risk assessment of pollutant treatment options.
[0104] The above describes the pollutant conversion evaluation method associated with cross-scale spectral and chromatographic data in the embodiment of the present application. The following describes the pollutant conversion evaluation system associated with cross-scale spectral and chromatographic data in the embodiment of the present application. Figure 2 In one embodiment of the present application, a pollutant conversion assessment system for associating cross-scale spectral and chromatographic data includes:
[0105] The acquisition module is used to collect reaction parameters in the gangue degradation system through multi-point sensing, normalize the acquired spectral patterns and chromatographic peaks, and obtain a pollutant characteristic data matrix;
[0106] A mapping module is used to fuse and map data of different scales using multi-layer correlation tensor decomposition based on the pollutant characteristic data matrix to obtain a pollutant structure-activity relationship map;
[0107] A construction module is used to construct a correspondence rule between molecular changes and treatment conditions based on the pollutant structure-activity relationship map through a deep self-attention neural network to generate a coal gangue degradation efficiency predictor;
[0108] The simulation module is used to simulate and evaluate multiple groups of treatment parameters based on the coal gangue degradation efficiency predictor to form an optimal treatment solution report card.
[0109] Through the collaborative efforts of these components, multi-point sensing is used to collect reaction parameters in the gangue degradation system and normalize spectral profiles and chromatographic peaks, achieving unified representation of multi-source data. This effectively addresses the difficulty in directly comparing data obtained by different measurement methods and lays a standardized foundation for subsequent analysis. Multi-layer correlation tensor decomposition is used to fuse and map data at different scales, overcoming the limitations of traditional methods in independent data analysis. This establishes intrinsic connections between data across scales and forms a pollutant structure-activity relationship map containing rich structural information. A deep self-attention neural network constructs rules for mapping molecular changes to treatment conditions. The innovative introduction of graph-structured data processing and attention mechanisms enables the model to automatically identify key active sites in molecular structures and reaction patterns under different treatment conditions, improving prediction accuracy and interpretability. Simulations and evaluations of multiple treatment parameter sets based on the gangue degradation efficiency predictor significantly reduce the number of experiments and resource consumption. A multi-objective comprehensive evaluation more comprehensively considers multidimensional indicators such as degradation efficiency, economic cost, and treatment time. The resulting optimal treatment solution report card provides intuitive and clear guidance for practical applications. This method excels in analyzing degradation mechanisms and optimizing treatment conditions, particularly for processing complex environmental samples and rapidly screening treatment conditions. From the perspective of artificial intelligence algorithm application, this approach innovatively combines tensor decomposition technology with deep neural networks, leveraging the advantages of tensor decomposition for high-dimensional data reduction and feature extraction, as well as the ability of deep self-attention neural networks to capture complex nonlinear relationships. The multi-layer correlation tensor decomposition algorithm offers unique advantages for revealing potential correlations between heterogeneous data from multiple sources, while the self-attention mechanism accurately captures structure-activity relationships by dynamically weighting different molecular structural sites. These algorithmic features contribute to this approach by: first, enabling intelligent processing throughout the entire process from data acquisition to output; second, providing data-driven mechanism analysis capabilities; and third, possessing predictive generalization capabilities for emerging pollutants. Overall, this approach effectively integrates multi-source data through algorithmic innovation, significantly improving the evaluation efficiency and optimization accuracy of the gangue / persulfate treatment system, providing a more scientific and efficient technical solution for environmental pollutant remediation.
[0110] above Figure 2 The pollutant conversion evaluation system associated with mid-scale spectrum and chromatographic data in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The pollutant conversion evaluation device associated with cross-scale spectrum and chromatographic data in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0111] Figure 3: This is a schematic structural diagram of a pollutant conversion assessment device associated with cross-scale spectral and chromatographic data provided by an embodiment of the present invention. The pollutant conversion assessment device 300 associated with cross-scale spectral and chromatographic data may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more massive storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be temporary storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), each module may include a series of instruction operations in the pollutant conversion assessment device 300 associated with cross-scale spectral and chromatographic data. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the pollutant conversion assessment device 300 associated with cross-scale spectral and chromatographic data to implement the steps of the pollutant conversion assessment method associated with cross-scale spectral and chromatographic data.
[0112] The pollutant conversion assessment device 300 for correlating cross-scale spectral and chromatographic data may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the pollutant conversion evaluation device associated with cross-scale spectral and chromatographic data shown does not constitute a limitation of the pollutant conversion evaluation device associated with cross-scale spectral and chromatographic data provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0113] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of a pollutant conversion assessment method that associates cross-scale spectral and chromatographic data.
[0114] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a pollutant conversion assessment device (which can be a personal computer, server, or network device, etc.) that associates cross-scale spectral and chromatographic data to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.
[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pollutant transformation assessment method based on cross-scale spectral and chromatographic data correlation, characterized in that: The method comprises: The reaction parameters in the gangue degradation system are collected through multi-point sensing, and the obtained spectra and chromatographic peaks are normalized to obtain the pollutant characteristic data matrix; According to the pollutant characteristic data matrix, multi-layer correlation tensor decomposition is used to fuse and map data of different scales to obtain a pollutant structure-activity relationship map; Based on the pollutant structure-activity relationship map, a deep self-attention neural network is used to construct the correspondence rules between molecular changes and treatment conditions to generate a coal gangue degradation efficiency predictor; Based on the gangue degradation efficiency predictor, multiple groups of treatment parameters are simulated and evaluated to form an optimal treatment solution report card.
2. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 1, characterized in that: The reaction parameters in the gangue degradation system are collected by multi-point sensing, and the obtained spectral maps and chromatographic peaks are normalized to obtain a pollutant characteristic data matrix, including: The adsorption temperature, adsorption time, waste concentration, pH value and gangue particle size in the gangue degradation system were monitored in real time through multiple channels to obtain the original environmental parameter set; The FTIR spectrum data was subjected to wavelet transform denoising and baseline correction to obtain the infrared characteristic fingerprint region signal; The UV-Vis and fluorescence three-dimensional excitation-emission matrix data were subjected to scattered light elimination and inner filter effect correction to obtain purified spectral signals. Perform peak identification and peak area integration calculation on UPLC-MS / MS chromatograms to obtain quantitative retention time-response intensity pairs; Based on synchronized timestamps, the multimodal monitoring data are aligned in time dimensions and missing values are interpolated to obtain a complete time series dataset. Apply the minimum-maximum normalization algorithm to the complete time series data set to obtain a multi-source data matrix with unified dimensions; The multi-source data matrix is reorganized according to the time-feature-sample three-dimensional structure to obtain the pollutant feature data matrix.
3. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 1, characterized in that: The method comprises: performing fusion mapping on data of different scales by using multi-layer correlation tensor decomposition based on the pollutant characteristic data matrix to obtain a pollutant structure-activity relationship map, including: Applying tensor low-rank representation and nuclear norm minimization to the pollutant characteristic data matrix, a refined data tensor is obtained that eliminates redundancy and noise. The refined data tensor is input into the improved CANDECOMP / PARAFAC decomposition algorithm for three-dimensional decomposition operation to obtain a core factor matrix group containing time dimension, feature dimension and sample dimension; Apply L1 / L2 mixed norm regularization and non-negative constraint optimization to the core factor matrix group to obtain a set of characteristic component expressions with physical meaning; Based on the characteristic component expression set, the Kendall correlation coefficient matrix of cross-spectral modalities and cross-chromatographic modalities was calculated to obtain the characteristic dependency network reflecting the correlation between multiple monitoring methods; Applying hierarchical clustering and edge weight threshold filtering to the feature dependency network, we obtain highly correlated feature clusters that reveal the characteristics of pollutants. Cross-mapping the highly correlated feature clusters with the molecular vibrational mode database calculated based on density functional theory to obtain a spectral feature-molecular structural unit correspondence table; Based on the spectral feature-molecular structure unit correspondence table, a weighted bidirectional correlation graph structure containing active site markers was constructed to obtain a pollutant structure-activity relationship map describing the degradation mechanism.
4. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 1, characterized in that: Based on the pollutant structure-activity relationship map, a correspondence rule between molecular changes and treatment conditions is constructed through a deep self-attention neural network to generate a coal gangue degradation efficiency predictor, including: Converting the pollutant structure-activity relationship map into a node feature matrix and a topological structure adjacency matrix to obtain standardized input graph data; Performing spatial feature extraction on the standardized input graph data through a multi-layer graph convolutional network to obtain a structure-aware node representation vector; Inputting the node representation vector into a self-attention calculation unit composed of multiple attention heads to obtain a weighted context feature representation; The features are weighted based on the dot product attention calculation method and fused with the reaction condition parameters for conditional encoding to obtain a conditionally perceived attention feature map, wherein the reaction condition parameters include temperature, pH value, particle size, and dosage; Passing the conditionally perceived attention feature map through a multi-layer residual connection module to obtain a deep feature representation; Performing dimensionality reduction mapping on the deep feature representation through a multi-layer feedforward neural network with a regularization mechanism to obtain a pollutant degradation efficiency prediction vector; Error calculation and parameter optimization adjustment are performed on the pollutant degradation efficiency prediction vector and experimental label data to obtain a coal gangue degradation efficiency predictor with reasoning ability.
5. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 4, characterized in that: The deep feature representation is subjected to dimensionality reduction mapping through a multi-layer feedforward neural network with a regularization mechanism to obtain a pollutant degradation efficiency prediction vector, including: Performing batch normalization on the deep feature representation to obtain standardized features with a mean of zero and a variance of one; Input the standardized features into the first fully connected layer for linear transformation to obtain an intermediate hidden layer representation; Applying a nonlinear activation function to the intermediate hidden layer representation to obtain a hidden layer activation value having nonlinear characteristics; Applying a random dropout operation with a dropout rate control to the hidden layer activation values to obtain a sparse representation that prevents overfitting; The sparse representation is passed through a second fully connected layer with an L2 norm penalty term to perform feature extraction to obtain a reduced-dimensional feature vector; Associating and matching the dimension-reduced feature vector with historical experimental data to obtain a set of candidate degradation efficiency values; Probability density estimation is performed based on the candidate degradation efficiency value set to obtain a pollutant degradation efficiency prediction vector including the predicted value and its uncertainty quantification.
6. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 1, characterized in that: The gangue degradation efficiency predictor is used to simulate and evaluate multiple sets of treatment parameters to form an optimal treatment solution report card, including: Based on historical experimental data, a parameter search space was constructed, which included four treatment parameter dimensions: temperature, pH value, gangue particle size, and persulfate dosage, and a parameter combination matrix was obtained. Inputting each set of parameters in the parameter combination matrix into the gangue degradation efficiency predictor to obtain a corresponding degradation efficiency prediction result set; Performing a multi-objective evaluation function calculation based on the degradation efficiency prediction result set, comprehensively considering the three indicators of degradation efficiency, economic cost and processing time, and obtaining a comprehensive score table of the scheme; Applying the analytic hierarchy process to the comprehensive score table of the scheme to optimize the weights and obtain the weighted performance score of the balanced multi-dimensional evaluation index; Sorting and screening the parameter combinations based on the weighted performance scores to obtain the top three candidate optimal processing solutions; Conduct sensitivity analysis and stability testing on the candidate optimal treatment solutions to obtain a solution detail card containing performance boundaries and reliability evaluation; The solution detail cards are integrated into an optimal treatment solution report card including treatment parameter configuration, expected degradation effect, economic evaluation and engineering implementation suggestions.
7. The pollutant conversion assessment method based on cross-scale spectral and chromatographic data correlation according to claim 6, characterized in that: Each set of parameters in the parameter combination matrix is input into the gangue degradation efficiency predictor to obtain a corresponding degradation efficiency prediction result set, including: Performing Latin hypercube sampling on the parameter combination matrix to obtain a uniformly distributed parameter combination subset; Normalizing each set of parameter vectors in the parameter combination subset to obtain a normalized parameter vector; fusing the normalized parameter vector with pollutant characteristic information to obtain a model input data packet; Sending the model input data packets into the gangue degradation efficiency predictor in batches to obtain initial prediction results; Calculating confidence intervals and screening outliers on the initial prediction results to obtain prediction data with credibility marks; Performing Monte Carlo sampling and multiple predictions based on the predicted data of the credibility mark to obtain a prediction distribution feature; The predicted distribution features are reorganized and arranged according to the parameter combination index to obtain a complete degradation efficiency prediction result set.
8. A pollutant conversion assessment system that associates cross-scale spectral and chromatographic data, characterized in that: A pollutant conversion assessment method for implementing cross-scale spectral and chromatographic data correlation according to any one of claims 1 to 7, wherein the pollutant conversion assessment system for cross-scale spectral and chromatographic data correlation comprises: The acquisition module is used to collect reaction parameters in the gangue degradation system through multi-point sensing, normalize the acquired spectral patterns and chromatographic peaks, and obtain a pollutant characteristic data matrix; A mapping module is used to fuse and map data of different scales using multi-layer correlation tensor decomposition based on the pollutant characteristic data matrix to obtain a pollutant structure-activity relationship map; A construction module is used to construct a correspondence rule between molecular changes and treatment conditions based on the pollutant structure-activity relationship map through a deep self-attention neural network to generate a coal gangue degradation efficiency predictor; The simulation module is used to simulate and evaluate multiple groups of treatment parameters based on the coal gangue degradation efficiency predictor to form an optimal treatment solution report card.
9. A pollutant conversion assessment device that associates cross-scale spectral and chromatographic data, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the pollutant conversion assessment method based on the association of cross-scale spectral and chromatographic data according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the pollutant conversion assessment method based on correlation of cross-scale spectral and chromatographic data according to any one of claims 1 to 7.
Citation Information
Cited By
Propolis component intelligent identification and traceability system and method
CN120948407A
Pipe network sludge component detection method, device, medium and equipment
CN121558669A