Aquaculture tail water pollution load prediction method and system

By generating high-order exponential features through multimodal sensing data and feature fusion algorithms, and combining them with wastewater treatment process parameters, the accuracy and adaptability issues of aquaculture wastewater pollution load prediction are solved, achieving high-precision and long-term pollution load prediction and supporting intelligent control of wastewater treatment.

CN121809773APending Publication Date: 2026-04-07ZHONGKAI UNIV OF AGRI & ENG +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for predicting pollution loads from aquaculture wastewater rely on data from a single physicochemical sensor, which cannot comprehensively characterize the water quality status. They also suffer from insufficient data fusion, poor adaptability of prediction models, and difficulty in achieving high-precision, rapid-response, and long-term pollution load predictions.

Method used

By acquiring physical, visual, and spectral multimodal sensor data, high-order exponential features are generated using feature extraction and fusion algorithms. Combined with effluent treatment process parameters, dynamic time-series features and static auxiliary features are constructed, and a pollution load index prediction model is used for multi-step prediction.

Benefits of technology

It achieves high-precision, rapid-response, and long-term pollution load prediction, improves the model's adaptability and prediction accuracy, and provides a basis for precise regulation of effluent treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809773A_ABST
    Figure CN121809773A_ABST
Patent Text Reader

Abstract

The invention discloses an aquaculture tail water pollution load prediction method and system. The method comprises the following steps: acquiring sensing data of multiple modes of a tail water acquisition point; based on a preset feature extraction algorithm, a sub-modal feature corresponding to each piece of sensing data is extracted, and according to the sub-modal features and a preset feature fusion enhancement algorithm, high-order index features of tail water pollution are generated; generating dynamic time sequence characteristics from the high-order index characteristics based on time sequence processing, and determining static auxiliary characteristics and time sequence auxiliary characteristics of the dynamic time sequence characteristics according to preset tail water treatment process parameters and sensing data acquisition time; and according to the dynamic time sequence characteristics, the static auxiliary characteristics, the time sequence auxiliary characteristics and a preset pollution load index prediction model, outputting a prediction result corresponding to the pollution load index in the future time period. Therefore, high-precision, quick-response and long-time prediction urgently needed by the aquaculture tail water pollution load can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wastewater treatment technology, and in particular to a method and system for predicting pollution load in aquaculture wastewater. Background Technology

[0002] With the rapid development of intensive and factory-style aquaculture, the efficient treatment and compliant discharge of aquaculture wastewater have become core requirements for the industry's sustainable development. Aquaculture wastewater is characterized by complex pollutant composition, drastic fluctuations in water quality and quantity, and strong spatiotemporal heterogeneity. The dynamic changes in its pollution load (such as the concentrations of key indicators like chemical oxygen demand, ammonia nitrogen, and total phosphorus) directly affect the operating efficiency and control effect of the treatment process. Accurately predicting the trend of pollution load changes over a future period is crucial for achieving proactive control of the treatment system, optimizing reagent and energy consumption, and ensuring stable compliance of effluent water quality.

[0003] However, current technologies for predicting pollution loads from aquaculture wastewater have significant shortcomings, which are reflected in the following aspects: (1) Existing prediction methods mostly rely on data from single physicochemical sensors (such as pH, DO, ammonia nitrogen, etc.), which are single and isolated in terms of data dimension and cannot fully depict the microscopic characteristics and evolution mechanism of water quality. For example, the instantaneous value of ammonia nitrogen alone cannot distinguish between natural fluctuations and sudden impacts of pollution load, and lacks the capture of key information such as the morphology of suspended particulate matter and the composition of dissolved organic matter, resulting in one-sided input features of the prediction model and difficulty in reflecting the true change law of pollution load; (2) Although some methods integrate multiple types of sensing devices, the utilization of heterogeneous data composed of physicochemical data, image data and spectral data is only at the level of simple parallel or independent processing, without deep integration. The potential spatiotemporal correlation and causal logic between data are not fully explored and utilized, and it is impossible to generate high-order features that comprehensively reflect the state of pollution load, resulting in poor adaptability of the model to complex water quality changes. (3) The performance of the prediction model is limited and it is difficult to meet the actual needs: Existing prediction methods mostly use traditional time series models or simple machine learning algorithms, which lack effective capture of the dependence of long time series data and cannot take into account the combined effects of dynamic water quality characteristics and static process constraints. This results in low prediction accuracy, slow response speed and limited prediction time, making it difficult to achieve fine load prediction in the next few hours and unable to provide reliable data support for the precise control of tailwater pollution treatment.

[0004] Therefore, the core problems with existing technologies for wastewater treatment based on load forecasting lie in insufficient data utilization, inadequate feature fusion, and poor adaptability of the prediction model. This results in pollution load forecasts failing to meet the practical application requirements for accuracy, speed, and foresight. Clearly, existing technologies have shortcomings that urgently need to be addressed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for predicting pollution load in aquaculture wastewater, which can achieve the high-precision, rapid-response and long-term prediction of pollution load in aquaculture wastewater that is urgently needed.

[0006] To address the aforementioned technical problems, the first aspect of this invention discloses a method for predicting pollution load in aquaculture wastewater, the method comprising: Multiple modal sensor data from tailwater collection points are acquired, and the modal types of the sensor data include physical modal, visual modal, and spectral modal. Based on a preset feature extraction algorithm, the submodal features corresponding to each of the sensor data are extracted. Based on the submodal features and the preset feature fusion enhancement algorithm, a high-order index feature of wastewater pollution is generated. Based on the temporal arrangement, the higher-order exponential features are used to generate dynamic temporal features. According to the preset tailwater treatment process parameters and sensor data acquisition time, the static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined. Based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and the preset pollution load index prediction model, the prediction results corresponding to the pollution load index in the future time period are output.

[0007] As an optional implementation, in the first aspect of the present invention, among the multiple modal sensing data, the physicochemical modal sensing data is obtained by collecting water quality parameters corresponding to dissolved oxygen sensor, pH sensor, redox potential sensor, ammonia nitrogen sensor, turbidity sensor and conductivity sensor. The sensor data of the visual modality is collected by a microscope / flow cytometer to collect the visual parameters of suspended matter in the tailwater, wherein the visual parameters include the size, shape and quantity distribution of suspended matter. The spectral mode sensing data is obtained by collecting the spectral parameters corresponding to the components of the tailwater using an ultraviolet-visible spectral probe / fluorescence spectrometer. The spectral parameters include the characteristics of dissolved organic matter and the concentration of target pollutants.

[0008] As an optional implementation, in the first aspect of the present invention, before extracting the submodal features corresponding to each of the sensing data based on a preset feature extraction algorithm, the following steps are included: The sampling frequency of the sensor data is preset, and the sampling frequency is synchronized for sensors of any modal type; The sensing data is preprocessed, including moving average filtering / wavelet denoising to eliminate sensor noise and environmental interference, identifying and removing outliers in the sensing data based on statistical methods / rule bases, and normalizing the original sensing data to the same interval to eliminate differences in dimensions and / or ranges. Timestamp alignment is performed on sensor data with different sampling frequencies, and missing values ​​are filled in using linear interpolation / GAN.

[0009] As an optional implementation, in the first aspect of the present invention, based on a preset feature extraction algorithm, the modal features corresponding to each of the sensing data are extracted, including: For the physicochemical submodal sensing data, based on statistical feature extraction and trend feature analysis, the statistical features and variation features of the basic parameters corresponding to water quality are extracted. The physicochemical submodal features are determined according to the statistical features and variation features. The statistical features of the basic parameters include the mean and variance of the water quality parameters corresponding to the physicochemical submodal sensing data. The variation features of the basic parameters include the slope of the curve of the change trend calculated by linear regression. For visual modal sensing data, morphological features of suspended particles are extracted based on a lightweight neural network for image processing, and texture features of suspended particles are extracted based on a gray-level co-occurrence matrix. Visual modal features are determined based on the morphological and texture features, wherein the morphological features include the average particle size, density, and number concentration of suspended particles, and the texture features include energy and entropy. For spectral modal sensing data, based on spectral feature engineering and peak position identification analysis, UV-Vis spectral features and fluorescence spectral features are extracted. Spectral modal features are determined based on the UV-Vis spectral features and fluorescence spectral features. The UV-Vis spectral features include the absorbance ratio between 254 nm and 436 nm wavelengths, the absorbance at 350 nm wavelength, and the full width at half maximum (FWHM) of the characteristic peaks. The fluorescence spectral features include the fluorescence index and humification index calculated from protein-like fluorescence peaks and humic substance-like fluorescence peaks.

[0010] As an optional implementation, in the first aspect of the present invention, based on the modal features and a preset feature fusion enhancement algorithm, a high-order index feature of wastewater pollution is generated, including: All submodal features are concatenated based on the target dimension to generate an initial fusion vector; Principal component analysis with domain knowledge constraints is used to reduce the dimensionality of the initial fusion vector and retain the principal components with a cumulative variance contribution rate greater than 95%. The domain knowledge is the knowledge of wastewater treatment. The corresponding constraint is that the weight of the pollution load index contained in the first principal component in the principal component analysis is greater than the preset constraint value. The pollution load index includes chemical oxygen demand concentration, ammonia nitrogen concentration and total phosphorus concentration. The initial fusion vector after dimensionality reduction is enhanced with features based on an autoencoder, and the domain knowledge is fused with the enhanced initial fusion vector to generate high-order exponential features of wastewater pollution. The autoencoder includes an input layer, an output layer, a bottleneck layer, and two hidden layers. The input of the input layer is the fusion vector after dimensionality reduction. The output layer is set with reconstruction error constraints. The bottleneck layer is used to generate the higher-order exponential features. The hidden layers include the ReLU activation function. The higher-order index features include the comprehensive pollution load index, the sludge microbial activity index, and the characteristic pollutant fingerprint; The comprehensive pollution load index is expressed as: ; in, The comprehensive pollution load index is used to characterize the pollution load level. This represents the standardized chemical oxygen demand. This indicates the maximum permissible concentration of chemical oxygen demand. This represents the standardized ammonia nitrogen concentration. This indicates the maximum permissible concentration of ammonia nitrogen. This represents the absorbance ratio between the 254nm and 436nm wavelengths. This represents the maximum reference value for the absorbance ratio. This represents the standardized number concentration of suspended solids. The weighting of chemical oxygen demand. The weighting of ammonia nitrogen concentration. The weighting of absorbance ratio is indicated. The weight representing the number concentration of suspended solids, and , , and satisfy ; The sludge microbial activity index is expressed as: ; in, The sludge microbial activity index is used to characterize the combined dissolved oxygen trend, sludge floc structure, and microbial metabolite characteristics, thereby determining the activity status of the microorganisms. This represents the slope of the linear regression of the standardized dissolved oxygen time series data. This represents the sludge floc density extracted from sensor data of visual modalities. This represents the standardized microbial protein characteristic values; The characteristic pollutant fingerprint is represented as follows: ; in, The fingerprint represents a characteristic pollutant, specifically a three-dimensional vector composed of spectral modal features. This indicates the molecular weight of organic compounds as characterized by the absorbance ratio between 254 nm and 436 nm. The fluorescence index is used to determine the source of organic matter, and the humification index is used to determine the degree of decay of organic matter. Each type of pollutant corresponds one-to-one with the three-dimensional vector of the FPF.

[0011] As an optional implementation, in the first aspect of the present invention, dynamic temporal features are generated from the higher-order exponential features based on temporal arrangement. Static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined according to preset wastewater treatment process parameters and sensor data acquisition time, including: The higher-order exponential features are arranged in time series using a preset historical time step to generate dynamic time series features that capture the dynamic change patterns of water quality and their association with multimodal sensor data. The preset effluent treatment process parameters are normalized to generate static auxiliary features that provide constraints for the process scenario. The effluent treatment process parameters include the volume of the biological treatment tank, the area of ​​the aeration zone, the residence time of the coagulation tank, the membrane pore size, and the sensor deployment node type. The sensor data acquisition time is processed by 1-hot encoding and normalization to generate a time-series auxiliary feature of the pollution load time periodic pattern. The dimensions of the sensor data acquisition time include hours, minutes, aquaculture status and weather status. The dynamic time-series features, static auxiliary features, and time-series auxiliary features are used as inputs to the pollution load index prediction model in tensor form, which is expressed as follows: ; in, Indicates input, Represents the set of real numbers. Indicates batch size. This indicates the historical time step. Dimensions representing dynamic temporal characteristics The dimension representing the static auxiliary feature. The dimension representing the temporal auxiliary features.

[0012] As an optional implementation, in the first aspect of the present invention, the preset pollution load index prediction model includes: The input embedding layer is used to map different types of dynamic temporal features, static auxiliary features, and temporal auxiliary features in the tensor form of input data to the same dimensional space; A variable selection network layer is used to filter the target information of the spliced ​​features output by the input embedding layer, and to compress the feature dimension to adapt to the position encoding layer and the temporal attention layer, wherein the target information is used to enhance the contribution of higher-order exponential features to the prediction results. The location encoding layer is used to inject time step location information into the output of the variable selection network layer, so that the prediction model can perceive the degree of influence of time sequence on historical data; The temporal attention layer is used to capture long-range dependencies in historical temporal data output by the positional encoding layer through a self-attention mechanism, so as to identify the correlation strength between the features of the target time point and the predicted target result. The gated fusion layer is used to dynamically balance the contribution ratio between the temporal attention features output by the temporal attention layer and the static auxiliary features broadcast within the model, and to capture the interaction dependency between the temporal attention features and the static auxiliary features. Then, based on the contribution ratio and interaction dependency, the fusion weight is determined and the fusion feature is output. A feedforward network layer is used to extract the implicit cross-modal correlations within the fusion features output by the gated fusion layer, so as to enhance the discriminative expression of the fusion features for higher-order exponential features in the prediction task. The decoder layer is used to generate a continuous change prediction curve of the pollution load index in future time periods from the output of the feedforward network layer by time step, wherein the time step process includes using the prediction result of the previous time step as the input of the current time step. The output prediction layer is used to convert the high-dimensional time-series features of the decoding layer output, which contain continuous concentration change curves, into prediction results corresponding to pollution load indicators. The prediction results are the concentration values ​​of pollution load indicators at any time in the future period.

[0013] As an optional implementation, in the first aspect of the present invention, based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and a preset pollution load index prediction model, the predicted results corresponding to the pollution load index in the future time period are output, including: The dynamic temporal features, static auxiliary features, and temporal auxiliary features are mapped to a high-dimensional space of the same dimension through corresponding embedding layers. After the static auxiliary features are broadcast to the temporal dimension, the features of the same dimension are concatenated to generate concatenated features. Each dimension of the spliced ​​features is assigned an importance weight and the features are sorted by feature priority. The weighted spliced ​​features are then subjected to linear dimensionality reduction and / or non-linear enhancement through projection to generate a standardized feature vector. The standardized feature vector is injected with time step position information using sine-cosine position coding, wherein the sine-cosine position coding is calculated as follows: ; ; The encoding dimension is the same as the standardized feature vector dimension. This represents the location code, and pos represents the time step index. This represents the encoding dimension, where i represents the dimension index; The standardized feature vector of position encoding is computed in parallel through a multi-head attention mechanism to generate multiple sets of local attention features corresponding to the number of attention heads. Then, multiple sets of local features are output through linear projection and scaled dot product attention calculation. The multiple sets of local features are integrated and restored according to the dimension of the standardized feature vector. Temporal attention features are output through residual connection and normalization processing. The static auxiliary features broadcast are averaged and pooled to obtain static global features. The input temporal attention features are processed by a gating unit with the static global features as the initial state, retaining the temporal output and the last step output of the gating unit. The last step output is input into a fully connected layer network and an activation function to determine the gating weights. After broadcasting the static global features to the target time step dimension, the temporal output and the adjusted dimension static global features are weighted, fused, and normalized based on the gating weights to output the fused features. The fused features are expanded in dimension by a first fully connected network to capture the implicit correlation features between multimodal features in the fused features, and then passed through a second fully connected network to retain the effective correlation features in the implicit correlation features. The fused features are then enhanced by using the effective correlation features through residual connections and normalization. Initialize the target time step based on the target future time period corresponding to the prediction result. Determine the initial predicted value of the fusion feature through a fully connected layer based on the last step feature. Extract the single-step feature fragments related to the current time step from the initial predicted value and the fusion feature and concatenate them into a single-step fusion feature. Input the single-step fusion feature into the autoregressive decoding unit to output the single-step prediction feature. Input the single-step prediction feature and the fusion feature into the cross-attention unit and calculate the single-step enhanced feature through scaling dot product attention. Map the single-step enhanced feature through a fully connected layer to determine the current concentration value. Update the predicted value of the previous time step to the current concentration value until the current time step meets the target time step. Concatenate all the single-step enhanced features to generate a high-dimensional temporal feature. The high-dimensional time-series features are reduced in dimensionality based on the feature projection layer, and the effective features of the associated predicted values ​​are enhanced by the activation function to output low-dimensional time-series features. The low-dimensional time-series features are then regressed using a linear activation function through a concentration prediction head, mapping the low-dimensional features to continuous predicted values ​​of pollution load index concentrations. The continuous predicted values ​​of pollution load index concentrations are then restored to true concentration values ​​through inverse normalization, and the true concentration values ​​of the pollution load indexes are output as the prediction results. The pollution load indexes include chemical oxygen demand concentration, ammonia nitrogen concentration, and total phosphorus concentration, and the future time period corresponds to the target time step of the pollution load indexes.

[0014] A second aspect of this invention discloses an aquaculture wastewater pollution load prediction system, the system comprising: The data acquisition module is configured to acquire sensor data of multiple modes from the tailwater collection point. The sub-modal types of the sensor data include physical sub-modal, visual sub-modal, and spectral sub-modal. The data fusion module is configured to extract the submodal features corresponding to each of the sensor data based on a preset feature extraction algorithm, and generate high-order index features of wastewater pollution based on the submodal features and the preset feature fusion enhancement algorithm. The data processing module is configured to generate dynamic time-series features from the higher-order exponential features based on the time-series arrangement, and to determine the static auxiliary features and time-series auxiliary features of the dynamic time-series features according to the preset tailwater treatment process parameters and the sensor data acquisition time. The data prediction module is configured to output the prediction results of the pollution load index in the future period based on the dynamic time series characteristics, static auxiliary characteristics, time series auxiliary characteristics and the preset pollution load index prediction model.

[0015] As an optional implementation, in the second aspect of the present invention, the sensing data of the multiple modes, the sensing data of the physicochemical mode are obtained by collecting water quality parameters corresponding to the dissolved oxygen sensor, pH sensor, redox potential sensor, ammonia nitrogen sensor, turbidity sensor and conductivity sensor. The sensor data of the visual modality is collected by a microscope / flow cytometer to collect the visual parameters of suspended matter in the tailwater, wherein the visual parameters include the size, shape and quantity distribution of suspended matter. The spectral mode sensing data is obtained by collecting the spectral parameters corresponding to the components of the tailwater using an ultraviolet-visible spectral probe / fluorescence spectrometer. The spectral parameters include the characteristics of dissolved organic matter and the concentration of target pollutants.

[0016] As an optional implementation, in the second aspect of the present invention, before extracting the submodal features corresponding to each of the sensing data based on a preset feature extraction algorithm, the following steps are included: The sampling frequency of the sensor data is preset, and the sampling frequency is synchronized for sensors of any modal type; The sensing data is preprocessed, including moving average filtering / wavelet denoising to eliminate sensor noise and environmental interference, identifying and removing outliers in the sensing data based on statistical methods / rule bases, and normalizing the original sensing data to the same interval to eliminate differences in dimensions and / or ranges. Timestamp alignment is performed on sensor data with different sampling frequencies, and missing values ​​are filled in using linear interpolation / GAN.

[0017] As an optional implementation, in the second aspect of the present invention, based on a preset feature extraction algorithm, the modal features corresponding to each of the sensing data are extracted, including: For the physicochemical submodal sensing data, based on statistical feature extraction and trend feature analysis, the statistical features and variation features of the basic parameters corresponding to water quality are extracted. The physicochemical submodal features are determined according to the statistical features and variation features. The statistical features of the basic parameters include the mean and variance of the water quality parameters corresponding to the physicochemical submodal sensing data. The variation features of the basic parameters include the slope of the curve of the change trend calculated by linear regression. For visual modal sensing data, morphological features of suspended particles are extracted based on a lightweight neural network for image processing, and texture features of suspended particles are extracted based on a gray-level co-occurrence matrix. Visual modal features are determined based on the morphological and texture features, wherein the morphological features include the average particle size, density, and number concentration of suspended particles, and the texture features include energy and entropy. For spectral modal sensing data, based on spectral feature engineering and peak position identification analysis, UV-Vis spectral features and fluorescence spectral features are extracted. Spectral modal features are determined based on the UV-Vis spectral features and fluorescence spectral features. The UV-Vis spectral features include the absorbance ratio between 254 nm and 436 nm wavelengths, the absorbance at 350 nm wavelength, and the full width at half maximum (FWHM) of the characteristic peaks. The fluorescence spectral features include the fluorescence index and humification index calculated from protein-like fluorescence peaks and humic substance-like fluorescence peaks.

[0018] As an optional implementation, in a second aspect of the invention, based on the modal features and a preset feature fusion enhancement algorithm, a high-order index feature of wastewater pollution is generated, including: All submodal features are concatenated based on the target dimension to generate an initial fusion vector; Principal component analysis with domain knowledge constraints is used to reduce the dimensionality of the initial fusion vector and retain the principal components with a cumulative variance contribution rate greater than 95%. The domain knowledge is the knowledge of wastewater treatment. The corresponding constraint is that the weight of the pollution load index contained in the first principal component in the principal component analysis is greater than the preset constraint value. The pollution load index includes chemical oxygen demand concentration, ammonia nitrogen concentration and total phosphorus concentration. The initial fusion vector after dimensionality reduction is enhanced with features based on an autoencoder, and the domain knowledge is fused with the enhanced initial fusion vector to generate high-order exponential features of wastewater pollution. The autoencoder includes an input layer, an output layer, a bottleneck layer, and two hidden layers. The input of the input layer is the fusion vector after dimensionality reduction. The output layer is set with reconstruction error constraints. The bottleneck layer is used to generate the higher-order exponential features. The hidden layers include the ReLU activation function. The higher-order index features include the comprehensive pollution load index, the sludge microbial activity index, and the characteristic pollutant fingerprint; The comprehensive pollution load index is expressed as: ; in, The comprehensive pollution load index is used to characterize the pollution load level. This represents the standardized chemical oxygen demand. This indicates the maximum permissible concentration of chemical oxygen demand. This represents the standardized ammonia nitrogen concentration. This indicates the maximum permissible concentration of ammonia nitrogen. This represents the absorbance ratio between the 254nm and 436nm wavelengths. This represents the maximum reference value for the absorbance ratio. This represents the standardized number concentration of suspended solids. The weighting of chemical oxygen demand. The weighting of ammonia nitrogen concentration. The weighting of absorbance ratio is indicated. The weight representing the number concentration of suspended solids, and , , and satisfy ; The sludge microbial activity index is expressed as: ; in, The sludge microbial activity index is used to characterize the combined dissolved oxygen trend, sludge floc structure, and microbial metabolite characteristics, thereby determining the activity status of the microorganisms. This represents the slope of the linear regression of the standardized dissolved oxygen time series data. This represents the sludge floc density extracted from sensor data of visual modalities. This represents the standardized microbial protein characteristic values; The characteristic pollutant fingerprint is represented as follows: ; in, The fingerprint represents a characteristic pollutant, specifically a three-dimensional vector composed of spectral modal features. This indicates the molecular weight of organic compounds as characterized by the absorbance ratio between 254 nm and 436 nm. The fluorescence index is used to determine the source of organic matter, and the humification index is used to determine the degree of decay of organic matter. Each type of pollutant corresponds one-to-one with the three-dimensional vector of the FPF.

[0019] As an optional implementation, in a second aspect of the invention, the higher-order exponential features are used to generate dynamic temporal features based on temporal arrangement. Static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined according to preset wastewater treatment process parameters and sensor data acquisition time, including: The higher-order exponential features are arranged in time series using a preset historical time step to generate dynamic time series features that capture the dynamic change patterns of water quality and their association with multimodal sensor data. The preset effluent treatment process parameters are normalized to generate static auxiliary features that provide constraints for the process scenario. The effluent treatment process parameters include the volume of the biological treatment tank, the area of ​​the aeration zone, the residence time of the coagulation tank, the membrane pore size, and the sensor deployment node type. The sensor data acquisition time is processed by 1-hot encoding and normalization to generate a time-series auxiliary feature of the pollution load time periodic pattern. The dimensions of the sensor data acquisition time include hours, minutes, aquaculture status and weather status. The dynamic time-series features, static auxiliary features, and time-series auxiliary features are used as inputs to the pollution load index prediction model in tensor form, which is expressed as follows: ; in, Indicates input, Represents the set of real numbers. Indicates batch size. This indicates the historical time step. Dimensions representing dynamic temporal characteristics The dimension representing the static auxiliary feature. The dimension representing the temporal auxiliary features.

[0020] As an optional implementation, in a second aspect of the present invention, the preset pollution load index prediction model includes: The input embedding layer is used to map different types of dynamic temporal features, static auxiliary features, and temporal auxiliary features in the tensor form of input data to the same dimensional space; A variable selection network layer is used to filter the target information of the spliced ​​features output by the input embedding layer, and to compress the feature dimension to adapt to the position encoding layer and the temporal attention layer, wherein the target information is used to enhance the contribution of higher-order exponential features to the prediction results. The location encoding layer is used to inject time step location information into the output of the variable selection network layer, so that the prediction model can perceive the degree of influence of time sequence on historical data; The temporal attention layer is used to capture long-range dependencies in historical temporal data output by the positional encoding layer through a self-attention mechanism, so as to identify the correlation strength between the features of the target time point and the predicted target result. The gated fusion layer is used to dynamically balance the contribution ratio between the temporal attention features output by the temporal attention layer and the static auxiliary features broadcast within the model, and to capture the interaction dependency between the temporal attention features and the static auxiliary features. Then, based on the contribution ratio and interaction dependency, the fusion weight is determined and the fusion feature is output. A feedforward network layer is used to extract the implicit cross-modal correlations within the fusion features output by the gated fusion layer, so as to enhance the discriminative expression of the fusion features for higher-order exponential features in the prediction task. The decoder layer is used to generate a continuous change prediction curve of the pollution load index in future time periods from the output of the feedforward network layer by time step, wherein the time step process includes using the prediction result of the previous time step as the input of the current time step. The output prediction layer is used to convert the high-dimensional time-series features of the decoding layer output, which contain continuous concentration change curves, into prediction results corresponding to pollution load indicators. The prediction results are the concentration values ​​of pollution load indicators at any time in the future period.

[0021] As an optional implementation, in a second aspect of the present invention, based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and a preset pollution load index prediction model, the predicted results corresponding to the pollution load index in the future time period are output, including: The dynamic temporal features, static auxiliary features, and temporal auxiliary features are mapped to a high-dimensional space of the same dimension through corresponding embedding layers. After the static auxiliary features are broadcast to the temporal dimension, the features of the same dimension are concatenated to generate concatenated features. Each dimension of the spliced ​​features is assigned an importance weight and the features are sorted by feature priority. The weighted spliced ​​features are then subjected to linear dimensionality reduction and / or non-linear enhancement through projection to generate a standardized feature vector. The standardized feature vector is injected with time step position information using sine-cosine position coding, wherein the sine-cosine position coding is calculated as follows: ; ; The encoding dimension is the same as the standardized feature vector dimension. This represents the location code, and pos represents the time step index. This represents the encoding dimension, where i represents the dimension index; The standardized feature vector of position encoding is computed in parallel through a multi-head attention mechanism to generate multiple sets of local attention features corresponding to the number of attention heads. Then, multiple sets of local features are output through linear projection and scaled dot product attention calculation. The multiple sets of local features are integrated and restored according to the dimension of the standardized feature vector. Temporal attention features are output through residual connection and normalization processing. The static auxiliary features broadcast are averaged and pooled to obtain static global features. The input temporal attention features are processed by a gating unit with the static global features as the initial state, retaining the temporal output and the last step output of the gating unit. The last step output is input into a fully connected layer network and an activation function to determine the gating weights. After broadcasting the static global features to the target time step dimension, the temporal output and the adjusted dimension static global features are weighted, fused, and normalized based on the gating weights to output the fused features. The fused features are expanded in dimension by a first fully connected network to capture the implicit correlation features between multimodal features in the fused features, and then passed through a second fully connected network to retain the effective correlation features in the implicit correlation features. The fused features are then enhanced by using the effective correlation features through residual connections and normalization. Initialize the target time step based on the target future time period corresponding to the prediction result. Determine the initial predicted value of the fusion feature through a fully connected layer based on the last step feature. Extract the single-step feature fragments related to the current time step from the initial predicted value and the fusion feature and concatenate them into a single-step fusion feature. Input the single-step fusion feature into the autoregressive decoding unit to output the single-step prediction feature. Input the single-step prediction feature and the fusion feature into the cross-attention unit and calculate the single-step enhanced feature through scaling dot product attention. Map the single-step enhanced feature through a fully connected layer to determine the current concentration value. Update the predicted value of the previous time step to the current concentration value until the current time step meets the target time step. Concatenate all the single-step enhanced features to generate a high-dimensional temporal feature. The high-dimensional time-series features are reduced in dimensionality based on the feature projection layer, and the effective features of the associated predicted values ​​are enhanced by the activation function to output low-dimensional time-series features. The low-dimensional time-series features are then regressed using a linear activation function through a concentration prediction head, mapping the low-dimensional features to continuous predicted values ​​of pollution load index concentrations. The continuous predicted values ​​of pollution load index concentrations are then restored to true concentration values ​​through inverse normalization, and the true concentration values ​​of the pollution load indexes are output as the prediction results. The pollution load indexes include chemical oxygen demand concentration, ammonia nitrogen concentration, and total phosphorus concentration, and the future time period corresponds to the target time step of the pollution load indexes.

[0022] A third aspect of this invention discloses another aquaculture wastewater pollution load prediction system, the system comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps in the aquaculture wastewater pollution load prediction method disclosed in the first aspect of the present invention.

[0023] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the aquaculture wastewater pollution load prediction method disclosed in the first aspect of the present invention.

[0024] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: This invention constructs a multi-dimensional water quality characterization system by simultaneously acquiring sensor data in physical, visual, and spectral sub-modalities, expanding the dimensions of input features and comprehensively capturing microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, it generates high-order index features of wastewater pollution. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability enables the model to identify complex correlations ignored by traditional methods. The high-order index features are arranged in a time series and combined with wastewater treatment process parameters to generate static and temporal auxiliary features. Through a pollution load index prediction model combined with dynamic temporal features and auxiliary features, multi-step prediction is performed. Addressing the core defects of existing wastewater pollution load prediction technologies, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, this invention achieves high-precision, highly adaptable, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. It forms a closed-loop design for the entire process of aquaculture wastewater pollution prediction. The accurate prediction of pollution load can be widely applied in the field of aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating a method for predicting pollution load in aquaculture wastewater, as disclosed in an embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram of the structure of an aquaculture wastewater pollution load prediction system disclosed in an embodiment of the present invention.

[0028] Figure 3 This is a schematic diagram of another aquaculture wastewater pollution load prediction system disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of the network architecture of an exemplary pollution load index prediction model of the present invention; Figure 5 yes Figure 4 A schematic diagram of an exemplary prediction process under the network architecture of the pollution load index prediction model; Figure 6 This is a schematic diagram of an exemplary aquaculture wastewater pollution load prediction system according to the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0031] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0032] This invention discloses a method and system for predicting pollution load in aquaculture wastewater. It simultaneously acquires sensor data in physical-chemical, visual, and spectral modalities, constructing a multi-dimensional water quality characterization system to expand the dimensions of input features and comprehensively capture microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, it generates high-order index features of wastewater pollution. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability allows the model to identify complex correlations ignored by traditional methods. The high-order index features are arranged in a time series, and static and temporal auxiliary features are generated by combining wastewater treatment process parameters. Multi-step prediction is performed by combining the pollution load index prediction model with dynamic temporal features and auxiliary features. Addressing the core shortcomings of existing wastewater pollution load prediction technologies, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, this invention achieves high accuracy, high adaptability, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. A closed-loop design for predicting pollution in aquaculture wastewater is established. Accurate prediction of pollution loads enables its widespread application in aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes. These are explained in detail below.

[0033] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for predicting pollution load in aquaculture wastewater, as disclosed in an embodiment of the present invention. Figure 1 The described method for predicting pollution loads from aquaculture wastewater can be applied to data processing systems / data processing equipment / data processing servers (including local processing servers or cloud processing servers). For example... Figure 1 As shown, the method for predicting pollution load in aquaculture wastewater may include the following operations: 101. Acquire sensor data of multiple modes at the tailwater collection point, wherein the sub-mode types of the sensor data include physical sub-mode, visual sub-mode, and spectral sub-mode.

[0034] Optionally, among the multiple modal sensing data, the physicochemical modal sensing data is obtained by collecting water quality parameters corresponding to the dissolved oxygen sensor, pH sensor, redox potential sensor, ammonia nitrogen sensor, turbidity sensor, and conductivity sensor. The sensor data of the visual modality is collected by a microscope / flow cytometer to collect the visual parameters of suspended matter in the tailwater, wherein the visual parameters include the size, shape and quantity distribution of suspended matter. The spectral mode sensing data is obtained by collecting the spectral parameters corresponding to the components of the tailwater using an ultraviolet-visible spectral probe / fluorescence spectrometer. The spectral parameters include the characteristics of dissolved organic matter and the concentration of target pollutants.

[0035] Specifically, multimodal sensors, capable of collecting data from multiple modes, are deployed at key nodes in the wastewater treatment process, such as the inlet of the equalization tank, the outlet of the biological treatment tank and sedimentation tank, and the final discharge outlet. These sensors collect multi-dimensional water quality data, providing comprehensive data support for pollution load prediction. The configuration includes three types of sensor modules: a physicochemical modal sensor group, a visual image modal acquisition group, and a spectral modal acquisition group. This can be understood as the physicochemical modal proposed in this application being a combination of physical and chemical modal sensor data. The physicochemical modal sensor group is used to acquire conventional water quality parameters, and the acquisition equipment includes dissolved oxygen sensors, pH sensors, redox potential sensors, ammonia nitrogen sensors, turbidity sensors, and conductivity sensors. The visual image modal acquisition group is used to acquire visual information about suspended matter, capturing the size, shape, and quantity distribution of suspended particles, and the acquisition equipment includes underwater microscopic imaging devices or flow cytometry. The spectral modal acquisition group is used to acquire spectral information of water components, detect the characteristics of dissolved organic matter and the concentration of specific pollutants, and the acquisition equipment includes ultraviolet-visible spectral probes or fluorescence spectrometers.

[0036] Specifically, in addition to the multimodal sensing unit, it also includes an edge computing unit, which serves as a local intelligent node and features low latency, high reliability, and strong adaptability. Deployed at the aquaculture wastewater treatment site, it connects directly to the multimodal sensing unit and is responsible for real-time data acquisition, preprocessing, feature extraction, and multi-source data fusion to generate high-order state features that can be used for intelligent decision-making. The edge computing gateway is responsible for the real-time acquisition of data from each sensor and preprocesses the data, including filtering, noise reduction, outlier detection and cleaning, and data standardization, to achieve time-series alignment and interpolation, real-time computing, and lightweight inference.

[0037] 102. Based on a preset feature extraction algorithm, extract the sub-modal features corresponding to each of the sensor data, and generate high-order index features of wastewater pollution based on the sub-modal features and a preset feature fusion enhancement algorithm.

[0038] Specifically, before entering the multimodal data fusion process, secondary optimization is required based on the preprocessing results of the edge computing gateway to ensure the consistency and validity of the input data: Data standardization is enhanced by using Z-Score standardization to process physicochemical parameters and spectral characteristic values, as shown in the following expression: ; In the formula, For parameters Historical average, Standard deviation; The image features are normalized using Min-Max and mapped to the [0,1] interval, as shown in the following expression: ; For spatiotemporal alignment refinement, a timestamp interpolation alignment algorithm is used to construct a unified temporal framework with 100ms as the minimum time unit, taking into account the differences in sampling frequency of different modalities. For missing data, linear interpolation is used for short-term missing data (≤3 sampling points), and GAN generation and completion are used for long-term missing data (>3 sampling points) based on adjacent time points of the same modality.

[0039] 103. Based on the temporal arrangement, the higher-order exponential features are used to generate dynamic temporal features. According to the preset tailwater treatment process parameters and sensor data acquisition time, the static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined.

[0040] 104. Based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and the preset pollution load index prediction model, output the prediction results corresponding to the pollution load index in the future time period.

[0041] Specifically, the pollution load prediction model, based on multimodal fusion features and process and time information, accurately predicts the continuous concentration change curves of key pollution indicators (COD, ammonia nitrogen, and total phosphorus) in aquaculture wastewater over the next few hours to obtain the specific concentration value at any given time, providing forward-looking data support for subsequent treatment process control.

[0042] As an example, this model employs a Temporal Fusion Transformer, which possesses the ability to capture long-term temporal dependencies and dynamically balance the contributions of multiple feature types. The model inputs three core feature classes: 256-dimensional multimodal high-order fusion features (including CPLI, SMAI, and FPF), 16-dimensional static process parameters, and 4-dimensional timestamp auxiliary features, comprehensively covering water quality dynamics and scenario constraints. Figure 4 As shown, the model contains eight core components, which achieve end-to-end accurate prediction through a variable selection network to filter key features, a temporal attention layer to capture long-range temporal correlations, a gating fusion layer to balance dynamic and static features, and an autoregressive decoder to generate continuous prediction curves.

[0043] As can be seen, the above-described embodiments of the invention, by simultaneously acquiring sensor data of physical, visual, and spectral submodalities, construct a multi-dimensional water quality characterization system, expanding the dimensions of input features and comprehensively capturing microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, high-order index features of wastewater pollution are generated. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability enables the model to identify complex correlations ignored by traditional methods. The high-order index features are arranged in a time series, and static and temporal auxiliary features are generated by combining wastewater treatment process parameters. Through a pollution load index prediction model combined with dynamic temporal features and auxiliary features, multi-step prediction is performed. Addressing the core defects of existing technologies in wastewater pollution load prediction, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, the invention achieves high accuracy, high adaptability, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. This forms a closed-loop design for the entire process of aquaculture wastewater pollution prediction. The accurate prediction of pollution load can be widely applied in the field of aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes.

[0044] As an optional embodiment, before extracting the submodal features corresponding to each of the sensing data based on a preset feature extraction algorithm, the above steps include the following steps: The sampling frequency of the sensor data is preset, and the sampling frequency is synchronized for sensors of any modal type; The sensing data is preprocessed, including moving average filtering / wavelet denoising to eliminate sensor noise and environmental interference, identifying and removing outliers in the sensing data based on statistical methods / rule bases, and normalizing the original sensing data to the same interval to eliminate differences in dimensions and / or ranges. Timestamp alignment is performed on sensor data with different sampling frequencies, and missing values ​​are filled in using linear interpolation / GAN.

[0045] Specifically, the edge computing gateway synchronously collects data from various sensors at a user-defined sampling frequency to ensure spatiotemporal consistency. It supports breakpoint resumption and local caching, temporarily storing data in case of network anomalies and automatically retransmitting it upon recovery.

[0046] Specifically, the edge computing gateway uses methods such as moving average filtering and wavelet denoising to eliminate sensor noise and environmental interference; it identifies and removes abnormal data points based on statistical methods (such as the 3σ principle) or rule bases (such as exceeding the range, abnormal mutations, etc.); and it normalizes the original data of different dimensions and ranges to a unified interval to facilitate subsequent fusion and modeling.

[0047] Specifically, the edge computing gateway timestamps data from different sampling frequencies and uses linear interpolation or GAN methods to fill in missing time data.

[0048] As can be seen, through the above optional embodiments, by presetting a unified sampling frequency and forcibly synchronizing each submodal sensor, and by eliminating the time offset of cross-modal data through a timestamp alignment algorithm, the problem of temporal misalignment caused by the difference in sampling period of multimodal data is solved, thereby reducing the spatiotemporal correlation modeling error of pollutant evolution process, providing a strictly aligned input sequence for subsequent cross-modal attention mechanism, and improving the accuracy of high-order feature fusion. In response to the core problems of low quality of heterogeneous data, poor temporal consistency, and large dimensional differences in aquaculture tailwater pollution load prediction, the standardized preprocessing of multimodal sensor data lays a high-quality data foundation for subsequent feature extraction and model prediction.

[0049] As an optional embodiment, the step above, extracting the submodal features corresponding to each of the sensing data based on a preset feature extraction algorithm, includes: For the physicochemical submodal sensing data, based on statistical feature extraction and trend feature analysis, the statistical features and variation features of the basic parameters corresponding to water quality are extracted. The physicochemical submodal features are determined according to the statistical features and variation features. The statistical features of the basic parameters include the mean and variance of the water quality parameters corresponding to the physicochemical submodal sensing data. The variation features of the basic parameters include the slope of the curve of the change trend calculated by linear regression. For visual modal sensing data, morphological features of suspended particles are extracted based on a lightweight neural network for image processing, and texture features of suspended particles are extracted based on a gray-level co-occurrence matrix. Visual modal features are determined based on the morphological and texture features, wherein the morphological features include the average particle size, density, and number concentration of suspended particles, and the texture features include energy and entropy. For spectral modal sensing data, based on spectral feature engineering and peak position identification analysis, UV-Vis spectral features and fluorescence spectral features are extracted. Spectral modal features are determined based on the UV-Vis spectral features and fluorescence spectral features. The UV-Vis spectral features include the absorbance ratio between 254 nm and 436 nm wavelengths, the absorbance at 350 nm wavelength, and the full width at half maximum (FWHM) of the characteristic peaks. The fluorescence spectral features include the fluorescence index and humification index calculated from protein-like fluorescence peaks and humic substance-like fluorescence peaks.

[0050] As an example, a feature-level fusion strategy is adopted to transform three heterogeneous data types—physicochemical, visual image, and spectral—into unified high-order index features, integrating the Pollution Load Index (CPLI), Sludge Microbial Activity Index (SMAI), and Featured Pollutant Fingerprint (FPF). Features are extracted from the physicochemical, visual image, and spectral modalities, respectively. For the physicochemical modality, statistical features such as mean and variance, as well as trend slope features based on linear regression, are extracted, totaling 18 dimensions. For the visual image modality, morphological features such as average particle size, compactness, and number concentration, as well as texture features such as energy and entropy, are extracted through a combination of lightweight CNN and traditional image processing, totaling 8 dimensions. For the spectral modality, features such as absorbance ratio and characteristic peak half-width are extracted from UV-Vis spectra, and features such as fluorescence index and humification index are extracted from fluorescence spectra, totaling 10 dimensions. The three types of modal features are concatenated into a 36-dimensional initial feature vector. Based on domain knowledge constraints and an improved PCA algorithm, the dimensionality is reduced, and the principal components with a cumulative variance contribution rate of ≥95% are retained. Then, feature enhancement is performed through an autoencoder to finally generate a 256-dimensional fused feature vector, which includes high-order index features with semantics such as the Comprehensive Pollution Load Index (CPLI), the Sludge Microbial Activity Index (SMAI), and the Featured Pollutant Fingerprint (FPF).

[0051] Specifically, physicochemical modal feature extraction involves using statistical feature extraction and trend analysis methods to extract the basic state of water quality and core pollution indicators. The basic statistical features include the mean. ,variance Regarding trend characteristics, the slope was calculated using linear regression. To determine the upward / downward trend of parameters.

[0052] Further, visual image feature extraction: A method combining a lightweight CNN network and traditional image processing was used to extract the morphological features of suspended matter and the state features of sludge flocs. Among these, morphological features included: average particle size. ( The equivalent diameter of the particle. (particle area) and density L is the particle perimeter, and the number concentration is ( Texture features: Energy extraction via gray-level co-occurrence matrix. ,entropy .

[0053] Furthermore, spectral modal feature extraction: Spectral feature engineering and peak position identification methods are used to identify organic composition and extract characteristic pollutant fingerprints. UV-Vis spectroscopy: The absorbance ratio at 254nm / 436nm (E2 / E4), absorbance at 350nm (characterizing humic content), and full width at half maximum (FWHM) of characteristic peaks are extracted. Fluorescence spectroscopy: Protein-like fluorescence peaks (Ex280nm / Em350nm) and humic-like fluorescence peaks (Ex350nm / Em450nm) are identified, and the fluorescence index is calculated. Humus index .

[0054] As can be seen, through the above optional embodiments, the stability and fluctuation range of pollutant concentration are reflected by calculating the mean and variance of water quality parameters; the slope of the pollutant concentration change curve is captured by linear regression; the dynamic evolution trend of pollution load is captured; the average particle size, compactness, and number concentration of suspended matter are calculated based on a lightweight neural network to quantify the physical morphology of particulate matter; the spatial distribution characteristics of suspended matter are characterized by calculating energy (reflecting texture uniformity) and entropy (reflecting texture complexity) through the gray-level co-occurrence matrix (GLCM); morphological features can distinguish suspended matter from different sources (such as algal particles or feed residue), improving classification accuracy; and texture features reflect the aggregation / dispersion process of suspended matter. This approach shortens the model's response speed to sudden changes in water turbidity. It extracts the 254nm / 436nm absorbance ratio to characterize the molecular weight distribution of organic matter, the 350nm absorbance to reflect the humic content, and the full width at half maximum (FWHM) of characteristic peaks to indicate the aromaticity of organic matter. It provides macroscopic parameters (such as ammonia nitrogen concentration) through physicochemical features, supplements the physical properties of particulate matter with visual features, and delves into the molecular level with spectral features. The fusion of these three features expands the model's representation of pollutants. Furthermore, through refined feature extraction from multimodal sensor data, it constructs a microscopic representation system of water pollution load from physicochemical, visual, and spectral dimensions, significantly improving the feature expression depth, dynamic modeling capability, and adaptability to complex scenarios in pollution load prediction.

[0055] As an optional embodiment, the step above, generating high-order index features of wastewater pollution based on the modal features and a preset feature fusion enhancement algorithm, includes: All submodal features are concatenated based on the target dimension to generate an initial fusion vector; Principal component analysis with domain knowledge constraints is used to reduce the dimensionality of the initial fusion vector and retain the principal components with a cumulative variance contribution rate greater than 95%. The domain knowledge is the knowledge of wastewater treatment. The corresponding constraint is that the weight of the pollution load index contained in the first principal component in the principal component analysis is greater than the preset constraint value. The pollution load index includes chemical oxygen demand concentration, ammonia nitrogen concentration and total phosphorus concentration. The initial fusion vector after dimensionality reduction is enhanced with features based on an autoencoder, and the domain knowledge is fused with the enhanced initial fusion vector to generate high-order exponential features of wastewater pollution. The autoencoder includes an input layer, an output layer, a bottleneck layer, and two hidden layers. The input of the input layer is the fusion vector after dimensionality reduction. The output layer is set with reconstruction error constraints. The bottleneck layer is used to generate the higher-order exponential features. The hidden layers include the ReLU activation function. The higher-order index features include the comprehensive pollution load index, the sludge microbial activity index, and the characteristic pollutant fingerprint; The comprehensive pollution load index is expressed as: ; in, The comprehensive pollution load index is used to characterize the pollution load level. This represents the standardized chemical oxygen demand. This indicates the maximum permissible concentration of chemical oxygen demand. This represents the standardized ammonia nitrogen concentration. This indicates the maximum permissible concentration of ammonia nitrogen. This represents the absorbance ratio between the 254nm and 436nm wavelengths. This represents the maximum reference value for the absorbance ratio. This represents the standardized number concentration of suspended solids. The weighting of chemical oxygen demand. The weighting of ammonia nitrogen concentration. The weighting of absorbance ratio is indicated. The weight representing the number concentration of suspended solids, and , , and satisfy ; The sludge microbial activity index is expressed as: ; in, The sludge microbial activity index is used to characterize the combined dissolved oxygen trend, sludge floc structure, and microbial metabolite characteristics, thereby determining the activity status of the microorganisms. This represents the slope of the linear regression of the standardized dissolved oxygen time series data. This represents the sludge floc density extracted from sensor data of visual modalities. This represents the standardized microbial protein characteristic values; The characteristic pollutant fingerprint is represented as follows: ; in, The fingerprint represents a characteristic pollutant, specifically a three-dimensional vector composed of spectral modal features. This indicates the molecular weight of organic compounds as characterized by the absorbance ratio between 254 nm and 436 nm. The fluorescence index is used to determine the source of organic matter, and the humification index is used to determine the degree of decay of organic matter. Each type of pollutant corresponds one-to-one with the three-dimensional vector of the FPF.

[0056] As an example, feature vector construction involves concatenating the features extracted from each modality along fixed dimensions to form an initial fusion vector. There are 36 dimensions in total, among which, (18 dimensions): Includes the mean, variance, and trend slope of 6 types of physicochemical parameters; (8-dimensional): Includes 3 types of morphological features, 3 types of texture features, and 2 types of distribution features; (10-dimensional): Includes 5 types of UV-Vis features and 5 types of fluorescence spectral features.

[0057] Further exemplified, feature dimensionality reduction and enhancement: An improved PCA algorithm (with domain knowledge constraints) is used to reduce the 36-dimensional features to 25-28 dimensions, retaining principal components with a cumulative variance contribution rate ≥95%. The constraint is that the first principal component must contain weights ≥0.6 for core pollution indicators such as ammonia nitrogen and COD. After dimensionality reduction, feature enhancement is performed using an autoencoder, the structure of which is as follows: Input layer: 36-dimensional initial feature vector (18-dimensional physicochemical modes + 8-dimensional optical image modes + 10-dimensional spectral modes); Hidden layer 1: 24-dimensional (activation function ReLU); Bottleneck layer: 12-dimensional (high-order feature core) Hidden layer 2: 24-dimensional (ReLU activation function) Output layer: 36-dimensional (reconstruction error constraint ≤ 0.05).

[0058] Specifically, each pollutant type corresponds to a unique FPF vector, and the pollutant database can be matched using cosine similarity to achieve rapid identification of characteristic pollutants.

[0059] As can be seen, through the above optional embodiments, by unifying the physical, visual, and spectral modal features to the target dimension, an initial fusion vector is generated, eliminating data silos between modalities, retaining principal components with a cumulative variance contribution rate >95%, and forcibly constraining the weights of chemical oxygen demand, ammonia nitrogen, and total phosphorus in the first principal component to be higher than a preset threshold, ensuring that key pollution indicators dominate the feature space. PCA dimensionality reduction compresses the feature dimension by more than 50%, reducing computational complexity. Domain knowledge constraints increase the proportion of core indicators such as chemical oxygen demand, ammonia nitrogen, and total phosphorus in the feature space, shortening the response speed of the prediction model to key pollutants. This is achieved through the input layer, hidden layer (ReLU activation), and bottleneck... The nonlinear mapping of the layer and output layer enhances the features of the fusion vector after dimensionality reduction and introduces reconstruction error constraints to improve feature robustness. High-order exponential features (CPLI, SMAI, FPF) are generated in the bottleneck layer, and physicochemical indicators, microbial activity and pollutant fingerprints are deeply integrated to improve the model's ability to identify complex pollution patterns. Thus, through multimodal feature fusion and domain knowledge-driven high-order exponential construction, the core problems of feature redundancy, insufficient representation of key indicators and fuzzy identification of pollution types in aquaculture tailwater pollution load prediction are solved, thereby realizing multi-dimensional quantitative assessment of pollution load, dynamic modeling optimization and improved pollutant source tracing capabilities.

[0060] As an optional embodiment, the above steps, generating dynamic temporal features from the higher-order exponential features based on temporal arrangement, and determining the static auxiliary features and temporal auxiliary features of the dynamic temporal features according to preset tailwater treatment process parameters and sensor data acquisition time, include: The higher-order exponential features are arranged in time series using a preset historical time step to generate dynamic time series features that capture the dynamic change patterns of water quality and their association with multimodal sensor data. The preset effluent treatment process parameters are normalized to generate static auxiliary features that provide constraints for the process scenario. The effluent treatment process parameters include the volume of the biological treatment tank, the area of ​​the aeration zone, the residence time of the coagulation tank, the membrane pore size, and the sensor deployment node type. The sensor data acquisition time is processed by 1-hot encoding and normalization to generate a time-series auxiliary feature of the pollution load time periodic pattern. The dimensions of the sensor data acquisition time include hours, minutes, aquaculture status and weather status. The dynamic time-series features, static auxiliary features, and time-series auxiliary features are used as inputs to the pollution load index prediction model in tensor form, which is expressed as follows: ; in, Indicates input, Represents the set of real numbers. Indicates batch size. This indicates the historical time step. Dimensions representing dynamic temporal characteristics The dimension representing the static auxiliary feature. The dimension representing the temporal auxiliary features.

[0061] As an example, as shown in Table 1, the input data of the pollution load prediction model includes dynamic time-series features, static auxiliary features, and time-series auxiliary features. Among the three types of data input to the model, the dynamic time-series features are the time-series integration of multimodal features (including CPLI, SMAI, and FPF high-order quantization features) after processing by feature extraction algorithms and feature fusion enhancement algorithms. The static auxiliary features and time-series auxiliary features are two types of supplementary features, which respectively provide constraints on the tailwater treatment process scenario and time-series period constraints. The generation process of the dynamic time-series features can be summarized as follows: the multi-dimensional data collected by the multimodal sensing unit is extracted by the modal feature extraction in step 102 to form 36-dimensional initial features. The initial features are first reduced to 25-28 dimensions by the improved PCA, and then enhanced by the autoencoder and embedded with the three types of high-order features CPLI, SMAI, and FPF to generate 256-dimensional single-step fusion features. Finally, the single-step fusion features are arranged in the historical 24-hour (8640 steps) time sequence to form the dynamic time-series features (8640×256 dimensions) required by the model.

[0062] Table 1 - Details of Pre-input Data

[0063] Specifically, the dynamic time-series features, static auxiliary features, and time-series auxiliary features are organized in tensor form as input data for the preset pollution load index prediction model: ,in, Batch size (set during training) During reasoning ); The length of the historical time series; For dynamic temporal feature dimensions; For static auxiliary feature dimensions; As the temporal auxiliary feature dimension, the total input dimension is 358.

[0064] As can be seen, through the above optional embodiments, dynamic time-series features are generated by arranging higher-order exponential features based on a preset historical time step, capturing the nonlinear changes of water quality parameters over time. The time-series arrangement preserves the continuous information of pollutant migration, transformation, and degradation, improving the model's adaptability to complex water quality changes. Furthermore, the fusion of dynamic time-series features allows the model to identify the chain reaction of "sudden load increase - microbial inactivation - organic matter enrichment," reducing prediction errors. Normalization of effluent treatment process parameters generates static auxiliary features, which serve as process scenario constraints for model input, enabling the model to automatically match the operating characteristics of different treatment units and adjusting the sensor data acquisition time. One-hot encoding and normalization are performed to generate time-series auxiliary features, explicitly modeling the time-periodic pattern of pollution load. Short-term fluctuations in pollution load are captured by hourly / minute-level time encoding, shortening the model's response speed to sudden load changes. Dynamic time-series features, static auxiliary features, and time-assisted features are computed in parallel in tensor form to improve efficiency. Through the collaborative modeling of dynamic time-series features, static auxiliary features, and time-assisted features, combined with tensor input structure, the tailwater pollution load prediction model achieves deep integration of water quality dynamic evolution law, process constraints, and time-periodic law, significantly improving prediction accuracy, system adaptability, and control foresight.

[0065] As an optional embodiment, such as Figure 4 As shown, the preset pollution load index prediction model in the above steps includes: The input embedding layer is used to map different types of dynamic temporal features, static auxiliary features, and temporal auxiliary features in the tensor form of input data to the same dimensional space; A variable selection network layer is used to filter the target information of the spliced ​​features output by the input embedding layer, and to compress the feature dimension to adapt to the position encoding layer and the temporal attention layer, wherein the target information is used to enhance the contribution of higher-order exponential features to the prediction results. The location encoding layer is used to inject time step location information into the output of the variable selection network layer, so that the prediction model can perceive the degree of influence of time sequence on historical data; The temporal attention layer is used to capture long-range dependencies in historical temporal data output by the positional encoding layer through a self-attention mechanism, so as to identify the correlation strength between the features of the target time point and the predicted target result. The gated fusion layer is used to dynamically balance the contribution ratio between the temporal attention features output by the temporal attention layer and the static auxiliary features broadcast within the model, and to capture the interaction dependency between the temporal attention features and the static auxiliary features. Then, based on the contribution ratio and interaction dependency, the fusion weight is determined and the fusion feature is output. A feedforward network layer is used to extract the implicit cross-modal correlations within the fusion features output by the gated fusion layer, so as to enhance the discriminative expression of the fusion features for higher-order exponential features in the prediction task. The decoder layer is used to generate a continuous change prediction curve of the pollution load index in future time periods from the output of the feedforward network layer by time step, wherein the time step process includes using the prediction result of the previous time step as the input of the current time step. The output prediction layer is used to convert the high-dimensional time-series features of the decoding layer output, which contain continuous concentration change curves, into prediction results corresponding to pollution load indicators. The prediction results are the concentration values ​​of pollution load indicators at any time in the future period.

[0066] As an example, such as Figure 5 As shown in Table 2, the input embedding layer is divided into three sub-layers for different input types: dynamic embedding layer, static embedding layer and temporal embedding layer. This maps different types of input features to a high-dimensional space of the same dimension, eliminating differences in scale and distribution.

[0067] Table 2 - Structure of the Input Embedding Layer

[0068] As an example, in order to concatenate the three types of embedded features, the static embedded features ( Broadcast to the time-series dimension to obtain The dimension of the feature. After concatenating the three features, the dimension of the output feature is... .

[0069] As an example, the core objective of the variable selection network layer is to automatically filter key information in the 576-dimensional spliced ​​features, suppress redundant or noisy features, enhance the contribution of key features (such as CPLI and feature pollutant fingerprints) to the prediction results, and compress the feature dimension from 576 dimensions to 256 dimensions to adapt to the computational requirements of subsequent position encoding and attention layers.

[0070] Furthermore, as shown in Table 3, the weight calculation branch assigns importance weights to each dimension of the 576-dimensional features. This process prioritizes features. Feature weighting operations enhance or suppress the original features according to their weights; that is, important features (high weights) are preserved or amplified, while redundant features (low weights) are weakened or masked. The output projection layer performs linear dimensionality reduction and nonlinear enhancement on the weighted features, compressing 576 dimensions to 256 dimensions to generate standardized feature vectors.

[0071] Table 3 - Variable Selection Network Layer Structure

[0072] As an example, the role of the positional encoding layer is to inject time step location information, allowing the model to perceive the influence of the temporal order of historical data, such as the greater impact of recent data on prediction compared to older data. Sinusoidal Positional Encoding is used, with an encoding dimension of 256, consistent with the output dimension of the variable selection network. The calculation method for positional encoding is as follows: ; ; In the formula, For time step index (0~8639). , For dimension index (0~127).

[0073] The output dimension of the position coding layer is The 256-dimensional features include feature vectors and positional encoding information.

[0074] As an example, as shown in Table 4, the temporal attention layer captures long-range dependencies in historical time-series data through a self-attention mechanism. From 256 dimensions of features across 8640 time steps (24 hours), it accurately identifies the correlation strength between key time-point features and the current prediction target. Simultaneously, it dynamically allocates attention weights, assigning high weights to recent and strongly correlated features, and low weights to distant irrelevant and noisy features. By integrating the weighted temporal features, it generates an enhanced feature vector containing long-range dependency information, providing global temporal correlation support for subsequent gating fusion and prediction. The multi-head self-attention module computes 8 sets of attention weights in parallel, capturing multiple types of temporal dependencies and generating 8 sets of local attention features. (Number of attention heads) Each head dimension The input features (256 dimensions) are mapped to Q, K, and V (8×32 dimensions each) via linear projection. Scaling dot product attention is independently computed for each head group. The multi-head fusion layer outputs a set of 32-dimensional local features. Attention Dropout (p=0.1) is used to suppress overfitting and avoid excessive reliance on features from a single head. The multi-head fusion layer integrates the eight sets of local attention features into a unified 256-dimensional global feature, restoring the target dimension. The GELU activation function is applied to strengthen the non-linear correlation of features. Residual connections and normalization stabilize the training process, alleviate gradient vanishing, and ensure a stable feature distribution. Residual connections add the output of the multi-head fusion layer element-wise to the original input of this layer, preserving the original feature information. Layer normalization is applied to average the summed features. ),variance( Normalization eliminates distribution bias, thus adapting to long sequence training with a step length of 8640 and avoiding gradient decay caused by increasing network depth.

[0075] Table 4 - Temporal Attention Layer Structure

[0076] As an example, the core objective of the gated fusion layer is to dynamically balance the contribution ratio of temporal attention features and static auxiliary features, while capturing the interaction dependency between the two types of features, generating enhanced features that combine temporal dynamism and process constraints. Specific functions include: (1) Extract the global information of the process corresponding to the static auxiliary features, such as the constraint effect of fixed parameters such as the volume of the biochemical tank and the membrane pore size, to avoid the static features being ignored as simple splicing items.

[0077] (2) The weights are adjusted in real time through the gating mechanism. When the water quality is stable, the static process constraints are strengthened, and when the water quality changes abruptly, the temporal characteristics are emphasized.

[0078] (3) Integrate the features after interaction and maintain dimensional consistency. This provides high-quality input for the nonlinear transformation of subsequent feedforward network layers.

[0079] Furthermore, as shown in Table 5, the gated fusion layer consists of four components. Through global pooling of static features, global constraint information is extracted from the broadcasted static features, compressing the dimensionality to... The Gating Unit (GRU) processes temporal attention features, capturing long-range temporal dependencies, while also being guided by static global features. The gating weight calculation module generates dynamic gating weights, quantifying the contribution ratio of temporal features to static features. Feature fusion is then performed, fusing the two types of features according to their gating weights.

[0080] Table 5 - Gating Fusion Layer Structure

[0081] As an example, the core objective of the feedforward network layer is to uncover the complex nonlinear relationships among the gated fused features, enhance feature representation capabilities, and provide more discriminative high-order features for subsequent decoder prediction tasks. As shown in Table 6, FFN utilizes a two-stage fully connected network to expand and compress the 256-dimensional features after gated fusion. First, the first fully connected layer expands the 256-dimensional features to 1024 dimensions, capturing implicit relationships between multimodal features in the high-dimensional space, such as the potential connection between sludge floc morphology in optical images and recalcitrant organic matter in spectra. Then, the second fully connected layer compresses the features back to 256 dimensions, retaining effective correlated features and eliminating redundant information in the high-dimensional space. Both fully connected layers employ the GELU activation function (Gaussian error linear unit). Compared to ReLU, GELU provides smoother activation of feature values, supports gradient propagation of small-amplitude feature values, and better adapts to continuous fluctuations in water quality features, such as the gradual changes in ammonia nitrogen concentration. Dropout (p=0.2) is added after two fully connected layers to randomly deactivate some neurons and suppress the model's over-reliance on specific feature combinations. The residual connection directly adds the input and output of the FFN to alleviate the gradient vanishing problem in deep networks and ensures that long sequence training with a length of 8640 steps can converge stably.

[0082] Table 6 - Feedforward Network Layer Structure

[0083] As an example, the goal of the decoder layer is to generate continuous pollution load prediction curves for the next 1-4 hours (360 / 720 / 1440 steps) step by step based on the encoded time-series features (24 hours, step size 8640). As shown in Table 7, the decoder layer consists of an autoregressive decoding unit, a cross-attention module, and a target time step setting module. The decoder ensures the continuity of the prediction sequence through an autoregressive approach, using the prediction results of the previous time step as the input of the current time step to avoid abrupt fluctuations in the prediction values, which conforms to the physical law of continuous and gradual changes in water quality. By associating historical encoded features through cross-attention, it ensures that the prediction does not deviate from the support of historical data. Through a configurable target time step ( It supports configurable prediction durations of 1 / 2 / 4 hours, dynamically adjusts the number of generation steps, and outputs concentration change curves for three indicators: COD, ammonia nitrogen, and total phosphorus, meeting the forecasting needs of different decision-making scenarios.

[0084] Table 7 - Decoder Layer Structure

[0085] As an example, the output prediction layer can convert the high-dimensional temporal features output by the decoder layer ( The pollution load is transformed into a specific pollution load index that can be directly used for decision-making, achieving a precise mapping from high-dimensional features to concentration values. As shown in Table 8, the output prediction layer consists of a feature projection layer, a concentration prediction head, and an inverse normalization operation. The feature projection layer reduces the dimensionality of the 256-dimensional high-dimensional features generated by the decoder to 64 dimensions through a fully connected layer, and strengthens the effective features through GELU activation to improve mapping accuracy. The concentration prediction head uses a fully connected layer to map the low-dimensional features to normalized prediction values ​​of three types of pollution loads, achieving the core mapping from features to concentration. Since the pollution load concentration is a continuous value, a linear activation function is used for regression prediction to avoid prediction value offset caused by nonlinear activation. During the training phase, Z-score normalization is performed on the pollution load concentration to improve the model convergence speed. During the prediction phase, an inverse normalization operation is required. To restore the true concentration.

[0086] Table 8 - Output Prediction Layer Structure

[0087] As can be seen, through the above optional embodiments, the input embedding layer maps dynamic temporal features, static auxiliary features, and temporal auxiliary features to the same dimensional space, eliminating the dimensional and semantic differences between heterogeneous data and improving feature compatibility and model generalization ability; the variable selection network layer filters the target information in the spliced ​​features output by the input embedding layer, and adapts subsequent modules through dimensional compression, which can eliminate redundant information and strengthen key features; the positional encoding layer injects time step position information into the historical temporal data output by the variable selection layer, enabling the model to perceive the impact of temporal order on pollution load, and then the temporal attention layer captures long-range dependencies through a self-attention mechanism, identifying the correlation strength between the target time point features and the prediction results. Long-range dependency modeling reduces pollution load prediction errors while improving temporal sensitivity; the gated fusion layer dynamically balances temporal attention features and Static auxiliary features are used, and fusion weights are calculated through interactive dependencies to increase process adaptability and cross-modal information correlation mining. This not only reveals the synergistic effect of process parameters, time cycles, and load changes, but also improves the accuracy of prediction. Cross-modal implicit correlations in gated fusion features are extracted through a feedforward network layer to enhance the discriminative power of higher-order exponential features. The continuous change curve of pollution load in future periods is recursively generated by the decoder layer based on the prediction results of the previous time step, simulating the physicochemical laws of pollutant migration, transformation, and degradation. This provides the ability to mine implicit laws and make continuous predictions. Through a pollution load index prediction model with multimodal feature embedding and deep time series modeling, combined with variable selection, location encoding, self-attention mechanism, and gated fusion, high-precision dynamic prediction of effluent pollution load, long-cycle dependency modeling, and cross-modal feature interactive optimization are achieved.

[0088] As an optional embodiment, the step described above, outputting the prediction results of the pollution load index for the future period based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and the preset pollution load index prediction model, includes: The dynamic temporal features, static auxiliary features, and temporal auxiliary features are mapped to a high-dimensional space of the same dimension through corresponding embedding layers. After the static auxiliary features are broadcast to the temporal dimension, the features of the same dimension are concatenated to generate concatenated features. Each dimension of the spliced ​​features is assigned an importance weight and the features are sorted by feature priority. The weighted spliced ​​features are then subjected to linear dimensionality reduction and / or non-linear enhancement through projection to generate a standardized feature vector. The standardized feature vector is injected with time step position information using sine-cosine position coding, wherein the sine-cosine position coding is calculated as follows: ; ; The encoding dimension is the same as the standardized feature vector dimension. This represents the location code, and pos represents the time step index. This represents the encoding dimension, where i represents the dimension index; The standardized feature vector of position encoding is computed in parallel through a multi-head attention mechanism to generate multiple sets of local attention features corresponding to the number of attention heads. Then, multiple sets of local features are output through linear projection and scaled dot product attention calculation. The multiple sets of local features are integrated and restored according to the dimension of the standardized feature vector. Temporal attention features are output through residual connection and normalization processing. The static auxiliary features broadcast are averaged and pooled to obtain static global features. The input temporal attention features are processed by a gating unit with the static global features as the initial state, retaining the temporal output and the last step output of the gating unit. The last step output is input into a fully connected layer network and an activation function to determine the gating weights. After broadcasting the static global features to the target time step dimension, the temporal output and the adjusted dimension static global features are weighted, fused, and normalized based on the gating weights to output the fused features. The fused features are expanded in dimension by a first fully connected network to capture the implicit correlation features between multimodal features in the fused features, and then passed through a second fully connected network to retain the effective correlation features in the implicit correlation features. The fused features are then enhanced by using the effective correlation features through residual connections and normalization. Initialize the target time step based on the target future time period corresponding to the prediction result. Determine the initial predicted value of the fusion feature through a fully connected layer based on the last step feature. Extract the single-step feature fragments related to the current time step from the initial predicted value and the fusion feature and concatenate them into a single-step fusion feature. Input the single-step fusion feature into the autoregressive decoding unit to output the single-step prediction feature. Input the single-step prediction feature and the fusion feature into the cross-attention unit and calculate the single-step enhanced feature through scaling dot product attention. Map the single-step enhanced feature through a fully connected layer to determine the current concentration value. Update the predicted value of the previous time step to the current concentration value until the current time step meets the target time step. Concatenate all the single-step enhanced features to generate a high-dimensional temporal feature. The high-dimensional time-series features are reduced in dimensionality based on the feature projection layer, and the effective features of the associated predicted values ​​are enhanced by the activation function to output low-dimensional time-series features. The low-dimensional time-series features are then regressed using a linear activation function through a concentration prediction head, mapping the low-dimensional features to continuous predicted values ​​of pollution load index concentrations. The continuous predicted values ​​of pollution load index concentrations are then restored to true concentration values ​​through inverse normalization, and the true concentration values ​​of the pollution load indexes are output as the prediction results. The pollution load indexes include chemical oxygen demand concentration, ammonia nitrogen concentration, and total phosphorus concentration, and the future time period corresponds to the target time step of the pollution load indexes.

[0089] As can be seen, through the above optional embodiments, heterogeneous features are mapped to the same dimensional space by the unified embedding of dynamic temporal, static auxiliary, and temporal auxiliary features, eliminating dimensional differences. Through importance weight ranking and linear / nonlinear projection, core features are retained and redundant dimensions are compressed, improving feature utilization while optimizing computational complexity. Sine-cosine positional encoding and temporal attention enhance the model's ability to perceive temporal order while accurately capturing long-range dependencies. By transforming process parameters into static global features as the initial state of the gating unit, the contribution ratio of temporal output and static features is dynamically adjusted through the gating mechanism, strengthening the guidance of process constraints on prediction results. Gating fusion enables the model to automatically adapt to different processing units to improve prediction accuracy. The accuracy is improved, and the gating weight reveals the synergistic effect of process parameters, time cycle, and load changes, supporting the generation of process optimization suggestions. By using an autoregressive decoding unit and a cross-attention mechanism, a continuous curve of pollution load for future periods is recursively generated based on the predicted value of the previous time step. The interaction between single-step prediction features and fused features strengthens the correlation modeling between local details and global trends, thereby improving the ability to discover complex correlations of implicit patterns and thus enhancing the continuous prediction capability of effluent pollution load. Through a pollution load index prediction method that combines unified embedding of multimodal features, temporal attention modeling, and gating fusion enhancement, along with autoregressive decoding and cross-attention mechanisms, high-precision dynamic prediction, long-cycle dependency modeling, and cross-modal feature interaction optimization of effluent pollution load are achieved.

[0090] In one specific implementation scheme, an aquaculture wastewater pollution load prediction system is provided based on the technical solution of the present invention, addressing the following deficiencies of the existing technology: 1. Existing prediction methods mostly rely on data from single physicochemical sensors (such as pH, DO, ammonia nitrogen, etc.), which are single and isolated in dimension and cannot fully characterize the microscopic features and evolution mechanism of water quality.

[0091] 2. Although some methods integrate multiple types of sensing devices, the utilization of heterogeneous data composed of physicochemical data, image data, and spectral data is limited to simple parallel or independent processing, without deep integration, and the potential spatiotemporal correlations and causal logic between the data are not fully explored and utilized.

[0092] 3. Limited performance of prediction models, making it difficult to meet actual needs: Existing prediction methods mostly use traditional time series models or simple machine learning algorithms, which lack effective capture of long-term time series data dependencies and cannot take into account the combined effects of dynamic water quality characteristics and static process constraints. This results in low prediction accuracy, slow response speed, and limited prediction time, making it difficult to achieve refined load prediction for the next few hours and unable to provide reliable data support for the precise control of tailwater pollution treatment.

[0093] This prediction system addresses the shortcomings of existing technologies, specifically insufficient data utilization, inadequate feature fusion, and poor model adaptability in tailwater pollution load prediction. It employs a technical approach involving multimodal data acquisition, high-order feature fusion, and dynamic time-series modeling. This approach overcomes the limitations of existing technologies, providing high-precision, highly adaptable, and forward-looking tailwater pollution load predictions.

[0094] To meet the characteristics of long-term, multimodal, and attention-based models in the prediction system, a training scheme was designed to ensure stable training, accurate fitting, and strong generalization ability. An example of this training scheme is as follows: I. Training Environment Configuration The training hardware environment requires an NVIDIA A100 GPU (≥32GB VRAM) or V100 GPU (≥16GB VRAM), a CPU with ≥16 cores (Intel Xeon 8375C), ≥64GB of RAM, and ≥1TB of storage (SSD). The software environment uses Ubuntu 20.04 LTS as the operating system, PyTorch 1.18+ as the deep learning framework, Python 3.8+, and dependencies including NumPy, Pandas, Scikit-learn, H5Py, and OpenCV. For tool support, TensorBoard is used for training monitoring, and MLflow is used for model management.

[0095] II. Hyperparameter Settings In the model structure hyperparameters, the batch size (B) is set to 32 during training and 1 during inference, and the historical time step (B) is set to 1. The prediction time step is 8640. The default is 1440 steps (configurable). Embedding layer dimensions are 256 for dynamic, 256 for static, and 64 for temporal; the number of attention heads is 8 for temporal attention and 4 for cross-attention. The GRU layers are 2 in both the gating fusion layer and the decoder, with a hidden dimension of 256 in both. The Dropout probability is 0.2 in the variable selection layer, 0.1 in the attention layer, and 0.2 in the GRU layer. In the training hyperparameters, the initial learning rate (… )for The training epochs are 100, and the weight decay is [value missing]. The gradient pruning threshold is 1.0, and the optimizer selected is AdamW( , The learning rate scheduling uses cosine annealing (...). , ).

[0096] III. Loss Function Design A hybrid loss function is adopted to balance prediction accuracy, outlier resistance, and temporal continuity. The formula is as follows: The loss terms are defined as follows: Mean squared error (MSE) is It is used to penalize large errors and improve overall accuracy; The mean absolute error (MAE) is This can reduce the impact of outliers; Dynamic Time Warping (DTW) is This ensures the continuity of the time series curve. The weighting coefficients α=0.5, β=0.3, and γ=0.2 were determined through 5-fold cross-validation and can be fine-tuned based on the performance of the validation set.

[0097] IV. Training Steps S1. Data Loading and Preprocessing: A batch loading strategy is adopted, using H5Py to read the dataset, loading one batch (B=32) at a time, and applying data augmentation in real time, such as noise injection and feature perturbation. When formatting the input data, dynamic time-series features, static process features, and timestamp auxiliary features are concatenated in tensor format (shape=(32,8640,358)), and the label format is (32,1440,3). S2. Model Initialization: For parameter initialization, fully connected layers (embedding layers, variable selection network) are initialized using a Xavier normal distribution, attention layers and GRU layers are initialized using orthogonal initialization to avoid gradient vanishing, and the output layer is initialized using a normal distribution (mean = 0, variance = 0.01). During model warm-up, pre-trained weights can be loaded, or the first 3 epochs of the embedding layer can be frozen after initialization, and only subsequent layers can be trained (for initial training of a stable model). S3, Core Training: During the warm-up training phase, a low learning rate is used for the first 5 epochs ( ), gradually upgraded to To avoid model oscillation caused by an excessively high initial learning rate, the input data passes through the embedding layer, variable selection network, position encoding, temporal attention, gated fusion, feedforward network, decoder, and finally to the output layer, generating predicted values. During backpropagation, after calculating the mixed loss, the gradient is calculated using automatic differentiation, and gradient clipping (threshold = 1.0) is applied to prevent gradient explosion. The parameters are then updated using the AdamW optimizer. An early stopping strategy with a patience of 15 is used; if the validation set loss does not decrease for 15 consecutive epochs, training is stopped, and the historical best model is loaded. S4. Validation Testing: Validation set evaluation is performed every 5 epochs. MSE, MAE, MAPE, DTW, and Pearson correlation coefficients are calculated for the validation set. The optimal model is recorded (based on the minimum MAPE on the validation set). After training, the indicators are calculated on the test set to ensure overall accuracy reaches MAPE ≤ 5%, RMSE ≤ 0.1 mg / L (COD / ammonia nitrogen / total phosphorus), and temporal continuity meets the requirements of Pearson correlation coefficient ≥ 0.95 and DTW distance ≤ 5. Under extreme conditions, MAPE ≤ 8% for high pollution load scenarios and ≤ 7% for low activity scenarios.

[0098] V. Training Optimization Strategies Regularization and generalization capabilities are improved: L2 regularization (weight_decay=1e-5) is applied to all fully connected layers to suppress overfitting. Dropout is added to the variable selection network, temporal attention layer, and feedforward network layer to randomly deactivate some neurons. Four data augmentation strategies are applied in real time during training: temporal stretching, noise injection, scene generation, and feature perturbation, to improve the model's generalization ability.

[0099] Learning rate and gradient optimization: Cosine annealing scheduling is used, with the learning rate decaying according to a cosine curve every 20 epochs. At the end of the decay period (e.g., the 40th, 60th, and 80th epochs), the learning rate is reset to 1e-5 to help the model escape local optima. If GPU memory is insufficient, gradient accumulation is used (gradients are accumulated once every 4 mini-batches), which is equivalent to a batch size of 128, improving training stability.

[0100] Model distillation and deployment optimization: After training, using the complex TFT model as the teacher model, the number of attention heads was reduced to 4, the GRU hidden dimension was reduced to 128, and a lightweight model was distilled. The distillation loss was KL divergence. Improve inference speed. Quantize the optimal model into FP16 format and deploy it to an edge computing gateway to ensure inference latency ≤100ms / sample.

[0101] VI. Model Saving and Version Management Model weights, including model parameters, optimizer state, and learning rate scheduler state, are saved every 5 epochs. The model with the smallest MAPE on the validation set and the largest Pearson correlation coefficient on the test set is selected as the final model. Hyperparameters, training data, and evaluation metrics for each model version are recorded using MLflow to build a model version library for subsequent iterative optimization.

[0102] like Figure 6 As shown, the prediction system includes a multimodal sensing unit, an edge computing and data fusion unit, and a pollution load prediction engine. Through a closed-loop design of the entire process of data acquisition and preprocessing, multimodal data fusion, deep prediction, and result output, it achieves accurate prediction of pollution load.

[0103] In summary, the technical solutions disclosed in the embodiments of the present invention have the following advantages: This invention constructs a multi-dimensional water quality characterization system by simultaneously acquiring sensor data in physical, visual, and spectral sub-modalities, expanding the dimensions of input features and comprehensively capturing microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, it generates high-order exponential features of wastewater pollution. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability allows the model to identify complex correlations ignored by traditional methods. The high-order exponential features are arranged in a time series, and static and temporal auxiliary features are generated by combining wastewater treatment process parameters. Through a pollution load index prediction model combined with dynamic temporal features and auxiliary features, multi-step prediction is performed. Addressing the core shortcomings of existing wastewater pollution load prediction technologies, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, this invention achieves high accuracy, high adaptability, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. This forms a closed-loop design for the entire process of aquaculture wastewater pollution prediction. Accurate pollution load prediction can be widely applied in the field of aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes.

[0104] Example 2 Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an aquaculture wastewater pollution load prediction system disclosed in an embodiment of the present invention. Figure 2 The described aquaculture wastewater pollution load prediction system can be applied to data processing systems / data processing equipment / data processing servers (wherein, the server includes local processing servers or cloud processing servers). For example... Figure 2 As shown, the aquaculture wastewater pollution load prediction system may include: The data acquisition module 201 is configured to acquire sensor data of multiple modes from the tailwater collection point. The sub-modal types of the sensor data include physical sub-modal, visual sub-modal, and spectral sub-modal. The data fusion module 202 is configured to extract the sub-modal features corresponding to each of the sensor data based on a preset feature extraction algorithm, and generate high-order index features of wastewater pollution based on the sub-modal features and the preset feature fusion enhancement algorithm. The data processing module 203 is configured to generate dynamic time-series features from the higher-order exponential features based on the time-series arrangement, and to determine the static auxiliary features and time-series auxiliary features of the dynamic time-series features according to the preset tailwater treatment process parameters and the sensor data acquisition time. The data prediction module 204 is configured to output the prediction results of the pollution load index in the future period based on the dynamic time series characteristics, static auxiliary characteristics, time series auxiliary characteristics and the preset pollution load index prediction model.

[0105] As can be seen, the above-described embodiments of the invention, by simultaneously acquiring sensor data of physical, visual, and spectral submodalities, construct a multi-dimensional water quality characterization system, expanding the dimensions of input features and comprehensively capturing microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, high-order index features of wastewater pollution are generated. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability enables the model to identify complex correlations ignored by traditional methods. The high-order index features are arranged in a time series, and static and temporal auxiliary features are generated by combining wastewater treatment process parameters. Through a pollution load index prediction model combined with dynamic temporal features and auxiliary features, multi-step prediction is performed. Addressing the core defects of existing technologies in wastewater pollution load prediction, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, the invention achieves high accuracy, high adaptability, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. This forms a closed-loop design for the entire process of aquaculture wastewater pollution prediction. The accurate prediction of pollution load can be widely applied in the field of aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes.

[0106] As an optional embodiment, among the multiple modal sensing data, the physicochemical modal sensing data is obtained by collecting water quality parameters corresponding to dissolved oxygen sensor, pH sensor, redox potential sensor, ammonia nitrogen sensor, turbidity sensor and conductivity sensor. The sensor data of the visual modality is collected by a microscope / flow cytometer to collect the visual parameters of suspended matter in the tailwater, wherein the visual parameters include the size, shape and quantity distribution of suspended matter. The spectral mode sensing data is obtained by collecting the spectral parameters corresponding to the components of the tailwater using an ultraviolet-visible spectral probe / fluorescence spectrometer. The spectral parameters include the characteristics of dissolved organic matter and the concentration of target pollutants.

[0107] As an optional embodiment, before extracting the submodal features corresponding to each of the sensing data based on a preset feature extraction algorithm, the following steps are included: The sampling frequency of the sensor data is preset, and the sampling frequency is synchronized for sensors of any modal type; The sensing data is preprocessed, including moving average filtering / wavelet denoising to eliminate sensor noise and environmental interference, identifying and removing outliers in the sensing data based on statistical methods / rule bases, and normalizing the original sensing data to the same interval to eliminate differences in dimensions and / or ranges. Timestamp alignment is performed on sensor data with different sampling frequencies, and missing values ​​are filled in using linear interpolation / GAN.

[0108] As can be seen, through the above optional embodiments, by presetting a unified sampling frequency and forcibly synchronizing each submodal sensor, and by eliminating the time offset of cross-modal data through a timestamp alignment algorithm, the problem of temporal misalignment caused by the difference in sampling period of multimodal data is solved, thereby reducing the spatiotemporal correlation modeling error of pollutant evolution process, providing a strictly aligned input sequence for subsequent cross-modal attention mechanism, and improving the accuracy of high-order feature fusion. In response to the core problems of low quality of heterogeneous data, poor temporal consistency, and large dimensional differences in aquaculture tailwater pollution load prediction, the standardized preprocessing of multimodal sensor data lays a high-quality data foundation for subsequent feature extraction and model prediction.

[0109] As an optional embodiment, based on a preset feature extraction algorithm, the modal features corresponding to each of the sensing data are extracted, including: For the physicochemical submodal sensing data, based on statistical feature extraction and trend feature analysis, the statistical features and variation features of the basic parameters corresponding to water quality are extracted. The physicochemical submodal features are determined according to the statistical features and variation features. The statistical features of the basic parameters include the mean and variance of the water quality parameters corresponding to the physicochemical submodal sensing data. The variation features of the basic parameters include the slope of the curve of the change trend calculated by linear regression. For visual modal sensing data, morphological features of suspended particles are extracted based on a lightweight neural network for image processing, and texture features of suspended particles are extracted based on a gray-level co-occurrence matrix. Visual modal features are determined based on the morphological and texture features, wherein the morphological features include the average particle size, density, and number concentration of suspended particles, and the texture features include energy and entropy. For spectral modal sensing data, based on spectral feature engineering and peak position identification analysis, UV-Vis spectral features and fluorescence spectral features are extracted. Spectral modal features are determined based on the UV-Vis spectral features and fluorescence spectral features. The UV-Vis spectral features include the absorbance ratio between 254 nm and 436 nm wavelengths, the absorbance at 350 nm wavelength, and the full width at half maximum (FWHM) of the characteristic peaks. The fluorescence spectral features include the fluorescence index and humification index calculated from protein-like fluorescence peaks and humic substance-like fluorescence peaks.

[0110] As can be seen, through the above optional embodiments, the stability and fluctuation range of pollutant concentration are reflected by calculating the mean and variance of water quality parameters; the slope of the pollutant concentration change curve is captured by linear regression; the dynamic evolution trend of pollution load is captured; the average particle size, compactness, and number concentration of suspended matter are calculated based on a lightweight neural network to quantify the physical morphology of particulate matter; the spatial distribution characteristics of suspended matter are characterized by calculating energy (reflecting texture uniformity) and entropy (reflecting texture complexity) through the gray-level co-occurrence matrix (GLCM); morphological features can distinguish suspended matter from different sources (such as algal particles or feed residue), improving classification accuracy; and texture features reflect the aggregation / dispersion process of suspended matter. This approach shortens the model's response speed to sudden changes in water turbidity. It extracts the 254nm / 436nm absorbance ratio to characterize the molecular weight distribution of organic matter, the 350nm absorbance to reflect the humic content, and the full width at half maximum (FWHM) of characteristic peaks to indicate the aromaticity of organic matter. It provides macroscopic parameters (such as ammonia nitrogen concentration) through physicochemical features, supplements the physical properties of particulate matter with visual features, and delves into the molecular level with spectral features. The fusion of these three features expands the model's representation of pollutants. Furthermore, through refined feature extraction from multimodal sensor data, it constructs a microscopic representation system of water pollution load from physicochemical, visual, and spectral dimensions, significantly improving the feature expression depth, dynamic modeling capability, and adaptability to complex scenarios in pollution load prediction.

[0111] As an optional embodiment, based on the submodal features and a preset feature fusion enhancement algorithm, high-order index features of wastewater pollution are generated, including: All submodal features are concatenated based on the target dimension to generate an initial fusion vector; Principal component analysis with domain knowledge constraints is used to reduce the dimensionality of the initial fusion vector and retain the principal components with a cumulative variance contribution rate greater than 95%. The domain knowledge is the knowledge of wastewater treatment. The corresponding constraint is that the weight of the pollution load index contained in the first principal component in the principal component analysis is greater than the preset constraint value. The pollution load index includes chemical oxygen demand concentration, ammonia nitrogen concentration and total phosphorus concentration. The initial fusion vector after dimensionality reduction is enhanced with features based on an autoencoder, and the domain knowledge is fused with the enhanced initial fusion vector to generate high-order exponential features of wastewater pollution. The autoencoder includes an input layer, an output layer, a bottleneck layer, and two hidden layers. The input of the input layer is the fusion vector after dimensionality reduction. The output layer is set with reconstruction error constraints. The bottleneck layer is used to generate the higher-order exponential features. The hidden layers include the ReLU activation function. The higher-order index features include the comprehensive pollution load index, the sludge microbial activity index, and the characteristic pollutant fingerprint; The comprehensive pollution load index is expressed as: ; in, The comprehensive pollution load index is used to characterize the pollution load level. This represents the standardized chemical oxygen demand. This indicates the maximum permissible concentration of chemical oxygen demand. This represents the standardized ammonia nitrogen concentration. This indicates the maximum permissible concentration of ammonia nitrogen. This represents the absorbance ratio between the 254nm and 436nm wavelengths. This represents the maximum reference value for the absorbance ratio. This represents the standardized number concentration of suspended solids. The weighting of chemical oxygen demand. The weighting of ammonia nitrogen concentration. The weighting of absorbance ratio is indicated. The weight representing the number concentration of suspended solids, and , , and satisfy ; The sludge microbial activity index is expressed as: ; in, The sludge microbial activity index is used to characterize the combined dissolved oxygen trend, sludge floc structure, and microbial metabolite characteristics, thereby determining the activity status of the microorganisms. This represents the slope of the linear regression of the standardized dissolved oxygen time series data. This represents the sludge floc density extracted from sensor data of visual modalities. This represents the standardized microbial protein characteristic values; The characteristic pollutant fingerprint is represented as follows: ; in, The fingerprint represents a characteristic pollutant, specifically a three-dimensional vector composed of spectral modal features. This indicates the molecular weight of organic compounds as characterized by the absorbance ratio between 254 nm and 436 nm. The fluorescence index is used to determine the source of organic matter, and the humification index is used to determine the degree of decay of organic matter. Each type of pollutant corresponds one-to-one with the three-dimensional vector of the FPF.

[0112] As can be seen, through the above optional embodiments, by unifying the physical, visual, and spectral modal features to the target dimension, an initial fusion vector is generated, eliminating data silos between modalities, retaining principal components with a cumulative variance contribution rate >95%, and forcibly constraining the weights of chemical oxygen demand, ammonia nitrogen, and total phosphorus in the first principal component to be higher than a preset threshold, ensuring that key pollution indicators dominate the feature space. PCA dimensionality reduction compresses the feature dimension by more than 50%, reducing computational complexity. Domain knowledge constraints increase the proportion of core indicators such as chemical oxygen demand, ammonia nitrogen, and total phosphorus in the feature space, shortening the response speed of the prediction model to key pollutants. This is achieved through the input layer, hidden layer (ReLU activation), and bottleneck... The nonlinear mapping of the layer and output layer enhances the features of the fusion vector after dimensionality reduction and introduces reconstruction error constraints to improve feature robustness. High-order exponential features (CPLI, SMAI, FPF) are generated in the bottleneck layer, and physicochemical indicators, microbial activity and pollutant fingerprints are deeply integrated to improve the model's ability to identify complex pollution patterns. Thus, through multimodal feature fusion and domain knowledge-driven high-order exponential construction, the core problems of feature redundancy, insufficient representation of key indicators and fuzzy identification of pollution types in aquaculture tailwater pollution load prediction are solved, thereby realizing multi-dimensional quantitative assessment of pollution load, dynamic modeling optimization and improved pollutant source tracing capabilities.

[0113] As an optional embodiment, dynamic temporal features are generated from the higher-order exponential features based on the temporal arrangement. Static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined according to preset wastewater treatment process parameters and sensor data acquisition time, including: The higher-order exponential features are arranged in time series using a preset historical time step to generate dynamic time series features that capture the dynamic change patterns of water quality and their association with multimodal sensor data. The preset effluent treatment process parameters are normalized to generate static auxiliary features that provide constraints for the process scenario. The effluent treatment process parameters include the volume of the biological treatment tank, the area of ​​the aeration zone, the residence time of the coagulation tank, the membrane pore size, and the sensor deployment node type. The sensor data acquisition time is processed by 1-hot encoding and normalization to generate a time-series auxiliary feature of the pollution load time periodic pattern. The dimensions of the sensor data acquisition time include hours, minutes, aquaculture status and weather status. The dynamic time-series features, static auxiliary features, and time-series auxiliary features are used as inputs to the pollution load index prediction model in tensor form, which is expressed as follows: ; in, Indicates input, Represents the set of real numbers. Indicates batch size. This indicates the historical time step. Dimensions representing dynamic temporal characteristics The dimension representing the static auxiliary feature. The dimension representing the temporal auxiliary features.

[0114] As can be seen, through the above optional embodiments, dynamic time-series features are generated by arranging higher-order exponential features based on a preset historical time step, capturing the nonlinear changes of water quality parameters over time. The time-series arrangement preserves the continuous information of pollutant migration, transformation, and degradation, improving the model's adaptability to complex water quality changes. Furthermore, the fusion of dynamic time-series features allows the model to identify the chain reaction of "sudden load increase - microbial inactivation - organic matter enrichment," reducing prediction errors. Normalization of effluent treatment process parameters generates static auxiliary features, which serve as process scenario constraints for model input, enabling the model to automatically match the operating characteristics of different treatment units and adjusting the sensor data acquisition time. One-hot encoding and normalization are performed to generate time-series auxiliary features, explicitly modeling the time-periodic pattern of pollution load. Short-term fluctuations in pollution load are captured by hourly / minute-level time encoding, shortening the model's response speed to sudden load changes. Dynamic time-series features, static auxiliary features, and time-assisted features are computed in parallel in tensor form to improve efficiency. Through the collaborative modeling of dynamic time-series features, static auxiliary features, and time-assisted features, combined with tensor input structure, the tailwater pollution load prediction model achieves deep integration of water quality dynamic evolution law, process constraints, and time-periodic law, significantly improving prediction accuracy, system adaptability, and control foresight.

[0115] As an optional embodiment, the preset pollution load index prediction model includes: The input embedding layer is used to map different types of dynamic temporal features, static auxiliary features, and temporal auxiliary features in the tensor form of input data to the same dimensional space; A variable selection network layer is used to filter the target information of the spliced ​​features output by the input embedding layer, and to compress the feature dimension to adapt to the position encoding layer and the temporal attention layer, wherein the target information is used to enhance the contribution of higher-order exponential features to the prediction results. The location encoding layer is used to inject time step location information into the output of the variable selection network layer, so that the prediction model can perceive the degree of influence of time sequence on historical data; The temporal attention layer is used to capture long-range dependencies in historical temporal data output by the positional encoding layer through a self-attention mechanism, so as to identify the correlation strength between the features of the target time point and the predicted target result. The gated fusion layer is used to dynamically balance the contribution ratio between the temporal attention features output by the temporal attention layer and the static auxiliary features broadcast within the model, and to capture the interaction dependency between the temporal attention features and the static auxiliary features. Then, based on the contribution ratio and interaction dependency, the fusion weight is determined and the fusion feature is output. A feedforward network layer is used to extract the implicit cross-modal correlations within the fusion features output by the gated fusion layer, so as to enhance the discriminative expression of the fusion features for higher-order exponential features in the prediction task. The decoder layer is used to generate a continuous change prediction curve of the pollution load index in future time periods from the output of the feedforward network layer by time step, wherein the time step process includes using the prediction result of the previous time step as the input of the current time step. The output prediction layer is used to convert the high-dimensional time-series features of the decoding layer output, which contain continuous concentration change curves, into prediction results corresponding to pollution load indicators. The prediction results are the concentration values ​​of pollution load indicators at any time in the future period.

[0116] As can be seen, through the above optional embodiments, the input embedding layer maps dynamic temporal features, static auxiliary features, and temporal auxiliary features to the same dimensional space, eliminating the dimensional and semantic differences between heterogeneous data and improving feature compatibility and model generalization ability; the variable selection network layer filters the target information in the spliced ​​features output by the input embedding layer, and adapts subsequent modules through dimensional compression, which can eliminate redundant information and strengthen key features; the positional encoding layer injects time step position information into the historical temporal data output by the variable selection layer, enabling the model to perceive the impact of temporal order on pollution load, and then the temporal attention layer captures long-range dependencies through a self-attention mechanism, identifying the correlation strength between the target time point features and the prediction results. Long-range dependency modeling reduces pollution load prediction errors while improving temporal sensitivity; the gated fusion layer dynamically balances temporal attention features and Static auxiliary features are used, and fusion weights are calculated through interactive dependencies to increase process adaptability and cross-modal information correlation mining. This not only reveals the synergistic effect of process parameters, time cycles, and load changes, but also improves the accuracy of prediction. Cross-modal implicit correlations in gated fusion features are extracted through a feedforward network layer to enhance the discriminative power of higher-order exponential features. The continuous change curve of pollution load in future periods is recursively generated by the decoder layer based on the prediction results of the previous time step, simulating the physicochemical laws of pollutant migration, transformation, and degradation. This provides the ability to mine implicit laws and make continuous predictions. Through a pollution load index prediction model with multimodal feature embedding and deep time series modeling, combined with variable selection, location encoding, self-attention mechanism, and gated fusion, high-precision dynamic prediction of effluent pollution load, long-cycle dependency modeling, and cross-modal feature interactive optimization are achieved.

[0117] As an optional embodiment, based on the dynamic time-series features, static auxiliary features, time-series auxiliary features, and a preset pollution load index prediction model, the predicted results corresponding to the pollution load index in the future time period are output, including: The dynamic temporal features, static auxiliary features, and temporal auxiliary features are mapped to a high-dimensional space of the same dimension through corresponding embedding layers. After the static auxiliary features are broadcast to the temporal dimension, the features of the same dimension are concatenated to generate concatenated features. Each dimension of the spliced ​​features is assigned an importance weight and the features are sorted by feature priority. The weighted spliced ​​features are then subjected to linear dimensionality reduction and / or non-linear enhancement through projection to generate a standardized feature vector. The standardized feature vector is injected with time step position information using sine-cosine position coding, wherein the sine-cosine position coding is calculated as follows: ; ; The encoding dimension is the same as the standardized feature vector dimension. This represents the location code, and pos represents the time step index. This represents the encoding dimension, where i represents the dimension index; The standardized feature vector of position encoding is computed in parallel through a multi-head attention mechanism to generate multiple sets of local attention features corresponding to the number of attention heads. Then, multiple sets of local features are output through linear projection and scaled dot product attention calculation. The multiple sets of local features are integrated and restored according to the dimension of the standardized feature vector. Temporal attention features are output through residual connection and normalization processing. The static auxiliary features broadcast are averaged and pooled to obtain static global features. The input temporal attention features are processed by a gating unit with the static global features as the initial state, retaining the temporal output and the last step output of the gating unit. The last step output is input into a fully connected layer network and an activation function to determine the gating weights. After broadcasting the static global features to the target time step dimension, the temporal output and the adjusted dimension static global features are weighted, fused, and normalized based on the gating weights to output the fused features. The fused features are expanded in dimension by a first fully connected network to capture the implicit correlation features between multimodal features in the fused features, and then passed through a second fully connected network to retain the effective correlation features in the implicit correlation features. The fused features are then enhanced by using the effective correlation features through residual connections and normalization. Initialize the target time step based on the target future time period corresponding to the prediction result. Determine the initial predicted value of the fusion feature through a fully connected layer based on the last step feature. Extract the single-step feature fragments related to the current time step from the initial predicted value and the fusion feature and concatenate them into a single-step fusion feature. Input the single-step fusion feature into the autoregressive decoding unit to output the single-step prediction feature. Input the single-step prediction feature and the fusion feature into the cross-attention unit and calculate the single-step enhanced feature through scaling dot product attention. Map the single-step enhanced feature through a fully connected layer to determine the current concentration value. Update the predicted value of the previous time step to the current concentration value until the current time step meets the target time step. Concatenate all the single-step enhanced features to generate a high-dimensional temporal feature. The high-dimensional time-series features are reduced in dimensionality based on the feature projection layer, and the effective features of the associated predicted values ​​are enhanced by the activation function to output low-dimensional time-series features. The low-dimensional time-series features are then regressed using a linear activation function through a concentration prediction head, mapping the low-dimensional features to continuous predicted values ​​of pollution load index concentrations. The continuous predicted values ​​of pollution load index concentrations are then restored to true concentration values ​​through inverse normalization, and the true concentration values ​​of the pollution load indexes are output as the prediction results. The pollution load indexes include chemical oxygen demand concentration, ammonia nitrogen concentration, and total phosphorus concentration, and the future time period corresponds to the target time step of the pollution load indexes.

[0118] As can be seen, through the above optional embodiments, heterogeneous features are mapped to the same dimensional space by the unified embedding of dynamic temporal, static auxiliary, and temporal auxiliary features, eliminating dimensional differences. Through importance weight ranking and linear / nonlinear projection, core features are retained and redundant dimensions are compressed, improving feature utilization while optimizing computational complexity. Sine-cosine positional encoding and temporal attention enhance the model's ability to perceive temporal order while accurately capturing long-range dependencies. By transforming process parameters into static global features as the initial state of the gating unit, the contribution ratio of temporal output and static features is dynamically adjusted through the gating mechanism, strengthening the guidance of process constraints on prediction results. Gating fusion enables the model to automatically adapt to different processing units to improve prediction accuracy. The accuracy is improved, and the gating weight reveals the synergistic effect of process parameters, time cycle, and load changes, supporting the generation of process optimization suggestions. By using an autoregressive decoding unit and a cross-attention mechanism, a continuous curve of pollution load for future periods is recursively generated based on the predicted value of the previous time step. The interaction between single-step prediction features and fused features strengthens the correlation modeling between local details and global trends, thereby improving the ability to discover complex correlations of implicit patterns and thus enhancing the continuous prediction capability of effluent pollution load. Through a pollution load index prediction method that combines unified embedding of multimodal features, temporal attention modeling, and gating fusion enhancement, along with autoregressive decoding and cross-attention mechanisms, high-precision dynamic prediction, long-cycle dependency modeling, and cross-modal feature interaction optimization of effluent pollution load are achieved.

[0119] The technical solutions disclosed in the embodiments of the present invention have the following advantages: This invention constructs a multi-dimensional water quality characterization system by simultaneously acquiring sensor data in physical, visual, and spectral sub-modalities, expanding the dimensions of input features and comprehensively capturing microscopic changes in water quality. Based on a preset feature extraction algorithm combined with a feature fusion enhancement algorithm, it generates high-order exponential features of wastewater pollution. Dynamic evolution modeling adapts the model to water quality fluctuation scenarios, and the improved feature expression capability allows the model to identify complex correlations ignored by traditional methods. The high-order exponential features are arranged in a time series, and static and temporal auxiliary features are generated by combining wastewater treatment process parameters. Through a pollution load index prediction model combined with dynamic temporal features and auxiliary features, multi-step prediction is performed. Addressing the core shortcomings of existing wastewater pollution load prediction technologies, such as insufficient data utilization, inadequate feature fusion, and poor model adaptability, this invention achieves high accuracy, high adaptability, and forward-looking control of pollution load prediction through a technical path of multi-modal data acquisition, high-order feature fusion, and dynamic temporal modeling. This forms a closed-loop design for the entire process of aquaculture wastewater pollution prediction. Accurate pollution load prediction can be widely applied in the field of aquaculture wastewater treatment, providing a solid foundation for intelligent control of subsequent treatment processes.

[0120] Example 3 Please see Figure 3 , Figure 3 This is another aquaculture wastewater pollution load prediction system disclosed in the embodiments of the present invention. Figure 3 The described aquaculture wastewater pollution load prediction system is applied in a data processing system / data processing equipment / data processing server (wherein, the server includes a local processing server or a cloud processing server). For example... Figure 3 As shown, the aquaculture wastewater pollution load prediction system may include: Memory 301 storing executable program code; Processor 302 coupled to memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the aquaculture wastewater pollution load prediction method described in Embodiment 1.

[0121] Example 4 This invention discloses a computer read storage medium that stores a computer program for electronic data interchange, wherein the computer program causes a computer to execute the steps of the aquaculture wastewater pollution load prediction method described in Embodiment 1.

[0122] Example 5 This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps of the aquaculture wastewater pollution load prediction method described in Embodiment 1.

[0123] The foregoing has described specific embodiments of this specification; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0124] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0125] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0126] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0131] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0132] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0133] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0134] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0135] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0136] Finally, it should be noted that the aquaculture wastewater pollution load prediction method and system disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting pollution load in aquaculture wastewater, characterized in that, The method includes: Multiple modal sensor data from tailwater collection points are acquired, and the modal types of the sensor data include physical modal, visual modal, and spectral modal. Based on a preset feature extraction algorithm, the submodal features corresponding to each of the sensor data are extracted. Based on the submodal features and the preset feature fusion enhancement algorithm, a high-order index feature of wastewater pollution is generated. Based on the temporal arrangement, the higher-order exponential features are used to generate dynamic temporal features. According to the preset tailwater treatment process parameters and sensor data acquisition time, the static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined. Based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and the preset pollution load index prediction model, the prediction results corresponding to the pollution load index in the future time period are output.

2. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Among the multiple modes of sensing data, the physicochemical mode of sensing data is obtained by collecting water quality parameters corresponding to dissolved oxygen sensor, pH sensor, redox potential sensor, ammonia nitrogen sensor, turbidity sensor and conductivity sensor. The sensor data of the visual modality is collected by a microscope / flow cytometer to collect the visual parameters of suspended matter in the tailwater, wherein the visual parameters include the size, shape and quantity distribution of suspended matter. The spectral mode sensing data is obtained by collecting the spectral parameters corresponding to the components of the tailwater using an ultraviolet-visible spectral probe / fluorescence spectrometer. The spectral parameters include the characteristics of dissolved organic matter and the concentration of target pollutants.

3. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Before extracting the modal features corresponding to each of the sensing data based on the preset feature extraction algorithm, the following steps are included: The sampling frequency of the sensor data is preset, and the sampling frequency is synchronized for sensors of any modal type; The sensing data is preprocessed, including moving average filtering / wavelet denoising to eliminate sensor noise and environmental interference, identifying and removing outliers in the sensing data based on statistical methods / rule bases, and normalizing the original sensing data to the same interval to eliminate differences in dimensions and / or ranges. Timestamp alignment is performed on sensor data with different sampling frequencies, and missing values ​​are filled in using linear interpolation / GAN.

4. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Based on a preset feature extraction algorithm, the modal features corresponding to each of the sensing data are extracted, including: For the physicochemical submodal sensing data, based on statistical feature extraction and trend feature analysis, the statistical features and variation features of the basic parameters corresponding to water quality are extracted. The physicochemical submodal features are determined according to the statistical features and variation features. The statistical features of the basic parameters include the mean and variance of the water quality parameters corresponding to the physicochemical submodal sensing data. The variation features of the basic parameters include the slope of the curve of the change trend calculated by linear regression. For visual modal sensing data, morphological features of suspended particles are extracted based on a lightweight neural network for image processing, and texture features of suspended particles are extracted based on a gray-level co-occurrence matrix. Visual modal features are determined based on the morphological and texture features, wherein the morphological features include the average particle size, density, and number concentration of suspended particles, and the texture features include energy and entropy. For spectral modal sensing data, based on spectral feature engineering and peak position identification analysis, UV-Vis spectral features and fluorescence spectral features are extracted. Spectral modal features are determined based on the UV-Vis spectral features and fluorescence spectral features. The UV-Vis spectral features include the absorbance ratio between 254 nm and 436 nm wavelengths, the absorbance at 350 nm wavelength, and the full width at half maximum (FWHM) of the characteristic peaks. The fluorescence spectral features include the fluorescence index and humification index calculated from protein-like fluorescence peaks and humic substance-like fluorescence peaks.

5. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Based on the aforementioned modal characteristics and a preset feature fusion enhancement algorithm, high-order index features of effluent pollution are generated, including: All submodal features are concatenated based on the target dimension to generate an initial fusion vector; Principal component analysis with domain knowledge constraints is used to reduce the dimensionality of the initial fusion vector and retain the principal components with a cumulative variance contribution rate greater than 95%. The domain knowledge is the knowledge of wastewater treatment. The corresponding constraint is that the weight of the pollution load index contained in the first principal component in the principal component analysis is greater than the preset constraint value. The pollution load index includes chemical oxygen demand concentration, ammonia nitrogen concentration and total phosphorus concentration. The initial fusion vector after dimensionality reduction is enhanced with features based on an autoencoder, and the domain knowledge is fused with the enhanced initial fusion vector to generate high-order exponential features of wastewater pollution. The autoencoder includes an input layer, an output layer, a bottleneck layer, and two hidden layers. The input of the input layer is the fusion vector after dimensionality reduction. The output layer is set with reconstruction error constraints. The bottleneck layer is used to generate the higher-order exponential features. The hidden layers include the ReLU activation function. The higher-order index features include the comprehensive pollution load index, the sludge microbial activity index, and the characteristic pollutant fingerprint; The comprehensive pollution load index is expressed as: ; in, The comprehensive pollution load index is used to characterize the pollution load level. This represents the standardized chemical oxygen demand. This indicates the maximum permissible concentration of chemical oxygen demand. This represents the standardized ammonia nitrogen concentration. This indicates the maximum permissible concentration of ammonia nitrogen. This represents the absorbance ratio between the 254nm and 436nm wavelengths. This represents the maximum reference value for the absorbance ratio. This represents the standardized number concentration of suspended solids. The weighting of chemical oxygen demand. The weighting of ammonia nitrogen concentration. The weighting of absorbance ratio is indicated. The weight representing the number concentration of suspended solids, and , , and satisfy ; The sludge microbial activity index is expressed as: ; in, The sludge microbial activity index is used to characterize the combined dissolved oxygen trend, sludge floc structure, and microbial metabolite characteristics, thereby determining the activity status of the microorganisms. This represents the slope of the linear regression of the standardized dissolved oxygen time series data. This represents the sludge floc density extracted from sensor data of visual modalities. This represents the standardized microbial protein characteristic values; The characteristic pollutant fingerprint is represented as follows: ; in, The fingerprint represents a characteristic pollutant, specifically a three-dimensional vector composed of spectral modal features. This indicates the molecular weight of organic compounds as characterized by the absorbance ratio between 254 nm and 436 nm. The fluorescence index is used to determine the source of organic matter, and the HIX index is used to determine the degree of decay of organic matter. Each type of pollutant corresponds one-to-one with the three-dimensional vector of the FPF.

6. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Based on the temporal arrangement, the higher-order exponential features are used to generate dynamic temporal features. According to preset effluent treatment process parameters and sensor data acquisition time, static auxiliary features and temporal auxiliary features of the dynamic temporal features are determined, including: The higher-order exponential features are arranged in time series using a preset historical time step to generate dynamic time series features that capture the dynamic change patterns of water quality and their association with multimodal sensor data. The preset effluent treatment process parameters are normalized to generate static auxiliary features that provide constraints for the process scenario. The effluent treatment process parameters include the volume of the biological treatment tank, the area of ​​the aeration zone, the residence time of the coagulation tank, the membrane pore size, and the sensor deployment node type. The sensor data acquisition time is processed by 1-hot encoding and normalization to generate a time-series auxiliary feature of the pollution load time periodic pattern. The dimensions of the sensor data acquisition time include hours, minutes, aquaculture status and weather status. The dynamic time-series features, static auxiliary features, and time-series auxiliary features are used as inputs to the pollution load index prediction model in tensor form, which is expressed as follows: ; in, Indicates input, Represents the set of real numbers. Indicates batch size. This indicates the historical time step. Dimensions representing dynamic temporal characteristics The dimension representing the static auxiliary feature. The dimension representing the temporal auxiliary features.

7. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, The preset pollution load index prediction model includes: The input embedding layer is used to map different types of dynamic temporal features, static auxiliary features, and temporal auxiliary features in the tensor form of input data to the same dimensional space; A variable selection network layer is used to filter the target information of the spliced ​​features output by the input embedding layer, and to compress the feature dimension to adapt to the position encoding layer and the temporal attention layer, wherein the target information is used to enhance the contribution of higher-order exponential features to the prediction results. The location encoding layer is used to inject time step location information into the output of the variable selection network layer, so that the prediction model can perceive the degree of influence of time sequence on historical data; The temporal attention layer is used to capture long-range dependencies in historical temporal data output by the positional encoding layer through a self-attention mechanism, so as to identify the correlation strength between the features of the target time point and the predicted target result. The gated fusion layer is used to dynamically balance the contribution ratio between the temporal attention features output by the temporal attention layer and the static auxiliary features broadcast within the model, and to capture the interaction dependency between the temporal attention features and the static auxiliary features. Then, based on the contribution ratio and interaction dependency, the fusion weight is determined and the fusion feature is output. A feedforward network layer is used to extract the implicit cross-modal correlations within the fusion features output by the gated fusion layer, so as to enhance the discriminative expression of the fusion features for higher-order exponential features in the prediction task. The decoder layer is used to generate a continuous change prediction curve of the pollution load index in future time periods from the output of the feedforward network layer by time step, wherein the time step process includes using the prediction result of the previous time step as the input of the current time step. The output prediction layer is used to convert the high-dimensional time-series features of the decoding layer output, which contain continuous concentration change curves, into prediction results corresponding to pollution load indicators. The prediction results are the concentration values ​​of pollution load indicators at any time in the future period.

8. The method for predicting pollution load in aquaculture wastewater according to claim 1, characterized in that, Based on the dynamic time-series characteristics, static auxiliary characteristics, time-series auxiliary characteristics, and the preset pollution load index prediction model, the predicted results of the pollution load index for future periods are output, including: The dynamic temporal features, static auxiliary features, and temporal auxiliary features are mapped to a high-dimensional space of the same dimension through corresponding embedding layers. After the static auxiliary features are broadcast to the temporal dimension, the features of the same dimension are concatenated to generate concatenated features. Each dimension of the spliced ​​features is assigned an importance weight and the features are sorted by feature priority. The weighted spliced ​​features are then subjected to linear dimensionality reduction and / or non-linear enhancement through projection to generate a standardized feature vector. The standardized feature vector is injected with time step position information using sine-cosine position coding, wherein the sine-cosine position coding is calculated as follows: ; ; The encoding dimension is the same as the standardized feature vector dimension. This represents the location code, and pos represents the time step index. This represents the encoding dimension, where i represents the dimension index; The standardized feature vector of position encoding is computed in parallel through a multi-head attention mechanism to generate multiple sets of local attention features corresponding to the number of attention heads. Then, multiple sets of local features are output through linear projection and scaled dot product attention calculation. The multiple sets of local features are integrated and restored according to the dimension of the standardized feature vector. Temporal attention features are output through residual connection and normalization processing. The static auxiliary features broadcast are averaged and pooled to obtain static global features. The input temporal attention features are processed by a gating unit with the static global features as the initial state, retaining the temporal output and the last step output of the gating unit. The last step output is input into a fully connected layer network and an activation function to determine the gating weights. After broadcasting the static global features to the target time step dimension, the temporal output and the adjusted dimension static global features are weighted, fused, and normalized based on the gating weights to output the fused features. The fused features are expanded in dimension by a first fully connected network to capture the implicit correlation features between multimodal features in the fused features, and then passed through a second fully connected network to retain the effective correlation features in the implicit correlation features. The fused features are then enhanced by using the effective correlation features through residual connections and normalization. Initialize the target time step based on the target future time period corresponding to the prediction result. Determine the initial predicted value of the fusion feature through a fully connected layer based on the last step feature. Extract the single-step feature fragments related to the current time step from the initial predicted value and the fusion feature and concatenate them into a single-step fusion feature. Input the single-step fusion feature into the autoregressive decoding unit to output the single-step prediction feature. Input the single-step prediction feature and the fusion feature into the cross-attention unit and calculate the single-step enhanced feature through scaling dot product attention. Map the single-step enhanced feature through a fully connected layer to determine the current concentration value. Update the predicted value of the previous time step to the current concentration value until the current time step meets the target time step. Concatenate all the single-step enhanced features to generate a high-dimensional temporal feature. The high-dimensional time-series features are reduced in dimensionality based on the feature projection layer, and the effective features of the associated predicted values ​​are enhanced by the activation function to output low-dimensional time-series features. The low-dimensional time-series features are then regressed using a linear activation function through a concentration prediction head, mapping the low-dimensional features to continuous predicted values ​​of pollution load index concentrations. The continuous predicted values ​​of pollution load index concentrations are then restored to true concentration values ​​through inverse normalization, and the true concentration values ​​of the pollution load indexes are output as the prediction results. The pollution load indexes include chemical oxygen demand concentration, ammonia nitrogen concentration, and total phosphorus concentration, and the future time period corresponds to the target time step of the pollution load indexes.

9. A pollution load prediction system for aquaculture wastewater, characterized in that, The system includes: The data acquisition module is configured to acquire sensor data of multiple modes from the tailwater collection point. The sub-modal types of the sensor data include physical sub-modal, visual sub-modal, and spectral sub-modal. The data fusion module is configured to extract the submodal features corresponding to each of the sensor data based on a preset feature extraction algorithm, and generate high-order index features of wastewater pollution based on the submodal features and the preset feature fusion enhancement algorithm. The data processing module is configured to generate dynamic time-series features from the higher-order exponential features based on the time-series arrangement, and to determine the static auxiliary features and time-series auxiliary features of the dynamic time-series features according to the preset tailwater treatment process parameters and the sensor data acquisition time. The data prediction module is configured to output the prediction results of the pollution load index in the future period based on the dynamic time series characteristics, static auxiliary characteristics, time series auxiliary characteristics and the preset pollution load index prediction model.

10. A pollution load prediction system for aquaculture wastewater, characterized in that, The system includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the aquaculture wastewater pollution load prediction method as described in any one of claims 1-8.