Intelligent regulation and control method for three-stage constructed wetland recirculating aquaculture system

By deploying multiple types of sensors in a three-level artificial wetland system and building a pollution evolution trend prediction model, combined with a reinforced learning control strategy network, the problem of inaccurate regulation in the existing technology is solved, accurate perception and real-time optimization of the dynamic environment are achieved, and the system's regulation efficiency and intelligence level are improved.

CN120406176AActive Publication Date: 2025-08-01YELLOW SEA FISHERIES RES INST CHINESE ACAD OF FISHERIES SCI

Patent Information

Application Number
CN202510912168.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The existing artificial wetland circulating aquaculture system has problems such as relying on experience, lagging regulation response, inaccurate control, and difficulty in achieving accurate perception and real-time optimization of the dynamic environment. Especially under complex and dynamically changing water quality conditions, the regulation effect fluctuates greatly.

Method used

By deploying multiple types of environmental sensors in a three-level artificial wetland system, building a multi-source heterogeneous data set, and using a pollution evolution trend prediction model to predict water quality, combined with a reinforced learning control strategy network, the optimal control parameters are generated to achieve multi-target optimization with the largest pollutant removal rate, stable water quality indicators and minimum system energy consumption.

Benefits of technology

It significantly improves the system's perception and prediction performance of complex dynamic water quality evolution laws, automatically identify key control parameters, improves the accuracy and efficiency of regulatory strategies, and achieves collaborative optimization of multiple goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406176A_ABST
    Figure CN120406176A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent regulation and control method for a three-stage constructed wetland recirculating aquaculture system, and belongs to the technical field of machine learning. The method comprises the following steps: firstly, continuously collecting water quality data and operation control data of each control unit, and constructing a multi-source heterogeneous data set under a unified time scale; then, constructing a pollution evolution trend prediction model, and capturing a dynamic evolution trend of water quality along with time and control behavior changes; then, under the guidance of a prediction result, analyzing the similarity of historical states and the sensitivity of regulation and control response, automatically identifying key control parameters which influence the water quality change of the system at present, and reasoning the dynamic adjustable boundary of the key control parameters; and finally, constructing a reinforcement learning strategy network fusing state prediction, a parameter boundary and a control target, realizing multi-target tradeoff among pollutant removal efficiency, a water quality standard-reaching rate and operation energy consumption, and outputting an efficient and steady control strategy through continuous interactive training. According to the invention, efficient, accurate and robust operation of the wetland system can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to an intelligent regulation method for a three-stage constructed wetland recirculating aquaculture system. Background Art

[0002] With the continuous improvement of the requirements for ecological environment protection, as an economical and efficient means of water quality purification and ecological restoration, constructed wetlands have been widely used in the fields of aquaculture, rural sewage treatment, and agricultural non-point source pollution control. Especially in the recirculating aquaculture system, the three-stage constructed wetland realizes the efficient removal of pollutants such as nitrogen, phosphorus, and organic matter by constructing different functional areas in series and parallel, ensuring the ecological safety of water body recycling. However, limited by complex environmental disturbances, frequent fluctuations in water quality indicators, and non-linear coupling of control parameters, current wetland systems generally have problems such as regulation processes relying on experience, adjustment response lags, and rigid control strategies, making it difficult to achieve accurate perception of the dynamic environment and real-time optimal scheduling. Against this background, constructing an intelligent management method that can integrate multi-source environmental information, dynamically identify the key states of the system, and have an adaptive regulation ability has become the core requirement for improving the operation efficiency and intelligent level of wetland systems.

[0003] Currently, the regulation of the constructed wetland recirculating aquaculture system mainly includes the following methods: (1) Manual experience regulation method: The manual experience regulation method is still widely used in many current small and medium-sized farms. It mainly relies on on-site operators to make adjustments based on empirical judgments and historical water quality monitoring data. For example, when observing that the water body becomes turbid or the odor intensifies, the operator subjectively judges whether it is necessary to increase the reflux ratio, extend the hydraulic retention time, or adjust the pump operation frequency. This method has the advantages of simple operation and low input cost, but the regulation highly depends on individual experience, lacks data support, and is prone to problems such as response lags, inaccurate control, and non-replicable operations. Especially when facing complex and dynamically changing water quality conditions, the regulation effect fluctuates greatly; (2) Single-parameter feedback control method: The single-parameter feedback control method is based on the real-time monitoring of a certain key water quality indicator (such as ammonia nitrogen concentration or dissolved oxygen concentration). When the monitored value exceeds the set threshold, the system automatically triggers corresponding control instructions, such as adjusting the water inflow or starting the aeration device, thereby forming a simple closed-loop control. This method has a certain degree of real-time and automation, and can achieve rapid response. However, its control logic usually only focuses on a single variable, ignoring the multi-factor coupling between water quality indicators and the overall operation state of the system, and is prone to control imbalance, insufficient adjustment, or system oscillation. Especially in the scenario of co-evolution of multiple pollutants, it is difficult to ensure the stability of the system; (3) Rule base + PLC automatic control method: This method is a relatively mature engineering regulation means, which manages the operation parameters of the wetland system through the preset "condition-action" logic (such as IF-THEN rules). A typical practice is: "If the ammonia nitrogen is greater than 1.5 mg / L and the water temperature is higher than 25°C, then increase the pump operation frequency to 80%." Such methods automatically execute control commands with the help of programmable logic controllers, and have strong engineering adaptability and execution efficiency. However, this method relies on fixed rules set in advance, lacks the ability to adapt to new scenarios, and rule updates rely on manual adjustment. It is difficult to dynamically optimize or explain the complex logic in the control process. Once the system state deviates from the scope of rule application, it may fail or be miscontrolled, especially in highly non-linear and multi-variable interference scenarios, the performance is limited. Summary of the Invention

[0004] In view of the above problems, the present invention proposes an intelligent regulation method for a three-stage constructed wetland recirculating aquaculture system, including the following steps: S1. Based on multiple types of environmental sensors deployed in the key structural units of the three-stage constructed wetland system and the aquaculture pond body, continuously collect water quality time-series data and control time-series data, and construct multi-source heterogeneous data under a unified time scale; S2. Input the water quality time-series data and control time-series data collected in S1 into the constructed and trained pollution evolution trend prediction model, and output multi-dimensional water quality prediction results; S3. Perform average aggregation on the multi-dimensional water quality prediction results in S2 in the time dimension to obtain a state vector to be evaluated; perform unsupervised clustering operations on the historical water quality time-series data set to obtain cluster centers and corresponding clusters; perform state similarity matching and response sensitivity analysis on the state vector to be evaluated and the cluster centers, obtain the key control parameters with the most regulatory value at the current moment, and infer its safe and adjustable dynamic response interval; S4. Based on the constructed reinforcement learning control strategy network, use the water quality time-series data, control time-series data and multi-dimensional water quality prediction results as environmental state inputs, use the key control parameters as adjustable actions, use the dynamic response interval as the action constraint, design a comprehensive reward function under the multi-objective optimization goals of the maximum pollutant removal rate, stable compliance of water quality indicators, and minimum system energy consumption, and generate optimal control parameter adjustment actions through the reinforcement learning strategy network.

[0005] Furthermore, the parameters of the water quality time-series data include pH value , dissolved oxygen concentration , ammonia nitrogen concentration , nitrite concentration , water temperature and turbidity ; The parameters of the control time-series data include influent flow 、 Effluent flow rate 、 Hydraulic retention time 、 Wetland module activation status 、 Valve opening degree 、 Pump operation frequency and Recirculation ratio ; The multi-dimensional water quality prediction label parameters include pollutant indicators ; Reaction environment indicators ; System disturbance indicators ; Together they constitute the water quality prediction label combination .

[0006] Further, the data collection methods for model training and obtaining key control parameters and their dynamically adjustable response intervals include: Input water quality time series data collection: For each control unit, at time point Collect the surrounding water quality data of all control units to obtain the water quality data at this time point , where represents the water quality data collected at the time point, at the th control unit, and ; Based on this principle, a total of time points of complete input water quality time series data are collected , and ; Input control time series data collection: At time point Collect the control data corresponding to all control units ; Among them, represents the control data collected at the time point, at the th control unit, and ; Similarly, collect time points of complete input control time series data , and .

[0007] Further, the pollution evolution trend prediction model includes a water quality dynamic perception channel, a control parameter coding channel, and a fusion decoding prediction channel; The water quality dynamic perception channel introduces a three-branch parallel structure, with different branches designed with different receptive fields to enhance the sensitivity of the model to the changing trends of various water quality characteristics over time; The input is the collected water quality time series data, capturing short-term and medium- to long-term trends, and generating water quality dynamic coding features ; The input of the control parameter encoding channel is the collected control timing data. A gated fusion mechanism is introduced to perform weighted integration on multiple control dimensions, dynamically adjust the representational contributions of different control variables in specific states, and output gated fusion features. ; Subsequently, the control fusion features are fed into the first enhanced GRU unit for preliminary timing modeling; then, the output of the first GRU unit is concatenated with the gated fusion features and fed into the second enhanced GRU unit to depict the deep interaction relationships of control variables under state combination conditions; the output result of the second GRU unit is processed by layer normalization to obtain the final features of the first stacked unit ; Based on the above structure, five stacks are performed to extract multi-layer timing features from local response to global regulation layer by layer. Finally, the representational features of the control parameter encoding channel are obtained through the Sigmoid activation function. ; The fusion decoding prediction channel concatenates and fuses the water quality dynamic encoding features and the representational features of the control parameter encoding channel on the channel dimension, and outputs multi-dimensional water quality prediction results based on the dual Transformer decoding layer - gated fusion structure. .

[0008] Furthermore, the specific structure of the water quality dynamic perception channel includes: Short-term perception branch: The water quality timing data first passes through dilated convolution blocks for local fluctuation feature extraction, and then passes through dilated convolution blocks to further expand the short-term context perception range, obtaining short-term dynamic features ; Long-term perception branch: The water quality timing data sequentially passes through dilated convolution blocks and dilated convolution blocks, using a larger dilation rate to capture the trend evolution law across time slices, for modeling the cumulative impact of operating condition changes on pollutants, and outputting medium-term evolution features ; Original retention branch: Retain the original water quality timing data as a residual connection path to maintain low-level detail information, forming basic retention features ; After that, the output features of the three branches are concatenated on the channel dimension to form fusion features , and it is sent to the global average pooling layer to enhance the global representation ability. The fusion result is then input into the next round of the three-branch parallel structure for deep feature extraction. The above structure is stacked repeatedly for a total of 5 layers to construct a progressive multi-scale temporal modeling path. After the output of the final layer, a Sigmoid activation function is connected to enhance the non-linear expression ability, and finally, water quality dynamic coding features are generated. .

[0009] Furthermore, the specific structure of the fusion decoding prediction channel includes: First, the water quality dynamic coding features are concatenated with the representation features of the control parameter coding channel in the channel dimension. Subsequently, the concatenated features are sent into the gated fusion mechanism to obtain the interaction features ; The interaction features are sequentially input into two consecutive Transformer decoding layers to capture the long-range dependence relationship and potential sequence structure between multi-dimensional features. Each Transformer decoding layer includes a multi-head attention mechanism, a feed-forward neural network, and a residual connection structure. After that, the Transformer decoding result is sent into the gated fusion mechanism again; The dual Transformer decoding layer - gated fusion constitutes a decoding module unit. This decoding module unit is stacked repeatedly for 5 layers to form a multi-layer information abstraction and trend inference path. Finally, it is sent into the Dropout layer to enhance the generalization ability, and then processed by the Linear mapping layer and the activation function to output multi-dimensional water quality prediction results .

[0010] Furthermore, the specific process of S3 includes: Construction of the state vector to be evaluated: For the output multi-dimensional water quality prediction results , average aggregation is performed on each control unit in the time dimension to obtain the state vector to be evaluated at the current regulation moment ; Aggregation and analysis of historical state data: Based on the constructed water quality time series dataset, the water quality state data at each single moment in the past is extracted , and a historical water quality state sample set with a scale of is formed. Unsupervised clustering is performed on all historical water quality state vectors to obtain cluster centers , where represents the th cluster center, and ; Next, calculate the Euclidean distance between the current state vector to be evaluated and all cluster centers, and obtain a set of distance vectors ; among them represents the state vector to be evaluated and the Euclidean distance from the -th cluster center; After that, select the cluster center with the closest distance , and extract all historical water quality state data in its corresponding cluster as the candidate sample set; further calculate the cosine similarity between all samples in the cluster and ; finally, select the top historical sample water quality time series data with the highest similarity and the corresponding control time series data , and form a representative control sample set for response analysis ; Response sensitivity evaluation and key control parameter screening: For each set of control parameter data in the control sample set , conduct perturbation experiments respectively to calculate the sensitivity coefficient corresponding to each control parameter ; among them, and respectively represent the -th control parameter and its sensitivity coefficient; Finally, take the parameters that satisfy the sensitivity coefficient greater than the preset threshold as the key control parameter set at the current regulation moment ; Based on the key control parameter set, infer its safe and adjustable dynamic response range.

[0011] Furthermore, the process of inferring its safe and adjustable dynamic response range based on the key control parameter set is as follows: For each selected key parameter , extract the maximum and minimum values from its historical values in the control sample set to form the preliminary response range of the key parameter : , where respectively represent the minimum and maximum values of the parameter in the control sample set; After that, combine the system preset operation safety constraints to perform double-boundary clipping to obtain the dynamically adjustable response range of this parameter: ; Among them, are respectively the minimum and maximum values preset by the system for the parameter , and ​For maximum and minimum value operations; Is a key parameter The corresponding dynamically adjustable response interval, and form a structured control parameter-response interval pair ; Based on obtaining The dynamically adjustable response intervals corresponding to all key parameters in the set, to obtain a response interval set ; The complete set of control parameter-response interval pairs is , where , and .

[0012] Furthermore, the reinforcement learning control policy network constructed in S4 adopts an Actor-Critic dual-head architecture; First, splice the current water quality time series data , the control time series data and the predicted state into a unified state vector, and send it to the network input layer for standardization and feature embedding to extract basic information; The feature extraction layer consists of three fully connected networks and ReLU activation units, which perform high-dimensional feature transformation on the state vector to extract deep control signal features; The policy head constructs Gaussian distribution value outputs for each key control parameter in the action , finally outputs the expected value and corresponding variance of the control action, and obtains the corresponding value interval; that is, the value interval for each potential key variable is: the value interval of the influent flow rate , the value interval of the effluent flow rate , the value interval of the hydraulic retention time , the discrete value interval of the wetland module state , 1 means enabled, 0 means disabled; the value interval of the valve opening ; the value interval of the pump frequency , the value interval of the reflux ratio ; Then, further limit the action constraints of each key variable through the dynamic response interval and perform sub-sampling on all key parameter value intervals to obtain a set of potential actions ; The value head calculates the long-term cumulative expected reward that can be obtained by respectively executing the potential actions , so as to guide network optimization; the value head consists of two fully connected networks, and as the Critic end in the Actor-Critic framework, it provides a gradient estimation basis for policy optimization, guiding the policy network to converge to a better solution.

[0013] Further, the environmental state is jointly composed of the water quality time series data, control time series data, and multi-dimensional water quality prediction results of the control system at the current moment; it includes the current water quality data , which is the water quality time series data of all control units at time , the real-time control data , which is the executed control time series data corresponding to all control units at time ; the predicted state , which is the multi-dimensional water quality prediction results at all control units in the future steps ; Among them, the action is the adjustment amount of the key control parameters for each control unit ; the value range of the action is restricted by the dynamic response interval ; Among them, the reward function is a weighted multi-objective comprehensive reward mechanism, which encourages efficient removal of pollutants, stable compliance of indicators, and minimization of control energy consumption, that is, the multi-objective comprehensive reward function: ; Among them, is the pollutant removal rate, defined as the standardized average value of the changes in ammonia nitrogen and nitrite concentrations per unit time; is the water quality compliance rate, indicating the weighted accumulation of the number of times all water quality parameters meet the standards in each control unit; is the total system energy consumption, comprehensively represented by the normalized average values of the pump frequency and valve opening.

[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) Pollution evolution trend prediction mechanism based on dual-channel fusion modeling: A joint prediction structure integrating a "water quality dynamic perception channel" and a "control parameter coding channel" is designed to achieve forward-looking modeling of the pollutant concentration change trend; the expression ability of the model for mutations and lag responses is improved through parallel multi-scale dilated convolution and gated GRU structures, effectively enhancing the system's perception and prediction performance of complex dynamic water quality evolution laws; (2) Key control parameter screening method for state similarity and response sensitivity: A parameter screening strategy combining a clustering-matching mechanism and perturbation sensitivity analysis is proposed, which can automatically identify the key control parameters that are most sensitive to the current pollution evolution without introducing artificial experience, and infer their dynamic response intervals in combination with historical regulation boundaries, significantly improving the strategy search efficiency and the accuracy of control parameter selection; (3) Reinforcement learning control framework integrating prediction guidance and strategy optimization: A control strategy network based on the Actor-Critic structure was constructed, and the water quality prediction results were introduced into the state definition process to dynamically enhance the strategy input; at the same time, a multi-objective weighted reward function and a PPO optimization mechanism were adopted to jointly optimize the pollutant removal efficiency, water quality compliance, and energy consumption control, achieving efficient convergence of the control strategy and multi-objective coordination. Description of the Drawings

[0015] Figure 1 This is the overall flowchart of the intelligent control method of the present invention.

[0016] Figure 2 This is the network structure diagram of the pollution evolution trend prediction model.

[0017] Figure 3 This is the flowchart for screening key control parameters and generating response intervals.

[0018] Figure 4 This is the RMSE comparison result graph of different models in the present invention's embodiment for the pollutant prediction task.

[0019] Figure 5 This is the MAE comparison result graph of different models in the present invention's embodiment for the pollutant prediction task.

[0020] Figure 6 This is the heat map of the sensitivity of the key control parameters in the present invention's embodiment. Detailed Embodiments

[0021] The present invention proposes an intelligent control method for a three-stage constructed wetland recirculating aquaculture system. The overall technical path is as Figure 1 shown: S1. Construction of a multi-source environmental time-series data set: Deploy multiple types of environmental sensors in the key structural units and aquaculture ponds of the three-stage constructed wetland system to continuously collect water quality time-series data and control time-series data, and construct multi-source heterogeneous data under a unified time scale; provide comprehensive data support for the training and inference of subsequent models; S2. Construction of a pollution evolution trend prediction model: Input the water quality time-series data and control time-series data collected in S1 into the constructed and trained pollution evolution trend prediction model, and output multi-dimensional water quality prediction results, including pollutant indicators, reaction environment indicators, and system disturbance indicators; used for screening key control parameters in S3 and forward guidance of the control strategy in S4; S3. Key control parameter screening and response interval generation: Guided by the prediction results in S2, combined with the historical operation database, perform state similarity matching and response sensitivity analysis to identify the key control parameters with the most regulatory value at the current moment, and infer their dynamically adjustable and safe response intervals, providing parameter boundaries and search space constraints for subsequent strategy optimization; S4. Generation of optimal regulation strategy driven by reinforcement learning: Based on the multi-dimensional water quality prediction results output by the S2 module, the key control parameters determined by the S3 module, and their dynamically adjustable response intervals, construct a reinforcement learning control strategy network; Take the current system state (including real-time water quality, control data, and prediction information) as the environmental state input, use the key control parameters as adjustable actions, design a comprehensive reward function under the multi-objective optimization goals of maximizing pollutant removal rate, stably meeting water quality indicators, and minimizing system energy consumption, and generate optimal control parameter adjustment actions through the reinforcement learning strategy network to achieve intelligent collaborative scheduling and dynamic regulation strategy output for multiple control units.

[0022] The invention will be further described below in conjunction with specific embodiments.

[0023] I. Construction of multi-source environmental time series dataset The present invention aims to achieve dynamic prediction and intelligent regulation of key water quality indicators in wetland systems. Therefore, a standardized multi-source time series dataset integrating water quality parameters and operation status parameters is first constructed. The specific steps include: 1. Selection of water quality parameters and operation control parameters: According to the operation characteristics of the wetland system and the water quality evolution law, select the parameter types that have important impacts on pollutant changes and water body purification processes to form a water quality-operation joint parameter combination; (1) Selection of water quality parameters: Based on the "Surface Water Environment Quality Standard" and the pollutant transformation law in the water body purification process of typical constructed wetlands, select six of the most representative water quality monitoring indicators, including: pH value , dissolved oxygen concentration , ammonia nitrogen concentration , nitrite concentration , water temperature and turbidity ; Together, they form the water quality parameter combination ; (2) Selection of control parameters: For the common adjustment means and energy efficiency management mechanisms in the operation control of wetland systems, select seven parameters closely related to hydraulic process control and facility start-stop status, including: influent flow , effluent flow , hydraulic retention time , wetland module activation status , valve opening , pump operation frequency With the reflux ratio ; jointly constitute the operating parameter combination ; 2. Multi-structure unit input data acquisition: The wetland system is usually composed of multiple functional modules connected in series and parallel. There are spatial differences in the processing capabilities and water quality responses of each module. Therefore, corresponding environmental sensors are arranged at each main control unit (a total of units) to collect water quality parameter data. At the same time, the corresponding operating parameters of each control unit are monitored; specifically including: (1) Input water quality time series data acquisition: For each control unit, at the time point collect the surrounding water quality data of all control units to obtain the water quality data at this time point , where represents the water quality data collected at the th time point at the [[ID=2]]th control unit, and ; Based on this principle, a total of complete input water quality data at time points are collected , and ; (2) Input control time series data acquisition: At the time point collect the control data corresponding to all control units ; among them, represents the control data collected at the th time point at the th control unit, and ; Similarly, collect complete input control data at time points , and ; Therefore, the input data of the multi-source environmental time series dataset includes: the water quality data at all control units and the control data corresponding to all control units ; 3. Water quality prediction labels: The purposes of model prediction include: (1) to determine key control parameters; (2) to enhance the reinforcement learning state; therefore, the water quality prediction labels need to be sensitive to control parameters and cover the potential risk trends of the system. Therefore, the selected prediction label parameters include pollutant indicators ; reaction environmental indicators ; system disturbance indicators ; jointly constitute the water quality prediction label combination ; After that, set the prediction time step to , and collect the water quality data at all control units at the time point ​, where represents the water quality prediction label data collected at the th control unit at the time point, and , ; Finally, collect the water quality prediction label data corresponding to all time points to obtain the complete water quality prediction label , and use it as the output data of the multi-source environmental time series dataset; 4. Construction of multi-source environmental time series dataset: Align the collected input data and , as well as the constructed water quality prediction label output data one by one, and combine them to form a complete multi-source environmental input-output data pair. Based on the above method, a total of groups of input-output samples with a unified time reference are collected, and finally a multi-source environmental time series dataset is constructed.

[0024] II. Design of pollution evolution trend prediction model To realize the prediction and intelligent regulation of key water quality indicators in the wetland system, the present invention designs a pollution evolution trend prediction model for multi-source input and supporting multi-step prediction on the basis of constructing a multi-source environmental time series dataset; the model includes a water quality dynamic perception channel, a control parameter coding channel, and a fusion decoding prediction channel; the model structure is as Figure 2 shown; 1. Design of water quality dynamic perception channel To effectively capture the dynamic evolution law of water quality indicators in the wetland system at multiple time scales, the present invention designs a water quality dynamic perception channel to enhance the sensitivity of the model to the changing trends of various water quality characteristics over time. The input of this channel is the collected water quality time series data ; First of all, to fully explore the time series dependence relationship at different time scales, this channel introduces a three-branch parallel structure, and different branches have different receptive field designs to capture short-term and medium- to long-term trends. The structure is as follows: (1) Short-term perception branch: First, pass through the dilated convolution block for local fluctuation feature extraction, and then pass through the dilated convolution block to further expand the short-term context perception range, so as to depict the response pattern of water quality parameters under short-term perturbations, and finally obtain the short-term dynamic feature ; (2) Long-term perception branch: Sequentially pass through the dilated convolution block and The dilated convolutional block captures the trend evolution law across time slices with a larger dilation rate, which is used to model the cumulative impact of operating condition changes on pollutants and outputs the medium-term evolution features ; (3) The original retention branch: retains the original water quality time series data as the residual connection path, maintains the low-level detailed information, and forms the basic retention features ; After that, the output features of the three branches are concatenated in the channel dimension to form the fused features , and are fed into the Global Average Pooling (GAP) layer to enhance the global representation ability. The fused result is then input into the next round of the three-branch parallel structure for deep feature extraction; the above structure is stacked repeatedly for a total of 5 layers to construct a progressive multi-scale time series modeling path; after the output of the final layer, a Sigmoid activation function is connected to enhance the non-linear expression ability, and finally the water quality dynamic coding features are generated ; Among them, The dilated convolutional block consists of the following three parts: First is the , with a dilation rate of Dilated TCN layer, which is used to expand the receptive field while maintaining the computational efficiency, realize non-uniform time series dependence modeling, and then obtain the feature processing result through layer normalization and ReLU activation function; Therefore, this structure supports multi-scale parallel perception and is suitable for multi-layer extraction and dynamic modeling of time series features in complex environments; 2. Design of the control parameter coding channel Aiming at the characteristics of diverse dimensions and obvious regulation delay effects of operating parameters in the wetland system, an invention designs a control parameter response modeling channel to enhance the dynamic expression ability and regulation sensitivity of the model to key operating state changes. The input of this channel is the collected operating control time series data ; First, a gated fusion mechanism (Gated Fusion Module) is introduced to make full use of the combined effects of various operating parameters at different time periods. Through this module, multiple control dimensions are weighted and integrated, so as to dynamically adjust the representation contributions of different control variables in a specific state and output the gated fusion features ; Subsequently, the control fusion features are fed into the first enhanced GRU unit for preliminary time series modeling, so as to use the time series network with a memory mechanism to cope with the non-linear dependence characteristics in the operation process of the wetland system; After that, to further enhance the joint modeling ability between the control behavior and the system response, the output of the first GRU unit is combined with the gated fusion features It is cascaded and fed into the second enhanced GRU unit to characterize the deep interaction relationship of the control variables under the state combination conditions; the output result of the second GRU unit is processed by layer normalization to obtain the final features of the first stacked unit ; To achieve a progressive understanding and deep abstract expression of the evolution law of control parameters, this module is stacked five times based on this structure, layer by layer extracting multi-layer temporal features from local response to global regulation, and finally obtaining the characterization features of the control parameter encoding channel through the Sigmoid activation function ; Among them, the enhanced GRU unit consists of the following parts: First, use the standard GRU network to implement time series dependence modeling, then embed the temporal pooling module to enhance the representation ability for sudden and slow-changing control modes, and finally combine the ReLU activation function to enhance the non-linear modeling ability for complex control responses; this structure can improve the compressed expression of control inputs while maintaining the coherence of the time structure; 3. Design of the fusion decoding prediction channel To achieve a high-precision prediction of the future water quality state of the wetland system, this module designs a fusion decoding prediction channel for jointly modeling the deep association between water quality evolution features and operation control dynamics; First, the water quality dynamic coding features and the characterization features of the control parameter encoding channel are concatenated in the channel dimension; subsequently, the concatenated features are fed into the gated fusion mechanism to adaptively adjust the contribution weights of water quality and control information and strengthen their key interaction paths to obtain the interaction features ; On this basis, the interaction features are sequentially input into two consecutive Transformer decoding layers to capture the long-distance dependence relationship and potential sequence structure between multi-dimensional features; each Transformer decoding layer contains a multi-head attention mechanism, a feed-forward neural network, and a residual connection structure; then, the Transformer decoding result is fed into the gated fusion mechanism again to further enhance the response consistency and prediction reliability of the key feature dimensions; The above structure (double Transformer decoding layer - gated fusion) constitutes a decoding module unit with feature collaborative processing ability; and this decoding module unit is stacked 5 layers repeatedly to form a multi-layer information abstraction and trend inference path, and finally fed into the Dropout layer to enhance the generalization ability, and then processed by the Linear mapping layer and the activation function to output multi-dimensional water quality prediction results ; 4. Model Training: Based on the multi-source environmental time-series dataset constructed, the present invention uses the Mean Squared Error (MSE) as the loss function to measure the water quality prediction results output by the model and the actual water quality labels The deviation between them; the model parameters are iteratively updated through the SGD (Stochastic Gradient Descent) algorithm, and the training is terminated after reaching the preset maximum number of training epochs, and finally a prediction model with the ability to model the pollution evolution trend is obtained.

[0025] III. Screening of Key Control Parameters and Generation of Response Intervals To achieve the balance between the regulation efficiency and execution accuracy of the three-stage constructed wetland circulation system in dynamic control, the present invention proposes a method for screening key control parameters and generating response intervals based on historical state similarity and response sensitivity analysis; this process is guided by the multi-dimensional water quality prediction results output by the model, and the overall process is as Figure 3 shown; 1) Construction of the state vector to be evaluated: For the multi-dimensional water quality prediction results output by the model , perform average aggregation on each control unit in the time dimension, so as to obtain the state vector to be evaluated at the current regulation moment ; 2) Aggregation analysis of historical state data: Based on the water quality data in the constructed multi-source environmental time-series dataset, extract the water quality state data at each single moment in the past ( ), and form a historical water quality state sample set with a scale of ; then, to improve the similarity matching efficiency, perform unsupervised clustering on all historical water quality state vectors to obtain clustering centers , where represents the th clustering center, and ; ; Next, calculate the Euclidean distance between the current state vector to be evaluated and all clustering centers, and obtain a set of distance vectors ; where represents the Euclidean distance between the state vector to be evaluated and the th clustering center; After that, select the clustering center with the closest distance , and extract all the historical water quality state data in its corresponding cluster as the candidate sample set; further calculate the cosine similarity between all the samples in the cluster and ; finally, select the control parameter data corresponding to the top historical water quality sample data with the highest similarity , which constitutes a representative control sample set for response analysis ; 3) Response sensitivity assessment and key control parameter screening: for control sample sets Each set of control parameter data in , which corresponds to the specific values of 7 control parameters (including water flow , water flow , hydraulic retention time , wetland module enabled status , valve opening , Pump operating frequency Reflux ratio ); Afterwards, disturbance experiments were performed to calculate each control parameter The corresponding sensitivity coefficient Specifically, the gradient approximation is used to calculate the cumulative response intensity of each control parameter as the sensitivity coefficient ;in, and Respectively represent control parameters and their sensitivity coefficients, and , corresponding to the above 7 control parameters respectively; Finally, the sensitivity coefficient will be satisfied Greater than the preset threshold The parameters are used as the key control parameter set at the current control moment The parameters in this set are considered to have the most significant impact on the current water quality evolution trend and are the core variables in the design of subsequent control strategies; 4) Dynamically adjustable response interval reasoning: For each selected key parameter , from the control sample set Extract the maximum and minimum values from the historical values to form key parameters Initial response range: ,in Represents parameters respectively The minimum and maximum values in the control sample set; Afterwards, combined with the system preset operation safety constraints Perform double boundary clipping to obtain the dynamically adjustable response range of this parameter: ; in, Parameters for the system The preset minimum and maximum values, and To obtain the maximum and minimum value operations; Key parameters The corresponding dynamically adjustable response intervals, and form a structured control parameter-response interval pair ; Based on this method, all the dynamically adjustable response intervals corresponding to the key parameters in the set are obtained, and a set of response intervals is obtained ; Therefore, the complete set of control parameter-response interval pairs is , where , and .

[0026] IV. Generation of Optimal Regulation Strategies Driven by Reinforcement Learning Based on the predicted results of future water quality output and the screened key control parameters and their response intervals , a network for generating optimal strategies driven by reinforcement learning (RL) is constructed to dynamically output an optimal regulation action sequence for multiple control units, realizing multi-objective collaborative optimization of maximizing pollutant removal efficiency, stably meeting water quality indicators, and minimizing system operation energy consumption; 1. Definition of the reinforcement learning framework: The core is a reinforcement learning controller based on policy optimization, and its key elements are designed as follows: (1) Definition of the state: The environmental state is jointly composed of the water quality time series data, control time series data, and multi-dimensional water quality prediction results of the control system at the current moment. Specifically, it includes: the current water quality data , which is the water quality data of all control units at time . The real-time control data , which is the executed control data corresponding to all control units at time ; The predicted state , which is the predicted water quality results at all control units for the future steps output, used to assist forward-looking regulation; (2) Definition of the action : It is the adjustment amount for the key control parameter of each control unit. The value range of the action is restricted by the dynamically adjustable response interval provided by the S3 module; (3) Reward function : Design a weighted multi-objective comprehensive reward mechanism to encourage efficient pollutant removal, stable achievement of indicators, and minimization of control energy consumption. That is, the multi-objective comprehensive reward function: ; Among them, is the pollutant removal rate, defined as the normalized average value of ammonia nitrogen and nitrite concentration changes per unit time; is the water quality compliance rate, representing the weighted cumulative number of times all water quality parameters meet the standards in each control unit; is the total system energy consumption, comprehensively represented by the normalized average value of pump frequency and valve opening; 2. Policy architecture network design: To enhance the expression ability and generalization performance of policy generation, the policy network structure adopts the Actor-Critic dual-head architecture; specifically: (1) First, concatenate the current water quality data , real-time control data , and the predicted state into a unified state vector, and send it to the network input layer for normalization and feature embedding to extract basic information; (2) The feature extraction layer consists of three fully connected networks and ReLU activation units, which perform high-dimensional feature transformation on the state vector to extract deep control signal features; (3) The policy head constructs Gaussian distribution value outputs for each key control parameter in the action , finally outputs the expected value and corresponding variance of the control action, and obtains the corresponding value range; that is, the value range of each potential key variable is: the value range of the influent flow rate , the value range of the effluent flow rate , the value range of the hydraulic retention time , the discrete value range of the wetland module state , 1 means enabled, 0 means disabled; the value range of the valve opening ; the value range of the pump frequency , the value range of the reflux ratio ; then, further limit the action constraints of each key variable through the dynamic response interval and perform times of sampling on all key parameter value ranges, so as to obtain a set of potential actions ; (4) The value head calculates the long-term cumulative expected reward that can be obtained by respectively executing the potential action , so as to guide network optimization; specifically, the value head consists of two fully connected networks, and as the Critic end in the Actor-Critic framework, it provides a gradient estimation basis for policy optimization and guides the policy network to converge to a better solution; 3. Design of Policy Network Training Mechanism: To achieve robust training and efficient convergence of the policy network in the complex wetland system operation environment, an end-to-end reinforcement learning training mechanism based on the Actor-Critic architecture is constructed, and the improved Proximal Policy Optimization (PPO) algorithm is used to optimize the parameters of the policy network. The specific training and implementation technical route is as follows: (1) The system constructs the environmental interaction process through the simulator, and uses the current state (including real-time water quality data , control data and future water quality prediction information ) to generate control actions , and obtains rewards and the next state according to the environmental feedback to construct the training trajectory; the policy network continuously optimizes the action output according to the principle of maximizing the cumulative reward; (2) The training objectives include two parts: the policy head uses the PPO loss function in the form of the clipping probability ratio to improve the policy stability; while the value head estimates the expected return of the current state by minimizing the temporal difference error; both jointly drive the policy to converge in the direction of high pollutant removal rate, good water quality stability and low energy consumption; (3) Policy entropy regularization is introduced during the training process to maintain policy exploration, and an early stopping criterion is set to prevent overfitting, further improving the convergence speed and generalization ability of the policy; (4) After the policy network training is completed, the system can output the optimal adjustment action combination of the subsequent key control parameters according to the environmental state in each control cycle , and finally, the mean value of the continuous sampling action set output by the policy is used as the best control action of the system , realizing the collaborative scheduling and dynamic optimization of multiple control units.

[0027] V. Experimental Result Analysis To verify the effectiveness of the "Intelligent Regulation Method for the Three-Level Constructed Wetland Recirculating Aquaculture System" of the present invention, two types of comparative experiments are designed, corresponding to: (1) Evaluation of the prediction accuracy of the pollution evolution trend; (2) Sensitivity analysis of key control parameters. The comparative models include the traditional time series modeling methods LSTM and GRU to highlight the prediction advantages of the multi-modal modeling structure of the present invention and the control mechanism interpretation ability.

[0028] 1. Evaluation of the Prediction Accuracy of Pollutant Concentrations Based on the constructed multi-source environmental time-series dataset, the LSTM, GRU, and the pollution evolution trend prediction model of the present invention are respectively used to perform multi-step prediction on the changes in water quality data in the future period. All prediction and control experiments are modeled and evaluated based on the monitoring data of 8 typical control units. Each unit covers different wetland structure modules, with strong heterogeneity and different operation responses, which can effectively verify the adaptability and stability of the proposed method in multi-condition and multi-structure scenarios. The error between the model output result and the true observed value is compared, and the evaluation indicators include the Root Mean Square Error (RMSE) and the Mean Absolute Error (MAE). The results are as Figure 4 and Figure 5 shown; The statistical results show that the method of the present invention achieves the lowest error level for all pollutant indicators. Among them, the RMSE is overall controlled between 0.098 - 0.114, and the average value is about 0.106; the MAE is stable between 0.078 - 0.088, and the average value is about 0.083, which is significantly better than the comparison models. In contrast, the RMSE range of the LSTM model is 0.132 -  0.150, and the MAE is 0.101 - 0.113; the GRU model has higher errors, with the RMSE reaching 0.157 - 0.176 and the MAE being between 0.122 - 0.134. In the medium and high complexity scenarios of multiple control units, the proposed method is superior to the traditional sequence models in terms of error volatility and convergence stability. This verifies the effectiveness and adaptability of the multi-scale perception structure and fusion decoding mechanism in the proposed model in processing multi-source environmental time-series data, and has higher generalization ability and engineering practical value.

[0029] 2. Analysis of the sensitivity heat map of key control parameters To further verify the scientificity of the key control parameter screening mechanism, a control parameter perturbation experiment is designed to quantitatively calculate the response intensity of seven operating parameters under multiple regulation states. The selected parameters include: (1) influent flow rate; (2) effluent flow rate; (3) hydraulic retention time; (4) wetland module activation status; (5) valve opening; (6) pump operating frequency; (7) reflux ratio; The sensitivity coefficient of each parameter is measured by the average influence degree of its perturbation on the change of pollutant prediction value, and the results are summarized as a sensitivity heat map, as Figure 6 shown; ]> Figure 6 The sensitivity heat map shown is statistically calculated based on the historical operation data of 8 control units during the complete regulation cycle. The sensitivity coefficient of each control parameter is quantified by the average influence of its perturbation on the change of pollutant concentration, and is shown in the figure in the form of a heat map.

[0030] As can be observed from the results in the figure, there are certain fluctuations in the sensitivity performance of different control parameters under different units, indicating that the system has local dependencies and state coupling relationships in its response to various parameters. Overall, the "reflux ratio" and "pump operating frequency" exhibit relatively higher sensitivity levels in most units, and the sensitivity coefficients of some units even exceed 0.8, verifying their core status in the system stability control. In addition, under specific regulation scenarios, parameters such as "inlet flow rate" and "valve opening" also exhibit characteristics of high local response, suggesting their phased regulatory effects on the system dynamic changes. Therefore, the control parameter sensitivity analysis method proposed in the present invention can dynamically identify the most valuable parameter combinations for regulation by combining the system historical state distribution and pollutant response trends, avoiding redundant control inputs, and improving the efficiency and operability of strategy deployment.

[0031] Therefore, through the experiments on pollutant prediction accuracy evaluation and control parameter sensitivity analysis in the system design, the modeling ability and intelligent regulation performance of the present invention in a multi-source heterogeneous environment are comprehensively verified. In the prediction link, this method demonstrates better error control ability and stability than traditional models under multiple control units and complex working conditions, fully reflecting the advantages of its multi-scale perception structure and fusion decoding mechanism in dealing with the dynamic evolution of pollutants; while in the control link, through parameter perturbation analysis and heatmap visualization, the operating parameter combinations with key regulatory values are accurately identified. In particular, the "reflux ratio" and "pump operating frequency" show significant influence in most system states, verifying the interpretability and adaptability of the method for obtaining key control parameters in complex regulation scenarios.

[0032] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0033] Although the specific implementation manners of the present invention are described above, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. An intelligent regulation method for a three - stage artificial wetland circulating water aquaculture system, characterized in that, It includes the following steps: S1. Based on multi-type environmental sensors deployed for the key structural units of the three-stage constructed wetland system and the aquaculture pond body, continuously collect water quality time-series data and control time-series data, and construct multi-source heterogeneous data under a unified time scale; S2. Input the water quality time-series data and control time-series data collected in S1 into the constructed and trained pollution evolution trend prediction model, and output multi-dimensional water quality prediction results; S3. Perform average aggregation on the multi-dimensional water quality prediction results in S2 in the time dimension to obtain the state vector to be evaluated; perform unsupervised clustering operations on the historical water quality time-series data set to obtain the clustering centers and corresponding clusters; perform state similarity matching and response sensitivity analysis on the state vector to be evaluated and the clustering centers, obtain the key control parameters with the most regulatory value at the current moment, and infer their safe and adjustable dynamic response intervals; S4. Based on the constructed reinforcement learning control policy network, use the water quality time-series data, control time-series data, and multi-dimensional water quality prediction results as environmental states as input, use the key control parameters as adjustable actions, use the dynamic response interval as the action constraint, design a comprehensive reward function under the multi-objective optimization goals of the maximum pollutant removal rate, stable compliance of water quality indicators, and minimum system energy consumption, and generate optimal control parameter adjustment actions through the reinforcement learning policy network.

2. The intelligent regulation method for a three-stage constructed wetland recirculating aquaculture system according to claim 1, characterized in that: The parameters of the water quality time series data include pH value , dissolved oxygen concentration , ammonia nitrogen concentration , nitrite concentration , water temperature and turbidity ; The parameters for controlling the timing data include the influent flow rate , the effluent flow rate , the hydraulic retention time , the wetland module enabling status , the valve opening degree , the pump operating frequency and the reflux ratio ; The multi-dimensional water quality prediction label parameters include pollutant indicators ; reaction environmental indicators ; system disturbance indicators ; Collectively form a water quality prediction label combination .

3. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 2, characterized in that: The data collection method for the data set used for model training and obtaining the key control parameters and their dynamically adjustable response intervals includes: Input water quality time series data collection: For each control unit, at time point Collect the surrounding water quality data of all control units to obtain the water quality data at this time point , where Indicates at time point , the water quality data collected at the -th control unit, and ; Based on this principle, a total of complete input water quality time series data at time points are collected , and ; Input control timing data acquisition: At time point Collect the control data corresponding to all control units ; among them, Indicates the control data collected at the th control unit at time point , and ; similarly, collect the complete input control timing data at time points , and .

4. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 1, characterized in that: The pollution evolution trend prediction model includes a water quality dynamic perception channel, a control parameter encoding channel, and a fusion decoding prediction channel; The water quality dynamic perception channel introduces a three-branch parallel structure, and different branches are designed with different receptive fields to enhance the sensitivity of the model to the changing trends of various water quality characteristics over time; The input is the collected time-series water quality data, which captures short-term and medium- to long-term trends and generates dynamic coding features of water quality ; The input of the control parameter encoding channel is the collected control timing data. A gating fusion mechanism is introduced to perform weighted integration on multiple control dimensions, dynamically adjust the representation contributions of different control variables in specific states, and output gating fusion features ; Subsequently, the control fusion feature is fed into the first enhanced GRU cell for preliminary temporal modeling; then, the output of the first GRU cell and the gated fusion feature are concatenated and fed into the second enhanced GRU cell to characterize the deep interaction relationship of the control variables under the condition of the state combination; The output result of the second GRU unit is processed by layer normalization to obtain the final feature of the first stacked unit ; Based on the above structure, five stacks are performed to extract multi-layer temporal features from local response to global regulation layer by layer. Finally, the characteristic features of the control parameter encoding channel are obtained through the Sigmoid activation function. ; The fusion decoding prediction channel combines the water quality dynamic coding features with the representation features of the control parameter coding channel and splices and fuses them in the channel dimension. Based on the dual Transformer decoding layer-gated fusion structure, it outputs multi-dimensional water quality prediction results .

5. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 4, characterized in that: The specific structure of the water quality dynamic perception channel includes: Short-term perception branch: The water quality time series data first passes through the dilated convolutional block for local fluctuation feature extraction, and then passes through the dilated convolutional block to further expand the short-term context perception range, obtaining short-term dynamic features ; Long-term perception branch: The water quality time series data sequentially passes through dilated convolutional blocks and dilated convolutional blocks, which capture the trend evolution law across time slices with a larger dilation rate, are used to model the cumulative impact of operating condition changes on pollutants, and output the medium-term evolution features ; Original retention branch: Retain the original water quality time series data as the residual connection path, maintain low-level detail information, and form the basic retention features ; After that, the output features of the three branches are concatenated in the channel dimension to form fused features , and are fed into the global average pooling layer to enhance the global representation ability. The fused result is then input into the next round of the three-branch parallel structure for deep feature extraction. The above structure is repeatedly stacked for a total of 5 layers to construct a progressive multi-scale temporal modeling path. After the output of the final layer, a Sigmoid activation function is connected to enhance the non-linear expression ability, and finally, water quality dynamic coding features are generated .

6. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 4, characterized in that: The specific structure of the fusion decoding prediction channel includes: First, the water quality dynamic coding features and the characterization features of the control parameter coding channels are concatenated in the channel dimension; subsequently, the concatenated features are fed into the gating fusion mechanism to obtain the interaction features ; Interaction feature Two consecutive Transformer decoding layers are input in sequence to capture the long-distance dependencies and potential sequence structures between multi-dimensional features; each Transformer decoding layer includes a multi-head attention mechanism, a feed-forward neural network, and a residual connection structure; then, the Transformer decoding result is sent to the gated fusion mechanism again; The double-Transformer decoding layer-gating fusion constitutes a decoding module unit, and this decoding module unit is repeatedly stacked 5 layers to form a multi-layer information abstraction and trend inference path; finally, it is sent to the Dropout layer to enhance the generalization ability, and then processed by the Linear mapping layer and the activation function to output the multi-dimensional water quality prediction result 。 7. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 1, characterized in that: The specific process of S3 includes: Construction of the state vector to be evaluated: For the multi-dimensional water quality prediction results output , perform average aggregation on each control unit in the time dimension to obtain the state vector to be evaluated at the current regulation moment ; Historical state data aggregation analysis: Based on the constructed water quality time series dataset, extract the water quality state data at each single moment in the past , and form a historical water quality state sample set with a scale of ; Perform unsupervised clustering on all historical water quality state vectors to obtain cluster centers , where represents the th cluster center, and ; Next, calculate the current state vector to be evaluated and the Euclidean distances from all cluster centers to obtain a set of distance vectors ; where represents the Euclidean distance between the state vector to be evaluated and the th cluster center; After that, select the nearest clustering center , and extract all historical water quality status data in its corresponding cluster as the candidate sample set; further calculate the cosine similarity between all samples in the cluster and ; finally, select the control time series data corresponding to the first historical sample water quality time series data with the highest similarity , and use this to form a representative control sample set for response analysis ; Response sensitivity assessment and key control parameter screening: for control sample sets For each set of control parameter data in the , a disturbance experiment is performed to calculate each control parameter The corresponding sensitivity coefficient ;in, and Respectively represent control parameters and their sensitivity coefficients; Finally, the parameters that satisfy the sensitivity coefficient greater than the preset threshold are used as the key control parameter set at the current regulation moment .

8. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system according to claim 7, characterized in that: Based on the set of key control parameters, infer its safe and adjustable dynamic response interval, and the specific process is: For each selected key parameter , extract the maximum and minimum values from its historical values in the control sample set to form the preliminary response range of the key parameter : , where respectively represent the minimum value and the maximum value of the parameter in the control sample set; After that, combined with the system's preset operating safety constraints perform double-boundary clipping to obtain the dynamically adjustable response interval of this parameter: ; wherein, are respectively the minimum and maximum values preset by the system for the parameter , and are the operations of taking the maximum and minimum values; is the key parameter corresponding to the dynamically adjustable response interval, and constitutes a structured control parameter-response interval pair .

9. The intelligent regulation method of a three - stage constructed wetland circulating water aquaculture system as described in claim 1, characterized in that: The reinforcement learning control policy network constructed in S4 adopts an Actor-Critic dual-head architecture; First, the current water quality time series data , the control time series data and the prediction state are concatenated into a unified state vector and sent to the input layer of the network for standardization and feature embedding to extract basic information; The feature extraction layer consists of three fully connected networks and ReLU activation units, performs high-dimensional feature transformation on the state vector, and extracts deep control signal features; Policy head pairs with actions For each key control parameter in, a Gaussian distribution value output is constructed respectively, and finally the expected value and corresponding variance of the control action are output, and the corresponding value range is obtained; then, through the dynamic response range The action constraints of each key variable are defined, and among all the value ranges of the key parameters Sub-sampling is performed to obtain a set of potential actions ; The value head calculates the potential actions separately The long-term cumulative expected reward that can be obtained , so as to guide network optimization; the value head consists of two fully connected networks, and as the Critic part in the Actor-Critic framework, it provides a basis for gradient estimation for policy optimization and guides the policy network to converge to a better solution.

10. The intelligent regulation method for a three-stage constructed wetland recirculating aquaculture system according to claim 9, characterized in that: Among them, The environmental state is jointly composed of the water quality time series data, control time series data, and multi-dimensional water quality prediction results of the control system at the current moment; it includes the current water quality data , which is the water quality time series data of all control units at time , and the real-time control data , which is the executed control time series data corresponding to all control units at time ; the prediction state , which is the multi-dimensional water quality prediction results at all control units for the next steps Among them, the action is the adjustment amount of the key control parameter for each control unit . The value range of the action is restricted by the dynamic response interval . Among them, the reward function is a weighted multi-objective comprehensive reward mechanism that rewards efficient pollutant removal, stable compliance with indicators, and minimization of energy consumption.

Citation Information

Patent Citations

  • A framework for automatic calibration of crop varietal parameters in crop growth period model under uncertain conditions

    CN109472320A

  • Water quality early warning method based on improved meta learning

    CN116628444A

  • Method for realizing high-efficiency low-consumption micro-aerobic hydrolytic acidification of petrochemical wastewater by regulating and controlling aeration rate in combination with space-time diagram neural network and reinforced learning

    CN119430460A

  • Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

    CN119494521A

  • Full-period traffic flow prediction method based on spatial-temporal feature deep fusion

    CN119541195A

Cited By

  • DCS automatic control system and method for odor treatment of sewage plant

    CN120762386A

  • Self-adaptive water level adjusting method, system and equipment

    CN120909352A

  • Intelligent fish tank control system, control method and intelligent fish tank

    CN120993788A

  • Sewage treatment strategy adaptive optimization method and system based on reinforcement learning

    CN121020687A

  • A sewage treatment strategy self-adaptive optimization method and system based on reinforcement learning

    CN121020687B