Fuzzy self-adaptive Q learning control method and system for sewage treatment

By integrating fuzzy control and online Q-learning methods, and combining particle swarm optimization to optimize network weights, the problem of insufficient accuracy and adaptability in dissolved oxygen concentration tracking and control in wastewater treatment systems was solved. This resulted in high-precision and stable dissolved oxygen concentration control, improving wastewater treatment efficiency and energy management.

CN121800319APending Publication Date: 2026-04-07BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing wastewater treatment systems suffer from low accuracy in dissolved oxygen concentration tracking and control, insufficient self-adaptability, low weight update efficiency, and difficulty in adapting to complex dynamic characteristics.

Method used

An online Q-learning method integrating fuzzy control and particle swarm optimization is proposed. By adjusting the network weights through fuzzy adaptive PID parameters and using an online Q-learning framework, and updating the weights using a particle swarm algorithm, high-precision tracking control of dissolved oxygen concentration is achieved.

Benefits of technology

It improves the accuracy of dissolved oxygen concentration tracking and control, enhances system operational stability, adapts to complex dynamic characteristics, reduces energy consumption, and improves wastewater treatment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121800319A_ABST
    Figure CN121800319A_ABST
Patent Text Reader

Abstract

The invention provides a fuzzy self-adaptive Q learning control method and system for sewage treatment, and the method comprises the steps: obtaining the dissolved oxygen concentration as a system state, and constructing a nonlinear optimization problem; performing dynamic adjustment on proportion, integral and differential coefficients by adopting a Mamdani type fuzzy inference rule in combination with fuzzy logic, performing defuzzification through a centroid method to obtain adaptive PID parameters, and forming a fuzzy adaptive control strategy; an online Q learning framework is further constructed, a Q function is approximated by using the evaluation network, a network generation strategy adjustment amount is executed, and network weight is optimized based on a Bellman equation and a particle swarm algorithm; finally, fuzzy control and Q learning output are fused, control input is generated through a coupling coefficient to adjust an oxygen transfer coefficient, and tracking control over the dissolved oxygen concentration is achieved. According to the invention, the tracking control precision of the dissolved oxygen concentration and the system operation stability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wastewater treatment technology, and in particular to a fuzzy adaptive Q-learning control method and system for wastewater treatment, used to achieve precise tracking and control of dissolved oxygen concentration during wastewater treatment. Background Technology

[0002] With the accelerating pace of urbanization and the continuous expansion of the population, urban sewage discharge has experienced explosive growth, posing a severe threat to water quality and ecological security. As a core component in maintaining water resource circulation and improving the ecological environment, sewage treatment efficiency and operational stability are directly related to the achievement of "dual carbon" targets and the construction of an ecological civilization, and therefore have received widespread attention.

[0003] As a typical industrial control system, the control methods for wastewater treatment processes have continuously evolved with technological advancements. In wastewater treatment, dissolved oxygen concentration is a key parameter affecting biochemical reaction efficiency, effluent quality compliance, and operational energy consumption. Its precise control is crucial for ensuring the efficient, stable, and low-energy operation of the wastewater treatment system. However, wastewater treatment systems possess complex characteristics such as strong nonlinearity, large time delays, time-varying parameters, and frequent disturbances (e.g., fluctuations in influent flow rate, water quality, and water temperature). This places extremely high demands on the precise control of dissolved oxygen concentration.

[0004] Within the framework of classical control theory, methods such as feedforward control, PID control, and fuzzy control have been effectively applied in wastewater treatment plants, demonstrating satisfactory control performance. However, with the increasing complexity of industrial systems and the ever-increasing demands on control performance, traditional methods, while simple to implement, rely on precise system models and struggle to adapt to the complex dynamic characteristics of wastewater treatment systems. Especially when facing complex processes with strong nonlinearity and time-varying characteristics, the adaptive capabilities of classical control strategies are insufficient, prompting researchers to explore more intelligent control methods.

[0005] To overcome these limitations, data-driven intelligent control methods have attracted much attention in recent years. Innovative methods based on neural networks, such as fuzzy neural networks, deep neural networks, and model predictive control, have shown significant advantages in various industrial control systems. Among them, adaptive evaluative control algorithms, with their advantages of not requiring an accurate system model and possessing online learning and optimization capabilities, have shown great potential in the control of unknown nonlinear systems.

[0006] Existing ensemble methods, in the weight update process within an adaptive dynamic programming framework, still largely employ traditional iterative approaches such as gradient descent. These methods are prone to getting trapped in local optima, resulting in slow weight convergence and limited adaptability to dynamic disturbances under complex operating conditions. This makes it difficult to further improve the control accuracy of dissolved oxygen concentration and the stability of system operation, thus limiting their application in high-precision, high-requirement wastewater treatment scenarios. How to retain the advantages of fuzzy logic and online Q-learning ensemble methods while optimizing the weight update mechanism of the adaptive dynamic programming framework using intelligent algorithms such as particle swarm optimization (PSO) to enhance its global optimization capability and convergence performance, thereby better addressing the complex dynamic characteristics of wastewater treatment systems and achieving high-precision, high-stability tracking control of dissolved oxygen concentration, has become a critical problem urgently needing to be solved in the field of intelligent wastewater treatment control. Summary of the Invention

[0007] The purpose of this application is to provide a fuzzy adaptive Q-learning control method and system for wastewater treatment, aiming to solve the problems of low accuracy, insufficient adaptive capability, and low weight update efficiency in existing wastewater treatment systems. This application improves the accuracy of dissolved oxygen concentration tracking and control and the stability of system operation by integrating fuzzy control and particle swarm optimization into an online Q-learning method.

[0008] To achieve the above objectives, this application provides the following technical solution: This application provides a fuzzy adaptive Q-learning control method for wastewater treatment, including: Obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and establish a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective; Based on the system state and the preset dissolved oxygen concentration setting, the tracking error and error increment are calculated, the tracking error and error increment are normalized, and the fuzzy set is divided by the fuzzy logic system to obtain the normalized error input. Based on the normalized error input, the proportional, integral and derivative coefficients are dynamically adjusted using Mamdani-type fuzzy inference rules. Adaptive PID parameters are obtained by defuzzification using the centroid method, and a fuzzy adaptive control strategy is generated. Based on the aforementioned nonlinear system optimization problem, an online Q-learning framework is constructed, comprising a judgment network and an execution network. The judgment network approximates the Q-function based on the input layer to hidden layer weight matrix and the adjustable hidden layer to output layer weight vector, while the execution network generates control policy adjustment quantities through the input layer weight matrix and the adjustable output layer weight matrix. Based on the online Q-learning framework, the Bellman equation is derived using the time difference analysis method. The evaluation network prediction error and the execution network prediction error are calculated based on the Bellman equation, and their respective optimization objective functions are constructed. Based on the aforementioned optimization objective function, the particle swarm optimization algorithm is applied to optimize and update the hidden layer to output layer weight vector of the evaluation network and the output layer weight matrix of the execution network, respectively, so as to minimize their respective optimization objective functions and obtain the optimized network weight parameters. The control strategy adjustment amount generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters is used to generate the final control input through a preset coupling coefficient, thereby adjusting the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration.

[0009] Preferably, the step of obtaining the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state and establishing a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective includes: Obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment, represent the wastewater treatment system as a nonlinear system, where the system state is the dissolved oxygen concentration, the control input is the oxygen transfer coefficient, the nonlinear system function is an unknown function, and establish a nonlinear system model. Based on the aforementioned nonlinear system model, the dynamic trajectory of dissolved oxygen concentration is set according to engineering experience, and the concentration values ​​are set to 1 mg / L, 2.2 mg / L and 1.8 mg / L at different time periods to generate the dissolved oxygen concentration set value. Based on the nonlinear system model and the dissolved oxygen concentration setpoint, the tracking error is defined as the difference between the system state and the setpoint, thus obtaining the nonlinear system optimization problem.

[0010] Preferably, the step of calculating the tracking error and error increment based on the system state and a preset dissolved oxygen concentration setpoint, normalizing the tracking error and error increment, and dividing the fuzzy set using a fuzzy logic system to obtain the normalized error input includes: Based on the system state and the preset dissolved oxygen concentration setting, the tracking error at the current moment is calculated, and the error increment is calculated based on the tracking error at the current moment and the previous moment, thus obtaining the tracking error and the error increment. The tracking error and the error increment are normalized by using the hyperbolic tangent function, which compresses the input value to the range of negative 1 to positive 1, reduces the numerical instability caused by large error fluctuations, and obtains the normalized tracking error and error increment. Based on the normalized tracking error and error increment, each normalized input variable is divided into five fuzzy sets with Gaussian membership functions, namely negative large, negative small, zero, positive small, and positive large. The membership function of each fuzzy set adopts the form of a Gaussian function to obtain the normalized error input.

[0011] Preferably, the step of dynamically adjusting the proportional, integral, and derivative coefficients using Mamdani-type fuzzy inference rules based on the normalized error input, obtaining adaptive PID parameters through centroid defuzzification, and generating a fuzzy adaptive control strategy includes: Based on the normalized error input, 25 Mamdani-type fuzzy inference rules are used to adjust the proportional coefficient, integral coefficient, and differential coefficient respectively. The proportional coefficient adjustment rules, integral coefficient adjustment rules, and differential coefficient adjustment rules adopt different rule mapping relationships to obtain the fuzzy inference results. Based on the fuzzy inference results, the centroid method is used to defuzzify the proportional coefficient, integral coefficient, and differential coefficient, and the weighted average value of each coefficient is calculated as the final parameter adjustment value to obtain the adaptive PID parameters. Based on the adaptive PID parameters, the defuzzified parameters are integrated through an incremental PID controller. The control law increment is calculated based on the current error, the error of the previous time step, and the errors of the previous two time steps to generate the fuzzy adaptive control strategy.

[0012] Preferably, an online Q-learning framework comprising a judge network and an execution network is constructed, wherein the judge network approximates the Q-function based on the input-to-hidden-layer weight matrix and the adjustable hidden-to-output-layer weight vector, including: Based on the aforementioned nonlinear system optimization problem, the Q function is defined as the cumulative sum of future utility functions including a discount factor, wherein the utility function is constructed based on the sum of squares of the tracking error and the control policy adjustment, thus obtaining the definition of the Q function; Based on the definition of the Q function, the Bellman equation is derived through time difference analysis, and a recursive relationship is established between the Q function at the current time, the Q function at the next time, and the instantaneous utility function, thus obtaining the Bellman equation. Based on the Bellman equation, a judgment network is constructed to approximate the Q function. The judgment network contains a fixed input layer to hidden layer weight matrix and an adjustable hidden layer to output layer weight vector. It adopts a hyperbolic tangent activation function, takes the tracking error and control policy adjustment as input, and outputs the approximate value of the Q function to obtain the online Q learning framework.

[0013] Preferably, the execution network generates control policy adjustment amounts through the input layer weight matrix and the adjustable output layer weight matrix, including: Based on the online Q-learning framework, the execution network is constructed to approximate and track the control input. The execution network includes a fixed input layer weight matrix and an adjustable output layer weight matrix, resulting in the execution network structure. According to the execution network structure, the tracking error is used as the input of the execution network, and a linear transformation is performed through the input layer weight matrix to obtain the transformed signal; Based on the transformed signal, a hyperbolic tangent activation function is used to perform nonlinear processing on the transformed signal, and the control strategy adjustment amount is generated through the adjustable output layer weight matrix.

[0014] Preferably, the step of calculating the evaluation network prediction error and the execution network prediction error based on the Bellman equation, and constructing their respective optimization objective functions, includes: Based on the Bellman equation, the prediction error of the evaluation network is defined as the difference between the current Q-function approximation value and the sum of the instantaneous utility function and the next Q-function approximation value, thus obtaining the prediction error of the evaluation network. Based on the Bellman equation, the prediction error of the execution network is defined as the gradient of the Q function with respect to the adjustment of the control policy, where the utility function value corresponding to the final optimization objective is zero, thus obtaining the prediction error of the execution network. Each of the optimization objective functions is constructed based on the squares of the prediction error of the evaluation network and the prediction error of the execution network, respectively.

[0015] Preferably, the particle swarm optimization algorithm is used to optimize and update the hidden-to-output layer weight vectors of the evaluation network and the output layer weight matrix of the execution network, respectively, minimizing their respective optimization objective functions to obtain optimized network weight parameters, including: Based on the optimization objective function, the particle swarm optimization (PSO) algorithm parameters are initialized, including the number of particles, the maximum number of iterations, the inertia weight, and the acceleration coefficient, to obtain the PSO algorithm parameters. Based on the particle swarm optimization (PSO) parameters, for the evaluation network, the weight vector from the hidden layer to the output layer is used as the particle position, and the optimization objective function of the evaluation network is used as the fitness function. The optimal weight vector is found iteratively through the PSO algorithm to obtain the optimal weight vector of the evaluation network. Based on the particle swarm optimization (PSO) parameters and the optimal weight vector of the evaluation network, for the execution network, the output layer weight matrix is ​​used as the particle position, and the optimization objective function of the execution network is used as the fitness function. The optimal weight matrix is ​​found iteratively through the PSO algorithm to obtain the optimized network weight parameters.

[0016] Preferably, the control strategy adjustment amount generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters, and generating the final control input through a preset coupling coefficient, adjusts the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration, including: Based on the fuzzy adaptive control strategy and the optimized network weight parameters, the control strategy adjustment amount is generated through the optimized network weight parameters to obtain the control strategy adjustment amount; The fuzzy adaptive control strategy and the control strategy adjustment amount are weighted and combined using a preset coupling coefficient to obtain the final control input; Based on the final control input, the oxygen transfer coefficient in the wastewater treatment system is adjusted to change the dissolved oxygen concentration in the wastewater treatment system, thereby achieving tracking control of the dissolved oxygen concentration.

[0017] This application also provides a fuzzy adaptive Q-learning control system for wastewater treatment, including: The data acquisition module is used to obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and to establish a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective. The fuzzy processing module is used to calculate the tracking error and error increment based on the system state and the preset dissolved oxygen concentration setting, normalize the tracking error and the error increment, and divide the fuzzy set through the fuzzy logic system to obtain the normalized error input. The fuzzy control module is used to dynamically adjust the proportional, integral, and derivative coefficients according to the normalized error input using Mamdani-type fuzzy inference rules, obtain adaptive PID parameters by defuzzification using the centroid method, and generate a fuzzy adaptive control strategy. An online learning module is used to construct an online Q-learning framework containing an evaluation network and an execution network based on the nonlinear system optimization problem. The evaluation network approximates the Q-function based on the input layer to hidden layer weight matrix and the adjustable hidden layer to output layer weight vector, and the execution network generates control policy adjustment quantities through the input layer weight matrix and the adjustable output layer weight matrix. The error calculation module is used to derive the Bellman equation using the time difference analysis method based on the online Q-learning framework, calculate the evaluation network prediction error and the execution network prediction error based on the Bellman equation, and construct their respective optimization objective functions. The intelligent optimization module is used to optimize and update the hidden layer to output layer weight vector of the evaluation network and the output layer weight matrix of the execution network based on the optimization objective function, thereby minimizing their respective optimization objective functions and obtaining the optimized network weight parameters. The control execution module is used to combine the fuzzy adaptive control strategy and the control strategy adjustment amount generated by the optimized network weight parameters, generate the final control input through a preset coupling coefficient, and adjust the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration.

[0018] Compared with the prior art, the embodiments of this application have at least the following beneficial effects: 1. This application's embodiments, by integrating fuzzy control and online Q-learning, combined with swarm intelligence optimization, construct an adaptive control method that does not require a precise system model, suitable for highly nonlinear and time-varying wastewater treatment systems. This application uses particle swarm optimization instead of the traditional gradient descent method for network weight updates, improving global optimization capabilities, accelerating convergence speed, and avoiding the problem of getting trapped in local optima.

[0019] 2. Based on the experimental results, the embodiments of this application exhibit the best tracking control performance under three typical operating conditions: sunny days, rainy days, and heavy rain, which is significantly better than traditional PID control and conventional fuzzy online Q-learning control methods. Therefore, this application solves the problems of insufficient accuracy and poor robustness in the existing dissolved oxygen concentration control of sewage treatment, which is conducive to improving sewage treatment efficiency, reducing energy consumption, and has significant economic and environmental benefits. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a structural diagram of the wastewater treatment system based on the swarm intelligence-guided fuzzy online Q-learning (PSO-FOQL) algorithm of this application; Figure 2 This is a structural framework diagram of the PSO-FOQL algorithm in this application; Figure 3 This is a flowchart of the fuzzy adaptive Q-learning control method for wastewater treatment in this application; Figure 4 This is a diagram illustrating the effect of dissolved oxygen concentration tracking under sunny weather conditions in this application. Figure 5 This is a diagram illustrating the effect of tracking dissolved oxygen concentration under rainy weather conditions in this application. Figure 6 This is a graph illustrating the effect of tracking dissolved oxygen concentration during heavy rain in this application. Figure 7 The oxygen transfer coefficient (K) under three weather conditions in this application La,5 The effect diagram of the control. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0023] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S11, S12, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0024] It will be understood by those skilled in the art that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application’s specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein may include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0025] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Throughout the description, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0027] Figure 1The structure of a wastewater treatment system based on the PSO-FOQL algorithm is shown. The PSO-FOQL algorithm is a hybrid intelligent algorithm integrating particle swarm optimization and fuzzy Q-learning. This wastewater treatment system structure includes a biological reactor, where urban wastewater flows in from the left side and sequentially passes through an aerobic zone 1.2 (comprising units 1.4 and 1.5), an anaerobic zone 1.3 (comprising units 1.6, 1.7, and 1.8), and then into a secondary sedimentation tank 1.9 for solid-liquid separation, further separating the supernatant from the sludge. The supernatant is discharged from the top of the system, while the sludge is divided into two parts: one part is transported to the front of the anaerobic zone via a return pipe to participate in the reaction again, and the other part is discharged from the bottom of the system as residual sludge. The controller 2.1 acquires the dissolved oxygen concentration S at the end of the aerobic zone in real time. 0,5 After processing by the PSO-FOQL algorithm, a control signal is output, which ultimately changes the oxygen transfer coefficient K. La,5 This enables closed-loop tracking control of dissolved oxygen concentration against the set value.

[0028] like Figure 2 As shown, Figure 2 The structural framework diagram of the PSO-FOQL algorithm is shown. The structural framework of the PSO-FOQL algorithm includes the following steps: Step 1: Establish the dissolved oxygen concentration tracking control optimization problem; Step 2: Generate the initial control strategy using a fuzzy adaptive controller; Step 3: Optimize the dissolved oxygen concentration control strategy online through the evaluation network and execution network under the online Q-learning framework; Step 4: Train the evaluation network and execution network using a swarm intelligence algorithm.

[0029] like Figure 3 As shown, this application provides a fuzzy adaptive Q-learning control method for wastewater treatment, including: Step S101: Obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and establish a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective; Step S102: Based on the system state and the preset dissolved oxygen concentration setting, calculate the tracking error and the error increment, normalize the tracking error and the error increment, and divide the fuzzy set through the fuzzy logic system to obtain the normalized error input; Step S103: Based on the normalized error input, the proportional, integral and derivative coefficients are dynamically adjusted using Mamdani-type fuzzy inference rules. Adaptive PID parameters are obtained by defuzzification using the centroid method, and a fuzzy adaptive control strategy is generated. Step S104: Based on the nonlinear system optimization problem, construct an online Q-learning framework including an evaluation network and an execution network. The evaluation network approximates the Q-function based on the input layer to hidden layer weight matrix and the adjustable hidden layer to output layer weight vector. The execution network generates the control policy adjustment amount through the input layer weight matrix and the adjustable output layer weight matrix. Step S105: Based on the online Q-learning framework, derive the Bellman equation using the time difference analysis method, calculate the evaluation network prediction error and the execution network prediction error based on the Bellman equation, and construct their respective optimization objective functions; Step S106: Based on the optimization objective function, the particle swarm optimization algorithm is applied to optimize and update the hidden layer to output layer weight vector of the evaluation network and the output layer weight matrix of the execution network respectively, so as to minimize their respective optimization objective functions and obtain the optimized network weight parameters. Step S107: Combine the control strategy adjustment amount generated by the fuzzy adaptive control strategy and the optimized network weight parameters, generate the final control input through the preset coupling coefficient, and adjust the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration.

[0030] In this embodiment, it is first necessary to obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and then establish a nonlinear system optimization problem with the goal of tracking and controlling the dissolved oxygen concentration.

[0031] Based on the acquired system state and the preset dissolved oxygen concentration setpoint, the tracking error and error increment are calculated. The tracking error and error increment are then normalized and partitioned into fuzzy sets using a fuzzy logic system to obtain the normalized error input. The tracking error refers to the difference between the actual dissolved oxygen concentration and the setpoint, while the error increment represents the change in error between the current and previous moments. These two parameters together reflect the dynamic response characteristics of the system. A fuzzy logic system is a control method based on fuzzy set theory and fuzzy inference. It simulates the decision-making process of human experts and can handle uncertainties and nonlinearities in the system. Normalization is the process of mapping error signals of different dimensions and ranges to a unified interval (e.g., [-1, 1]), which helps improve the numerical stability of the system. Fuzzy set partitioning divides the normalized error signal into multiple linguistic variables (e.g., "negative large", "negative small", "zero", "positive small", "positive large"), and the membership degree of each value to each linguistic variable is determined by a membership function.

[0032] Based on the normalized error input, Mamdani-type fuzzy inference rules are used to dynamically adjust the proportional, integral, and derivative coefficients. Adaptive PID parameters are obtained through centroid defuzzification, generating a fuzzy adaptive control strategy. Mamdani-type fuzzy inference is a commonly used fuzzy inference method that maps the input fuzzy set to the output fuzzy set using an "if-then" rule. In this application, the fuzzy inference rules are designed based on expert experience and control theory, and are used to dynamically adjust the three parameters of the PID controller according to the current error state: the proportional coefficient (control response speed), the integral coefficient (to eliminate steady-state error), and the derivative coefficient (to improve system stability). The centroid method is a commonly used defuzzification method that converts the fuzzy output into explicit control parameter values ​​by calculating the "centroid" position of the fuzzy inference result. The adaptive PID control strategy combines the simplicity of traditional PID control with the adaptability of fuzzy control, enabling dynamic adjustment of control parameters according to the system state to adapt to control requirements under different operating conditions.

[0033] Based on nonlinear system optimization problems, an online Q-learning framework comprising a judge network and an execution network is constructed. Online Q-learning is a type of reinforcement learning that optimizes long-term cumulative rewards through interactive learning without requiring an exact system model. The Q-function represents the expected value of the long-term cumulative reward obtained by taking a certain action in a specific state and is the core of the reinforcement learning algorithm. The judge network is a neural network structure used to approximate the Q-function, containing a fixed input-to-hidden-layer weight matrix and an adjustable hidden-to-output-layer weight vector. The execution network is another neural network structure used to generate the optimal control policy (i.e., the control policy adjustment), mapping the state to control actions through the input-layer weight matrix and the adjustable output-layer weight matrix. The two networks work together to form a complete online learning control framework.

[0034] Based on the online Q-learning framework, the Bellman equation is derived using temporal difference analysis. The evaluation network prediction error and the execution network prediction error are calculated based on the Bellman equation, and their respective optimization objective functions are constructed. Temporal difference analysis is a reinforcement learning technique combining dynamic programming and Monte Carlo methods. It does not require a complete environment model and can learn through sampling. The Bellman equation is a fundamental equation in reinforcement learning, describing the recursive relationship between the current state value and the next state value. The evaluation network prediction error refers to the difference between the network's output Q-value and the target Q-value calculated according to the Bellman equation, while the execution network prediction error measures the deviation between the current control policy and the optimal control policy. The optimization objective function typically uses the sum of squares of these errors, aiming to minimize the errors by adjusting the network parameters.

[0035] Based on the objective function, the particle swarm optimization (PSO) algorithm is applied to optimize and update the hidden-to-output layer weight vectors of the evaluation network and the output layer weight matrix of the execution network, minimizing their respective objective functions to obtain the optimized network weight parameters. PSO is a swarm intelligence optimization algorithm that simulates the foraging behavior of birds, finding the global optimum by having multiple "particles" search the solution space. Compared to traditional gradient descent, PSO has advantages such as strong global search capability, ease of implementation, and fast convergence speed, making it particularly suitable for handling nonlinear and nonconvex optimization problems. In this application, each particle represents a set of possible network weight parameters, and the fitness of the particle is determined by the objective function. By iteratively updating the position and velocity of the particles, the algorithm gradually approaches the optimal solution, thereby obtaining the optimal network weight parameters.

[0036] The control strategy adjustment amount, generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters, is used to generate the final control input through a preset coupling coefficient. This input adjusts the oxygen transfer coefficient to achieve precise control of dissolved oxygen concentration. The coupling coefficient is a parameter between 0 and 1, used to balance the contribution ratios of fuzzy control and reinforcement learning control. A larger coupling coefficient means greater reliance on the control strategy generated by reinforcement learning, while a smaller value indicates greater reliance on fuzzy control. The final control input directly acts on the aeration equipment (such as blowers) of the wastewater treatment system, adjusting the oxygen transfer coefficient to change the amount of oxygen entering the water body, thereby achieving precise control of dissolved oxygen concentration.

[0037] This application's control method integrates expert knowledge from fuzzy logic control and the adaptive capabilities of reinforcement learning, combined with the global search advantages of swarm intelligence optimization, to construct a highly efficient and robust wastewater treatment control system. This method does not rely on precise mathematical models, can adapt to the nonlinear characteristics of the system and environmental disturbances, and achieves high-precision tracking control of dissolved oxygen concentration. It significantly improves wastewater treatment efficiency and reduces energy consumption, providing an innovative solution for intelligent and green wastewater treatment technology.

[0038] The process of obtaining the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state and establishing a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective includes: obtaining the dissolved oxygen concentration in the wastewater treatment system at the current moment; representing the wastewater treatment system as a nonlinear system, where the system state is the dissolved oxygen concentration, the control input is the oxygen transfer coefficient, and the nonlinear system function is an unknown function; establishing a nonlinear system model; based on the nonlinear system model, setting the dynamic trajectory of the dissolved oxygen concentration according to engineering experience, setting the concentration values ​​to 1 mg / L, 2.2 mg / L, and 1.8 mg / L at different time periods, respectively, and generating a dissolved oxygen concentration setpoint; based on the nonlinear system model and the dissolved oxygen concentration setpoint, defining the tracking error as the difference between the system state and the setpoint, thus obtaining the nonlinear system optimization problem.

[0039] Wastewater treatment systems exhibit typical nonlinear and time-varying characteristics due to their complex biochemical reaction processes, variable influent water quality and load, temperature fluctuations, and dynamic changes in microbial communities. In this application, the system state is defined as dissolved oxygen concentration, which directly reflects the internal state of the system; the control input is defined as the oxygen transfer coefficient, which characterizes the efficiency of oxygen transfer from the gas phase to the liquid phase, measured in units of one hour, and is a variable that the control system can directly operate on, achieved by adjusting the blower speed or aeration rate.

[0040] Setting a dynamic trajectory for dissolved oxygen concentration based on engineering experience is key to optimized control. Different dissolved oxygen concentration values ​​are set for different time periods based on the operating patterns and treatment requirements of the wastewater treatment plant. Specifically, a low concentration setting of 1 mg / L is typically used during low-load periods at night (e.g., 0:00-6:00 AM), when influent flow is low, pollutant concentration is low, and microbial activity is reduced; a low dissolved oxygen setting can significantly reduce energy consumption. A high concentration setting of 2.2 mg / L is used during high-load periods in the morning and evening (e.g., 7:00-9:00 AM and 5:00-7:00 PM), when influent flow is high and pollutant concentration is high, requiring sufficient oxygen to maintain efficient biodegradation. A medium concentration setting of 1.8 mg / L is used during other normal load periods. This dynamic setting not only considers treatment efficiency but also fully considers energy consumption optimization.

[0041] Based on a nonlinear system model and a dissolved oxygen concentration setpoint, the tracking error is defined as the difference between the system state and the setpoint, thus deriving a nonlinear system optimization problem. Tracking error is an indicator in a control system that measures the deviation between the actual output and the desired output; in this system, it is represented by the difference between the actual dissolved oxygen concentration and the setpoint. This error signal is the primary basis for control system decisions and an important indicator for evaluating control performance. In wastewater treatment dissolved oxygen control, this optimization problem can be formulated as: finding the optimal oxygen transfer coefficient adjustment strategy to minimize the deviation between the actual dissolved oxygen concentration and the dynamic setpoint, while considering the constraint of minimizing energy consumption. This type of optimization problem typically involves a multi-objective balance, such as the trade-off between tracking accuracy and control stability, and between treatment efficiency and energy consumption.

[0042] In practical applications, solving optimization problems for nonlinear systems faces numerous challenges. Due to the uncertainty of the system model, the randomness of environmental disturbances, and the dynamic changes in the control objective, traditional model-based optimization methods often fail to achieve satisfactory results. Therefore, this application employs a method that integrates fuzzy logic and reinforcement learning, achieving effective control of nonlinear systems through online learning and adaptive adjustment. This data-driven control method does not rely on an accurate system model but learns the optimal control strategy from historical data and real-time feedback. It can effectively cope with the nonlinearity, time-varying characteristics, and external disturbances of the system, providing an intelligent solution for wastewater treatment process control.

[0043] Specifically, wastewater treatment systems can be represented as the following type of nonlinear system:

[0044] in, This represents the system state at time k, indicating the dissolved oxygen concentration in unit 5.1.8. , The control input at time k represents the oxygen transfer coefficient. ; This represents an unknown nonlinear system function.

[0045] Dissolved oxygen concentration set value Represented as:

[0046] in: This refers to setting the trajectory function. Based on engineering experience, the dynamic trajectory setting for dissolved oxygen concentration typically takes the following function form:

[0047] Tracking error between dissolved oxygen concentration and its set value Defined as:

[0048] In conventional adaptive evaluation algorithms, the calculated control policy is a direct control policy. This makes the control input To enhance the algorithm's anti-interference capability, this application uses an adaptive fuzzy controller to design a supplementary control strategy, such that the control input satisfies the following equation:

[0049] in, Indicated by the adaptive fuzzy controller Timing control strategies; This represents the control input generated iteratively by the online Q-learning algorithm; It is the coupling coefficient for adjusting and optimizing strength.

[0050] The process of calculating the tracking error and error increment based on the system state and the preset dissolved oxygen concentration setting, and normalizing the tracking error and error increment by using a fuzzy logic system and dividing them into fuzzy sets to obtain the normalized error input includes: calculating the tracking error at the current moment based on the system state and the preset dissolved oxygen concentration setting, and calculating the error increment based on the tracking error at the current moment and the previous moment to obtain the tracking error and error increment.

[0051] Based on the system state and the preset dissolved oxygen concentration setpoint, the tracking error at the current moment is calculated, and the error increment is calculated based on the tracking errors at the current moment and the previous moment, thus obtaining the tracking error and the error increment. This step is fundamental to the design of the control system; it compares the actual operating state of the system with the desired state, providing a basis for subsequent control decisions. The tracking error refers to the difference between the currently measured dissolved oxygen concentration and the preset target value, reflecting the degree to which the system deviates from the target.

[0052] In this process, the hyperbolic tangent function is used to normalize the tracking error and error increment, compressing the input values ​​into the range of -1 to +1. This reduces numerical instability caused by large error fluctuations, resulting in normalized tracking error and error increment. Normalization is the process of mapping data with different dimensions and ranges to a unified interval, and it is often used in control systems to improve the numerical stability and versatility of algorithms.

[0053] Based on the normalized tracking error and error increment, each normalized input variable is divided into five fuzzy sets with Gaussian membership functions: negative large, negative small, zero, positive small, and positive large. The membership function of each fuzzy set adopts the form of a Gaussian function, resulting in the normalized error input. A fuzzy set is a fundamental concept in fuzzy logic. Unlike traditional set theory, a fuzzy set allows elements to belong to a set to a certain degree (membership). The five fuzzy sets represent different linguistic descriptions of the error and error increment; for example, "negative large" indicates that the actual value is much lower than the set value, and "zero" indicates that the actual value is close to the set value. The membership function defines the degree of membership of an element to the fuzzy set, with values ​​ranging from [0,1], where 0 indicates that the element does not belong to the set at all, and 1 indicates that the element belongs to the set completely.

[0054] Normalized error input forms the basis for subsequent fuzzy inference. It converts continuous error signals into fuzzy quantities with linguistic descriptive properties, enabling the control system to simulate the decision-making process of human experts. For example, when the normalized tracking error is -0.7 and the error increment is 0.2, membership calculations show that the tracking error mainly belongs to the "negatively large" fuzzy set (membership approximately 0.8) and the "negatively small" fuzzy set (membership approximately 0.2), while the error increment mainly belongs to the "zero" fuzzy set (membership approximately 0.7) and the "positively small" fuzzy set (membership approximately 0.3). This fuzzy representation is closer to human thinking than simple numerical representation and can more naturally incorporate expert knowledge and experience.

[0055] The fineness of fuzzy set partitioning directly affects control accuracy and computational complexity. Five fuzzy sets are a common choice in practice, achieving a good balance between expressive power and computational efficiency. For more complex systems or applications requiring higher precision, increasing the number of fuzzy sets can be considered, but this will also increase the number and complexity of subsequent fuzzy rules. In systems like wastewater treatment, where control precision requirements are not extremely high but strong robustness is needed, five fuzzy sets can usually meet control requirements while maintaining the system's simplicity and practicality.

[0056] Specifically, based on the principle of fuzzy rule mapping, a two-input, three-output fuzzy logic system is established to track errors. and its increment The input is used to map the increments of the three PID parameters. The hyperbolic tangent function is used to normalize the input variables, compressing the input values ​​to the interval [-1, 1] to mitigate numerical instability caused by large error fluctuations. Each normalized input variable is divided into five fuzzy sets with Gaussian membership functions. : Negative large, negative small, zero, positive small, and positive large. The membership function of each fuzzy set is defined as:

[0057] in, , For tracking error and its increment Together as input, and These represent the center and width of the membership function, respectively.

[0058] In another example, based on the normalized error input, a Mamdani-type fuzzy inference system containing 25 rules is used to dynamically adjust the proportional, integral, and derivative coefficients. The adjustment rules for the three coefficients employ different mapping relationships. After obtaining the fuzzy inference results, the centroid method is used for defuzzification, and the weighted average of each coefficient is calculated as the adaptive PID parameter. Subsequently, the defuzzified parameters are input into an incremental PID controller. Combining the current error, the error from the previous time step, and the errors from the previous two time steps, the control law increment is calculated, thereby generating a fuzzy adaptive control strategy.

[0059] Based on the normalized error input, this system employs 25 Mamdani-type fuzzy inference rules to dynamically adjust the proportional, integral, and derivative coefficients of the PID controller. The input to this fuzzy inference system is the normalized tracking error and its rate of change, while the output is the adjustment amount of the three PID parameters. The 25 rules are derived from a complete combination of five fuzzy sets of the input variables.

[0060] In terms of rule design, the three parameters are adjusted using different mapping relationships. The proportional coefficient rule mainly focuses on the dynamic response of the system. When the error is large, the proportional coefficient is increased to speed up the response, and when the error is close to zero, the proportional coefficient is decreased to prevent overshoot. The integral coefficient rule focuses on eliminating steady-state error. When the error persists, the integral action is enhanced, and when the error changes rapidly, the integral action is weakened to avoid integral saturation. The derivative coefficient rule focuses on system stability. When the error changes drastically, the derivative action is enhanced to provide damping, and when the error changes gradually, the derivative action is weakened to suppress noise interference.

[0061] The fuzzy inference process consists of three key steps: First, the activation strength of each rule is calculated, and the minimum operator is used to process the "AND" logic; then, the corresponding output fuzzy set is truncated according to the activation strength; finally, the output fuzzy sets of all rules are synthesized by the maximum operator to obtain the comprehensive output fuzzy set.

[0062] After obtaining the fuzzy inference results, the system employs the centroid method for defuzzification. This method converts fuzzy quantities into precise parameter adjustment values ​​by calculating the weighted centroid of the output fuzzy set. Specifically, the output universe of discourse is discretized into a finite set of points, and the weighted average of the membership degree and position of each point is calculated. The centroid method can smoothly integrate the influence of all activation rules, producing a continuous and stable output, making it particularly suitable for processes like wastewater treatment that require high control stability.

[0063] This embodiment achieves precise tracking and control of dissolved oxygen concentration by combining fuzzy inference, defuzzification, incremental PID control, and adaptive dynamic programming, while maintaining the stability and robustness of the control process. Fuzzy inference and defuzzification dynamically adjust the core parameters of the PID controller based on error characteristics, incremental PID control ensures smooth output of the control law, and adaptive dynamic programming optimizes the weights of the evaluation and execution networks in the online Q-learning framework using a particle swarm optimization algorithm, continuously iteratively optimizing the control strategy adjustment, further enhancing the system's adaptability to the nonlinear and time-varying characteristics of wastewater treatment.

[0064] Specifically, in order to map the error to appropriate PID parameter adjustments, this embodiment designs a set of fuzzy inference rules based on expert experience and control theory principles. The fuzzy rule structures for adjusting the proportional, integral, and derivative coefficients are shown in Tables 1, 2, and 3.

[0065] Table 1: Rules for Adjusting Proportional Coefficients

[0066] Table 2: Rules for Adjusting Integral Coefficients

[0067] Table 3: Adjustment Rules for Differential Coefficients

[0068] Among them, the centroid method is used for defuzzification to ensure that the parameters are smoothly transformed according to the following formula:

[0069] in, ; These represent the numbers in Tables 1-3 respectively. The weights corresponding to the fuzzy rules. After defuzzification, the adjusted parameters are obtained through an incremental PID controller. , and Integrated into an incremental PID controller:

[0070] in, , and This is the learning rate.

[0071] The control law increment satisfies:

[0072] Therefore, the output of the fuzzy adaptive control input is given by the following equation:

[0073] This fuzzy adaptive incremental PID control law combines dynamic parameter tuning with incremental control logic, enabling it to adapt to system uncertainties and nonlinearities in real time.

[0074] In another example, the construction of an online Q-learning framework comprising a judge network and an execution network, wherein the judge network approximates the Q-function based on the input layer weight matrix and the output layer weight vector, includes: defining the Q-function as the cumulative sum of future utility functions including a discount factor, based on the nonlinear system optimization problem, wherein the utility function is constructed based on the sum of squares of the tracking error and the control policy adjustment, thus obtaining the Q-function definition; deriving the Bellman equation through time difference analysis based on the Q-function definition, establishing a recursive relationship between the Q-function at the current time step, the Q-function at the next time step, and the instantaneous utility function, thus obtaining the Bellman equation; and constructing the judge network to approximate the Q-function based on the Bellman equation, thus obtaining the online Q-learning framework.

[0075] The Q-function is a core concept in reinforcement learning. It represents the expected long-term cumulative reward of taking an action in the current state and then following a specific strategy. In wastewater treatment dissolved oxygen control systems, the Q-function provides a basis for long-term utility assessment of control decisions, enabling the control system to weigh the impact of current control actions on future system performance.

[0076] The utility function is an immediate indicator of system performance, constructed in this system as the sum of the squares of the tracking error (the difference between the dissolved oxygen concentration and the setpoint) and the control adjustment (the amplitude of the control signal change). This construction reflects two fundamental objectives of the control system: first, to minimize the tracking error, ensuring accurate tracking of the dissolved oxygen concentration to the setpoint; and second, to minimize the control adjustment amplitude, reducing excessive controller intervention, lowering energy consumption, and extending equipment life. The square term of the tracking error ensures that the control system pays equal attention to both positive and negative errors, while the square term of the control adjustment suppresses drastic fluctuations in the control signal. The two are typically balanced by weighting coefficients, reflecting the trade-off between control accuracy and control stability.

[0077] Understandably, the discount factor is an important parameter in the Q function. It is a value between 0 and 1, used to adjust the degree of emphasis the current decision places on future returns. A larger discount factor (closer to 1) means that the control system pays more attention to long-term benefits, which is suitable for handling wastewater treatment systems with long time lags; a smaller discount factor (closer to 0) makes the system pay more attention to immediate benefits, which is suitable for scenarios with high requirements for short-term control accuracy.

[0078] Based on the definition of the Q-function, the Bellman equation is derived through time-difference analysis. This establishes a recursive relationship between the Q-function at the current time step, the Q-function at the next time step, and the instantaneous utility function, thus obtaining the Bellman equation. The Bellman equation is the theoretical foundation of dynamic programming and reinforcement learning. It describes the recursive structure of optimization problems, decomposing complex long-term decision problems into a series of simple, single-step decision problems.

[0079] Temporal difference analysis (TDA) is a key method for deriving the Bellman equation, based on the fundamental principle that "the optimal decision at the current moment must consider its impact on future states." By comparing the Q-functions of adjacent moments, TDA reveals the recursive structure of the Q-function over time: the Q-value at the current moment equals the immediate utility gained at the current moment plus the expected value of the optimal Q-value at the next moment after discounting. This recursive relationship is the core of the Q-learning algorithm, enabling the system to progressively improve its estimation of the Q-function by observing state transitions and immediate rewards.

[0080] Based on the Bellman equation, a judge network is constructed to approximate the Q-function. The judge network contains a fixed input-to-hidden-layer weight matrix and a hidden-to-output-layer weight vector. It takes tracking error and control policy adjustments as inputs and outputs an approximate value of the Q-function, thus obtaining an online Q-learning framework. The judge network is a special neural network structure used to approximate the Q-function, and it plays a role in evaluating the value of control actions in reinforcement learning-based control systems.

[0081] Online Q-learning frameworks are real-time adaptive reinforcement learning methods that allow systems to continuously update their Q-function estimates as they interact with their environment, without requiring extensive pre-collection of training data. This framework is particularly well-suited for dynamic environments such as wastewater treatment, where system characteristics may change over time, and pre-trained models may quickly become ineffective. Online learning enables continuous adaptation to these changes, maintaining control performance.

[0082] In wastewater treatment dissolved oxygen control applications, the evaluation network's inputs include normalized dissolved oxygen tracking error and the control strategy adjustment from the previous step. The network's output is an assessment of the long-term value of the current state-action combination. This assessment guides the control system to strike a balance between immediate control accuracy and long-term energy consumption optimization, achieving more intelligent control decisions. Constructing the evaluation network is the first step in the online Q-learning framework, laying the foundation for subsequent execution network design and control strategy learning. This step enables the evaluation of the long-term value of control actions, providing a core component for reinforcement learning-based intelligent control.

[0083] Specifically, a judge network and an execution network are constructed within an online Q-learning framework. The Q-function is defined using the online Q-learning framework as follows:

[0084] in, Discount factor; utility function This application sets constant weights. ; Through time difference analysis, the Bellman equation is derived as follows:

[0085] Construct a judge network to approximate the Q function, and define the output of the judge network as:

[0086] in, The input layer-hidden layer weight matrix is ​​fixed. This represents the number of hidden layer neurons. This represents the hidden-output layer weight vector of an adjustable evaluation network; This represents the hyperbolic tangent activation function.

[0087] To simplify the representation, define The output formula can be simplified to:

[0088] In another example, the execution network generates control policy adjustments using an input layer weight matrix and an adjustable output layer weight matrix. This includes: constructing the execution network based on the online Q-learning framework to approximate the tracking control input; the execution network comprising a fixed input layer weight matrix and an adjustable output layer weight matrix, resulting in an execution network structure; and, based on the execution network structure, using the tracking error as input to the execution network, generating the control policy using the input layer weight matrix and the output layer weight matrix.

[0089] Based on the online Q-learning framework, an execution network is constructed to approximate and track control inputs. The execution network consists of a fixed input layer weight matrix and an adjustable output layer weight matrix, resulting in the execution network structure. The execution network works in conjunction with a judge network to form a complete architecture. In this architecture, the execution network is responsible for generating control actions, while the judge network evaluates the long-term value of these actions.

[0090] Unlike judgment networks, execution networks typically only receive the system state as input and do not require current action information. This is because the purpose of execution networks is to generate the optimal action directly from the state, rather than to evaluate the value of state-action pairs.

[0091] Furthermore, the output of the execution network can be combined with traditional control methods (such as PID control) to form a hybrid control strategy. This combines the advantages of both methods, utilizing the adaptability and ability to handle nonlinear systems of reinforcement learning while retaining the reliability and engineering practicality of traditional control methods.

[0092] Specifically, the execution network approximates and tracks the control input through the following mapping relationship. :

[0093] in, The input layer weight matrix is ​​fixed. This represents the output layer weight matrix of the adjustable execution network; The hyperbolic tangent function is also used; Specify the number of hidden layer neurons.

[0094] definition Then the output formula simplifies to:

[0095] In another example, the evaluation network prediction error and the execution network prediction error are calculated based on the Bellman equation, and their respective optimization objective functions are constructed. The optimization objective functions are constructed based on the squares of the evaluation network prediction error and the execution network prediction error, respectively.

[0096] The evaluation network's prediction error, also known as temporal difference error, is a core concept in reinforcement learning. It measures the difference between the evaluation network's estimate of the Q-function and the target value on the right-hand side of the Bellman equation. The Bellman equation plays a benchmark role in this step, providing the theoretical basis for evaluating the accuracy of the Q-function estimate. According to the Bellman equation, the Q-function value at the current time step should equal the sum of the instantaneous utility function (immediate reward) and the discounted Q-function value at the next time step. In practical computation, we use the output of the evaluation network as an approximation of the Q-function, and the difference between this approximation and the target value calculated based on the Bellman equation is defined as the prediction error.

[0097] The execution network prediction error is a metric measuring the difference between the current control policy and the optimal control policy. It guides the execution network to adjust its parameters to generate better control actions. The gradient of the Q function with respect to the control policy adjustment represents the sensitivity of the Q value to the control action, indicating how to fine-tune the control action to increase the Q value. In practice, this gradient is calculated through backpropagation of the evaluation network: first, the current state and the action output by the execution network are input into the evaluation network, and then the gradient of the output (Q estimate) with respect to the input (control action) is calculated. The direction of this gradient indicates the direction of the action adjustment to increase the Q value, while the magnitude of the gradient reflects the magnitude of the adjustment.

[0098] We construct separate optimization objective functions based on the squares of the evaluation network prediction error and the execution network prediction error, respectively. The optimization objective function is a mathematical expression that guides the network parameter updates, transforming the prediction error into an optimizable scalar objective.

[0099] The objective function for evaluating a network is typically defined as the expected value of the squared prediction error or the squared prediction error in a single step during online learning. Minimizing this objective function means that the Q-function estimate of the network gets closer and closer to the theoretical optimum satisfying the Bellman equation, thus providing a more accurate value assessment.

[0100] The objective function of the execution network is based on the square of the network's prediction error, which physically means maximizing the gradient ascent of the Q-function value. Since the prediction error of the execution network is the gradient of the Q-function with respect to the control action, minimizing this objective function actually involves moving in the control action space along the direction of increasing Q-value, gradually approaching the optimal control strategy.

[0101] In wastewater treatment dissolved oxygen control systems, the design of the optimization objective function requires careful consideration of the trade-off between control accuracy and energy consumption. Overemphasizing control accuracy may lead to over-control and energy waste, while overemphasizing energy conservation may sacrifice treatment efficiency. By appropriately setting the weights of tracking error and control adjustment in the utility function, and constructing the optimization objective function of the evaluation network and execution network based on this, a balance can be achieved, resulting in a control effect that satisfies treatment requirements while conserving energy.

[0102] Specifically, based on the Bellman optimality principle, the evaluation of the network's prediction error is designed as follows:

[0103] in, U k This represents the instantaneous utility function value. The evaluation network's optimization objective is to minimize the squared error function through swarm intelligence algorithm training. .

[0104] The prediction error of the execution network is defined as:

[0105] in: The ultimate optimization objective of the online Q-learning algorithm is to achieve a final utility function of 0.

[0106] By minimizing This can optimize the weights of the execution network.

[0107] In another example, the particle swarm optimization (PSO) algorithm is used to optimize and update the hidden-to-output layer weight vectors of the evaluation network and the output layer weight matrix of the execution network, respectively, minimizing their respective optimization objective functions to obtain optimized network weight parameters. This includes: initializing PSO parameters based on the optimization objective functions, including the number of particles, maximum number of iterations, inertia weights, and acceleration coefficients; and using the PSO parameters, for the evaluation network, taking the weight vectors as particle positions and the optimization objective function of the evaluation network as the fitness function, iteratively searching for the optimal weight vector through the PSO algorithm to obtain the optimal weight vector of the evaluation network.

[0108] Particle swarm optimization (PSO) is a global optimization method inspired by the behavior of natural populations. The number of particles determines the coverage and diversity of the search space. In this system, a value related to the number of weight parameters to be optimized is typically chosen, generally 20-50 particles, to balance the search breadth and computational burden. The maximum number of iterations limits the algorithm's runtime. Considering the real-time requirements of online control systems, it is generally set to 10-30 iterations to ensure that the algorithm can complete the computation within the control cycle.

[0109] Based on the particle swarm optimization (PSO) parameters, for the evaluation network, the weight vector is used as the particle position, and the evaluation network's objective function is used as the fitness function. The optimal weight vector is found iteratively through the PSO algorithm. The core of evaluating network optimization is finding the weight parameters that most accurately estimate the Q-function value, thus minimizing the prediction error.

[0110] During the optimization iteration process, the system first evaluates the fitness of each particle's position, that is, applies the weight vector represented by the particle to the evaluation network and calculates the squared prediction error under the current state-action pair. Then, based on the fitness value, the individual historical best position and the swarm global best position of each particle are updated. Next, the algorithm updates the velocity and position of each particle based on its current position, velocity, individual best position, and global best position, causing the particles to move towards more promising weight space regions. This process is repeated until the maximum number of iterations is reached or other termination conditions are met.

[0111] The particle swarm optimization (PSO) method for optimizing neural network weights has several advantages over traditional gradient-based methods: it does not rely on gradient information and can handle non-smooth objective functions; secondly, it is a global optimization algorithm, reducing the risk of getting trapped in local optima. These characteristics make it particularly suitable for handling complex, nonlinear, and noisy control scenarios such as wastewater treatment systems. Of course, this method also faces the challenge of computational complexity increasing rapidly with the number of parameters.

[0112] Specifically, swarm intelligence algorithms are used to update the evaluation network parameters. Using the swarm intelligence PSO algorithm, the parameters can be directly updated. E c,k This allows for descent without differentiation, and improves the optimization efficiency:

[0113] The PSO algorithm can improve the performance of weight updates in the execution network and avoid the derivative calculation of the original gradient descent method.

[0114] In another example, the control strategy adjustment amount generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters, and the final control input generated through a preset coupling coefficient, adjusts the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration. This includes: generating a control strategy adjustment amount based on the fuzzy adaptive control strategy and the optimized network weight parameters; weighting the fuzzy adaptive control strategy and the control strategy adjustment amount using a preset coupling coefficient to obtain the final control input; and adjusting the oxygen transfer coefficient in the wastewater treatment system based on the final control input to change the dissolved oxygen concentration in the wastewater treatment system, thereby achieving tracking control of dissolved oxygen concentration.

[0115] The fuzzy adaptive control strategy and the control strategy adjustment are weighted and combined using a preset coupling coefficient to obtain the final control input. The coupling coefficient is a weight value between 0 and 1, which determines the proportion of the fuzzy control output and the reinforcement learning adjustment in the final control signal, and is a core parameter of the control system's fusion strategy. In practical applications, the setting of the coupling coefficient needs to balance the stability of fuzzy control and the adaptability of reinforcement learning. A larger coupling coefficient means that the fuzzy control strategy relies more on empirical knowledge, which helps maintain system stability and reliability; a smaller coupling coefficient gives reinforcement learning more adjustment space, enhancing the system's ability to cope with complex changes. This weighted combination can be understood as a control architecture of "basic control + adaptive correction," where fuzzy control provides the basic control framework, and reinforcement learning provides targeted adjustments.

[0116] Based on the final control input, the oxygen transfer coefficient in the wastewater treatment system is adjusted to change the dissolved oxygen concentration, achieving precise control of the dissolved oxygen concentration. In the wastewater treatment system, the oxygen transfer coefficient can be changed by adjusting the operating parameters of the aeration equipment, such as adjusting the blower airflow, regulating the aerator opening, or altering the stirring intensity. The final control input generated by the control system, after appropriate signal conversion and drive circuitry, directly acts on these aeration devices to achieve precise adjustment of the oxygen transfer coefficient.

[0117] In this embodiment, to demonstrate the superior performance of the proposed PSO-FOQL algorithm, it is comprehensively compared with PID control algorithm and fuzzy online Q-learning algorithm in dissolved oxygen concentration tracking control. Experimental results are as follows: Figure 4 , Figure 5 , Figure 6 As shown, compared to the other two control algorithms, the PSO-FOQL algorithm proposed in this embodiment has higher control accuracy and smaller error fluctuations in dissolved oxygen concentration. Under the action of the PSO-FOQL algorithm, the change process of the oxygen transfer coefficient, which is the control input, is as follows: Figure 7 As shown.

[0118] To evaluate the tracking performance of the proposed control algorithm, the integral of absolute error (IAE) and integral of squared error (ISE) are used to provide quantitative measures of control accuracy. The comparative analysis is shown in Table 4.

[0119] Table 4: Comparison of control performance under three weather conditions

[0120] The experimental results in Table 4 demonstrate that the PSO-FOQL-based controller in this embodiment exhibits optimal tracking accuracy under all weather conditions. Specifically, under sunny conditions, its integral absolute error (IAE) and integral squared error (ISE) are 17.3 times and 305 times lower than the PID controller, respectively; it maintains the lowest error level even under rainy conditions; and under heavy rain conditions, the proposed algorithm further reduces IAE and ISE by 3.04 times and 1.33 times, respectively, compared to the OQL algorithm, significantly demonstrating the algorithm's strong robustness. This proves that the PSO-FOQL controller can maintain optimal tracking performance under highly variable environmental conditions, effectively improving the control effect of dynamic dissolved oxygen concentration tracking in wastewater treatment processes.

[0121] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A fuzzy adaptive Q-learning control method for wastewater treatment, characterized in that, include: Obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and establish a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective; Based on the system state and the preset dissolved oxygen concentration setting, the tracking error and error increment are calculated, the tracking error and error increment are normalized, and the fuzzy set is divided by the fuzzy logic system to obtain the normalized error input. Based on the normalized error input, the proportional, integral and derivative coefficients are dynamically adjusted using Mamdani-type fuzzy inference rules. Adaptive PID parameters are obtained by defuzzification using the centroid method, and a fuzzy adaptive control strategy is generated. Based on the aforementioned nonlinear system optimization problem, an online Q-learning framework is constructed, comprising a judgment network and an execution network. The judgment network approximates the Q-function based on the input layer to hidden layer weight matrix and the adjustable hidden layer to output layer weight vector, while the execution network generates control policy adjustment quantities through the input layer weight matrix and the adjustable output layer weight matrix. Based on the online Q-learning framework, the Bellman equation is derived using the time difference analysis method. The evaluation network prediction error and the execution network prediction error are calculated based on the Bellman equation, and their respective optimization objective functions are constructed. Based on the aforementioned optimization objective function, the particle swarm optimization algorithm is applied to optimize and update the hidden layer to output layer weight vector of the evaluation network and the output layer weight matrix of the execution network, respectively, so as to minimize their respective optimization objective functions and obtain the optimized network weight parameters. The control strategy adjustment amount generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters is used to generate the final control input through a preset coupling coefficient, thereby adjusting the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration.

2. The method according to claim 1, characterized in that, The process of obtaining the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state and establishing a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective includes: Obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment, represent the wastewater treatment system as a nonlinear system, where the system state is the dissolved oxygen concentration, the control input is the oxygen transfer coefficient, the nonlinear system function is an unknown function, and establish a nonlinear system model. Based on the aforementioned nonlinear system model, the dynamic trajectory of dissolved oxygen concentration is set according to engineering experience, and the concentration values ​​are set to 1 mg / L, 2.2 mg / L and 1.8 mg / L at different time periods to generate the dissolved oxygen concentration set value. Based on the nonlinear system model and the dissolved oxygen concentration setpoint, the tracking error is defined as the difference between the system state and the setpoint, thus obtaining the nonlinear system optimization problem.

3. The method according to claim 1, characterized in that, Based on the system state and a preset dissolved oxygen concentration setpoint, the tracking error and error increment are calculated. The tracking error and error increment are then normalized, and a fuzzy set is partitioned using a fuzzy logic system to obtain the normalized error input, including: Based on the system state and the preset dissolved oxygen concentration setting, the tracking error at the current moment is calculated, and the error increment is calculated based on the tracking error at the current moment and the previous moment, thus obtaining the tracking error and the error increment. The tracking error and the error increment are normalized by using the hyperbolic tangent function, which compresses the input value to the range of negative 1 to positive 1, reduces the numerical instability caused by large error fluctuations, and obtains the normalized tracking error and error increment. Based on the normalized tracking error and error increment, each normalized input variable is divided into five fuzzy sets with Gaussian membership functions, namely negative large, negative small, zero, positive small, and positive large. The membership function of each fuzzy set adopts the form of a Gaussian function to obtain the normalized error input.

4. The method according to claim 3, characterized in that, The process involves dynamically adjusting the proportional, integral, and derivative coefficients based on the normalized error input using Mamdani-type fuzzy inference rules, obtaining adaptive PID parameters through centroid defuzzification, and generating a fuzzy adaptive control strategy, including: Based on the normalized error input, 25 Mamdani-type fuzzy inference rules are used to adjust the proportional coefficient, integral coefficient, and differential coefficient respectively. The proportional coefficient adjustment rules, integral coefficient adjustment rules, and differential coefficient adjustment rules adopt different rule mapping relationships to obtain the fuzzy inference results. Based on the fuzzy inference results, the centroid method is used to defuzzify the proportional coefficient, integral coefficient, and differential coefficient, and the weighted average value of each coefficient is calculated as the final parameter adjustment value to obtain the adaptive PID parameters. Based on the adaptive PID parameters, the defuzzified parameters are integrated through an incremental PID controller. The control law increment is calculated based on the current error, the error of the previous time step, and the errors of the previous two time steps to generate the fuzzy adaptive control strategy.

5. The method according to claim 1, characterized in that, An online Q-learning framework is constructed, comprising a judge network and an execution network. The judge network approximates the Q-function based on the input-to-hidden-layer weight matrix and an adjustable hidden-to-output-layer weight vector, including: Based on the aforementioned nonlinear system optimization problem, the Q function is defined as the cumulative sum of future utility functions including a discount factor, wherein the utility function is constructed based on the sum of squares of the tracking error and the control policy adjustment, thus obtaining the definition of the Q function; Based on the definition of the Q function, the Bellman equation is derived through time difference analysis, and a recursive relationship is established between the Q function at the current time, the Q function at the next time, and the instantaneous utility function, thus obtaining the Bellman equation. Based on the Bellman equation, a judgment network is constructed to approximate the Q function. The judgment network contains a fixed input layer to hidden layer weight matrix and an adjustable hidden layer to output layer weight vector. It adopts a hyperbolic tangent activation function, takes the tracking error and control policy adjustment as input, and outputs the approximate value of the Q function to obtain the online Q learning framework.

6. The method according to claim 5, characterized in that, The execution network generates control policy adjustment amounts through the input layer weight matrix and the adjustable output layer weight matrix, including: Based on the online Q-learning framework, the execution network is constructed to approximate and track the control input. The execution network contains a fixed input layer weight matrix and an adjustable output layer weight matrix, resulting in the execution network structure. According to the execution network structure, the tracking error is used as the input of the execution network, and a linear transformation is performed through the input layer weight matrix to obtain the transformed signal; Based on the transformed signal, a hyperbolic tangent activation function is used to perform nonlinear processing on the transformed signal, and the control strategy adjustment amount is generated through the adjustable output layer weight matrix.

7. The method according to claim 1, characterized in that, The calculation of the evaluation network prediction error and the execution network prediction error based on the Bellman equation, and the construction of their respective optimization objective functions, include: Based on the Bellman equation, the prediction error of the evaluation network is defined as the difference between the current Q-function approximation value and the sum of the instantaneous utility function and the next Q-function approximation value, thus obtaining the prediction error of the evaluation network. Based on the Bellman equation, the prediction error of the execution network is defined as the gradient of the Q function with respect to the adjustment of the control policy, where the utility function value corresponding to the final optimization objective is zero, thus obtaining the prediction error of the execution network. Each of the optimization objective functions is constructed based on the squares of the prediction error of the evaluation network and the prediction error of the execution network, respectively.

8. The method according to claim 1, characterized in that, The particle swarm optimization algorithm is applied to optimize and update the hidden-to-output layer weight vectors of the evaluation network and the output layer weight matrix of the execution network, respectively, minimizing their respective optimization objective functions to obtain the optimized network weight parameters, including: Based on the optimization objective function, the particle swarm optimization (PSO) algorithm parameters are initialized, including the number of particles, the maximum number of iterations, the inertia weight, and the acceleration coefficient, to obtain the PSO algorithm parameters. Based on the particle swarm optimization (PSO) parameters, for the evaluation network, the weight vector from the hidden layer to the output layer is used as the particle position, and the optimization objective function of the evaluation network is used as the fitness function. The optimal weight vector is found iteratively through the PSO algorithm to obtain the optimal weight vector of the evaluation network. Based on the particle swarm optimization (PSO) parameters and the optimal weight vector of the evaluation network, for the execution network, the output layer weight matrix is ​​used as the particle position, and the optimization objective function of the execution network is used as the fitness function. The optimal weight matrix is ​​found iteratively through the PSO algorithm to obtain the optimized network weight parameters.

9. The method according to claim 1, characterized in that, The control strategy adjustment amount generated by combining the fuzzy adaptive control strategy and the optimized network weight parameters is used to generate the final control input through a preset coupling coefficient, and the oxygen transfer coefficient is adjusted to achieve tracking control of dissolved oxygen concentration, including: Based on the fuzzy adaptive control strategy and the optimized network weight parameters, the control strategy adjustment amount is generated through the optimized network weight parameters to obtain the control strategy adjustment amount; The fuzzy adaptive control strategy and the control strategy adjustment amount are weighted and combined using a preset coupling coefficient to obtain the final control input; Based on the final control input, the oxygen transfer coefficient in the wastewater treatment system is adjusted to change the dissolved oxygen concentration in the wastewater treatment system, thereby achieving tracking control of the dissolved oxygen concentration.

10. A fuzzy adaptive Q-learning control system for wastewater treatment, characterized in that, include: The data acquisition module is used to obtain the dissolved oxygen concentration in the wastewater treatment system at the current moment as the system state, and to establish a nonlinear system optimization problem with dissolved oxygen concentration tracking and control as the objective. The fuzzy processing module is used to calculate the tracking error and error increment based on the system state and the preset dissolved oxygen concentration setting, normalize the tracking error and the error increment, and divide the fuzzy set through the fuzzy logic system to obtain the normalized error input. The fuzzy control module is used to dynamically adjust the proportional, integral, and derivative coefficients according to the normalized error input using Mamdani-type fuzzy inference rules, obtain adaptive PID parameters by defuzzification using the centroid method, and generate a fuzzy adaptive control strategy. An online learning module is used to construct an online Q-learning framework containing an evaluation network and an execution network based on the nonlinear system optimization problem. The evaluation network approximates the Q-function based on the input layer to hidden layer weight matrix and the adjustable hidden layer to output layer weight vector, and the execution network generates control policy adjustment quantities through the input layer weight matrix and the adjustable output layer weight matrix. The error calculation module is used to derive the Bellman equation using the time difference analysis method based on the online Q-learning framework, calculate the evaluation network prediction error and the execution network prediction error based on the Bellman equation, and construct their respective optimization objective functions. The intelligent optimization module is used to optimize and update the hidden layer to output layer weight vector of the evaluation network and the output layer weight matrix of the execution network based on the optimization objective function, thereby minimizing their respective optimization objective functions and obtaining the optimized network weight parameters. The control execution module is used to combine the fuzzy adaptive control strategy and the control strategy adjustment amount generated by the optimized network weight parameters, generate the final control input through a preset coupling coefficient, and adjust the oxygen transfer coefficient to achieve tracking control of dissolved oxygen concentration.