Intelligent agent interpretability analysis method and device for sewage treatment control
By performing action-dynamic visualization, agent decision tree modeling, Sobol sensitivity and decision path analysis on the intelligent agent for wastewater treatment control, the problem of the invisible decision-making process of the intelligent agent is solved, and its application feasibility in wastewater treatment is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies lack interpretable analysis of reinforcement learning agents in wastewater treatment control, which makes it impossible to demonstrate their decision-making process and thus affects their feasibility of application in wastewater treatment.
The action-dynamic visualization analysis method, agent decision tree modeling analysis method, Sobol sensitivity analysis method, and decision path analysis method are used to conduct interpretability analysis on the target intelligent agent of wastewater treatment control and demonstrate its decision-making process.
The decision-making process of the target intelligent agent can be demonstrated through interpretability analysis results, thereby improving its feasibility for application in wastewater treatment.
Smart Images

Figure CN121787597A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment technology, and in particular to a method and apparatus for interpretability analysis of reinforcement learning agents for wastewater treatment control. Background Technology
[0002] In the process of wastewater treatment, adjusting relevant operating parameters (such as dissolved oxygen setpoint, internal reflux ratio, and external carbon source dosage) is crucial for adaptive operation and multi-objective optimization of the process.
[0003] With the development of artificial intelligence technology, existing technologies have initially applied reinforcement learning agents, which are capable of perceiving the environment, autonomously analyzing, making decisions, and executing actions, to wastewater treatment control. However, due to the current lack of interpretability analysis for these agents, their decision-making process in wastewater treatment control cannot be demonstrated, thus hindering feasibility analysis and impeding further application of these agents in wastewater treatment technology. Summary of the Invention
[0004] This invention proposes a method and apparatus for interpretability analysis of reinforcement learning agents for wastewater treatment control. The method employs action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis to perform interpretability analysis on the target agent for wastewater treatment control, yielding interpretability analysis results. Since these results can demonstrate the target agent's decision-making process for wastewater treatment control, feasibility analysis of the agent's approach to wastewater treatment control can be achieved, facilitating further application of the agent in wastewater treatment technology.
[0005] To achieve the above objectives, the present invention adopts the following technical solution, which is discussed in the following five aspects.
[0006] In a first aspect, the present invention provides an agent interpretability analysis method for wastewater treatment control, comprising: acquiring wastewater treatment control information of a target agent over a period of time; wherein the target agent is a reinforcement learning agent based on wastewater treatment; the wastewater treatment control information includes, for each of multiple time points within a period of time, the following corresponding to each time point: influent data, effluent data, the target agent's control strategy for wastewater treatment, and the control strategy reward value; the control strategy reward value indicates the wastewater treatment effect under the control strategy. Based on the wastewater treatment control information, an action-dynamic visualization analysis method, an agent decision tree modeling analysis method, a Sobol sensitivity analysis method, and a decision path analysis method are used to determine the interpretability analysis results of the target agent's generation of control strategies over a period of time; the interpretability analysis results are used to demonstrate the process of the target agent generating control strategies.
[0007] The agent-based interpretability analysis method for wastewater treatment control provided by this invention first acquires the wastewater treatment control information of the target agent. Then, it employs four analysis methods—action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis—to analyze this wastewater treatment control information, obtaining interpretability analysis results that demonstrate the process by which the target agent generates control strategies. As can be seen from the above, the interpretability analysis results obtained by the above four analysis methods can be used to conduct feasibility analysis of the agent's control over wastewater treatment, thereby facilitating the further application of the agent in wastewater treatment technology.
[0008] In one implementation of the first aspect, the interpretability analysis results of the target agent's generation control strategy over a period of time are determined, including: Based on wastewater treatment control information and action-dynamic visualization analysis methods, a control behavior visualization diagram of the target agent generating control strategies is determined over a certain period. This visualization diagram displays the target agent's response process to influent wastewater data over a given period. Based on wastewater treatment control information and surrogate decision tree modeling analysis methods, a surrogate decision tree for the target agent's control strategies is determined over a certain period. This surrogate decision tree displays the wastewater treatment effect under the target agent's control strategies over a given period. Based on wastewater treatment control information and the Sobol sensitivity analysis method, the Sobol sensitivity index for the target agent's control strategies is determined over a certain period. The Sobol sensitivity index represents the correlation between the target agent's control strategies and influent data, effluent data, and control strategy reward values over a given period. Based on wastewater treatment control information and decision path analysis methods, the decision path for the target agent's control strategies is determined over a certain period. This decision path indicates the change process of the target agent's control strategies over a given period. The interpretability analysis results include the control behavior visualization diagram, surrogate decision tree, Sobol sensitivity index, and decision path.
[0009] In one implementation of the first aspect, determining the control behavior visualization diagram of the target agent generating control strategies over a period of time includes: dividing the wastewater treatment influent data corresponding to each time point in multiple time points within a period of time into influent data under multiple influent scenarios based on the influent scenario division conditions. For each influent scenario, a k-means clustering algorithm is used to cluster the control strategies of the target agent for wastewater treatment corresponding to the influent data under the influent scenario, obtaining multiple cluster centers; projecting the multiple cluster centers into a multi-dimensional strategy space, obtaining the control behavior visualization diagram under the influent scenario; wherein, the dimensions of the multi-dimensional strategy space include dissolved oxygen setpoint, internal recirculation ratio, and external carbon source dosage. The control behavior visualization diagram under multiple influent scenarios is determined as the control behavior visualization diagram of the control strategy.
[0010] In one implementation of the first aspect, determining the proxy decision tree for the target agent to generate control strategies over a period of time includes: using the influent data of wastewater treatment corresponding to each time point in multiple time points within a period of time as input features, using the reward value of the control strategy corresponding to each time point in multiple time points within a period of time as output features, and using the mean square error as the regression splitting criterion to construct the proxy decision tree for the control strategy.
[0011] In one implementation of the first aspect, determining the Sobol sensitivity index of the target agent's generation control strategy over a period of time includes: calculating the Sobol sensitivity index between the target agent's control strategy and the influent data, effluent data, and control strategy reward value over a period of time, and determining it as the Sobol sensitivity index of the control strategy; the Sobol sensitivity index includes a first-order sensitivity index and a total sensitivity index.
[0012] In one implementation of the first aspect, determining the decision path for the target agent to generate control strategies over a period of time includes: dividing the wastewater treatment influent data corresponding to each time point within a period of time into influent data under multiple influent scenarios based on influent scenario segmentation conditions. For each influent scenario, the control strategy for wastewater treatment of the target agent corresponding to the influent data under the influent scenario is projected into a multi-dimensional strategy space to obtain multiple strategy points; and the multiple strategy points are connected in chronological order to obtain the decision path under the influent scenario; wherein the dimensions of the multi-dimensional strategy space include dissolved oxygen setpoint, internal reflux ratio, and external carbon source dosage. The decision path under the multiple influent scenarios is determined as the decision path for the target agent to generate control strategies over a period of time.
[0013] In one implementation of the first aspect, based on the influent scenario classification conditions, the influent data of the wastewater treatment corresponding to each time point within a period of time is divided into influent data under multiple influent scenarios, including: for each time point within a period of time: when the influent flow rate of the wastewater treatment influent data corresponding to the time point is greater than a flow threshold, the wastewater treatment influent data corresponding to the time point is classified as influent data under a high-flow influent scenario. When the influent flow rate of the wastewater treatment influent data corresponding to the time point is less than or equal to the flow threshold, the wastewater treatment influent data corresponding to the time point is classified as influent data under a low-flow influent scenario.
[0014] In one implementation of the first aspect, the target agent is a reinforcement learning agent based on wastewater treatment.
[0015] In one implementation of the first aspect, the influent data for wastewater treatment includes influent flow rate, influent chemical oxygen demand (COD), and / or total Kjeldahl nitrogen (TKN); the control strategy includes the dissolved oxygen setpoint, internal recirculation ratio, and external carbon source dosage for wastewater treatment; and the effluent data for wastewater treatment includes effluent flow rate, effluent COD, and / or effluent TKN.
[0016] Secondly, this invention provides an agent interpretability analysis device for wastewater treatment control, comprising an information acquisition module and an interpretability analysis module. The information acquisition module acquires wastewater treatment control information of a target agent over a period of time. This information includes, for each of multiple time points within the period, the following data corresponding to each time point: influent data, effluent data, the target agent's control strategy for wastewater treatment, and the control strategy reward value. The control strategy reward value indicates the wastewater treatment effect under the control strategy. The interpretability analysis module, based on the wastewater treatment control information, uses action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis to determine the interpretability analysis results of the target agent's generation of control strategies over a period of time. The interpretability analysis results are used to demonstrate the process of the target agent generating control strategies.
[0017] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method described in the first aspect above or any implementation thereof.
[0018] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0019] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0020] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here. Attached Figure Description
[0021] Figure 1 This is one of the schematic diagrams of an agent interpretability analysis method for wastewater treatment control provided in the embodiments of this application; Figure 2 This is a second schematic diagram of an agent interpretability analysis method for wastewater treatment control provided in an embodiment of this application; Figure 3 This is the third schematic diagram of an agent interpretability analysis method for wastewater treatment control provided in the embodiments of this application; Figure 4 This is the fourth schematic diagram of an agent interpretability analysis method for wastewater treatment control provided in the embodiments of this application; Figure 5 This is the fifth schematic diagram of an intelligent agent interpretability analysis method for wastewater treatment control provided in the embodiments of this application; Figure 6 This is the sixth schematic diagram of an agent interpretability analysis method for wastewater treatment control provided in the embodiments of this application; Figure 7 This is a visualization of the control behavior of the target intelligent agent provided in the embodiments of this application; Figure 8 This is a schematic diagram of the agent decision tree of the target intelligent agent provided in the embodiments of this application; Figure 9 This is a schematic diagram of the Sobol sensitivity index of the target intelligent agent provided in the embodiments of this application; Figure 10 This is a schematic diagram of the decision-making path of the target intelligent agent provided in the embodiments of this application; Figure 11 This is a schematic diagram of an intelligent agent interpretability analysis device for wastewater treatment control provided in an embodiment of this application. Detailed Implementation
[0022] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.
[0023] In the embodiments of this application, "and / or" indicates a relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist simultaneously.
[0024] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0025] In the description of this invention, unless otherwise stated, "a plurality of" means two or more. For example, "a plurality of time points" means two or more time points.
[0026] The methods and apparatus provided in this application relate to wastewater treatment and can be used to perform interpretability analysis on intelligent agents to obtain interpretability analysis results that demonstrate the decision-making process of intelligent agents for wastewater treatment control.
[0027] Understandably, Artificial Intelligence (AI) is a technological system that simulates human intelligence through algorithms and data. The four core technologies of AI are perception, reasoning, learning, and action. An AI agent, as a specific application of AI, is an intelligent entity capable of perceiving its environment, making autonomous decisions, and executing actions. Its goal is to complete specific tasks through interaction with the outside world. Because the process of generating strategies for AI agents is a kind of technological black box, it is invisible to users, thus hindering the analysis of the effectiveness of the generated strategies.
[0028] To address the problem in the background art where the lack of interpretability analysis for intelligent agents prevents the demonstration of the decision-making process of intelligent agents in wastewater treatment control, thus hindering feasibility analysis and further application of intelligent agents in wastewater treatment technology, this application provides a method and apparatus for interpretability analysis of intelligent agents in wastewater treatment control. It employs action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis to perform interpretability analysis on the target intelligent agent used for wastewater treatment control, obtaining interpretability analysis results. Since the interpretability analysis results can demonstrate the decision-making process of the target intelligent agent in wastewater treatment control, the feasibility of intelligent agents in wastewater treatment applications is further improved.
[0029] For example, the intelligent agent interpretability analysis method for wastewater treatment control provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware of the computer may include: a processor, memory, network interface, user interface, communication bus, etc.
[0030] The processor controls the electronic device to perform related processing and computation tasks, such as acquiring wastewater treatment control information of the target intelligent agent over a period of time, and determining the interpretability analysis results of the target intelligent agent's generation control strategy over a period of time. The processor may include a central processing unit (CPU) or other processors, and the processor may be single-core or multi-core; for example, the processor may include multiple CPUs.
[0031] Memory is used to store computer instructions and related data, such as wastewater treatment control information and interpretable analysis results. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.
[0032] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).
[0033] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.
[0034] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.
[0035] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.
[0036] The following explains the technical terms used in the embodiments of this application.
[0037] 1. Intelligent agent.
[0038] An AI agent, as a specific application of artificial intelligence, typically comprises four core modules: a perception module, a decision-making module (also known as a technology module), an action module, and a memory module. The perception module is used to acquire information from data input and transform that information into a understandable format. The decision-making module performs logical reasoning, task planning, or strategy generation based on the information acquired by the perception module. The action module executes the decision results. The memory module stores data, supporting long-term reasoning and contextual understanding.
[0039] 2. Reinforce learning agents.
[0040] A reinforcement learning agent is a decision-making entity that learns how to achieve a goal or maximize long-term gains through autonomous interaction with its environment. The reinforcement learning (RL) agent is the core decision-making unit in reinforcement learning; it learns the optimal policy through interaction with its environment to maximize long-term cumulative rewards. The core task of reinforcement learning agents is to learn to select actions that maximize long-term cumulative rewards during interactions with the environment. Specifically, reinforcement learning agents learn through trial and error, optimizing their strategies by employing an exploration-exploitation trade-off.
[0041] In the embodiments of this application, the reinforcement learning agent can be a commonly used reinforcement learning agent in the prior art (such as a DQN agent), or a reinforcement learning agent with Bayesian optimization (BO).
[0042] 3. Action-dynamic visualization analysis method.
[0043] Action-dynamic visualization analysis refers to a research method that uses computer graphics technology to transform the action information of entities into dynamic visual representations, and then performs quantitative analysis and qualitative evaluation based on these representations. In the action-dynamic visualization analysis method used in this application, the entity refers to the target intelligent agent, the action information of the entity refers to the wastewater treatment control information of the target intelligent agent, and the result obtained by the action-dynamic visualization analysis method is a visualization diagram of the control behavior of the target intelligent agent in generating control strategies.
[0044] 4. Proxy decision tree modeling and analysis method.
[0045] Proxy decision tree modeling is a modeling technique that assists decision-making by constructing a tree-like structure. Its core idea is to decompose complex problems into multiple simple decision nodes, ultimately outputting a prediction result through the leaf nodes. Each Proxy Decision Tree obtained through this method includes: a root node (the initial decision point, selecting the best feature for the first split), multiple non-leaf nodes (intermediate decision processes, partitioning data subsets based on feature values), and multiple leaf nodes (the final decision result, outputting classification labels or regression values).
[0046] The proxy decision tree modeling and analysis method used in this application embodiment analyzes the wastewater treatment control information of the target intelligent agent, and obtains the proxy decision tree of the target intelligent agent generating control strategies through the proxy decision tree modeling and analysis method.
[0047] 4. Sobol sensitivity analysis method.
[0048] Sobol sensitivity analysis is a global sensitivity analysis method used to assess the independent effects and interactions of input variables on a model's output. The core principle of Sobol sensitivity analysis is based on variance decomposition. Specifically, it quantifies the importance of different factors by decomposing the model's output variance into variance components contributing to each input variable and their interactions.
[0049] In this embodiment, the Sobol sensitivity analysis method described above is used to analyze the wastewater treatment control information of the target intelligent agent to obtain the Sobol sensitivity index of the target intelligent agent's generation control strategy.
[0050] 5. Decision path analysis method.
[0051] Decision path analysis refers to tracking and visualizing the changes in control decisions of a reinforcement learning agent under different states over time, revealing how its control actions evolve, converge, and adjust with the system state when responding to dynamic influent conditions or operational disturbances. This analysis can reflect the agent's strategy stability and the balance mechanism between exploration and utilization in multi-objective optimization, thereby helping to understand its dynamic decision-making logic and process response mechanism. The decision path analysis method used in the embodiments of this application analyzes the wastewater treatment control information of the target agent to obtain the decision path for the target agent to generate control strategies.
[0052] For example, such as Figure 1 As shown in the embodiment of this application, an agent interpretability analysis method for wastewater treatment control includes S1-S2.
[0053] S1. Obtain the wastewater treatment control information of the target intelligent agent within a certain period of time.
[0054] The aforementioned wastewater treatment control information is time-series data, including the following for each time point within a given period: influent data, effluent data, the target agent's control strategy for wastewater treatment, and the control strategy reward value. The given period can be 10 days or 15 days; the time interval between two adjacent time points can be the same or different, and can be 10 minutes or 15 minutes. This embodiment does not limit the duration of the given period or the duration of the time interval between adjacent time points.
[0055] Optionally, the influent data for the aforementioned wastewater treatment may include influent flow rate, influent chemical oxygen demand (COD), and / or influent total Kjeldahl nitrogen (TKN). The effluent data for the aforementioned wastewater treatment may include effluent flow rate, effluent chemical oxygen demand (COD), and / or effluent total Kjeldahl nitrogen (TKN). The aforementioned control strategies may include the dissolved oxygen (DO) setpoint, internal recirculation ratio (IMLR), and external carbon source dosage for wastewater treatment. The reward value for the aforementioned control strategy indicates the wastewater treatment effect under the control strategy. The reward value for the aforementioned control strategy can be calculated based on parameters such as the wastewater treatment cost index, water quality index, and energy index.
[0056] It should be noted that the aforementioned wastewater treatment cost index, water quality index, and energy index are commonly used indices in the field of wastewater treatment, and this application embodiment does not limit the method of obtaining these three indices. Furthermore, since the control strategy reward value is a commonly used technique in reinforcement learning, this application embodiment will not elaborate on the calculation method of the aforementioned control strategy reward value.
[0057] In this embodiment, the target agent is a reinforcement learning agent based on wastewater treatment. This target agent can generate a control strategy based on influent and / or effluent data from wastewater treatment, and calculate the control strategy reward value based on the control strategy. The wastewater treatment control information can be wastewater treatment control information in a simulated environment or wastewater treatment control information from actual wastewater treatment equipment. This embodiment does not limit the working environment for the target agent to generate the control strategy.
[0058] It should be noted that the aforementioned reinforcement learning agent based on wastewater treatment refers to a reinforcement learning agent trained in a simulated wastewater treatment environment. Therefore, the aforementioned target agent can also be considered a type of reinforcement learning agent. The aforementioned simulated wastewater treatment environment can be constructed based on knowledge in this technical field or based on actual wastewater treatment data. The aforementioned wastewater treatment technology can be biological nitrogen and phosphorus removal (BNR) technology or other technologies involved in wastewater treatment. This application does not limit the reinforcement learning process of the aforementioned target agent or the aforementioned wastewater treatment technology.
[0059] Furthermore, the aforementioned target intelligent agent can be a reinforcement learning intelligent agent based on wastewater treatment (also known as an RL agent) or a Bayesian optimized reinforcement learning intelligent agent (also known as an RL+BO agent). This application embodiment does not further limit the optimization method of the aforementioned target intelligent agent.
[0060] S2. Based on the wastewater treatment control information, the interpretability analysis results of the target agent generation control strategy are determined by using the action-dynamic visualization analysis method, the agent decision tree modeling analysis method, the Sobol sensitivity analysis method, and the decision path analysis method.
[0061] Specifically, the results of the above interpretability analysis are used to visualize the process by which the target intelligent agent generates control strategies.
[0062] Optionally, the results of the above interpretability analysis may include a visualization of control behavior, an agent decision tree, the Sobol sensitivity index, and decision paths. Combined with... Figure 1 ,like Figure 2 As shown, S2 above includes S201-S204.
[0063] S201. Based on wastewater treatment control information and action-dynamic visualization analysis methods, determine the visualization diagram of the control behavior of the target intelligent agent in generating control strategies over a period of time.
[0064] The above visualization of control behavior shows the response process of the target intelligent agent to the influent data of sewage treatment over a period of time.
[0065] In one application scenario of the aforementioned S201, combined with Figure 2 ,like Figure 3 As shown, S201 above includes S2011-S2013.
[0066] S2011. Based on the influent scenario classification conditions, the influent data of sewage treatment corresponding to each time point in a period of time is divided into influent data under multiple influent scenarios.
[0067] In some embodiments of the above application scenarios, S2011 includes the following.
[0068] For each of multiple time points within a given period: when the influent flow rate in the wastewater treatment data corresponding to that time point is greater than a flow rate threshold, the wastewater treatment influent data corresponding to that time point is classified as influent data under a high-flow-rate scenario. When the influent flow rate in the wastewater treatment influent data corresponding to that time point is less than or equal to the flow rate threshold, the wastewater treatment influent data corresponding to that time point is classified as influent data under a low-flow-rate scenario. It is understood that the value of the aforementioned flow rate threshold can be set according to the working environment of the target intelligent agent or based on experience. This application embodiment does not limit the value or setting method of the aforementioned flow rate threshold.
[0069] In one embodiment of the above embodiments, when two scenarios, high-flow-rate water intake and low-flow-rate water intake, are obtained based on the above water intake scenario classification conditions, the high-flow-rate water intake and low-flow-rate water intake scenarios are further classified by the carbon-nitrogen ratio (C / N) threshold, resulting in four water intake scenarios: high-flow-high C / N water intake scenario, low-flow-high C / N water intake scenario, high-flow-low C / N water intake scenario, and low-flow-low C / N water intake scenario.
[0070] The process of dividing the high-flow-rate water inlet scenario and the low-flow-rate water inlet scenario in the above embodiments is as follows: For influent data in high-flow-rate scenarios, if the influent C / N ratio is greater than the carbon-nitrogen ratio threshold, the influent data is classified as high-flow-high C / N influent data; if the influent C / N is less than or equal to the carbon-nitrogen ratio threshold, the influent data is classified as high-flow-low C / N influent data. For influent data in low-flow-rate scenarios, if the influent C / N is greater than the carbon-nitrogen ratio threshold, the influent data is classified as low-flow-high C / N influent data; if the influent C / N is less than or equal to the carbon-nitrogen ratio threshold, the influent data is classified as low-flow-low C / N influent data. The influent C / N can be calculated using the influent COD; the calculation process for the influent C / N is not detailed in this embodiment.
[0071] S2012. For each of the multiple water inflow scenarios, the k-means clustering algorithm is used to cluster the control strategies of the target intelligent agent corresponding to the water inflow data in the water inflow scenario for wastewater treatment, resulting in multiple cluster centers. The multiple cluster centers are projected into the multi-dimensional policy space to obtain a visualization of the control behavior in the water inflow scenario.
[0072] The dimensions of the aforementioned multidimensional strategy space include the dissolved oxygen setpoint, the internal reflux ratio, and the amount of external carbon source added. Since the k-means clustering algorithm is a commonly used technique in this technical field, the clustering process of the k-means clustering algorithm will not be described in detail in this embodiment.
[0073] S2013. The control behavior visualization diagrams under various water ingress scenarios are identified as the control behavior visualization diagrams of the control strategy.
[0074] As shown above, the action-dynamic visualization analysis method can be used to perform interpretability analysis on the target agent, revealing the agent's response patterns under different influent conditions from a process kinetics perspective. By comparing and analyzing control actions such as influent flow rate, C / N ratio, DO, IMLR, and carbon source dosage, it can be intuitively demonstrated whether the behavior of the target agent conforms to known biochemical mechanisms of wastewater treatment (such as denitrification and nitrification).
[0075] S202. Based on wastewater treatment control information and agent decision tree modeling and analysis methods, determine the agent decision tree for the target agent to generate control strategies within a certain period of time.
[0076] The above agent decision tree demonstrates the sewage treatment effect under the control strategy of the target agent over a period of time.
[0077] In one application scenario of the aforementioned S202, combined with Figure 3 ,like Figure 4 As shown, S202 above includes S2021.
[0078] S2021. Using the influent data of wastewater treatment corresponding to each time point in a period of time as the input feature, the control strategy reward value corresponding to each time point in a period of time as the output feature, and the mean square error as the regression splitting criterion, a proxy decision tree of the control strategy is constructed.
[0079] For example, the above mean square error MSE The calculation formula is shown below.
[0080] in, n The number of times. i For the i-th time point among multiple time points, For the water inflow data at time point i, This represents the average of influent data at multiple time points.
[0081] Understandably, since the method of constructing the agent decision tree is a common technique in this technical field, the embodiments of this application will not elaborate further on the above-mentioned agent decision tree construction process.
[0082] As can be seen from the above, using the agent decision tree modeling and analysis method to perform interpretability analysis on the target agent can transform the complex process of the target agent generating control strategies into simplified if-then rules from a logical structure perspective, thereby revealing the internal decision-making path of the target agent. This method maps the high-dimensional and continuous decision-making process of the target agent into a readable tree structure, enabling relevant technical personnel to understand the correspondence between the target agent's decisions and traditional control logic (such as DO thresholds).
[0083] S203. Based on wastewater treatment control information and the Sobol sensitivity analysis method, determine the Sobol sensitivity index of the target agent generation control strategy over a period of time.
[0084] The Sobol sensitivity index mentioned above indicates the degree of correlation between the target agent's control strategy and the influent data, effluent data, and control strategy reward value over a period of time.
[0085] In one application scenario of the aforementioned S203, combined with Figure 4 ,like Figure 5 As shown, S203 above includes S2031.
[0086] S2031. Calculate the Sobol sensitivity index between the control policy of the target agent and the influent data, effluent data and control policy reward value over a period of time, and determine it as the Sobol sensitivity index of the control policy.
[0087] In one embodiment of the above application scenario, the Sobol sensitivity index includes a first-order sensitivity index and a total sensitivity index. The aforementioned first-order sensitivity index... and total sensitivity index The calculation formula is shown below.
[0088] in, Let i be the i-th variable in the input variables of the target agent. The control strategy output for the target intelligent agent. Let be the set of all variables in the input variables except for the i-th input variable. To output the result given the value of the i-th variable. Y Conditional expectation, Output the result when all input variables except the i-th variable take values. Y Conditional expectation, For the output results Y The total variance.
[0089] The first-order sensitivity index in the Sobol sensitivity index reflects the direct impact of a single input variable on the target agent's output control strategy; the total sensitivity index considers both the direct effect of a single input variable and the interaction effect of other input variables. Therefore, Sobol sensitivity analysis can be used to comprehensively evaluate the importance of each input variable of the target agent and its role in the target agent's control decision-making, helping to explain how the target agent generates control strategies based on different input information.
[0090] As can be seen from the above, using the Sobol sensitivity analysis method to perform interpretability analysis on the target agent can quantitatively analyze the main and interaction effects of different state variables (such as flow rate, COD, NH4⁺-N, and cost index) on control performance (reward value) from the perspective of variable contribution, and clarify which variables dominate the decision-making and policy evolution of the target agent.
[0091] S204. Based on wastewater treatment control information and decision path analysis methods, determine the decision path for the target intelligent agent generation control strategy within a certain period of time.
[0092] The decision path described above indicates the process of changes in the control strategy of the target agent over a period of time.
[0093] In one application scenario of S204 mentioned above, combined with Figure 5 ,like Figure 6As shown, S204 includes S2041-S2043.
[0094] S2041. Based on the influent scenario classification conditions, the influent data of sewage treatment corresponding to each time point in a period of time is divided into influent data under multiple influent scenarios.
[0095] It is understood that the specific implementation process of S2041 above refers to the relevant description in one embodiment of an application scenario of S201 above, and will not be repeated here in the embodiments of this application.
[0096] S2042. For each of the multiple water inflow scenarios, the control strategy of the target intelligent agent for wastewater treatment corresponding to the water inflow data in the water inflow scenario is projected into the multi-dimensional strategy space to obtain multiple strategy points; and the multiple strategy points are connected in chronological order to obtain the decision path in the water inflow scenario.
[0097] The dimensions of the multidimensional strategy space include dissolved oxygen setpoint, internal reflux ratio, and external carbon source dosage.
[0098] S2043. Determine the decision paths under various water ingress scenarios as the decision paths for the target intelligent agent generation control strategy over a period of time.
[0099] As can be seen from the above, by using the decision path analysis method to perform interpretability analysis on the target intelligent agent, the path of the target intelligent agent's control strategy changes over time can be characterized from the perspective of time evolution. This reveals the stability, convergence speed, and adaptive process of the target intelligent agent in generating control strategies, thus demonstrating the dynamic mechanism by which the target intelligent agent learns to control.
[0100] In one implementation of this application, the target intelligent agent performs wastewater treatment control in two scenarios as follows.
[0101] The influent data in the wastewater treatment environment includes influent flow rate, influent COD, and influent TKN; the effluent data includes effluent flow rate, effluent COD, and effluent TKN; the control strategies include DO setpoint, IMLR, and external carbon source dosage; and the influent scenarios are divided into four types: high flow rate-high C / N influent, low flow rate-high C / N influent, high flow rate-low C / N influent, and low flow rate-low C / N influent.
[0102] Using the aforementioned RL agent and RL+BO agent as target agents respectively, the aforementioned RL agent and RL+BO agent are used to control the wastewater treatment environment in the aforementioned scenario, thereby obtaining the wastewater treatment control information of the aforementioned RL agent over a certain period of time and the wastewater treatment control information of the aforementioned RL+BO agent over a certain period of time, which correspond to the scenario.
[0103] The interpretability analysis process and results of the target intelligent agent in the above scenario are as follows.
[0104] The action-dynamic visualization analysis method in S201 above is used to analyze the wastewater treatment control information of the RL agent over a period of time and the wastewater treatment control information of the RL+BO agent over a period of time. The resulting control behavior visualization diagrams of the RL agent and the RL+BO agent are as follows: Figure 7 As shown.
[0105] refer to Figure 7 , Figure 7 (A) demonstrates the control behavior of RL agents and RL+BO agents in a high-flow-with-high-C / N influent scenario. Figure 7 (B) Demonstrates the control behavior of RL agents and RL+BO agents in a high-flow-low-C / N influent scenario. Figure 7 (C) demonstrates the control behavior of the RL agent and the RL+BO agent in a low-infuent flow with high C / N inflow scenario. Figure 7 (D) illustrates the control behavior of RL agents and RL+BO agents in a low infuent flow with low C / N scenario.
[0106] Depend on Figure 7 It can be seen that under high flow rate and high C / N influent conditions, the RL+BO aerator slightly increased the intrinsic reaction rate (IMLR) (from 31.49 L / day to 38.46 L / day) and significantly reduced the external carbon dosage (from 2.50 L / day to 1.06 L / day), while maintaining a similar DO setpoint. From an engineering perspective, this reflects a shift towards utilizing endogenous carbon sources, i.e., enhancing biomass contact by increasing the IMLR to promote nitrification without the need for additional external carbon sources. The resulting higher return (2.50 vs. 1.23) indicates that the RL+BO aerator achieves more cost-effective nitrogen removal under high flow rate influent conditions rich in organic matter.
[0107] The agent decision tree modeling and analysis method in S202 above is used to analyze the wastewater treatment control information of the RL agent over a period of time and the wastewater treatment control information of the RL+BO agent over a period of time, respectively. The resulting agent decision trees of the RL agent and the RL+BO agent are as follows: Figure 8 As shown. Reference Figure 8 , Figure 8 (A) shows the agent decision tree of the RL agent. Figure 8 (B) shows the agent decision tree of the RL+BO agent.
[0108] Depend on Figure 8 It can be seen that the agent decision tree structure of the RL agent is more complex, containing multiple split thresholds, but the reward values of many leaf nodes are low (e.g., ≤ 2.06), reflecting its limited exploration range of potential control options. In contrast, the agent decision tree structure of the RL+BO agent is simpler and quickly converges to high-reward leaf nodes (e.g., > 4.0), demonstrating a superior overall control strategy. Furthermore, Figure 8 (A) The agent decision tree of the RL agent mainly focuses on the effluent quality index and the cost index; while Figure 8 (B) The agent decision tree of the RL+BO agent achieves a more balanced trade-off between effluent quality, operating costs and influent variability, and obtains a higher reward value, which reflects the improvement of multi-objective optimization performance.
[0109] The Sobol sensitivity analysis method in S203 above was used to analyze the wastewater treatment control information of the RL agent over a period of time and the wastewater treatment control information of the RL+BO agent over a period of time, respectively. The obtained Sobol sensitivity indices of the RL agent and the RL+BO agent are as follows: Figure 9 As shown. Reference Figure 9 , Figure 9 (C) shows the first-order sensitivity index of the RL agent and the RL+BO agent (corresponding to S1 in the figure). Figure 9 (D) shows the total sensitivity index of the RL agent and the RL+BO agent (corresponding to S in the figure). t ).
[0110] Depend on Figure 9 It can be seen that the high S1 and S2 of the control strategy of the RL agent are related to the control strategy of the RL t The values are mainly concentrated around the Effluent Quality Index and the Cost Index, indicating that its optimization objective is relatively narrow and easily overlooks the dynamic changes in influent conditions; conversely, the control strategy of the RL+BO agent focuses on S1 and S2. tThe analysis showed that they were more sensitive to influent flow, indicating that they have a stronger ability to respond to operational disturbances related to flow rate changes.
[0111] The decision path analysis method in S204 above is used to analyze the wastewater treatment control information of the RL agent over a period of time and the wastewater treatment control information of the RL+BO agent over a period of time, respectively. The resulting decision paths of the RL agent and the RL+BO agent are as follows: Figure 10 As shown.
[0112] refer to Figure 10 , Figure 10 (A) illustrates the decision-making paths of the RL agent and the RL+BO agent in a high-flow-with-high-C / N influent scenario. Figure 10 (B) illustrates the decision-making paths of RL agents and RL+BO agents in a high-flow-with-low-C / N inflow scenario. Figure 10 (C) illustrates the decision paths of the RL agent and the RL+BO agent in a low-infuent flow with high C / N inflow scenario. Figure 10 (D) illustrates the decision paths of RL agents and RL+BO agents in a low infuent flow with low C / N scenario.
[0113] Depend on Figure 10It can be seen that in the high flow-high C / N influent scenario, the decision trajectories of the RL agent and the RL+BO agent are generally consistent. However, the RL+BO agent exhibits a smoother transition in the changes of DO setpoint and IMLR ratio, reflecting its more stable and generalizable decision logic, and achieving comparable or higher rewards with less fluctuation. In the high flow-low C / N influent scenario, the control strategies of the RL agent and the RL+BO agent differ more significantly. The RL agent's adjustments to DO and carbon addition are more dispersed, while the RL+BO agent can quickly converge to a compact and optimal reward trajectory with minimal fluctuation. This indicates that BO helps optimize the strategy under carbon-constrained conditions, enabling the agent to avoid suboptimal actions. In the low flow-high C / N influent scenario, the RL agent's control strategy exhibits oscillating behavior in control actions, especially in the external carbon addition amount; in contrast, the RL+BO agent maintains a narrower trajectory range in DO and IMLR, demonstrating more efficient utilization of carbon sources and a simpler and more stable control logic, which helps improve operational reliability. In low-flow-low-C / N influent scenarios, the RL agent exhibits significant fluctuations in external carbon injection and IMLR adjustment, demonstrating insufficient robustness; while the RL+BO agent shows clear and monotonous trends across all control dimensions, maintaining a high reward level even with limited resources.
[0114] From the above Figures 7-10 The analysis reveals that the controller action-influent dynamic analysis method directly reflects the correspondence between agent decision-making and process mechanisms, and is the only method that can verify the rationality of control logic from a process visualization perspective. The surrogate decision tree modeling analysis method can transform continuous black-box strategies into concise and operable rule structures, possessing the highest engineering portability. The Sobol sensitivity analysis method provides a unique quantitative interpretation dimension, calculating the importance and interaction strength of input variables, providing a scientific basis for subsequent sensor deployment and feature selection; the decision trajectory analysis method is the only method that can demonstrate the real-time response and convergence behavior of the agent under dynamic disturbances, revealing the balance between strategy stability and exploration / utilization.
[0115] Based on the combined results of the four analysis methods, a multi-level, operable, and engineering-oriented explanation of the target intelligent agent can be achieved from the following four aspects.
[0116] From a mechanistic perspective: the actions of the RL agent are verified through controller action-influent dynamic analysis to determine whether they conform to the laws of biological reaction kinetics, such as the rationality of reducing DO and increasing IMLR when the flow rate is high and the C / N ratio is low, thus proving that the algorithm decision is consistent with the process mechanism.
[0117] From a logical perspective (Decision Transparency): Agent decision tree modeling simplifies complex strategies into "if-then" rules, clearly defining key thresholds (such as effluent quality index of 4454 and flow rate of 17.6 m³). 3 / d etc.) to achieve the transformation from "black box decision-making" to "readable logic".
[0118] From a quantitative attribution perspective: Sobol analysis quantitatively reveals the key driving variables of control behavior and their interaction effects, providing a quantifiable "causal explanation" for agents and a basis for operators to optimize monitoring points and control weights.
[0119] From a dynamic perspective: Decision trajectory analysis demonstrates how the actions of the RL agent evolve with water inflow disturbances, verifying its stability and adaptability, and revealing the path mechanism by which RL+BO achieves high reward, low energy consumption, and low oscillation under dynamic conditions.
[0120] Furthermore, these four types of analysis work together to form a multidimensional interpretability framework that can map the control logic of intelligent agents from the abstract algorithm layer to the specific engineering behavior layer; transform complex "black box strategies" into consistent, verifiable, and executable control rules; enable operators to understand, audit, and trust the decision-making process of intelligent agents; and provide a transparent and regulatory compliance foundation for the industrial deployment of RL control.
[0121] In summary, the agent interpretability analysis method for wastewater treatment control provided in this application first acquires the wastewater treatment control information of the target agent, and then uses four analysis methods—action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis—to analyze the aforementioned wastewater treatment control information, obtaining interpretability analysis results that demonstrate the process by which the target agent generates control strategies. As can be seen from the above, the interpretability analysis results obtained by the above four analysis methods can be used to conduct feasibility analysis of the agent's control over wastewater treatment, thereby facilitating the further application of the agent in wastewater treatment technology.
[0122] Accordingly, embodiments of this application provide an intelligent agent interpretability analysis device for wastewater treatment control, such as... Figure 11 As shown, it includes an information acquisition module 501 and an interpretability analysis module 502.
[0123] The information acquisition module 501 is used to acquire the wastewater treatment control information of the target intelligent agent over a period of time. The wastewater treatment control information includes, for each of the multiple time points within that period, the following data corresponding to each time point: influent data, effluent data, the target intelligent agent's control strategy for wastewater treatment, and the control strategy reward value. The control strategy reward value indicates the wastewater treatment effect under the control strategy. For example, the information acquisition module 501 is used to implement step S1 of the above method.
[0124] The interpretability analysis module 502 is used to determine the interpretability analysis results of the target agent's generation control strategy over a period of time, based on wastewater treatment control information, using action-dynamic visualization analysis, agent decision tree modeling analysis, Sobol sensitivity analysis, and decision path analysis. The interpretability analysis results are used to demonstrate the process of the target agent generating the control strategy. For example, the interpretability analysis module 502 is used to implement step S2 of the above method.
[0125] Optionally, the interpretability analysis module 502 is specifically used for: determining a control behavior visualization diagram of the target agent generating control strategies over a period of time based on wastewater treatment control information and action-dynamic visualization analysis methods; the control behavior visualization diagram shows the response process of the target agent to the influent data of wastewater treatment over a period of time; determining a proxy decision tree of the target agent generating control strategies over a period of time based on wastewater treatment control information and proxy decision tree modeling analysis methods; the proxy decision tree shows the wastewater treatment effect under the action of the target agent's control strategies over a period of time; determining the Sobol sensitivity index of the target agent generating control strategies over a period of time based on wastewater treatment control information and Sobol sensitivity analysis methods; the Sobol sensitivity index represents the degree of correlation between the target agent's control strategies and influent data, effluent data, and control strategy reward values over a period of time; determining the decision path of the target agent generating control strategies over a period of time based on wastewater treatment control information and decision path analysis methods; the decision path indicates the change process of the target agent's control strategies over a period of time; wherein, the interpretability analysis results include the control behavior visualization diagram, the proxy decision tree, the Sobol sensitivity index, and the decision path. For example, the interpretability analysis module 502 is specifically used to implement S201-S204 of the above method.
[0126] The modules of the intelligent agent interpretability analysis device for wastewater treatment control described above can also be used to perform other steps in the above method embodiments. All relevant content involved in the above method embodiments can be referred to in the functional description of the corresponding functional module, and will not be repeated here.
[0127] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods described in the above embodiments. The processor can implement the information acquisition module 501 and the interpretability analysis module 502; the memory can also be used to store wastewater treatment control information and interpretability analysis results, etc.
[0128] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.
[0129] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described in the above embodiments.
[0130] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for interpretability analysis of intelligent agents for wastewater treatment control, characterized in that, include: Acquire wastewater treatment control information of a target intelligent agent over a period of time; wherein, the target intelligent agent is a reinforcement learning intelligent agent based on wastewater treatment; the wastewater treatment control information includes, for each of the multiple time points within the period of time, the following: influent data, effluent data, the target intelligent agent's control strategy for wastewater treatment, and the control strategy reward value; the control strategy reward value indicates the wastewater treatment effect under the control strategy. Based on the wastewater treatment control information, the action-dynamic visualization analysis method, the agent decision tree modeling analysis method, the Sobol sensitivity analysis method, and the decision path analysis method are used to determine the interpretability analysis results of the target agent generating the control strategy within the specified time period; the interpretability analysis results are used to demonstrate the process by which the target agent generates the control strategy.
2. The method as described in claim 1, characterized in that, The interpretability analysis results of the control strategy generated by the target agent within a certain period of time include: Based on the wastewater treatment control information and the action-dynamic visualization analysis method, a control behavior visualization diagram of the target intelligent agent generating the control strategy is determined within a certain period of time; the control behavior visualization diagram shows the response process of the target intelligent agent to the influent data of the wastewater treatment within the specified period of time. Based on the wastewater treatment control information and the agent decision tree modeling and analysis method, the agent decision tree for the control strategy generated by the target agent within a certain period of time is determined; the agent decision tree shows the wastewater treatment effect under the control strategy of the target agent within the certain period of time. Based on the wastewater treatment control information and the Sobol sensitivity analysis method, the Sobol sensitivity index of the target agent in generating the control strategy within a certain period of time is determined; the Sobol sensitivity index represents the degree of correlation between the target agent's control strategy and the influent data, effluent data, and the control strategy reward value within the certain period of time. Based on the wastewater treatment control information and decision path analysis method, the decision path for the target agent to generate the control strategy within a certain period of time is determined; the decision path indicates the change process of the target agent's control strategy within a certain period of time. The interpretability analysis results include the control behavior visualization, the agent decision tree, the Sobol sensitivity index, and the decision path.
3. The method as described in claim 2, characterized in that, The step of generating a visualization of the control behavior of the target agent within a certain period of time includes: Based on the criteria for classifying influent scenarios, the influent data of wastewater treatment corresponding to each time point within a certain period of time is divided into influent data under multiple influent scenarios. For each of the various influent scenarios, the k-means clustering algorithm is used to cluster the control strategies of the target agent corresponding to the influent data in the influent scenario for wastewater treatment, resulting in multiple cluster centers. The multiple cluster centers are then projected into a multi-dimensional strategy space to obtain a visualization of the control behavior in the influent scenario. The dimensions of the multi-dimensional strategy space include dissolved oxygen setpoint, internal reflux ratio, and external carbon source dosage. The control behavior visualization diagrams under the various water ingress scenarios are determined as the control behavior visualization diagrams of the control strategy.
4. The method as described in claim 2, characterized in that, The process of generating the agent decision tree for the control strategy by the target agent within a certain period of time includes: Using the influent data of the wastewater treatment corresponding to each time point within a certain period as the input feature, the control strategy reward value corresponding to each time point within a certain period as the output feature, and the mean square error as the regression splitting criterion, a proxy decision tree for the control strategy is constructed.
5. The method as described in claim 2, characterized in that, The determination of the Sobol sensitivity index of the target agent in generating the control strategy within a certain period of time includes: Calculate the Sobol sensitivity index between the control policy of the target agent and the influent data, effluent data and the reward value of the control policy over the specified time period, and determine it as the Sobol sensitivity index of the control policy; the Sobol sensitivity index includes the first-order sensitivity index and the total sensitivity index.
6. The method as described in claim 2, characterized in that, The step of determining the decision path for the target agent to generate the control strategy within a certain period of time includes: Based on the criteria for classifying influent scenarios, the influent data of wastewater treatment corresponding to each time point within a certain period of time is divided into influent data under multiple influent scenarios. For each of the various influent scenarios, the control strategy for wastewater treatment of the target agent corresponding to the influent data in the influent scenario is projected into a multi-dimensional strategy space to obtain multiple strategy points; and the multiple strategy points are connected in chronological order to obtain the decision path in the influent scenario; wherein, the dimensions of the multi-dimensional strategy space include dissolved oxygen setpoint, internal reflux ratio and external carbon source dosage. The decision paths under the various water ingress scenarios are determined as the decision paths for the target agent to generate the control strategy within the specified time period.
7. The method as described in any one of claims 3 or 6, characterized in that, The process involves classifying wastewater treatment influent data for each time point within a given period into multiple influent scenarios based on influent scenario classification conditions, including: For each of the multiple time points within the aforementioned time period: When the influent flow rate in the wastewater treatment influent data corresponding to the time point is greater than the flow rate threshold, the wastewater treatment influent data corresponding to the time point is classified as influent data under the high flow rate influent scenario. When the influent flow rate in the wastewater treatment influent data corresponding to the time point is less than or equal to the flow rate threshold, the wastewater treatment influent data corresponding to the time point is classified as influent data under the low flow rate influent scenario.
8. The method as described in claim 1, characterized in that, The influent data for the wastewater treatment includes influent flow rate, influent chemical oxygen demand (COD), and / or total Kjeldahl nitrogen (TKN); the control strategy includes dissolved oxygen setpoint, internal reflux ratio, and external carbon source dosage for the wastewater treatment; the effluent data for the wastewater treatment includes effluent flow rate, effluent COD, and / or effluent TKN.
9. An agent-based interpretability analysis device for wastewater treatment control, characterized in that, It includes an information acquisition module and an interpretability analysis module; The information acquisition module is used to acquire the wastewater treatment control information of the target intelligent agent within a certain period of time; the wastewater treatment control information includes the following for each time point in the multiple time points within the period of time: influent data, effluent data, the control strategy of the target intelligent agent for wastewater treatment, and the control strategy reward value; the control strategy reward value indicates the wastewater treatment effect under the control strategy. The interpretability analysis module is used to determine the interpretability analysis results of the control strategy generated by the target agent within the specified time period, based on the wastewater treatment control information, using action-dynamic visualization analysis method, agent decision tree modeling analysis method, Sobol sensitivity analysis method, and decision path analysis method. The interpretability analysis results are used to demonstrate the process by which the target agent generates the control strategy.
10. An electronic device, characterized in that, The device includes a processor and a memory coupled to the processor; the memory is used to store computer instructions, which, when the electronic device is running, are executed by the processor to cause the electronic device to perform the method as described in any one of claims 1 to 8.