Deep reinforcement learning control strategy evaluation method and device for urban drainage system

By evaluating deep reinforcement learning control strategies through a multi-dimensional evaluation system and a high-fidelity simulation platform, the problems of singularity and inconsistency in existing evaluation methods are solved, enabling a comprehensive and objective evaluation of urban drainage systems and improving the system's flood resistance, economy, and stability.

CN121806491APending Publication Date: 2026-04-07CHINA THREE GORGES CORPORATION +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing evaluation methods for urban drainage system control strategies are limited by their single evaluation dimension and inconsistent processes, failing to provide suitable comparative criteria and making it difficult to comprehensively and objectively evaluate the performance of deep reinforcement learning control strategies.

Method used

A multi-dimensional evaluation system is adopted, including the degree of achievement of control objectives, decision execution efficiency, policy robustness and system adaptability. A deep reinforcement learning control strategy is constructed and compared with traditional methods. The evaluation is carried out by simulating various rainfall scenarios through a high-fidelity simulation platform.

Benefits of technology

It enables a comprehensive and objective evaluation of deep reinforcement learning control strategies, improves the flood resistance of urban drainage systems, reduces pollution emissions, optimizes energy consumption and costs, and enhances system resilience and sustainability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121806491A_ABST
    Figure CN121806491A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of drainage control evaluation, in particular to a deep reinforcement learning control strategy evaluation method and device of an urban drainage system, and the method comprises the steps: constructing a deep reinforcement learning control strategy of the urban drainage system according to a control target of the urban drainage system; based on the target operation scene, analyzing the control target realization degree, decision execution efficiency, strategy robustness and system adaptability of the deep reinforcement learning control strategy and the pre-constructed parallel comparison control strategy by using a pre-constructed multi-dimensional evaluation system so as to generate analysis results of the reinforcement learning control strategy and the comparison control strategy; and evaluating the control performance of the deep reinforcement learning control strategy according to the analysis result. Therefore, the problem that the performance of the urban drainage system is difficult to evaluate effectively due to the fact that a current evaluation method is single in evaluation dimension, non-uniform in evaluation process and incapable of providing a proper comparison basis in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of drainage control evaluation technology, and in particular to a method and apparatus for evaluating deep reinforcement learning control strategies for urban drainage systems. Background Technology

[0002] With the acceleration of urbanization and the frequent occurrence of extreme weather events, urban drainage systems are facing increasing operational pressure. Traditional control models relying on manual experience are no longer adequate for complex and ever-changing operational scenarios. Automation and intelligent control strategies have become a key development direction for improving the operational efficiency of drainage systems. Against this backdrop, how to scientifically, comprehensively, and objectively evaluate the performance of various control strategies and provide a reliable basis for optimizing and selecting drainage system control schemes has become a core issue that urgently needs to be addressed in the fields of municipal engineering and intelligent control.

[0003] In related technologies, existing methods for evaluating urban drainage system control strategies generally suffer from prominent shortcomings such as limited evaluation dimensions, insufficient objectivity, and limited universality. Regarding evaluation indicators, traditional methods often focus on single operational performance indicators such as drainage volume and duration of waterlogging, frequently neglecting key performance dimensions such as the system's robustness under extreme rainfall, its adaptability to changes in weather patterns, and the speed of control command response. This results in an inability to form a comprehensive understanding of the overall capabilities of the control strategy. While some evaluation systems attempt to expand the scope of indicators, the lack of systematic correlation between indicators makes it difficult to form a scientific, composite evaluation logic and accurately reflect the practical application value of the control strategy.

[0004] In terms of evaluation processes and test scenario design, existing technologies lack unified standards and specifications. Different control strategies are often tested under varying operating conditions. For example, some evaluations are based on short-term historical rainfall data for specific regions, while others use simple simulated rainfall scenarios, resulting in a lack of repeatability and comparability in the evaluation results. At the same time, the scenario coverage of the test set is severely insufficient, mostly including only routine rainfall scenarios, and lacking consideration for extreme weather events such as short-term heavy rainfall and continuous rainstorms. This makes it difficult to generalize the evaluation results to complex and ever-changing real-world operating environments, greatly reducing their universality and representativeness.

[0005] Crucially, current evaluation methods have significant shortcomings in comparing different types of control strategies. With the increasing application of intelligent algorithms such as deep reinforcement learning in drainage system control, there is an urgent need to establish an effective comparison system between them and classic methods such as traditional empirical rule control and optimization computational control. However, existing evaluations either lack clear reference baselines or the selection of baseline strategies lacks scientific rigor, failing to provide quantitative comparative evidence for improving the performance of deep reinforcement learning strategies. This hinders the practical application and iterative optimization of intelligent control technology in drainage systems.

[0006] Therefore, developing an objective quantitative evaluation method with multi-dimensional evaluation capabilities, unified standard procedures, and broad scenario coverage has become an urgent need to promote the development of intelligent control technology for urban drainage systems, and has important practical significance for improving the operational safety and efficiency of drainage systems. Summary of the Invention

[0007] This application provides a method and apparatus for evaluating deep reinforcement learning control strategies for urban drainage systems, in order to solve the problems in related technologies, such as the difficulty in effectively evaluating the performance of urban drainage systems due to the current evaluation methods having a single evaluation dimension, inconsistent evaluation processes, and the inability to provide suitable comparison criteria.

[0008] The first aspect of this application provides a method for evaluating a deep reinforcement learning control strategy for an urban drainage system, comprising the following steps: constructing a deep reinforcement learning control strategy for the urban drainage system based on at least one control objective; analyzing the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability of the deep reinforcement learning control strategy based on at least one target operating scenario using a pre-constructed multi-dimensional evaluation system to generate analysis results for the reinforcement learning control strategy; and analyzing the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability of a pre-constructed parallel comparison control strategy using the evaluation system to generate analysis results for the comparison control strategy; and evaluating the control performance of the deep reinforcement learning control strategy based on the analysis results.

[0009] Through the aforementioned technical means, the embodiments of this application can first construct a deep reinforcement learning control strategy based on at least one control objective of the urban drainage system, and then, relying on a multi-dimensional evaluation system, analyze the performance of the deep reinforcement learning control strategy and a pre-constructed parallel comparison control strategy in four dimensions: the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability, respectively, for at least one objective operation scenario, and generate corresponding analysis results. Finally, the control performance of the deep reinforcement learning control strategy is evaluated by comparing the analysis results of the two types of strategies. This allows for a comprehensive, objective, and accurate verification of the advantages and disadvantages of the deep reinforcement learning control strategy compared to the comparison strategy, providing a reliable basis for the optimization and selection of urban drainage system control strategies.

[0010] Optionally, in one embodiment of this application, the construction of a deep reinforcement learning control strategy for an urban drainage system includes: setting control objectives for the drainage system based on environmental, economic, and operational dimensions, respectively; and guiding an agent to learn and train based on the control objectives through a predefined objective function until a preset convergence condition is met, so as to generate the control strategy.

[0011] Through the aforementioned technical means, the embodiments of this application can set up a deep reinforcement learning control strategy to construct optimization goals in three dimensions: environment, economy, and operation, and realize intelligent dynamic regulation with multi-objective collaboration. This can significantly reduce the risk of urban flooding and overflow pollution during rainfall, and prioritize environmental safety. Under this premise, optimized scheduling can effectively reduce pump station energy consumption and operating costs, and improve economic efficiency. At the same time, it smooths the frequency of equipment operation, reduces mechanical wear, thereby extending the life of facilities and enhancing the overall stability of system operation. Ultimately, this strategy can achieve long-term economical, reliable, and sustainable intelligent operation of the drainage system while ensuring environmental performance.

[0012] Optionally, in one embodiment of this application, the calculation formula for the control target is as follows:

[0013] in, P Env , P Eco and P Oper These represent environmental objectives, economic objectives, and operational objectives, respectively. and These represent the flood overflow, pollution load, and node water depth of the drainage system, respectively. C ele and C che These represent the energy consumption and drug consumption of the system facilities, respectively. Indicates control facilities i At any moment t The state of action; This represents the weighted overall control objective of the system. and These represent the weights of the three control objectives.

[0014] Through the above-mentioned technical means, the embodiments of this application can achieve collaborative optimization by integrating environmental, economic and operational multi-dimensional objectives, and guide the intelligent agent to learn autonomously with the help of objective functions, and finally generate dynamic control strategies, thereby realizing intelligent adaptive regulation of system operation. While improving flood resistance and reducing pollution emissions, it optimizes energy consumption and costs, and significantly enhances the overall resilience, economy and sustainability of the drainage system.

[0015] Optionally, in one embodiment of this application, the formula for calculating the evaluation index of decision execution efficiency is:

[0016] in, Indicates the decision-making time. This represents the total simulation time for the model with policy optimization. This represents the simulation computation time for the model without policy optimization. This indicates the total number of time steps in the control time domain.

[0017] Through the aforementioned technical means, the embodiments of this application can evaluate the efficiency of drainage system control strategies, with its core value lying in verifying the effectiveness of control commands from generation to implementation. Efficient execution ensures the system responds quickly during critical rainfall windows, thereby fully converting the theoretical potential of algorithm optimization (such as pre-lowering water levels and staggered discharge) into actual environmental benefits (reducing flooding and overflow pollution). Simultaneously, efficient and stable execution is itself an important operational objective, significantly reducing losses and failure risks caused by frequent equipment start-ups and shutdowns, ensuring long-term stable system operation. Therefore, this indicator is a key benchmark for measuring whether a control strategy can move from "simulation effectiveness" to "engineering practicality."

[0018] Optionally, in one embodiment of this application, the robustness evaluation index is calculated using the following formula:

[0019] in, This indicates the robustness of each control indicator; This represents the performance value of each control indicator under different events; Indicates the baseline value; i To represent different events; n This indicates the total number of events.

[0020] Through the above-mentioned technical means, the embodiments of this application can set robustness indicators by introducing random disturbances into rainfall intensity, constructing simulated scenarios with prediction bias, designing test groups based on actual rainfall events, and using root mean square error to calculate the deviation of the performance values ​​of each control indicator under different events relative to the benchmark value. This allows for accurate evaluation of the stability and expected control performance maintenance capability of urban drainage system control strategies under external disturbances such as rainfall forecast errors. Furthermore, this evaluation method can be extended to various uncertainty scenarios of the system, possessing strong universality and applicability.

[0021] Optionally, in one embodiment of this application, the formula for calculating the adaptability evaluation index is:

[0022] in, This represents the average performance value of the evaluation model under various scenarios; This represents the standard deviation of the evaluation model under various scenarios. j These represent different performance metrics; These represent environmental performance, economic performance, and operational performance, respectively.

[0023] A second aspect of this application provides a deep reinforcement learning control strategy evaluation device for an urban drainage system, comprising: a construction module for constructing a deep reinforcement learning control strategy for the urban drainage system based on at least one control objective; a first analysis module for analyzing the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability of the deep reinforcement learning control strategy based on at least one target operating scenario using a pre-constructed multi-dimensional evaluation system, to generate analysis results of the reinforcement learning control strategy; a second analysis module for analyzing the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability of a pre-constructed parallel comparison control strategy using the evaluation system, to generate analysis results of the comparison control strategy; and an evaluation module for evaluating the control performance of the deep reinforcement learning control strategy based on the analysis results.

[0024] Through the aforementioned technical means, the embodiments of this application can first construct a deep reinforcement learning control strategy based on at least one control objective of the urban drainage system, and then, relying on a multi-dimensional evaluation system, analyze the performance of the deep reinforcement learning control strategy and a pre-constructed parallel comparison control strategy in four dimensions: the degree of control objective achievement, decision execution efficiency, strategy robustness, and system adaptability, respectively, for at least one objective operation scenario, and generate corresponding analysis results. Finally, the control performance of the deep reinforcement learning control strategy is evaluated by comparing the analysis results of the two types of strategies. This allows for a comprehensive, objective, and accurate verification of the advantages and disadvantages of the deep reinforcement learning control strategy compared to the comparison strategy, providing a reliable basis for the optimization and selection of urban drainage system control strategies.

[0025] Optionally, in one embodiment of this application, the construction module includes: a setting unit, configured to set control objectives for the drainage system based on environmental, economic, and operational dimensions respectively; and a generation unit, configured to guide the agent to learn and train based on the control objectives using a predefined objective function until a preset convergence condition is met, so as to generate the control strategy.

[0026] Through the aforementioned technical means, the embodiments of this application can set up a deep reinforcement learning control strategy to construct optimization goals in three dimensions: environment, economy, and operation, and realize intelligent dynamic regulation with multi-objective collaboration. This can significantly reduce the risk of urban flooding and overflow pollution during rainfall, and prioritize environmental safety. Under this premise, optimized scheduling can effectively reduce pump station energy consumption and operating costs, and improve economic efficiency. At the same time, it smooths the frequency of equipment operation, reduces mechanical wear, thereby extending the life of facilities and enhancing the overall stability of system operation. Ultimately, this strategy can achieve long-term economical, reliable, and sustainable intelligent operation of the drainage system while ensuring environmental performance.

[0027] Optionally, in one embodiment of this application, the calculation formula for the control target is as follows:

[0028] in, P Env , P Eco and P Oper These represent environmental objectives, economic objectives, and operational objectives, respectively. and These represent the flood overflow, pollution load, and node water depth of the drainage system, respectively. C ele and C che These represent the energy consumption and drug consumption of the system facilities, respectively. Indicates control facilities i At any moment t The state of action; This represents the weighted overall control objective of the system. and These represent the weights of the three control objectives.

[0029] Through the above-mentioned technical means, the embodiments of this application can achieve collaborative optimization by integrating environmental, economic and operational multi-dimensional objectives, and guide the intelligent agent to learn autonomously with the help of objective functions, and finally generate dynamic control strategies, thereby realizing intelligent adaptive regulation of system operation. While improving flood resistance and reducing pollution emissions, it optimizes energy consumption and costs, and significantly enhances the overall resilience, economy and sustainability of the drainage system.

[0030] Optionally, in one embodiment of this application, the formula for calculating the evaluation index of decision execution efficiency is:

[0031] in, Indicates the decision-making time. This represents the total simulation time for the model with policy optimization. This represents the simulation computation time for the model without policy optimization. This indicates the total number of time steps in the control time domain.

[0032] Through the aforementioned technical means, the embodiments of this application can evaluate the efficiency of drainage system control strategies, with its core value lying in verifying the effectiveness of control commands from generation to implementation. Efficient execution ensures the system responds quickly during critical rainfall windows, thereby fully converting the theoretical potential of algorithm optimization (such as pre-lowering water levels and staggered discharge) into actual environmental benefits (reducing flooding and overflow pollution). Simultaneously, efficient and stable execution is itself an important operational objective, significantly reducing losses and failure risks caused by frequent equipment start-ups and shutdowns, ensuring long-term stable system operation. Therefore, this indicator is a key benchmark for measuring whether a control strategy can move from "simulation effectiveness" to "engineering practicality."

[0033] Optionally, in one embodiment of this application, the robustness evaluation index is calculated using the following formula:

[0034] in, This indicates the robustness of each control indicator; This represents the performance value of each control indicator under different events; Indicates the baseline value; i To represent different events; n This indicates the total number of events.

[0035] Through the above-mentioned technical means, the embodiments of this application can set robustness indicators by introducing random disturbances into rainfall intensity, constructing simulated scenarios with prediction bias, designing test groups based on actual rainfall events, and using root mean square error to calculate the deviation of the performance values ​​of each control indicator under different events relative to the benchmark value. This allows for accurate evaluation of the stability and expected control performance maintenance capability of urban drainage system control strategies under external disturbances such as rainfall forecast errors. Furthermore, this evaluation method can be extended to various uncertainty scenarios of the system, possessing strong universality and applicability.

[0036] Optionally, in one embodiment of this application, the formula for calculating the adaptability evaluation index is:

[0037] in, This represents the average performance value of the evaluation model under various scenarios; This represents the standard deviation of the evaluation model under various scenarios. j These represent different performance metrics; These represent environmental performance, economic performance, and operational performance, respectively.

[0038] Through the above-mentioned technical means, the embodiments of this application can set adaptive indicators by generating a variety of rainfall event combinations with different rainfall durations, spatial distribution patterns and intensity ratios, focusing on "unseen conditions" outside of the training data, eliminating the interference of observation errors or forecast biases, and calculating based on the average performance value and standard deviation of the evaluation model under various scenarios. This can accurately measure the generalization ability, transfer ability and adaptive ability of urban drainage system control strategies (especially deep reinforcement learning controllers) in new environments, and effectively evaluate the stable performance and generalizability of urban drainage system control strategies under cross-condition conditions.

[0039] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the deep reinforcement learning control strategy evaluation method for urban drainage systems as described in the above embodiments.

[0040] A fourth aspect of this application provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described deep reinforcement learning control strategy evaluation method for an urban drainage system.

[0041] A fifth aspect of this application provides a computer program product that stores a computer program that, when executed by a processor, implements the above-described deep reinforcement learning control strategy evaluation method for urban drainage systems.

[0042] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0043] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a deep reinforcement learning control strategy evaluation method for an urban drainage system according to an embodiment of this application. Figure 2 This is a schematic diagram of a comprehensive simulation model structure according to a specific embodiment of this application; Figure 3 This is a schematic diagram of the framework of a deep reinforcement learning control method for an urban drainage system according to a specific embodiment of this application; Figure 4 This is a schematic diagram of an evaluation index system framework according to a specific embodiment of this application; Figure 5This is a flowchart of a deep reinforcement learning control strategy evaluation method for an urban drainage system according to a specific embodiment of this application; Figure 6 This is a schematic diagram of the structure of a deep reinforcement learning control strategy evaluation device for an urban drainage system according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0044] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0045] Patent CN119047742A discloses a method and system for integrated optimization scheduling of urban drainage systems, networks, and rivers. This technology proposes an integrated optimization scheduling method for urban drainage systems, networks, and rivers. By constructing a hydraulic and water quality model and combining pollution source strength estimation with topological data, it achieves multi-objective scheduling to reduce overflow risk and improve the intelligent early warning level of urban drainage. However, this technology has shortcomings in strategy performance evaluation, mainly in the lack of testing the generalization ability of different scheduling schemes under complex actual operating conditions.

[0046] Patent CN107292527A discloses a method for evaluating the performance of urban drainage systems. This method quantitatively characterizes the resilience and sustainability of urban drainage systems based on simulation results from urban drainage system models under extreme rainfall scenarios by designing resilience and sustainability indicators. However, this method focuses on assessing resilience and sustainability and is not suitable for evaluating the effectiveness of urban drainage system scheduling schemes and control strategies, making it difficult to achieve multi-dimensional quantitative evaluation of different scheduling and control methods.

[0047] Therefore, in view of the above-mentioned technical defects, this application proposes a method and apparatus for evaluating deep reinforcement learning control strategies for urban drainage systems.

[0048] The following description, with reference to the accompanying drawings, describes a method and apparatus for evaluating deep reinforcement learning control strategies for urban drainage systems according to embodiments of this application. Addressing the problems mentioned in the background section regarding the difficulty in effectively evaluating the performance of urban drainage systems due to the current evaluation methods' single evaluation dimensions, inconsistent evaluation processes, and lack of suitable comparative criteria, this application provides a method for evaluating deep reinforcement learning control strategies for urban drainage systems. This method establishes a unified framework for evaluating drainage system control strategies, incorporating empirical rule-based, optimization-based, and deep reinforcement learning methods into the same evaluation system, enabling reproducible and quantifiable comparisons of multiple algorithms under consistent conditions. It proposes a comprehensive evaluation system including multiple indicators such as environmental performance, energy consumption cost, and control efficiency, extending to robustness and adaptability dimensions, thereby systematically revealing the stability and generalization ability of control strategies in multiple scenarios. Furthermore, by utilizing a high-fidelity urban drainage system simulation platform, the decision-making process of different control strategies can be repeatedly executed, combined with diverse rainfall events and system boundary condition configurations, realistically reproducing actual operational differences, providing scientific support for algorithm optimization and engineering applications. This solves the problems in related technologies, such as the difficulty in effectively evaluating the performance of urban drainage systems due to the current evaluation methods having only one evaluation dimension, inconsistent evaluation processes, and the inability to provide suitable comparison criteria.

[0049] Specifically, Figure 1 This is a flowchart illustrating a deep reinforcement learning control strategy evaluation method for an urban drainage system provided in an embodiment of this application.

[0050] like Figure 1 As shown, the deep reinforcement learning control strategy evaluation method for this urban drainage system includes the following steps: In step S101, a deep reinforcement learning control strategy for the urban drainage system is constructed based on at least one control objective of the urban drainage system.

[0051] In practical implementation, this application embodiment can construct an intelligent control framework for a drainage system based on deep reinforcement learning, enabling the system to achieve adaptive scheduling under various operating scenarios and providing a reproducible decision-making process for subsequent performance analysis. This framework can be configured to consist of three main parts: control objective setting, policy learning process during the training phase, and policy execution process during the decision-making phase.

[0052] Optionally, in one embodiment of this application, constructing a deep reinforcement learning control strategy for an urban drainage system includes: setting control objectives for the drainage system based on environmental, economic, and operational dimensions, respectively; and guiding an agent to learn and train based on the control objectives through a predefined objective function until a preset convergence condition is met, so as to generate a control strategy.

[0053] In this application, the control objectives of the urban drainage system can be clearly defined first.

[0054] Optionally, in one embodiment of this application, the formula for calculating the control target is as follows:

[0055] Specifically, the control objectives are set from three dimensions: environmental, economic, and operational. Environmental performance is the primary concern, requiring the minimization of urban flooding risk and system overflow (including overflow volume and external pollution load) during rainfall events. Secondly, economic performance measures system operating costs, aiming to minimize energy and chemical consumption at pumping stations and related facilities. Finally, operational performance assesses the impact of control strategies on facility stability, requiring reduced equipment start-up and shutdown frequency to extend their service life. These multi-dimensional indicators can be mathematically described using formulas (1) to (3): (1) (2) (3)

[0056] in, P Env , P Eco and P Oper These are environmental objectives, economic objectives, and operational objectives; and These are the flood overflow, pollution load, and node water depth of the drainage system, which constitute the independent variables for calculating the environmental objective function value. C ele and C che This indicates the energy and pharmaceutical consumption of the system facilities; This indicates the operational state of control facility i at time t. The frequency of operation can be characterized by the difference in operational state before and after the facility. The weighted overall control objective of the system. and These are the weights for the three objectives, which can be designed based on specific urban drainage system management objectives and preferences.

[0057] This deep reinforcement learning control strategy constructs optimization objectives across three dimensions: environment, economy, and operation. It achieves intelligent dynamic regulation with multi-objective collaboration, thereby significantly reducing the risk of urban flooding and overflow pollution during rainfall and prioritizing environmental safety. Under this premise, it effectively reduces pump station energy consumption and operating costs through optimized scheduling, improving economic efficiency. At the same time, it smooths out equipment operation frequency, reduces mechanical wear, thereby extending facility life and enhancing the overall stability of system operation. Ultimately, this strategy can achieve long-term, economical, reliable, and sustainable intelligent operation of the drainage system while ensuring environmental performance.

[0058] During the training phase, a comprehensive simulation model of the urban drainage system was first established based on EPASWMM-ASM-RWQM to characterize the dynamic response processes of the pipe network, sewage treatment plant, and receiving water body under different rainfall scenarios and the coupling mechanism of water quantity and quality. The model takes rainfall data, initial water level, inflow boundary conditions, pollutant concentration, and facility operating status as inputs, and outputs the system's time-series water level changes, flow distribution, and water quality parameter evolution. Its model structure is as follows: Figure 2 As shown.

[0059] At each time step, the agent acquires real-time system status information, including water levels, flow rates, water quality indicators, and equipment start-up / shutdown status at each node. Based on the current status, it selects appropriate control actions, such as adjusting pump station operation or changing gate opening. The selected actions are input into the simulation model to drive the hydrodynamic-water quality response at the next time step, resulting in an updated system state and corresponding performance indicators. Subsequently, an immediate reward is calculated using a predefined objective function, where the reward function is set to the negative of the objective function. ( )= ( This guides the agent to learn the optimal control strategy that can simultaneously reduce environmental risks, reduce operating energy consumption, and maintain stable water quality in the system.

[0060] During training, interaction data such as states, actions, rewards, and subsequent states are stored in an experience replay pool. The agent continuously optimizes its policy parameters through random sampling, thereby gradually improving the estimation accuracy of state-action values. This process is iteratively carried out under various typical rainfall conditions until the control strategy achieves convergence and robustness in both water quantity regulation and water quality maintenance dimensions.

[0061] In the decision-making phase, the trained strategy is fixed and embedded into the urban drainage system model. For each typical operating condition, the system executes the strategy based on the real-time status, outputting control commands to achieve coordinated scheduling of pumping stations and gates, driving the model to simulate the dynamic response of the system. The entire process records the water quantity and quality changes and corresponding control behaviors at each time step, and comprehensively evaluates the strategy performance from three dimensions: environment, economy, and operation. No further parameter updates are performed in this phase; the output results serve as the basis for subsequent performance evaluation and comparative analysis to ensure the repeatability and objectivity of the evaluation results. The water quantity-quality coordinated control framework for urban drainage systems based on deep reinforcement learning is as follows: Figure 3 As shown.

[0062] Through the above-mentioned technical means, the embodiments of this application can achieve collaborative optimization by integrating environmental, economic and operational multi-dimensional objectives, and guide the intelligent agent to learn autonomously with the help of objective functions, and finally generate dynamic control strategies, thereby realizing intelligent adaptive regulation of system operation. While improving flood resistance and reducing pollution emissions, it optimizes energy consumption and costs, and significantly enhances the overall resilience, economy and sustainability of the drainage system.

[0063] In step S102, based on at least one target operating scenario, a pre-constructed multi-dimensional evaluation system is used to analyze the degree of control objective achievement, decision execution efficiency, policy robustness, and system adaptability of the deep reinforcement learning control strategy, so as to generate analysis results of the reinforcement learning control strategy; and The evaluation system is used to analyze the degree of achievement of control objectives, decision execution efficiency, strategy robustness, and system adaptability of pre-constructed parallel comparison control strategies, so as to generate analysis results of comparison control strategies.

[0064] To comprehensively evaluate the effectiveness of the proposed deep reinforcement learning control method, this step selects two representative control strategies as comparative references: one is a control strategy based on human experience rules, and the other is a control strategy based on optimization computation.

[0065] Specifically, this step aims to construct a multi-dimensional performance evaluation system for analyzing the comprehensive performance of different real-time control strategies under various operating conditions. This system unfolds from four aspects: the degree of achievement of control objectives, decision execution efficiency, strategy robustness, and system adaptability, forming an evaluation framework with a clear logical structure and engineering feasibility, such as... Figure 4 As shown.

[0066] First, at the level of control objectives, the evaluation system uses three categories of indicators—environmental, economic, and operational—to quantitatively measure the overall performance of the system. Environmental indicators reflect the level of drainage and overflow risks, focusing on the effectiveness of urban flood control and pollution reduction during rainfall events; economic indicators are mainly used to characterize energy consumption and operating costs; and operational indicators assess the start-up and shutdown frequency, operational stability, and maintenance requirements of facilities. All of the above indicators can be obtained from high-precision numerical simulations or field monitoring data, and are calculated based on a high-fidelity model of the urban drainage system that includes the Saint-Venant equation, to ensure that the hydraulic and water quality changes of the system under dynamic operating conditions can be accurately captured. The calculation method of the control objective indicators is consistent with the form defined in formulas (1) to (3).

[0067] Secondly, decision efficiency indicators are used to quantify the response speed and execution efficiency of control strategies. In this application, embodiments can compare the system simulation time before and after strategy execution, and perform standardization based on time steps, reflecting the computational performance and responsiveness of the strategy in actual operation, thereby verifying its feasibility in real-time control scenarios. These indicators are typically expressed as the average strategy computation time per unit time step, and their mathematical expression is as follows:

[0068] in, Indicates the decision-making time. This represents the total simulation time for the model with policy optimization. This represents the simulation computation time for the model without policy optimization. This indicates the total number of time steps in the control time domain.

[0069] Through the aforementioned technical means, the embodiments of this application can evaluate the efficiency of drainage system control strategies, with its core value lying in verifying the effectiveness of control commands from generation to implementation. Efficient execution ensures the system responds quickly during critical rainfall windows, thereby fully converting the theoretical potential of algorithm optimization (such as pre-lowering water levels and staggered discharge) into actual environmental benefits (reducing flooding and overflow pollution). Simultaneously, efficient and stable execution is itself an important operational objective, significantly reducing losses and failure risks caused by frequent equipment start-ups and shutdowns, ensuring long-term stable system operation. Therefore, this indicator is a key benchmark for measuring whether a control strategy can move from "simulation effectiveness" to "engineering practicality."

[0070] Furthermore, robustness indices are used to evaluate the stability of control strategies under external disturbances, particularly the impact of uncertainties such as rainfall forecast errors. By introducing random disturbances into rainfall intensity and constructing a simulated scenario with prediction bias, it is possible to test whether the control strategy can maintain the expected control performance and operational stability in the face of external uncertainties. This evaluation method can be extended to various types of uncertainty scenarios in urban drainage systems, demonstrating strong universality and applicability.

[0071]

[0072] in, i To represent different events; n Represents the total number of events, and the robustness of each control metric is... By calculating its performance values ​​under different events Relative to the baseline value The root mean square error (RMSE) is used to determine this. In calculating this index, the evaluation model is applied to a set of design events constructed based on actual rainfall events to assess the stability of the control strategy under uncertain conditions.

[0073] Through the above-mentioned technical means, the embodiments of this application can set robustness indicators by introducing random disturbances into rainfall intensity, constructing simulated scenarios with prediction bias, designing test groups based on actual rainfall events, and using root mean square error to calculate the deviation of the performance values ​​of each control indicator under different events relative to the benchmark value. This allows for accurate evaluation of the stability and expected control performance maintenance capability of urban drainage system control strategies under external disturbances such as rainfall forecast errors. Furthermore, this evaluation method can be extended to various uncertainty scenarios of the system, possessing strong universality and applicability.

[0074] Furthermore, the adaptability index focuses on measuring the generalization and transfer capabilities of the control strategy under "unseen conditions," emphasizing whether the deep reinforcement learning controller can maintain excellent performance in situations outside the training data. By generating various combinations of rainfall events with different variations, such as adjusting rainfall duration, spatial distribution patterns, and intensity ratios, the system response and regulation effectiveness of the control strategy in new environments are evaluated. This evaluation does not consider the impact of observation errors or forecast biases, but focuses on the generalizability and adaptability of the strategy under different operating conditions. The adaptability index is calculated based on the average performance values ​​obtained by the evaluation model under various scenarios (…). ) and standard deviation ( This is used to characterize the stable performance of the control strategy under various operating conditions. The calculation method is shown in formulas (6) to (7):

[0075]

[0076] in, j Representing different performance metrics, such as These represent environmental performance, economic performance, and operational performance, respectively.

[0077] Through the above-mentioned technical means, the embodiments of this application can set adaptive indicators by generating a variety of rainfall event combinations with different rainfall durations, spatial distribution patterns and intensity ratios, focusing on "unseen conditions" outside of the training data, eliminating the interference of observation errors or forecast biases, and calculating based on the average performance value and standard deviation of the evaluation model under various scenarios. This can accurately measure the generalization ability, transfer ability and adaptive ability of urban drainage system control strategies (especially deep reinforcement learning controllers) in new environments, and effectively evaluate the stable performance and generalizability of urban drainage system control strategies under cross-condition conditions.

[0078] In step S103, the control performance of the deep reinforcement learning control strategy is evaluated based on the analysis results.

[0079] In actual implementation, the embodiments of this application can analyze the control strategies of urban drainage systems, control strategies based on human experience rules, and control strategies based on optimization calculations through the above-mentioned multi-dimensional evaluation index model, generate corresponding analysis results, and then evaluate the performance of the control strategies of urban drainage systems based on the analysis results.

[0080] Furthermore, to systematically evaluate the decision-making efficiency and target achievement effects of different control strategies under various rainfall conditions, this application proposes a hybrid rainfall event construction method. This method can comprehensively utilize historical observed rainfall and virtual designed rainfall to generate an evaluation sample set that balances typicality and extreme scenarios. Historical rainfall events have realistic time series and intensity distribution characteristics, reflecting the response and scheduling effects of control strategies in actual operation. Virtual rainfall events, by adjusting the morphological characteristics and intensity distribution of rainfall events, construct diverse test conditions to verify the applicability and stability of control strategies under unconventional situations. Based on the constructed event set, the evaluation model is used to simulate and calculate the control strategies under various rainfall conditions, quantifying their comprehensive performance and response efficiency in environmental, economic, and operational dimensions, thereby forming a multi-strategy-multi-scenario performance matrix, providing a quantitative basis for the selection and optimization of control schemes.

[0081] In terms of robustness analysis, the embodiments of this application can construct evaluation scenarios by applying perturbations to actual rainfall data. Specifically, typical historical rainfall events can be selected as basic samples. While maintaining their temporal distribution structure, random perturbations are introduced into the rainfall intensity for each period to simulate the uncertainty caused by rainfall forecast errors. The perturbation amplitude can be set to a limited range of fluctuations based on statistical analysis results. By generating multiple perturbation sample event sets, the performance stability and reliability of the control strategy in the presence of external perturbations can be evaluated. This method is simple to implement and can realistically reproduce the uncertain conditions commonly encountered in the actual operation of urban drainage systems, providing a basis for the quantitative evaluation of strategy robustness.

[0082] In terms of adaptive analysis, this application's embodiments can design various rainfall patterns and intensity variation schemes for unseen operating conditions based on a set of real rainfall events to construct a diverse test event set. Rainfall patterns include typical forms such as uniform distribution, unimodal distribution (peak value located at the beginning, middle, or end of the rainfall), and bimodal distribution; changes in rainfall intensity are achieved by scaling the original rainfall sequence to form event sets of different intensity levels. These scenarios do not involve prediction errors or observation noise and can be used to individually evaluate the adaptability and transfer performance of control strategies under novel rainfall conditions, making them particularly suitable for verifying the generalization characteristics of intelligent control algorithms such as deep reinforcement learning. This method has strong scalability and can be further customized or expanded according to application requirements to achieve a more comprehensive strategy performance evaluation.

[0083] In summary, this application proposes a performance evaluation method for urban drainage systems based on deep reinforcement learning, such as... Figure 5 As shown, the method may include the following four steps: Step 1, constructing a control strategy for an urban drainage system based on deep reinforcement learning as the core control method; Step 2, constructing parallel control strategies for comparative analysis, such as rule-based control or optimization control, to provide a reference baseline for performance evaluation; Step 3, designing various operating scenarios, including different rainfall intensities, drainage loads, and system boundary conditions, to simulate real complex environments and test the adaptability and robustness of the control strategy; Step 4, establishing a strategy performance index system, and performing performance calculations and comparative analyses on different control strategies under various operating scenarios, thereby achieving multi-dimensional performance evaluation of the urban drainage system control strategy.

[0084] Its corresponding technical effects can be summarized as follows: 1. This method can perform quantitative analysis and comparison of any type of control strategy, improving the objectivity and versatility of the evaluation.

[0085] 2. This method can comprehensively evaluate the overall capability of control strategies by constructing a multi-dimensional composite index system that fully covers aspects such as operational performance, robustness, adaptability, and decision-making efficiency.

[0086] 3. This method can improve the reliability of experimental results by designing a unified control and evaluation process to ensure that different strategies can be tested repeatedly and comparablely under the same operating conditions.

[0087] 4. This method can significantly improve the representativeness and universality of the evaluation results by combining historical rainfall and tectonic rainfall events to form a test set covering typical and extreme scenarios.

[0088] 5. This method can introduce a control method based on empirical rules and optimization calculation as a reference baseline to construct a comparison system between deep reinforcement learning strategies and traditional control strategies, providing a quantitative basis for performance improvement.

[0089] The deep reinforcement learning control strategy evaluation method for urban drainage systems proposed in this application can achieve reproducible and quantifiable comparisons of multiple algorithms under consistent conditions by establishing a unified evaluation framework for drainage system control strategies, incorporating empirical rule-based, optimization-based, and deep reinforcement learning-based methods into the same evaluation system. Furthermore, it can systematically reveal the stability and generalization ability of control strategies in multiple scenarios by proposing a comprehensive evaluation system that includes multiple indicators such as environmental performance, energy consumption cost, and control efficiency, and extending it to robustness and adaptability dimensions. Moreover, it can utilize a high-fidelity urban drainage system simulation platform to repeatedly execute the decision-making process of different control strategies, combining diverse rainfall events and system boundary condition configurations to realistically reproduce actual operational differences, providing scientific support for algorithm optimization and engineering applications. Therefore, it solves the problems in related technologies where current evaluation methods suffer from limited evaluation dimensions, inconsistent evaluation processes, and the inability to provide suitable comparison criteria, leading to difficulties in effectively evaluating the performance of urban drainage systems.

[0090] Next, refer to the appendix. Figure 6 This application describes a deep reinforcement learning control strategy evaluation apparatus for urban drainage systems, based on embodiments thereof.

[0091] Figure 6 This is a block diagram of a deep reinforcement learning control strategy evaluation device for an urban drainage system according to an embodiment of this application.

[0092] like Figure 6 As shown, the deep reinforcement learning control strategy evaluation device 10 for the urban drainage system includes: a construction module 100, a first analysis module 200, a second analysis module 300, and an evaluation module 400.

[0093] The construction module 100 is used to construct a deep reinforcement learning control strategy for the urban drainage system based on at least one control objective of the urban drainage system.

[0094] The first analysis module 200 is used to analyze the degree of control objective achievement, decision execution efficiency, policy robustness, and system adaptability of a deep reinforcement learning control strategy based on at least one target operating scenario using a pre-constructed multi-dimensional evaluation system, in order to generate analysis results for the reinforcement learning control strategy. The second analysis module 300 is used to analyze the degree of control objective achievement, decision execution efficiency, policy robustness, and system adaptability of a pre-constructed parallel comparison control strategy using the evaluation system, in order to generate analysis results for the comparison control strategy.

[0095] Evaluation module 400 is used to evaluate the control performance of deep reinforcement learning control strategies based on the analysis results.

[0096] Optionally, in one embodiment of this application, the construction module 100 includes: a setting unit and a generation unit; wherein, the setting unit is used to set control objectives for the drainage system based on environmental, economic and operational dimensions respectively; the generation unit is used to guide the agent to learn and train based on the control objectives through a predefined objective function until a preset convergence condition is met, so as to generate a control strategy.

[0097] Optionally, in one embodiment of this application, the formula for calculating the control target is as follows:

[0098] in, P Env , P Eco and P Oper These represent environmental objectives, economic objectives, and operational objectives, respectively. and These represent the flood overflow, pollution load, and node water depth of the drainage system, respectively. C ele and C che These represent the energy consumption and drug consumption of the system facilities, respectively. Indicates control facilities i At any moment t The state of action; This represents the weighted overall control objective of the system. and These represent the weights of the three control objectives.

[0099] Optionally, in one embodiment of this application, the formula for calculating the evaluation index of decision execution efficiency is as follows:

[0100] in, Indicates the decision-making time. This represents the total simulation time for the model with policy optimization. This represents the simulation computation time for the model without policy optimization. This indicates the total number of time steps in the control time domain.

[0101] Optionally, in one embodiment of this application, the robustness evaluation index is calculated using the following formula:

[0102] in, This indicates the robustness of each control indicator; This represents the performance value of each control indicator under different events; Indicates the baseline value; i To represent different events; n This indicates the total number of events.

[0103] Optionally, in one embodiment of this application, the formula for calculating the adaptability evaluation index is:

[0104] in, This represents the average performance value of the evaluation model under various scenarios; This represents the standard deviation of the evaluation model under various scenarios. j These represent different performance metrics; These represent environmental performance, economic performance, and operational performance, respectively.

[0105] It should be noted that the explanation of the aforementioned embodiment of the deep reinforcement learning control strategy evaluation method for urban drainage systems also applies to the deep reinforcement learning control strategy evaluation device for urban drainage systems in this embodiment, and will not be repeated here.

[0106] The deep reinforcement learning control strategy evaluation device for urban drainage systems proposed in this application can establish a unified evaluation framework for drainage system control strategies, incorporating empirical rule-based, optimization-based, and deep reinforcement learning methods into the same evaluation system. This enables reproducible and quantifiable comparisons of multiple algorithms under consistent conditions. Furthermore, by proposing a comprehensive evaluation system that includes multiple indicators such as environmental performance, energy consumption cost, and control efficiency, and extending it to robustness and adaptability, the stability and generalization ability of control strategies in various scenarios can be systematically revealed. Moreover, by utilizing a high-fidelity urban drainage system simulation platform, the decision-making process of different control strategies can be repeatedly executed, combined with diverse rainfall events and system boundary condition configurations, to realistically reproduce actual operational differences, providing scientific support for algorithm optimization and engineering applications. Therefore, this solves the problems in related technologies where current evaluation methods suffer from limited evaluation dimensions, inconsistent evaluation processes, and the inability to provide suitable comparison criteria, leading to difficulties in effectively evaluating the performance of urban drainage systems.

[0107] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 701, the processor 702, and the computer program stored on the memory 701 and executable on the processor 702.

[0108] When the processor 702 executes the program, it implements the deep reinforcement learning control strategy evaluation method for urban drainage systems provided in the above embodiments.

[0109] Furthermore, electronic devices also include: Communication interface 703 is used for communication between memory 701 and processor 702.

[0110] The memory 701 is used to store computer programs that can run on the processor 702.

[0111] The memory 701 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0112] If the memory 701, processor 702, and communication interface 703 are implemented independently, then the communication interface 703, memory 701, and processor 702 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0113] Optionally, in a specific implementation, if the memory 701, processor 702, and communication interface 703 are integrated on a single chip, then the memory 701, processor 702, and communication interface 703 can communicate with each other through an internal interface.

[0114] The processor 702 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0115] This application also provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described deep reinforcement learning control strategy evaluation method for urban drainage systems.

[0116] This application also provides a computer program product storing a computer program that, when executed by a processor, implements the above-described deep reinforcement learning control strategy evaluation method for urban drainage systems.

[0117] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0118] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0119] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0120] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0121] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0122] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0123] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0124] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for evaluating deep reinforcement learning control strategies in urban drainage systems, characterized in that, Includes the following steps: Construct a deep reinforcement learning control strategy for the urban drainage system based on at least one control objective of the urban drainage system; Based on at least one target operating scenario, the degree of achievement of the control objective, decision execution efficiency, policy robustness, and system adaptability of the deep reinforcement learning control strategy are analyzed using a pre-constructed multi-dimensional evaluation system to generate the analysis results of the reinforcement learning control strategy. and The evaluation system is used to analyze the degree of achievement of control objectives, decision execution efficiency, strategy robustness, and system adaptability of the pre-constructed parallel comparison control strategy, so as to generate the analysis results of the comparison control strategy. The control performance of the deep reinforcement learning control strategy is evaluated based on the analysis results.

2. The method according to claim 1, characterized in that, The deep reinforcement learning control strategy for constructing the urban drainage system includes: The control objectives of the drainage system are set based on environmental, economic, and operational dimensions, respectively. Based on the control objective, the agent is guided to learn and train through a predefined objective function until a preset convergence condition is met, so as to generate the control strategy.

3. The method according to claim 2, characterized in that, The formula for calculating the control target is as follows: in, P Env , P Eco and P Oper These represent environmental objectives, economic objectives, and operational objectives, respectively. and These represent the flood overflow, pollution load, and node water depth of the drainage system, respectively. C ele and C che These represent the energy consumption and drug consumption of the system facilities, respectively. Indicates control facilities i At any moment t The state of action; This represents the weighted overall control objective of the system. and These represent the weights of the three control objectives.

4. The method according to claim 1, characterized in that, The formula for calculating the evaluation index of decision execution efficiency is as follows: in, Indicates the decision-making time. This represents the total simulation time for the model with policy optimization. This represents the simulation computation time for the model without policy optimization. This indicates the total number of time steps in the control time domain.

5. The method according to claim 1, characterized in that, The formula for calculating the robustness evaluation index is as follows: in, This indicates the robustness of each control indicator; This represents the performance value of each control indicator under different events; Indicates the baseline value; i To represent different events; n This indicates the total number of events.

6. The method according to claim 1, characterized in that, The formula for calculating the adaptability assessment index is as follows: in, This represents the average performance value of the evaluation model under various scenarios; This represents the standard deviation of the evaluation model under various scenarios. j These represent different performance metrics; These represent environmental performance, economic performance, and operational performance, respectively.

7. A deep reinforcement learning control strategy evaluation device for an urban drainage system, characterized in that, include: Module building; Used to construct a deep reinforcement learning control strategy for the urban drainage system based on at least one control objective of the urban drainage system; The first analysis module is used to analyze the degree of achievement of the control objective, decision execution efficiency, policy robustness and system adaptability of the deep reinforcement learning control strategy based on at least one target operation scenario using a pre-constructed multi-dimensional evaluation system, so as to generate the analysis results of the reinforcement learning control strategy. and The second analysis module is used to analyze the degree of achievement of control objectives, decision execution efficiency, strategy robustness and system adaptability of the pre-constructed parallel comparison control strategy using the evaluation system, so as to generate the analysis results of the comparison control strategy. An evaluation module is used to evaluate the control performance of the deep reinforcement learning control strategy based on the analysis results.

8. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the deep reinforcement learning control strategy evaluation method for urban drainage systems as described in any one of claims 1-6.

9. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the deep reinforcement learning control strategy evaluation method for urban drainage systems as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, The computer program is executed to implement the deep reinforcement learning control strategy evaluation method for urban drainage systems as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Urban sewerage and drainage system performance evaluation method

    CN107292527A

  • Urban plant-network-river-based integrated optimal scheduling method and system

    CN119047742A