A method, system and device for pumped storage site selection based on reinforcement learning

CN121303538BActive Publication Date: 2026-09-15POWERCHINA HUADONG ENG CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511379404.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-09-15
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明提供了一种基于强化学习的抽水蓄能站点选择方法、系统及设备,以解决传统方法中依赖倾向性样本导致的主观性强、场景适应性差的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303538B_ABST
    Figure CN121303538B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of pumped storage power station selection, and discloses a pumped storage power station site selection method, system and equipment based on reinforcement learning, which acquires multi-dimensional original feature data of candidate sites; the original data is preprocessed to generate a standardized multi-dimensional feature matrix and a feature importance weight vector; a reinforcement learning framework is constructed, the state space of the reinforcement learning framework contains candidate site features, selected site information and a scene constraint vector, the action space supports site selection or sorting, a reward function comprehensively considers water energy, construction and environmental benefits, and the weights are dynamically adjusted according to the feature importance; a deep reinforcement learning network is introduced, an initial decision strategy is generated through experience replay and parameter optimization; an online learning and scene self-adaptive mechanism is established to incrementally update a model and optimize parameters; finally, an optimized site list is output based on the optimized strategy, multi-dimensional evaluation is carried out, and a final result is output, so that the site selection process is scientific and intelligent, and strong support is provided for pumped storage power station site decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy storage power station selection technology, specifically to a method, system, and equipment for selecting pumped storage power stations based on reinforcement learning. Background Technology

[0002] Currently, pumped storage is a key means of regulating fluctuations in new energy power generation and ensuring the stability of the power system. Its site selection requires comprehensive consideration of complex factors across multiple dimensions, including location, hydropower resources, construction, and the environment. Traditional site selection methods often rely on expert experience or single-objective models, which suffer from strong subjectivity and insufficient balance of multiple objectives—for example, focusing solely on hydropower parameters while neglecting environmental impact, or relying on fixed weights that are difficult to adapt to different regional development scenarios, easily leading to a disconnect between site selection results and actual needs.

[0003] Meanwhile, traditional methods lack dynamic learning capabilities. When faced with fluctuations in candidate site data or changes in development constraints, model parameters need to be manually adjusted, resulting in low efficiency and difficulty in ensuring decision consistency. Furthermore, existing technologies offer only one dimension for evaluating site selection results, focusing primarily on economic or technical indicators. They have not formed a comprehensive verification system covering energy utilization, construction feasibility, environmental compatibility, and strategy stability, making it difficult to fully support scientific decision-making for pumped storage site development. Therefore, there is an urgent need for an intelligent, adaptive, and multi-dimensionally optimized site selection technology solution. Summary of the Invention

[0004] In view of this, the present invention provides a method, system and equipment for selecting pumped storage sites based on reinforcement learning, so as to solve the problems of strong subjectivity and poor scene adaptability caused by the reliance on biased samples in traditional methods.

[0005] In a first aspect, the present invention provides a method for selecting pumped storage sites based on reinforcement learning, the method comprising: Obtain multi-dimensional raw feature data of multiple candidate sites. The multi-dimensional raw feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment. The original feature data is preprocessed to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, which serve as the standardized feature dataset. Based on the standardized feature dataset, a reinforcement learning framework is constructed with state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit reward, construction feasibility reward, and environmental compatibility reward, and the weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector. By introducing a deep reinforcement learning network and through experience replay and parameter optimization, an initial decision-making strategy with the ability to evaluate the comprehensive advantages of a site is generated. Establish an online learning and scenario adaptation mechanism to absorb new decision data in real time and incrementally update the model to optimize policy network parameters; construct an online experience replay pool to store the scenario feature vector Φ at time t. t Interaction samples (S) t A t ,R t, S t+1 , Φ t ), and assign higher sampling weights to the new samples, where A t Let S be the action performed at time t. t Let S be the state vector. t+1 This is the next state vector; Update policy network parameters based on incremental samples and stochastic gradient descent algorithm; The region type, development stage, and key constraints are transformed into a standardized scenario vector Φ. Based on this scenario vector, the weight coefficients λ of the reward function are adjusted using the following formula:

[0006] Where, λ 0,k γ is the initial weight, and γ is the scene influence factor. Let be the scene feature value related to the k-th type of reward at time t; pass The quadratic weighted state vector completes the optimization of the policy network parameters, where... This is the state vector after being weighted by the feature importance weights; Based on the optimized decision-making strategy, a list of preferred sites is output and evaluated from multiple dimensions. The final output includes the list of preferred sites, an evaluation report, and a statement on the reliability of the strategy.

[0007] The reinforcement learning-based pumped storage site selection method provided in this invention provides a rich and comprehensive information foundation for pumped storage site selection by comprehensively covering four categories of multi-dimensional original feature data: location, hydropower, construction, and external environment. The preprocessing stage generates a standardized feature matrix and feature importance weight vectors, eliminating differences in data dimensions and allowing different types of features to participate fairly in subsequent decision-making. The weight vectors reflect the degree of influence of features on site selection. The constructed reinforcement learning framework integrates candidate and selected sites and scenario constraints in the state space, supports site selection and ranking in the action space, and incorporates multiple benefits in the reward mechanism, making the decision more aligned with actual needs. The deep reinforcement learning network, combined with experience replay and parameter optimization, can autonomously learn and generate initial decision strategies, possessing the ability to evaluate the comprehensive advantages of sites. Online learning and scenario adaptation mechanisms can absorb new data in real time and dynamically optimize the model, allowing the strategy to adapt to different regional development scenarios and dynamically changing conditions. Finally, multi-dimensional evaluation can verify the rationality of the selected sites and the reliability of the strategy, providing a strong and credible basis for decision-making in subsequent engineering practices. Overall, it realizes the scientific and intelligent process of site selection from data processing and strategy generation to dynamic optimization and result verification.

[0008] In one alternative implementation, the location condition includes: distance d from the load center. 负荷中心,i Distance d from the substation 变电站,i ; distance d from the new energy cluster 新能源集群,i ; The hydropower parameters include: installed capacity: P i Continuous full-load hours t 满发,i Average head H i Horizontal distance L between upper and lower warehouses i Distance to height ratio r L / H,i ; The construction conditions include: maximum dam height: upper reservoir height h 上库,i 、Lower warehouse h 下库,i Geological lithology code C 地质,i Traffic Conditions Score 交通,i ; The external environment includes: the number N of potentially significant and sensitive objects. 敏感对象,i , number of people submerged N 人口,i .

[0009] This invention refines various parameters to make site selection more precise. Location parameters reflect the convenience of power consumption, grid connection, and synergy with new energy sources; hydropower parameters are the core of evaluating power generation capacity; construction condition parameters relate to project feasibility and cost; and external environmental parameters reflect ecological and social impacts. These refined parameters provide accurate input for subsequent data processing and modeling, making site selection more targeted in specific assessments such as location suitability and hydropower utilization, thus helping to scientifically select advantageous sites.

[0010] In one optional implementation, the preprocessing of the original feature data to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, as a standardized feature dataset, includes: Use the interquartile range method to identify and correct outliers; Missing values ​​were handled using the mean of similar sites; Numerical mapping is performed on geological lithology classification variables to convert them into quantitative scores; The min-max normalization method is used to map all continuous features to the [0,1] interval; The calculation includes derived characteristics such as distance-to-height ratio, normalized installed capacity, location comprehensive index, and environmental impact score; wherein, the location comprehensive index is calculated using the following formula: S 区位,i = α1·(1 d 负荷中心,标准化,i ) + α2·(1 d 变电站,标准化,i ) + α3·(1 d 新能源集群,标准化,i ) Wherein, α1, α2, and α3 are weighting coefficients, which respectively reflect the importance of the distance from the i-th load center, substation, and new energy cluster to the location advantage, and satisfy α1 + α2 + α3 = 1; The environmental impact score is calculated using the following formula: ; in and These respectively reflect the importance of the number of potentially major sensitive objects and the number of people to be inundated to the environmental impact, and ; The feature importance weight vector is obtained by calculating the variance contribution of each standardized feature; The preprocessed features are integrated into a standardized multi-dimensional feature matrix to form the feature vector for each site.

[0011] In the preprocessing stage, this invention employs multiple methods to ensure data quality: interquartile range correction corrects outliers to avoid interference from extreme data; mean values ​​from similar sites fill in missing values, improving the data while preserving distribution characteristics; numerical mapping of geological lithology classification variables transforms qualitative data into quantitative data; and min-max standardization unifies the data range. Derived feature calculations comprehensively consider multiple factors, intuitively reflecting locational advantages and environmental impacts. The feature importance weight vector objectively reflects the feature's influence, making subsequent reward function weight adjustments more reasonable, providing a high-quality dataset for reinforcement learning, and contributing to more scientific decision-making.

[0012] In one alternative implementation, the reward function R t Represented as: R t = λ1· R 水能 ,t + λ2· R 建设 ,t + λ3· R 环境 ,t λ4· C t Where λ1, λ2, λ3, λ4 are reward weight coefficients, and satisfy λ1 + λ2 + λ3 + λ4 = 1, and their values ​​are dynamically adjusted by the feature importance weight W; Hydropower Benefit Reward R 水能,t Represented as: R 水能 ,t= α · P 标准化,k + β · t 满发,标准化,k + ψ · (1 r L / H,标准化,k ) Among them, P 标准化,k For the standardized installed capacity of the selected site, t 满发,标准化,k To standardize the number of consecutive full-load hours, r L / H,标准化,k To standardize the distance-to-height ratio, the coefficients α, β, and ψ reflect the priority of the hydropower sub-targets; Feasibility Study Bonus R 建设 , t Represented as:

[0013] Among them, C 地质,标准化,k To standardize geological lithology scoring, S 交通,标准化,k To standardize traffic condition scoring, h 标准化,k The maximum height of the standardized main dam for the k-th upper or lower reservoir is given by the reward value, where a higher reward value indicates lower construction difficulty. These are weighting coefficients, which respectively reflect the importance of geological lithology, transportation conditions, and the maximum dam height to the feasibility of construction, and satisfy the following conditions: ; Environment compatibility reward R 环境,t Quantifying environmental impact: R 环境,t = η· S 环境,k Wherein, η is the weighting coefficient, which reflects the importance of environmental impact in the overall reward; Penalty item C t To avoid ineffective decisions, C is used when the agent selects duplicate sites or sites that do not meet the constraints. t =1, otherwise C t = 0.

[0014] The reward function in this invention integrates multiple benefits and includes a penalty term to guide the agent in comprehensively weighing site selection. Hydropower rewards ensure power generation potential, construction rewards reduce engineering difficulty, environmental rewards consider ecological and social effects, and the penalty term avoids ineffective decisions. The weights of each reward are dynamically adjusted based on feature importance to adapt to the influence of different features, making the strategies generated by reinforcement learning more aligned with multi-objective collaborative optimization needs and improving the overall efficiency of site selection.

[0015] In one alternative implementation, a deep reinforcement learning network is introduced to generate an initial decision-making strategy with the ability to evaluate the comprehensive advantages of a site through experience replay and parameter optimization, including: The deep reinforcement learning network uses a deep Q-network to approximate the action value function, and the expression for the action value function is as follows:

[0016] Where μ is the discount factor and , For the current network parameters, - For target network parameters; The interaction samples of the intelligent agent are stored through an experience playback mechanism (S). t A t , R t , S t+1 The algorithm minimizes the loss function using stochastic gradient descent, iteratively updates the network parameters, and finally generates an initial decision strategy that can output the comprehensive advantage evaluation results of candidate sites.

[0017] This invention utilizes a deep Q-network that combines experience replay with stochastic gradient descent optimization to efficiently learn the optimal strategy. An action value function evaluates action value, experience replay makes training more stable, and stochastic gradient descent iteratively optimizes parameters, enabling the initial decision strategy to accurately assess the overall advantages of a site. Compared to traditional methods, this approach can more accurately and efficiently identify advantageous sites.

[0018] In one optional implementation, the establishment of an online learning and scenario adaptation mechanism to absorb new decision data in real time and incrementally update the model, and optimize policy network parameters, includes: Construct an online experience replay pool to store the scene feature vector Φ at time t. t Interaction samples (S) t A t , R t, S t+1 , Φ t And assign higher sampling weights to new samples; Update policy network parameters based on incremental samples and stochastic gradient descent algorithm; The region type, development stage, and key constraints are transformed into a standardized scenario vector Φ. Based on this scenario vector, the weight coefficients λ of the reward function are adjusted using the following formula:

[0019] Where, λ 0,k γ is the initial weight, and γ is the scene influence factor. Let be the scene feature value related to the k-th type of reward at time t; pass The quadratic weighted state vector completes the optimization of the policy network parameters, where... This is the state vector after being weighted by the importance of features.

[0020] This invention utilizes an online experience replay pool to prioritize new data, enabling the model to quickly absorb new scenario information. Incremental sample updates update parameters, allowing the model to adapt to data changes in real time. Standardized scenario vectors combined with reward weight adjustment formulas make the strategy adaptable to different scenarios. Secondary weighted state vectors enhance the fusion of scenario and site features, making the optimization of policy network parameters more realistic and improving the policy's scenario adaptability and decision accuracy.

[0021] In an optional embodiment, the establishment of the online learning and scenario adaptation mechanism further includes a dynamic exploration rate adjustment mechanism, where the exploration rate ε t The calculation formula is:

[0022] Where ε0 is the initial exploration rate, ε min The minimum exploration rate is given by κ, which is the decay coefficient. Let be the average reward value of the N steps preceding time t; Furthermore, an exploration diversity reward is set up, with the reward formula as follows:

[0023] Where, N visitedM represents the number of explored sites, M represents the total number of candidate sites, and δ represents the exploration reward coefficient.

[0024] This invention employs a dynamic exploration rate adjustment mechanism to balance strategies. When performance is good, the exploration rate is reduced to maintain stability; when performance is poor, the exploration rate is increased to encourage new attempts. Exploration diversity rewards incentivize the exploration of more sites, avoiding local optima. This ensures that the site selection strategy both stably utilizes existing knowledge and explores potentially better sites, improving global optimality and robustness.

[0025] In an optional embodiment, a preferred site list is output, including: Calculate the comprehensive strategy value of each station, which is the weighted sum of the optimal action value output by the strategy network and the objective value of standardized features; The Pareto optimal solution set for four objectives—hydropower utilization efficiency, construction feasibility, environmental compatibility, and strategy stability—was selected using a non-dominated sorting method. The K stations with the highest comprehensive scores are selected from the Pareto optimal solution set to form the final preferred list.

[0026] In an optional embodiment, the comprehensive strategy value calculation takes into account both model decision-making and objective feature advantages, resulting in a more comprehensive evaluation. Non-dominated ranking is used to screen a multi-objective Pareto optimal solution set, balancing objectives such as hydropower, construction, environment, and strategy stability. The K sites with the highest comprehensive scores are selected from the solution set to ensure that the sites in the final list perform well across multiple dimensions, improving the overall optimality of site selection and providing a better combination of sites for the project.

[0027] In an optional embodiment, the multi-dimensional evaluation includes: Hydropower utilization efficiency assessment, calculating the average hydropower efficiency index of the preferred sites; Feasibility assessment is conducted to calculate the comprehensive score of the construction conditions for the preferred site; Environmental compatibility assessment, calculating the average environmental impact score of the preferred site; The superiority of the strategy is evaluated by calculating the overlap rate and performance improvement rate with the results of the preset traditional method. Strategy stability is verified by calculating the standard deviation of the overall performance score and the rate of change of the results through cross-validation and sensitivity analysis.

[0028] This invention comprehensively verifies the optimal site selection and strategy through multi-dimensional evaluation. Hydropower assessment clarifies power generation capacity, construction assessment determines engineering difficulty, environmental assessment demonstrates eco-friendliness, strategy superiority assessment highlights methodological advantages, and strategy stability verification ensures the strategy remains superior even when data or conditions change. This comprehensive approach guarantees the reliability of site selection results and the effectiveness of the method, providing a credible basis for engineering decisions.

[0029] Secondly, the present invention provides a pumped storage site selection system based on reinforcement learning, the system comprising: The raw feature data acquisition module is used to acquire multi-dimensional raw feature data of multiple candidate sites. The multi-dimensional raw feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment. The data preprocessing module is used to preprocess the original feature data to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, which serve as a standardized feature dataset. The reinforcement learning site selection decision model construction module is used to construct a reinforcement learning framework based on the standardized feature dataset, with a state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit rewards, construction feasibility rewards, and environmental compatibility rewards. The weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector. The initial decision strategy acquisition module is used to introduce a deep reinforcement learning network and generate an initial decision strategy with the ability to evaluate the comprehensive advantages of a site through experience playback and parameter optimization. The optimization module is used to establish an online learning and scene adaptation mechanism, which absorbs new decision data in real time and incrementally updates the model to optimize the policy network parameters; it also constructs an online experience replay pool to store the scene feature vector Φ at time t. t Interaction samples (S) t A t , R t, S t+1 , Φ t ), and assign higher sampling weights to the new samples, where A t Let S be the action performed at time t. t Let S be the state vector. t+1 This is the next state vector; Update policy network parameters based on incremental samples and stochastic gradient descent algorithm; The region type, development stage, and key constraints are transformed into a standardized scenario vector Φ. Based on this scenario vector, the weight coefficients λ of the reward function are adjusted using the following formula:

[0030] Where, λ 0,k γ is the initial weight, and γ is the scene influence factor. Let be the scene feature value related to the k-th type of reward at time t; pass The quadratic weighted state vector completes the optimization of the policy network parameters, where... This is the state vector after being weighted by the feature importance weights; The optimization result output and evaluation module is used to output a list of preferred sites based on the optimized decision-making strategy and conduct multi-dimensional evaluation. The final result includes a list of preferred sites, an evaluation report, and a statement of strategy reliability.

[0031] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the reinforcement learning-based pumped storage site selection method described in the first aspect or any corresponding embodiment thereof.

[0032] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the reinforcement learning-based pumped storage site selection method of the first aspect or any corresponding embodiment described above.

[0033] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the reinforcement learning-based pumped storage site selection method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0034] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is a flowchart illustrating the reinforcement learning-based pumped storage site selection method according to an embodiment of the present invention. Figure 2 This is a structural block diagram of a pumped storage site selection system based on reinforcement learning according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] This invention provides an embodiment of a pumped storage site selection method based on reinforcement learning. By using reinforcement learning technology, the method achieves intelligent optimization of pumped storage sites, solving the problems of strong subjectivity and poor scenario adaptability caused by reliance on biased samples in traditional methods. Through self-learning and dynamic optimization mechanisms, it outputs sites with better comprehensive conditions. Figure 1 This is a flowchart of a reinforcement learning-based pumped storage site selection method according to an embodiment of the present invention. It should be noted that the steps shown in the flowchart can be executed in a computer device such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that presented here. Figure 1 As shown, the process includes the following steps: Step S1: Obtain multi-dimensional original feature data of multiple candidate sites. The multi-dimensional original feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment.

[0038] Specifically, the location conditions in this embodiment of the invention include: distance d from the load center. 负荷中心,i Distance d from the substation 变电站,i ; distance d from the new energy cluster 新能源集群,i ; Hydropower parameters include: Installed capacity: P i Continuous full-load hours t 满发,i Average head H i Horizontal distance L between upper and lower warehouses i Distance r L / H,i ; Construction conditions include: Maximum dam height: h (upper reservoir height) 上库,i 、Lower warehouse h 下库,i Geological lithology code C 地质,i Traffic Conditions Score 交通,i ; The external environment includes: the number N of potentially significant and sensitive objects. 敏感对象,i , number of people submerged N 人口,i .

[0039] This invention achieves comprehensive coverage of the multi-dimensional attributes of pumped storage sites by clearly defining four core parameters: location conditions (including distance from load centers, substations, and new energy clusters), hydropower parameters (including installed capacity, continuous full-load hours, etc.), construction conditions (including watershed area, normal water level, etc.), and external environment (including the number of potentially major sensitive objects and the number of people to be submerged). This avoids decision-making biases caused by missing or one-sided parameters in traditional site selection, provides more complete and objective basic data support for site selection, and significantly improves the matching degree between site selection results and actual development needs.

[0040] Step S2: Preprocess the original feature data to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, which serve as a standardized feature dataset.

[0041] First, this embodiment of the invention uses the interquartile range (IQR) method to identify and correct outliers. Specifically, the IQR method is used to truncate values ​​that exceed a reasonable range, as expressed in the following formula:

[0042] Where Q1 and Q3 are the lower and upper quartiles of the feature, respectively, and IQR = Q3 Q1, ensure the rationality of data distribution.

[0043] For missing values, imputation is performed using the mean of similar sites, and the estimated value of the missing feature is calculated using the following formula:

[0044] in, This is a set of sites with no missing features, ensuring data integrity.

[0045] Numerical mapping is performed on geological lithology classification variables to convert them into quantitative scores. The numerical mapping is implemented through the following logic, transforming qualitative descriptions into quantitative scores:

[0046] in sk Scoring for lithological categories, δik This is an indicator function used to solve the quantification problem of categorical data.

[0047] To eliminate the dimensional differences between different features, the min-max normalization method is used to map all continuous features to the [0,1] interval, as expressed by the formula:

[0048] Where, x ij For the original eigenvalues, min(x) j ), max(x j) represent the minimum and maximum values ​​of this feature, respectively, x ij,标准化 This is a standardized result.

[0049] Based on this, key derived characteristics are further calculated, including distance ratio, installed capacity, location comprehensive index, and environmental impact score. The distance-to-height ratio integrates the correlation between horizontal distance and head using the following formula: .

[0050] Installed capacity is normalized using the following formula: .

[0051] The location comprehensive index in this invention is calculated using the following formula, which transforms distance characteristics into a positive score: S 区位,i = α1 ·(1 d 负荷中心,标准化,i ) + α2 ·(1 d 变电站,标准化,i ) + α3 ·(1 d 新能源集群,标准化,i ) The environmental impact score is calculated using the following formula:

[0052] To quantify the importance of each feature, the feature importance weight vector uses variance contribution to calculate the initial weights, as expressed in the following formula:

[0053] in, w represents the variance of the standardized features. j This reflects the explanatory power of the feature for the overall differences; finally, all preprocessed features are integrated into a standardized multi-dimensional feature matrix X. i = [x i1,标准化 , x i2,标准化 , ..., x ij,标准化 ,..., x in,标准化 This forms the feature vector for each site.

[0054] During preprocessing, the interquartile range method corrects outliers, avoiding interference from extreme data in subsequent analysis and ensuring data quality. Using the mean of similar sites to fill missing values ​​improves data integrity while preserving distribution characteristics as much as possible. Numerical mapping of geological lithology classification variables transforms qualitative information into quantitative scores, facilitating quantitative evaluation. Min-max standardization unifies data intervals, enabling effective comparison and calculation of different features. Calculating derived features such as the location comprehensive index integrates multiple locational factors, intuitively reflecting the locational advantages of a site. Environmental impact scoring quantifies environmental constraints. Feature importance weight vectors are determined through variance contribution, objectively reflecting the influence of each feature on site selection, allowing for more reasonable adjustment of reward function weights in subsequent reinforcement learning. Overall, this improves data preprocessing quality and provides a higher-quality, more targeted standardized feature dataset for subsequent reinforcement learning frameworks.

[0055] The feature matrix X and feature importance weight vector W obtained in this embodiment of the invention provide a comprehensive and objective input for the construction of the state space of the subsequent reinforcement learning model, ensuring that the state can accurately reflect the core attributes of the site; the latter provides a basis for the design of the reinforcement learning reward mechanism, enabling the reward function to dynamically adjust the weights based on feature importance, avoiding human-set bias, and laying a solid data foundation for the agent to learn a reasonable point selection strategy.

[0056] Step S3: Construct a reinforcement learning framework based on the standardized feature dataset, with state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit reward, construction feasibility reward, and environmental compatibility reward. The weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector.

[0057] The preprocessing result of this embodiment of the invention, state vector S t Defined as a high-dimensional vector containing features of the current candidate site set, information of selected sites, and scenario constraints, with the specific expression as follows: S t = [X t , Θ t , Γ t ] Among them, X t This is the standardized feature matrix of the remaining candidate sites at time t (dimension is the number of remaining sites × number of features), containing the location-standardized features of each site (e.g., d). 负荷中心,标准化,i d 变电站,标准化,i ), hydropower standardization characteristics (such as P) 标准化,i r L / H,标准化,i), construction standardization features (such as h) 标准化,i、S交通,标准化,i ) and environmental standardization characteristics (such as S 环境,i ); Θt is the set of feature vectors of the selected sites, which records the attributes of the sites that have been included in the selection range during the decision-making process to avoid duplicate selection; Γt is the scenario constraint vector, which includes external constraint parameters such as the water energy demand threshold and environmental sensitivity level of the development area to ensure that the decision meets the requirements of the actual scenario.

[0058] The dimensional design of the state space needs to cover all core feature dimensions of the module, and the state vectors are weighted using feature importance weights W, i.e. This allows intelligent agents to focus more on highly important features (such as installed capacity, geological lithology, etc.) and improve the effectiveness of state representation.

[0059] The action space design needs to enable the agent to make decisions regarding candidate sites, balancing the flexibility of site selection with the feasibility of the decisions. Action A t Defined as a discrete action by which an agent selects a site or adjusts the site priority from the current set of candidate sites at time t, it is specifically divided into two categories: 1. Single-site selection action A sel,t This indicates that one site is selected from the remaining candidate sites to be included in the preferred set, expressed as A. sel,t = k (where k is the index of the remaining candidate sites, k ∈ [1, M)). t ], M t (The number of remaining stations at time t). 2. Multi-site sorting action A rank,t This indicates that the current candidate sites are prioritized based on their overall advantages. Output sorted result vector A rank,t = [r1, r2, ..., rM t (where r) i To assign a priority score to the i-th site, r i (∈ [0, 1], the higher the score, the more significant the overall advantage).

[0060] The size of the action space is dynamically adjusted according to the number of candidate sites. By setting an action mask mechanism, the agent is prevented from selecting already selected sites or sites that do not meet the scene constraints, thus ensuring the rationality of the action.

[0061] The reward mechanism is the core of guiding the agent to learn effective strategies. It needs to transform the development goals of pumped storage sites (such as high hydropower utilization, low construction costs, and minimal environmental impact) into quantifiable reward signals. Based on the feature importance weights W of the modules, the reward function Rt adopts a multi-dimensional weighted summation form, comprehensively considering the three core objectives of hydropower efficiency, construction feasibility, and environmental compatibility. The specific expression is as follows: Rt = λ1· R 水能 ,t + λ2· R 建设 ,t + λ3· R 环境 ,t λ4· C t Wherein, λ1, λ2, λ3, λ4 are reward weight coefficients, and satisfy λ1+ λ2+ λ3+ λ4= 1. Their values ​​are dynamically adjusted by the feature importance weight W (e.g., if the installed capacity weight in W is high, then λ1 will increase accordingly). Hydropower Benefit Reward R 水能,t Represented as: R 水能 ,t= α · P 标准化,k + β · t 满发,标准化,k + ψ · (1 r L / H,标准化,k ) Among them, R 水能,t P represents the hydropower benefit reward for a pumped storage hydroelectric power station at time t. 标准化,k Let t be the standardized installed capacity of the k-th selected site. 满发,标准化,k Let r be the number of consecutive full-load hours of the kth standardized period. L / H,标准化,k For the k-th standardized distance-to-height ratio, the coefficients α, β, and ψ reflect the priority of the hydropower sub-targets; Feasibility Study Bonus R 建设 , t Represented as:

[0062] Among them, C 地质,标准化,k For the standardized geological lithology score of the k-th selected site, S 交通,标准化,k For the standardized traffic condition score of the k-th selected station, h 标准化,k The maximum height of the standardized main dam for the k-th upper or lower reservoir is given by the reward value, where a higher reward value indicates lower construction difficulty. These are weighting coefficients, which respectively reflect the importance of geological lithology, transportation conditions, and the maximum dam height to the feasibility of construction, and satisfy the following conditions: A higher reward value indicates a lower construction difficulty. Environment compatibility reward R 环境,t Quantifying environmental impact: R 环境,t = η· S 环境,k Here, η is the weighting coefficient, which reflects the importance of environmental impact in the overall reward.

[0063] Penalty item C t To avoid ineffective decisions, C is used when the agent selects duplicate sites or sites that do not meet the constraints. t =1, otherwise C t = 0.

[0064] The reward function provided in this invention integrates hydropower, construction, and environmental benefits, and includes penalty terms. This guides the reinforcement learning agent to comprehensively weigh power generation capacity, construction feasibility, and environmental compatibility when selecting a site, avoiding the neglect of other important factors due to an emphasis on a single objective. The hydropower benefit reward considers installed capacity and full-capacity operation time to ensure the site's power generation potential; the construction feasibility reward focuses on geology and transportation to reduce the difficulty of engineering implementation; the environmental compatibility reward quantifies environmental impact, taking into account both ecological and social effects; and the penalty terms avoid ineffective decisions and improve decision-making effectiveness. The weights of each reward are dynamically adjusted based on feature importance, enabling the reward mechanism to adapt to the impact of different features on the site. This makes the reinforcement learning-generated strategy more aligned with the needs of multi-objective collaborative optimization in actual site selection, improving the comprehensive benefits of site selection in economic, technological, and ecological aspects.

[0065] Step S4: Introduce a deep reinforcement learning network and generate an initial decision-making strategy with the ability to evaluate the comprehensive advantages of a site through experience playback and parameter optimization.

[0066] To enable the agent to efficiently learn the point selection strategy, this embodiment of the invention introduces a deep reinforcement learning network architecture, using a deep Q-network (DQN) to approximate the action value function Q(S). t A t ; θ), where θ are network parameters. The value function is defined as the value of the agent in state S. t Next, execute action A t The cumulative discount reward obtained later is:

[0067] in, Discount factor ( This is used to balance immediate rewards and long-term rewards. θ represents the current network parameters. These are the target network parameters.

[0068] The interaction samples of the intelligent agent are stored through an experience playback mechanism (S). t A t , Rt , S t+1 And the stochastic gradient descent algorithm is used to minimize the loss function:

[0069] Iterative updates of network parameters are achieved, and the output is the agent's policy network, which can be based on the input state vector S. t Output optimal action This output represents the site selection or ranking result that best combines overall advantages. This output directly provides an initial decision model for subsequent dynamic learning and strategy iteration. Through objective feature data and the autonomous decision-making capability of reinforcement learning, it ensures that the site selection strategy can reflect the actual attribute differences of the sites and dynamically adapt to development goals and scenario constraints, providing intelligent decision support for the optimal selection of pumped storage sites.

[0070] Step S5: Establish an online learning and scenario adaptation mechanism to absorb new decision data in real time and incrementally update the model to optimize the policy network parameters.

[0071] This invention addresses the problem of policy failure in traditional methods when scenarios change by constructing an online learning mechanism, a dynamic scenario adaptation model, and a strategy evaluation and iteration framework. It enables the autonomous evolution of point selection strategies under different development scenarios and regional characteristics. Based on an initial policy network, and combined with real-time interactive data and changes in scenario characteristics, the decision model parameters are dynamically adjusted to ensure that the point selection strategy always aligns with actual development needs.

[0072] The core of dynamic learning lies in establishing a continuous interaction mechanism between the agent and the environment to achieve online policy updates. Based on a deep Q-network architecture, an online experience replay pool and an incremental learning mechanism are introduced, enabling the agent to absorb new decision feedback data in real time. The online experience replay pool stores the latest samples generated by the agent in actual point selection decisions.

[0073] Where, Φ t Let be the scene feature vector at time t (containing dynamic information such as regional water energy demand and environmental constraint levels). Compared with historical samples, the new sample is assigned a higher sampling weight ωt, and the weight calculation formula is:

[0074] Where α is the weight coefficient (α ∈ [0, 1]), β is the decay factor, t0 is the sample generation time, and t is the current training time. This formula ensures that recent samples have a greater impact on policy updates, thereby improving the model's adaptability to the latest scenarios.

[0075] In incremental learning, the update of the policy network parameters θ no longer depends on the entire set of samples, but is instead iteratively optimized using stochastic gradient descent on new samples. The parameter update formula is:

[0076] Where η is the learning rate. The loss function is calculated based on the new sample at time t. It has the same form as the loss function but only uses incremental samples, which greatly reduces the computational cost of repeated training.

[0077] This invention achieves dynamic strategy adaptation by quantifying changes in scene features. First, a scene feature extraction module is constructed to transform qualitative scene information, such as region type (e.g., mountainous / plain), development stage (e.g., early / mid-stage planning), and key constraints (e.g., economic priority / environmental priority), into quantitative vectors.

[0078] in, The standardized value of the m-th scene feature ( ∈ [0, 1]).

[0079] The weight coefficient λ of the reward function is adjusted based on the scene vector Φ, and the adjustment formula is as follows:

[0080] Where, λ 0,k λ is the initial reward weight (determined by the feature importance weight W). t,k Let γ be the weight of the k-th type of reward at time t, and let γ be the scene influence factor (γ ∈ [0, 0.5]). Let be the scene feature value related to the k-th type of reward at time t.

[0081] For example, in areas rich in new energy sources, the location reward weight λ1, which is related to the distance to new energy clusters, will increase accordingly; in ecologically sensitive areas, the environmental reward weight λ3 will increase significantly, making the strategy more focused on site selection with low environmental impact. Simultaneously, the scenario adaptation mechanism will dynamically adjust the feature weights of the state space, through... The state vector is weighted twice to enhance the feature signals that match the current scene. This is the state vector after being weighted by the importance of features.

[0082] This invention assigns higher sampling weights to new samples through an online experience replay pool, prioritizing the use of the latest decision data and enabling the model to quickly absorb information from new scenarios. Updating parameters based on incremental samples allows the model to adapt to data changes in real time. Standardized scenario vectors combined with reward function weight adjustment formulas can dynamically adjust reward weights according to region type, development stage, etc., making the strategy adaptable to different scenario requirements. The double-weighted state vector strengthens the fusion of scenario features and site features, making the optimization of strategy network parameters more closely aligned with actual scenarios. This allows the pumped storage site selection strategy to flexibly address differences in different regions and development stages, improving the strategy's scenario adaptability and decision-making accuracy.

[0083] To balance the stability and exploratory nature of the strategy, this embodiment also provides a dynamic exploration rate adjustment mechanism to prevent the agent from getting trapped in local optima. Exploration rate ε t (That is, the probability of the agent choosing a random action) changes dynamically with the training process and policy performance, and is calculated using the following formula:

[0084] Where ε0 is the initial exploration rate, ε min The minimum exploration rate is given by κ, which is the decay coefficient. Let be the average reward value of the N steps preceding time t. When the policy performance improves ( When ε increases t Automatically reduce ineffective exploration; when the strategy stagnates ( When ε remains constant over a long period of time, t Maintain a high level of activity and encourage agents to try new actions and discover better strategies.

[0085] Simultaneously, an exploration diversity reward is introduced. When the agent selects a combination of sites that has not been tried before, an additional exploration reward is given. The reward formula is as follows:

[0086] Where, N visited M represents the number of explored sites, M represents the total number of candidate sites, and δ represents the exploration reward coefficient.

[0087] The strategy evaluation and iterative optimization mechanism of this invention ensures continuous improvement in model performance, and sets up a dual-period evaluation framework: short-term evaluation calculates the average reward of the most recent T steps using a sliding window approach.

[0088] when If there is no improvement for K consecutive periods, a strategy fine-tuning is triggered; long-term evaluation calculates the strategy's comprehensive point selection score across all candidate sites.

[0089] in Let i be the state vector of the i-th station. For the optimal action, when S long If the value falls below a preset threshold, policy retraining is initiated.

[0090] During retraining, a policy distillation mechanism is introduced to transfer knowledge of the historical best policy to the new policy, minimizing the KL divergence between the old and new policies.

[0091] Retaining effective decision-making experience accelerates the convergence of new strategies, and the output is the optimized policy network parameters. Set of parameters for scene adaptation . in, These are the optimal strategy parameters after dynamic learning and scenario adaptation, enabling more accurate site selection decisions in the current and similar scenarios. The scenario adaptation parameters record the strategy adjustment rules under different scenarios, providing a transferable adaptation mechanism for site selection decisions across regions and stages. The output is the module's optimal result. The output and evaluation provide high-performance decision model support, ensuring that the final optimal result not only has static comprehensive advantages but also can dynamically adapt to changes in development scenarios, realizing intelligent and long-term optimization of pumped storage sites.

[0092] Step S6: Based on the optimized decision-making strategy, output a list of preferred sites and conduct multi-dimensional evaluation, outputting the final results including the list of preferred sites, evaluation report and strategy reliability statement.

[0093] In this invention, the output calculates the comprehensive strategy value for each station, which is a weighted sum of the optimal action value output by the strategy network and the objective value of standardized features, as shown in the formula:

[0094] in, For the optimal action value at the i-th station, ω str For strategy value weights; x ij,标准化 For the standardized eigenvalues, w j ω represents the feature importance weight. feat The feature value weights are ω. str + ω feat = 1.

[0095] The above formula integrates the decision-making value of strategy learning with the objective value of original features, avoiding bias from a single evaluation dimension. Based on this, a non-dominated ranking method is used to rank the sites across multiple objectives, with hydropower efficiency, construction feasibility, and environmental compatibility as core objectives. This selects the Pareto optimal solution set that is not dominated by other sites in any of the three objective dimensions. Finally, based on the scale of development needs (e.g., planning to build 1-3 sites), the K sites with the highest comprehensive scores are selected from the Pareto optimal solution set, forming a list of preferred results. O k The index and core feature parameters of the k-th preferred site.

[0096] The construction of a multi-dimensional evaluation system is the key to verifying the validity of the results, covering four dimensions: hydropower utilization efficiency, construction feasibility, environmental compatibility, and strategy stability.

[0097] Hydropower utilization efficiency assessment focuses on the energy output potential of key sites and calculates the average hydropower efficiency index E for preferred sites. 水能 :

[0098] Where P ok t 满发,ok H ok These are the installed capacity, continuous full-load hours, and average head of the preferred sites, respectively, max(E 水能,all The index represents the maximum hydropower efficiency value among all candidate sites. The closer the index is to 1, the higher the hydropower utilization efficiency.

[0099] The feasibility assessment quantifies the difficulty of developing and implementing the site, and uses a comprehensive score (E) based on the construction conditions. 建设 measure:

[0100] Where C 地质,ok For geological lithology scoring, S 交通,ok To score traffic conditions, h ok,标准化 To standardize the height of the main dam, a higher score indicates a lower construction difficulty.

[0101] Environmental compatibility assessment focuses on ecological impact control and calculates the environmental impact index Eenvironment.

[0102] A higher environmental impact index indicates a smaller negative environmental impact.

[0103] To verify the superiority of the reinforcement learning strategy, this embodiment of the invention introduces a comparative evaluation mechanism, performing a difference analysis between the optimized results and the results of traditional methods (such as clustering recommendation methods). The overlap rate R between the results of the two methods is calculated.overlap :

[0104] Among them O trad The result is the optimal choice for traditional methods, with the overlap rate reflecting the consistency of the strategy; simultaneously, the performance improvement R of the reinforcement learning result relative to the traditional method is calculated. impr :

[0105] Where E total (O) = 0.4 · E 水能 + 0.3 · E 建设 + 0.3 · E 环境 The improvement rate directly reflects the advantages of reinforcement learning strategies when used to calculate the overall performance score.

[0106] Strategy stability verification ensures the reliability of results under data fluctuations or scenario changes, achieved through cross-validation and sensitivity analysis. In this embodiment, cross-validation employs a K-fold partitioning method, dividing the candidate site dataset into K subsets, and using K fold partitioning steps each time. Train the model on one subset and test it on the remaining subset. Calculate the standard deviation σ of the results of K tests. cv :

[0107] Where E total,k The overall performance score for the k-th fold. The smaller the standard deviation, the better the stability of the strategy, representing the average score.

[0108] Sensitivity analysis, on the other hand, fine-tunes the core eigenvalues.

[0109] Where O′ represents the optimal result after feature fine-tuning, with a change rate lower than a preset threshold.

[0110] The final output consists of three parts: 1. A list of preferred sites, including the core parameters of each site (such as installed capacity, distance-to-height ratio, environmental impact score, etc.) and overall ranking; 2. A multi-dimensional assessment report, covering quantitative scores for hydropower efficiency, construction feasibility, and environmental compatibility, as well as comparative analysis with traditional methods; 3. Strategy reliability description, including stability indicators such as cross-validation standard deviation and sensitivity analysis results. These outputs provide comprehensive technical support for the subsequent exploration, design, and investment decisions of pumped storage sites, ensuring that the optimal selection results not only conform to objective data patterns but also meet the diverse needs of actual development, achieving an organic unity between scientific site selection and efficient development.

[0111] This embodiment also provides a reinforcement learning-based pumped storage site selection system, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0112] This embodiment provides a pumped storage site selection system based on reinforcement learning, such as... Figure 2 As shown, it includes: The original feature data acquisition module 21 is used to acquire multi-dimensional original feature data of multiple candidate sites. The multi-dimensional original feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment. Data preprocessing module 22 is used to preprocess the original feature data to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, as a standardized feature dataset; The reinforcement learning site selection decision model construction module 23 is used to construct a reinforcement learning framework based on the standardized feature dataset, with state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit reward, construction feasibility reward, and environmental compatibility reward. The weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector. The initial decision strategy is obtained modulo 24 and used to introduce a deep reinforcement learning network. Through experience replay and parameter optimization, an initial decision strategy with the ability to evaluate the comprehensive advantages of a site is generated. Optimization module 25 is used to establish an online learning and scenario adaptation mechanism, absorb new decision data in real time and incrementally update the model to optimize the strategy network parameters; The optimization result output and evaluation module 26 is used to output a list of preferred sites based on the optimized decision-making strategy and conduct multi-dimensional evaluation, and output the final results including the list of preferred sites, evaluation report and strategy reliability statement.

[0113] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0114] The reinforcement learning-based pumped storage site selection system in this embodiment is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0115] This invention also provides a computer device having the above-described features. Figure 2 The system shown is a pumped storage site selection system based on reinforcement learning.

[0116] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 3 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 3 Take a processor 10 as an example.

[0117] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0118] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0119] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0120] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0121] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0122] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0123] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0124] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for selecting pumped storage sites based on reinforcement learning, characterized in that, include: Obtain multi-dimensional raw feature data of multiple candidate sites. The multi-dimensional raw feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment. The original feature data is preprocessed to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, which serve as the standardized feature dataset. Based on the standardized feature dataset, a reinforcement learning framework is constructed with state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit reward, construction feasibility reward, and environmental compatibility reward, and the weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector. By introducing a deep reinforcement learning network and through experience replay and parameter optimization, an initial decision-making strategy with the ability to evaluate the comprehensive advantages of a site is generated. Establish an online learning and scenario adaptation mechanism to absorb new decision data in real time and incrementally update the model to optimize policy network parameters; construct an online experience replay pool to store the scenario feature vector Φ at time t. t Interaction samples (S) t A t , R t, S t+1 , Φ t ), and assign higher sampling weights to the new samples, where A t Let S be the action performed at time t. t Let S be the state vector. t+1 This is the next state vector; Update policy network parameters based on incremental samples and stochastic gradient descent algorithm; The region type, development stage, and key constraints are transformed into a standardized scenario vector Φ. Based on this scenario vector, the weight coefficients λ of the reward function are adjusted using the following formula: Where, λ 0,k γ is the initial weight, and γ is the scene influence factor. Let be the scene feature value related to the k-th type of reward at time t; pass The quadratic weighted state vector completes the optimization of the policy network parameters, where... This is the state vector after being weighted by the feature importance weights; Based on the optimized decision-making strategy, a list of preferred sites is output and evaluated from multiple dimensions. The final output includes the list of preferred sites, an evaluation report, and a statement on the reliability of the strategy.

2. The method according to claim 1, characterized in that, The location conditions include: distance d from the load center. 负荷中心,i Distance d from the substation 变电站,i Distance d from the new energy cluster 新能源集群,i ; The hydropower parameters include: installed capacity P. i Continuous full-load hours t 满发,i Average head H i Horizontal distance L between upper and lower warehouses i Distance to height ratio r L / H,i ; The construction conditions include: maximum dam height: upper reservoir height h 上库,i 、Lower warehouse h 下库,i Geological lithology code C 地质,i Traffic Conditions Score 交通,i ; The external environment includes: the number N of potentially significant and sensitive objects. 敏感对象,i , number of people submerged N 人口,i .

3. The method according to claim 2, characterized in that, The preprocessing of the original feature data to generate a standardized multi-dimensional feature matrix and feature importance weight vectors, as a standardized feature dataset, includes: Use the interquartile range method to identify and correct outliers; Missing values ​​were handled using the mean of similar sites; Numerical mapping is performed on geological lithology classification variables to convert them into quantitative scores; The min-max normalization method is used to map all continuous features to the [0,1] interval; The calculation includes derived characteristics such as distance-to-height ratio, normalized installed capacity, location comprehensive index, and environmental impact score; wherein, the location comprehensive index is calculated using the following formula: S 区位,i = α1·(1 d 负荷中心,标准化,i ) + α2·(1 d 变电站,标准化,i ) + α3·(1 d 新能源集群,标准化,i ) Wherein, α1, α2, and α3 are weighting coefficients, which respectively reflect the importance of the distance from the i-th load center, substation, and new energy cluster to the location advantage, and satisfy α1 + α2 + α3 = 1; The environmental impact score is calculated using the following formula: ; in and These respectively reflect the importance of the number of potentially major sensitive objects and the number of people to be submerged to the environmental impact, and ; The feature importance weight vector is obtained by calculating the variance contribution of each standardized feature; The preprocessed features are integrated into a standardized multi-dimensional feature matrix to form the feature vector for each site.

4. The method according to claim 2 or 3, characterized in that, The reward function R t Represented as: R t = λ1· R 水能 , t + λ2· R 建设 , t + λ3· R 环境 , t λ4· C t Among them, R t Let λ1, λ2, λ3, λ4 be the total reward at time t, and let λ1, λ2, λ3, λ4 be the reward weight coefficients, satisfying λ1 + λ2 + λ3 + λ4 = 1. Their values ​​are dynamically adjusted by the feature importance weight W. Hydropower Benefit Reward R 水能,t Represented as: R 水能 , t = a · P 标准化,k + β t 满发,标准化,k + ψ (1 r L / H,标准化,k ) Among them, R 水能,t P represents the hydropower benefit reward for a pumped storage hydroelectric power station at time t. 标准化,k Let t be the standardized installed capacity of the k-th selected site. 满发,标准化,k Let r be the number of consecutive full-load hours of the kth standardized period. L / H,标准化,k For the k-th standardized distance-to-height ratio, the coefficients α, β, and ψ reflect the priority of the hydropower sub-targets; Construction Feasibility Incentive R 建设 , t Represented as: Among them, C 地质,标准化,k For the standardized geological lithology score of the k-th selected site, S 交通,标准化,k For the standardized traffic condition score of the k-th selected station, h 标准化,k The maximum height of the standardized main dam for the k-th upper or lower reservoir is given by the reward value, where a higher reward value indicates lower construction difficulty. These are weighting coefficients, which respectively reflect the importance of geological lithology, transportation conditions, and the maximum dam height to the feasibility of construction, and satisfy the following conditions: ; Environment compatibility bonus R 环境,t Quantifying environmental impact: R 环境,t = n· S 环境,k Wherein, η is the weighting coefficient, which reflects the importance of environmental impact in the overall reward; Penalty item C t To avoid ineffective decisions, C is used when the agent selects duplicate sites or sites that do not meet the constraints. t = 1, otherwise C t = 0.

5. The method according to claim 1, characterized in that, By introducing a deep reinforcement learning network and optimizing parameters through experience replay, an initial decision-making strategy with the ability to evaluate the comprehensive advantages of a site is generated, including: The deep reinforcement learning network uses a deep Q-network to approximate the action value function, and the expression of the action value function is as follows: Where μ is the discount factor and , θ represents the current network parameters. S represents the target network parameters. t Let S be the state vector at time t. t+1 Let A be the next state vector. t Let A be the action performed at time t. t+1 The action to be performed in the next moment; The interaction samples of the intelligent agent are stored through an experience playback mechanism (S). t A t , R t , S t+1 The algorithm minimizes the loss function using stochastic gradient descent, iteratively updates the network parameters, and finally generates an initial decision strategy that can output the comprehensive advantage evaluation results of candidate sites.

6. The method according to claim 1, characterized in that, The established online learning and scenario adaptation mechanism also includes a dynamic exploration rate adjustment mechanism, where the exploration rate ε t The calculation formula is: Where ε0 is the initial exploration rate, ε min The minimum exploration rate is given by κ, which is the decay coefficient. Let be the average reward value of the N steps preceding time t; Furthermore, an exploration diversity reward is set up, with the reward formula as follows: Where, N visited M represents the number of explored sites, M represents the total number of candidate sites, and δ represents the exploration reward coefficient.

7. The method according to claim 1, characterized in that, Output a list of preferred sites, including: Calculate the comprehensive strategy value of each station, which is the weighted sum of the optimal action value output by the strategy network and the objective value of standardized features; The Pareto optimal solution set for four objectives—hydropower utilization efficiency, construction feasibility, environmental compatibility, and strategy stability—was selected using a non-dominated sorting method. The K stations with the highest comprehensive scores are selected from the Pareto optimal solution set to form the final preferred list.

8. The method according to claim 1 or 7, characterized in that, Multi-dimensional assessment includes: Hydropower utilization efficiency assessment, calculating the average hydropower efficiency index of the preferred sites; Feasibility assessment is conducted to calculate the comprehensive score of the construction conditions for the preferred site; Environmental compatibility assessment, calculating the average environmental impact score of the preferred site; The superiority of the strategy is evaluated by calculating the overlap rate and performance improvement rate with the results of the preset traditional method. Strategy stability is verified by calculating the standard deviation of the overall performance score and the rate of change of the results through cross-validation and sensitivity analysis.

9. A pumped storage site selection system based on reinforcement learning, characterized in that, The system includes: The raw feature data acquisition module is used to acquire multi-dimensional raw feature data of multiple candidate sites. The multi-dimensional raw feature data includes at least four categories of parameters: location conditions, hydropower parameters, construction conditions, and external environment. The data preprocessing module is used to preprocess the original feature data to generate a standardized multi-dimensional feature matrix and a feature importance weight vector, which serve as a standardized feature dataset. The reinforcement learning site selection decision model construction module is used to construct a reinforcement learning framework based on the standardized feature dataset, with a state space, action space, and reward mechanism as its core. The state space includes current candidate site features, selected site information, and scenario constraint vectors. The action space includes actions such as selecting sites from candidate sites or prioritizing candidate sites. The reward function includes at least hydropower benefit rewards, construction feasibility rewards, and environmental compatibility rewards, and the weight coefficients of the reward function are dynamically adjusted by the feature importance weight vector. The initial decision strategy acquisition module is used to introduce a deep reinforcement learning network and generate an initial decision strategy with the ability to evaluate the comprehensive advantages of a site through experience playback and parameter optimization. The optimization module is used to establish an online learning and scene adaptation mechanism, which absorbs new decision data in real time and incrementally updates the model to optimize the policy network parameters; it also constructs an online experience replay pool to store the scene feature vector Φ at time t. t Interaction samples (S) t A t , R t, S t+1 , Φ t ), and assign higher sampling weights to the new samples, where A t Let S be the action performed at time t. t Let S be the state vector. t+1 This is the next state vector; Update policy network parameters based on incremental samples and stochastic gradient descent algorithm; The region type, development stage, and key constraints are transformed into a standardized scenario vector Φ. Based on this scenario vector, the weight coefficients λ of the reward function are adjusted using the following formula: Where, λ 0,k γ is the initial weight, and γ is the scene influence factor. Let be the scene feature value related to the k-th type of reward at time t; pass The quadratic weighted state vector completes the optimization of the policy network parameters, where... This is the state vector after being weighted by the feature importance weights; The optimization result output and evaluation module is used to output a list of preferred sites based on the optimized decision-making strategy and conduct multi-dimensional evaluation. The final output includes a list of preferred sites, an evaluation report, and a statement of strategy reliability.

Citation Information

Patent Citations

  • Pumped storage power station multi-element multi-target intelligent online site selection method and device and storage medium

    CN119784108A

  • Diversified recommendation method based on graph neural network and reinforcement learning

    CN120086441A