Wind power installation potential assessment method based on marginal constraint model
By combining the marginal constraint model and random forest algorithm with the ecological network intelligent avoidance algorithm, the problems of subjectivity and ecological impact in the existing wind power installed capacity potential assessment are solved, realizing accurate and transparent wind farm installed capacity potential assessment and ecological protection, which is applicable to multi-level energy spatial layout.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-05
AI Technical Summary
Existing methods for assessing wind power installed capacity potential are highly subjective, lack causal explanations, and ignore ecological impacts, leading to unstable assessment results and resource waste. They are difficult to meet the requirements of scientific rigor and transparency, and overlapping ecological protection areas can cause project adjustments or cancellations.
We employ a marginal constraint model and random forest algorithm, combined with multi-source data from wind farm site selection candidate areas, to construct feature vectors. We then use the random forest model to predict suitability probabilities, combine constraint line equations and utilization coefficients to assess installed capacity potential, and introduce an ecological network intelligent avoidance algorithm to protect the ecological environment.
It enables more accurate and efficient assessment of wind power installed capacity potential, reduces subjective factors, improves the accuracy and transparency of assessment, ensures the scientific nature and ecological safety of wind farm construction, and is applicable to the optimization of wind farm layout under different terrain and policy conditions.
Smart Images

Figure CN121980633A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of assessing the maximum exploitable installed capacity of wind farms, and specifically relates to a method for assessing wind power installed capacity potential based on a marginal constraint model. Background Technology
[0002] Wind energy, with its abundant resources, mature technology, and great potential for large-scale development, has become a core support for building a new energy system. However, the development and utilization of wind energy resources is not easy, as its site selection and construction are subject to both complex natural geographical conditions and diverse social constraints. Therefore, a scientific and accurate assessment of a region's wind power installation potential requires a comprehensive consideration of resource endowment, techno-economic feasibility, and ecological and environmental impacts, forming a complex system decision-making process.
[0003] Current mainstream evaluation methods primarily rely on resource endowment conditions such as wind speed and wind power density, combined with assessments of natural topography and land cover types to evaluate regional installed capacity potential. The core technology lies in constructing restrictive maps using meteorological station data, reanalysis data, or numerical simulation results, and leveraging the spatial overlay and quantitative analysis functions of Geographic Information Systems (GIS). Furthermore, based on preset slope thresholds and land use restrictions, different levels of installed capacity potential are classified. These methods offer advantages in data processing and spatial visualization, providing relatively simple operation and intuitive results.
[0004] However, existing evaluation systems have significant limitations: First, traditional wind power capacity potential assessments (such as multi-index weighted superposition models) suffer from strong subjectivity, unreasonable index weighting, and a lack of causal explanation, making them difficult to adapt to the higher requirements of scientific rigor, transparency, and adaptability in modern wind power development. Second, an ecological perspective is lacking. Existing methods mainly focus on resource and topographical factors, generally neglecting the potentially severe disturbances to regional ecosystems that large-scale wind farm construction may cause (such as habitat fragmentation, soil erosion, noise and electromagnetic interference, and landscape disruption).
[0005] This directly leads to two prominent problems: First, the evaluation system relies too heavily on subjective assignment and lacks comparison with existing data, resulting in unstable and unreproducible model results. The actual installed capacity differs significantly from the theoretical installed capacity, making it difficult to provide convincing technical evidence for planning approval and policy formulation. Second, many areas initially deemed "suitable" actually overlap significantly with important ecological protection areas or habitats of rare and endangered species. Projects are forced to halt or be adjusted during the subsequent rigorous Environmental Impact Assessment (EIA) phase due to ecological conflicts, resulting in wasted initial investment and delays in development. Therefore, there is an urgent need to introduce an intelligent installed capacity potential evaluation method that can integrate multi-dimensional influencing factors, possess predictive capabilities, and demonstrate causal interpretability. This is especially crucial given the rigid constraints on ecological protection, to achieve coordinated and efficient collaboration between wind power development and ecological security. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a method for evaluating wind power installed capacity potential based on a marginal constraint model. This invention innovatively combines a constraint line model and a random forest algorithm, providing a more accurate and efficient way to evaluate the installed capacity potential of wind farms, and aims to provide decision support for wind energy development.
[0007] This invention proposes a method for evaluating wind power installed capacity potential based on a marginal constraint model, comprising:
[0008] Construct feature vectors for the grid corresponding to the candidate sites for wind farms;
[0009] The feature vector is input into a preset random forest model to obtain the predicted suitability probability value of the grid corresponding to the wind farm site selection candidate area for wind power development, and then the final candidate area is selected.
[0010] Based on the predicted suitability probability of the grid corresponding to the final candidate area, the theoretical maximum installed capacity of the final candidate area is obtained by using a preset constraint line equation that reflects the theoretical maximum installed capacity corresponding to the suitability probability.
[0011] Obtain the utilization coefficient of the land use type and slope grade combination corresponding to the final candidate area;
[0012] The theoretical maximum installed capacity of the final candidate area is multiplied by the utilization coefficient to obtain the installed capacity potential assessment result of the final candidate area.
[0013] In a specific embodiment of the present invention, the construction of the feature vector of the grid corresponding to the wind farm site selection candidate area includes:
[0014] 1) Collect multi-source raster and vector data of wind farm site selection candidate areas and unify the spatial coordinate system;
[0015] 2) Based on the results of step 1), calculate the characteristic indicators of the grid corresponding to the wind farm site selection candidate area, including:
[0016] 2-1) Calculate wind energy potential indicators, which include: wind power density indicators and wind resource stability indicators;
[0017] The formula for calculating the wind power density index is as follows:
[0018]
[0019] in, For wind power density, For air density, Hourly wind speed at a height of 100 meters;
[0020] The formula for calculating the wind resource stability index is as follows:
[0021]
[0022]
[0023] in, As an indicator of wind resource stability. An event consisting of wind speeds below 3 m / s or above 25 m / s. The number of hours in a year. The total number of events in a year with wind speeds below 3 m / s or above 25 m / s;
[0024] 2-2) Calculate socioeconomic indicators, including: GDP, distance from roads, and distance from power grids;
[0025] 2-3) Calculate the geographical environment indicators, which include: slope, elevation, and frequency of extreme temperatures;
[0026] 2-4) Calculate the disaster risk level indicators, which include: geological disaster index and meteorological disaster index;
[0027] 2-5) Calculate ecological security indicators, which include: proximity score, vegetation cover score, and biodiversity score. The specific steps are as follows:
[0028] 2-5-1) Calculate the proximity score, including:
[0029] The grids corresponding to ecological source areas and ecological corridors are identified. A maximum influence distance of 10 kilometers is set for either the ecological corridor or the ecological source area. A negative exponential function is used to calculate the proximity score between the sample's corresponding grid and the ecological source area and ecological corridor, as shown in the following expression:
[0030]
[0031] In the formula, For the calculated distance value, The attenuation coefficient;
[0032] 2-5-2) Vegetation coverage score;
[0033] The value of the raster corresponding to each sample in the vegetation cover raster data product is the vegetation cover score of that sample, and is directly used as the vegetation cover feature value of that sample.
[0034] 2-5-3) Calculate the biodiversity score;
[0035] The value of the raster corresponding to each sample in the spatial distribution data product of biodiversity index is the biodiversity score of that sample, and is directly used as the biodiversity feature value of that sample.
[0036] 3) Normalize the index values of each grid that are not in the range of 0-1 and convert them into feature values. Then, combine all the feature values of each grid to form the feature vector of that grid.
[0037] In a specific embodiment of the present invention, before inputting the feature vector into a preset random forest model, the method further includes:
[0038] Train the random forest model;
[0039] Training the random forest model includes:
[0040] 1) Construct a sample set consisting of positive samples and negative samples, wherein the positive samples are the grids corresponding to the built wind farms, the negative samples are the grids of the unbuilt wind farms, and the number of positive samples and negative samples is the same;
[0041] Set the label value corresponding to the positive sample to 1, indicating that it is suitable to build a wind farm; set the label value corresponding to the negative sample to 0, indicating that it is not suitable to build a wind farm.
[0042] 2) Collect multi-source raster data and vector data for each sample in the sample set, and unify the spatial coordinate system;
[0043] 3) Based on the results of step 2), calculate the feature index of the raster corresponding to each sample in the sample set;
[0044] 4) Normalize the indicator values of each sample's corresponding grid that are not in the range of 0-1 and convert them into feature values. Then, combine all the feature values of each sample's corresponding grid to form the feature vector of that sample.
[0045] 5) Construct a data sample set by combining the feature vectors and corresponding labels of each sample, and then randomly divide the data sample set into a training set and a test set according to a preset ratio;
[0046] 6) Construct a random forest model, train the model using the training set, and test the trained model using the test set to obtain the final random forest model.
[0047] In one specific embodiment of the present invention, it further includes:
[0048] Based on the random forest model outputting the corresponding grid of the wind farm site selection candidate area, a wind farm site selection suitability probability distribution grid map is generated.
[0049] In one specific embodiment of the present invention, it further includes:
[0050] The feature vectors of the grid corresponding to the existing wind farms are obtained and input into the random forest model. Then, the minimum value is selected from the suitability probability prediction values of the existing wind farms output by the random forest model as the candidate area screening threshold. If the suitability probability prediction value of any candidate wind farm site is not less than the threshold, then the candidate wind farm site is selected as the final candidate area.
[0051] In one specific embodiment of the present invention, it further includes:
[0052] Using the wind farm site suitability probability distribution grid map, a scatter plot reflecting the relationship between suitability probability and installed capacity density is generated based on the data of existing wind farms. The scatter plot uses the suitability probability of each existing wind farm as the horizontal axis and the measured installed capacity corresponding to the existing wind farm as the vertical axis.
[0053] In one specific embodiment of the present invention, it further includes:
[0054] In the scatter plot, the X-axis is divided into 30 equally spaced intervals; within each interval, the point corresponding to the 95th percentile of the installation density among all sample points is found; a smooth curve is used to connect all the 95th percentile points, and this curve is the constraint line.
[0055] In one specific embodiment of the present invention, it further includes:
[0056] The ratio of the measured installed capacity density of each existing wind farm to the theoretical maximum installed capacity density corresponding to the constraint line is calculated as the residual of that wind farm.
[0057] Each existing wind farm is grouped according to its corresponding land use type and slope level combination; for each land use type and slope combination, the average or median of the residuals of all wind farms under that combination is calculated as the utilization coefficient corresponding to that combination.
[0058] The features and beneficial effects of this invention are as follows:
[0059] This invention reveals the maximum possible installed capacity in a region under optimal conditions by establishing a marginal constraint model, and determines the installed capacity by adding other limiting factors. Specifically, by combining advanced machine learning techniques with constraint line models and utilizing various geographical and socio-economic factors, including wind energy resources, ecological security, geographical conditions, and socio-economic development, it provides a more accurate and efficient way to assess the installed capacity potential of wind farms, aiming to provide decision support for wind energy development.
[0060] A significant advantage of this invention lies in its use of a marginal constraint model. This model incorporates the impact of geographical factors such as land use and slope on installed capacity, taking into account the actual distribution of existing wind farms. Traditional assessment methods often neglect these factors, potentially leading to situations where the expected installed capacity cannot be achieved in areas rich in wind energy resources. By using the marginal constraint model, this invention can accurately assess the actual installed capacity of each region, comprehensively considering multiple factors such as topography and land use, thus avoiding the situation of "high suitability but low installed capacity." This improvement provides more scientific indicators for subsequent development and construction, ensuring the rationality of wind farm installed capacity and the feasibility of construction, thereby avoiding resource waste.
[0061] This invention utilizes an interpretable machine learning model, comprehensively considering multiple key indicators. While reducing subjective factors, it automatically processes large amounts of high-dimensional data and extracts potential nonlinear relationships, thereby improving the accuracy and consistency of the evaluation. Because machine learning models possess high transparency and interpretability, they not only provide decision-makers with objective analytical results but also quantify the importance of each evaluation indicator, further enhancing the transparency and reliability of the evaluation process. The application of this method significantly reduces the bias that may arise from human judgment, making the evaluation process more scientific and precise.
[0062] This invention also takes an ecological perspective, introducing a multi-constraint dynamic elimination mechanism and an ecological network intelligent avoidance algorithm to address the ecological impact of wind farm construction. Through these innovative methods, key ecological indicators are dynamically quantified and transformed into rigid decision parameters and spatial access rules, ensuring that wind power project construction does not negatively impact the ecological environment. Especially during the wind farm site selection process, the ecological network intelligent avoidance algorithm can automatically identify and avoid ecologically sensitive areas, preventing wind farm construction from damaging important biological habitats or ecological corridors. This mechanism ensures efficient wind power project construction while maximizing the protection of the ecological environment.
[0063] Finally, the "ecological-intelligent" fusion model proposed in this invention possesses excellent regional transferability and parameter reconstruction flexibility. Based on a unified data architecture and Python language modeling process, this invention can be quickly replicated and applied to wind farm layout optimization tasks under different terrains, ecological backgrounds, or policy constraints. This invention is applicable not only to national-scale wind power development planning but also to site selection for specific prefecture-level cities or concentrated wind power development zones, providing precise technical support for multi-level energy spatial layout. Attached Figure Description
[0064] Figure 1 This is an overall flowchart of a wind power installed capacity potential assessment method based on a marginal constraint model according to an embodiment of the present invention.
[0065] Figure 2 This is a wind farm site selection suitability probability distribution grid map generated using the method described in this embodiment of the present invention.
[0066] Figure 3 This is a schematic diagram of fitting constraint line equations to a scatter plot reflecting the relationship between suitability and installed capacity density of an existing wind farm in a specific embodiment of the present invention.
[0067] Figure 4 This is a wind power installed capacity density distribution map obtained by using the method described in this embodiment of the present invention. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of this application clearer, the application will be described in further detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.
[0069] Conversely, this application covers any alternatives, modifications, equivalent methods, and schemes made within the spirit and scope of this application as defined by the claims. Furthermore, to provide the public with a better understanding of this application, certain specific details are described in detail below. However, this application can be fully understood by those skilled in the art even without these detailed descriptions.
[0070] This invention proposes a method for evaluating wind power installed capacity potential based on a marginal constraint model. The overall process is as follows: Figure 1 As shown, it includes:
[0071] Construct feature vectors for the grid corresponding to the candidate sites for wind farms;
[0072] The feature vector is input into a preset random forest model to obtain the predicted suitability probability value of the grid corresponding to the wind farm site selection candidate area for wind power development, and then the final candidate area is selected.
[0073] Based on the predicted suitability probability of the grid corresponding to the final candidate area, the theoretical maximum installed capacity of the final candidate area is obtained by using a preset constraint line equation that reflects the theoretical maximum installed capacity corresponding to the suitability probability.
[0074] Obtain the utilization coefficient of the land use type and slope grade combination corresponding to the final candidate area;
[0075] The theoretical maximum installed capacity of the final candidate area is multiplied by the utilization coefficient to obtain the installed capacity potential assessment result of the final candidate area.
[0076] In a specific embodiment of the present invention, the wind power installed capacity potential assessment method based on a marginal constraint model includes the following steps:
[0077] 1) Training phase.
[0078] 1-1) Construct a sample set.
[0079] In this embodiment, the sample set is preferably constructed using the method of "measured locations at wind farms + spatially random negative samples" to ensure the representativeness and discriminative power of the model. The specific steps are as follows:
[0080] 1-1-1) Selection of positive samples.
[0081] In this embodiment, the precise geographic coordinates (latitude and longitude) of existing wind farms within the study area are obtained based on a global power plant database (such as the Global Power Plant Database) or publicly available data from national energy authorities. Each wind farm is mapped to a geographic raster cell with a preset resolution (e.g., 1km × 1km). If multiple wind farms fall into the same raster cell, they are considered as a single positive sample to avoid sample bias caused by spatial autocorrelation. A label value of 1 is assigned to this type of raster cell, indicating "suitable for establishing a wind farm".
[0082] 1-1-2) Negative sample selection.
[0083] In this embodiment, to ensure the model's discriminative ability, an equal number of negative samples are randomly generated across the entire area, equal to the number of positive samples. Negative samples must meet the following conditions: the grid corresponding to a negative sample does not overlap with any wind farm in the grid corresponding to a positive sample; the grid corresponding to a negative sample does not fall into a restricted area. In a specific embodiment of this invention, negative samples are selected using a random sampling and spatial filtering method with GeoPandas + rasterio in Python, ensuring uniform spatial distribution. A label value of 0 is assigned to this type of grid cell, indicating "not suitable for establishing a wind farm".
[0084] 1-1-3) Combine the positive samples obtained in step 1-1-1) and the negative samples obtained in step 1-1-2) to form a sample set.
[0085] 1-2) Based on the results of step 1-1), obtain and preprocess the multi-source geospatial data of each raster in the sample set; the specific steps are as follows:
[0086] 1-2-1) Collect multi-source raster data (such as digital elevation model data, wind speed distribution, land use data classification, built-up area boundaries) and vector data (such as bird habitats, nature reserves, potential geological hazard points and potential meteorological hazard points, etc.) corresponding to each grid in the sample set, and unify the spatial coordinate system.
[0087] 1-2-2) After converting the raster data and vector data collected in step 1-2-1) into spatial coordinate system 1, the raster data is resampled by bilinear interpolation, and the vector data is rasterized by assigning field values (e.g., nature reserve = 0, non-protection zone = 1). Finally, all data are unified to the same resolution.
[0088] In this embodiment, to improve data continuity, 3×3 mean interpolation is performed on missing pixels.
[0089] 1-3) Based on the results of step 1-2), calculate the feature index of the raster corresponding to each sample in the sample set.
[0090] In this embodiment, for each sample corresponding to a raster, based on the preprocessed raw data, the values of various indicators in the multidimensional suitability evaluation system are calculated, including traditional condition factors such as wind energy potential, socio-economic factors, geographical factors, and disaster risks, as well as ecological security indicators. The specific steps are as follows:
[0091] 1-3-1) Calculate wind energy potential indicators, which include: wind power density indicators and wind resource stability indicators.
[0092] The formula for calculating the wind power density index is as follows:
[0093]
[0094] in, For wind power density, For air density, The hourly wind speed is measured at a height of 100 meters.
[0095] When the wind speed is below 3 m / s or above 25 m / s, the wind turbine will not operate. In this embodiment, this situation is defined as a breach of stability. The calculation expression for the wind resource stability index is as follows:
[0096]
[0097]
[0098] in, As an indicator of wind resource stability. An event consisting of wind speeds below 3 m / s or above 25 m / s. The number of hours in a year is 8760. This represents the total number of events in a year where wind speeds are either below 3 m / s or above 25 m / s.
[0099] 1-3-2) Calculate socioeconomic indicators, including: GDP, distance from roads, and distance from power grids.
[0100] In this embodiment, the GDP spatial distribution dataset is used to extract the GDP index value of the corresponding raster for each sample.
[0101] Input the vector line data of the main roads (highways, national highways, and provincial highways) into the GIS, use the Euclidean distance tool in the GIS to calculate the straight-line distance from the center point of each grid to the nearest road, generate a distance grid map, and obtain the index value of the distance between the grid and the road for each sample.
[0102] Input the vector line data of the main power grid (power grid substations and transmission lines) into the GIS, use the Euclidean distance tool in the GIS to calculate the straight-line distance from the center point of each grid to the nearest line or power grid substation, generate a distance grid map, and obtain the index value of the distance of each sample grid from the power grid.
[0103] 1-3-3) Calculate the geographical environment indicators, which include: slope, elevation, and frequency of extreme temperatures.
[0104] In this embodiment, the digital elevation model raster data itself is the elevation value, which can be directly used as a feature value.
[0105] By using the slope tool in GIS, the terrain slope of each sample corresponding to the raster data of the digital elevation model can be directly calculated.
[0106] Calculation of frequency index for extreme temperature events:
[0107] In this embodiment, when calculating the frequency index of extreme temperature events, the input data consists of 40 consecutive years of daily maximum (TX) and daily minimum (TN) temperature raster data for each sample's corresponding raster. In a specific embodiment of this invention, a relative threshold method is used, that is, defining "extreme events" based on the climatological percentiles of long-term historical data. If the maximum temperature of a certain day exceeds the 95th percentile threshold of that raster point itself, then an extreme high-temperature event is defined as having occurred on that day. Similarly, calculating the 5th percentile (P5) of its multi-year daily minimum temperature time series defines an extreme low-temperature event as having occurred on that day. The number of days with extreme high-temperature and extreme low-temperature events occurring within a year for that raster is counted. Two preliminary raster layers are generated: an annual extreme high-temperature event frequency map and an annual extreme low-temperature event frequency map. Since extreme high temperatures and extreme low temperatures have different impact mechanisms on wind turbine operation, but both exhibit negative effects, the frequencies of extreme high-temperature events and extreme low-temperature events corresponding to each raster in these two frequency maps are summed to obtain the frequency of extreme temperature events occurring in that raster, which serves as the frequency index of extreme temperature events.
[0108] 1-3-4) Calculate the disaster risk level indicators, which include: geological disaster index and meteorological disaster index.
[0109] Calculation of the geological disaster index:
[0110] In this embodiment, vector point data of historical geological disaster points (including debris flows, landslides, and mudslides) are input, with each point representing the location of a historical disaster event. Kernel density analysis is performed using ArcGIS's Kernel Density feature, with a search radius set to 10 kilometers and a quadratic kernel function used. This generates a continuous raster layer covering the entire country, where each pixel represents the density of geological disaster points per unit area. Higher density values indicate a more unstable geological environment and a higher risk. The pixel value of the corresponding raster for each sample on the raster layer is obtained; this pixel value represents the geological disaster index for that sample.
[0111] Calculation of the meteorological disaster index:
[0112] In this embodiment, vector point data of historical meteorological disaster points (including floods and typhoons) are input, with each point representing the location of a historical disaster event. The calculation method is the same as that for the geological disaster index, but the search radius is larger, at 20 kilometers. A continuous raster layer is generated, where the value of each pixel represents the density of meteorological disaster points per unit area at that location. High-value areas indicate regions with frequent meteorological disasters or concentrated paths. The pixel value of the corresponding raster for each sample on the raster layer is obtained; this pixel value represents the meteorological disaster index for that sample.
[0113] 1-3-5) Calculate ecological security indicators.
[0114] In this embodiment, the core of calculating ecological security indicators lies in assessing the spatial relationship between the selected site and key elements of the ecological network (source areas, corridors), as well as its own ecological value (vegetation cover and biodiversity). The ecological security indicators include: proximity score, vegetation cover score, and biodiversity score. The specific steps are as follows:
[0115] 1-3-5-1) Calculate the proximity score. The specific steps are as follows:
[0116] 1-3-5-1-1) Determine the ecological source area.
[0117] In this embodiment, high-precision land cover data is used, with woodlands, grasslands, and water bodies designated as the "foreground," and urban areas, farmland, bare land, and other human activity areas designated as the "background." The morphological spatial pattern analysis module in Guidos Toolbox software is used to perform MSPA analysis on the foreground. An appropriate edge width is set (e.g., 100 meters based on species dispersal capacity) to identify core area patches with intact internal habitat conditions as candidate ecological source areas.
[0118] Farmland, construction land, and roads were set as threat sources with weights of 0.7, 1, and 0.6, respectively, and maximum stress distances of 4 km, 6 km, and 2 km, respectively. These were then input into the habitat quality module of the InVEST model to obtain a habitat quality index raster map (value range 0-1). The higher the habitat quality index value of the raster, the better the habitat quality of that raster.
[0119] Spatial overlay analysis (intersection) was performed between the core area identified by MSPA analysis and the top 20% of areas ranked by habitat quality index. Based on the minimum feasible patch theory, patches with an area greater than 20 km² were selected from the intersection and identified as the final ecological source areas. These patches possess both high habitat quality and core ecological functions.
[0120] 1-3-5-1-2) Determine ecological corridors.
[0121] Land use type, elevation, slope, and distance from road are selected as resistance factors. Each factor is assigned the same weight. The "weighted overlay" tool is used in GIS to overlay the raster of each resistance factor according to its weight, generating a continuous ecological resistance surface raster.
[0122] Input the ecological source areas (as source points) and ecological resistance surfaces determined in steps 1-3-5-1-1) into the "Build Corridors" module of the LinkageMapper toolkit to calculate the minimum cost path for outward expansion from each source area. These paths are the potential optimal channels for species migration or ecological flow, which are ecological corridors.
[0123] 1-3-5-1-3) Calculate the proximity score.
[0124] This embodiment sets the maximum impact distance to 10 kilometers from ecological corridors or ecological source areas based on ecological criteria (such as species activity range and edge effect range). Beyond this distance, no further impact occurs. Because the negative exponential decay function decays rapidly and more sensitively reflects the high risk at close range, it is used to calculate the original proximity score between the sample's corresponding raster and the ecological source area and ecological corridor. The closer the distance, the greater the harm. The calculation expression is as follows:
[0125]
[0126] In the formula, For the calculated distance value, The attenuation coefficient is set to 0.2.
[0127] 1-3-5-2) Vegetation coverage score.
[0128] This embodiment utilizes vegetation cover raster data products, where the value of each sample's corresponding raster (typically 0-100% or 0-1) is the vegetation cover score for that sample, which can be directly used as the vegetation cover feature value for that sample. The higher the vegetation cover feature value, the better the vegetation cover status of the corresponding raster, and the higher its ecological conservation value.
[0129] 1-3-5-3) Calculate the biodiversity score.
[0130] In this embodiment, existing biodiversity index spatial distribution data products are directly used. The value of each sample's corresponding raster in the biodiversity index spatial distribution data product is the biodiversity score of that sample, which can be directly used as the biodiversity characteristic value of that sample. The higher the biodiversity characteristic value, the higher the species richness and evenness of the corresponding raster, and the higher the stability and conservation value of the ecosystem.
[0131] 1-4) Based on the results of steps 1-3), normalize the index values of each grid that are not in the range of 0-1 and convert them into feature values. Then, combine all the feature index values of each grid to form the feature vector of that grid.
[0132] In this embodiment, to prevent features with excessively large numerical ranges from dominating model training, all continuous features are subjected to min-max normalization, which linearly transforms them to the [0,1] interval, as shown in the following expression:
[0133]
[0134] Among them, the vegetation coverage score and biodiversity score have values between 0 and 1, so no further normalization is required.
[0135] 1-5) Based on the results of steps 1-4), the feature vectors and corresponding labels of the grids corresponding to each sample are used to form a data sample set. Then, the data sample set is randomly divided into a training set and a test set according to a preset ratio (7:3 in a specific embodiment of the present invention), which are used for model parameter training and independent performance evaluation, respectively.
[0136] 1-4) Train a random forest model using the training set obtained in steps 1-3), and validate it using the validation set; the specific steps are as follows:
[0137] 1-4-1) Initialize the random forest model.
[0138] Set the initial hyperparameters of the random forest model, including: the number of decision trees, which is set to 200 in this embodiment; the maximum depth of the trees, which can be set to 10 in this embodiment, or set to None to allow the trees to grow fully; the minimum number of samples required for internal nodes to be further divided, which can be set to 2 in the initial value; and the minimum number of samples required for leaf nodes, which can be set to 1 in the initial value.
[0139] 1-4-2) Use the training set obtained in step 1-3) to train the random forest model.
[0140] After initializing the random forest model, it is essential to perform systematic optimization and rigorous validation to ensure excellent generalization ability, stability, and reliability, and to avoid overfitting or underfitting. This step includes hyperparameter optimization, performance validation, and robustness analysis.
[0141] Define the hyperparameter search space: Based on the principles of the random forest algorithm and preliminary experiments, determine the key hyperparameters that have the greatest impact on model performance and their candidate value ranges to be searched. The search space defined in this embodiment is as follows: Number of decision trees: [100, 200, 300, 400]; Maximum tree depth: [5, 10, 15, 20, None] (None represents no depth restriction); Minimum number of samples required for internal node re-split: [2, 5, 10]; Minimum number of samples for leaf nodes: [1, 2, 4]; Number of features considered when finding the best split: [sqrt, log2, 0.5, 0.7].
[0142] Grid search cross-validation is performed: K-fold cross-validation (K=10 in this embodiment) combined with grid search is used for hyperparameter optimization. The specific process is as follows: The training set (70% of the total set) is randomly divided into 10 mutually exclusive subsets (folds) of roughly equal size. For each hyperparameter combination in the search space, 10 rounds of training and validation are performed sequentially: In each round, 9 folds of data from the training set are used as a sub-training set for model training. The remaining 1 folds of data are used as a sub-validation set, and the model's performance score on this validation set is calculated (AUC is used as the main evaluation metric in this embodiment). The average AUC score obtained from 10 rounds of validation is calculated for each hyperparameter combination, and this average cross-validation performance is used as the cross-validation average performance of that combination. After traversing all hyperparameter combinations, the hyperparameter combination with the highest average cross-validation performance is selected as the optimal configuration.
[0143] Finally, using the entire training set (70% of the data) and this optimal hyperparameter combination, the random forest model is retrained to obtain the trained random forest model.
[0144] 1-4-3) Use the validation set obtained in step 1-3) to test the random forest model trained in step 1-4-2).
[0145] In this embodiment, the trained model is evaluated using a test set. The evaluation is primarily based on AUC (Area Under the Curve); the closer the AUC value is to 1, the stronger the model's classification ability. This embodiment requires the final model's AUC value to be no less than 0.85.
[0146] It should be noted that in this embodiment, during the hyperparameter search space process, the range of parameters will be continuously changed to ensure that the AUC value is not lower than 0.85; if changing the range still fails to make the model pass the validation set test, the candidate range of hyperparameters such as the number of decision trees and the maximum depth of the trees will be further expanded.
[0147] 2) Evaluation phase.
[0148] 2-1) Obtain and preprocess multi-source geospatial data of wind farm site selection candidate areas. The specific steps are as follows:
[0149] 2-1-1) Selecting candidate sites for wind farms.
[0150] Wind farm site selection involves multiple dimensions of factors, including economic, technological, natural, social, and environmental considerations. In this embodiment, the candidate wind farm site refers to the areas identified through preliminary screening during the macro-assessment and preliminary planning stage of wind energy resources, demonstrating potential development conditions. These areas initially meet the natural, socio-economic, and technological requirements for wind farm construction and require further detailed evaluation to determine the optimal site.
[0151] In some embodiments, candidate wind farm sites are determined based on restrictive principles. Specifically, this involves comprehensively considering geographical location, economic factors, and environmental influences to determine exclusion criteria for wind farm sites, identifying unsuitable sites, and deducting these unsuitable sites from the possible sites within the study area to obtain candidate wind farm sites. This embodiment, in accordance with national and local policies and technical limitations on wind power development, eliminates areas lacking the conditions for wind power construction, specifically including three types of restrictions:
[0152] Policy restrictions include: integrating national and local ecological protection policies, land regulations and special plans, and clearly defining the areas where wind power development is prohibited, including: permanent basic farmland, ecological protection red lines, nature reserves, forest parks, wetland parks, geological parks, scenic spots, natural heritage sites, national parks and other protected areas.
[0153] Natural limitations include: areas with an average annual wind speed of < 4 m / s; areas with an altitude of > 4000 meters; and areas with a terrain slope of > 30°.
[0154] Social impact restrictions include a 1km buffer zone around built-up areas, industrial and mining areas, schools, hospitals, and some densely populated areas or sensitive facilities.
[0155] If any of the three restrictions mentioned above are met, the area is defined as a terrain-unsuitable area and excluded from the initial candidate areas. Layer overlay technology of Geographic Information System (GIS) is used to generate candidate wind power sites.
[0156] 2-1-2) Collect multi-source raster data (such as digital elevation model data, wind speed distribution, land use data classification, and built-up area boundaries) and vector data (such as bird habitats, nature reserves, potential geological hazard points and potential meteorological hazard points, etc.) of the wind farm site selection candidate area obtained in step 2-1-1), and unify the spatial coordinate system.
[0157] 2-1-3) After converting the raster data and vector data collected in step 2-1-2) into spatial coordinate system one, the raster data is resampled by bilinear interpolation, and the vector data is rasterized by assigning field values (e.g., nature reserve = 0, non-protection zone = 1). Finally, all data are unified to the same resolution.
[0158] 2-2) Based on the results of step 2-1), repeat steps 1-3-1) to 1-3-5) to calculate the characteristic index of the grid corresponding to the wind farm site selection candidate area.
[0159] 2-3) Normalize the feature indicators of the grid corresponding to the wind farm site selection candidate area that are not in the range of 0-1. After normalization, form the feature vector of the grid by combining all the feature indicators of the grid corresponding to the wind farm site selection candidate area.
[0160] 2-4) Input the feature vector obtained in step 2-3) into the random forest model obtained in step 1). The random forest model outputs the probability prediction value of the suitability of the grid corresponding to the wind farm site selection candidate area for wind power development. The higher the probability prediction value, the more suitable the grid is for wind power development.
[0161] Furthermore, in this embodiment, based on the model output, a spatially continuous raster map of the wind farm site suitability probability distribution can be generated. This raster map has the same spatial reference and resolution as the input data. Its pixel value represents the predicted probability, with a value range of [0, 1].
[0162] 2-4) Based on the results of step 2-3), select the final candidate area from the wind farm site selection candidate area.
[0163] In this embodiment, the minimum value of the suitability probability prediction value of the grid corresponding to the existing wind farm in the study area output by the random forest model is used as the candidate area screening threshold. Then, based on the suitability probability prediction value of the grid corresponding to the candidate wind farm site obtained in steps 2-3), candidate areas with probability prediction values not lower than the threshold are selected as the final candidate wind farm site and included in the installed capacity assessment scope.
[0164] 2-5) Based on the results of step 2-4), the wind power installation potential of the final candidate areas is assessed.
[0165] This step aims to transform the suitability probability of wind farm site selection into a specific spatial distribution plan for installed capacity that can guide planning and construction. By introducing constraint line models and on-site correction coefficients, it achieves a precise assessment from "where is it suitable to build" to "how much can be built." The specific steps are as follows:
[0166] 2-5-1) Data collection and processing of existing wind farms.
[0167] Collect spatial boundaries (polygon vector data), total installed capacity (unit: MW), and site area (unit: km²) of all existing wind farms within the study area.
[0168] For each existing wind farm, calculate its measured installed capacity density (unit: MW / km²): Installed capacity density = Total installed capacity divided by the area of the farm area.
[0169] Based on the vector data of existing wind farms, the polygons corresponding to each existing wind farm (the polygons do not need to be generated manually; instead, the vector data (in shp format) of the existing wind farms is directly obtained, which already contains the spatial boundary information of the actual construction area of the wind farm and can be directly used for subsequent spatial overlay analysis with the suitability probability map) are spatially overlaid with the wind farm site selection suitability probability distribution raster map generated in steps 2-3) to extract the average suitability probability within the area of each existing wind farm. Thus, the suitability probability and measured installed capacity density corresponding to each existing wind farm are obtained.
[0170] 2-5-2) Based on the results of step 2-5-1), a scatter plot reflecting the relationship between suitability probability and installed capacity is drawn, and constraint line fitting is performed. The specific steps are as follows:
[0171] In this embodiment, based on the results of step 2-5-1), a scatter plot reflecting the relationship between suitability and installed capacity density is drawn. The scatter plot uses the suitability probability of each existing wind farm as the horizontal axis (X-axis) and the corresponding measured installed capacity as the vertical axis (Y-axis).
[0172] Then, the X-axis (suitability probability, ranging from 0 to 1) is divided into 30 equally spaced intervals. Within each interval, the 95th quantile of the installed capacity (Y-value) among all sample points is identified. These points represent the higher installed capacity actually achieved at the corresponding suitability level. All 95th quantile points are connected using a smooth curve (such as a quadratic polynomial curve, exponential curve, or spline curve). This fitted curve is the constraint line, which defines the theoretical maximum installed capacity corresponding to a specific suitability probability under ideal conditions.
[0173] 2-5-3) Based on the results of step 2-5-2), calculate the ratio of the measured installed capacity density of each existing wind farm to the theoretical maximum installed capacity density corresponding to the constraint line as the residual of the wind farm.
[0174] 2-5-4) Group each existing wind farm according to its combination of land use type (bare land, grassland, woodland, cultivated land) and slope grade (e.g., <3°, 3°-6°, 6°-15°, >15°). For each combination of land use type and slope grade, calculate the average or median of the residuals of all wind farms under that combination as the utilization factor corresponding to that combination.
[0175] 2-5-5) Calculate the installed capacity potential of the final candidate areas. The specific steps are as follows:
[0176] 2-5-5-1) Based on the suitability probability value (Ps) of the final candidate area, and combined with the constraint line equation, calculate the theoretical maximum installed capacity (Dmax) of the final candidate area.
[0177] 2-5-5-2) Obtain the combination of land use type and slope grade for the grid corresponding to the final candidate area. Based on the utilization coefficient corresponding to this combination obtained in step 2-5-4), calculate the actual developable installed capacity density of the grid:
[0178] Dactual = Dmax × UC
[0179] Where UC represents the utilization coefficient.
[0180] Furthermore, the method described in this embodiment also includes:
[0181] The calculated Dactual value is assigned to each grid cell to generate a spatially continuous wind farm installed capacity distribution map.
[0182] The method described in this embodiment will be further explained below with reference to a specific example.
[0183] This embodiment uses the core coverage area of the North China Power Grid as an example for verification. This region is rich in wind energy resources (average annual wind speed greater than 5 m / s) and includes diverse land use types such as farmland, towns, and ecological protection zones. The terrain is mainly plains and hills, with some mountainous areas, making it a typical area for testing the comprehensive capabilities of the method described in this embodiment. All calculation processes were completed on a workstation equipped with an Intel Core i7-12700K processor and 64GB of memory, based on the Python 3.9 platform, mainly utilizing libraries such as scikit-learn, shap, geopandas, and rasterio.
[0184] In this embodiment, multi-source data of the target area is collected, including: 100-meter resolution DEM data, meteorological reanalysis wind speed data, MODIS land use data, GDP spatialized kilometer grid data, road and power grid vector data, nature reserve and ecological protection red line boundaries, and historical geological and meteorological disaster point data. All raster and vector data are unified to the CGCS2000 coordinate system and resampled to a 1km×1km resolution to form a standardized basic database.
[0185] In this embodiment, the following unsuitable areas are eliminated in sequence: Policy-restricted areas: permanent basic farmland, ecological protection red lines, and national nature reserves. Natural-restricted areas: areas with an average annual wind speed <4m / s, areas with an altitude >4000m, and areas with a slope >30°. Socially-restricted areas: built-up areas and a 1km buffer zone. After the above eliminations, approximately 85% of the total study area is obtained as the scope for subsequent refined evaluation of wind farm site selection.
[0186] During the training phase, based on a global power plant database, 127 existing wind farm sites within the region were identified as positive samples (label=1). 127 non-wind farm sites were randomly generated within the candidate area as negative samples (label=0) to ensure uniform spatial distribution. These were then randomly divided into training and test sets in a 7:3 ratio.
[0187] The random forest model was trained using the training set, and the hyperparameters were optimized using grid search and 10-fold cross-validation. The optimal parameter combination was obtained as follows: n_estimators=300, max_depth=15, min_samples_leaf=2.
[0188] The model performance was validated on an independent test set. The results showed that the model accuracy was 83.5% and the AUC value reached 0.89, indicating that the model has excellent classification and discrimination ability and generalization ability.
[0189] During the evaluation phase, a raster map of the probability distribution of wind farm site suitability is generated using the trained model, such as... Figure 2 As shown, it represents the suitability values for different regions. The results show that highly suitable areas (probability > 0.7) account for approximately 25% of the candidate area, mainly distributed in regions with superior wind resources, gentle slopes, and that avoid core ecological areas.
[0190] In this embodiment, SHAP (SHapley Additive exPlanations) is used for model interpretation. Global importance analysis shows that the top five key factors and their average |SHAP| values are shown in Table 1.
[0191] Table 1. Top five key factors and their average |SHAP| values in a specific embodiment of the present invention.
[0192] Ranking Variable name Average SHAP contribution 1 Distance from the ecological corridor 0.181 2 vegetation coverage 0.153 3 Geological disaster risk 0.144 4 Wind power density 0.083 5 GDP 0.065
[0193] Based on the suitability probability of existing wind farms in the region and the measured installed capacity, this embodiment obtains the constraint line equation through fitting. Figure 3 This is a schematic diagram of fitting constraint line equations to a scatter plot reflecting the relationship between suitability and installed capacity density of an existing wind farm in this embodiment. The rhombuses represent boundary points above the 95th percentile, and the remaining points represent points below the 95th percentile. The curve is generated by fitting these points, and the circular points represent the installed capacity under different land use types and restrictions.
[0194] In this embodiment, the utilization coefficients for different land types and slope combinations are shown in Table 2:
[0195] Table 2. Utilization coefficients for different land use types and slope combinations in a specific embodiment of the present invention.
[0196] bare land woodland arable land grassland <3% 0.967 0.231 0.289 0.787 3%≤slope<6% 0.482 0.122 0.153 0.392 6%≤slope<15% 0.271 0.072 0.084 0.235 15%≤slope 0.198 0.041 0.053 0.132
[0197] By analyzing the utilization coefficients of different land types and slope combinations, this embodiment corrects the theoretical installed capacity density, ultimately generating a wind power installed capacity density distribution map for the region, as shown below. Figure 4 As shown, where Figure 4 The grid value is the installed capacity density, and the average installed capacity density in this area is 2.73 W / m². 2 . Figure 4 It clearly shows the specific developable capacity per unit area in different regions, providing direct and quantitative spatial decision support for the energy sector to plan wind power projects.
Claims
1. A method for evaluating the installed capacity potential of wind power based on a marginal constraint model, characterized in that, include: Construct feature vectors for the grid corresponding to the candidate sites for wind farms; The feature vector is input into a preset random forest model to obtain the predicted suitability probability value of the grid corresponding to the wind farm site selection candidate area for wind power development, and then the final candidate area is selected. Based on the predicted suitability probability of the grid corresponding to the final candidate area, the theoretical maximum installed capacity of the final candidate area is obtained by using a preset constraint line equation that reflects the theoretical maximum installed capacity corresponding to the suitability probability. Obtain the utilization coefficient of the land use type and slope grade combination corresponding to the final candidate area; The theoretical maximum installed capacity of the final candidate area is multiplied by the utilization coefficient to obtain the installed capacity potential assessment result of the final candidate area.
2. The method according to claim 1, characterized in that, The feature vector for constructing the grid corresponding to the candidate wind farm site includes: 1) Collect multi-source raster and vector data of wind farm site selection candidate areas and unify the spatial coordinate system; 2) Based on the results of step 1), calculate the characteristic indicators of the grid corresponding to the wind farm site selection candidate area, including: 2-1) Calculate wind energy potential indicators, which include: wind power density indicators and wind resource stability indicators; The formula for calculating the wind power density index is as follows: in, For wind power density, For air density, Hourly wind speed at a height of 100 meters; The formula for calculating the wind resource stability index is as follows: in, As an indicator of wind resource stability. An event consisting of wind speeds below 3 m / s or above 25 m / s. The number of hours in a year. The total number of events in a year with wind speeds below 3 m / s or above 25 m / s; 2-2) Calculate socioeconomic indicators, including: GDP, distance from roads, and distance from power grids; 2-3) Calculate the geographical environment indicators, which include: slope, elevation, and frequency of extreme temperatures; 2-4) Calculate the disaster risk level indicators, which include: geological disaster index and meteorological disaster index; 2-5) Calculate ecological security indicators, which include: proximity score, vegetation cover score, and biodiversity score. The specific steps are as follows: 2-5-1) Calculate the proximity score, including: The grids corresponding to ecological source areas and ecological corridors are identified. A maximum influence distance of 10 kilometers is set for either the ecological corridor or the ecological source area. A negative exponential function is used to calculate the proximity score between the sample's corresponding grid and the ecological source area and ecological corridor, as shown in the following expression: In the formula, For the calculated distance value, The attenuation coefficient; 2-5-2) Vegetation coverage score; The value of the raster corresponding to each sample in the vegetation cover raster data product is the vegetation cover score of that sample, and is directly used as the vegetation cover feature value of that sample. 2-5-3) Calculate the biodiversity score; The value of the raster corresponding to each sample in the spatial distribution data product of biodiversity index is the biodiversity score of that sample, and is directly used as the biodiversity feature value of that sample. 3) Normalize the index values of each grid that are not in the range of 0-1 and convert them into feature values. Then, combine all the feature values of each grid to form the feature vector of that grid.
3. The method according to claim 2, characterized in that, Before inputting the feature vector into a preset random forest model, the method further includes: Train the random forest model; Training the random forest model includes: 1) Construct a sample set consisting of positive samples and negative samples, wherein the positive samples are the grids corresponding to the built wind farms, the negative samples are the grids of the unbuilt wind farms, and the number of positive samples and negative samples is the same; Set the label value corresponding to the positive sample to 1, indicating that it is suitable to build a wind farm; set the label value corresponding to the negative sample to 0, indicating that it is not suitable to build a wind farm. 2) Collect multi-source raster data and vector data for each sample in the sample set, and unify the spatial coordinate system; 3) Based on the results of step 2), calculate the feature index of the raster corresponding to each sample in the sample set; 4) Normalize the indicator values of each sample's corresponding grid that are not in the range of 0-1 and convert them into feature values. Then, combine all the feature values of each sample's corresponding grid to form the feature vector of that sample. 5) Construct a data sample set by combining the feature vectors and corresponding labels of each sample, and then randomly divide the data sample set into a training set and a test set according to a preset ratio; 6) Construct a random forest model, train the model using the training set, and test the trained model using the test set to obtain the final random forest model.
4. The method according to claim 3, characterized in that, Also includes: Based on the random forest model outputting the corresponding grid of the wind farm site selection candidate area, a wind farm site selection suitability probability distribution grid map is generated.
5. The method according to claim 4, characterized in that, Also includes: The feature vectors of the grid corresponding to the existing wind farms are obtained and input into the random forest model. Then, the minimum value is selected from the suitability probability prediction values of the existing wind farms output by the random forest model as the candidate area screening threshold. If the suitability probability prediction value of any candidate wind farm site is not less than the threshold, then the candidate wind farm site is selected as the final candidate area.
6. The method according to claim 5, characterized in that, Also includes: Using the wind farm site suitability probability distribution grid map, a scatter plot reflecting the relationship between suitability probability and installed capacity density is generated based on the data of existing wind farms. The scatter plot uses the suitability probability of each existing wind farm as the horizontal axis and the measured installed capacity corresponding to the existing wind farm as the vertical axis.
7. The method according to claim 6, characterized in that, Also includes: In the scatter plot, the X-axis is divided into 30 equally spaced intervals; Within each interval, find the point corresponding to the 95th percentile of the installation density among all sample points; connect all 95th percentile points with a smooth curve, which is the constraint line.
8. The method according to claim 7, characterized in that, Also includes: The ratio of the measured installed capacity density of each existing wind farm to the theoretical maximum installed capacity density corresponding to the constraint line is calculated as the residual of that wind farm. Each existing wind farm is grouped according to its corresponding land use type and slope level combination; for each land use type and slope combination, the average or median of the residuals of all wind farms under that combination is calculated as the utilization coefficient corresponding to that combination.
Citation Information
Patent Citations
High-resolution land wind power generation resource evaluation method and system
CN115841265A
Wind-solar power station optimization site selection method and system based on multi-dimensional evaluation
CN118134024A
Extreme low water change attribution and low water resource potential assessment method
CN119377915A