A fast general selection method, device and storage medium for pumped-storage power station sites based on hierarchical clustering and parallel computing
The method of hierarchical clustering and parallel computing optimizes the site selection process for pumped storage power stations, improving efficiency and reducing errors in site identification.
Patent Information
- Application Number
- CN202510355165.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In the prior art, the site selection efficiency of pumped storage power stations is not high, the analysis time is too long, and the site selection efficiency needs to be optimized.
Using a method based on hierarchical clustering and parallel computing, the digital elevation model data of the target point selection area is obtained, preprocessing and hierarchical clustering is performed, and multi-process parallel calculations are used to meet the reservoir target constraints, and candidate points that meet the conditions are quickly selected.
It improves the efficiency of point selection, reduces the probability of wrong selection and missed selection, and provides a scientific basis for the planning of pumped storage power stations.
Smart Images

Figure CN119863043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pumped storage planning and site selection, and in particular to a rapid general site selection method, device and storage medium for pumped storage power stations based on hierarchical clustering and parallel computing. Background Art
[0002] As an integral part of a new power system dominated by new energy, pumped storage plays a crucial role in the development of new energy. The site selection of power stations is an important preliminary task for pumped storage projects. Traditional power station site selection mainly relies on manual efforts, consuming a large amount of time and energy for on-site surveys and data analysis. This site selection method not only involves a large workload, wastes manpower and material resources, but also has a relatively limited analysis scope, making it difficult to meet the overall planning requirements of pumped storage power stations in a large-scale area.
[0003] With the continuous in-depth research on the site selection of pumped storage power stations by domestic and foreign scholars, various studies on site selection efficiency, site selection scientificity, etc. have emerged in an endless stream. These research directions are mainly manifested as: the site selection schemes are becoming more and more multi-dimensional, the evaluation indicators are becoming more abundant, and the technologies adopted are becoming more diverse. With the gradual improvement of digital site selection schemes, more and more pumped storage projects have started to change their power station site selection methods. However, in actual application scenarios, the site selection range of pumped storage power stations is relatively wide. Although the digital site selection method improves efficiency compared with traditional manual site selection, there is still a problem of too long analysis time in actual applications, and the site selection efficiency needs to be continuously optimized. Therefore, in view of the problem of low site selection efficiency in existing research in actual applications, the present invention proposes a rapid general site selection method for pumped storage power station sites based on hierarchical clustering and parallel computing.
[0004] In view of this, it is necessary to propose a rapid general site selection method, device and storage medium for pumped storage power station sites based on hierarchical clustering and parallel computing to solve or at least alleviate the above defects. Summary of the Invention
[0005] The main object of the present invention is to provide a rapid general site selection method, device and storage medium for pumped storage power station sites based on hierarchical clustering and parallel computing, so as to solve the technical problems of low site selection efficiency, too long analysis time and the need for continuous optimization of site selection efficiency in the prior art for pumped storage power stations.
[0006] To achieve the above object, the present invention provides a rapid general site selection method for pumped storage power station sites based on hierarchical clustering and parallel computing, including the following steps:
[0007] S1, obtaining digital elevation model data of a target site selection area, and preprocessing the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set;
[0008] S2. Perform hierarchical clustering on the candidate point information to obtain a candidate point classification result data set;
[0009] S3. Perform multi-process parallel computing on the candidate point classification result data set according to the pre-established reservoir target constraints to obtain a final candidate point set that meets the reservoir target constraints.
[0010] Preferably, step S1 specifically includes the following steps:
[0011] S11. Extract contour data according to the preset contour interval and the digital elevation model data, and extract river network data according to the water system flow threshold and the digital elevation model data;
[0012] S12. Obtain the river network line segments in the river network data, traverse all the river network line segments, and then extract the midpoints of each river network line segment and put them into the candidate point data set P; where, , represents the i-th candidate point, and n represents the number of candidate points;
[0013] S13. Calculate the elevation of each candidate point in the candidate point data set P, and collect the elevations of all candidate points to obtain a candidate point elevation data set.
[0014] Preferably, step S2 specifically includes the following steps:
[0015] S21. Traverse the candidate point data set , calculate the distances between each candidate point in the candidate point data set to obtain a distance matrix E; where, the expression of the distance matrix E is:
[0016] ;
[0017] Where, ; represents the distance between the i-th cluster and the j-th cluster, is the Euclidean distance, when i = j, , each candidate point in the candidate point data set is a separate cluster;
[0018] S22. Find the two closest clusters according to the distance matrix E. If the distance between the two closest clusters is less than the distance threshold, then merge these two clusters, otherwise do not merge;
[0019] S23. Update the distance matrix E;
[0020] S24. Repeat steps S22 - S23 until the distances between all clusters are greater than or equal to the distance threshold, and obtain the candidate point classification result data set ; wherein, , represents the j-th cluster, , represents the t-th candidate point in, m represents the number of clusters, and z represents the number of candidate points in the cluster.
[0021] Preferably, the reservoir target constraint condition in step S3 is:
[0022]
[0023] wherein, is the dam length constraint, is the dam height constraint, is the reservoir capacity constraint, is the reservoir dam line length of the i-th candidate point, is the reservoir dam height of the i-th candidate point, is the reservoir capacity of the i-th candidate point, is the maximum reservoir dam length allowed by the project, is the maximum reservoir dam height allowed by the project, is the minimum reservoir capacity allowed by the project.
[0024] Preferably, step S3 specifically includes the following steps:
[0025] S31, start multiple processes based on the preset number of sub-processes U ; wherein, , represents the t-th sub-process;
[0026] S32, traverse the candidate point classification result data set , and sequentially assign the clusters in the candidate point classification result data set to idle sub-processes for multi-process parallel calculation to obtain a final candidate point set that meets the reservoir target constraint conditions; wherein, one cluster is one task. When the number of tasks is greater than the number of sub-processes U, all sub-processes are in a working state, and unassigned tasks need to wait until the sub-process status becomes idle before being assigned.
[0027] Preferably, in step S32, the clusters in the candidate point classification result data set are sequentially assigned to idle sub-processes for multi-process parallel calculation to obtain a final candidate point set that meets the reservoir target constraint conditions; wherein, for the sub-process in , the following steps are performed:
[0028] S321, traverse the set , and according to the candidate point elevation , the maximum dam height of the reservoir , screen the contour data by elevation to obtain the screened contour set , where represents the f-th contour, g represents the number of contours in the set, represents the elevation of satisfying ;
[0029] S322, obtain the planar coordinates of the candidate point and obtain the river network segment where the candidate point is located, and calculate the straight line passing through the candidate point and intersecting the river network segment;
[0030] S323, obtain the required line segment located on the straight line and solve the coordinates of the two endpoints of the required line segment ;
[0031] S324, based on the preset rotation angle , calculate the total number of rotations c ;
[0032] S325, traverse c times, rotate the required line segment at different angles to determine whether there is a reservoir dam line that meets the reservoir target constraint conditions for the currently calculated candidate point; specifically, it includes the following steps:
[0033] S3251, calculate the slope of the straight line after the u-th rotation angle, and solve the coordinates of the two endpoints of the required line segment at this angle;
[0034] S3252, generate a line feature layer according to the coordinates of the two endpoints, and use the intersect analysis tool of arcpy to extract the intersection points with the contours to obtain the reservoir dam line; among them, in the intersection layer attribute table, filter out the geometric elements of the single-point type and retain the geometric elements of the multi-point type;
[0035] S3253, calculate the lengths of each reservoir dam line, and obtain the set of dam lines whose reservoir dam line lengths meet the dam length constraint;
[0036] S3254, sort the reservoir dam lines according to the elevations of the endpoints of the reservoir dam line from high to low, select the dam line B with the maximum elevation, and according to the dam line B with the maximum elevation and the contour where its endpoints are located, obtain the closed polygon formed by the dam line B with the maximum elevation and the contour ;
[0037] S3255, determine the reservoir capacity according to the closed polygon and digital elevation model data ;
[0038] S3256, if the reservoir capacity , then return to step S325 to perform the next rotation angle; otherwise, there is a reservoir dam line that meets the reservoir target constraint conditions for the currently calculated candidate point, traverse end, put the candidate point into the final candidate point set P'', and the subprocess becomes idle.
[0039] Preferably, the equation expression of the straight line in step S322 is specifically:
[0040] When k exists, ;
[0041] When k does not exist, ; where k is the slope, y is the dependent variable, and x is the independent variable.
[0042] Preferably, step S323 specifically includes the following steps:
[0043] Set the length of the required line segment to 2 , and the solution process of the coordinates of the two endpoints is as follows:
[0044] When k exists, use the distance formula between two points to solve the coordinates of the two endpoints of the line segment, that is:
[0045]
[0046] The coordinates of the two endpoints of the line segment are obtained through the derived formula ( , ), ( , ) as:
[0047] ;
[0048] When k does not exist, the coordinates of the two endpoints of the line segment are respectively ( , ), ( , ).
[0049] The present invention also provides a rapid general selection device for pumped storage power station sites based on hierarchical clustering and parallel computing, including:
[0050] A preprocessing unit for obtaining digital elevation model data of a target point selection area and preprocessing the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set;
[0051] A hierarchical clustering processing unit for performing hierarchical clustering processing on the candidate point information to obtain a candidate point classification result data set;
[0052] A parallel computing unit for performing multi-process parallel computing on the candidate point classification result data set according to pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions.
[0053] The present invention also provides a storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned rapid general selection method for pumped storage power station sites based on hierarchical clustering and parallel computing are implemented.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The present invention provides a rapid general selection method, device and storage medium for pumped storage power station sites based on hierarchical clustering and parallel computing. By obtaining digital elevation model data of a target point selection area and preprocessing the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set, performing hierarchical clustering processing on the candidate point information to obtain a candidate point classification result data set, and performing multi-process parallel computing on the candidate point classification result data set according to pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions.
[0056] The present invention considers spatial autocorrelation and point selection efficiency, clusters the midpoints of river network line segments based on the hierarchical clustering method, and uses the multi-process parallel technology to regard each cluster as a point search task and allocate it to idle sub-processes. Through the preset reservoir target constraint conditions, power stations that meet the conditions are rapidly generally selected. Compared with the traditional manual point selection, the present invention, as a digital general selection power station method, not only improves the efficiency of point selection in the actual application scenario, but also reduces the probability of misselection and omission, providing a scientific basis for the planning of pumped storage power stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on the structures shown in these drawings without creative efforts.
[0058] Figure 1 It is a schematic flowchart of an embodiment of the present invention;
[0059] Figure 2 It is a schematic diagram of the station site general election result of the target site selection area in an embodiment of the present invention.
[0060] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0061] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0063] In the present invention, the descriptions involving "right part", "middle part", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "right part", "middle part" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0064] Please refer to the attached Figures 1 to 2 A fast general election method for pumped storage power station sites based on hierarchical clustering and parallel computing in an embodiment provided by the present invention includes the following steps:
[0065] S1. Obtain the digital elevation model data of the target site selection area, and preprocess the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set;
[0066] Specifically, the digital elevation model data is DEM (Digital Elevation Model) data, which is a digital expression of the terrain surface morphology and contains a large amount of rich geographical information. The DEM data plays an important role in digital site selection. Preferably, the arcpy site package is used to call the ArcGIS analysis tool to preprocess the DEM data to obtain candidate point information.
[0067] S2. Perform hierarchical clustering on the candidate point information to obtain a candidate point classification result dataset;
[0068] It should be noted that in this application, candidate points are classified based on the hierarchical clustering method. Starting from the law that closer things are more relevant proposed by the first law of geography, it is considered that generally, the attributes of similar candidate points on the same river or adjacent rivers are similar. Therefore, when there are already candidate power stations that meet the construction conditions within a certain distance, other candidate points within the range are regarded as homogeneous power stations, and redundant analysis and calculation are no longer performed. Therefore, a distance threshold D constraint is added during clustering in this application.
[0069] S3. Perform multi-process parallel calculation on the candidate point classification result dataset according to the pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions.
[0070] In order to improve the power station site selection efficiency and quickly conduct a global selection of the area of interest to form digital assets. This application maximally utilizes computer resources, adopts a parallel acceleration calculation method for the clustering results, controls multiple child processes, allocates the clusters to be calculated to each idle child process. When the number of clusters to be calculated is greater than the number of idle child processes, the unallocated clusters wait until the status of the child process becomes idle and then continue to be allocated. Each child process analyzes and calculates the candidate points in the cluster based on the reservoir target constraint conditions, so as to screen out the candidate points that meet the conditions to obtain the final candidate point set.
[0071] In the solution of this application, by considering spatial autocorrelation and site selection efficiency, clustering is performed on the midpoints of river network line segments based on the hierarchical clustering method, and the multi-process parallel technology is used. Each cluster is regarded as a point search task and allocated to idle child processes. Through the preset reservoir target constraint conditions, power stations that meet the conditions are quickly selected. Compared with the traditional manual site selection, as a digitalized power station selection method, the present invention not only improves the site selection efficiency in actual application scenarios, but also reduces the probability of misselection and omission, providing a scientific basis for the planning of pumped-storage power stations.
[0072] As a preferred embodiment, the step S1 specifically includes the following steps:
[0073] S11. Extract contour data according to the preset contour interval and the digital elevation model data, and extract river network data according to the water system flow threshold and the digital elevation model data; according to the preset contour interval and water system flow threshold, use the contour analysis tool and hydrological analysis tool of arcpy (including filling depressions, flow direction, flow rate, conditional function, river chain, river classification, river vectorization, etc.) to extract the contour data and river network data of the target site selection area respectively.
[0074] S12. Obtain the river network line segments in the river network data, traverse all the river network line segments, and then extract the midpoints of each river network line segment and put them into the candidate point dataset P; where, , represents the i-th candidate point, and n represents the number of candidate points;
[0075] S13. Calculate the elevation of each candidate point in the candidate point dataset P, and collect the elevations of all candidate points to obtain the candidate point elevation dataset. Specifically, the elevation of each candidate point can be calculated using the arcpy Extract Values to Points tool.
[0076] As a preferred embodiment, the step S2 specifically includes the following steps:
[0077] S21. Traverse the candidate point dataset , and calculate the distances between each candidate point in the candidate point dataset to obtain the distance matrix E; where, the expression of the distance matrix E is:
[0078] ;
[0079] where, ; represents the distance between the i-th cluster and the j-th cluster, is the Euclidean distance. When i = j, , each candidate point in the candidate point dataset is a separate cluster;
[0080] S22. According to the distance matrix E, find the two clusters with the closest distance. If the distance between the two clusters with the closest distance is less than the distance threshold, then merge these two clusters, otherwise do not merge;
[0081] S23. Update the distance matrix E; in order to make the distance between the data points within the cluster less than the distance threshold D constraint, the inter-cluster distance calculation method adopts the complete-link method, that is, the distance threshold D between any two points between two clusters.
[0082] S24. Repeat steps S22~S23 until the distance between all clusters is greater than or equal to the distance threshold, and obtain the candidate point classification result dataset ; where, , represents the j-th cluster, , represents the t-th candidate point in, m represents the number of clusters, and z represents the number of candidate points in the cluster.
[0083] Furthermore, the reservoir target constraint condition in the step S3 is:
[0084]
[0085] Among them, is the dam length constraint, is the dam height constraint, is the reservoir capacity constraint, is the reservoir dam line length of the i-th candidate point, is the reservoir dam height of the i-th candidate point, is the reservoir capacity of the i-th candidate point, is the maximum reservoir dam length allowed by the project, is the maximum reservoir dam height allowed by the project, is the minimum reservoir capacity allowed by the project.
[0086] As a preferred embodiment, the step S3 specifically includes the following steps:
[0087] S31, start multiple processes based on the preset number of subprocesses U ; Among them, , represents the t-th subprocess;
[0088] S32, traverse the candidate point classification result data set , and sequentially assign the clusters in the candidate point classification result data set to the idle subprocesses for multi-process parallel computing to obtain the final candidate point set that meets the reservoir target constraint conditions; where a cluster is a task. When the number of tasks is greater than the number of subprocesses U, all subprocesses are in the working state, and the unassigned tasks need to wait until the subprocess state becomes idle before being assigned.
[0089] As a preferred embodiment, in the step S32, the clusters in the candidate point classification result data set are sequentially assigned to the idle subprocesses for multi-process parallel computing to obtain the final candidate point set that meets the reservoir target constraint conditions; where, for the subprocess in , execute the following steps:
[0090] S321, traverse the set , and screen the contour data by elevation according to the elevation of the candidate point and the maximum reservoir dam height to obtain the screened contour set , where represents the f-th contour, g represents the number of contours in the set, represents 's elevation, satisfies ; Under this condition, the reservoir dam height calculated subsequently will all meet the dam height constraint in the reservoir target constraint conditions.
[0091] S322, obtain candidate points of the planar point coordinates and obtain the river network segment where the candidate point is located, and calculate the straight line passing through the candidate point and intersecting with the river network segment;
[0092] Furthermore, the equation expression of the straight line in step S322 is specifically:
[0093] When k exists, ;
[0094] When k does not exist, ; where k is the slope, y is the dependent variable, and x is the independent variable.
[0095] S323, obtain the required line segment located on the straight line and solve the coordinates of the two endpoints of the required line segment ;
[0096] It should be noted that in order to obtain the line segment required for subsequent intersection analysis , this application will solve the coordinates of the two endpoints of the required line segment according to the straight line equation, where , taking the candidate point as the midpoint of the required line segment , and considering that when the required line segment intersects with the contour line at only one intersection point, by adopting the "one increases while the other decreases" method for the two side line segments after bisecting the midpoint of the line segment, it is possible to make the required line segment intersect with the contour line to find the intersection point pair. Therefore, the length of the required line segment is set to 2 , and the coordinates of the two endpoints of the required line segment are solved as follows:
[0097] When k exists, use the distance formula between two points to solve the coordinates of the two endpoints of the line segment, that is:
[0098]
[0099] The coordinates of the two endpoints of the line segment are obtained through the derived formula as ( , ) and ( , ) as:
[0100] ;
[0101] When k does not exist, the coordinates of the two endpoints of the line segment are respectively ( , ), ([[]]END]] , ).
[0102] S324. Calculate the total number of rotations c based on a preset rotation angle .
[0103] In an actual reservoir construction project, on the premise of ensuring the water retaining condition, the reservoir dam line can be rotated by multiple angles to find the maximum reservoir dam line formed with the mountain body. Therefore, in the embodiment of the present application, based on a preset rotation angle , calculate the total number of rotations c.
[0104]
[0105]
[0106] where floor is rounding down, is the candidate point on the slope of the river network segment where it is located, is the angle (in radians) between the river network segment and the positive x-axis direction of the rectangular coordinate system.
[0107] S325. Traverse c times and rotate the required line segment by different angles to determine whether there is a reservoir dam line that meets the reservoir target constraint conditions for the currently calculated candidate point; specifically, it includes the following steps:
[0108] S3251. Calculate the slope of the line after the u-th rotation angle, and solve the coordinates of the two endpoints of the required line segment at this angle;
[0109] Specifically, the slope of the line after the u-th ( ) rotation angle can be calculated according to the following formula, and the coordinates of the two endpoints of the required line segment can be obtained according to step S323.
[0110]
[0111] ; where A is the angle (in radians) after the line rotation, and Z represents the set of integers. In other cases, it is non-existent.
[0112] S3252. Generate a line feature layer based on the coordinates of two endpoints, and use the intersection analysis tool of arcpy to extract the intersection points with the contour lines to obtain the reservoir dam line. Among them, in the attribute table of the intersection layer, filter out the geometric elements of the single-point type and retain the geometric elements of the multi-point type. It should be noted that the connection line between each pair of intersection points is the reservoir dam line.
[0113] It should be noted that in the attribute table of the intersection layer, there may be two geometric types, "single point" and "multi point". Since only the "multi point" geometric elements can generate a closed polygon formed by the reservoir dam line and the contour lines, the geometric elements of "single point" need to be filtered out. In addition, there may be more than 2 intersection points on the same contour line. For this abnormal situation, in the embodiment of the present application, the two intersection points closest to the river are used as the endpoints of the reservoir dam line. Suppose 、 are the two endpoints of the reservoir dam line formed by the intersection of the f-th contour line and the line feature, then and is located on the reservoir dam line.
[0114] S3253. Calculate the length of each reservoir dam line, and obtain a set of dam lines whose reservoir dam line lengths meet the dam length constraint; calculate the distance between the two endpoints of each reservoir dam line (i.e., the length of the reservoir dam line) according to the Euclidean distance formula, and obtain a set of dam lines whose reservoir dam line lengths meet the dam length constraint.
[0115] S3254. Sort the reservoir dam lines according to the elevation of the endpoints of the reservoir dam line from high to low, select the dam line B with the maximum elevation, and according to the dam line B with the maximum elevation and the contour line where its endpoints are located obtain the closed polygon formed by the dam line B with the maximum elevation and the contour line ;
[0116] Considering that the reservoir capacity decreases as the elevation of the contour line decreases, that is, when the reservoir capacity formed by the contour line with the maximum elevation value and the dam line does not meet the reservoir capacity constraint, then the contour line with a smaller elevation value also does not meet. Therefore, in order to avoid unnecessary calculations, in the embodiment of the present application, the reservoir dam lines are sorted according to the elevation of the endpoints of the reservoir dam line from high to low, and the dam line B with the maximum elevation is selected. In addition, the feature to polygon tool of arcpy can be used, and the dam line B with the maximum elevation and the contour line where its endpoints are located are used as the input parameters of the tool, and the closed polygon formed by the reservoir dam line and the contour line can be generated (reservoir surface).
[0117] S3255. Determine the reservoir capacity according to the closed polygon and the digital elevation model data ;
[0118] Specifically, based on the closed polygon Using the Extract by Mask tool in arcpy with DEM data to obtain the terrain raster data within the range, and then calculating the reservoir storage capacity using the Surface Volume tool in arcpy based on the terrain raster data and the plane elevation ( elevation value). .
[0119] S3256, if the reservoir storage capacity , then return to step S325 to perform the next rotation angle; otherwise, there is a reservoir dam line at the currently calculated candidate point that meets the reservoir target constraint conditions, traverse end, put the candidate point into the final candidate point set P'', and the subprocess becomes idle.
[0120] To further illustrate the technical solution of the present application, the present application also provides a specific embodiment for those skilled in the art to understand. The 30-meter DEM data of a certain research area in Sichuan Province is used for specific illustration as follows:
[0121] Since the calculation process of the present application involves parameters such as length and area, the data in the embodiment is described in steps using the projected coordinate system (CGCS2000). The coordinates of the circumscribed rectangle of the area are (646032.171753, 3225012.113918, 687402.224616, 3262461.666722), the area of the area is 1505.49 square kilometers, the minimum elevation of the area is 2901 meters, and the maximum elevation is 4942 meters. The CPU of the developed hardware environment is 16 cores or more.
[0122] In this embodiment, the contour interval adopted is 30 meters and the water system flow threshold is 300. The extracted contour layer attribute table contains 3739 records, and the water system layer attribute table contains 2692 records.
[0123] The candidate point data set extracted from the river network data generated according to step S11 .
[0124] The candidate point elevation data set obtained by using the Extract Values to Points tool in arcpy .
[0125] In this embodiment, the distance threshold D = 500 meters is adopted, and the candidate points are clustered into 1292 clusters, that is .
[0126] In this embodiment, = 600 meters, = 100 meters, = 5 million cubic meters, α = 15°, the number of subprocesses U = 15.
[0127] There are a total of 15 subprocesses started in this embodiment, that is .
[0128] Traverse , and allocate each cluster to different subprocesses. During the experiment, there are candidate points in cluster that meet the constraint conditions. Therefore, cluster is selected for the description of the embodiment.
[0129] Traverse . The candidate points in the current cluster meet the reservoir target constraint conditions. Therefore, the following steps further elaborate on the solution of this application by calculating the candidate points . In this embodiment = 3883. Therefore, the contour lines in the elevation range of [3883, 3983] are screened to obtain the contour line set .
[0130] In this embodiment, the plane coordinates of the candidate point are (684008.728937, 3261368.698813), and the straight line equation is:
[0131]
[0132] In this embodiment, the ID of the river segment where the candidate point is located is 5. After calculation, the slope of the river segment is obtained , c = 12, .
[0133] Traverse 12 times. Take the result of the first traversal for the subsequent step description. It is calculated that , slope , and the coordinates of the two endpoints of the line segment are (684603.423877, 3261448.309916) and (683414.033997, 3261289.08771) respectively. In this embodiment, the plane coordinates are uniformly reserved to 6 significant digits after the decimal point.
[0134] Generate a line feature layer according to the coordinates of the two endpoints, and use the intersection analysis tool of arcpy to extract the intersection points with the contour lines to obtain an intersection point layer. The layer attribute table contains 5 records, and all belong to the "multipoint" geometry type.
[0135] Sort the reservoir dam line according to the elevation of the endpoints of the reservoir dam line from high to low. At this time, the length of the dam line with the highest elevation is 357.55 meters, and the dam height is 97 meters.
[0136] Using the feature to polygon tool in arcpy, the reservoir dam line and the contour line where it is located (with an elevation value of 3980 meters) are used as the input parameters of the tool to generate the reservoir surface .
[0137] Using the extract by mask tool in arcpy to obtain the terrain raster data within the range. Based on the terrain raster data and the plane elevation, the reservoir storage capacity is calculated using the surface volume tool in arcpy ten thousand cubic meters.
[0138] , which meets the storage capacity constraint in the reservoir target constraint conditions. Traverse 、 end, and put into the final candidate point set P'', and the current subprocess becomes idle.
[0139] In this embodiment, a total of 209 candidate points that meet the reservoir target constraint conditions are calculated, and it takes 726.25 s.
[0140] Furthermore, in order to more clearly highlight the advantages of the present invention in the site selection efficiency, two sets of comparative experiments are conducted on the two variables of classification and parallelism using the control variable method in this embodiment. As shown in Table 1, the experimental results show that compared with the site selection method without classification and parallelism, the fast general site selection method for pumped storage power stations based on hierarchical clustering and parallel computing proposed by the present invention improves the site selection efficiency to a certain extent.
[0141] Table 1 Comparison table of time spent in different calculation methods
[0142]
[0143] The present invention also provides a fast general site selection device for pumped storage power stations based on hierarchical clustering and parallel computing, including:
[0144] A preprocessing unit for obtaining the digital elevation model data of the target site selection area and preprocessing the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set;
[0145] A hierarchical clustering processing unit for performing hierarchical clustering processing on the candidate point information to obtain a candidate point classification result data set;
[0146] A parallel computing unit for performing multi-process parallel computing on the candidate point classification result data set according to the pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions.
[0147] The present invention also provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned rapid general selection method of pumped storage power station sites based on hierarchical clustering and parallel computing are realized. It can be understood that when executed by the processor, the above-mentioned rapid general selection method of pumped storage power station sites based on hierarchical clustering and parallel computing is realized. Therefore, all embodiments of the above method are applicable to this storage medium, and all can achieve the same or similar beneficial effects.
[0148] The above are only the preferred embodiments of the present invention, and do not limit the protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the protection scope of the present invention.
Claims
1. A rapid general selection method for pumped-storage power station sites based on hierarchical clustering and parallel computing, characterized in that, It includes the following steps: S1. Obtain the digital elevation model data of the target site selection area, and preprocess the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set; S2. Perform hierarchical clustering processing on the candidate point information to obtain a candidate point classification result data set; S3. Perform multi-process parallel calculation on the candidate point classification result data set according to the pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions; The specific steps of step S1 include the following steps: S11. Extract contour data according to a preset contour interval and the digital elevation model data, and extract river network data according to a water system flow threshold and the digital elevation model data; S12. Obtain the river network line segments in the river network data, traverse all the river network line segments, and then extract the midpoints of each river network line segment and put them into the candidate point data set P; where, , represents the i-th candidate point, and n represents the number of candidate points; S13. Calculate the elevation of each candidate point in the candidate point data set P, and collect the elevations of all candidate points to obtain a candidate point elevation data set; The specific steps of step S2 include the following steps: S21, traverse the candidate point dataset , calculate the distances between each pair of candidate points in the candidate point dataset to obtain a distance matrix E; wherein, the expression of the distance matrix E is: ; Among them, ; represents the i th cluster and the j th cluster, is the Euclidean distance. When i = j, , each candidate point in the candidate point dataset is a separate cluster; S22. Find the two closest clusters according to the distance matrix E. If the distance between the two closest clusters is less than the distance threshold, then merge these two clusters, otherwise do not merge; S23. Update the distance matrix E; S24. Repeat steps S22 - S23 until the distance between all clusters is greater than or equal to the distance threshold, and a candidate point classification result dataset is obtained. ; where , represents the j -th cluster, , represents the -th candidate point in t . m represents the number of clusters, and z represents the number of candidate points in a cluster.
2. The method for quickly and generally selecting pumped storage power station sites based on hierarchical clustering and parallel computing according to claim 1, characterized in that The reservoir target constraint conditions in step S3 are: Among them, is the dam length constraint, is the dam height constraint, is the reservoir capacity constraint, is the i reservoir dam line length of the th candidate point, i is the reservoir dam height of the th candidate point, i is the reservoir capacity of the th candidate point, is the maximum reservoir dam length allowed by the project, is the minimum reservoir capacity allowed by the project.
3. The rapid general selection method of pumped - storage power station sites based on hierarchical clustering and parallel computing according to claim 2, wherein, The specific steps of step S3 include the following steps: S31, start multiple processes based on the preset number of subprocesses U ; where , represents the t-th subprocess; S32. Traverse the candidate point classification result data set , and sequentially assign the clusters in the candidate point classification result data set to idle child processes for multi-process parallel computing to obtain a final candidate point set that meets the reservoir target constraint conditions; where a cluster is a task. When the number of tasks is greater than the number of child processes U, all child processes are in a working state, and unassigned tasks need to wait until the child process status becomes idle before being assigned.
4. The rapid general selection method of pumped storage power station sites based on hierarchical clustering and parallel computing according to claim 3, wherein In the step S32, the clusters in the candidate point classification result data set are sequentially assigned to idle child processes for multi-process parallel computing to obtain a final candidate point set that meets the reservoir target constraint conditions; among them, for the child process in the , the following steps are executed: S321, traverse the set , according to the candidate points ' elevation , the maximum reservoir dam height , screen the contour data by elevation to obtain the screened contour set , where represents the f-th contour, g represents the number of contours in the set, represents 's elevation, satisfies ; S322, Obtain candidate points The planar point coordinates and obtain the candidate points where the river network segment is located, and calculate the straight line passing through the candidate points and intersecting with the river network segment; S323, obtain the required line segment located on the said straight line and solve the said required line segment for the coordinates of its two end points; S324, based on a preset rotation angle , calculate the total number of rotations c ; S325, traverse c times, rotate the required line segment at different angles to determine whether there is a reservoir dam line that meets the reservoir target constraint conditions for the currently calculated candidate points; specifically, it includes the following steps: S3251, calculate the slope of the line after the u-th rotation angle , and solve for the required line segment coordinates of the two endpoints at this angle; S3252. Generate a line feature layer according to the coordinates of two endpoints, and use the intersection analysis tool of arcpy to extract the intersection points with the contour lines to obtain the reservoir dam line; wherein, in the intersection layer attribute table, filter out the geometric elements of the single-point type and retain the geometric elements of the multi-point type; S3253. Calculate the length of each reservoir dam line, and obtain a set of dam lines whose reservoir dam line lengths meet the dam length constraints; S3254, sort the reservoir dam lines according to the elevation of the endpoints of the reservoir dam line from high to low, select the dam line B with the maximum elevation, and according to the dam line B with the maximum elevation and the contour line where its endpoints are located Obtain the closed polygon formed by the dam line B with the maximum elevation and the contour line ; S3255, determining the reservoir capacity according to the closed polygon and digital elevation model data ; S3256, if the reservoir storage capacity , return to step S325 for the next rotation angle; otherwise, there is a reservoir dam line among the currently calculated candidate points that meets the reservoir target constraint conditions. Traverse End, put the candidate point into the final candidate point set P'', and the subprocess becomes idle.
5. The rapid general selection method of pumped storage power station sites based on hierarchical clustering and parallel computing according to claim 4, characterized in that, The specific equation expression of the straight line in step S322 is: In the presence of k, ; When k does not exist, ; where k is the slope, y is the dependent variable, and x is the independent variable.
6. The rapid general selection method for pumped storage power station sites based on hierarchical clustering and parallel computing according to claim 4, characterized in that The specific steps of step S323 include the following steps: Set the length of the required line segment to 2 , and the solution process for the coordinates of the two endpoints is as follows: When k exists, use the distance formula between two points to solve the coordinates of the two endpoints of the line segment, that is: The coordinates of the two endpoints of the line segment are obtained by deriving the formula ( , ), ( , ) as follows: ; When k does not exist, the coordinates of the two endpoints of the line segment are respectively ( , ), ( , ).
7. A rapid general selection device for pumped storage power station sites based on hierarchical clustering and parallel computing, characterized in that, It includes: A preprocessing unit, configured to obtain the digital elevation model data of the target site selection area, and preprocess the digital elevation model data to obtain candidate point information; wherein, the candidate point information includes a candidate point data set and a candidate point elevation data set; specifically includes the following steps: S11. Extract contour data according to a preset contour interval and the digital elevation model data, and extract river network data according to a water system flow threshold and the digital elevation model data; S12. Obtain the river network line segments in the river network data, traverse all the river network line segments, and then extract the midpoints of each river network line segment and put them into the candidate point data set P; where , represents the i-th candidate point, and n represents the number of candidate points; S13. Calculate the elevation of each candidate point in the candidate point data set P, and collect the elevations of all candidate points to obtain a candidate point elevation data set; A hierarchical clustering processing unit, configured to perform hierarchical clustering processing on the candidate point information to obtain a candidate point classification result data set; specifically includes the following steps: S21, traverse the candidate point dataset , calculate the distances between each pair of candidate points in the candidate point dataset to obtain a distance matrix E; where the expression of the distance matrix E is: ; Among them, ; represents the distance between the i-th cluster and the j-th cluster, is the Euclidean distance. When i = j, , and each candidate point in the candidate point data set is an individual cluster; S22. Find the two closest clusters according to the distance matrix E. If the distance between the two closest clusters is less than the distance threshold, then merge these two clusters, otherwise do not merge; S23. Update the distance matrix E; S24. Repeat steps S22 - S23 until the distances between all clusters are greater than or equal to the distance threshold, then obtain the candidate point classification result data set ; where , represents the j-th cluster, , represents the t-th candidate point in , m represents the number of clusters, and z represents the number of candidate points in a cluster; A parallel calculation unit, configured to perform multi-process parallel calculation on the candidate point classification result data set according to the pre-established reservoir target constraint conditions to obtain a final candidate point set that meets the reservoir target constraint conditions.
8. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for quickly and generally selecting pumped-storage power station sites based on hierarchical clustering and parallel computing according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target position point determination method, device and equipment, and storage medium
CN110493333A
Pumped storage power station site selection method, terminal equipment and storage medium
CN116451860A