Power grid dispatching data mining method, system and device and storage medium
By applying the density peak algorithm of mixed density and microcluster aggregation, the Bray-Curtis algorithm and pruning tree algorithm of hybrid density and microcluster aggregation in power grid scheduling, the complex problems of traditional grid scheduling data management and analysis are solved, data mining efficiency and accuracy are improved, and the reliability of power grid scheduling is improved.
Patent Information
- Application Number
- CN202411661766.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-02
AI Technical Summary
The traditional power system scheduling model is difficult to meet the real-time regulation needs, and the massive scheduling data has a wide variety of types, complex structures, and uneven data quality, resulting in complex data management and analysis, affecting the security of power grid scheduling.
The density peak algorithm based on mixed density and microcluster aggregation is used to screen the grid scheduling data, and the data is classified through the Bray-Curtis algorithm, and the information mining is used to obtain the target effective information.
It improves the efficiency and accuracy of grid scheduling data mining, eliminates joint errors in data extraction, and provides better quality data, thereby improving the reliability of grid scheduling.
Smart Images

Figure CN119917940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, system, device and storage medium for mining power grid dispatching data. Background Art
[0002] In the process of accelerating the construction of new power systems, the complexity of the power grid structure continues to increase, which puts forward more stringent requirements for the safe operation and power supply reliability of the power grid. As a key link in the production and operation of the power grid, power grid dispatching involves the coordinated participation of multi-level dispatching agencies. Dispatching agencies at all levels need to obtain and process power system operation data covering various links such as generation, transmission, distribution, and use in real time and accurately in order to make quick and accurate decisions. However, with the increasing proportion of renewable energy and the increasingly complex and changeable power grid operation environment, the traditional power system dispatching mode has been difficult to meet the needs of real-time regulation. The large amount of data accumulated over a long period of operation makes the data management and analysis of the power dispatching system more complicated. These data are not only of various types and complex structures, but also may be affected by multiple factors such as communication failures and equipment defects, resulting in uneven data quality. These problems not only increase the difficulty of data management, but also seriously affect the effective analysis and utilization of data by the dispatching system, thereby threatening the safety of power grid dispatching. Data mining and big data analysis are conducive to the dispatching system to accurately and reliably complete real-time decisions. Many studies have constructed dispatching decision management based on big data from the aspects of timeliness and reliability of data acquisition. In fact, facing the massive amount of dispatching data, and the data has the characteristics of diversified states and complex variable types, it increases the difficulty of power grid dispatching and makes real-time decision-making very difficult. Therefore, how to effectively mine and extract data from power systems under different operating conditions is the key to intelligent dispatching. Summary of the invention
[0003] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0004] To this end, an object of an embodiment of the present invention is to provide a power grid dispatching data mining method, which improves the efficiency and accuracy of power grid dispatching data mining, thereby improving the reliability of power grid dispatching.
[0005] Another object of an embodiment of the present invention is to provide a power grid dispatching data mining system.
[0006] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0007] In a first aspect, an embodiment of the present invention provides a method for mining power grid dispatching data, comprising the following steps:
[0008] The first power grid dispatching data is screened by a density peak algorithm based on hybrid density and micro-cluster aggregation to obtain second power grid dispatching data;
[0009] Classifying the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types;
[0010] The third power grid dispatching data is mined through information mining using a pruning tree algorithm to obtain target effective information.
[0011] Further, in one embodiment of the present invention, the first power grid dispatching data is screened by a density peak algorithm based on mixed density and micro-cluster aggregation to obtain the second power grid dispatching data, which specifically includes:
[0012] Determining the local density of each data sample of the first power grid dispatching data;
[0013] Determining a sample relative distance of each of the data samples according to the local density;
[0014] Determine the reverse K nearest neighbor set of each data sample by a reverse K nearest neighbor algorithm according to the relative distance of the samples;
[0015] Determining the absolute density of each of the data samples according to the reverse K nearest neighbor set, and determining the attribution relationship between the data samples according to the absolute density;
[0016] Determine a mixed density of each of the data samples according to the absolute density and the attribution relationship;
[0017] The first power grid dispatching data is screened according to the mixed density to obtain the second power grid dispatching data.
[0018] Further, in one embodiment of the present invention, the local density is determined by the following formula:
[0019]
[0020] or
[0021]
[0022] Among them, ρ i Represents data sample x i The local density of the first power grid dispatching data, D represents the data sample set, dist(x i ,x j ) represents the data sample x i With data sample x j The Euclidean distance, d c Represents the cutoff distance.
[0023] Further, in one embodiment of the present invention, the reverse K nearest neighbor set is determined by the following formula:
[0024] rnn K (x i )={x j ∈D|x i ∈KNN(x j )}
[0025] KNN(x j )={x∈D|d(x j ,x)<d(x j ,x j-Kth )}
[0026] Among them, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set, KNN(x j ) represents the data sample x j The K nearest neighbor set of d(x j ,x) represents the data sample x j The distance from the data sample x, x j-Kth Table distance data sample x j The K-th most recent data sample, where K is a preset value.
[0027] Further, in one embodiment of the present invention, the absolute density is determined by the following formula:
[0028]
[0029] Among them, P i Represents data sample x i The absolute density, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set, dist(x i ,x j ) represents the data sample x i With data sample x j The Euclidean distance.
[0030] Further, in one embodiment of the present invention, the second power grid dispatching data is classified by the Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types, which specifically includes:
[0031] Construct data dissimilarity matrix;
[0032] According to the data dissimilarity matrix, calculating the data type dissimilarity between each data sample of the second power grid dispatching data by using a Bray-Curtis algorithm;
[0033] The second power grid dispatching data is classified according to the data type difference to obtain the third power grid dispatching data of multiple data types.
[0034] Further, in one embodiment of the present invention, the information mining of the third power grid dispatching data by using a pruning tree algorithm to obtain target valid information specifically includes:
[0035] Determine a plurality of sample vectors according to the third power grid dispatching data;
[0036] Classifying the sample vectors by a pruning tree algorithm, and determining a decision center probability value for cluster mining according to the classification result and the data type;
[0037] The effective information category is determined according to the decision center probability value, and then the target effective information is obtained by screening from the third power grid dispatching data according to the effective information category.
[0038] In a second aspect, an embodiment of the present invention provides a power grid dispatching data mining system, including:
[0039] A data screening module, used for screening the first power grid dispatching data by using a density peak algorithm based on mixed density and micro-cluster aggregation to obtain second power grid dispatching data;
[0040] A data classification module, used for classifying the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types;
[0041] The information mining module is used to perform information mining on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information.
[0042] In a third aspect, an embodiment of the present invention provides a power grid dispatching data mining device, comprising:
[0043] at least one processor;
[0044] at least one memory for storing at least one program;
[0045] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned power grid scheduling data mining method.
[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to execute the above-mentioned power grid dispatching data mining method when executed by the processor.
[0047] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:
[0048] The embodiment of the present invention screens the first power grid dispatching data through a density peak algorithm based on mixed density and micro-cluster aggregation to obtain second power grid dispatching data, classifies the second power grid dispatching data through a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types, and performs information mining on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information. The embodiment of the present invention screens data by using a density peak algorithm based on mixed density and micro-cluster aggregation, takes into account the differences in power grid dispatch data types and divides the data into different numbers of micro-clusters according to the differences, and then aggregates the micro-clusters in combination with the similarities between the micro-clusters until the number of micro-clusters reaches the real number of clusters, effectively solving the problem of selection errors in massive dispatch data, eliminating the associated errors in data extraction, and providing better quality data for subsequent data mining; data type dissimilarity calculation based on the Bray-Curtis algorithm adopts the method of constructing data structure and dissimilarity matrix to obtain the similarity between data, thereby realizing reasonable classification of data types from massive data, and facilitating targeted analysis according to different data types; in order to solve the shortcomings of slow mining speed and low accuracy of traditional clustering algorithms, information mining of data is performed based on the pruning tree algorithm, which can effectively mine potential effective information of different data types, improve the efficiency and accuracy of power grid dispatch data mining, and thus improve the reliability of power grid dispatch. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 A flowchart of a method for mining power grid dispatching data provided by an embodiment of the present invention;
[0051] Figure 2 A structural block diagram of a power grid dispatching data mining system provided by an embodiment of the present invention;
[0052] Figure 3 A structural block diagram of a power grid dispatching data mining device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0054] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.
[0055] At present, the research on data mining is mainly from the perspective of extracting outliers from dispatching data, and less consideration is given to data classification and screening based on the characteristic types of the data itself. Through the application of data mining technology, we can extract valuable information from massive data in order to provide faster and more effective decision support. This method uses density clustering technology to screen data and Bray-Curtis-based data type dissimilarity calculation, and uses the pruning tree optimization algorithm to extract effective information and its correlation from massive data, improve data processing speed and calculation accuracy, and proposes a power grid dispatching data mining method based on the density peak algorithm of mixed density and micro-cluster aggregation.
[0056] The embodiment of the present invention takes into account the diversity and complexity of scheduling data, which is not conducive to statistical analysis. It adopts a density peak algorithm of mixed density and micro-cluster aggregation to screen and mine the data, improves the adaptability of the traditional density peak algorithm to data sets with uneven density distribution, thereby effectively dealing with the complex structure of scheduling data, low processing efficiency caused by abnormal data interference, and frequent errors, thereby improving the intelligent analysis capability of scheduling data.
[0057] Reference Figure 1 The embodiment of the present invention provides a power grid dispatching data mining method, which specifically includes the following steps:
[0058] S101. Filter first power grid dispatching data by using a density peak algorithm based on mixed density and micro-cluster aggregation to obtain second power grid dispatching data.
[0059] Specifically, the present invention takes into account the huge amount of power grid dispatching data. The main idea of the density peak (DPC) method is to cluster based on the similarity principle within a specific range, which can effectively filter the noise data, and control the granularity and accuracy of clustering by setting the density parameter as the termination condition to adapt to different clustering scenarios. However, the data types between different power types and different lines in the power grid dispatching data are relatively similar, and the data distinction is not obvious. Moreover, the algorithm cannot well reflect the differences in the structure of various clusters of data with uneven density distribution, and it is easy to ignore the cluster center of the low-density cluster; on the other hand, a large amount of false data and abnormal data in the smart grid rush in, and the use of the traditional density peak algorithm is prone to cause the phenomenon of error association. In order to solve the problem of selection errors in massive dispatching data, the data screening of the density peak algorithm based on mixed density and micro-cluster aggregation first divides the data into different micro-clusters according to the difference in data type, and then defines the similarity between micro-clusters, and aggregates the micro-clusters according to the similarity until the number of micro-clusters reaches the real number of clusters. This method can effectively eliminate the associated errors in data extraction and provide better quality data for subsequent data mining.
[0060] As an optional implementation, the first power grid dispatching data is screened by a density peak algorithm based on mixed density and micro-cluster aggregation to obtain the second power grid dispatching data, which specifically includes:
[0061] S1011, determining the local density of each data sample of the first power grid dispatching data;
[0062] S1012, determining the sample relative distance of each data sample according to the local density;
[0063] S1013, determining the reverse K nearest neighbor set of each data sample by using the reverse K nearest neighbor algorithm according to the relative distance of the samples;
[0064] S1014, determining the absolute density of each data sample according to the reverse K nearest neighbor set, and determining the attribution relationship between the data samples according to the absolute density;
[0065] S1015, determining the mixed density of each data sample according to the absolute density and the attribution relationship;
[0066] S1016. Filter the first power grid dispatching data according to the mixed density to obtain the second power grid dispatching data.
[0067] As a further optional embodiment, the local density is determined by the following formula:
[0068]
[0069] or
[0070]
[0071] Among them, ρ i Represents data sample x i The local density of the first power grid dispatching data, D represents the data sample set, dist(x i ,x j ) represents the data sample x i With data sample x j The Euclidean distance, d c Represents the cutoff distance.
[0072] As an optional implementation, the reverse K nearest neighbor set is determined by the following formula:
[0073] rnn K (x i )={x j ∈D|x i ∈KNN(x j )}
[0074] KNN(x j )={x∈D|d(x j ,x)<d(x j ,x j-Kth )}
[0075] Among them, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set, KNN(x j ) represents the data sample x j The K nearest neighbor set of d(x j ,x) represents the data sample x j The distance from the data sample x, x j-Kth Table distance data sample x j The K-th most recent data sample, where K is a preset value.
[0076] As a further optional embodiment, the absolute density is determined by the following formula:
[0077]
[0078] Among them, P i Represents data sample x i The absolute density, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set, dist(x i ,xj ) represents the data sample x i With data sample x j The Euclidean distance.
[0079] Specifically, the steps of data screening in the embodiment of the present invention are as follows:
[0080] Step 1: Local density calculation: The DPC algorithm gives the relative distance of samples according to the local density and sample distance, selects the cluster center based on the local density and relative distance, and finally allocates non-center samples.
[0081] Let x i and x j are two data samples in power grid dispatching, and the local density of the samples is calculated as:
[0082]
[0083] In the formula, dist(x i ,x j ) is the sample x i and x j The Euclidean distance, d c is the cutoff distance.
[0084] It can be seen that there are two ways to calculate local density. When the data sample belongs to a small-scale data set (the data sample size is less than the preset threshold), the local density is calculated by formula (1); otherwise, it is calculated by formula (2).
[0085] Step 2: Relative distance calculation. Sample x i The relative distance δ i for:
[0086]
[0087] For samples with non-maximum density, the relative distance is the distance between the sample and the sample with a larger density and the closest distance to it. For samples with the largest local density, the relative distance is the distance between the farthest samples in the data.
[0088] Step 3: Reverse K nearest neighbor. For data sets with large density differences, the selection of valid data is prone to errors. Therefore, an absolute density is defined based on the reverse K nearest neighbor to describe the relationship between samples. Sample x i The reverse K nearest neighbor set is:
[0089] rnn K (x i )={x j ∈D|x i ∈KNN(x j )} (4)
[0090] KNN(x j )={x∈D|d(x j ,x)<d(x j ,x j-Kth )}
[0091] In the formula, KNN represents sample x i The K neighboring sets of x j-Kth is the distance x j The K-th most recent sample. rnn K (x i ) can be adaptively adjusted according to local information, so the samples in the core area have more reverse K nearest neighbors than the samples in the edge area. This method enables the reverse K nearest neighbors to better reflect local information.
[0092] Step 4: Absolute density calculation. Based on reverse K nearest neighbors, sample x i The absolute density p i for
[0093]
[0094] In the formula, x i The number of reverse K nearest neighbor samples. i The more reverse neighbors there are and the closer they are to the reverse neighbor samples, the i The greater the absolute density.
[0095] Step 5: Determine the attribution relationship. If the sample x j X i Absolute density is high and far from x i The sample with the closest distance is called x i For x j A subordinate sample of .
[0096] Step 6: Calculate the mixed density. Based on the absolute density and the relationship between samples, the mixed density HP i Calculated as
[0097]
[0098] In the formula, α i is the density weight value, α i =exp(-P i / max(P));dn(x i ) is the variable x i The natural logarithm of the differential; mean(P) is the vector P i The average value of .
[0099] In the face of large-scale heterogeneous data types in power grid dispatching, the DPC algorithm is subject to the influence of differences in different data types on the clustering effect, so that low-density sample clusters can also be selected as cluster centers, resulting in poor final clustering effect. The HMDPC algorithm of the present invention reduces the influence of interference information on effective information extraction through density weight values, thereby eliminating the erroneous selection of cluster centers for data sets with uneven density distribution and realizing intelligent screening of massive data.
[0100] S102: Classify the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types.
[0101] As an optional implementation, the second power grid dispatching data is classified by the Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types, which specifically includes:
[0102] S1021, constructing a data dissimilarity matrix;
[0103] S1022. Calculate the data type dissimilarity between each data sample of the second power grid dispatching data by using the Bray-Curtis algorithm according to the data dissimilarity matrix;
[0104] S1023. Classify the second power grid dispatching data according to data type differences to obtain third power grid dispatching data of multiple data types.
[0105] Specifically, in order to reasonably classify data types from massive data and conduct targeted analysis based on different data types, the data structure and dissimilarity matrix are constructed to obtain the similarity between data, so as to achieve classification processing by data type. The following is the constructed data dissimilarity matrix:
[0106]
[0107] The data dissimilarity matrix is a data structure used to store the dissimilarity or degree of difference between n objects. In the formula: n represents the number of data objects, s and d represent the difference values between data, and p represents the attribute. When the difference value is a positive number, if s and d are closer to 0, the attribute value p will be larger, which means that in this case, the similarity between objects s and d is low, that is, they are not similar; on the contrary, if the values of s and d are less than 0, the value of the attribute value p will be smaller. This means that the similarity between objects s and d is high.
[0108] Based on the above matrix, the Bray-Curtis algorithm is used to calculate the data type dissimilarity between data objects. This algorithm considers the absolute deviation and average value between samples to calculate the dissimilarity:
[0109]
[0110] In the formula, s f Indicates the absolute deviation of the variable value; m f Represents the absolute average value of f. The data type dissimilarity is calculated based on formula (7):
[0111]
[0112] In the formula, d(i,j) is used to represent the dissimilarity between objects i and j, and the dissimilarity is usually non-negative. If objects i and j are more similar, the dissimilarity is closer to 0; otherwise, the larger the value, the greater the dissimilarity, and d(i,j) = d(i,j), d(i,j) = 0. Data type dissimilarity is calculated as follows:
[0113] W=d(i,j)*k l (9)
[0114] In the formula, k l To represent the amount of data for cluster analysis. So far, data screening and data type dissimilarity calculation have been implemented, and the next step is to implement fast mining of scheduling data.
[0115] S103. Perform information mining on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information.
[0116] As an optional implementation, information mining is performed on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information, which specifically includes:
[0117] S1031. Determine a plurality of sample vectors according to the third power grid dispatching data;
[0118] S1032, classifying the sample vectors by a pruning tree algorithm, and determining a decision center probability value for cluster mining according to the classification result and the data type;
[0119] S1033. Determine the effective information category according to the decision center probability value, and then filter the target effective information from the third power grid dispatching data according to the effective information category.
[0120] Specifically, a fast mining method for power dispatching data based on pruning trees is an efficient data mining technology designed to process the large amount of data generated in power system dispatching. In power dispatching, historical and real-time data need to be analyzed to optimize grid operation, predict load, implement fault diagnosis, and formulate effective power distribution strategies. Since power data is usually large-scale and high-dimensional, fast and efficient data mining methods are needed to process these data.
[0121] Pruning trees is a common strategy in data mining. This strategy is used to reduce the search space and improve data mining efficiency. In the pruning tree method, a series of criteria are used to decide whether to keep a node, thereby eliminating unimportant or irrelevant data branches and focusing on processing data with more potential value.
[0122] The input sample vector consists of attribute values and category labels and is defined as (v1, v2, …, v i , c), where v i Represents the attribute values of the input sample, and c represents the category label of the sample. This vector is a representation of real-world data records, which also constitute the training data set. Among them, the label can also be used as input training data. After the classification is completed, the pruning tree algorithm can be introduced for data mining and prediction accuracy: first, learn or extract knowledge from the provided training data; then, use the trained decision tree to classify the new input data. For each data attribute value v in turn i Test and record until the potential information contained in the data is mined by finding the class to which the record belongs. Pruning tree algorithm expression:
[0123] C ost (M,D)=C ost (DM)+B cost (M) (10)
[0124] In the formula, C ost (M,D) is the encoding code, C ost (DM) is the coding cost, B sost (M) is the total number of classification errors.
[0125] After pruning the data using the pruning tree constructed by formula (10), the decision center probability value of cluster mining in power dispatching data is calculated:
[0126]
[0127] Where a represents the scheduling parameter of the decision center; x k represents the dynamic inertia weight; Indicates the valid information category.
[0128] According to the calculation of central probability, effective information in the data can be mined:
[0129]
[0130] According to the above formula, determine the sample data x that meets the conditions iThe corresponding valid information category is used to filter the target valid information from the third power grid dispatching data; n represents a positive integer; x k+1 Represents the k+1th element of a sequence.
[0131] The method steps of the embodiment of the present invention are described above. It can be recognized that the embodiment of the present invention screens data by using a density peak algorithm based on mixed density and micro-cluster aggregation, takes into account the differences in power grid dispatch data types and divides the data into different numbers of micro-clusters according to the differences, and then aggregates the micro-clusters in combination with the similarity between the micro-clusters until the number of micro-clusters reaches the real number of clusters, effectively solving the selection error problem in massive dispatch data, eliminating the associated errors in data extraction, and providing better quality data for subsequent data mining; based on the Bray-Curtis algorithm, the data type dissimilarity calculation is carried out, and the similarity between data is obtained by constructing a data structure and a dissimilarity matrix, so as to realize the reasonable classification of data types from massive data, and facilitate targeted analysis according to different data types; in order to solve the shortcomings of slow mining speed and low accuracy of traditional clustering algorithms, information mining of data based on the pruning tree algorithm can effectively mine potential effective information of different data types, improve the efficiency and accuracy of power grid dispatch data mining, and thus improve the reliability of power grid dispatch.
[0132] Reference Figure 2 , an embodiment of the present invention provides a power grid dispatching data mining system, including:
[0133] A data screening module, used for screening the first power grid dispatching data by using a density peak algorithm based on mixed density and micro-cluster aggregation to obtain second power grid dispatching data;
[0134] A data classification module, used for classifying the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types;
[0135] The information mining module is used to perform information mining on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information.
[0136] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0137] Reference Figure 3 , an embodiment of the present invention provides a power grid dispatching data mining device, comprising:
[0138] at least one processor;
[0139] at least one memory for storing at least one program;
[0140] When the at least one program is executed by the at least one processor, the at least one processor implements the power grid scheduling data mining method.
[0141] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0142] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned power grid dispatching data mining method.
[0143] A computer-readable storage medium according to an embodiment of the present invention can execute a power grid dispatching data mining method provided by an embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0144] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.
[0145] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0146] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0147] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0148] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0149] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.
[0150] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0151] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0152] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0153] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A power grid dispatching data mining method, characterized in that: The following steps are involved: The first power grid dispatching data is screened by a density peak algorithm based on hybrid density and micro-cluster aggregation to obtain second power grid dispatching data; Classifying the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types; The third power grid dispatching data is mined through information mining using a pruning tree algorithm to obtain target effective information.
2. A power grid dispatching data mining method according to claim 1, characterized in that: The first power grid dispatching data is screened by a density peak algorithm based on mixed density and micro-cluster aggregation to obtain the second power grid dispatching data, which specifically includes: Determining the local density of each data sample of the first power grid dispatching data; Determining a sample relative distance of each of the data samples according to the local density; Determine the reverse K nearest neighbor set of each data sample by a reverse K nearest neighbor algorithm according to the relative distance of the samples; Determining the absolute density of each of the data samples according to the reverse K nearest neighbor set, and determining the attribution relationship between the data samples according to the absolute density; Determine a mixed density of each of the data samples according to the absolute density and the attribution relationship; The first power grid dispatching data is screened according to the mixed density to obtain the second power grid dispatching data.
3. A power grid dispatching data mining method according to claim 2, characterized in that: The local density is determined by the following formula: or Among them, ρ i Represents data sample x i The local density of D represents the data sample set of the first power grid dispatching data. dist(x i ,x j ) represents the data sample x i With data sample x j The Euclidean distance, d c Represents the cutoff distance.
4. A power grid dispatching data mining method according to claim 2, characterized in that: The reverse K nearest neighbor set is determined by the following formula: rnn K (x i )={x j ∈Dx i ∈KNN(x j )} KNN(x j )={x∈Dd(x j ,x)<d(x j ,x j-Kth )} Among them, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set, KNN(x j ) represents the data sample x j The K nearest neighbor set of d(x j ,x) represents the data sample x j The distance from the data sample x, x j-Kth Table distance data sample x j The K-th most recent data sample, where K is a preset value.
5. A power grid dispatching data mining method according to claim 2, characterized in that: The absolute density is determined by the following formula: Among them, P i Represents data sample x i The absolute density, rnn K (x i ) represents the data sample x i The reverse K nearest neighbor set of dist(x i ,x j ) represents the data sample x i With data sample x j The Euclidean distance.
6. A power grid dispatching data mining method according to claim 1, characterized in that: The second power grid dispatching data is classified by the Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types, which specifically includes: Construct data dissimilarity matrix; According to the data dissimilarity matrix, calculating the data type dissimilarity between each data sample of the second power grid dispatching data by using a Bray-Curtis algorithm; The second power grid dispatching data is classified according to the data type difference to obtain the third power grid dispatching data of multiple data types.
7. A power grid dispatching data mining method according to claim 1, characterized in that: The information mining of the third power grid dispatching data by using a pruning tree algorithm to obtain target valid information specifically includes: Determine a plurality of sample vectors according to the third power grid dispatching data; Classifying the sample vectors by a pruning tree algorithm, and determining a decision center probability value for cluster mining according to the classification result and the data type; The effective information category is determined according to the decision center probability value, and then the target effective information is obtained by screening from the third power grid dispatching data according to the effective information category.
8. A power grid dispatching data mining system, characterized in that: include: A data screening module, used for screening the first power grid dispatching data by using a density peak algorithm based on mixed density and micro-cluster aggregation to obtain second power grid dispatching data; A data classification module, used for classifying the second power grid dispatching data by using a Bray-Curtis algorithm to obtain third power grid dispatching data of multiple data types; The information mining module is used to perform information mining on the third power grid dispatching data through a pruning tree algorithm to obtain target effective information.
9. A power grid dispatching data mining device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a power grid dispatching data mining method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to execute a power grid dispatching data mining method as claimed in any one of claims 1 to 7 when executed by the processor.