Power grid planning risk identification method and apparatus based on big data technique, and device and storage medium
By performing dimensionality reduction and grey relational analysis on power grid planning risk data, a risk identification model was constructed, which solved the problem of insufficient utilization of historical data in power grid planning and achieved efficient and accurate risk identification.
Patent Information
- Application Number
- PCT/CN2024/125085
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2024-10-15
- Publication Date
- 2025-12-04
AI Technical Summary
Existing power grid planning neglects the changing patterns reflected in historical data of the distribution network, resulting in insufficient objectivity, accuracy, and reliability in identifying planning problems.
By collecting power grid planning risk data and performing dimensionality reduction processing, the grey relational analysis method is used to mine the correlation between power grid implementation planning projects and power grid planning data and power grid operation data, and a pre-set risk identification model is constructed for risk identification.
This improves the objectivity, accuracy, and reliability of power grid planning, ensuring that risk identification data reflects the changing patterns of historical data in the distribution network, and achieving efficient and high-precision risk identification.
Smart Images

Figure CN2024125085_04122025_PF_FP_ABST
Abstract
Description
A method, device, equipment, and storage medium for identifying power grid planning risks based on big data technology. Technical Field
[0001] This invention relates to the field of power grid risk identification technology, and in particular to a method, apparatus, equipment and storage medium for power grid planning risk identification based on big data technology. Background Technology
[0002] Currently, commonly used methods for power grid risk identification and risk analysis include: expert survey, brainstorming, fault tree analysis, and SWOT analysis. Expert survey involves consulting with experts, posing relevant questions, synthesizing and summarizing their responses, and then feeding the results back to each expert for further feedback until a relatively consistent and scientific conclusion is reached. Brainstorming is a special form of group meeting where participants freely express their ideas and concepts, fostering a chain reaction of new ideas through mutual inspiration and association, ultimately achieving a complementary and synergistic effect, thus making predictions and identifications more accurate. Fault tree analysis uses diagrams to break down large risks into smaller risks or decompose the causes of various risks, arranging project risks hierarchically from large to small, from coarse to fine, making it easier to identify all influencing factors. SWOT analysis is a method for analyzing risk environments. It is a systematic tool for analyzing risk factors, and its main purpose is to analyze and identify the strengths and weaknesses, opportunities and threats of risk from multiple perspectives.
[0003] However, traditional power grid planning methods have failed to fully utilize the value of power big data. The discovery of transmission and distribution network problems and the establishment of correlations between related influencing factors rely on the professional knowledge of planners, ignoring the changing patterns reflected in the historical data of the distribution network itself, which restricts the objectivity of finding planning problems.
[0004] Therefore, there is an urgent need for a method that can improve the objectivity, accuracy, and credibility of power distribution network planning.
[0005] Summary of the Invention
[0006] This invention provides a method, device, equipment, and storage medium for identifying risks in power grid planning based on big data technology. This addresses the technical problem in existing technologies that neglect the changing patterns reflected in historical data of the distribution network itself, which restricts the objectivity, accuracy, and reliability of planning problem identification.
[0007] To address the aforementioned technical problems, embodiments of the present invention provide a power grid planning risk identification method based on big data technology, comprising:
[0008] Collect power grid planning risk data and perform dimensionality reduction on the power grid planning risk data to obtain data to be associated; wherein, the data to be associated includes power grid planning data after dimensionality reduction of power grid planning risk data, power grid operation data, and power grid implementation planning project data;
[0009] By calculating the correlation between the power grid planning data and the power grid operation data corresponding to the data of each power grid implementation planning project, the correlation between the data of the power grid implementation planning project and the power grid planning data and the power grid operation data is obtained. Based on the correlation between the data of the power grid implementation planning project and the power grid planning data and the power grid operation data, the data with the greatest correlation with the data of the power grid implementation planning project in the power grid planning data and the power grid operation data respectively are obtained as risk identification data.
[0010] Based on the preset risk identification model, the risks of power grid implementation planning projects are identified by inputting the risk identification data.
[0011] As a preferred embodiment, the correlation calculation is performed on the power grid planning data and power grid operation data corresponding to the data of each power grid implementation planning project to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and power grid operation data, specifically including:
[0012] The power grid planning risk data is subjected to feature dimensionality reduction to obtain the data to be associated, and the feature data in the data to be associated is extracted.
[0013] The feature data is standardized to obtain feature values of the power grid planning data and power grid operation data corresponding to each project;
[0014] Based on the power grid implementation planning project data, a first sample data sequence of corresponding power grid implementation planning project data is obtained, and based on the feature values of the power grid planning data and power grid operation data corresponding to each project, a second sample data sequence of corresponding projects is obtained, composed of feature values selected from the power grid planning data and power grid operation data.
[0015] Based on the first sample data sequence and the second sample data sequence, a grey relational degree is calculated, and based on the grey relational degree, the correlation between the power grid implementation planning project data, power grid planning data, and power grid operation data corresponding to each project is obtained.
[0016] As a preferred embodiment, the step of performing feature dimensionality reduction on the power grid planning risk data to obtain data to be associated, and extracting feature data from the data to be associated, specifically involves:
[0017] The power grid planning risk data is preprocessed by using an outlier detection method to remove duplicate, abnormal and missing data in the preset dataset, thereby obtaining preprocessed power grid planning risk data.
[0018] Based on the power grid implementation planning project data in the power grid planning risk data, corresponding projects are selected from the power grid planning data and power grid operation data in the preprocessed power grid planning risk data to form a power grid planning risk data matrix.
[0019] The power grid planning risk data matrix is subjected to feature dimensionality reduction to calculate the average value of each data in the power grid planning risk data matrix. Then, the obtained power grid planning risk data matrix is decentralized based on the average value to obtain a feature data matrix, which is used as the feature data in the preset dataset.
[0020] As a preferred embodiment, the step of performing data standardization processing on the feature data to obtain feature values of the power grid planning data and power grid operation data corresponding to each project is specifically as follows:
[0021] The covariance matrix is calculated based on the feature data matrix.
[0022] The feature data is standardized using the covariance matrix to calculate the eigenvectors and eigenvalues corresponding to the feature data matrix.
[0023] As a preferred embodiment, the first sample data sequence of the corresponding power grid implementation planning project data is obtained based on the power grid implementation planning project data, and a second sample data sequence of each project is obtained based on the feature values of the power grid planning data and power grid operation data corresponding to each project, composed of feature values selected from the power grid planning data and power grid operation data. Specifically:
[0024] Based on the power grid implementation planning project data, a first sample data sequence of the corresponding power grid implementation planning project data is obtained;
[0025] Based on the magnitude of the characteristic values of power grid planning data and power grid operation data, a preset first number of characteristic values is calculated as the information contribution rate, and based on the information contribution rate, a preset second number of characteristic values is obtained.
[0026] Based on the obtained second set of feature values, a second sample data sequence corresponding to each project is obtained, consisting of power grid planning data and / or power grid operation data.
[0027] As a preferred embodiment, the calculation of the grey relational degree based on the first sample data sequence and the second sample data sequence specifically includes: X0=[x0(1), x0(2), ...x0(n)] X i =[x i (1), x i (2), ...x i (n)]
[0028] Where X0 is the first sample data sequence, x0(n) is the feature value of the nth item in the first sample data sequence, and X i Let x be the first sample data sequence. i (n) represents the feature value of the i-th data point of the n-th item. For the normalized sequence X i The eigenvalue of the k-th term; x i,max x i,min These represent the maximum and minimum values of the feature values of the i-th data point, respectively, where ζ is the resolution coefficient, and γ(X0, X...) i ) represents the grey relational degree.
[0029] As a preferred embodiment, the method for constructing the preset risk identification model includes:
[0030] An initial preset risk identification model is constructed, and based on a preset training dataset, the initial preset risk identification model is iteratively trained to obtain a preset risk identification model; wherein, the preset training dataset includes data of each historical power grid implementation planning project with a gray correlation degree greater than a preset value, as well as the corresponding historical power grid planning data and historical power grid operation data.
[0031] Accordingly, the present invention also provides a power grid planning risk identification device, comprising: a data acquisition module, an association module, and an identification module;
[0032] The acquisition module is used to acquire power grid planning risk data and reduce the dimensionality of the power grid planning risk data to obtain data to be associated; wherein, the data to be associated includes power grid planning data after dimensionality reduction of power grid planning risk data, power grid operation data, and power grid implementation planning project data;
[0033] The correlation module is used to calculate the correlation degree through the power grid planning data and power grid operation data corresponding to the power grid implementation planning project data, to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, and to obtain the data with the greatest correlation with the power grid implementation planning project data in the power grid planning data and the power grid operation data respectively, based on the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, as risk identification data.
[0034] The identification module is used to identify the risks of power grid implementation planning projects based on a preset risk identification model and the input of the risk identification data.
[0035] Accordingly, the present invention also provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the power grid planning risk identification method based on big data technology as described above.
[0036] Accordingly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the power grid planning risk identification method based on big data technology as described above.
[0037] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0038] The technical solution of this invention reduces the dimensionality of collected power grid planning risk data, thereby avoiding the curse of dimensionality caused by high-dimensional data during the calculation process. It also removes irrelevant data from the dataset, improving computational efficiency. Based on the dimensionality-reduced data, a grey relational analysis method is used to mine the correlation between relevant feature data of power grid implementation planning projects and power grid planning data and power grid operation data. This improves the objectivity, accuracy, and reliability of the collected big data on power grid planning and operation data, providing a valid basis for the formulation of power grid planning projects. This enhances the efficiency and precision of power grid planning and risk identification, ensuring that the power grid planning risk data used for risk identification reflects the changing patterns of historical distribution network data. Furthermore, a pre-set risk identification model is used to identify risks in power grid planning based on correlated data, thus achieving efficient and high-precision risk identification in power grid planning based on big data technology. Attached Figure Description
[0039] Figure 1: A flowchart of the steps of a power grid planning risk identification method based on big data technology provided in an embodiment of the present invention;
[0040] Figure 2: A structural diagram of a power grid planning risk identification device based on big data technology provided in an embodiment of the present invention. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Example 1
[0043] Please refer to Figure 1, which illustrates a power grid planning risk identification method based on big data technology according to an embodiment of the present invention, including the following steps S101-S103:
[0044] Step S101: Collect power grid planning risk data and reduce the dimensionality of the power grid planning risk data to obtain data to be associated; wherein, the data to be associated includes power grid planning data after dimensionality reduction of power grid planning risk data, power grid operation data, and power grid implementation planning project data.
[0045] In this embodiment, to improve the efficiency and accuracy of power grid planning risk identification, the data is preprocessed to remove duplicate, abnormal, and missing data. After collecting power grid planning risk data, initial dimensionality reduction is performed on the collected data. Preferably, the Z-score method is used for outlier detection, and the calculation formula is as follows:
[0046] Where, x i denoted as z; μ is the mean of all data; δ is the standard deviation of all data values. i The absolute value of |z| represents the distance between the score within the standard deviation and the population mean. i If the value is greater than the threshold, it can be identified as an outlier.
[0047] It should be noted that, regarding the identification of risk sources in power grid planning, these risks can be categorized into three types: technical risk sources, economic risk sources, and management risk sources. Technical risk sources include the accuracy of load forecasting, the impact and regulations of distributed energy and energy storage construction, and the applicability of standards and new technologies. Economic risk sources include operational risks arising from unreasonable transmission and distribution pricing, and fluctuations in equipment and material procurement prices. Management risk sources include the professional competence of planning staff, the number of personnel allocated to planning positions, and planning contingency plans.
[0048] Furthermore, regarding the collection of power grid planning risk data, since power grid planning risk identification is based on the analysis of power grid planning data, power grid operation data, and planning project implementation data to identify risks in future planning, it is necessary to collect relevant data for power grid planning risk identification. The relevant data for power grid planning risk identification is mainly divided into three categories: power grid planning data, power grid operation data, and power grid fault data. In this embodiment, since power grid faults also generate power grid fault data and simultaneously increase power grid risks, power grid fault data represents the fault phenomena that occur after power grid risks exist. Therefore, power grid planning data and power grid operation data are mainly used as the key data for power grid planning risk identification.
[0049] In this embodiment, the power grid planning data includes weather data, geographic location data, regional characteristic data, personnel allocation data, power grid project quotation data, and planned construction scale data; the power grid operation data includes distribution transformer capacity data, transformer operation data, transmission line operation data, cable operation data, and load switch operation data; and the power grid implementation planning project data includes data on projects already in operation, projects that have been postponed, and projects that have been cancelled.
[0050] Step S102: Calculate the correlation between the power grid implementation planning project data and the power grid operation data corresponding to the power grid implementation planning project data, and obtain the correlation between the power grid implementation planning project data and the power grid planning data and the power grid operation data. Based on the correlation between the power grid implementation planning project data and the power grid planning data and the power grid operation data, obtain the data with the highest correlation with the power grid implementation planning project data in the power grid planning data and the power grid operation data, respectively, and use it as risk identification data.
[0051] As a preferred embodiment, the step of calculating the correlation between the power grid implementation planning project data and the power grid planning data and power grid operation data corresponding to the power grid implementation planning project data specifically includes:
[0052] The power grid planning risk data is subjected to feature dimensionality reduction to obtain data to be associated, and feature data is extracted from the data to be associated. The feature data is then subjected to data standardization to obtain feature values of the power grid planning data and power grid operation data corresponding to each project. Based on the power grid implementation planning project data, a first sample data sequence of the corresponding power grid implementation planning project data is obtained, and based on the feature values of the power grid planning data and power grid operation data corresponding to each project, a second sample data sequence of the corresponding projects is obtained, composed of feature values selected from the power grid planning data and power grid operation data. Based on the first sample data sequence and the second sample data sequence, a grey relational degree is calculated, and based on the grey relational degree, the association relationship between the power grid implementation planning project data and the power grid planning data and power grid operation data corresponding to each project is obtained.
[0053] As a preferred embodiment, the step of performing feature dimensionality reduction on the power grid planning risk data to obtain data to be associated, and extracting feature data from the data to be associated, specifically involves:
[0054] The power grid planning risk data is preprocessed using an outlier detection method to remove duplicate, abnormal, and missing data from the preset dataset, resulting in preprocessed power grid planning risk data. Based on the power grid implementation planning project data within the preprocessed data, corresponding projects are selected from the power grid planning data and power grid operation data to form a power grid planning risk data matrix. Feature dimensionality reduction is then performed on the power grid planning risk data matrix to calculate the average value of each data point. This average value is then used to decentralize the resulting power grid planning risk data matrix, yielding a feature data matrix, which serves as the feature data in the preset dataset.
[0055] In this embodiment, since power grid data is mostly high-dimensional data, the resulting dataset has a large number of features, including potentially redundant fault feature variables and weakly correlated variables. Feature dimensionality reduction can effectively eliminate irrelevant and redundant features, improve identification efficiency, and enhance the understandability of correlated features. Preferably, principal component analysis is mainly used to extract the main features from the data.
[0056] In this embodiment, the collected power grid planning risk data is preprocessed to remove duplicate, abnormal, and missing data. Then, n items are selected from the power grid planning risk data, and each item contains m types of data, forming a power grid planning risk data matrix, represented by matrix A:
[0057] Among them, a nmLet be the m-th data value of the n-th project. The power grid planning data, power grid operation data, and power grid implementation planning project data can all be represented in the form of matrix A.
[0058] Next, the average value of each data point in matrix A is calculated using the following formula:
[0059] in, Let a be the average value of the i-th feature data across n planning projects; ji Let i be the i-th feature data in the j-th planning project.
[0060] Therefore, matrix A is decentered to obtain the feature data matrix, represented by matrix B, and the calculation formula is as follows:
[0061] As a preferred embodiment, the step of performing data standardization processing on the feature data to obtain feature values of the power grid planning data and power grid operation data corresponding to each project specifically involves:
[0062] Based on the feature data matrix, the covariance matrix is calculated; the feature data is then standardized using the covariance matrix to calculate the eigenvectors and eigenvalues corresponding to the feature data matrix.
[0063] In this embodiment, the covariance matrix C is solved from matrix B, and the calculation formula is as follows:
[0064] The eigenvalues and corresponding eigenvectors of matrix C are calculated using the following formula: (λE-C)v=0
[0065] Where λ is the eigenvalue and v is the eigenvector corresponding to λ.
[0066] Then, select a predetermined first number (preferably k) of eigenvalues from all eigenvalues in descending order, and use the eigenvectors corresponding to the k eigenvalues as the eigenvector matrix X. e The row vectors are then used as the subsequent sample data sequence to calculate the eigenvalue λ. i The information contribution rate of (i = 1, 2, ..., k) is calculated using the following formula:
[0067] Where, β i For the eigenvalue λ i The contribution rate; The sum of eigenvalues, and then based on each eigenvalue λ iThe information contribution rate of (i = 1, 2, ..., k) is used to select a preset second number (preferably, the preset second number is Q) of optimal data as variable data in the sequence.
[0068] As a preferred embodiment, the step of obtaining a first sample data sequence of power grid implementation planning project data based on the power grid implementation planning project data, and obtaining a second sample data sequence of each project composed of feature values selected from the power grid planning data and power grid operation data based on the feature values of each project's corresponding power grid planning data and power grid operation data, specifically involves:
[0069] Based on the power grid implementation planning project data, a first sample data sequence of corresponding power grid implementation planning project data is obtained; according to the size of the feature values of the power grid planning data and power grid operation data, a preset first number of feature values is calculated as the information contribution rate, and a preset second number of feature values is obtained according to the information contribution rate; according to the obtained preset second number of feature values, a second sample data sequence of corresponding projects is obtained, consisting of power grid planning data and / or power grid operation data.
[0070] It should be noted that the power grid implementation planning project data includes data on projects already in operation, projects that have been postponed, and projects that have been cancelled. Therefore, based on one or more of the data on projects already in operation, projects that have been postponed, and / or projects that have been cancelled, the first sample data sequence of the corresponding power grid implementation planning project data can be obtained. For example, taking the number of days of project implementation delay from the postponed project data as an example, the first sample data sequence is sequence X0: X0 = [x0(1), x0(2), ...x0(n)]
[0071] Where x0(n) is the number of days the implementation of the nth project is delayed.
[0072] For example, the data selected from power grid planning data and power grid operation data can be denoted as the second sample data sequence X. i The calculation formula is as follows: X i =[x i (1), x i (2), ...x i (n)]
[0073] Where x i (n) represents the i-th data value of the n-th item.
[0074] As a preferred embodiment, the step of calculating the grey relational degree based on the first sample data sequence and the second sample data sequence specifically includes: X0=[x0(1), x0(2), ...x0(n)] X i =[xi (1), x i (2), ...x i (n)]
[0075] Where X0 is the first sample data sequence, x0(n) is the feature value of the nth item in the first sample data sequence, and X i Let x be the first sample data sequence. i (n) represents the feature value of the i-th data point of the n-th item. For the normalized sequence X i The eigenvalue of the k-th term; x i,max x i,min These represent the maximum and minimum values of the feature values of the i-th data point, respectively, where ζ is the resolution coefficient, and γ(X0, X...) i ) represents the grey relational degree.
[0076] In this embodiment, since the influence of each type of data is related to its value range, the variable data is normalized to calculate the normalized second sample data sequence X. i of
[0077] In this embodiment, the grey relational degree of each data sequence obtained above is sorted. The higher the grey relational degree, the greater the correlation between the data and the number of days the project implementation is delayed.
[0078] In this embodiment, the correlation between power grid planning data, operation data and power grid planning project implementation data is mined. By measuring the degree of grey relational analysis between the data, the relationship between the implementation risks of power grid planning projects and planning and operation data can be identified. Subsequently, the correlation between relevant feature data of power grid planning project implementation and relevant electrical feature indicators can be mined using the grey relational analysis method to obtain the corresponding power grid planning risk data, i.e. risk identification data, and thus identify the implementation risks of power grid planning projects.
[0079] Step S103: Based on the preset risk identification model, the risks of the power grid implementation planning project are identified by inputting the risk identification data.
[0080] As a preferred embodiment, the method for constructing the preset risk identification model includes:
[0081] An initial preset risk identification model is constructed, and based on a preset training dataset, the initial preset risk identification model is iteratively trained to obtain a preset risk identification model; wherein, the preset training dataset includes data of each historical power grid implementation planning project with a gray correlation degree greater than a preset value, as well as the corresponding historical power grid planning data and historical power grid operation data.
[0082] In this embodiment, the initial preset risk identification model can be a convolutional neural network model. By collecting data related to the power grid operation status as a preset training dataset, that is, by pre-collecting and calculating the gray correlation degree greater than a preset value for each historical power grid implementation planning project data, as well as the corresponding historical power grid planning data and historical power grid operation data, the preset training dataset is used as the training data for the convolutional neural network model in the initial preset risk identification model, thereby training the preset risk identification model.
[0083] Furthermore, during the iterative training of the initial pre-set risk identification model, model validation is also required. The trained model is validated using a test set to evaluate its generalization ability and the accuracy of risk identification. Subsequently, the iteratively trained pre-set risk identification model is optimized; that is, based on the model validation results, the model is further adjusted and optimized to obtain the final pre-set risk identification model.
[0084] In this embodiment, power grid planning data, as the research object, is characterized by multiple data types, large data volume, and ambiguous relationships between data. Considering that risk identification modeling based on only two or three feature data is ineffective and lacks interpretability, as much information as possible should be collected during the data acquisition process to accurately describe the problem and enhance the persuasiveness of the research results. However, large-scale data collection can easily generate a large amount of high-dimensional data. Principal component analysis can be used to reduce the dimensionality of high-dimensional power grid data, avoid the curse of dimensionality caused by high-dimensional data during the calculation process, and remove irrelevant data from the dataset to improve computational efficiency.
[0085] Based on the selected feature data, the grey relational analysis method is used to mine the correlation between the feature data related to the implementation of power grid planning projects and the relevant electrical feature indicators, providing an effective basis for the formulation of power grid planning projects, thereby improving the efficiency and precision of power grid planning, and ensuring that the method of power grid planning risk identification based on big data technology can be objective, accurate and reliable.
[0086] Implementing the above embodiments has the following effects:
[0087] The technical solution of this invention reduces the dimensionality of collected power grid planning risk data, thereby avoiding the curse of dimensionality caused by high-dimensional data during the calculation process. It also removes irrelevant data from the dataset, improving computational efficiency. Based on the dimensionality-reduced data, a grey relational analysis method is used to mine the correlation between relevant feature data of power grid implementation planning projects and power grid planning data and power grid operation data. This improves the objectivity, accuracy, and reliability of the collected big data on power grid planning and operation data, providing a valid basis for the formulation of power grid planning projects. This enhances the efficiency and precision of power grid planning and risk identification, ensuring that the power grid planning risk data used for risk identification reflects the changing patterns of historical distribution network data. Furthermore, a pre-set risk identification model is used to identify risks in power grid planning based on correlated data, thus achieving efficient and high-precision risk identification in power grid planning based on big data technology.
[0088] Example 2
[0089] Please refer to Figure 2, which shows the power grid planning risk identification device based on big data technology provided by the present invention, including: a data acquisition module 201, an association module 202, and an identification module 203;
[0090] The acquisition module 201 is used to acquire power grid planning risk data and reduce the dimensionality of the power grid planning risk data to obtain data to be associated; wherein, the data to be associated includes power grid planning data after dimensionality reduction of power grid planning risk data, power grid operation data, and power grid implementation planning project data;
[0091] The association module 202 is used to calculate the correlation degree through the power grid planning data and power grid operation data corresponding to the power grid implementation planning project data, to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, and to obtain the data with the greatest correlation with the power grid implementation planning project data in the power grid planning data and the power grid operation data respectively, based on the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, as risk identification data.
[0092] The identification module 203 is used to identify the risks of power grid implementation planning projects based on a preset risk identification model and the input of the risk identification data.
[0093] As a preferred embodiment, the correlation calculation is performed on the power grid planning data and power grid operation data corresponding to the data of each power grid implementation planning project to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and power grid operation data, specifically including:
[0094] The power grid planning risk data is subjected to feature dimensionality reduction to obtain the data to be associated, and the feature data in the data to be associated is extracted.
[0095] The feature data is standardized to obtain feature values of the power grid planning data and power grid operation data corresponding to each project;
[0096] Based on the power grid implementation planning project data, a first sample data sequence of corresponding power grid implementation planning project data is obtained, and based on the feature values of the power grid planning data and power grid operation data corresponding to each project, a second sample data sequence of corresponding projects is obtained, composed of feature values selected from the power grid planning data and power grid operation data.
[0097] Based on the first sample data sequence and the second sample data sequence, a grey relational degree is calculated, and based on the grey relational degree, the correlation between the power grid implementation planning project data, power grid planning data, and power grid operation data corresponding to each project is obtained.
[0098] As a preferred embodiment, the step of performing feature dimensionality reduction on the power grid planning risk data to obtain data to be associated, and extracting feature data from the data to be associated, specifically involves:
[0099] The power grid planning risk data is preprocessed by using an outlier detection method to remove duplicate, abnormal and missing data in the preset dataset, thereby obtaining preprocessed power grid planning risk data.
[0100] Based on the power grid implementation planning project data in the power grid planning risk data, corresponding projects are selected from the power grid planning data and power grid operation data in the preprocessed power grid planning risk data to form a power grid planning risk data matrix.
[0101] The power grid planning risk data matrix is subjected to feature dimensionality reduction to calculate the average value of each data in the power grid planning risk data matrix. Then, the obtained power grid planning risk data matrix is decentralized based on the average value to obtain a feature data matrix, which is used as the feature data in the preset dataset.
[0102] As a preferred embodiment, the step of performing data standardization processing on the feature data to obtain feature values of the power grid planning data and power grid operation data corresponding to each project is specifically as follows:
[0103] The covariance matrix is calculated based on the feature data matrix.
[0104] The feature data is standardized using the covariance matrix to calculate the eigenvectors and eigenvalues corresponding to the feature data matrix.
[0105] As a preferred embodiment, the first sample data sequence of the corresponding power grid implementation planning project data is obtained based on the power grid implementation planning project data, and a second sample data sequence of each project is obtained based on the feature values of the power grid planning data and power grid operation data corresponding to each project, composed of feature values selected from the power grid planning data and power grid operation data. Specifically:
[0106] Based on the power grid implementation planning project data, a first sample data sequence of the corresponding power grid implementation planning project data is obtained;
[0107] Based on the magnitude of the characteristic values of power grid planning data and power grid operation data, a preset first number of characteristic values is calculated as the information contribution rate, and based on the information contribution rate, a preset second number of characteristic values is obtained.
[0108] Based on the obtained second set of feature values, a second sample data sequence corresponding to each project is obtained, consisting of power grid planning data and / or power grid operation data.
[0109] As a preferred embodiment, the calculation of the grey relational degree based on the first sample data sequence and the second sample data sequence specifically includes: X0=[x0(1), x0(2), ...x0(n)] X i =[x i (1), x i (2), ...x i (n)]
[0110] Where X0 is the first sample data sequence, x0(n) is the feature value of the nth item in the first sample data sequence, and X i Let x be the first sample data sequence. i (n) represents the feature value of the i-th data point of the n-th item. For the normalized sequence X i The eigenvalue of the k-th term; x i,max x i,min These represent the maximum and minimum values of the feature values of the i-th data point, respectively, where ζ is the resolution coefficient, and γ(X0, X...) i ) represents the grey relational degree.
[0111] As a preferred embodiment, the method for constructing the preset risk identification model includes:
[0112] An initial preset risk identification model is constructed, and based on a preset training dataset, the initial preset risk identification model is iteratively trained to obtain a preset risk identification model; wherein, the preset training dataset includes data of each historical power grid implementation planning project with a gray correlation degree greater than a preset value, as well as the corresponding historical power grid planning data and historical power grid operation data.
[0113] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0114] Implementing the above embodiments has the following effects:
[0115] The technical solution of this invention reduces the dimensionality of collected power grid planning risk data, thereby avoiding the curse of dimensionality caused by high-dimensional data during the calculation process. It also removes irrelevant data from the dataset, improving computational efficiency. Based on the dimensionality-reduced data, a grey relational analysis method is used to mine the correlation between relevant feature data of power grid implementation planning projects and power grid planning data and power grid operation data. This improves the objectivity, accuracy, and reliability of the collected big data on power grid planning and operation data, providing a valid basis for the formulation of power grid planning projects. This enhances the efficiency and precision of power grid planning and risk identification, ensuring that the power grid planning risk data used for risk identification reflects the changing patterns of historical distribution network data. Furthermore, a pre-set risk identification model is used to identify risks in power grid planning based on correlated data, thus achieving efficient and high-precision risk identification in power grid planning based on big data technology.
[0116] Example 3
[0117] Accordingly, the present invention also provides a terminal device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the power grid planning risk identification method based on big data technology as described in any of the above embodiments.
[0118] The terminal device of this embodiment includes a processor, a memory, and a computer program and computer instructions stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps in Embodiment 1 above, such as steps S101 to S103 shown in FIG1. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above device embodiment, such as the associated module 202.
[0119] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the terminal device. For example, the association module 202 is used to calculate the correlation between the power grid implementation planning project data and the power grid planning data and power grid operation data corresponding to the power grid implementation planning project data, and based on the correlation between the power grid implementation planning project data and the power grid planning data and power grid operation data, obtain the data with the highest correlation to the power grid implementation planning project data in the power grid planning data and power grid operation data, respectively, as risk identification data.
[0120] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the schematic diagram is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0121] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0122] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0123] Wherein, if the modules / units integrated in the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the various method embodiments described above. Wherein, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0124] Example 4
[0125] Accordingly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the power grid planning risk identification method based on big data technology as described in any of the above embodiments.
[0126] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A power grid planning risk identification method based on big data technology, characterized in that, The method comprises the following steps: Collect power grid planning risk data, and reduce the dimension of the power grid planning risk data to obtain to-be-associated data; wherein the to-be-associated data comprises power grid planning data, power grid operation data and power grid implementation planning project data obtained by reducing the dimension of the power grid planning risk data; Calculate the correlation degree of the power grid planning data and the power grid operation data corresponding to each power grid implementation planning project data, obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, and obtain the data most relevant to the power grid implementation planning project data in the power grid planning data and the power grid operation data respectively as risk identification data according to the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data. According to a preset risk identification model, the risk of the power grid implementation planning project is identified through input of the risk identification data.
2. The power grid planning risk identification method based on big data technology according to claim 1, characterized in that, The correlation degree calculation of the power grid planning data and the power grid operation data corresponding to each power grid implementation planning project data to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data comprises the following steps: Feature dimension reduction is performed on the power grid planning risk data to obtain to-be-associated data, and feature data in the to-be-associated data is extracted; Data standardization processing is performed on the feature data to obtain feature values of the power grid planning data and the power grid operation data corresponding to each project; Based on the power grid implementation planning project data, a first sample data sequence corresponding to the power grid implementation planning project data is obtained, and based on the feature values of the power grid planning data and the power grid operation data corresponding to each project, a second sample data sequence corresponding to each project is obtained, which is composed of feature values selected from the power grid planning data and the power grid operation data; According to the first sample data sequence and the second sample data sequence, the gray correlation degree is calculated, so that the correlation relationship between the power grid implementation planning project data corresponding to each project and the power grid planning data and the power grid operation data is obtained based on the gray correlation degree.
3. The power grid planning risk identification method based on big data technology according to claim 2, characterized in that, The feature dimension reduction of the power grid planning risk data to obtain to-be-associated data and the extraction of feature data in the to-be-associated data comprises the following steps: Through an outlier detection method, data preprocessing is performed on the power grid planning risk data, so that repeated, abnormal and missing data in the preset data set are eliminated to obtain preprocessed power grid planning risk data; According to the power grid implementation planning project data in the power grid planning risk data, the power grid planning data and the power grid operation data in the preprocessed power grid planning risk data are selected according to corresponding projects to form a power grid planning risk data matrix; Feature dimension reduction is performed on the power grid planning risk data matrix, so that the average value of each data in the power grid planning risk data matrix is calculated, and then the power grid planning risk data matrix is decentered according to the average value to obtain a feature data matrix as the feature data in the preset data set.
4. The power grid planning risk identification method based on big data technology according to claim 3, characterized in that, The data standardization processing of the feature data to obtain the feature values of the power grid planning data and the power grid operation data corresponding to each project comprises the following steps: According to the feature data matrix, a covariance matrix is calculated; The feature data is subjected to data standardization processing through the covariance matrix, so that a feature vector corresponding to the feature data matrix and a corresponding feature value are calculated.
5. The power grid planning risk identification method based on big data technology according to claim 4, characterized in that, The first sample data sequence corresponding to the power grid implementation planning project data is obtained based on the power grid implementation planning project data, and the second sample data sequence corresponding to each project is obtained based on the feature values of the power grid planning data and the power grid operation data selected from the power grid planning data and the power grid operation data, specifically: The first sample data sequence corresponding to the power grid implementation planning project data is obtained based on the power grid implementation planning project data; According to the size of the feature values of the power grid planning data and the power grid operation data, the information contribution rate of the preset first number of feature values is calculated, and the preset second number of feature values is obtained according to the information contribution rate. According to the obtained preset second number of feature values, the second sample data sequence corresponding to each project composed of the power grid planning data and / or the power grid operation data is obtained.
6. The power grid planning risk identification method based on big data technology according to claim 5, characterized in that, The grey correlation degree is calculated according to the first sample data sequence and the second sample data sequence, and specifically includes: X0=[x0(1), x0(2),...x0(n)] i X i 1=[x i 1(1), x i 1(2),...x wherein X0is the first sample data sequence, x0(n) is a feature value of an nth item in the first sample data sequence, X i is the first sample data sequence, x i (n) is a feature value of an ith data of the nth item, Xk is the kth term of the normalized sequence X i Xk is the kth term of the normalized sequence X i,max Xk is the kth term of the normalized sequence X i,min Xk is the kth term of the normalized sequence X i Xk is the kth term of the normalized sequence X 7. The power grid planning risk identification method based on big data technology according to claim 6, characterized in that, The construction method of the preset risk identification model comprises: An initial preset risk identification model is constructed, and the initial preset risk identification model is iteratively trained based on a preset training data set, so that the preset risk identification model is obtained; wherein the preset training data set comprises each historical power grid implementation planning project data whose gray correlation degree is greater than a preset value, and corresponding historical power grid planning data and historical power grid operation data.
8. A power grid planning risk identification device based on big data technology, characterized in that, It comprises: An acquisition module, an association module, and an identification module; The acquisition module is configured to acquire power grid planning risk data, and perform dimension reduction on the power grid planning risk data to obtain to-be-associated data; wherein the to-be-associated data comprises power grid planning data, power grid operation data, and power grid implementation planning project data obtained by performing dimension reduction on the power grid planning risk data; The association module is configured to calculate the correlation degrees of the power grid planning data and the power grid operation data corresponding to each power grid implementation planning project data, to obtain the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data, and to obtain, as risk identification data, the data in the power grid planning data and the power grid operation data that are most relevant to the power grid implementation planning project data according to the correlation relationship between the power grid implementation planning project data and the power grid planning data and the power grid operation data; The identification module is configured to identify the risk of the power grid implementation planning project according to a preset risk identification model through input of the risk identification data.
9. A terminal device, comprising: The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the power grid planning risk identification method based on big data technology according to any one of claims 1 to 7 when the computer program runs.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the power grid planning risk identification method based on big data technology according to any one of claims 1 to 7 when the computer program runs.
Citation Information
Patent Citations
Power grid planning risk evaluation system and method based on grey correlation degree TOPSIS (Technique for Order Preference by Similarity to an Ideal Solution)
CN105023065A
Risk identification method and device, equipment and storage medium
CN114202337A
Enterprise risk quantitative evaluation method and device, electronic equipment and storage medium
CN116468535A
Power grid planning risk identification method and device based on big data technology, equipment and storage medium
CN118278750A
Distribution network risk identification system and method and computer storage medium
US20190305589A1