A power distribution network planning evaluation method based on big data analysis

By performing spatiotemporal consistency cleaning and fusion of multi-dimensional data in distribution network planning and evaluation, and combining a planning scenario deduction model based on physical constraints and data-driven rules, the problems of inconsistent data integration and insufficient model simulation in existing technologies have been solved, enabling accurate evaluation and risk identification of the distribution network.

CN122114756APending Publication Date: 2026-05-29MINJIANG UNIVERSITY +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MINJIANG UNIVERSITY
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In the current distribution network planning and evaluation, multi-dimensional data is not integrated, spatiotemporal consistency is poor, standardized data carriers cannot be generated, the model does not have an iterative solution module, and potential weak links and risk patterns cannot be accurately identified.

Method used

By acquiring multi-dimensional data from sources, grids, loads, and storage, and performing spatiotemporal consistency cleaning and multi-scale fusion, a standardized spatiotemporal fusion data cube is generated. Combined with a planning scenario deduction model that couples physical constraints and data-driven rules, future development scenarios are simulated, iterative solutions and pattern matching are performed, and weak links and risk patterns are identified.

Benefits of technology

It enables accurate assessment of power distribution network planning, identifies potential weaknesses and risk patterns, provides in-depth analysis, avoids data adaptation conflicts, and simulates operating conditions in accordance with actual working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114756A_ABST
    Figure CN122114756A_ABST
Patent Text Reader

Abstract

The application discloses a power distribution network planning evaluation method based on big data analysis, and relates to the technical field of power distribution network planning big data, which comprises the following steps: acquiring source network load storage multi-dimensional historical and real-time data of a planning area; constructing an original data pool covering distributed power output sequence, load change curve, network topology connection relationship, equipment operation state log and meteorological environment information; performing space-time consistency cleaning and multi-scale fusion processing; and generating a standardized space-time fusion data cube supporting reasoning calculation and state inversion. Key feature index sets of planning are extracted to input a planning scene deduction model coupled with physical constraints and data-driven rules; implicit equations of energy flow, equipment action and network evolution are iteratively solved through a reasoning calculation module; and future multi-scene operation states are simulated and matched with historical typical state modes. The method can accurately identify weak links and risk modes of the power distribution network, and improve the accuracy and reliability of planning evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of big data technology for power distribution network planning, specifically a power distribution network planning and evaluation method based on big data analysis. Background Technology

[0002] Current distribution network planning and assessment work relies on distributed generation output sequences, load change curves, network topology connections, equipment operation status logs, and meteorological information as analytical bases. Data collection and analysis are fragmented and independent, failing to integrate multi-dimensional historical and real-time data from sources, grids, loads, and storage into a unified raw data pool. Data application is limited to discrete retrieval and localized analysis. Existing data processing methods do not perform spatiotemporal consistency cleaning for multi-source heterogeneous data, resulting in biases, redundancy, and missing data across different spatiotemporal dimensions. Furthermore, multi-scale fusion processing is lacking, failing to generate standardized data carriers suitable for inference calculations and state inversion. Consequently, the data's standardization and completeness are insufficient to meet the needs of refined distribution network assessment.

[0003] Existing distribution network planning scenario simulation models rely solely on single physical constraints or data-driven rules. These models lack iterative reasoning and computation modules, making it impossible to simulate future distribution network operation under various scenarios by solving implicit equations related to energy flow, equipment behavior, and network evolution. Current assessment methods do not perform pattern matching between simulated and historical typical operating states, hindering the accurate identification of potential weaknesses in distribution network planning and making it difficult to recognize hidden risk patterns. This invention aims to achieve accurate assessment of distribution network planning through iterative reasoning and state matching, by completing spatiotemporal fusion processing of multi-dimensional data, constructing a standardized spatiotemporal fusion data cube, and building a simulation model coupled with dual rules. Summary of the Invention

[0004] This invention aims to solve at least one of the technical problems existing in the prior art; Therefore, this invention proposes a distribution network planning and evaluation method based on big data analysis, comprising: The system acquires multi-dimensional historical and real-time data on power generation, grid, load and storage in the planning area to form a raw data pool for distribution network assessment. The raw data pool for distribution network assessment includes distributed power output sequences, load change curves, network topology connections, equipment operation status logs and meteorological information. Spatiotemporal consistency cleaning and multi-scale fusion processing are performed on the original data pool of the power distribution network assessment to generate a standardized spatiotemporal fusion data cube. The standardized spatiotemporal fusion data cube is used to support subsequent inference calculations and state inversion processes. From the standardized spatiotemporal fusion data cube, an index set for characterizing key features of distribution network planning is extracted, and the index set is input into the planning scenario inference model, which operates based on coupled physical constraints and data-driven rules. Within the planning scenario simulation model, an inference calculation module is used to simulate the operating state of the power distribution network under different future development scenarios. The inference calculation module achieves this by iteratively solving a set of implicit equations describing energy flow, equipment operation, and network evolution. The simulated operating status output by the inference computing module is matched with the historical typical operating status extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns.

[0005] Furthermore, spatiotemporal consistency cleaning and multi-scale fusion processing are performed on the original data pool for the distribution network assessment to generate a standardized spatiotemporal fusion data cube, including: For the distributed power generation output sequence and the load change curve, the sliding window interpolation method is used to fill in the missing data points, and statistical tests are used to remove abnormal fluctuation points to form a continuous and smooth time series curve. Based on the network topology connection relationship and the device operation status log, establish a mapping relationship table between devices and nodes, lines and branches, and align all device event logs to a unified time coordinate axis; Features strongly correlated with the operation of the power distribution network are extracted from the meteorological environment information, including temperature, humidity, wind speed and light intensity, and then resampled to the same time resolution as the output sequence of the distributed power source. A multi-source data association engine is established, which associates and splices the processed distributed power output sequence, the load change curve, the network topology connection relationship, the equipment operation status log and the meteorological environment features based on the same timestamp and the same spatial location encoding. The concatenated data is then reorganized according to the time dimension, spatial location dimension, and data category dimension to construct a standardized spatiotemporal fusion data cube with a unified index.

[0006] Furthermore, within the planning scenario simulation model, an inference calculation module is used to simulate the operating state of the distribution network under different future development scenarios, including: Define a set of future development scenarios, each of which is uniquely determined by a combination of load growth rate, distributed power penetration rate, new equipment commissioning plan and policy control parameters; In the planning scenario simulation model, the complete model of the current distribution network is loaded, including the node admittance matrix, transformer tap parameters, line impedance parameters and protection device settings; The typical daily load curve and the typical output curve of distributed power sources extracted from the standardized spatiotemporal fusion data cube are scaled and shaped according to the parameters of each future development scenario to generate predictive input data for the future development scenario. The reasoning and calculation module is driven by the predicted input data and constrained by the complete model of the distribution network. It solves a set of equations including power flow balance equation, equipment operation limit equation and protection action logic equation to obtain the voltage distribution, branch power flow distribution and equipment load rate of the whole network in each future time period. For each defined future development scenario, the operations of loading the distribution network model, generating predictive input data, and solving the operating state equations are executed sequentially. After traversing all future development scenarios, the complete simulated operating state under each scenario is stored.

[0007] Furthermore, the simulated operating state output by the inference computing module is pattern matched with the historical typical operating states extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns, including: From the simulated operating status output by the inference calculation module, a series of system-level and equipment-level status indicators are calculated, including node voltage over-limit duration, line overload probability, transformer load rate extreme value, and system network loss level. From the standardized spatiotemporal fusion data cube, operational data for periods during which alarms or faults occurred in the past are extracted, and the system-level and device-level status indicators of the same set are calculated to form a set of historical risk status indicators. Cluster analysis is used to compare the simulated state indicators with the historical risk state indicator set. Simulated states whose indicator combination features fall into the same cluster are identified as having similar risks. The identified simulated states with similar risks are located to specific power grid equipment, lines, and nodes, and these equipment, lines, and nodes are marked as potential weak points. The conditions for occurrence and the evolution process of the simulated states with similar risks are analyzed, and they are summarized into one or more specific risk patterns. A descriptive label is assigned to each risk pattern.

[0008] Furthermore, it also includes: Based on the identified potential weaknesses and risk patterns, a state inversion algorithm is initiated. The goal of the state inversion algorithm is to backtrack from the current network state to derive the key planning decision sequence and combination of external conditions that led to the potential weaknesses and risk patterns. The state inversion algorithm completes the backtracking derivation by constructing and solving an inverse optimization problem with planning and decision as optimization variables and approximating the actual observed state as the objective. The state inversion algorithm outputs a planning decision impact tracing report, which details the impact paths and contributions of each historical planning decision on the current network state. By combining the forward-looking simulation results of the planning scenario model with the retrospective analysis conclusions of the planning decision impact tracing report, a distribution network planning scheme to be evaluated is cross-validated in multiple rounds and from multiple perspectives. Based on the problems and conflicts exposed during the cross-validation process, a list of quantitative evaluation conclusions and structured adjustment suggestions for the power distribution network planning scheme is generated. The state inversion algorithm is initiated, and its objective is to deduce, from the current network state, the key planning decision sequences and combinations of external conditions that lead to the potential weaknesses and risk patterns, including: The specific network state corresponding to the identified potential weak link is defined as the target inversion state of the state inversion algorithm; A planning decision knowledge base is established, which stores the power grid planning decisions implemented in the planning area in a timeline format, including adding new lines, upgrading equipment, adjusting operation mode and changing protection settings. The state inversion algorithm constructs a reverse causal graph, the endpoint of which is the target inversion state, and the starting point is the previous planning decisions and observable external conditions. In the reverse causal graph, using the data in the standardized spatiotemporal fusion data cube, the conditional probability of events occurring and the relationship between time delay are calculated from each starting point to the ending point; By maximizing the joint conditional probability from the starting point to the ending point, one or more causal chains leading to the occurrence of the target inversion state are searched. Each causal chain consists of a set of key planning decision sequences ordered by time and the external conditions.

[0009] Furthermore, the state inversion algorithm completes the backtracking derivation by constructing and solving an inverse optimization problem with planning decisions as optimization variables and approximating the actual observed state as the objective, including: Construct an objective function that aims to minimize the difference between the historical network state derived by the state inversion algorithm and the historical network state actually recorded in the standardized spatiotemporal fusion data cube; Define decision variables, which are whether to adopt a certain planning decision recorded in the planning decision knowledge base in various historical time periods; Construct constraints, including mutual exclusivity between planning decisions, temporal sequence logic, and limits on total investment. A heuristic optimization algorithm is used to solve the constrained optimization problem. A set of planning decision variables is searched out so that the simulated historical state trajectory output by the network simulation model is closest to the actual historical state trajectory under the drive of the corresponding planning decision variable combination. The combination of planning decision variables obtained from the optimization solution is decoded into a series of specific planning decision behaviors and the time points in which they occur, thus forming a reconstruction of historical decisions.

[0010] Furthermore, the state inversion algorithm outputs a planning decision impact tracing report, including: For each causal chain derived by the state inversion algorithm, each planning and decision node on the chain is analyzed; Query the planning decision knowledge base to obtain detailed information for each planning decision node, including decision type, implementation scope, investment scale, and technical parameters; The contribution of each planning decision node to the final target inversion state is calculated, and the contribution is quantified by comparing the difference in the probability of occurrence of the target inversion state in the two cases: including the corresponding planning decision node and not including the corresponding planning decision node. The key planning decision sequence, the corresponding combination of external conditions, the detailed content of each planning decision node and its contribution are organized in reverse chronological order. Fill the organized information into the preset template to generate the planning decision impact tracing report, which includes text, data tables, and causal relationship diagrams.

[0011] Furthermore, combining the forward-looking simulation results of the planning scenario extrapolation model with the retrospective analysis conclusions of the planning decision impact tracing report, a multi-round, multi-angle cross-validation of a distribution network planning scheme to be evaluated is conducted, including: The proposed power distribution network planning scheme is analyzed to extract the planned implementation measures, expected targets, and investment plans. The planning measures are used as new inputs and injected into the planning scenario deduction model to perform a new round of reasoning calculations, so as to obtain the predicted operating status of each scenario after the implementation of the power distribution network planning scheme. From the aforementioned planning decision impact tracing report, extract historical negative planning decision cases that led to similar weaknesses or risk patterns as negative references; From the aforementioned planning decision impact tracing report, extract historical positive planning decision cases that have successfully improved network status and enhanced operational levels as positive references; The predicted operational status of the new scheme is compared horizontally with the results of historical positive and negative cases to evaluate the performance of the new scheme in avoiding historical mistakes and inheriting historical successes, thus completing a round of cross-validation. By modifying key parameters in the planning scheme, and repeating the process of injection, simulation, and comparison, multiple rounds of cross-validation are conducted.

[0012] Furthermore, based on the problems and conflicts exposed during the cross-validation process, a list of quantitative evaluation conclusions and structured adjustment suggestions for the distribution network planning scheme is generated, including: The results of multiple rounds of cross-validation were summarized, and the frequency and severity of the failure of various operational indicators of the proposed power distribution network planning scheme to be evaluated under different scenarios were statistically analyzed. Identify the main categories of reasons that lead to failure to meet the targets, including improper timing of planning decisions, insufficient capacity of selected equipment, overestimation of the distributed power absorption capacity, or underestimation of load growth. For each of the main cause categories mentioned above, and based on the historical experience in the planning decision impact tracing report, one or more specific planning adjustment suggestions are proposed. The adjustment suggestions include advancing or postponing an investment, changing the equipment model, adjusting the network structure, or adding a voltage regulating device. The original target indicators, simulated prediction indicators, gap values, main cause categories, and corresponding adjustment suggestions of the plan are compiled into a structured list of evaluation items. Based on the comprehensive scores of all evaluation items, a summary quantitative evaluation conclusion is formed, which clearly indicates the overall feasibility level of the plan.

[0013] Furthermore, the planning scenario simulation model is constructed, including: Construct a three-layer model framework comprising a physical layer, a rule layer, and a driver layer; In the physical layer, based on the network topology connections, equipment parameters and geographic information contained in the standardized spatiotemporal fusion data cube, a digital twin network model that can accurately reflect the electrical connection relationships and equipment characteristics of the power distribution network in the planning area is established. In the rule layer, explicit rules from scheduling operation procedures, equipment technical standards and safety guidelines are integrated, as well as implicit rules about load transfer, distributed power absorption and fault recovery mined from historical operation data through machine learning, to form a composite rule library that drives the evolution of the digital twin network model. In the driving layer, a multi-source input interface is designed to receive real-time data from the standardized spatiotemporal fusion data cube, processed predictive input data, and external input future policies and development boundary conditions. The digital twin network model of the physical layer, the composite rule base of the rule layer, and the multi-source input interface of the driving layer are coupled together so that the digital twin network model can simulate the changes in the power grid state with time and external conditions under the combined effect of the constraints of the composite rule base and the input data of the driving layer, thereby completing the construction of the planning scenario inference model.

[0014] Compared with the prior art, the beneficial effects of the present invention are: Spatiotemporal consistency cleaning is performed on the original data pool for distribution network assessment, which consists of historical and real-time data from multiple dimensions of sources, grids, loads, and storage. This process removes misaligned, redundant, and invalid duplicate data, maintaining the data's consistent attributes across time and space. Multi-scale fusion processing integrates data of different scales and types, such as distributed power generation output sequences, load change curves, network topology connections, equipment operation status logs, and meteorological information, into a standardized spatiotemporal fusion data cube. This cube can directly match the parameter formats for inference calculations and state inversion, avoiding compatibility conflicts between multi-source heterogeneous data in subsequent applications and eliminating inherent spatiotemporal errors in the data itself.

[0015] The planning scenario simulation model couples physical constraints with data-driven rule operation, simultaneously aligning with the physical operating laws of the distribution network and the characteristics of actual data changes. The inference and calculation module iteratively solves implicit equations describing energy flow, equipment actions, and network evolution, fully reconstructing the distribution network's operation under different future development scenarios and generating simulated operating states that closely match actual conditions. By matching the simulated operating states with historical typical operating states, abnormal states in the distribution network can be directly located, potential weaknesses that conventional assessment methods cannot detect can be identified, and risk patterns within the distribution network can be pinpointed, completing an in-depth analysis of the distribution network's planned state. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the steps of a power distribution network planning and evaluation method based on big data analysis as described in this invention. Figure 2 A flowchart for spatiotemporal consistency cleaning and fusion; Figure 3 A flowchart for identifying weaknesses and risk patterns. Detailed Implementation

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] See Figure 1 This invention provides a distribution network planning and evaluation method based on big data analysis. The method includes: acquiring multi-dimensional historical and real-time data on sources, networks, loads, and storage in the planning area. These data collectively constitute the original data pool for distribution network evaluation, specifically covering distributed power generation output sequences, load change curves, network topology connections, equipment operation status logs, and meteorological information. The original data pool is then subjected to spatiotemporal consistency cleaning and multi-scale fusion processing to generate a standardized spatiotemporal fusion data cube. This data cube provides a unified data foundation for subsequent inference calculations and state inversion processes. From the standardized spatiotemporal fusion data cube, a set of indicators characterizing key features of distribution network planning is extracted and input into a planning scenario inference model. The model operates based on coupled physical constraints and data-driven rules. Within the planning scenario inference model, an inference calculation module simulates the operating state of the distribution network under different future development scenarios. The inference calculation module achieves simulation by iteratively solving a set of implicit equations describing energy flow, equipment actions, and network evolution. The simulated operating status output by the inference computing module is matched with the historical typical operating status extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns.

[0019] In one embodiment of the present invention, spatiotemporal consistency cleaning and multi-scale fusion processing are performed on the original data pool for distribution network assessment to generate a standardized spatiotemporal fusion data cube. (See also...) Figure 2 For distributed generation output sequences and load change curves, a sliding window interpolation method is used to fill in missing data points, and statistical tests are used to remove abnormal fluctuations, forming a continuous and smooth time series curve. For network topology connections and equipment operation status logs, a mapping table between equipment and nodes, and lines and branches is established, and all equipment event logs are aligned to a unified time axis. Features strongly correlated with distribution network operation, including temperature, humidity, wind speed, and light intensity, are extracted from meteorological environmental information and resampled to the same time resolution as the distributed generation output sequences. A multi-source data association engine is established, based on the same timestamp and spatial location encoding, to associate and stitch together the processed distributed generation output sequences, load change curves, network topology connections, equipment operation status logs, and meteorological environmental features. The associated and stitched data is then reorganized according to time, spatial location, and data category dimensions to construct a standardized spatiotemporal fusion data cube with a unified index.

[0020] In practical implementation, the original data pool for distribution network assessment undergoes spatiotemporal consistency cleaning and multi-scale fusion processing to generate a standardized spatiotemporal fusion data cube. Taking a city distribution network planning area as an example, this area covers multiple photovoltaic power stations, load nodes, and meteorological monitoring points. Distributed power generation output sequences are collected at 15-minute intervals, but there are random missing points due to communication failures. Load change curves are recorded at the same intervals, but include instantaneous peaks caused by instrument malfunctions. The sliding window interpolation method is applied to the distributed power generation output sequences and load change curves to fill in the missing data points. The sliding window size is set to 5 time points, meaning that for missing points, the interpolated output... The calculation is as follows: in: This is the radius of the sliding window, set to 2, corresponding to a window size of 5. It is the weight coefficient of the data point at time index j, defined as , It is the original distributed power output value at time index j. Through this calculation, the missing positions are filled with the smooth weighted values ​​of data from adjacent time points. Statistical tests are then performed to remove abnormal fluctuation points. The Z-score method is used to calculate the deviation of each data point from the mean. Load data points with an absolute Z-score value exceeding 3 are removed and replaced with interpolated results to form a continuous and smooth time series curve.

[0021] To address network topology connectivity and device operation status logs, a mapping table is established between devices and nodes, and lines and branches. This mapping table is stored in structured tabular form; for example, transformer device numbers are mapped to distribution network node identifiers, and feeder line identifiers are mapped to electrical branch numbers. The timestamps of all device event logs, such as circuit breaker trip records and transformer overload alarms, are aligned to a unified time axis. This unified time axis uses Coordinated Universal Time (UTC) and is accurate to milliseconds. Device connection status change events in the network topology connectivity are also synchronized to the same time axis. A multi-source data association engine is established. Based on the same timestamp and spatial location code, it associates and concatenates the processed distributed power output sequence, load change curve, network topology connectivity, device operation status logs, and meteorological environmental characteristics. The timestamp accuracy is unified to the second level, and the spatial location code uses a geographic grid coding system; for example, each distribution network node is associated with a grid code. The multi-source data association engine matches and merges data records from different sources using a combination of timestamps and grid codes. In implementation, the multi-source data association engine is implemented using a distributed computing framework to handle large-scale data streams. In practical implementation, the associative stitching process ensures that each spatiotemporal point contains complete data attributes. For example, under the time "2023-07-15 14:00:00" and grid "G-1024", it associates photovoltaic output, load, topology connection status, equipment event flags, and meteorological characteristic values. It can be understood that the multi-source data association engine eliminates heterogeneity between data sources. The associative stitched data is reorganized according to time, spatial location, and data category dimensions. The time dimension forms a time-series index with 15-minute intervals as the basic unit. The spatial location dimension uses distribution network node numbers and line identifiers as hierarchical indexes. The data category dimension includes multiple data categories such as distributed power output, load, topology connection relationships, equipment operation status logs, and meteorological environmental characteristics, forming a standardized spatiotemporal fusion data cube with a unified index. The data cube is stored in an in-memory database as a multidimensional array, supporting slicing analysis along any dimension. Optionally, the data cube can be compressed for storage to save space.

[0022] In one embodiment of the present invention, within the planning scenario simulation model, an inference calculation module simulates the operating state of the distribution network under different future development scenarios. A set of future development scenarios is defined, each scenario being uniquely determined by a combination of load growth rate, distributed generation penetration rate, new equipment commissioning plans, and policy control parameters. In the planning scenario simulation model, a complete model of the current distribution network is loaded, including node admittance matrices, transformer tap parameters, line impedance parameters, and protection device settings. Typical daily load curves extracted from the standardized spatiotemporal fusion data cube and typical distributed generation output curves are scaled and shaped according to the parameters of each future development scenario to generate predictive input data for that future development scenario. Driven by the predictive input data and constrained by the complete distribution network model, the inference calculation module solves a set of equations including power flow balance equations, equipment operation limit equations, and protection action logic equations to obtain the voltage distribution, branch power flow distribution, and equipment load rate of the entire network in each future time period. For each defined development scenario, the operations of loading the distribution network model, generating predictive input data, and solving the operating state equations are executed sequentially. After traversing all future development scenarios, the complete simulated operating state under each scenario is stored. The simulated operating status output by the inference computing module is pattern-matched with historical typical operating statuses extracted from standardized spatiotemporal fusion data cubes to identify potential weaknesses and risk patterns. (See also...) Figure 3 From the simulated operating status output by the inference calculation module, a series of system-level and equipment-level status indicators are calculated, including node voltage over-limit duration, line overload probability, transformer load rate extremes, and system network loss level. From the standardized spatiotemporal fusion data cube, operating data from historically occurring alarm or fault periods are extracted, and system-level and equipment-level status indicators for the same set are calculated, forming a historical risk status indicator set. Using cluster analysis, the simulated status indicators are compared with the historical risk status indicator set, and simulated states whose indicator combinations fall into the same cluster are identified as having similar risks. The identified simulated states with similar risks are located at specific power grid equipment, lines, and nodes, and these are marked as potential weak points. The occurrence conditions and evolution processes of simulated states with similar risks are analyzed, summarized into one or more specific risk patterns, and a descriptive label is assigned to each risk pattern.

[0023] In practical implementation, the simulated operating status output by the inference computing module is pattern matched with historical typical operating statuses extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns. From the simulated operating status output by the inference computing module, a series of system-level and equipment-level status indicators are calculated. System-level status indicators include node voltage over-limit duration and system network loss level; equipment-level status indicators include line overload probability and transformer load rate extremes. Node voltage over-limit duration is calculated as the total time during which the voltage exceeds the upper limit of 0.95 pu or falls below the lower limit of 0.90 pu in a day. Line overload probability is the percentage of time the line current exceeds its long-term allowable current carrying capacity. From the standardized spatiotemporal fusion data cube, operating data from historically occurring alarm or fault periods are extracted, and system-level and equipment-level status indicators for the same set are calculated to form a historical risk status indicator set. This historical risk status indicator set is a multi-dimensional vector set, with each vector representing a characteristic indicator of a historical risk event. Cluster analysis is employed to compare simulated state indicators with historical risk state indicator sets. The K-means algorithm is used to identify simulated states whose indicator combinations fall into the same cluster as having similar risks. The number of clusters is determined using the elbow rule. In practice, the similarity between two state vectors is measured using Euclidean distance. It can be understood that falling into a high-risk cluster implies similarity to historical adverse operating patterns. The identified simulated states with similar risks are then located at specific power grid equipment, lines, and nodes, which are marked as potential weak points. For example, if a simulated operating state is identified as having similar risks, and its state indicators show that node N15 experiences excessively long voltage exceedances and line L7 has a high probability of overload, then node N15 and line L7 are marked. The analysis examines the occurrence conditions and evolution processes of simulated states with similar risks, summarizing them into one or more specific risk patterns. Each risk pattern is assigned a descriptive label. For example, if the analysis reveals that when the output of distributed power sources suddenly decreases while the load is at its peak, voltage drops and line overloads occur in specific areas, this can be summarized as the risk pattern of "high-penetration distributed power source reverse peak shaving leading to localized overload." Optionally, the descriptive labels for risk patterns are stored in a knowledge base. In some embodiments, the analysis process can incorporate decision tree methods. It is understood that identification and labeling help to quickly locate problems. When calculating system-level and device-level state indicators, the probability of line overload is considered. Defined as: in: This refers to the total duration during which the line current exceeds its long-term allowable current carrying capacity. It is the total observation time, during the simulation operation. Typically, a typical 24-hour day is used. Optionally, the total observation time can also be set according to different analytical needs.

[0024] In one embodiment of the present invention, a state inversion algorithm is initiated based on the identified potential weaknesses and risk patterns. The goal of the state inversion algorithm is to deduce the key planning decision sequences and external condition combinations that lead to the potential weaknesses and risk patterns from the current network state. The specific network state corresponding to the identified potential weaknesses is defined as the target inversion state of the state inversion algorithm. A planning decision knowledge base is established, which stores the power grid planning decisions implemented in the planning area in a timeline format, including adding lines, upgrading equipment, adjusting operating modes, and changing protection settings. The state inversion algorithm constructs a reverse causal graph, with the target inversion state as the endpoint and the previous planning decisions and observable external conditions as the starting points. In the reverse causal graph, data from a standardized spatiotemporal fusion data cube is used to calculate the conditional probability and time delay relationship of events from each starting point to the endpoint. By maximizing the joint conditional probability from the starting point to the endpoint, one or more causal chains leading to the target inversion state are searched. Each causal chain consists of a set of key planning decision sequences and external condition combinations ordered by time.

[0025] The state inversion algorithm completes backtracking derivation by constructing and solving an inverse optimization problem with planning decisions as optimization variables and the goal of approximating the actual observed state. An objective function is constructed to minimize the difference between the historical network state inverted by the state inversion algorithm and the actual historical network state recorded in the standardized spatiotemporal fusion data cube. Decision variables are defined as whether to adopt a planning decision recorded in the planning decision knowledge base at various historical time periods. Constraints are constructed, including mutual exclusivity between planning decisions, temporal sequence logic, and a limit on the total investment. A heuristic optimization algorithm is used to solve the constrained optimization problem, searching for a set of planning decision variable combinations such that, driven by the corresponding planning decision variable combinations, the simulated historical state trajectory output by the network simulation model most closely approximates the actual historical state trajectory. The optimized combination of planning decision variables is decoded into a series of specific planning decision behaviors and their occurrence times, forming a reconstruction of historical decisions. The state inversion algorithm outputs a planning decision impact tracing report. For each causal chain inverted by the state inversion algorithm, each planning decision node in the chain is analyzed. The system queries the planning decision-making knowledge base to obtain detailed information for each planning decision node, including decision type, implementation scope, investment scale, and technical parameters. It calculates the contribution of each planning decision node to the final target inversion state, quantifying this contribution by comparing the probability of the target inversion state occurring with and without the corresponding planning decision node. The system organizes the key planning decision sequences, corresponding combinations of external conditions, detailed information for each planning decision node, and their contribution in reverse chronological order. Finally, it inputs the organized information into a pre-defined template to generate a planning decision impact tracing report containing text, data tables, and causal relationship diagrams.

[0026] In practical implementation, based on the identified potential weaknesses and risk patterns, a state inversion algorithm is initiated. The goal of the state inversion algorithm is to deduce the key planning decision sequences and external condition combinations that led to the potential weaknesses and risk patterns from the current network state. Taking a risk pattern identified as "local voltage exceeding limits at midday in summer" as an example, the specific network state corresponding to it is that the voltage of node N15 is consistently below 0.90 pu during midday in summer during a specific historical period. This network state is defined as the target inversion state of the state inversion algorithm. A planning decision knowledge base is established, which stores the power grid planning decisions implemented in the planning area in a timeline format. The storage structure of the planning decision knowledge base includes decision timestamps, decision types, decision content descriptions, affected equipment or area identifiers, and investment scale fields. For example, a decision record is "2022-03-10, new line, erecting a 10kV line Lnew between node N12 and node N18, with an investment scale of 1.5 million yuan". In practical implementation, the planning decision knowledge base is imported from the power grid management information system. The state inversion algorithm constructs a reverse causal graph. The endpoint of the reverse causal graph is the target inversion state, namely the voltage limit exceeding of node N15. The starting points are the planning decisions and observable external conditions, including "PV1 power plant commissioned in 2021", "Lnew new line added in 2022", and "daily maximum temperature exceeding 38℃ in summer 2023". In the reverse causal graph, data from the standardized spatiotemporal fusion data cube is used to calculate the conditional probability and time delay relationship of events from each starting point to the endpoint. For example, given the decision that "PV1 power plant commissioned in 2021", the probability of the voltage limit exceeding of node N15 is calculated. The time delay relationship is statistically analyzed by analyzing the time interval distribution between the two events. By maximizing the joint conditional probability from the starting point to the end point, one or more causal chains leading to the occurrence of the target inversion state are searched. Each causal chain consists of a set of key planning decision sequences ordered by time and external conditions. For example, a causal chain obtained by searching is "PV1 photovoltaic power station put into operation in 2021" -> "natural load growth in 2022" -> "high temperature in summer of 2023" -> "voltage limit exceeded at node N15".

[0027] The state inversion algorithm completes backtracking derivation by constructing and solving an inverse optimization problem with planning decisions as optimization variables and approximating the actual observed state as the objective. An objective function is constructed, aiming to minimize the difference between the historical network state inverted by the state inversion algorithm and the historical network state actually recorded in the standardized spatiotemporal fusion data cube. The objective function is defined as the norm of the historical network state difference. Decision variables are defined as whether to adopt a certain planning decision recorded in the planning decision knowledge base at each historical time period. Decision variables are binary variables. Constraints are constructed, including mutual exclusivity between planning decisions, temporal sequence logic, and a limit on the total investment amount. Mutually exclusive constraints include that "line renovation" and "line reconstruction" cannot be performed on the same line in the same year. Temporal sequence logic constraints include that "adding a transformer" must be done after "expanding a substation." The total investment amount is limited to the sum of investments in each time period not exceeding the historical actual investment budget. A heuristic optimization algorithm is employed to solve the constrained optimization problem. This heuristic algorithm utilizes a genetic algorithm, which searches the decision variable space through selection, crossover, and mutation operations to find a set of planning decision variables. This combination ensures that, driven by the corresponding planning decision variable combination, the simulated historical state trajectory output by the network simulation model most closely approximates the actual historical state trajectory. The network simulation model uses the same power flow calculation kernel as the planning scenario extrapolation model. The optimized planning decision variable combination is decoded into a series of specific planning decision behaviors and their occurrence times, forming a reconstruction of historical decisions. For example, the decoded result could be "Decision D5 (PV grid connection) executed in Q2 2021, Decision D8 (load switching) executed in Q3 2022". In some embodiments, the heuristic optimization algorithm can also employ simulated annealing. Essentially, the optimization process achieves reverse extrapolation of the historical decision sequence.

[0028] The state inversion algorithm outputs a planning decision impact tracing report. For each causal chain derived by the algorithm, it analyzes each planning decision node along the chain. For example, the above causal chain analyzes three nodes: "Commissioning Photovoltaic Power Station PV1", "Natural Load Growth", and "Summer High Temperature". The planning decision knowledge base is queried to obtain detailed information for each planning decision node, including decision type, implementation scope, investment scale, and technical parameters. For the "Commissioning Photovoltaic Power Station PV1" node, the query reveals its decision type as "Distributed Power Supply Access", implementation scope as "Node N08", investment scale as "2 million yuan", and technical parameters as "Rated Capacity 2MW". The contribution of each planning decision node to the final target inversion state is calculated. The contribution is quantified by comparing the probability of the target inversion state occurring with and without the corresponding planning decision node. The contribution calculation formula is: in: This represents the contribution of the i-th planning decision node Di. Indicates the presence of decision nodes The probability of the target inversion state S occurring at that time. This indicates that no decision nodes are included. The probability of the target inversion state S occurring is calculated through correlations in a statistically standardized spatiotemporal fusion data cube or through numerous simulation scenarios. It can be understood that contribution quantifies the impact of a single decision. The key planning decision sequences, corresponding combinations of external conditions, detailed information of each planning decision node, and their contribution are organized in reverse chronological order, from most recent events to earlier events. The organized information is then entered into a preset template to generate a planning decision impact tracing report containing text, data tables, and causal relationship diagrams. The data tables in the report display detailed information about the key decision sequences; see Table 1.

[0029] Table 1: Excerpt from the Report on the Impact of Planning Decisions In one embodiment of the present invention, a multi-round, multi-angle cross-validation is performed on a distribution network planning scheme to be evaluated by combining the forward-looking simulation results of the planning scenario inference model with the retrospective analysis conclusions of the planning decision impact tracing report. The distribution network planning scheme to be evaluated is analyzed to extract the planned implementation measures, expected indicators, and investment plans. The planning measures are then used as new inputs to inject into the planning scenario inference model, performing a new round of inference calculations to obtain the predicted operating status under various scenarios after implementing the distribution network planning scheme. Negative planning decision cases that led to similar weaknesses or risk patterns in the past are extracted from the planning decision impact tracing report as negative references. Positive planning decision cases that successfully improved network conditions and operational levels in the past are also extracted from the planning decision impact tracing report as positive references. The predicted operating status of the new scheme is compared horizontally with the results of historical positive and negative cases to evaluate the performance of the new scheme in avoiding historical errors and inheriting historical successes, completing one round of cross-validation. Multiple rounds of cross-validation are performed by modifying key parameters in the planning scheme and repeating the injection, simulation, and comparison process. Based on the problems and conflicts exposed during cross-validation, a list of quantitative evaluation conclusions and structured adjustment suggestions for the distribution network planning scheme is generated. The results of multiple rounds of cross-validation are summarized, and the frequency and severity of non-compliance of various operational indicators for the distribution network planning scheme under different scenarios are statistically analyzed. The main categories of reasons for non-compliance are identified, including inappropriate timing of planning decisions, insufficient capacity of selected equipment, overestimation of distributed power absorption capacity, or underestimation of load growth. For each main category of cause, based on historical experience from the planning decision impact tracing report, one or more specific planning scheme adjustment suggestions are proposed. These suggestions include advancing or delaying an investment, changing equipment models, adjusting network structure, or adding voltage regulating devices. The original target indicators, simulated prediction indicators, gap values, main cause categories, and corresponding adjustment suggestions are compiled into a structured evaluation item list. Based on the comprehensive score of all evaluation items, a summary quantitative evaluation conclusion is formed, which clearly indicates the overall feasibility level of the scheme.

[0030] In practice, planning measures, expected indicators, and investment plans are automatically extracted from the scheme document in the form of structured data tables. These planning measures are then used as new inputs and injected into the planning scenario simulation model to perform a new round of inference calculations. This yields the predicted operating status under various future scenarios after the implementation of the distribution network planning scheme. For example, the measure of "adding transformer T3" is injected by modifying the physical layer network topology and parameters in the planning scenario simulation model. Simulations are then performed under two future development scenarios with load growth rates of 5% and 8%, respectively, to obtain predicted operating status data such as voltage distribution and load rate for each of the next three years. In some embodiments, the injection process is completed automatically by calling the application programming interface of the planning scenario simulation model. From the impact tracing report of planning decisions, negative planning decision cases that led to similar weaknesses or risk patterns in the past are extracted as negative references. For example, a historical case was extracted from the report showing that "in 2019, only a small-scale line renovation was carried out in the N20 area without simultaneous expansion of the main transformer, resulting in main transformer overload and voltage drop during the peak load period in the summer of 2020." This case was identified as a negative planning decision case of "insufficient investment and mismatched timing." From the impact tracing report of planning decisions, positive planning decision cases that successfully improved network conditions and enhanced operational levels in the past are extracted as positive references.

[0031] Based on the problems and conflicts exposed during cross-validation, a list of quantitative evaluation conclusions and structured adjustment suggestions for distribution network planning schemes is generated. The results of multiple rounds of cross-validation are summarized, and the frequency and severity of non-compliance of various operational indicators for the distribution network planning schemes under different scenarios are statistically analyzed. For example, in a load growth rate scenario of 8%, two out of three simulations showed that the "node voltage qualification rate" indicator did not reach the expected 99.9%, while in a load growth rate scenario of 5%, this indicator met the target in all simulations. The main categories of reasons for non-compliance are identified, including improper timing of planning decisions, insufficient equipment capacity, overestimation of distributed power absorption capacity, or underestimation of load growth. For each major cause category, based on historical experience from the planning decision impact tracing report, propose one or more specific planning scheme adjustment suggestions. For the category of "insufficient equipment selection capacity", based on the experience of using forward-looking configurations in historical positive planning decision cases, propose an adjustment suggestion: "Increase the capacity of transformer T3 in the plan from 10MVA to 12.5MVA". For the category of "inappropriate timing of planning decisions", propose an adjustment suggestion: "Advance the renovation time of feeder F5 from 2024 to 2023 to better match it with the peak load growth."

[0032] The original target indicators, simulated prediction indicators, gap values, main cause categories, and corresponding adjustment suggestions of the plan are compiled into a structured evaluation item list. This list is presented in tabular form, with each row corresponding to one planning measure or indicator being evaluated. Based on the comprehensive score of all evaluation items, a summary quantitative evaluation conclusion is formed. The comprehensive score is calculated by weighting the indicator gap values ​​and frequency of non-compliance for each evaluation item. The quantitative evaluation conclusion clearly indicates the overall feasibility level of the plan, such as "basically feasible, but requiring local optimization." See Table 2 for the evaluation item list. Table 2: Structured Evaluation Checklist for Distribution Network Planning Schemes Overall score The calculation formula is: in: For the comprehensive evaluation of the proposal, To assess the number of entries, This is the weight coefficient for the k-th item, reflecting its importance. Let be the difference between the simulated predicted index and the original target index for the k-th item. For the original target index value of the k-th item, the function This represents the absolute value of the difference; the formula quantifies how closely the simulation results approximate the planning target. Optional, weighting coefficients. It can be determined through expert scoring. In practical implementation, the feasibility level can be based on a comprehensive score. Threshold interval division.

[0033] In one embodiment of the present invention, a planning scenario simulation model is constructed. A three-layer model framework comprising a physical layer, a rule layer, and a driving layer is built. In the physical layer, based on the network topology connections, equipment parameters, and geographic information contained in the standardized spatiotemporal fusion data cube, a digital twin network model is established that can accurately reflect the electrical connections and equipment characteristics of the distribution network in the planning area. In the rule layer, explicit rules from scheduling operation procedures, equipment technical standards, and safety guidelines are integrated, along with implicit rules about load transfer, distributed power source absorption, and fault recovery mined from historical operation data through machine learning, forming a composite rule library that drives the evolution of the digital twin network model. In the driving layer, a multi-source input interface is designed to receive real-time data from the standardized spatiotemporal fusion data cube, processed predictive input data, and external input future policies and development boundary conditions. The digital twin network model of the physical layer, the composite rule library of the rule layer, and the multi-source input interface of the driving layer are coupled, so that the digital twin network model, under the constraints of the composite rule library and the input data of the driving layer, simulates the changes in the power grid state with time and external conditions, thereby completing the construction of the planning scenario simulation model.

[0034] In practical implementation, a planning scenario simulation model is constructed, creating a three-layer model framework comprising a physical layer, a rule layer, and a driving layer. Taking a new urban area's power distribution network as an example planning area, the physical layer, rule layer, and driving layer each undertake different modeling and data interaction functions. In the physical layer, based on the network topology connections, equipment parameters, and geographic information contained in the standardized spatiotemporal fusion data cube, a digital twin network model is established that accurately reflects the electrical connections and equipment characteristics of the power distribution network in the planning area. The digital twin network model specifically includes the π-type equivalent circuit model of all lines, the on-load tap changer model of transformers, the inverter interface model of distributed power sources, and the static voltage characteristic model of each node load. Its electrical connections are described through a node-branch correlation matrix. Equipment parameters such as line resistance, reactance, transformer turns ratio, and maximum output of distributed power sources are directly extracted and instantiated from the corresponding fields of the standardized spatiotemporal fusion data cube. Geographic information is used to determine the spatial layout of equipment and lines. The topology of the digital twin network model is consistent with the actual physical power grid. At the rule layer, explicit rules from scheduling operation procedures, equipment technical standards, and safety guidelines are integrated with implicit rules about load transfer, distributed power absorption, and fault recovery mined from historical operation data through machine learning. This forms a composite rule base that drives the evolution of the digital twin network model. Explicit rules are encoded in the form of "condition-action" logical statements. The provisions on the long-term overload capacity of transformers in the equipment technical standards are transformed into constraint inequalities in the model. The requirements for N-1 verification in the safety guidelines are transformed into an automated simulation verification process. Implicit rules are mined from historical operation data stored in a standardized spatiotemporal fusion data cube through machine learning. For example, the random forest algorithm is used to analyze historical load transfer records and learn the optimal load transfer path under different load levels and different distributed power output combinations, forming a data-driven implicit rule: "When the weather pattern is A and the main transformer load rate is greater than 80%, the probability of switching load L1 to feeder F2 is 85%." The composite rule base is a collection of all explicit and implicit rules.

[0035] In practical implementation, implicit patterns are updated through periodic model training. In the driver layer, a multi-source input interface is designed to receive real-time data from a standardized spatiotemporal fusion data cube, processed predictive input data, and external inputs of future policies and development boundary conditions. The multi-source input interface defines standardized data formats and communication protocols. For example, it receives real-time load measurement data pushed by the standardized spatiotemporal fusion data cube via a message queue, loads processed predictive input data on future load growth via a file interface, and receives external inputs of future policies and development boundary conditions, such as new electricity pricing policy texts and key land use change information in urban development plans, via a manual configuration interface. In some embodiments, the multi-source input interface supports real-time access to streaming data. Optionally, future policies and development boundary conditions can be input in the form of structured scripts. The physical layer's digital twin network model, the rule layer's composite rule base, and the driver layer's multi-source input interface are coupled. This allows the digital twin network model to simulate the changes in power grid state over time and external conditions under the combined influence of the composite rule base's constraints and the driver layer's input data. The coupling mechanism is implemented through a master-controlled simulation loop. At each simulation step, the driver layer's multi-source input interface injects real-time or predicted data into the digital twin network model. The model then calculates the power flow based on these inputs, obtaining the instantaneous states of voltage at each node and power in each branch. The rule layer's composite rule base then scans these states, matching and triggering relevant rules. For example, if a voltage limit violation rule is triggered, the composite rule base generates a control command to adjust the transformer tap changer. This command is fed back to the digital twin network model for execution, updating the model's state accordingly. This process repeats continuously. The state evolution of the digital twin network model can be described as follows: in: This represents the state vector of the digital twin network model at time t, containing the voltages of all nodes and the power of each branch. This represents the input data vector injected by the driver layer at time t. This indicates a composite rule base. Indicated in the composite rule base Under the constraints of the current state and input Calculate the next state The evolution function is used to construct the planning scenario simulation model. This model can be understood as a dynamic, closed-loop simulation system. Optionally, the step size of the master simulation loop can be configured. The evolution function... It encapsulates the physical laws and operating rules of the power grid.

[0036] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A distribution network planning and evaluation method based on big data analysis, characterized in that, include: The system acquires multi-dimensional historical and real-time data on power generation, grid, load and storage in the planning area to form a raw data pool for distribution network assessment. The raw data pool for distribution network assessment includes distributed power output sequences, load change curves, network topology connections, equipment operation status logs and meteorological information. Spatiotemporal consistency cleaning and multi-scale fusion processing are performed on the original data pool of the power distribution network assessment to generate a standardized spatiotemporal fusion data cube. The standardized spatiotemporal fusion data cube is used to support subsequent inference calculations and state inversion processes. From the standardized spatiotemporal fusion data cube, an index set for characterizing key features of distribution network planning is extracted, and the index set is input into the planning scenario inference model, which operates based on coupled physical constraints and data-driven rules. Within the planning scenario simulation model, an inference calculation module is used to simulate the operating state of the power distribution network under different future development scenarios. The inference calculation module achieves this by iteratively solving a set of implicit equations describing energy flow, equipment operation, and network evolution. The simulated operating status output by the inference computing module is matched with the historical typical operating status extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns.

2. The distribution network planning and evaluation method based on big data analysis as described in claim 1, characterized in that, The original data pool for the distribution network assessment is subjected to spatiotemporal consistency cleaning and multi-scale fusion processing to generate a standardized spatiotemporal fusion data cube, including: For the distributed power generation output sequence and the load change curve, the sliding window interpolation method is used to fill in the missing data points, and statistical tests are used to remove abnormal fluctuation points to form a continuous and smooth time series curve. Based on the network topology connection relationship and the device operation status log, establish a mapping relationship table between devices and nodes, lines and branches, and align all device event logs to a unified time coordinate axis; Features strongly correlated with the operation of the power distribution network are extracted from the meteorological environment information, including temperature, humidity, wind speed and light intensity, and then resampled to the same time resolution as the output sequence of the distributed power source. A multi-source data association engine is established, which associates and splices the processed distributed power output sequence, the load change curve, the network topology connection relationship, the equipment operation status log and the meteorological environment features based on the same timestamp and the same spatial location encoding. The concatenated data is then reorganized according to the time dimension, spatial location dimension, and data category dimension to construct a standardized spatiotemporal fusion data cube with a unified index.

3. The distribution network planning and evaluation method based on big data analysis as described in claim 2, characterized in that, Within the planning scenario simulation model, an inference and calculation module is used to simulate the operating state of the power distribution network under different future development scenarios, including: Define a set of future development scenarios, each of which is uniquely determined by a combination of load growth rate, distributed power penetration rate, new equipment commissioning plan and policy control parameters; In the planning scenario simulation model, the complete model of the current distribution network is loaded, including the node admittance matrix, transformer tap parameters, line impedance parameters and protection device settings; The typical daily load curve and the typical output curve of distributed power sources extracted from the standardized spatiotemporal fusion data cube are scaled and shaped according to the parameters of each future development scenario to generate predictive input data for the future development scenario. The reasoning and calculation module is driven by the predicted input data and constrained by the complete model of the distribution network. It solves a set of equations including power flow balance equation, equipment operation limit equation and protection action logic equation to obtain the voltage distribution, branch power flow distribution and equipment load rate of the whole network in each future time period. For each defined future development scenario, the operations of loading the distribution network model, generating predictive input data, and solving the operating state equations are executed sequentially. After traversing all future development scenarios, the complete simulated operating state under each scenario is stored.

4. The distribution network planning and evaluation method based on big data analysis as described in claim 3, characterized in that, The simulated operating status output by the inference computing module is matched with historical typical operating statuses extracted from the standardized spatiotemporal fusion data cube to identify potential weaknesses and risk patterns, including: From the simulated operating status output by the inference calculation module, a series of system-level and equipment-level status indicators are calculated, including node voltage over-limit duration, line overload probability, transformer load rate extreme value, and system network loss level. From the standardized spatiotemporal fusion data cube, operational data for periods during which alarms or faults occurred in the past are extracted, and the system-level and device-level status indicators of the same set are calculated to form a set of historical risk status indicators. Cluster analysis is used to compare the simulated state indicators with the historical risk state indicator set. Simulated states whose indicator combination features fall into the same cluster are identified as having similar risks. The identified simulated states with similar risks are located to specific power grid equipment, lines, and nodes, and these equipment, lines, and nodes are marked as potential weak points. The conditions for occurrence and the evolution process of the simulated states with similar risks are analyzed, and they are summarized into one or more specific risk patterns. A descriptive label is assigned to each risk pattern.

5. The distribution network planning and evaluation method based on big data analysis as described in claim 4, characterized in that, Also includes: Based on the identified potential weaknesses and risk patterns, a state inversion algorithm is initiated. The goal of the state inversion algorithm is to backtrack from the current network state to derive the key planning decision sequence and combination of external conditions that led to the potential weaknesses and risk patterns. The state inversion algorithm completes the backtracking derivation by constructing and solving an inverse optimization problem with planning and decision as optimization variables and approximating the actual observed state as the objective. The state inversion algorithm outputs a planning decision impact tracing report, which details the impact paths and contributions of each historical planning decision on the current network state. By combining the forward-looking simulation results of the planning scenario model with the retrospective analysis conclusions of the planning decision impact tracing report, a distribution network planning scheme to be evaluated is cross-validated in multiple rounds and from multiple perspectives. Based on the problems and conflicts exposed during the cross-validation process, a list of quantitative evaluation conclusions and structured adjustment suggestions for the power distribution network planning scheme is generated. The state inversion algorithm is initiated, and its objective is to deduce, from the current network state, the key planning decision sequences and combinations of external conditions that lead to the potential weaknesses and risk patterns, including: The specific network state corresponding to the identified potential weak link is defined as the target inversion state of the state inversion algorithm; A planning decision knowledge base is established, which stores the power grid planning decisions implemented in the planning area in a timeline format, including adding new lines, upgrading equipment, adjusting operation mode and changing protection settings. The state inversion algorithm constructs a reverse causal graph, the endpoint of which is the target inversion state, and the starting point is the previous planning decisions and observable external conditions. In the reverse causal graph, using the data in the standardized spatiotemporal fusion data cube, the conditional probability of events occurring and the relationship between time delay are calculated from each starting point to the ending point; By maximizing the joint conditional probability from the starting point to the ending point, one or more causal chains leading to the occurrence of the target inversion state are searched. Each causal chain consists of a set of key planning decision sequences ordered by time and the external conditions.

6. The distribution network planning and evaluation method based on big data analysis as described in claim 5, characterized in that, The state inversion algorithm completes the backtracking derivation by constructing and solving an inverse optimization problem with planning decisions as optimization variables and approximating the actual observed state as the objective, including: Construct an objective function that aims to minimize the difference between the historical network state derived by the state inversion algorithm and the historical network state actually recorded in the standardized spatiotemporal fusion data cube; Define decision variables, which are whether to adopt a certain planning decision recorded in the planning decision knowledge base in various historical time periods; Construct constraints, including mutual exclusivity between planning decisions, temporal sequence logic, and limits on total investment. A heuristic optimization algorithm is used to solve the constrained optimization problem. A set of planning decision variables is searched out so that the simulated historical state trajectory output by the network simulation model is closest to the actual historical state trajectory under the drive of the corresponding planning decision variable combination. The combination of planning decision variables obtained from the optimization solution is decoded into a series of specific planning decision behaviors and the time points in which they occur, thus forming a reconstruction of historical decisions.

7. The distribution network planning and evaluation method based on big data analysis as described in claim 6, characterized in that, The state inversion algorithm outputs a planning decision impact tracing report, including: For each causal chain derived by the state inversion algorithm, each planning and decision node on the chain is analyzed; Query the planning decision knowledge base to obtain detailed information for each planning decision node, including decision type, implementation scope, investment scale, and technical parameters; The contribution of each planning decision node to the final target inversion state is calculated, and the contribution is quantified by comparing the difference in the probability of occurrence of the target inversion state in the two cases: including the corresponding planning decision node and not including the corresponding planning decision node. The key planning decision sequence, the corresponding combination of external conditions, the detailed content of each planning decision node and its contribution are organized in reverse chronological order. Fill the organized information into the preset template to generate the planning decision impact tracing report, which includes text, data tables, and causal relationship diagrams.

8. The distribution network planning and evaluation method based on big data analysis as described in claim 7, characterized in that, Combining the forward-looking simulation results of the planning scenario model with the retrospective analysis conclusions of the planning decision impact tracing report, a distribution network planning scheme to be evaluated is cross-validated in multiple rounds and from multiple perspectives, including: The proposed power distribution network planning scheme is analyzed to extract the planned implementation measures, expected targets, and investment plans. The planning measures are used as new inputs and injected into the planning scenario deduction model to perform a new round of reasoning calculations, so as to obtain the predicted operating status of each scenario after the implementation of the power distribution network planning scheme. From the aforementioned planning decision impact tracing report, extract historical negative planning decision cases that led to similar weaknesses or risk patterns as negative references; From the aforementioned planning decision impact tracing report, extract historical positive planning decision cases that have successfully improved network status and enhanced operational levels as positive references; The predicted operational status of the new scheme is compared horizontally with the results of historical positive and negative cases to evaluate the performance of the new scheme in avoiding historical mistakes and inheriting historical successes, thus completing a round of cross-validation. By modifying key parameters in the planning scheme, and repeating the process of injection, simulation, and comparison, multiple rounds of cross-validation are conducted.

9. The distribution network planning and evaluation method based on big data analysis as described in claim 8, characterized in that, Based on the problems and conflicts exposed during the cross-validation process, a list of quantitative evaluation conclusions and structured adjustment suggestions for the distribution network planning scheme is generated, including: The results of multiple rounds of cross-validation were summarized, and the frequency and severity of the failure of various operational indicators of the proposed power distribution network planning scheme to be evaluated under different scenarios were statistically analyzed. Identify the main categories of reasons that lead to failure to meet the targets, including improper timing of planning decisions, insufficient capacity of selected equipment, overestimation of the distributed power absorption capacity, or underestimation of load growth. For each of the main cause categories mentioned above, and based on the historical experience in the planning decision impact tracing report, one or more specific planning adjustment suggestions are proposed. The adjustment suggestions include advancing or postponing an investment, changing the equipment model, adjusting the network structure, or adding a voltage regulating device. The original target indicators, simulated prediction indicators, gap values, main cause categories, and corresponding adjustment suggestions of the plan are compiled into a structured list of evaluation items. Based on the comprehensive scores of all evaluation items, a summary quantitative evaluation conclusion is formed, which clearly indicates the overall feasibility level of the plan.

10. The distribution network planning and evaluation method based on big data analysis as described in claim 9, characterized in that, Constructing the planning scenario simulation model includes: Construct a three-layer model framework comprising a physical layer, a rule layer, and a driver layer; In the physical layer, based on the network topology connections, equipment parameters and geographic information contained in the standardized spatiotemporal fusion data cube, a digital twin network model that can accurately reflect the electrical connection relationships and equipment characteristics of the power distribution network in the planning area is established. In the rule layer, explicit rules from scheduling operation procedures, equipment technical standards and safety guidelines are integrated, as well as implicit rules about load transfer, distributed power absorption and fault recovery mined from historical operation data through machine learning, to form a composite rule library that drives the evolution of the digital twin network model. In the driving layer, a multi-source input interface is designed to receive real-time data from the standardized spatiotemporal fusion data cube, processed predictive input data, and external input future policies and development boundary conditions. The digital twin network model of the physical layer, the composite rule base of the rule layer, and the multi-source input interface of the driving layer are coupled together so that the digital twin network model can simulate the changes in the power grid state with time and external conditions under the combined effect of the constraints of the composite rule base and the input data of the driving layer, thereby completing the construction of the planning scenario inference model.