Power distribution network power supply partition division method, electronic equipment, storage medium and system
By co-training a two-layer optimization model and a time-representation model, the problem of static and dynamic disconnect in the power supply division of the distribution network was solved, functional complementarity identification and self-balancing capability were improved, and the regional division and energy storage scheduling of the distribution network were optimized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-20
- Publication Date
- 2026-03-27
Smart Images

Figure CN121749136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power transmission planning technology, specifically to a method for dividing power distribution network into power supply zones, electronic equipment, storage media, and system. Background Technology
[0002] With the dense integration of high-proportion distributed photovoltaic (PV), flexible loads, and battery energy storage systems into distribution networks, the distribution system is transforming from a traditional unidirectional power supply network into a dense and multidirectional energy ecosystem. This transformation has fundamentally changed the operating mode of the power grid, whose balancing capacity, flexibility, and resilience increasingly rely on the dynamic interaction and real-time coordination among distributed source-load-storage resources. This transformation has created a fundamental contradiction between the traditional distribution network power allocation methods based on static topology, geographical location, or administrative boundaries, and the high volatility, multidirectional power flow, and time-varying behavior brought about by distributed PV, energy storage, electric vehicles, and flexible loads.
[0003] In related technologies, power distribution network partitioning methods can be categorized into three types. The first type uses techniques such as dynamic time warping and spectral clustering to cluster node behavior based on historical data. This type of method implicitly assumes that historical behavior patterns will stably repeat in the future, and cannot effectively cope with non-stationary and sudden fluctuations driven by weather and user behavior. The second type of method constructs graph models based on electrical distance and power flow sensitivity to identify electrical communities in the network. This type of method focuses on static electrical connections and ignores behavioral states that evolve over time. The third type of method introduces probabilistic models to generate uncertain scenarios, attempting to consider volatility in partitioning. However, its scenario generation is usually independent of the partitioning mechanism, and its representation of uncertainty is superficial, treating time fluctuations as exogenous noise and lacking a deep characterization of the underlying structured and state-dependent dynamic processes.
[0004] In summary, the relevant technologies suffer from a disconnect between static topology partitioning and dynamic grid operation, an inability to identify and utilize the functional complementarity that naturally forms across multiple time scales due to the dynamics of distributed photovoltaics, flexible loads, and energy storage, and a disconnect between the predictions of the relevant models and structural planning decisions such as regional partitioning and network reconfiguration. As a result, the partitioning results cannot truly reflect the future dynamic evolution of the system. Summary of the Invention
[0005] To achieve the above objectives, the present invention employs the following technical solution: The first aspect of this invention provides a method for dividing power supply zones in a power distribution network, the method comprising: A two-level optimization model for distribution network partitioning is established. The upper level of the two-level optimization model is configured to optimize the partitioning decision with node partitioning as the decision variable and maximizing the self-balancing capability of the partition as the objective. The lower level of the two-level optimization model is configured to simulate the operational feasibility of the partitioning decision through a set of constraints. Based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network, an attention mechanism is used to construct a time representation model. The time representation model is used to extract the time-series features of each node at multiple time scales and generate the potential state embedding vector of each node. The dynamic operating characteristics of nodes represented by the latent state embedding vector are used as input parameters of the bi-level optimization model. The bi-level optimization model is solved to determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation techniques and gradient feedback mechanisms to co-train the time representation model and the bi-level optimization model, while updating the parameters of the time representation model and the partitioning decision of the bi-level optimization model until the preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.
[0006] Furthermore, in some embodiments of this disclosure, a two-layer optimization model for distribution network zoning is established, including: Define an objective function that minimizes the weighted sum of the loss of self-balancing ability of nodes within a partition, the mutual assistance demand of nodes between partitions, cross-scenario impact, cross-regional energy exchange cost, and partition stability penalty. The configuration constraints include at least: node allocation constraints, regional base constraints, electrical adjacency constraints, and multi-timescale balance feasibility constraints. Among them, the node allocation constraints are used to ensure that each node is uniquely assigned to a partition, the regional base constraints are used to constrain the range of the number of nodes in each partition, the electrical adjacency constraints are used to constrain the electrical coherence within each partition through impedance amplitude, inverse reactivity, and spatial attenuation factor, and the multi-timescale balance feasibility constraints are used to ensure that the net load deviation of nodes in the same partition, the aggregate imbalance of photovoltaic fluctuations and energy storage regulation in multiple scenarios is lower than the preset imbalance threshold.
[0007] In some embodiments of this disclosure, the constraints further include at least one of the following functional consistency constraints: region boundary consistency constraint, node temporal correlation constraint, multi-scale feasibility constraint, and intra-partition functional similarity constraint, wherein, The regional boundary consistency constraint is used to calculate a boundary penalty term for any two nodes divided into different partitions based on the differences in equivalent impedance, electrical distance and demand response capability between them. The constraint ensures that the total value of the boundary penalty term after integration and summation under the preset disturbance scenario does not exceed the preset penalty threshold. The node temporal correlation constraint ensures that the weighted correlation coefficient of the net balance deviation predicted for any two nodes under different potential variation scenarios is not lower than a preset deviation threshold. Multi-scale feasibility constraints are used to ensure that the weighted sum of the deviations between the operational status indicators of any node under multiple preset operating mechanisms and multiple time scales and the average status indicators of the node in its respective partition is less than a preset node-level feasibility threshold. The functional similarity constraint within a partition is used to constrain any two nodes assigned to the same partition to have a combined difference in predicted imbalance, photovoltaic power output fluctuation patterns, and dynamic characteristics of energy storage regulation trajectories that is lower than the maximum difference threshold allowed by the partition.
[0008] In some embodiments of this disclosure, extracting the temporal features of each node operating at multiple time scales to generate node temporal embedding vectors includes: An attention mechanism encoder is used to process time-series operational data, capture the long-term dependencies of the operational behavior of each node in the distribution network, and output the preliminary time-series feature vectors of each node. The initial temporal feature vector is fused with the corresponding node's running state information, and a latent state embedding vector is generated through nonlinear mapping to characterize the node's dynamic running features.
[0009] In some embodiments of this disclosure, a time representation model and a bilayer optimization model are co-trained using differentiable relaxation techniques and gradient feedback mechanisms, while simultaneously updating the parameters of the time representation model and the partitioning decision of the bilayer optimization model, including: A classification function with a temperature parameter is used to perform a continuous approximation on the decision variable of node partitioning, resulting in a continuous probability distribution. In each iteration of the collaborative training of the time representation model and the bi-layer optimization model, the gradient of the objective function with respect to the probability distribution is calculated. The gradient is then backpropagated to the parameters of the time representation model using the chain rule to generate a pseudo gradient. The pseudo gradient is used to characterize the sensitive direction of the objective function's value to changes in the latent state embedding vector. The parameters of the time representation model are updated by combining the pseudo gradient, the prediction error loss of the time representation model, and the cross-layer consistency loss. The cross-layer consistency loss is the error between the probability distribution and the expected value of the decision variable. Based on the sensitive direction of the pseudo-gradient representation, the time-series running data is perturbed in a directional manner to generate enhanced time-series running data. The enhanced time-series running data is used as the training data for the time representation model in the next iteration to drive the optimization of the partitioning decision in the bi-layer optimization model.
[0010] In some embodiments of this disclosure, the preset convergence condition is that the change in the comprehensive projection value of all partitions between two adjacent iterations is less than a preset change threshold. The comprehensive projection value is a comprehensive prediction score of the internal coordination and self-balancing potential of the partitions in all time periods.
[0011] A second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described power distribution network partitioning method.
[0012] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for dividing power distribution network zones.
[0013] A fourth aspect of the present invention provides a computer program product, the computer program product including computer instructions, the computer instructions instructing a computer to execute the above-described power distribution network zoning method.
[0014] The fifth aspect of this invention provides a power distribution network zoning system, comprising: The optimization layer construction module is used to establish a two-layer optimization model for the partitioning of the power distribution network. The upper layer of the two-layer optimization model is configured to optimize the partitioning decision with node partitioning as the decision variable and maximizing the self-balancing capability of the partition as the objective. The lower layer of the two-layer optimization model is configured to simulate the operational feasibility of the partitioning decision through a set of constraints. The learning layer construction module and the optimization layer construction module are used to construct a time representation model based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network. The attention mechanism is used to extract the time-series features of each node at multiple time scales and generate the potential state embedding vector of each node. The collaborative training module uses the node dynamic operation characteristics represented by the latent state embedding vector as input parameters of the bi-level optimization model to solve the bi-level optimization model and determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation techniques and gradient feedback mechanisms to collaboratively train the time representation model and the bi-level optimization model, while updating the parameters of the time representation model and the partitioning decision of the bi-level optimization model until the preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.
[0015] The present invention discloses a method, electronic device, and system for dividing power distribution network zones, which, compared with related technologies, have the following advantages: 1. This invention combines deep learning-based temporal embedding with optimization-driven region partitioning. By constructing a time-series representation model, it directly extracts potential features from the operating curves of photovoltaics, loads, and energy storage, identifying complementary behavioral patterns such as ramp-up synchronicity, surplus overlap, and unbalance curvature. Region partitioning is dynamically generated based on deep temporal embedding features, enabling region boundaries to evolve dynamically according to operational feasibility, rather than relying on static electrical connections, fundamentally solving the problem of the disconnect between static partitioning and dynamic operation. 2. This invention can identify and utilize the functional complementarity naturally formed by distributed photovoltaics, flexible loads and energy storage dynamics at multiple time scales, aligning regional division with temporal complementary structures, effectively reducing net load fluctuations within the region, improving balance stability across multiple time periods, and effectively suppressing cross-regional energy exchange during periods of high fluctuations, making energy storage dispatch more smooth and efficient, and significantly reducing the imbalance index during the midday surge of photovoltaics and the evening load peak. 3. This invention establishes an iterative feedback mechanism between the time representation learning layer and the region partitioning optimization layer through differentiable relaxation technology. Gradient information guides the learning system to adjust towards a representation direction that better aligns with the operational feasibility of the region, while the optimization layer optimizes the partitioning based on the newly learned representations. This bidirectional collaborative relationship unifies deep behavioral prediction and structural planning decisions, providing a systematic methodological support for building more resilient, self-balancing, and flexible power distribution networks. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of a power distribution network partitioning method provided by the present invention; Figure 2 An example diagram of a three-dimensional multi-timescale net load map provided by the present invention; Figure 3 An example diagram of a photovoltaic-load combined density distribution provided by the present invention; Figure 4 An example diagram illustrating a patterned potential space for photovoltaic-load dynamics provided by the present invention; Figure 5 An example diagram of a multi-timescale net load mosaic for potential runtime embedding provided by the present invention; Figure 6 An example diagram of the seasonal envelope of photovoltaic power output provided by the present invention; Figure 7 An example diagram illustrating the evolution of inter-node complementarity with curvature differences across multiple time scales, provided by this invention; Figure 8 An example diagram illustrating the optimization of a two-layer loss function between prediction embedding and partitioning provided by this invention; Figure 9 An example diagram of a partition structure in a learned spatiotemporal behavior embedding provided by the present invention; Figure 10 This invention provides a block diagram of a power distribution network zoning system. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and marked in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0020] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0021] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper," "lower," "horizontal," or "inner" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of the invention is in use, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, terms such as "first" and "second" are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] Furthermore, the use of the term "horizontal" does not imply that the component must be absolutely horizontal, but rather that it can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0023] In the description of the embodiments of the present invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention according to the specific circumstances.
[0024] The present invention will now be described in further detail with reference to the accompanying drawings: See Figure 1 This invention discloses a method for dividing power distribution network into zones, comprising the following steps: S101. Establish a two-layer optimization model for the partitioning of the distribution network. The upper layer of the two-layer optimization model is configured to optimize the partitioning decision with node partitioning as the decision variable and maximizing the self-balancing capability of the partition as the objective. The lower layer of the two-layer optimization model is configured to simulate the operational feasibility of the partitioning decision through a set of constraints.
[0025] In some embodiments of this disclosure, the two-layer optimization model aims to formalize the operational mechanism by which distributed photovoltaic, time-varying loads, and energy storage devices jointly influence the self-balancing capacity of each candidate region. Given that the structural characteristics of a region are closely related to the temporal behavior patterns of its constituent nodes, the model construction should not only describe instantaneous power relationships but also characterize dynamic interactions across multiple time scales, such as ramp-up synchronicity, intraday complementarity, energy storage recovery cycles, and imbalance propagation.
[0026] In some embodiments, the distribution network is represented as a set of regions, where the behavior of each region is defined by its aggregated power flow, internal resource distribution, and time-varying operating status. This decision structure integrates the time series of hourly generation, load, and energy storage charging / discharging states, thereby quantifying the ability of a set of nodes to maintain power balance without external support. At the core of the model is an objective function that evaluates the overall quality of region partitioning, comprehensively assessing partitioning schemes from dimensions such as temporal consistency, balance stability, and minimization of cross-regional energy exchange. Around this objective, the model introduces a series of constraints to describe power conservation, energy storage dynamics, operational feasibility, and temporal coordination relationships between different time levels.
[0027] In some embodiments, a series of mathematical models are constructed by combining the modeling elements of the above-mentioned two-layer optimization model, as shown in Equations 1-16, so that the solution of the region partitioning problem can reflect the inherent behavioral characteristics of distributed resources over time, rather than relying solely on the static topological relationship of their electrical connections.
[0028] In some embodiments, an objective function is defined to minimize the weighted sum of the loss of self-balancing ability of nodes within a partition, the mutual assistance demand of nodes between partitions, cross-scenario impact, cross-regional energy exchange cost, and partition stability penalty. The objective function is modeled as shown in Equation 1: (Formula 1); Formula 1 establishes a system-level objective for multi-time, state-regulated systems, where the predicted self-balancing capability of each node is defined. The necessity of mutual support Cross-influence with state dependencies , Woven into all time periods A single function standard. Based on region embedding. The coupling penalty of the squared Euclidean norm stabilizes the region in time, ensuring that abrupt boundary shifts are discouraged unless strongly demonstrated by the predicted equilibrium sensitivity. Parameters , , , The interactions between them provide a rich set of multi-weighted designs, where physical feasibility, spatiotemporal coherence, and predictive functional complementarity simultaneously influence the resulting mesh generation. This represents the set of global decision variables, which includes the time-series characteristic-related decision parameters of all nodes.
[0029] in, It represents the set of global operating state variables, including the power balance state of nodes, photovoltaic fluctuation response state, etc. This represents the set of variables related to global constraints, including parameters associated with constraints such as partition size and electrical connectivity. It indicates a specific time period, covering all runtime periods. This represents the set of all time periods to be considered. This represents a single node in a power distribution network. It represents the set of all nodes in the distribution network. This indicates an auxiliary node subscript, used to indicate its relationship to the node. Other nodes that interact with each other. This indicates the specific operating scenario. This represents the set of all operational scenarios to be considered. Indicates time period The weighting coefficient of "self-balancing capacity loss" reflects the degree of importance attached to the node's autonomous balancing capacity at different time periods. Indicates time period The weighting coefficient of "mutual assistance demand" is adjusted to regulate the intensity of the penalty for nodes' external support demands. Representing a scene The weighting coefficient for "cross-scenario impact" reflects the priority of extreme or typical scenarios. Indicates time period The weighting coefficient of "inter-regional energy exchange" controls the penalty for energy flow between regions. Indicates time period The weighting coefficient of "partition stability" is adjusted to regulate the penalty intensity for sudden changes in partition boundaries. Point During the period The self-balancing capability loss function has a larger value, indicating that the node's ability to autonomously maintain power balance is weaker. Represents a node During the period The predicted net load deviation (unit: kW) reflects the real-time fluctuation of the difference between load and photovoltaic output. Represents a node During the period The photovoltaic power fluctuation coefficient is used to quantify the intermittent intensity of photovoltaic output. Represents a node During the period The mutual support demand function has a larger value, indicating that the node needs more external support to maintain balance. Represents a node During the period The energy storage regulation potential (unit: kW) reflects the upper limit of the energy storage equipment's ability to smooth out fluctuations. Represents a node During the period Scene The cross-scenario impact function quantifies the interference of different scenarios on node operation. Represents a node and During the period The cost function for cross-regional energy exchange is such that the larger the value, the worse the economic efficiency of cross-regional power supply between two nodes. Represents a node During the period The predicted net load deviation. Represents a node and Electrical distance factor. Indicates time period The partition embedding vector below is an optimization variable. , , The comprehensive representation reflects the structural characteristics of the partition. Indicates the previous period The partition embedding vector.
[0030] Furthermore, the configuration constraints include at least: node allocation constraints, regional base constraints, electrical adjacency constraints, and multi-timescale balance feasibility constraints. Node allocation constraints ensure that each node is uniquely assigned to a single partition; regional base constraints limit the range of node numbers within each partition; electrical adjacency constraints constrain electrical continuity within each partition using impedance amplitude, inverse reactivity, and spatial attenuation factors; and multi-timescale balance feasibility constraints ensure that the net load deviation, photovoltaic fluctuations, and aggregate imbalance of energy storage regulation within the same partition are below a preset imbalance threshold across multiple scenarios. The modeling of node allocation constraints, regional base constraints, electrical adjacency constraints, and multi-timescale balance feasibility constraints is shown in Equation 2-5. (Formula 2); Among them, the node allocation constraint requires each node Must be through a binary selector It belongs to one and only one region. Although this structure appears simple on the surface, its underlying structure is profoundly significant: every predicted balance resilience, every energy storage recovery trajectory, and every peak coincidence signal generated by the lower-level learning model must ultimately converge on a unique discrete choice for each node. This uniformity requires enforcing a clear partition geometry, avoiding fractional uncertainty, and forcing the optimization process to negotiate combined conflicts between spatial proximity, electrical connectivity, and future balance compatibility, thereby enabling the learned representations to meaningfully shape topology decisions. This represents a binary allocation variable (0-1 variable) used to define the affiliation between nodes and partitions. A value of 1 indicates a node... Assigned to partition When the value is 0: node Not assigned to a partition . This represents a single functional zone that is ultimately divided in the distribution network. This represents the complete set of all functional partitions to be divided.
[0031] (Formula 3); Among them, the regional cardinality constraint is achieved through a lower limit and an upper limit. and The allowable size of each region is controlled by a specific structure. This structure prevents the emergence of degenerate regions and forces the composite layer to allocate nodes to ensure that each region has a meaningful balancing potential. Notably, there is a subtle interaction between these size constraints and the self-balancing embedding predicted by the Transformer: regions with too many highly volatile nodes will violate stability constraints, while regions with too few nodes may not fully utilize beneficial complementarity. Thus, the cardinality boundary creates a tension between electrical adjacency and functional equivalence, forcing optimization to navigate through high-dimensional potential compatibility. Indicates partition Minimum number of nodes allowed. Indicates partition The maximum number of nodes allowed.
[0032] (Formula 4); Among them, electrical adjacency constraints are achieved by using impedance amplitude. Reverse reactivity and spatial decay factor To enforce region-level connectivity. This is achieved through physical distance. An exponential decay is introduced to ensure that electrically distant nodes are not prioritized unless their predicted balance embedding can justify such grouping. Threshold Effectively ensuring an electrically consistent backbone network for each region aligns predicted functional similarity with the underlying power flow topology. This creates a multi-layered consistency condition where physical line attributes work in conjunction with learned cross-node similarities. It represents the set of all electrical connection branches (such as lines and transformers) in the distribution network, reflecting the physical connectivity between nodes. Represents a node and There is a direct electrical connection between them. Represents a node and The impedance magnitude (unit: Ω) of the electrical connection between them reflects the strength of the electrical connection. Represents a node and Reactance value of the electrical connection between them (unit: Ω) It is the reciprocal of the reactance. This represents a preset constant that controls the degree to which physical distance attenuates the weight of electrical connectivity. Represents a node and The actual geographical distance between them. Indicates partition Minimum electrical connectivity requirements that must be met.
[0033] (Formula 5); Multi-timescale balancing feasibility conditions will predict net load deviations. Photovoltaic fluctuations and energy storage recovery adjustment Integration in a space of potential change In the middle. By using The weighted integral captures how fluctuations across operational microstates propagate to nodal-level equilibrium stresses. The upper bound on the right... This reflects an acceptable level of imbalance, while involving , , The summation of these factors enables nodes within the same region to compensate for each other's fluctuations. This formula embodies the core of the proposed region partitioning concept: the grouping of nodes is not merely based on topological or geographical proximity, but on their collective ability to absorb, mitigate, and restabilize multi-timescale imbalance trajectories predicted by the learning model. This represents a single running micro-scene. This represents a collection of subdivided operational micro-scenes. Represents a node During the period The overall imbalance across multiple scenarios. Representing micro-scenes Importance weights. Represents a node During the period Micro-scenes The predicted net load deviation (unit: kW) is the difference between the load and the reference output. Represents a node During the period Micro-scenes The photovoltaic power fluctuation (unit: kW) reflects the intermittency of photovoltaic power. Represents a node During the period Micro-scenes The energy storage charging and discharging regulation (unit: kW) is adjusted to smooth out fluctuations. Indicates time period The maximum unbalance threshold acceptable to the next node (preset based on the grid frequency regulation capability and energy storage capacity). Indicates the node under time period τ and The complementary adjustment weight (the larger the value, the stronger the complementary cancellation effect). This represents a binary allocation variable (0-1 variable) used to define the affiliation between nodes and partitions. A value of 1 indicates a node... Assigned to partition When the value is 0: node Not assigned to a partition .
[0034] In some embodiments of this disclosure, the constraints also include at least one of the following functional consistency constraints: region boundary consistency constraint, node temporal correlation constraint, multi-scale feasibility constraint, and intra-partition functional similarity constraint, modeled as shown in Equations 6-9: (Formula 6); The above inequalities construct a high-resolution representation of the region boundary consistency. Whenever a node... and When located in different regions, mismatched terms Activated, based on factors similar to impedance. Inverse weighting and distance decay effect This results in a penalty contribution. To incorporate temporal dynamics, the flexibility of the predicted demand response varies. Additional functional dissimilarity penalties are introduced. Furthermore, the potential perturbation domain... Integrals on the integral introduce boundary tensions under uncertain conditions. , and by Weighted summation. The entire summation result. The upper bound ensures that the generated region boundaries maintain electrical and functional consistency under fluctuating operating modes. This represents a set of potential disturbance scenarios. This represents a single potential disturbance scenario, used to cover scenarios that are not being run. Includes sudden interference. Represents a node and The equivalent impedance factor between the two points reflects the electrical conduction characteristics at the boundary. Represents a node and Boundary weight coefficients between Use its reciprocal to adjust the boundary importance. This represents a preset constant that controls the degree to which physical distance attenuates boundary compatibility. Represents a node and The actual geographical distance. The boundary compatibility factor represents a boundary that decreases exponentially with increasing geographical distance, reflecting the "reduced requirements for functional consistency at long-distance boundaries". Represents a node and The weighting of demand response coordination between different parties adjusts the penalty intensity for differences in demand response. Represents a node During the period The predictive demand response capability (unit: kW) reflects the flexibility of load regulation. Represents a node and The difference in demand response flexibility and the difference in quantifiable functional complementarity. Indicates a disturbance scenario The importance weighting is higher for scenarios with extreme disturbances. Indicates a disturbance scenario Next, node (belonging to the zone) )and (Not part of the zone) The boundary tension reflects the boundary stability. Indicates partition The maximum permissible boundary tension threshold.
[0035] (Formula 7); This cross-correlation constraint is based on two nodes. and Predicting future balance residuals and applying them to a rich space of potential variations. The above is used for integration. The weighted product of the residuals is obtained through... Capture, measuring whether two nodes experience fluctuations of synchronization, desynchronization, or structural orthogonality. Lower bound. This prevents dangerous negative correlations that could amplify imbalances during periods of stress. (Variable) Threshold relaxation is allowed only when two nodes are located in the same region, thus ensuring that region clustering is consistent with the predicted cross-temporal correlation patterns. Essentially, this constraint ensures that nodes within a region are grouped according to their unbalanced trajectories sharing compatible temporal characteristics. This represents a refined set of potential variation scenarios. This represents a single potential variation scenario, used to capture the impact of fine-grained temporal fluctuations on node correlation. The importance weight of the variation scenario (xi) is represented by the higher weight of the high volatility scenario, which strengthens the constraint of temporal correlation. Represents a node In the mutation scenario The net balance deviation below. Represents a node In the mutation scenario The net balance deviation below. Indicates time period The minimum allowable correlation threshold (negative value) for the next node pair is set to prevent excessive negative correlation from amplifying the imbalance. Indicates time period The correlation threshold relaxation is applied only when the node and Effective within the same region.
[0036] (Formula 8); Multi-timescale feasibility requires the use of state weights. Aggregate node-level imbalance risks into a set of operational states. The squared imbalance reflects the magnitude of the deviation from the predicted balance. This is further expressed by the index relative to the regional average. Time layer deviation Supplementing and weighting by time scale Adjustments will be made. (By) The upper bound constraint ensures that each node maintains a feasible equilibrium state across multiple forward-looking time scales. Importantly, the combination of squared state integrals and time scale biases creates a multi-axis feasibility criterion that captures both dynamic and equilibrium structural behavior. This represents the set of operating mechanisms of the distribution network (such as routine operation, standby dispatch, emergency response, etc.). This represents a single operating mechanism used to cover node behavior under different scheduling modes. Indicates the operating mechanism The importance weighting is higher for key mechanisms such as emergency response. It represents a set with multiple time scales. It represents a single time scale and is used to capture the node balance characteristics at different durations. The importance weight of the time scale h is indicated, and the weight of long-term time scales (such as 24 hours) can be appropriately increased. Represents a node During the period Time scale The following are time-series characteristic indicators. Indicates partition z in time period Time scale The average time-series characteristic index reflects the overall behavior of the partition. Represents a node With partitions The difference in average temporal characteristics quantifies the consistency of behavior between nodes and partitions. Indicates time period The maximum permissible multi-mechanism-multi-scale balance deviation threshold for the next node.
[0037] (Formula 9); This constraint introduces a functional similarity requirement within each region. When a node and When they share a region, the weighted sum of their prediction equilibrium variances is: Photovoltaic diffusion Differences in storage trajectory Must be kept at the regional threshold The following standard guarantees that grouped nodes have a compatible variability structure, thus preventing regions from combining nodes with completely mismatched dynamic behaviors. PV differences in the variable domain. Integrals on the time plane introduce deep nonlinear time coupling, enriching the expressive power of constraints. Indicates time period The weighting coefficient of the imbalance of the next node is used to adjust the priority of balance stability in functional similarity. Indicates time period The weighting coefficients for differences in energy storage trajectories are used to strengthen the consistency requirements of energy storage regulation behavior. Indicates partition During the period The maximum permissible functional difference threshold.
[0038] In some embodiments, the constraints further include the following constraints that enhance behavioral coordination and resilience: behavioral synchronicity constraints, dynamic response matching constraints, multi-scale evolutionary consistency constraints, complementary coordination strength constraints, dynamic instability suppression constraints, extreme scenario consistency constraints, and variable domain constraints, modeled sequentially as shown in Equations 10-16: (Formula 10); The above inequality enforces the condition across the composite manifold. Synchronization conditions for node-wise time intensity integrands. Each integrand combines photovoltaic injection, net load characteristics, and storage-response trajectories into a single functional expression scaled by institutional sensitivity weights. Minimum synchronization boundary To prevent destructive divergences between node configuration files. When nodes belong to the same region (by... (Capture), synchronization threshold is relaxed This enables the region to capture the beneficial complementarity between nodes that support joint self-balancing by capturing their dynamic behavior. It represents a set of manifolds that integrate PV, load, and energy storage dynamics, and is used to comprehensively evaluate the time-series synchronization of multiple resources. It represents a single composite scenario, simultaneously covering the coordinated changes in PV, load, and energy storage. Representing a composite scene The importance weighting is higher for scenarios with high collaboration requirements (such as evening rush hour). Indicates time period The minimum allowed synchronization threshold (negative value) for the next node pair is set to prevent excessive divergence from disrupting the balance. Indicates time period The synchronicity threshold relaxation is applied only when nodes are in the same region, allowing for moderate asynchrony.
[0039] (Formula 11); Among these, the sensitivity differences in PV skewness or net load curvature often reveal local dynamics prone to instability. Using the state domain... The first and second derivatives on the first and second derivatives are used to measure the different responses of two nodes to changes in operating conditions. Additional headroom is allowed only when nodes share a region. This allows partitioning to favor node pairs that are naturally aligned in terms of slant amplitude and curvature, while still giving them the flexibility to enhance regional stability through complementary behavior. This represents the set of response mechanisms to changes in operating conditions. This represents a single response mechanism used to capture the dynamic response characteristics of a node to a specific disturbance. Indicating response mechanism The weighting coefficients for PV ramping differences are used to adjust the ramping consistency priority. Indicating response mechanism The weighting coefficient for the difference in net load curvature is used to reinforce the requirement for consistency in the trend of curve changes. Represents a node In response mechanism The first derivative of the PV output is used to quantify the speed and direction of the ramp. Represents a node In response mechanism The first derivative of the PV output is used to quantify the speed and direction of the ramp. Represents a node In response mechanism The second derivative of the net load is used to quantify the steepness of load change. Represents a node In response mechanism The second derivative of the net load is used to quantify the steepness of load change. Indicates time period The maximum permissible response difference between the next node pair. Indicates time period The response difference relaxation amount only takes effect when nodes are in the same region, allowing complementary responses.
[0040] (Formula 12); For regions that must withstand transitions between short-term, hourly, and diurnal states, multi-level consistency across features is crucial. Crucial. In the context manifold Integrating the deviation of the average trajectory over the region yields a uniform measure of dissimilarity. Constraining this composite quantity ensures that each region contains nodes that maintain consistency in its multi-scale evolution, preventing unstable mismatches that could disrupt region-level equilibrium over longer periods. It represents a set of scene manifolds at multiple scales, integrating the characteristics of operating scenes at different time scales. It represents a single multi-scale scene, covering short, medium, and long-term operational status. Representing time scale manifold scene The joint weighting coefficients adjust the importance of different scales. Represents a node During the period Time scale h, manifold scene The temporal characteristics below. Indicates partition During the period Time scale manifold scene The average time series characteristics under the given conditions. Indicates partition During the period The maximum permissible multiscale difference threshold.
[0041] (Formula 13); When the product of the unbalanced components between nodes is in a potential load-photovoltaic-storage scenario When favorable alignment is observed, pairwise cooperation strength between nodes emerges. This is achieved through... The weighted integral determines whether two nodes reinforce or cancel each other out in terms of volatility. Minimum buffer. Ensuring resilience, while maintaining a cooperation threshold. Clustering nodes are encouraged, as their predicted trajectories provide complementary stabilizing effects. This represents a set of scenarios that integrate the synergistic effects of PV, load, and energy storage, and is used to evaluate the complementary stability effect of node pairs. It represents a single collaborative scenario, covering the node's operational status under the combined effect of multiple resources. Representing collaborative scenarios Importance weights. Indicates time period The minimum cooperative resilience buffer ensures that node pairs have basic stability even without strong complementarity. Indicates time period The required level of collaboration between nodes is only effective when nodes are in the same region, and it mandates positive complementarity.
[0042] (Formula 14); Among them, the index imbalance index and For prediction in the perturbation domain Nodes exhibiting high deviations incur a sharply increased penalty. (This is followed by a seemingly unrelated sentence fragment: "In passing...") This ensures that only nodes capable of exhibiting bounded imbalance strength under different states are allowed to cluster together. Partition architecture. This represents a set of scenarios covering various disturbances (such as equipment fluctuations and sudden load changes), used to test node stability. This represents a single disturbance scenario, simulating the instability factors faced by a node. Indicates a disturbance scenario Next node The penalty coefficient for unbalanced intensity. Indicates a disturbance scenario Next node The penalty coefficient for unbalanced intensity. Represents a node The exponential penalty factor for imbalance intensity amplifies the penalty for high imbalance. Represents a node The exponential penalty factor for imbalance intensity strengthens the constraint on extreme volatility. Indicates partition z in time period The maximum allowable total amount of dynamic instability is determined to ensure the overall stability of the partition.
[0043] (Formula 15); Among them, the extreme event domain The underlying pattern typically determines the most severe stress conditions under which a region must survive. The squared difference between the PV load residual and the storage trajectory is determined by... and Weighting quantifies the discrepancies among high-impact scenarios. These discrepancies are then controlled within a threshold. The following approach only allows nodes with fully aligned extreme behavioral characteristics to remain in groups, ensuring robust regional performance under rare but significant disturbances. It represents a set of extreme scenarios that include strong disturbances such as extreme weather, sudden loads, and equipment failures. This represents a single extreme scenario, simulating the most severe operating conditions faced by the power grid. Indicating extreme scenarios The penalty coefficient for the combined deviation between PV and load amplifies the impact of extreme supply-demand imbalances. Indicating extreme scenarios The penalty coefficient for the difference in energy storage regulation trajectory is used to strengthen the requirements for coordinated energy storage response. Indicates time period The maximum permissible difference between the next node pair in extreme scenarios is determined to ensure consistent behavior.
[0044] (Formula 16); The domain restrictions of this formula on all variables ensure the feasibility of partitioning decisions and compatibility with the functional domain derived from learning. Binary discreteness of the assignment. The positivity of the main model parameters and the square integrability of the predicted imbalance components jointly ensure mathematical well-posedness. Furthermore, the smoothness of the region embedding... It prevents unrealistic fluctuations in regional configurations across consecutive layers and supports stable convergence of the two-layer learning optimization workflow. This indicates that the variable takes the value of a positive real number, ensuring the non-negativity of the physical quantity. This indicates that the variable belongs to the space of square-integrable functions, ensuring that integral calculations of imbalance, fluctuation, etc., are meaningful. Indicates time period The maximum allowable change in the sub-partition embedding vector controls the adjustment range of the partition boundary. Represents a node During the period Time scale The core time-series characteristic indicators.
[0045] S102. Based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network, an attention mechanism is used to construct a time representation model. The time representation model is used to extract the time-series features of each node at multiple time scales and generate the potential state embedding vector of each node.
[0046] It should be noted that by extracting potential time-series features from the operation curves of photovoltaics, loads and energy storage, complementary operational behaviors that are difficult to reveal by traditional trajectory analysis or clustering methods can be captured. These time-series embeddings constitute the information basis for the two-level decision-making process.
[0047] In some embodiments of this disclosure, extracting the temporal features of each node operating at multiple time scales and generating node temporal embedding vectors includes: processing the temporal operation data using an attention mechanism encoder to capture the long-term dependencies of the operating behavior of each node in the distribution network, and outputting preliminary temporal feature vectors for each node; fusing the preliminary temporal feature vectors with the corresponding node's operating state information, and generating a latent state embedding vector to characterize the dynamic operating characteristics of the node through nonlinear mapping. The modeling of the time representation model is shown in Equations 17 and 18 below: (Formula 17); Among them, rich multilevel representation The integral from this focus time, where the query key-value operator reflects cross-variable domains It exhibits Transformer-style associative behavior. Weighting involving the first-order feature derivative enhances the model's sensitivity to short-term transformations. Overall, this entity forms the predictive backbone of the forward-looking balanced embeddings for each node. Represents a node During the period Multi-time-series representation vectors. This represents a Transformer-style query-key-value attention operator that captures long-range dependencies in node time-series data. Represents a node In response mechanism Time period The query vector below is used for attention calculation. Represents a node In response mechanism Time period The key vector is matched with the query vector to calculate the attention weight. Represents a node In response mechanism Time period The features are obtained by summing the value vectors under attention weights. Represents a node In response mechanism Time period The original time-series feature vector. Represents the response mechanism of the original feature vector. The first derivative is used to capture short-term mutations and dynamic response trends of features. and Represents a node The feature transformation weight matrix is used to map the original features and derivative features to the same dimensional space.
[0048] (Formula 18); Among them, latent state embedding This stems from a hybrid approach combining the Transformer output with a state-conditional nonlinear projection. This integrated construction forces the predicted behavior to align with the underlying operating pattern while penalizing deviations from historical benchmark features. This hybrid approach significantly enriches the model's state awareness. Represents a node During the period The operating mechanism is embedded in the vector. This represents the set of core operating mechanisms of the power distribution network. This represents a single operating mechanism used to distinguish the behavioral characteristics of nodes under different scheduling modes. and The transformation weight matrix represents the temporal representation, which maps the high-dimensional temporal representation to the runtime mechanism embedding space. This represents the activation function of the rectified linear unit, introducing a nonlinear transformation to enhance the model's ability to express complex operating modes. Indicates the operating mechanism The bias term below, adjust The activation threshold is adapted to the feature distribution of different mechanisms. Represents a node During the period Time scale The historical time series characteristics are used as a benchmark for deviation comparison. Representing time scale The historical baseline deviation weight is used to adjust the constraint strength of historical features on the embedding vector.
[0049] S103 uses the node dynamic operation characteristics represented by the latent state embedding vector as the input parameters of the bi-level optimization model to solve the bi-level optimization model in order to determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation techniques and gradient feedback mechanisms to co-train the time representation model and the bi-level optimization model, while updating the parameters of the time representation model and the partitioning decision of the bi-level optimization model until the preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.
[0050] In some embodiments of this disclosure, a time-representation model and a bi-layer optimization model are co-trained using differentiable relaxation techniques and gradient feedback mechanisms, while simultaneously updating the parameters of the time-representation model and the partitioning decisions of the bi-layer optimization model: a classification function with a temperature parameter is used to approximate the decision variable based on node partitioning, resulting in a continuous probability distribution; in each iteration of the co-training of the time-representation model and the bi-layer optimization model, the gradient of the objective function with respect to the probability distribution is calculated, and the gradient is backpropagated to the parameters of the time-representation model using the chain rule to generate a pseudo-gradient. The pseudo-gradient is used to characterize the sensitive direction of the objective function's value to changes in the latent state embedding vector; the parameters of the time-representation model are updated by combining the pseudo-gradient, the prediction error loss of the time-representation model, and the cross-layer consistency loss, where the cross-layer consistency loss is the error between the probability distribution and the expected value of the decision variable; the time-series running data is directionally perturbed according to the sensitive direction represented by the pseudo-gradient to generate enhanced time-series running data, which is used as the training data for the time-representation model in the next iteration to drive the optimization of the partitioning decisions in the bi-layer optimization model.
[0051] Specifically, the modeling formulas for the differentiable relaxation process, gradient feedback calculation, generation of enhanced training samples, cross-layer consistency loss, and collaborative updating of model parameters are shown in Equations 19-23: (Formula 19); Among them, node-level feasibility score The predicted imbalance magnitude is combined with a state-sensitive latent projection. The squared component reinforces the penalty for large or unstable behavioral features, while the weighted latent interactions produce a comprehensive measure of expected stability under future perturbations. This score then guides region aggregation. Represents a node During the period Multi-scenario feasibility assessment. Representing micro-scenes The embedded vector weights are used to project the runtime mechanism embedded vectors into the scene adaptation space. This represents the dot product of the embedding vector of the operating mechanism and the scene weights, quantizing the embedding vector for the micro-scene. The degree of compatibility.
[0052] (Formula 20); Among them, partitioned projection It integrates node feasibility, PV pattern coherence within the region, and potential embedding uniformity. The regions generated by these hierarchical integrals maintain internal consistency across multiple behavioral dimensions, amplifying the synergistic effect of nodes with harmonious prediction characteristics. Indicates partition During the period The comprehensive projection value (the smaller the value, the stronger the synergy between partitions) is used as the core indicator for partition optimization. Indicates partition During the period The operating mechanism is embedded in the mean vector, reflecting the overall operating mode characteristics of the partition. It represents the square of the Euclidean norm, quantifying the difference between the node embedding vector and the partition mean, i.e., embedding homogeneity. Indicates the operating mechanism The embedding difference weight is used to adjust the priority of embedding homogeneity in partition projection.
[0053] Understandably, this disclosure introduces partitioned projection values. This transforms the discrete combinatorial decision problem of partitioning into a differentiable probabilistic allocation problem based on partition quality scores.
[0054] (Formula 21); Among them, random relaxation By temperature parameters The controlled softmax transformation introduces differentiability. This construction enables a smooth approximation of discrete assignments, allows gradients to flow through the partitioning process, and promotes consistent learning optimization and co-evolution. Represents a node Assigned to partition The probability (within the range [0,1]) replaces the discrete 0-1 variable to achieve differentiability. This means mapping the partitioned projection values to positive real numbers, providing a basis for probability normalization. This indicates the degree of smoothness controlling the probability distribution: The larger the value, the more concentrated the probability (close to a 0-1 discrete distribution); The smaller the value, the more even the probability. This represents an auxiliary identifier for traversing all partitions, used to calculate the normalized denominator.
[0055] (Formula 22); Among them, pseudo gradient It acts as a bridge from the composition layer back to the learning layer. This is achieved through the variable domain. The integration of this chain derivative captures how changes in underlying feature patterns affect the softened partition probability. This feedback enables the neural model to internalize region-level constraints. Represents a node Home partition The pseudo gradient reflects the direction and intensity of the influence of changes in the original features on the partition probability. It represents the partial derivative of the probabilistic attribution variable with respect to the embedding vector of the operating mechanism, quantifying the impact of the embedding vector on the partition probability. It represents the partial derivative of the runtime embedding vector with respect to the temporal representation vector, quantifying the impact of the temporal representation on the embedding quality. It represents the partial derivative of the temporal representation vector with respect to the original feature vector, quantifying the impact of the original features on the representation quality.
[0056] (Formula 23); Among them, the reinforcement rules based on residuals Synthesize new training samples aligned with the pseudo-gradient direction. Perturbation-weighted. and randomly Expanding the dataset to the surrounding decision-sensitive areas and enhancing the feasibility of learning partitions is the most vulnerable aspect. Indicates the first The enhanced original feature vectors are used to expand the training dataset. This indicates the index of the augmented sample, with a value ranging from 1 to... . This indicates the preset number of augmented samples, controlling the scale of dataset expansion. Indicates the first The pseudo-gradient direction perturbation weights of each enhanced sample control the feature adjustment magnitude. Indicates extraction of pseudo gradient The direction (positive → 1, negative → -1, zero → 0) guides the adjustment of the feature direction. Indicates the first The random perturbation intensity of each enhanced sample is introduced to avoid overfitting by incorporating appropriate randomness. This represents a random variable that follows a uniform distribution (usually taking values [-1, 1]), generating the direction of random perturbations.
[0057] (Formula 24); Among them, cross-layer consistency loss Penalize the deviation between the relaxed partition assignment and the final discrete decision. (In the scene manifold) The superposition of this bias forces the Transformer to learn a representation aligned with the boundaries of the stable region, thus avoiding overfitting to transient fluctuations. This represents the loss of consistency across layers. Related to the previous text They have the same meaning and serve as a unified identifier for partitioned sets in formula writing. This represents the final hard allocation result (0-1 variables), i.e., the node. Does it actually belong to a partition? . Representing a manifold scene The loss weights are adjusted, with higher weights for extreme scenarios, to enhance consistency in key scenarios.
[0058] (Formula 25); Among them, model parameters The iterative update combines the prediction loss gradient and the consistency loss gradient. This dual descent mechanism ensures that learning improves prediction accuracy and structural consistency in partition decisions, thereby producing uniform convergence behavior. This represents all trainable parameters of the learning model. Indicates the first The set of model parameters after the next iteration. Indicates the current iteration number. This indicates the next iteration. Indicates the first The set of model parameters after the next iteration is used for the next round of training. This indicates the step size for updating control parameters, avoiding oscillations caused by excessively large step sizes or slow convergence caused by excessively small step sizes. This represents the prediction error loss of the model based on time-series data such as PV, load, and energy storage.
[0059] In some embodiments of this disclosure, the preset convergence condition is that the change in the comprehensive projection value of all partitions between two adjacent iterations is less than a preset change threshold. The comprehensive projection value is a comprehensive prediction score of the internal coordination and self-balancing potential of the partitions across all time periods. The convergence model is as follows: (Formula 26); Convergence is indicated when successive partition projections exhibit negligible deviations during iteration. Consistent boundary conditions are also considered. This ensures the global stability of the partitioning structure and guarantees that the learning module and the combined decision layer are synchronized to a consistent balance. Indicates the first Sub-iteration, time period The projection value of the lower partition z. Indicates the first Sub-iteration, time period The projection value of the lower partition z. This means taking the maximum difference of all partition projection values to ensure global convergence. This represents the preset minimum allowable difference; if the difference is less than this value, convergence is considered.
[0060] Figures 2-9This paper presents the results and analysis of a case study demonstrating the application of the power distribution zoning method provided by this invention to a distribution network. The case study is based on a real medium-voltage distribution system. This system is an extension of the widely adopted IEEE 123-node feeder, comprising 142 nodes. By introducing additional branches and rooftop photovoltaic clusters, it simulates the spatial density characteristics of a California suburban power grid with high distributed energy penetration. The network includes 57 residential-commercial mixed load nodes, 31 photovoltaic-dominated nodes, and 12 nodes configured with residential energy storage systems ranging from 9.6 kWh to 36.8 kWh. Load data was collected at a 15-minute resolution, covering the entire operating cycle, and sourced from smart meter datasets calibrated with PGE and CAISO demand curves. Each node's time series has 35,040 points. The photovoltaic power generation time series is synthesized using the Perez transpose model based on NRELLSNRDB irradiance data, and the system capacity ranges from 4.2 kW residential photovoltaic strings to 112 kW commercial rooftop systems. To capture the multi-state dynamics required for potential embedding learning, the dataset was divided into six operating states: clear skies, partially cloudy, rapidly climbing clouds, evening rush hour, weekend load shifting, and heatwave stress conditions, each lasting approximately 900 to 2,400 time steps. To train the Transformer-based temporal learning module, all node-level sequences underwent normalization, seasonal trend decomposition, and gradient-based fluctuation extraction. Each node used a 6-layer encoder (8 attention heads per layer) to generate a 512-dimensional embedding vector, which was further integrated by a 320-dimensional state-aware fusion layer. The training set contained 80% of the time data throughout the year, 10% for validation, and 10% for out-of-sample testing. Multi-timescale prediction windows covered 15 minutes, 1 hour, 6 hours, and 24 hours, enabling the learning model to simultaneously predict short-term imbalances and cumulative energy storage recovery trends. The energy storage model was set with a round-trip efficiency of 92%, and maximum charge / discharge power ranging from 3.3kW to 9.9kW depending on system size. A 1-minute resolution was used to simulate the state of charge trajectory during high-climb photovoltaic events. The complementarity properties between nodes are calculated by cross-correlation of cross-spectral components, heliotropic alignment factors and multi-timescale unbalanced curvature, generating a 64-dimensional complementarity descriptor for each pair of nodes.
[0061] Figure 2This data manifold displays the spatiotemporal structure of net load dynamics throughout the year (35,040 data points), integrating high-resolution smart meter measurement data, NREL-derived photovoltaic irradiance curves, and multi-state behavioral load patterns. The vertical axis represents the net load amplitude, and the horizontal axis corresponds to the time of day (0-24 hours) and the number of days in a year (1-365 days). The Z-axis represents the net load value, ranging from 0 to 4.5, indicating the intensity of grid load fluctuations at different times and dates. Color changes represent different net load values, with darker colors indicating higher net load values. The specific numerical range of the colors corresponds to the color scale on the right, from red (high fluctuation area) to blue (low fluctuation area). Several key features can be observed from the data manifold: 1) During the summer months, strong radiation and dense rooftop photovoltaic installations create significant midday lows; 2) The evening rush hour has a prominent uphill climb, corresponding to the surge in electric vehicle charging and the high fluctuations caused by air conditioning load; 3) The seasonal base load exhibits asymmetry, with shallower troughs and weaker photovoltaic compensation in winter; 4) Rapid short-term fluctuations are mainly driven by cloud transients and load mutations caused by distributed energy sources.
[0062] The resulting multi-scale non-stationary structure indicates that the behavior of power distribution systems cannot be fully described by a single time scale. Instead, regional partitioning and operational learning models need to internalize nested fluctuation levels ranging from minute-level photovoltaic ramp-up to monthly irradiance cycles—which is precisely the motivation behind this study's proposal of a multi-time-scale Transformer embedding and two-layer learning-optimization framework.
[0063] Figure 3This study presents the joint density distribution of photovoltaic (PV) power generation and load demand under typical operating conditions, providing a high-resolution view of their interaction. Probabilistic mass is primarily concentrated in the region of 35-75 kW PV output and 128-155 kW load, indicating a strongly coupled operating baseline range for most sampling periods. The stretched morphology of the high-density region reflects a noisy linear co-variance, with a visually estimated regression slope of approximately 0.5 kW increase in load for every 1 kW increase in PV output, consistent with the intrinsic compositional relationships within the dataset. The distribution exhibits significant anisotropy: the highest density contour lines surround a region with a PV span of approximately 30 kW and a load span of approximately 20 kW, revealing the variability characteristic of midday operation—when cloud cover changes cause minute-level PV fluctuations, which are followed by responses from localized air conditioning, heating, and electric vehicle charging loads. Another notable feature appears in the mid-range region of 45-60 kW PV output, where the load distribution expands to a width of approximately 40 kW, roughly ranging from 120-160 kW. This extension reveals multiple behavioral patterns inherent at the same photovoltaic (PV) output level: these include residential clusters responding to time-of-use pricing, small commercial establishments operating on a fixed schedule, and EV charging stations with random start-stop cycles. This heterogeneity is directly relevant to the design of the learning framework, as it demonstrates that the similarity of the original time series alone is insufficient to establish a consistent regional division. The multi-leaf morphology of low-density areas further supports this observation: when cloudy days cause PV output to drop below 20kW, the load distribution becomes dispersed, ranging from 105kW to 150kW, indicating a decoupling between demand-side behavior and generation levels during periods dominated by fluctuations. This deviation is precisely the temporal irregularity that the underlying representation layer needs to learn and internalize.
[0064] Figure 4This study presents the compressed potential spatial representation of 35,040 photovoltaic-load sequences throughout the year, revealing six distinct operating states that emerge naturally from the Transformer temporal encoder. The clear-sky state presents as compact, low-variance clusters, with samples concentrated in the first potential dimension (15-10), exhibiting minimal dispersion in other dimensions, reflecting the stability of operating patterns under unobstructed irradiance conditions. Partially cloudy states significantly extend along potential dimension 2, often exceeding 50 units in span, reflecting high-frequency fluctuations caused by intermittent cloud changes. Rapidly climbing cloud events manifest as a curved transition structure extending nearly 40 units along potential dimension 1, corresponding to steep photovoltaic power changes caused by drastic irradiance fluctuations. Evening peak states are concentrated in the upper region of potential dimension 3 (values greater than 60), forming dense and rising clusters, typically associated with synchronous surges in residential and commercial loads. Weekend load shift states occupy a lower, horizontally wide strip of approximately 35 units, characterizing the slow, planned changes in consumption patterns relative to weekdays. Heat wave stress conditions extend along the negative depth of potential dimension 3, often exceeding 70 units, reflecting the extreme and sustained increase in temperature-driven air conditioning loads. These geometrical separations indicate that the potential space internalizes behavioral heterogeneity across multiple time scales, enabling the differentiation of patterns that are difficult to discern from raw measurement data alone, such as meteorological conditions, demand-driven changes, and fluctuation-dominated events.
[0065] Figure 5A multi-timescale net load mosaic plot reveals the structural fluctuation characteristics required to support learning-based region partitioning. The plot shows steep inclines in the morning and evening, ranging from 3.8 to 4.2 units, while the midday trough dominated by PV is approximately 0.8 to 1.1 units, highlighting the short-term volatility that must be captured through high-resolution embedding. Within 24 hours, the upward incline slope often exceeds 0.25 units / hour, while the PV-driven downward slope approaches 0.35 units / hour, reflecting the operational nonlinearity that the region partitioning algorithm must anticipate. The weekly-scale subplot presents a wider fluctuation envelope: the peak-to-trough difference can reach a maximum of 5.7 units on weekends, while the average weekday fluctuation is approximately 3.9 units. The seasonal panel highlights long-term differences: in summer, the midday trough due to PV occurs as low as 0.6 units, while in winter, the daytime trajectory is relatively flat, with net load rarely falling below 2.2 units. These patterns collectively demonstrate that the formation of functional power supply regions cannot rely on static indicators but requires a learning architecture to analyze the nested fluctuation structure from minutes to months. A comprehensive interpretation of the mosaic diagram reveals that no single time scale dominates the system's equilibrium pattern. For example, nodes A and B might both peak at 4.0 units on a typical daytime evening, but node A could jump to 6.1 units on the weekend, while node B only rises to 4.3 units, suggesting they are in different behavioral states. Therefore, a model capable of simultaneously fusing changes across multiple time scales is needed to extract potential time clusters representing such nested operational states. The staggered time scales presented in the diagram provide an intuitive theoretical basis for the Transformer-based embedding layer in the region partitioning framework.
[0066] Figure 6 This diagram illustrates the seasonal envelope of photovoltaic (PV) output, derived from irradiance simulations or measured data combined with a PV conversion model. The horizontal axis represents the PV output level, with the upper envelope peaking at approximately 8.5 kW around day 180, while the lower envelope typically falls below 1.2 kW during the winter trough (around day 355). The shaded area between the two curves reflects intraseasonal fluctuations caused by cloud cover, aerosol concentration, and stochastic atmospheric patterns. This shaded area exhibits uneven width distribution: for example, its vertical span often exceeds 3.0 kW between days 140 and 200, while shrinking to below 1.8 kW between days 250 and 300. This asymmetric fluctuation is significant in modeling because it directly impacts the Transformer-based architecture's analysis of PV state transitions. The envelope plot highlights the non-stationary seasonal characteristics: high output and high volatility in summer, and moderate output with highly irregular peaks in spring and autumn. These characteristics necessitate a learning mechanism with multi-timescale, highly expressive modeling capabilities, rather than relying on shallow statistical summaries. Furthermore, the figure shows that it is extremely rare for a photovoltaic system to operate at 7.0kW or higher for more than 30 consecutive days. This challenges the assumption of long-term stability and confirms the necessity of introducing time segmentation into the learning process.
[0067] Figure 7 The evolution of complementarity between two nodes with respect to their curvature differences across multiple timescales is presented, characterizing interactions at short, medium, and long-term scales. Complementarity values cover the entire range from 0 to 1, forming a smooth but strongly nonlinear manifold. Within the curvature difference range of approximately 3-5, the complementarity score drops below 0.15, constituting a high-conflict zone, indicating a severe misalignment between nodes in terms of ramp direction and timing. This region typically corresponds to significant deviations in daytime and sub-daytime load curvature, often caused by rapid photovoltaic load drops or asynchronous electric vehicle charging surges. When the curvature difference exceeds 8, the surface rapidly rises and enters a broad plateau region, with complementarity consistently above 0.85. These high-value regions represent a high degree of consistency in the nodes' dynamics across multiple timescales, such as synchronized evening peaks, consistent weekend load shifts, or shared smooth clear-sky photovoltaic conditions, making them mutually stable units within the region. The ridge-like structure visible along the diagonal reflects the underlying symmetry learned by the complementarity model: when the curvature amplitude increases synchronously across time scales, the model interprets it as structural reinforcement behavior even with different temporal patterns. Overall, the geometry of the surface suggests that complementarity is not a monotonic function of temporal differences, but rather a state-dependent metric whose value is influenced by the alignment or conflict between fluctuation patterns, diurnal curvature, and seasonal rhythms at the nodes.
[0068] Figure 8 This paper presents the optimization landscape of the two-layer loss function between the prediction embedding layer and the region partitioning layer, revealing how the training process gradually converges from initial inconsistency to a jointly feasible state. When the prediction embedding consistency is below 0.3 and the region balance feasibility is below 0.2, the loss value is in the high-energy region (greater than 4.5), indicating that the temporal embedding cannot support feasible region construction at this point. As the embedding consistency improves to the range of 0.6-0.8, the loss decreases sharply along a ridge-like descent channel, forming a pseudo-gradient feedback path that drives learning towards structural compatibility. This path illustrates that a small improvement in embedding consistency can trigger adjustments in the region partitioning layer, which in turn prompts the embedding model to make inverse corrections, forming the sloping valley structure unique to the coupled optimization mechanism. The deepest valley of loss occurs when the embedding consistency is close to 0.9-1.0 and the region feasibility is between 0.2-0.35, at which point the loss value drops below 0.4. This convergence region marks the system reaching an equilibrium point: the embedding can reflect temporally consistent node behavior, while the region partitioning layer can identify node groups that can maintain balance across multiple time scales. The curvature characteristics of the valley floor indicate that perfect embedding alignment is not a sufficient condition for global feasibility; feasibility can only be stably achieved when the two layers cooperate to adapt to their internal representations. Therefore, the topology of this loss landscape intuitively presents the core principle of this framework: the predictive and structural components are not optimized in isolation, but rather descend cooperatively into a coupled optimization landscape whose geometry encodes the system's long-term balancing ability.
[0069] Figure 9 The final region segmentation results obtained through a two-layer learning framework are presented, where feeder electrical distance, potential behavioral similarity, and balance intensity jointly determine the functional boundaries of each region. Region 1, located at the high end of the intensity axis, has a stable balance index between 0.75 and 0.95, forming a behavior-driven cluster: although its nodes are spatially separated, they possess highly consistent photovoltaic-load curves and stable multi-timescale operating states. Region 2 forms a broad medium-intensity platform, with standardized feeder distances between 45 and 85, and potential behavioral indices distributed between -2 and +2, indicating the simultaneous existence of weekday-weekend differences and photovoltaic-driven ramp-up behavior within this region. Region 3, located in the low-intensity, high-fluctuation range, has balance indices generally below 0.30 and spatial distances concentrated below 40, encompassing nodes susceptible to surges in electric vehicle charging and cloud fluctuations. Region 4 exhibits a compact and geographically localized cluster structure with intensity values above 0.8, indicating that some feeders with shorter electrical distances can maintain strong internal consistency even under multi-timescale perturbations. Region 5 covers the widest geographical area, with feeder distances exceeding 70-110 km, while maintaining a balance index of 0.6-0.85. This indicates that despite the physical dispersion of nodes, their underlying behaviors exhibit high complementarity at the temporal level, enabling the formation of stable regions. The geometry of the embedded space reveals that balance regions do not strictly adhere to physical proximity. The shape and orientation of each polyhedron actually reflect the interaction between behavioral states, fluctuation characteristics, and temporal complementarity, thus forming stable multi-node groupings. The results demonstrate that the regional division of the power distribution system is a product of the fusion of spatial topology and deep temporal behavior, rather than simply relying on geographical proximity.
[0070] The case study results show that regions formed based on latent temporal embedding exhibit stronger internal consistency, resulting in smoother net load curves, lower volatility, and a significant reduction in dependence on inter-regional power exchange during photovoltaic fluctuations or load surges. Empirical analysis demonstrates that when temporal diversity is properly quantified and embedded in the model, it can be transformed into a structural asset of the network. In regions where the synchronicity between photovoltaic surplus and load ramping is captured through learned representations, energy storage units respond more efficiently, charge-discharge cycles are optimized, and energy availability during peak hours is correspondingly improved. Imbalance analysis further reveals that the partitioned regions exhibit stronger intrinsic stability during midday photovoltaic peaks and evening load surges, while traditional partitioning methods often result in under-matched or weakly coupled clusters. These advantages remain stable across different operating scenarios, demonstrating the robustness of the learning-enhanced region partitioning method and its adaptability to typical and extreme operating conditions.
[0071] One embodiment of the present invention, such as Figure 10As shown, a power distribution network zoning system is provided, including: The optimization layer construction module 210 is used to establish a two-layer optimization model for the partitioning of the distribution network. The upper layer of the two-layer optimization model is configured to optimize the partitioning decision with the node partitioning assignment as the decision variable and the goal of maximizing the self-balancing capability of the partition as the objective. The lower layer of the two-layer optimization model is configured to simulate the operational feasibility of the partitioning decision through a set of constraints. The learning layer construction module 220 is used to construct a time representation model based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network, using an attention mechanism. The time representation model is used to extract the time-series features of each node's operation at multiple time scales and generate the potential state embedding vector of each node. The collaborative training module 230 is used to use the node dynamic operation characteristics represented by the latent state embedding vector as input parameters of the bi-level optimization model to solve the bi-level optimization model in order to determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation technology and gradient feedback mechanism to collaboratively train the time representation model and the bi-level optimization model, while updating the parameters of the time representation model and the partitioning decision of the bi-level optimization model until the preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.
[0072] It should be noted that the specific implementation methods of the embodiments disclosed herein are different from those of the embodiments in this paper. Figure 1 The principle of the embodiments shown is the same, and will not be repeated here.
[0073] One embodiment of the present invention provides an electronic device including a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a power distribution network zoning method.
[0074] In one embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). It should be noted that more specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CDROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0075] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0076] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0077] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the power distribution zone division method in the above embodiments.
[0078] One embodiment of the present invention provides a computer program product containing computer instructions for guiding a computer to execute a power distribution network zoning method. By integrating functions such as environmental data acquisition, model training, and leakage current prediction, the program product achieves effective monitoring of surge arrester leakage current.
[0079] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for dividing power supply zones in a power distribution network, characterized in that, The method includes: A two-layer optimization model for power distribution network partitioning is established. The upper layer of the two-layer optimization model is configured to optimize partitioning decisions with node partitioning as the decision variable and maximizing the self-balancing capability of the partitions as the objective. The lower layer of the two-layer optimization model is configured to simulate the operational feasibility of partitioning decisions through a set of constraints. Based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network, an attention mechanism is used to construct a time representation model. The time representation model is used to extract the time-series features of each node at multiple time scales and generate the potential state embedding vector of each node. The node dynamic operation characteristics represented by the latent state embedding vector are used as input parameters of the two-layer optimization model to solve the two-layer optimization model in order to determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation techniques and gradient feedback mechanisms to co-train the time representation model and the two-layer optimization model, while updating the parameters of the time representation model and the partitioning decision of the two-layer optimization model, until a preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.
2. The method for dividing power distribution network zones according to claim 1, characterized in that, The establishment of the two-level optimization model for distribution network zoning includes: Define an objective function that minimizes the weighted sum of the loss of self-balancing ability of nodes within a partition, the mutual assistance demand of nodes between partitions, cross-scenario impact, cross-regional energy exchange cost, and partition stability penalty. The constraints configured include at least: node allocation constraints, regional base constraints, electrical adjacency constraints, and multi-timescale balance feasibility constraints. The node allocation constraints are used to ensure that each node is uniquely assigned to a partition. The regional base constraints are used to constrain the range of the number of nodes in each partition. The electrical adjacency constraints are used to constrain the electrical coherence within each partition through impedance amplitude, inverse reactivity, and spatial attenuation factor. The multi-timescale balance feasibility constraints are used to ensure that the net load deviation of nodes within the same partition, the aggregate imbalance of photovoltaic fluctuations and energy storage regulation under multiple scenarios is lower than a preset imbalance threshold.
3. The method for dividing power distribution network zones according to claim 2, characterized in that, The constraints also include at least one of the following functional consistency constraints: region boundary consistency constraint, node temporal correlation constraint, multi-scale feasibility constraint, and intra-partition functional similarity constraint, wherein... The regional boundary consistency constraint is used to calculate a boundary penalty term for any two nodes divided into different partitions based on the differences in equivalent impedance, electrical distance and demand response capability between them, and to constrain the total value of the boundary penalty term after integration and summation under a preset disturbance scenario to not exceed a preset penalty threshold. The node temporal correlation constraint ensures that the weighted correlation coefficient of the net balance deviation predicted for any two nodes under different potential variation scenarios is not lower than a preset deviation threshold. The multi-scale feasibility constraint is used to ensure that the weighted sum of the deviations between the operating status indicators of any node under multiple preset operating mechanisms and multiple time scales and the average status indicators of the node in its respective partition is less than a preset node-level feasibility threshold. The functional similarity constraint within the partition is used to constrain any two nodes assigned to the same partition to have a combined difference in predicted imbalance, photovoltaic power output fluctuation patterns, and dynamic characteristics of energy storage regulation trajectories that is lower than the maximum difference threshold allowed by the partition.
4. The method for dividing power distribution network zones according to claim 1, characterized in that, The step of extracting the temporal features of each node operating at multiple time scales and generating node temporal embedding vectors includes: An attention mechanism encoder is used to process the time-series operation data, capture the long-term dependencies of the operation behavior of each node in the distribution network, and output the preliminary time-series feature vector of each node. The preliminary temporal feature vector is fused with the corresponding node's running state information, and a potential state embedding vector is generated through nonlinear mapping to characterize the node's dynamic running features.
5. The method for dividing power distribution network zones according to claim 2, characterized in that, The step of co-training the time representation model and the bi-layer optimization model using differentiable relaxation techniques and gradient feedback mechanisms, while simultaneously updating the parameters of the time representation model and the partitioning decision of the bi-layer optimization model, includes: A classification function with a temperature parameter is used to perform a continuous approximation on the node partitioning decision variable, resulting in a continuous probability distribution. In each iteration of the collaborative training of the time representation model and the two-layer optimization model, the gradient of the objective function with respect to the probability distribution is calculated, and the gradient is backpropagated to the parameters of the time representation model through the chain rule to generate a pseudo gradient. The pseudo gradient is used to characterize the sensitive direction of the value of the objective function to changes in the latent state embedding vector. The parameters of the time representation model are updated by combining the pseudo gradient, the prediction error loss of the time representation model, and the cross-layer consistency loss, wherein the cross-layer consistency loss is the error between the probability distribution and the expected value of the decision variable. The time-series running data is perturbed in a directional manner according to the sensitive direction of the pseudo-gradient representation to generate enhanced time-series running data. The enhanced time-series running data is used as the training data of the time representation model in the next iteration to drive the optimization of the partitioning decision in the two-layer optimization model.
6. The method for dividing power distribution network zones according to claim 5, characterized in that, The preset convergence condition is that the change in the comprehensive projection value of all partitions between two adjacent iterations is less than a preset change threshold. The comprehensive projection value is a comprehensive prediction score of the internal coordination and self-balancing potential of the partition in all time periods.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the power distribution zone division method of any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the power distribution zone division method according to any one of claims 1 to 6.
9. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the power distribution zone division method according to any one of claims 1 to 6.
10. A power distribution network zoning system, characterized in that, include: An optimization layer construction module is used to establish a two-layer optimization model for power distribution network partitioning. The upper layer of the two-layer optimization model is configured to optimize partitioning decisions with node partitioning as the decision variable and maximizing the self-balancing capability of the partitions as the objective. The lower layer of the two-layer optimization model is configured to simulate the operational feasibility of partitioning decisions through a set of constraints. The learning layer construction module and the optimization layer construction module are used to construct a time representation model based on the time-series operation data of photovoltaic output, load demand and energy storage status of each node in the distribution network, using an attention mechanism. The time representation model is used to extract the time-series features of each node's operation at multiple time scales and generate the potential state embedding vector of each node. A collaborative training module is used to use the node dynamic operation features represented by the latent state embedding vector as input parameters of the two-layer optimization model to solve the two-layer optimization model in order to determine the distribution network partitioning scheme. The solution process includes: using differentiable relaxation techniques and gradient feedback mechanisms to collaboratively train the time representation model and the two-layer optimization model, while updating the parameters of the time representation model and the partitioning decision of the two-layer optimization model, until a preset convergence condition is reached, and outputting the partitioning assignment of each node in the distribution network.