Power distribution network multi-measurement-point load prediction method and device based on controlled channel communication, storage medium and computer equipment

CN122823403APending Publication Date: 2026-09-25NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611282009.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请提供了一种基于受控通道通信的配电网多测点负荷预测方法及装置、存储介质、计算机设备,通过通道独立编码、分组局部通信、原型中继全局通信和受控残差融合的受控通信架构,解决了现有配电网多测点负荷预测方法中全局无约束混合建模存在的干扰传播与计算爆炸问题、以及对完整拓扑信息的依赖问题

Benefits of technology

[0007]依据本申请又一个方面,提供了一种存储介质,其上存储有计算机程序,所述程序被处理器执行时实现上述基于受控通道通信的配电网多测点负荷预测方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823403A_ABST
    Figure CN122823403A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of power distribution network load prediction, and particularly discloses a power distribution network multi-measuring-point load prediction method and device based on controlled channel communication, a storage medium and a computer device. The method comprises the following steps: obtaining historical load data of each measuring point of a power distribution network, and constructing load time sequence data samples corresponding to each measuring point according to the historical load data; based on the load time sequence data samples of each measuring point, a plurality of training samples are constructed by using a sliding window method; each training sample comprises load time subsequence samples of each measuring point in a same historical time window and real load labels of each measuring point in a prediction time window; a pre-constructed multi-measuring-point load prediction model is trained based on the plurality of training samples; load time subsequences of each measuring point in a target historical time window are input into the trained multi-measuring-point load prediction model, and load result prediction values of each measuring point in a target prediction time window are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power distribution network load forecasting technology, and in particular to a method and device for multi-point load forecasting of power distribution networks based on controlled channel communication, as well as storage media and computer equipment. Background Technology

[0002] As a critical link in the power system directly facing end users, the distribution network's operational status directly affects power supply reliability, electricity safety, and dispatch efficiency. With the widespread deployment of monitoring equipment such as smart meters, distribution transformer monitoring terminals, feeder monitoring devices, charging station access points, distributed energy access points, and energy storage access points, the number of load monitoring points that can be collected in the distribution network continues to increase. Accurately predicting load changes at multiple monitoring points within a future time window is crucial for overload risk identification, inspection plan development, load transfer, load limiting operation, and auxiliary dispatch decisions. However, distribution network load data is characterized by strong time-series dependence, strong coupling among multiple monitoring points, and complex fluctuation patterns, placing high demands on the accuracy and stability of prediction methods.

[0003] Existing methods for load forecasting in power distribution networks mainly include traditional time series forecasting methods based on statistical models, machine learning-based methods, and deep learning-based methods. Statistical model-based methods typically utilize autoregressive models, exponential smoothing models, or time series decomposition models for forecasting, but their ability to map complex nonlinear load fluctuations and sudden peak-valley changes is limited. While machine learning-based methods can learn the mapping relationship between historical loads and external variables, they often rely on manual feature construction when dealing with high-dimensional multi-measuring-point load sequences, making it difficult to fully model the dynamic correlations between multiple measuring points. In recent years, prediction methods based on recurrent neural networks, Transformers, and graph neural networks have made some progress, but they still have the following shortcomings in the application of synchronous prediction of multiple measuring points in distribution networks: First, some existing methods are mainly aimed at single measuring point or overall regional load prediction, and it is difficult to output the future load values ​​of multiple measuring points simultaneously, so they cannot directly support measuring point-level risk screening; Second, when performing global hybrid modeling of all measuring point channels, information from weakly correlated or abnormal measuring points may be incorrectly propagated to other measuring points, affecting prediction stability; Third, the computational overhead of global pairwise interactions increases rapidly with the number of measuring points, which is not conducive to deployment in large-scale measuring point scenarios; Fourth, graph neural network-based methods usually rely on a relatively accurate distribution network topology or dynamic adjacency matrix, but in multi-source measuring point scenarios on the user side, complete and real-time updated topology information is often difficult to obtain. Summary of the Invention

[0004] In view of this, this application provides a method and device for load forecasting of multiple measurement points in distribution networks based on controlled channel communication, as well as a storage medium and computer equipment. Through a controlled communication architecture that integrates independent channel coding, grouped local communication, prototype relay global communication, and controlled residual fusion, it solves the problems of interference propagation and computational explosion, as well as the dependence on complete topology information, that exist in existing methods for load forecasting of multiple measurement points in distribution networks using global unconstrained hybrid modeling.

[0005] According to one aspect of this application, a method for multi-measuring-point load forecasting in a distribution network based on controlled channel communication is provided, comprising: Historical load data of each measuring point in the distribution network is obtained, and load time series data samples corresponding to each measuring point are constructed based on the historical load data. Based on the load time series data samples of each measuring point, multiple training samples are constructed using the sliding window method. Each training sample includes load time subsequence samples of each measuring point in the same historical time window, as well as the actual load labels of each measuring point in the prediction time window. Based on the multiple training samples, a pre-constructed multi-measurement point load prediction model is trained. The multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to perform temporal coding on the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point and, based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; a prototype relay-type global communication module, used to perform information aggregation and backhaul between all measurement points and all global operating state prototypes based on a preset number of built-in learnable global operating state prototypes and the channel-independent representation, and construct global communication messages based on the aggregation and backhaul results; a controlled residual fusion module, used to perform residual fusion on the channel-independent representation, the local communication messages, and the global communication messages to obtain an updated representation; and a prediction output module, used to output the predicted load result value of each measurement point in the prediction time window based on the updated representation, so as to iteratively train the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label. The load time subsequence of each measuring point within the target historical time window is input into the trained multi-measuring-point load prediction model to obtain the predicted load result value of each measuring point within the target prediction time window.

[0006] According to another aspect of this application, a multi-point load forecasting device for distribution networks based on controlled channel communication is provided, comprising: The data acquisition module is used to acquire historical load data of each measuring point in the distribution network, and to construct load time series data samples corresponding to each measuring point based on the historical load data. The training sample construction module is used to construct multiple training samples based on the load time series data samples of each measuring point using the sliding window method. Each training sample includes load time subsequence samples of each measuring point in the same historical time window, as well as the actual load labels of each measuring point in the prediction time window. The model training module is used to train a pre-constructed multi-measurement point load prediction model based on the multiple training samples. The multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to perform temporal coding on the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point and, based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; and a prototype relay-style global communication module, used to train a pre-constructed multi-measurement point load prediction model based on a built-in preset number of data points. The system comprises a learnable global operating state prototype and a channel-independent representation, and performs information aggregation and backhaul between all measurement points and all global operating state prototypes. A global communication message is constructed based on the aggregation and backhaul results. A controlled residual fusion module is used to perform residual fusion of the channel-independent representation, the local communication message, and the global communication message to obtain an updated representation. A prediction output module is used to output the predicted load result value for each measurement point within the prediction time window based on the updated representation, and to iteratively train the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label. The prediction module is used to input the load time subsequence of each measuring point within the target historical time window into the trained multi-measuring-point load prediction model to obtain the predicted load result value of each measuring point within the target prediction time window.

[0007] According to another aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described method for multi-point load forecasting in a distribution network based on controlled channel communication.

[0008] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for multi-point load forecasting of a distribution network based on controlled channel communication.

[0009] By employing the above technical solutions, this application provides a method and device for load forecasting of multiple measuring points in a distribution network based on controlled channel communication, a storage medium, and a computer device. This method independently encodes each measuring point using a channel-independent timing coding module, preserving the load variation patterns of each measuring point while avoiding contamination from abnormal measuring point information, thus solving the problem of mutual interference between weakly correlated channels in traditional global hybrid coding. Furthermore, a local direct communication module restricts direct interaction between measuring points to within groups, allowing measuring points within the same group to share local collaborative information. This captures the mutual influence of neighboring measuring points and effectively suppresses noise propagation. Simultaneously, a prototype relay-type global communication module aggregates information from all measuring points to a preset number of learnable global operating state prototypes for relay interaction and backhaul, preventing the computational overhead of global communication from increasing dramatically with the number of measuring points, overcoming the resource bottleneck of large-scale global pairwise interaction. The above process is based solely on historical load data, without relying on the distribution network topology or adjacency matrix, and can be directly deployed in scenarios lacking complete topology information. Ultimately, the multi-measuring-point load forecasting model can simultaneously output the future load values ​​of all measuring points within the forecast time window, providing direct data support for overload risk screening and auxiliary dispatching decisions of the distribution network, while taking into account forecast accuracy, stability, and feasibility of engineering deployment.

[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This paper illustrates a flowchart of a multi-measuring-point load forecasting method for distribution networks based on controlled channel communication, provided in an embodiment of this application. Figure 2 This illustration shows a structural schematic diagram of a multi-measuring-point load prediction model provided in an embodiment of this application; Figure 3 This illustration shows a schematic diagram of the internal workflow of a local direct communication module provided in an embodiment of this application; Figure 4 This paper illustrates a schematic diagram of the internal working process of a prototype relay-type global communication module provided in an embodiment of this application. Figure 5 This illustration shows a diagram of the maximum risk value and the thresholds for attention, early warning, and criticality under different prediction time windows provided in an embodiment of this application. Figure 6 This illustration shows a schematic diagram of the high-risk measurement point event ranking results provided in an embodiment of this application; Figure 7 This illustration shows a comparison diagram of the future predicted load curve and capacity threshold of a measuring point corresponding to the highest risk event, provided in an embodiment of this application. Figure 8 This illustration shows a structural schematic diagram of a multi-measuring-point load forecasting device for a distribution network based on controlled channel communication, provided in an embodiment of this application. Figure 9 A schematic diagram of the device structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation

[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0013] This embodiment provides a method for multi-measuring-point load forecasting in distribution networks based on controlled channel communication, such as... Figure 1 As shown, the method includes: Step 101: Obtain historical load data for each measuring point in the distribution network, and construct load time series data samples corresponding to each measuring point based on the historical load data.

[0014] Step 102: Based on the load time series data samples of each measuring point, construct multiple training samples using the sliding window method; each training sample includes load time subsequence samples of each measuring point in the same historical time window, and the actual load label of each measuring point in the prediction time window.

[0015] Step 103: Based on the multiple training samples, train a pre-constructed multi-measurement point load prediction model; the multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to perform temporal coding on the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point, and based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; a prototype relay-style global communication module, used to train a pre-constructed multi-measurement point load prediction model based on a preset number of... The system includes a learnable global operating state prototype and a channel-independent representation, which perform information aggregation and backhaul between all measurement points and all global operating state prototypes, and construct a global communication message based on the aggregation and backhaul results; a controlled residual fusion module, which performs residual fusion of the channel-independent representation, the local communication message, and the global communication message to obtain an updated representation; and a prediction output module, which outputs the predicted load result value of each measurement point in the prediction time window based on the updated representation, and iteratively trains the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label.

[0016] Step 104: Input the load time subsequence of each measuring point within the target historical time window into the trained multi-measuring-point load prediction model to obtain the predicted load result value of each measuring point within the target prediction time window.

[0017] This application provides a method for multi-point load forecasting of distribution networks based on controlled channel communication. First, historical load data of each measuring point in the distribution network (such as smart meters, transformer monitoring terminals, etc.) are collected, and the historical load data of each measuring point are organized into independent load time series data samples in chronological order.

[0018] Next, a sliding window method is used to segment a large number of training samples from the aforementioned load time series data samples. Specifically, a fixed-length historical time window and a subsequent prediction time window can be set, and the window is gradually slid along the time axis. Each slide extracts a segment of the historical sequence as a load time subsequence sample, and the corresponding future sequence as the true load label, thus generating multiple training samples. It is important to note that each training sample includes load time subsequence samples from all measuring points within the same historical time window, as well as the true load label within the prediction time window.

[0019] A multi-point load prediction model containing five core modules can be pre-built. For example... Figure 2The diagram illustrates the structure of a multi-measurement point load prediction model provided in this embodiment. The first module is a channel-independent time-series encoding module, which is responsible for independently extracting features from the load time-series samples of each measurement point, generating their respective channel-independent sub-representations, and using these channel-independent sub-representations as the final channel-independent representation. Here, "channel" refers to the data dimension corresponding to each measurement point, and "independent" means that the information of each measurement point is not mixed during the encoding process. The advantage of this design is that it preserves the unique load change pattern of each measurement point and avoids mutual interference between different measurement points.

[0020] The second module is the Local Direct Communication module. It first divides all measuring points into multiple groups according to preset rules, and then allows measuring points within each group to interact through a multi-head attention mechanism. This allows the model to observe the correlation between measuring points simultaneously from multiple different perspectives. The output of this module is a local communication message, which enables measuring points within the same group to share information on local coordinated changes. For example, users in the same transformer area or on the same feeder often exhibit similar load fluctuations.

[0021] The third module is the prototype relay-style global communication module, which introduces multiple learnable global operational state prototypes. These prototypes can be understood as typical load patterns (such as morning peak mode, evening peak mode, holiday mode, etc.) automatically extracted by the model during training. This module first aggregates information from all measurement points onto these prototypes, allowing the prototypes to exchange information to form a unified understanding of the global situation. Then, the fused global information is sent back to each measurement point. In this way, measurement points can indirectly obtain globally shared information without direct pairwise communication, thus transmitting the global situation while significantly reducing computational overhead.

[0022] The fourth module is the controlled residual fusion module, which weights and fuses the outputs of the first three modules according to a preset intensity ratio. Channel-independent representations maintain their dominant position, while local and global communication messages are added as correction terms with smaller weights. This effectively prevents excessive interference from external information while introducing collaborative information between measurement points, ensuring the stability of the prediction.

[0023] The fifth module is the prediction output module, which maps the updated representation to the predicted load results for each measurement point within the future prediction time window, and updates the model parameters by calculating the training loss by comparing the predicted load results with the actual load labels. The entire training process iterates repeatedly until the model converges, thereby obtaining a multi-measurement point load prediction model with good predictive capabilities.

[0024] After obtaining the trained multi-measurement point load prediction model, the load time subsequence of each measurement point within the target historical time window before the current time can be input into the trained model. After processing by the aforementioned modules, the model synchronously outputs the load result prediction values ​​of all measurement points in the future target prediction time window, realizing collaborative prediction of multiple measurement points.

[0025] By applying the technical solution of this embodiment, each measuring point is independently encoded through a channel-independent timing coding module. This preserves the load variation patterns of each measuring point while avoiding contamination from abnormal measuring point information, thus solving the problem of mutual interference between weakly correlated channels in traditional global hybrid coding. Furthermore, a local direct communication module restricts direct interaction between measuring points to within a group, allowing only measuring points within the same group to share local collaborative information. This captures the mutual influence of neighboring measuring points and effectively suppresses noise propagation. Simultaneously, a prototype relay-type global communication module aggregates information from all measuring points onto a preset number of learnable global operating state prototypes for relay interaction and backhaul. This prevents the computational overhead of global communication from increasing dramatically with the number of measuring points, overcoming the resource bottleneck of large-scale global pairwise interaction. The above process is based solely on historical load data and does not rely on the distribution network topology or adjacency matrix, enabling direct deployment even in scenarios lacking complete topology information. Ultimately, the multi-measuring-point load forecasting model can simultaneously output the future load values ​​of all measuring points within the forecast time window, providing direct data support for overload risk screening and auxiliary dispatching decisions of the distribution network, while taking into account forecast accuracy, stability, and feasibility of engineering deployment.

[0026] In this embodiment, optionally, the channel-independent temporal coding module outputs a channel-independent representation based on the following steps: dividing the load time subsequence samples of each measurement point in the training samples according to a preset time length to obtain multiple sequence segments, and performing linear embedding mapping on each sequence segment to obtain a segment representation sequence corresponding to each measurement point; performing the following operation on the segment representation sequence of each measurement point to obtain the channel-independent subrepresentation corresponding to the measurement point: performing temporal feature mapping on the segment representation sequence of the measurement point through a shared temporal feature extraction function to obtain a corresponding first intermediate representation; and mapping the segment representation sequence of the measurement point with the first intermediate representation. The intermediate representations are residually connected, and a random deactivation operation is performed on the residual connection result to obtain the corresponding second intermediate representation; a layer normalization operation is performed on the second intermediate representation corresponding to the measurement point to obtain the corresponding third intermediate representation; the third intermediate representation corresponding to the measurement point is input into the feedforward network for nonlinear transformation to obtain the corresponding fourth intermediate representation; the third intermediate representation corresponding to the measurement point and the fourth intermediate representation are residually connected, and a layer normalization operation is performed on the residual connection result to obtain the corresponding channel-independent sub-representation; the channel-independent sub-representation includes the representation value corresponding to each sequence segment; the channel-independent sub-representations corresponding to all measurement points are used as the channel-independent representation.

[0027] In this embodiment, firstly, the load time subsequence samples for each measurement point in each training sample can be fragmented. For each measurement point's load time subsequence sample, it can be divided into multiple consecutive sequence segments according to a preset time length. For example, the load time subsequence samples from the past 720 time points can be divided into 45 sequence segments, each containing data from 16 time points. Then, linear embedding mapping is performed on the values ​​within each sequence segment to transform them into fixed-length feature vectors. The purpose of this is to convert the load time subsequence samples into a vector form that the model can process, while preserving the temporal information within each load time subsequence sample.

[0028] Next, the independent encoding process for each measurement point begins. Specifically, the segment representation sequence of each measurement point is processed sequentially. First, a shared temporal feature extraction function is used to perform a holistic temporal feature mapping on the segment representation sequences of each measurement point, obtaining the first intermediate representation corresponding to each measurement point. Here, all measurement points use the same set of temporal feature extraction functions for feature extraction, which not only significantly reduces the number of model parameters but also enables the model to learn general temporal patterns rather than overfitting to noise specific to each measurement point.

[0029] Subsequently, the fragment representation sequence is residually connected to the first intermediate representation, i.e., the two are directly added together, and a random deactivation operation is performed on the result to obtain the second intermediate representation. The purpose of the residual connection is to allow information to be smoothly transmitted in the deep network, avoiding training difficulties caused by gradient vanishing; while random deactivation randomly shuts down some neurons during training to prevent the model from over-relying on certain specific features and causing overfitting.

[0030] After obtaining the second intermediate representation, layer normalization can be performed to obtain the third intermediate representation. Then, the obtained third intermediate representation is fed into a feedforward network for nonlinear transformation. The feedforward network can consist of multiple fully connected layers and nonlinear activation functions, enabling higher-dimensional abstraction and reorganization of features, capturing complex nonlinear relationships in the load data, and obtaining the fourth intermediate representation.

[0031] Subsequently, residual connections and layer normalization are performed again. Specifically, the third intermediate representation is added to the fourth intermediate representation, and then layer normalization is performed to obtain the final channel-independent sub-representation of the measurement point. This step echoes the previous residual structure, ensuring that the original information is still preserved after the nonlinear transformation of the feedforward network, while layer normalization further stabilizes the feature distribution. Each channel-independent sub-representation contains the representation values ​​corresponding to all sequence segments of the measurement point; that is, each sequence segment corresponds to an independent representation value in this channel-independent sub-representation.

[0032] Finally, the channel-independent sub-representations corresponding to all measurement points are combined according to the measurement point dimension to form a holistic channel-independent representation. This representation fully preserves the feature information of each measurement point in each sequence segment, and the representation values ​​of each measurement point are independent of each other and do not mix, providing a stable and reliable input foundation for subsequent local and global communication.

[0033] Compared to directly encoding all measurement points together, this embodiment of the application uses independent channel encoding. While maintaining the load time-series characteristics of each measurement point, it improves training efficiency through parameter sharing, ensures the stability of deep networks through residual connections and layer normalization, and enhances the generalization ability of the model through random deactivation. Ultimately, it achieves high-quality independent encoding of multi-measurement point load time subsequence samples, providing clean and reliable feature representations for subsequent modules, and fundamentally ensuring the stability and accuracy of the entire prediction framework.

[0034] In this embodiment of the application, optionally, the local direct communication module outputs local communication messages based on the following steps: dividing all measurement points into multiple groups according to a preset grouping rule; determining the time segment window corresponding to each sequence segment; for each time segment window, obtaining the representation values ​​of each measurement point in each group corresponding to the time segment window from each channel independent sub-representation included in the channel independent representation; for each group, combining the representation values ​​of each measurement point in the group according to the measurement point order to obtain a two-dimensional data matrix, and using the two-dimensional data matrix as the multi-head attention input of the group in the time segment window; for each attention head, through the... The three learnable projection matrices corresponding to the attention head are used to map the two-dimensional data matrix into a query matrix, a key matrix, and a value matrix, respectively. The product of the query matrix and the transpose of the key matrix is ​​calculated to obtain a weight matrix representing the pairwise correlation between test points within a group. The weight matrix is ​​multiplied by the value matrix to obtain the attention result corresponding to the attention head. The attention results corresponding to each attention head are concatenated, and a linear transformation is performed on the concatenated result to obtain the intra-group local communication result of each test point within the group in the time segment window. The intra-group local communication results of each time segment window and each group are combined according to the order of the test points to obtain the local communication message.

[0035] In this embodiment, all measuring points are first grouped. Specifically, all measuring points can be divided into multiple groups according to preset grouping rules, with each group containing multiple measuring points. Measuring points within the same group can exchange information with each other, while direct communication between different groups is not allowed. The purpose of grouping is to limit the scope of direct communication between measuring points. This not only captures the cooperative change relationships between local measuring points but also effectively prevents weakly correlated or abnormal measuring points from propagating interference information globally, while also significantly reducing the computational load of direct communication.

[0036] After the load time subsequence samples of each measuring point are divided into multiple sequence segments, each sequence segment actually corresponds to a time segment window. Because the load time subsequence samples of different measuring points are divided according to the same preset time length, different measuring points correspond to a sequence segment within the same time segment window, ensuring the physical meaning of subsequent interactions. Then, for each time segment window, the representation value of each channel-independent sub-representation within that time segment window is extracted from the channel-independent sub-representations, i.e., the representation value of the sequence segment corresponding to that time segment window, preparing data for subsequent intra-group interactions.

[0037] Furthermore, for each group, the representation values ​​of each measurement point within that group are arranged and combined into a two-dimensional data matrix according to the original measurement point order. The rows of this matrix correspond to different measurement points, and the columns correspond to the representation values. This two-dimensional data matrix is ​​then used as input to the multi-head attention mechanism. The purpose of organizing the data into a two-dimensional data matrix is ​​to meet the computational requirements of the attention mechanism. Each row is treated as an independent object, and subsequent calculations of the correlation between rows enable information exchange between measurement points.

[0038] The multi-head attention computation then proceeds, allowing multiple attention heads to be activated simultaneously. Each attention head possesses three independent learnable projection matrices, mapping the input two-dimensional data matrix into a query matrix, a key matrix, and a value matrix, respectively. For each attention head, the product of the query matrix and the transpose of the key matrix is ​​calculated, resulting in a square matrix where each element represents the degree of association or similarity between two corresponding measure points. This association matrix is ​​then scaled and normalized to obtain a weight matrix, which is multiplied by the value matrix, transforming the output of each measure point into a weighted mixture of the other measure points in the same group. Finally, the attention results from each attention head are concatenated and fused through a linear transformation to form the final intra-group local communication result. The design of multiple attention heads allows the model to capture interaction patterns between measure points from different perspectives; for example, some heads may focus on peak synchronization, while others may focus on delayed following.

[0039] After completing the local communication calculations for all groups across all time segments, the data can be reassembled. Specifically, the local communication results within each time segment and each group can be reassembled according to the original measurement point order to restore the data arrangement consistent with the input structure, thus obtaining the final local communication messages, which can be directly used by subsequent modules.

[0040] In a specific embodiment, such as Figure 3 The diagram illustrates the internal workflow of a local direct communication module. Specifically, the module receives the channel-independent sub-representations of all C measurement point channels (i.e., C measurement points) as input. Then, it sequentially divides all measurement points into multiple groups according to a preset group size g. Figure 3The example demonstrates that group G1 contains measurement points 1 to g, group G2 contains measurement points g+1 to 2g, and so on until group M, GM, contains the remaining measurement points. Grouping is a prerequisite for subsequent local communication. Grouping can divide the potentially hundreds or thousands of measurement points into several manageable groups, with the number of measurement points within each group limited to g. Within each group, direct communication is performed, meaning that measurement points within the same group can interact through a multi-head attention mechanism to share local collaborative information. When updating its own representation, each measurement point can also refer to the characteristic information of other measurement points within the group, thereby capturing the collaborative change relationships between measurement points in the same group. For example, users under the same feeder or the same transformer area often exhibit similar load fluctuation patterns. At the same time, different groups do not communicate directly with each other; each group is independent and does not interfere with the others. This design structurally prevents information from weakly correlated or abnormal measurement points from being propagated to irrelevant measurement points, effectively suppressing the spread of interference. Meanwhile, since each group only needs to handle interactions between g measurement points within the group, the computational load is significantly reduced from global pairwise interactions, greatly alleviating the computational burden. After the direct intra-group communication is completed across all groups and all time-segment windows, the intra-group local communication results of each group are restored according to the original measurement point channel order. The restored data has the same measurement point arrangement order as the input, facilitating direct use by subsequent modules. The final output local communication message M loc Its dimensions are completely consistent with the independent representation of channels.

[0041] Compared to directly performing global pairwise interactions on all measurement points, the local direct communication scheme in this application restricts the communication range to a local area through a grouping mechanism, which significantly reduces computational complexity and effectively blocks the propagation of interference from weakly correlated and abnormal measurement points. Meanwhile, the multi-head attention mechanism captures the complex collaborative change patterns between measurement points within a group in parallel from multiple subspaces, ensuring both the sufficiency of information interaction and the controllability of the communication range, achieving a good balance between prediction accuracy and computational efficiency.

[0042] Optionally, in this embodiment, the prototype relay-type global communication module outputs a global communication message based on the following steps: pooling each independent sub-representation of the channel-independent representation to obtain a measurement point channel summary corresponding to each measurement point; acquiring a preset number of learnable global operating state prototypes; using each global operating state prototype as a common query and each measurement point channel summary as a common key and value, performing a cross-attention operation from the measurement point to the global operating state prototype to obtain a first updated prototype corresponding to each global operating state prototype; using each first updated prototype itself as a common query, key, and value, performing a self-attention operation between global operating state prototypes to obtain a second updated prototype corresponding to each first updated prototype; using each measurement point channel summary as a common query and each second updated prototype as a common key and value, performing a cross-attention operation from the global operating state prototype to the measurement point to obtain a global summary message corresponding to each measurement point; and broadcasting the global summary message corresponding to each measurement point to each corresponding sequence segment position to obtain the global communication message.

[0043] In this embodiment, the channel-independent representations are first compressed. Since each measurement point's channel-independent sub-representation contains representation values ​​from multiple time-segment windows, directly using them for global interaction would impose a significant computational burden. Therefore, pooling operations can be performed on these features. Specifically, the representation values ​​of all sequence segments for each measurement point can be averaged or summed along the time-segment window dimension, condensing multiple representation values ​​into a single vector. This condensed vector is the measurement point channel summary. Through this step, each measurement point transforms from a "detailed feature sequence across multiple time points" into a "concise summary reflecting the overall state of the measurement point," preserving the core state information of the measurement point while significantly reducing the computational load of subsequent global communication.

[0044] Next, the information aggregation stage begins. This stage acquires a pre-defined number of learnable global operating state prototypes. These prototypes can be understood as typical load patterns (such as morning peak, evening peak, and holiday patterns) automatically extracted by the model during training. They are part of the model parameters and are continuously optimized during training. In this step, the global operating state prototype acts as the query party, and the channel summaries of all measuring points serve as the key-value database being queried, performing a cross-attention operation. The core idea of ​​cross-attention is that each global operating state prototype "queries" the channel summaries of all measuring points, extracting the most relevant information based on its own needs to update its own state. After this step, each global operating state prototype absorbs the overall situational information of all measuring points, forming the first updated prototype. At this point, each prototype is no longer an isolated pattern template but a cognitive node carrying the current global operating state information of the power grid.

[0045] Subsequently, these first-update prototypes interact with each other. Specifically, each first-update prototype simultaneously acts as a query, key, and value, performing a self-attention operation. Self-attention allows each first-update prototype to focus on information from all other first-update prototypes, enabling consensus and a more unified global understanding among different first-update prototypes. After this step, connections are established between the previously independently updated first-update prototypes, and the typical pattern information they represent is mutually calibrated and supplemented, forming second-update prototypes. At this point, the set of second-update prototypes has become a highly condensed and coordinated global situational representation.

[0046] After integrating information between prototypes, the global situational awareness is fed back to each measurement point. In this step, the measurement point channel summary of each measurement point acts as the querying party, and the second updated prototype acts as the key-value library being queried, performing cross-attention operation again. Reversing the previous "measurement point → prototype" direction, this time it's "prototype → measurement point": each measurement point, based on its current measurement point channel summary, selectively extracts the most relevant global information from each of the second updated prototypes representing the global situation, thus obtaining a unique global summary message for itself. In this way, each measurement point can obtain global information as needed without introducing noise from irrelevant measurement points, achieving precise global information distribution.

[0047] Finally, the global summary message corresponding to each measurement point is broadcast to each corresponding sequence segment position. Since we obtain a condensed global summary message for each measurement point, while the subsequent fusion module requires a data structure with the same dimension as the channel-independent representation, we can copy the global summary message for each measurement point multiple times and fill it into each sequence segment position of that measurement point, so that each time segment window carries the same global summary information. At this point, we obtain a complete global communication message with a dimension completely consistent with the channel-independent representation, which can be directly used for subsequent residual fusion.

[0048] In a specific embodiment, such as Figure 4The diagram illustrates the internal workflow of a prototype relay-based global communication module. Specifically, the module first compresses the independent channel representations to obtain the channel summaries corresponding to each measurement point. Then, the channel summaries corresponding to each measurement point are used as input. The model can include K learnable global operating state prototypes, labeled P1 to Pk. These prototypes are built-in learnable parameters of the model. As the training process continues to optimize, they can be understood as K typical load operation modes automatically extracted by the model, such as the morning peak climbing mode, the midday smooth mode, and the evening peak mode. These prototypes constitute relay stations for global information exchange. All measurement points share the same set of prototypes, achieving indirect global communication through prototypes. During global communication, the aggregation from measurement points to prototypes is first performed. In this step, the channel summaries of each measurement point serve as keys and values, and the K prototypes together act as queries, performing a cross-attention operation. This allows each prototype to selectively extract the most relevant state information from the channel summaries of all measurement points based on its representative typical mode, thus forming the first updated prototype corresponding to each prototype. After aggregation, each prototype carries global information related to its own mode within the current overall power grid situation. Subsequently, K prototypes interact with each other. Each first-update prototype uses itself simultaneously as the query, key, and value to perform self-attention operations, enabling information exchange between different prototypes. Through this step, connections are established between the prototypes, and their information is mutually calibrated and supplemented, forming second-update prototypes. After completing the interaction between prototypes, a prototype-to-measurement point feedback operation is performed. In this step, the measurement point channel summary is used as the query, and the second-update prototype as the key and value, again performing cross-attention operations. This allows each measurement point to selectively extract the most relevant global context information from the K second-update prototypes representing global consensus, based on its current state, thus forming a global summary message corresponding to each measurement point. After feedback, each measurement point obtains its own unique global summary message. Finally, the global summary messages corresponding to each measurement point are broadcast to all corresponding sequence segment positions, ultimately forming a dimensionally complete global communication message M. glob .

[0049] Compared to direct pairwise interaction among all measurement points, the prototype relay scheme in this application achieves indirect communication through a preset number of learnable global operational state prototypes: all measurement points first aggregate information onto the prototypes, and after the prototypes interact to form a global consensus, the information is then transmitted back to each measurement point. In this way, regardless of the increase in the number of measurement points, the computational load and information capacity of global communication are controlled within a fixed range determined by the number of prototypes, completely solving the computational explosion problem of global pairwise interaction. Simultaneously, each measurement point extracts global information on demand through a cross-attention mechanism, avoiding noise pollution from irrelevant measurement points and achieving efficient and stable global information sharing.

[0050] Optionally, in this embodiment, the controlled residual fusion module outputs an updated representation based on the following steps: receiving the channel-independent representation, the local communication message, and the global communication message; weighting the local communication message according to a preset local injection strength to obtain a weighted local communication message; weighting the global communication message according to a preset global injection strength to obtain a weighted global communication message; adding the channel-independent representation, the weighted local communication message, and the weighted global communication message to obtain a fused representation; performing a layer normalization operation on the fused representation to obtain a normalized fused representation; inputting the normalized fused representation into a feedforward network for nonlinear transformation to obtain a feedforward output; performing a residual connection between the normalized fused representation and the feedforward output, and performing a layer normalization operation on the residual connection result to obtain the updated representation.

[0051] In this embodiment, the outputs of the preceding three modules are first received: channel-independent representations, local communication messages, and global communication messages. These three data sets have identical dimensional structures, thus providing the basic conditions for fusion. The channel-independent representations carry the original timing characteristics of each measuring point, the local communication messages carry the collaborative information between measuring points in the same group, and the global communication messages condense the global situational information of the entire distribution network. The three complement each other and each has its own focus, providing a complete information foundation for subsequent fusion.

[0052] Next, the two communication messages are weighted separately. Specifically, the local communication message is scaled according to a preset local injection strength, and the global communication message is scaled according to a preset global injection strength. Here, the injection strength is a positive number less than 1; the smaller the injection strength, the weaker the impact of external messages on the final result. This design ensures that the local and global communication messages are merely corrections to the original timing representation, rather than replacements, thus effectively preventing excessive interference from external information.

[0053] After weighting, the channel-independent representation, weighted local communication messages, and weighted global communication messages can be added element-by-element to obtain a fused representation. The channel-independent representation retains the load variation patterns of each measuring point, the weighted local communication messages supplement the coordinated variation information of the same group of measuring points, and the weighted global communication messages inject the macroscopic situational information of the entire distribution network. The superposition of these three representations maintains the individual characteristics of each measuring point while incorporating common local and global information.

[0054] Subsequently, layer normalization is performed on the fused representation. Layer normalization calculates the mean and variance of the fused representation along the feature dimensions and uses these statistics to readjust the data to a standard distribution with a mean of 0 and a variance of 1. Since the fused representation is derived from the sum of three different sources, its numerical distribution may fluctuate significantly. Layer normalization stabilizes the data distribution, accelerates the convergence of subsequent networks, and ensures smooth information transfer between different modules.

[0055] After normalization, the normalized fused representation is input into a feedforward network for nonlinear transformation. The feedforward network can consist of multiple fully connected layers and nonlinear activation functions, enabling higher-dimensional abstraction and reorganization of features, capturing complex nonlinear relationships in the load data. This step further enhances the model's expressive power, allowing the fused features to be better mapped to the final prediction target.

[0056] Finally, the normalized fused representation is residually connected to the feedforward output, and layer normalization is performed on the connection result to obtain the final updated representation. The residual connection directly adds the input and output of the feedforward network, ensuring that the original information is retained after nonlinear transformation, effectively solving the gradient vanishing and degradation problems in deep networks. Layer normalization further stabilizes the feature distribution of the output, providing high-quality input for subsequent prediction output modules.

[0057] Compared to simply stacking the three elements indiscriminately, the controlled residual fusion scheme in this embodiment precisely controls the degree of correction of the original representation by external information, enabling the model to flexibly adjust the contribution of local and global information according to the actual scenario. At the same time, the structural design of layer normalization and residual connection ensures the stability of the training process and the integrity of feature transfer, so that the model has both strong nonlinear expression capabilities and maintains the stability of single-point temporal modeling, achieving a balance between the richness and controllability of information fusion.

[0058] Optionally, in this embodiment, step 103, "training a pre-constructed multi-point load prediction model based on the multiple training samples," includes: inputting each training sample into the current multi-point load prediction model, outputting the predicted load result value of each measurement point within the prediction time window corresponding to each training sample after forward propagation; calculating the time domain error between the predicted load result value and the corresponding real load label; transforming the predicted load result value and the corresponding real load label to the frequency domain respectively, and calculating the difference in the transformation result in the frequency domain to obtain the frequency domain error; weighted summing the time domain error and the frequency domain error of each training sample to obtain the total training loss; updating the parameters of the current multi-point load prediction model using an adaptive optimizer with the goal of minimizing the total training loss, re-outputting the predicted load result value of each measurement point within the prediction time window corresponding to each training sample based on the updated multi-point load prediction model, and returning to the step of calculating the time domain error between the predicted load result value and the corresponding real load label, until a preset training termination condition is met, thus obtaining a trained multi-point load prediction model.

[0059] In this embodiment, firstly, each training sample is input into the multi-measurement point load prediction model under the current state. The data is processed sequentially through modules such as channel-independent time-series coding, local direct communication, prototype relay-type global communication, and controlled residual fusion. Finally, the prediction output module generates the predicted load result for each measurement point within the prediction time window. This step completes the full information flow from input to output, providing a foundation for subsequent error calculation.

[0060] After obtaining the load forecast, the time domain error can be calculated. The time domain is the original time axis. The time domain error can be determined by directly comparing the difference between the load forecast and the actual load indicated by the actual load label at each time step. Specifically, the mean absolute error can be used to measure it. The closer the load forecast is to the actual value, the smaller the time domain error.

[0061] Besides time-domain errors, the predicted and actual load results can also be transformed into the frequency domain for calculation. Frequency domain analysis examines the components of a signal at different frequencies, using a real-number Fast Fourier Transform to decompose the time-series signal into combinations of different frequency components. Load data typically exhibits clear periodic patterns (such as daily or weekly cycles). Frequency-domain errors can constrain the model's accuracy in terms of periodic structures. Even if the predicted load results deviate slightly from the actual values ​​on the time axis, frequency-domain errors can capture these waveform differences, thus enabling the model to more accurately learn the periodic fluctuations of the load.

[0062] Subsequently, the two types of errors from each training sample are weighted and summed according to preset weights to obtain the total training loss. The time domain error ensures the accuracy of the predicted numerical values, while the frequency domain error ensures the reasonableness of the predicted waveform shape. The combination of the two allows the model to focus on both "whether the specific values ​​are correct" and "whether the trend of change is similar," avoiding the prediction bias that may be caused by a single error target.

[0063] Finally, with the goal of minimizing the total training loss, an adaptive optimizer is used to update all learnable parameters of the model. The adaptive optimizer can automatically adjust the learning rate of each parameter based on the historical gradient of the parameters, making the training process more stable and efficient. After the update is completed, the predicted load results are recalculated based on the updated model and returned to the error calculation step. This process is repeated iteratively until the preset training termination condition is met (such as the validation set error no longer decreasing), and finally the trained multi-point load prediction model is obtained.

[0064] Compared to traditional methods that use only time-domain errors for training, the training scheme of this application introduces frequency-domain errors to construct a dual constraint mechanism: time-domain errors ensure the accuracy of predicted values, and frequency-domain errors ensure the accuracy of load change cycles. The two complement each other, making the model more comprehensive in learning complex load fluctuation patterns. At the same time, the use of an adaptive optimizer and the iterative training strategy ensure that the model can efficiently converge to the optimal state, ultimately obtaining a high-quality prediction model that combines numerical accuracy and trend fitting ability.

[0065] In this embodiment of the application, optionally, the target prediction time window includes multiple time steps; after step 104, the method further includes: for the load result prediction value corresponding to each measuring point, performing the following operations: calculating the risk value of each time step based on the predicted load and corresponding operating limit of the measuring point in each time step within the target prediction time window, and taking the maximum value among the risk values ​​of each time step as the comprehensive risk value of the measuring point; the risk value is used to characterize the degree of occupancy of the predicted load relative to the operating limit; counting the number of time steps in which the predicted load of each time step of the measuring point exceeds the operating limit, as the overload duration step number of the measuring point; determining the risk level of the measuring point according to the numerical range of the comprehensive risk value and a preset risk level classification rule, and generating corresponding operating suggestions based on the risk level and the overload duration step number.

[0066] In this embodiment, since the target prediction time window contains multiple time steps, and each measuring point has a corresponding predicted load at each time step, these predicted loads can be analyzed one by one to assess the potential overload risk in the future.

[0067] Specifically, for each measuring point, the risk value for each time step within the target prediction time window is first calculated. The predicted load for each time step can be divided by the corresponding operating limit for that measuring point; the resulting ratio is the risk value for that time step. This ratio directly reflects the degree of occupancy of the predicted load relative to the equipment's capacity. The closer the ratio is to 1, the closer the predicted load is to the upper limit; a ratio exceeding 1 indicates that the predicted load has exceeded the equipment's rated capacity, posing an overload risk. Subsequently, the maximum value among these risk values ​​is selected as the comprehensive risk value for that measuring point, characterizing the most severe overload that the measuring point may face throughout the entire target prediction time window.

[0068] In addition to focusing on the peak value of the risk, we can also count the number of time steps within which the predicted load at a given measuring point exceeds the operating limit within the target prediction time window, as the overload duration steps. This indicator measures the persistence of the overload state. If a measuring point only briefly exceeds the limit, its urgency is relatively low; however, if it continuously exceeds the limit for a longer period, even if the peak value is not too high, it means that the equipment needs to withstand overload operation for an extended period, and the actual risk may be more severe. Therefore, the overload duration steps provide supplementary information in the time dimension for risk assessment.

[0069] Then, based on the numerical range of the comprehensive risk value, the risk level of the measuring point is determined according to the preset risk level classification rules. For example, when the comprehensive risk value is below a certain threshold, it is considered normal; when it is between two thresholds, it enters a state of attention or warning; and when it exceeds a higher threshold, it is judged to be a critical state. Each level reflects a different degree of overload urgency.

[0070] After determining the risk level, corresponding operational recommendations are generated by comprehensively considering the risk level and the number of overload duration steps. These recommendations are executable instructions for operations and maintenance personnel, and their content depends on the severity of the risk level. For normal conditions, only routine monitoring is required; for conditions of concern and warning, it is recommended to strengthen inspections, shorten inspection cycles, or consider load shifting; for critical conditions, immediate load limiting and on-site handling are necessary. The number of overload duration steps serves as an auxiliary correction factor, enabling the system to differentiate recommendations more finely within the same risk level. For example, under the same risk level, measurement points with longer durations can receive stronger warning wording or higher priority in handling.

[0071] In this embodiment, the post-processing flow transforms abstract numerical prediction results into intuitive and actionable risk assessments and operational recommendations, bridging the gap between "data prediction" and "engineering intervention." This flow not only provides a judgment on the severity of peak loads but also introduces a continuous consideration through the number of overload duration steps, making the risk assessment more comprehensive and multi-dimensional. The final operational recommendations directly serve the daily operation and maintenance and scheduling decisions of the distribution network, realizing a complete closed loop from load forecasting to risk warning and operational guidance, greatly enhancing the application value of load forecast values ​​in practical engineering.

[0072] In one specific embodiment, before training the multi-measuring-point load prediction model, this application can preprocess the acquired historical load data from multiple measuring points. Specifically, after parsing and aligning the historical load data based on the time index, load time series data samples corresponding to each measuring point are formed. Then, invalid or low-quality measuring points are removed. Specifically, for each measuring point, the proportion of non-zero records of all its sampling points can be calculated based on its load time series data samples. When this proportion is lower than a preset threshold or the historical maximum load value is zero, it indicates that the measuring point has severe data loss or has been in a long-term shutdown state, and the corresponding measuring point is removed, thereby selecting a set of valid measuring points with qualified data quality for subsequent modeling.

[0073] To enhance the applicability of this application in engineering environments, outlier detection and correction can be performed on the load time series data samples of the screened measurement points. Specifically, the three-standard-deviation criterion can be used to identify outliers. That is, when the load value of a sampling point deviates from the mean of the measurement point by more than three standard deviations, it is judged as an outlier and replaced with the mean of the adjacent time window to suppress the extreme noise impact caused by communication anomalies or short-term jumps. For a small number of missing values ​​caused by missing data acquisition or format conversion, linear interpolation is used to fill them in. If the missing value is located at the end of the sequence, the most recent valid value is used to fill it in. Finally, to eliminate the difference in load magnitude between different measurement points, each measurement point is standardized based on the mean and standard deviation of the training set to make the load data of all measurement points of the same order of magnitude. The statistical parameters of the training set are reused in the verification and testing phases to avoid future information leakage.

[0074] Furthermore, the method of this application embodiment was verified on the Electricity Load Diagrams 2011-2014 public dataset. This dataset contains historical load data of multiple customers or electricity consumption points, and each customer or electricity consumption point can be regarded as a measurement point. In the experiment, 321 measurement points after validity screening were selected as prediction objects, and historical time windows were used as input to simultaneously predict the load values ​​of each measurement point in multiple future time steps.

[0075] In this embodiment, mean squared error (MSE) and mean absolute error (MAE) are used as evaluation indicators for prediction performance; the lower the values ​​of both, the smaller the prediction error. Table 1 shows the performance comparison results of the method of this embodiment and several representative prediction models on the Electricity dataset. The results in the table are expressed as MSE / MAE, with lower values ​​indicating better prediction performance. The relative error reduction rate is calculated as "(error of the comparison method - error of the present invention) / error of the comparison method × 100%".

[0076] Table 1: Comparison of load forecasting performance on the Electricity dataset

[0077] Table 1 shows that the first model is TimeMixer, the second model is SimpleTM, the third model is DLinear, the fourth model is iTransformer, and the fifth model is TQNet. As can be seen from Table 1, the method in this application maintains low prediction errors at prediction step sizes of 96, 192, 336, and 720. Compared with TimeMixer (the first model), the average MSE of this application's method is reduced by 16.8%, and the average MAE is reduced by 13.4%. Compared with SimpleTM (the second model), DLinear (the third model), iTransformer (the fourth model), and TQNet (the fifth model), this application's method also achieves varying degrees of error reduction. These results demonstrate that the method of this application is well-suited for medium- and high-dimensional multi-measurement point load prediction scenarios.

[0078] The effectiveness of this method primarily stems from its controlled channel communication structure. Specifically, the channel-independent timing coding module maintains stable modeling of the load sequence at a single measurement point; the local direct communication module captures the collaborative change relationships between measurement points within the same group; the prototype relay-style global communication module transmits cross-group shared operational status information; and the controlled residual fusion module limits the disturbance of communication messages to the original timing representation. Thus, the model can utilize multi-measurement point collaborative information while avoiding unconstrained pairwise mixing among all measurement points.

[0079] In the risk assessment demonstration, 10 prediction time windows were selected for verification. Each window contains 321 measurement points, and each measurement point predicts 96 future time steps. Figures 5-7 The results of an example of multi-point load prediction and overload risk assessment are shown. Among them, Figure 5 It displays the maximum risk value and the thresholds for attention, early warning, and criticality under different forecast time windows; Figure 6 The results of the ranking of high-risk measurement point events are displayed; Figure 7The paper presents a comparison of the future predicted load curves and capacity thresholds for the measurement points corresponding to the highest-risk events. Experimental results show that the system identified 3203 normal events and 7 events of concern, with no warning or critical events occurring. The highest-risk event occurred at the 75th measurement point within the 9th prediction time window, with a maximum risk index of 0.817. The event first reached the concern threshold at the 82nd future time step. These results demonstrate that the method described in this application can automatically perform measurement point-level risk calculation, high-risk event ranking, and risk curve display after obtaining future load prediction results.

[0080] The above results demonstrate that this application can not only output future load forecasts from multiple measurement points, but also further transform the forecast results into risk indicators, risk levels, and operational recommendations, forming a complete process of "multi-point load forecasting—overload risk assessment—operational recommendation generation—result recording." In actual engineering deployments, Figures 5-7 The simulated rated capacity shown can be replaced with the actual rated capacity of the equipment, the user contract capacity, or the equivalent power limit calculated from the feeder current limit and the switchgear current limit, so as to be used for actual distribution network operation monitoring and overload early warning.

[0081] Furthermore, as Figure 1 In terms of specific implementation, this application provides a multi-point load forecasting device for distribution networks based on controlled channel communication, such as... Figure 8 As shown, the device includes: The data acquisition module is used to acquire historical load data of each measuring point in the distribution network, and to construct load time series data samples corresponding to each measuring point based on the historical load data. The training sample construction module is used to construct multiple training samples based on the load time series data samples of each measuring point using the sliding window method. Each training sample includes load time subsequence samples of each measuring point in the same historical time window, as well as the actual load labels of each measuring point in the prediction time window. The model training module is used to train a pre-constructed multi-measurement point load prediction model based on the multiple training samples. The multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to perform temporal coding on the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point and, based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; and a prototype relay-style global communication module, used to train a pre-constructed multi-measurement point load prediction model based on a built-in preset number of data points. The system comprises a learnable global operating state prototype and a channel-independent representation, and performs information aggregation and backhaul between all measurement points and all global operating state prototypes. A global communication message is constructed based on the aggregation and backhaul results. A controlled residual fusion module is used to perform residual fusion of the channel-independent representation, the local communication message, and the global communication message to obtain an updated representation. A prediction output module is used to output the predicted load result value for each measurement point within the prediction time window based on the updated representation, and to iteratively train the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label. The prediction module is used to input the load time subsequence of each measuring point within the target historical time window into the trained multi-measuring-point load prediction model to obtain the predicted load result value of each measuring point within the target prediction time window.

[0082] Optionally, the channel-independent timing coding module is used for: The load time subsequence samples of each measurement point in the training sample are divided according to a preset time length to obtain multiple sequence segments. Then, a linear embedding mapping is performed on each sequence segment to obtain the segment representation sequence corresponding to each measurement point. For each measurement point, the following operation is performed on the segment representation sequence to obtain the channel-independent sub-representation corresponding to the measurement point: By using the shared temporal feature extraction function of each measurement point, the segment representation sequence of the measurement point is mapped with temporal features to obtain the corresponding first intermediate representation; The fragment representation sequence of the measurement point is residually joined with the first intermediate representation, and a random deactivation operation is performed on the residual joining result to obtain the corresponding second intermediate representation; Perform layer normalization on the second intermediate representation corresponding to the measurement point to obtain the corresponding third intermediate representation; The third intermediate representation corresponding to the measurement point is input into the feedforward network and subjected to nonlinear transformation to obtain the corresponding fourth intermediate representation; The third intermediate representation and the fourth intermediate representation corresponding to the measurement point are joined by residual connection, and the residual connection result is subjected to layer normalization to obtain the corresponding channel-independent sub-representation; the channel-independent sub-representation includes the representation value corresponding to each sequence segment; The channel-independent sub-representation corresponding to all measurement points is taken as the channel-independent representation.

[0083] Optionally, the local direct communication module is used for: All measurement points are divided into multiple groups according to the preset grouping rules; The time segment window corresponding to each sequence segment is determined. For each time segment window, the representation value of each measurement point in each group corresponding to the time segment window is obtained from each channel independent sub-representation included in the channel independent representation. For each group, the representation values ​​of each measurement point within the group are combined in the order of the measurement points to obtain a two-dimensional data matrix, and the two-dimensional data matrix is ​​used as the multi-head attention input of the group in the time segment window; For each attention head, the two-dimensional data matrix is ​​mapped into a query matrix, a key matrix, and a value matrix respectively through the three learnable projection matrices corresponding to the attention head. The product of the query matrix and the transpose of the key matrix is ​​calculated to obtain a weight matrix that represents the pairwise correlation between test points within the group. The weight matrix is ​​multiplied by the value matrix to obtain the attention result corresponding to the attention head. The attention results corresponding to each attention head are concatenated, and a linear transformation is performed on the concatenated result to obtain the intra-group local communication result of each test point within the group in the time segment window. The local communication results within each time segment window and each group are combined according to the measurement point order to obtain the local communication message.

[0084] Optionally, the prototype relay-type global communication module is used for: Each independent sub-representation of the channel in the independent channel representation is subjected to pooling processing to obtain the channel summary of each measurement point; Get a preset number of learnable global running state prototypes, use each global running state prototype as a query and each measurement point channel summary as a key and value, perform cross attention operation from measurement points to global running state prototypes, and obtain the first updated prototype corresponding to each global running state prototype. Using each first updated prototype as a query, key, and value, perform self-attention operations between global runtime prototypes to obtain the second updated prototype corresponding to each first updated prototype. Using the channel summaries of each measurement point as the query and the second update prototypes as the key and value, perform a cross-attention operation from the global running state prototype to the measurement point to obtain the global summary message corresponding to each measurement point. The global summary message corresponding to each measurement point is broadcast to each corresponding sequence segment position to obtain the global communication message.

[0085] Optionally, the controlled residual fusion module is used for: Receive the channel-independent representation, the local communication message, and the global communication message; The local communication messages are weighted according to a preset local injection strength to obtain weighted local communication messages; The global communication messages are weighted according to a preset global injection strength to obtain weighted global communication messages; The independent channel representation, the weighted local communication message, and the weighted global communication message are added together to obtain the fused representation; Perform a layer normalization operation on the fused representation to obtain a normalized fused representation; The normalized fused representation is input to the feedforward network and subjected to a nonlinear transformation to obtain the feedforward output; The normalized fused representation is residually concatenated with the feedforward output, and the residual concatenation result is subjected to layer normalization to obtain the updated representation.

[0086] Optionally, the model training module is used for: Each training sample is input into the current multi-point load prediction model, and after forward propagation, the predicted load result of each measurement point within the prediction time window corresponding to each training sample is output. Calculate the time domain error between the predicted load result and the corresponding actual load label; The predicted load results and the corresponding actual load labels are transformed to the frequency domain, and the difference between the transformation results in the frequency domain is calculated to obtain the frequency domain error. The total training loss is obtained by weighted summing of the time domain error and frequency domain error of each training sample. With the goal of minimizing the total training loss, an adaptive optimizer is used to update the parameters of the current multi-point load prediction model. Based on the updated multi-point load prediction model, the predicted load results for each measurement point within the prediction time window corresponding to each training sample are re-output. The process then returns to the step of calculating the time domain error between the predicted load results and the corresponding real load label, until the preset training termination condition is met, thus obtaining the trained multi-point load prediction model.

[0087] Optionally, the target prediction time window includes multiple time steps; the device further includes a running suggestion generation module; the running suggestion generation module is used to: After obtaining the predicted load results for each measuring point within the target prediction time window, perform the following operations for each predicted load result for each measuring point: Based on the predicted load and corresponding operating limits of the measuring point at each time step within the target prediction time window, the risk value of each time step is calculated, and the maximum value among the risk values ​​of each time step is taken as the comprehensive risk value of the measuring point; the risk value is used to characterize the degree of occupancy of the predicted load relative to the operating limits. The number of time steps in which the predicted load at each time step of the measuring point exceeds the operating limit is counted, and this number is taken as the overload duration step of the measuring point. Based on the numerical range of the comprehensive risk value, the risk level of the measuring point is determined according to the preset risk level classification rules, and corresponding operation suggestions are generated based on the risk level and the overload duration steps.

[0088] It should be noted that other corresponding descriptions of the functional units involved in the multi-point load forecasting device for distribution networks based on controlled channel communication provided in this application embodiment can be found in the following references. Figures 1 to 7 The corresponding descriptions in the method will not be repeated here.

[0089] This application also provides a computer device, which may specifically be a personal computer, a server, a network device, etc. Figure 9 As shown, the computer device includes a bus, a processor, memory, and a communication interface, and may also include an input / output interface and a display device. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores load data, model parameters, prediction results, and other related data. The network interface allows communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0090] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0091] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0092] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0093] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0094] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0095] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.

Claims

1. A method for multi-point load forecasting in distribution networks based on controlled channel communication, characterized in that, include: Historical load data of each measuring point in the distribution network is obtained, and load time series data samples corresponding to each measuring point are constructed based on the historical load data. Based on the load time series data samples of each measuring point, multiple training samples are constructed using the sliding window method. Each training sample includes load time subsequence samples of each measuring point in the same historical time window, as well as the actual load labels of each measuring point in the prediction time window. Based on the multiple training samples, a pre-constructed multi-point load prediction model is trained. The multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to temporally encode the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point and, based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; a prototype relay-type global communication module, used to perform information aggregation and backhaul between all measurement points and all global operating state prototypes based on a preset number of built-in learnable global operating state prototypes and the channel-independent representation, and construct global communication messages based on the aggregation and backhaul results; a controlled residual fusion module, used to perform residual fusion of the channel-independent representation, the local communication messages, and the global communication messages to obtain an updated representation; and a prediction output module, used to output the predicted load result value of each measurement point in the prediction time window based on the updated representation, so as to iteratively train the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label. By inputting the load time subsequence of each measuring point within the target historical time window into the trained multi-measuring-point load prediction model, the predicted load result value of each measuring point within the target prediction time window is obtained.

2. The method according to claim 1, characterized in that, The channel-independent timing coding module outputs a channel-independent representation based on the following steps: The load time subsequence samples of each measurement point in the training sample are divided according to a preset time length to obtain multiple sequence segments. Then, a linear embedding mapping is performed on each sequence segment to obtain the segment representation sequence corresponding to each measurement point. For each measurement point, the following operation is performed on the segment representation sequence to obtain the channel-independent sub-representation corresponding to the measurement point: By using the shared temporal feature extraction function of each measurement point, the segment representation sequence of the measurement point is mapped with temporal features to obtain the corresponding first intermediate representation; The fragment representation sequence of the measurement point is residually joined with the first intermediate representation, and a random deactivation operation is performed on the residual joining result to obtain the corresponding second intermediate representation; Perform layer normalization on the second intermediate representation corresponding to the measurement point to obtain the corresponding third intermediate representation; The third intermediate representation corresponding to the measurement point is input into the feedforward network and subjected to nonlinear transformation to obtain the corresponding fourth intermediate representation; The third intermediate representation and the fourth intermediate representation corresponding to the measurement point are residually joined, and the residual joining result is subjected to layer normalization to obtain the corresponding channel-independent sub-representation; the channel-independent sub-representation includes the representation value corresponding to each sequence segment; The channel-independent sub-representation corresponding to all measurement points is taken as the channel-independent representation.

3. The method according to claim 2, characterized in that, The local direct communication module outputs local communication messages based on the following steps: All measurement points are divided into multiple groups according to the preset grouping rules; The time segment window corresponding to each sequence segment is determined respectively. For each time segment window, the representation value of each measurement point in each group and the time segment window is obtained from each channel independent sub-representation included in the channel independent representation. For each group, the representation values ​​of each measurement point within the group are combined in the order of the measurement points to obtain a two-dimensional data matrix, and the two-dimensional data matrix is ​​used as the multi-head attention input of the group in the time segment window; For each attention head, the two-dimensional data matrix is ​​mapped into a query matrix, a key matrix, and a value matrix respectively through the three learnable projection matrices corresponding to the attention head. The product of the query matrix and the transpose of the key matrix is ​​calculated to obtain a weight matrix that represents the pairwise correlation between test points within the group. The weight matrix is ​​multiplied by the value matrix to obtain the attention result corresponding to the attention head. The attention results corresponding to each attention head are concatenated, and a linear transformation is performed on the concatenated result to obtain the intra-group local communication result of each test point within the group in the time segment window. The local communication results within each time segment window and each group are combined according to the measurement point order to obtain the local communication message.

4. The method according to claim 2, characterized in that, The prototype relay-type global communication module outputs global communication messages based on the following steps: Each independent sub-representation of the channel in the independent channel representation is subjected to pooling processing to obtain the channel summary of each measurement point; Get a preset number of learnable global running state prototypes, use each global running state prototype as a query and each measurement point channel summary as a key and value, perform cross-attention operation from measurement points to global running state prototypes, and obtain the first updated prototype corresponding to each global running state prototype. Using each first updated prototype as a query, key, and value, perform self-attention operations between global runtime prototypes to obtain the second updated prototype corresponding to each first updated prototype. Using the channel summaries of each measurement point as the query and the second update prototypes as the key and value, perform a cross-attention operation from the global running state prototype to the measurement point to obtain the global summary message corresponding to each measurement point. The global summary message corresponding to each measurement point is broadcast to each corresponding sequence segment position to obtain the global communication message.

5. The method according to claim 1, characterized in that, The controlled residual fusion module outputs an updated representation based on the following steps: Receive the channel-independent representation, the local communication message, and the global communication message; The local communication messages are weighted according to a preset local injection strength to obtain weighted local communication messages; The global communication messages are weighted according to a preset global injection strength to obtain weighted global communication messages; The independent channel representation, the weighted local communication message, and the weighted global communication message are added together to obtain the fused representation; Perform a layer normalization operation on the fused representation to obtain a normalized fused representation; The normalized fused representation is input to the feedforward network and subjected to a nonlinear transformation to obtain the feedforward output; The normalized fused representation is residually concatenated with the feedforward output, and the residual concatenation result is subjected to layer normalization to obtain the updated representation.

6. The method according to claim 1, characterized in that, The process of training a pre-constructed multi-point load prediction model based on the multiple training samples includes: Each training sample is input into the current multi-point load prediction model, and after forward propagation, the predicted load result of each measurement point within the prediction time window corresponding to each training sample is output. Calculate the time domain error between the predicted load result and the corresponding actual load label; The predicted load results and the corresponding actual load labels are transformed to the frequency domain, and the difference between the transformation results in the frequency domain is calculated to obtain the frequency domain error. The total training loss is obtained by weighted summing of the time domain error and frequency domain error of each training sample. With the goal of minimizing the total training loss, an adaptive optimizer is used to update the parameters of the current multi-point load prediction model. Based on the updated multi-point load prediction model, the predicted load results for each measurement point within the prediction time window corresponding to each training sample are re-output. The process then returns to the step of calculating the time domain error between the predicted load results and the corresponding real load label, until the preset training termination condition is met, thus obtaining the trained multi-point load prediction model.

7. The method according to claim 1, characterized in that, The target prediction time window includes multiple time steps; after obtaining the predicted load results for each measuring point within the target prediction time window, the method further includes: For the predicted load result value corresponding to each measuring point, perform the following operations: Based on the predicted load and corresponding operating limits of the measuring point at each time step within the target prediction time window, the risk value of each time step is calculated, and the maximum value among the risk values ​​of each time step is taken as the comprehensive risk value of the measuring point; the risk value is used to characterize the degree of occupancy of the predicted load relative to the operating limits. The number of time steps in which the predicted load at each time step of the measuring point exceeds the operating limit is counted, and this number is taken as the overload duration step of the measuring point. Based on the numerical range of the comprehensive risk value, the risk level of the measuring point is determined according to the preset risk level classification rules, and corresponding operation suggestions are generated based on the risk level and the overload duration steps.

8. A multi-point load forecasting device for distribution networks based on controlled channel communication, characterized in that, include: The data acquisition module is used to acquire historical load data of each measuring point in the distribution network, and to construct load time series data samples corresponding to each measuring point based on the historical load data. The training sample construction module is used to construct multiple training samples based on the load time series data samples of each measuring point using the sliding window method. Each training sample includes load time subsequence samples of each measuring point in the same historical time window, as well as the actual load labels of each measuring point in the prediction time window. The model training module is used to train a pre-built multi-point load prediction model based on the multiple training samples. The multi-measurement point load prediction model includes: a channel-independent temporal coding module, used to temporally encode the load time subsequence samples of each measurement point in the training samples to obtain a channel-independent representation; a local direct communication module, used to group each measurement point and, based on the channel-independent representation, perform multi-head attention interaction within each group to obtain intra-group local communication results, and construct local communication messages based on the intra-group local communication results of each group; a prototype relay-type global communication module, used to perform information aggregation and backhaul between all measurement points and all global operating state prototypes based on a preset number of built-in learnable global operating state prototypes and the channel-independent representation, and construct global communication messages based on the aggregation and backhaul results; a controlled residual fusion module, used to perform residual fusion of the channel-independent representation, the local communication messages, and the global communication messages to obtain an updated representation; and a prediction output module, used to output the predicted load result value of each measurement point in the prediction time window based on the updated representation, so as to iteratively train the multi-measurement point load prediction model based on the predicted load result value and the corresponding real load label. The prediction module is used to input the load time subsequence of each measuring point within the target historical time window into the trained multi-measuring-point load prediction model to obtain the predicted load result value of each measuring point within the target prediction time window.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.