Power distribution collection terminal cooperative anomaly detection privacy protection method and system
Patent Information
- Application Number
- CN202610976003.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-02
AI Technical Summary
[0006]针对现有技术中配电采集终端协同异常检测无法同时满足拓扑关联异常的跨终端协同检测需求与终端数据隐私保护及资源约束的技术问题,本发明提供一种配电采集终端协同异常检测的隐私保护方法及系统
[0022]This invention achieves a balance between collaborative detection of cross-terminal topological association anomalies and terminal data privacy protection by generating privacy-preserving sketches through random projection dimensionality reduction based on the Johnson-Lindenstrauss lemma at the terminal side, and performing topological association detection based on graph attention networks at the concentrator side in the irreversible sketch space rather than the original feature space. The irreversible dimensionality reduction property of random projection ensures that the sketch of a single terminal does not contain enough information to reconstruct the original collected data. The superimposed differential privacy noise further provides a formal privacy guarantee, resolving the contradiction between collaborative anomaly detection and terminal data privacy protection in existing technologies. This invention utilizes the mathematical property that random projection dimensionality reduction naturally reduces sketch sensitivity, compressing the amplitude of the required superimposed differential privacy noise to half that in the original feature space under the same privacy budget, thus mitigating the impact of privacy-preserving noise on the detection accuracy of extremely low-frequency association anomalies. This invention adopts a three-layer architecture of terminal-concentrator-master station, which naturally matches the existing communication architecture of the power distribution network from terminal to concentrator to master station. The terminal only performs two low-computational-complexity operations: lightweight temporal encoding and matrix-vector multiplication. The amount of data in the sketch is only one-quarter of that in the original features. The computational and communication loads are controlled within the capacity of the embedded terminal. The graph attention computation required for topology association detection is deployed in the concentrator layer, which has relatively abundant computing power. The computational load of each layer of devices is matched with its actual resource capabilities. This invention compresses the high-expressive teacher encoder trained by the master station into a deep separable convolutional lightweight temporal encoder that can be run on the terminal through knowledge distillation. Under the strict computing power constraints of the terminal, the quality of temporal feature encoding is maintained. The federated aggregation process only involves the graph attention network parameters of the concentrator layer. The terminal does not need to upload any model parameters or gradients and does not participate in the model training process at all, further eliminating the computational, communication, and privacy risks brought about by the terminal's participation in federated learning training. This invention enhances the privacy and security of the system under long-term continuous operation by periodically changing the pseudo-random seed of the projection matrix and setting a transition window between the old and new seeds, while ensuring the continuity of detection and preventing attackers from using the time correlation of historical sketch sequences to carry out long-term cumulative attacks.
Smart Images

Figure CN122533858B_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence algorithms, deep learning, and the Internet of Things, specifically to a privacy protection method and system for collaborative anomaly detection of power distribution data acquisition terminals. Background Technology
[0002] As a critical link in the power system directly facing users, the safe and stable operation of the distribution network directly impacts power supply quality and user experience. With the widespread application of IoT technology in distribution networks, numerous data acquisition terminals are deployed at various nodes of distribution lines to monitor multi-dimensional electrical parameters such as voltage, current, power factor, and harmonic content in real time. Anomaly detection based on these massive amounts of collected data using deep learning and artificial intelligence algorithms has become an important technical means to ensure the safe operation of the distribution network. However, the physical topology of the distribution network determines that many typical anomalies do not occur in isolation at a single terminal, but rather propagate along the topological path of the distribution network between multiple terminals according to specific spatiotemporal patterns. For example, voltage sags extend from the fault point to both ends along the feeder, harmonic distortion propagates step by step between adjacent nodes, and multi-terminal cascade responses triggered by line faults. Accurate identification of these topology-related anomalies requires comprehensive analysis of the time-series data and topological relationships of multiple terminals on the same or adjacent feeders.
[0003] Federated learning technology has been applied to anomaly detection in smart grids to train a global anomaly detection model without centralizing raw data. In a typical federated learning framework, each participant independently trains its local anomaly detection model, and then a global model is generated through a federated averaging algorithm. However, this approach defines anomalies as statistical deviations from the local data of a single participant. The federated aggregation process only performs a weighted average of the model parameters, completely discarding the topological adjacency information between participants. The global model cannot encode the spatiotemporal propagation patterns along specific feeder paths, making it difficult to identify topologically related anomalies requiring cross-terminal collaborative judgment.
[0004] Graph neural network (GNN) technology provides an effective approach for detecting associated anomalies using the topology of power distribution networks. By modeling the physical topology of the distribution network as a graph, it captures spatial correlation features between nodes using mechanisms such as graph attention. However, existing topology-aware anomaly detection schemes based on GNNs require centralized access and processing of data collected from all nodes, making them unsuitable for data isolation. Distribution data acquisition terminals belong to different management areas, and the raw data collected by each terminal involves sensitive information such as user electricity consumption behavior. Directly aggregating raw data or intermediate feature representations for graph structure inference faces data security and compliance constraints. If an attempt is made to introduce intermediate feature or attention weight exchange between terminals within a federated learning framework to compensate for the lack of topology information, the exchanged intermediate representations can be reverse-engineered to reveal the sensitive features of the original data, creating new privacy leakage channels. There is an irreconcilable contradiction between the collaborative anomaly detection capability of topology-aware systems and the protection of terminal data privacy.
[0005] Furthermore, power distribution data acquisition terminals are typically based on embedded processors, with significantly lower memory capacity and computing power than edge computing nodes. Communication between these terminals and higher-level systems often relies on narrowband carrier waves or public wireless networks, resulting in limited available bandwidth. Relying on robust cryptographic techniques such as homomorphic encryption and secure multi-party computation to protect information exchange during collaboration is computationally and communicationally prohibitive for such resource-constrained terminals. Simultaneously, due to differences in the types of users served and topological locations, the data distribution of acquisition terminals on different feeders and in different transformer areas exhibits significant non-independent and identically distributed characteristics. Under these conditions, global model aggregation under the standard federated learning framework can easily lead to local detection performance degradation. In privacy-preserving scenarios, the superposition of differential privacy noise further amplifies this degradation effect. For topologically correlated anomaly samples with extremely low occurrence frequencies, the negative impact of noise on detection accuracy is particularly pronounced. Existing differential privacy schemes face a sharp decline in detection accuracy in such highly non-independent and identically distributed power distribution network scenarios with extremely rare anomalies. Summary of the Invention
[0006] To address the technical problem that existing technologies for collaborative anomaly detection in power distribution data acquisition terminals cannot simultaneously meet the requirements of cross-terminal collaborative detection of topology association anomalies while protecting terminal data privacy and mitigating resource constraints, this invention provides a privacy-preserving method and system for collaborative anomaly detection in power distribution data acquisition terminals. This method generates an irreversible privacy-preserving sketch by performing random projection dimensionality reduction on temporal features at the terminal side. At the concentrator side, topology association anomaly detection is performed in the sketch space using a graph attention network. At the master station side, cross-regional association determination and model federated updates are conducted. This achieves collaborative detection capability for cross-terminal topology association anomalies without disclosing the original features of the terminal's acquired data.
[0007] One aspect of the present invention provides a privacy protection method for collaborative anomaly detection of power distribution acquisition terminals, comprising the following steps: Step S1, the master station trains a lightweight timing encoder through knowledge distillation and deploys it to each power distribution acquisition terminal, trains a graph attention network and deploys it to a concentrator; Step S2, each terminal runs the lightweight timing encoder to encode the multidimensional electrical parameter timing data within a sliding window into a d-dimensional feature vector, generates a random projection matrix based on a pseudo-random seed shared with the concentrator, performs dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector, superimposes differential privacy noise to generate a privacy protection sketch and uploads it to the concentrator, where k is less than d; Step S3, the concentrator constructs a local detection graph with each terminal as a node, the power distribution network topology connection as an edge, and the privacy protection sketch as a node attribute, runs the graph attention network, calculates the anomaly score of each terminal in the sketch space of the local detection graph and determines the topology association anomaly; Step S4, the master station receives the anomaly detection reports from each concentrator, performs cross-regional association adjudication, and performs weighted federated average aggregation on the graph attention network parameters of each concentrator according to a preset period before distributing them.
[0008] The method includes the following steps.
[0009] Step S1: The main station trains a lightweight temporal encoder using knowledge distillation and deploys it to each power distribution acquisition terminal, and trains a graph attention network and deploys it to the concentrator. Specifically, the knowledge distillation process includes: first training a teacher encoder with a multi-layer dilated causal convolutional structure, and then training the lightweight temporal encoder with a depthwise separable one-dimensional convolutional structure using the output features of the teacher encoder as the supervision target. The distillation loss function is a weighted sum of mean squared error loss and cosine similarity loss. Further, the teacher encoder adopts a multi-layer temporal convolutional network structure, containing multiple dilated causal convolutional layers and residual connections, encoding the multi-dimensional acquisition data sequence of the terminal within a preset time window into a d-dimensional feature vector, where d takes the value of sixty-four. The lightweight temporal encoder consists of four depthwise separable convolutional layers and one global average pooling layer. Each convolutional layer contains two sub-operations: depthwise convolution and pointwise convolution. The depthwise convolution kernel size is five, and the pointwise convolution kernel size is one. The number of output channels for each layer is sixteen, thirty-two, sixty-four, and sixty-four, respectively. The lightweight timing encoder has no more than 30,000 parameters and its floating-point operations in a single inference run do not exceed 500,000 multiply-accumulate operations. It is compatible with embedded ARM (Advanced RISC Machine) processors with a clock frequency in the hundreds of megahertz range. After training, the lightweight timing encoder is deployed to each power distribution acquisition terminal via a concentrator.
[0010] Furthermore, the graph attention network comprises two graph attention convolutional layers. The first layer is configured with multiple attention heads, and the outputs of each attention head are concatenated and processed by an activation function. The second layer is configured with a single attention head, and its output is mapped to the anomaly score of each terminal by a sigmoid activation function. During the training phase, the historical data is subjected to random projection and differential privacy noise simulation with the same parameters as in the online phase, which serves as the training input for the graph attention network. Specifically, the first layer is configured with four attention heads, and the output dimension of each attention head is eight. For the target node and each of its topological neighbor nodes, their respective feature vectors are concatenated after undergoing a learnable linear transformation. The attention score is calculated through the shared attention vector, and the attention coefficients are obtained by LeakyReLU activation and softmax normalization. The negative slope parameter of LeakyReLU is set to 0.2. The outputs of the four attention heads are concatenated and activated by ELU (Exponential Linear Unit) as the thirty-two-dimensional output of the first layer. The second layer takes a 32-dimensional input, which, after single-head attention computation, outputs a one-dimensional scalar. This scalar is then mapped to the zero-to-one interval using a sigmoid activation function, serving as the anomaly score for each terminal node. The main site trains the graph attention network using a historical dataset containing normal operation data and labeled topologically related anomaly event data. The loss function is binary cross-entropy loss, the optimizer is Adam (Adaptive Moment Estimation), the initial learning rate is 0.001, and a cosine annealing strategy is employed.
[0011] In step S2, each terminal runs the lightweight timing encoder to encode the multi-dimensional electrical parameter timing data within the sliding window into a d-dimensional feature vector. Based on a pseudo-random seed shared with the concentrator, a random projection matrix is generated to perform dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector. Differential privacy noise is then superimposed to generate a privacy-protected sketch, which is uploaded to the concentrator, where k is less than d. Specifically, each distribution acquisition terminal segments the locally acquired multi-dimensional electrical parameter timing data, such as voltage, current, power factor, and harmonic content, according to a preset sliding time window. The length of the time window is determined based on the sampling period of the distribution network and the typical propagation time of associated anomalies. At the end of each time window, the terminal runs the lightweight timing encoder to encode the multi-dimensional timing data within that window into a sixty-four-dimensional feature vector.
[0012] Furthermore, the random projection matrix has a dimension of k rows and d columns, and its elements are independently sampled from a Gaussian distribution with a mean of zero and a variance of 1 / d. The global sensitivity of the privacy-preserving sketch is reduced by the square root of k divided by d compared to the sensitivity directly calculated on the d-dimensional eigenvector. The standard deviation of the differential privacy noise is calibrated based on the reduced sensitivity and privacy budget parameters. Specifically, the terminal and its concentrator share a pseudo-random number seed in advance, and both parties independently generate the same random projection matrix using a deterministic pseudo-random number generation algorithm based on this seed. The random projection matrix has a dimension of 16 rows and 64 columns, with a compression ratio of one-quarter. The terminal uses this random projection matrix to multiply the 64-dimensional eigenvector on the left to obtain a 16-dimensional projection vector. Based on the theoretical guarantee of the Johnson-Lindenstrauss lemma, the above random projection approximates the Euclidean distance relationship between any two eigenvectors with a high probability. For a typical distribution network scenario where the number of terminals in the distribution area does not exceed one hundred, when the projection dimension k is 16, the maximum distortion of the Euclidean distance between any two terminal sketches relative to the corresponding distance in the original feature space does not exceed 50%. Meanwhile, since the sixteen-dimensional projection vector is much lower than the sixty-four-dimensional original feature vector, reconstructing the original feature vector from the projection vector is an underdetermined problem with no unique solution. Therefore, the sketch of a single terminal is insufficient to reveal the specific features of its original acquired data.
[0013] Preferably, d is set to 64, k to 16, and the differential privacy noise is generated using a Gaussian mechanism. The privacy budget parameter epsilon ranges from one to ten, and delta is the negative first power of the total number of participating terminals. Specifically, the standard deviation sigma of the Gaussian noise is calibrated based on the global sensitivity of the sketch and the privacy budget parameter. The calculation method is: sigma equals the sketch sensitivity multiplied by two, multiplied by the natural logarithm 1.25, divided by the square root of delta, and then divided by epsilon. The global sensitivity of the sketch is defined as the maximum L2 norm change of the projection vector caused by the maximum permissible change in the original acquired data of a single terminal between adjacent datasets. When k equals 16 and d equals 64, the sketch sensitivity is half of the original feature sensitivity. Therefore, the required superimposed noise standard deviation is reduced to half compared to directly applying differential privacy in the original feature space, significantly mitigating the impact of privacy-preserving noise on anomaly detection accuracy. The terminal transmits the 16-dimensional privacy-preserving sketch with the terminal's topological location identifier after superimposing differential privacy noise to its respective concentrator via a narrowband communication channel. Compared to transmitting the original 64-dimensional feature vector, the amount of communication data for the sketch is reduced to one-quarter, satisfying the bandwidth constraints of narrowband communication in power distribution terminals.
[0014] Step S3: The concentrator constructs a local detection graph using each terminal as a node, the distribution network topology connection as an edge, and the privacy-preserving sketch as a node attribute. The graph attention network is then run to calculate the anomaly score for each terminal in the sketch space of the local detection graph and determine topology association anomalies. Specifically, after receiving the privacy-preserving sketch uploaded by all distribution acquisition terminals within its jurisdiction, the concentrator constructs a local detection graph based on the distribution network topology adjacency relationship. This local detection graph uses each terminal as a node, the physical connection relationship of the distribution lines between terminals as edges, and the 16-dimensional privacy-preserving sketch uploaded by each terminal as node attributes. The distribution network topology adjacency relationship is maintained by the master station and sent to the concentrator for updating when the topology changes. When a switching operation in the distribution network causes a topology change, the concentrator updates the edge connection relationship of the local detection graph accordingly. The concentrator runs the graph attention network on the local detection graph. The input of the graph attention network is the sixteen-dimensional sketch vector of each node. Since the node input used by the graph attention network during the training phase is the sketch vector of the original features after the same random projection and simulated noise superposition, the network has learned the ability to extract topological association features in the sketch space.
[0015] Further, the determination of topological association anomalies includes: marking terminals with anomaly scores exceeding a single-node threshold as suspected anomaly nodes; checking whether the suspected anomaly nodes constitute a connected subgraph on the local detection graph with a node count reaching a minimum association scale threshold; if so, performing delay mode verification on the peak time of the anomaly scores of each node in the connected subgraph; and determining a topological association anomaly when the peak time of each node along the topological path shows a delay pattern that gradually increases from the source node outwards. Specifically, the concentrator jointly determines the anomaly scores of each terminal output by the graph attention network. When the anomaly score of a terminal exceeds a preset single-node threshold, the terminal is marked as a suspected anomaly node. The concentrator further checks whether the suspected anomaly nodes constitute a connected subgraph on the topological graph, i.e., whether there are multiple topologically adjacent suspected anomaly nodes forming a connected anomaly region. If the number of nodes in the connected subgraph formed by the suspected anomaly nodes reaches a preset minimum association scale threshold, then delay mode verification is initiated. The specific method for verifying the time delay pattern is as follows: For each suspected abnormal node in the connected subgraph, extract the time series of its abnormal score within the most recent time windows, determine the time when the abnormal score of each node reaches its peak, and if the sequence of peak times of each node conforms to a time delay pattern that gradually increases from a certain source node outward along the topological path, that is, the terminal farther away from the source node has a later peak time, and the time delay difference between adjacent nodes is within a reasonable range, then the abnormal event in the connected subgraph is determined to have topological correlation. After the concentrator confirms the topological correlation anomaly, it generates an anomaly detection report containing information such as the event type, the set of involved terminals, the estimated location of the propagation source node, and the direction of the propagation path, and reports it to the main station.
[0016] Preferably, the single-node threshold is 0.7, the minimum association size threshold is 3, and the maximum allowable delay deviation in the delay pattern verification is the length of two detection time windows.
[0017] Step S4: The master station receives anomaly detection reports from each concentrator, performs cross-regional association adjudication, and distributes the graph attention network parameters of each concentrator after weighted federated average aggregation according to a preset period. Specifically, after receiving the association anomaly detection reports reported by each concentrator, the master station performs cross-regional association adjudication on association anomalies involving adjacent substation boundaries. When concentrators in two adjacent substations sharing a boundary detect topology association anomaly events within a short time window, and both events involve terminal sets that include terminals located at the substation boundary, the master station merges the two events into a single cross-regional topology association anomaly event and reassesses the propagation source node and propagation range by integrating the report information from both concentrators.
[0018] Furthermore, each concentrator fine-tunes the graph attention network by constructing a pseudo-label dataset based on local high-confidence detection results, and then uploads the model parameters to the main station. The weights of the weighted federated average aggregation are proportional to the number of terminals under the jurisdiction of each concentrator. Specifically, within each update cycle, each concentrator fine-tunes its graph attention network based on locally accumulated detection samples. The training data used for fine-tuning is constructed in a pseudo-label manner: the concentrator labels the set of sketch samples corresponding to the high-confidence association anomaly judgment results output during the detection process as positive samples, and randomly samples an equal number of samples from the sketch samples recorded during historical normal operation as negative samples, thus forming the local fine-tuning dataset. The fine-tuning adopts the binary cross-entropy loss function, with a learning rate of one-tenth of the initial training learning rate, and fine-tunes one training epoch within each update cycle. After fine-tuning, each concentrator uploads the model parameters of its graph attention network to the main station. The main station aggregates the model parameters of each concentrator using a weighted federated average algorithm, and distributes the aggregated global graph attention network parameters to each concentrator. Each concentrator replaces its local model with the global model for subsequent detection.
[0019] Furthermore, the master station updates the lightweight temporal encoder on the terminal side at a frequency lower than the federated update cycle of the graph attention network. The master station retrains the teacher encoder using the newly accumulated collected data, generates updated student encoder parameters through knowledge distillation, and distributes them to each terminal through the concentrator. After receiving the new encoder parameters, the terminal replaces the local encoder and uses the updated encoder from the next detection time window.
[0020] Furthermore, the pseudo-random seed is changed at a preset period. During the change, the concentrator generates a new seed and distributes it to each terminal through an encrypted channel. An overlapping transition window is set between the old and new seeds. Within the overlapping transition window, each terminal simultaneously uses both the old and new seeds to generate two drafts for upload. The concentrator maintains two sets of projection matrices for detection. After the overlapping transition window ends, the new seed is used. Specifically, the periodic seed change ensures that even if an attacker intercepts multiple consecutive time windows of transmitted draft sequences, the changed projection matrices prevent the attacker from using the temporal correlation between drafts to infer the original data information through accumulated observations, thus enhancing the system's privacy and security under long-term operating conditions.
[0021] Another aspect of the present invention provides a privacy protection system for collaborative anomaly detection of a power distribution acquisition terminal, comprising a power distribution acquisition terminal, a concentrator, and a master station; the power distribution acquisition terminal includes a data acquisition unit, a timing encoding module, and a sketch generation module; the data acquisition unit is used to acquire local electrical parameter timing data; the timing encoding module is used to run a lightweight timing encoder to encode the multidimensional electrical parameter timing data within a sliding window into a d-dimensional feature vector; the sketch generation module is used to generate a random projection matrix based on a pseudo-random seed shared with the concentrator, perform dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector, and superimpose differential privacy noise to generate a privacy-protected sketch, which is then uploaded to the concentrator, wherein k is small. The concentrator includes a graph construction module and an association detection module. The graph construction module is used to construct a local detection graph with each terminal as a node, the distribution network topology connection as an edge, and the privacy protection sketch as a node attribute. The association detection module is configured with a graph attention network, which is used to run the graph attention network in the sketch space of the local detection graph to calculate the anomaly score of each terminal and determine the topology association anomaly. The master station includes a cross-region association module and a federated update module. The cross-region association module is used to receive the anomaly detection reports of each concentrator and perform cross-region association adjudication. The federated update module is used to perform weighted federated average aggregation on the graph attention network parameters of each concentrator according to a preset period and then distribute them. Specifically, the data acquisition unit collects local electrical parameter time-series data according to a preset sampling period. The time-series encoding module is configured with a lightweight time-series encoder trained by knowledge distillation, which encodes the multi-dimensional time-series data in the sliding window into a fixed-dimensional feature vector. The sketch generation module includes a random projection unit and a differential privacy noise superposition unit. The random projection unit generates a random projection matrix based on a pseudo-random seed shared with the concentrator and performs dimensionality reduction projection on the feature vector. The differential privacy noise superposition unit adds Gaussian noise calibrated according to privacy budget parameters to each component of the projection result, generating a privacy-preserving sketch with topology location identifiers and transmitting it to the concentrator. The graph construction module receives privacy-preserving sketches uploaded by each terminal within the distribution area and constructs a local detection graph based on the distribution network topology adjacency relationship issued by the master station, with terminals as nodes, physical connections as edges, and sketches as node attributes. The association detection module is configured with a graph attention network, calculates the anomaly score of each terminal node on the local detection graph, and determines topology association anomalies through connected subgraph analysis and time delay pattern verification, generating a detection report and reporting it to the master station. The cross-region association module receives anomaly detection reports from each concentrator and performs cross-regional merging and adjudication of association anomalies at the boundaries of adjacent distribution areas. The federated update module collects the graph attention network parameters of each concentrator at a preset period, performs weighted federated average aggregation, and distributes the aggregated global model parameters to each concentrator.
[0022] This invention achieves a balance between collaborative detection of cross-terminal topological association anomalies and terminal data privacy protection by generating privacy-preserving sketches through random projection dimensionality reduction based on the Johnson-Lindenstrauss lemma at the terminal side, and performing topological association detection based on graph attention networks at the concentrator side in the irreversible sketch space rather than the original feature space. The irreversible dimensionality reduction property of random projection ensures that the sketch of a single terminal does not contain enough information to reconstruct the original collected data. The superimposed differential privacy noise further provides a formal privacy guarantee, resolving the contradiction between collaborative anomaly detection and terminal data privacy protection in existing technologies. This invention utilizes the mathematical property that random projection dimensionality reduction naturally reduces sketch sensitivity, compressing the amplitude of the required superimposed differential privacy noise to half that in the original feature space under the same privacy budget, thus mitigating the impact of privacy-preserving noise on the detection accuracy of extremely low-frequency association anomalies. This invention adopts a three-layer architecture of terminal-concentrator-master station, which naturally matches the existing communication architecture of the power distribution network from terminal to concentrator to master station. The terminal only performs two low-computational-complexity operations: lightweight temporal encoding and matrix-vector multiplication. The amount of data in the sketch is only one-quarter of that in the original features. The computational and communication loads are controlled within the capacity of the embedded terminal. The graph attention computation required for topology association detection is deployed in the concentrator layer, which has relatively abundant computing power. The computational load of each layer of devices is matched with its actual resource capabilities. This invention compresses the high-expressive teacher encoder trained by the master station into a deep separable convolutional lightweight temporal encoder that can be run on the terminal through knowledge distillation. Under the strict computing power constraints of the terminal, the quality of temporal feature encoding is maintained. The federated aggregation process only involves the graph attention network parameters of the concentrator layer. The terminal does not need to upload any model parameters or gradients and does not participate in the model training process at all, further eliminating the computational, communication, and privacy risks brought about by the terminal's participation in federated learning training. This invention enhances the privacy and security of the system under long-term continuous operation by periodically changing the pseudo-random seed of the projection matrix and setting a transition window between the old and new seeds, while ensuring the continuity of detection and preventing attackers from using the time correlation of historical sketch sequences to carry out long-term cumulative attacks. Attached Figure Description
[0023] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0024] Figure 1 This is a schematic diagram of the overall process of the privacy protection method for collaborative anomaly detection of power distribution acquisition terminals provided in this embodiment of the invention.
[0025] Figure 2 This is a schematic diagram of the data flow of the terminal-side privacy protection sketch generation process provided in an embodiment of the present invention.
[0026] Figure 3This is a schematic diagram of the concentrator-side topology association anomaly detection process provided in an embodiment of the present invention.
[0027] Figure 4 This is a structural block diagram of the privacy protection system for collaborative anomaly detection of power distribution acquisition terminals provided in an embodiment of the present invention.
[0028] Figure 5 This is a visualization of the Johnson-Lindenstrauss random projection dimensionality reduction and privacy protection mechanism provided in this embodiment of the invention.
[0029] Figure 6 This is a visualization of the sketch space graph attention reasoning mechanism provided in the embodiments of the present invention.
[0030] Figure 7 This is a timing diagram of the federated learning aggregation and seed cycle replacement mechanism provided in an embodiment of the present invention.
[0031] Figure 8 This is a scenario diagram of the physical deployment and communication link of the power distribution network area provided in the embodiment of the present invention.
[0032] Figure 9 This is a spatial topology diagram for cross-regional anomaly propagation and detection provided in an embodiment of the present invention.
[0033] Figure 10 This is a comparison chart of ROC curves and PR curves for node-level anomaly detection in various methods provided in the embodiments of the present invention.
[0034] Figure 11 This is a thermogram of ablation experiment performance indicators provided in an embodiment of the present invention.
[0035] Figure 12 This is a multi-dimensional comparison chart of communication efficiency and computational latency provided in an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives and technical solutions of this invention clearer, the embodiments of this invention will be further described in detail below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this invention and should not be construed as limiting the scope of protection of this invention. Example 1
[0037] like Figure 1 As shown, this embodiment of the invention provides a privacy protection method for collaborative anomaly detection of power distribution acquisition terminals. The method is based on a three-layer division of labor architecture of terminal-concentrator-master station, and includes the following steps S1 to S4.
[0038] Step S1: The main station trains a lightweight temporal encoder through knowledge distillation and deploys it to each power distribution acquisition terminal, and trains a graph attention network and deploys it to the concentrator.
[0039] In this embodiment, the main station first constructs a training dataset based on historical data. This training dataset contains multi-dimensional electrical parameter time-series data for multiple transformer substations over at least a six-month operating period. Each data sample is extracted according to a preset sliding time window, with a window length of fifteen minutes, corresponding to thirty sampling points under a typical sampling interval of the distribution network. The training dataset includes normal operation data and labeled topology-related anomaly event data. The anomaly events are manually labeled by maintenance experts based on historical fault work orders and equipment alarm records. The labeling granularity is the binary label of each terminal within each time window, as well as the terminal set and propagation path information of the associated anomaly events. The training dataset is divided into a training set, a validation set, and a test set in a 7:1:2 ratio. During the division, stratified sampling is performed on a transformer substation basis to ensure a consistent proportion of anomaly events in each subset.
[0040] The main station uses a teacher-student knowledge distillation framework to train a lightweight temporal encoder. The main station first trains the teacher encoder, which employs a multi-layer temporal convolutional network structure containing six dilated causal convolutional layers and residual connections. The dilation factors are 1, 2, 4, 8, 16, and 32, respectively. Each layer has a kernel size of 3 and 128 channels. The teacher encoder encodes the multi-dimensional data sequence acquired by the terminal within a time window into a 64-dimensional feature vector. The teacher encoder is trained using a contrastive learning paradigm, with encoding results from the same terminal in adjacent time windows as positive sample pairs and encoding results from different terminals as negative sample pairs. The loss function is InfoNCE (Noise Contrastive Estimation) loss, and the temperature parameter is set to 0.07. Training was performed using the AdamW (Adaptive Moment Estimator with Weight Decay) optimizer, with an initial learning rate of 4.7 × 10^-4, a weight decay factor of 0.01, a batch size of 256, and a training duration of 120 epochs. The learning rate was gradually decayed from the initial value to one percent of the initial value using a cosine annealing strategy. Training was completed on a workstation equipped with four GPUs (Graphics Processing Units).
[0041] After the teacher encoder is trained, the main station uses the output features of the teacher encoder as the supervision target to train a lightweight temporal encoder with a depthwise separable one-dimensional convolutional structure as the student encoder. The student encoder consists of four depthwise separable convolutional layers and one global average pooling layer. Each convolutional layer contains two sub-operations: depthwise convolution and pointwise convolution. The depthwise convolution kernel size is five, and the pointwise convolution kernel size is one. The number of output channels in each layer is sixteen, thirty-two, sixty-four, and sixty-four, respectively. The distillation loss function is a weighted sum of mean squared error loss and cosine similarity loss, with weights of 0.6 and 0.4, respectively. Distillation training uses the Adam optimizer with an initial learning rate of 7.8 × 10^-4, a batch size of 512, and a training duration of 80 epochs. The learning rate also uses a cosine annealing strategy. During distillation, the parameters of the teacher encoder are frozen, and only the student encoder is updated. After training, the number of parameters in the student encoder does not exceed 30,000, the floating-point operations in a single inference do not exceed 500,000 multiply-accumulate operations, and it is compatible with embedded ARM processors with a clock frequency in the hundreds of megahertz range. After distillation, the cosine similarity between the output features of the student encoder and the teacher encoder is evaluated on the test set. A mean cosine similarity of 0.93 or higher is considered acceptable for distillation quality. The main station then distributes the student encoder parameters that have passed quality verification to each power distribution acquisition terminal via a concentrator.
[0042] The main station uses the same training dataset to train the graph attention network. During the training phase, the main station performs random projection and differential privacy noise simulation on the original 64-dimensional feature vectors of each terminal in the historical data, with the same parameters as in the online phase, to generate simulated sketch vectors as the training input of the graph attention network. This makes the network robust to information distortion caused by noise and dimensionality compression in the sketch space. The graph attention network consists of two graph attention convolutional layers. The first layer is configured with four attention heads, each with an output dimension of eight. For the target node and each of its topological neighbor nodes, their respective feature vectors are concatenated after a learnable linear transformation. Attention scores are calculated using the shared attention vectors, and attention coefficients are obtained by LeakyReLU activation and softmax normalization. The negative slope parameter of LeakyReLU is set to 0.2. The outputs of the four attention heads are concatenated and activated by the ELU function to serve as the 32-dimensional output of the first layer. The second layer is configured with a single attention head, with a 32-dimensional input. After attention calculation, the output is a one-dimensional scalar, which is mapped to the interval between zero and one by the sigmoid activation function as the anomaly score for each terminal. The training loss function is binary cross-entropy loss, the optimizer is Adam, the initial learning rate is 1.2 × 10^-3 with cosine annealing, the batch size is 32 graph samples, and the training duration is 200 epochs. An early stopping strategy is used: training terminates when the AUC (Area Under Curve) on the validation set fails to improve for 15 consecutive epochs. To address class imbalance between normal and abnormal samples, the loss for abnormal samples is weighted fivefold. After training, the graph attention network is deployed to each concentrator. The detection performance of the graph attention network is evaluated on the test set, using the area under the receiver operating characteristic (ROC) curve and the area under the precision-recall curve as the primary evaluation metrics.
[0043] In step S2, each terminal runs the lightweight timing encoder to encode the multidimensional electrical parameter timing data within the sliding window into a d-dimensional feature vector. Based on the pseudo-random seed shared with the concentrator, a random projection matrix is generated to perform dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector. Differential privacy noise is superimposed to generate a privacy-preserving sketch, which is then uploaded to the concentrator, where k is less than d.
[0044] like Figure 2As shown, each power distribution acquisition terminal segments the locally acquired multi-dimensional electrical parameter time-series data, such as voltage, current, power factor, and harmonic content, according to a preset sliding time window. In this embodiment, the time window length is set to fifteen minutes, and there is no overlap between adjacent time windows. At the end of each time window, the terminal runs the lightweight time encoder deployed in step S1 to encode the time-series data matrix containing dimensions such as effective voltage value, effective three-phase current value, active power, reactive power, power factor, and total harmonic distortion rate into a 64-dimensional feature vector. The encoding process is completed locally on the terminal. The encoder input is a two-dimensional matrix formed by splicing the multi-dimensional electrical parameters at each sampling time within the window. The time-domain features are extracted layer by layer through four layers of depthwise separable convolution, and finally aggregated along the time dimension by a global average pooling layer to output a 64-dimensional feature vector.
[0045] The terminal performs random projection dimensionality reduction on the 64-dimensional feature vector obtained from the encoding. The terminal and its concentrator share a pseudo-random number seed beforehand via a secure key negotiation protocol. Based on this seed, both parties independently generate a 16x64 random projection matrix using a deterministic pseudo-random number generation algorithm. Each element in the matrix is independently sampled from a Gaussian distribution with zero mean and 1 / 64 variance. The terminal uses this random projection matrix to left-multiply the 64-dimensional feature vector to obtain a 16-dimensional projection vector, achieving a compression ratio of 1 / 4. Based on the theoretical guarantee of the Johnson-Lindenstrauss lemma, for a typical distribution network scenario with no more than one hundred terminals in the area, when the projection dimension is 16, the maximum distortion of the Euclidean distance between any two terminal sketches relative to the corresponding distance in the original feature space does not exceed 50%. Since the 16-dimensional projection vector is much lower than the 64-dimensional original feature vector, reconstructing the original feature vector from the projection vector is an underdetermined problem with no unique solution. A single terminal sketch is insufficient to reveal the specific characteristics of its original collected data.
[0046] Building upon random projection, the terminal further superimposes differential privacy noise onto each component of the sixteen-dimensional projection vector to provide formal privacy guarantees. The noise is generated using a Gaussian mechanism, satisfying the definition of relaxed differential privacy. In this embodiment, the privacy budget parameter epsilon is set to three, and delta is set to the negative first power of the total number of participating terminals N. The standard deviation sigma of the Gaussian noise is calculated as follows: sigma equals the sketch sensitivity multiplied by two, multiplied by the natural logarithm 1.25, divided by the square root of delta, and then divided by epsilon. The global sensitivity of the sketch is defined as the maximum L2 norm change in the projection vector caused by the maximum permissible variation of the raw data acquired by a single terminal between adjacent datasets. When the projection dimension k equals sixteen and the original feature dimension d equals sixty-four, the sketch sensitivity is half of the original feature sensitivity, meaning the sensitivity is reduced proportionally to the square root of k divided by d. Therefore, the required superimposed noise standard deviation is reduced to half compared to directly applying differential privacy in the original feature space. The terminal transmits a 16-dimensional privacy-preserving sketch with differential privacy noise superimposed on it, along with its own topological location identifier, to its concentrator via a narrowband communication channel. This reduces the amount of communication data to one-quarter compared to transmitting the original 64-dimensional feature vector.
[0047] like Figure 5 As shown, the privacy protection mechanism of this invention achieves synergy between feature dimensionality reduction and privacy protection based on the Johnson-Lindenstrauss random projection principle. The left side of the figure shows the distribution of feature points of multiple terminals in a 64-dimensional feature space, where normal terminals and abnormal terminals maintain a clear distance relationship d1 in the high-dimensional space. The middle part shows the structure of the 16×64-dimensional random projection matrix R, which is generated based on a pseudo-random seed shared by each terminal and the concentrator.
[0048] After random projection, the high-dimensional feature vectors are compressed into a 16-dimensional sketch space, as shown on the right side of the figure. According to the Johnson-Lindenstrauss lemma, the distance d' between point pairs after projection is approximately the same as the original distance d1, with a compression ratio of 1 / 4. Differential privacy Gaussian noise is superimposed on the projected sketch vectors, and the noise perturbation range is represented by gray halos in the figure. The sensitivity reduction formula is delta_sketch = delta_orig × sqrt(k / d), meaning that the projection dimensionality reduction itself reduces the sensitivity to half of the original value, thus significantly reducing the noise amplitude under the same privacy budget. This dimensionality reduction process is irreversible; since the projection matrix is an underdetermined system, an attacker cannot uniquely recover the original feature vectors from the sketch.
[0049] Step S3: The concentrator constructs a local detection graph with each terminal as a node, the distribution network topology connection as an edge, and the privacy protection sketch as a node attribute. The graph attention network is then run to calculate the anomaly score of each terminal in the sketch space of the local detection graph and determine the topology association anomaly.
[0050] like Figure 3 As shown, after receiving the privacy-protected sketches uploaded by all distribution acquisition terminals within its jurisdiction, the concentrator constructs a local detection graph based on the distribution network topology adjacency relationship. This local detection graph uses each terminal as a node, the physical connection relationship of the distribution lines between terminals as edges, and the 16-dimensional privacy-protected sketches uploaded by each terminal as node attributes. In this embodiment, a typical distribution area contains 30 to 80 distribution acquisition terminals, and the average node degree of the local detection graph is two to four, reflecting the characteristics of the tree-like or weak-loop topology of the distribution network. The topology adjacency relationship of the distribution network is maintained by the master station and sent to the concentrator for updating when the topology changes. When a switching operation in the distribution network causes a topology change, the concentrator updates the edge connection relationship of the local detection graph accordingly. The concentrator independently generates a random projection matrix identical to that on the terminal side using the same pseudo-random seed deployed in step S1, used to verify the consistency of the sketch format, but not for reverse recovery of the original features.
[0051] The concentrator runs the graph attention network deployed in step S1 on the local detection graph, taking the 16-dimensional sketch vector of each node as input. In the first layer of graph attention convolution, for each target node, the graph attention network first obtains the 16-dimensional sketch vector of the node and all its topological neighbors, maps them to an 8-dimensional intermediate representation through the learnable linear transformation corresponding to each attention head, concatenates the intermediate representations of the target node and its neighbors into a 16-dimensional vector, and projects it onto the shared attention vector to obtain a scalar attention score. After the attention score is activated by LeakyReLU, softmax normalization is performed among all neighbors of the target node to obtain attention coefficients. The output of the target node under this attention head is the weighted sum of the transformed features of the neighbor nodes according to the attention coefficients. The 8-dimensional outputs of the four attention heads are concatenated into a 32-dimensional vector and activated by ELU. The second layer calculates a 1-dimensional scalar output through single-head attention, maps it to the zero-to-one interval through sigmoid as the anomaly score of the terminal.
[0052] The concentrator performs a three-stage joint judgment on the anomaly scores of each terminal output by the graph attention network to determine whether there are topologically related anomalies. The first stage is single-node screening, marking terminals with anomaly scores exceeding 0.7 as suspected anomaly nodes. The second stage is spatial clustering verification, where the concentrator checks whether suspected anomaly nodes form a connected subgraph with three or more nodes in the local detection graph, meaning at least three topologically adjacent terminals simultaneously exhibit high anomaly scores. If the minimum association scale threshold is not met, it is only recorded as an isolated single-point anomaly and does not trigger association anomaly judgment. The third stage is time delay pattern verification, where for each suspected anomaly node within a connected subgraph that meets the spatial clustering condition, the time series of its anomaly score within the most recent several time windows is extracted to determine the time when each node's anomaly score reaches its peak. If the sequence of peak times of each node conforms to a time delay pattern that gradually increases along the topological path from a source node outwards (i.e., the terminal farther from the source node has a later peak time), and the time delay difference between adjacent nodes does not exceed the length of two detection time windows, then the anomaly event in the connected subgraph is determined to have topological correlation. In the time-delay mode, the source node is estimated to be the node with the earliest peak time, and the propagation direction is the topology path direction that increases with the peak time. After the concentrator confirms the topology association anomaly, it generates an anomaly detection report containing information such as event type, set of involved terminals, estimated propagation source node location, and propagation path direction, and reports it to the main station.
[0053] like Figure 6 As shown, this invention employs a multi-head graph attention network on the concentrator side to perform anomaly reasoning on terminal nodes in the sketch space. The left side of the figure shows a distribution network topology subgraph, where the attributes of each terminal node are 16-dimensional privacy-preserving sketches. Target node i and its neighboring nodes j1, j2, and j3 are selected. First, the 16-dimensional sketch is projected onto an 8-dimensional attention space using a linear transformation matrix W (8×16). Then, the transformation results of the target node and each neighboring node are concatenated, and the attention score is calculated using the attention vector a.
[0054] After softmax normalization, attention weights are obtained, such as a(i,j1)=0.5, a(i,j2)=0.3, and a(i,j3)=0.2, reflecting the contribution of each neighboring node to the anomaly detection of the target node. This invention employs four attention heads for parallel computation, each independently focusing on different feature patterns, and the outputs are concatenated into a 32-dimensional vector. After processing by the second-layer graph attention network, anomaly scores ranging from 0 to 1 are output through the sigmoid activation function. The figure below illustrates the time-delay propagation pattern of anomalies along the topological path, with the peak time showing an increasing trend along the topological path.
[0055] In step S4, the main station receives the anomaly detection reports from each concentrator, performs cross-regional association adjudication, and distributes the graph attention network parameters of each concentrator after weighted federated average aggregation according to a preset period.
[0056] After receiving the correlation anomaly detection reports from each concentrator, the main station performs cross-regional correlation adjudication on correlation anomalies involving adjacent substation boundaries. When concentrators in two adjacent substations sharing a boundary detect topology correlation anomaly events within a time difference of no more than three detection time windows, and both events involve terminals located at the substation boundary, the main station merges the two events into a single cross-regional topology correlation anomaly event and reassesses the propagation source node and propagation range by integrating the report information from both concentrators. The main station classifies the severity of the merged cross-regional event according to its propagation range and the number of terminals involved, and generates a comprehensive alarm report containing complete propagation link information.
[0057] The main station performs weighted federated average aggregation of the graph attention network model parameters of each concentrator according to a preset federated update cycle. In this embodiment, the federated update cycle is set to seven days. Within each update cycle, each concentrator fine-tunes its graph attention network based on locally accumulated detection samples. The training data used for fine-tuning is constructed using pseudo-labels: the concentrator marks the set of sketch samples corresponding to the high-confidence associated anomaly judgment results output during the detection process as positive samples. Specifically, sketches corresponding to associated anomaly events with anomaly scores exceeding 0.9 and verified through the time-delay mode are used as positive samples, and an equal number of samples are randomly sampled from the sketch samples recorded during historical normal operation and marked as negative samples, thus forming the local fine-tuning dataset. Fine-tuning uses the binary cross-entropy loss function, with a learning rate of one-tenth of the initial training learning rate, i.e., 1.2 × 10^-4, and the optimizer is Adam. Fine-tuning is performed for one training epoch within each update cycle. After fine-tuning, each concentrator uploads its graph attention network model parameters to the main station. The main station uses a weighted federated average algorithm to aggregate the model parameters of each concentrator. The aggregation weight is proportional to the number of terminals under the jurisdiction of each concentrator. That is, if the i-th concentrator governs n_i terminals, its aggregation weight is n_i divided by the total number of terminals in all concentrators. The aggregated global graph attention network parameters are distributed to each concentrator, and each concentrator replaces its local model with the global model for subsequent detection.
[0058] The master station updates the lightweight temporal encoder on the terminal side at a frequency lower than the federated update cycle of the graph attention network; in this embodiment, the encoder update cycle is set to thirty days. The master station retrains the teacher encoder using newly accumulated data within the update cycle, generates updated student encoder parameters through knowledge distillation, and distributes them to each terminal via a concentrator. Upon receiving the new encoder parameters, the terminal replaces its local encoder and uses the updated encoder starting from the next detection time window.
[0059] Furthermore, the pseudo-random seed of the random projection matrix is changed according to a preset seed update cycle, which is set to fourteen days in this embodiment. When the seed is changed, the concentrator generates a new pseudo-random seed and distributes it to all terminals under its jurisdiction via an encrypted communication channel. Two overlapping transition windows of varying detection time window lengths are set between the old and new seeds. Within the transition window, the terminal simultaneously uses both the old and new seeds to generate two drafts for upload. The concentrator maintains two sets of projection matrices for detection simultaneously. After the transition window ends, it switches to using only the new seed. This periodic seed change ensures that even if an attacker intercepts multiple consecutive time window transmissions of draft sequences, the changed projection matrix prevents them from inferring the original data information through accumulated observations based on the temporal correlation between the drafts, thus enhancing the system's privacy and security under long-term operating conditions.
[0060] like Figure 7 As shown, the system operation of this invention involves a coordination mechanism with three different cycles. At the terminal layer, each acquisition terminal uploads a 16-dimensional privacy-preserving sketch to its respective concentrator every 15 minutes. At the concentrator layer, each concentrator continuously receives the terminal sketches and runs a graph attention network to detect topological association anomalies.
[0061] A federated update cycle is performed every 7 days: after each concentrator completes local fine-tuning, it uploads the model parameters to the main station. The main station performs weighted federated average aggregation and then distributes the global model back to each concentrator. A seed replacement cycle is performed every 14 days. During the replacement transition window, the old and new seeds run in parallel, and the terminals simultaneously upload two copies of the draft image generated based on the two seeds, ensuring that detection continuity is not affected by seed switching. An encoder update cycle is performed every 30 days. The main station retrains the system using knowledge distillation based on the latest data and distributes the updated lightweight encoder parameters to each terminal. The nested coordination of these three cycles ensures that the system maintains both detection accuracy and privacy protection levels during continuous operation. Example 2
[0062] In a preferred embodiment of the present invention, the process of topological association anomaly detection performed by a graph attention network in sketch space is described in detail. This embodiment focuses on sketch space inference on the concentrator side, and provides the specific working mechanism and parameter configuration for the graph attention network to learn topological association features in a low-dimensional projected space.
[0063] After receiving the 16-dimensional privacy-preserving sketches uploaded by N power distribution acquisition terminals within the distribution area, the concentrator constructs a local detection graph G=(V, E, X), where V is the set of terminal nodes, E is the set of edges formed by the topological connections of the power distribution lines, and X is a node attribute matrix with each node's sketch vector as the row, and X has a dimension of N rows and sixteen columns. The first layer of the graph attention network performs attention calculations on each target node i and its neighboring nodes j in its topological neighbor set N(i). The specific process is as follows: First, the 16-dimensional sketch vectors of node i and node j are projected into an eight-dimensional space through a shared linear transformation matrix W. The linear transformation matrix W has a dimension of eight rows and sixteen columns, and its parameters are obtained through training. Then, the two projected eight-dimensional vectors are concatenated into a 16-dimensional vector, and an inner product operation is performed with the learnable attention vector a. The attention vector a has a dimension of sixteen. The inner product result is processed by the LeakyReLU activation function to obtain the original attention score e(i,j), where the negative slope parameter of LeakyReLU is set to 0.2. Softmax normalization is performed on the raw attention scores of all topological neighbors of target node i to obtain normalized attention coefficients alpha(i,j). These coefficients reflect the relative importance of the sketch information of neighbor node j to the inference of the abnormal state of target node i. The first-layer output feature of target node i is the result of the weighted sum of the linearly transformed features of all its topological neighbors according to the attention coefficients. The first layer is configured with four independent attention heads, each using an independent linear transformation matrix and attention vector. The eight-dimensional vectors output by the four attention heads are concatenated into a thirty-two-dimensional vector, which is then processed by the ELU activation function as the final output of the first layer.
[0064] Specifically, the feasibility of performing the above attention calculation in sketch space is based on the following conditions: the Johnson-Lindenstrauss (hereinafter referred to as JL) lemma guarantees that random projections approximate the Euclidean distance between vectors with a high probability, while the attention score of the graph attention mechanism essentially depends on the similarity measure between node feature vectors. When the projection dimension k is sixteen and the original feature dimension d is sixty-four, for a distribution network scenario where the number of terminals N in the distribution area does not exceed one hundred, according to the classical form of the JL lemma, the relative error between the inner product between any two node sketches and the corresponding inner product in the original feature space is kept within the distortion parameter ε_JL with a probability not less than one minus two divided by the square of N, where the relationship between ε_JL and k is that ε_JL is equal to a constant multiplied by the square root of the natural logarithm N divided by k. When N equals one hundred and k equals sixteen, the value of ε_JL is approximately 0.54, meaning that although the calculation of the attention score has a certain order of magnitude of distortion, it is still distinguishable for differentiating the topological association feature patterns between normal and abnormal nodes. Furthermore, since the input samples used by the graph attention network during the training phase are also sketch vectors processed with the same parameters of random projection and differential privacy noise, the network's attention parameters have been adapted to the feature distribution and distortion characteristics in the sketch space during the training process. The learning objective of the attention vector a and the linear transformation matrix W is to maximize the detection and discrimination ability of topological association anomalies under the feature distribution conditions of the sketch space. Therefore, the trained network has inherent robustness to projection distortion in the sketch space.
[0065] Preferably, the training process of the graph attention network employs the following strategy to enhance detection capabilities in the sketch space. The training dataset contains time-series electrical parameter data during historical normal operation and during labeled topology-related anomaly events. The ratio of normal to anomaly samples is typically 20:1 to 50:1, reflecting the low-frequency characteristics of associated anomaly events in the distribution network. To address the severe imbalance in the ratio of positive to negative samples, a weighted binary cross-entropy loss function is used, with the anomaly class weight set to ten to twenty times the normal class weight. Training uses the Adam optimizer with an initial learning rate of 0.001, employing a cosine annealing strategy, with the minimum learning rate decaying to one percent of the initial value. The number of training epochs is one hundred to two hundred. Batch size is in graph units, with each batch containing four to eight local detection graphs. During training, random projection and differential privacy noise stacking with the same parameters as in the online phase are performed on each training sample. Within each training epoch, different random projection matrix instances and different noise samples are used for the same sample, enabling the network to learn feature representations that generalize to both projection matrix instances and noise instances. Specifically, for each batch of samples in each training epoch, a seed is randomly selected from a predefined candidate seed pool to generate the projection matrix. The size of the candidate seed pool is fifty to one hundred seeds, thus simulating the change in the projection matrix after the periodic replacement of seeds in online deployment. When the number of labeled associated anomaly events in the training data is less than twenty, a data augmentation strategy is adopted: a small Gaussian perturbation is applied to the sketch vector of the existing anomaly events to generate augmented samples. The standard deviation of the perturbation is five to ten percent of the L2 norm of the original sketch. At the same time, a non-critical edge is randomly discarded or a non-existent edge is randomly added from the topology graph to simulate the scenario of minor topological changes. When the number of terminals in the transformer area is too small, resulting in fewer than ten nodes in the topology graph, the graph attention network has limited topological association patterns that can be learned. Under this condition, the single-node anomaly threshold can be appropriately reduced to improve the recall rate.
[0066] Furthermore, the input to the second-layer graph attention convolution is the 32-dimensional node features output from the first layer. A single attention head is configured, and the linear transformation matrix has a dimension of one row and thirty-two columns. This projects the 32-dimensional input to a one-dimensional output, which is then mapped by a sigmoid activation function to scalar values in the range of zero to one, serving as the anomaly score for each node. For a local detection graph with N nodes within a distribution area, the computational complexity of the second layer is O(|E| multiplied by 32), where |E| is the number of edges. For a typical radial distribution network topology, |E| is approximately equal to N minus one. Therefore, the computational cost of a single inference is approximately linearly related to the number of terminals, making it suitable for the ARM processor of the concentrator.
[0067] After calculating node-level anomaly scores, the concentrator performs topology association determination. Nodes with anomaly scores exceeding 0.7 are marked as suspected anomaly nodes. The largest connected subgraph formed by suspected anomaly nodes is searched on the topology graph. If a connected subgraph contains three or more suspected anomaly nodes, a time-delay pattern verification is initiated for that subgraph. The time-delay pattern verification extracts the time series of anomaly scores for each suspected anomaly node within the most recent ten detection time windows, determining the moment when the anomaly score first exceeds 0.5 as the anomaly start time for that node. It then verifies whether the anomaly start times of each node in the subgraph increase outwards from the source node along the topology path. Specifically, the node with the earliest anomaly start time in the connected subgraph is marked as a candidate source node. A breadth-first traversal is performed along the topology path starting from the candidate source node, verifying whether the anomaly start time of each node reached in each traversal step is later than the anomaly start time of its topology parent node. The maximum allowed time-delay deviation is two detection time windows. When all topology paths pass the time-delay increment verification, the anomaly event corresponding to the connected subgraph is determined to have topology association, and an anomaly detection report is generated.
[0068] The technical solution of this embodiment has the following beneficial effects. Regarding privacy protection, the entire inference process of the graph attention network is completed in a 16-dimensional sketch space. The concentrator never accesses the original 64-dimensional feature vector of any terminal. Even if the concentrator itself is compromised, the attacker only obtains the irreversibly projected low-dimensional sketch, unable to reconstruct the original collected data such as the terminal's voltage, current, and power factor. Regarding detection accuracy, by using multi-instance random projection and noise simulation during the training phase, the attention parameters of the graph attention network are adapted to the feature distribution of the sketch space. The weighted binary cross-entropy loss function corrects for the imbalance of positive and negative samples in low-frequency associated anomalies, enabling the network to learn effective topological association features even under sample scarcity conditions. Regarding computational efficiency, the 16-dimensional input of the sketch space, compared to the original 64-dimensional features, reduces the number of linear transformation parameters in the first layer of the graph attention network to one-quarter of the original. This correspondingly reduces the floating-point operation volume of a single inference operation on the concentrator side, meeting the real-time processing requirements of the concentrator's ARM processor. Regarding communication efficiency, the amount of sketch data is one-quarter of the original features, allowing for uploading within a single sampling period under narrowband carrier communication conditions in the power distribution terminal.
[0069] From a theoretical perspective, the technical feasibility of graph attention inference in sketch space in this embodiment stems from the following two mutually supporting technical factors. The Johnson-Lindenstrauss lemma mathematically guarantees that random projection approximately preserves the Euclidean distance between vector pairs in high-dimensional space. Furthermore, the calculation of attention coefficients in the graph attention mechanism relies on the concatenated inner product of node feature vectors after linear transformation. The result of the inner product operation is directly related to the distance and angle between vectors. Therefore, the distance preservation characteristic ensures that the ranking of attention coefficients in sketch space is roughly consistent with the ranking in the original feature space. That is, neighboring nodes that contribute significantly to the anomaly detection of the target node in the original space tend to receive higher attention weights in the sketch space. Simultaneously, the training process of the graph attention network itself constitutes an adaptive learning of the feature distribution in sketch space: training samples are input in sketch form, and attention parameters are backpropagated to find the optimal discrimination direction on the feature manifold of sketch space. This end-to-end training implicitly compensates for the perturbations introduced by projection distortion and differential privacy noise, and its compensation capability is reflected in the learned values of the attention vector and the linear transformation matrix. The synergistic effect of the two factors keeps the degradation of graph attention reasoning in sketch space in terms of detection capability within an acceptable range compared to the original feature space. Example 3
[0070] In another embodiment of the present invention, a privacy protection scheme for collaborative anomaly detection in power distribution acquisition terminals based on sparse binary random projection and graph isomorphic networks is provided. This embodiment solves the same technical problem as the aforementioned embodiments, but adopts different technical paths in the selection of privacy projection mechanism and topology inference network.
[0071] In this embodiment, the privacy-preserving projection on the terminal side uses a sparse binary random projection matrix instead of the Gaussian random projection matrix described in the previous embodiment. The sparse binary random projection matrix still has a dimension of k rows and d columns, where k is 16 and d is 64. However, the matrix elements are no longer sampled from a Gaussian distribution, but from a discrete distribution: each element has a one-sixth probability of being the reciprocal of the square root of a positive s, a one-sixth probability of being the reciprocal of the square root of a negative s, and a two-thirds probability of being zero, where s equals three. Approximately two-thirds of the elements in the resulting projection matrix are zero, and the non-zero elements take only two symmetric values. Compared to the Gaussian random projection matrix, approximately two-thirds of the matrix-vector multiplication operations in the sparse binary projection matrix can be skipped directly. Multiplication at non-zero positions is simplified to addition and subtraction operations after division by a constant, without involving floating-point multiplication. The computational load of the terminal's matrix-vector projection is reduced to approximately one-third that of the Gaussian projection, further reducing the demand on the limited computing power of the embedded terminal. Sparse binary random projection also satisfies the distance preservation condition of the Johnson-Lindenstrauss lemma, and its distance preservation accuracy is theoretically equivalent to that of Gaussian projection, thus not affecting the feature discriminability of subsequent topological inference. The terminal still shares the pseudo-random seed with the concentrator and manages the seed according to the same periodic replacement strategy. The differential privacy noise superposition process is consistent with the aforementioned implementation method.
[0072] On the concentrator side, topological association anomaly detection uses a graph isomorphic network instead of the graph attention network in the aforementioned implementation. Each layer of the graph isomorphic network performs the following operations: For target node i, firstly, the sketch features of all its topological neighbors are summed and aggregated to obtain the neighborhood aggregated features. Then, the target node's own sketch features are multiplied by a learnable parameter η_l, incremented by one, and added to the neighborhood aggregated features, where η_l is the learnable scalar parameter of the l-th layer, initially set to zero, and the increment operation ensures the basic preservation of the target node's own features. The result after addition is input into a two-layer multilayer perceptron for nonlinear transformation. The hidden dimension of the first layer is twice the input dimension, and the activation function is ReLU. The output dimension of the second layer is the same as the input dimension. This implementation configures a three-layer graph isomorphic network with an input dimension of sixteen and a hidden dimension of thirty-two for each layer. The output features of the three-layer graph isomorphic network are summarized by a global readout function and input into a linear classification layer to obtain the anomaly score of each node.
[0073] Graph isomorphic networks differ from graph attention networks in their expressive power. Graph isomorphic networks aggregate neighborhood features using a summation operation, possessing discriminative power equivalent to the Weisfeiler-Leman graph isomorphism test from a graph theory perspective, capable of distinguishing different graph structures and node neighborhood patterns. In contrast, graph attention networks assign different weights to the features of different neighbors through attention coefficients. While this may provide finer-grained discrimination in scenarios with significant differences in node features, for topological association anomalies in distribution networks where multiple nodes within the neighborhood simultaneously deviate from normal patterns, the summation aggregation of graph isomorphic networks can effectively capture the cumulative effect of overall neighborhood anomalies. Specifically, when the sketch features of multiple neighboring nodes deviate from the normal distribution, summation aggregation amplifies the deviation signals, which is beneficial for detecting multi-node collaborative anomaly patterns.
[0074] Regarding the determination of association anomalies, this implementation adopts the same connected subgraph detection and latency pattern verification strategies as the aforementioned implementations. The settings for parameters such as single-node threshold, minimum association size threshold, and latency deviation tolerance are consistent with the aforementioned implementations. The federated update and seed management mechanisms also remain consistent with the aforementioned methods; the main station periodically collects the graph isomorphic network parameters of each concentrator, performs weighted federated average aggregation, and then distributes them.
[0075] As an alternative, this implementation may have relative advantages in the following scenarios: When the processor computing power of the power distribution acquisition terminal is more limited and can only support integer or fixed-point operations, the skipping of zero elements and the addition and subtraction of non-zero positions in the sparse binary projection matrix can completely avoid floating-point multiplication, realizing privacy-preserving projection under pure integer operation paths; when the anomaly of interest is characterized by multi-node synchronous deviation rather than inter-node differential responses, the summation and aggregation of graph isomorphic networks are well-suited to capturing such patterns. When the number of terminals in the distribution area is large, leading to an increase in the size of the topology graph, graph isomorphic networks have a certain advantage in computational complexity compared to graph attention networks because they do not involve calculating the attention coefficients between node pairs one by one. The time complexity of each layer is O(|V| multiplied by the hidden dimension plus |E| multiplied by the input dimension), where the overhead of attention coefficient calculation is eliminated. However, in scenarios where the terminal node features are highly differentiated and fine-grained distinction of different neighbor contributions is required, graph attention networks may exhibit better detection accuracy due to their adaptive attention weighting mechanism. Example 4
[0076] like Figure 4 As shown, this embodiment of the invention provides a privacy protection system for collaborative anomaly detection of power distribution acquisition terminals. The system includes three levels of devices: power distribution acquisition terminals, concentrators, and master stations. The devices at each level are connected through existing communication links in the power distribution network.
[0077] The power distribution data acquisition terminal includes a data acquisition unit, a timing encoding module, and a sketch generation module. The data acquisition unit collects multi-dimensional electrical parameter timing data, such as local voltage, current, power factor, and harmonic content, according to a preset sampling period, and segments and buffers the collected data according to a preset sliding time window. The timing encoding module is equipped with a lightweight timing encoder trained through knowledge distillation, used to encode the multi-dimensional electrical parameter timing data within the sliding window into a d-dimensional feature vector. The lightweight timing encoder consists of four depthwise separable convolutional layers and one global average pooling layer. Each convolutional layer contains two sub-operations: depthwise convolution and pointwise convolution. The number of parameters does not exceed 30,000, and the floating-point operations in a single inference do not exceed 500,000 multiply-accumulate operations, making it compatible with embedded ARM processors with a clock frequency in the hundreds of megahertz range. The sketch generation module includes a random projection unit and a differential privacy noise superposition unit. It generates a random projection matrix based on a pseudo-random seed shared with the concentrator, performs dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector, and superimposes differential privacy noise to generate a privacy-preserving sketch, which is then uploaded to the concentrator. Here, k is less than d. The random projection unit generates a k-row, d-column random projection matrix using a deterministic pseudo-random number generation algorithm based on the pseudo-random seed shared with the concentrator, and performs matrix-vector multiplication on the feature vector to obtain the k-dimensional projection vector. The differential privacy noise superposition unit adds Gaussian noise calibrated according to the privacy budget parameters to each component of the projection vector. The noise standard deviation is calculated based on the global sensitivity of the sketch and the privacy budget parameters epsilon and delta. The generated privacy-preserving sketch includes the topology location identifier of this terminal and is transmitted to the concentrator via a narrowband communication channel.
[0078] The concentrator includes a graph construction module and an association detection module. The graph construction module receives privacy-preserving sketches uploaded by terminals within the distribution area and constructs a local detection graph using each terminal as a node, the distribution network topology connections as edges, and the privacy-preserving sketches as node attributes. The distribution network topology adjacency relationships are maintained by the master station and sent to the concentrator when the topology changes. The graph construction module updates the edge connections of the local detection graph accordingly when the topology changes. The association detection module is configured with a graph attention network, which is used to calculate the anomaly score of each terminal and determine topology association anomalies in the sketch space of the local detection graph. The graph attention network contains two graph attention convolutional layers. The first layer has four attention heads, and the outputs of each attention head are concatenated and processed by the ELU activation function. The second layer has a single attention head, and its output is mapped to the anomaly score of each terminal by the sigmoid activation function. The association detection module performs a three-stage joint judgment: terminals with anomaly scores exceeding the single-node threshold are marked as suspected anomaly nodes; it checks whether the suspected anomaly nodes constitute a connected subgraph in the local detection graph with a node count reaching the minimum association scale threshold; if so, it performs time delay pattern verification on the peak time of the anomaly score of each node in the connected subgraph; when the peak time of each node shows a time delay pattern that gradually increases from the source node outward along the topological path, it is determined to be a topological association anomaly. After confirming the topological association anomaly, the association detection module generates an anomaly detection report containing information such as event type, the set of involved terminals, the estimated location of the propagation source node, and the direction of the propagation path, and reports it to the main station.
[0079] The main station includes a cross-region association module and a federated update module. The cross-region association module receives anomaly detection reports from each concentrator and performs cross-region association rulings. When concentrators in two adjacent zones sharing a boundary detect topology association anomalies within a short time window, and both events involve terminals located at the zone boundary, the two events are merged into a single cross-region topology association anomaly event. The propagation source node and propagation range are then reassessed based on the reports from both concentrators. The federated update module performs weighted federated average aggregation of the graph attention network parameters of each concentrator at a preset period before distributing them. The aggregation weight is proportional to the number of terminals under the jurisdiction of each concentrator. The federated update module is also responsible for updating the lightweight temporal encoder on the terminal side at a frequency lower than the graph attention network federated update cycle. It uses newly accumulated data to generate updated student encoder parameters through knowledge distillation and distributes them to each terminal via the concentrator.
[0080] At the hardware deployment level, the power distribution acquisition terminal uses an embedded ARM processor with a clock frequency in the hundreds of megahertz range and a memory capacity in the hundreds of kilobytes range, performing two computationally low-computational-complexity operations: lightweight temporal encoding and matrix-vector projection. The concentrator uses an ARM processor with relatively ample computing power and a memory capacity in the megabytes range, undertaking the computational tasks of graph attention network inference and association anomaly detection. The main station is deployed on a server equipped with a GPU, responsible for computationally intensive tasks such as teacher encoder training, knowledge distillation, initial training of the graph attention network, and federated aggregation. Communication between devices at each layer follows the existing layered communication architecture of the power distribution network from terminal to concentrator to main station. Narrowband carrier communication channels are used between the terminal and the concentrator, while broadband communication channels are used between the concentrator and the main station. The computational load of each device is matched with its actual resource capabilities. The terminal only needs to perform lightweight temporal encoder inference with no more than 30,000 parameters and vector multiplication of a 16x64 matrix. The amount of communication data for the sketch is only one-quarter of the original features, meeting the bandwidth constraints of narrowband communication. Example 5
[0081] This embodiment illustrates the implementation process and technical effects of the above-mentioned privacy protection method for collaborative anomaly detection of power distribution acquisition terminals through a specific application case.
[0082] The application scenario involves three adjacent transformer substations in a distribution network area, with a total of 150 distribution data acquisition terminals deployed. Each substation contains 45, 52, and 53 terminals respectively. Each terminal collects eight electrical parameters at one-minute intervals: RMS voltage, RMS three-phase current, active power, reactive power, power factor, and total harmonic distortion (THD). The sliding time window is set to 15 minutes. The distribution network topology is radial, with concentrators in the three substations managing their respective terminals and cross-regional information exchange via a master station. In this application scenario, due to the complex distribution network topology, large number of terminals, and limited narrowband communication bandwidth, traditional centralized anomaly detection methods require each terminal to upload complete raw data to the master station for global analysis, resulting in high communication overhead and the risk of terminal data privacy leakage. Furthermore, methods based on independent detection of a single terminal cannot capture cross-terminal topology-related anomaly patterns.
[0083] like Figure 8 As shown, the privacy-protected collaborative anomaly detection system of this invention is deployed in a typical power distribution area scenario. The center of the area is a distribution transformer, and power lines radiate from the transformer in a tree-like topology. Power acquisition terminals are installed at each monitoring node of the power lines. Each terminal has a built-in ARM processor responsible for collecting multi-dimensional electrical parameters such as voltage, current, power factor, and harmonic content.
[0084] Each data acquisition terminal uploads a 16-dimensional privacy-preserving sketch to a concentrator located in the distribution room via narrowband carrier or wireless public network communication links, with a single upload data volume of only 64 bytes. The concentrator connects to the main station via a broadband communication link and is responsible for reporting anomaly detection reports and model parameters. A typical distribution area contains 30 to 80 data acquisition terminals, with an average topology node degree of 2 to 4. The figure also illustrates the propagation process of voltage sag anomalies along the feeder, with the anomalous signal propagating step by step along the distribution line topology from the source node to the downstream terminal.
[0085] The main station first uses twelve months of historical operational data for offline model training. This historical data includes 123 manually labeled topological association anomalies and a large amount of normal operational data, with an anomaly-to-normal data sample ratio of approximately 1:35. Following step S1, the main station trains the teacher encoder and the lightweight student encoder. After distillation, the mean cosine similarity between the student encoder output features and the teacher encoder output features reaches 0.94. The main station then trains the graph attention network using the sketch vectors (simulated by random projection and differential privacy noise) of the original features from each terminal in the training data as input. During training, the loss weight for the anomaly class is set to five times that for the normal class. After training, the student encoder is deployed to each terminal, and the graph attention network is deployed to three concentrators.
[0086] During the online detection phase, each terminal runs a student encoder at the end of each 15-minute time window to generate a 64-dimensional feature vector. This vector is then dimensionality-reduced using a 16x64 random projection matrix and superimposed with Gaussian difference privacy noise with a privacy budget parameter of epsilon equal to 3, generating a 16-dimensional privacy-preserving sketch which is uploaded to the concentrator. Each concentrator constructs a local detection graph and runs a graph attention network to calculate the anomaly score for each terminal. It then performs connected subgraph analysis and latency pattern verification to determine topological association anomalies.
[0087] During the three-month testing period, a total of 37 topology association anomalies were detected across three transformer substations, which were subsequently manually verified. Of these, 34 were intra-substation topology association anomalies, and 3 involved cross-regional topology association anomalies at the boundaries of adjacent substations. Using the area under the receiver operating characteristic (AUC-ROC) curve and the area under the precision-recall (AUC-PR) curve as the main evaluation metrics, the method of this invention achieved an AUC-ROC of 0.917 and an AUC-PR of 0.683 for node-level anomaly detection. For event-level topology association anomaly determination, using precision, recall, and F1 score as evaluation metrics, the method of this invention achieved a precision of 0.857, a recall of 0.811, and an F1 score of 0.833. Of the three cross-regional topology association anomalies, the main station correctly merged and identified two through cross-regional association adjudication. The third was missed because the association anomaly score of one substation was slightly below the threshold and was not reported by the concentrator on that side.
[0088] like Figure 9 As shown, this invention supports anomaly propagation detection and association determination across distribution zones. The figure illustrates the spatial topology of three adjacent distribution zones: Zone 1 contains 45 terminals, Zone 2 contains 52 terminals, and Zone 3 contains 53 terminals. Each zone has a radial distribution topology with a concentrator device at its center. Shared boundary terminals exist at the zone boundaries, indicated by two-color markings representing their special locations simultaneously affected by two zones.
[0089] The diagram illustrates an abnormal event path propagating from area 1 across the boundary to area 2. The source node is marked with a star. The abnormal signal passes through time nodes t1, t2, t3, and t4 sequentially along the topological path, with the peak value showing an increasing trend, consistent with the physical propagation law. When concentrators in areas 1 and 2 simultaneously report abnormal events involving boundary terminals, the main station performs cross-regional correlation determination, merges the detection results from both sides, determines the anomaly to be a cross-area propagation event, and outputs the complete propagation path and source node location. Each concentrator is connected to the main station via a communication link, enabling the aggregation of detection reports and collaborative updates of the global model.
[0090] To verify the contributions of each key technical component of this invention, comparisons were made with the following methods on the same test data. Method A is a centralized graph attention detection method without privacy protection; each terminal directly uploads its 64-dimensional original feature vector to the concentrator, which then runs a graph attention network with the same structure in the original feature space for detection. Method B is a privacy-protected method using only differential privacy without random projection dimensionality reduction; each terminal directly superimposes Gaussian noise with the same privacy budget onto its 64-dimensional original feature vector before uploading. Method C is an isolated detection method where each terminal independently uses autoencoder reconstruction errors for anomaly detection without considering topological association information.
[0091] In terms of node-level anomaly detection performance, Method A has an AUC-ROC of 0.938 and an AUC-PR of 0.721; Method B has an AUC-ROC of 0.874 and an AUC-PR of 0.586; and Method C has an AUC-ROC of 0.812 and an AUC-PR of 0.437. The AUC-ROC of the method in this invention is 0.917, only 0.021 lower than Method A, indicating that random projection dimensionality reduction has a relatively small impact on detection accuracy. The graph attention network effectively compensates for projection distortion through sketch space adaptation learning during the training phase. Compared to Method B, the AUC-ROC of the method in this invention is improved by 0.043 and the AUC-PR by 0.097. This is because random projection reduces the sensitivity to half, requiring less noise amplitude to be superimposed under the same privacy budget, thus preserving more feature information that can be used for detection. Compared with method C, the method of the present invention improves AUC-ROC and AUC-PR by 0.105 and 0.246 respectively, demonstrating the significant gain of topological association information on anomaly detection capability.
[0092] In terms of performance in event-level topological association anomaly detection, Method A has a precision of 0.875, a recall of 0.838, and an F1 score of 0.856; Method B has a precision of 0.813, a recall of 0.703, and an F1 score of 0.754. Method C, lacking topological association detection capabilities, cannot be directly compared at the event level. The method of this invention achieves a 0.023 decrease in F1 score compared to Method A and a 0.079 improvement compared to Method B.
[0093] To verify the performance of the method of this invention in node-level anomaly detection tasks, comparative experiments were conducted on a power distribution network simulation dataset. The comparison methods included: Method A, a centralized scheme without privacy protection; Method B, a scheme using only differential privacy without random projection; and Method C, a scheme where each terminal detects anomalies in isolation and does not utilize topology information. The evaluation metrics used were the area under the ROC curve (AUC-ROC) and the area under the PR curve (AUC-PR).
[0094] like Figure 10As shown, the left figure compares the ROC curves of each method, and the right figure compares the PR curves. The AUC-ROC of the method of this invention reaches 0.917, only 2.1 percentage points lower than the centralized method A without privacy protection, indicating that the impact of random projection dimensionality reduction and differential privacy noise superposition on detection performance is small. Compared with method B which only uses differential privacy, the AUC-ROC of the method of this invention is improved by 4.3 percentage points, indicating that random projection dimensionality reduction not only reduces communication overhead, but also reduces the amount of privacy noise required through sensitivity reduction effect. The AUC-ROC of method C is only 0.812, which is 10.5 percentage points lower than that of the method of this invention, verifying the significant role of topological association detection in improving anomaly detection capability. In terms of the PR curve, the AUC-PR of the method of this invention is 0.683, which is a more significant improvement compared to the 0.437 of method C, indicating that the introduction of topological information makes an important contribution to reducing the false alarm rate.
[0095] Further ablation experiments were conducted to verify the technical contributions of each core component. Ablation scheme one removed the random projection dimensionality reduction step and directly superimposed differential privacy noise on the original 64-dimensional features before performing graph attention detection. This scheme corresponds to the configuration of method B. Ablation scheme two removed the differential privacy noise superposition step and only performed random projection dimensionality reduction before directly uploading the 16-dimensional projection vector to the concentrator. The node-level AUC-ROC of this scheme was 0.925, and the AUC-PR was 0.701, which is a slight improvement compared to the complete scheme. This indicates that differential privacy noise has a certain impact on detection accuracy, but the improvement is limited, indicating that random projection dimensionality reduction itself already provides the main privacy protection capability. Ablation scheme three replaced the graph attention network with a multilayer perceptron that does not consider the topology structure and directly calculated the anomaly score independently for the 16-dimensional sketch of each terminal. The event-level F1 score of this scheme decreased to 0.521, which is 0.312 lower than the complete scheme. This indicates that the graph attention network's use of topological neighborhood information for association reasoning is the key source of the topological association anomaly detection capability. Ablation scheme four replaces knowledge distillation training with direct training of a lightweight temporal encoder. This scheme reduces the mean cosine similarity of the output features of the student encoder and the teacher encoder to 0.81 on the test set, and the node-level AUC-ROC to 0.883, indicating that the knowledge distillation mechanism plays an important role in maintaining coding quality under terminal resource constraints.
[0096] To evaluate the contribution of each core component of this invention to the overall system performance, ablation experiments were designed to remove the random projection module, differential privacy noise module, topology information, and knowledge distillation mechanism, respectively. The detection performance of each ablation scheme was compared under the same test conditions. Evaluation metrics included node-level AUC-ROC and AUC-PR, as well as event-level precision, recall, and F1 score.
[0097] like Figure 11As shown, the complete scheme achieves superior performance across all five metrics, with an AUC-ROC of 0.917 and an event-level F1 score of 0.833. Ablation 3, after removing topological information, shows the most significant performance degradation, with the event-level F1 score decreasing from 0.833 to 0.521 and the AUC-PR from 0.683 to 0.498, indicating that the graph attention network's utilization of topological association information is the core source of the invention's detection capability. Ablation 1, after removing random projection, lowers the event-level F1 score to 0.754, demonstrating that random projection not only serves privacy protection but also helps reduce noise interference through its compact representation after dimensionality reduction. Ablation 4, after removing knowledge distillation, lowers the AUC-ROC to 0.883, verifying the role of teacher model knowledge transfer in improving the feature quality of the lightweight encoder. Ablation 2, after removing differential privacy noise, shows a slight performance improvement, as expected, indicating that the performance loss caused by the privacy protection mechanism is within an acceptable range.
[0098] Regarding communication efficiency, the data uploaded by each terminal in each time window in the method of this invention is 16 floating-point numbers, or 64 bytes, which is reduced to one-quarter compared to the 64 floating-point numbers, or 256 bytes, uploaded by methods A and B. With a scale of 150 terminals across three distribution zones, the total uplink communication data volume for each 15-minute detection cycle is approximately 9,400 bytes, satisfying the bandwidth constraints of the narrowband carrier communication channel for the power distribution terminal. In terms of computational efficiency, the single inference time of the lightweight timing encoder on the terminal side is approximately 15 milliseconds on a 100 MHz-level ARM processor, and the matrix vector projection and noise superposition operations take approximately 0.3 milliseconds. The total computational overhead on the terminal side is far less than the 15-minute detection window cycle. The single inference time of the graph attention network on the concentrator side on the largest distribution zone containing 53 nodes is approximately 8 milliseconds, also meeting the real-time detection requirements.
[0099] To verify the feasibility of deploying the method of this invention on resource-constrained power distribution data acquisition terminals, comparative tests were conducted on the communication data volume, computational latency, and concentrator-side inference latency at the terminal side. The test hardware consisted of an ARM Cortex-M4 processor with a clock speed of 100MHz, 256KB of RAM, and narrowband carrier communication. Method A directly uploaded the original 64-dimensional feature vector, while Method B uploaded differential privacy noise superimposed on the 64-dimensional feature vector.
[0100] like Figure 12 As shown above, the method of this invention uploads only 64 bytes of data per cycle, a 75% reduction compared to the 256 bytes of methods A and B, thus meeting the narrowband carrier single-frame transmission constraint. Regarding terminal-side computational latency, the total time of this invention is approximately 15.3 milliseconds, with approximately 15 milliseconds for the encoding stage and only 0.3 milliseconds for random projection. This is essentially the same as the 15 milliseconds of method A and the 15.1 milliseconds of method B, indicating that the additional computational overhead from the random projection operation is negligible. Figure 12As shown in the lower part, in terms of inference latency on the concentrator side, the method of the present invention compresses the input dimension from 64 dimensions to 16 dimensions, and the graph attention network inference latency is only about 8 milliseconds, which is 55.6% lower than the 18 milliseconds of method A and method B, significantly improving the concentrator's real-time response capability to multi-terminal anomalies.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A privacy protection method for collaborative anomaly detection in power distribution data acquisition terminals, characterized in that, Includes the following steps: Step S1: The main station trains a lightweight temporal encoder through knowledge distillation and deploys it to each power distribution acquisition terminal; it also trains a graph attention network and deploys it to the concentrator. Step S2: Each terminal runs the lightweight timing encoder to encode the multidimensional electrical parameter timing data in the sliding window into a d-dimensional feature vector. Based on the pseudo-random seed shared with the concentrator, a random projection matrix is generated to perform dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector. Differential privacy noise is superimposed to generate a privacy protection sketch and upload it to the concentrator, where k is less than d. Step S3: The concentrator constructs a local detection graph with each terminal as a node, the distribution network topology connection as an edge, and the privacy protection sketch as a node attribute. The graph attention network is then run to calculate the anomaly score of each terminal in the sketch space of the local detection graph and determine the topology association anomaly. In step S4, the main station receives the anomaly detection reports from each concentrator, performs cross-regional association adjudication, and distributes the graph attention network parameters of each concentrator after weighted federated average aggregation according to a preset period.
2. The method according to claim 1, characterized in that, In step S1, the knowledge distillation process includes: first training a teacher encoder with a multi-layer dilated causal convolutional structure, and then training a lightweight temporal encoder with a depth separable one-dimensional convolutional structure using the output features of the teacher encoder as the supervision target. The distillation loss function is a weighted sum of mean square error loss and cosine similarity loss.
3. The method according to claim 1, characterized in that, In step S2, the random projection matrix has a dimension of k rows and d columns, and the matrix elements are independently sampled from a Gaussian distribution with a mean of zero and a variance of 1 / d. The global sensitivity of the privacy-preserving sketch is reduced by the ratio of k divided by the square root of d compared to the sensitivity directly calculated on the d-dimensional feature vector. The standard deviation of the differential privacy noise is calibrated based on the reduced sensitivity and privacy budget parameters.
4. The method according to claim 3, characterized in that, The value of d is 64, the value of k is 16, the differential privacy noise is generated using a Gaussian mechanism, the privacy budget parameter epsilon ranges from one to ten, and the value of delta is the negative first power of the total number of participating terminals.
5. The method according to claim 1, characterized in that, In step S3, the determination of topological association anomalies includes: marking terminals with anomaly scores exceeding the single-node threshold as suspected abnormal nodes; checking whether the suspected abnormal nodes constitute a connected subgraph on the local detection graph with the number of nodes reaching the minimum association scale threshold; if so, performing time delay mode verification on the peak time of the anomaly scores of each node in the connected subgraph; and determining that the topological association anomalies occur when the peak time of each node shows a time delay pattern that gradually increases from the source node outward along the topological path.
6. The method according to claim 5, characterized in that, The single-node threshold is set to 0.7, the minimum association size threshold is set to 3, and the maximum allowed time delay deviation in the time delay mode verification is the length of two detection time windows.
7. The method according to claim 1, characterized in that, The pseudo-random seed is changed according to a preset period. When changing, the concentrator generates a new seed and distributes it to each terminal through an encrypted channel. An overlapping transition window is set between the old and new seeds. Within the overlapping transition window, each terminal uses both the old and new seeds to generate two drafts and upload them. The concentrator maintains two sets of projection matrices for detection. After the overlapping transition window ends, the new seed is used.
8. The method according to claim 1, characterized in that, In step S1, the graph attention network includes two graph attention convolutional layers. The first layer is configured with multiple attention heads, and the outputs of each attention head are concatenated and processed by an activation function. The second layer is configured with a single attention head, and the output is mapped to the abnormal score of each terminal by a sigmoid activation function. During the training phase, the same random projection and differential privacy noise simulation with the same parameters as the online phase are performed on historical data as the training input for the graph attention network.
9. The method according to claim 1, characterized in that, In step S4, each concentrator constructs a pseudo-label dataset based on the local high-confidence detection results, fine-tunes the graph attention network, and then uploads the model parameters to the main station. The weight of the weighted federated average aggregation is proportional to the number of terminals under the jurisdiction of each concentrator.
10. A privacy protection system for collaborative anomaly detection in power distribution data acquisition terminals, characterized in that, This includes power distribution data acquisition terminals, concentrators, and master stations; The power distribution acquisition terminal includes a data acquisition unit, a timing encoding module, and a sketch generation module. The data acquisition unit is used to acquire local electrical parameter timing data. The timing encoding module is used to run a lightweight timing encoder to encode the multidimensional electrical parameter timing data within a sliding window into a d-dimensional feature vector. The sketch generation module is used to generate a random projection matrix based on a pseudo-random seed shared with the concentrator, perform dimensionality reduction projection on the d-dimensional feature vector to obtain a k-dimensional projection vector, and superimpose differential privacy noise to generate a privacy-protected sketch, which is then uploaded to the concentrator, where k is less than d. The concentrator includes a graph construction module and an association detection module. The graph construction module is used to construct a local detection graph with each terminal as a node, the distribution network topology connection as an edge, and the privacy protection sketch as a node attribute. The association detection module is configured with a graph attention network, which is used to run the graph attention network in the sketch space of the local detection graph to calculate the abnormal score of each terminal and determine the topology association abnormality. The main station includes a cross-region association module and a federated update module. The cross-region association module is used to receive anomaly detection reports from each concentrator and perform cross-region association rulings. The federated update module is used to perform weighted federated average aggregation on the graph attention network parameters of each concentrator according to a preset period and then distribute them.
Citation Information
Patent Citations
Intelligent electric meter anomaly detection method and system based on federal differential privacy and attention mechanism
CN120670825A
Distribution communication network distributed cloud edge collaborative anomaly detection method and system
CN122268613A