Power distribution room anomaly detection system based on cloud side-end cooperation

Through the cloud-edge-device collaborative architecture, real-time data processing and efficient operation and maintenance of the intelligent power distribution room system have been realized, solving the problems of low data transmission efficiency, insufficient edge computing capabilities, low anomaly detection accuracy, and high operation and maintenance response latency, thereby improving the system's detection accuracy and response speed.

CN121417469APending Publication Date: 2026-01-27BEIHANG UNIV
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202511393915.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing intelligent power distribution room systems suffer from problems such as low data transmission efficiency, insufficient edge computing capabilities, low anomaly detection accuracy, difficulty in multi-source data fusion, and high operation and maintenance response delays, making it difficult to achieve real-time, accurate anomaly detection and efficient operation and maintenance.

Method used

Employing a cloud-edge-device collaborative architecture, the system achieves local real-time data processing and global cloud optimization through the collaborative work of the data acquisition layer, edge computing layer, and cloud decision layer. The data acquisition layer collects data in real time and adjusts the sampling frequency according to edge commands; the edge computing layer performs local real-time processing and model training; and the cloud decision layer optimizes the model and provides decision support. Federated learning and multi-level anomaly detection models are used to improve detection accuracy and response speed.

Benefits of technology

It significantly improves the accuracy and response speed of anomaly detection, reduces network bandwidth pressure, improves system reliability and operation and maintenance efficiency, realizes the identification of complex anomaly patterns and the predictability of hidden faults, and enhances the system's adaptability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121417469A_ABST
    Figure CN121417469A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution room anomaly detection system based on cloud side-end cooperation, and belongs to the technical field of intelligent power grids. In order to solve the problems of high network bandwidth pressure, insufficient edge computing capability, low anomaly detection accuracy, difficulty in multi-source data fusion and the like caused by the adoption of an end-cloud direct connection architecture in an existing power distribution room monitoring system, the system comprises: a data acquisition layer configured with various heterogeneous sensors to acquire operating parameters and environmental data in real time; the edge storage and calculation layer carries out local real-time processing, anomaly detection, model training and visual display, an anomaly detection module of the edge storage and calculation layer carries out research and judgment on real-time data to generate early warning information, and a prediction and detection linkage module monitors an anomaly probability trend and adjusts a sampling frequency; the edge gateway realizes protocol conversion and data forwarding; and the cloud decision-making layer aggregates multi-edge node data, optimizes a global model through federal learning, and issues and updates a local model. The system is used for improving the accuracy, real-time performance and reliability of anomaly detection of the power distribution room, reducing the operation and maintenance cost and realizing intelligent operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart grid technology. More specifically, this invention relates to a power distribution room anomaly detection system based on cloud-edge-device collaboration. Background Technology

[0002] As a core facility of smart power distribution networks, low-voltage intelligent substations face numerous technical challenges in their intelligent operation and maintenance. Current technical solutions exhibit the following significant problems and shortcomings: Low data transmission efficiency: Existing systems generally adopt a direct "end-to-cloud" architecture, requiring all raw monitoring data to be directly uploaded to the cloud for processing. Due to the large volume and high sampling frequency of monitoring data in power distribution rooms, network bandwidth consumption is severe, especially in wireless transmission scenarios where data packet loss and transmission delays frequently occur, seriously affecting the real-time performance of monitoring.

[0003] Insufficient edge computing capabilities: Existing edge devices mostly use low-performance processors, possessing only simple data collection and forwarding functions. Anomaly detection algorithms rely on cloud execution, making local real-time analysis impossible. When network connectivity is unstable, the system will lose its anomaly detection capabilities, posing a security risk.

[0004] Low anomaly detection accuracy: Traditional systems mainly rely on threshold alarms and rule-based judgments, which have limited ability to identify complex anomaly patterns. In actual operation, the false alarm rate is relatively high, leading to fatigue of maintenance personnel and a decrease in system reliability.

[0005] Multi-source data fusion is difficult: power distribution room monitoring involves various heterogeneous sensor data such as temperature, humidity, switch status, and high-frequency sampling. Existing systems lack effective data correlation analysis mechanisms, making it difficult to mine potential fault characteristics from multi-dimensional data.

[0006] High operation and maintenance response delay: From the occurrence of an anomaly to the intervention of operation and maintenance personnel, the existing system has a high response delay, which is difficult to meet the requirements of critical power distribution rooms for rapid fault handling. This may expand the scope of the fault and affect key business operations such as operation and maintenance, fault repair and customer service.

[0007] While some improvement solutions have been proposed, they have failed to fundamentally address key issues such as cloud-edge collaboration, data fusion, and real-time processing. Therefore, there is an urgent need to develop a new cloud-edge-device collaborative anomaly detection platform to improve the intelligent operation and maintenance level of power distribution rooms. Summary of the Invention

[0008] This invention provides a power distribution room anomaly detection system based on cloud-edge-device collaboration. Through the collaborative architecture of the data acquisition layer, edge computing layer and cloud decision layer, it can realize local real-time processing and cloud-based global optimization of power distribution room monitoring data, effectively improving the accuracy, response speed and system reliability of anomaly detection, while reducing network bandwidth pressure and operation and maintenance costs.

[0009] To achieve these objectives and other advantages of the present invention, a power distribution room anomaly detection system based on cloud-edge-device collaboration is provided, comprising: The data acquisition layer is equipped with various heterogeneous sensors for real-time acquisition of operating parameters and environmental data of electrical equipment in the power distribution room; the data acquisition layer receives control commands from the edge storage and computing layer and dynamically adjusts its own data sampling frequency. An edge computing layer, which is communicatively connected to the data acquisition layer, is used for local real-time processing, anomaly detection, model training, and visualization of received sensor data. The edge computing layer has a built-in anomaly detection module and a prediction detection linkage module. The anomaly detection module is equipped with an anomaly detection model, which is used to analyze real-time monitoring data and generate early warning information. The prediction detection linkage module is used to monitor the changing trend of anomaly probability and generate control commands based on the changing trend, which are sent to the data acquisition layer through the edge gateway to dynamically adjust the sampling frequency. An edge gateway, integrated into the edge computing layer, is used to realize protocol conversion and data forwarding between the data acquisition layer and the edge computing layer, and to maintain bidirectional communication with the cloud decision-making layer; The cloud-based decision layer, connected to the edge computing layer network, aggregates data from multiple edge nodes for centralized storage, model optimization, and decision support. The cloud-based decision layer has a built-in federated model training module, which periodically aggregates the local model parameters trained by each edge computing layer, generates an optimized global model through a federated learning algorithm, and distributes the global model parameters to each edge computing layer to update its local anomaly detection model.

[0010] Preferably, the edge computing layer further includes an early warning push module, a visualization module, and an edge data storage and management module; the early warning push module is configured to classify early warning information into multi-level alarm levels according to preset rules and send it to the visualization module and the edge data storage and management module; The prediction and detection linkage module built into the edge computing layer is configured to continuously monitor the slope of the change trend of the anomaly probability. When the slope of the change trend exceeds a first threshold and the anomaly probability continues to exceed a second threshold, an instruction to increase the sampling frequency of the data acquisition layer is generated. When the anomaly probability drops below a third threshold and remains stable for more than a preset time, an instruction to decrease the sampling frequency is generated. The third threshold is lower than the second threshold.

[0011] Preferably, the anomaly detection model is a comprehensive anomaly detection model, which includes a sensor fault self-verification unit, a dynamic business threshold determination unit, a machine learning model determination unit, and a collaborative adjudication unit connected in sequence. The collaborative adjudication unit specifically includes: First, the threshold comparison result from the dynamic service threshold determination unit and the preliminary anomaly confidence level output from the machine learning model determination unit are received. When the threshold comparison result indicates an anomaly and the initial anomaly confidence level is lower than a preset confidence threshold, the collaborative decision unit accesses a local historical event database to query whether it is normal operation rather than a real fault under similar equipment operating conditions. When the threshold comparison result does not indicate an anomaly but the preliminary anomaly confidence level is higher than the preset confidence threshold, the collaborative adjudication unit initiates the interpretability analysis program of the machine learning model judgment unit to obtain the key feature sequence that leads to the high confidence level, and matches the feature sequence with a known latent fault feature library. The collaborative adjudication unit outputs the final anomaly determination conclusion based on the above query or matching results; The comprehensive anomaly detection model is also configured with a parameter adaptive engine, which is used to adjust the decision tree model branch weights in the dynamic business threshold determination unit or the loss function of the machine learning model determination unit in reverse based on the confirmation result after the operation and maintenance personnel confirm the false alarm or missed alarm through the visualization module.

[0012] Preferably, the federated model training module built into the cloud decision layer specifically performs the following process: The cloud server initializes a global anomaly detection model and sends its parameters to each edge storage and computing layer. Each edge computing layer uses local data to train the global anomaly detection model, and then encrypts the model gradient parameters obtained after training using an encryption algorithm before uploading them to the cloud; The cloud server aggregates encrypted model gradient parameters from multiple edge computing layers and uses a federated averaging algorithm to perform weighted averaging calculations to generate parameters for a new generation of global models. The parameters of the new generation global model are distributed to all edge computing layers to update their local anomaly detection models.

[0013] Preferably, the edge computing layer is configured to automatically trigger the data upload process when the overall anomaly confidence level output by its anomaly detection module is lower than a preset threshold. The data upload process includes: encapsulating the multi-source sensor monitoring data and associated local feature vectors within a set time window before and after the trigger time into a data packet and uploading it to the cloud decision layer through an edge gateway; The cloud-based decision layer is configured to perform the following operations upon receiving the data packet: a) invoke its analysis module to perform correlation analysis and deep learning model inference on the multi-source data within the data packet; b) perform pattern matching between the analysis results and the historical anomaly feature database from other edge nodes; c) generate a feedback data packet containing diagnostic conclusions and treatment instructions. The feedback data packet is simultaneously sent via the network to the source edge node and its adjacent edge nodes determined according to the power grid topology; the data storage and management module of the edge storage and computing layer is configured to receive and store the feedback data packet to update the local early warning rules.

[0014] Preferably, the edge computing layer further includes an anomaly detection model training module, which is configured to perform local augmentation. The learning process involves optimizing the anomaly detection model using incremental updates based on historical and real-time monitoring data from the edge data storage and management module. The hyperparameters of the anomaly detection model are dynamically adjusted through feedback from operations and maintenance personnel or cloud commands. The edge data storage and management module in the edge computing layer is configured to receive and store monitoring data from the data acquisition layer, early warning information sent by the early warning push module, and abnormal processing results and hyperparameters of the abnormal detection model returned by the visualization module. It also periodically uploads abnormal early warning information and processing results to the cloud decision layer and accesses multi-source heterogeneous sensor data through wired or wireless protocols. The edge data storage and management module has a built-in time-series data alignment and cleaning engine, which is configured to perform time synchronization, outlier filtering and format standardization on multi-source sensor data from different sampling frequencies and protocols to form a regular time-series dataset.

[0015] Preferably, the visualization module in the edge computing layer provides a graphical human-machine interface through an integrated high-definition industrial display panel; The graphical human-computer interaction interface includes at least: The real-time data cockpit dynamically displays the current values ​​and change curves of key equipment operating parameters and environmental data. The abnormal alarm control panel presents alarm information in a priority sorted list, with each piece of information associated with the device location, alarm type, and initial handling guidelines. An interactive feedback window is used to receive the anomaly handling process and result tags entered by the operation and maintenance personnel. The result tags will be sent back to the edge data storage and management module as training feedback data.

[0016] The edge computing layer also includes a statistical analysis module, which is connected to the visualization module. The statistical analysis module is configured to generate multi-dimensional statistical analysis reports at preset intervals, including pie charts of abnormal event type distribution and device health trend charts. The reports are displayed through the statistical information section of the visualization module. The statistical analysis module has a built-in root cause correlation analyzer that performs correlation analysis on concurrent abnormal events based on the Apriori algorithm and outputs the most relevant root cause inference results to the visualization module.

[0017] Preferably, the cloud-based decision-making layer further includes: a cloud-based data storage and management module, a cockpit module, a predictive maintenance module, and a power grid dispatching module; The cloud data storage and management module constructs a feature repository across edge nodes for archiving and indexing feature vectors and alarm events uploaded from the edge side; The cockpit module provides an overview of the regional power grid's operational status and is equipped with a global control interface for the hyperparameters of the anomaly detection model. The predictive maintenance module uses an LSTM network to predict the remaining effective life of the equipment based on historical data and real-time stress, and generates a pre-maintenance plan. The power grid dispatching module receives the equipment pre-maintenance time window and availability status information output by the predictive maintenance module, and calculates the regional load dispatching scheme using a linear programming algorithm as the core, with the minimum system operating cost as the objective function and line capacity, node voltage deviation and equipment availability status as constraints.

[0018] Preferably, the machine learning model decision unit includes a cross-modal dynamic graph fusion network, which comprises, in sequence: The multimodal feature extraction branch is configured to use a temporal convolutional network to process electrical quantity time-series data and a gated recurrent unit to process non-electrical quantity time-series data. The graph structure learning layer takes the feature vectors output by the multimodal feature extraction branch as its input, dynamically infers the correlation strength between the feature vectors of each monitoring point through a trainable probabilistic generation model, and outputs a sparse adjacency matrix to represent the device association graph at the current moment. The spatiotemporal graph attention layer takes the feature vector and the adjacency matrix as input. This layer includes a spatial attention sublayer and a temporal attention sublayer. The spatial attention sublayer calculates the attention coefficient between each node in the graph and its neighboring nodes. The temporal attention sublayer calculates the attention coefficient between each node at the current time point and historical time points within a sliding time window, and fuses the features. The normalized flow density estimator estimates the probability density of node features after fusion by the spatiotemporal graph attention layer and outputs the anomaly confidence level.

[0019] Preferably, it also includes a full-link trusted execution and dynamic resource scheduling mechanism, which specifically includes: A hardware-level trusted execution environment is integrated into the processors of the edge computing layer and the cloud decision layer to store anomaly detection model parameters and federated learning gradients; When transmitting data between the data acquisition layer and the edge computing layer, the AES-128 algorithm is used for encryption; when transmitting data between the edge computing layer and the cloud decision layer, an elliptic curve-based asymmetric encryption algorithm is used, and the key is rotated every 24 hours. The edge computing layer also integrates a dynamic resource scheduler, which monitors the computing load and the comprehensive anomaly confidence level output by the anomaly detection model in real time. When the overall anomaly confidence level is lower than the first confidence level threshold, the dynamic resource scheduler allocates 70% of the computing resources to the anomaly detection model training task and triggers a CPU frequency reduction instruction to enter power saving mode. When the overall confidence level exceeds the second confidence level threshold which is higher than the first confidence level threshold, the dynamic resource scheduler immediately interrupts the anomaly detection model training task, reallocates 90% of the computing resources to the anomaly detection module, and removes the CPU frequency reduction instruction.

[0020] The present invention has at least the following beneficial effects: First, this invention achieves efficient collaboration between data acquisition, local processing, and cloud optimization by constructing a collaborative architecture comprising a data acquisition layer, an edge computing layer, an edge gateway, and a cloud decision-making layer. The data acquisition layer can dynamically adjust the sampling frequency based on edge commands, significantly reducing network bandwidth pressure; the edge computing layer possesses local real-time anomaly detection and model training capabilities, reducing reliance on the cloud and improving system response speed and robustness; the edge gateway performs protocol conversion and data forwarding, ensuring smooth network connectivity; and the cloud decision-making layer aggregates knowledge from multiple nodes through federated learning, continuously optimizing the global model and improving anomaly detection accuracy and generalization ability. The overall system, while ensuring real-time performance, improves data utilization efficiency and operational intelligence.

[0021] Secondly, this invention utilizes an early warning push module to implement multi-level alarm classification, enabling maintenance personnel to quickly identify the urgency of anomalies and improve handling efficiency. The predictive detection linkage module intelligently adjusts the sampling frequency based on the slope of the anomaly probability change trend, increasing data detail capture when anomalies rise and reducing sampling to conserve resources under normal conditions, achieving adaptive data acquisition. This mechanism not only optimizes edge computing resource allocation but also improves the system's sensitivity to anomaly changes and response accuracy, reducing the risk of false alarms and missed alarms.

[0022] Third, the comprehensive anomaly detection model of this invention significantly improves the accuracy and reliability of anomaly detection through multi-level processing including sensor fault self-verification, dynamic threshold determination, machine learning model determination, and collaborative adjudication unit. When the threshold and model results are inconsistent, the collaborative adjudication unit performs secondary analysis by combining historical event databases and latent fault feature databases, effectively distinguishing between real faults and normal operations, and reducing the false alarm rate. The parameter adaptive engine dynamically adjusts model parameters based on operational feedback, enabling the system to continuously learn and optimize, adapt to different operating conditions, and improve the system's intelligence level.

[0023] Fourth, the federated model training module of this invention achieves global model collaborative optimization while protecting the data privacy of each edge node through encrypted gradient uploading and a federated averaging algorithm. This method avoids the security risks associated with centralized data transmission and enhances the model's generalization ability by utilizing distributed data features. The cloud periodically aggregates and distributes model parameters, ensuring that each edge node always maintains the latest detection capabilities, solving the data silo problem, and enhancing the overall detection performance and adaptability of the system.

[0024] Fifth, by automatically triggering the low-confidence data upload process, this invention enables the system to perform in-depth cloud-based analysis of potential anomalies, uncovering complex features that are difficult to identify at the edge. Feedback data packets are simultaneously sent to the source node and adjacent nodes, achieving distributed sharing of anomaly knowledge and updating of early warning rules, thus improving regional collaborative diagnostic capabilities. This mechanism effectively avoids the spread of local anomalies and enhances the system's predictability and efficiency in handling latent faults.

[0025] Sixth, this invention standardizes multi-source heterogeneous data through a time-series data alignment and cleaning engine, improving data quality and model training effectiveness. The anomaly detection model training module supports incremental learning and, combined with dynamic hyperparameter adjustment, achieves efficient model iteration and personalized optimization. The edge data storage and management module uniformly manages monitoring data, early warning information, and feedback results, providing complete data support for local processing and improving the system's data governance capabilities and operational efficiency.

[0026] Seventh, the visualization module of this invention intuitively displays equipment status, alarm information, and statistical reports through a graphical human-computer interaction interface, greatly improving the ease of operation and decision-making speed for maintenance personnel. The statistical analysis module performs root cause correlation analysis based on the Apriori algorithm, helping to quickly locate the source of anomalies and reduce troubleshooting time. The interactive feedback window collects maintenance processing results, forming a closed-loop learning mechanism to continuously optimize system performance.

[0027] Eighth, this invention constructs a cross-node feature warehouse through a cloud-based data storage and management module, supporting big data archiving and rapid retrieval. The predictive maintenance module accurately predicts equipment lifespan based on LSTM networks, enabling a shift from scheduled maintenance to predictive maintenance and reducing maintenance costs. The power grid dispatching module optimizes load allocation using linear programming algorithms, minimizing operating costs while ensuring power grid safety, thereby improving the economy and reliability of the regional power grid.

[0028] Ninth, the cross-modal dynamic graph fusion network of this invention effectively integrates the spatiotemporal correlation features of electrical and non-electrical quantity data through multimodal feature extraction, graph structure learning, and spatiotemporal attention mechanisms, thereby improving the ability to identify complex anomaly patterns. The normalized flow density estimator outputs anomaly confidence, providing a reliable basis for collaborative decision-making and significantly improving detection accuracy and sensitivity to latent faults.

[0029] Tenth, the end-to-end trusted execution environment and multi-layer encryption mechanism of this invention ensure the security of data transmission and storage, preventing the leakage of sensitive information. The dynamic resource scheduler adjusts the allocation of computing resources in real time based on the anomaly confidence level, optimizing energy efficiency and extending equipment lifespan while ensuring real-time detection. This system achieves comprehensive optimization in terms of security, energy efficiency, and reliability, making it suitable for demanding industrial scenarios.

[0030] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the relational structure of the power distribution room anomaly detection system based on cloud-edge-device collaboration according to the present invention. Detailed Implementation

[0032] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.

[0033] It should be understood that terms such as “having,” “comprising,” and “including” as used herein do not exclude the presence or addition of one or more other elements or combinations thereof.

[0034] like Figure 1 As shown, this invention provides a power distribution room anomaly detection system based on cloud-edge-device collaboration. The system achieves efficient collaboration between data acquisition, local processing, and cloud optimization through a layered architecture. The system consists of four main parts: a data acquisition layer, an edge computing layer, an edge gateway, and a cloud decision layer.

[0035] The data acquisition layer, as the foundation of the system, is equipped with various heterogeneous sensors, including but not limited to temperature sensors, humidity sensors, current transformers, voltage sensors, vibration sensors, ultrasonic / ground wave combined partial discharge sensors, and gas sensors. These sensors are deployed at different key locations in the power distribution room to collect real-time operating parameters of electrical equipment and environmental data. For example, voltage and current sensors are installed at key nodes of the distribution cabinet to monitor parameters such as current, voltage, and power factor; temperature sensors are arranged around equipment such as transformers and switchgear, as well as at key environmental locations, to obtain information on temperature changes in the equipment and environment. The data acquisition layer includes a data acquisition module and a data upload module. The data acquisition module connects various electrical equipment and environmental sensors in the power distribution room, including sensing terminals such as temperature sensors, humidity sensors, ultrasonic / ground wave combined partial discharge sensors, and electrical energy thermometers, to achieve comprehensive monitoring of the power distribution room's operating status. The data upload module preprocesses the acquired data and uploads it to the edge computing layer. After high-frequency data acquisition, data interaction, and simple data filtering and analysis, the sensors transmit the acquired data to the edge gateway of the edge computing layer via wired communication methods (such as RS485 bus and Ethernet) and wireless communication methods (such as LoRa and WiFi). For example, the humidity sensor and gas sensor transmit the acquired current, voltage, and other data to the edge gateway of the edge computing layer via RS485 bus and DI interface. The voltage, current, and temperature sensors, partial discharge sensors, and wireless temperature sensors send device temperature and ambient temperature and humidity data to a non-standard protocol support plug-in via LoRa wireless communication. The plug-in then transmits the acquired data to the edge gateway of the edge computing layer via RS485 bus.

[0036] The data acquisition layer is not only responsible for acquiring raw data but also has the ability to dynamically adjust the sampling frequency. Its principle lies in receiving control commands from the edge computing layer and intelligently adjusting the sampling rate based on the changing trends of anomaly probability: when the system detects an increased anomaly risk, the sampling frequency can be increased to tens or even hundreds of times per second to capture more detailed data features; under normal operating conditions, the sampling frequency can be reduced to a few times per second to save energy and network resources. For example, the sampling frequency can be dynamically adjusted between 1Hz and 100Hz, with specific thresholds set according to the actual application scenario. During operation, sensors continuously collect data and transmit it to the edge computing layer via wired or wireless protocols (such as RS-485, ZigBee, or 4G / 5G). This adaptive sampling mechanism effectively avoids data overload while ensuring that critical anomaly events are not missed, providing a high-quality data source for subsequent processing.

[0037] The edge computing layer includes edge intelligent devices and edge gateways. The edge intelligent devices are deployed close to the data acquisition layer and include a high-performance processor, a network cable interface, and a storage unit of a certain capacity. The processor can be a Central Processing Unit (CPU), and in this example, it is a low-power, high-computing-power ARM architecture domestic chip with a rated power of 48W, providing a maximum CPU computing power of 1.6 TFLOPS. It supports the deployment, training, and online real-time inference of time-series data anomaly detection algorithms and can be customized and extended to other AI algorithms, achieving real-time data analysis and generating early warning reports within one second. The network cable interface in this example supports multiple wired / wireless protocols for accessing multi-source heterogeneous sensor data. An external industrial display screen runs an Android system, enabling visual interaction through an Android application. The edge computing layer communicates directly with the data acquisition layer, undertaking multiple tasks such as local real-time processing, anomaly detection, anomaly detection model training, and visualization. The edge computing layer incorporates an anomaly detection module and a predictive detection linkage module, forming a closed-loop processing flow. The anomaly detection module deploys an advanced anomaly detection model, which can quickly analyze real-time monitoring data and generate early warning information. The underlying principle typically involves machine learning algorithms, such as deep learning models based on time-series data, which identify anomalies by analyzing data patterns. The prediction and detection linkage module continuously monitors the changing trends of anomaly probabilities, particularly the trend slope. When the slope exceeds a preset threshold and the anomaly probability remains high, the module generates an instruction to increase the sampling frequency; conversely, when the anomaly probability decreases and stabilizes, the instruction adjusts to decrease the sampling frequency. For example, the first threshold can be set to a slope above 0.5, the second threshold to an anomaly probability of 70%, and the third threshold to 30%, and these values ​​can be adjusted according to the actual scenario. In implementation, after data flows into the edge computing layer, it undergoes preprocessing (such as filtering and normalization), and then the anomaly detection model outputs a confidence score, which the prediction module uses to dynamically adjust the data acquisition layer. This collaboration allows the system to complete most of the analysis at the edge, reducing reliance on the cloud, improving response speed, and extending device lifespan through resource optimization.

[0038] The edge gateway, integrated into the edge computing layer, acts as a hub for protocol conversion and data forwarding. Its core function is to achieve seamless connectivity between the data acquisition layer and the edge computing layer, and to maintain bidirectional communication with the cloud decision-making layer. The gateway supports multiple industrial protocols (such as Modbus, OPC UA, or MQTT), enabling it to convert data from heterogeneous sensors in different formats into a unified standard, ensuring data compatibility. During data transmission, the gateway is also responsible for encryption and compression to ensure security and efficiency. For example, AES-128 encryption can be used in local communication, while asymmetric encryption is used for cloud-edge communication. In operation, the gateway monitors the data stream in real time, performs protocol parsing and forwarding, and simultaneously receives cloud commands and distributes them to the edge layer. This design solves the complexity of multi-source device access, ensuring stability and low latency throughout the entire communication chain.

[0039] The cloud-based decision-making layer connects to the edge computing layer via a network, focusing on aggregating data from multiple edge nodes for centralized storage, model optimization, and decision support. This layer incorporates a federated model training module based on federated learning technology, enabling collaborative model optimization while protecting data privacy. In practice, the cloud server first initializes a global anomaly detection model and distributes its parameters to each edge node. Each edge node trains its model using local data, uploading only encrypted gradient parameters. The cloud then aggregates these parameters, generates a new generation of global model using a federated averaging algorithm, and distributes it to update the edge models. This periodic aggregation and distribution process, for example, every 24 hours, allows the system to leverage distributed data to improve model generalization capabilities while avoiding the security risks associated with centralized data. The cloud also provides centralized storage and advanced analytics, such as trend prediction based on historical data, to support operational decisions.

[0040] Traditional power distribution room monitoring systems often employ a direct "edge-to-cloud" architecture, resulting in high network bandwidth pressure, high response latency, and insufficient edge computing capabilities, making real-time anomaly detection difficult. This invention, through a cloud-edge-device collaborative architecture, distributes data processing tasks to the edge layer, significantly reducing network transmission load while improving anomaly detection response speed through local real-time processing. Compared to existing technologies, the system can more accurately identify complex anomaly patterns, significantly reduces false alarm rates, and possesses adaptive optimization capabilities to adapt to different operating conditions. Overall, the system demonstrates substantial improvements in reliability, real-time performance, and intelligence, providing an efficient solution for power distribution room operation and maintenance.

[0041] In one specific implementation, the edge computing layer further integrates an early warning push module, a visualization module, and an edge data storage and management module. These modules together constitute the core of local processing and interaction. The early warning push module receives early warning information generated by the anomaly detection module and classifies it into multi-level alarm levels according to preset rules. For example, it classifies anomalies into different levels such as emergency, important, and general based on their severity. The emergency level can correspond to high-risk events such as equipment overheating or sudden current surges, the important level can involve humidity anomalies or slight fluctuations, and the general level is used to alert routine maintenance personnel. Alarms can also be set to Level 1, Level 2, and Level 3. For example, sensor anomalies are Level 1 alarms, temperature anomalies are Level 2 alarms, and network anomalies are Level 3 alarms. The alarm information is labeled with its alarm level and sent to the visualization module and the edge data storage and management module. This early warning push module dynamically adjusts the alarm threshold through a built-in logic judgment unit, combined with a historical event database and real-time data characteristics, ensuring the accuracy and adaptability of alarm classification. The classified early warning information is sent in parallel to the visualization module and the edge data storage and management module, achieving real-time display and persistent storage of the information. The visualization module typically provides a graphical interface through an integrated high-definition industrial display panel, presenting alarm details, equipment status, and trend curves in an intuitive way, facilitating rapid anomaly identification by maintenance personnel. The edge data storage and management module employs time-series databases or distributed storage technology to uniformly manage early warning information, raw monitoring data, and processing results, supporting rapid querying and backtracking. These modules work collaboratively through an internal data bus: the early warning push module ensures the structured flow of alarm information, the visualization module provides a human-machine interaction interface, and the storage module provides support for basic data, thus forming a closed-loop processing chain at the edge, improving the completeness and efficiency of local decision-making.

[0042] The prediction and detection linkage module, serving as the core of intelligent control in the edge computing layer, continuously monitors the slope of the anomaly probability trend to dynamically adjust the sampling frequency of the data acquisition layer. This module quantifies the speed and direction of anomaly changes by calculating the derivative of the anomaly probability or the slope of a linear fit within a sliding window in real time. For example, the slope is calculated by differentiating recent anomaly probability values, and the sliding window length can be set to 5 to 10 sampling points to balance response speed and stability. When the slope exceeds a first threshold (e.g., a value between 0.5 and 1.0, indicating a rapid increase in anomaly probability) and the anomaly probability itself consistently exceeds a second threshold (e.g., 60% to 80%, indicating a relatively certain anomaly), the module generates an instruction to increase the sampling frequency of the data acquisition layer, for example, from the usual 1Hz to 50Hz or higher, to capture more refined data features for in-depth analysis. Conversely, when the anomaly probability drops below the third threshold (e.g., 20% to 40%, below the second threshold to provide hysteresis) and remains stable for more than a preset time (e.g., 3 to 5 minutes to avoid frequent fluctuations), the prediction and detection linkage module will trigger an instruction to reduce the sampling frequency, restoring the sampling rate to the basic level to reduce resource consumption. This process is implemented based on a state machine mechanism: the prediction and detection linkage module is initially in a monitoring state; once the slope and probability conditions are met, it switches to an instruction generation state and sends control signals to the data acquisition layer through the edge gateway. This adaptive adjustment mechanism effectively couples anomaly risk with data acquisition intensity, enhancing monitoring granularity during high-risk periods and optimizing energy efficiency during stable periods, reflecting the real-time performance and economy of edge computing.

[0043] In actual operation, the threshold settings of the predictive detection linkage module can be dynamically adjusted according to the actual working conditions of the power distribution room. For example, the first threshold can be set to 0.7 based on the historical failure rate of the equipment, the second threshold to 75% based on maintenance experience, and the third threshold to 25%, with a preset duration of 4 minutes. These values ​​can be optimized through the cloud decision-making layer or the local configuration interface. The predictive detection linkage module integrates lightweight time series analysis algorithms, such as exponentially weighted moving average or Kalman filtering, to smooth the anomaly probability sequence and reduce noise interference. After the command is generated, the edge gateway sends the sampling frequency adjustment command to the sensor controller of the data acquisition layer through standard industrial protocols (such as Modbus TCP or MQTT) to achieve seamless linkage. In addition, the predictive detection linkage module also records each adjustment event to the edge data storage and management module for subsequent performance evaluation and parameter tuning. The entire process runs in an event-driven manner: the real-time update of the anomaly probability triggers the slope calculation, and the threshold comparison logic is executed periodically (e.g., once per second) to ensure timely and reliable system response. Through this refined dynamic sampling strategy, the system significantly reduces the computational and communication overhead on the edge side while ensuring the sensitivity of anomaly detection, thus extending the service life of the equipment.

[0044] In one specific implementation, the anomaly detection model is a comprehensive anomaly detection model. This model constructs a multi-stage, cascaded intelligent judgment pipeline, with its entry point being a sensor fault self-verification unit. This unit first verifies the source reliability of the input monitoring data. Its principle is based on designing self-checking rules based on the sensor's physical characteristics or historical data patterns. For example, for a temperature sensor, preliminary diagnosis can be made by checking whether the reading is within a reasonable physical range (e.g., between -10℃ and 150℃), whether there are contradictions between adjacent sensor values, or whether the signal fluctuation is abnormally smooth (potentially indicating a sensor malfunction). The sensor fault self-verification unit outputs a data quality identifier, such as "reliable," "suspected," or "failed," and appends this identifier to the data stream for transmission to subsequent units. A dynamic business threshold determination unit follows. Its core feature is that the threshold is not fixed but dynamically adjusted according to the equipment's operating conditions (e.g., load rate, ambient temperature, time period). For example, when a transformer operates under high load in summer, its temperature alarm threshold automatically increases from the usual 85℃ to 90℃ to accommodate the additional temperature rise generated during normal operation. This dynamic business threshold determination unit can maintain a threshold mapping table based on a decision tree or rule engine. By querying real-time operating condition variables, it obtains the corresponding dynamic threshold range and compares the monitored data with these thresholds to generate a preliminary binary determination result (normal or abnormal). This design enables the system to adapt to normal fluctuations in equipment under different conditions, reducing false alarms caused by changes in operating conditions.

[0045] As the core decision-making component of the integrated anomaly detection model, the collaborative decision-making unit receives threshold comparison results from the dynamic business threshold judgment unit and preliminary anomaly confidence levels (a probability value between 0% and 100%) from the machine learning model judgment unit. Its decision-making logic addresses two typical inconsistency scenarios. When the threshold comparison result indicates an anomaly, such as current exceeding the dynamic upper limit, but the preliminary confidence level output by the machine learning model is low (e.g., below a preset confidence threshold, which can be set at a value between 60% and 80%, such as 70%, depending on the system's tolerance for false alarms), the collaborative decision-making unit will not immediately classify it as a fault. It accesses a local historical event database stored on the edge side, which records a large number of past device operation logs and corresponding system states. The collaborative decision-making unit queries whether, under similar combinations of power, temperature, and load conditions, similar threshold alarms, rather than actual device faults, have been triggered in the past due to normal operations (such as planned capacitor switching or equipment startup inrush current). If the query results indicate a high probability of normal operation, the collaborative adjudication unit classifies the event as "operational interference," thereby suppressing the alarm. Conversely, in another scenario, if the threshold comparison results do not show any anomalies, but the machine learning model gives a high confidence level (higher than the preset confidence threshold), the collaborative adjudication unit will activate the interpretability analysis program built into the machine learning model, such as SHAP or LIME methods, to extract which key feature sequences (such as a sustained, slight increase in a specific harmonic component of a phase current) led to the model's high confidence judgment. Subsequently, it matches this feature sequence with a known latent fault feature library, which may contain fault mode features that are difficult to capture by simple thresholds, such as "early winding loosening" or "slow degradation of insulation materials." If the match is successful, the adjudication unit adopts the machine learning model's judgment and outputs an anomaly conclusion. Through this cross-validation and knowledge base query mechanism, the collaborative adjudication unit significantly improves the accuracy of anomaly judgment in complex scenarios.

[0046] The parameter adaptive engine is a key component for the comprehensive anomaly detection model to achieve continuous self-optimization, forming a closed loop from operational feedback to model parameters. When operations personnel confirm through the visualization module interface that an alarm is a false alarm (the system alarms but there is no actual fault) or a missed alarm (a fault actually occurs but the system does not alarm), the confirmation result, along with relevant data samples, is recorded and sent back to the edge data storage and management module. The parameter adaptive engine is then triggered, and it adjusts the internal parameters of the model in reverse based on this feedback information. For example, for a confirmed false alarm, the engine will reduce the branch weight of a decision tree rule in the dynamic business threshold judgment unit that caused the false alarm, thus weakening the influence of that path in future decisions; or, it will adjust specific weight parameters in the loss function used by the machine learning model's judgment unit, increasing the penalty for the feature pattern that caused the false alarm, making the model more inclined to ignore such patterns in subsequent training. The adjustment magnitude can be controlled by the learning rate parameter, for example, fine-tuning by 0.1% to 1% each time, to avoid overfitting to a single feedback. This process is essentially a small-batch supervised learning, which uses high-value feedback from maintenance personnel as labels to enable the anomaly detection model to continuously adapt to the specific operating characteristics of the power distribution room and emerging fault modes, thus becoming more and more accurate with use.

[0047] In one specific implementation, the federated model training module built into the cloud decision layer first performs the initialization process of the global anomaly detection model by the cloud server. This global anomaly detection model is typically a pre-designed machine learning model structure, such as a time-series prediction model based on deep neural networks or a support vector machine model, whose parameters include trainable variables such as weight matrices and bias vectors. During initialization, the cloud server sets initial values ​​for the parameters based on historical datasets or standard configurations. These initial values ​​can be achieved using random initialization methods such as Xavier initialization or by loading from a pre-trained model to ensure the model possesses basic anomaly detection capabilities. After initialization, the cloud server distributes the complete parameter set of the global anomaly detection model to each edge computing layer via a secure network communication protocol, such as HTTPS based on transport layer security protocols or a dedicated virtual private network tunnel. This distribution process can be performed periodically, such as every 24 hours or weekly, with the specific frequency adjusted according to actual operational needs and network bandwidth to ensure that edge nodes always use a unified model baseline. After the parameters are distributed, the edge computing layer loads them into its local anomaly detection module as the base model for subsequent local training. This initialization step ensures that all edge nodes start learning from a consistent starting point, avoiding model bias caused by differences in initial states, and laying the framework foundation for subsequent collaborative optimization. In implementation, the cloud server can maintain a model version management system to record each parameter version issued, facilitating tracking and rollback.

[0048] After receiving the global anomaly detection model parameters, each edge computing layer trains the model using locally stored monitoring data. This data includes real-time operating parameters and environmental data collected by various heterogeneous sensors within the power distribution room, such as time-series information on temperature, humidity, current, and voltage. The training process employs supervised or semi-supervised learning, adjusting model parameters through optimization algorithms such as stochastic gradient descent or adaptive moment estimation to better adapt the global anomaly detection model to local data characteristics. After training, the edge computing layer does not upload the raw data but only extracts the gradient parameters calculated during model training. These gradient parameters reflect the direction and magnitude of the local data's influence on the model parameters. To protect data privacy and security, the gradient parameters are encrypted using encryption algorithms. Encryption methods can include symmetric encryption algorithms such as Advanced Encryption Standard 128-bit or asymmetric encryption algorithms such as the Rivest-Shamir-Adleman algorithm. Encryption keys are distributed uniformly in the cloud or managed based on public key infrastructure. The encrypted gradient parameters are uploaded to the cloud server via the edge gateway. The upload frequency can be set according to network conditions and computing resources, such as once every 4 hours or once a day, to avoid bandwidth congestion. Edge nodes can be trained using different hyperparameter configurations, such as batch size of 32 or 64 and learning rate between 0.001 and 0.01. These hyperparameters can be dynamically adjusted via cloud commands to balance training efficiency and stability. This process ensures the privacy of local data while allowing distributed knowledge to be aggregated in the cloud.

[0049] After receiving encrypted gradient parameters from multiple edge computing layers, the cloud server first decrypts the data and then aggregates it using a federated averaging algorithm. The core principle of the federated averaging algorithm is to calculate a weighted average of the gradient parameters uploaded by all edge nodes. Weights can be assigned based on the amount of data, data quality, or device importance of each node; for example, nodes with larger data volumes have higher weights to reflect their contribution. This aggregation process generates parameters for a new generation of global anomaly detection models. These parameters integrate the learning results from multiple edge nodes, improving the model's generalization ability and robustness. Subsequently, the cloud server distributes the updated parameters to all edge computing layers. Upon receiving the parameters, the edge nodes automatically replace their local model parameters, completing the model update. The distribution process can use incremental updates, transmitting only changed parameters to reduce network load, or a full update to ensure consistency. The entire aggregation and distribution process can be executed cyclically at regular intervals, such as every 48 hours or twice a week, ensuring continuous model optimization. This mechanism avoids the risks of centralized data storage through distributed collaboration while leveraging the local data from edge nodes to improve the adaptability and accuracy of the global model. During operation, the cloud server can integrate a monitoring module to track the aggregation effect and adjust the federated learning strategy.

[0050] In one specific implementation, in a cloud-edge-device collaborative power distribution room anomaly detection system, the edge computing layer plays a core role in local intelligent processing. Its anomaly detection module continuously analyzes real-time monitoring data through a comprehensive anomaly detection model and outputs a comprehensive anomaly confidence score. This comprehensive anomaly confidence score is a probability value between 0% and 100%, reflecting the degree to which the current data deviates from the normal pattern. When this comprehensive anomaly confidence score falls below a preset threshold, the system automatically triggers the data upload process. The preset threshold can be flexibly set according to actual application scenarios and operational needs; for example, it can be set between 40% and 60%. 50% is a common reference value used to define ambiguous states that are difficult for the edge layer to clearly determine but pose potential risks. The triggering mechanism is based on the state machine principle: the edge computing layer continuously monitors the confidence score output; once it detects that the confidence score value has fallen below the threshold and remains there for a period of time to avoid instantaneous fluctuations, it immediately initiates the upload logic. The data upload process involves the system extending a predefined time window forward and backward from the trigger time. The window length can be adjusted based on data characteristics and network conditions, for example, a value between 5 and 30 minutes, such as 10 minutes, to ensure sufficient contextual information is captured. Within this time window, multi-source sensor monitoring data, such as time-series readings of temperature, humidity, current, and voltage, as well as local feature vectors pre-extracted by the edge layer, such as statistical features, frequency domain features, or model intermediate layer outputs, are uniformly encapsulated into a structured data packet. The encapsulation process uses standard formats such as JSON or Protocol Buffers and adds metadata such as timestamps and device identifiers. The data packet is then uploaded to the cloud decision layer via an edge gateway integrated in the edge computing layer. The edge gateway is responsible for protocol conversion and data compression, such as converting local protocols like Modbus to cloud-friendly MQTT or HTTPS, and protecting the data packet with lightweight encryption such as AES-128 before transmission. This automated triggering and uploading process ensures that when there is high uncertainty in the edge-side judgment, key data fragments can be sent to the cloud for deeper analysis in a timely manner, avoiding the omission of potential anomalies.

[0051] Upon receiving data packets from the edge computing layer, the cloud-based decision layer immediately activates its built-in analysis module for in-depth processing. The analysis module first decapsulates the data packets, extracting multi-source sensor monitoring data and local feature vectors. Subsequently, it performs correlation analysis and deep learning model inference. Correlation analysis aims to discover inherent connections between different sensor data, such as analyzing whether temperature increases and current fluctuations occur synchronously in time, or the correlation between humidity changes and equipment insulation status. This is based on time alignment and correlation coefficient calculations, such as the Pearson correlation coefficient or distance-based similarity metrics. Deep learning model inference involves more complex machine learning models, such as using deep autoencoders for anomaly feature extraction, or employing clustering algorithms like DBSCAN to group data patterns to identify hidden anomaly patterns in multi-dimensional data. After preliminary analysis, the cloud-based decision layer performs pattern matching between the analysis results and a centrally stored historical anomaly feature database. This database aggregates historical alarm events and their corresponding feature patterns from multiple edge nodes across the network, forming a rich knowledge base. The matching process can employ similarity-based search algorithms, such as k-nearest neighbors or cosine similarity-based retrieval, to determine whether the current data pattern is highly similar to known fault patterns, such as early characteristics of winding overheating or insulation degradation trends. Based on the results of association analysis, deep learning model inference, and pattern matching, the cloud-based decision layer generates a structured feedback data packet. This packet contains diagnostic conclusions, such as "suspected local overheating, infrared inspection recommended" or "low match between features and historical hidden fault database, continuous observation required," as well as specific handling instructions, such as "adjust sampling frequency to a higher level" or "notify maintenance personnel for on-site verification." The process of generating the feedback data packet integrates a rule engine and model inference, ensuring that the diagnostic conclusions are both data-supported and actionable. The entire processing flow runs on a cloud server, utilizing its powerful computing resources to complete the parsing and decision-making of complex data packets within a short time, such as seconds to minutes, thereby providing timely and accurate guidance to the edge side.

[0052] The generated feedback data packets are simultaneously sent to two targets via network links: the source edge node serving as the data source, and adjacent edge nodes determined based on the power grid topology. The power grid topology is typically defined by a network connection graph maintained in the cloud, where nodes represent substations or monitoring stations, and edges represent electrical connections or communication links. The principle of determining adjacent nodes is based on the adjacency concept in graph theory; for example, nodes directly electrically connected to the source node or located in the same power supply corridor are considered adjacent nodes. The sending process employs reliable transmission protocols, such as TCP, to ensure complete delivery of the data packets. The feedback data packets are encapsulated in a specific message format and include priority identifiers to ensure that critical instructions are processed first. Upon receiving the feedback data packet, the data storage and management module of the edge computing layer immediately parses it and stores it in the local database. The stored content includes not only diagnostic conclusions and handling instructions but also key parameters or logical conditions used to update local early warning rules. For example, if the feedback indicates that a specific current harmonic mode is a sign of a latent fault, the local early warning rules will be updated, increasing the monitoring threshold for that harmonic mode or adding new judgment rules. The update process can be incremental, achieved by modifying the rule engine's configuration file or updating the feature weights of the machine learning model. The data storage and management module timestamps and tags each feedback message for easy subsequent querying and auditing. Simultaneously, this feedback information is also used in the local model's retraining process, forming a closed knowledge loop from the cloud to the edge. By simultaneously distributing feedback to adjacent nodes, the system achieves rapid dissemination and sharing of anomaly knowledge, enabling potential risks discovered by one node to alert surrounding nodes in advance. This builds a collaborative defense system at the regional level, enhancing the overall resilience of the power grid.

[0053] In one specific implementation, the edge computing layer further includes an anomaly detection model training module. This module is designed to execute a local incremental learning process. Its core function is to continuously optimize the anomaly detection model through incremental updates based on historical and real-time monitoring data accumulated in the edge data storage and management module. Essentially, the anomaly detection model training module is a machine learning engine embedded in the edge device. It employs incremental learning algorithms, such as online gradient descent or mini-batch learning, allowing the model to progressively adjust using newly arrived data without retraining the entire dataset. The principle is that incremental learning iteratively updates model parameters to adapt to changes in data distribution, thereby avoiding model aging and improving detection accuracy. In specific implementations, the anomaly detection model training module periodically or triggerably initiates the training process, for example, every few hours or when the amount of new data reaches a certain scale, such as after accumulating 1000 new samples, it automatically loads recent data from the edge data storage and management module, including sensor readings such as temperature, current, and voltage, and corresponding anomaly labels. During training, the anomaly detection model training module calculates the model loss function and adjusts the network weights using optimization algorithms such as stochastic gradient descent. The adjustment magnitude is controlled by the learning rate, which can be set between 0.001 and 0.01, with the specific value dynamically selected based on data volatility. Hyperparameters such as batch size and learning rate are not fixed but dynamically adjusted based on feedback from maintenance personnel or cloud commands. For example, after maintenance personnel confirm a false alarm through a visual interface, the system automatically reduces the learning rate to stabilize training, or the cloud issues commands to adjust the batch size from 32 to 64 based on global model performance to balance training speed and accuracy. This module works closely with the edge data storage and management module. Training data is retrieved from the storage module in real time, and training results are fed back to update local model parameters, forming a closed-loop optimization. The entire process, including data loading, model forward computation, loss assessment, parameter backpropagation, and model update, is completed at the edge, reducing reliance on the cloud. This design enables the system to quickly adapt to local changes in power distribution room equipment, such as seasonal changes or equipment aging, improving model robustness.

[0054] The edge data storage and management module serves as the data hub of the edge computing layer. It is configured to receive and store data streams from multiple sources, including real-time monitoring data from the data acquisition layer, early warning information sent by the early warning push module, and anomaly handling results and hyperparameters of the anomaly detection model returned by the visualization module. This edge data storage and management module is typically implemented using an embedded database or time-series database, such as InfluxDB or SQLite, possessing efficient read / write and compression capabilities for persistent storage of structured and unstructured data. Its principle is based on a data bus architecture, unifying the access and management of heterogeneous data to ensure data consistency and accessibility. In usage, the module directly accesses multi-source heterogeneous sensor data via wired or wireless protocols, such as RS-485, ModbusTCP, or 4G / 5G networks. For example, a temperature sensor transmits data at a frequency of 1Hz, while a current sensor transmits at a frequency of 10Hz. The module receives and buffers this data in real time. Simultaneously, it periodically uploads anomaly early warning information and processing results to the cloud-based decision-making layer. The upload frequency can be set to once every 30 minutes or hour, with the specific interval adjusted according to network conditions to avoid bandwidth congestion. Internally, the module implements data classification and storage. For example, monitoring data is indexed by timestamp, early warning information is archived by level, and hyperparameters are saved in configuration files. During operation, data undergoes preliminary verification upon inflow, such as checking data integrity, before being distributed to different storage areas. When the visualization module returns processing results to operations personnel, such as confirming anomalies or marking false alarms, the module immediately updates relevant data records and synchronizes this feedback to the anomaly detection model training module for model optimization. Various technical features work collaboratively through data flow: the data acquisition layer provides raw input, the early warning push module adds semantic tags, the visualization module introduces human feedback, all enriching the data content, while the storage module serves as a unified repository, supporting querying, backtracking, and data sharing, providing a complete data foundation for local processing.

[0055] The edge data storage and management module incorporates a time-series data alignment and cleaning engine. This engine is specifically responsible for preprocessing multi-source sensor data from different sampling frequencies and protocols, including time synchronization, outlier filtering, and format standardization, to form a well-organized time-series dataset. Essentially, the time-series data alignment and cleaning engine is a data preprocessing pipeline, its principles involving signal processing and data fusion techniques, aiming to solve the inconsistency problem between heterogeneous data sources. Time synchronization is achieved through interpolation or resampling algorithms. For example, for temperature data with a sampling frequency of 1Hz and current data with a sampling frequency of 10Hz, the engine aligns the low-frequency data to a high-frequency timestamp through linear interpolation, or uniformly downsamples it to a common frequency such as 5Hz. The synchronization window size can be set from several seconds to several minutes, such as a 5-second window, to ensure time alignment. Outlier filtering uses statistical methods, such as based on Z-score or IQR rules. The threshold can be set so that data points with a Z-score greater than 3 or less than -3 are considered outliers and removed. The specific threshold can be adjusted according to the historical data distribution. Format standardization converts data from different protocols, such as Modbus messages and MQTT messages, into a unified structured format, such as JSON or CSV, eliminating protocol differences. During implementation, the engine monitors incoming data in real time, first parsing the protocol header to extract the payload, then performing timestamp correction, followed by applying filtering rules to remove noise, and finally outputting well-organized data for subsequent modules. For example, when sensor data arrives, the engine checks for timestamp deviation; if the deviation exceeds a tolerance value such as 100 milliseconds, adjustments are made. The filtering stage removes suddenly changing readings; for example, a temperature change exceeding 20 degrees Celsius is considered an anomaly. The engine integrates with the storage module, directly storing the cleaned data into the database to ensure data quality. These combined technical features resolve time-series inconsistencies, improve data cleanliness through filtering, and enhance compatibility through standardization, collectively forming a high-quality time-series dataset that provides reliable input for anomaly detection models.

[0056] In one specific implementation, the visualization module, as a key component of the edge computing layer, provides maintenance personnel with an intuitive graphical human-machine interface through an integrated high-definition industrial display panel. This display panel typically employs an industrial-grade design, featuring high resolution, wide temperature adaptability, and anti-interference capabilities. For example, the screen size can be between 10 and 15 inches, with a resolution of 1920x1080 or higher to ensure clear readability of information in complex environments such as power distribution rooms. The core principle of the visualization module is to transform multi-source heterogeneous data into visual elements and dynamically update the interface using graphics rendering technology, helping maintenance personnel quickly perceive the system status. The graphical human-machine interface first includes a real-time data dashboard, which continuously and dynamically displays the current values ​​and historical change curves of key equipment operating parameters and environmental data. For example, parameters such as current, voltage, temperature, and humidity are presented in the form of digital meters and trend charts. The update frequency can be adjusted according to the data source, such as updating once per second or every few seconds. Curve plotting uses a sliding time window method, with the window length set between 5 and 30 minutes to balance real-time performance and historical context. Meanwhile, the anomaly alarm control panel organizes alarm information using a priority-based sorting list. Priorities are dynamically assigned based on alarm severity; for example, emergency alarms correspond to high-risk events such as equipment overload or sudden temperature rises, important alarms involve parameter fluctuations, and general alarms alert to routine anomalies. Each alarm message is associated with an equipment location identifier, alarm type classification, and preliminary handling guidelines. For instance, location information can be accurate to the distribution cabinet number, alarm types include "overcurrent" and "temperature rise," and handling guidelines provide operational suggestions such as "check the ventilation system" or "reduce the load." The sorting algorithm can combine alarm level and timestamp to ensure high-risk alarms are displayed first. During operation, the data dashboard and alarm control panel update in parallel: data streams are pulled in real-time from the edge data storage and management module, parsed, and used to refresh interface elements; alarm information is pushed by the early warning push module, and the control panel automatically scrolls to display new alarms and highlights key items. This design allows maintenance personnel to grasp the overall operational status at a glance and quickly locate anomalies, reducing manual screening time.

[0057] The interactive feedback window is a crucial interface for the visualization module, allowing maintenance personnel to directly input anomaly handling processes and result labels. This window is typically integrated into the graphical interface as a modal dialog box or side panel, supporting touch input or external keyboard operation for quick recording by on-site personnel. Its principle is based on an event-driven architecture. When an alarm is triggered or maintenance personnel take initiative, the window pops up and guides the user to input processing details, such as anomaly confirmation results, descriptions of handling measures, and classification labels. Result labels can include predefined options such as "false alarm," "confirmed fault," and "repaired," and can also support free text supplementation. After input, the data is sent back to the edge data storage and management module via an internal message queue. The feedback delay can be controlled within a few seconds to ensure timely feedback. In terms of usage, the feedback window is linked to the alarm console: when maintenance personnel click on an alarm, the window automatically associates with the event context and pre-fills device information; processing records can be stored in a structured manner, such as timestamps, operators, and processing steps. After receiving the feedback, the edge data storage and management module associates it with the original monitoring data for archiving and triggers the parameter adaptation engine or model training module for optimization and adjustment. For example, if multiple false alarms are flagged, the system can automatically relax the rule weights of the dynamic threshold judgment unit. During implementation, the feedback window also supports uploading attachments, such as on-site photos or log files, to enrich the feedback content. With these technical features working together, the visualization module provides an input interface, the storage module persists data, and the training module utilizes a feedback iterative model to form a closed-loop flow from human experience to system optimization, enhancing the system's adaptive capabilities.

[0058] The statistical analysis module is closely integrated with the visualization module, generating multi-dimensional statistical analysis reports at preset intervals. These intervals can be set daily, weekly, or monthly, flexibly adjusted according to operational strategies. Report content includes visual charts such as pie charts showing the distribution of abnormal event types and equipment health trend graphs. For example, pie charts display the percentage of various anomalies, such as temperature anomalies accounting for 30% and current anomalies accounting for 40%. Trend graphs show the changes in equipment health indices over time in line graph form, and health indices can be calculated based on parameter deviations. Reports are displayed through a dedicated statistical information area in the visualization module, which can exist as a standalone page or a dashboard component, supporting chart zooming and filtering. The statistical analysis module incorporates a root cause correlation analyzer, using the Apriori algorithm to perform correlation analysis on concurrent abnormal events. The Apriori algorithm scans a historical event database to uncover frequently co-occurring anomaly patterns; for example, if temperature and humidity anomalies often occur simultaneously, environmental factors are inferred as the root cause. The analyzer outputs the most relevant root cause inferences, such as "decreased cooling system efficiency leads to temperature rise," and presents them through text areas or prompts in the visualization module. During operation, the statistical analysis module periodically pulls historical anomaly data from the edge data storage and management module, performs data aggregation and pattern mining, and pushes the generated report to the visualization module. Root cause analysis can be based on concurrent events within a sliding time window, with the window length set to several hours to several days, to capture temporal correlations. In terms of the synergy of various technical features, the statistical module provides in-depth analysis capabilities, while the visualization module provides a display channel, enabling operations and maintenance personnel to not only view real-time status but also gain insights into long-term trends and root causes of faults, thus improving the depth of decision-making.

[0059] In one specific implementation, the cloud data storage and management module in the cloud decision layer constructs a distributed feature warehouse across edge nodes. This module is essentially a high-performance data management system, employing time-series databases such as InfluxDB or distributed storage systems such as Hadoop, to archive and index feature vectors and alarm events uploaded from various edge nodes. Feature vectors are typically mathematical representations extracted from sensor data, such as statistical features, frequency domain features, or model intermediate layer outputs, while alarm events include information such as anomaly trigger time, device identifier, and severity level. The principle is based on big data indexing and compression technology, enabling unified storage and rapid retrieval of massive heterogeneous data. Efficient queries by time range, device type, or anomaly category are achieved through the construction of inverted indexes or B-tree indexes. In usage, this module receives data streams from edge gateways in real time, cleans, deduplicates, and standardizes the data before storing it in shards across different storage nodes. Index updates are performed in near real-time, such as every few seconds or minutes, to ensure data accessibility. During implementation, data is first verified after it flows in, such as checking data integrity and timestamp consistency, and then archived according to a preset strategy. The archiving period can be set to daily or weekly, and long-term data can be transferred to cold storage to save costs.

[0060] The cockpit module provides a comprehensive overview of the regional power grid's operational status. This graphical user interface, typically deployed as a web dashboard, dynamically displays key indicators for the entire regional power grid, such as equipment online rate, anomaly distribution heatmaps, and health scores. Its principle is a data visualization engine that uses chart libraries like ECharts to transform complex data into intuitive graphics. Furthermore, the cockpit module features a global hyperparameter adjustment interface for the anomaly detection model, allowing operations personnel to centrally adjust model parameters for all edge nodes, such as learning rate, confidence threshold, or training period. The learning rate can be adjusted between 0.001 and 0.01 using a slider, and the confidence threshold can be selected between 50% and 90%. These values ​​are for reference only and can be flexibly set according to actual operational needs. During operation, the cockpit module periodically pulls aggregated data from the cloud data storage and management module, updating the view every 5 minutes for example. The adjustment interface sends user-input parameters to the federated model training module via API, enabling global optimization of model parameters. These two modules work together: the storage module provides a unified and reliable data foundation for the system, ensuring the traceability of historical data and real-time features; the dashboard module empowers maintenance personnel with macro-level monitoring and fine-tuning capabilities, making the regional power grid status clear at a glance and facilitating model adjustments. Specifically, the dashboard module provides maintenance personnel with a graphical interface showcasing the overall operational status of the power distribution room. The interface design employs a modular layout to ensure clear information presentation and easy access. Equipment operating status is displayed in the form of dynamic skeuomorphic diagrams, intuitively reflecting the normal, warning, or fault status of equipment through combinations of different colors and text. Statistical parameters are presented in real-time pie charts, clearly showing the changing trends of alarm types such as temperature and humidity. Abnormal alarm information is prominently displayed in pop-up windows, including key information such as event type, occurrence time, equipment location, and suggested handling measures, and is prioritized according to alarm level to ensure that important alarms attract the attention of maintenance personnel immediately. Maintenance personnel can monitor the operation of the power distribution room in real time through the cloud-based dashboard, quickly locating abnormal equipment and the root cause of problems. The cloud-based dashboard module offers powerful search and filtering capabilities, allowing maintenance personnel to quickly query relevant data based on equipment name, location, anomaly type, and other criteria. The system provides detailed data analysis reports and historical trend comparison functions, helping maintenance personnel gain a more comprehensive understanding of the equipment's operational status and potential problems. Based on the system's suggested solutions, maintenance personnel can take timely action, such as remotely adjusting equipment parameters or arranging on-site maintenance, thereby effectively preventing the further escalation of the fault.

[0061] The predictive maintenance module is a crucial component of the cloud-based decision-making layer. Based on historical equipment data and real-time stress information, it utilizes a Long Short-Term Memory (LSTM) network to predict the remaining effective lifespan of power distribution equipment and generate pre-maintenance plans. Historical data includes long-term collected equipment operating parameters such as temperature, current, vibration, and load curves, covering periods from months to years, with sampling frequencies ranging from hourly to daily. Real-time stress refers to the physical stress experienced by the equipment under current operating conditions, such as overload current, abnormal temperature rise, or mechanical vibration. This data is uploaded in real-time from the edge, with update frequencies ranging from every minute to every few minutes. The LSTM network is a recurrent neural network that uses gating mechanisms to memorize long-term temporal dependencies, enabling it to capture the gradual degradation trend of equipment performance. In practice, this module first obtains the equipment's entire lifecycle data from the cloud data storage and management module, performs preprocessing such as missing value imputation and outlier filtering, and then trains the LSTM model to map the input sequence to remaining lifespan values. During training, network hyperparameters, such as the number of hidden layer units, can be set between 64 and 256, and the training cycle can be between 200 and 500 epochs. These values ​​can be adjusted based on the amount of data and computing resources. During the prediction process, the module takes into account a recent historical time series, such as data from the past 30 days. The sliding window length can be 100 to 500 time points, and the output is an estimated remaining lifespan, such as 1000 hours of equipment remaining operation. Based on the prediction results, the module generates a pre-maintenance plan, including suggested maintenance time windows, maintenance types, and resource requirements. For example, when the remaining lifespan is below a threshold such as 200 hours, the plan can suggest scheduling an inspection or replacement within the following week. In implementation, the predictive maintenance module runs periodically, for example, performing batch predictions every morning. The results are stored in a cloud database, and alarms are pushed to maintenance personnel. This module is tightly integrated with the cloud data storage and management module to ensure reliable data sources, while also providing crucial input to the power grid dispatching module.

[0062] The power grid dispatch module receives equipment pre-maintenance time windows and availability status information from the predictive maintenance module. Using a linear programming algorithm as its core, with the objective function of minimizing system operating cost and constraints such as line capacity, node voltage deviation, and equipment availability status, it calculates a regional load dispatch scheme. The equipment pre-maintenance time window refers to the start and end times of planned equipment maintenance, and the availability status information indicates whether the equipment is online or dispatchable. Linear programming is a mathematical optimization method that solves for the optimal solution of the objective function under linear equality and inequality constraints. Here, the objective function is to minimize system operating cost, including generation cost, network loss cost, and maintenance cost; the constraints involve line capacity, such as setting the maximum current limit of transmission lines to 1000 amperes based on the rated value, and allowing node voltage deviation within ±5% of the nominal voltage. Equipment availability status is used as a binary variable in the modeling. In practice, this module collects power grid topology, real-time load demand, electricity price information, and the output of the predictive maintenance module to construct a linear programming model. During implementation, decision variables such as the output value of each generator unit are first defined. The objective function can be expressed as minimizing the total cost = ∑(output × unit cost). Constraints include ensuring that the line power flow does not exceed the capacity limit, the node voltage is within the allowable range, and load transfer restrictions when pre-maintenance equipment is unavailable. The solver uses the simplex method or interior point method for calculation, and the solution time can be controlled within a few seconds to meet real-time requirements. During operation, the grid dispatch module performs optimization periodically, for example, every 15 minutes or hour. Recalculation is also triggered when the predictive maintenance module updates information. The output dispatch scheme includes load allocation instructions, switching operation suggestions, and cost estimates. These instructions are sent to the edge execution unit or grid control system through the cloud interface. This module works in conjunction with the predictive maintenance module to ensure that the grid can still operate safely and economically during equipment maintenance, avoiding overload or voltage exceeding limits. Specifically, when the grid dispatch module comprehensively analyzes sensor data and the equipment loss status of the distribution rooms, it will access operating data such as current, voltage, and power factor from each distribution room, as well as the equipment loss assessment results. For example, when it is detected that the power supply capacity of a certain area substation A is reduced due to the aging of its internal equipment, while the adjacent substation B has a certain power supply margin, the power grid dispatch module calculates the optimal load dispatch scheme, issues instructions to adjust the opening and closing status of the power supply lines between the two substations, realizes the dynamic balance of the power grid load, improves the efficiency and reliability of power grid operation, and ensures the continuity and stability of power supply.

[0063] In one specific implementation, the cross-modal dynamic graph fusion network included in the machine learning model determination unit first performs preliminary processing on heterogeneous sensor data in the power distribution room through a multimodal feature extraction branch. This branch uses a temporal convolutional network for feature extraction of electrical time-series data such as current and voltage. The temporal convolutional network captures long-term dependencies in the data through multi-layer causal convolution and dilated convolution operations. Its kernel size can be set to 3 or 5, and the dilation factor can be periodically adjusted according to the data, such as 1, 2, 4, etc., to gradually expand the receptive field without increasing the number of parameters, thereby efficiently extracting local and global temporal patterns. For non-electrical time-series data such as temperature, humidity, and vibration, a gated recurrent unit is used for processing. The gated recurrent unit selectively remembers or forgets historical information through update and reset gate mechanisms. The hidden state dimension can be set to 64 or 128 to capture dynamic changes in the sequence. During operation, the multimodal feature extraction branch receives real-time data streams from sensors in parallel. Electrical quantity data is first preprocessed through standardization and then input into a temporal convolutional network, outputting high-dimensional feature vectors. Non-electrical quantity data is also normalized and then fed into the gate control recurrent unit to generate another set of feature vectors. These two sets of feature vectors respectively carry information on the operating status of electrical equipment and environmental impact, providing a foundation for subsequent fusion analysis. The branch may contain residual connections or batch normalization layers to ensure training stability. The feature extraction process is executed multiple times per second to meet real-time monitoring requirements. By dividing the processing of different modalities, the technical features ensure the specificity and richness of the feature representation, providing high-quality input for the graph structure learning layer.

[0064] Next, the graph structure learning layer receives the feature vectors output from the multimodal feature extraction branch and dynamically infers the association strength between each monitoring point using a trainable probabilistic generative model. This layer treats each monitoring point as a node in the graph, constructing a dynamically changing device association graph by learning the similarity or causal relationships between node features. The probabilistic generative model can be implemented based on a graph attention network or a variational autoencoder. It calculates the association probability between node pairs using a trainable parameter matrix and outputs a sparse adjacency matrix, where non-zero elements represent strongly associated node pairs. The threshold can be set to 0.1 or 0.2, retaining only connections with probabilities higher than this value to control computational complexity. During implementation, the graph structure learning layer receives feature vectors in real time and updates the association inference every second or every few seconds. For example, when a node's features fluctuate abnormally, the model recalculates its association strength with neighboring nodes. The sparsity of the adjacency matrix can be maintained between 10% and 30% to balance expressive power and efficiency. Subsequently, the spatiotemporal graph attention layer simultaneously inputs the feature vectors and the adjacency matrix. This layer contains spatial attention sublayers and temporal attention sublayers. The spatial attention sublayer calculates the attention coefficients of each node with its first- or second-order neighbors based on the topological structure defined by the adjacency matrix. The number of attention heads can be set to 4 or 8 to capture spatial dependencies from multiple perspectives. The temporal attention sublayer, for each node, calculates the attention coefficients between the current time point and historical time points within a sliding time window, such as the past 10 or 30 time points. The window length can be adjusted according to the data sampling frequency. After weighted aggregation of spatial and temporal features by the two sublayers, they are fused by concatenation or averaging to form an enhanced node representation. Through dynamic graph construction and spatiotemporal attention mechanisms, the various technical features effectively capture the collaborative change patterns and temporal evolution laws between devices, providing context-aware feature representations for anomaly detection.

[0065] Finally, the normalized flow density estimator estimates the probability density of the node features fused by the attention layer of the spatiotemporal graph and outputs the anomaly confidence score. Normalized flow is a generative model that maps complex data distributions to simpler distributions, such as Gaussian distributions, through a series of reversible transformations, thereby estimating the probability density of input features. This normalized flow density estimator can be composed of multiple coupled layers or autoregressive flow layers stacked together, with the number of layers set to 5 to 10, and the number of hidden layer neurons to 256 or 512, to ensure the model's expressive power and computational efficiency. During operation, the normalized flow density estimator receives the fused feature vector of each node, first calculating its probability value in the latent space through a forward transformation, and then comparing it with the baseline distribution learned under normal operating conditions to obtain the log-likelihood or probability density score. The anomaly confidence score is finally converted to a value between 0% and 100% through normalization or a sigmoid function; a higher value indicates a greater probability of an anomaly. The threshold can be set at 70% or 80% based on historical data as an alarm trigger point. In implementation, the density estimator is periodically trained using normal data samples to update the baseline distribution. The training frequency can be daily or weekly to ensure that the model adapts to changes in device status. Each technical feature maps node features to interpretable confidence indices through probability density estimation, so that anomaly detection is based not only on feature bias but also on overall distribution characteristics, thereby improving the reliability of detection.

[0066] In one specific implementation, the end-to-end trusted execution and dynamic resource scheduling mechanism first integrates a hardware-level trusted execution environment (TEX) into the processors of the edge computing layer and the cloud decision layer. This is a secure area built through physical isolation or hardware encryption technology, ensuring that sensitive data is protected from external attacks or unauthorized access during storage and processing. Its principle is based on the processor's built-in security extension functions, such as ARM TrustZone or Intel SGX, to create an independent execution space that only allows trusted code to run. In this system, the TEX is specifically used to store key parameters of the anomaly detection model and gradient information generated during federated learning, preventing model leakage or tampering. In terms of usage, when the edge or cloud processor needs to access or update the model, it must be verified through the TEX to ensure the legitimacy of the operation. During implementation, the TEX is initialized upon system startup, a security certificate is loaded, and model parameters are encrypted and stored within it. During the federated learning aggregation phase, gradient data is sealed in the TEX before uploading and is only decrypted for cloud computing. Each technical feature, through the combination of hardware-level security foundation and software processes, provides underlying protection for the trustworthiness of data across the entire chain. The specific implementation of the trusted execution environment may vary depending on the processor model, but the core goal remains the same: to enhance the overall security of the system through isolation mechanisms.

[0067] Regarding data transmission security, the system employs a layered encryption strategy for different links. Communication between the data acquisition layer and the edge computing layer is typically based on the local network, involving large data volumes but requiring high latency. Therefore, the AES-128 symmetric encryption algorithm is used. This algorithm uses a 128-bit key to encrypt data in blocks, balancing efficiency and security. Its principle is to quickly encrypt the data stream using permutation and obfuscation operations. The key can be dynamically generated and securely distributed by the edge gateway. In cross-network transmission between the edge computing layer and the cloud decision-making layer, due to the involvement of the public network environment and higher risks, an asymmetric encryption algorithm based on elliptic curves, such as ECDH or ECDSA, is used. Its principle relies on the intractability of elliptic curve mathematical problems, using public-key encryption and private-key decryption to ensure that even if the public key is leaked, the private key cannot be derived in reverse. To further enhance security, the system performs a key rotation every 24 hours. The rotation cycle can be adjusted according to actual security needs; for example, it can be shortened to 12 hours in high-risk scenarios. The rotation process is automatically completed through the key management service, and the new key takes effect immediately after the old key expires. In implementation, the data sender first obtains the currently valid key, encrypts the data packet, and then transmits it; the receiver verifies the key validity and then decrypts it. Each encryption feature, through algorithm selection adapted to different network scenarios and regular key updates, constructs a dynamic defense system. The AES-128 key length is fixed at 128 bits, while elliptic curve algorithms can use standard curves such as secp256r1. These choices are for reference only and can be adjusted according to the actual security strategy.

[0068] The dynamic resource scheduler integrated into the edge computing layer is responsible for real-time monitoring of the system's computational load and the comprehensive anomaly confidence level output by the anomaly detection module. This confidence level is a probability value ranging from 0% to 100%, reflecting the current risk of data anomalies. The scheduler dynamically allocates CPU computing resources based on a priority strategy. Its principle is to monitor CPU utilization and confidence level changes through an operating system-level interface and trigger resource reallocation instructions. When the comprehensive anomaly confidence level falls below a first confidence threshold (e.g., 40%, which can be set between 30% and 50%), indicating a low-risk state, the scheduler immediately allocates approximately 70% of the computing resources to the background anomaly detection model training task to optimize the model during idle periods. Simultaneously, it triggers CPU downclocking instructions to reduce the processor frequency and enter energy-saving mode, reducing energy consumption. Conversely, when the overall anomaly confidence level exceeds the second confidence threshold (which is higher than the first threshold and can be set between 70% and 90%, e.g., 80%), it indicates a significantly increased anomaly risk. The scheduler immediately interrupts all training tasks, reallocates approximately 90% of resources to the real-time anomaly detection module to ensure rapid response in the detection process, and removes the CPU frequency reduction instruction, restoring full-frequency operation. During implementation, the scheduler continuously samples confidence data at second-level intervals, compares thresholds through decision logic, and adjusts task priorities and CPU governor settings via system calls. These technical features, through real-time monitoring and flexible resource allocation, enable the system to adaptively balance safety and energy efficiency. The resource allocation ratios and threshold values ​​are merely examples and can be optimized based on equipment performance. It should be noted that the overall anomaly confidence level is the final output confidence level of the anomaly detection module, the result of multi-level processing by the collaborative adjudication unit. It integrates information from multiple sources, including: threshold comparison results from the dynamic business threshold judgment unit; preliminary anomaly confidence from the machine learning model judgment unit; query results from the local historical event database (used to distinguish between normal operation and real faults); and matching results from the latent fault feature database (extracting key features through interpretability analysis). Therefore, the comprehensive anomaly confidence is the final, more reliable anomaly probability assessment derived from the preliminary anomaly confidence, after multi-source information fusion and secondary analysis by the collaborative adjudication unit. Specifically, the system compares the preliminary anomaly confidence directly output by the machine learning model with the rule-based judgment results based on dynamic thresholds. When the two are inconsistent, the collaborative adjudication unit queries the historical event database to distinguish normal operation interference or initiates interpretability analysis to match the latent fault feature database, thereby correcting or confirming the preliminary result. When the two judgments are consistent, the final adjudication in this process is calculated using a simple weighted average or logical AND operation. Therefore, the comprehensive anomaly confidence is no longer simply a model output, but an intelligent decision-making result that integrates rules, historical experience, and fault mechanism knowledge, improving the accuracy and reliability of detection.

[0069] The number of devices and processing scale described herein are for the purpose of simplifying the description of the invention. Applications, modifications, and variations of the invention will be readily apparent to those skilled in the art.

[0070] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details.

Claims

1. A power distribution room anomaly detection system based on cloud-edge-device collaboration, characterized in that, include: The data acquisition layer is equipped with various heterogeneous sensors for real-time acquisition of operating parameters and environmental data of electrical equipment in the power distribution room; the data acquisition layer receives control commands from the edge storage and computing layer and dynamically adjusts its own data sampling frequency. An edge computing layer, which is communicatively connected to the data acquisition layer, is used for local real-time processing, anomaly detection, model training, and visualization of received sensor data. The edge computing layer has a built-in anomaly detection module and a prediction detection linkage module. The anomaly detection module is equipped with an anomaly detection model, which is used to analyze real-time monitoring data and generate early warning information. The prediction detection linkage module is used to monitor the changing trend of anomaly probability and generate control commands based on the changing trend, which are sent to the data acquisition layer through the edge gateway to dynamically adjust the sampling frequency. An edge gateway, integrated into the edge computing layer, is used to realize protocol conversion and data forwarding between the data acquisition layer and the edge computing layer, and to maintain bidirectional communication with the cloud decision-making layer; The cloud-based decision layer, connected to the edge computing layer network, aggregates data from multiple edge nodes for centralized storage, model optimization, and decision support. The cloud-based decision layer has a built-in federated model training module, which periodically aggregates the local model parameters trained by each edge computing layer, generates an optimized global model through a federated learning algorithm, and distributes the global model parameters to each edge computing layer to update its local anomaly detection model.

2. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 1, characterized in that, The edge computing layer also includes an early warning push module, a visualization module, and an edge data storage and management module; the early warning push module is configured to classify early warning information into multi-level alarm levels according to preset rules and send it to the visualization module and the edge data storage and management module. The prediction and detection linkage module built into the edge computing layer is configured to continuously monitor the slope of the change trend of the anomaly probability. When the slope of the change trend exceeds a first threshold and the anomaly probability continues to exceed a second threshold, an instruction to increase the sampling frequency of the data acquisition layer is generated. When the anomaly probability drops below a third threshold and remains stable for more than a preset time, an instruction to decrease the sampling frequency is generated. The third threshold is lower than the second threshold.

3. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 2, characterized in that, The anomaly detection model is a comprehensive anomaly detection model, which includes a sensor fault self-verification unit, a dynamic business threshold determination unit, a machine learning model determination unit, and a collaborative adjudication unit connected in sequence. The collaborative adjudication unit specifically includes: First, the threshold comparison result from the dynamic service threshold determination unit and the preliminary anomaly confidence level output from the machine learning model determination unit are received. When the threshold comparison result indicates an anomaly and the initial anomaly confidence level is lower than a preset confidence threshold, the collaborative decision unit accesses a local historical event database to query whether it is normal operation rather than a real fault under similar equipment operating conditions. When the threshold comparison result does not indicate an anomaly but the preliminary anomaly confidence level is higher than the preset confidence threshold, the collaborative adjudication unit initiates the interpretability analysis program of the machine learning model judgment unit to obtain the key feature sequence that leads to the high confidence level, and matches the feature sequence with a known latent fault feature library. The collaborative adjudication unit outputs the final anomaly determination conclusion based on the above query or matching results; The comprehensive anomaly detection model is also configured with a parameter adaptive engine, which is used to adjust the decision tree model branch weights in the dynamic business threshold determination unit or the loss function of the machine learning model determination unit in reverse based on the confirmation result after the operation and maintenance personnel confirm the false alarm or missed alarm through the visualization module.

4. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 1, characterized in that, The built-in federated model training module in the cloud-based decision layer specifically performs the following process: The cloud server initializes a global anomaly detection model and sends its parameters to each edge storage and computing layer. Each edge computing layer uses local data to train the global anomaly detection model, and then encrypts the model gradient parameters obtained after training using an encryption algorithm before uploading them to the cloud; The cloud server aggregates encrypted model gradient parameters from multiple edge computing layers and uses a federated averaging algorithm to perform weighted averaging calculations to generate parameters for a new generation of global models. The parameters of the new generation global model are distributed to all edge computing layers to update their local anomaly detection models.

5. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 1, characterized in that, The edge storage and computing layer is configured to automatically trigger the data upload process when the comprehensive anomaly confidence level output by its anomaly detection module is lower than a preset threshold. The data upload process includes: encapsulating the multi-source sensor monitoring data and associated local feature vectors within a set time window before and after the trigger time into a data packet and uploading it to the cloud decision layer through an edge gateway; The cloud-based decision layer is configured to perform the following operations upon receiving the data packet: a) invoke its analysis module to perform correlation analysis and deep learning model inference on the multi-source data within the data packet; b) perform pattern matching between the analysis results and the historical anomaly feature database from other edge nodes; c) generate a feedback data packet containing diagnostic conclusions and treatment instructions. The feedback data packet is simultaneously sent via the network to the source edge node and its adjacent edge nodes determined according to the power grid topology; the data storage and management module of the edge storage and computing layer is configured to receive and store the feedback data packet to update the local early warning rules.

6. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 2, characterized in that, The edge storage and computing layer also includes an anomaly detection model training module, which is configured to execute a local incremental learning process: based on historical and real-time monitoring data in the edge data storage and management module, the anomaly detection model is optimized in an incremental update manner, wherein the hyperparameters of the anomaly detection model training are dynamically adjusted through feedback from operation and maintenance personnel or cloud commands. The edge data storage and management module in the edge computing layer is configured to receive and store monitoring data from the data acquisition layer, early warning information sent by the early warning push module, and abnormal processing results and hyperparameters of the abnormal detection model returned by the visualization module. It also periodically uploads abnormal early warning information and processing results to the cloud decision layer and accesses multi-source heterogeneous sensor data through wired or wireless protocols. The edge data storage and management module has a built-in time-series data alignment and cleaning engine, which is configured to perform time synchronization, outlier filtering and format standardization on multi-source sensor data from different sampling frequencies and protocols to form a regular time-series dataset.

7. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 6, characterized in that, The visualization module in the edge computing layer provides a graphical human-machine interface through an integrated high-definition industrial display panel. The graphical human-computer interaction interface includes at least: The real-time data cockpit dynamically displays the current values ​​and change curves of key equipment operating parameters and environmental data. The abnormal alarm control panel presents alarm information in a priority sorted list, with each piece of information associated with the device location, alarm type, and initial handling guidelines. An interactive feedback window is used to receive the anomaly handling process and result tags entered by the operation and maintenance personnel. The result tags will be sent back to the edge data storage and management module as training feedback data. The edge computing layer also includes a statistical analysis module, which is connected to the visualization module. The statistical analysis module is configured to generate multi-dimensional statistical analysis reports at preset intervals, including pie charts of abnormal event type distribution and device health trend charts. The reports are displayed through the statistical information section of the visualization module. The statistical analysis module has a built-in root cause correlation analyzer that performs correlation analysis on concurrent abnormal events based on the Apriori algorithm and outputs the most relevant root cause inference results to the visualization module.

8. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 1, characterized in that, The cloud-based decision-making layer also includes: a cloud-based data storage and management module, a cockpit module, a predictive maintenance module, and a power grid dispatching module; The cloud data storage and management module constructs a feature repository across edge nodes for archiving and indexing feature vectors and alarm events uploaded from the edge side; The cockpit module provides an overview of the regional power grid's operational status and is equipped with a global control interface for the hyperparameters of the anomaly detection model. The predictive maintenance module uses an LSTM network to predict the remaining effective life of the equipment based on historical data and real-time stress, and generates a pre-maintenance plan. The power grid dispatching module receives the equipment pre-maintenance time window and availability status information output by the predictive maintenance module, and calculates the regional load dispatching scheme using a linear programming algorithm as the core, with the minimum system operating cost as the objective function and line capacity, node voltage deviation and equipment availability status as constraints.

9. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 3, characterized in that, Place The machine learning model decision unit includes a cross-modal dynamic graph fusion network, which comprises, in sequence: The multimodal feature extraction branch is configured to use a temporal convolutional network to process electrical quantity time-series data and a gated recurrent unit to process non-electrical quantity time-series data. The graph structure learning layer takes the feature vectors output by the multimodal feature extraction branch as its input, dynamically infers the correlation strength between the feature vectors of each monitoring point through a trainable probabilistic generation model, and outputs a sparse adjacency matrix to represent the device association graph at the current moment. The spatiotemporal graph attention layer takes the feature vector and the adjacency matrix as input. This layer includes a spatial attention sublayer and a temporal attention sublayer. The spatial attention sublayer calculates the attention coefficient between each node in the graph and its neighboring nodes. The temporal attention sublayer calculates the attention coefficient between each node at the current time point and historical time points within a sliding time window, and fuses the features. The normalized flow density estimator estimates the probability density of node features after fusion by the spatiotemporal graph attention layer and outputs the anomaly confidence level.

10. The power distribution room anomaly detection system based on cloud-edge-device collaboration according to claim 1, characterized in that, It also includes end-to-end trusted execution and dynamic resource scheduling mechanisms, which specifically include; A hardware-level trusted execution environment is integrated into the processors of the edge computing layer and the cloud decision layer to store anomaly detection model parameters and federated learning gradients; When transmitting data between the data acquisition layer and the edge computing layer, the AES-128 algorithm is used for encryption; when transmitting data between the edge computing layer and the cloud decision layer, an elliptic curve-based asymmetric encryption algorithm is used, and the key is rotated every 24 hours. The edge computing layer also integrates a dynamic resource scheduler, which monitors the computing load and the comprehensive anomaly confidence level output by the anomaly detection model in real time. When the overall anomaly confidence level is lower than the first confidence level threshold, the dynamic resource scheduler allocates 70% of the computing resources to the anomaly detection model training task and triggers a CPU frequency reduction instruction to enter power saving mode. When the overall confidence level exceeds the second confidence level threshold which is higher than the first confidence level threshold, the dynamic resource scheduler immediately interrupts the anomaly detection model training task, reallocates 90% of the computing resources to the anomaly detection module, and removes the CPU frequency reduction instruction.

Citation Information

Cited By

  • Power operation data monitoring method and system based on cloud network

    CN121663813A

  • Temperature monitoring system and method thereof

    CN121877206A

  • Temperature monitoring system and method thereof

    CN121877206B

  • Underground coal mine multi-parameter self-adaptive sensing and anomaly recognition sensor network driven by edge calculation

    CN121884567A

  • Industrial edge data acquisition gateway system and implementation method

    CN122027397A