SDNN-based power distribution fusion terminal real-time video reasoning method
By adopting SDNN technology on the power distribution converged terminal, real-time processing of video data and optimization of model, the delay and bandwidth consumption problems caused by underutilization of computing resources and video processing relying on the cloud in the existing technology are solved, and efficient and real-time video surveillance and abnormal detection are achieved.
Patent Information
- Application Number
- CN202510196643.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
Smart Images

Figure HDA0005281693570000011 
Figure HDA0005281693570000021
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart grid monitoring, and more particularly to a real-time video inference method for a distribution fusion terminal based on SDNN. Background Art
[0002] In recent years, the rapid development of artificial intelligence (AI) technology has led to its increasingly widespread application in the field of video surveillance. In particular, the breakthroughs in deep learning and computer vision technologies have enabled the video surveillance system to gradually evolve from the traditional manual patrol and video playback modes towards the intelligent and automated directions. In the traditional video surveillance system, the storage and playback of video data are the core functions, which usually rely on a central server for centralized processing. However, with the increase in the number of surveillance cameras, the data volume has increased exponentially, and the manual patrol method is difficult to meet the real-time analysis requirements of large-scale video data. At the same time, the traditional system mostly uses fixed thresholds and rule matching algorithms, and has limited capabilities for target recognition and anomaly detection in complex environments, and is easily affected by environmental light changes, occlusion, and noise interference, resulting in false alarms and missed alarms.
[0003] To improve the intelligence level of video surveillance, AI technology has been gradually introduced and has played an important role in target detection, behavior analysis, and anomaly event recognition. AI algorithms rely on deep learning models to automatically learn the features in video data, identify abnormal behaviors, and improve the efficiency of security management. Currently, the application of AI in video surveillance mainly relies on the cloud computing architecture, that is, transmitting video data to a cloud server, and a high-performance computing cluster executes the deep learning inference task. This mode is suitable for large-scale data processing, but is limited by network bandwidth, transmission delay, and the allocation efficiency of cloud computing resources, and is difficult to meet the application scenarios with high real-time requirements. In addition, the data transmission in the cloud computing mode requires a long-term dependence on the Internet connection, which has potential data security risks.
[0004] To make up for the limitations of the cloud computing mode in video surveillance applications, edge computing technology has gradually been applied to video inference tasks. The core concept of edge AI video inference technology is to deploy an AI model at a location close to the data source (such as a camera, sensor, embedded terminal, or edge server) to achieve real-time and low-latency data analysis and decision-making. Compared with cloud computing, edge AI can reduce the data transmission volume, reduce network dependence, and at the same time improve the response speed and privacy protection ability of the video surveillance system. Currently, edge AI has been applied in the fields of intelligent security, industrial monitoring, and smart cities. For example, by deploying a lightweight target detection algorithm at the camera end, local face recognition, vehicle detection, and anomaly behavior monitoring can be achieved, thereby reducing the dependence on cloud computing resources.
[0005] In the field of power grid distribution, monitoring systems are usually used for the status monitoring of key infrastructures such as substations and distribution substations. Although existing integrated terminals for distribution substations already have certain data acquisition and computing capabilities, their AI computing power resources have not been fully utilized, and most video data still needs to be transmitted to the cloud or data center for processing. Due to the complex network conditions and limited bandwidth resources in the power grid environment, remote transmission may cause high latency, affecting the real-time monitoring ability of the system. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention discloses a real-time video inference method for a distribution integration terminal based on SDNN, aiming to solve the problems existing in the existing technology, such as the underutilization of computing power resources of existing distribution terminals, high latency and high bandwidth consumption caused by video processing relying on the cloud, and poor adaptability of the model in complex power grid scenarios.
[0007] The present invention adopts the following technical solutions:
[0008] A real-time video inference method for a distribution integration terminal based on SDNN, comprising the following steps:
[0009] Step S1: Collect original video data, screen target key frames by using a time-domain key frame extraction and local feature description method, mark the target area by using a semi-automatic image segmentation and semantic annotation method, and construct a power grid sample library through a stratified sampling mechanism, including training samples, verification samples and test samples;
[0010] Step S2: Based on the constructed power grid sample library, through a convolutional neural network combined with a transfer learning strategy, perform model training through a gradient descent optimization algorithm and a data augmentation mechanism, and use cross-validation and grid search methods for hyperparameter tuning to obtain a basic model;
[0011] Step S3: Adopt an intermediate representation conversion mechanism to convert the trained basic model into the ONNX format, and perform compatibility verification on each operator through an operator fusion and kernel reconstruction algorithm. If the operator verification fails, use a tensor operation approximation algorithm for replacement, and optimize the operator execution order by using a subgraph segmentation mechanism;
[0012] Step S4: Based on the sparsity optimization strategy of SDNN, prune and quantize the converted ONNX format model to low-bit precision, compile the ONNX format model into an SDNN inference model; and deploy the SDNN format model to the distribution integration terminal through a heterogeneous computing acceleration framework, and the heterogeneous computing acceleration framework optimizes memory access through a hierarchical caching mechanism;
[0013] Step S5: Deploy the compiled and optimized SDNN inference model on the edge device of the distribution fusion terminal, perform frame-by-frame processing on the real-time video stream through the adaptive frame sampling, noise suppression, and dynamic range compression preprocessing mechanism, and use the statistical threshold and anomaly detection algorithm to judge the target behavior in the video in real time. If it is determined to be abnormal, send a warning signal; otherwise, continue to process the subsequent video frames.
[0014] Step S6: Monitor the real-time performance of the SDNN inference model through the adaptive feedback mechanism optimized based on reinforcement learning; if a decrease in inference accuracy or an excessive response delay is detected, retrain and update the model through the incremental learning and adaptive hyperparameter adjustment mechanism.
[0015] Preferably, the working method of step S1 is as follows: First, calculate the pixel motion vector field between consecutive video frames through the time-domain gradient joint detection algorithm, and combine the inter-frame structural similarity index to construct a key-frame screening model based on dynamic energy entropy: When the proportion of the optical flow mutation region between adjacent frames exceeds the preset threshold, capture the transient features of the device including but not limited to mechanical vibration and arc flash through the key-frame intercepting mechanism; Based on the captured transient features, adopt the local feature description method, introduce the octree segmentation strategy of the Gaussian difference pyramid when constructing the scale space in the key frame, divide the image into hierarchical grid units, and perform direction histogram clustering on the feature points in each unit to generate a multi-dimensional feature vector with spatial topological constraints; Then, execute the target region extraction through the semi-supervised adversarial image segmentation network; Finally, establish a dynamic hierarchical category based on the operating parameters of the power distribution equipment, and the operating parameters of the power distribution equipment include but not limited to the load rate and the temperature rise rate, and select samples through Latin hypercube sampling within each layer to output the power grid sample library.
[0016] Preferably, the working method of step S2 is as follows: First, initialize the model parameters by using the convolutional neural network, and adopt the transfer learning strategy to extract general features from the power grid sample library output in step S1; Then, use the gradient descent optimization algorithm to iteratively update the model parameters, calculate the cross-entropy or mean square error loss function in each training cycle, and use the backpropagation mechanism to adjust the weights; At the same time, apply operations including but not limited to rotation, scaling, brightness adjustment, and noise injection through the data augmentation mechanism during the input data processing; Then, divide the sample library into multiple subsets through the cross-validation method, and use k-fold cross-validation to train and validate the model; Finally, use the grid search method to adjust the learning rate, batch size, regularization coefficient, and convolutional layer number parameters within the preset hyperparameter space, and evaluate the performance of the validation set under each parameter combination, and select the configuration that minimizes the validation error.
[0017] Preferably, the working method of step S3 is as follows: First, through the heterogeneous computing graph intermediate representation conversion mechanism, when converting the basic model to the ONNX format, a dynamic operator mapping table is embedded. Then, the graph structure parsing algorithm is used to extract the model computation flow graph, and the topological sorting is used to identify the adjacent nodes of the batch normalization layer and the convolutional layer. If an unaligned memory access pattern is detected, the convolution, batch normalization, and activation function sequences are merged into a single composite operator through the cross-layer operator fusion strategy, and the kernel memory layout is reconstructed to match the vector register bit width of the hardware. Subsequently, the tensor semantic compatibility verification algorithm is used to verify the operator compatibility. The tensor semantic compatibility verification algorithm defines the constraint conditions based on the ONNX operator protocol, and simulates the input and output tensor dimensions and data types of each operator through symbolic execution. If a dynamic shape tensor is detected, the non-supported operator is numerically equivalently replaced through polynomial expansion. Finally, the depth-first search is used to traverse the computation graph to construct the data dependency chain. When it is detected that there are shared weight tensors in the parallel branches, the computation graph is split into multiple independently executable subgraph modules through the subgraph segmentation strategy, and a dedicated cache area is allocated for each subgraph by the hardware resource allocator.
[0018] Preferably, the SDNN-based sparsity optimization strategy first calculates the product of the L1 norm and the gradient of each convolutional kernel weight through the gradient sensitivity analysis algorithm. If the score is lower than the preset threshold, redundant convolutional kernels are removed at the channel granularity through the structured pruning mechanism, and the information loss is compensated through the residual connection. Then, the mixed-precision quantization mechanism is used to perform low-bit conversion: The mixed-precision quantization mechanism analyzes the distribution characteristics of the activation values of each layer using the KL divergence, allocates 8-bit quantization bit width to the high-dynamic range layer, reduces the dimension of the smooth response layer to 4 bits, and generates a dynamic scaling factor through the layer-by-layer calibration algorithm. Subsequently, the convolution operator is disassembled into a vector multiply-accumulate instruction sequence supported by the hardware through the instruction set mapping table. When it is detected that the tensor dimension does not match the register bit width, the input feature map is block-aligned according to the hardware cache line size through the data rearrangement strategy, and a transpose instruction is inserted to eliminate the memory access fragmentation.
[0019] Preferably, the hierarchical cache cooperation mechanism constructs a three-level cache structure in the heterogeneous acceleration framework, including D1 cache, D2 cache, and D3 cache; the D1 cache is used to store the high-frequency access blocks of the weight tensors, the D2 cache is used to pre-store the input feature maps of the next inference cycle, and the D3 cache is used to retain the intermediate calculation results for multi-model reuse; when a memory access request arrives, if it is detected that the target data exists in the D2 cache and the hit rate is lower than 85%, the cross-layer shared data block is retained based on the principle of spatial and temporal locality through the cache replacement algorithm.
[0020] Preferably, the working method of step S5 is as follows: Based on the original video stream collected by the edge device, first, an adaptive frame sampling mechanism is adopted to calculate the key frame index according to the inter-frame motion amplitude and waveform distribution; after sampling, Wiener filtering is used to process the missing key frames; in addition, non-linear mapping of pixel value distribution is performed through dynamic range compression; the frame-by-frame images after deletion are input into the SDNN inference model, and the chi-square test method is used to evaluate the behavior target features in real time; during the detection process, based on the prior statistical model and historical feature distribution for comparative analysis, if the deviation exceeds the preset threshold, an abnormal signal is sent using the MQTT protocol.
[0021] Preferably, the adaptive feedback mechanism optimized based on reinforcement learning statistically infers the standard deviation of the accuracy and the average response delay through a sliding window. When the accuracy drops by more than 2% for 3 consecutive windows or the delay exceeds the preset threshold, a state space is constructed through a decision engine driven by reinforcement learning, including accuracy deviation, delay level, and hardware resource occupancy rate; and the optimal actions for model update, parameter adjustment, and resource reallocation are calculated through a double deep Q network.
[0022] Preferably, when the model triggers retraining and update in step S6, first, frames with a confidence level lower than 0.6 are extracted from the real-time video stream through an incremental adversarial training method and combined with edge cases synthesized by the generative adversarial network to form an incremental data set; and the parameters of the fully connected layer are updated through an elastic weight consolidation algorithm; at the same time, a Gaussian process regression model is constructed based on historical tuning records to predict the optimal combination of learning rates. If the validation loss does not decrease for 5 consecutive iterations, an early stopping mechanism is triggered and rolled back to the optimal parameter snapshot; finally, a parallel computing sandbox is established outside the model inference thread. After the incremental training is completed, the output similarity of the new and old models is compared. If the standard is met, the model weights in the memory are replaced through an atomic write operation.
[0023] Based on the above technical solutions, the positive and beneficial effects of the present invention are as follows:
[0024] 1. Through the time-domain key frame extraction and hierarchical sampling mechanism in step S1, key frame data with high information density are effectively screened out. Combining with the adaptive frame sampling technology in step S5 to dynamically adjust the video acquisition frequency, the transmission volume of redundant video data is greatly reduced; this technical chain directly solves the bandwidth pressure caused by full-volume video transmission in the existing solutions, enabling stable transmission of multiple video streams with limited network resources; at the same time, model pruning and low-bit quantization in step S4 further compress the model volume and reduce the network load during model update.
[0025] 2. The operator fusion and kernel reconstruction in step S3 optimize the model calculation efficiency. Combined with the hierarchical caching mechanism in step S4, it optimizes the memory access pattern and significantly shortens the single-frame inference time. In addition, the anomaly detection algorithm in step S5 can complete the fault determination in a short time through temporal feature fusion and dynamic threshold judgment. Compared with the traditional cloud backhaul scheme, it significantly shortens the system response delay and meets the real-time monitoring requirements of power grid key equipment for transient events.
[0026] 3. The transfer learning strategy in step S2 combined with the heterogeneous computing acceleration framework in step S4 makes full use of the dedicated computing units (such as NPUs) of the terminal hardware and significantly improves the utilization rate of computing power resources. The reinforcement learning feedback mechanism in step S6 realizes the adaptive matching of computing tasks and hardware capabilities through dynamic resource allocation and hot-plug model updates, enabling the terminal device to process multiple types of AI tasks in parallel and breaking through the excessive dependence on cloud computing power in the traditional scheme.
[0027] 4. The semi-automatic image segmentation and semantic annotation in step S1 generate a high-quality sample library. Combined with the multi-modal data augmentation technology in step S2, it improves the model's adaptability to complex scenarios such as light changes and noise interference. In addition, the dynamic range compression and noise suppression algorithm in step S5 effectively improves the input video quality, enabling the model to maintain stable detection accuracy under conditions such as electromagnetic interference and bad weather. Brief Description of the Drawings
[0028] Figure 1 It is a step schematic diagram of a real-time video inference method for a distribution fusion terminal based on SDNN of the present invention;
[0029] Figure 2 It is a working method framework diagram of step S1 of the present invention. Detailed Embodiments
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] In the embodiment, as Figure 1 shown: The steps of a real-time video inference method for a distribution fusion terminal based on SDNN are as follows:
[0032] Step S1: Collect the original video data, screen the target key frames by using the time-domain key frame extraction and local feature description methods, mark the target area through the semi-automatic image segmentation and semantic annotation methods, and construct the power grid sample library through the stratified sampling mechanism, including training samples, validation samples, and test samples; Further, as Figure 2 shown, the working method of step S1 is as follows: First, calculate the pixel motion vector field between consecutive video frames through the time-domain gradient joint detection algorithm, and combine the structural similarity index between frames to construct a key frame screening model based on dynamic energy entropy: When the proportion of the optical flow mutation area between adjacent frames exceeds the preset threshold, capture the transient features of the equipment including but not limited to mechanical vibration and arc flash through the key frame intercepting mechanism; Based on the captured transient features, adopt the local feature description method, introduce the octree segmentation strategy of the Gaussian difference pyramid when constructing the scale space in the key frame, divide the image into hierarchical grid units, and perform direction histogram clustering on the feature points in each unit to generate a multi-dimensional feature vector with spatial topological constraints; Then, execute the target area extraction through the semi-supervised adversarial image segmentation network; Finally, establish a dynamic hierarchical category based on the operating parameters of the distribution equipment, and the operating parameters of the distribution equipment include but not limited to the load rate and the temperature rise rate, and select samples through Latin hypercube sampling within each layer to output the power grid sample library.
[0033] Step S2: Based on the constructed power grid sample library, through the convolutional neural network combined with the transfer learning strategy, perform model training through the gradient descent optimization algorithm and the data augmentation mechanism, and use the cross-validation and grid search methods to optimize the hyperparameters to obtain the basic model; Further, the working method of step S2 is as follows: First, initialize the model parameters by using the convolutional neural network, and adopt the transfer learning strategy to extract the general features from the power grid sample library output in step S1; Then, use the gradient descent optimization algorithm to iteratively update the model parameters, calculate the cross-entropy or mean square error loss function in each training cycle, and use the backpropagation mechanism to adjust the weights; At the same time, apply operations including but not limited to rotation, scaling, brightness adjustment, and noise injection through the data augmentation mechanism during the input data processing; Then, divide the sample library into multiple subsets through the cross-validation method, and use k-fold cross-validation to train and validate the model; Finally, use the grid search method to adjust the learning rate, batch size, regularization coefficient, and convolutional layer number parameters within the preset hyperparameter space, and evaluate the performance of the validation set under each parameter combination, and select the configuration that minimizes the validation error.
[0034] Step S3: Use an intermediate representation conversion mechanism to convert the trained basic model into the ONNX format, and perform compatibility verification on each operator through an operator fusion and kernel reconstruction algorithm. If the operator verification fails, use a tensor operation approximation algorithm for replacement, and optimize the operator execution order using a subgraph segmentation mechanism. Further, the working method of step S3 is as follows: First, when converting the basic model into the ONNX format through a heterogeneous computation graph intermediate representation conversion mechanism, embed a dynamic operator mapping table. Then, use a graph structure parsing algorithm to extract the model computation flow graph, and identify adjacent nodes of the batch normalization layer and the convolutional layer through topological sorting. If an unaligned memory access pattern is detected, merge the convolution, batch normalization, and activation function sequences into a single composite operator through a cross-layer operator fusion strategy, and reconstruct the kernel memory layout to match the vector register bit width of the hardware. Subsequently, use a tensor semantic compatibility verification algorithm to verify operator compatibility. The tensor semantic compatibility verification algorithm defines constraint conditions based on the ONNX operator protocol, and simulates the input and output tensor dimensions and data types of each operator through symbolic execution. If a dynamic shape tensor is detected, perform a numerical equivalent substitution on the non-supported operator through polynomial expansion. Finally, use a depth-first search to traverse the computation graph to construct a data dependency chain. When a shared weight tensor is detected in a parallel branch, split the computation graph into multiple independently executable subgraph modules through a subgraph segmentation strategy, and allocate a dedicated cache area for each subgraph through a hardware resource allocator.
[0035] Step S4: Based on the sparsity optimization strategy of SDNN, prune and perform low-bit precision quantization on the converted ONNX format model, and compile the ONNX format model into an SDNN inference model; and deploy the SDNN format model to the power distribution fusion terminal through a heterogeneous computing acceleration framework, and the heterogeneous computing acceleration framework optimizes memory access through a hierarchical caching mechanism; further, based on the sparsity optimization strategy of SDNN, first calculate the product of the L1 norm and gradient of each convolutional kernel weight through the gradient sensitivity analysis algorithm. If the score is lower than the preset threshold, remove redundant convolutional kernels at the channel granularity through the structured pruning mechanism, and compensate for information loss through residual connection; then, adopt a mixed-precision quantization mechanism to perform low-bit conversion: the mixed-precision quantization mechanism uses KL divergence to analyze the distribution characteristics of the activation values of each layer, assigns 8-bit quantization bit width to high-dynamic-range layers, reduces the dimension of smooth response layers to 4 bits, and generates dynamic scaling factors through a layer-by-layer calibration algorithm; subsequently, disassemble the convolutional operator into a vector multiply-accumulate instruction sequence supported by the hardware through an instruction set mapping table. When it is detected that the tensor dimension does not match the register bit width, align the input feature map in blocks according to the hardware cache line size through a data rearrangement strategy, and insert transpose instructions to eliminate memory access fragmentation; in addition, a hierarchical caching cooperation mechanism constructs a three-level cache structure in the heterogeneous acceleration framework, including a D1 cache, a D2 cache, and a D3 cache; the D1 cache is used to store the high-frequency access blocks of the weight tensor, the D2 cache is used to pre-store the input feature map of the next inference cycle, and the D3 cache is used to retain intermediate calculation results for multi-model reuse; when a memory access request arrives, if it is detected that the target data exists in the D2 cache and the hit rate is lower than 85%, retain the cross-layer shared data block based on the principle of spatio-temporal locality through a cache replacement algorithm.
[0036] Step S5: Deploy the compiled and optimized SDNN inference model on the edge device of the power distribution fusion terminal, perform frame-by-frame processing on the real-time video stream through an adaptive frame sampling, noise suppression, and dynamic range compression preprocessing mechanism, and use a statistical threshold and an anomaly detection algorithm to judge the target behavior in the video in real time. If it is determined to be abnormal, send a warning signal; otherwise, continue to process subsequent video frames; further, the working method of step S5 is as follows: based on the original video stream collected by the edge device, first adopt an adaptive frame sampling mechanism to calculate the key frame index according to the inter-frame motion amplitude and waveform distribution; after sampling, use Wiener filtering to process the missing key frames; in addition, perform a non-linear mapping on the pixel value assignment through dynamic range compression; the processed frame-by-frame images are input into the SDNN inference model, and the behavior target features are evaluated in real time through the chi-square test method; during the detection process, based on a priori statistical models and historical feature distributions for comparative analysis, if the deviation exceeds the preset threshold, send an anomaly signal using the MQTT protocol.
[0037] Step S6: Monitor the real-time performance of the SDNN inference model through an adaptive feedback mechanism optimized based on reinforcement learning; if a decrease in inference accuracy or an excessive response delay is detected, retrain and update the model through an incremental learning and adaptive hyperparameter adjustment mechanism; further, the adaptive feedback mechanism optimized based on reinforcement learning statistically calculates the standard deviation of inference accuracy and the average response delay through a sliding window. When the accuracy drops by more than 2% in 3 consecutive windows or the delay exceeds the preset threshold, a state space is constructed through a decision engine driven by reinforcement learning, including accuracy deviation, delay level, and hardware resource occupancy rate; and the optimal actions for model update, parameter adjustment, and resource reallocation are calculated through a double deep Q-network; in addition, when the model triggers retraining and update in step S6, first, frames with a confidence level lower than 0.6 are extracted from the real-time video stream through an incremental adversarial training method, and together with the edge cases synthesized by the generative adversarial network, they form an incremental data set; and the parameters of the fully connected layer are updated through an elastic weight consolidation algorithm; at the same time, a Gaussian process regression model is constructed based on historical tuning records to predict the optimal combination of learning rates. If the validation loss does not decrease for 5 consecutive iterations, an early stopping mechanism is triggered and rolled back to the optimal parameter snapshot; finally, a parallel computing sandbox is established outside the model inference thread. After the incremental training is completed, the output similarity of the new and old models is compared. If it meets the standard, the model weights in the memory are replaced through an atomic write operation.
[0038] In step S1 of the above embodiment, the core of the time-domain gradient joint detection algorithm lies in constructing a joint discriminant function of the optical flow field and structural similarity (SSIM). Its mathematical essence is to non-linearly couple the mutation energy of the optical flow vector field and the attenuation rate of SSIM; the optical flow field energy calculation is based on the variational optimization solution of the Horn-Schunck optical flow equation. By minimizing the brightness constancy constraint and the smoothness term constraint, the pixel-level motion vector field is solved. The SSIM attenuation model quantifies the temporal change intensity of video content by calculating the similarity differences in the three dimensions of brightness, contrast, and structure between adjacent frames. The construction of the dynamic energy entropy adopts the information entropy theory, and the energy distribution in the optical flow mutation region and the SSIM attenuation rate are weighted and fused. When the entropy value exceeds the threshold, key frame extraction is triggered. Through the dynamic balance mechanism of energy entropy, this model preferentially captures the periodic motion patterns caused by mechanical vibrations and the non-steady light intensity changes of arc flashes, effectively distinguishing normal working conditions from abnormal transient events.
[0039] The octree segmentation strategy of the Difference of Gaussian (DoG) pyramid is a multi-scale spatial decomposition method. In the scale space construction stage, an image pyramid is generated through multi-scale convolution of the Gaussian kernel function, and each layer corresponds to the feature response of a specific spatial frequency. Octree segmentation divides each layer of the image into 8×8 grid cells, and dynamically adjusts the grid density through a recursive quadtree splitting mechanism (extended to three dimensions is the octree): if the density of feature points within a cell exceeds the threshold, the grid is further subdivided to improve the local description accuracy. Direction histogram clustering uses an improved K-means++ algorithm, and introduces a spatial topology weight factor in the 128-dimensional feature space - the closer the Euclidean distance of the feature point to the center of the grid, the higher the clustering weight of its Histogram of Oriented Gradients (HOG). This mechanism enables the feature description of local micro-targets such as wire joints and insulator cracks to have spatial position awareness, enhancing the discriminative robustness of subsequent classifiers.
[0040] The Semi-Supervised Adversarial Image Segmentation Network (Semi-AGS) consists of a generator and a discriminator to form a dynamic game system. The generator is based on the U-Net architecture and introduces a channel attention mechanism (ECA-Net) to dynamically allocate the feature weights of each convolutional layer, preferentially strengthening the feature responses of device edges and texture details. The discriminator adopts the structure of a Markov discriminator (PatchGAN), evaluates the difference between the generated segmentation mask and the real annotation at the image patch level, and forces the generator to optimize the continuity of the segmentation boundary through the backpropagation of the adversarial loss. For labeled data, Dice loss + cross-entropy loss is used, and for unlabeled data, it is jointly constrained by the adversarial loss and consistency regularization (such as CutMix data augmentation) to achieve a balance between annotation efficiency and segmentation accuracy.
[0041] The hierarchical mechanism based on Latin Hypercube Sampling (LHS) achieves the optimal coverage of the sample space through the following steps:
[0042] R1: Dynamically divide the sample space according to the joint distribution of device operating parameters (such as the two-dimensional probability density function of load rate - temperature rise rate), and each layer corresponds to a specific operating condition interval (such as high load - high temperature rise, low load - low temperature rise, etc.).
[0043] R2: LHS performs uniform spatial sampling within each layer: divide the sample space into equal-probability hypercubes, ensuring that only one sample point falls into each interval in each dimension, thus avoiding the clustering bias of traditional random sampling.
[0044] The mathematical essence of this method is to maximize the statistical independence of the sample set through the theory of orthogonal experimental design, enabling the training samples to cover the extreme combination operating conditions of device operating parameters (such as sudden load increase accompanied by local overheating), and significantly improving the generalization ability of the model to rare fault modes.
[0045] In actual implementation, the present invention uses an integrated fusion terminal as the hardware platform. This terminal integrates a high-definition camera, an image acquisition module, an embedded processor (such as a good working platform for GPU / FPGA), and a dedicated data acquisition card for monitoring. Data is transmitted at high speed between the camera and the embedded processor through high-speed serial interfaces (such as USB 3.0, PCIe). The camera continuously acquires power grid monitoring videos. After the pre-conversion of analog signals to digital signals, the videos are transmitted to the time-domain joint detection module running in the embedded processor. This module calculates the optical flow vector field and the structural similarity index between consecutive frames based on a preset algorithm, and realizes hardware processing through hardware acceleration to ensure real-time analysis of high-frame-rate data streams. Then, the software module embedded in the processor realizes the construction of the Gaussian sparsity index pyramid and the octree image segmentation. This module realizes partial acceleration using FPGA during the repair process, generates structured grid data, calculates the histogram of the feature point directions within each grid, and outputs multi-dimensional feature processing. The key frames and feature data obtained through the local storage unit (such as SSD or eMMC) are temporarily stored, and then the built-in semi-supervised image segmentation network performs the target area extraction task on the embedded GPU. At the same time, the real-time data acquisition interface is connected to the upper-layer connection device through the CAN bus or the grid to form a monitoring system connection, and the power grid operation parameters are obtained in real time; these parameters are normalized and feature-processed through a specially designed data fusion module to form a dynamic hierarchical category data structure. Through the software module based on the Latin hypercube algorithm, targeted samples are selected within each category, and after integration with the image feature data, a power grid sample library is constructed. The hardware and software modules of the entire system are interconnected through an internal high-speed network, and a dedicated protocol is used between each functional module to achieve data synchronization and status feedback, realizing the full-process connection from video acquisition, feature extraction, target segmentation to sample library construction. This method efficiently and fully utilizes the local computing resources of the terminal, embeds image processing algorithms, deep learning models, and data sampling strategies into the embedded platform, realizes full-process edge computing, avoids the problem of uploading a large amount of data to the cloud in traditional systems, and improves the real-time performance and system stability through hardware acceleration and hardware processing.
[0046] The implementation of step S1 of the present invention enables the fusion terminal to directly perform real-time power supply and consumption feature realization on the power grid monitoring video at the edge, significantly improving the accuracy of data collection marking and the target area. Compared with the prior art, traditional systems usually rely on cloud centralized processing. Due to the complexity of the network environment and bandwidth determination in the power grid, the video data transmission delay is serious, and it is difficult to effectively realize key features in low-quality videos. However, the present invention captures key frames by calculating the inter-frame optical flow in real time inside the terminal and using the dynamic energy entropy model, enabling transient features (such as mechanical vibration and arc trigger) to be quickly captured, and generating high-dimensional feature warnings through multi-pixel local feature description and orientation histograms, ensuring the sufficiency of image information. Constructing a dynamic hierarchical category model in combination with power grid operation parameters makes the collected sample data more conform to the actual operation state, so that the abnormal conditions of power grid equipment can be better reflected during model training. Therefore, this solution reduces the interference of sparse data during the consumption stage, optimizes the structure of the sample library, and improves the effect of subsequent model training. In addition, since all the work is completed in the peripheral part, the delay and bandwidth consumption caused by remote data transmission are avoided, significantly improving the real-time response ability and stability of the system.
[0047] In step S2 of the above embodiment, the core of transfer learning lies in solving the problem of data distribution differences between the pre-trained model (such as ResNet trained on the ImageNet dataset) and the power grid scenario. Domain adaptation is achieved through the feature space alignment module. Among them, KL divergence-driven parameter fine-tuning calculates the feature distribution difference (KL divergence) between the source domain (ImageNet) and the target domain (power grid sample library). When the difference exceeds the threshold, the channel attention weights of the convolutional layer are dynamically adjusted to strengthen the feature response to the texture of power grid equipment (such as insulator skirts and conductor oxide layers) and suppress the interference of irrelevant backgrounds (such as sky and vegetation). In addition, deformable convolution is introduced into the deep convolutional layer of the backbone network, and the receptive field is adaptively matched to the geometric morphology of equipment components (such as arc edges and irregular cracks) through the offset learning mechanism, improving the geometric robustness of local feature extraction.
[0048] The traditional gradient descent algorithm is prone to falling into local optima in the high-noise scenario of power grid data. This solution uses the AdamGDA optimizer with momentum direction perception. By calculating the included angle θ of the gradient vectors of adjacent iteration steps, when θ>90°, it is determined as parameter oscillation, and the momentum factor β1 is dynamically decayed (such as from 0.9 to 0.7) to suppress the oscillation. In addition, the learning rate is adjusted periodically according to the cosine function, using a larger learning rate at the beginning of training for rapid convergence and fine-tuning at the later stage to improve the accuracy, avoiding the problem of premature convergence.
[0049] Based on traditional geometric transformations (rotation, cropping), a conditional generative adversarial network (cGAN) is introduced to synthesize fault features: by cross-modal feature fusion, visible light images and infrared thermal images are input into the cGAN generator to synthesize mixed images of abnormal states such as overheating points and arc lights; the discriminator network enforces the feature distribution of the generated images to be consistent with that of real fault data through feature similarity losses (such as perceptual loss), ensuring the physical rationality of the enhanced data.
[0050] A hierarchical cross-validation mechanism is designed to divide the samples into mutually exclusive subsets according to the device operating states (normal / overloaded / faulty) to address the class imbalance problem in power grid samples (such as scarce fault samples), ensuring that each fold of the validation set contains representative samples of all state classes; and a Gaussian process regression model is constructed based on a tree-structured Parzen estimator to predict the optimal combination of hyperparameters such as learning rate and batch size, and the search efficiency is improved through an early stopping mechanism (terminating if there is no improvement in the validation loss for 3 consecutive rounds).
[0051] In the actual system deployment, this embodiment uses an embedded aggregation terminal equipped with a GPU as the training platform. The terminal internally integrates a high-definition camera, an image acquisition card, an embedded CPU, and a high-speed memory module, and realizes data transmission between hardware modules through a high-speed bus (such as PCIe, USB 3.0). The key frames and their feature data collected from the power grid monitoring efficient system are stored in a hard disk (SSD) through the local data unit and additionally transmitted to the embedded GPU cluster. The pre-trained CNN model is pre-loaded on the GPU cluster, the model is executed through a dedicated deep learning framework (such as TensorFlow, PyTorch), and the built-in data augmentation module is used to supplement the real-time collected images online. During the corresponding training process, the system adopts a diversified storage and computing method, divides the large-scale sample library into multiple data blocks, and conducts accuracy training and performance evaluation on each data block through k-fold cross-validation; the grid search output module automatically opens the mechanism hyperparameter combination in the background and optimizes the parameter configuration through the statistical validation set table. After training is completed, the generated basic model is saved in the local memory and seamlessly docked with the subsequent model conversion and SDNN deployment module through a dedicated data interface to form a closed-loop data stream. The entire training and tuning are scheduled by the embedded operating system and supported by a dedicated acceleration hardware unit to ensure high ARM process infringement, ensuring that the system operates efficiently and with low power consumption in the power grid real-time monitoring scenario.
[0052] Compared with the prior art, in this embodiment, model training and optimization are implemented on the edge side, which can make full use of local computing resources, thus effectively avoiding the bandwidth bottleneck and high latency problems caused by remote data transmission. By using transfer learning and data augmentation techniques, the obtained basic model shows higher robustness and generalization ability in the power grid monitoring scenario, and can maintain a high recognition accuracy under complex working conditions such as indication, occlusion, and noise. The introduction of cross-validation and grid search techniques makes hyperparameter tuning more refined and systematic, significantly reducing the training task and excessive risk; at the same time, the application of centralized training and hardware acceleration techniques greatly shortens the model training cycle, realizing real-time data input and online model update.
[0053] In step S3 of the above embodiment, the technical principle
[0054] Step S3 of this embodiment is based on a number of cutting-edge technologies, and efficiently converts the basic model into the ONNX format adapted to the SDNN platform. Its core lies in the deep integration of multiple technical features and the optimization of the internal mechanism. First, a heterogeneous computational graph intermediate representation conversion mechanism is adopted. The principle is to abstract the computational graph of the deep learning model from the original framework into a unified intermediate representation form, so that different hardware platforms can share the same operator definition; in this process, a dynamic operator mapping table is embedded, and through the pre-constructed mapping relationship, each operator is matched with the instruction set supported by the hardware underlying layer to achieve dynamic adaptation. This mechanism is based on the formal semantics theory to mathematically describe the operator function, thus ensuring semantic consistency during the conversion process.
[0055] Secondly, the converted computational graph is traversed through a graph structure parsing algorithm. This algorithm uses principles such as the shortest path, connected components, and topological sorting in graph theory to accurately identify the dependency relationships between operators in the model; especially when detecting the adjacent relationship between the batch normalization layer and the convolutional layer, the topological sorting method is used to analyze the node hierarchy and reveal the discontinuity of the memory access pattern. When the system detects that operations such as convolution and normalization cannot cooperate efficiently due to memory alignment problems, the cross-layer operator fusion strategy is triggered. This strategy is based on the theory of numerical stability and matrix operation optimization, combines multiple consecutive operations into a single composite operator, reduces the number of memory reads and writes through joint calculation, and at the same time reconstructs the kernel memory layout method to make the data arrangement meet the bit width requirements of the hardware vector register, thereby improving the parallel operation efficiency.
[0056] Furthermore, regarding the operator compatibility issue, this embodiment introduces a tensor semantic compatibility verification algorithm. Its basic principle is to perform formal verification of the dimensions and data types of the input and output tensors of each operator through symbolic execution technology, and perform static analysis on the computational graph using the predefined constraint conditions in the ONNX operator protocol; when encountering a dynamically shaped tensor, the system adopts a polynomial expansion method to transform the complex non-linear relationship of the non-supported operator into an approximate linear combination, thereby achieving numerical equivalent substitution. This method relies on the numerical approximation theory and ensures the tolerance of the replaced operator in terms of accuracy through approximation error control.
[0057] Finally, this embodiment uses depth-first search to traverse the computational graph, constructs a data dependency chain, and detects the existence of shared weight tensors in the parallel branches. Based on this, the subgraph segmentation strategy is used to split the large computational graph into multiple logically independent and computationally intensive subgraph modules. This segmentation process relies on a tree-like data structure and a graph cut algorithm, which can maximize the local parallelism while ensuring the integrity of data dependencies. Subsequently, the hardware resource allocator allocates independent cache areas for each subgraph according to the subgraph characteristics, realizing data locality optimization and cache utilization. This process is based on the cache coherence protocol and memory management theory to ensure that each subgraph avoids resource competition and data conflicts during parallel execution.
[0058] In a specific implementation, this solution is deployed on a distribution fusion terminal with heterogeneous computing capabilities. Its hardware components include a high-definition surveillance camera, an image acquisition card, a multi-core CPU, a high-performance embedded GPU, an FPGA accelerator, as well as dedicated high-speed buses (such as PCIe and USB 3.0) and a solid-state cache module. The collected basic model data is first transmitted to the embedded server, and the heterogeneous computing graph intermediate representation conversion is performed under the cooperation of the CPU and GPU through a dedicated conversion software module, and a dynamic operator mapping table is embedded. The conversion module uses a customized graph parsing library to read the model calculation flow graph, and detects the dependency relationship between the convolutional layer and the batch normalization layer through topological sorting; if a memory access pattern mismatch is detected, the system calls the cross-layer operator fusion strategy to merge consecutive operators and uses the FPGA acceleration unit to reconstruct the kernel memory layout so that the data format meets the requirements of the hardware vector register. Then, a dedicated tensor semantic verification module uses symbolic execution technology on the embedded CPU to simulate the tensor dimensions and data types of each operator, and calls a polynomial expansion substitution algorithm for dynamic shape operators. Subsequently, the depth-first search algorithm is used to traverse the calculation graph to construct a data dependency chain, identify and split the parallel branches of the shared weight tensors. The system divides the large graph into multiple subgraphs through the subgraph segmentation module, and the hardware resource allocator allocates independent cache areas for each subgraph in the GPU or FPGA cache. Data synchronization and error checking are achieved between modules through a customized communication protocol, and the entire conversion and optimization process is scheduled by the embedded operating system, realizing an automated and efficient model deployment process. All processes are implemented using a mixed programming of C++ and CUDA to ensure meeting the requirements of low latency and high throughput rate during real-time inference.
[0059] Different from the prior art that relies on simple conversion or cloud centralized processing, this embodiment realizes comprehensive verification and optimization of the compatibility of each operator after the model is converted into the ONNX format by embedding a dynamic operator mapping table and symbolic execution simulation, significantly reducing the performance bottleneck caused by operator mismatch. The cross-layer operator fusion strategy and kernel reconstruction effectively merge multiple consecutive operations into a composite operator and optimize the memory layout, making full use of the hardware vectorization ability, thereby greatly improving the operator execution efficiency. Using polynomial expansion to replace the dynamic shape operator ensures that the model still maintains numerical accuracy and stability when processing non-fixed shape data. Through subgraph segmentation and exclusive cache allocation, data locality is improved, parallel execution efficiency is enhanced, and memory access latency is further reduced. Overall, this solution realizes low-latency and high-throughput real-time video inference on the SDNN platform, gives full play to the local computing resources of the power distribution fusion terminal, and avoids the bandwidth bottleneck and high latency problems caused by traditional remote data transmission. In addition, the automated dynamic scheduling and error detection mechanisms implemented at the hardware and software levels of the system ensure continuous and stable operation in complex power grid monitoring scenarios, improve the real-time performance and accuracy of the status monitoring of key power grid infrastructure, and have significant industrial application value and market competitive advantages.
[0060] In step S4 of the above embodiment, by analyzing the joint sensitivity of the absolute value of the weight and the backpropagation gradient, the importance of each convolutional channel is quantified. The magnitude of the weight itself reflects its static contribution, while the gradient change characterizes its dynamic impact on model training. The product of the two is used as the pruning score to accurately identify redundant channels. After pruning the low-score channels, learnable residual connection parameters are introduced to dynamically compensate for the key information lost due to pruning, ensuring that the model accuracy is lossless under a high compression rate. In addition, based on the KL divergence metric in information theory, the distribution differences of the activation values of each layer of the neural network are analyzed. For layers with a large dynamic range and uneven distribution of activation values (such as the input convolutional layer), a higher bit width (8 bits) is retained to maintain detailed information; for deep networks with a smooth distribution (such as fully connected layers), low-bit (4 bits) quantization is used to reduce the computational overhead. By dynamically calibrating the scaling factor, the quantization error is controlled within an allowable range. The convolutional calculation is disassembled into a vectorized instruction sequence supported by the hardware, and fragmented access is eliminated through data chunking and memory alignment strategies. For example, the input feature map is chunked according to the hardware cache line size to ensure that each memory read can completely load continuous data, reducing the number of redundant accesses. Transpose instructions are inserted to adjust the data layout format to match the efficient read mode of the hardware computing unit. Secondly, a three-level cache system is constructed to store high-frequency weights, pre-stored feature maps, and shared intermediate results respectively. Based on the principle of spatio-temporal locality of data access, the data blocks required for future calculations are predicted and pre-loaded into the cache. When the cache hit rate drops, an improved replacement algorithm is used to preferentially retain data that is shared across layers and may be reused recently, significantly improving the cache utilization rate.
[0061] In actual deployment, an integrated power distribution fusion terminal is selected in this embodiment. The terminal hardware integrates a high-definition surveillance camera, an image acquisition card, an embedded CPU, a multi-core GPU, and a dedicated FPGA accelerator. High-speed data transmission is achieved between modules through a high-speed PCIe bus and an internal interconnection network. First, the ONNX format model after basic training is loaded by the embedded server, and the heterogeneous computing graph conversion is started through a customized software module. This module realizes the parsing and reconstruction of each operator of the model, and embeds a dynamic operator mapping table during the conversion process. Subsequently, the gradient sensitivity evaluation module runs on the CPU to calculate the parameter contribution degree of each convolution kernel, uses a preset threshold to eliminate inefficient channels, and realizes information compensation through the residual connection module. Immediately afterwards, the mixed-precision quantization subsystem executes on the GPU, determines the appropriate quantization bit width for each layer through the KL divergence and layer-by-layer calibration algorithm, and generates the corresponding dynamic scaling factor. After the converted model is parsed by the instruction set mapping module, the high-level operators are converted into vectorized instructions supported by the hardware; when data mismatch is detected, the data rearrangement module automatically blocks and reorganizes the input feature map on the FPGA and inserts transpose instructions to achieve data alignment. The built-in cache controller of the system constructs a three-level cache network, which are respectively configured in SRAM, DRAM, and shared memory, and automatically optimizes the cache replacement policy through the spatial and temporal locality algorithm. The entire process is uniformly scheduled by the embedded operating system, and each sub-module is interconnected through a standard API interface to ensure the full automation of the entire process from model pruning, quantization, instruction mapping to cache scheduling. Data and commands are transmitted between each hardware unit through a low-latency communication protocol (such as NVLink, PCIe Gen4), realizing the optimal utilization of computing resources.
[0062] After deploying the model optimization and acceleration technology on the edge side in this step, the parameter redundancy and computational complexity of the deep learning model are significantly reduced, and efficient and low-power real-time video inference is realized. Structured pruning ensures that the network maintains high performance while eliminating inefficient channels through quantitative evaluation and residual compensation; mixed-precision quantization dynamically adjusts the bit width according to the activation distribution, greatly reducing the storage and operation requirements while maintaining data characteristics and accuracy. The introduction of the instruction set mapping and data rearrangement strategies enables the model operation to fully match the hardware vectorization characteristics, effectively eliminating the memory access bottleneck caused by data misalignment; while the hierarchical cache coordination mechanism minimizes the memory access latency and improves the overall system throughput through multi-level cache management. Generally speaking, compared with the traditional solution that relies on a single GPU or cloud processing, this method achieves lower latency, higher energy efficiency, and stronger robustness on the SDNN platform, and is particularly suitable for monitoring tasks with extremely high requirements for real-time performance and reliability in the power grid distribution scenario, which can effectively alleviate the problems of limited network bandwidth and high transmission delay, while greatly reducing the system energy consumption and improving the overall operation efficiency and monitoring accuracy of the power distribution fusion terminal.
[0063] In step S5 of the above embodiment, in the real-time video inference stage on the edge device, a technical means combining multi-level signal preprocessing and statistical detection is adopted to convert the original video data into high-quality feature inputs and achieve abnormal behavior detection under low-latency conditions. Specifically, the adaptive frame sampling module uses time-domain signal processing methods to extract motion amplitude and waveform characteristic information from consecutive video frames, dynamically calculates the motion index of each frame, which reflects the inter-frame energy change, and real-time screens out key frames with high information content through threshold determination, thereby reducing unnecessary redundant data; then, noise suppression uses a Wiener filter constructed based on the classical least mean square error principle to recover the missing or distorted parts in the selected key frames, estimates the true signal according to the noise statistical model, and effectively improves the image signal-to-noise ratio; at the same time, the dynamic range compression module redistributes the pixel values through a non-linear mapping function (such as logarithmic or gamma function), adjusts the gray-scale distribution of the image, and enhances the low-contrast regions. The preprocessed image data is then sent to the SDNN inference model, and the target behavior features output by the model are evaluated in real time by the statistical detection module using the chi-square distribution theory. This method calculates the deviation index by comparing the current frame feature distribution with the preset historical statistical model, and then determines whether there is an abnormal fluctuation; after the abnormal determination is triggered, the system immediately generates and sends a warning signal based on a lightweight communication protocol. Each link is based on mathematical theories and algorithms such as signal detection, statistical inference, non-linear mapping, and optimal filtering to achieve efficient preprocessing of video data and real-time abnormal evaluation.
[0064] In practical applications, this solution is deployed on an edge computing platform integrated into a distribution fusion terminal. The platform includes a high-definition surveillance camera, a data acquisition card, a high-performance embedded processor (cooperative work of multi-core CPU and dedicated GPU), a high-speed storage unit, and a dedicated communication module. The camera collects real-time video streams through a high-speed interface (such as USB 3.0 or GigE). The video signal enters the embedded system through a preprocessing module. The system has a dedicated software module built-in to implement adaptive frame sampling: the edge device performs real-time motion energy and waveform statistics on each frame of data, and dynamically selects key frames based on the calculated key frame index. Subsequently, the Wiener filtering module uses a filtering algorithm accelerated on a dedicated DSP or GPU to recover the areas affected by noise interference or data loss in the key frames. Immediately afterwards, the dynamic range compression module performs normalization and redistribution operations on the pixel data using a preset non-linear mapping function, and outputs an image with improved quality. After preprocessing, each frame of image is subjected to feature extraction and target behavior detection by an SDNN inference model. The model output data is subjected to real-time statistical analysis through a built-in chi-square test module, which runs in the embedded system and calculates the deviation value by comparing the current feature data with the historical standard distribution. When the detection result exceeds the set threshold, the built-in MQTT communication module (supporting low-power wireless or wired networks) in the system immediately sends an abnormal alarm message to the monitoring center; otherwise, the system continues to collect and process subsequent video frames. Each hardware module is connected through a high-speed PCIe or internal interconnection bus. The data stream is uniformly scheduled and high-speed cache managed inside the system to ensure that each preprocessing, inference, and communication module works in cooperation with low latency.
[0065] This solution integrates multiple image pre - processing techniques such as adaptive frame sampling, noise suppression, and dynamic range compression on the edge side, effectively improving the quality and information density of the input data, thus providing a high - precision feature basis for the SDNN inference model. Compared with the traditional monitoring system that relies on cloud - based centralized processing, this embodiment makes full use of local computing resources to achieve in - place data processing, significantly reducing the amount of network - transmitted data and the system response time. The adaptive frame sampling can dynamically adjust the sampling rate to ensure the capture of key frames while reducing redundant data. The application of Wiener filtering effectively eliminates noise interference, and the non - linear dynamic range compression improves the image contrast, providing a more reliable input signal for subsequent model inference. In addition, the statistical detection method based on chi - square test can capture the abnormal fluctuations of behavioral characteristics in real time, improving the sensitivity and accuracy of anomaly detection. Using the MQTT protocol to transmit warning signals greatly shortens the alarm delay, ensuring more timely and reliable security monitoring of the key infrastructure of the power grid. Generally speaking, this solution realizes full - process image pre - processing and abnormal behavior detection on edge devices, reducing the dependence on bandwidth and cloud resources, while improving the real - time performance and stability of the system, providing a solid technical guarantee for the efficient application of distribution fusion terminals in power grid monitoring.
[0066] In step S6 of the above - mentioned embodiment, the adaptive feedback mechanism is based on the reinforcement learning theory and realizes the continuous stability of the model performance through dynamic monitoring and decision optimization. This mechanism uses the sliding window method to statistically analyze the inference accuracy and response delay data within a continuous time period, calculates the standard deviation of the accuracy fluctuation and the average delay index as the quantitative basis for real - time performance. When constructing the state space, multi - dimensional indicators such as model error, delay level, and hardware resource occupancy rate are integrated to form a continuous and discretized state representation. Using the double - deep Q - network, this system evaluates the possible actions (including model parameter update, hyper - parameter adjustment, resource re - allocation, etc.) in each state, and selects the optimal strategy by maximizing the future cumulative reward value. On the other hand, in terms of model re - training, an incremental adversarial training framework is introduced. Samples with confidence lower than a predetermined threshold are screened out from the real - time video stream and fused with the edge cases constructed by the generative adversarial network to construct an incremental data set covering the latest environmental changes. At the same time, while maintaining the previous knowledge through the elastic weight consolidation method, the parameters of the fully - connected layer are fine - tuned in a targeted manner, and the Gaussian process regression model is used to predict the optimal learning rate combination to achieve adaptive hyper - parameter adjustment. When the verification index fails to improve in multiple consecutive iteration cycles, the system will activate the early - stopping strategy and automatically roll back to the parameter snapshot with the optimal performance. Finally, the parallel computing sandbox is used to evaluate the similarity of the outputs of the old and new models to ensure that the updated model meets the predetermined accuracy and robustness requirements before deployment, and seamlessly replaces the model weights running in memory in an atomic operation.
[0067] In a specific implementation, this solution is deployed on an edge computing platform integrated in a distribution fusion terminal. The platform includes high-speed acquisition devices, an embedded multi-core processor, a dedicated GPU and an FPGA acceleration unit, a high-speed cache, and a low-latency communication interface. The acquisition module obtains real-time video data from a surveillance camera. After preprocessing, a preliminary output is generated by the SDNN inference module. Subsequently, the built-in performance monitoring unit in the system samples and statistically analyzes the inference results through a sliding window, and records the accuracy fluctuation and response delay in real time. Data is efficiently transmitted between the processor, GPU, and FPGA through a high-speed bus (such as PCIe or NVLink), and multi-threaded parallel operations are implemented by embedded software. On this platform, the reinforcement learning decision engine runs on a dedicated scheduler. Its state acquisition module obtains hardware load and memory occupancy information in real time through built-in sensors and monitoring software, and constructs a multi-dimensional state vector. The double deep Q-network then makes online decisions on the edge side using a pre-trained model to guide the adaptive update of the model. At the same time, the incremental adversarial training module constructs an incremental data set by screening low-confidence frames and using a preset generative adversarial network model to generate edge samples. This process is completed in a dedicated storage medium, and the elastic weight consolidation module performs local updates of the fully connected layer parameters under the cooperation of the CPU and GPU. The Gaussian process regression module predicts the learning rate using historical tuning data and monitors the validation error during the iteration process. When the validation error has not decreased for several consecutive times, the system automatically triggers an early stopping and rollback mechanism. The updated model is subjected to an output comparison test in an independent sandbox environment, and the seamless switching of the final model weights is achieved through memory atomic write operations, ensuring that the overall update process does not interrupt the existing inference service.
[0068] In this step, by introducing an adaptive feedback mechanism optimized by reinforcement learning, the dynamic self-regulation of the model on the edge side is effectively achieved. It can automatically identify performance degradation and quickly initiate the retraining process when the environment and data distribution change. Compared with traditional static model update methods, this embodiment uses sliding window statistics and double deep Q-network decision-making, enabling the system to accurately capture performance fluctuation signals, thereby reducing false alarms and missed detections, and significantly improving the timeliness and accuracy of anomaly detection. The combination of incremental adversarial training and elastic weight consolidation technology not only ensures the rapid adaptability of the model to new data but also prevents catastrophic forgetting, thus maintaining the stable inheritance of the overall knowledge of the model. By predicting the optimal learning rate through Gaussian process regression, the automation and refinement of hyperparameter adjustment are realized, avoiding long-term ineffective training. The establishment of a parallel computing sandbox and the introduction of atomic write operations ensure seamless switching during the model update process and the continuity of system services. Overall, this solution realizes the real-time self-optimization of the model on edge devices, greatly reducing the dependence on the central server, alleviating the network transmission delay problem, while reducing energy consumption and hardware resource consumption, and significantly improving the response speed, stability, and accuracy of the power grid monitoring system, providing a more efficient and secure condition monitoring solution for critical infrastructure in practical applications.
[0069] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that these specific implementation manners are only illustrative examples. Without departing from the principles and essence of the present invention, those skilled in the art can make various omissions, substitutions, and changes to the details of the above methods and systems. For example, combining the above method steps, thereby performing substantially the same function according to a substantially the same method to achieve substantially the same result, belongs to the scope of the present invention. Therefore, the scope of the present invention is only defined by the appended claims.
Claims
1. A real-time video reasoning method for distribution fusion terminal based on SDNN, characterized by: The following steps are involved: Step S1, collecting original video data, using time domain key frame extraction and local feature description methods to screen target key frames, marking the target area through semi-automatic image segmentation and semantic annotation methods, and building a power grid sample library through a stratified sampling mechanism, including training samples, verification samples and test samples; Step S2: Based on the constructed power grid sample library, a convolutional neural network is combined with a transfer learning strategy, a gradient descent optimization algorithm and a data enhancement mechanism are used to perform model training, and cross-validation and grid search methods are used to perform hyperparameter tuning to obtain a basic model; Step S3: Use the intermediate representation conversion mechanism to convert the trained basic model into the ONNX format, and perform compatibility verification on each operator through the operator fusion and kernel reconstruction algorithm. If the operator verification fails, use the tensor operation approximation algorithm to replace it, and use the subgraph partitioning mechanism to optimize the operator execution order; Step S4: based on the sparsity optimization strategy of SDNN, prune and quantize the converted ONNX format model with low bit precision, and compile the ONNX format model into an SDNN inference model; and deploy the SDNN format model to the distribution fusion terminal through a heterogeneous computing acceleration framework, wherein the heterogeneous computing acceleration framework optimizes memory access through a hierarchical cache mechanism; Step S5: deploy the compiled and optimized SDNN inference model on the edge device of the distribution fusion terminal, process the real-time video stream frame by frame through adaptive frame sampling, noise suppression and dynamic range compression preprocessing mechanism, and use statistical thresholds and anomaly detection algorithms to judge the target behavior in the video in real time. If it is judged to be abnormal, send a warning signal, otherwise, continue to process subsequent video frames; Step S6: monitoring the real-time performance of the SDNN inference model through an adaptive feedback mechanism based on reinforcement learning optimization; If it is detected that the inference accuracy has dropped or the response delay has exceeded the standard, the model will be retrained and updated through incremental learning and adaptive hyperparameter adjustment mechanisms.
2. According to the SDNN-based real-time video reasoning method for distribution fusion terminal according to claim 1, it is characterized by: The working method of step S1 is as follows: first, the pixel motion vector field between consecutive video frames is calculated by the time domain gradient joint detection algorithm, and a key frame screening model based on dynamic energy entropy is constructed in combination with the inter-frame structural similarity index. When the proportion of the optical flow mutation area between adjacent frames exceeds a preset threshold, the transient characteristics of the device including but not limited to mechanical vibration and arc flicker are captured by the key frame interception mechanism; Based on the captured transient features, a local feature description method is adopted. When constructing the scale space in the key frame, an octree segmentation strategy of Gaussian difference pyramid is introduced to divide the image into hierarchical grid units. The feature points in each unit are clustered by directional histogram to generate a multidimensional feature vector with spatial topological constraints. Then, the target area is extracted through a semi-supervised adversarial image segmentation network. Finally, dynamic hierarchical categories are established based on the operating parameters of the distribution equipment, which include but are not limited to load rate and temperature rise rate. Samples are selected in each layer through Latin hypercube sampling to output the power grid sample library.
3. According to the SDNN-based real-time video reasoning method for distribution fusion terminal according to claim 1, it is characterized by: The working method of step S2 is as follows: first, the model parameters are initialized by using a convolutional neural network, and the common features are extracted from the power grid sample library output by step S1 by using a transfer learning strategy; then, the model parameters are iteratively updated by using a gradient descent optimization algorithm, the cross entropy or mean square error loss function is calculated in each training cycle, and the weights are adjusted by using a back propagation mechanism; at the same time, operations including but not limited to rotation, scaling, brightness adjustment, and noise injection are applied in the input data processing process through a data enhancement mechanism; then, the sample library is divided into multiple subsets by a cross-validation method, and the model is trained and validated by using k-fold cross-validation; finally, the learning rate, batch size, regularization coefficient, and convolution layer number parameters are adjusted in a preset hyperparameter space by using a grid search method, and the validation set performance under each set of parameter combinations is evaluated, and the configuration that minimizes the validation error is selected.
4. According to the SDNN-based real-time video reasoning method for distribution fusion terminal according to claim 1, it is characterized by: The working method of step S3 is as follows: first, a dynamic operator mapping table is embedded when converting the basic model into the ONNX format through the intermediate representation conversion mechanism of the heterogeneous computational graph, and then the model computational flow graph is extracted by using the graph structure parsing algorithm, and the adjacent nodes of the batch normalization layer and the convolution layer are identified by topological sorting. If an unaligned memory access pattern is detected, the convolution, batch normalization and activation function sequences are merged into a single composite operator through a cross-layer operator fusion strategy, and the kernel memory arrangement is reconstructed to match the vector register bit width of the hardware; Subsequently, a tensor semantic compatibility verification algorithm is used to verify the compatibility of operators. The tensor semantic compatibility verification algorithm defines constraints based on the ONNX operator protocol, and simulates the input and output tensor dimensions and data types of each operator through symbolic execution. If a dynamic shape tensor is detected, the non-supported operator is numerically replaced by polynomial expansion; finally, the computational graph is traversed using depth-first search to build a data dependency chain. When a shared weight tensor is detected in a parallel branch, the computational graph is split into multiple independently executed subgraph modules through a subgraph partitioning strategy, and an exclusive cache area is allocated to each subgraph through a hardware resource allocator.
5. According to the SDNN-based real-time video reasoning method for distribution fusion terminal according to claim 1, it is characterized by: The SDNN-based sparsity optimization strategy first calculates the L1 norm and gradient product of each convolution kernel weight through a gradient sensitivity analysis algorithm. If the score is lower than a preset threshold, the redundant convolution kernels are removed according to the channel granularity through a structured pruning mechanism, and the information loss is compensated through residual connections. Then, a mixed precision quantization mechanism is used to perform low-bit conversion: the mixed precision quantization mechanism uses KL divergence to analyze the distribution characteristics of the activation values of each layer, allocates 8-bit quantization bit width to the high dynamic range layer, reduces the dimension of the smooth response layer to 4 bits, and generates a dynamic scaling factor through a layer-by-layer calibration algorithm. Subsequently, the convolution operator is disassembled into a vector multiplication and addition instruction sequence supported by the hardware through the instruction set mapping table. When it is detected that the tensor dimension does not match the register bit width, the input feature map is aligned in blocks according to the hardware cache line size through a data rearrangement strategy, and a transpose instruction is inserted to eliminate memory access fragmentation.
6. According to the SDNN-based real-time video reasoning method for distribution fusion terminal of claim 1, it is characterized by: The hierarchical cache coordination mechanism builds a three-level cache structure in the heterogeneous acceleration framework, including D1 cache, D2 cache and D3 cache; the D1 cache is used to store high-frequency access blocks of weight tensors, the D2 cache is used to pre-store input feature maps for the next reasoning cycle, and the D3 cache is used to retain intermediate calculation results for multi-model reuse; When a memory access request arrives, if it is detected that the target data exists in the D2 cache and the hit rate is lower than 85%, a cache replacement algorithm is used to retain cross-layer shared data blocks based on the principle of spatiotemporal locality.
7. According to the SDNN-based real-time video reasoning method for distribution fusion terminal of claim 1, it is characterized by: The working method of step S5 is as follows: based on the original video stream collected by the edge device, an adaptive frame sampling mechanism is first used to calculate the key frame index according to the inter-frame motion amplitude and waveform distribution; after sampling, the key frame is missing processed by using Wiener filtering; in addition, the pixel value distribution is nonlinearly mapped by dynamic range compression; The deleted frame-by-frame images are input into the SDNN inference model, and the behavioral target characteristics are evaluated in real time using the chi-square test method. During the detection process, a comparative analysis is performed based on the prior statistical model and the historical feature distribution. If the deviation exceeds the preset threshold, an abnormal signal is sent using the MQTT protocol.
8. The real-time video reasoning method for distribution fusion terminal based on SDNN according to claim 1 is characterized by: The adaptive feedback mechanism based on reinforcement learning optimization uses a sliding window to statistically infer the standard deviation of accuracy and the average response delay. When the accuracy drops by more than 2% for three consecutive windows or the delay exceeds a preset threshold, a decision engine driven by reinforcement learning is used to build a state space including accuracy deviation, delay level, and hardware resource occupancy rate; and the optimal actions for model update, parameter adjustment, and resource reallocation are calculated through a dual deep Q network.
9. The real-time video reasoning method for distribution fusion terminal based on SDNN according to claim 1 is characterized by: In step S6, when the model triggers retraining and updating, firstly, frames with confidence lower than 0.6 are extracted from the real-time video stream through the incremental adversarial training method, and the frames and the edge cases synthesized by the generative adversarial network are used to form an incremental data set; and the fully connected layer parameters are updated through the elastic weight solidification algorithm; at the same time, a Gaussian process regression model is constructed based on the historical tuning records to predict the optimal combination of learning rates. If the verification loss does not decrease for 5 consecutive iterations, an early stopping mechanism is triggered and the optimal parameter snapshot is rolled back; finally, a parallel computing sandbox is established outside the model inference thread, and the similarity of the new and old model outputs is compared after the incremental training is completed. If the standard is met, the model weights in the memory are replaced by atomic write operations.
Citation Information
Cited By
Relay state prediction and fault early warning method and system based on deep learning
CN120763779A
Intelligent water pollution monitoring method and system based on edge calculation
CN120892927A
Multi-scene interaction control method and system in IPTV education platform
CN120935385A
Text-to-voice hardware acceleration system supporting dynamic input adaptation
CN120977286A
Distributed algorithm model compression method for micro-grid group collaborative optimization, electronic equipment and storage medium
CN121257631A