Method and device for identifying state of power equipment driven by lightweight residual network
Through lightweight residual networks and efficient data enhancement strategies, the problem of high model complexity in power equipment status recognition is solved, efficient recognition on edge devices is achieved, and recognition accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510876228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-16
AI Technical Summary
Existing power equipment status recognition methods have problems such as high model complexity, large computing resource requirements, overfitting and insufficient generalization ability, making it difficult to effectively deploy and identify small-scale remote sensing image datasets on edge devices.
A lightweight residual network is adopted, combined with a dynamic learning rate strategy and mixed precision training. By alternately stacking Skip-ResBlock and Basic-ResBlock, embedding a global average pooling layer and a channel attention module, multi-dimensional data enhancement and real-time inference are performed to optimize network parameters.
While ensuring recognition accuracy, it significantly reduces model complexity, improves data adaptability and training efficiency, is suitable for edge devices, and realizes real-time power equipment status recognition.
Smart Images

Figure CN120656130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for identifying the state of power equipment driven by a lightweight residual network, and belongs to the technical field of power equipment detection and deep learning. Background Art
[0002] With the rapid development of computer vision technology, the number and resolution of high-resolution images have increased significantly, and scene recognition has become a core task in intelligent image processing. Traditional scene recognition methods rely on manual feature extraction, such as texture and shape, but their feature expression capabilities are limited and they are difficult to adapt to the diversity of complex scenes.
[0003] In recent years, deep learning-based convolutional neural networks (CNNs) have made significant progress in image scene recognition, but some challenges remain. For example, typical networks such as ResNet50 and VGG16 have large parameter counts and high computational resource requirements, making them difficult to deploy on edge devices. Furthermore, existing models are prone to overfitting or insufficient generalization for smaller remote sensing image datasets. Furthermore, traditional methods have long training cycles and fail to fully utilize data augmentation techniques to optimize model robustness.
[0004] How to reduce model complexity, improve data adaptability, and improve training efficiency while ensuring recognition accuracy has become a technical problem that needs to be solved in the field of power equipment state recognition. To this end, the present invention proposes a method and apparatus for power equipment state recognition driven by a lightweight residual network. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a method and device for power equipment state identification driven by a lightweight residual network, which can improve the accuracy and efficiency of power equipment state identification.
[0006] The technical solution adopted by the present invention to solve its technical problems is: In a first aspect, an embodiment of the present invention provides a method for identifying power equipment status driven by a lightweight residual network, comprising the following steps: S1, deploys a monocular camera in the power equipment monitoring scene to periodically collect equipment images; S2, standardizes and multi-dimensionally enhances the acquired original device image, including pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement; S3, builds a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections, with a stacking ratio of 1:1 and a total depth of 10 layers, embedding a global average pooling layer and a channel attention module; S4 uses a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network; S5 uses the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
[0007] As a possible implementation of this embodiment, in step S1, a monocular camera is deployed in the power equipment monitoring scene to ensure that the equipment image in the target area is clear and unobstructed, and the equipment image with a resolution of 1920×1080 is collected at a frequency of 1 second / frame.
[0008] As a possible implementation of this embodiment, the standardization and multi-dimensional enhancement processing of the collected original device image includes the following steps: Pixel normalization to the [0,1] interval: Subtract the minimum pixel value from each pixel value in the image, and then divide it by the difference between the maximum and minimum pixel values to achieve linear scaling; Spatial domain enhancement: The image data augmentation tool performs random rotation of ±45 degrees, ±15% translation in width and height, shear strength of 0.2, horizontal flip probability of 0.5, and 0.8-1.2 times scaling; Frequency domain enhancement: Gaussian noise with a standard deviation of 0.05 is superimposed, and high-pass and low-pass mixed filtering is performed through Fourier transform. The high-pass filtering retains the components in the high-frequency range, while the low-pass filtering retains the components in the low-frequency range.
[0009] As a possible implementation of this embodiment, the construction of a lightweight residual network includes the following steps: Residual blocks with skip connections and residual blocks without skip connections are stacked alternately in a 1:1 ratio, with a total depth of 10 layers. The number of channels at the entrance of each residual block with skip connections is reduced to a quarter of the original number of channels through 1×1 convolution; The global average pooling layer is used at the end to replace the fully connected layer, so that the number of model parameters does not exceed 3.5 million; The channel attention module is embedded. First, the feature map of each channel is globally average pooled to obtain the global feature vector of the channel. The global feature vector is then passed through a fully connected layer with a compression ratio of 16 and activated by ReLU. Finally, the global feature vector is restored to the original number of channels through a fully connected layer and activated by Sigmoid to generate the channel attention weight.
[0010] As a possible implementation of this embodiment, the step S4, using a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network, includes the following steps: Initialize the SGD optimizer: momentum = 0.9, weight decay = 0.0005, initial learning rate = 0.001; The cosine annealing strategy is used to dynamically adjust the learning rate, with cycles = 100 and minimum learning rate = 0.0001; Enable mixed precision training (AMP), batch size = 32, training epochs = 200; Use the cross entropy loss function and introduce label smoothing (smoothing=0.1) to alleviate overfitting; The device image after standardization and multi-dimensional enhancement is input into the lightweight residual network for training to obtain the optimized lightweight residual network.
[0011] As a possible implementation of this embodiment, the cosine annealing strategy is: the learning rate starts from an initial 0.001 and decays to 0.0001 according to the cosine function law, and completes a decay cycle every 100 rounds; The mixed precision training is as follows: forward propagation uses 16-bit floating point calculations to speed up operations, and backward propagation gradient accumulation uses 32-bit floating point numbers to ensure accuracy. The batch size is set to 32, and training is performed for 200 rounds.
[0012] As a possible implementation of this embodiment, mixed precision training uses FP16 for forward propagation calculations and FP32 for gradient accumulation. The batch size is set to 32 and the training cycle is 200 rounds, so that the edge device inference speed is ≥25 frames / second. FP16 is used for convolution and activation calculations, and FP32 is used for BN layer statistics calculation and gradient accumulation.
[0013] As a possible implementation of this embodiment, in step S5, real-time reasoning is achieved through the video frame interception thread, the reasoning thread and the result refresh thread, with an end-to-end delay of no more than 50 milliseconds, and the device status classification result is output.
[0014] As a possible implementation of this embodiment, the video frame interception thread is as follows: reading the video stream at 30fps through OpenCV, and triggering 1 second / frame sampling according to the timestamp; The inference thread is as follows: loading a pre-trained model with an input size of 224×224 pixels, and the inference time for a single frame does not exceed 40 milliseconds; The result refresh thread is to update the interface in real time through Tkinter to display the four classification results of normal, overheating, discharge, and mechanical failure.
[0015] In a second aspect, an embodiment of the present invention provides a device for identifying power equipment status driven by a lightweight residual network, comprising: The data acquisition module is used to deploy a monocular camera in the power equipment monitoring scene and periodically collect equipment images; The pre-processing module is used to perform standardization and multi-dimensional enhancement on the acquired original device images. The standardization and multi-dimensional enhancement processing includes pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement. The network construction module is used to construct a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections in a 1:1 stacking ratio, with a total depth of 10 layers, and embedding a global average pooling layer and a channel attention module; Adaptive training module, used to optimize the network parameters of lightweight residual networks using dynamic learning rate strategy and mixed precision training; The recognition execution module is used to use the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the method for identifying the state of an electric power device driven by any lightweight residual network as described above.
[0017] In a fourth aspect, an embodiment of the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, executes the steps of the method for identifying the state of an electric power device driven by any lightweight residual network as described above.
[0018] The beneficial effects of the technical solutions of the embodiments of the present invention are as follows: By constructing a lightweight residual network and combining it with efficient data augmentation strategies and training techniques, this method significantly reduces model complexity while ensuring recognition accuracy, improving data adaptability and training efficiency. This method improves the accuracy and efficiency of power equipment status recognition, making it suitable for resource-constrained scenarios such as edge devices, and enhancing the practicality and reliability of power equipment status recognition.
[0019] By utilizing the method of the present invention, operation and maintenance personnel can intuitively obtain relevant equipment information through computer equipment, identify and classify the status of power equipment in operation and maintenance sites, and perform identification and classification operations on the equipment status by displaying substation operation and maintenance personnel information on computer equipment, thereby realizing digital management and control of substation equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of a method for identifying power equipment status driven by a lightweight residual network according to an exemplary embodiment; Figure 2This is a structural diagram of a device for identifying power equipment status driven by a lightweight residual network according to an exemplary embodiment; Figure 3 The figure is a schematic diagram showing the structure of a lightweight residual network according to an exemplary embodiment. DETAILED DESCRIPTION
[0021] In order to more clearly illustrate the technical features of the present invention, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.
[0022] like Figure 1 As shown, an embodiment of the present invention provides a method for identifying the state of an electric power device driven by a lightweight residual network, comprising the following steps: S1, deploys a monocular camera in the power equipment monitoring scene to periodically collect equipment images; S2, standardizes and multi-dimensionally enhances the acquired original device image, including pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement; S3, builds a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections, with a stacking ratio of 1:1 and a total depth of 10 layers, embedding a global average pooling layer and a channel attention module; S4 uses a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network; S5 uses the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
[0023] As a possible implementation of this embodiment, in step S1, a monocular camera is deployed in the power equipment monitoring scene to ensure that the equipment image in the target area is clear and unobstructed, and the equipment image with a resolution of 1920×1080 is collected at a frequency of 1 second / frame.
[0024] As a possible implementation of this embodiment, the standardization and multi-dimensional enhancement processing of the collected original device image includes the following steps: Pixel normalization to the [0,1] interval: Subtract the minimum pixel value from each pixel value in the image, and then divide it by the difference between the maximum and minimum pixel values to achieve linear scaling; Spatial domain enhancement: The image data augmentation tool performs random rotation of ±45 degrees, ±15% translation in width and height, shear strength of 0.2, horizontal flip probability of 0.5, and 0.8-1.2 times scaling; Frequency domain enhancement: Gaussian noise with a standard deviation of 0.05 is superimposed, and high-pass and low-pass mixed filtering is performed through Fourier transform. The high-pass filtering retains the components in the high-frequency range, while the low-pass filtering retains the components in the low-frequency range.
[0025] As a possible implementation of this embodiment, the construction of a lightweight residual network includes the following steps: Residual blocks with skip connections and residual blocks without skip connections are stacked alternately in a 1:1 ratio, with a total depth of 10 layers. The number of channels at the entrance of each residual block with skip connections is reduced to a quarter of the original number of channels through 1×1 convolution; The global average pooling layer is used at the end to replace the fully connected layer, so that the number of model parameters does not exceed 3.5 million; The channel attention module is embedded. First, the feature map of each channel is globally average pooled to obtain the global feature vector of the channel. The global feature vector is then passed through a fully connected layer with a compression ratio of 16 and activated by ReLU. Finally, the global feature vector is restored to the original number of channels through a fully connected layer and activated by Sigmoid to generate the channel attention weight.
[0026] As a possible implementation of this embodiment, the step S4, using a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network, includes the following steps: Initialize the SGD optimizer: momentum = 0.9, weight decay = 0.0005, initial learning rate = 0.001; The cosine annealing strategy is used to dynamically adjust the learning rate, with cycles = 100 and minimum learning rate = 0.0001; Enable mixed precision training (AMP), batch size = 32, training epochs = 200; Use the cross entropy loss function and introduce label smoothing (smoothing=0.1) to alleviate overfitting; The device image after standardization and multi-dimensional enhancement is input into the lightweight residual network for training to obtain the optimized lightweight residual network.
[0027] As a possible implementation of this embodiment, the cosine annealing strategy is: the learning rate starts from an initial 0.001 and decays to 0.0001 according to the cosine function law, and completes a decay cycle every 100 rounds; The mixed precision training is as follows: forward propagation uses 16-bit floating point calculations to speed up operations, and backward propagation gradient accumulation uses 32-bit floating point numbers to ensure accuracy. The batch size is set to 32, and training is performed for 200 rounds.
[0028] As a possible implementation of this embodiment, mixed precision training uses FP16 for forward propagation calculations and FP32 for gradient accumulation. The batch size is set to 32 and the training cycle is 200 rounds, so that the edge device inference speed is ≥25 frames / second. FP16 is used for convolution and activation calculations, and FP32 is used for BN layer statistics calculation and gradient accumulation.
[0029] As a possible implementation of this embodiment, in step S5, real-time reasoning is achieved through the video frame interception thread, the reasoning thread and the result refresh thread, with an end-to-end delay of no more than 50 milliseconds, and the device status classification result is output.
[0030] As a possible implementation of this embodiment, the video frame interception thread is as follows: reading the video stream at 30fps through OpenCV, and triggering 1 second / frame sampling according to the timestamp; The inference thread is as follows: loading a pre-trained model with an input size of 224×224 pixels, and the inference time for a single frame does not exceed 40 milliseconds; The result refresh thread is to update the interface in real time through Tkinter to display the four classification results of normal, overheating, discharge, and mechanical failure.
[0031] As a possible implementation of this embodiment, the method of using the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results includes the following steps: Use load_model in tensorflow.keras.models to load the optimized lightweight residual network (trained model); Create the main window through Tkinter and use filedialog.askopenfilename to pop up the file selection dialog box for users to select the image file; Use the Pillow library's Image.open() method to open the selected image, and use Image.resize() to adjust the image size to fit the model input. Use TensorFlow's image.img_to_array() to convert the Pillow image object to a NumPy array; Run the model for classification, output the predicted category, and display the results on the interface.
[0032] like Figure 2 As shown, an embodiment of the present invention provides a device for identifying the state of power equipment driven by a lightweight residual network, including: The data acquisition module is used to deploy a monocular camera in the power equipment monitoring scene and periodically collect equipment images; The pre-processing module is used to perform standardization and multi-dimensional enhancement on the acquired original device images. The standardization and multi-dimensional enhancement processing includes pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement. The network construction module is used to construct a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections in a 1:1 stacking ratio, with a total depth of 10 layers, and embedding a global average pooling layer and a channel attention module; Adaptive training module, used to optimize the network parameters of lightweight residual networks using dynamic learning rate strategy and mixed precision training; The recognition execution module is used to use the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
[0033] The dynamic learning rate strategy described in this invention adopts a step-wise annealing learning rate strategy, reducing the learning rate at fixed intervals according to the training stage. At the beginning of training, a relatively large initial learning rate, such as 0.01, is set. As training progresses, the learning rate is multiplied by a decay factor (such as 0.1) after every certain number of epochs (for example, every 10 epochs). This results in a step-wise decrease in the learning rate. In the early stages of training, a relatively large learning rate enables the model to quickly converge to a preferred solution space. As training progresses, gradually reducing the learning rate allows the model to fine-tune more stably as it approaches the optimal solution, avoiding oscillation near the optimal solution caused by an excessively large learning rate, thereby improving the model's ultimate performance.
[0034] As a possible implementation of this embodiment, the data acquisition module may include: Camera configuration unit: Call the OpenCV library (cv2.VideoCapture) through a Python script to initialize the camera device and set the resolution (such as 1920×1080) and frame rate (30fps); Permission management unit: Use the subprocess module to execute system commands (such as v4l2-ctl) to obtain camera recording permissions. If permission acquisition fails, an automatic retry loop is triggered (while loop combined with try-except exception capture); Image extraction unit: read the video stream frame by frame through cv2.VideoCapture.read(), combine the timestamp judgment (time.time()) to extract images at a frequency of 1 second / frame, and save them in JPEG format (cv2.imwrite).
[0035] As a possible implementation of this embodiment, the preprocessing module may include: Pixel normalization unit: normalizes image pixel values to the [0, 1] range through TensorFlow's tf.image.convert_image_dtype function; Spatial enhancement unit: call tensorflow.keras.preprocessing.image.ImageDataGenerator class; Frequency domain enhancement unit: This unit implements frequency domain enhancement based on the SciPy library. For example, it uses scipy.fftpack.fft2 to perform Fourier transform, superimposes Gaussian noise (numpy.random.normal(0, 0.05, size)), and then performs inverse transform to restore the image.
[0036] As a possible implementation of this embodiment, the network construction module constructs a residual block through tensorflow.keras.layers, alternately stacks Skip-ResBlock and Basic-ResBlock, with a total depth of 10 layers, replaces the fully connected layer with GlobalAveragePooling2D, calls tensorflow.keras.layers.Multiply and Dense layers to implement SE Block, and dynamically adjusts the channel weights.
[0037] As a possible implementation of this embodiment, the adaptive training module configures the optimizer, uses tensorflow.keras.optimizers.schedules.CosineDecay to implement cosine annealing, and enables tf.keras.mixed_precision.set_global_policy('mixed_float16') to accelerate training.
[0038] As a possible implementation of this embodiment, the recognition execution module loads the pre-trained model through tensorflow.keras.models.load_model, uses Tkinter to build a GUI, and combines multi-threading (threading module) to achieve real-time video frame interception and asynchronous inference to ensure smooth interface response.
[0039] The present invention achieves multi-dimensional data enhancement through data collection and preprocessing, extracts features through a lightweight residual network, and performs adaptive training to optimize parameter convergence efficiency. It realizes real-time reasoning through engineering design, forming a complete technical chain.
[0040] The core technology of this invention is to construct a lightweight residual network, such as Figure 3As shown in Figure 2, the lightweight residual network mainly includes the following parts.
[0041] 1) Input layer. The input layer receives raw image data and performs preliminary normalization, adjusting the pixel value range to [0, 1], in preparation for subsequent convolution operations. Its structure is relatively simple, primarily performing data preprocessing and format conversion.
[0042] 2) Skip-ResBlock and Basic-ResBlock modules. The Basic-ResBlock module consists of two convolutional layers. The first convolutional layer uses a 3x3 kernel size and is used to extract local features of the image. Assuming the number of input channels is C1, the number of output channels after this convolutional layer becomes C2. The convolution operation is followed by a Batch Normalization (BN) layer to normalize the data, accelerate network convergence, and prevent overfitting. The ReLU activation function is then used to increase the network's nonlinear representation capabilities. The second convolutional layer also uses a 3x3 kernel, maintaining the number of channels at C2, and then passes through a BN layer and ReLU activation function again. The Skip-ResBlock module is similar to the Basic-ResBlock module, but with the addition of skip connections. The internal convolutional layer configuration is the same as the Basic-ResBlock, except that the output of the convolution operation within the module is directly added to the input at the output to achieve cross-layer feature transfer. This skip connection helps solve the gradient vanishing problem in deep networks, allowing the network to learn richer features.
[0043] 3) Frequency domain enhancement module (frequency domain hybrid filtering).
[0044] This module performs a frequency domain transform on the feature map processed by the residual block, using a Fourier transform to convert spatial domain features to the frequency domain. In the frequency domain, a hybrid filter is designed to process different frequency components. Specifically, the high-pass filter cutoff frequency range is set to [ωh1, ωh2] to enhance high-frequency components and highlight edges and details in the image; the low-pass filter cutoff frequency range is set to [ωl1, ωl2] to retain low-frequency components, smooth the image, and reduce noise interference. After hybrid filtering, the frequency domain features are converted back to the spatial domain through an inverse Fourier transform and fused with the original feature map to obtain the enhanced feature map.
[0045] 4) SE Block. The SE Block is connected to the residual block as follows: the feature map output by the residual block serves as the input to the SE Block. First, the feature map is compressed in the spatial dimension through global average pooling to obtain channel descriptors. The channel descriptors are then input into a multi-layer perceptron (MLP) consisting of two fully connected layers. The first fully connected layer compresses the number of channels by a factor of r (the channel compression ratio is 1 / r) and uses the ReLU activation function. The second fully connected layer restores the number of channels to the original number and uses the Sigmoid activation function to obtain a weight coefficient for each channel. Finally, this weight coefficient is multiplied by the feature map output by the residual block in the channel dimension to achieve adaptive weighting of different channel features, enhancing the feature expression of important channels and suppressing the features of unimportant channels.
[0046] 5) Output layer. The output layer consists of a global average pooling layer and a fully connected layer. The global average pooling layer further compresses the feature maps processed by the previous modules in the spatial dimension, generating a fixed-length feature vector. This feature vector is input to the fully connected layer, which uses a linear transformation of the weight matrix to map the features to the final number of classification categories or regression target dimensions, and outputs the prediction result.
[0047] In the initial stages of lightweight residual network construction, the input image passes through a convolutional layer with a 7x7 kernel size, a stride of 2, and 64 channels. This layer is used to quickly extract preliminary image features and reduce the image resolution. This is followed by a max pooling layer with a 3x3 kernel size and a stride of 2 to further reduce the spatial dimensionality of the feature map.
[0048] In the alternating stacking of Skip-ResBlock and Basic-ResBlock, assume the initial input channel count is 64. In the first Basic-ResBlock, the first 3x3 convolutional layer increases the channel count from 64 to 128, while the second 3x3 convolutional layer maintains the channel count at 128. In the following Skip-ResBlock, two 3x3 convolutional layers similarly increase the channel count from 128 to 256 and then maintain it at 256. Subsequent Basic-ResBlock and Skip-ResBlock layers are stacked alternately following a similar pattern, with the channel count increasing by a certain multiple each time, for example, from 256 to 512, then from 512 to 1024, and so on. The convolution kernel size remains at 3x3 throughout the alternating stacking process to continuously extract local image features.
[0049] The Fast Fourier Transform (FFT) algorithm is used to convert the spatial domain feature map to the frequency domain to obtain the frequency domain feature representation. FFT can efficiently map the image from the spatial domain to the frequency domain, making it possible to process different frequency components in the frequency domain.
[0050] The input and global average pooling process is as follows: the feature map output by the residual block is used as the input of the SE block. The feature map size is H × W × C, where H and W are the height and width of the feature map, respectively, and C is the number of channels. Through the global average pooling operation, the spatial information of each channel is compressed into a scalar, resulting in a channel descriptor of size 1 × 1 × C. This step effectively aggregates the information in the spatial dimension into a description in the channel dimension, providing a foundation for subsequent learning of channel weights.
[0051] The channel descriptors are input into a multi-layer perceptron (MLP). The MLP consists of two fully connected layers. The first fully connected layer compresses the number of channels by a factor of r (the channel compression ratio is 1 / r). For example, when r = 16, the number of channels increases from C to C / 16. This layer uses the ReLU activation function to increase the nonlinearity of the network, enabling the model to learn more complex channel weight relationships. The second fully connected layer restores the number of channels to the original number C and uses the Sigmoid activation function to map the output values to the range [0, 1]. This generates the weight coefficient for each channel, forming a weight vector of size .
[0052] The resulting weight vector is multiplied by the original feature map output by the residual block along the channel dimension to achieve adaptive weighting of different channel features. Specifically, each pixel value in each channel of the original feature map is multiplied by the corresponding weight coefficient. The weighted feature map serves as the output of the SE Block and is then connected to the subsequent network structure, passing on the features enhanced by channel attention, improving the network's focus on important features and its ability to express them.
[0053] During mixed-precision training, most calculations in the model's forward propagation process use the FP16 data type. This includes operations such as convolution, activation function calculations, and various linear transformations. Because FP16 occupies less memory and performs faster, it can significantly accelerate the forward propagation process and improve training efficiency. For example, in convolution operations, the data types of the input feature map, convolution kernel, and output feature map are all FP16, significantly reducing memory usage and calculation time while maintaining acceptable computational accuracy. However, for operations with higher precision requirements, such as mean and variance calculations in BatchNormalization, the FP32 data type is used. This is because the calculation results of BatchNormalization have a significant impact on the stability and convergence of the model. Using FP32 ensures computational accuracy and avoids model performance degradation due to precision loss.
[0054] When calculating gradients during backpropagation, the majority of these calculations also use the FP16 data type. Starting from the output layer, the gradients of each layer are calculated sequentially, following the reverse process of forward propagation. Throughout this process, such as the gradient calculations of the convolutional layers and the backpropagation calculations of the activation functions, FP16 is used to leverage its computational speed. However, the FP32 data type is used for gradient accumulation. Because gradients are continuously accumulated and updated during training, using FP32 reduces precision loss during the accumulation process, ensuring the accuracy of gradient information and thus facilitating stable model convergence.
[0055] When updating model parameters, use the FP32 data type. Gradients (FP32), processed by the gradient accumulation and optimizer, are applied to the model parameters (FP32) to update the model's weights and biases. This is because the accuracy of model parameters is critical to model performance. Using FP32 ensures the accuracy of parameter updates and avoids parameter update bias caused by lower-precision data types, which can affect model training and final performance.
[0056] The present invention utilizes adaptive optimizers such as Adagrad, Adadelta, RMSProp, and Adam, which can automatically adjust the learning rate of each parameter based on the update history of model parameters. Taking the Adam optimizer as an example, it combines the advantages of Adagrad and RMSProp, maintaining an adaptive learning rate for each parameter during training. The Adam optimizer dynamically adjusts the learning rate of each parameter by calculating the first-order moment estimate (i.e., mean) and second-order moment estimate (i.e., variance) of the gradient. For frequently changing parameters, a smaller learning rate is assigned to prevent overly drastic parameter updates; for slowly changing parameters, a larger learning rate is assigned to accelerate parameter updates. This adaptive learning rate adjustment enables more efficient model learning at different training stages and for different parameters, improving training stability and convergence speed. It is particularly suitable for training complex deep learning models and large-scale datasets.
[0057] During the implementation of the present invention, the lightweight residual network can also be replaced by introducing a Ghost module and depthwise separable convolution.
[0058] The Ghost module performs a linear transformation on the feature maps generated by the original convolution operation, generating a series of "ghost" feature maps to reduce the computational complexity and parameter requirements of the convolution operation. In the network architecture, some traditional convolutional layers can be replaced with Ghost modules. For example, the first convolutional layer after the input layer, which originally uses a standard 3x3 convolution, can be replaced with a Ghost module. The Ghost module first generates a small number of "original" feature maps through a convolution operation with a small kernel size (such as 1x1). It then generates more "ghost" feature maps through a series of simple linear operations (such as element-wise addition and element-wise multiplication). These feature maps collectively constitute the output of the Ghost module. Compared to traditional convolution, the Ghost module significantly reduces computational complexity and memory usage while maintaining feature representation capabilities, making it particularly suitable for edge devices with limited computing resources.
[0059] Depthwise separable convolution decomposes the traditional convolution operation into depthwise convolution and pointwise convolution. Depthwise convolution performs convolution on each input channel separately, extracting only local spatial features within the channel without changing the number of channels. Pointwise convolution uses a 1x1 convolution kernel to adjust the number of channels and fuse features on the output of the depthwise convolution. Depthwise separable convolution is widely used in networks, such as the MobileNet series, to build lightweight models. In the lightweight networks described in this patent, some convolutional layers in the Skip-ResBlock and Basic-ResBlock can be replaced with depthwise separable convolution. For example, in the first convolution layer of the Basic-ResBlock, the original 3x3 traditional convolution can be replaced with a 3x3 depthwise convolution, followed by a 1x1 pointwise convolution to adjust the number of channels. This replacement significantly reduces computational complexity and parameter count without significantly degrading model performance, improving network efficiency on edge devices.
[0060] This invention utilizes model compression technology to reduce the model's memory usage and computational complexity by converting the model's weights and activations from high-precision data types (such as FP32) to low-precision data types (such as INT8). When applying quantization technology on edge devices, the trained model must first be quantized offline. This can be achieved by employing a post-training quantization approach. After model training, the model is inferred using a small calibration dataset to collect activation distribution information for each layer in the model. Based on this information, quantization parameters, such as the quantization scale factor and zero offset, are determined. The model's weights and activations are then converted to a low-precision data type according to the quantization parameters. During inference, the edge device's hardware accelerator (such as some chips that support INT8 computation) can directly and efficiently compute the quantized model, accelerating inference speed. Furthermore, because low-precision data types consume less memory, the model's storage requirements on the edge device are significantly reduced, making it easier to deploy the model on resource-constrained devices.
[0061] Compared with the prior art, the present invention has the following characteristics: 1. Lightweight architecture design: By alternately stacking Skip-ResBlock and Basic-ResBlock in a 1:1 ratio (total depth of 10 layers), and replacing fully connected layers with global average pooling, the number of parameters is reduced to 3.5 million (an 86% reduction compared to ResNet50), adapting to the computing power limitations of edge devices. 2. Real-time inference optimization: Mixed-precision training (FP16 forward propagation + FP32 gradient accumulation) combined with a multi-threaded asynchronous architecture achieves inference speeds exceeding 25 FPS and end-to-end latency less than 50ms. 3. Small sample adaptability: Three-level data enhancement (pixel normalization + spatial enhancement + frequency domain enhancement) works in conjunction with label smoothing (smoothing = 0.1), achieving a recognition accuracy of ≥ 90% when the proportion of fault samples is ≤ 10%.
[0062] This paper proposes a lightweight residual network architecture, combined with an efficient data enhancement strategy, which significantly reduces the model complexity while ensuring recognition accuracy.
[0063] An embodiment of the present invention provides an electronic device, including a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the method for power equipment state identification driven by any lightweight residual network as described above.
[0064] Specifically, the above-mentioned memory and processor can be general-purpose memory and processor, which are not specifically limited here. When the processor runs the computer program stored in the memory, it can execute the above-mentioned method of power equipment state identification driven by the lightweight residual network.
[0065] Those skilled in the art will understand that the structure of the computer device does not constitute a limitation of the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently.
[0066] Corresponding to the method for starting the above-mentioned application, an embodiment of the present invention also provides a storage medium on which a computer program is stored. When the computer program is run by a processor, the steps of the method for identifying the state of an electric power device driven by any lightweight residual network are executed.
[0067] The startup device of the application provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in the embodiment of the present application are the same as those of the aforementioned method embodiment. For the sake of brief description, for any part not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.
[0068] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0069] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0070] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the purpose of this embodiment based on actual needs.
[0071] In addition, each functional module in the embodiments provided in the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for identifying power equipment status driven by a lightweight residual network, characterized in that: The steps include: S1, deploys a monocular camera in the power equipment monitoring scene to periodically collect equipment images; S2, standardizes and multi-dimensionally enhances the acquired original device image, including pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement; S3, builds a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections, with a stacking ratio of 1:1 and a total depth of 10 layers, embedding a global average pooling layer and a channel attention module; S4 uses a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network; S5 uses the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
2. The method for identifying power equipment status driven by a lightweight residual network according to claim 1, characterized in that: In step S1, a monocular camera is deployed in the power equipment monitoring scene to ensure that the equipment image in the target area is clear and unobstructed, and the equipment image with a resolution of 1920×1080 is collected at a frequency of 1 second / frame.
3. The method for identifying power equipment status driven by a lightweight residual network according to claim 1, characterized in that: The standardization and multi-dimensional enhancement processing of the collected original device image includes the following steps: Pixel normalization to the [0,1] interval: Subtract the minimum pixel value from each pixel value in the image, and then divide it by the difference between the maximum and minimum pixel values to achieve linear scaling; Spatial domain enhancement: The image data augmentation tool performs random rotation of ±45 degrees, ±15% translation in width and height, shear strength of 0.2, horizontal flip probability of 0.5, and 0.8-1.2 times scaling; Frequency domain enhancement: Gaussian noise with a standard deviation of 0.05 is superimposed, and high-pass and low-pass mixed filtering is performed through Fourier transform. The high-pass filtering retains the components in the high-frequency range, while the low-pass filtering retains the components in the low-frequency range.
4. The method for identifying power equipment status driven by a lightweight residual network according to claim 1, characterized in that: The construction of the lightweight residual network includes the following steps: Residual blocks with skip connections and residual blocks without skip connections are stacked alternately in a 1:1 ratio, with a total depth of 10 layers. The number of channels at the entrance of each residual block with skip connections is reduced to a quarter of the original number of channels through 1×1 convolution; The global average pooling layer is used at the end to replace the fully connected layer, so that the number of model parameters does not exceed 3.5 million; The channel attention module is embedded. First, the feature map of each channel is globally average pooled to obtain the global feature vector of the channel. The global feature vector is then passed through a fully connected layer with a compression ratio of 16 and activated by ReLU. Finally, the global feature vector is restored to the original number of channels through a fully connected layer and activated by Sigmoid to generate the channel attention weight.
5. The method for identifying power equipment status driven by a lightweight residual network according to claim 1, characterized in that: The S4 uses a dynamic learning rate strategy and mixed precision training to optimize the network parameters of the lightweight residual network, including the following steps: Initialize the SGD optimizer: momentum = 0.9, weight decay = 0.0005, initial learning rate = 0.001; The cosine annealing strategy is used to dynamically adjust the learning rate, with cycles = 100 and minimum learning rate = 0.0001; Enable mixed precision training (AMP), batch size = 32, training epochs = 200; Use the cross entropy loss function and introduce label smoothing (smoothing=0.1) to alleviate overfitting; The device image after standardization and multi-dimensional enhancement is input into the lightweight residual network for training to obtain the optimized lightweight residual network.
6. The method for power equipment state identification driven by a lightweight residual network according to any one of claims 1 to 5, characterized in that: In step S5, real-time reasoning is achieved through the video frame interception thread, reasoning thread and result refresh thread, with end-to-end delay not exceeding 50 milliseconds, and the device status classification result is output.
7. The method for identifying power equipment status driven by a lightweight residual network according to claim 6, characterized in that: The video frame interception thread is as follows: read the video stream at 30fps through OpenCV and trigger 1 second / frame sampling according to the timestamp; The inference thread is as follows: loading a pre-trained model with an input size of 224×224 pixels, and the inference time for a single frame does not exceed 40 milliseconds; The result refresh thread is to update the interface in real time through Tkinter to display the four classification results of normal, overheating, discharge, and mechanical failure.
8. A device for identifying the state of power equipment driven by a lightweight residual network, characterized in that: include: The data acquisition module is used to deploy a monocular camera in the power equipment monitoring scene and periodically collect equipment images; The pre-processing module is used to perform standardization and multi-dimensional enhancement on the acquired original device images. The standardization and multi-dimensional enhancement processing includes pixel normalization to the [0, 1] interval, spatial enhancement, and frequency domain enhancement. The network construction module is used to construct a lightweight residual network by alternately stacking Skip-ResBlock with skip connections and Basic-ResBlock without skip connections in a 1:1 stacking ratio, with a total depth of 10 layers, and embedding a global average pooling layer and a channel attention module; Adaptive training module, used to optimize the network parameters of lightweight residual networks using dynamic learning rate strategy and mixed precision training; The recognition execution module is used to use the optimized lightweight residual network to capture video frames in real time and perform asynchronous inference to output device status classification results.
9. An electronic device, characterized in that: It includes a processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the method for power equipment state identification driven by a lightweight residual network as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for identifying the state of an electric power device driven by a lightweight residual network as claimed in any one of claims 1 to 7.