Pump room anomaly detection method based on lightweight model

By building a lightweight model Yolov5s-NSFnet, using Fasternet and channel attention mechanism SE, and replacing the loss function with NWD, the problem of insufficient memory and slow detection speed of deep learning models running on embedded devices is solved, and stable operation and real-time detection on embedded devices is achieved.

CN120014540APending Publication Date: 2025-05-16SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510008808.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing deep learning object detection model is running on embedded devices due to the large amount of parameters and high calculations, resulting in insufficient memory and inability to detect in real time.

Method used

By building a lightweight model Yolov5s-NSFnet, using Fasternet as the backbone network, adding the channel attention mechanism SE, and replacing the loss function with NWD, to reduce the amount of model parameters and calculations, and at the same time improve detection accuracy.

Benefits of technology

The ability to operate stably on embedded devices and detect targets in real time is realized, the amount of parameters and calculations is reduced, the detection speed is improved, and the detection effect is shown on the pump room data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014540A_ABST
    Figure CN120014540A_ABST
Patent Text Reader

Abstract

The invention discloses a pump room anomaly detection method based on a lightweight model. The method comprises the following steps: constructing a Yov5s-NSFnet, training the lightweight model Yov5s-NSFnet on a pump room data set, applying the trained Yov5s-NSFnet on pump room embedded equipment, and detecting an anomaly image recorded by a pump room monitoring camera; wherein a backbone network in the Yolov5s-NSFnet network structure is a lightweight class network structure, namely, Fasnet, and the backbone network in the Yolov5s-NSFnet network structure is a The method comprises the following steps: adding a channel attention mechanism SE in a first Fasnet block, and extracting the weight of a feature map; by replacing the backbone network, the model parameter quantity and the calculation quantity are reduced, and the detection speed is increased; sE channel attention and a replacement loss function are added to be NWD, so that the precision of the lightweight model is ensured, and the improved model can stably run on embedded equipment and detect a target in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection based on deep learning, and in particular to a pump room anomaly detection method based on a lightweight model. Background Art

[0002] In recent years, with the rapid economic development and continuous improvement of urbanization level in my country, the number of medium and high-rise buildings in domestic cities has continued to increase, and secondary water supply for high-rise buildings has become an important part of urban water supply. Secondary water supply for urban tap water or drinking water refers to the water supply method of taking water from the public water supply pipeline in the city or town, pressurizing it in the water storage facilities, and then supplying it to the end of the pipe network, that is, the user. The intrusion of outsiders or animals may cause the secondary water supply equipment to fail to operate normally. In order to provide users with more stable and safe drinking water, abnormal monitoring in the secondary water supply pump room is very important.

[0003] The monitoring system of the pump room is usually an embedded device. The embedded system is specially designed to provide services for specific needs. The system has the characteristics of tailorable software and hardware, small size and high real-time processing. However, the current deep learning target detection model requires large-scale cluster computing and neural network training, which is too large for edge devices. If it is directly deployed to embedded devices, it will face problems such as inability to run and insufficient storage. Therefore, it is crucial to study small and efficient lightweight target detection models and combine them with embedded devices.

[0004] Therefore, in order to enable the deep learning model to run stably and detect in real time on embedded devices such as pump room monitoring, the present invention proposes a pump room anomaly detection method based on a lightweight model. Summary of the invention

[0005] In view of this, the present invention provides a pump room anomaly detection method based on a lightweight model to solve the problems of large memory occupation and slow detection speed of the original Yolov5s model on embedded devices.

[0006] The technical solution of the present invention to solve the above technical problems is as follows: A pump room abnormality detection method based on a lightweight model, comprising:

[0007] Build Yolov5s-NSFnet, train the lightweight model Yolov5s-NSFnet on the pump room dataset, apply the trained Yolov5s-NSFnet on the embedded devices in the pump room to detect abnormal images recorded by the pump room surveillance cameras;

[0008] The backbone network in the Yolov5s-NSFnet network structure is a lightweight network structure Fasternet;

[0009] The network structure Fasternet has four hierarchical levels, each of which is preceded by an embedding layer or a merging layer for spatial downsampling and channel number expansion;

[0010] Each level has a Fasternet block for network feature extraction;

[0011] Each Fasternet block has one PConv layer, two PWConv or Conv 1×1 layers; the PConv or Conv 1×1 layer in the middle layer has a bn layer and a Relu layer;

[0012] Add the channel attention mechanism SE to the first Fasternet block to extract the weight of the feature map.

[0013] Specifically, the method for obtaining the pump room data set is:

[0014] Collect abnormal images in real time at the pump room through the monitoring camera of the pump room; abnormal images include images with helmets and smoking objects;

[0015] The helmets and smoking objects in the images are labeled to create a data set.

[0016] Specifically, the channel attention mechanism SE is added to the first Fasternet block to extract the weight of the feature map:

[0017] The input feature map passes through the SE channel attention module to obtain the output feature map and the weights of each channel;

[0018] Then arrange the weights in descending order and reorder the feature maps of the channels corresponding to each weight;

[0019] Finally, when passing PConv, the cp used for spatial feature extraction acts on the first cp channels with larger weights, and the remaining c-cp channels are spliced ​​to obtain the final output feature map.

[0020] Specifically, the loss function of the Yolov5s network structure is adjusted to NWD to improve the detection accuracy of small targets;

[0021] The NWD is:

[0022] The bounding box is modeled as a 2D Gaussian distribution, and the similarity between the bounding boxes is calculated through the Gaussian distribution corresponding to the bounding box; assuming that the bounding box R = (cx, cy, w, h), for two bounding boxes R a =(cx a ,cy a ,w a,h a ), R b =(cx b ,cy b ,w b ,h b ) Its second-order Wasserstein distance is defined as:

[0023]

[0024] The normalized exponential is used to obtain the normalized Wasserstein distance:

[0025]

[0026] Where C is a constant related to the data set, and the loss function L based on NWD NWD :

[0027] L NWD =1-NWD(N a ,N b ).

[0028] The present invention provides a pump room anomaly detection method based on a lightweight model. The method reduces the model parameters and calculation amount by replacing the backbone network, thereby improving the detection speed. SE channel attention is added and the loss function is replaced with NWD to ensure the accuracy of the lightweight model. The improved model of the present invention can run stably on embedded devices and detect targets in real time.

[0029] The lightweight model Yolov5s-NSFnet has a good detection effect on the pump room dataset, and on embedded devices such as Jetson, it solves the problems of insufficient memory and inability to detect in real time that may result from directly deploying deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flow chart of a pump room abnormality detection method based on a lightweight model provided by the present invention;

[0031] Figure 2 It is a schematic diagram of the specific structure of Fasternet of the present invention;

[0032] Figure 3 This is a schematic diagram of the structure of the block after adding SE in the embodiment disclosed in the present invention;

[0033] Figure 4 It is a schematic diagram of the structure of the improved network Yo l ov5s-NSFnet in the disclosed embodiment of the present invention;

[0034] Figure 5It is the detection data of different scenes of the pump room in the actual project in the embodiment disclosed in the present invention;

[0035] Figure 6 This is a result image detected in an actual project using the method of the present invention. DETAILED DESCRIPTION

[0036] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0037] like Figure 1 As shown, the present invention provides a pump room anomaly detection method based on a lightweight model, including: constructing Yolov5s-NSFnet, training the lightweight model Yolov5s-NSFnet on a pump room data set, applying the trained Yolov5s-NSFnet on a pump room embedded device, and detecting abnormal images recorded by a monitoring camera in the pump room;

[0038] The specific steps include:

[0039] S1: Collect real-time abnormal images of the pump room through the pump room monitoring camera:

[0040] The surveillance video in the secondary water supply pump room is frame-extracted, and images with safety helmets and smoking targets are collected.

[0041] S2: Label the helmet and smoking target in the image to create a data set:

[0042] The helmets and smoking targets were labeled to create a data set, which was divided into training set, validation set and test set in a ratio of 7:2:1 for the subsequent S6 model training and testing part.

[0043] S3: To solve the problem that the pump room monitoring equipment has limited computing power and cannot detect in real time, the backbone network in the Yolov5s network structure is replaced with a lightweight network structure Fasternet to improve the model detection speed:

[0044] For the backbone network structure of Yolov5s network, the backbone network of the original Yolov5s network is replaced with the lightweight network structure Fasternet, which reduces the number of parameters and computation of the overall network. Fasternet is a new neural network proposed based on partial convolution PConv. It runs much faster than other networks on various devices, and has a simple and non-cumbersome architecture, making it hardware-friendly overall.

[0045] The overall structure of Fasternet is as follows Figure 2As shown in the figure, Fasternet has four hierarchical levels, each of which is preceded by an embedding layer (embedding) with a 4×4 standard convolution with a stride of 4 (4×4 standard convolution with a stride of 4) or a merging layer (merging) (2×2 standard convolution with a stride of 2) for spatial downsampling and channel number expansion. Each stage has a Fasternet block for network feature extraction. The Fasternet block structure is shown in the figure below. Figure 3 As shown, the information from all channels is fully and effectively utilized. The Fasternet block is shown together as an inverted residual block, where the intermediate layers have an expanded number of channels and shortcut connections are placed to reuse input features. Normalization layers and activation layers are also essential for high-performance neural networks. However, excessive use of normalization layers and activation layers throughout the network may limit the diversity of features, thereby harming performance and slowing down overall computation. Therefore, the Fasternet block only uses bn layers and activation layers in the PWConv of the intermediate layer to maintain feature diversity and achieve lower latency.

[0046] S4: To address the detection difficulties in the pump room due to dim light and small targets, the SE channel attention mechanism is added to the first block of Fasternet to improve the model detection efficiency;

[0047] After replacing the backbone network of Yolov5s with Fasternet, the number of parameters and the amount of calculation are greatly reduced, but the model detection accuracy is also reduced. Since PConv only uses the first or last cp channel for spatial feature extraction, the feature map with relatively large weights in the middle may not be used, resulting in the loss of feature map information. Therefore, in response to this problem, the present invention extracts the weight of the feature map through the channel attention mechanism SE (SqueezeExcite), and applies the cp in PConv to the feature map with larger weights to reduce the loss of model accuracy. The structural diagram of the channel attention mechanism SE is shown in the figure. Figure 3 shown.

[0048] The input feature map first passes through the SE channel attention module to obtain the output feature map and the weights of each channel, then the weights are arranged in descending order, and the feature maps of the channels corresponding to each weight are also reordered (channel rangement), so that the channels with larger weights will be in front and the channels with smaller weights will be in the back. Finally, when passing through PConv, the cp used for spatial feature extraction will act on the first cp channels with larger weights, and the remaining c-cp channels will be spliced ​​to obtain the final output feature map. The improved PConv can effectively apply the cp of the spatial extraction part to more important channels, ensuring the detection accuracy and effect of the network after lightweighting with PConv.

[0049] S5: Replace the original Yolov5s loss function with NWD;

[0050] The Intersection over Union (IoU) based metric is very sensitive to positional deviations of tiny objects and can significantly degrade detection performance when used in anchor-based detectors. The sensitivity of IoU to objects of different scales varies greatly. For tiny objects with low pixels, a small positional deviation can lead to a significant drop in IoU, resulting in inaccurate label assignment. However, for normal objects with high pixels, the IoU changes slightly with the same positional deviation. This can lead to a decrease in the accuracy of the network on small object detection in the pump room dataset, so for small object detection in the pump room, the NWD (Normalized Wasserstein Distance) loss function is used instead of IoU to improve the detection accuracy of small objects.

[0051] NWD is a new metric to calculate the similarity between boxes. First, the bounding box is modeled as a 2D Gaussian distribution, and the similarity between the bounding boxes is calculated through the Gaussian distribution corresponding to the bounding box. Assuming the bounding box R = (cx, cy, w, h), for two bounding boxes, the second-order Wasserstein distance can be defined as:

[0052]

[0053] Use the normalized exponent to get a new metric - the normalized Wasserstein distance:

[0054]

[0055] Based on NWD loss:

[0056] L NWD =1-NWD(N p ,N g )

[0057] Compared with traditional IOU, NWD is insensitive to the scale of the target, has a smooth change in position difference, and has the ability to measure the similarity of non-intersecting boxes. It is more stable for detecting small targets and helps improve the accuracy of small target detection. Therefore, the loss function of the original YOLOv5s is replaced with NWD to improve the detection accuracy of small targets.

[0058] S6: Train the lightweight model Yolov5s-NSFnet on the pump room dataset; apply the trained Yolov5s-NSFnet on the embedded device of the pump room to detect abnormal images recorded by the monitoring camera of the pump room;

[0059] This implementation also uses multiple indicators to evaluate the above lightweight model Yolov5s-NSFnet:

[0060] (1) Mean Average Precision (mAP)

[0061] mAP is a common and important evaluation indicator in target detection. mAP refers to the average AP of all categories in all images. Its calculation formula is as follows:

[0062]

[0063] Among them, AP represents the accuracy of a single category, which can be used to measure the detection performance of the algorithm on a single category. The calculation formula of AP is as follows:

[0064]

[0065] To calculate the AP value for a certain category, you need to first calculate the detection precision and recall. Precision refers to the ratio of the actual number of positive samples in the predicted sample to the total number of positive samples. The Precision calculation formula is as follows:

[0066]

[0067] In the formula, TP represents the number of predicted samples that are judged as positive examples and are actually positive examples, and FP represents the number of predicted samples that are judged as positive examples but are actually negative examples.

[0068] Recall refers to the ratio of the actual number of positive samples in the predicted samples to all predicted samples. The calculation formula of Recall is as follows:

[0069]

[0070] In the formula, FN represents the number of positive examples judged as negative examples in the prediction samples.

[0071] This article will use mAP 0.5 (mAP when IOU is greater than 0.5) and mAP 0.5:0.95 (The average mAP when IOU is greater than 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, and 0.95 respectively) are used as two indicators to measure the detection effect.

[0072] (2) Parameters

[0073] Parameter quantity refers to the number of parameters that need to be learned in the model, usually including weights and biases. These parameters are adjusted during the model training process to gradually improve the accuracy of the task. This article will analyze the complexity and performance of the model through the number of model parameters.

[0074] (3) Computational Amount - Floating Points Operations (FLOPs)

[0075] FLOPs, or floating point operations per second, is one of the important indicators for measuring the computational complexity of neural networks. In a neural network, each neuron needs to perform certain floating point operations, including addition, multiplication, activation functions, and other operations. In convolutional neural networks, operations such as convolutional layers and pooling layers also require a large number of floating point operations. Therefore, the number of floating point operations can reflect the computational complexity of a neural network and is of great significance for evaluating the performance of a neural network.

[0076] (4) Frame rate (FPS)

[0077] FPS refers to the number of image frames that the network can process per unit time, usually expressed in "frames per second". FPS is one of the important indicators to measure the detection speed of the target detection network. The higher it is, the faster the network processing speed is, and the faster the target detection can be performed on the image.

[0078] The detection speed of the target detection network is usually obtained by actually testing the network on a specific hardware platform (such as CPU, GPU, FPGA, etc.). In the test, the image sequence to be detected needs to be input into the network, and then the time required for the network to process these images is recorded, and finally the FPS is obtained by calculating the number of image frames that can be processed per second. It should be noted that different hardware platforms and image sizes will affect the FPS, so when conducting FPS tests, it is necessary to ensure the consistency of the test conditions so that the detection speeds of different networks under the same conditions can be compared.

[0079] The method provided by the present invention is applied to practical scenarios, such as Figure 4As shown in Figure 1, the experimental data comes from the real-time scene detection data of the pump room in the actual project. The pump room magnetic leakage image dataset used in the experiment contains 6400 infrared fill-in images and color images in different scenes.

[0080] Figure 5 The following is the detection result of the improved network Yolov5s-NSFnet. Compared with the original YOLOv5s, the lightweight network structure designed in this implementation scheme achieves a 78.5% decrease in parameters and an 80.4% decrease in computational complexity while maintaining a reasonable accuracy loss on the pump room dataset. The detection speed is increased from 17.3FPS to 24.8FPS on Jetson, meeting the requirements of real-time detection and making it easier to deploy and run on embedded devices with limited computing power.

[0081] The above method provided by the present invention is applied in the actual environment of pump room abnormality detection, and comprises the following steps:

[0082] 1. Data collection: collect abnormal images of the pump room and annotate them to form a data set for subsequent model training and testing.

[0083] 2. Design a lightweight model structure and train the model on Windows.

[0084] 3. Configure the operating environment in the pump room monitoring embedded device system and deploy the trained model to the device.

[0085] 4. Input the images in the test set.

[0086] 5. Apply the model to detect the input image, obtain the detection results, and complete the model test.

[0087] 6. Connect the pump room camera and apply the model to the images collected by the camera for real-time detection.

[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A pump room anomaly detection method based on a lightweight model, characterized in that: include: Build Yolov5s-NSFnet, train the lightweight model Yolov5s-NSFnet on the pump room dataset, apply the trained Yolov5s-NSFnet on the embedded devices in the pump room to detect abnormal images recorded by the pump room surveillance cameras; The backbone network in the Yolov5s-NSFnet network structure is a lightweight network structure Fasternet; The network structure Fasternet has four hierarchical levels, each of which is preceded by an embedding layer or a merging layer for spatial downsampling and channel number expansion; Each level has a Fasternet block for network feature extraction; Each Fasternet block has one PConv layer, two PWConv or Conv 1×1 layers; the PConv or Conv 1×1 layer in the middle layer has a bn layer and a Relu layer; Add the channel attention mechanism SE to the first Fasternet block to extract the weight of the feature map.

2. A pump room abnormality detection method based on a lightweight model according to claim 1, characterized in that: The method for obtaining the pump room data set is: Collect abnormal images in real time at the pump room through the monitoring camera of the pump room; abnormal images include images with helmets and smoking objects; The helmets and smoking objects in the images are labeled to create a data set.

3. A pump room abnormality detection method based on a lightweight model according to claim 1, characterized in that: The channel attention mechanism SE is added to the first Fasternet block to extract the weight of the feature map: The input feature map passes through the SE channel attention module to obtain the output feature map and the weights of each channel; Then arrange the weights in descending order and reorder the feature maps of the channels corresponding to each weight; Finally, when passing PConv, the cp used for spatial feature extraction acts on the first cp channels with larger weights, and the remaining c-cp channels are spliced ​​to obtain the final output feature map.

4. A pump room abnormality detection method based on a lightweight model according to claim 1, characterized in that: Adjust the loss function of the Yolov5s network structure to NWD to improve the detection accuracy of small targets; The NWD is: The bounding box is modeled as a 2D Gaussian distribution, and the similarity between the bounding boxes is calculated through the Gaussian distribution corresponding to the bounding box; assuming that the bounding box R = (cx, cy, w, h), for two bounding boxes R a =(cx a ,cy a ,w a ,h a ), R b =(cx b ,cy b ,w b ,h b ), its second-order Wasserstein distance is defined as: The normalized exponential is used to obtain the normalized Wasserstein distance: Where C is a constant related to the data set, and the loss function L based on NWD NWD : L NWD =1-NWD(N p ,OF g )。