LightYOLOv4-based termite detection device and method for dams
By improving the backbone network, fusion layer, and detection head of the YOLOv4 model, and combining it with a termite attracting device, efficient and low-cost termite detection of dams was achieved. This solved the problems of high false alarm rate and low efficiency of existing detection methods, and improved detection accuracy and robustness.
Patent Information
- Application Number
- CN202311582638.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-11-24
AI Technical Summary
Existing termite detection methods suffer from high false alarm rates, high costs, long detection times, and require professional operation. Furthermore, the general-purpose YOLOv4 model is not efficient and accurate enough in termite detection, making it difficult to meet the needs of termite detection in dams.
The LightYOLOV4 network model is adopted. By improving the backbone network, fusion layer and detection head, and combining it with a termite attracting device, real-time image acquisition and recognition are performed. The lightweight MobileNetv1 network is used as the backbone network to construct a lightweight PANet network. Softpool is used to replace the max pooling operation in the fusion layer. A detection head module is added to improve the accuracy and efficiency of termite recognition.
It achieves efficient and low-cost termite detection with an accuracy rate of over 95%, reduces the number of model parameters, improves detection efficiency and robustness, is suitable for multi-point detection of dams, and has wireless transmission and long battery life capabilities.
Smart Images

Figure CN117581841B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of termite control technology, and more particularly to termite detection technology. Background Technology
[0002] I. The harm caused by termites.
[0003] The damage caused by termites to dikes:
[0004] (1): Termites weaken and damage the structure of dams by gnawing on the wooden structure and tree roots. Termites prefer to live in humid and warm environments, so the bottom and surrounding areas of reservoir dams are particularly vulnerable to termite infestation.
[0005] (2) Termites destroy the integrity of the dam by building a large number of tunnels and nests, causing the soil and materials inside the dam to gradually loosen, which may trigger disasters such as mudslides, landslides and collapses.
[0006] (3) Termites consume cellulose in the wood structure, reducing the dam's earthquake resistance and stability, making the dam more vulnerable to natural disasters such as earthquakes.
[0007] (4) Termite activity also has a negative impact on the ecological environment around the dam. They destroy a large number of vegetation roots, leading to soil erosion and water loss, and causing damage to the ecosystem.
[0008] Therefore, termite damage to reservoir dams not only directly threatens the structure and stability of the dams but also has a long-term impact on the surrounding ecological environment. Strengthening termite detection, prevention, and control is of great significance for ensuring the safe operation of reservoir dams and maintaining ecological balance.
[0009] II. Existing termite detection technologies.
[0010] Termite detection is the starting point for prevention and control. If termites cannot be detected in time, it will be difficult to carry out targeted prevention and control measures. Existing termite detection methods include:
[0011] (1) Electronic detectors: Using electronic devices (such as termite detectors) to detect termites is a common method. Some electronic devices can detect vibrations and sounds produced inside termites, thereby determining their location and activity. However, electronic detection devices suffer from problems such as false alarms, short lifespan, and high cost.
[0012] (2) High-precision geophysical methods: This method uses geophysical instruments (such as seismographs, radar, infrared scanners, etc.) to directly locate the main nest of termites. This method is generally suitable for detecting dams and can provide accurate location information. This type of method requires specialized equipment and skilled personnel to operate and analyze the data, which places high demands on both equipment and personnel. Moreover, it is less effective in detecting termites in some complex underground media.
[0013] (3) Extracellular matrix detection: Methods for detecting termites using extracellular matrix are under research. This method determines the presence and distribution of termites by analyzing extracellular substances produced by termites, such as chemical substances and volatile substances. It is a technology under research and is not yet mature.
[0014] These detection methods each have their own advantages and limitations. Many methods require the cooperation of professional personnel, have low robustness, high cost, and the detection results are not intuitive, and the detection time is long.
[0015] Current termite control techniques often use attractants (such as wood, which termites like to eat) to lure out and kill the termites. While attracting termites to detect them is a good approach, attractants often also attract other organisms that are interested in the same attractant, such as woodlice, beetles, borers, and even some birds that also feed on wood. Therefore, relying solely on the presence of such organisms to determine the presence of termites often leads to misjudgments.
[0016] Therefore, it is necessary to introduce machine vision to identify whether the lured creatures are termites.
[0017] III. Introduction to the existing machine vision object detection model YOLOv4.
[0018] YOLOv4 is a popular object detection model consisting of a backbone network, a fusion layer, and a head.
[0019] The fusion layer (Neck) contains an SPP layer, which is used to address the problem of inconsistent input image sizes, increase the model's receptive field, and enhance the model's ability to understand the target.
[0020] The YOLOv4 processing procedure is as follows: 1. The image to be detected is passed as input to the backbone network, which extracts features from the image to obtain the extracted feature map. 2. The feature map extracted by the backbone network is passed to the fusion layer, which fuses feature maps from different levels and scales to capture target information at different scales and generate a fused feature map. 3. The fused feature map is passed to the detection head, which further processes the fused feature map to generate the bounding box and class probability of the target, and outputs the recognition result, including the target's class and location information.
[0021] IV. Problems with the existing YOLOv4 target detection model in termite detection.
[0022] The general-purpose YOLOv4 target detection model is widely applicable to various target detection methods. However, it has a large number of parameters, and some of these parameters may still consume computational resources when used for termite detection, even if they are not used. Since it has not been optimized for specific tasks such as termite detection, both its detection efficiency and accuracy need to be improved.
[0023] The technical idea of this invention is to make targeted improvements to the YOLOv4 target detection model to detect termites on dams, thereby reducing detection costs and personnel requirements, and improving detection efficiency and robustness of the detection method. Summary of the Invention
[0024] The purpose of this invention is to provide a termite detection device for dams based on LYOLOV4, which reduces detection costs and improves detection accuracy and robustness of the detection method.
[0025] To achieve the above objectives, the LightYOLOV4-based termite detection device for dams of the present invention includes a housing and a computer. The housing includes a top plate, a bottom plate, and side plates connected circumferentially between the top plate and the bottom plate.
[0026] The enclosure consists of multiple sets, each with the same structure. Multiple entry points for termites and insects are spaced apart at the bottom of the side panels. Termite attractants are placed on the bottom plate inside the enclosure. The lower surface of the top plate is equipped with a light source facing the bottom plate, a camera facing the bottom plate, a first wireless communication module for communication with a computer, and a power module. The power module is connected to the light source, camera, and first wireless communication module and supplies power to all three. A solar panel is fixedly connected to the top plate, and the solar panel is connected to a solar controller. The solar controller is connected to the power module and supplies power to it.
[0027] The computer is connected to a second wireless communication module. The first wireless communication module of each enclosure is connected to the second wireless communication module of the computer and is used to transmit the images captured by the camera of each enclosure to the computer. The computer stores a LightYOLOV4 network model for identifying termites. The computer is also connected to a monitor and speakers.
[0028] This invention also discloses a corresponding termite detection method, which uses the aforementioned LightYOLOV4-based dam termite detection device and proceeds according to the following steps:
[0029] The first step is modeling and training to obtain the trained LightYOLOV4 network model.
[0030] The second step is to use multiple boxes to simultaneously acquire and transmit images; each box is numbered, and the numbering information of each box is stored in the computer.
[0031] The third step is to use the trained LightYOLOV4 network model to identify images captured in real time by the camera, and to issue an alert through a speaker when termites are detected.
[0032] The first step includes the following sub-steps:
[0033] The first sub-step is to create a labeled dataset, in which the location and category of each termite are labeled in each image of the labeled dataset;
[0034] The second sub-step is data preprocessing, which involves image scaling and data augmentation;
[0035] The third sub-step is to construct a LightYOLOV4 network model for real-time detection of surface defects in termites.
[0036] The fourth sub-step is to train the LightYOLOv4 network model;
[0037] The fifth sub-step is to save the LightYOLOv4 network model.
[0038] In the first sub-step of the first step, the termite images are manually labeled using the open-source deep learning annotation tool LabelImg;
[0039] The location of each termite is marked in each image of the labeled dataset by using the rectangle annotation tool in the LabelImg toolbar to mark the area where the termite is located.
[0040] When labeling categories, there are two preset categories: Termite and Others;
[0041] Each termite is manually labeled to distinguish its category, forming a labeled dataset, and finally the labeled file is saved.
[0042] In the second sub-step of the first step, the image is scaled up specifically as follows:
[0043] The width and height of each collected termite image were scaled up by the same ratio and filled with RGB 128 colors. The final scaled images had the same width and height and the image content was not distorted. Considering the graphics card and video memory, the width and height were scaled to 416×416, which was used as the input size of the LightYOLOV4 model.
[0044] In the second sub-step of the first step, the data augmentation of the image specifically involves:
[0045] The Mosaic data augmentation method is used to randomly select four images and stitch them together to form a new image, which contains the corresponding label information.
[0046] Image data augmentation involves three steps:
[0047] The first step is to randomly select four images from the labeled dataset.
[0048] The second step is to flip and scale the four images respectively, and then place the four processed images in the top left, bottom left, top right and bottom right positions of the four original-size canvases.
[0049] The third step is to use a matrix to extract the image regions from the four images and stitch them together to form a new image. The label information of the four images is also recalculated and merged into the label information of the new image.
[0050] The third sub-step of the first step is specifically: to modify the existing YOLOv4 object detection model, including:
[0051] ① For the backbone network, MobileNetv1 is used as the backbone network of the LightYOLOv4 network model, and its core is depthwise separable convolution.
[0052] ② For the fusion layer, the standard convolutional and depthwise separable convolutional alternating stacked modules CD_3 and CD_5 are used to replace the continuous standard convolutional modules CBL×3 and CBL×5 in PANet as the fusion layer of the LightYOLOV4 network model.
[0053] ③ For the fusion layer, i.e., the Neck, use Softpool to construct the SPP layer within it;
[0054] ④ For the detection head, the original YOLO Head 52×52 and YOLO Head 26×26 modules in the YOLOv4 target detection model are retained, and a YOLO Head 104×104 module is added. The YOLO Head 13×13 module is deleted. The fusion layer inputs the fused feature maps of 104×104, 52×52 and 26×26 sizes into the detection head to improve the head's performance in locating and classifying termites.
[0055] The experimental training parameters in the fourth sub-step of the first step are configured as follows: the image input size is 416×416 pixels, and the pre-trained model is MobileNetv1_1_0_224_tf.h5 trained based on ImageNet; the training process is divided into two stages: freezing and not freezing the backbone network, and the parameter configurations for these two stages are shown in the table below:
[0056]
[0057] In the first to fourth sub-steps, the LightYOLOV4 network model was created using Python 3.6 on the PyCharm platform for a termite surface defect detection program. The deep learning object detection network, namely the LightYOLOV4 network model, was built using TensorFlow 1.31.1, OpenCV 4.5.5.64, CUDA 11.2, and cuDNN 8.0.1. In the fifth sub-step, the final structure and parameters of the LightYOLOV4 network model were saved as a .h5 file.
[0058] The second step is as follows:
[0059] Place multiple test boxes at different locations on the dam to be tested, turn on the computer, and perform the following operations for each test box:
[0060] First, place termite attractants on the bottom plate inside each box;
[0061] Next, turn on the lights and wait for the lighting conditions to stabilize before turning on the camera; the images captured by the camera are transmitted to the computer through the first wireless communication module and the second wireless communication module.
[0062] The third step is as follows:
[0063] For images captured by the camera, data augmentation is performed to resize the images to a uniform size, followed by normalization preprocessing, which serves as the input image for the LightYOLOV4 network model. Then, after a series of convolutional operations and downsampling, feature maps calculated by the last four depthwise separable convolutions with a stride of 1 in the backbone network are extracted. These four feature maps are then input into the fusion layer (Neck) for multi-scale feature fusion. The fused 104×104, 52×52, and 26×26 feature maps are then input into the detection head (Head) to obtain the location coordinates and quantity information of termites. When termites are present in the image, the computer emits an alarm sound through the speaker and displays the corresponding box number information on the monitor to alert the staff.
[0064] The present invention has the following advantages:
[0065] (1) Termites and other insects are classified using deep learning, with a detection accuracy of over 95%.
[0066] (2) The camera captures images wirelessly and the terminal computer receives the images, performs automatic recognition and detection, which greatly improves the convenience and intuitiveness of operation.
[0067] (3) Termite detection devices can be deployed at multiple locations along the embankment, offering advantages such as wireless transmission, low cost, intuitive image capture, and long battery life (solar power). This invention improves the efficiency of termite detection by simultaneously acquiring images from multiple enclosures at different locations along the embankment using solar power.
[0068] (4) In the backbone and fusion layer of LightYOLOV4, the present invention replaces the standard convolution operation with a depth-separable convolution operation. That is, a lightweight MobileNetv1 network is used as the backbone and a lightweight PANet network is constructed as the neck, which greatly reduces the number of model parameters and the size of the network model file, thus achieving model lightweighting.
[0069] (5) In the SPP structure of LightYOLOv4, this invention uses Softpool to construct the SPP, thereby obtaining richer multi-scale defect feature information. The core of Softpool is to pool the feature map using an exponentially weighted summation method. It calculates the weights of each position in the local feature map based on the information of the local feature map of the pooling kernel size, and then calculates the pooling value of the local feature map by weighted summation of each position. The pooling kernel traverses the entire feature map according to the stride, and finally obtains the pooled feature map. Therefore, it not only reduces the feature map size but also integrates more termite surface feature information, which is beneficial for more defect semantic information to participate in subsequent operations.
[0070] (6) Improvements to the detection head (Head) include the addition and deletion of the detection head module, which avoids the problem of defects being less than one pixel, making the algorithm more compatible with the task of termite identification, and is conducive to the fusion of features by the Neck and the classification and localization of image objects by the Head.
[0071] (7) Compared with the standard YOLOv4 network model, the LightYOLOv4 network model of this invention significantly improves the accuracy and efficiency of termite identification, while greatly reducing the model size to 39 MB. This achieves lightweighting of the model while meeting the requirements of termite surface detection on dams. Compared with the 245 MB file size of the general YOLOv4 network model before the improvement, this invention effectively achieves lightweighting of the network model, which is beneficial to improving the network model's running speed and identification efficiency. Attached Figure Description
[0072] Figure 1 This is a three-dimensional structural diagram of the casing of the termite detection device for dams based on LightYOLOV4;
[0073] Figure 2 This is a schematic diagram of the main structure of the casing of a termite detection device for dams based on LightYOLOV4;
[0074] To clearly express the structure, Figure 1 and Figure 2 The enclosure and solar panels were made transparent, meaning only their outlines were visible, without obstructing the view of the internal structure. This allowed the camera, first wireless communication module, power module, lighting, and termite attractants hidden inside the enclosure to be displayed. Figure 1 and Figure 2 These structures should not be visible in the text. Figure 2 In the ant and insect entry point 4, some termite attractants should be visible.
[0075] Figure 3 This is a flowchart of the overall termite detection method.
[0076] Figure 4 This is the network structure diagram of the final constructed LightYOLOV4 network model.
[0077] Figure 5 This is a schematic diagram of the principle of depth-separable convolution in the third sub-step of the first step.
[0078] Figure 6 This is a structural diagram of the SPP layer constructed using Softpool in the third sub-step of the first step.
[0079] Figure 7The network structure obtained by lightweighting and improving Backbone and Neck in the first and third steps (is the result of...) Figure 4 (The intermediate network structure before). Detailed Implementation
[0080] like Figures 1 to 7 As shown, the dam termite detection device based on LightYOLOV4 of the present invention includes a box and a computer. The box includes a top plate 1, a bottom plate 2 and a side plate 3 connected circumferentially between the top plate 1 and the bottom plate 2.
[0081] The enclosure consists of multiple sets, each with the same structure. Multiple entry points 4 for termites and insects are spaced apart at the bottom of the side panels 3. Termite attractants 5 are placed on the bottom plate 2 inside the enclosure. A lighting lamp 6 facing the bottom plate 2, a camera 7 facing the bottom plate 2, a first wireless communication module 8 (using a Wi-Fi, Bluetooth, or Zigbee module; for longer distances from the computer, a 4G or 5G module) and a power module 9 are installed on the lower surface of the top plate 1. The power module 9 is connected to the lighting lamp 6, camera 7, and first wireless communication module 8 and supplies power to all three. A solar panel 11 is fixedly connected to the top plate 1. The solar panel 11 is connected to a solar controller (a standard accessory of the solar panel 11, not shown in the figure). The solar controller is connected to the power module 9 and supplies power to it. The camera 7 is fixed to the lower surface of the top plate 1 via a mounting plate 10.
[0082] The computer is connected to a second wireless communication module. The first wireless communication module 8 of each enclosure is connected to the computer's second wireless communication module and is used to transmit images captured by the cameras 7 of each enclosure to the computer. The computer stores a LightYOLOv4 network model for identifying termites. The computer is connected to a monitor and speakers. The computer, monitor, speakers, and wireless communication module are all conventional technologies; the computer, monitor, speakers, and second wireless communication module are not shown in the figure.
[0083] In this embodiment, the side plate 3 includes a left side plate 3, a right side plate 3, a front side plate 3, and a rear side plate 3, all of which have a rectangular horizontal cross-section. The lengths of the front side plate 3 and the rear side plate 3 are greater than the lengths of the left side plate 3 and the right side plate 3. Multiple termite entry points 4 are evenly spaced at the bottom of the front side plate 3 and the rear side plate 3. The termite attractant 5 can be one or more of the termite-preferred foods such as wood, pulp, cellulose, and protein.
[0084] This invention also provides a corresponding termite detection method, which uses the above-mentioned LightYOLOV4-based dam termite detection device and proceeds according to the following steps:
[0085] The first step is modeling and training to obtain the trained LightYOLOv4 network model; the LightYOLOv4 network model constructed in the first step is as follows: Figure 4 As shown.
[0086] The second step is to use multiple boxes to simultaneously acquire and transmit images; each box is numbered, and the numbering information of each box is stored in the computer.
[0087] The third step involves using the trained LightYOLOv4 network model to identify the images captured in real time by camera 7. When termites are detected, an alert is issued via speaker. Staff must constantly monitor the computer monitor and, upon hearing the alert, promptly take follow-up actions after the termites are discovered, such as carrying out termite extermination work as needed.
[0088] This invention improves the efficiency of termite detection on dams by simultaneously acquiring images from multiple boxes at different locations on the dam.
[0089] The first step includes the following sub-steps:
[0090] The first sub-step is to create a labeled dataset, in which the location and category of each termite are labeled in each image of the labeled dataset;
[0091] The second sub-step is data preprocessing, which involves image scaling and data augmentation;
[0092] The third sub-step is to construct a LightYOLOV4 network model for real-time detection of surface defects in termites.
[0093] The fourth sub-step is to train the LightYOLOv4 network model;
[0094] The fifth sub-step is to save the LightYOLOv4 network model.
[0095] In the first sub-step of the first step, the termite images are manually labeled using the open-source deep learning annotation tool LabelImg;
[0096] The location of each termite is marked in each image of the labeled dataset by using the rectangle annotation tool in the LabelImg toolbar to mark the area where the termite is located.
[0097] When labeling categories, there are two preset categories: Termite and Others.
[0098] Each termite is manually labeled to distinguish its category (i.e., the two categories mentioned above), forming a labeled dataset. Finally, the labeled file is saved. The saved file format is XML, which contains the coordinates of the top left and bottom right corners of each termite rectangle, as well as the category information, i.e., the label information, corresponding to that rectangle.
[0099] In the second sub-step of the first step, the image is scaled up specifically as follows:
[0100] The width and height of each collected termite image were scaled up by the same ratio and filled with RGB 128 colors. The final scaled images had the same width and height and the image content was not distorted. Considering the graphics card and video memory, the width and height were scaled to 416×416, which was used as the input size of the LightYOLOV4 model.
[0101] In the second sub-step of the first step, the data augmentation of the image specifically involves:
[0102] The Mosaic data augmentation method is used to randomly select four images and stitch them together to form a new image, which contains the corresponding label information.
[0103] Image data augmentation involves three steps:
[0104] The first step is to randomly select four images from the labeled dataset.
[0105] The second step is to flip and scale the four images respectively, and then place the four processed images in the top left, bottom left, top right and bottom right positions of the four original-size canvases.
[0106] Thirdly, the image regions from the four images are extracted using a matrix and stitched together to form a new image. The label information of the four images is also recalculated and incorporated into the label information of the new image. The resulting composite image contains rectangular box labels.
[0107] The third sub-step of the first step is specifically: to modify the existing YOLOv4 object detection model, including:
[0108] ① For the backbone network (also known as the feature extraction network, i.e., the backbone), MobileNetv1 (a lightweight neural network) is used as the backbone network (i.e., the backbone) of the LightYOLOv4 network model. Its core is depthwise separable convolution; depthwise separable convolution uses depthwise convolution and pointwise convolution to separately compute the channels and spatial regions of the feature map. Depthwise convolution computes one channel per kernel, while pointwise convolution is 1×1×C_in (number of channels in the upper layer feature map)×C_out (number of channels in the output feature map). For example... Figure 5 The image shows a depthwise separable convolution.
[0109] ② For the fusion layer (also known as the feature fusion network, or Neck), the consecutive standard convolutional modules CBL×3 and CBL×5 in PANet (path aggregation network) are replaced by alternating stacked standard convolutional and depthwise separable convolutional modules CD_3 and CD_5 as the fusion layer (Neck) of the LightYOLOv4 network model; for example Figure 7 The diagram shows the network structure after lightweight improvements to Backbone and Neck.
[0110] ③ For the fusion layer, i.e., the Neck, the SPP layer is constructed using Softpool (replacing the original max pooling operation); Softpool is a pooling operation and is an existing technology; the structure of the constructed SPP layer is as follows. Figure 6 As shown.
[0111] Softpool's core principle is to pool feature maps using an exponentially weighted summation method. It calculates the weights of each position in the local feature map based on information from the pooling kernel size, and then calculates the pooling value by summing the weights at each position. The pooling kernel traverses the entire feature map according to the stride, ultimately yielding the pooled feature map. Therefore, it not only reduces the feature map size but also integrates more surface feature information from termites, allowing for the participation of more defect semantic information in subsequent operations.
[0112] ④ For the detection head, the original YOLO Head 52×52 (pixel, the same below) module and YOLO Head 26×26 module in the YOLOv4 target detection model are retained, and a YOLO Head 104×104 module is added. The YOLO Head 13×13 module is deleted. The fusion layer inputs the fused feature maps of 104×104, 52×52 and 26×26 sizes into the detection head to improve the head's performance in locating and classifying termites.
[0113] This invention improves the performance of the detection head in locating and classifying termites by adding and removing detection head modules, making the overall algorithm more compatible with the termite identification task.
[0114] The necessity of improving the detection head lies in the fact that termites occupy a very small proportion of the input image (416×416 pixels) of the LightYOLOv4 network model. After downsampling by 8x, 16x, and 32x on the backbone, the termite pixels are 5×5, 2×3, and 1×1, respectively. After 32x downsampling, the defect is less than one pixel. A defect of less than one pixel hinders the Neck's feature fusion and the Head's classification and localization of termites. Therefore, the Head part needs improvement.
[0115] The improvements to the head avoid the problem of defects being less than one pixel, which is beneficial for the Neck to fuse features and for the head to classify and locate termites.
[0116] The experimental training parameters in the fourth sub-step of the first step are configured as follows: the image input size is 416×416 pixels, and the pre-trained model is MobileNetv1_1_0_224_tf.h5 trained on ImageNet; the training process is divided into two stages: freezing and not freezing the backbone network. The parameter configurations for these two stages are shown in the table below (Table 1 is the model training parameter configuration table):
[0117] Table 1
[0118]
[0119] In the first to fourth sub-steps, the LightYOLOV4 network model was created using Python 3.6 on the PyCharm platform for termite surface defect detection. The deep learning object detection network, LightYOLOV4, was built using TensorFlow 1.31.1, OpenCV 4.5.5.64, CUDA 11.2, and cuDNN 8.0.1. In the fifth sub-step, the final structure and parameters of the LightYOLOV4 network model were saved as a .h5 file, with a file size of 45.2 MB. Compared to the 245 MB file size of the previous general YOLOV4 network model, this represents a significant reduction in network size, which is beneficial for improving the network model's running speed and recognition efficiency.
[0120] The second step is as follows:
[0121] Place multiple test boxes at different locations on the dam to be tested, turn on the computer, and perform the following operations for each test box:
[0122] First, place termite attractants 5 on the bottom plate 2 inside each box; if there are termites on the embankment near the box, they will be attracted into the box.
[0123] Next, turn on the lighting 6, and after the lighting conditions stabilize, turn on the camera 7; the image captured by the camera 7 is transmitted to the computer through the first wireless communication module 8 and the second wireless communication module.
[0124] The third step is as follows:
[0125] For the images captured by camera 7, the image size is changed to a uniform size through data augmentation, and then normalized and preprocessed as the input image for the LightYOLOV4 network model. After a series of convolution operations and downsampling, the feature maps calculated by the last four depthwise separable convolutions with a stride of 1 in the backbone network are extracted. These four feature maps are respectively input into the fusion layer (Neck) for multi-scale feature fusion. Then, the fused 104×104, 52×52 and 26×26 (all pixels) feature maps are input into the detection head (Head) to obtain the location coordinates of termites (i.e., the coordinates of objects in the image identified as Termites) and their quantity information. When the network model determines that there are termites in the image, the computer emits an alarm sound (such as a buzzer) through the speaker and displays the corresponding box number information on the monitor to remind the staff to pay attention, so as to arrange subsequent investigation and extermination work in a timely manner.
[0126] Table 2 is a comparison table of model size, mAP, and FPS between the LightYOLOv4 network model of this invention and the general YOLOv4 network model.
[0127] mAP is the average of the recognition accuracy for the two classes (Termite and Others); FPS is used to measure the detection speed of the optimized model.
[0128] Table 2
[0129] Network Model Model size mAP FPS YOLOv4 245MB 72.60% 24.27 LightYOLOV4 45.2MB 92.71% 36.14 .
[0130] Table 3 is a comparison table of F1-score and AP of the LightYOLOV network model and the general YOLOv4 network model of this invention.
[0131] AP is used to measure the accuracy of recognition of two classes (Termite and Others), while F1-score is the harmonic mean of precision P and recall R, reflecting the model's detection accuracy for each class (Termite and Others, i.e., termites and others).
[0132] Table 3 Comparison of F1-score and AP
[0133]
[0134] As shown in Tables 2 and 3, the LightYOLOV network model of the present invention has a significantly improved accuracy in identifying termites and non-termites (other) compared to the existing YOLOv4 network model, and its size (file size) is greatly reduced, achieving lightweight design.
[0135] The above embodiments are only used to illustrate and not limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention without departing from the spirit and scope of the present invention. Any modifications or partial substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A termite detection method, characterized in that: A termite detection device for dams based on LightYOLOV4 is adopted. The termite detection device for dams based on LightYOLOV4 includes a box and a computer. The box includes a top plate, a bottom plate and side plates connected circumferentially between the top plate and the bottom plate. The enclosure consists of multiple sets, each with the same structure. Multiple entry points for termites and insects are spaced apart at the bottom of the side panels. Termite attractants are placed on the bottom plate inside the enclosure. The lower surface of the top plate is equipped with a light source facing the bottom plate, a camera facing the bottom plate, a first wireless communication module for communication with a computer, and a power module. The power module is connected to the light source, camera, and first wireless communication module and supplies power to all three. A solar panel is fixedly connected to the top plate, and the solar panel is connected to a solar controller. The solar controller is connected to the power module and supplies power to it. The computer is connected to a second wireless communication module. The first wireless communication module of each enclosure is connected to the second wireless communication module of the computer and is used to transmit the images captured by the camera of each enclosure to the computer. The computer stores a LightYOLOV4 network model for identifying termites. The computer is connected to a monitor and speakers. The termite detection method is performed according to the following steps: The first step is modeling and training to obtain the trained LightYOLOV4 network model. The second step is to use multiple boxes to simultaneously acquire and transmit images; each box is numbered, and the numbering information of each box is stored in the computer. The third step is to use the trained LightYOLOV4 network model to identify images captured by the camera in real time, and to issue an alert through a speaker when termites are detected. The first step includes the following sub-steps: The first sub-step is to create a labeled dataset, in which the location and category of each termite are labeled in each image of the labeled dataset; The second sub-step is data preprocessing, which involves image scaling and data augmentation; The third sub-step is to construct a LightYOLOV4 network model for real-time detection of surface defects in termites. The fourth sub-step is to train the LightYOLOv4 network model; The fifth sub-step is to save the LightYOLOv4 network model; The third sub-step of the first step is specifically: to modify the existing YOLOv4 object detection model, including: ① For the backbone network, MobileNetv1 is used as the backbone network of the LightYOLOv4 network model, and its core is depthwise separable convolution. ② For the fusion layer, standard convolution and depthwise separable convolution modules CD_3 and CD_5 are used in alternating stacks to replace the corresponding modules. In PANet, consecutive standard convolutional modules CBL×3 and CBL×5 are used as fusion layers in the LightYOLOV4 network model; ③ For the fusion layer, i.e., the Neck, use Softpool to construct the SPP layer within it; ④ For the detection head, retain the original YOLO Head52×52 module and YOLO from the YOLOv4 object detection model. The Head26×26 module is added, and the YOLO Head104×104 module is added. The YOLO Head13×13 module is deleted. The fusion layer inputs the fused feature maps of 104×104, 52×52 and 26×26 sizes into the detection head, which improves the head's performance in locating and classifying termites.
2. The termite detection method according to claim 1, characterized in that: In the first sub-step of the first step, the termite images are manually labeled using the open-source deep learning annotation tool LabelImg; The location of each termite is marked specifically in each image of the labeled dataset, using the LabelImg tool. The rectangular box in the toolbar marks the area where the termites are located; When labeling categories, there are two preset categories: Termite and Others; Each termite is manually labeled to distinguish its category, forming a labeled dataset, and finally the labeled file is saved.
3. The termite detection method according to claim 2, characterized in that: In the second sub-step of the first step, the image is scaled up specifically as follows: The width and height of each collected termite image were scaled up by the same ratio and filled with RGB 128 colors. The final scaled images had the same width and height and the image content was not distorted. Considering the graphics card and video memory, the width and height were scaled to 416×416, which was used as the input size of the LightYOLOV4 model. In the second sub-step of the first step, the data augmentation of the image specifically involves: The Mosaic data augmentation method is used to randomly select four images and stitch them together to form a new image, which contains the corresponding label information. Image data augmentation involves three steps: The first step is to randomly select four images from the labeled dataset. The second step is to flip and scale the four images respectively, and then place the four processed images in the top left, bottom left, top right and bottom right positions of the four original-size canvases. The third step is to use a matrix to extract the image regions from the four images and stitch them together to form a new image. The label information of the four images is also recalculated and merged into the label information of the new image.
4. The termite detection method according to claim 3, characterized in that: The experimental training parameters in the fourth sub-step of the first step are configured as follows: the image input size is 416×416 pixels, the pre-trained model is MobileNetv1_1_0_224_tf.h5 trained on ImageNet; the training process is divided into two stages: freezing and not freezing the backbone network.
5. The termite detection method according to claim 4, characterized in that: In the first through fourth sub-steps, the LightYOLOv4 network model was implemented using Python 3.6 on the PyCharm platform. The termite surface defect detection program was created using TensorFlow 1.31.1, OpenCV 4.5.5.64, CUDA 11.2, and cuDNN 8.0.1 to build a deep learning object detection network, namely the LightYOLOV4 network model. In the fifth sub-step, the final structure and parameters of the LightYOLOV4 network model were saved as a .h5 file.
6. The termite detection method according to claim 5, characterized in that: The second step is as follows: Place multiple boxes at different locations on the dam to be tested, turn on the computer, and perform the following operations for each box: First, place termite attractants on the bottom plate inside each box; Next, turn on the lights and wait for the lighting conditions to stabilize before turning on the camera; the images captured by the camera are transmitted to the computer through the first wireless communication module and the second wireless communication module.
7. The termite detection method according to claim 6, characterized in that: The third step is as follows: For images captured by the camera, data augmentation is performed to resize the images to a uniform size, followed by normalization preprocessing, which serves as the input image for the LightYOLOV4 network model. Then, after a series of convolutional operations and downsampling, feature maps calculated by the last four depthwise separable convolutions with a stride of 1 in the backbone network are extracted. These four feature maps are then input into the fusion layer (Neck) for multi-scale feature fusion. The fused 104×104, 52×52, and 26×26 feature maps are then input into the detection head (Head) to obtain the location coordinates and quantity information of termites. When termites are present in the image, the computer emits an alarm sound through the speaker and displays the corresponding box number information on the monitor to alert the staff.
Citation Information
Patent Citations
Integrated management system for termite control
CN103262836A
Facial expression recognition method based on deep learning feature fusion
CN113516047A
Solenopsis invicta monitoring method and system
CN113749068A
Rapid pest detection method based on improved YOLO V4
CN114220035A