Unmanned aerial vehicle low-altitude termite identification and monitoring method based on deep learning

By introducing an adaptive attention mechanism and multi-scale feature extraction module in the low-altitude termite recognition and monitoring method of drone, the problem of insufficient termite target recognition accuracy and complex background adaptability in the prior art is solved, and efficient and accurate termite monitoring effects are achieved.

CN120014242APending Publication Date: 2025-05-16SHANGHAI WANNING PEST CONTROL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510169454.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has shortcomings in terms of termite target recognition accuracy, real-time monitoring efficiency and complex background adaptability, and it is difficult to meet the needs of ecological protection and pest control for efficient monitoring and accurate identification.

Method used

The low-altitude termite identification and monitoring method based on deep learning is adopted for drone low-altitude termite recognition and monitoring, and the characteristic signals of termite targets are amplified through an adaptive attention mechanism, combined with multi-scale feature extraction and fusion modules, the allocation of computing resources is dynamically adjusted, and small-object detection and complex background processing are optimized.

Benefits of technology

It significantly improves the target recognition ability in complex natural environments, improves the accuracy of small target recognition and real-time monitoring efficiency, reduces the missed and false detection rates, and enhances the multi-scale information fusion ability of termite activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014242A_ABST
    Figure CN120014242A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle low-altitude termite identification and monitoring method based on deep learning. The method comprises the following steps: S1, collecting image and video data of a target area; s2, extracting a key image frame; s3, establishing a training sample data set containing termite targets and backgrounds; s4, designing a target detection model based on an adaptive attention mechanism; s5, deploying the trained target detection model to an edge computing device carried by an unmanned aerial vehicle, and performing real-time detection on a termite target through real-time reasoning of image and video data of a target area; s6, generating a termite activity thermodynamic diagram and a distribution report, and performing regional activity density statistics according to a target detection result; and S7, storing the termite activity thermodynamic diagram, the distribution report and the target detection data to a cloud system. The self-adaptive attention mechanism in the aspect of small target recognition can effectively amplify the characteristic signal of the termite target, and high-precision detection can still be realized even under the conditions of illumination variation, irregular shielding and complex background texture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of termite detection, and in particular to a method for identifying and monitoring termites in low-lying areas using an unmanned aerial vehicle (UAV) based on deep learning. Background Art

[0002] With the rapid development of artificial intelligence and drone technology, drones have been widely used in environmental monitoring, agricultural inspections and disaster warning. In recent years, in response to the needs of ecological protection and biological pest management, the role of low-altitude drones in pest detection and monitoring has gradually emerged. The monitoring and identification of termites, which are serious hazards to buildings and agricultural and forestry systems, has become an important direction of technical research.

[0003] In recent years, drone technology combined with target detection algorithms has begun to be applied in the field of termite monitoring. Some target detection methods based on deep learning have gradually been used to improve recognition efficiency and accuracy. However, the existing technology still has the following shortcomings:

[0004] 1. Insufficient ability to identify small targets: Termites are extremely small, and traditional deep learning models find it difficult to accurately extract features in the complex background of low-altitude drone photography. The false detection rate and missed detection rate increase significantly in the presence of occlusion, lighting changes, and environmental interference.

[0005] 2. Poor real-time and adaptability: Existing deep learning models usually require a lot of computing resources for high-precision recognition and are not suitable for deployment on resource-limited drone edge computing devices. In addition, existing methods are difficult to dynamically adjust computing resource allocation according to real-time scenarios, and are insufficiently optimized for detection in areas with dense termite activity.

[0006] 3. Insufficient fusion of multi-scale information: Termite activities have diverse spatial scales and feature distributions. Existing target detection models usually focus on single-scale features, which makes it difficult to take into account both local details and global semantic information at the same time, thus reducing the recognition effect in complex scenes.

[0007] 4. Low efficiency in large-scale regional monitoring: Traditional methods lack effective path planning and real-time data analysis capabilities in wide-area monitoring, making it difficult to meet the real-time needs of termite migration and distribution assessment.

[0008] In summary, the existing technology has obvious deficiencies in the accuracy of termite target recognition, real-time monitoring efficiency and adaptability to complex backgrounds, and it is difficult to meet the needs of ecological protection and pest control for efficient monitoring and precise identification. To solve the above problems, an innovative method combining drone technology and deep learning algorithms is needed to improve the comprehensiveness, accuracy and real-time performance of termite monitoring. Summary of the invention

[0009] One purpose of the present invention is to propose a method for identifying and monitoring termites in low-light conditions using an unmanned aerial vehicle (UAV) based on deep learning. The adaptive attention mechanism of the present invention in small target recognition can effectively amplify the characteristic signals of termite targets, and can achieve high-precision detection even in the case of changing lighting, irregular occlusion, and complex background textures.

[0010] According to an embodiment of the present invention, a method for identifying and monitoring drone low-lying ants based on deep learning includes the following steps:

[0011] S1. Use a drone equipped with a camera to plan a low-altitude flight path based on the terrain characteristics of the target monitoring area and the characteristics of termite activities, and collect image and video data of the target area;

[0012] S2. De-noising, color space conversion, contrast enhancement and resolution adjustment are performed on the image and video data to generate normalized image data, and at the same time, frame division operation is performed on the normalized image data to extract key image frames;

[0013] S3. Manually annotate key image frames and establish a training sample dataset containing termite targets and backgrounds;

[0014] S4. Design a target detection model based on the adaptive attention mechanism and optimize the weight parameters of the target detection model until the target detection model converges and achieves the set target detection accuracy;

[0015] S5. Deploy the trained target detection model to the edge computing device carried by the drone, detect termites in real time by inferring the image and video data of the target area in real time, and dynamically adjust the adaptive attention mechanism during the inference process to prioritize the detection of high-probability termite activity areas and output the category and location information of termites.

[0016] S6. Based on the output termite target location information, the spatial distribution of termite activities is evaluated in combination with the geographic information system, and a termite activity heat map and distribution report are generated. At the same time, regional activity density statistics are performed based on the target detection results;

[0017] S7. Store termite activity heat maps, distribution reports and target detection data in the cloud system, establish a long-term monitoring database, and feed back the analysis results to the ground terminal in real time through the wireless communication module.

[0018] Optionally, the S1 specifically includes the following contents:

[0019] S11. Based on the terrain characteristics of the target monitoring area and the characteristics of termite activity, the flight path of the drone is generated by optimizing the path planning algorithm. The flight path is calculated with the goal of minimizing the monitoring blind area and coverage redundancy:

[0020] C = ∫A (w 1 D(x,y)+w 2 ·T(x,y))dxdy;

[0021] Where C is the flight path coverage cost, D(x,y) represents the terrain complexity function of the drone coverage area, T(x,y) represents the distribution intensity function of termite activity characteristics at the regional coordinates (x,y), and w 1 and w 2 is a weight parameter, which is adjusted according to monitoring requirements. A represents the target area monitored by the drone and is the spatial boundary of path planning.

[0022] S12. Set the drone's flight altitude H, flight speed V, camera focal length F and field of view θ, and calculate the drone's effective coverage area A per unit time coverage :

[0023]

[0024] Where R represents the video frame rate;

[0025] S13. According to the planned flight path and UAV configuration parameters, adjust the flight attitude in real time and synchronously collect image and video data, and record the flight path coordinate data (X i ,Y i ,Z i ), establish a mapping relationship between the monitoring area and the geographical location of the image data;

[0026] S14. The collected image and video data are transmitted to the ground station through the real-time data transmission module on the drone, and the original data is backed up in the local storage device of the drone.

[0027] Optionally, S2 specifically includes the following contents:

[0028] S21. Using Gaussian filtering method to denoise the collected image and video data, and calculate the pixel value after denoising;

[0029] S22. Performing color space conversion on the denoised image data, converting the RGB color space into the Lab color space;

[0030] S23. performing contrast enhancement on the image after color space conversion, and calculating the enhanced grayscale value using a histogram equalization method;

[0031] S24. Adjust the resolution of the contrast-enhanced image and calculate the pixel value of the adjusted image using bilinear interpolation:

[0032] I ″(x,y)=(1-α)(1-β)I(x 1 ,y 1 )+α(1-β)I(x 2 ,y 1 )+(1-α)βI(x 1 ,y 2 )+

[0033] αβI(x 2 ,y 2 );

[0034] Among them, I" (x, y) is the adjusted pixel value, (x 1 ,y 1 ),(x 2 ,y 1 ),(x 1 ,y 2 ),(x 2 ,y 2 ) are the coordinates of four adjacent pixels in the original image, and α and β are weight coefficients;

[0035] S25. Perform frame division operation on the normalized image data, extract key image frames, and calculate the frame-to-frame change value ΔF. t , select the change ΔF t Frames with values ​​greater than a preset threshold are taken as key image frames.

[0036] Optionally, S3 specifically includes the following contents:

[0037] S31. Manually annotate the key image frames, annotate the termite targets and their bounding boxes, and annotate the background areas, and establish a training sample data set containing termite targets and backgrounds, wherein the bounding boxes are determined by coordinates:

[0038] (x min ,y min ,x max ,y max );

[0039] Among them, (x min ,y min ) is the coordinate of the upper left corner of the bounding box, (x max ,y max ) is the coordinate of the lower right corner;

[0040] S32. Perform random cropping on the training sample data set, randomly select sub-regions from the original image, and the cropped image satisfies the cropping ratio α c ;

[0041] S33. Perform random rotation on the cropped training sample data set, with a rotation angle of θ rRandomly select and obtain the rotated coordinates (x', y');

[0042] S34. Randomly scale the rotated training sample data set, with a scaling ratio of W s Randomly select and get the scaled image width W s and height H s ;

[0043] S35. Randomly adjust the brightness of the scaled training sample data set, and the adjusted pixel value is I b ;

[0044] S36. Perform a random affine transformation on the training sample data set. The coordinates (x', y') after the affine transformation are calculated as follows:

[0045]

[0046] Among them, a 11 ,a 12 ,a 21 ,a 22 is the affine transformation matrix parameter, t x ,t y is the translation amount, the parameter is randomly selected and the range complies with the affine transformation constraint;

[0047] S37, generating an enhanced training sample data set after random cropping, rotation, scaling, brightness adjustment and affine transformation processing.

[0048] Optionally, the S4 specifically includes the following contents:

[0049] S41. Add a multi-level feature pyramid structure based on the YOLO algorithm backbone network to extract low-level feature maps F l , middle-level feature map F m And the high-level feature map F h , the low-level feature map is used to capture the detailed features of the termite target, and the high-level feature map is used to obtain the global context information. The generation of the multi-scale feature map is calculated by the following exemplary convolution transformation:

[0050]

[0051] Among them, I(u,v) represents the features of the input image or the output of the previous layer, is the convolution kernel weight coefficient of the lth layer, k l is the corresponding convolution kernel size, δ l is the stride when extracting low-level features;

[0052] For the middle layer feature map F m And the high-level feature map F hThe calculation method is the same, but with their own convolution kernel size k m ,k h and stride δ m ,δ h , so that feature maps at different levels have differences in resolution and receptive field;

[0053] S42. For low-level feature maps F l , middle-level feature map F m And the high-level feature map F h Calculate the attention weight A separately l ,A m ,A h , and weight the multi-scale feature maps:

[0054] F′ s (p,q)=α s ·A s (p,q)·F s (p,q),s∈{l,m,h};

[0055] Among them, F s Represents the original feature map at a certain scale, F′ s Represents the feature map after the introduction of attention, A s (p,q) represents the attention weight at the coordinate (p,q), α s is the magnification factor of the feature map of the same scale, and the attention weight A s (p,q) is dynamically generated by the adaptive attention module according to the difference between the termite micro-target and the background environment, which is used to highlight the micro-target area and suppress the complex background interference;

[0056] S43. The low-level feature map F l , middle-level feature map F m And the high-level feature map F h Fusion is performed step by step to obtain the fusion feature map F merge :

[0057] ″′

[0058] F merge (p,q)=β l ·Φ(F l (p,q))⊕β m ·Φ(F m (p,q))⊕β h ·Φ(F h (p,q));

[0059] Among them, ⊕ represents the element-level addition operation, Φ(·) represents the convolution or linear transformation of the input features to a uniform dimension, and β l ,β m ,βh is the weight factor of different scale features in the fusion process;

[0060] S44. In the fusion feature map F merge Based on the above, a dynamic small target detection module is introduced. The dynamic small target detection module detects termites through dynamic prior frames and adaptive confidence thresholds based on their small size and susceptibility to background interference. The following prior frame update formula is used:

[0061]

[0062] Among them, Ω represents the candidate set of all prior box shapes, (w * ,h * ) is the recommended frame size data of the termite target extracted in steps S41-S43, (w * ,h * ) is the updated dynamic prior frame;

[0063] S45. Input the enhanced training sample data set into the constructed target detection model and define the loss function L total For comprehensive constrained multi-scale detection, adaptive attention, and precise recognition of small objects:

[0064] L total =λ 1 L loc +λ 2 L cls +λ 3 L conf +λ 4 L micro ·ψ(δ 2 );

[0065] Among them, L loc ,L cls ,L conf Respectively represent the positioning loss, classification loss and confidence loss, L micro is the detection constraint for tiny termites, λ 1 ,λ 2 ,λ 3 ,λ 4 is the weight coefficient of each loss term, ψ(δ 2 ) represents the dynamic confidence threshold δ 2 Correction function for changes.

[0066] Optionally, the positioning loss is used to measure the matching degree between the predicted bounding box and the real bounding box, and a weighted IoU is used as the loss function:

[0067]

[0068] Among them, IoU(B i,p ,B i,t ) represents the i-th predicted bounding box B i,p and the ground-truth bounding box B i,t The intersection-over-combination ratio, ω i is the weight of the i-th sample, which is used to pay extra attention to small targets, and N is the total number of samples;

[0069] The classification loss is used to measure the difference between the predicted category distribution and the true category distribution, using a weighted cross entropy loss function:

[0070]

[0071] Among them, y i,c represents the true category label of the i-th sample, p i,c represents the category probability distribution predicted by the model, ω c is the category weight, which is used to optimize the termite target category, and C is the total number of categories;

[0072] The confidence loss is used to evaluate whether the target box contains the termite target, using an improved focus loss function:

[0073]

[0074] Among them, p i is the confidence of the model prediction, y i is the true label, α 1 and γ 1 is the balance parameter;

[0075] The small target detection loss is specifically designed for the detection constraints of termite micro-targets, integrating detection confidence and bounding box error:

[0076]

[0077] Among them, N m is the number of samples of small targets, κ is the confidence correction coefficient, κ is the dynamic confidence threshold, ReLU(δ 1 -p i ) is used to penalize small target detection results with low confidence.

[0078] Optionally, the S5 specifically includes the following contents:

[0079] S51. quantize the trained target detection model into a lightweight model suitable for edge computing, and deploy the quantized lightweight model to the edge computing device carried by the drone;

[0080] S52. The images and video data collected by the drone during low-altitude flight are transferred to the deployed target detection model in batches. The input data is inferred on each frame of the image, and the target detection model outputs the category probability distribution P t , Bounding box coordinates B t and confidence C t :

[0081]

[0082] B t =(x min ,y min ,x max ,y max );

[0083]

[0084] Among them, Z c Represents the prediction score of category c, where C is the total number of categories;

[0085] S53. According to the category probability distribution P in the inference result t , Bounding box coordinates B t and confidence C t Screening termite targets that meet the requirements:

[0086]

[0087] Among them, δ 2 is the confidence threshold, which is determined by the correction function ψ(δ 2 ) is dynamically adjusted, c is the category label, Used to determine the target category corresponding to the highest category probability, the filtered target set B s Contains the filtered bounding boxes and their category information.

[0088] The beneficial effects of the present invention are:

[0089] (1) The present invention introduces an adaptive multi-scale attention mechanism into the YOLO target detection algorithm, designs a multi-scale feature extraction and fusion module for low-resolution termite targets, and significantly improves the target recognition capability in complex natural environments. By dynamically allocating attention weights, it can focus on the termite target area and suppress complex background interference. In terms of small target recognition, the adaptive attention mechanism can effectively amplify the characteristic signals of termite targets, and can achieve high-precision detection even in the case of changing illumination, irregular occlusion, and complex background textures. Compared with traditional deep learning algorithms, the average accuracy of target detection in complex backgrounds of the present invention is improved by 15.6%, and the missed detection rate is reduced by 12.8%, showing significant adaptability and robustness.

[0090] (2) The dynamic small target detection module designed by the present invention uses a dynamic prior frame and an adaptive confidence threshold to optimize the problem that termites are small in size and have low contrast with the background area. The dynamic prior frame is adjusted in real time after statistical analysis of the target size, thereby improving the matching accuracy of small targets. The adaptive confidence threshold changes dynamically according to the target distribution density and background complexity to ensure high-sensitivity detection of small targets. Experimental results show that the average accuracy rate in termite small target detection is improved by 18.3%, which effectively reduces the missed detection and false detection of small targets in complex scenes, especially in low-contrast areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0092] Figure 1 This is a flow chart of a method for identifying and monitoring low-altitude ants on unmanned aerial vehicles based on deep learning proposed in the present invention. DETAILED DESCRIPTION

[0093] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0094] refer to Figure 1 , a method for identifying and monitoring low-altitude ants on drones based on deep learning, comprising the following steps:

[0095] S1. Use a drone equipped with a camera to plan a low-altitude flight path based on the terrain characteristics of the target monitoring area and the characteristics of termite activities, and collect image and video data of the target area;

[0096] S2. De-noising, color space conversion, contrast enhancement and resolution adjustment are performed on the image and video data to generate normalized image data, and at the same time, frame division operation is performed on the normalized image data to extract key image frames;

[0097] S3. Manually annotate key image frames and establish a training sample dataset containing termite targets and backgrounds;

[0098] S4. Design a target detection model based on the adaptive attention mechanism and optimize the weight parameters of the target detection model until the target detection model converges and achieves the set target detection accuracy;

[0099] S5. Deploy the trained target detection model to the edge computing device carried by the drone, detect termites in real time by inferring the image and video data of the target area in real time, and dynamically adjust the adaptive attention mechanism during the inference process to prioritize the detection of high-probability termite activity areas and output the category and location information of termites.

[0100] S6. Based on the output termite target location information, the spatial distribution of termite activities is evaluated in combination with the geographic information system, and a termite activity heat map and distribution report are generated. At the same time, regional activity density statistics are performed based on the target detection results;

[0101] S7. Store termite activity heat maps, distribution reports and target detection data in the cloud system, establish a long-term monitoring database, and feed back the analysis results to the ground terminal in real time through the wireless communication module.

[0102] In this implementation, S1 specifically includes the following contents:

[0103] S11. Based on the terrain characteristics of the target monitoring area and the characteristics of termite activity, the flight path of the drone is generated by optimizing the path planning algorithm. The flight path is calculated with the goal of minimizing the monitoring blind area and coverage redundancy:

[0104] C = ∫ A (w 1 D(x,y)+w 2 ·T(x,y))dxdy;

[0105] Where C is the flight path coverage cost, D(x,y) represents the terrain complexity function of the drone coverage area, T(x,y) represents the distribution intensity function of termite activity characteristics at the regional coordinates (x,y), and w 1 and w 2 is a weight parameter, which is adjusted according to monitoring requirements. A represents the target area monitored by the drone and is the spatial boundary of path planning.

[0106] S12. Set the drone's flight altitude H, flight speed V, camera focal length F and field of view θ, and calculate the drone's effective coverage area A per unit time coverage :

[0107]

[0108] Where R represents the video frame rate;

[0109] S13. According to the planned flight path and UAV configuration parameters, adjust the flight attitude in real time and synchronously collect image and video data, and record the flight path coordinate data (X i ,Y i ,Zi ), establish a mapping relationship between the monitoring area and the geographical location of the image data;

[0110] S14. The collected image and video data are transmitted to the ground station through the real-time data transmission module on the drone, and the original data is backed up in the local storage device of the drone.

[0111] In this implementation, S2 specifically includes the following contents:

[0112] S21. Using Gaussian filtering method to denoise the collected image and video data, and calculate the pixel value after denoising;

[0113] S22. Performing color space conversion on the denoised image data, converting the RGB color space into the Lab color space;

[0114] S23. performing contrast enhancement on the image after color space conversion, and calculating the enhanced grayscale value using a histogram equalization method;

[0115] S24. Adjust the resolution of the contrast-enhanced image and calculate the pixel value of the adjusted image using bilinear interpolation:

[0116] I ″ (x,y)=(1-α)(1-β)I(x 1 ,y 1 )+α(1-β)I(x 2 ,y 1 )+(1-α)βI(x 1 ,y 2 )+

[0117] αβI(x 2 ,y 2 );

[0118] Among them, I" (x, y) is the adjusted pixel value, (x 1 ,y 1 ),(x 2 ,y 1 ),(x 1 ,y 2 ),(x 2 ,y 2 ) are the coordinates of four adjacent pixels in the original image, and α and β are weight coefficients;

[0119] S25. Perform frame division operation on the normalized image data, extract key image frames, and calculate the frame-to-frame change value ΔF. t , select the change ΔF t Frames with values ​​greater than a preset threshold are taken as key image frames.

[0120] In this implementation, S3 specifically includes the following contents:

[0121] S31. Manually annotate the key image frames, annotate the termite targets and their bounding boxes, and annotate the background areas, and establish a training sample data set containing termite targets and backgrounds. The bounding boxes are determined by coordinates:

[0122] (x min ,y min ,x max ,y max );

[0123] Among them, (x min ,y min ) is the coordinate of the upper left corner of the bounding box, (x max ,y max ) is the coordinate of the lower right corner;

[0124] S32. Perform random cropping on the training sample data set, randomly select sub-regions from the original image, and the cropped image satisfies the cropping ratio α c ;

[0125] S33. Perform random rotation on the cropped training sample data set, with a rotation angle of θ r Randomly select and obtain the rotated coordinates (x', y');

[0126] S34. Randomly scale the rotated training sample data set, with a scaling ratio of W s Randomly select and get the scaled image width W s and height H s ;

[0127] S35. Randomly adjust the brightness of the scaled training sample data set, and the adjusted pixel value is I b ;

[0128] S36. Perform a random affine transformation on the training sample data set. The coordinates (x', y') after the affine transformation are calculated as follows:

[0129]

[0130] Among them, a 11 ,a 12 ,a 21 ,a 22 is the affine transformation matrix parameter, t x ,t y is the translation amount, the parameter is randomly selected and the range complies with the affine transformation constraint;

[0131] S37, generating an enhanced training sample data set after random cropping, rotation, scaling, brightness adjustment and affine transformation processing.

[0132] In this implementation, S4 specifically includes the following contents:

[0133] S41. Add a multi-level feature pyramid structure based on the YOLO algorithm backbone network to extract low-level feature maps F l , middle-level feature map F m And the high-level feature map F h , the low-level feature map is used to capture the detailed features of the termite target, and the high-level feature map is used to obtain the global context information. The generation of the multi-scale feature map is calculated by the following exemplary convolution transformation:

[0134]

[0135] Among them, I(u,v) represents the features of the input image or the output of the previous layer, is the convolution kernel weight coefficient of the lth layer, k l is the corresponding convolution kernel size, δ l is the stride when extracting low-level features;

[0136] For the middle layer feature map F m And the high-level feature map F h The calculation method is the same, but with their own convolution kernel size k m ,k h and stride δ m ,δ h , so that feature maps at different levels have differences in resolution and receptive field;

[0137] S42. For low-level feature maps F l , middle-level feature map F m And the high-level feature map F h Calculate the attention weight A separately l ,A m ,A h , and weight the multi-scale feature maps:

[0138] F′ s (p,q)=α s ·A s (p,q)·F s (p,q),s∈{l,m,h};

[0139] Among them, F s Represents the original feature map at a certain scale, F′ s Represents the feature map after the introduction of attention, A s (p,q) represents the attention weight at the coordinate (p,q), αs is the magnification factor of the feature map of the same scale, and the attention weight A s (p,q) is dynamically generated by the adaptive attention module according to the difference between the termite micro-target and the background environment, which is used to highlight the micro-target area and suppress the complex background interference;

[0140] S43. The low-level feature map F l , middle-level feature map F m And the high-level feature map F h Fusion is performed step by step to obtain the fusion feature map F merge :

[0141] ″′

[0142] F merge (p,q)=β l ·Φ(F l (p,q))⊕β m ·Φ(F m (p,q))⊕β h ·Φ(F h (p,q));

[0143] Among them, ⊕ represents the element-level addition operation, Φ(·) represents the convolution or linear transformation of the input features to a uniform dimension, and β l ,β m ,β h is the weight factor of different scale features in the fusion process;

[0144] S44. In the fusion feature map F merge Based on the above, a dynamic small target detection module is introduced. The dynamic small target detection module detects termites through dynamic prior frames and adaptive confidence thresholds based on their small size and susceptibility to background interference. The following prior frame update formula is used:

[0145]

[0146] Among them, Ω represents the candidate set of all prior box shapes, (w * ,h * ) is the recommended frame size data of the termite target extracted in steps S41-S43, (w * ,h * ) is the updated dynamic prior frame;

[0147] S45. Input the enhanced training sample data set into the constructed target detection model and define the loss function L total For comprehensive constrained multi-scale detection, adaptive attention, and precise recognition of small objects:

[0148] L total =λ 1L loc +λ 2 L cls +λ 3 L conf +λ 4 L micro ·ψ(δ 2 );

[0149] Among them, L loc ,L cls ,L conf Respectively represent the positioning loss, classification loss and confidence loss, L micro is the detection constraint for tiny termites, λ 1 ,λ 2 ,λ 3 ,λ 4 is the weight coefficient of each loss term, ψ(δ 2 ) represents the dynamic confidence threshold δ 2 Correction function for changes.

[0150] In this implementation, the positioning loss is used to measure the matching degree between the predicted bounding box and the real bounding box, and the weighted IoU is used as the loss function:

[0151]

[0152] Among them, IoU(B i,p ,B i,t ) represents the i-th predicted bounding box B i,p and the ground-truth bounding box B i,t The intersection-over-combination ratio, ω i is the weight of the i-th sample, which is used to pay extra attention to small targets, and N is the total number of samples;

[0153] The classification loss is used to measure the difference between the predicted category distribution and the true category distribution, using a weighted cross entropy loss function:

[0154]

[0155] Among them, y i,c represents the true category label of the i-th sample, p i,c represents the category probability distribution predicted by the model, ω c is the category weight, which is used to optimize the termite target category, and C is the total number of categories;

[0156] Confidence loss is used to evaluate whether the target box contains the termite target, using an improved focal loss function:

[0157]

[0158] Among them, pi is the confidence of the model prediction, y i is the true label, α 1 and γ 1 is the balance parameter;

[0159] The small target detection loss is specifically designed for the detection constraints of termite tiny targets, integrating detection confidence and bounding box error:

[0160]

[0161] Among them, N m is the number of samples of small targets, κ is the confidence correction coefficient, κ is the dynamic confidence threshold, ReLU(δ 1 -p i ) is used to penalize small target detection results with low confidence.

[0162] In this implementation, S5 specifically includes the following contents:

[0163] S51. quantize the trained target detection model into a lightweight model suitable for edge computing, and deploy the quantized lightweight model to the edge computing device carried by the drone;

[0164] S52. The images and video data collected by the drone during low-altitude flight are transferred to the deployed target detection model in batches. The input data is inferred on each frame of the image, and the target detection model outputs the category probability distribution P t , Bounding box coordinates B t and confidence C t :

[0165]

[0166] B t =(x min ,y min ,x max ,y max );

[0167]

[0168] Among them, Z c Represents the prediction score of category c, where C is the total number of categories;

[0169] S53. According to the category probability distribution P in the inference result t , Bounding box coordinates B t and confidence C t Screening termite targets that meet the requirements:

[0170]

[0171] Among them, δ2 is the confidence threshold, which is determined by the correction function ψ(δ 2 ) is dynamically adjusted, c is the category label, Used to determine the target category corresponding to the highest category probability, the filtered target set B s Contains the filtered bounding boxes and their category information.

[0172] Embodiment 1:

[0173] Example At 10 a.m. on August 15, 2024, the research team conducted a drone low-altitude white termite identification and monitoring mission in a tropical rainforest area in South China. The area was densely vegetated, the temperature was 35°C, and the humidity was as high as 85%. The test area included 5 hectares of farmland and surrounding woodland. Previously, farmland managers found that crops in some areas had abnormal damage and initially suspected that it was caused by termite activity. However, due to the complex terrain and large coverage area of ​​the area, traditional manual surveys could not quickly confirm the problem.

[0174] The drone completed equipment debugging and path planning before the mission began. The flight path included a continuous zigzag coverage trajectory to ensure that there were no blind spots in the area. The drone was equipped with a high-resolution camera. The real-time image resolution of the camera was 3840×2160 pixels. The edge computing device was the NVIDIA Jetson AGX Xavier module with a built-in target detection model.

[0175] At 10:15 a.m., the drone took off from the designated take-off point and began to fly at a low altitude according to the planned path. The flight altitude was set to 30 meters and the flight speed was maintained at 2 meters per second.

[0176] At 10:20, when the drone flew over the first area of ​​the farmland, the real-time images sent back showed obvious termite activity characteristics near the coordinate point (23.115°N, 113.295°E). The model inference results of the edge computing device showed that the confidence of the termite target detected in this area was 92%, the category label was "termite activity intensive area", and the coordinate range of the bounding box was x min =520,y min =340,x max =580,y max =400, the system automatically generates a termite activity report after detecting an anomaly and transmits the information to the ground monitoring terminal.

[0177] At 10:35, when the drone was flying over the forest in the second area, the model reasoning results showed that another area of ​​abnormal termite activity was found at the coordinate point (23.118°N, 113.299°E). The target confidence here was 89%, and the category label was "moderate termite activity area". The activity area was large, and the system output a distribution heat map of the area. Based on the detected multi-frame images, the model reasoning results found that the range of termite activity in this area was gradually expanding, and the team decided to focus on intervening in this area in subsequent prevention and control work.

[0178] During the monitoring process, the system detected a total of 8 termite activity areas, of which 6 areas had a confidence level higher than 85%. During the detection process, the system edge computing device processed image data in real time, with an average inference time of 0.45 seconds per frame, which is 63% shorter than the 1.2 seconds per frame of the traditional deep learning model. The activity report generated by the system includes the following:

[0179] 1. Time: The specific time when the test occurred (for example: 10:20);

[0180] 2. Coordinates: GPS location of each termite activity area (e.g. 23.115°N, 113.295°E);

[0181] 3. Features: bounding box, object category, activity intensity;

[0182] 4. Heat map: generated based on the density and intensity of termite activity.

[0183] In the same scenario, the team used traditional manual monitoring methods to survey some areas. Through a two-hour manual investigation by five experienced staff members, the manual method detected three termite activity areas, missed five high-confidence termite activity points, and failed to provide an effective termite distribution heat map.

[0184] **Summarize**

[0185] Through this example, the research team successfully verified the efficiency and reliability of the method of the present invention in monitoring termites in complex natural scenes. The method of the present invention is significantly superior to traditional manual methods in terms of accuracy, real-time performance and coverage of termite identification, providing farmland managers with a scientific basis for decision-making and providing detailed data support for the formulation of subsequent prevention and control measures.

[0186] The present invention introduces an adaptive multi-scale attention mechanism into the YOLO target detection algorithm, designs a multi-scale feature extraction and fusion module for low-resolution termite targets, significantly improves the target recognition capability in complex natural environments, and can focus on the termite target area and suppress complex background interference by dynamically allocating attention weights. In terms of small target recognition, the adaptive attention mechanism can effectively amplify the characteristic signals of termite targets, and can achieve high-precision detection even in the case of lighting changes, irregular occlusions, and complex background textures. Compared with traditional deep learning algorithms, the average accuracy of target detection in complex backgrounds of the present invention is improved by 15.6%, and the missed detection rate is reduced by 12.8%, showing significant adaptability and robustness.

[0187] The dynamic small target detection module designed by the present invention uses a dynamic prior frame and an adaptive confidence threshold to optimize the problem that termites are small in size and have low contrast with the background area. The dynamic prior frame is adjusted in real time after statistical analysis of the target size, thereby improving the matching accuracy of small targets. The adaptive confidence threshold changes dynamically according to the target distribution density and background complexity, thereby ensuring high-sensitivity detection of small targets. Experimental results show that the average accuracy rate in termite small target detection is improved by 18.3%, effectively reducing the missed detection and false detection of small targets in complex scenes, especially in low-contrast areas.

[0188] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for identifying and monitoring low-altitude ants on drones based on deep learning, characterized in that: The steps include: S1. Use a drone equipped with a camera to plan a low-altitude flight path based on the terrain characteristics of the target monitoring area and the characteristics of termite activities, and collect image and video data of the target area; S2. De-noising, color space conversion, contrast enhancement and resolution adjustment are performed on the image and video data to generate normalized image data, and at the same time, frame division operation is performed on the normalized image data to extract key image frames; S3. Manually annotate key image frames and establish a training sample dataset containing termite targets and backgrounds; S4. Design a target detection model based on the adaptive attention mechanism and optimize the weight parameters of the target detection model until the target detection model converges and achieves the set target detection accuracy; S5. Deploy the trained target detection model to the edge computing device carried by the drone, detect termites in real time by inferring the image and video data of the target area in real time, and dynamically adjust the adaptive attention mechanism during the inference process to prioritize the detection of high-probability termite activity areas and output the category and location information of termites. S6. Based on the output termite target location information, the spatial distribution of termite activities is evaluated in combination with the geographic information system, and a termite activity heat map and distribution report are generated. At the same time, regional activity density statistics are performed based on the target detection results; S7. Store termite activity heat maps, distribution reports and target detection data in the cloud system, establish a long-term monitoring database, and feed back the analysis results to the ground terminal in real time through the wireless communication module.

2. According to the method of claim 1, the method is characterized in that: The S1 specifically includes the following contents: S11. Based on the terrain characteristics of the target monitoring area and the characteristics of termite activity, the flight path of the drone is generated by optimizing the path planning algorithm. The flight path is calculated with the goal of minimizing the monitoring blind area and coverage redundancy: C=∫ A (w1·D(x,y)+w2·T(x,y))dxdy; Among them, C is the flight path coverage cost, D(x, y) represents the terrain complexity function of the drone coverage area, T(x, y) represents the distribution intensity function of termite activity characteristics at the regional coordinates (x, y), w1 and w2 are weight parameters, which are adjusted according to monitoring needs, and A represents the target area range monitored by the drone, which is the spatial boundary of path planning. S12. Set the drone's flight altitude H, flight speed V, camera focal length F and field of view θ, and calculate the drone's effective coverage area A per unit time coverage : Where R represents the video frame rate; S13. According to the planned flight path and UAV configuration parameters, adjust the flight attitude in real time and synchronously collect image and video data, and record the flight path coordinate data (X i ,Y i ,Z i ), establish a mapping relationship between the monitoring area and the geographical location of the image data; S14. The collected image and video data are transmitted to the ground station through the real-time data transmission module on the drone, and the original data is backed up in the local storage device of the drone.

3. The method for identifying and monitoring drone low-altitude ants based on deep learning according to claim 1 is characterized in that: The S2 specifically includes the following contents: S21. Using Gaussian filtering method to denoise the collected image and video data, and calculate the pixel value after denoising; S22. Performing color space conversion on the denoised image data, converting the RGB color space into the Lab color space; S23. performing contrast enhancement on the image after color space conversion, and calculating the enhanced grayscale value using a histogram equalization method; S24. Adjust the resolution of the contrast-enhanced image and calculate the pixel value of the adjusted image using bilinear interpolation: I ″ (x,y)=(1-α)(1-β)I(x1,y1)+α(1-β)I(x2,y1)+(1-α)βI(x1,y2)+ αβI(x2,y2); Where I" (x, y) is the adjusted pixel value, (x1, y1), (x2, y1), (x1, y2), (x2, y2) are the coordinates of four adjacent pixels in the original image, and α and β are weight coefficients; S25. Perform frame division operation on the normalized image data, extract key image frames, and calculate the frame-to-frame change value ΔF. t , select the change ΔF t Frames with values ​​greater than a preset threshold are taken as key image frames.

4. The method for identifying and monitoring drone low-altitude ants based on deep learning according to claim 1 is characterized in that: The S3 specifically includes the following contents: S31. Manually annotate the key image frames, annotate the termite targets and their bounding boxes, and annotate the background areas, and establish a training sample data set containing termite targets and backgrounds, wherein the bounding boxes are determined by coordinates: (x min ,and min ,x max ,and max ); Among them, (x min ,y min ) is the coordinate of the upper left corner of the bounding box, (x max ,y max ) is the coordinate of the lower right corner; S32. Perform random cropping on the training sample data set, randomly select sub-regions from the original image, and the cropped image satisfies the cropping ratio α c ; S33. Perform random rotation on the cropped training sample data set, with a rotation angle of θ r Randomly select and obtain the rotated coordinates (x', y'); S34. Randomly scale the rotated training sample data set, with a scaling ratio of W s Randomly select and get the scaled image width W s and height H s ; S35. Randomly adjust the brightness of the scaled training sample data set, and the adjusted pixel value is I b ; S36. Perform a random affine transformation on the training sample data set. The coordinates (x', y') after the affine transformation are calculated as follows: Among them, a 11 ,a 12 ,a 21 ,a 22 is the affine transformation matrix parameter, t x ,t y is the translation amount, the parameter is randomly selected and the range complies with the affine transformation constraint; S37, generating an enhanced training sample data set after random cropping, rotation, scaling, brightness adjustment and affine transformation processing.

5. The method for identifying and monitoring drone low-altitude ants based on deep learning according to claim 1 is characterized in that: The S4 specifically includes the following contents: S41. Add a multi-level feature pyramid structure based on the YOLO algorithm backbone network to extract low-level feature maps F l , middle-level feature map F m And the high-level feature map F h , the low-level feature map is used to capture the detailed features of the termite target, and the high-level feature map is used to obtain the global context information. The generation of the multi-scale feature map is calculated by the following exemplary convolution transformation: Among them, I(u,v) represents the features of the input image or the output of the previous layer, is the convolution kernel weight coefficient of the lth layer, k l is the corresponding convolution kernel size, δ l is the stride when extracting low-level features; For the middle layer feature map F m And the high-level feature map F h The calculation method is the same, but with their own convolution kernel size k m ,k h and stride δ m ,δ h , so that feature maps at different levels have differences in resolution and receptive field; S42. For low-level feature maps F l , middle-level feature map F m And the high-level feature map F h Calculate the attention weight A separately l ,A m ,A h , and weight the multi-scale feature maps: F′ s (p,q)=α s ·A s (p,q)·F s (p,q),s∈{l,m,h}; Among them, F s Represents the original feature map at a certain scale, F′ s Represents the feature map after the introduction of attention, A s (p,q) represents the attention weight at the coordinate (p,q), α s is the magnification factor of the feature map of the same scale, and the attention weight A s (p,q) is dynamically generated by the adaptive attention module according to the difference between the termite micro-target and the background environment, which is used to highlight the micro-target area and suppress the complex background interference; S43. The low-level feature map F l , middle-level feature map F m And the high-level feature map F h Fusion is performed step by step to obtain the fusion feature map F merge : in, represents the element-wise addition operation, Φ(·) represents the convolution or linear transformation of the input features with uniform dimensional mapping, and β l ,β m ,β h is the weight factor of different scale features in the fusion process; S44. In the fusion feature map F merge Based on the above, a dynamic small target detection module is introduced. The dynamic small target detection module detects termites through dynamic prior frames and adaptive confidence thresholds based on their small size and susceptibility to background interference. The following prior frame update formula is used: Among them, Ω represents the candidate set of all prior box shapes, (w * ,h * ) is the recommended frame size data of the termite target extracted in steps S41-S43, (w * ,h * ) is the updated dynamic prior frame; S45. Input the enhanced training sample data set into the constructed target detection model and define the loss function L total For comprehensive constrained multi-scale detection, adaptive attention, and precise recognition of small objects: L total =λ1L loc +λ2L cls +λ3L conf +λ4L micro ·ψ(δ2); Among them, L loc ,L cls ,L conf Respectively represent the positioning loss, classification loss and confidence loss, L micro is the detection constraint term for tiny termites, λ1,λ2,λ3,λ4 are the weight coefficients of each loss term, and ψ(δ2) represents the correction function that changes with the dynamic confidence threshold δ2.

6. The method for identifying and monitoring drone low-altitude ants based on deep learning according to claim 5 is characterized in that: The positioning loss is used to measure the matching degree between the predicted bounding box and the real bounding box, and the weighted IoU is used as the loss function: Among them, IoU(B i,p ,B i,t ) represents the i-th predicted bounding box B i,p and the ground-truth bounding box B i,t The intersection-over-combination ratio, ω i is the weight of the i-th sample, which is used to pay extra attention to small targets, and N is the total number of samples; The classification loss is used to measure the difference between the predicted category distribution and the true category distribution, using a weighted cross entropy loss function: Among them, y i,c represents the true category label of the i-th sample, p i,c represents the category probability distribution predicted by the model, ω c is the category weight, which is used to optimize the termite target category, and C is the total number of categories; The confidence loss is used to evaluate whether the target box contains the termite target, using an improved focus loss function: Among them, p i is the confidence of the model prediction, y i is the true label, α1 and γ1 are balance parameters; The small target detection loss is specifically designed for the detection constraints of termite micro-targets, integrating detection confidence and bounding box error: Among them, N m is the number of samples of small targets, κ is the confidence correction coefficient, κ is the dynamic confidence threshold, ReLU (δ1-p i ) is used to penalize small target detection results with low confidence.

7. The method for identifying and monitoring drone low-altitude ants based on deep learning according to claim 5 is characterized in that: The S5 specifically includes the following contents: S51. quantize the trained target detection model into a lightweight model suitable for edge computing, and deploy the quantized lightweight model to the edge computing device carried by the drone; S52. The images and video data collected by the drone during low-altitude flight are transferred to the deployed target detection model in batches. The input data is inferred on each frame of the image, and the target detection model outputs the category probability distribution P t , Bounding box coordinates B t and confidence C t : B t =(x min ,y min ,x max ,y max ); Among them, Z c Represents the prediction score of category c, where C is the total number of categories; S53. According to the category probability distribution P in the inference result t , Bounding box coordinates B t and confidence C t Screening termite targets that meet the requirements: Among them, δ2 is the confidence threshold, which is dynamically adjusted by the correction function ψ(δ2), c is the category label, Used to determine the target category corresponding to the highest category probability, the filtered target set B s Contains the filtered bounding boxes and their category information.

Citation Information

Cited By

  • AI-based security and protection monitoring video analysis method and system

    CN120495963A