Wild animal dynamic multi-scale feature detection method under complex night condition based on improved DCAM-YOLO

The improved DCAM-YOLO model solves the problems of low recognition rate of small targets and poor adaptability to complex environments in wildlife detection, and realizes high-precision wildlife detection with low computational resources, which is suitable for intelligent monitoring of nature reserves and decision support for ecological protection.

CN120876887APending Publication Date: 2025-10-31JILIN INST OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510997778.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-19
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing deep learning models suffer from problems in wildlife detection, such as low recognition rate of small targets, poor adaptability to complex environments, and high computational resource requirements. In particular, their detection accuracy is not high in low light conditions at night and when targets are occluded.

Method used

An improved DCAM-YOLO model is adopted, which enhances the adaptability to targets of different scales by introducing a deformable large kernel attention mechanism (D-LKA). The SCINet module is combined to optimize the image quality in low light, the SEAM module enhances the ability to identify occluded targets, and a lightweight CGNet module is designed to integrate local details and global contextual information.

Benefits of technology

It significantly improves the accuracy and robustness of wildlife detection, especially in complex nighttime environments for detecting small targets and occlusions. The detection accuracy is improved by more than 30%, the false negative rate is reduced by 45%, the recall rate is improved by 25%, and the computational efficiency is improved by 40%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876887A_ABST
    Figure CN120876887A_ABST
Patent Text Reader

Abstract

The invention provides a high-precision wild animal detection method based on complex conditions at night, and belongs to the field of computer vision. The DCAM-YOLO combines dynamic multi-scale feature fusion and lightweight design, makes up for the defect of insufficient detection precision of a traditional method in a complex environment at night, and especially has outstanding performance when processing challenges such as target shielding, low illumination and dense distribution of small targets. According to the method, a deformable large kernel attention mechanism (D-LKA) is innovatively introduced, and the shape and size of a convolution kernel are dynamically adjusted to adapt to diversified target features; the low-illumination image quality is optimized through an SCINet module, and the response loss of a shielding target is compensated in combination with an SEAM module; and a lightweight downsampling module CGNet is designed, so that local details and global context information are effectively integrated. Experiments show that the detection precision and the real-time performance of the method are obviously superior to those of an existing model in a complex field environment, and more accurate and reliable technical support is provided for intelligent monitoring of wild animals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision, specifically relating to a high-precision wildlife detection method based on an improved DCAM-YOLO, applicable to scenarios such as intelligent monitoring of nature reserves, biodiversity research, anti-poaching patrols, wildlife behavior analysis, and ecological conservation decision support. This method is particularly designed for wildlife detection in complex nighttime environments, effectively addressing challenges such as low light, target occlusion, and dense distribution of small targets, providing an intelligent technological solution for wildlife conservation and research. Background Technology

[0002] With the rapid development of computer vision and deep learning technologies, intelligent monitoring systems are increasingly being used in ecological conservation, especially in automated wildlife detection. Wildlife detection, as a fundamental technology for biodiversity monitoring, habitat assessment, and ecological conservation decision-making, directly impacts the accuracy of subsequent population statistics and behavioral analysis. Accurate wildlife detection can improve the efficiency of protected area supervision and provide reliable data support for scientific research.

[0003] In complex outdoor environments, low light at night and extreme weather conditions (such as rain, snow, fog, and haze) can severely impact target detection performance. Furthermore, species diversity (such as differences in body size, coat color, and behavioral patterns) leads to unstable detection results. Occlusion phenomena (such as vegetation cover and animal herding) also frequently affect detection accuracy, especially under dense occlusion conditions, where detection algorithms may fail to accurately identify individual targets, resulting in missed or false detections.

[0004] Traditional object detection methods include feature engineering-based algorithms (such as Haar features, HOG features + SVM) and statistical learning-based algorithms (such as the DPM model). With the rise of deep learning, convolutional neural networks (CNNs) have become the mainstream method for object detection, especially the YOLO series and Faster R-CNN architectures, which have greatly promoted the development of this field. However, existing deep learning models still have problems in wildlife detection, such as low recognition rate of small targets, poor adaptability to complex environments, and high computational resource requirements. Summary of the Invention

[0005] To address the problems existing in current wildlife detection technologies, this invention proposes a high-precision wildlife detection method based on an improved DCAM-YOLO. This method enhances the model's adaptability to targets of different scales by introducing a deformable large kernel attention mechanism (D-LKA) to dynamically adjust the shape and size of the convolutional kernel; it improves nighttime detection performance by optimizing low-light image quality through the SCINet module; it enhances the model's ability to identify occluded targets by combining the attention network and repulsion loss of the SEAM module; and it designs a lightweight CGNet module to effectively integrate local details and global contextual information, improving the accuracy of small target detection.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A high-precision wildlife detection method based on an improved DCAM-YOLO includes the following steps:

[0008] Step 1, Construct the network model: The entire network consists of three main parts: the backbone network, the feature processing module, and the detection output module. The backbone network adopts an improved DarkNet-53 architecture, the feature processing module includes four core sub-modules: SCINet, D-LKA, SEAM, and CGNet, and the detection output module adopts a multi-scale feature fusion mechanism.

[0009] Step 2: Obtain a dataset of wildlife images;

[0010] Step 3: Input the wildlife dataset prepared in Step 2 into the DCMI-YOLO network model constructed in Step 1 for training to obtain the trained wildlife detection model.

[0011] The process of building the wildlife detection model includes:

[0012] Based on the initial feature map of the input image, low-light enhancement processing is performed through the SCINet module to extract the illumination-optimized feature map;

[0013] The illumination-optimized feature map is input into the D-LKA module, and multi-scale features are extracted through deformable large kernel convolution to obtain an adaptive receptive field feature map.

[0014] Based on the adaptive receptive field feature map, the spatial-channel dual attention mechanism of the SEAM module is used to generate the occlusion compensation feature map.

[0015] The occlusion compensation feature map is input into the CGNet module, and through multi-level downsampling (1 / 2, 1 / 4, 1 / 8) and residual connections, local details and global context information are fused to obtain a multi-scale fused feature map;

[0016] The multi-scale fused feature map is input into the detection head, and the final detection result is output through classification and regression branches.

[0017] The backbone network of the wildlife detection model adopts an improved DarkNet-53 architecture, using deformable convolutions in the last three convolutional layers to maintain the feature map resolution of 1 / 8 of the input size, while enhancing the feature extraction capability for irregular targets.

[0018] The wildlife detection model also includes:

[0019] 1) The SCINet module employs a cascaded illumination estimation network and a self-calibrating unit, comprising four cascaded stages and a weight-sharing mechanism;

[0020] 2) The D-LKA module integrates deformable convolution (adjustable kernels from 3×3 to 9×9) and a large kernel attention mechanism (7×7), and achieves adaptive feature extraction through dynamic offset learning;

[0021] 3) The SEAM module contains three parallel CSMM sub-modules at scales 6, 7, and 8, which combine repulsion loss (RepGT+RepBox) to handle occlusion problems;

[0022] 4) The CGNet module adopts a three-stage architecture (M=3, N=2), with each stage containing a context guidance unit and a feature preservation strategy, and only performs 3 downsampling operations to preserve spatial information.

[0023] Step 4: Select the minimum loss function and the optimal evaluation metric: Minimize the RepGT and RepBox loss functions until the number of training iterations reaches a set threshold or the value of the loss function reaches a set range. This indicates that the model parameters have been pre-trained and the model parameters are saved. At the same time, select mAP@0.5:0.95 and FPS as evaluation metrics to measure the performance of the algorithm.

[0024] Step 5, fine-tune the model: Train and fine-tune the model using a wildlife image dataset for a specific scenario to obtain stable and usable model parameters, further improving the model's detection capabilities;

[0025] Step 6, Save the model: Solidify the finalized model parameters. When wildlife detection is needed, simply input the image into the network to obtain the final detection result.

[0026] The backbone network of the wildlife detection model adopts an improved DarkNet-53 architecture, using deformable convolutions in the last three convolutional layers to maintain the ability to extract features from targets of different shapes.

[0027] The wildlife detection model also includes: a cascaded illumination estimation and self-calibration mechanism in the SCINet module; a deformable large kernel convolution (3×3 to 9×9 adjustable receptive field) and dynamic attention mechanism in the D-LKA module; a spatial-channel dual attention network and a repulsion loss compensation mechanism in the SEAM module; and a three-stage downsampling structure and multi-scale feature fusion strategy in the CGNet module.

[0028] The beneficial effects of this invention are:

[0029] The beneficial effects of this invention lie in significantly improving the accuracy and robustness of wildlife detection by introducing a Deformable Large Kernel Attention (D-LKA) mechanism, a SCINet low-light enhancement module, and a SEAM occlusion compensation module. Specifically, the D-LKA module enhances the model's adaptability to targets of different scales by dynamically adjusting the shape and size of the convolutional kernel, performing particularly well in the detection of small and distant targets. The SCINet module significantly improves image quality under low-light conditions through cascaded illumination learning and self-correction mechanisms, increasing the accuracy of wildlife detection at night by more than 30%. The SEAM module effectively solves the occlusion problem when animals are clustered, reducing the false negative rate in occluded scenes by 45% through an attention network and a repulsion loss function. The CGNet module integrates local details and global contextual information through a multi-scale feature fusion strategy, improving the recall rate of small target detection by 25%. The improved NMS algorithm, combined with an occlusion awareness mechanism, effectively filters out redundant detection boxes during the inference stage while retaining partially occluded targets, improving the accuracy of the detection results by 15%. This method demonstrates outstanding computational efficiency, reducing the number of model parameters by 40% compared to the traditional YOLOv8 while maintaining high accuracy, making it more suitable for deployment on edge computing devices. This invention maintains stable performance even in complex field environments such as extreme lighting, complex backgrounds, target occlusion, and scale variations, exhibiting broad applicability. In fields such as intelligent monitoring in nature reserves, biodiversity research, anti-poaching patrols, and wildlife behavior analysis, it can provide more accurate and reliable detection support for ecological conservation efforts. This technology has been deployed and tested in actual protected areas, successfully achieving real-time monitoring of 17 species of wild animals, including Siberian tigers and sika deer, providing crucial data support for ecological conservation decision-making, and demonstrating significant social benefits and market application prospects.

[0030] This invention has the following innovative features:

[0031] (1) Deformable Large Kernel Attention Module (D-LKA): By dynamically adjusting the shape of the convolution kernel and the size of the receptive field, the model's adaptability to targets of different scales is enhanced, especially in complex environments at night, where the detection performance of small targets and distant wild animals is significantly improved.

[0032] (2) SCINet Low Light Enhancement Module: It adopts a cascaded illumination learning and self-correction mechanism to effectively improve the brightness and contrast of nighttime images, solves the problem of detail loss in extremely dark environments in traditional methods, and improves the accuracy of nighttime detection by more than 35%.

[0033] (3) SEAM Occlusion Compensation Module: Through the spatial-channel dual attention mechanism and the repulsion loss function, the model’s ability to handle densely occluded scenes is significantly improved, reducing the false detection rate when animals are clustered by 40%, while effectively suppressing false detections.

[0034] (4) CGNet lightweight feature fusion module: adopts a multi-scale context-guided mechanism to integrate local detailed features and global semantic information. It significantly improves the detection effect of small target wild animals (such as mink and weasel) and increases the recall rate by 28%. Attached Figure Description

[0035] Figure 1 Flowchart of the detection method;

[0036] Figure 2 A schematic diagram of the DCAM-YOLO network structure of this invention;

[0037] Figure 3 A structural diagram of the deformable large-kernel attention module (D-LKA) of the present invention;

[0038] Figure 4 Comparison of the processing effects of the SCINet low-light enhancement module of the present invention;

[0039] Figure 5 A schematic diagram illustrating the working principle of the SEAM occlusion compensation module of the present invention;

[0040] Figure 6 Structure diagram of the CGNet lightweight feature fusion module of the present invention;

[0041] Figure 7 Comparison of wildlife detection results of the present invention (showing a comparison of detection results under different lighting and shading conditions);

[0042] Figure 8 Comparison chart of detection performance indicators of the present invention (comparison of key indicators such as mAP and FPS with other methods). Detailed Implementation

[0043] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0044] like Figure 1As shown, this invention provides a high-precision wildlife detection method based on an improved DCAM-YOLO, comprising the following steps:

[0045] Step 1: Data Acquisition and Preprocessing.

[0046] This invention collects infrared image datasets containing various wild animals and annotates each image with bounding boxes to generate label files. Data sources include nighttime images captured by infrared cameras in nature reserves (such as Changbai Mountain and Hunchun Nature Reserve) and publicly available datasets such as iWildCam, covering different seasons, weather conditions, and occlusion scenarios. During image preprocessing, adaptive histogram equalization is used to enhance contrast, and data augmentation operations such as random cropping, rotation, and flipping are applied to improve the model's adaptability to complex environments.

[0047] Step 2: Model building.

[0048] like Figure 2 , Figure 3 , Figure 4 , Figure 5 As shown, the model building phase consists of four core modules:

[0049] 1) SCINet Low-Light Enhancement Module: This module significantly improves nighttime image quality through cascaded illumination estimation and self-correction mechanisms. It includes: an illumination estimation sub-network that extracts illumination features via depthwise separable convolutions; a self-correction unit that establishes a bridge between initial observations and stage outputs; and a nonlinear mapping layer that uses the LeakyReLU activation function to preserve details in dark areas. Experiments show that this module can improve the PSNR of nighttime images by 8.2 dB, significantly improving subsequent detection performance.

[0050] 2) D-LKA Deformable Large Kernel Attention Module: Dynamically adjusts the shape and size of the convolution kernel to adapt to targets of different scales; This module enables the 7×7 large kernel convolution to adapt to changes in target shape through dynamic offset learning, improving the AP of small target detection by up to 15.6% while maintaining a large receptive field.

[0051] 3) SEAM Occlusion Compensation Module: Combining spatial-channel dual attention and repulsion loss, this module addresses the animal cluster occlusion problem. It includes three parallel CSMM sub-modules (patch-6 / 7 / 8), a spatial-channel dual attention mechanism, and a repulsion loss function (RepGT+RepBox). On the test set, this module reduced the false negative rate in dense scenes from 32.1% to 17.8%.

[0052] 4) CGNet Lightweight Feature Fusion Module: Multi-scale context-guided approach enhances small target detection capabilities. Through a three-stage feature preservation strategy, it maintains the ability to detect small targets (such as minks) with only three downsampling operations, achieving a 9.3% improvement in mAP@0.5.

[0053] The backbone network employs an improved DarkNet-53 architecture, introducing deformable convolutions in the last three convolutional layers to maintain 1 / 8 of the feature map resolution while enhancing adaptability to irregular animal poses. The Feature Pyramid Network (FPN) uses bidirectional cross-scale connections to deeply fuse low-level detailed features with high-level semantic features.

[0054] Step 3: Design the loss function.

[0055] This invention employs an improved combination of loss functions:

[0056] 1) RepGT loss: This loss moves the predicted bounding box away from the surrounding ground truth bounding boxes. Its function is to make the current bounding box as far away from the surrounding ground truth bounding boxes as possible. Here, "surrounding ground truth bounding boxes" refers to the bounding box that has the largest IOU with the face label other than the bounding box to be predicted.

[0057] 2) RepBox loss: Reduces overlap between predicted boxes; its function is to keep the current bounding box as far away as possible from the surrounding ground truth bounding boxes. Here, "surrounding ground truth bounding boxes" refers to the bounding box that has the largest loU with the face label other than the bounding box to be predicted.

[0058] 3) CIoU loss: Optimizes bounding box regression accuracy.

[0059] By optimizing the loss through multi-task collaboration, the detection performance of the model in complex scenarios can be improved.

[0060] Step 4: Model training and optimization.

[0061] During model training, the DCMI-YOLO model was trained end-to-end using a preprocessed wildlife dataset. A multi-task loss function (including classification loss, localization loss, and confidence loss) was employed to calculate the overall loss, and network parameters were updated via backpropagation. A dynamic learning rate adjustment strategy was implemented during training: the initial learning rate was set to 0.01, and a cosine annealing scheduler was used for periodic adjustments. A warm-up strategy was also used to gradually increase the learning rate over the first five epochs to ensure stability in the early stages of training. The optimizer used was SGD (Stochastic Gradient Descent) with Nesterov momentum (momentum = 0.9), and the weight decay coefficient was set to 0.0005 to control model complexity. Specifically, for the nighttime detection task, a curriculum learning strategy was introduced, training on normal lighting samples first and then gradually adding low-light samples to allow the model to gradually adapt to complex environments. Furthermore, mixed precision training (AMP) was employed, which improved training speed by 1.8 times and reduced memory usage by 40% while maintaining model accuracy, significantly improving training efficiency. A comprehensive evaluation is performed on the validation set every 10 epochs of training. The optimal model is saved based on the mAP@0.5:0.95 metric to ensure the robustness of the final model in complex field environments.

[0062] Step 5: Post-processing and reasoning.

[0063] An improved occlusion-aware NMS algorithm is adopted, which adjusts the IoU threshold by calculating the occlusion ratio of candidate boxes. The formula is: IoU_occ = IoU × (1 - α·OccRatio), where OccRatio is the occlusion ratio and α is the adjustment coefficient. This algorithm effectively preserves partially occluded targets while suppressing redundant detection boxes.

[0064] like Figure 8 As shown, the present invention achieves an mAP of 86.3% on the test set @ 0.5:0.95, which is 6.5% higher than the baseline model, while maintaining an inference speed of 35 FPS, meeting the needs of practical applications.

[0065] The above provides a detailed description of a high-precision wildlife detection method based on an improved DCAM-YOLO, as provided by this invention. The specific embodiments are described only to aid in understanding the method and its core principles. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims.

Claims

1. A high-precision wildlife detection method based on an improved DCAM-YOLO, characterized in that, Includes the following steps: S1. Collect wildlife image datasets: The dataset should include a variety of complex nighttime scenes, different species, occlusion conditions and lighting conditions to ensure the generalization ability of the model. Each image needs to be labeled with bounding boxes, and the labeling content includes animal category and location information. S2. Image preprocessing: Adaptive image enhancement technology is used to perform brightness correction and noise suppression on low-light images to avoid information loss. Operations such as random cropping, rotation, flipping, and color dithering are used to enhance the diversity of training data. S3. Construct an improved DCAM-YOLO model: Introduce a deformable large kernel attention mechanism (D-LKA) to dynamically adjust the shape and size of the convolution kernel to adapt to the features of different targets; optimize low-light image quality through the SCINet module; combine the SEAM module to compensate for the response loss of occluded targets; design a lightweight downsampling module CGNet to fuse local details and global contextual information; S4. Model Training: RepGT and RepBox loss functions are used to handle occlusion issues, combined with CIoU loss to optimize bounding box regression, and SGD optimizer is used with momentum set to 0.9 and weight decay set to 0.0005 to prevent overfitting; S5. Improved NMS algorithm: By calculating the occlusion-aware IoU of candidate boxes, redundant boxes are filtered out, improving the accuracy of detection results.

2. The method according to claim 1, characterized in that, The Deformable Large Kernel Attention Module (D-LKA) dynamically adjusts the sampling position of the convolution kernel, uses a large-size convolution kernel to extract wide-area receptive field features, and combines the dynamic adaptability of deformable convolution to enhance the model's ability to detect blurred targets at night and multi-scale targets.

3. The method according to claim 1, characterized in that, The SCINet module estimates and optimizes illumination components through cascaded illumination learning and weight sharing mechanisms, and combines self-calibration mapping to reduce computational burden and improve the quality of low-light images.

4. The method according to claim 1, characterized in that, The SEAM module compensates for the response loss of occluded targets through an attention network and a repulsion loss function, thereby enhancing the model's robustness to densely distributed and occluded scenes.

5. The method according to claim 1, characterized in that, The CGNet module integrates local features, neighborhood context, and global semantic information through a multi-scale feature fusion strategy, and uses residual learning to optimize gradient propagation, thereby improving the detection accuracy of small targets.

6. The method according to claim 1, characterized in that, The formula for calculating the RepGT loss function is as follows: ,in, For the predicted bounding box, GRep is the area surrounding the bounding box with the largest value. The true bounding box, This represents the proportion of the intersection of the predicted bounding box and the ground truth bounding box to the total ground truth bounding box.

7. The method according to claim 1, characterized in that, The The formula for calculating the loss function is: ,in, and Different prediction boxes are used to reduce overlap between them.

8. The method according to claim 1, characterized in that, The improved NMS algorithm uses occlusion detection. The similarity of candidate boxes is calculated using the following formula: ,in, α is the occlusion ratio, an adjustment factor used to suppress highly overlapping but occluded redundant boxes.