Wildlife detection method based on feature fusion reconstruction and lightweight design
By introducing FeatureGuideFPN and Dyhead feature fusion modules and dynamic detection heads in the YOLOv5 model, combined with model pruning and knowledge distillation technology, the problems of insufficient detection of traditional models in complex scenarios and difficulty in deploying edge devices are solved, and more efficient wildlife detection performance and lightweight design are achieved.
Patent Information
- Application Number
- CN202411359161.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-27
AI Technical Summary
When facing complex wildlife detection scenarios, traditional YOLOv5 models are difficult to accurately locate and identify partially obstructed targets, resulting in high missed detection rates and false detection rates. At the same time, the complex model structure is not conducive to the deployment of edge equipment.
Wildlife detection methods based on feature fusion reconstruction and lightweight design are adopted, including the construction of a new feature fusion module FeatureGuideFPN and dynamic detection head Dyhead, combining model pruning and knowledge distillation technology to optimize network structure and performance.
It improves the network's feature extraction and fusion capabilities, optimizes the positioning and classification of animal targets, realizes the lightweight design of the model, facilitates the deployment of edge devices, and improves the detection performance in complex environments.
Smart Images

Figure CN119229474B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection technology, and in particular to a wild animal detection method based on feature fusion reconstruction and lightweight design. Background Art
[0002] The protection and monitoring of wildlife is an important measure to protect biodiversity and maintain the balance of ecosystems. With the sharp decline in the number of wildlife worldwide, efficient detection technology has become essential. Traditional wildlife detection methods include manual surveys, GPS collar technology, and sound detection methods. By detecting specific sound signals emitted by animals, they can achieve real-time monitoring under long-term and changing climate conditions. However, these methods have the problems of high cost, high labor intensity, and great interference to animals. The limited information contained in the sound signal and the high noise in the natural environment make it complicated to extract enough information from the sound for accurate detection.
[0003] Compared with sound detection and traditional detection technologies, image-based wildlife detection methods provide richer and more intuitive information, and have significant advantages such as non-invasiveness, all-weather monitoring capabilities, high degree of automation, high cost-effectiveness and wide-area coverage. This method can reduce interference with the natural behavior of animals and reduce manpower requirements. Despite challenges such as image quality, data processing and hidden target detection, with the advancement of computer vision and machine learning technologies, image-based detection methods are gradually becoming an important tool for wildlife monitoring, providing strong support for biodiversity conservation and ecosystem maintenance.
[0004] Current target detection algorithms are mainly divided into two categories: one-stage algorithms and two-stage algorithms. One-stage target detection algorithms mainly include SSD and YOLO series, which have the advantage of fast detection speed, but lower accuracy than two-stage algorithms. Two-stage target detection algorithms mainly include FastR-CNN, FasterR-CNN, etc., which have the advantage of high detection accuracy, but slower detection speed. In recent years, one-stage target detection algorithms represented by the YOLO series have made significant progress in detection accuracy and speed.
[0005] Among image detection methods, the YOLOv5 algorithm has the characteristics of fast detection speed, lightweight model and high detection accuracy. However, although YOLOv5 performs well in general target detection tasks, it is not suitable for diverse and complex wildlife detection scenarios, such as in forests with complex backgrounds and changing lighting conditions, where animals may be obscured by leaves or only part of their bodies are exposed in groups. In this case, the traditional YOLOv5 model has difficulty in accurately locating and identifying partially obscured targets, and may result in high missed detection and false detection rates. Existing wildlife detection algorithms are complex in structure and are not conducive to edge device deployment.
[0006] In the traditional YOLOv5, PANet is used as a feature fusion module. The fused features have the problem of homogenizing importance and lack of context information; they may contain redundant information. The traditional Concat operation cannot distinguish the importance of different features, cannot reduce the impact of redundant information, and cannot better capture cross-channel and cross-layer context information.
[0007] The traditional YOLOv5 Head uses fixed convolution operations and feature fusion strategies when fusing features, which may not fully utilize feature information at different levels, especially when dealing with multi-scale targets, the effect may not be as good as the dynamically adjusted strategy. The traditional YOLOv5 Head usually does not integrate the attention mechanism and cannot effectively improve the feature representation of specific areas. This may lead to performance degradation in detection tasks with complex backgrounds or small objects. One of the design goals of YOLOv5 is real-time performance, so its Head part is relatively simple in design and has a relatively small number of parameters. However, this simplified design may limit the representation ability of the network, making it perform worse than a more complex Head when dealing with complex scenes.
[0008] In order to solve the above problems, the present invention proposes a wildlife detection method based on feature fusion reconstruction and lightweight design. Summary of the invention
[0009] The purpose of the present invention is to propose a wildlife detection method based on feature fusion reconstruction and lightweight design to solve the problems raised in the background technology:
[0010] PANet used in the traditional YOLOv5 network cannot effectively extract important feature information of the image; the traditional detection head cannot make full use of information at different levels and cannot effectively improve the feature representation of specific areas; there is an imbalance between the model's computational complexity and performance.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] The wildlife detection method based on feature fusion reconstruction and lightweight design includes the following steps:
[0013] S1: Collect wildlife images and perform preprocessing operations;
[0014] S2: Build the YOLOv5 network, including the feature extraction module, feature fusion module and detection head;
[0015] S3: Build a new feature fusion module FeatureGuideFPN to replace the original feature fusion module, and replace the original detection head with a dynamic detection head Dyhead that combines scale, space, and channel attention;
[0016] S4: Use the model pruning method to perform structured pruning on the improved network to lightweight the improved model;
[0017] S5: Use model distillation method to fine-tune the pruned model;
[0018] S6: Wildlife detection based on the improved YOLOv5 network.
[0019] Preferably, the feature extraction module of the YOLOv5 network in S2 includes a Conv module, a C3 module and a SPPF module.
[0020] Preferably, the FeatureGuideFPN in S3 is obtained by replacing the Concat module in the original feature fusion network with a FGM module.
[0021] Preferably, the FGM module includes a feature fusion unit, a feature weighting unit, a weighted feature unit and a feature summing unit;
[0022] The feature fusion unit fuses two different feature inputs, feature 1 and feature 2, together through a connection operation to form a new feature representation;
[0023] The feature weighting unit passes the fused features through an average pooling layer to globally average the features of each channel; passes through a fully connected layer with a ReLU activation function; and then uses a Hard-sigmoid function to limit the values in the feature vector to a certain range;
[0024] The weighted feature unit uses the obtained weight vector to perform channel-by-channel weighting on the original feature 1 and the original feature 2;
[0025] The feature summing unit performs an element-by-element addition operation on the weighted features to fuse the original features with the weighted features.
[0026] Preferably, the dynamic detection head Dyhead in S4 uses an attention mechanism to enhance the feature representation and learning ability of the target detection model.
[0027] Preferably, the attention mechanism used by the dynamic detection head Dyhead includes scale-aware attention, space-aware attention and task-aware attention;
[0028] The attention mechanism W(F) is specifically as follows:
[0029] W(F)=π C (π S (π L (F)·F)·F)·F
[0030] Among them, F∈R L×S×C is the feature tensor, L, S and C are the scale dimension, spatial dimension and channel dimension respectively; π L (·), π S (·) and π C (·) is the attention function corresponding to the three different dimensions of scale, space, and channel;
[0031] The scale-aware attention π L (F) are as follows:
[0032]
[0033] Where f(·) is a linear function approximated by a 1×1 convolution; σ is a Hard-sigmoid function;
[0034] The spatial perception attention π S (F) are as follows:
[0035]
[0036] Where K is the number of sparse sampling positions; p k is the sampling position; Δp k is the position change; Δm k is at position p k An important scalar for self-learning;
[0037] The task perceives attention π C (F) are as follows:
[0038] π C (F)·F=max(α 1 (F)·F c +β 1 (F),α 2 (F)·F c +β 2 (F)
[0039] Among them, F cis the feature slice of the cth channel in the feature tensor F, and the control parameter α of the activation function is learned through the hyperfunction θ(·) 1 , β 1 , α 2 , β 2 .
[0040] Preferably, the S4 uses the Group-Taylor pruning method for offline channel pruning in structured pruning, by removing redundant channels in the convolutional layer to compress the width of the model, and whether the channel is redundant is determined by the importance score of the channel;
[0041] The importance score is calculated by Taylor series; the change of the loss function after removing the specified neuron is calculated using the first-order and second-order Taylor expansions, thereby obtaining the importance score of the neuron; specifically, as follows:
[0042] First-order Taylor expansion for:
[0043]
[0044] Among them, w m is the parameter vector; is the loss function E and parameter w m The partial derivative of
[0045] Second-order Taylor expansion for:
[0046]
[0047] Among them, g m is the gradient; T is the transpose of the vector; H m is the mth row of the Hessian matrix of the loss function;
[0048] It also reduces the complexity and computational cost of the network by iteratively removing the neurons that contribute the least to the final loss;
[0049] The contribution of removing a particular neuron to the final loss is defined by calculating the square of the change in the loss function after removing a parameter; the formula is as follows:
[0050]
[0051] Among them, E(D,W) represents the loss function on the entire data set D; W is the parameter set of the network; Indicates that the parameter w is removed m The loss function after removing the parameter w m The square of the change in the loss function caused by this is used as a measure of the importance of this parameter.
[0052] Preferably, in S5, a channel knowledge distillation method is used to fine-tune the pruned model;
[0053] The model before model pruning is selected as the teacher model, and the model after model pruning is selected as the student model; the output of the teacher network is the soft label; the loss function for calculating the difference between the student network output and the soft label is the distillation loss function, and the loss function for calculating the difference between the student network output and the true label is the student loss function; the total loss function calculation formula is:
[0054] Loss=student loss+βdistillation loss
[0055] Among them, β is the proportion of distillation loss;
[0056] The distillation loss of the channel knowledge distillation method is calculated using the channel distillation loss, which measures the difference between the outputs of two models by calculating the KL divergence; the activation map is converted into a probability distribution using the softmax function, and the temperature coefficient T is introduced to control the smoothness of the probability distribution. The formula is as follows:
[0057]
[0058] Among them, y c represents the activation map of channel c in the teacher network or student network; H×W represents the spatial dimension of the activation map; φ(y c ) is the probability distribution of channel c obtained by converting the softmax function;
[0059] The KL divergence is calculated as follows:
[0060]
[0061] Among them, φ(φ(y T ),φ(y S )) is the calculated KL divergence value; φ(y T ) and φ(y S ) are the probability distributions of the teacher network and the student network obtained by the softmax function conversion; φ(y T,c,i ) and φ(y S,c,i ) are the activation values of the teacher network and the student network at position i in channel c respectively;
[0062] By minimizing the channel distillation loss, the student network is made to learn to imitate the channel-level feature representation of the teacher network.
[0063] Compared with the prior art, the present invention provides a wildlife detection method based on feature fusion reconstruction and lightweight design, which has the following beneficial effects:
[0064] The present invention uses a new feature fusion network of the YOLOv5 model to improve the network's feature extraction and feature fusion capabilities; introduces a dynamic detection head based on a multiple attention mechanism into its head to optimize the positioning and classification of animal targets; and also prunes the improved model and uses knowledge distillation to adjust the model performance to ensure that the improved model is lightweight while maintaining the same performance. The present invention improves its detection performance in complex environments by introducing a new model structure and data enhancement technology, and performs a lightweight design on the model to facilitate the deployment of the model on edge devices, so as to provide more reliable support for wildlife protection and research. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a schematic diagram of the model structure of YOLOv5 based on feature fusion reconstruction and lightweight design mentioned in Example 1 of the present invention;
[0066] Figure 2 It is a structural diagram of the FeatureGuideModule (FGM) mentioned in Example 1 of the present invention;
[0067] Figure 3 This is a schematic diagram of the Dyhead structure mentioned in Example 1 of the present invention;
[0068] Figure 4 This is a schematic diagram of the pruning principle mentioned in Example 1 of the present invention;
[0069] Figure 5 This is a schematic diagram of the channel distillation mentioned in Example 1 of the present invention;
[0070] Figure 6 This is the knowledge distillation flow chart mentioned in Example 1 of the present invention. DETAILED DESCRIPTION
[0071] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0072] The present invention builds a new feature fusion network of the YOLOv5 model to improve the network's feature extraction and feature fusion capabilities; introduces a dynamic detection head based on a multiple attention mechanism into its head to optimize the positioning and classification of animal targets; and prunes the improved model and uses knowledge distillation to adjust the model performance to ensure that the improved model is lightweight while maintaining unchanged performance. The present invention improves its detection performance in complex environments by introducing a new model structure and data enhancement technology, and performs a lightweight design on the model to facilitate the deployment of the model on edge devices, so as to provide more reliable support for wildlife protection and research. Specifically, it includes the following contents.
[0073] Embodiment 1:
[0074] See also Figure 1-6 The present invention provides a wildlife detection method based on feature fusion reconstruction and lightweight design, comprising the following steps:
[0075] S1: Collect wildlife images and perform preprocessing operations; the details are as follows:
[0076] For the animal targets that need to be detected, the labelimg annotation tool is used to annotate these images, and 16,554 images with detailed annotation information are obtained.
[0077] The existing dataset is enhanced. The data enhancement methods include adding noise, changing brightness, cropping, translation, rotation, mirroring, cutout, and flipping. The enhanced dataset contains a total of 30,694 animal images.
[0078] Divide the dataset. The training set contains 21,485 animal images. The validation set contains 3,069 images. The test set contains 6,140 images.
[0079] S2: Build the YOLOv5 network, which includes the feature extraction module, feature fusion module and detection head; the details are as follows:
[0080] The YOLOv5 network is mainly composed of three parts, namely Backbone, Neck, and Head, which correspond to the feature extraction, feature fusion, and detection head of the network. The Backbone is New CSP-Darknet53, and the main structures in the Backbone are Conv module, C3 module, and SPPF module.
[0081] S3: Build a new feature fusion module FeatureGuideFPN to replace the original feature fusion module, and replace the original detection head with a dynamic detection head Dyhead that combines scale, space, and channel attention; the details are as follows:
[0082] Neck uses a new feature fusion network FeatureGuideFPN to improve the network's feature extraction and feature fusion capabilities. Secondly, a dynamic detection head (Dyhead) based on a multi-attention mechanism is introduced to the head to optimize the model's positioning and classification of animal targets.
[0083] FeatureGuideFPN is obtained by replacing the Concat module in the original feature fusion network with the FGM module. Figure 2 , the FGM module realizes the dynamic adjustment of the importance of different channel features through feature fusion and weighting, so that the model can more flexibly highlight key features while retaining comprehensive information. The FGM module includes a feature fusion unit, a feature weighting unit, a weighted feature unit and a feature summing unit; the feature fusion unit fuses two different feature inputs, feature 1 and feature 2, through a connection operation to form a new feature representation; the feature weighting unit passes the fused features through an average pooling layer to globally average the features of each channel; passes through a fully connected layer with a ReLU activation function; and then limits the values in the feature vector to a certain range through the Hard-sigmoid function; the weighted feature unit uses the obtained weight vector to weight the original feature 1 and the original feature 2 channel by channel; the feature summing unit performs element-by-element addition operation on the weighted features to fuse the original features with the weighted features.
[0084] The dynamic detection head Dyhead uses the attention mechanism to enhance the feature representation and learning ability of the target detection model. Figure 3 ,The attention mechanisms used by the dynamic detection head Dyhead include scale-aware attention, space-aware attention, and task-aware attention;
[0085] The attention mechanism W(F) is as follows:
[0086] W(F)=π C (π S (π L (F)·F)·F)·F
[0087] Among them, F∈R L×S×C is the feature tensor, L, S and C are the scale dimension, spatial dimension and channel dimension respectively; π L (·), π S (·) and π C (·) is the attention function corresponding to the three different dimensions of scale, space, and channel;
[0088] Scale-aware attention: DyHead uses a scale-aware attention module to dynamically adjust the importance of features at different levels. This attention mechanism allows the model to adaptively highlight features that are more important for a specific target scale, thereby improving the detection ability of multi-scale targets. L (F) are as follows:
[0089]
[0090] Where f(·) is a linear function approximated by a 1×1 convolution; σ is a Hard-sigmoid function;
[0091] Spatial-aware attention: The spatial-aware attention module enables the model to focus on specific areas in the image that are related to the object. In this way, the model can better understand the spatial layout and shape changes of the object. Spatial-aware attention π S (F) are as follows:
[0092]
[0093] Where K is the number of sparse sampling positions; p k is the sampling position; Δp k is the position change; Δm k is at position p k An important scalar for self-learning;
[0094] Task-aware attention: The task-aware attention module allows DyHead to adjust the response of feature channels according to different detection tasks (e.g., classification, bounding box regression). This mechanism enables the model to assign different attention weights to different tasks. Task-aware attention π C (F) are as follows:
[0095] π C (F)·F=max(α 1 (F)·F c +β 1 (F),α 2 (F)·F c +β 2 (F)
[0096] Among them, F c is the feature slice of the cth channel in the feature tensor F, and the control parameter α of the activation function is learned through the hyperfunction θ(·) 1 , β 1 , α 2 , β 2 .
[0097] Unified framework of attention mechanisms: DyHead integrates scale-aware, space-aware, and task-aware attention into a unified framework, so that these attention mechanisms can complement each other and jointly improve the performance of object detection.
[0098] Using the feature weighting mechanism of the FGM (FeatureGuideModule) structure, dynamic weights can be assigned to each feature, thereby enhancing the expression of important features and weakening the influence of unimportant features. This mechanism helps the network better focus on key information and improves the effectiveness of feature extraction. By recalibrating channel features, this structure can capture more contextual information and cross-channel dependencies. This is particularly useful for dealing with complex object detection tasks because it can help the model better identify and distinguish similar objects or backgrounds. The feature weighting mechanism helps suppress irrelevant features or noise and reduce their negative impact on the final detection results. This can improve the robustness of the model and perform better when dealing with complex scenes or noisy data. The dynamic weight mechanism can better capture fine-grained information. For example, in object detection, small objects or parts with details may be more easily emphasized, thereby improving the accuracy of detection.
[0099] Input the preprocessed images into the model for training. Set the number of training times to 500, the initial learning rate to 0.01, and the batch size of each training image to 32. During the training process, update the training weights to minimize the loss. In each round of training, update the training weights of this round with the best training weights so far.
[0100] S4: Use the model pruning method to perform structured pruning on the improved network to lightweight the improved model; the details are as follows:
[0101] Reference Figure 4 , Group Taylor pruning technology is introduced to reduce the number of model parameters and the amount of calculation. Model pruning is mainly used to compress and accelerate the model, and to reduce the computational burden and model size by removing unimportant parts of the model. This embodiment uses the Group-Taylor pruning method as offline channel pruning in structured pruning, which compresses the width of the model by removing redundant channels in the convolutional layer, and whether the channel is redundant is judged by the importance score of the channel.
[0102] The importance score is estimated by Taylor series. The first-order and second-order Taylor expansions are used to approximate the change in the loss function after removing the specified neuron, thereby obtaining the importance score of the neuron. The details are as follows:
[0103] First-order Taylor expansion for:
[0104]
[0105] Among them, w m is the parameter vector; is the loss function E and parameter w m The partial derivative of
[0106] Second-order Taylor expansion for:
[0107]
[0108] Among them, g m is the gradient; T is the transpose of the vector; H m is the mth row of the Hessian matrix of the loss function;
[0109] The complexity and computational cost of the network are reduced by iteratively removing the neurons that contribute the least to the final loss. The contribution of a particular neuron to the final loss after removal is defined by calculating the square of the change in the loss function after removing a parameter (or a group of parameters, such as a neuron). The formula is as follows:
[0110]
[0111] Among them, E(D,W) represents the loss function on the entire data set D; W is the parameter set of the network; Indicates that the parameter w is removed m The loss function after removing the parameter w m The square of the change in the loss function caused by this is used as a measure of the importance of this parameter.
[0112] A trained network is fed into the pruning program and pruned during iterative fine-tuning. The following steps are repeated in each cycle:
[0113] 1) For each mini-batch, the parameter gradients are calculated and the network weights are updated by gradient descent. At the same time, the Taylor series is used to approximate the importance score of each neuron.
[0114] 2) After a predetermined number of mini-batches, the importance scores are accumulated and calculated, the importance scores of each neuron are averaged, and the N neurons with the smallest importance scores are removed.
[0115] 3) Continue fine-tuning and pruning until the target number of neurons is pruned or the training loss exceeds the maximum acceptable loss.
[0116] S5: Use the model distillation method to fine-tune the pruned model. The details are as follows:
[0117] The knowledge distillation technique is used to transfer the knowledge of the original unpruned model to the pruned model to further improve its performance.
[0118] Knowledge distillation is a model compression technology that trains a lightweight student network to imitate the behavior of a well-trained, more complex and better-performing teacher network. In order to improve the performance of the pruned model as much as possible without increasing the number of parameters and computational complexity, knowledge distillation is performed on the pruned model. This article uses the channel knowledge distillation method (CWD). The schematic diagram of channel distillation is shown in the figure. Figure 5 shown.
[0119] The schematic diagram of the knowledge distillation process is as follows: Figure 6 As shown. The model before model pruning is selected as the teacher model, and the model after model pruning is selected as the student model; the output of the teacher network is called soft labels. The loss function that calculates the difference between the student network output and the soft labels is the distillation loss function, and the loss function that calculates the difference between the student network output and the true labels (hard labels) is the student loss function. The total loss function calculation formula is:
[0120] Loss=student loss+βdistillation loss
[0121] Among them, β is the proportion of distillation loss;
[0122] The calculation method of the distillation loss in this embodiment uses the channel distillation loss (CWD-Loss), which measures the difference between the outputs of two models by calculating the KL divergence. In the CWD method, the softmax function is used to convert the activation map into a probability distribution, and the temperature coefficient T is introduced to control the smoothness of the probability distribution. The formula is as follows:
[0123]
[0124] Among them, y c represents the activation map of a channel c in the teacher network or the student network; H×W represents the spatial dimension of the activation map; φ(y c ) is the probability distribution of channel c obtained by softmax function transformation; T is the temperature coefficient, which controls the "sharpness" or "smoothness" of the probability distribution. The KL divergence is calculated as follows:
[0125]
[0126] Among them, φ(φ(y T ),φ(y S )) is the calculated KL divergence value; φ(y T ) and φ(y S) are the probability distributions of the teacher network and the student network obtained by the softmax function conversion; φ(yT,c,i) and φ(y S,c,i ) are the activation values of the teacher network and the student network at position i in channel c, respectively; by minimizing the channel distillation loss, the student network learns to imitate the channel-level feature representation of the teacher network.
[0127] S6: Wildlife detection based on the improved YOLOv5 network. The details are as follows:
[0128] The improved YOLOv5 network is used to detect wild animals with YOLOv5n, YOLOv7-tiny, YOLOv8, and YOLOv3-tiny, and the detection results are compared, as shown in Table 1:
[0129] Table 1 Comparison of detection results of different detection algorithms
[0130] Model P R mAP@0.5 mAP@0.5:0.95 GFLOPS YOLOv5n 0.941 0.918 0.955 0.687 4.5 YOLOv7-tiny 0.95 0.91 0.961 0.65 13.2 YOLOv8 0.965 0.957 0.977 0.775 8.1 YOLOv3-tiny 0.913 0.873 0.929 0.593 13.0 Improving YOLOv5 0.973 0.964 0.978 0.762 4.2
[0131] As can be seen from the table, the accuracy of the improved algorithm in this embodiment is improved by 3.2%, 2.3%, 0.8%, and 6%, respectively, and the recall rate is improved by 4.6%, 5.4%, 0.7%, and 9.1%, respectively. mAP@0.5 is also improved compared with the other four algorithms. mAP@0.5:0.95 is improved compared with YOLOv5n, YOLOv7-tiny, and YOLOv3-tiny. Compared with YOLOv8, mAP@0.5:0.95 is reduced by 1.3%, but the model in the present invention is 48% lower than YOLOv8 in terms of computational complexity. Since the equipment used in wildlife detection is generally an edge device with limited computing power, it is more advantageous when the accuracy is similar. From the above analysis, it can be seen that the improved model can better meet the needs of wildlife detection.
[0132] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A wildlife detection method based on feature fusion reconstruction and lightweight design, characterized in that: The steps include: S1: Collect wildlife images and perform preprocessing operations; S2: Build the YOLOv5 network, including the feature extraction module, feature fusion module and detection head; S3: Build a new feature fusion module FeatureGuideFPN to replace the original feature fusion module, and replace the original detection head with a dynamic detection head Dyhead that combines scale, space, and channel attention; the FeatureGuideFPN is obtained by replacing the Concat module in the original feature fusion network with the FGM module; the FGM module includes a feature fusion unit, a feature weighting unit, a weighted feature unit, and a feature summing unit; The feature fusion unit fuses two different feature inputs, feature 1 and feature 2, together through a connection operation to form a new feature representation; The feature weighting unit passes the fused features through an average pooling layer to globally average the features of each channel; passes through a fully connected layer with a ReLU activation function; and then uses a Hard-sigmoid function to limit the values in the feature vector to a certain range; The weighted feature unit uses the obtained weight vector to perform channel-by-channel weighting on the original feature 1 and the original feature 2; The feature summing unit performs an element-by-element addition operation on the weighted features to merge the original features with the weighted features; S4: Use the model pruning method to perform structured pruning on the improved network to lightweight the improved model; S5: Fine-tune the pruned model using model distillation. S6: Wildlife detection based on the improved YOLOv5 network.
2. The method for detecting wild animals based on feature fusion reconstruction and lightweight design according to claim 1 is characterized in that: The feature extraction module of the YOLOv5 network in S2 includes a Conv module, a C3 module and a SPPF module.
3. The method for detecting wild animals based on feature fusion reconstruction and lightweight design according to claim 1 is characterized in that: The dynamic detection head Dyhead in S4 uses an attention mechanism to enhance the feature representation and learning ability of the target detection model.
4. The method for detecting wild animals based on feature fusion reconstruction and lightweight design according to claim 3 is characterized in that: The attention mechanism used by the dynamic detection head Dyhead includes scale-aware attention, space-aware attention and task-aware attention; The attention mechanism W(F) is specifically as follows: W(F)=π C (p S (p L (F)·F)·F)·F Among them, F∈R L×S×C is the feature tensor, L, S and C are the scale dimension, spatial dimension and channel dimension respectively; π L (·), π S (·) and π C (·) is the attention function corresponding to the three different dimensions of scale, space, and channel; The scale-aware attention π L (F) are as follows: Where f(·) is a linear function approximated by a 1×1 convolution; σ is a Hard-sigmod function; The spatial perception attention π S (F) are as follows: Where K is the number of sparse sampling positions; p k is the sampling position; Δp k is the position change; Δm k is at position p k An important scalar for self-learning; The task perceives attention π C (F) are as follows: p C (F)·F=max(α 1 (F)·F c +b 1 (F),a 2 (F)·F c +b 2 (F)) Among them, F c is the feature slice of the cth channel in the feature tensor F, and the control parameter α of the activation function is learned by the hyperfunction θ(·) 1 , β 1 , α 2 , β 2 .
5. The method for detecting wild animals based on feature fusion reconstruction and lightweight design according to claim 1 is characterized in that: The S4 uses the Group-Taylor pruning method for offline channel pruning in structured pruning, which compresses the width of the model by removing redundant channels in the convolutional layer, and whether the channel is redundant is determined by the importance score of the channel; The importance score is calculated by Taylor series; the change of the loss function after removing the specified neuron is calculated using the first-order and second-order Taylor expansions, thereby obtaining the importance score of the neuron; specifically, as follows: First-order Taylor expansion for: Among them, w m is the parameter vector; is the loss function E and parameter w m The partial derivative of Second-order Taylor expansion for: Among them, g m is the gradient; T is the transpose of the vector; H m is the mth row of the Hessian matrix of the loss function; It also reduces the complexity and computational cost of the network by iteratively removing the neurons that contribute the least to the final loss; The contribution of removing a particular neuron to the final loss is defined by calculating the square of the change in the loss function after removing a parameter; the formula is as follows: Among them, E(D,W) represents the loss function on the entire data set D; W is the parameter set of the network; Indicates that the parameter w is removed m The loss function after removing the parameter w m The square of the change in the loss function caused by this is used as a measure of the importance of this parameter.
6. The method for detecting wild animals based on feature fusion reconstruction and lightweight design according to claim 1, characterized in that: In S5, the channel knowledge distillation method is used to fine-tune the pruned model; The model before model pruning is selected as the teacher model, and the model after model pruning is selected as the student model; the output of the teacher network is the soft label; the loss function for calculating the difference between the student network output and the soft label is the distillation loss function distillationloss, and the loss function for calculating the difference between the student network output and the true label is the student loss function studentloss; the total loss function calculation formula is: Loss=studentloss+βdistillationloss Among them, β is the proportion of distillation loss; The distillation loss of the channel knowledge distillation method is calculated using the channel distillation loss, which measures the difference between the outputs of the two models by calculating the KL divergence; the activation map is converted into a probability distribution using the softmax function, and the temperature coefficient T is introduced to control the smoothness of the probability distribution. The formula is as follows: Among them, y c represents the activation map of channel c in the teacher network or student network; H×W represents the spatial dimension of the activation map; φ(y c ) is the probability distribution of channel c obtained by converting the softmax function; The KL divergence is calculated as follows: Among them, φ(φ(y T ),φ(y S )) is the calculated KL divergence value; φ(y T ) and φ(y S ) are the probability distributions of the teacher network and the student network obtained by the softmax function conversion; φ(y T,c,i ) and φ(y S,c,i ) are the activation values of the teacher network and the student network at position i in channel c respectively; By minimizing the channel distillation loss, the student network is made to learn to imitate the channel-level feature representation of the teacher network.
Citation Information
Patent Citations
Target detection method and system based on multi-scale feature map reconstruction and knowledge distillation
CN111626330A
Knowledge distillation-based YOLOv5 target detection method
CN115631396A