Small-Sample Wildlife Detection Method Based on Improved YOLOv5

By improving the YOLOv5 network model and combining the coordinate attention module CA and two-stage training method, the problem of insufficient wildlife samples is solved, high-precision wildlife detection is achieved, and labor costs are reduced.

CN115393618BActive Publication Date: 2025-07-22ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211018085.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-07-22
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Existing wildlife detection methods are difficult to achieve high-precision detection when there are insufficient samples, especially the detection effect of rare species is poor.

Method used

The improved YOLOv5 network model is adopted, and the coordinate attention module CA is added and a two-stage training method is adopted to screen and annotate the real wildlife data set, and the YOLOv5-CA-TL network model is constructed for detection.

Benefits of technology

It improves the accuracy of detection of small samples of wildlife, reduces labor costs, and improves the generalization ability of the network, which is better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393618B_ABST
    Figure CN115393618B_ABST
Patent Text Reader

Abstract

The present invention relates to a small-sample wild animal detection method based on improved YOLOv5, including: screening the AwA2 animal dataset according to the collected real wild animal dataset of the area to be detected, and using the screened images as the experimental dataset; screening the collected real wild animal dataset of the area to be detected to obtain a small-sample experimental dataset; annotating the experimental dataset and the small-sample experimental dataset; adding a coordinate attention module CA to the YOLOv5 network model to obtain a YOLOv5-CA network model; adopting a two-stage training method to obtain a YOLOv5-CA-TL network model, detecting the real wild animals in the area to be detected and conducting feasibility verification. By taking the collected real wild animal dataset of the area to be detected as the research object, the present invention effectively solves the problem of low detection accuracy of some other algorithms by using the coordinate attention module CA, and greatly reduces the labor cost compared with the traditional method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and object detection and classification, and in particular to a small-sample wild animal detection method based on improved YOLOv5. Background Art

[0002] Biological resources are the natural foundation for humans to achieve sustainable development and a powerful guarantee for the balance and stability of the ecosystem. Therefore, it is necessary to continuously monitor and protect wild animals. The wild environment is complex and changeable, with many unknown risks, making it difficult to carry out protection actions for wild animals. With the development of technology, various modern technologies have been developed for wild animal monitoring, including radio tracking, wireless sensor network tracking, satellite and global positioning system (GPS) tracking, and monitoring through motion-sensitive cameras. With the progress of digital technology, infrared cameras are widely used in wild animal detection, facilitating shooting without disturbing the normal rest of animals.

[0003] Infrared cameras automatically collect images and videos in the wild for a long time, generating a large amount of image and video data. Manually processing such a large amount of images is also very time-consuming. In the field of object detection, the dataset has a direct impact on the detection effect. A dataset with a rich and balanced number of samples is conducive to completing the detection and classification tasks. However, some wild animal nature reserves are vast, and many wild animals are rare species on the verge of extinction. The number of rare animals is small and their whereabouts are erratic. It is difficult for infrared cameras to capture clear images of animals. Therefore, the number of wild animal samples that can be truly used for research in the images captured by infrared cameras is very small.

[0004] Currently, existing wild animal detection and classification methods at home and abroad are all based on sufficient samples. However, in the case of insufficient wild animal samples, these methods are difficult to well complete the wild animal detection and classification tasks. How to improve the detection accuracy of wild animals in the case of insufficient samples is a technical problem that urgently needs to be solved in the current wild animal detection field. Summary of the Invention

[0005] The purpose of the present invention is to provide a small-sample wild animal detection method based on improved YOLOv5 that can solve the problem of insufficient wild animal samples, can better improve the overall performance of object detection, and can effectively improve the object detection accuracy of small-sample wild animals.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A small-sample wild animal detection method based on improved YOLOv5, the method includes the following steps in sequence:

[0007] (1) Download the AwA2 animal dataset, and screen the AwA2 animal dataset according to the collected real wild animal dataset of the area to be detected, and select the images similar to the real wild animal dataset of the area to be detected as the experimental dataset, and divide the experimental dataset into the first training set and the validation set;

[0008] (2) Screen the collected real wild animal dataset of the area to be detected, and select the photos with clear animal pictures and high pixel quality as the small sample experimental dataset, and divide the small sample experimental dataset into the second training set and the test set;

[0009] (3) Label the experimental dataset and the small sample experimental dataset;

[0010] (4) Build a YOLOv5 network model, add a coordinate attention module CA to the YOLOv5 network model to obtain a YOLOv5-CA network model, and input the first training set into the YOLOv5-CA network model for training;

[0011] (5) Use a two-stage training method to obtain a YOLOv5-CA-TL network model, input the labeled test set into the YOLOv5-CA-TL network model, detect the wild animals in the picture, and conduct a feasibility verification.

[0012] In step (3), the labeling refers to using the Labelimg tool to label the animals in the pictures of the experimental dataset and the small sample experimental dataset, and the labeling format is the yolo format.

[0013] The specific steps of step (4) include the following steps:

[0014] (4a) Replace the first C3 module and the last C3 module of the backbone network of the YOLOv5 network model with the coordinate attention module CA. Then, encode the feature maps respectively to form two feature maps, which are sensitive to direction perception and position respectively; Take any intermediate tensor X = [x1, x2, x3..., x C ∈ R C×H×W as the input, and output a tensor Y = [y1, y2, y3..., y c of the same length. Use pooling kernels of size (H, 1) and (1, W) on X to encode each channel along the horizontal and vertical directions. The output of the c-th channel with height h is as follows:

[0015]

[0016] The output of the c-th channel with width w is as follows:

[0017]

[0018] Wherein, H is the height of the pooling kernel, W is the width of the pooling kernel, X C is the tensor of the c-th channel, and h is the height of X;

[0019] Formulas (1) and (2) are two transformations of feature aggregation. They perform aggregation along two spatial directions respectively and return two direction-aware attention maps. The coordinate attention module CA generates two feature layers before concatenation and then shares a 1×1 convolution operation for transformation, as shown in Formula (3):

[0020] f = δ(F1([Z h , Z w )) (3)

[0021] In the formula, δ is a non-linear activation function, F1 is a 1×1 convolution transformation, f is an intermediate feature map, which is the result of feature encoding of spatial information in the horizontal and vertical directions; then f is decomposed into 2 separate tensors along the spatial dimension, f h ∈R C / r×H and f w ∈R C / r×W , and then two other 1×1 convolution transformations F h and F w are used to transform f h and f w into tensors with the same number of feature layers as X respectively, obtaining:

[0022] g h = σ(F h (f h )) (4)

[0023] g w = σ(F w (f w )) (5)

[0024] In the formula, F h , F w are both 1×1 convolution transformations, σ is a sigmoid activation function. During the transformation process, the reduction ratio r is used to reduce the number of channels of f division, and then the outputs g h and g w are expanded and used as weighting weights respectively. The final output of the coordinate attention module CA is as shown in Formula (6):

[0025]

[0026] In the formula, y c (i, j) represents the value of (i, j) in the c-th channel of the output, and x c (i, j) represents the value of (i, j) in the c-th channel, represents the i-th value at height h in the c-th channel, represents the j-th value at width w in the c-th channel;

[0027] (4b) Adjust the first training set to an image of 640×640 and input it into the YOLOv5-CA network model to obtain 1024 1×1 output matrices.

[0028] The YOLOv5-CA network model in step (4) specifically includes:

[0029] The first layer: This layer is the Focus layer of the deep neural network. It converts the 640×640 image into three RGB channels. The Focus layer performs an operation of taking one value every other pixel on the feature map of each channel to obtain 4 independent feature layers. These 4 feature layers are stacked. At this time, the information in the width and height dimensions is converted to the channel dimension, and the input channels are quadrupled. Then, through feature extraction, 64 output matrices with a size of 320×320 are obtained as the input matrices for the second layer;

[0030] The second layer: This layer defines 128 2×2 convolutional kernels. Each convolutional kernel has the function of a filter and is used for learning and feature extraction in the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 128 output matrices with a size of 160×160 as the input matrices for the third layer;

[0031] The third layer: This layer is the Coordinate Attention module CA, and 128 output matrices with a size of 160×160 are obtained as the input matrices for the fourth layer;

[0032] The fourth layer: This layer is a convolutional layer that defines 256 2×2 convolutional kernels. Each convolutional kernel has the function of a filter and is used for learning and feature extraction in the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 256 output matrices with a size of 80×80 as the input matrices for the fifth layer;

[0033] The fifth layer: This layer is the C3 module, which contains 3 standard convolutional layers and multiple Bottleneck modules. The C3 module is the main module for learning residual features. It is divided into two branches. One branch uses multiple Bottleneck modules stacked and 3 standard convolutional layers, and the other branch only passes through 1 basic convolutional layer. Finally, the two branches are concatenated operationally in the channels to obtain 256 output matrices with a size of 80×80 as the input matrices for the sixth layer;

[0034] The sixth layer: This layer is a convolutional layer, defining 512 convolution kernels of 2×2. Each convolution kernel acts as a filter for the learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 512 output matrices of size 40×40 as the input matrix for the seventh layer;

[0035] The seventh layer: This layer is a C3 module. The input matrix is divided into two branches along the channel dimension. One branch passes through multiple Bottleneck stacks and 3 standard convolutional layers, and the other branch only passes through 1 basic convolutional layer. Finally, the two branches are concatenated operationally in the channel dimension to obtain 512 output matrices of size 40×40 as the input matrix for the eighth layer;

[0036] The eighth layer: This layer is a convolutional layer, defining 1024 convolution kernels of 2×2. Each convolution kernel acts as a filter for the learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 1024 output matrices of size 20×20 as the input matrix for the ninth layer;

[0037] The ninth layer: This layer is an SPP module. The input matrix is serially passed through multiple MaxPool layers of size 5×5 to obtain 1024 output matrices of size 20×20 as the input matrix for the tenth layer;

[0038] The tenth layer: This layer is a coordinate attention module CA, obtaining 1024 output matrices of size 1×1.

[0039] Step (5) specifically includes the following steps:

[0040] (5a) In the first stage, the first training set is input into the YOLOv5-CA network model for training to obtain the optimal weights;

[0041] (5b) In the second stage, the optimal weights obtained in the first stage are used as the pre-training weights of the YOLOv5-CA network model, and the second training set is input into the YOLOv5-CA network model for training to obtain the YOLOv5-CA-TL network model;

[0042] (5c) The labeled test set is input into the YOLOv5-CA-TL network model for detection to detect wild animals in the picture and conduct a feasibility verification.

[0043] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, by using the real wild animal dataset of the to-be-detected area collected as the research object, the present invention effectively solves the problem of low detection accuracy of some other algorithms by using the Coordinate Attention module CA, and greatly reduces the labor cost compared with the traditional method; Second, the present invention adopts a two-stage training method to improve the generalization ability of the network and solve the problem of insufficient wild animal samples; Third, compared with algorithms such as Faster R-CNN, SSD, YOLOv3-spp, YOLOX, and PP-YOLOE, the present invention can better improve the overall performance of object detection and effectively improve the object detection accuracy of small-sample wild animals. Description of the Drawings

[0044] Figure 1 is the flowchart of the method of the present invention;

[0045] Figure 2 is the structural diagram of the improved YOLOv5-CA network model;

[0046] Figure 3 is the schematic diagram of the two-stage training method. Detailed Embodiments

[0047] As Figure 1 shown, a small-sample wild animal detection method based on improved YOLOv5, the method includes the following steps in sequence:

[0048] (1) Download the AwA2 animal dataset, screen the AwA2 animal dataset according to the real wild animal dataset of the to-be-detected area collected, and select the images similar to the real wild animal dataset of the to-be-detected area as the experimental dataset, and divide the experimental dataset into the first training set and the validation set; Specifically, download the AwA2 animal dataset through the Internet, screen the AwA2 animal dataset according to the real wild animal dataset of the to-be-detected area, that is, the real wild animal dataset of Qilian Mountain used in the experiment, and select a total of 2398 images of 5 types of animals similar to the real wild animal dataset of Qilian Mountain as the experimental dataset, classify the experimental dataset, and randomly allocate it according to the ratio of eight to two, and divide the data into the first training set and the validation set;

[0049] (2) Screen the real wild animal dataset of the to-be-detected area collected, select the photos with clear animal pictures and high pixel quality as the small-sample experimental dataset, and divide the small-sample experimental dataset into the second training set and the test set; Specifically, screen the photos transmitted back by the Qilian Mountain wildlife protection camera, select 282 photos of 5 types with clear animal pictures and high pixel quality as the small-sample experimental dataset, and randomly allocate them according to the ratio of eight to two, and divide the data into the second training set and the test set;

[0050] (3) Annotate the experimental dataset and the small-sample experimental dataset;

[0051] (4) Construct a YOLOv5 network model, add a coordinate attention module CA to the YOLOv5 network model to obtain a YOLOv5-CA network model, and input the first training set into the YOLOv5-CA network model for training;

[0052] (5) Use a two-stage training method to obtain a YOLOv5-CA-TL network model, input the annotated test set into the YOLOv5-CA-TL network model, detect wild animals in the picture, and conduct a feasibility verification.

[0053] In step (3), the annotation refers to using the Labelimg tool to annotate the animals in the pictures of the experimental dataset and the small-sample experimental dataset, and the annotation format is the yolo format.

[0054] The specific steps of step (4) are as follows:

[0055] (4a) Replace the first C3 module and the last C3 module of the backbone network of the YOLOv5 network model with the coordinate attention module CA. Then, encode the feature maps to form two feature maps, which are sensitive to direction perception and position respectively; Take any intermediate tensor X = [x1, x2, x3..., x C ∈ R C×H×W as the input, and output a tensor Y = [y1, y2, y3..., y c of the same length. Use pooling kernels of size (H, 1) and (1, W) on X to encode each channel along the horizontal and vertical directions. The output of the c-th channel with height h is as follows:

[0056]

[0057] The output of the c-th channel with width w is as follows:

[0058]

[0059] where H is the height of the pooling kernel, W is the width of the pooling kernel, X C is the tensor of the c-th channel, and h is the height of X;

[0060] Formulas (1) and (2) are two transformations of feature aggregation. They aggregate along two spatial directions respectively and return two direction-aware attention maps; The coordinate attention module CA generates two feature layers before concatenation, and then shares a 1×1 convolutional operation transformation, as shown in formula (3):

[0061] f = δ(F1([Zh ,Z w )) (3)

[0062] In the formula, δ is a non - linear activation function, F1 is a 1×1 convolution transformation, f is an intermediate feature map, which is the result of encoding spatial information in both horizontal and vertical directions; then f is decomposed into 2 separate tensors along the spatial dimension, f h ∈R C / r×H and f w ∈R C / r×W , and then two other 1×1 convolution transformations F h and F w are used to transform f h and f w into tensors with the same number of feature layers to X, obtaining:

[0063] g h =σ(F h (f h )) (4)

[0064] g w =σ(F w (f w )) (5)

[0065] In the formula, F h 、F w are both 1×1 convolution transformations, σ is the sigmoid activation function. During the transformation process, the reduction ratio r is used to reduce the number of channels of the f split, and then the outputs g h and g w are expanded and used as weighted weights respectively. The final output of the coordinate attention module CA is shown in formula (6):

[0066]

[0067] In the formula, y c (i,j) represents the value of (i,j) in the c - th channel of the output, x c (i,j) represents the value of (i,j) in the c - th channel, represents the i - th value with height h in the c - th channel, represents the j - th value with width w in the c - th channel;

[0068] (4b) Adjust the first training set to an image of 640×640 and input it into the YOLOv5 - CA network model to obtain 1024 1×1 output matrices.

[0069] As Figure 2 shown, the YOLOv5 - CA network model in step (4) specifically includes:

[0070] The first layer: This layer is the Focus layer of the deep neural network. It converts the 640×640 image into three RGB channels. The Focus layer performs an operation of taking one value every other pixel on the feature map of each channel to obtain 4 independent feature layers. These 4 feature layers are stacked. At this time, the information in the width and height dimensions is converted to the channel dimension, and the input channels are quadrupled. Then, through feature extraction, 64 output matrices with a size of 320×320 are obtained as the input matrix of the second layer;

[0071] The second layer: This layer defines 128 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. And 128 convolutional kernels can help the system extract enough features. This layer uses BN normalization and the SiLU activation function to obtain 128 output matrices with a size of 160×160 as the input matrix of the third layer;

[0072] The third layer: This layer is the Coordinate Attention module CA, obtaining 128 output matrices with a size of 160×160 as the input matrix of the fourth layer;

[0073] The fourth layer: This layer is a convolutional layer, defining 256 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. And 256 convolutional kernels can help the system extract enough features. This layer uses BN normalization and the SiLU activation function to obtain 256 output matrices with a size of 80×80 as the input matrix of the fifth layer;

[0074] The fifth layer: This layer is the C3 module, which contains 3 standard convolutional layers and multiple Bottleneck modules. The C3 module is the main module for learning residual features. It is divided into two branches. One branch uses multiple stacked Bottleneck modules and 3 standard convolutional layers, and the other branch only passes through 1 basic convolutional layer. Finally, the two branches are concatenated operationally in the channels to obtain 256 output matrices with a size of 80×80 as the input matrix of the sixth layer;

[0075] The sixth layer: This layer is a convolutional layer, defining 512 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. And 512 convolutional kernels can help the system extract enough features. This layer uses BN normalization and the SiLU activation function to obtain 512 output matrices with a size of 40×40 as the input matrix of the seventh layer;

[0076] Layer 7: This layer is the C3 module, which divides the input matrix into two branches in the channel dimension. One branch goes through multiple Bottleneck stacks and 3 standard convolutional layers, and the other branch only goes through 1 basic convolutional layer. Finally, the two branches are concatenated in the channel dimension to obtain an output matrix of 512 with a size of 40×40 as the input matrix for the eighth layer;

[0077] Layer 8: This layer is a convolutional layer, defining 1024 2×2 convolutional kernels. Each convolutional kernel has the function of a filter for learning and extracting features in the deep neural network, and 1024 convolutional kernels can help the system extract enough features. This layer uses BN normalization and the SiLU activation function to obtain an output matrix of 1024 with a size of 20×20 as the input matrix for the ninth layer;

[0078] Layer 9: This layer is the SPP module, which serially passes the input matrix through multiple 5×5 MaxPool layers. It should be noted here that serially passing two 5×5 MaxPool layers has the same calculation result as a 9×9 MaxPool layer, and serially passing three 5×5 MaxPool layers has the same calculation result as a 13×13 MaxPool layer, obtaining an output matrix of 1024 with a size of 20×20 as the input matrix for the tenth layer;

[0079] Layer 10: This layer is the Coordinate Attention module CA, obtaining an output matrix of 1024 with a size of 1×1.

[0080] As Figure 3 shown, step (5) specifically includes the following steps:

[0081] (5a) In the first stage, the first training set is input into the YOLOv5-CA network model for training to obtain the optimal weights;

[0082] (5b) In the second stage, the optimal weights obtained in the first stage are used as the pre-training weights of the YOLOv5-CA network model, and the second training set is input into the YOLOv5-CA network model for training to obtain the YOLOv5-CA-TL network model;

[0083] (5c) The labeled test set is input into the YOLOv5-CA-TL network model for detection to detect wild animals in the picture for feasibility verification. Feasibility verification is carried out according to the detection accuracy. If the detection accuracy is high, it means it is feasible, otherwise, it means it is not feasible. For example, if the wild animal in the picture is a tiger and the final detection result is a tiger, it means the detection result is correct. In multiple pictures, the detection accuracy is calculated based on multiple detection results.

[0084] In summary, the present invention takes the real wild animal dataset collected from the area to be detected as the research object, and effectively solves the problem of low detection accuracy of some other algorithms by using the coordinate attention module CA. Compared with the traditional method, the labor cost is greatly reduced; the present invention adopts a two-stage training method to improve the generalization ability of the network and solves the problem of insufficient wild animal samples; compared with algorithms such as Faster R-CNN, SSD, YOLOv3-spp, YOLOX, and PP-YOLOE, the present invention can better improve the overall performance of object detection and effectively improve the object detection accuracy of small-sample wild animals.

Claims

1. A small-sample wildlife detection method based on improved YOLOv5, characterized in that: The method includes the following steps in sequence: (1) Download the AwA2 animal dataset, screen the AwA2 animal dataset according to the collected real wild animal dataset of the area to be detected, and select the images similar to the real wild animal dataset of the area to be detected as the experimental dataset. Divide the experimental dataset into a first training set and a validation set; (2) Screen the collected real wild animal dataset of the area to be detected, and select the photos with clear animal pictures and high pixel quality as the small sample experimental dataset. Divide the small sample experimental dataset into a second training set and a test set; (3) Label the experimental dataset and the small sample experimental dataset; (4) Construct a YOLOv5 network model, add a coordinate attention module CA to the YOLOv5 network model to obtain a YOLOv5-CA network model, and input the labeled first training set into the YOLOv5-CA network model for training; (5) Use a two-stage training method to obtain a YOLOv5-CA-TL network model, input the labeled test set into the YOLOv5-CA-TL network model, detect the wild animals in the picture, and conduct a feasibility verification; The specific steps of step (4) include the following: (4a) Replace the first C3 module and the last C3 module of the backbone network of the YOLOv5 network model with the coordinate attention module CA. Then, encode the feature maps respectively to form two feature maps, which are sensitive to direction perception and position respectively; Take any intermediate tensor X = [x1, x2, x3..., x C ∈ R C×H×W as input and output a tensor Y = [y1, y2, y3..., y c of the same length. Use pooling kernels of size (H, 1) and (1, W) on X to encode each channel along the horizontal and vertical directions. The output of the c-th channel with height h is as follows: The output of the c-th channel with width w is as follows: where H is the height of the pooling kernel, W is the width of the pooling kernel, X C is the tensor of the c-th channel, and h is the height of X; Formulas (1) and (2) are two transformations for feature aggregation. They perform aggregation along two spatial directions respectively and return two direction-aware attention maps; the coordinate attention module CA generates two feature layers before cascading, and then shares a 1×1 convolution operation transformation, as shown in formula (3): f = δ(F1([Z h , Z w )) (3) where δ is a non-linear activation function, F1 is a 1×1 convolutional transformation, f is an intermediate feature map, which is the result of encoding spatial information in both horizontal and vertical directions; then f is decomposed into 2 separate tensors along the spatial dimension, f h ∈R C / r×H and f w ∈R C / r×W , and then two other 1×1 convolutional transformations F h and F w are used to transform f h and f w into tensors with the same number of feature layers to X, obtaining: g h = σ(F h (f h )) (4) g w = σ(F w (f w )) (5) Where F h and F w are both 1×1 convolution transforms, σ is the sigmoid activation function. During the transformation, the reduction ratio r is used to reduce the number of channels of the f branch, and then the outputs g h and g w are expanded and used as weighted weights respectively. The final output of the coordinate attention module CA is shown in formula (6): where y c (i, j) represents the value at (i, j) in the c-th output channel, and x c (i, j) represents the value at (i, j) in the c-th channel, represents the i-th value at height h in the c-th channel, represents the j-th value at width w in the c-th channel; (4b) Adjust the labeled first training set to an image of 640×640, and input it into the YOLOv5-CA network model to obtain 1024 output matrices of 1×1.

2. The small-sample wildlife detection method based on the improved YOLOv5 according to claim 1, wherein: In step (3), the labeling refers to using the Labelimg tool to label the animals in the pictures of the experimental dataset and the small sample experimental dataset, and the labeling format is the yolo format.

3. The small-sample wildlife detection method based on improved YOLOv5 according to claim 1, wherein: The YOLOv5-CA network model in step (4) specifically includes: The first layer: This layer is the Focus layer of the deep neural network. It converts the 640×640 image into three RGB channels. The Focus layer performs an operation of taking one value every other pixel on the feature map of each channel to obtain 4 independent feature layers. Stack these 4 feature layers. At this time, the information in the width and height dimensions is converted to the channel dimension, the input channels are expanded four times, and then feature extraction is performed to obtain 64 output matrices with a size of 320×320 as the input matrix of the second layer; The second layer: This layer defines 128 2×2 convolutional kernels. Each convolutional kernel has the function of a filter and is used for learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 128 output matrices with a size of 160×160 as the input matrix of the third layer; The third layer: This layer is the coordinate attention module CA, and obtains 128 output matrices with a size of 160×160 as the input matrix of the fourth layer; The fourth layer: This layer is a convolutional layer, defining 256 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 256 output matrices of size 80×80 as the input matrix for the fifth layer; The fifth layer: This layer is a C3 module, containing 3 standard convolutional layers and multiple Bottleneck modules. The C3 module is the main module for learning residual features, divided into two branches. One branch uses multiple Bottleneck modules stacked and 3 standard convolutional layers, and the other branch only passes through 1 basic convolutional layer. Finally, the two branches are concatenated operation on the channels to obtain 256 output matrices of size 80×80 as the input matrix for the sixth layer; The sixth layer: This layer is a convolutional layer, defining 512 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 512 output matrices of size 40×40 as the input matrix for the seventh layer; The seventh layer: This layer is a C3 module, dividing the input matrix into two branches in the channel dimension. One branch passes through multiple Bottleneck stacks and 3 standard convolutional layers, and the other branch only passes through 1 basic convolutional layer. Finally, the two branches are concatenated operation on the channels to obtain 512 output matrices of size 40×40 as the input matrix for the eighth layer; The eighth layer: This layer is a convolutional layer, defining 1024 2×2 convolutional kernels. Each convolutional kernel acts as a filter for the learning and feature extraction of the deep neural network. This layer uses BN normalization and the SiLU activation function to obtain 1024 output matrices of size 20×20 as the input matrix for the ninth layer; The ninth layer: This layer is an SPP module, serially passing the input matrix through multiple 5×5 MaxPool layers to obtain 1024 output matrices of size 20×20 as the input matrix for the tenth layer; The tenth layer: This layer is a coordinate attention module CA, obtaining 1024 output matrices of size 1×1.

4. The small-sample wildlife detection method based on improved YOLOv5 according to claim 1, characterized in that: The specific steps of step (5) include the following steps: (5a) In the first stage, the labeled first training set is input into the YOLOv5-CA network model for training to obtain the optimal weights; (5b) In the second stage, the optimal weights obtained in the first stage are used as the pre-training weights of the YOLOv5-CA network model, and the labeled second training set is input into the YOLOv5-CA network model for training to obtain the YOLOv5-CA-TL network model; (5c) The labeled test set is input into the YOLOv5-CA-TL network model for detection to detect wild animals in the picture and conduct feasibility verification.

Citation Information

Patent Citations

  • Marine organism detection method and system and equipment

    CN113963251A

  • Human body security check image detection method and system based on improved YOLOv5s

    CN114862837A