Deep Learning-based Method for Identifying Clouded Leopard Individuals
Through the deep learning method of using residual networks and relationship-aware global attention modules, the problem of time-consuming and laborious and reduced accuracy of Jinqian Leopard individual recognition in the prior art is solved, and efficient and accurate recognition in the wild environment is achieved.
Patent Information
- Application Number
- CN202211333232.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-28
AI Technical Summary
The existing method of individual identification of leopards is time-consuming and laborious and interferes with the ecological environment. The existing deep learning model has decreased the recognition accuracy of large-scale leopard data, making it difficult to meet the needs.
The residual network is used as the backbone network, and the network is trained by combining triple loss and cross entropy loss, and a relationship-aware global attention module (RGA) is added to the convolutional neural network. Through data augmentation and feature extraction, the rapid and accurate identification of the leopard individual is achieved.
It improves the accuracy and applicability of the identification of leopard individuals, and can efficiently identify leopard individuals in wild environments.
Smart Images

Figure CN115880540B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing methods, and in particular, to a method for individual recognition of leopards based on deep learning. Background Art
[0002] The leopard is an endangered wild animal. As the top predator in the ecosystem, it plays a crucial role in the entire ecosystem. Once the leopard disappears, it will affect the entire food chain and ultimately lead to the collapse of the entire ecosystem. The individual recognition of leopards can measure the population situation of leopards, thereby estimating various ecological indicators such as diversity, balance, abundance, relative abundance, and environmental carrying capacity.
[0003] Currently, the methods for individual recognition of leopards mainly include marking, scars, stripes, DNA analysis, and fecal analysis. Although these methods are accurate, they are very time-consuming and laborious for researchers. Moreover, the above methods all require collecting materials in the wild, which has a certain interference on the living environment of leopards. The literature "Research and Implementation of Individual Recognition of Leopards Based on Deep Learning Models" built a CIFAR-10 network model to recognize leopards. However, since the leopard data used is the picture data of the middle part of the leopard's body side, it is difficult to be promoted in practical applications. Although the accuracy rate in the experiment reached 93%, it trained with 17 leopard data. As the number of leopard individuals increases, the CIFAR-10 network structure and softmax loss function it built are difficult to meet the needs of individual recognition tasks. Summary of the Invention
[0004] The present invention provides a method for individual recognition of leopards based on deep learning. A residual network is used as the backbone network for leopard feature extraction. Different features are extracted through two branches, and the network is jointly trained using triplet loss and cross-entropy loss, enabling the network model to automatically and quickly and accurately recognize leopard individuals.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: A method for individual recognition of leopards based on deep learning, which is characterized by including the following steps:
[0006] S1: Use cameras arranged in the wild to collect leopard image data. After preprocessing the data, divide the data into a training set, a test set, and a query set;
[0007] S2: Perform data augmentation on the training set, including horizontal flipping and random erasing;
[0008] S3: Use PyTorch to build a convolutional neural network. Input the training set into the convolutional neural network, and use the cross-entropy loss function and the triplet loss function for iterative training. After iterating multiple times, obtain a weight file;
[0009] S4: Input the query set images into the trained convolutional neural network, which will predict the individuals of leopards.
[0010] A further technical solution lies in that the step S1 includes the following steps: Write the code for video-to-image conversion using python, extract one image from every 50 frames of the video, use the labelImg tool to label each leopard in the image to generate an xml file, crop out the complete leopard image according to the coordinate information in the xml, and name the cropped image. Divide the leopard data according to the ratio of training set:test set being 3:1. At the same time, extract 2 - 5 image data from each leopard data in the test set to form the query set.
[0011] A further technical solution lies in that the images in the data are named with 16 - digit numbers. Among them, the 1st - 4th digits represent which leopard, the 5th - 8th digits represent the leopard data captured by which camera, the 9th - 14th digits represent which frame of the video data the picture is, and the 15th - 16th digits are reserved for later pre - selection.
[0012] A further technical solution lies in that in the step S2: Horizontal flipping means flipping the picture 180 degrees from left to right. The specific method of random erasing is for the input image, randomly select a rectangular area in the image and fill this part of the rectangular area with random values, which is equivalent to occluding this part of the area in the image.
[0013] A further technical solution lies in that the specific structure of the convolutional neural network in the step S3 is as follows: The main part of the network adopts the ResNet50 network structure, and an RGA module is added after each stage. After the fifth stage, the network is divided into two branches. The first branch uses global average pooling to obtain a 2048 - dimensional feature vector, and this feature vector passes through a 1*1 convolutional layer, a batch normalization layer, and a ReLU activation function layer to be further reduced to 512. In the second branch, the same rectangular part of the feature maps extracted in the same batch is set to zero, and global max pooling is used to obtain a 2048 - dimensional feature vector, which is further reduced to 1024 using a 1*1 convolutional layer, a batch normalization layer, and a ReLU activation function layer. The feature vectors of both branches use triplet loss and cross - entropy loss to train the network.
[0014] A further technical solution lies in that the calculation process of the spatial domain RGA (Relation - Aware Global Attention) module is as follows:
[0015] Regarding the vectors at each spatial position of the feature matrix as feature vectors, each feature vector has C dimensions, and there are a total of N = H × W feature vectors. Denote the feature vectors as x, and the i-th feature vector as x i and the j-th feature vector x j The correlation r i,j can be expressed as:
[0016] r i,j = f s (x i , x j ) = θ s (x i ) T φ s (x i ) (1)
[0017] where θ s and φ s represent two embedding functions, including a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer;
[0018] θ s (x i ) = ReLU(W θ x i ) (2)
[0019] θ s (x i ) = ReLU(W θ x i ) (3)
[0020] Similarly, the correlation r j between x i and x j,i can be obtained. Use (r i,j , r j,i ) to represent the bidirectional relationship between the feature vectors x i and x j . Define an affinity matrix to represent the pairwise relationships between all feature vectors; for the i-th feature vector x i , stack the pairwise relationships between it and all other feature vectors in a certain order to obtain an association-aware feature vector r i = [R s (i, :), To learn the attention of the i-th feature vector, in addition to the association-aware feature vector r i , it is also necessary to utilize the feature vector itself to mine the global-scale structural information and local original information related to this feature. Combine x i with r iEmbed and splice them together respectively to obtain the spatially associated perception feature
[0021]
[0022] Where ψ s and represent the embedding function, including a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer, and pool c (·) represents global average pooling to reduce the channel dimension to 1.
[0023] A further technical solution lies in: including rich structural information in the global-range relationship. In order to use a learnable model to mine valuable information therein, the spatial attention value of the feature vector is obtained through the following function:
[0024]
[0025] where W1 and W2 include a 1×1 convolutional layer and a batch normalization layer;
[0026] The calculation method of the channel-domain RGA is similar to that of the spatial-domain RGA. The correlation r i between the i-th feature vector x j and the j-th feature vector x i,j is calculated, and then the bidirectional relationship (r i , r j ) between the feature vectors x i,j and x j,i is calculated. The pairwise relationships of all feature vectors are calculated to form an affinity matrix For the i-th feature vector x i , its pairwise relationships with all other feature vectors are stacked in a certain order to obtain an association perception feature vector r i =[R s (i, :), to represent the global structural information;
[0027] The channel-domain RGA embeds the input feature matrix X into a feature X' of 1×1×C, and splices it with the association perception feature vector r to form an association perception feature for calculating the channel-domain attention a of the feature vector.
[0028] The beneficial effects produced by adopting the above technical solution are as follows: The method of the present invention adds an RGA module to the batch feature erasing network, improves the recognition accuracy and the applicability of individual leopard recognition in different environments, and can effectively achieve the accurate recognition of individual leopards in the wild environment. Brief Description of the Drawings
[0029] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0030] Figure 1 is the overall flowchart of the method described in the embodiments of the present invention;
[0031] Figure 2 is the processing flowchart of the method described in the embodiments of the present invention; Specific embodiments
[0032] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0034] As Figure 1 shown, the embodiments of the present invention disclose a method for identifying individual leopards based on deep learning, including the following steps:
[0035] S1: Use cameras arranged in the wild to collect leopard image data. After preprocessing the data, divide the data into a training set, a test set, and a query set;
[0036] S2: Perform data augmentation on the training set, including horizontal flipping and random erasing;
[0037] S3: Use pytorch to build a convolutional neural network. Input the training set into the convolutional neural network, and use the cross-entropy loss function and the triplet loss function for iterative training. After iterating multiple times, obtain a weight file;
[0038] S4: Input the query set images into the trained convolutional neural network, and the convolutional neural network will predict the individuals of the leopards.
[0039] Further, as Figure 2 shown, the method includes the following steps:
[0040] S01: Use cameras arranged in the wild to collect leopard image data. Write code for converting video to images using python, extract one image every 50 frames from the video, and manually remove the images that do not contain leopard individuals.
[0041] S02: Use the labelImg tool to label each clouded leopard in the image, generate an xml file, crop the clouded leopard images completely according to the coordinate information in the xml, and name the cropped images. The images in the data are named with 16 digits. Among them, the 1st - 4th digits represent which clouded leopard it is, the 5th - 8th digits represent the clouded leopard data captured by which camera, the 9th - 14th digits represent which frame of the video data the picture is, and the 15th - 16th digits are reserved for later pre - selection. For example, it can be named as: 0001_c1s1_000050_01.jpg, where 0001 represents the 1st clouded leopard, c1s1 represents the clouded leopard data captured by the first camera, 000050 represents the 50th frame of the video data, and 01 is reserved for later pre - selection.
[0042] S03: Divide the clouded leopard data according to the ratio of training set: test set of 3:1. At the same time, extract 2 - 5 image data from each clouded leopard data in the test set to form a query set.
[0043] S04: Use pytorch to build a convolutional neural network. The backbone network adopts the ResNet50 network structure. The first stage of the network is a 7*7 convolutional layer with a stride of 2. The second stage of the network is a 3*3 max - pooling with a stride of 2. Then there are 3 bottleneck residual modules. Each residual module is composed of a 1*1 convolution, a 3*3 convolution, and a 1*1 convolution. Among them, the number of channels of the 1*1 convolution is 64, the number of channels of the 3*3 convolution is 64, and the number of channels of the 1*1 convolution is 256. The third stage of the network is 4 bottleneck residual modules. Each residual module is composed of a 1*1 convolution, a 3*3 convolution, and a 1*1 convolution. Among them, the number of channels of the 1*1 convolution is 128, the number of channels of the 3*3 convolution is 128, and the number of channels of the 1*1 convolution is 512. The fourth stage of the network is 6 bottleneck residual modules. Each residual module is composed of a 1*1 convolution, a 3*3 convolution, and a 1*1 convolution. Among them, the number of channels of the 1*1 convolution is 256, the number of channels of the 3*3 convolution is 256, and the number of channels of the 1*1 convolution is 1024. The fifth stage is 3 bottleneck residual modules. Each residual module is composed of a 1*1 convolution, a 3*3 convolution, and a 1*1 convolution. Among them, the number of channels of the 1*1 convolution is 512, the number of channels of the 3*3 convolution is 512, and the number of channels of the 1*1 convolution is 2048. After each stage, add the RGA (Relation - Aware Global Attention) module. The RGA module is divided into the spatial - domain RGA and the channel - domain RGA. The calculation process of the spatial - domain RGA is as follows: regard the vector at each spatial position of the feature matrix as a feature vector. Each feature vector has C dimensions. There are a total of N = H×W feature vectors. Denote the feature vector as x, the i - th feature vector x i and the j - th feature vector xj The correlation r i,j can be expressed as:
[0044] r i,j = f s (x i , x j ) = θ s (x i ) T φ s (x i ) (1)
[0045] where θ s and φ s represent two embedding functions, which are composed of a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer.
[0046] θ s (x i ) = ReLU(W θ x i ) (2)
[0047] θ s (x i ) = ReLU(W θ x i ) (3)
[0048] Similarly, the correlation r j between x i and x j,i can also be obtained. Using (r i,j , r j,i ) to represent the bidirectional relationship between the feature vectors x i and x j , an affinity matrix is defined to represent the pairwise relationships between all feature vectors. For the i-th feature vector x i , the pairwise relationships between it and all other feature vectors are stacked in a certain order to obtain an association-aware feature vector r i = [R s (i, :), To learn the attention of the i-th feature vector, in addition to the association-aware feature vector r i , it is also necessary to use the feature vector itself to mine the global-scale structural information and local raw information related to this feature. Embed and splice x i and r i together to obtain the spatial association-aware feature
[0049]
[0050] where ψ s and represent the embedding function, which consists of a 1×1 convolutional layer, a ReLU activation function layer, and a batch normalization layer. pool c (·) represents global average pooling to reduce the channel dimension to 1.
[0051] Rich structural information is contained in the global-scale relationship. To use a learnable model to mine valuable information therein, the spatial attention value of the feature vector is obtained through the following function:
[0052]
[0053] where W1 and W2 are composed of a 1×1 convolutional layer and a batch normalization layer.
[0054] The RGA calculation method in the channel domain is similar. The correlation r i between the i-th feature vector x j and the j-th feature vector x i,j is calculated, and then the bidirectional relationship (r i , r j ) between x i,j and x j,i is calculated. The pairwise relationships of all feature vectors are calculated to form an affinity matrix For the i-th feature vector x i , its pairwise relationships with all other feature vectors are stacked in a certain order to obtain an association-aware feature vector r i = [R s (i, :), to represent the global structural information.
[0055] The channel-domain RGA embeds the input feature matrix X into a feature X′ of 1×1×C, and splices it with the association-aware feature vector r to form an association-aware feature for calculating the channel-domain attention a of the feature vector.
[0056] After the fourth stage, the network branches into two branches. The first branch uses global average pooling to obtain a 2048-dimensional feature vector, which is further reduced to 512 through a 1*1 convolutional layer, a batch normalization layer, and a ReLU activation function layer. In the second branch, the same rectangular part of the feature maps extracted in the same batch is set to zero, and global max pooling is used to obtain a 2048-dimensional feature vector, which is further reduced to 1024 using a 1*1 convolutional layer, a batch normalization layer, and a ReLU activation function layer.
[0057] S05: Rescale the training data to a size of 256*256, and after horizontal flipping and random erasing in sequence, feed it into the convolutional neural network, and use triplet loss and cross-entropy loss to train the network.
[0058] S06: Input the picture, and the trained network will make predictions on individual leopards.
[0059] The method described in the present invention adds an RGA module to the batch feature erasing network, improves the recognition accuracy and the applicability of individual leopard recognition in different environments, and can effectively achieve accurate recognition of individual leopards in the wild environment.
Claims
1. A method for individual recognition of clouded leopards based on deep learning, characterized in that It includes the following steps: S1: Collect the image data of leopards using cameras deployed in the wild. After preprocessing the data, divide the data into a training set, a test set, and a query set; S2: Perform data augmentation on the training set, including horizontal flipping and random erasing; S3: Use pytorch to build a convolutional neural network. Input the training set into the convolutional neural network, and use the cross-entropy loss function and the triplet loss function for iterative training. After iterating multiple times, obtain the weight file; S4: Input the query set images into the trained convolutional neural network, and the convolutional neural network will predict the individuals of leopards; In the step S3, the specific structure of the convolutional neural network is that the backbone part of the network adopts the ResNet50 network structure, and the RGA, that is, the Relation-Aware Global Attention module, is added after each stage. After the fifth stage, the network is divided into two branches. The first branch uses global average pooling to obtain a 2048-dimensional feature vector, and this feature vector passes through a convolutional layer, a batch normalization layer, and a ReLU activation function layer to be further reduced to 512; in the second branch, the same rectangular part of the feature maps extracted in the same batch is set to zero, and global max pooling is used to obtain a 2048-dimensional feature vector, and a convolutional layer, a batch normalization layer, and a ReLU activation function layer are used to further reduce it to 1024; the feature vectors of both branches use triplet loss and cross-entropy loss to train the network; The calculation process of the spatial domain RGA module is as follows: Regarding the vectors at each spatial position of the feature matrix as feature vectors, each feature vector has C dimensions, and there are a total of feature vectors. The feature vectors are called . The i-th feature vector and the j-th feature vector correlation can be expressed as: (1) Among them, and represent two embedding functions, including a convolutional layer, a ReLU activation function layer, and a batch normalization layer; (2) (3) Similarly, we can obtain the correlation with using to represent the feature vector and define an affinity matrix to represent the pairwise relationships between all feature vectors; for the i-th feature vector stack the pairwise relationships between it and all other feature vectors in a certain order to obtain an association-aware feature vector ; to learn the attention of the i-th feature vector, in addition to the association-aware feature vector it is also necessary to use the feature vector itself to mine the global-scale structural information and local raw information related to this feature, andembed and splice and together to obtain the spatial association-aware feature : (4) wherein and represent embedding functions, including a convolutional layer, a ReLU activation function layer, and a batch normalization layer, represents global average pooling to reduce the channel dimension to 1; Rich structural information is contained in the global-scale relationship. In order to use a learnable model to mine the valuable information therein, obtain the spatial attention value of the feature vector through the following function: (5) Among them and include a convolutional layer and a batch normalization layer; The calculation method of the channel-domain RGA is similar to that of the spatial-domain RGA, and the i-th eigenvector is calculated and the j-th eigenvector for correlation and then the pairwise relationship between the eigenvector is calculated All pairwise relationships of the eigenvectors are calculated to form an affinity matrix For the i-th eigenvector its pairwise relationships with all other eigenvectors are stacked in a certain order to obtain a correlation-aware eigenvector to represent the global structural information; The channel domain RGA embeds the input feature matrix as features , and splices it with the correlation-aware feature vector to form the correlation-aware feature , which is used to calculate the channel domain attention of the feature vector .
2. The individual recognition method of leopard based on deep learning according to claim 1, wherein The step S1 includes the following steps: Write the code for converting video to image using python. Extract one image from every 50 frames of the video. Use the labelImg tool to annotate each leopard in the image to generate an xml file. Crop the leopard images completely according to the coordinate information in the xml, and name the cropped images. Divide the leopard data according to the ratio of training set: test set of 3:
1. At the same time, extract 2-5 image data from each leopard data in the test set to form the query set.
3. The individual recognition method of leopard based on deep learning according to claim 2, characterized in that: The images in the data are named with 16-digit numbers. Among them, the 1st-4th digits represent which leopard, the 5th-8th digits represent the leopard data captured by which camera, the 9th-14th digits represent which frame of the video data the picture is, and the 15th-16th digits are reserved for later pre-selection purposes.
4. The individual recognition method of leopard based on deep learning according to claim 1, characterized in that In the step S2: Horizontal flipping means flipping the picture 180 degrees from left to right. The specific method of random erasing is for the input image, randomly select a rectangular area in the image and fill this part of the rectangular area with random values, which is equivalent to occluding this part of the area in the image.
Citation Information
Patent Citations
Panda wild-training method
CN106719387A
A pedestrian re-identification method based on a hole convolution and attention learning mechanism
CN109784197A