Fine-grained Image Classification Method and Device for Wild Protected Animal Recognition
Through fine-grained image classification methods, identifying and classifying images of wild protected animals, the problem of insufficient technology application in the field of wild animal protection is solved, and efficient wild animal identification and classification is achieved, saving labor costs and improving work efficiency.
Patent Information
- Application Number
- CN202111386330.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The application of fine-grained image classification technology in the field of wildlife protection is relatively scarce, mainly because wild animals live in deep primitive jungles, difficult experimental conditions, and artificial intelligence technology has less demand in the field of animal protection.
A fine-grained image classification method for wild protected animals is provided, including obtaining images of wild protected animals to be classified, identifying the location of valid information in the image, framing and cropping to remove interference information, repeatedly extracting information that is not concerned by the preset network but is valuable for the classification results, and inputting it into the preset network to obtain the classification results.
It solves the problem of the gap in application of fine-grained image classification in the field of wildlife protection, saves the labor cost of wildlife protection, improves the efficiency of wildlife protection, and provides reliable technical support for scientific research on wildlife protection.
Smart Images

Figure CN114154568B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wind power generation, and particularly relates to a fine-grained image classification method and device for wild protected animal recognition. Background Art
[0002] The state strongly supports the development of artificial intelligence technology, investing a large amount of manpower and material resources in related fields, providing strong support for the breakthrough of artificial intelligence technology. Under this background, Chinese scientific research personnel have been constantly working hard and made major breakthroughs in the fields of deep learning and computer vision. A large number of advanced scientific and technological achievements have emerged, and fine-grained image classification is one of them. However, most of the current research in the field of fine-grained image classification still stays in the theoretical stage. It is an urgent task for us to enable the fine-grained image classification technology to enter the public life. The uses of fine-grained image classification are very extensive. In the field of intelligent driving, the fine-grained image classification technology can identify the vehicles, obstacles, and pedestrians in the front and rear, which is very helpful for the development of autonomous driving technology; in the field of mobile phone R & D, the fine-grained image recognition technology has been applied to the research of face recognition. Even if a person wears a mask, the fine-grained image classification can be used to identify the person's appearance and determine whether the identified person is a legitimate user of the mobile phone; in the field of health care, the fine-grained image technology has been applied to the work of health and epidemic prevention. By taking pictures with a camera, it can identify whether the relevant personnel wear masks, saving labor costs. However, the application of fine-grained image classification in the field of animal protection is very scarce, mainly due to the following reasons: First, most wild animals live deep in the primeval jungle and are far from humans. To apply the fine-grained image classification technology to wild animal protection, researchers need to enter the jungle, and the experimental conditions are harsh. Second, the work of wild animal protection is mainly carried out by relevant government departments. The general public has little contact with the work of wild animal protection, resulting in little demand for artificial intelligence technology in the field of animal protection. Therefore, there is very little relevant research.
[0003] In view of the above problems, it is necessary to propose a fine-grained image classification method and device for wild protected animal recognition. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art, and provides a method and device for recording wind turbine fault data.
[0005] One aspect of the present invention provides a fine-grained image classification method for wild protected animal recognition, the method comprising:
[0006] Obtaining an image of a wild protected animal to be classified, and making a data set according to the classification information of the image;
[0007] Identify the positions of the valid information in the image, frame and crop the image, and remove the interference information in the image;
[0008] Repeatedly extract the information in the image that is not concerned by the preset network but is valuable for the classification result, and input the information into the preset network;
[0009] Start training the parameters for the image, detect the positions of wild animals in the image, crop the image and then input it into the preset network to obtain the classification result.
[0010] Optionally, the starting to train the parameters for the image, detecting the positions of wild animals in the image, cropping the image and then inputting it into the preset network to obtain the classification result includes:
[0011] In each round of training, shuffle the training order of the images and randomly extract the same number of the images to input into the preset network.
[0012] Optionally, after the starting to train the parameters for the image, detecting the positions of wild animals in the image, cropping the image and then inputting it into the preset network to obtain the classification result, the method further includes:
[0013] After each round of training, verify the classification accuracy of the image and save the training parameters.
[0014] Optionally, after the starting to train the parameters for the image, detecting the positions of wild animals in the image, cropping the image and then inputting it into the preset network to obtain the classification result, the method further includes:
[0015] After the training is over, select the model with the highest classification accuracy during training, load it into the network, and test the classification accuracy.
[0016] Optionally, the formula for the classification result is:
[0017] where L is the final classification result, L image is the classification result of the image obtained after being processed by R-CNN, is the classification result of the image obtained after the first masking, is the classification result of the image obtained after the nth masking, and λ and β are hyperparameters.
[0018] Optionally, the repeatedly extracting the information in the image that is not concerned by the preset network but is valuable for the classification result and inputting the information into the preset network includes:
[0019] Obtain the part of the image with the highest score considered by the preset network and mask it;
[0020] to the part of the image that the preset network considers to have the second-highest score, and then mask it, and repeat this process, so that the preset network can extract useful features in the image;
[0021] The formula for the masking is:
[0022]
[0023] where M(i,j) is the mask obtained after masking, θ*A is the threshold, and F(i,j) is the pixel value of the image at (i,j).
[0024] Optionally, the data set is divided into a training set, a validation set, and a test set according to a preset ratio.
[0025] Another aspect of the present invention provides a fine-grained image classification device for wild protected animal recognition. The device includes an acquisition module, a target detection module, a visual attention module, and a training module.
[0026] The acquisition module is used to acquire an image of the wild protected animal;
[0027] The target detection module is used to identify the position of the valid information in the image, frame and crop the image, and remove the interference information in the image;
[0028] The visual attention module is used to repeatedly extract the information in the image that is not concerned by the preset network but is valuable for the classification result, and input the information into the preset network;
[0029] The training module is used to start training parameters for the image, detect the position of the wild animal in the image, and input the cropped image into the preset network to obtain a classification result.
[0030] Optionally, the device further includes a verification module and a test module;
[0031] The verification module is used to verify the classification accuracy of the image and save the training parameters after each round of training;
[0032] The test module is used to load the model with the highest classification accuracy during training into the network and test the classification accuracy after the training is completed.
[0033] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the image classification method described above.
[0034] The fine-grained image classification method and device for wild protected animal recognition of the present invention. The method includes: obtaining an image of a wild protected animal to be classified, and making a data set according to the classification information of the image; identifying the position of the valid information in the image, and framing and cropping the image to remove the interference information of the image; repeatedly extracting the information in the image that is not concerned by the preset network but is valuable for the classification result, and inputting the information into the preset network; starting to train the parameters of the image, detecting the position of the wild animal in the image, and inputting the cropped image into the preset network to obtain the classification result. The fine-grained image classification method for wild protected animal recognition solves the problem of the application blank of image classification in the field of wild animal protection, saves the labor cost of wild animal protection, improves the work efficiency of wild animal protection, and provides reliable technical support for scientific research in the aspect of wild animal protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic flowchart of a fine-grained image classification method for wild protected animal recognition according to an embodiment of the present invention;
[0036] Figure 2 It is a schematic structural diagram of a fine-grained image classification device for wild protected animal recognition according to another embodiment of the present invention;
[0037] Figure 3 It is a schematic diagram of the improved ResNet convolutional neural network structure diagram according to another embodiment of the present invention;
[0038] Figure 4 It is a visual attention module diagram according to another embodiment of the present invention;
[0039] Figure 5 It is a schematic structural diagram of an electronic device according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0041] As Figure 1 shown, one aspect of the present invention provides a fine-grained image classification method S100 for wild protected animal recognition. The method S100 includes:
[0042] S110. Obtain an image of a wild protected animal to be classified, and make a data set according to the classification information of the image.
[0043] Specifically, as Figure 2As shown, in this embodiment, the acquisition module 110 obtains images of wild protected animals under real wild conditions from the wildlife protection department. The acquired image types include wild boar, golden monkey, Chinese alligator, Amur tiger, snow leopard, black muntjac, Asian elephant, black-necked crane, Tibetan antelope, and crested ibis, a total of 10 types, with a total of 12,000 pictures. Correct classification information is labeled for the pictures, and a dataset is made. The pixels of the dataset are uniformly converted to 225×225 and divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0044] S120. Identify the positions of the valid information in the image, frame and crop the image, and remove the interference information in the image.
[0045] Specifically, in this embodiment, after comparing several commonly used neural networks such as AlexNet, VGG, and ResNet residual network, we selected the ResNet residual neural network as the basic network of the present invention. That is to say, improvements were made on the basis of the ResNet convolutional neural network. This network algorithm has high classification accuracy for images, and higher classification accuracy can be obtained by making improvements on its basis.
[0046] As Figure 2 shown, first, a target detection module 120 is connected before the ResNet convolutional neural network. In this embodiment, the target detection module is an R-CNN target detection module. The R-CNN target detection module detects the position of wild animals in the image. Because the images captured by the camera in the wild cannot ensure that wild animals are all in the center of the image, or only a certain part of the animal's body is captured, and the remaining parts of the image are background information and interference noise. If the specific position of the animal is not determined, it is very likely to lead to incorrect final classification results. By connecting the R-CNN target detection module, the wild animal pictures in the dataset can obtain specific positions, be framed and cropped, and the background information and useless information are discarded.
[0047] S130. Repeatedly extract the information in the image that is not concerned by the preset network but is valuable for the classification result, and input the information into the preset network.
[0048] Specifically, after introducing the R-CNN target detection module to obtain the specific position of the animal, since the ResNet convolutional neural network often pays too much attention to high-value classification targets and ignores other parts that can participate in classification, so as Figure 2As shown, a visual attention module 130 is introduced. The visual attention module 130 can enable the information in the image that is not concerned by the network but valuable for the classification result to be re-extracted by the preset network, and input the extracted valuable information into the preset network. It should be noted that in this embodiment, the preset network is a network obtained by adding an R-CNN object detection module 120 and a visual attention module 130 to the ResNet convolutional neural network.
[0049] It should be noted that in this embodiment, the visual attention module 130 adopts CAM attention. An average pooling layer is added to the ResNet convolutional neural network to obtain the part with the highest score considered by the network, and then it is masked. Then, through the attention mechanism, the part with the second highest score is obtained and masked again, and so on, so that the improved ResNet convolutional neural network can fully extract the useful features in the image. Among them, the masking formula is:
[0050]
[0051] Among them, M(i,j) is the mask obtained after masking, θ*A is the threshold, and F(i,j) is the pixel value of the image at (i,j).
[0052] S140. Start training the parameters of the image, detect the position of the wild animal in the image, and after cropping the image, input it into the preset network to obtain a classification result.
[0053] First, during each round of training, the training order of the images needs to be shuffled, and the same number of the images are randomly selected and input into the preset network.
[0054] Specifically, in order to prevent overfitting during the training process, during each round of training, the training order of the images needs to be shuffled, and the same number of images are selected from the training module 140 and sent into the improved ResNet convolutional neural network for training.
[0055] Then, start training the parameters of the image, detect the position of the wild animal in the image, and crop the image. The cropped image is sent into the improved ResNet convolutional neural network with a visual attention module 130 to obtain a classification result. The formula for the classification result is:
[0056] L is the final classification result, L image is the classification result of the image obtained after being processed by R-CNN, is the classification result of the image obtained after the first masking, is the classification result of the image obtained after the nth masking, and λ and β are hyperparameters.
[0057] Exemplarily, after the starting training parameters detect the position of wild animals in the image, crop the image and input it into a preset network to obtain a classification result, the method further includes:
[0058] After each round of training, verify the classification accuracy of the image and save the training parameters.
[0059] Specifically, after each round of training, verify the classification accuracy of the image on the verification module 150 and save the training parameters.
[0060] Exemplarily, after the starting training parameters detect the position of wild animals in the image, crop the image and input it into a preset network to obtain a classification result, the method further includes:
[0061] After training, select the model with the highest classification accuracy during training, load it into the network, and test the classification accuracy.
[0062] Specifically, after training, select the model with the highest classification accuracy during training, load it into the network, and use the test module 160 to test the classification accuracy.
[0063] As Figure 2 shown, another aspect of the present invention provides a fine-grained image classification device 100 for wild protected animal recognition. The device 100 includes an acquisition module 110, a target detection module 120, a visual attention module 130, and a training module 140.
[0064] The acquisition module 110 is used to acquire images of the wild protected animals and make a data set according to the classification information of the images.
[0065] Specifically, as Figure 2 shown, in this embodiment, the acquisition module 110 obtains images of wild protected animals under real wild conditions from the wild animal protection department. The acquired image types are wild boar, golden monkey, Chinese alligator, Amur tiger, snow leopard, black muntjac, Asian elephant, black-necked crane, Tibetan antelope, and crested ibis, a total of 10 types, with a total of 12,000 pictures. Mark the correct classification information for the pictures and make them into a data set.
[0066] The target detection module 120 is used to identify the position of valid information in the image, frame and crop the image, and remove the interference information in the image. In this embodiment, the target detection module 120 is connected in front of the ResNet convolutional neural network, and the target detection module is an R-CNN target detection module.
[0067] The visual attention module 130 is used to repeatedly extract information in the image that is not concerned by the preset network but valuable for the classification result, and input the information into the preset network. In this embodiment, the visual attention module 130 adopts CAM attention, and the visual attention module is also connected to the ResNet convolutional neural network.
[0068] The training module 130 is used to start training parameters for the image, detect the position of wild animals in the image, and after cropping the image, input it into the preset network to obtain a classification result.
[0069] It should be noted that in order to prevent overfitting during the training process, the training order of the images should be shuffled in each round of training, and the same number of images should be extracted from the training module 130 and sent into the improved network for training, that is, the ResNet convolutional neural network with the target detection module 120 and the visual attention module 130 added.
[0070] Exemplarily, the device 100 further includes a verification module 150 and a testing module 160.
[0071] The verification module 150 is used to verify the classification accuracy of the image and save the training parameters after each round of training.
[0072] The testing module 160 is used to select the model with the highest classification accuracy during training and load it into the network after training to test the classification accuracy.
[0073] The fine-grained image classification device for wild protected animal recognition of the present invention introduces a target detection module and a visual attention module into the ResNet convolutional neural network to classify wild animal protection, filling the application blank problem of fine-grained image classification in the field of wild animal protection at the present stage, saving the labor cost of wild animal protection, improving the work efficiency of wild animal protection, and providing reliable technical support for scientific research in the aspect of wild animal protection.
[0074] As Figure 5 shown, another aspect of the present invention provides an electronic device 200, including:
[0075] One or more processors 210, one or more storage units 220, and one or more storage units 220 are used to store one or more programs. When one or more programs are executed by one or more processors 210, one or more processors can implement the data recording method described above. The electronic device 200 further includes one or more input units 230 and one or more output units 240, etc. These components of the electronic device 200 are interconnected through a bus system 250 and / or other forms of connection mechanisms. It should be noted that Figure 3The components and structures of the electronic device 200 shown are merely exemplary and not restrictive. As needed, the electronic device 200 may also have other components and structures.
[0076] The processor 210 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 200 to perform desired functions.
[0077] The storage unit 220 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor may run the program instructions to implement the client functions (implemented by the processor) in the embodiments of the present invention described below and / or other desired functions. Various application programs and various data may also be stored in the computer-readable storage media. For example, various data used and / or generated by the application programs, etc.
[0078] The input unit 230 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch button, and a touch screen, etc.
[0079] The output unit 240 may output various information (such as images or sounds) to the outside (such as a user), and may include one or more of a display, a speaker, etc.
[0080] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the data recording method described above.
[0081] Among them, the computer-readable medium may be included in the device, equipment, and system of the present invention, or may exist alone.
[0082] Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be an electrical, magnetic, optical, electromagnetic, infrared, semiconductor system, device, or equipment. More specific examples include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, an optical fiber, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0083] Among them, the computer-readable storage medium can also include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Specific examples thereof include, but are not limited to, electromagnetic signals, optical signals, or any suitable combination thereof.
[0084] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention. However, the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered within the protection scope of the present invention.
Claims
1. A fine-grained image classification method for wild protected animal recognition, characterized in that, The method includes: Obtaining an image of a wild protected animal to be classified, and making a data set according to the classification information of the image; Identifying the positions of valid information in the image, framing and cropping the image, and removing the interference information of the image; Repeatedly extracting information in the image that is not concerned by the preset network but is valuable for the classification result, and inputting the information into the preset network; wherein, The preset network includes a ResNet convolutional neural network as the basic network, and an object detection module and a visual attention module connected to the ResNet convolutional neural network; Connecting the object detection module to the ResNet convolutional neural network, and the object detection module is used to detect the specific position of the wild protected animal in the image; Introducing the visual attention module into the ResNet convolutional neural network, and the visual attention module is used to repeatedly extract information in the image that is not concerned by the preset network but is valuable for the classification result, and inputting the information into the preset network; specifically including: Obtaining the part of the image with the highest score considered by the preset network, and masking it; Obtaining the part of the image with the second highest score considered by the preset network, and masking it again, and so on in a cycle, so that the preset network can extract useful features in the image; The formula for the masking is: where M(i,j) is the mask obtained after masking, θ*A is the threshold, and F(i,j) is the pixel value of the image at (i,j); Starting to train the parameters of the image, detecting the position of the wild animal in the image, and inputting the cropped image into the preset network to obtain a classification result.
2. The method according to claim 1, characterized in that, The starting to train the parameters of the image, detecting the position of the wild animal in the image, and inputting the cropped image into the preset network to obtain a classification result includes: During each round of training, the training order of the images needs to be shuffled, and the same number of the images are randomly selected and input into the preset network.
3. The method according to claim 2, characterized in that, After the starting to train the parameters of the image, detecting the position of the wild animal in the image, and inputting the cropped image into the preset network to obtain a classification result, the method further includes: After each round of training, verifying the classification accuracy of the image and saving the training parameters.
4. The method according to claim 3, characterized in that, After the starting to train the parameters of the image, detecting the position of the wild animal in the image, and inputting the cropped image into the preset network to obtain a classification result, the method further includes: After the training is over, select the model with the highest classification accuracy during training and load it into the network, and test the classification accuracy.
5. The method according to claim 2, characterized in that, The formula for the classification result is: Among them, L is the final classification result, L image is the classification result of the image obtained after R-CNN processing, is the classification result of the image obtained after the first masking, is the classification result of the image obtained after the nth masking, and λ and β are hyperparameters.
6. The method according to claim 1, characterized in that, The data set is divided into a training set, a validation set, and a test set according to a preset ratio.
7. An image classification device for wild protected animal recognition, characterized in that, For the method according to any one of claims 1 to 6, the device includes an acquisition module, an object detection module, a visual attention module, and a training module. The acquisition module is used to acquire the image of the wild protected animal; The object detection module is used to identify the positions of valid information in the image, frame and crop the image, and remove the interference information of the image; The visual attention module is used to repeatedly extract the information in the image that is not concerned by the preset network but valuable for the classification result, and input the information into the preset network; The training module is used to start training parameters for the image, detect the position of wild animals in the image, and input the cropped image into the preset network to obtain a classification result.
8. The device according to claim 7, characterized in that, The device further includes a verification module and a testing module; The verification module is used to verify the classification accuracy of the image and save the training parameters after each round of training; The testing module is used to select the model with the highest classification accuracy during training and load it into the network to test the classification accuracy after the training ends.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that,When the computer program is executed by a processor, it can implement the image classification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Fine-grained image classification method based on multi-scale repeated attention mechanism
CN111191737A