Fine-grained rodent recognition method based on convolutional neural network
By constructing a fine-grained mouse recognition model based on a three-branch convolutional neural network, and combining the characteristics of the subject perspective, erased viewing angle and cropped viewing angle for recognition, the problems of low recognition efficiency, weak generalization ability of the model and low recognition accuracy in the existing technology are solved, and high-precision fine-grained mouse recognition is achieved.
Patent Information
- Application Number
- CN202211089355.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-09-07
AI Technical Summary
The prior art has problems with low recognition efficiency, weak generalization ability of model and low recognition accuracy in fine-grained mouse recognition. Especially when dealing with different types of mouse, it is difficult to capture distinctive local areas.
A fine-grained mouse recognition method based on three-branch convolutional neural network is adopted. By constructing an image feature extraction module, a subject area selection module, an erased viewing angle image generation module, a cropped viewing angle image generation module, a global average pooling module, a classification network and a recognition result fusion module, the characteristics of the subject viewing angle, an erased viewing angle and a cropped viewing angle are identified, and the final recognition result is obtained.
Effectively filter background noise, focus on important areas in mouse images, enhance the characterization ability of local features, thereby improving the accuracy of fine-grained mouse recognition, reducing background interference, and improving recognition accuracy.
Smart Images

Figure CN116310617B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rodent identification, and more specifically, relates to a fine-grained rodent identification method based on convolutional neural network. Background Art
[0002] In addition to destroying crops, wasting food and damaging daily necessities, rats also carry many pathogenic microorganisms. At the same time, many small organisms living on rats (such as fleas and ticks) are also messengers and media for the spread of diseases. Rats can spread 57 diseases such as plague through saliva (rat bites), contact (flea bites), and excrement. The spread of these rodent-borne diseases seriously threatens the lives and health of the general public. Therefore, rodent control is of great significance to strengthening disease prevention and control capabilities and improving people's health. Rodent control usually requires effective response measures based on different types of rodents and their distribution. However, the grassroots staff engaged in rodent identification are unstable, lack professional experience, and generally have low ability to identify rodents. Potential misinformation will bring many difficulties to decision-making.
[0003] In order to assist grassroots staff to make judgments quickly and accurately, some well-known methods use image processing and machine learning techniques for rodent identification. For example, Tong Sun et al. (<Patent 202110495695.5>, 2021) use mobile detection technology to identify moving objects in the test area. When the minimum circumscribed rectangular area of the moving object is within the preset rodent area, the image of the test area is input into the YOLO v3 (You Only Look Once v3) model to achieve rodent identification. This method can monitor rodents in real time, but it cannot subdivide different categories of rodents. Due to the differences in the activity habits, disease transmission, and extermination schemes of different rodents, rodent control work needs to adopt corresponding prevention and control methods according to different rodents in order to implement precise prevention and control.
[0004] Traditional fine-grained rodent identification relies on manual observation of rodent features or bite marks, claw marks, rodent feces and other traces left by rodents, and comprehensive analysis based on domain knowledge and work experience to determine their categories. However, this method is highly subjective and has low recognition efficiency. At the same time, rodents are erratic and often active at night, making image acquisition difficult. Although most existing rodent image data are specimen images taken in specific laboratory scenes, there is still a lack of strict and unified shooting standards. The lack of high-quality rodent image data means that models trained based on existing machine learning methods can only focus on smaller local areas and are easily affected by environmental factors such as background, lighting, and shooting angles, resulting in weak model generalization capabilities. In addition, rodents of different categories have similar morphological features, and it is necessary to rely on subtle local differences in the rodent body to correctly subdivide different rodents. However, general recognition models are difficult to capture distinguishing local areas, resulting in low model recognition accuracy. In response to the above problems, Qiu Xueya et al. (<Patent 201911007638.7>, 2019) obtained images of rat feces through patrol robots and used image recognition models to determine the type of rats based on the shape of rat feces. This method overcomes the difficulty and inefficiency of manual identification. However, this type of trace-based rat identification method is susceptible to interference from external factors, which may lead to the destruction of related traces, and does not directly use the characteristics of rats from a morphological perspective, and cannot provide effective technical support for research, monitoring, and prevention and control based on the characteristics of rats. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a fine-grained rodent recognition method based on a convolutional neural network. By constructing a fine-grained rodent recognition model based on a three-branch convolutional neural network, the background noise can be effectively filtered out, and the important areas in the rodent image can be focused on, thereby enhancing the characterization ability of local features and improving the accuracy of fine-grained rodent recognition.
[0006] To achieve the above-mentioned object of the invention, the fine-grained rodent identification method based on convolutional neural network of the present invention comprises the following steps:
[0007] S1: Determine N mouse categories according to actual needs, collect a number of image samples for each mouse category and pre-process them using a preset method, classify and label each image sample according to the mouse category, and obtain a mouse image sample set D;
[0008] S2: Construct a fine-grained mouse recognition model based on a three-branch convolutional neural network, including an image feature extraction module, a subject area selection module, an erased view image generation module, a cropped view image generation module, a global average pooling module, a classification network, and a recognition result fusion module, where:
[0009] The image feature extraction module is used to extract the input mouse image Irow , the erased perspective image I sent by the erased perspective image generation module erase , the cropped view image I sent by the cropped view image generation module crop Perform feature extraction respectively to obtain a feature map F of size H×W×C i , where H and W represent the feature map F i The height and width of C represent the feature map F i The number of channels, i∈{row,erase,crop}, record the feature map F i The set of all feature points in is P i ={p i,1 ,p i,2 ,...,p i,H×W}, p i,j ∈R C Represents the feature map F i The feature vector of the j-th feature point, j = 1, 2, ..., H × W, R represents the real number domain; then the feature map F row 、F erase Send it to the main area selection module and convert the feature map F crop Send to the global average pooling module;
[0010] The main region selection module is used to select the feature map F row 、F erase Select the main area and obtain the feature point set P′ in the feature map as the main area row , P′ erase , the mouse image I row The feature point set P′ of the main area row The erased view image I is sent to the mouse erasing view generation module, the mouse cropping view generation module and the global average pooling module. erase The feature point set P′ of the main area erase Sent to the global average pooling module; the main area selection module includes a fully connected layer and a main feature point screening module, where:
[0011] The fully connected layer is used to transform the feature map F i′ Each feature point p i′,j Each feature point p is classified as an independent regional feature and predicted, i′∈{row,erase}. i′,j The rodent confidence vector f i′,j ∈R N , and then output to the subject feature point screening module;
[0012] The main feature point screening module receives all feature points p i′,j Predicted rat confidence vector f i′,jThen, from each mouse confidence vector f i′,j Select the maximum confidence value As the discriminant weight of the feature point representing the local area, all feature points are sorted from large to small according to the discriminant weight, and the set P′ of feature points corresponding to the first K discriminant weights is selected. i′ As the main area, the value of K is determined according to actual needs;
[0013] The mouse erasing perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding erased perspective image I erase The specific method is: randomly select the feature point set P′ row Select M feature points c from m , m=1,2,...,M, the value of M is determined according to actual needs and 1≤M≤K; then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map these M feature points back to the mouse image I row Get M pixels c′ m , with pixel c′ m Centered from the mouse image I row Erase M preset shape areas in the image to obtain the erased perspective image I erase And output to the image feature extraction module;
[0014] The mouse cropping perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding cropped view image I crop The specific method is: randomly select the feature point set P′ row Select H feature points from The value of H is determined according to actual needs and 1≤H≤K, and then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map the coordinates of the H feature points back to the mouse image I. row Get H pixels in In pixels As the center, according to the preset side length from the mouse image I row The H rectangular areas are cut out from the image, and the H rectangular areas are combined into one image according to the preset combination method, and then enlarged to the mouse image I row Size, get the cropped view image I crop And output to the image feature extraction module;
[0015] The global average pooling module is used to pool the feature point set P′ row , P′ erase The regional features corresponding to all feature points in the global average pooling are used to obtain the subject perspective feature f row ∈R C and erase the viewing angle feature f erase ∈R C ; For the cropped view image I crop The feature map F crop Perform global average pooling to obtain the cropping perspective feature f crop ∈R C ; Then each viewing feature f i Output to the classification network;
[0016] The classification network is used to classify the view features f i Classify and obtain the recognition result s under this perspective i ∈R N And output to the recognition result fusion module, s i The nth element in is the mouse image I row Confidence of belonging to the nth mouse category;
[0017] The recognition result fusion module is used to combine the recognition results of the three perspectives i Perform weighted fusion to obtain the final recognition result s pre :
[0018] s pre =λ row ×s row +λ erase ×s erase +λ crop ×s crop
[0019] Among them, λ row , erase , crop Respectively represent the preset weights of the corresponding perspectives, and λ row +λ erase +λ crop =1;
[0020] The identification results pre The mouse category corresponding to the maximum confidence value is taken as the mouse image I row The category to which it belongs;
[0021] S3: using the rodent image sample set D in step S1 to train the fine-grained rodent recognition model to obtain a trained fine-grained rodent recognition model;
[0022] S4: The mouse image to be identified is preprocessed using the same preprocessing method as step S1, and then input into the trained fine-grained mouse identification model to obtain the identification result.
[0023] The present invention is a fine-grained rodent recognition method based on a convolutional neural network. A rodent image sample set is collected, and a fine-grained rodent recognition model based on a three-branch convolutional neural network is constructed. The fine-grained rodent recognition model includes an image feature extraction module, a subject area selection module, an erased perspective image generation module, a cropped perspective image generation module, a global average pooling module, a classification network, and a recognition result fusion module. The features of the extracted subject perspective, erased perspective, and cropped perspective are recognized and fused to obtain a final recognition result. The rodent image sample set is used to train the fine-grained rodent recognition model, and the trained fine-grained rodent recognition model is used to recognize the rodent image to be recognized.
[0024] The present invention has the following beneficial effects:
[0025] 1) The present invention adopts a fine-grained rodent recognition model based on a three-branch convolutional neural network, which captures the features in the main area of the rodent image that are helpful for fine-grained recognition by screening the feature points in the main area, thereby effectively reducing background interference and improving recognition accuracy;
[0026] 2) The fine-grained rodent recognition model based on the three-branch convolutional neural network of the present invention erases and crops the main area of the rodent image, so that the model can learn more discriminative local features from different perspectives, and integrate the recognition results of the three perspectives in the recognition process to further improve the recognition accuracy;
[0027] 3) The present invention also proposes a loss function based on the classification loss function of each branch and the central loss function. Through the joint constraints of the cross entropy loss function of each branch and the central loss function, the model can identify mice according to the fine-grained features of the local area in the image, thereby improving the problem of difficulty in fine-grained recognition due to the similar features of mice of different categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flowchart of a specific implementation of the fine-grained rodent identification method based on a convolutional neural network of the present invention;
[0029] Figure 2 is a structural diagram of a fine-grained rodent recognition model based on a three-branch convolutional neural network in this embodiment;
[0030] Figure 3 It is a heat map of visualization analysis of some mouse image samples in this embodiment. DETAILED DESCRIPTION
[0031] The specific implementation of the present invention is described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0032] Example
[0033] Figure 1 FIG. 1 is a flow chart of a specific implementation of the fine-grained rodent identification method based on a convolutional neural network of the present invention. Figure 1 As shown, the specific steps of the fine-grained rodent identification method based on convolutional neural network of the present invention include:
[0034] S101: Obtaining rodent image samples:
[0035] Determine N mouse categories according to actual needs, collect several image samples for each mouse category and use a preset method to preprocess them, classify and label each image sample according to the mouse category, and obtain a mouse image sample set D.
[0036] In practical applications, the preprocessing method can be set according to the actual situation. The specific method of preprocessing the mouse image sample in this embodiment includes image enhancement and image normalization:
[0037] 1) Image enhancement:
[0038] In order to enable the convolutional neural network to learn more mouse feature information and improve the generalization ability of the network, in this embodiment, an image enhancement method is first used to process each acquired image sample. The image enhancement method includes image scaling, center cropping, horizontal mirroring, vertical flipping, brightness adjustment, contrast adjustment, and saturation adjustment.
[0039] 2) Image normalization:
[0040] For the image samples after image enhancement, each pixel value is divided by 255 to ensure that the pixel value range is between 0 and 1, and then normalized according to the following formula:
[0041]
[0042] Among them, x and x′ represent the pixel value vectors composed of all channel pixel values of the pixels before and after normalization, μ represents the mean vector of the pixel value vectors of all pixels in the image sample, and σ represents the variance vector of the pixel value vectors of all pixels in the image sample.
[0043] In practical applications, it is also necessary to perform necessary format conversion on the image samples according to the input format requirements of the subsequent fine-grained rodent recognition model. For example, the first module of the fine-grained rodent recognition model in this embodiment is the ResNet-50 convolutional neural network. Therefore, the input format of the image sample H0×W0×C0 needs to be converted to C0×H0×W0, where H0 and W0 represent the height and width of the mouse image, respectively, and C0 represents the number of channels of the mouse image. When the mouse image is an RGB image, C0=3.
[0044] S202: Constructing a fine-grained mouse recognition model based on a three-branch convolutional neural network:
[0045] In order to integrate multiple features and improve the accuracy of fine-grained rodent identification, the present invention constructs a fine-grained rodent identification model based on a three-branch convolutional neural network, wherein the first branch is used for rodent subject perspective recognition, which effectively reduces background interference by capturing the subject area in the rodent image; the second branch is used for rodent erasing perspective recognition, which enables the model to learn more distinguishing local features by randomly erasing some areas in the subject area; the third branch is used for rodent cropping perspective recognition, which enhances the model's ability to characterize local detail features by cropping and amplifying some features in the subject area. The three branches share the parameters of the backbone image feature extraction network, the subject area selection module, and the classification network.
[0046] Figure 2 is a structural diagram of the fine-grained rodent recognition model based on a three-branch convolutional neural network in this embodiment. Figure 2 As shown, the fine-grained rodent recognition model based on the three-branch convolutional neural network in this embodiment includes an image feature extraction module, a main area selection module, an erased perspective image generation module, a cropped perspective image generation module, a global average pooling module, a classification network and a recognition result fusion module. The following describes in detail each component module of the fine-grained rodent recognition model.
[0047] The image feature extraction module is used to extract the input mouse image I row , the erased perspective image I sent by the erased perspective image generation module erase , the cropped view image I sent by the cropped view image generation module crop Perform feature extraction respectively to obtain a feature map F of size H×W×C i , where H and W represent the feature map F i The height and width of C represent the feature map F i The number of channels, i∈{row,erase,crop}, record the feature map F i The set of all feature points in is P i ={p i,1 ,p i,2,...,p i,H×W}, p i,j ∈R C Represents the feature map F i The feature vector of the j-th feature point, j = 1, 2, ..., H × W, R represents the real number field, and then the feature map F row 、F erase Send it to the main area selection module and convert the feature map F crop Sent to the global average pooling module.
[0048] In this embodiment, the ResNet-50 convolutional neural network is selected as the image feature extraction network of the fine-grained mouse recognition model, that is, the three branches in the fine-grained mouse recognition model share the same ResNet-50 to extract the features of mouse images. ResNet-50 uses residual block technology to enhance gradient fluidity to alleviate the problem of network degradation and ensure that the network can extract discriminative fine-grained features from mouse images. Table 1 is the network structure information table of ResNet-50.
[0049]
[0050] Table 1
[0051] As shown in Table 1, ResNet-50 consists of one stage of convolutional layers and four stages of residual blocks. Stage 1 consists of a 7×7 convolutional layer, and stages 2, 3, 4, and 5 consist of 3, 4, 6, and 3 residual blocks, respectively. Each residual block consists of three convolutional layers of 1×1, 3×3, and 1×1 connected in series. In stages 3, 4, and 5, the first residual block downsamples the mouse image feature map with a step size of 2 to obtain the final feature map.
[0052] The main region selection module is used to select the feature map F row 、F erase Select the main area and obtain the feature point set P′ in the feature map as the main area row , P′ erase , the mouse image I row The feature point set P′ of the main area row The erased view image I is sent to the mouse erasing view generation module, the mouse cropping view generation module and the global average pooling module. erase The feature point set P′ of the main area erase Send to the global average pooling module. Figure 2 As shown, the subject area selection module in the present invention includes a fully connected layer and a subject feature point screening module, wherein:
[0053] The fully connected layer is used to transform the feature map Fi′ Each feature point p i′,j Each feature point p is classified as an independent regional feature and predicted, i′∈{row,erase}. i′,j The rodent confidence vector f i′,j ∈R N , and then output to the subject feature point screening module.
[0054] The main feature point screening module receives all feature points p i′,j Predicted rat confidence vector f i′,j Then, from each mouse confidence vector f i′,j Select the maximum confidence value As the discriminant weight of the feature point representing the local area, all feature points are sorted from large to small according to the discriminant weight, and the set P′ of feature points corresponding to the first K discriminant weights is selected. i′ As the main area, the value of K is determined according to actual needs, and the remaining feature points are regarded as background candidate areas that are useless for rodent identification.
[0055] The mouse erasing perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding erased perspective image I erase The specific method is: randomly select the feature point set P′ row Select M feature points c from m , m=1,2,...,M, the value of M is determined according to actual needs and 1≤M≤K. Then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map these M feature points back to the mouse image I row Get M pixels c′ m , with pixel c′ m Centered from the mouse image I row Erase M preset shape areas in the image to obtain the erased perspective image I erase And output to the image feature extraction module.
[0056] Mouse Image I row and feature map F row The mapping relationship of the pixels in the image can be determined according to the parameters of the image feature extraction module. As shown in Table 1, the network structure information of ResNet-50 shows that the image feature extraction module in this embodiment extracts the original mouse image I row Downsampling was performed 5 times with a factor of 2. 5 =32, so the mapping relationship can be expressed by the following formula:
[0057] x′ m =(x m +0.5)×32
[0058] y′ m =(y m +0.5)×32
[0059] Among them, (x m ,y m ) represents the feature map F row The coordinates of the feature points in (x′ m ,y′ m ) represents the feature point (x m ,y m ) is mapped back to the mouse image I row The parameter 0.5 is used to make the mapped pixel coordinates located at the center of the area corresponding to the feature point.
[0060] The shape of the erased area can be circular, rectangular, etc., and the size can be set according to the actual situation. It can be seen that the erasing perspective is to erase some key areas based on the main perspective, so that the fine-grained rodent recognition model can learn more features that are helpful for fine-grained recognition in the remaining areas.
[0061] The mouse cropping perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding cropped view image I crop The specific method is: randomly select the feature point set P′ row Select H feature points from The value of H is determined according to actual needs and 1≤H≤K, and then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map the coordinates of the H feature points back to the mouse image I. row Get H pixels in In pixels As the center, according to the preset side length from the mouse image I row The H rectangular areas are cut out from the image, and the H rectangular areas are combined into one image according to the preset combination method, and then enlarged to the mouse image I row Size, get the cropped view image I crop And output to the image feature extraction module.
[0062] It can be seen that the cropping perspective is to cut out some important areas from the main perspective, and then re-input them into the recognition model after enlarging them, so that the fine-grained mouse recognition model can learn more details in these areas.
[0063] The global average pooling module is used to pool the feature point set P′ row , P′ erase The regional features corresponding to all feature points in the global average pooling are used to obtain the subject perspective feature f row ∈R C and erase the viewing angle feature f erase ∈R C ; For the cropped view image I crop The feature map F crop Perform global average pooling to obtain the cropping perspective feature f crop ∈R C ; Then each viewing feature f i Output to the classification network.
[0064] The classification network is used to classify the view features f i Classify and obtain the recognition result s under this perspective i ∈R N And output to the recognition result fusion module, s i The nth element in is the mouse image I row The confidence that the n-th mouse category is included. The classification network in this embodiment is implemented using a conventional fully connected layer + softmax layer.
[0065] The recognition result fusion module is used to combine the recognition results of the three perspectives i Perform weighted fusion to obtain the final recognition result s pre :
[0066] s pre =λ row ×s row +λ erase ×s erase +λ crop ×s crop
[0067] Among them, λ row , erase , crop Respectively represent the preset weights of the corresponding perspectives, and λ row +λ erase +λ crop =1.
[0068] The identification results pre The mouse category corresponding to the maximum confidence value is taken as the mouse image I row The category to which it belongs.
[0069] According to the above description, the fine-grained rodent recognition model of the present invention includes three branches:
[0070] 1) Original mouse image I rowAfter the feature map is extracted by the image feature extraction module, the feature points of the subject area are determined by the subject area selection module, and then the feature vectors of the feature points of the subject area are globally averaged and pooled, and the subject view recognition result is obtained by the classification network;
[0071] 2) The mouse erasure perspective recognition module is based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding erased perspective image I erase , erase the view image I erase After the feature map is extracted by the image feature extraction module, the subject area selection module determines the erasing view image I again. erase The feature points of the main area are then globally averaged and pooled, and the classification network is used to obtain the erased view recognition result.
[0072] 3) The mouse cropping perspective recognition module is based on the mouse image I row The feature point set P′ of the main area row Generate mouse image I row The corresponding cropped view image I crop , cropped view image I crop After the feature map is extracted by the image feature extraction module, the feature map is globally averaged pooled, and then the cropping perspective recognition result is obtained by the classification network.
[0073] Finally, the three perspective recognition results obtained by the three branches are fused to obtain the final recognition result.
[0074] S103: Training a fine-grained rodent recognition model.
[0075] The rodent image sample set D in step S1 is used to train the fine-grained rodent recognition model to obtain a trained fine-grained rodent recognition model.
[0076] In the training of the fine-grained rodent recognition model, the setting of the loss function is very important. Since the fine-grained rodent recognition model of the present invention includes three branches, and the three branches share the backbone image feature extraction module, the global average pooling module and the classification network, it is necessary to learn the three branches together so that the constructed fine-grained rodent recognition model can reduce background interference, capture the discriminative local areas in the rodent image, and improve the recognition accuracy. Therefore, the following method is used in this embodiment to calculate the loss function L of the current image sample: total :
[0077] For each branch network, the classification loss function L of each view recognition result is calculated separately using the cross entropy loss function i :
[0078] L i =-s real logs i
[0079] Among them, s real Represents the one-hot encoding of the labeled category according to the image sample.
[0080] In view of the similarity of the characteristics of different categories of mice, the central loss function L center Constrain the distance between each mouse image sample feature and the central feature of its category, so that the mouse features of the same category are gathered near the central feature of its category, enhancing the distinguishability between the features of mice of different categories, thereby improving the accuracy of fine-grained mouse recognition. Center loss function L center The calculation formula is as follows:
[0081]
[0082]
[0083] Among them, || ||2 means to find the second norm, c n Represents the feature center of the nth category, which is iteratively updated as follows:
[0084]
[0085] Among them, c′ n It represents the category feature center before updating, and its initial value is 0 vector. α represents the preset learning rate, which is used to control the update rate of the feature center.
[0086] The joint loss function L used in this example total It is the weighted sum of the classification loss function and the center loss function of the three branches. The classification loss function of the three branches describes the gap between the labeled category and the model recognition result. The smaller the classification loss function value of the three branches, the closer the labeled category is to the model recognition result. The center loss function describes the distance between the mouse feature and the feature center of its category. The smaller the center loss function value, the closer the distance between the mouse feature and the feature center of its category. Through the joint constraints of the classification loss function of the three branches and the center loss function, the fine-grained mouse recognition model can recognize mice based on the fine-grained features of the local area in the mouse image. Joint loss function L total The calculation formula is as follows:
[0087] L total =γ row ×L row +γ erase ×L erase +γcrop ×L crop +γ center ×L center
[0088] Among them, γ row , γ erase , γ crop , γ center They respectively represent the preset weights of the corresponding loss functions.
[0089] The parameters of the fine-grained rodent recognition model are iteratively updated. In this embodiment, the Stochastic Gradient Descent with Momentum algorithm is used to iteratively update the parameters of the model. The algorithm adds the model parameter update amount of the previous step to the update calculation of the current model parameters, so that the update of the model parameters no longer depends only on the parameter gradient of the current iteration, thereby accelerating the convergence speed of the iterative update.
[0090] S104: Fine-grained rodent identification:
[0091] The mouse image to be identified is preprocessed using the same preprocessing method as in step S101, and then input into the trained fine-grained mouse identification model to obtain an identification result.
[0092] In order to better illustrate the technical solution of the present invention, a specific example is used to experimentally verify the present invention. In this embodiment, an identification experiment is conducted on common rodents in a certain province, and the mouse categories and mouse image samples are provided by the local endemic disease prevention and control institute. In this embodiment, the mouse categories include yellow-breasted rat (0), brown rat (1), big-footed rat (2), Smith's house mouse (3), house mouse (4), Carter's mouse (5), large woolly mouse (6), Qi's field mouse (7), plate-toothed rat (8), stinky shrew (9), gray musk shrew (10), tree shrew (11), community rat (12), needle-haired rat (13) and Chinese hairy hedgehog (14), a total of 15 categories, and the numbers in brackets are the indexes of the corresponding categories. In this embodiment, there are a total of 2003 mouse image samples. After the mouse image samples are pre-processed, they are divided into a training set and a test set in a ratio of 8:2.
[0093] The preprocessing process of the mouse image sample in this embodiment is: scale the mouse image sample to 512×512, then crop it from the center to 448×448, mirror it horizontally with a probability of 50%, flip it vertically with a probability of 50%, randomly fluctuate the brightness by 30% on the original basis, randomly fluctuate the contrast by 30% on the original basis, and randomly fluctuate the saturation by 20% on the original basis, to obtain the mouse image sample after data enhancement.
[0094] In addition, according to the input requirements of the ResNet-50 convolutional neural network in this embodiment, the mouse image sample is format converted to 3×448×448.
[0095] Then, each mouse image sample in the training set is input into the constructed fine-grained mouse recognition model based on the three-branch convolutional neural network to train the fine-grained mouse recognition model. The processing of the three branches is as follows:
[0096] 1) Original mouse image I row After the image feature extraction module extracts a feature map of size 2048×14×14, the subject area selection module selects 49 feature points with larger discriminant weights as feature points of the subject area, and then performs global average pooling on the feature vectors of these 49 subject area feature points to obtain the mouse subject perspective feature f row ∈R 2048 , and then the subject perspective recognition result s is obtained by the classification network row ∈R 15 .
[0097] 2) The mouse erasure perspective recognition module is based on the mouse image I row The feature point set P′ of the main area row Select 7 feature points and map them back to the mouse image I row Table 2 is a feature point mapping table in the process of generating the erased view image in this embodiment.
[0098] serial number Discriminant weight Feature map coordinates Original image coordinates 58 0.7428 (4,1) (144,48) 87 0.6537 (6,2) (208,80) 102 0.5897 (7,3) (240,112) 76 0.6145 (5,5) (176,176) 162 0.8567 (11,7) (368,240) 97 0.5623 (6,12) (208,400) 36 0.8714 (2,7) (80,240)
[0099] Table 2
[0100] Then, seven rectangular areas of size 64×64 are erased with the seven mapped pixels as the center to generate the mouse image I row The corresponding erased perspective image I erase . Erase view image I erase After the feature map is extracted by the image feature extraction module, the subject area selection module determines the erasing view image I again. erase The feature points of the main area are then globally averaged and pooled to obtain the mouse erasure perspective feature f erase ∈R 2048 , and then the classification network obtains the erased view recognition result s erase ∈R 15 ;
[0101] 3) The mouse cropping perspective recognition module is based on the mouse image I row The feature point set P′ of the main area row Select 4 feature points and map them back to the mouse image I rowTable 3 is a feature point mapping table in the process of generating the cropped view image in this embodiment.
[0102] serial number Discriminant weight Feature map coordinates Original image coordinates 92 0.7714 (6,7) (208,240) 118 0.7076 (8,5) (272,176) 147 0.8427 (10,6) (336,208) 80 0.6231 (5,9) (176,304)
[0103] Table 3
[0104] And take the mapped pixel point as the center from the mouse image I row Cut out four 112×112 rectangular areas from the top and place them at the upper left corner, lower left corner, upper right corner and lower right corner of the bottom plate respectively. Then enlarge the combined image to the mouse image I row The size is 3×448×448, and the cropped view image I is obtained crop . Cropped view image I crop After the feature map is extracted by the image feature extraction module, the feature map is globally averaged and pooled to obtain the mouse erasure perspective feature f crop ∈R 2048 , and then the classification network obtains the cropping perspective recognition result s crop ∈R 15 .
[0105] Finally, let the weight λ row =λ erase =λ crop =1 / 3, the three perspective recognition results obtained by the three branches are merged to obtain the final recognition result s pre ∈R 15 .
[0106] In this embodiment, during the training of the fine-grained rodent recognition model, the weight for calculating the joint loss function is set to γ row =γ erase =γ crop =1,γ center =0.001, feature center learning rate α =0.01, learning rate η =0.005 in the stochastic gradient descent algorithm with momentum, momentum factor ρ =0.9, iteratively update the model parameters to complete the model training.
[0107] Finally, the mouse image samples in the test set are input into the trained fine-grained mouse recognition model to verify the recognition accuracy. Taking a mouse image sample as an example, the recognition results obtained by the three branches are:
[0108] s row =[0,0,0,0,0.001,0.008,0,0,0.001,0.08,0.053,0.856,0.001,0,0.001]
[0109] s erase=[0.001,0.001,0.003,0,0,0.033,0,0.034,0,0.124,0.083,0.702,0.006,0,0.012]
[0110] s crop =[0,0,0,0,0.002,0.013,0,0.003,0.001,0.016,0.011,0.951,0.003,0,0.001]
[0111] The final recognition result obtained by fusion is:
[0112] s pre =[0,0,0.001,0,0.001,0.018,0,0.012,0,0.073,0.049,0.836,0.003,0,0.004]
[0113] The index 11 (tree shrew) corresponding to the maximum confidence value of 0.826 is selected as the recognition result of the mouse image to be identified, which is the same as the real label of the mouse image sample and the recognition is accurate.
[0114] In order to illustrate the technical effect of the present invention, two evaluation indicators, accuracy (ACC) and macro F1 score (Macro-F1), are used to verify the impact of different modules of the present invention on performance. Table 4 is a comparison table of the impact of different modules of the present invention on performance in this embodiment.
[0115] Model ACC Macro-F1 ResNet-50 0.781 0.717 ResNet-50+ main perspective branch 0.816 0.747 ResNet-50+ three branches 0.831 0.782 ResNet-50+three branches+center loss 0.852 0.816
[0116] Table 4
[0117] As shown in Table 4, different modules of the present invention improve the accuracy of fine-grained mouse recognition. After adding the three-branch module and the center loss function, the accuracy is improved by 7 percentage points, and the Macro-F1 score is improved by nearly 10 percentage points.
[0118] In order to more intuitively demonstrate the advantages of the present invention in reducing background interference, a visualized thermal analysis is performed on the feature extraction of mouse image samples by the present invention and the original ResNet-50 model. Figure 3 : is a heat map of visualization analysis of some mouse image samples in this embodiment. Figure 3 As shown in the figure, due to the lack of high-quality mouse image data, the original ResNet-50 model can only focus on a small local area and is easily affected by the background area. The method proposed in the present invention can effectively reduce background interference and focus on important areas in mouse images (such as head, tail, limbs, etc.), thereby improving the accuracy of fine-grained mouse recognition.
[0119] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. A fine-grained rodent identification method based on convolutional neural network, characterized in that: The following steps are involved: S1: Determine N mouse categories according to actual needs, collect a number of image samples for each mouse category and pre-process them using a preset method, classify and label each image sample according to the mouse category, and obtain a mouse image sample set D; S2: Construct a fine-grained mouse recognition model based on a three-branch convolutional neural network, including an image feature extraction module, a subject area selection module, an erased view image generation module, a cropped view image generation module, a global average pooling module, a classification network, and a recognition result fusion module, where: The image feature extraction module is used to extract the input mouse image I row , the erased perspective image I sent by the erased perspective image generation module erase , the cropped view image I sent by the cropped view image generation module crop Perform feature extraction respectively to obtain a feature map F of size H×W×C i , where H and W represent the feature map F i The height and width of C represent the feature map F i The number of channels, i∈{row,erase,crop}, record the feature map F i The set of all feature points in is P i ={p i,1 ,p i,2 ,...,p i,H×W }, p i,j ∈R C Represents the feature map F i The feature vector of the j-th feature point, j = 1, 2, ..., H × W, R represents the real number domain; then the feature map F row 、F erase Send it to the main area selection module and convert the feature map F crop Send to the global average pooling module; The main region selection module is used to select the feature map F row 、F erase Select the main area and obtain the feature point set P as the main area in the feature map r ' ow , P e ' rase , the mouse image I row The feature point set P of the main area r ' ow The erased view image I is sent to the mouse erasing view generation module, the mouse cropping view generation module and the global average pooling module. erase The feature point set P of the main area e ' rase Sent to the global average pooling module; the main area selection module includes a fully connected layer and a main feature point screening module, where: The fully connected layer is used to transform the feature map F i′ Each feature point p i′,j Each feature point p is classified as an independent regional feature and predicted, i′∈{row,erase}. i′,j The rodent confidence vector f i′,j ∈R N , and then output to the subject feature point screening module; The main feature point screening module receives all feature points p i′,j Predicted rat confidence vector f i′,j Then, from each mouse confidence vector f i′,j Select the maximum confidence value As the discriminant weight of the feature point representing the local area, all feature points are sorted from large to small according to the discriminant weight, and the set P of feature points corresponding to the first K discriminant weights is selected. i ″As the main area, the value of K is determined according to actual needs; The mouse erasing perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P of the main area r ' ow Generate mouse image I row The corresponding erased perspective image I erase The specific method is: randomly select the feature point set P r ' ow Select M feature points c from m , m=1,2,...,M, the value of M is determined according to actual needs and 1≤M≤K; then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map these M feature points back to the mouse image I row Get M pixels c′ m , with pixel c′ m Centered from the mouse image I row Erase M preset shape areas in the image to obtain the erased perspective image I erase And output to the image feature extraction module; The mouse cropping perspective generation module is used to generate a mouse image based on the mouse image I row The feature point set P of the main area r ' ow Generate mouse image I row The corresponding cropped view image I crop The specific method is: randomly select the feature point set P r ' ow Select H feature points from h=1,2,...,H, the value of H is determined according to actual needs and 1≤H≤K, then according to the mouse image I row and feature map F row The mapping relationship of the pixels in the image is used to map the coordinates of the H feature points back to the mouse image I. row Get H pixels in In pixels As the center, according to the preset side length from the mouse image I row The H rectangular areas are cut out from the image, and the H rectangular areas are combined into one image according to the preset combination method, and then enlarged to the mouse image I row Size, get the cropped view image I erase And output to the image feature extraction module; The global average pooling module is used to pool the feature point set P r ' ow , P e ' rase The regional features corresponding to all feature points in the global average pooling are used to obtain the subject perspective feature f row ∈R C and erase the viewing angle feature f erase ∈R C ; For the cropped view image I crop The feature map F crop Perform global average pooling to obtain the cropping perspective feature f crop ∈R C ; Then each viewing feature f i Output to the classification network; The classification network is used to classify the view features f i Classify and obtain the recognition result s under this perspective i ∈R N And output to the recognition result fusion module, s i The nth element in is the mouse image I row Confidence of belonging to the nth mouse category; The recognition result fusion module is used to combine the recognition results of the three perspectives i Perform weighted fusion to obtain the final recognition result s pre : s pre =λ row ×s row +λ erase ×s erase +λ crop ×s crop Among them, λ row , erase , crop Respectively represent the preset weights of the corresponding perspectives, and λ row +λ erase +λ crop =1; The identification results pre The mouse category corresponding to the maximum confidence value is taken as the mouse image I row The category to which it belongs; S3: Use the rodent image sample set D in step S1 to train the fine-grained rodent recognition model to obtain a trained fine-grained rodent recognition model. During the training, the loss function L of each image sample is calculated using the following method: total : For each branch network, the classification loss function L of each view recognition result is calculated separately using the cross entropy loss function i : L i =-s real logs i Among them, s real Represents the one-hot encoding of the labeled category according to the image sample; Then the center loss function L is calculated using the following formula center : Among them, c n Represents the feature center of the nth category, which is iteratively updated as follows: Among them, c′ n represents the category feature center before updating, its initial value is 0 vector, and α represents the preset learning rate; The joint loss function L is calculated using the following formula: total : L total =c row ×L row +g erase ×L erase +g crop ×L crop +g center ×L center Among them, γ row , γ erase , γ crop , γ center Respectively represent the preset weights of the corresponding loss functions; S4: The mouse image to be identified is preprocessed using the same preprocessing method as step S1, and then input into the trained fine-grained mouse identification model to obtain the identification result.
2. The fine-grained rodent identification method according to claim 1, characterized in that: The preprocessing method for the image sample in step S1 is: First, the image samples are enhanced, including image scaling, center cropping, horizontal mirroring, vertical flipping, brightness adjustment, contrast adjustment, and saturation adjustment. Then, for the enhanced image samples, each pixel value is divided by 255, and then normalized according to the following formula: Among them, x and x′ represent the pixel value vectors composed of all channel pixel values of the pixels before and after normalization, μ represents the mean vector of the pixel value vectors of all pixels in the image sample, and σ represents the variance vector of the pixel value vectors of all pixels in the image sample.
Citation Information
Patent Citations
Mouse type identification method and system
CN110705522A
Computer room rodent image recognition method, system and storage medium
CN113312981B
Image recognition model training method and device and electronic equipment
CN111368788A
Biological category identification method and device, storage medium and electronic equipment
CN111950344A