Animal Recognition Method Based on Dual-Path Feature Fusion
The CNN-LBP dual-path model addresses the challenges of texture feature loss and background complexity in animal image recognition by integrating LBP and CNNs, enhancing feature extraction and regression, resulting in improved accuracy and robustness.
Patent Information
- Application Number
- CN202210346833.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-31
AI Technical Summary
The prior art has low image data processing efficiency and difficulty in manual screening and classification in wildlife monitoring, and the texture features of convolutional neural networks are lost during deep learning, which affects the model convergence performance.
The CNN-LBP dual-path feature extraction mode is adopted, combined with convolutional neural network and local binary mode, and through data preprocessing, improved LBP feature extraction, channel attention mechanism and border prediction module, image quality and feature fusion are enhanced and feature expression is optimized.
It improves the accuracy and recall of image recognition, solves the problems of low image quality and complex background, enhances feature extraction capabilities, and achieves more efficient animal recognition.
Smart Images

Figure CN115273131B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and specifically, relates to an animal image recognition and classification method integrating CNN and LBP. Background Art
[0002] With the continuous progress of social technology, computer vision has become a hot research field today. Computer vision can replace the human visual processing system through a computer, and can simulate the human visual system to understand and process the content in the image. As one of the popular fields, object detection and recognition aims to extract the features of the target of interest from the image and identify its category.
[0003] Wildlife resources play a crucial role in maintaining the natural ecological balance. In order to ensure the effective protection of wildlife and collect a large amount of rich wildlife monitoring information, a large number of studies on information-based wildlife monitoring technologies have been carried out at home and abroad. Big data of wildlife is the basis for formulating wildlife protection strategies, and how to use emerging technologies such as artificial intelligence to empower animal information processing is the current research focus and difficulty.
[0004] At present, infrared induction cameras and wireless image sensors have been widely used in many nature reserves to monitor wildlife. Compared with the traditional manual monitoring method, this method greatly improves the monitoring efficiency. However, the amount of monitored image data collected by the above method is relatively large, and currently mainly relies on manual screening and classification, which greatly reduces the data processing efficiency. In the past decade, the huge development and breakthrough of artificial intelligence technologies, such as surveillance video detection, pedestrian detection, ship recognition, face recognition, etc., have brought many conveniences to people's lives and improved the intelligent level of cities. The rapid development of deep learning has also provided a better solution for the rapid and accurate automatic recognition of animal images. Therefore, the research on automatic recognition of wild animals based on convolutional neural network has high practical significance.
[0005] Feature extraction based on Local Binary Pattern (LBP) can better preserve the texture features of images. By introducing convolutional neural networks, a machine learning model under deep supervised learning, the accuracy of target recognition has been greatly improved and has become the mainstream technology in the field of image research. The existing technologies have the following disadvantages: (1) In traditional target detection and recognition methods, artificial feature extraction is carried out, which has a high rate of missed detection, false detection and low efficiency for image targets in practical applications; moreover, for the problem that there is a high similarity between classes and it is difficult for the human eye to distinguish, traditional machine learning methods often cannot effectively extract the subtle differences between similar objects. (2) The target detection and recognition algorithm based on convolutional neural network can automatically extract the target features in the image, and can also achieve a certain degree of translation, rotation, tilt and scale invariance. Multiple convolutional layers can understand and learn multi-level information of the output, and can better extract abstract features as the number of network layers increases. It is an efficient target detection and recognition algorithm, but as the network depth deepens, many low-frequency information such as texture features will be lost, affecting the convergence performance of the model. Summary of the Invention
[0006] The object of the present invention is to design a CNN-LBP dual-channel fusion feature extraction mode by using the features obtained by convolutional neural network in view of the deficiencies of the existing technical solutions. Compared with manual extraction, it can better represent the features of the target and understand and learn the abstract features of the target. And to a certain extent, it solves the problem that the shallow information such as texture features becomes shallower as the convolutional layer deepens, and can capture detail information while obtaining global abstract features; increases the preprocessing of low-quality images and enhances the quality of the monitored images; adds a prediction bounding box regression module to improve the problem that the distance between the real box and the prediction box cannot be predicted when they do not intersect. At the same time, the LBP feature extraction algorithm is optimized to solve the problem of detail loss when the central pixel value is too large or too small, and reduce the influence of noise points. Finally, through feature fusion based on channel attention mechanism, not only can the neural network automatically learn the importance between feature channels, but also it helps to reduce the number of channels and cross-channel information interaction to obtain the final detection result.
[0007] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0008] (1) Data preprocessing: Data augmentation is a commonly used method to reduce overfitting and improve the generalization ability of machine learning models, which can solve the problem of limited data. This method simultaneously adopts image enhancement methods such as image flipping, transformation, splicing and fusion. Each processing obtains a transformation matrix, and then all the transformation matrices are multiplied together to obtain the final transformation matrix. At the same time, aiming at the problem of low quality of the collected images, cubic interpolation and sharpening are used to improve the quality of the images.
[0009] (2) Define the CNN-LBP dual-channel model structure: The network model mainly consists of a backbone convolutional neural network, an LBP feature extraction channel, a feature fusion module, and a bounding box prediction regression network. Analyze that when the convolutional neural network extracts features, it adopts a method of extracting layer by layer. The output size of each module is different. Usually, the output features of the highest layer of the network are selected as the recognition basis. And in terms of practical applications, many factors such as environmental light sources need to be considered. Therefore, the high-level feature semantic information in the convolutional neural network has strong representation ability, but lacks texture detail features; the low-level features are rich in geometric information, but the semantic information representation ability is weak. Define the CNN-LBP dual-channel model, extract image information through the convolutional neural network and the LBP feature extraction channel respectively, and then fuse the features of different network layers through multi-scale feature fusion to improve the feature expression ability of the network.
[0010] (3) LBP feature extraction: Based on the traditional LBP algorithm, considering the influence of the central pixel value and the neighborhood pixel values, compare the sum of squares C of the differences between each neighborhood pixel value and the central pixel value with the limit value W. If C is within the limit range, select the central pixel value as the threshold to calculate the LBP value, fully considering the role of the central pixel value and the neighborhood pixel values to more accurately describe the local image features; if C is not within the limit range, then select the median of the neighborhood pixels and the central pixel as the threshold for calculation to reduce the influence of noise points.
[0011] (4) YOLOV5 model settings: The backbone network is designed as a 152-layer residual network, divided into 5 parts, where each part has grouped convolutional layers. The feature output feature map is extracted through 4 downsampling operations of the backbone network. Introduce the extraction of Region of Interest (ROI) features to reduce the influence of complex backgrounds on target recognition.
[0012] (5) Feature fusion based on the improved channel attention mechanism: Aiming at the disadvantage that the original concat fusion method will increase the number of channels of the image itself during the feature fusion process, for the processed upper and lower branch feature channels, this method uses the proposed channel attention feature fusion module (CAFF) for information integration. It can not only aggregate the feature information from each channel to achieve weight allocation, but also contribute to channel dimensionality reduction and cross-channel information interaction. Then fuse the shallow and deep layer features to avoid the problem of information loss caused by high-resolution features after multiple convolutions.
[0013] (6) Bounding box prediction loss module: The EIOU_Loss loss function is used in the bounding box prediction module, adding a penalty item for the aspect ratio to solve the problem that the gradient cannot adaptively change during the training process in the general loss function.
[0014] (7) Model prediction: After completing the model training, test data of any size can be input to obtain the species and location of the animals. The evaluation metrics used in the experiment are recall rate and mean average precision (mAP) to detect the model. Compared with the commonly used single-stage and two-stage algorithms, it has a higher recall rate and accuracy.
[0015] Advantages of the present invention compared with the prior art:
[0016] The present invention provides an image recognition and classification method integrating CNN and LBP. Compared with the conventional method using convolutional neural network, this method can retain and extract more texture features. After preprocessing the image by normalization, on the first channel, a trainable convolutional kernel is used to extract implicit features; on the second channel, an improved LBP feature operator is used to extract the LBP features of the image and output the feature vector.
[0017] In order to focus on the information more critical to the current task in a large amount of input information, that is, to select the channel most relevant to image recognition, the present invention proposes to improve the above CNN-LBP model by adding a channel attention mechanism. The attention mechanism has been proven to be a way to enhance the deep CNN network, making the final model have a good classification effect. At the same time, a penalty item for aspect ratio is added to the bounding box prediction module, which improves the problem that the distance between the ground truth box and the predicted box cannot be predicted when they do not intersect, and improves the accuracy of bounding box prediction. Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the overall network structure of the method proposed by the present invention.
[0019] Figure 2 It is a flowchart of LBP recognition proposed by the present invention. C is the sum of the squares of the differences between the neighborhood pixel values and the central pixel value; W is the limit value of C. Detailed Embodiments
[0020] The specific process of implementing the present invention is as shown in the appendix Figure 1 and the following is a detailed description of the specific implementation manners.
[0021] (1) Image preprocessing: Using the formula
[0022]
[0023] where src is the left image, dst is the right image after perspective transformation, xy are the horizontal and vertical coordinates, M is the transformation matrix, and M[2,0], M[2,1] are the transformation parameters. Image flipping, conversion, stitching and fusion are performed according to the above formula respectively. In image fusion, the target bounding box and all target bounding boxes in the filled image will be calculated to ensure that the loss value < 0.3.
[0024] Cubic interpolation and sharpening are adopted to improve the image quality: The gray values of 16 points around the point to be sampled are used for cubic interpolation, and a magnification effect closer to the high-resolution image can be obtained. An interpolation basis function is selected to fit the data. The interpolation basis function is as follows:
[0025]
[0026] After the low-quality pixel expansion is completed, there may still be a problem of image blurring. Therefore, image sharpening is used for reprocessing: First, median filtering is used to denoise the image. This method can eliminate isolated noise points and also retain the edges of the image. Then, the denoised image is sharpened using the edge detection method based on the sobel operator to detect and output the edge information of the image. The detected edge image is combined with the original image to complete image sharpening, and the detail information of the output image is significantly enhanced.
[0027] (2) CNN-LBP model setting and training: The network model is mainly composed of a backbone convolutional neural network, an LBP feature extraction path, a feature fusion module, and a bounding box prediction regression network. After the picture is preprocessed, it is input into the backbone convolutional neural network and the LBP feature extraction path to extract features. Then, the output feature map is convolved with a 3×3 convolution to increase the receptive field of the network, and then reduced in dimension through a 1×1 convolution. The backbone network performs 4 downsamplings until it is sampled to 1 / 32 of the original image; based on the result of this downsampling, convolution and upsampling are performed, and the output result is concatenated with the previous downsampling result. Then, after C3 convolution and upsampling, the output result is concatenated with the previous 1 / 8 downsampling result again. The output result is convolved to output features Figure 1 ; at the same time, a downsampling operation is performed, and it is concatenated with the previous output 1 / 16 upsampling result, and features are output through a convolution operation Figure 2 ; similarly, feature map 3 can be obtained. In order to reduce the number of convolution parameters and the computational amount, the above convolutions all use grouped convolutions.
[0028] (3) LBP Feature Extraction: In the traditional LBP algorithm, the central pixel value is directly used as the threshold for calculation. If only the influence of the central pixel is considered, it is greatly affected by noise, and details are likely to be submerged when the central pixel value is too large or too small. Therefore, an improved LBP algorithm is proposed, which takes into account the influence of both the central pixel value and the neighborhood pixel values. This method calculates the sum of squares C of the differences between each pixel value in the neighborhood and the central pixel value, and determines the threshold based on the value of C: If C is within the specified range, the central pixel value is selected as the threshold to calculate the LBP value, which can fully consider the roles of the central pixel value and the neighborhood pixel values, and can more accurately describe the local image features, effectively removing the influence of too large or too small central pixel values on image feature extraction; otherwise, the median of the neighborhood pixels and the central pixel is selected as the threshold for comparison to reduce the influence of noise points. The specific process is as follows: Construct a 3×3 pixel window, and calculate the sum of squares C of the differences between the neighborhood pixel values and the central pixel value in the window. The formula is:
[0029]
[0030] where g p is the neighborhood pixel value, and g i,j is the central pixel value. Let the specified value of C be W, and judge the size of C and W. If C ≤ W, the central pixel value is selected as the threshold to calculate the LBP value, that is, the LBP value is calculated using the above formula; when C > W, the median of the 9 pixel values is selected as the threshold to calculate the LBP value. The formula is:
[0031]
[0032] where p is the neighborhood pixel point, g p is the neighborhood pixel value, g m is the median of the 9 pixel values, and s(x) is a binary function. Thus, the LBP value can be calculated to obtain the LBP feature image. Then, by counting the number of occurrences of the LBP value, the LBP histogram is obtained. After normalization, the statistical histograms of each cell are connected to form a feature vector.
[0033] (4) YOLOV5 Model Settings: First, the image is input into the network for prediction. The network output is a tensor of S×S×(B×5 + C), where S is the number of divided grids, B is the number of network predicted bounding boxes, and C is the number of categories. Then, the entire image is divided into S×S grids, and each grid predicts B detection bounding boxes. The calculation formula for the confidence of each box is as follows:
[0034] Confidence = Pr(Object) * IOU
[0035] Among them, Pr(Object) is the probability that the bounding box contains the target object; in the algorithm, each detection bounding box includes four parameters: x, y, w, and h in addition to the confidence. Among them, x and y represent the specific position of the detection bounding box, and w and h are the width and height of the detection bounding box. Here, the sum-squared error is used as the loss function, which is only used for regression of the recognized object area and does not perform specific class prediction.
[0036] (5) Feature fusion based on the improved channel attention mechanism: The channel attention mechanism is introduced into the feature fusion module to construct a channel attention feature fusion module. As a lightweight module, SENet provides weight parameters for the channel correlation of features, realizing the attention mechanism between channels, but also ignoring the dimensionality reduction ability between networks. This method improves SENet and uses the CAFF module.
[0037] First, feature maps of different scales are unified in size by upsampling or downsampling, and after the concat operation, a feature map is obtained. Then, through the squeeze operation, the two-dimensional features of each channel are compressed into a one-dimensional vector through global pooling. Then, through a fully connected layer, the feature dimension is reduced, and through an activation layer, the non-linear relationship between feature channels is learned. Then, through a fully connected layer, the dimension is increased to the original dimension. Then, through the excitation operation, it is transformed into a normalized weight between 0 and 1 through the sigmoid function. Then, through the rescale operation, the normalized weight is weighted to each feature channel of the input feature map. Finally, a convolution operation is performed, and the obtained feature map is reduced in dimension as the input for subsequent operations. This module can not only enable the neural network to automatically learn the importance between feature channels, but also contribute to channel dimensionality reduction and cross-channel information interaction.
[0038] (6) Bounding box prediction loss module: The minimum bounding rectangle of the predicted box and the ground truth box is introduced, which improves the problem that the distance between the ground truth box and the predicted box cannot be predicted when they do not intersect. GIOU_Loss is used as the loss function for the object bounding box. The smaller the value, the smaller the gap between the predicted box and the ground truth box, and the better the prediction effect can reach the expected effect. The formula is as follows:
[0039]
[0040]
[0041] Among them, A is the target bounding box, B is the predicted bounding box, and C is the smallest enclosing box of A and B.
[0042] As the loss function of the prediction box, GIOU_Loss has improved the problems existing in the case where the real box and the prediction box do not intersect and has a certain measurement ability for the deviation trend. However, when the prediction box is within the real box of the target, i.e., A∩B = B, or the borders of A and B are in a vertical or horizontal position, its loss value will not change. Therefore, it cannot well measure the position state of the prediction box. Thus, EIOU_Loss is introduced, which integrates the overlapping part, the distance between the center points, and the aspect ratio of the two boxes into the calculation of the loss function.
[0043] Since A∩B = B, let the diagonal distance of the minimum bounding rectangle C of A and B be c, and the distance between the center points of the prediction box B and the real box A be r. It can be obtained that the length and width differences between the target box and the prediction box are h and w, and the length and width of the minimum bounding rectangle covering the prediction box and the target box are C h and C w , then the definition of EIOU_Loss is as follows:
[0044] h = |h1 - h2|
[0045] w = |w1 - w2|
[0046]
[0047] EIOU loss = 1 - EIOU
[0048] When the EIOU value is larger, it indicates that the prediction box is closer to the position of the real box, and its loss function value is smaller. However, the gradient of this loss function cannot adaptively change during the training process, thus affecting the training effect. To solve this problem, EIOU_Loss is further optimized, and a new bounding box regression function, namely IEIOU_Loss, is designed, which is defined as:
[0049] EIOU loss = 2×ln2 - 2×ln(1 + EIOU)
[0050] This method can achieve that both IEIOU_Loss and EIOU_Loss decrease as EIOU increases, and the absolute value of the gradient of IEIOU_Loss decreases as EIOU increases. When the distance between the prediction box and the real target box is far, the EIOU value is smaller, that is, it has a larger absolute value of the gradient.
[0051] (5) Model prediction and evaluation: In the experiment of this method, the ANIMALS public dataset is used to estimate the effect of this network, and the evaluation metrics used are recall rate and mean average precision (map) to detect the model.
[0052] The recall rate is defined as:
[0053] The average accuracy is defined as:
[0054] Among them, TP (true positives) is the number of correctly detected objects, FN (false negatives) is the number of undetected objects, and FP (false positives) is the number of false alarms. Among them, P is precision, R is recall, and P(R) is the precision-recall curve. In this experiment, the set detection threshold is 0.5. When the overlapping area between the detection box and the ground truth box exceeds 50%, the detection box is considered correct.
[0055] Table 1 Analysis of the prediction results of the method proposed in the present invention
[0056]
[0057] As shown in Table 1, on the ANIMALS dataset, the recall rate of the method of the present invention is increased by 0.9% compared with yolov5, and the accuracy rate is increased by 4.9%; compared with the LBP recognition, the recall rate is increased, and the accuracy rate is increased by 24.3%, and the accuracy rate is increased by 22.9%, obtaining better results. The experimental results prove that the algorithm of the present invention is effective and can more accurately identify objects in images. It can be seen that the improved attention mechanism module has a more significant impact on object detection on the ANIMALS dataset. The analysis is mainly because the extraction of texture features by the convolutional neural network is not clear enough. After adding the LBP path, the extraction of shallow semantic information is enhanced; EIOU makes the detection results more accurate because when predicting the bounding box regression, the overlapping part, the distance between the center points, and the aspect ratio of the two boxes are considered, and the considered factors are more comprehensive. Moreover, a channel attention mechanism is added to the feature fusion module, which is beneficial to channel dimensionality reduction and cross-channel information interaction, making the final effect better.
[0058] Compared with the traditional artificial feature extraction methods mainly based on LBP, HOG, etc. and the object detection algorithms based on deep convolutional neural networks, it has a higher recognition accuracy. This method combines the strong learning ability of the convolutional neural network and the advantage of LBP in effectively describing image texture features, and at the same time focuses on solving problems such as low image quality, complex background, and low accuracy of prediction boxes, and has achieved good recognition results on the ANIMALS public dataset.
[0059] The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Those skilled in the art should understand that various combinations, modifications or equivalent replacements of the technical solutions of the present invention do not depart from the scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An animal recognition method based on dual-channel feature fusion, characterized in that: The method includes the following steps: Data preprocessing: Image enhancement such as image flipping, transformation, splicing, and fusion is adopted during this process; a transformation matrix is obtained for each processing, and then all the transformation matrices are multiplied together to obtain the final transformation matrix; cubic interpolation and sharpening are used to improve the image quality. Defining the CNN-LBP dual-path model structure: It consists of a backbone convolutional neural network, an LBP feature extraction path, a feature fusion module, and a bounding box prediction regression network; it is analyzed that when the convolutional neural network extracts features, it adopts a way of extracting layer by layer in modules, and the output size of each module is different, so the output features of the highest layer of the network are extracted as the basis for recognition; the CNN-LBP dual-path model is defined, and image information is extracted through the convolutional neural network and the LBP feature extraction path respectively, and then the features of different network layers are fused through a multi-scale feature fusion method to improve the feature expression ability. LBP feature extraction: Considering the influence of the central pixel value and the neighborhood pixel values, by calculating the sum of squares C of the differences between each neighborhood pixel value and the central pixel value and comparing it with the limit value W, if C is within the limit range, the central pixel value is selected as the threshold to calculate the LBP value, fully considering the roles of the central pixel value and the neighborhood pixel values to accurately describe the local image features; if C is not within the limit range, then the median value of the neighborhood pixels and the central pixel is selected as the threshold for calculation to reduce the influence of noise points. YOLOV5 model setting: The backbone network is designed as a residual network with 152 layers in depth, divided into 5 parts, and grouped convolution is used for each part of the convolution. The feature output feature map is extracted through 4 downsampling operations of the backbone network; region of interest (ROI) feature extraction is introduced to reduce the influence of complex backgrounds on target recognition. Feature fusion based on the improved channel attention mechanism: Aiming at the disadvantage that the original concat fusion method will increase the number of channels of the image itself during the feature fusion process, for the processed upper and lower branch feature channels, the proposed channel attention feature fusion module CAFF is used for information integration. Bounding box prediction loss module: The EIOU_Loss loss function is used in the bounding box prediction module, and a penalty item for the aspect ratio is added. Model prediction: After completing the model training, test data of any size is input to obtain the species and location of the animal.