Pet nose print recognition algorithm based on improved ResNet50
By improving the ResNet50 model, a multi-scale attention fusion module and an adaptive feature pyramid are introduced, combined with a joint loss function and a two-stage training strategy, the invasiveness and recognition accuracy of traditional pet identity recognition methods are solved, and efficient and accurate pet nose pattern recognition is achieved.
Patent Information
- Application Number
- CN202510390371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional pet identity recognition methods are highly invasive and have limited recognition accuracy. The ResNet50 model has shortcomings in capturing the details of nose patterns and adapting to pose changes, and the model size and operating efficiency do not adapt to the resource limitations of the end-side equipment.
A multi-scale attention fusion module and an adaptive feature pyramid are introduced, combining joint loss functions and two-stage training strategies, the ResNet50 model is optimized, and resource limitations and real-time requirements are met through model compression technology.
It significantly improves the accuracy and generalization ability of pet nose patterns recognition. It is suitable for all types of pets, with wide applicability and easy integration and expansion.
Smart Images

Figure CN120299064A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pet identity recognition, and particularly relates to a pet nose print recognition algorithm based on improved ResNet50. Background Art
[0002] With the booming development of the pet industry, the demand for pet identity recognition is increasing day by day. Traditional methods such as chip implantation have deficiencies such as strong invasiveness and limited recognition accuracy, and are difficult to meet the needs of diverse scenarios. At the same time, deep learning technology has made remarkable progress in the field of image recognition, providing new ideas for pet nose print recognition. The ResNet50 model has good feature extraction ability, but when faced with complex texture features such as nose prints, there are still problems such as insufficient detail capture and poor adaptability to pose changes. In addition, in practical applications, the size and running efficiency of the model are also key considerations, and it is necessary to make it adapt to the resource limitations of edge devices through model compression technology without significantly reducing performance.
[0003] To solve the above problems, the present invention proposes a pet nose print recognition algorithm based on improved ResNet50. By introducing a multi-scale attention fusion module and an adaptive feature pyramid, it effectively captures the fine-grained features of nose prints and enhances the robustness to different poses and shooting directions. The joint loss function is used to optimize the feature space distribution, and the two-stage training strategy is combined to improve the model performance. At the same time, model compression technology is applied to meet the resource limitations and real-time requirements of edge devices while maintaining high accuracy, and is applicable to the nose print recognition of various types of pets, with wide applicability and the characteristics of being easy to integrate and expand. Summary of the Invention
[0004] In view of this, the present invention proposes a pet nose print recognition algorithm based on improved ResNet50. The present invention can effectively capture the fine-grained features of nose prints, enhance the robustness to different poses and shooting directions, significantly improve the recognition accuracy and generalization ability, and at the same time meet the resource limitations and real-time requirements of edge devices through model compression technology, and is applicable to the nose print recognition of various types of pets, with wide applicability and the characteristics of being easy to integrate and expand.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A pet nose print recognition algorithm based on improved ResNet50 provided by the present invention includes:
[0007] Pre-acquire a large number of nose print pictures of various types of pets, and perform preprocessing to obtain a training picture set and a verification picture set;
[0008] Build the ResNet50 model, utilize the multi-scale attention fusion module therein, and extract the fine-grained features of nose prints in the training dataset by cascading dilated convolution and dual attention mechanism;
[0009] Improve the ResNet50 model, design an adaptive feature pyramid to dynamically fuse shallow and deep features to enhance the robustness to different image shooting directions and postures;
[0010] Adopt a joint loss function to optimize the feature space distribution, and use the training dataset and validation dataset to train and validate the ResNet50 model through a two-stage training strategy until the model meets the preset criteria.
[0011] Preferably, the calculation formula of the multi-scale attention fusion module is:
[0012]
[0013] Among them, F out represents the output feature map, DConv r () represents the convolutional operation with a dilation rate of r, SE is the squeeze-and-excitation module for channel attention modeling, CBAM() is the spatial attention module for spatial attention modeling, and Fin is the input feature map;
[0014] The squeeze-and-excitation module SE performs the following operations:
[0015] Squeeze operation: Perform global average pooling on the input feature map to compress the two-dimensional feature map into a one-dimensional vector to capture global spatial information;
[0016] Excitation operation: Construct the dependence relationship between channels through two fully connected layers, and generate a weight vector with the same number of channels as the number of channels, which is used to recalibrate the importance of channel features;
[0017] The spatial attention module CBAM performs the following steps:
[0018] First, perform channel attention modeling, obtain features in the channel dimension through max pooling and average pooling, and then obtain the channel attention weights through a fully connected layer;
[0019] Then, perform spatial attention modeling, perform convolutional operations on the features processed by channel attention to obtain spatial attention weights, and finally combine the channel and spatial attention weights to achieve the enhancement of the input feature map.
[0020] Preferably, the weight generation method of the adaptive feature pyramid includes the following steps:
[0021] Perform global average pooling on the shallow and deep features respectively to obtain the deep feature vector v deep and the shallow feature vector vshallow ;
[0022] Output dynamic weights through a multi - layer perceptron network:
[0023] [α i , β i = Softmax(MLP([v deep ; v shallow ))
[0024] where α i and β i are the dynamic weights of the deep - layer features and shallow - layer features respectively, used to control the proportion of feature fusion;
[0025] The multi - layer perceptron network MLP() consists of two fully - connected layers. The number of neurons in the first fully - connected layer is twice the dimension of the input feature vector, and the activation function uses ReLU. The number of neurons in the second fully - connected layer is 2, which is used to output two weight values. The Softmax() function normalizes the output values to the range of 0 - 1 to ensure that the sum of the weights is 1.
[0026] Preferably, the expression of the joint loss function is:
[0027] L = 0.8·L ArcFace + 0.2·L Center
[0028]
[0029] where L represents the joint loss function, L ArcFace represents the ArcFace loss, L Center represents the center loss, s represents the scaling factor, represents the angle of the i - th target category y i , m represents the angle margin, θ j represents the angle between the feature vector of the j - th target category y j and the category weight vector, C is the total number of categories, K i is the feature set of the i - th target category, x is the feature vector, and k i is the feature center of the i - th target category.
[0030] Preferably, the two - stage training strategy includes the following steps:
[0031] The first stage: Pre-train the improved ResNet50 model based on the large-scale general image dataset ImageNet. Use the cross-entropy loss function to optimize the model parameters, learn the general image feature representation. The pre-training process uses the Stochastic Gradient Descent (SGD) optimizer with an initial learning rate of 0.1, a momentum of 0.9, a weight decay of 1e-4, a training period of 100 epochs, and the learning rate decays by 0.1 times every 30 epochs.
[0032] The second stage: Fine-tune the pre-trained model on a specific pet nose print dataset. Use the joint loss function to further optimize the model and conduct targeted learning for the special texture and morphological features of pet nose prints. The fine-tuning process uses the Adam optimizer with an initial learning rate of 1e-4, a training period of 50 epochs, and an early stopping mechanism to prevent overfitting. Training will stop early when the validation set loss does not decrease for 5 consecutive epochs.
[0033] Preferably, the model compression technology includes the following steps:
[0034] Replace some traditional convolutional layers with depthwise separable convolutions to reduce the number of model parameters and computational complexity. Depthwise separable convolutions decompose the standard convolution into a depthwise convolution and a pointwise convolution. The depthwise convolution performs convolution operations on each input channel separately, and the pointwise convolution then performs channel fusion on the output of the depthwise convolution through a 1x1 convolution.
[0035] Apply quantization-aware training to quantize the model weights and activation values from 32-bit floating-point numbers to 8-bit integers. By simulating quantization errors during the training process, the model can maintain high accuracy after quantization.
[0036] Utilize neural network pruning technology to prune unimportant channels according to the importance scores of channels, further reducing the model complexity. The pruning ratio is dynamically adjusted according to the characteristics of different layers and datasets to reduce the number of model parameters while ensuring that the model performance degradation is within an acceptable range.
[0037] The present invention has at least achieved the following beneficial effects:
[0038] 1. The present invention can effectively capture the fine-grained features of nose prints, enhance the robustness to different poses and shooting directions, significantly improve the recognition accuracy and generalization ability. At the same time, through the model compression technology, it meets the resource constraints and real-time requirements of edge devices, is applicable to the nose print recognition of various types of pets, and has the characteristics of wide applicability and easy integration and expansion.
[0039] Other advantages, objectives, and features of the present invention will be elaborated in the subsequent description, and to some extent, will be obvious to those skilled in the art, or those skilled in the art can obtain teachings from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer, the following drawings are provided for the description of the present invention:
[0041] Figure 1 It is a step flow chart of a pet nose print recognition algorithm based on improved ResNet50 in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0043] A pet nose print recognition algorithm based on improved ResNet50 provided by the present invention, referring to Figure 1 , includes:
[0044] Pre-acquire a large number of nose print pictures of various types of pets, and perform preprocessing to obtain a training atlas and a validation atlas;
[0045] Construct a ResNet50 model, use the multi-scale attention fusion module therein, and extract fine-grained features of nose prints in the training atlas through cascaded dilated convolution and dual attention mechanism;
[0046] Improve the ResNet50 model, design an adaptive feature pyramid to dynamically fuse shallow and deep features to enhance the robustness of different image shooting directions and postures;
[0047] Adopt a joint loss function to optimize the feature space distribution, and use the training atlas and the validation atlas to train and validate the ResNet50 model through a two-stage training strategy until the model meets the preset standard.
[0048] The working principle and beneficial effects of the above technical solution are as follows: Based on the improved ResNet50 model, the high-precision recognition of pet nose prints is mainly achieved through the following steps: First, a large number of nose print pictures of various types of pets are pre-acquired and preprocessed to obtain a training atlas and a validation atlas. Then, a ResNet50 model is constructed, and the multi-scale attention fusion module in it is used to extract the fine-grained features of the nose prints in the training atlas by cascading dilated convolution and dual attention mechanisms. Next, the ResNet50 model is improved by designing an adaptive feature pyramid to dynamically fuse shallow and deep features to enhance the robustness to different image shooting directions and postures. After that, a joint loss function is used to optimize the feature space distribution, and the ResNet50 model is trained and verified through a two-stage training strategy using the training atlas and the validation atlas until the model meets the preset standards. The working principle of this algorithm mainly includes data preprocessing and atlas construction, feature extraction and multi-scale attention fusion, adaptive feature pyramid and pose robustness enhancement, joint loss function and feature space distribution optimization, and two-stage training strategy and model optimization. Through the present invention, the fine-grained features of nose prints can be effectively captured, the robustness to different postures and shooting directions can be enhanced, the recognition accuracy and generalization ability can be significantly improved, and at the same time, the resource limitations and real-time requirements of end-side devices can be met through model compression technology, which is applicable to the nose print recognition of various types of pets and has the characteristics of wide applicability and easy integration and expansion.
[0049] In a preferred embodiment, the calculation formula of the multi-scale attention fusion module is:
[0050]
[0051] where, F out represents the output feature map, DConv r () represents the convolution operation with a dilation rate of r, SE is the squeeze-and-excitation module for channel attention modeling, CBAM() is the spatial attention module for spatial attention modeling, and F in is the input feature map;
[0052] The squeeze-and-excitation module SE performs the following operations:
[0053] Squeeze operation: Perform global average pooling on the input feature map to compress the two-dimensional feature map into a one-dimensional vector to capture global spatial information;
[0054] Excitation operation: Construct the interdependence between channels through two fully connected layers and generate a weight vector with the same number of channels as the number of channels to re-calibrate the importance of channel features;
[0055] The spatial attention module CBAM performs the following steps:
[0056] First, perform channel attention modeling. Obtain features in the channel dimension through max pooling and average pooling, and then get the channel attention weights through a fully connected layer.
[0057] Next, perform spatial attention modeling. Convolve the features processed by channel attention to obtain spatial attention weights. Finally, combine the channel and spatial attention weights to enhance the input feature map.
[0058] The working principle and beneficial effects of the above technical solution are as follows: The multi-scale attention fusion module realizes the enhancement of the input feature map through cascading dilated convolutions with different dilation rates and the dual attention mechanism. Its working principle is as follows: First, perform dilated convolutions with different dilation rates on the input feature map to obtain multi-scale features. Then, perform attention modeling on the channel dimension through the squeeze-and-excitation module. The squeeze operation compresses the two-dimensional feature map into a one-dimensional vector to capture global spatial information; the excitation operation constructs the dependence relationship between channels through two fully connected layers, generates a weight vector, and recalibrates the importance of channel features. Next, use the spatial attention module CBAM to further enhance the features. First, perform channel attention modeling, obtain channel features through max pooling and average pooling, and then get the channel attention weights through a fully connected layer; then perform spatial attention modeling, convolve the features processed by channel attention to obtain spatial attention weights. Finally, combine the channel and spatial attention weights to achieve a comprehensive enhancement of the input feature map. The multi-scale attention fusion module can effectively capture key information in the feature map, improve the robustness of the model to different scales and complex textures. The squeeze-and-excitation module enables the model to focus on important channel features through channel attention modeling, enhancing the feature expression ability. The spatial attention module CBAM further enhances the model's attention to local details, making the key regions of the feature map more prominent. The combination of this multi-scale and multi-dimensional attention mechanism not only improves the feature extraction ability of the model but also enhances the model's adaptability to complex scenes, thus significantly improving the recognition accuracy and generalization ability of the model.
[0059] In a preferred embodiment, the method for generating the weights of the adaptive feature pyramid includes the following steps:
[0060] Perform global average pooling on the deep and shallow features respectively to obtain the deep feature vector v deep and the shallow feature vector v shallow ;
[0061] Output dynamic weights through a multi-layer perceptron network:
[0062] [α i , β i = Softmax(MLP([v deep ; v shallow ))
[0063] where α i and β i are the dynamic weights of the deep features and shallow features respectively, which are used to control the proportion of feature fusion;
[0064] The multi-layer perceptron network MLP() consists of two fully connected layers. The number of neurons in the first fully connected layer is twice the dimension of the input feature vector, and the activation function uses ReLU. The number of neurons in the second fully connected layer is 2, which is used to output two weight values. The Softmax() function normalizes the output values to the range of 0-1 to ensure that the sum of the weights is 1.
[0065] The working principle and beneficial effects of the above technical solution are as follows: First, global average pooling is performed on the deep and shallow features respectively to obtain the deep feature vector and the shallow feature vector. Global average pooling averages the feature map in the spatial dimension, compresses the two-dimensional feature map into a one-dimensional vector, and captures the global spatial information, providing a basis for subsequent weight generation. Then, these two feature vectors are concatenated and used as the input of the multi-layer perceptron (MLP) network. The MLP network consists of two fully connected layers. The number of neurons in the first fully connected layer is twice the dimension of the input feature vector, and the activation function uses ReLU to increase the non-linear fitting ability of the model; the number of neurons in the second fully connected layer is 2, which is used to output two weight values. Finally, the output values are normalized to the range of 0-1 through the Softmax function to ensure that the sum of the weights is 1, and the dynamic weights of the deep features and shallow features are obtained, which are used to control the proportion of feature fusion. By learning the complex relationship between the deep and shallow features through the MLP network, more reasonable dynamic weights are generated, making the feature fusion process more flexible and accurate. Through the first embodiment of the present invention, one is to enhance the pose robustness of the model, which can automatically adapt to the proportion of feature fusion according to different input images, effectively solving the limitations of the traditional fixed weight fusion method in the face of diverse poses; the second is to improve the flexibility and accuracy of feature fusion. Compared with the traditional manually designed fixed weight fusion method, this data-driven dynamic weight generation method can better adapt to different data distributions and scene changes, improving the feature expression ability and recognition performance of the model; the third is that the normalization process ensures the stability and reasonableness of the weight values, avoiding numerical instability problems caused by too large or too small weights, and improving the stability of the model training and inference processes.
[0066] In a preferred embodiment, the expression of the joint loss function is:
[0067] L = 0.8·L ArcFace + 0.2·L Center
[0068]
[0069] Among them, \(L\) represents the combined loss function, \(L\) ArcFace represents the ArcFace loss, \(L\) Center represents the center loss, \(s\) represents the scaling factor, represents the angle of the \(i\)-th target class \(y\) i , \(m\) represents the angle margin, \(\theta\) j represents the angle between the feature vector of the \(j\)-th target class \(y\) j and the class weight vector, \(C\) is the total number of classes, \(K\) i is the feature set of the \(i\)-th target class, \(x\) is the feature vector, \(k\) i is the feature center of the \(i\)-th target class.
[0070] The working principle and beneficial effects of the above technical solution are as follows: The combined loss function is composed of the ArcFace loss and the center loss. By combining these two loss functions, the feature learning process of the model is optimized, thereby improving the recognition performance of the model. The ArcFace loss enhances the discriminative ability of features by adding an angle margin in the angular space, making samples of the same class closer in the feature space and samples of different classes farther apart. Specifically, when calculating the angle between the feature vector and the class weight vector, the ArcFace loss adds an angle margin to the true class, forcing the model to learn a more discriminative feature representation. The center loss minimizes the dispersion of intra-class features by constraining the distance between the sample features and their class centers, improving the aggregation of intra-class features. Specifically, the center loss calculates the Euclidean distance between each sample feature and its corresponding class center and optimizes to minimize these distances, thus ensuring that samples within the same class are closely aggregated in the feature space. Combining the ArcFace loss and the center loss forms the combined loss function. By reasonably setting the weight coefficients of the two, these two objectives can be optimized simultaneously during the model training process, enhancing both the discriminative ability of features and the compactness of intra-class features, thereby comprehensively improving the recognition performance of the model. This combined loss function first enhances the discriminative ability of features. By introducing an angle margin, the model pays more attention to the discrimination between different classes during the feature learning process, thereby improving the model's recognition ability for samples of different classes. Second, it improves the intra-class compactness. By constraining the sample features to approach the class center, the dispersion of intra-class features is reduced, making samples within the same class more concentrated in the feature space and enhancing the robustness and generalization ability of the model. Third, it improves the model performance, enabling better class separation and feature distribution in the feature space, thereby significantly improving the model's performance in various recognition tasks, especially in challenging scenarios such as complex backgrounds, illumination changes, and pose changes.
[0071] In a preferred embodiment, the two-stage training strategy includes the following steps:
[0072] The first stage: Pre-train the improved ResNet50 model based on the large-scale general image dataset ImageNet, optimize the model parameters using the cross-entropy loss function, learn the general image feature representation. The pre-training process uses the Stochastic Gradient Descent (SGD) optimizer with an initial learning rate of 0.1, a momentum of 0.9, a weight decay of 1e-4, a training period of 100 epochs, and the learning rate decays by 0.1 times every 30 epochs.
[0073] The second stage: Fine-tune the pre-trained model on the specific pet nose print dataset, further optimize the model using the joint loss function, and conduct targeted learning for the special texture and morphological features of pet nose prints. The fine-tuning process uses the Adam optimizer with an initial learning rate of 1e-4, a training period of 50 epochs, and an early stopping mechanism is adopted to prevent overfitting. Training stops early when the validation set loss does not decrease for 5 consecutive epochs.
[0074] The working principle and beneficial effects of the above technical solution are as follows: In the first stage, the improved ResNet50 model is pre-trained based on the large-scale general image dataset ImageNet, enabling the model to learn rich general image feature representations. The cross-entropy loss function is used to optimize the model parameters and learn general image feature representations. The Stochastic Gradient Descent (SGD) optimizer is used in the pre-training process. By using the SGD optimizer and the cross-entropy loss function, the model can effectively adjust the parameters, learn the basic feature patterns in the images. The initial learning rate is 0.1, the momentum is 0.9, and the weight decay is 1e-4, ensuring stable convergence during the training process, avoiding overfitting, and gradually optimizing the weights of the model. The training cycle is 100 epochs, and the learning rate decays by 0.1 times every 30 epochs, enabling the model to quickly learn in the initial stage of training and gradually reduce the learning step size in the later stage, finely adjusting the model parameters, and improving the generalization ability of the model. In the second stage, the pre-trained model is fine-tuned on a specific pet nose print dataset, and the joint loss function is used to further optimize the model. The joint loss function combines the ArcFace loss and the center loss, which can enhance the discriminability of features and the intra-class compactness simultaneously, enabling the model to better adapt to the special texture and morphological features of pet nose prints. Targeted learning is carried out for the special texture and morphological features of pet nose prints. The Adam optimizer is used in the fine-tuning process. This optimizer has the characteristic of adaptive learning rate, which can automatically adjust the learning step size of each parameter according to the magnitude of the gradient, helping the model to converge faster during the fine-tuning process and being able to more effectively optimize the model parameters when dealing with complex features. The initial learning rate is 1e-4, the training cycle is 50 epochs, and the early stopping mechanism is adopted to prevent overfitting. When the validation set loss does not decrease for 5 consecutive epochs, the training is stopped in advance. This mechanism can effectively prevent the model from overfitting, ensure that the model always maintains good generalization ability during the training process, and avoid losing adaptability to new data due to overfitting the training data. In the pre-training stage of the embodiment of the present invention, the model can learn general image feature representations, greatly reducing the time and computing resources required to train the model from scratch, and improving the initial performance of the model at the same time. Secondly, in the fine-tuning stage, optimization is carried out for specific tasks, enabling the model to better adapt to the special features of pet nose prints, and significantly improving the recognition accuracy of the model on this task. Experimental results show that the model using this two-stage training strategy has an accuracy about 12% higher than the model directly trained from scratch on the nose print dataset in the pet nose print recognition task, and shows stronger generalization ability and stability when dealing with actual scenarios such as complex backgrounds and lighting changes. In addition, by reasonably setting the optimizer parameters and the number of training epochs, a good balance can be achieved between the model performance and the training efficiency, meeting the requirements for fast training and efficient deployment in practical applications.
[0075] In a preferred embodiment, the model compression technology includes the following steps:
[0076] Replace some traditional convolutional layers with depthwise separable convolutions to reduce the number of model parameters and computational complexity. The depthwise separable convolution decomposes the standard convolution into a depthwise convolution and a pointwise convolution. The depthwise convolution performs convolution operations on each input channel separately, and the pointwise convolution fuses the channels of the output of the depthwise convolution through 1x1 convolution;
[0077] Apply quantization-aware training to quantize the model weights and activation values from 32-bit floating-point numbers to 8-bit integers. By simulating quantization errors during training, the model can maintain high accuracy after quantization;
[0078] Utilize neural network pruning techniques to prune unimportant channels based on the importance scores of the channels, further reducing the model complexity. The pruning ratio is dynamically adjusted according to the characteristics of different layers and datasets, reducing the number of model parameters while ensuring that the degradation of model performance is within an acceptable range.
[0079] The working principles and beneficial effects of the above technical solutions are as follows: First, replacing some traditional convolutional layers with depthwise separable convolutions decomposes the standard convolution into a depthwise convolution and a pointwise convolution. The depthwise convolution performs convolution operations on each input channel separately, effectively extracting local features of each channel; the pointwise convolution fuses the channels of the output of the depthwise convolution through 1x1 convolution to adjust the number of channels, thus reducing the number of parameters. Second, applying quantization-aware training quantizes the model weights and activation values from 32-bit floating-point numbers to 8-bit integers. By simulating quantization errors during training, the model adapts to the quantized environment and can still maintain high accuracy after quantization. Finally, utilizing neural network pruning techniques to prune unimportant channels based on the importance scores of the channels further reduces the model complexity. The pruning ratio is dynamically adjusted according to the characteristics of different layers and datasets to ensure that while reducing the number of parameters, the degradation of the model's performance is within an acceptable range. The comprehensive application of these technologies streamlines and optimizes the model from three dimensions: model structure optimization, numerical precision reduction, and parameter screening, enabling the model to have lower computational complexity and smaller storage requirements while maintaining high accuracy, significantly improving the deployment feasibility and operating efficiency of the model in practical applications, especially suitable for edge devices such as embedded systems and mobile devices with high resource requirements.
[0080] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. An improved ResNet50-based pet nose print recognition algorithm, characterized in that, Including: Pre-acquire a large number of nose print images of various types of pets in advance, and perform preprocessing to obtain a training atlas and a validation atlas; Construct a ResNet50 model, utilize the multi-scale attention fusion module therein, and extract the fine-grained features of nose prints in the training atlas through cascaded dilated convolution and dual attention mechanisms; Improve the ResNet50 model, design an adaptive feature pyramid to dynamically fuse shallow and deep layer features to enhance the robustness of different image shooting directions and postures; Adopt a joint loss function to optimize the feature space distribution, and use the training atlas and the validation atlas to train and validate the ResNet50 model through a two-stage training strategy until the model meets the preset criteria.
2. The pet nose print recognition algorithm based on the improved ResNet50 according to claim 1, wherein The calculation formula of the multi-scale attention fusion module is: Among them, F out represents the output feature map, DConv r () represents the convolutional operation with a dilation rate of r, SE is the squeeze-and-excitation module for channel attention modeling, CBAM() is the spatial attention module for spatial attention modeling, and F in is the input feature map; The squeeze-and-excitation module SE performs the following operations: Squeeze operation: Perform global average pooling on the input feature map to compress the two-dimensional feature map into a one-dimensional vector to capture global spatial information; Excitation operation: Construct the interdependence between channels through two fully connected layers, and generate a weight vector with the same number of channels as the number of channels, which is used to recalibrate the importance of channel features; The spatial attention module CBAM performs the following steps: First, perform channel attention modeling, obtain features in the channel dimension through max pooling and average pooling, and then obtain the channel attention weights through a fully connected layer; Then perform spatial attention modeling, perform convolution operations on the features processed by channel attention to obtain spatial attention weights, and finally combine the channel and spatial attention weights to achieve enhancement of the input feature map.
3. The pet nose print recognition algorithm based on the improved ResNet50 according to claim 1, characterized in that, The method for generating the weights of the adaptive feature pyramid includes the following steps: Global average pooling is performed on the deep and shallow features respectively to obtain the deep feature vector v deep and the shallow feature vector v shallow ; Output dynamic weights through a multi-layer perceptron network: [α i , β i = Softmax(MLP([v deep ; v shallow )) Among them, α i and β i are the dynamic weights of the deep features and the shallow features respectively, and are used to control the proportion of feature fusion; The multi-layer perceptron network MLP() consists of two fully connected layers. The number of neurons in the first fully connected layer is twice the dimension of the input feature vector, and the activation function uses ReLU. The number of neurons in the second fully connected layer is 2, which is used to output two weight values. The Softmax() function normalizes the output values to the range of 0-1 to ensure that the sum of the weights is 1.
4. The pet nose print recognition algorithm based on the improved ResNet50 according to claim 1, characterized in that, The expression of the joint loss function is: L = 0.8·L ArcFace + 0.2·L Center Among them, L represents the combined loss function, L ArcFace represents the ArcFace loss, L Center represents the center loss, s represents the scaling factor, θ yi represents the angle of the i-th target category y i of, m represents the angle margin, θ j represents the angle between the feature vector of the j-th target category y j and the category weight vector, C is the total number of categories, K i is the feature set of the i-th target category, x is the feature vector, k i is the feature center of the i-th target category.
5. The pet nose print recognition algorithm based on the improved ResNet50 according to claim 1, characterized in that, The two-stage training strategy includes the following steps: The first stage: Pre-train the improved ResNet50 model based on the large-scale general image dataset ImageNet, adopt the cross-entropy loss function to optimize the model parameters, learn the general image feature representation, use the stochastic gradient descent SGD optimizer in the pre-training process, set the initial learning rate to 0.1, the momentum to 0.9, the weight decay to 1e-4, the training period to 100 epochs, and the learning rate decays by 0.1 times every 30 epochs; The second stage: Fine-tune the pre-trained model on a specific pet nose print dataset, adopt the joint loss function to further optimize the model, and perform targeted learning for the special texture and morphological features of pet nose prints. Use the Adam optimizer in the fine-tuning process, set the initial learning rate to 1e-4, the training period to 50 epochs, and adopt an early stopping mechanism to prevent overfitting. Stop training in advance when the validation set loss does not decrease for 5 consecutive epochs.
6. The pet nose print recognition algorithm based on the improved ResNet50 according to claim 1, characterized in that, The model compression technology includes the following steps: Replace some traditional convolutional layers with depthwise separable convolutions to reduce the number of model parameters and computational complexity. The depthwise separable convolution decomposes the standard convolution into a depthwise convolution and a pointwise convolution. The depthwise convolution performs convolution operations on each input channel separately, and the pointwise convolution then fuses the channels of the output of the depthwise convolution through 1x1 convolution; Apply quantization-aware training to quantize the model weights and activation values from 32-bit floating-point numbers to 8-bit integers. By simulating quantization errors during training, the model can maintain high accuracy after quantization; Utilize neural network pruning techniques to prune unimportant channels based on the importance scores of the channels, further reducing the model complexity. The pruning ratio is dynamically adjusted according to the characteristics of different layers and datasets, reducing the number of model parameters while ensuring that the degradation of model performance is within an acceptable range.
Citation Information
Cited By
Face attribute recognition method and device, equipment and medium
CN116912896A
Dynamic space and spiral Mama fused depression image detection method
CN121482048A
A depression image detection method fusing dynamic spatial and spiral mamba
CN121482048B