Cow identity recognition method based on fine-grained information compensation loss

By using the feature extraction model of ResNet50 and ASCAM attention mechanism in cattle identity recognition, combining the generation of pseudo-label data sets and deep metric learning, the spatial distribution of feature is optimized, and the dependence problem of the model on calibrated data in the existing technology is solved, and unsupervised cattle identity recognition with high accuracy is achieved.

CN119942597APending Publication Date: 2025-05-06INNER MONGOLIA UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510093191.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing cattle identity recognition methods rely on calibration data, resulting in model training's dependence on data calibration, low accuracy, and difficult to effectively identify in unlabeled data scenarios.

Method used

A feature extraction model based on ResNet50 and ASCAM attention mechanism is adopted, and the generation and refinement of pseudo-label data sets are combined with deep metric learning, the spatial distribution of feature is optimized, and the fine-grained information compensation loss is calculated to realize the identification of the cattle identity under unsupervised learning.

Benefits of technology

It improves the model's learning ability on the label-free data set, reduces the impact of pseudo-label noise, enhances the accuracy of identity recognition, and solves the problem of model training dependence on calibrated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942597A_ABST
    Figure CN119942597A_ABST
Patent Text Reader

Abstract

The invention discloses a cattle identity recognition method based on fine-grained information compensation loss, and relates to the field of biological image recognition, and the method comprises the steps: building a ResNet-ASC feature extraction model based on ResNet 50 and an ASCAM attention mechanism; the method comprises the following steps: collecting a plurality of unmarked cattle face data sets, extracting global features by using a ResNet-ASC feature extraction model, and clustering the global features to generate a pseudo tag data set; inputting the pseudo tag data set into a ResNet-ASC feature extraction model for training, updating model parameters through back propagation and gradient descent optimization algorithms until the model converges, and obtaining a trained cattle face feature extraction model; and performing identity recognition on a to-be-detected cattle face image by using the trained cattle face feature extraction model to obtain a cattle identity recognition result. According to the method, the problem that model training depends on calibration data can be solved, and the accuracy of identity recognition under completely unsupervised learning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological image recognition, and more particularly to a cattle identification method based on fine-grained information compensation loss. Background Art

[0002] my country's animal husbandry production is transforming and upgrading towards informatization and intelligence. Precision animal husbandry refers to the management of livestock on an individual basis. Cattle identification is the prerequisite for precision animal husbandry cattle breeding and production, providing guarantees for cattle disease prevention and control, genetic improvement of breeds and individual refined breeding. It also provides technical support for meat product traceability and improving agricultural fake insurance claims.

[0003] Traditional cattle identification mainly relies on contact identification methods, which require a lot of manpower, have poor animal welfare, and low accuracy and reliability. With the continuous development of computer vision theory and technology, non-contact identification methods based on biological measurement features have attracted widespread attention. Cow mouth and nose pattern images, iris images, and retinal vascular distribution pattern images have all been used in cattle identification research. However, the above-mentioned biological features have problems such as difficulty in information collection and easy to cause stress reactions in cattle. Cow face images, as biological features of cattle appearance, have good uniqueness and are easy to collect. Therefore, recognition based on cattle facial features is currently the main research method.

[0004] At present, the recognition method based on cow face images mainly relies on supervised learning and a large amount of individual cow identification information. However, the identification of cow face images is labor-intensive and requires the assistance of professional breeders, which makes it difficult to produce large-scale labeled cow face datasets. However, in actual animal husbandry production scenarios, we are faced with larger-scale, unlabeled cow individual identification tasks in different environments.

[0005] Therefore, how to solve the problem of model training's dependence on data calibration and improve the accuracy of cattle identification under completely unsupervised learning is an urgent problem that technical personnel in this field need to solve. Summary of the invention

[0006] In view of this, the present invention provides a cattle identification method based on fine-grained information compensation loss. On the one hand, the feature extraction capability of the ResNet50 network is improved by adding the ASCAM attention mechanism; on the other hand, the pseudo-labels are refined by using local fine-grained information to reduce the influence of pseudo-label noise; on the other hand, the feature space distribution is optimized based on deep metric learning. The combined effect of the three helps to solve the problem of model training's dependence on calibration data, thereby improving the accuracy of identity identification under completely unsupervised learning.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A cattle identification method based on fine-grained information compensation loss includes the following steps:

[0009] S1: Based on ResNet50 and ASCAM attention mechanism, a ResNet-ASC feature extraction model is constructed;

[0010] S2: Collect several unlabeled cow face datasets, use the ResNet-ASC feature extraction model to extract global features, and generate a pseudo-label dataset by clustering the global features;

[0011] S3: Input the pseudo-label dataset into the ResNet-ASC feature extraction model for model training, and update the model parameters through back propagation and gradient descent optimization algorithm until the model converges to obtain a trained cow face feature extraction model;

[0012] S4: Use the trained cow face feature extraction model to perform identity recognition on the cow face image to be detected, and obtain the cow identity recognition result.

[0013] Optionally, in S1, the ResNet-ASC feature extraction model is constructed as follows:

[0014] Select ResNet50 as the base network, modify the parameters of the first Bottleneck in Layer-4, and set the step size of the second convolutional layer and the downsampling layer to 1;

[0015] For the global feature map output by Layer-4, add the ASCAM attention mechanism to obtain the global attention feature map, and then obtain the global feature through average pooling;

[0016] The global feature map output by Layer-4 is horizontally divided into three equal parts to obtain three local feature maps. The ASCAM attention mechanism is added to each local feature map to obtain a local attention feature map, and three local features are obtained through maximum pooling.

[0017] Optionally, in S1, the specific operation of the ASCAM attention mechanism is:

[0018] Calculate the activity level of each element in its channel using the following formula:

[0019]

[0020] Where: Represents the mean value of elements in a single channel, Represents the variance of elements within a single channel, x i represents the elements in the channel, N represents the sum of the number of elements in the current channel, and λ represents the regularization coefficient;

[0021] Calculate the activity of each element in the entire channel. The formula is as follows:

[0022]

[0023] Where: represents the mean value of all elements in the channel, represents the variance of all elements in the channel, x i Represents the elements in the channel, C represents the sum of the number of channels in the feature map, and λ represents the regularization coefficient;

[0024] The input feature map is adjusted by the calculated element activity value. The formula is as follows:

[0025]

[0026] Where: Represents the activity value calculated based on the mean value of elements on a single channel. Represents the activity value calculated based on the mean of all elements in the channel, x i Represents the elements in the channel, Represents ASCAM attention weighted features.

[0027] Optionally, in S2, the generation of the pseudo-label dataset specifically includes the following steps:

[0028] Collect several color digital images with cow faces centered and crop them to a fixed size to create an unlabeled cow face dataset ULCFDataSet;

[0029] Scale each image in the ULCFDataSet and normalize each channel to obtain the preprocessed RGB image.

[0030] The preprocessed RGB image is input into the ResNet-ASC model to extract the global feature f i g , and normalize the global features;

[0031] Obtain similar samples through K nearest neighbor search and K mutual nearest neighbor search, calculate the Jaccard distance of global features, and construct the Jaccard distance matrix;

[0032] Perform DBSCAN clustering on the Jaccard distance matrix, assign pseudo labels to the unlabeled data based on the clustering results, and retain the samples with valid clustering to generate a pseudo-label dataset.

[0033] Optionally, in S3, the model training of the ResNet-ASC feature extraction model specifically includes the following steps:

[0034] Input the pseudo-labeled dataset into the ResNet-ASC feature extraction model to extract global features and local features, and generate global feature sets and local feature sets;

[0035] Based on local features, dynamic label smoothing and deep metric learning are used to calculate the correlation detail reconstruction loss;

[0036] Based on global features, adaptive label refinement and deep metric learning are used to calculate the intra-class and extra-class distance constraint loss;

[0037] The relevance detail reconstruction loss and the in-class and out-class distance constraint loss are summed to calculate the fine-grained information compensation loss;

[0038] The loss is compensated by fine-grained information, and the model parameters are updated through back propagation and gradient descent optimization algorithms. The pseudo-label generation stage and the model training stage are performed alternately until the model converges to obtain a trained cow face feature extraction model.

[0039] Optionally, extracting global features and local features specifically includes the following steps:

[0040] Using the intra-class balanced sampling method, m categories are randomly sampled from the pseudo-label data set, and n samples are randomly selected from each category to form a batch sample of size m*n;

[0041] Perform data augmentation operations such as resizing, random flipping, image padding, random cropping, random erasing, and standardization on each sample image to obtain a training data set;

[0042] The training data set is input into the ResNet-ASC model, and the global feature map processed by the ASCAM attention mechanism is averaged and pooled to obtain the global feature f i g ;

[0043] The local feature map processed by the ASCAM attention mechanism is max-pooled to obtain the local features of each region.

[0044] Optionally, calculate the correlation detail reconstruction loss, which specifically includes the following steps:

[0045] Input the local features into the classifier composed of the fully connected layer and the softmax function, obtain the local prediction results, and calculate the cross entropy loss L between the local prediction results and the pseudo labels. l-ce , the formula is as follows:

[0046]

[0047] Where: y i represents the pseudo label of sample i, represents the local prediction result of the nth part of sample i, M is the total number of samples;

[0048] Create a uniform vector and calculate the KL divergence loss between the local prediction result and the uniform vector

[0049] Based on the feature matching coefficient, dynamically adjust the cross entropy loss L l-ce and KL divergence loss The weight of the dynamic label smoothing loss L is calculated dls , the formula is as follows:

[0050]

[0051] Where: Represents the feature matching coefficient of each local feature, represents the cross entropy loss calculated according to formula (4), u represents a uniform vector, represents KL divergence loss;

[0052] Based on local features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct local feature triplets;

[0053] Based on the local feature triples, the local feature ternary loss function is used to calculate the local ternary loss L l-tr , the formula is as follows:

[0054]

[0055] Where: M is the total number of samples, represents the local features of sample i, represents the local difficult positive sample of sample i, represents the local difficult negative sample of sample i, ||·|| represents the Euclidean norm;

[0056] Based on the local feature triples, the local feature density loss function is used to calculate the local density loss L l-densify , the formula is as follows:

[0057]

[0058] Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, represents the local features of sample i, represents the local difficult positive sample of sample i, Represents the local feature distance between sample i and its local difficult positive sample, Represents the average distance of local features in the category where sample i belongs;

[0059] Then, the correlation detail reconstruction loss L rdr The formula is as follows:

[0060] L rdr =L dls +L l-tr +L l-densify (8).

[0061] Optionally, calculate the intra-class and extra-class distance constraint loss, which specifically includes the following steps:

[0062] The feature matching coefficient is used to summarize the local prediction results and refine the pseudo labels. The formula is as follows:

[0063]

[0064] Where: represents the refined pseudo-label, y i is the pseudo label of sample i, represents the value of the feature matching coefficient calculated by the softmax function, β represents the refinement strength coefficient, Represents the local prediction result of the nth part of sample i;

[0065] Based on the refined pseudo-labels, the adaptive label refinement loss L is calculated using the cross entropy loss function alr , the formula is as follows:

[0066]

[0067] Where: represents the refined pseudo-label, Represents the global prediction result obtained by inputting the global features into the classifier composed of full connection and softmax function;

[0068] Based on the global features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct a global feature triplet;

[0069] Based on the global feature triples, the global feature triple loss function is used to calculate the global triple loss L g-tr , the formula is as follows:

[0070]

[0071] Where: M is the total number of samples, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the difficult negative sample of sample i, ||·|| represents the Euclidean norm;

[0072] Based on the global feature triples, the global feature dense loss function is used to calculate the global dense loss L g-densify , the formula is as follows:

[0073]

[0074] Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the global feature distance between sample i and its difficult positive sample, Represents the average distance of global features in the category where sample i belongs;

[0075] Then, the distance constraint loss L cdc The expression is:

[0076] L cdc =L alr +L g-tr +L g-densify (13).

[0077] Optionally, obtaining the feature matching coefficient includes the following steps:

[0078] For the global features of the sample image, a K-nearest neighbor search is performed on the global feature set to obtain a set of similar samples of the global features;

[0079] For each local feature of the sample image, perform K nearest neighbor search on the corresponding local feature set to obtain a similar sample set of each local feature;

[0080] According to the sample set obtained by the search, the feature matching coefficient of each local feature of the sample image relative to the global feature is calculated. The formula is as follows:

[0081]

[0082] Where: R i (f i g ,k) and are respectively the sets of samples with similar global features and samples with similar local features of the sample image, and |·| represents the cardinality after set operation.

[0083] Optionally, in S3, each round of training of the ResNet-ASC feature extraction model is based on the updated pseudo-labeled dataset until the model converges.

[0084] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a cattle identification method based on fine-grained information compensation loss, which can improve the learning ability of the model on unlabeled data sets, and achieve the purpose of extracting features from cattle face images and then confirming the cattle identity. Specifically:

[0085] (1) The present invention improves the feature extraction capability of the ResNet50 network by adding the ASCAM attention mechanism, so that the network focuses on the region of interest of the cow face during the feature extraction process and gives it a larger weight, so as to capture features that are more effective for identity recognition;

[0086] (2) Reliable local information is obtained by calculating the feature matching coefficient. The pseudo-labels are refined by taking advantage of the fact that the local information does not change significantly with the change of viewing angle and the posture of the cattle, so as to reduce the influence of noise on the pseudo-labels generated by global feature clustering.

[0087] (3) Based on deep metric learning, by calculating the ternary loss and dense loss for global features and local features, the distance between features of different categories is increased, while the distance between features of the same category is continuously compressed to optimize the spatial distribution of global and local features. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0089] Figure 1 The structure diagram of the ResNet-ASC feature extraction model provided by the present invention;

[0090] Figure 2 A flow chart for generating a pseudo-label dataset provided by the present invention;

[0091] Figure 3 A training flow chart of the ResNet-ASC feature extraction model provided by the present invention;

[0092] Figure 4 A feature extraction flow chart of the ResNet-ASC feature extraction model provided by the present invention;

[0093] Figure 5 A calculation flow chart of the characteristic matching coefficient provided by the present invention;

[0094] Figure 6 A calculation flow chart of fine-grained information compensation loss provided by the present invention;

[0095] Figure 7 This is an application flow chart of ResNet-ASC provided by the present invention. DETAILED DESCRIPTION

[0096] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0097] The embodiment of the present invention discloses a cattle identification method based on fine-grained information compensation loss, comprising the following steps:

[0098] S1. Model construction: Based on ResNet50 and ASCAM attention mechanism, a ResNet-ASC feature extraction model is constructed;

[0099] S2. Dataset generation: Collect several unlabeled cow face datasets, use the ResNet-ASC feature extraction model to extract global features, and generate a pseudo-label dataset by clustering the global features;

[0100] S3, model training: input the pseudo-label data set into the ResNet-ASC feature extraction model for model training, update the model parameters through back propagation and gradient descent optimization algorithm until the model converges, and obtain the trained cow face feature extraction model;

[0101] S4. Identity recognition: Use the trained cow face feature extraction model to perform identity recognition on the cow face image to be detected, and obtain the cow identity recognition result.

[0102] The above content discloses the main steps of the cattle identification method based on fine-grained information compensation loss. This technology uses cattle face images as identification marks, and has the advantages of convenient data collection, low operating risk factor, and high animal welfare. This method provides the steps for building a feature extraction model based on completely unsupervised learning, which solves the dependence of model training on calibration data.

[0103] In a specific embodiment, see Figure 1 The ResNet-ASC feature extraction model structure diagram shown in the figure shows that the model construction method of S1 is as follows:

[0104] Select ResNet50 as the base network, modify the parameters of the first Bottleneck in Layer-4, set the step size of the second convolutional layer and the downsampling layer to 1 to avoid reducing the size of the feature map and retain more local detail information;

[0105] For the global feature map output by Layer-4, add the ASCAM attention mechanism to obtain the global attention feature map, and then obtain the global feature through average pooling;

[0106] The global feature map output by Layer-4 is horizontally divided into three equal parts to obtain three local feature maps. The ASCAM attention mechanism is added to each local feature map to obtain a local attention feature map, and three local features are obtained through maximum pooling.

[0107] Furthermore, the ASCAM attention mechanism calculates the activity of a single element at two different levels, namely, the current channel and the entire channel, to measure the importance of the element and assign an activation value. The specific method is as follows:

[0108] Calculate the activity level of each element in its channel using the following formula:

[0109]

[0110] Where: Represents the mean value of elements in a single channel, Represents the variance of elements within a single channel, x i represents the elements in the channel, N represents the sum of the number of elements in the current channel, and λ represents the regularization coefficient;

[0111] Calculate the activity of each element in the entire channel. The formula is as follows:

[0112]

[0113] Where: represents the mean value of all elements in the channel, represents the variance of all elements in the channel, x i Represents the elements in the channel, C represents the sum of the number of channels in the feature map, and λ represents the regularization coefficient;

[0114] The input feature map is adjusted by the calculated element activity value. The formula is as follows:

[0115]

[0116] Where: Represents the activity value calculated based on the mean value of elements on a single channel. Represents the activity value calculated based on the mean of all elements in the channel, x i Represents the elements in the channel, Represents ASCAM attention weighted feature. The ASCAM attention mechanism achieves the refinement of the input feature map by generating a unique weight for each element in the feature map.

[0117] Furthermore, this embodiment uses a pseudo-label dataset generation method to extract global features from an unlabeled dataset, and clusters the features to generate a pseudo-label dataset, such as Figure 2 As shown, the specific steps include:

[0118] Collect several color digital images with cow faces centered and crop them to a fixed size to create an unlabeled cow face dataset ULCFDataSet; In this embodiment, a high-definition camera is used to collect color digital images of cow faces at different viewing angles, mainly cow frontal face images, with no less than 15 images for each cow, and the images are cropped to 500*500 and the cow faces are centered to create an unlabeled cow face dataset ULCFDataSet;

[0119] Each image in the ULCFDataSet is scaled and each channel is standardized to obtain a preprocessed RGB image. In this embodiment, the image can be scaled to 224*224, and then each pixel value is subtracted from the channel mean and divided by the channel standard deviation. The mean values ​​of the R, G, and B channels are 0.485, 0.456, and 0.406, and the standard deviations are 0.229, 0.224, and 0.225, respectively. The pixel value of each channel conforms to the distribution with a mean of 0 and a standard deviation of 1, which reduces the difference between channels and improves the convergence speed of the model.

[0120] The preprocessed RGB image is input into the ResNet-ASC model to extract the global feature f i g (dimension is 1*1*2048), and normalize the global features; specifically, through The feature vector f i g Each component x in j Divide by the Euclidean norm of the vector || f i g ||, where To reduce extreme values ​​in the eigenvector;

[0121] On the global feature set, similar samples are obtained through K nearest neighbor search and K mutual nearest neighbor search, the Jaccard distance of the global feature is calculated, and the Jaccard distance matrix is ​​constructed; specifically, first, for each sample i, 30 similar sample features are searched on the global feature set as the nearest neighbor set, secondly, the nearest neighbor set is expanded by finding samples that are mutually nearest neighbors, thirdly, the cosine similarity between sample i and the expanded nearest neighbor set is calculated, and normalized using the softmax function, and finally, the Jaccard distance matrix is ​​constructed based on the calculated similarity;

[0122] Perform DBSCAN clustering on the Jaccard distance matrix, assign pseudo labels to unlabeled data based on the clustering results, and retain samples of valid clustering to generate pseudo-labeled data sets. Specifically, set the hyperparameters in DBSCAN, set the maximum distance between two data points to 0.5, set the minimum number of data points contained in each cluster formed by clustering to 4, assign pseudo labels to unlabeled data based on the clustering results, and continuously update the pseudo labels after each round of training iterations, retain samples of valid clustering to generate pseudo-labeled data sets.

[0123] In order to realize cattle identification, this embodiment provides the steps of building a ResNet-ASC feature extraction model based on unsupervised learning to solve the dependence of model training on calibration data, such as Figure 3 As shown, the specific steps include:

[0124] Input the pseudo-labeled dataset into the ResNet-ASC feature extraction model to extract global features and local features, and generate global feature sets and local feature sets;

[0125] Based on local features, dynamic label smoothing and deep metric learning are used to calculate the correlation detail reconstruction loss;

[0126] Based on global features, adaptive label refinement and deep metric learning are used to calculate the intra-class and extra-class distance constraint loss;

[0127] The relevance detail reconstruction loss and the in-class and out-class distance constraint loss are summed to calculate the fine-grained information compensation loss;

[0128] The loss is compensated by fine-grained information, and the model parameters are updated through back propagation and gradient descent optimization algorithms. The pseudo-label generation stage and the model training stage are performed alternately until the model converges to obtain a trained cow face feature extraction model.

[0129] In a specific embodiment, the pseudo-labeled dataset is input into the ResNet-ASC feature extraction model to extract global features and local features, such as Figure 4 As shown, the specific steps include:

[0130] Using the intra-class balanced sampling method, m categories are randomly sampled from the pseudo-label data set, and n samples are randomly selected from each category to form a batch sample of size m*n; in this embodiment, m=8, n=4, forming a minimum batch sample of size 32;

[0131] The data enhancement operations of resizing, random flipping, image padding, random cropping, random erasing and standardization are performed on each sample image in the minimum batch sample to obtain the training data set; specifically, the cubic interpolation method is used to adjust the input cow face image to 224*224. On this basis, the data enhancement operations of randomly flipping the image with a probability of 50%, padding 10 pixels at each edge of the image, randomly cropping the image, and randomly erasing the image with a probability of 50% are performed, and each channel of the image is normalized, that is, the channel mean is subtracted from each pixel value, and then divided by the channel standard deviation. The means of the R, G, and B channels are 0.485, 0.456, and 0.406, and the standard deviations are 0.229, 0.224, and 0.225.

[0132] The training data set is input into the ResNet-ASC model, and the global feature map processed by the ASCAM attention mechanism is averaged and pooled to obtain the global feature f i g ; Among them, the dimension of the input training image is 224*224*3, and the global feature f i g The dimension is 1*1*2048;

[0133] The local feature map processed by the ASCAM attention mechanism is max-pooled to obtain a local feature map with a dimension of 1*1*2048 in each region.

[0134] Furthermore, the correlation detail reconstruction loss is calculated, which specifically includes the following steps:

[0135] In the first 5 rounds of the model, a fixed value is used to smooth the pseudo-labels, and the cross entropy loss is calculated based on the smoothed pseudo-labels and local prediction results to replace the dynamic label smoothing loss; specifically, the local features are input into the classifier composed of a fully connected layer and a softmax function to obtain the local prediction results. For pseudo labels, Smoothing is performed, where y i The pseudo-label represents the local feature, α represents the smoothing coefficient, which is set to 0.1, and K represents the total number of categories. Based on the smoothed pseudo-label and local prediction results, The cross entropy loss is calculated instead of the dynamic label smoothing loss.

[0136] When the number of model training rounds is greater than 5, the dynamic label smoothing loss is calculated; specifically, the local features are input into the classifier composed of a fully connected layer and a softmax function to obtain the local prediction results, and the cross entropy loss L between the local prediction results and the pseudo labels is calculated. l-ce , the formula is as follows:

[0137]

[0138] Where: y i represents the pseudo label of sample i, represents the local prediction result of the nth part of sample i, M is the total number of samples;

[0139] Create a uniform vector and calculate the KL divergence loss between the local prediction result and the uniform vector

[0140] Based on the feature matching coefficient, dynamically adjust the cross entropy loss L l-ce and KL divergence loss The weight of the dynamic label smoothing loss L is calculated dls , the formula is as follows:

[0141]

[0142] Where: Represents the feature matching coefficient of each local feature, represents the cross entropy loss calculated according to formula (4), u represents a uniform vector, that is, the probability of each category is equal, represents the KL divergence loss; when When it approaches 1, it indicates that the local feature information matches the global feature information well. The cross entropy loss is used to train the local prediction. When it approaches 0, it indicates that the local features contain unreliable supplementary information, and the KL divergence loss is used to guide the local prediction to approach the uniform vector;

[0143] Based on local features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct local feature triplets;

[0144] Based on the local feature triples, the local feature ternary loss function is used to calculate the local ternary loss L l-tr , the formula is as follows:

[0145]

[0146] Where: M is the total number of samples, represents the local features of sample i, represents the local difficult positive sample of sample i, represents the local difficult negative sample of sample i, ||·|| represents the Euclidean norm;

[0147] Based on the local feature triples, the local feature density loss function is used to calculate the local density loss L l-densify , the formula is as follows:

[0148]

[0149] Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, represents the local features of sample i, represents the local difficult positive sample of sample i, Represents the local feature distance between sample i and its local difficult positive sample, Represents the average distance of local features in the category where sample i belongs;

[0150] Then, the correlation detail reconstruction loss L rdr The formula is as follows:

[0151] L rdr =L dls +L l-tr +L l-densify (8).

[0152] Furthermore, the intra-class and extra-class distance constraint loss is calculated, which specifically includes the following steps:

[0153] The feature matching coefficient is used to summarize the local prediction results, and the pseudo-label is refined by supplementing reliable local fine-grained information. The formula is as follows:

[0154]

[0155] Where: represents the refined pseudo-label, y i is the pseudo label of sample i, represents the value of the feature matching coefficient calculated by the softmax function, β represents the refinement strength coefficient (set to 0.5), Represents the local prediction result of the nth part of sample i;

[0156] Based on the refined pseudo-labels, the adaptive label refinement loss L is calculated using the cross entropy loss function alr , the formula is as follows:

[0157]

[0158] Where: represents the refined pseudo-label, Represents the global prediction result obtained by inputting the global features into the classifier composed of full connection and softmax function;

[0159] Based on the global features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct global feature triples;

[0160] Based on the global feature triples, the global feature ternary loss function is used to calculate the global ternary loss L g-tr , the formula is as follows:

[0161]

[0162] Where: M is the total number of samples, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the difficult negative sample of sample i, ||·|| represents the Euclidean norm;

[0163] Based on the global feature triples, the global feature dense loss function is used to calculate the global dense loss L g-densify , the formula is as follows:

[0164]

[0165] Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the global feature distance between sample i and its difficult positive sample, Represents the average distance of global features in the category where sample i belongs;

[0166] Then, the distance constraint loss L cdc The expression is:

[0167] L cdc =L alr +L g-tr +L g-densify (13).

[0168] In a specific embodiment, the feature matching coefficient is used to measure the reliability of local information. Local features provide rich fine-grained information, but local features of the same image may contain information of different identity categories or noise information from global features. Therefore, it is necessary to evaluate the information contained in the local features, such as Figure 5 As shown, the specific steps include:

[0169] For the global features of the sample image, a K nearest neighbor search is performed on the global feature set to obtain a similar sample set of the global features; specifically, the number of search samples is set to 20;

[0170] For each local feature of the sample image, a K-nearest neighbor search is performed on the corresponding local feature set to obtain a similar sample set of each local feature; specifically, the number of search samples is set to 20. When a larger value is selected for the number of search samples, more data points with different identity categories will be present in the returned sample index, which will lead to an overall decrease in the feature matching coefficient. When a smaller value is selected for the number of search samples, the number of approximate samples returned is small and not universal, resulting in the calculation result of the feature matching coefficient being accidental.

[0171] According to the sample set obtained by the search, the feature matching coefficient of each local feature of the sample image relative to the global feature is calculated. The formula is as follows:

[0172]

[0173] Where: R i (f i g ,k) and are respectively the sets of samples with similar global features and samples with similar local features of the sample image, and |·| represents the cardinality after set operation.

[0174] In a specific embodiment, referring to Figure 6 , calculate the fine-grained information compensation loss:

[0175] L fgic =L rdr +L cdc (15);

[0176] Among them, L rdr =L dls +L l-tr +L l-densify represents the correlation detail reconstruction loss calculated based on local features using dynamic label smoothing and deep metric learning; L cdc =L alr +L g-tr +L g-densify Represents the intra-class and extra-class distance constraint loss calculated based on global features using adaptive label refinement and deep metric learning.

[0177] The model parameters are updated iteratively based on the fine-grained information compensation loss function. The optimizer uses Adam and the weight decay is set to 5x10. -4 , the initial value of the learning rate is set to 3.5x10 -4,The learning rate is reduced by 10 times after every 20 rounds, the number of training rounds is set to 50 rounds, and the pseudo-label generation stage and model training stage are performed alternately. Each round of training is based on the updated pseudo-label dataset until the model converges. Based on the above training process, as shown in Table 1, the construction and training optimization process of the ResNet-ASC feature extraction model can be summarized as follows.

[0178] Table 1 ResNet-ASC feature extraction model training process

[0179]

[0180]

[0181] In summary, this embodiment discloses a cattle identification method based on fine-grained information compensation loss. In practical application, there is no need to retrain the feature extraction model as the pasture changes. The specific application process is as follows: Figure 7 shown.

[0182] In the new pasture environment, a high-definition camera is used to collect at least 5 color digital images of the cow's face in the center of each cow in the pasture. The collected images are extracted using the ResNet-ASC feature extraction model trained on the ULCFDataSet dataset to extract global features and build a cow face feature library. The following steps are included: First, the collected cow face images are adjusted to 224*224 using the cubic interpolation method, and each channel of the image is normalized, that is, each pixel value is subtracted from the channel mean, and then divided by the channel standard deviation. The mean values ​​of the R, G, and B channels are 0.485, 0.456, and 0.406, and the standard deviations are 0.229, 0.224, and 0.225, respectively, to obtain the preprocessed RGB image; secondly, the preprocessed RGB image is input into the trained ResNet-ASC model to extract the global feature f i g , the dimension is 1*1*2048; again, through The feature vector f i g Each component x in j Divide by the Euclidean norm of the vector || f i g ||, where The acquired features are normalized, and finally, a cow face feature library is constructed based on the normalized global features.

[0183] In the identification process, a high-definition camera is used to collect a centered color digital image of the cow's face to be detected, and the cow's face image to be detected is input into the trained ResNet-ASC model to extract global features. Specifically, the cow's face image to be detected is scaled and channel normalized, and the pre-processed RGB image is input into the trained ResNet-ASC model to extract global features with a dimension of 1*1*2048 and normalized. The cow's face features to be detected and the constructed cow's face feature library are input into the K-NN classifier for distance measurement, and the cow's identity is output through a voting mechanism.

[0184] The present invention improves the feature extraction capability of the ResNet50 network by adding the ASCAM attention mechanism, so that the network focuses on the area of ​​interest of the cow face during the feature extraction process and gives it a greater weight to capture more effective features for identity recognition; obtains reliable local information by calculating the feature matching coefficient, and refines the pseudo-labels by using the characteristic that the local information will not change significantly with the change of the viewing angle and the posture of the cow, so as to reduce the influence of noise in the pseudo-labels generated by clustering; based on deep metric learning, the distance between features of different categories is increased by calculating the ternary loss and dense loss for global features and local features, while continuously compressing the distance between features of the same type to optimize the global and local feature space distribution. In summary, the present invention provides an unsupervised construction method for training a feature extraction model using an unlabeled data set, solves the problem of model training's dependence on calibration data, and achieves the purpose of extracting features from cow face images and then confirming the identity of the cow.

[0185] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0186] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cattle identification method based on fine-grained information compensation loss, characterized in that: The following steps are involved: S1: Based on ResNet50 and ASCAM attention mechanism, a ResNet-ASC feature extraction model is constructed; S2: Collect several unlabeled cow face datasets, use the ResNet-ASC feature extraction model to extract global features, and generate a pseudo-label dataset by clustering the global features; S3: Input the pseudo-label dataset into the ResNet-ASC feature extraction model for model training, and update the model parameters through back propagation and gradient descent optimization algorithm until the model converges to obtain a trained cow face feature extraction model; S4: Use the trained cow face feature extraction model to perform identity recognition on the cow face image to be detected, and obtain the cow identity recognition result.

2. A cattle identification method based on fine-grained information compensation loss according to claim 1, characterized in that: In S1, the ResNet-ASC feature extraction model is constructed as follows: Select ResNet50 as the base network, modify the parameters of the first Bottleneck in Layer-4, and set the step size of the second convolutional layer and the downsampling layer to 1; For the global feature map output by Layer-4, add the ASCAM attention mechanism to obtain the global attention feature map, and then obtain the global feature through average pooling; The global feature map output by Layer-4 is horizontally divided into three equal parts to obtain three local feature maps. The ASCAM attention mechanism is added to each local feature map to obtain a local attention feature map, and three local features are obtained through maximum pooling.

3. The cattle identification method based on fine-grained information compensation loss according to claim 1 is characterized in that: In S1, the specific operation of the ASCAM attention mechanism is: Calculate the activity level of each element in its channel using the following formula: Where: Represents the mean value of elements in a single channel, Represents the variance of elements within a single channel, x i represents the elements in the channel, N represents the sum of the number of elements in the current channel, and λ represents the regularization coefficient; Calculate the activity of each element in the entire channel. The formula is as follows: Where: represents the mean value of all elements in the channel, represents the variance of all elements in the channel, x i Represents the elements in the channel, C represents the sum of the number of channels in the feature map, and λ represents the regularization coefficient; The input feature map is adjusted by the calculated element activity value. The formula is as follows: Where: Represents the activity value calculated based on the mean value of elements on a single channel. Represents the activity value calculated based on the mean of all elements in the channel, x i Represents the elements in the channel, Represents ASCAM attention weighted features.

4. The cattle identification method based on fine-grained information compensation loss according to claim 1 is characterized in that: In S2, the generation of pseudo-label dataset includes the following steps: Collect several color digital images with cow faces centered and crop them to a fixed size to create an unlabeled cow face dataset ULCFDataSet; Scale each image in the ULCFDataSet and normalize each channel to obtain the preprocessed RGB image. The preprocessed RGB image is input into the ResNet-ASC model to extract the global feature f i g , and normalize the global features; Obtain similar samples through K nearest neighbor search and K mutual nearest neighbor search, calculate the Jaccard distance of global features, and construct the Jaccard distance matrix; Perform DBSCAN clustering on the Jaccard distance matrix, assign pseudo labels to the unlabeled data based on the clustering results, and retain the samples with valid clustering to generate a pseudo-label dataset.

5. The cattle identification method based on fine-grained information compensation loss according to claim 1 is characterized in that: In S3, the model training of the ResNet-ASC feature extraction model specifically includes the following steps: Input the pseudo-labeled dataset into the ResNet-ASC feature extraction model to extract global features and local features, and generate global feature sets and local feature sets; Based on local features, dynamic label smoothing and deep metric learning are used to calculate the correlation detail reconstruction loss; Based on global features, adaptive label refinement and deep metric learning are used to calculate the intra-class and extra-class distance constraint loss; The relevance detail reconstruction loss and the in-class and out-class distance constraint loss are summed to calculate the fine-grained information compensation loss; The loss is compensated by fine-grained information, and the model parameters are updated through back propagation and gradient descent optimization algorithms. The pseudo-label generation stage and the model training stage are performed alternately until the model converges to obtain a trained cow face feature extraction model.

6. A cattle identification method based on fine-grained information compensation loss according to claim 5, characterized in that: Extracting global features and local features includes the following steps: Using the intra-class balanced sampling method, m categories are randomly sampled from the pseudo-label data set, and n samples are randomly selected from each category to form a batch sample of size m*n; Perform data augmentation operations such as resizing, random flipping, image padding, random cropping, random erasing, and standardization on each sample image to obtain a training data set; The training data set is input into the ResNet-ASC model, and the global feature map processed by the ASCAM attention mechanism is averaged and pooled to obtain the global feature f i g ; The local feature map processed by the ASCAM attention mechanism is max-pooled to obtain the local features of each region.

7. The cattle identification method based on fine-grained information compensation loss according to claim 5 is characterized in that: Calculate the correlation detail reconstruction loss, which includes the following steps: Input the local features into the classifier composed of the fully connected layer and the softmax function, obtain the local prediction results, and calculate the cross entropy loss L between the local prediction results and the pseudo labels. l-ce , the formula is as follows: Where: y i represents the pseudo label of sample i, represents the local prediction result of the nth part of sample i, M is the total number of samples; Create a uniform vector and calculate the KL divergence loss between the local prediction result and the uniform vector Based on the feature matching coefficient, dynamically adjust the cross entropy loss L l-ce and KL divergence loss The weight of the dynamic label smoothing loss L is calculated dls , the formula is as follows: Where: Represents the feature matching coefficient of each local feature, represents the cross entropy loss calculated according to formula (4), u represents a uniform vector, represents KL divergence loss; Based on local features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct local feature triplets; Based on the local feature triples, the local feature ternary loss function is used to calculate the local ternary loss L l-tr , the formula is as follows: Where: M is the total number of samples, represents the local features of sample i, represents the local difficult positive sample of sample i, represents the local difficult negative sample of sample i, ||·|| represents the Euclidean norm; Based on the local feature triples, the local feature density loss function is used to calculate the local density loss L l-densify , the formula is as follows: Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, represents the local features of sample i, represents the local difficult positive sample of sample i, Represents the local feature distance between sample i and its local difficult positive sample, Represents the average distance of local features in the category where sample i belongs; Then, the correlation detail reconstruction loss L rdr The formula is as follows: L rdr =L dls +L l-tr +L l-densify (8) 8. The cattle identification method based on fine-grained information compensation loss according to claim 5 is characterized in that: Calculate the distance constraint loss inside and outside the class, specifically The following steps are involved: The feature matching coefficient is used to summarize the local prediction results and refine the pseudo labels. The formula is as follows: Where: represents the refined pseudo-label, y i is the pseudo label of sample i, represents the value of the feature matching coefficient calculated by the softmax function, β represents the refinement strength coefficient, Represents the local prediction result of the nth part of sample i; Based on the refined pseudo-labels, the adaptive label refinement loss L is calculated using the cross entropy loss function alr , the formula is as follows: Where: represents the refined pseudo-label, Represents the global prediction result obtained by inputting the global features into the classifier composed of full connection and softmax function; Based on the global features, the Euclidean distance between each sample and other samples in the current batch is calculated to obtain difficult positive samples and difficult negative samples, and to construct global feature triples; Based on the global feature triples, the global feature ternary loss function is used to calculate the global ternary loss L g-tr , the formula is as follows: Where: M is the total number of samples, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the difficult negative sample of sample i, ||·|| represents the Euclidean norm; Based on the global feature triples, the global feature dense loss function is used to calculate the global dense loss L g-densify , the formula is as follows: Where: M is the total number of samples, |M i | is the number of samples in the category of sample i, f i g represents the global feature of sample i, represents the difficult positive sample of sample i, represents the global feature distance between sample i and its difficult positive sample, Represents the average distance of global features in the category where sample i belongs; Then, the distance constraint loss L cdc The expression is: L cdc =L alr +L g-tr +L g-densify (13) 9. A cattle identification method based on fine-grained information compensation loss according to claim 7 or 8, characterized in that: The acquisition of feature matching coefficient includes the following steps: For the global features of the sample image, a K-nearest neighbor search is performed on the global feature set to obtain a set of similar samples of the global features; For each local feature of the sample image, perform K nearest neighbor search on the corresponding local feature set to obtain a similar sample set of each local feature; According to the sample set obtained by the search, the feature matching coefficient of each local feature of the sample image relative to the global feature is calculated. The formula is as follows: Where: R i (f i g ,k) and are respectively the sets of samples with similar global features and samples with similar local features of the sample image, and |·| represents the cardinality after set operation.

10. The cattle identification method based on fine-grained information compensation loss according to claim 1, characterized in that: In S3, each round of training of the ResNet-ASC feature extraction model is based on the updated pseudo-label dataset until the model converges.

Citation Information

Cited By

  • Cow individual identity recognition method and system based on lightweight deep learning network

    CN122024287A