A method for cow face recognition in a complex cattle farm environment

By adding patch-shift network layer and learnable Mask matrix to the Vision-Transformer model, the problem of difficulty in identifying cattle in complex cattle farm environments is solved, and a higher recognition rate and robustness are achieved.

CN115546828BActive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211155354.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-06-20
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

In complex cattle farm environments, cattle identification is difficult due to problems such as cattle herd covering, cows’ faces being dirty, and cows’ activity status and posture diversity.

Method used

A complex cattle face recognition algorithm based on Vision-Transformer is proposed. By adding patch-shift network layer and learnable Mask matrix to the VIT model, the correlation between the global features, local features and local features of the cow face image is fully learned, and the robustness of the recognition is improved.

Benefits of technology

It effectively alleviates the impact of dirty cow faces on recognition in complex cattle farm environments, improves the recognition rate and Top 1 sorting performance of cow faces in obstructed and dirty scenarios, and solves the problem of limited recognition performance of existing models in these scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546828B_ABST
    Figure CN115546828B_ABST
Patent Text Reader

Abstract

Aiming at the problem of difficult cattle identity recognition caused by problems such as cattle herd occlusion, dirty cattle faces, and the diversity of the activity states and postures of cattle in a complex cattle farm environment. The present invention discloses a cattle face recognition algorithm for a complex cattle farm environment based on Vision-Transformer. The present invention designs a brand-new feature fusion method for the complex cattle farm environment on the basis of the VIT model. A patch-shift network layer proposed by the present invention is added to the VIT model. The shift module in the patch-shift network layer fuses the information between features, not only using local features but also fusing the information between local features, alleviating the impact of dirt on recognition in the cattle farm environment. A learnable Mask matrix is added to enable the model to suppress the image patches without cattle face information, suppressing the interference of the image background, making the model pay more attention to the cattle face features in the image, and learning more robust cattle face features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biometrics, and particularly to a method for bovine face recognition in a complex cattle farm environment. Background Art

[0002] In large-scale cattle farms, to achieve individualized, automated, and information-based daily refined management, and to realize the tracking of the health status of each cow and the traceability of milk and meat products, the key lies in the identification of individual cows.

[0003] Traditional individual cow identification can be achieved by physically marking a certain part of the body or embedding a microchip, and differentiating individual cows through radio frequency ID. However, this method cannot prevent fraud, including copying marks and stealing equipment, and it will also cause harm to the cows themselves. Therefore, the method of using bovine faces for individual identification has begun to emerge.

[0004] Due to the rapid development of artificial intelligence technology in recent years, many new methods based on deep learning have also emerged in bovine face recognition. In terms of model design, the bovine face recognition technology based on Convolutional Neural Networks (CNN) has become increasingly mature. For example, the bovine face recognition algorithm based on incremental recognition constructs a sparse representation classification model using the features extracted by CNN, calculates the residuals of each category, and identifies individual cows according to the principle of the smallest residual. Or, redundant parameters of the VGG (Visual Geometry Group) network are removed, reducing the network parameters without affecting the recognition rate, providing a new idea for the cow recognition technology. There is also a method that establishes a bovine face feature extraction model based on VGG, calculates the similarity between bovine face features using Euclidean distance, and uses softmax loss and center loss as the loss functions for model training, increasing the inter-class distance of the features extracted by the model and reducing the intra-class distance, improving the recognition performance. Convolutional neural networks have achieved some results in the field of bovine face recognition due to their powerful feature expression ability. However, due to the limitation of its local receptive field, when extracting bovine face features, it often ignores the global context information of the bovine face image, cannot well represent the bovine face, and does not consider the impact of dirty bovine faces on recognition.

[0005] In terms of receptive fields, researchers have also begun to focus on solving the limitation of receptive fields before recognition. For example, the Vision-Transformer (VIT) model uses the characteristics of the global receptive field of Transformer to obtain better performance than CNN. After our tests, VIT only uses global features and ignores local features. Adding the global features of VIT directly to each piece of local features integrates the information between local and global features, but ignores the correlation between local features. Summary of the Invention

[0006] Aiming at the problem of difficult cattle identity recognition caused by problems such as cattle herd occlusion, dirty cattle faces, and the diversity of the activity states and postures of cattle in a complex cattle farm environment.

[0007] The present invention proposes a cattle face recognition algorithm based on Vision-Transformer for a complex cattle farm environment. The present invention designs a brand-new feature fusion method for the complex cattle farm environment based on the VIT model. A patch-shift network layer proposed by the present invention is added to the VIT model to fully learn the global features, local features, and the correlation between local features of cattle face images, effectively alleviating the impact of dirty cattle faces on recognition in a complex cattle farm environment, and a learnable Mask matrix is added to the algorithm to enable the model to suppress image blocks without cattle face information and learn more robust cattle face features. To solve the problem of difficult cattle identity recognition in a complex cattle farm environment,

[0008] The technical solution adopted by the present invention is as follows:

[0009] 1. A cattle face recognition method for a complex cattle farm environment, comprising the following steps:

[0010] S1. Collect cattle face data:

[0011] Under different lighting conditions and at the same shooting height, collect cattle face video data of three different cattle face postures, namely the front face, the left side face, and the right side face, simultaneously. Intercept the front face, left side face, and right side face picture data of each cattle from the video stream, and divide the cattle face picture data into a training set and a test set;

[0012] S2. Perform data processing on cattle faces in a complex cattle farm environment based on Vision-Transformer

[0013] S2-1 First, divide the input cattle face picture into N image blocks of the same size , and use the image block encoder E of Vision-Transformer to encode each image block into a feature vector with a dimension of D ;

[0014] S2-2 Then, in the matrix composed of N feature vectors , add a learnable classification vector x cls , and the classification vector x cls is used to represent the global feature after the cattle face image is encoded,

[0015] S2-3 Finally, add a position encoding containing spatial information to obtain the input sequence of the encoder:

[0016]

[0017] When S2-4 propagates forward to the (l-1)-th layer encoder before z0, the extracted cow face features are respectively input into the global branch and the local information fusion branch. Among them, the cow face features input into the global branch are used as the global branch input features, and the cow face features input into the local information fusion branch are used as the local branch input features;

[0018] S2-5 inputs the global branch input features into the l-th layer encoder in the global branch to extract global branch features;

[0019] S2-6 In the local information fusion branch, the patch-shift network layer is used to fuse the global features of the cow face and the local branch input features to obtain the features after information fusion of the patch-shift network layer;

[0020] S2-7 inputs the features after information fusion of the patch-shift network layer into the l-th layer encoder to obtain the final output feature S that includes the correlation between features, S = TransformerLayer(G M )

[0021] S2-8 Finally, the global branch features extracted by the global branch and the output features that include the correlation between features extracted by the local information fusion branch are input into the MLP for classification;

[0022] S3. Use the training set described in the step S1 to construct a loss function, and train the method for processing cow face data based on Vision-Transformer in the step S2. When the total loss drops to no more than 0.01, the training ends, and a trained cow face data processing method is obtained;

[0023] S4. Input the data in the test set described in the step S1 into the trained cow face data processing method to extract cow face image features and perform recognition and comparison.

[0024] Preferably, the loss function in the S3 includes: triplet loss L triplet and cross-entropy loss L softmax .

[0025] Preferably, in the S2-6, the structure of the local information fusion branch from bottom to top includes: adaptive average pooling layer, splicing layer, patch-shift network layer, l-th layer encoder, MLP classifier;

[0026] The specific process of pooling and splicing in the adaptive average pooling layer and the splicing layer includes the following steps:

[0027] First, the encoder output z of the (l - 1)-th layer l-1 Using average pooling, the N + 1 local branch input features are evenly divided into K parts, and then the K evenly divided local branch input features and the global feature are concatenated to obtain the input feature of the local information fusion branch That is:

[0028]

[0029] In the formula, γ is the adaptive average pooling layer, and ψ represents the concatenated local branch input feature and global feature after pooling;

[0030] The structure of the patch-shift network layer from bottom to top includes: a shift module, a convolutional layer Conv with a kernel size of 1; a learnable matrix Mask; an activation function ReLU; the operation process of the patch-shift network layer includes the following steps:

[0031] Input G0 into the M-layer patch-shift network layer for information fusion, and the output of the m-th layer patch-shift network layer is

[0032] G m = ReLU(Conv(shift(G m-1 ))⊙Mask + G m-1 ) m = 1, …, M

[0033] In the formula, G m-1 is the output of the (m - 1)-th layer patch-shift network layer, and the input of the first layer patch-shift is G0; shift is the shift module for fusing feature information proposed in this paper; Conv is a convolutional layer with a kernel size of 1; Mask is a learnable matrix for adaptively learning feature correlations; ReLU is an activation function.

[0034] Preferably, the shift module is used for inter-channel information fusion of the local branch input feature and the global feature; when the shift module fuses feature information, the j-th channel information value G m-1 of the i-th block feature of the feature G m-1 (i, j) is equal to the j-th channel information value G m-1 of the ((i + j) % (K + 1))-th block feature of the feature G m-1 ((i + j) % (K + 1), j), that is:

[0035] G m-1 (i, j) = G m-1 ((i + j) % (K + 1), j) i = 0, …, K; j = 0, …, D - 1

[0036] Among them, G m-1 is the input feature of the shift module in the m-th layer patch-shift, and is the global feature, and f i m-1 |i = 1, 2, …, K are the input features of the local branches.

[0037] Preferably, in the step S4, the recognition comparison is performed based on the cosine distance;

[0038] The cosine distance calculation formula is as follows:

[0039]

[0040] The larger the cosine distance, the higher the similarity of the surface features of the two cow faces. On the contrary, the lower the similarity of the surface features of the two cow faces extracted by the model;

[0041] The process of performing the recognition comparison based on the cosine distance includes the following steps:

[0042] Perform 1:1 different-class comparisons on the features obtained by extracting the test set cow face images and combining and normalizing them after being divided into 6, 4, and 2 equal parts to obtain the model comparison threshold T;

[0043] Then perform same-class comparisons on the images in the test set. When the comparison value of the combined features extracted from the same class is greater than T, it is regarded as a successful comparison; otherwise, it is considered a failed comparison.

[0044] The beneficial effects of the present invention are as follows:

[0045] The present invention proposes a cow face recognition algorithm based on Vision-Transformer. When using a Convolutional Neural Networks (CNN) to extract cow face features, the global context information is often ignored, and only local feature information of the cow face image can be extracted. The global receptive field characteristic of the Vision-Transformer (VIT) model can effectively improve the problem of the local receptive field of the CNN. However, the previous VIT recognition models did not consider the correlation between local features. By simply and violently dividing the local features, when there are large-scale occlusions, dirt, and multiple poses, the comparison value of the corresponding part will be relatively large when matching with the template, and the overall comparison value will be greater than the threshold, resulting in recognition failure. Therefore, the performance improvement brought by such a simple division in recognizing cows in complex scenarios is very limited. Compared with the original VIT model, the patch-shift network layer added in the local information fusion branch allows for full information interaction between the global features and local features of the cow face, enabling the model to fully learn the global features, local features, and the correlation between local features, thus making up for the defect of the VIT-based model ignoring the correlation between local features.

[0046] The previous VIT recognition models only simply perform noise reduction processing on the background noise in the image and do not consider the correlation between the background and the cow face information features. When there are strong lighting changes and matching with the template, the strong and dark lighting of the background will make the comparison value at the edge relatively large, and the overall comparison value will be greater than the threshold, resulting in recognition failure. By adding a learnable Mask matrix, the model can suppress the image patches without cow face information and learn more robust cow face features. Then, the features after information fusion are obtained by inputting them into the l-th encoder. Finally, the features extracted from the global branch and the local information fusion branch are input into the MLP for classification to accurately obtain the correct prediction result.

[0047] The present invention proposes a cow face recognition algorithm based on Vision-Transformer, which can effectively solve the problem of cow identity recognition in a complex cow farm environment and obtain excellent recognition performance in 1:1 comparison recognition in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is the structure of the Vision-Transformer cow face recognition algorithm

[0049] Figure 2 It is the model training flow chart

[0050] Figure 3 It is a partial cow face image of the COWYCTC-903 dataset

[0051] Figure 4 For the feature visualization heat map

[0052] Figure 5 For the shift module diagram

[0053] Figure 6 For the MLP module diagram

[0054] Figure 7 For the patch - shift network layer structure diagram

[0055] Figure 8 For the ROC curves of 4 test algorithms on 4 datasets Specific embodiments

[0056] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings.

[0057] This embodiment is a cow face recognition algorithm for complex cow farm environments based on Vision - Transformer, including the following steps:

[0058] S1. Construct a dataset

[0059] S1 - 1: When shooting a video, first, in order to improve the generalization ability of the cow face recognition model in the lighting scenarios of the cow farm, different lighting conditions are selected for shooting the cow face video during data collection; second, in order for the cow face recognition model to achieve multi - angle recognition of cow faces, while ensuring the same shooting height during data collection, 30 - second cow face videos are shot for each of the three different cow face postures of the front face, left face, and right face at each time period. According to the object detection algorithm, pictures of each cow are intercepted from the video stream and classified, numbered, and screened. The dataset altogether contains 903 classes of cow face image datasets COWYCTC - 903 with different angles and different lighting conditions. Among them, there are 868 classes of normal images. Each cow has 3 different postures of the left face, front face, and right face, and each posture contains 15 cow face images, so each cow has a total of 45 images; there are 35 classes of special images containing occlusions or dirt, with 10 images for each cow. The first 800 classes of the dataset are used as the training set.

[0060] S1 - 2: The training set obtained in the above steps is expanded 20 times through data augmentation methods such as rotation, translation, scaling, and brightness transformation. Finally, the number of training set pictures is 800×45×20 = 720000. The number of validation set pictures in the normal image library is 68×15×3 = 3060, with 1020 images for each of the front face, left face, and right face; the number of validation set pictures in the special image library is 10×35 = 350. The expanded images are normalized to 192*192 to obtain the training dataset. Some cow face images of COWYCTC - 903 are as Figure 3 shown.

[0061] S2. Design a cow face recognition algorithm based on Vision-Transformer for complex cow farm environments

[0062] The basic framework of the cow face recognition algorithm based on Vision-Transformer proposed in this invention is as follows Figure 1 shown

[0063] Figure 1 Among them, the input image is H, W, and C respectively represent the height, width, and number of channels of the input image. The image preprocessing method of this invention is the same as that of VIT

[0064] First, divide the input picture into N image patches of the same size and use the image patch encoder E of VIT to encode each image patch into a feature vector with a dimension of D

[0065] After that, add a learnable classification vector x to the N feature vectors cls , the classification vector x cls is used to represent the global feature after encoding the cow face image

[0066] Finally, add the position encoding containing spatial information to obtain the input sequence of the encoder

[0067]

[0068] Different from directly inputting the input sequence z0 into the Transformer encoder with l layers to obtain N + 1 feature vectors in VIT. In this invention, when z0 propagates forward to the (l - 1)-th layer encoder, the extracted cow face features are respectively input into the global branch and the local information fusion branch. Among them, the cow face features input into the global branch are used as the input features of the global branch, and the cow face features input into the local information fusion branch are used as the input features of the local branch

[0069] First, in the global branch, input the input features of the global branch into the l-th layer encoder to extract the global branch features

[0070] Secondly, in the local information fusion branch, use the patch-shift network layer to allow the global feature of the cow face and the input features of the local branch to perform sufficient information interaction to obtain the features after information fusion of the patch-shift network layer

[0071] After that, input the features after information fusion of the patch-shift network layer into the l-th layer encoder to obtain the final output features including the correlation between features

[0072] Finally, the global branch features extracted by the global branch and the output features containing the correlation between the local information fusion branch-extracted inclusion features are input into the MLP for classification.

[0073] We used the Vision-Transformer cow face recognition algorithm, VGG cow face recognition algorithm, ResNet-50, and the model generated by the LA-Transformer algorithm proposed based on VIT in the present invention to draw visual heatmaps of the features extracted by the visual models for the normal image test sets of the front face, left face, and right face of the same cow and the special image library with dirt, respectively, as Figure 4 shown. It can be seen from the figure that compared with the other three algorithms, the cow face recognition algorithm proposed in the present invention can correctly focus on the cow face part in the front face, left side face, and right side face of the normal test set, suppressing the interference of the background; at the same time, for the special images with dirt, it can also correctly focus on the cow face area by using the correlation between local features.

[0074] Global branch

[0075] For the global branch, directly input z l-1 into the l-th layer encoder to obtain and take the classification vector g containing the global information of the cow face cls as the global branch features extracted from the cow face image by the present invention.

[0076] Local information fusion branch

[0077] The local information fusion branch mainly uses the patch-shift network layer to fuse the information of the global features and the local branch input features. The structure of the patch-shift network layer is as Figure 7 shown. According to Figure 1 , first use average pooling to evenly divide the N + 1 local features into K parts for the encoder output z of the (l - 1)-th layer, and then concatenate the K evenly divided local branch input features and the global features l-1 to obtain the input features of the local information fusion branch i.e.: That is:

[0078]

[0079] In the formula, γ is the adaptive average pooling layer, and ψ represents the concatenated pooled local branch input features and global features.

[0080] Then input G0 into the M-layer patch-shift network layer for information fusion. The output of the m-th layer patch-shift network layer is

[0081] Gm = ReLU(Conv(shift(G m-1 )) ⊙ Mask + G m-1 ) m = 1,..., M (3)

[0082] Where G m-1 is the output of the (m - 1)-th layer patch-shift network layer. In particular, the input of the first layer patch-shift is G0; shift is the shift module proposed in this paper to fuse feature information; Conv is a convolutional layer with a kernel size of 1; Mask is a learnable matrix for adaptively learning feature correlations; ReLU is an activation function.

[0083] Finally, the feature G M after information fusion through M layers of patch-shift network layers is input into the l-th layer encoder to obtain the final output feature that includes the correlations between features:

[0084] S = TransformerLayer(G M ) (4)

[0085] shift module

[0086] To avoid the problem that the VIT-based model only focuses on the relationship between the global feature and the local branch input feature and ignores the correlation between the local branch input features, the shift module in the patch-shift network layer is used to fuse the information between features, which not only utilizes the local branch input features but also fuses the information between the local branch input features, alleviating the impact of dirt on recognition in the cattle farm environment.

[0087] The input feature of the shift module in the m-th layer patch-shift is the global feature, and f i m-1 | i = 1, 2,..., K are the local branch input features. The shift module fully utilizes the local branch input features and the global feature to perform inter-channel information fusion between the local branch input features and the global feature. When the shift module fuses the feature information, the j-th channel information value G m-1 of the i-th block feature is equal to the j-th channel information value G m-1 (i, j) of the feature G m-1 of the ((i + j) % (K + 1))-th block feature, that is: G m-1 ((i + j) % (K + 1), j), namely: G m-1 (i, j) = G m-1 ((i + j) % (K + 1), j) i = 0,..., K; j = 0,..., D - 1 (5)

[0088] After passing through the shift module, the K+1 blocks of features in G m-1 all contain the information of the other K blocks of features and themselves, fully integrating the information between the local branch input features and the global features, and solving the problem that the VIT-based model ignores the correlation between the local branch input features. Figure 5 describes the calculation process of the shift module when K = 4 and D = 20.

[0089] Figure 5 The vertical arrows in indicate that the information at the tail of the arrow is assigned to the head of the arrow. It can be seen from the feature map after shift in the figure that the local branch input features and the global features have fully integrated the information between each other in the channels.

[0090] Although the global features of the cow face and the local branch input features are fully integrated after passing through the shift module, the correlations between the global features, the local branch input features, and the local branch input features are different. Therefore, in this paper, a learnable Mask matrix is added after the shift module to allow the network to adaptively learn the correlations between the features. And by observing the process of dividing the cow face image into blocks, it can be found that the importance of each image block for the network learning is different. For example, some image blocks only contain background information without cow face images, and this type of image block may even affect the network learning. After adding the Mask matrix, the network can also adaptively focus on the image blocks with cow faces and suppress the interference of the image background.

[0091] Finally, the g of the global branch cls and the output feature S of the local information fusion branch are respectively input into the MLP module to calculate the prediction result of the model. The MLP structure is as Figure 6 shown.

[0092] S3. Train the entire model, as Figure 2 shown. First, construct the loss function. Since there are images with large similarities between different types of cow face images in the cow farm, in order to make the distance between the features of different types of cow face images as large as possible and the distance between the features of the same type of cow face images as small as possible during training, so that the network can learn the detailed differences of similar cow faces and obtain better cow face representation features. The present invention uses the triplet loss (hard triplet loss) L triplet and the cross-entropy loss L softmax to jointly train the network. The triplet loss is as follows:

[0093]

[0094] In the formula, N is the number of cow face images in a batch during training, is a sample of a cow face image randomly selected in a training batch, is in the same training batch as a sample of a cow face image belonging to the same class, is in the same training batch as a sample of a cow face image belonging to different classes, is and the Euclidean distance between the features extracted by the network model, is and the Euclidean distance between the features extracted by the network model, where a is a hyperparameter representing and the minimum interval between them, which is set to 0.3 in this paper. During training, the triplet loss narrows the distance between and and widens the distance between and to achieve the goal of having a close distance between cow face images of the same class and a far distance between cow face images of different classes, enabling the network model to finally extract cow face features with higher fineness.

[0095] When calculating the loss, the output features of the global branch and the local information fusion branch in the Vision-Transformer cow face recognition algorithm are concatenated as the final recognition features:

[0096] f all =[g cls ; s0; s1; s2; …; s K (7)

[0097] When calculating the loss, f all is used to calculate the final triplet loss and cross-entropy loss:

[0098]

[0099]

[0100] Total loss of the model:

[0101] E = E triplet + E softmax (10)

[0102] Iterate through the entire training set several times until the total loss of the model drops to around 0.01.

[0103] S4. Input the test set images into the trained model to extract the surface features of the cow face images and perform recognition and comparison.

[0104] The present invention uses the cosine distance as the standard for measuring the similarity of the surface features of cow faces. The cosine distance calculation formula is as follows:

[0105]

[0106] The larger the cosine distance, the higher the similarity of the surface features of two cow faces. Conversely, the lower the similarity of the surface features of two cows extracted by the model. The features obtained by extracting the test set cow images and combining and normalizing them after being equally divided into 6, 4, and 2 parts are compared one by one for different classes, and the model comparison threshold T is obtained. Then, the images in the test set are compared for the same class. When the comparison value of the combined features extracted for the same class is greater than T, it is considered a successful comparison. Conversely, it is considered a failed comparison.

[0107] Among them, different classes specifically refer to the surface features of cow faces extracted from the test set data according to the training model. The square value of the difference is calculated between the features of the cow face pictures of each cow and the features of the cow face pictures of other cows. The smallest value among all the finally obtained square values of the differences is determined as the model comparison threshold T.

[0108] The same class specifically refers to the surface features of cow faces extracted from the test set data according to the training model. The square value of the difference is calculated between the features of the cow face pictures of each cow and the features of its own other cow face pictures. If the square value of the difference is less than the model comparison threshold T, it is considered a successful comparison. Conversely, it is considered a failed comparison.

[0109] The experimental server GPU used in the present invention is NVIDIA TITAN RTX 3090, and the deep learning framework used is Pytorch. The input image is a 3-channel cow face image with a resolution of 224×224. The number of image patches N of the Transformer encoder is 196, and the feature dimension D after encoding the image patches is 768. Data augmentation methods such as random cropping, random horizontal flipping, and random erasing are used during training. The training batch size is 64, which includes 4 cow face pictures of 16 cows. The loss function is optimized by the Adaptive Momentum Estimation (ADAM) optimizer, and the learning rate of the optimizer is set to 3e-4, and the weight decay coefficient is 5e-4. The number of feature blocks K in the shift module is 4, and the number of patch-shift network layers M is 2.

[0110] The following is the experimental data analysis of the algorithm proposed in the present invention based on the divided test set database. The present invention compares the model performance with the VGG cow face recognition algorithm, ResNet-50, and LA-Transformer algorithm in the frontal face, left face, right face, and special image test sets with pollution in the COWYCTC-903 normal image test set respectively.

[0111] To verify the superiority of our model, we use ROC and Top1 ranking to compare the model performance. The discrimination rate compares the performance through the ROC curve. The abscissa is the False Acceptance Rate (FAR), and the ordinate is the False Rejection Rate (FRR). The false acceptance rate is the proportion of cattle face images of different categories that are judged to be of the same category when matched 1:1; the false rejection rate is the proportion of cattle face images of the same category that are judged to be of different categories when matched 1:1. The zero false acceptance and rejection rate is the rejection rate when the false acceptance rate is equal to 0.

[0112] The Top1 ranking performance is compared by counting the Top1 ranking success rate: select the first cattle face image of the same category as the template, and the remaining images within the category as the verification images. Compare the verification images with the template and the images outside the category, and count the proportion of the template ranked first.

[0113] We compare 3410 pieces of data in the test set. Draw the ROC characteristic curves simulated from the normal image library and the contaminated special image library in the COWYCTC-903 test set for the four models, as Figure 8 shown. For the test sets of the three poses of the frontal face, left face, and right face in the normal image library and the contaminated special image library, when the FAR is 0, the FRR of the cattle face recognition algorithm proposed in this paper is the lowest. Compared with the VGG cattle face recognition algorithm, it is reduced by 7.87%, 11.26%, 15.21%, and 18.96% respectively; compared with LA-Transformer, it is reduced by 0.53%, 0.7%, 0.78%, and 4.1% respectively. It effectively improves the recognition performance of cattle face contamination caused by the cattle farm environment.

[0114] The Top1 performance of the four algorithms on different pose data sets is shown in Table 1.

[0115] Table 1 Comparison of Top1 performance of different algorithms on cattle face data sets with different poses Unit: %

[0116]

[0117] As can be seen from Table 1, for the data sets of the three different poses of the frontal face, left face, and right face and the contaminated special image library, the Vision-Transformer cattle face recognition algorithm proposed in the present invention effectively improves the Top1 ranking success rate. Compared with the VGG cattle face recognition algorithm, it is increased by 9.19%, 10.77%, 13.5%, and 31.86% respectively; compared with the LA-Transformer algorithm, it is increased by 2.21%, 1.15%, 1.36%, and 4.52% respectively.

[0118] In view of the problem of difficult cattle identity recognition caused by problems such as cattle herd occlusion, dirty cattle faces, and the diversity of the activity states and postures of cattle in a complex cattle farm environment, a cattle face recognition algorithm based on Vision-Transformer is proposed, which fully integrates the information between the local features and global features of cattle face images. Compared with the current mainstream recognition models, the recognition rate and Top1 ranking performance of cattle faces in occlusion and dirty scenarios are effectively improved. This also demonstrates the effectiveness of the method proposed in the present invention.

[0119] The above embodiments of the present invention have been described in detail with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and these should also be regarded as the protection scope of the present invention.

Claims

1. A method for bovine face recognition in a complex cattle farm environment, characterized in that, The following steps are involved: S1. Collect cow face data: Under different lighting conditions and at the same shooting height, we simultaneously collected cow face video data of three different cow face postures: front face, left face, and right face. We intercepted the front face, left face, and right face image data of each cow from the video stream, and divided the cow face image data into a training set and a test set. S2. Data processing of cow faces in complex cattle farm environments based on Vision-Transformer S2-1 First, divide the input cattle face image into N image patches of the same size And use the patch encoder E of Vision-Transformer to encode each image patch into a feature vector with dimension D After S2-2, in the matrix composed of N eigenvectors add the learnable classification vector x cls , the classification vector x cls is used to represent the global features of the cow face image after encoding Finally, add the positional encoding containing spatial information at S2-3 to obtain the input sequence of the encoder: When S2-4 propagates forward to the (l-1)-th layer encoder before z0, the extracted cattle face features are respectively input into the global branch and the local information fusion branch. Among them, the cattle face features input into the global branch serve as the global branch input features, and the cattle face features input into the local information fusion branch serve as the local branch input features; S2-5 inputs the global branch input feature into the l-th layer encoder in the global branch to extract the global branch feature; S2-6 uses the patch-shift network layer in the local information fusion branch to fuse the global features of the cow face with the local branch input features to obtain the features after the patch-shift network layer information fusion; S2-7 Input the features after fusing the patch-shift network layer information into the l-th layer encoder to obtain the final output feature S that includes the correlation between features, S = TransformerLayer(G M ); S2-8 finally inputs the global branch features extracted by the global branch and the output features including the correlation between the features extracted by the local information fusion branch into the MLP for classification; S3, using the training set described in step S1, constructing a loss function, training the method for data processing of cow faces in a complex cattle farm environment based on Vision-Transformer in step S2, and terminating the training when the total loss drops to no more than 0.01, to obtain a trained cow face data processing method; S4, inputting the data in the test set described in step S1 into the trained cow face data processing method, extracting cow face image features and performing recognition and comparison.

2. The method for bovine face recognition in a complex cattle farm environment according to claim 1, characterized in that, The loss function in S3 includes: triplet loss L triplet and cross-entropy loss L softmax .

3. The method for bovine face recognition in a complex cattle farm environment according to claim 2, characterized in that, In S2-6, The structure of the local information fusion branch includes from bottom to top: an adaptive average pooling layer, a splicing layer, a patch-shift network layer, a first layer encoder, and an MLP classifier; The process of performing pooling and splicing in the adaptive average pooling layer and the splicing layer specifically includes the following steps: First, take the encoder output z of the (l - 1)-th layer l-1 Use average pooling to evenly divide the N + 1 local branch input features into K parts, and then combine the K evenly divided local branch input features and the global feature Concatenate them to obtain the input feature of the local information fusion branch That is: In the formula, γ is the adaptive average pooling layer, ψ represents the local branch input features and global features after concatenation and pooling; The structure of the patch-shift network layer includes, from bottom to top: a shift module, a convolution layer Conv with a convolution kernel size of 1; a learnable matrix Mask; an activation function ReLU; and the operation process of the patch-shift network layer includes the following steps: Input G0 into the patch-shift network layer of the Mth layer for information fusion. The output of the patch-shift network layer of the mth layer is G m = ReLU(Conv(shift(G m-1 )) ⊙ Mask + G m-1 ) for m = 1, …, M where G m-1 is the output of the (m - 1)-th layer patch-shift network layer, and the input of the first layer patch-shift is G0; shift is the shift module proposed in this paper to fuse feature information; Conv is a convolutional layer with a kernel size of 1; Mask is a learnable matrix for adaptively learning feature correlations; ReLU is an activation function.

4. The method for bovine face recognition in a complex cattle farm environment according to claim 3, characterized in that, The shift module is used to perform information fusion between channels for the local branch input features and the global features; when the shift module fuses the feature information, the feature G m-1 The j-th channel information value G of the i-th block of features m-1 (i, j) is equal to the feature G m-1 The information value G of the j-th channel of the ((i + j) % (K + 1))-th block of features m-1 ((i + j) % (K + 1), j), that is: G m-1 (i, j) = G m-1 ((i + j) % (K + 1), j) where i = 0, …, K; j = 0, …, D - 1 Among them, G m-1 is the input feature of the shift module in the m-th layer of patch-shift, is the global feature, is the input feature of the local branch.

5. The method for bovine face recognition in a complex cattle farm environment according to claim 1, characterized in that, In S4, the identification comparison is performed based on the cosine distance; The formula for calculating cosine distance is as follows: The larger the cosine distance is, the higher the similarity of the surface features of the two cow faces is. Conversely, the lower the similarity of the surface features of the two cow faces extracted by the model is. The process of identification and comparison based on cosine distance includes the following steps: The features extracted from the cow face images of the test set are divided into 6, 4, and 2 equal parts, and then combined and normalized for 1:1 comparison between different categories to obtain the model comparison threshold T; Then the images in the test set are compared with the same type. When the comparison value of the combined features extracted from the same type is greater than T, the comparison is considered successful; otherwise, the comparison is considered failed.

Citation Information

Patent Citations

  • Cow face detection and recognition method based on deep learning

    CN111368766A

  • High-accuracy face recognition improved algorithm

    CN112364809A