Cow face recognition method and system based on multi-scale feature interaction
By introducing multi-scale feature interaction, M_CBAM attention mechanism and Cure-Triplet Loss into the bull face recognition algorithm, the problem of low recognition accuracy of the existing bull face recognition algorithm in light changes, angle changes and complex backgrounds is solved, and higher recognition accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510201179.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-08
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
When faced with light changes, angle changes and complex backgrounds, the existing cattle face recognition algorithm has low recognition accuracy and complex structure, making it difficult to adapt to the actual pasture environment.
A method of face recognition based on multi-scale feature interaction is designed, and feature maps are extracted in different convolutional layers through multi-scale feature fusion modules, and shallow features and deep features are interacted with the deep features. At the same time, a new attention mechanism M_CBAM and cluster ternary loss Cure-Triplet Loss are introduced to improve the adaptability and recognition accuracy of the algorithm.
It improves the accuracy and robustness of cattle face recognition, can better adapt to complex pasture environments, reduce labor costs, and has good practical value.
Smart Images

Figure BDA0005283144960000031 
Figure BDA0005283144960000071 
Figure BDA0005283144960000081
Abstract
Description
Technical Field
[0001] The present invention relates to the field of livestock monitoring. Specifically, it relates to a cow face recognition method and a recognition system based on multi-scale feature interaction. Background Art
[0002] In recent years, with the development of the livestock industry, automation and informatization have been applied to cattle farm breeding, and cattle are monitored and managed through a precision breeding model. Cattle individual identification is the key to achieving precision breeding. Traditional identification methods mainly identify by wearing electronic ear tags based on radio frequency identification technology (RFID). However, this method has risks of being easily lost and forged, and is time-consuming, laborious, and has a high labor cost for identification. With the continuous development of biometric technology, researchers have begun to use traditional feature extraction operators and digital image processing technology to identify cow faces. However, when dealing with complex conditions such as light changes, angle changes, and complex backgrounds, the performance of these methods may be limited.
[0003] Currently, face recognition technology based on deep learning has become a research hotspot in the field of computer vision and is widely used in fields such as intelligent security, AI breeding, and self-service. As face recognition algorithms have become relatively mature in face recognition, livestock individual identification technology based on deep learning has also developed to a new stage, and various individual identification algorithms have emerged. However, there are still some problems to be further solved. First, since most ranch breedings are outdoors, the cow face is greatly affected by light. Second, considering that the cow face has non-planarity, and most existing cow face recognition algorithms use a direct addition aggregation method, which cannot highlight the characteristics of each cow well, so the existing recognition algorithms have a low recognition accuracy for cow faces. In addition, the structure of existing cow face recognition algorithms is still relatively complex. When the dataset is large, the corresponding computational complexity is high, and it is difficult to be transplanted to an embedded platform.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] In view of this, the present invention discloses a cow face recognition method and its recognition system based on multi-scale feature interaction. First, a multi-scale feature fusion module is designed to extract feature maps of different scales from different convolutional layers of the feature extraction network, and the extracted shallow features are made to interact with the deep features to achieve the purpose of accurately recognizing cow faces. Then, a new attention mechanism M_CBAM is designed at the output ends of the shallow network and the deep network respectively. This attention mechanism designs a parallel structure of channel and spatial attention on the basis of the traditional attention mechanism CBAM and integrates adaptive weight coefficients. The weight coefficients of each attention branch are adjusted according to the important features of the feature map, so as to automatically adjust the channel and spatial attention mechanisms by using dynamic weighting under different cow faces, and solve the problem of poor adaptability of the traditional CBAM to different cow faces, thereby improving the recognition accuracy of the algorithm. Finally, the clustering idea is introduced into the loss function, the training samples are clustered to obtain different clusters, each cluster is marked with a category, and then combined with Triplet Loss as the overall loss, thereby improving the recognition accuracy of the algorithm. This method is more accurate and has better effects than traditional recognition methods, and has good practical value.
[0006] Specifically, the present invention is realized through the following technical solutions:
[0007] In the first aspect, the present invention discloses a cow face recognition method based on multi-scale feature interaction, including the following steps:
[0008] Unify the size of the cow face image, and upsample the features extracted by the deep feature extraction network to be the same size as the features extracted by the shallow feature extraction network.
[0009] Construct the attention mechanism M_CBAM at the output ends of the deep feature extraction network and the shallow feature extraction network respectively to form a multi-scale feature fusion network, and realize the automatic adjustment of the channel and spatial attention mechanisms.
[0010] Use the clustering triplet loss Cure-Triplet Loss to assist the convergence of the network.
[0011] Preferably, the features output by the third basic unit in the extracted features are used as shallow features, and the finally output features are used as deep features; the above two types of features are respectively input into the M_CBAM attention mechanism, and the output features are concatenated.
[0012] Preferably, the concatenation step adopts the following calculation formula:
[0013]
[0014] In the formula: is the dimension of the shallow feature, is the dimension of the deep feature; H represents the height of the image; W represents the width of the image; C represents the number of channels of the image.
[0015] In a second aspect, the present invention discloses a cow face recognition system for complex scenes, including:
[0016] Sampling module: used to unify the size of cow face images, perform upsampling on the features extracted by the deep feature extraction network, and convert them into the same size as the features extracted by the shallow feature extraction network;
[0017] Construction module: used to construct the attention mechanism M_CBAM at the output ends of the deep feature extraction network and the shallow feature extraction network respectively, form a multi-scale feature fusion network, and realize the automatic adjustment of the channel and spatial attention mechanisms;
[0018] Convergence module: used to assist the convergence of the network by using the clustering triplet loss Cure-Triplet Loss.
[0019] In a third aspect, the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the cow face recognition method based on multi-scale feature interaction as in the first aspect are implemented.
[0020] In a fourth aspect, the present invention discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the cow face recognition method based on multi-scale feature interaction as in the first aspect are implemented.
[0021] The solution of the present invention can effectively address the complex and variable breeding environment in actual cattle farms and the non-planar characteristics of cattle face shapes, which often lead to insufficient recognition accuracy and poor robustness in the practical application of face recognition technology. By designing a multi-scale feature fusion module, feature maps at different scales are extracted from different convolutional layers of the feature extraction network, and the extracted shallow features are interacted with the deep features to achieve the purpose of accurately recognizing cattle faces. Subsequently, a new attention mechanism M_CBAM is constructed at the output ends of the shallow feature extraction network and the deep feature extraction network respectively. This attention mechanism designs a parallel structure of channel and spatial attention on the basis of the traditional attention mechanism CBAM and integrates adaptive weight coefficients. According to the important features of the feature map, the weight coefficients of each attention branch are adjusted, and dynamic weighting is used under different cattle faces to automatically adjust the channel and spatial attention mechanisms, improving the generalization ability of the model. Finally, a clustering triplet loss Cure-Triplet Loss is designed. By clustering the training samples into different clusters and labeling each cluster, the discrimination of the samples is improved, and then combined with Triplet Loss as the overall loss to enhance the recognition accuracy of the algorithm. Compared with some existing methods, this method can efficiently and accurately recognize cattle faces and has good application value.
[0022] Specifically, the implementation steps of the method of the present invention are mainly carried out in the following manner:
[0023] A. Cattle face feature interaction stage:
[0024] 1) Multi-scale feature interaction: Considering that the deepening of the number of network layers in a deep neural network will cause the original details in the cattle face image to become increasingly blurred or even lost after multiple convolutional operations; while the shallow network has a smaller receptive field and can enhance the expression ability of detail features. Therefore, a multi-scale feature fusion module is designed to extract feature maps at different scales from different convolutional layers of the feature extraction network, and the extracted shallow features are interacted with the deep features to achieve the purpose of accurately recognizing cattle faces.
[0025] 2) New attention mechanism: To enable the model to learn the key features in the feature map as efficiently as possible, the present invention constructs a new attention mechanism M_CBAM at the output ends of the shallow network and the deep network respectively. This attention mechanism designs a parallel structure of channel and spatial attention on the basis of the traditional attention mechanism CBAM and integrates adaptive weight coefficients. According to the important features of the feature map, the weight coefficients of each attention branch are adjusted, and dynamic weighting is used under different cattle faces to automatically adjust the channel and spatial attention mechanisms, solving the problem of poor adaptability of the traditional CBAM to different cattle faces, and thus improving the recognition accuracy of the algorithm.
[0026] B. Identity Recognition Phase:
[0027] 1) Clustering Triplet Loss Function: Triplet Loss is used as the loss function of the FaceNet network model to train the model so that the feature vectors have good discriminability. For a given reference image, the goal of the triplet loss function is to make the neural network learn that the Euclidean distance between different samples of the same class is less than the Euclidean distance between different classes, so as to distinguish the identity of cattle. However, considering that for traditional Triplet Loss, when the training samples have low discrimination, the recognition accuracy of the trained model is low. Therefore, inspired by the clustering idea, the present invention designs a clustering triplet loss Cure-Triplet Loss, which first clusters the training samples to obtain different clusters, and labels each cluster with a category, thereby improving the discrimination of the samples, and then combines Triplet Loss as the overall loss to improve the recognition accuracy of the algorithm. Brief Description of the Drawings
[0028] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0029] Figure 1 is the structural block diagram of the cattle face recognition method of the present invention;
[0030] Figure 2 is the schematic diagram of multi-scale feature fusion in the cattle face recognition method of the present invention;
[0031] Figure 3 is the schematic diagram of the M_CBAM novel attention mechanism structure;
[0032] Figure 4 is the schematic diagram of Cure-Triplet Loss clustering triplet loss;
[0033] Figure 5 is the result of different identity comparison tests;
[0034] Figure 6 is the visualization result of the cattle face heat map;
[0035] Figure 7 is the result of actual application tests;
[0036] Figure 8 is the flow schematic diagram of a computer device provided by an embodiment of the present invention. Detailed Embodiments
[0037] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0038] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used in the present invention refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0039] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0040] The present invention provides a cow face recognition method based on multi-scale feature interaction, including the following steps:
[0041] Unify the size of the cow face image, upsample the features extracted by the deep feature extraction network, and convert them to the same size as the features extracted by the shallow feature extraction network;
[0042] Construct an attention mechanism M_CBAM at the output ends of the deep feature extraction network and the shallow feature extraction network respectively to form a multi-scale feature fusion network, and realize the automatic adjustment of the channel and spatial attention mechanisms;
[0043] Use the clustering triplet loss Cure-Triplet Loss to assist the network in convergence.
[0044] Embodiment 1
[0045] In this embodiment, aiming at the problem of low accuracy of cattle face recognition caused by the complex and changeable breeding environment in the actual ranch and the non-planar characteristics of the cattle face shape, a deep neural network constructed using the Pytorch framework and a deep learning experimental platform built using the Ubuntu system are mainly used, and a self-built cattle face dataset is used to verify the effectiveness of the algorithm. This dataset was collected at the Outai Ranch in Hohhot, Inner Mongolia in March 2023. Under natural conditions, 1360 adult cattle were randomly selected for video shooting. The video duration of each cow was 6 - 8s, and the resolution was 1920×1080 (pixels). Open CV was used to perform frame splitting on the collected video stream, and the images of the same cow were organized into a folder, and similar images were manually removed. To improve the robustness and sample diversity of the model, the present invention expands the dataset by methods such as rotating by 5° and 10°, mirror inversion, and adding salt-and-pepper noise. After expansion, a total of 54400 cattle face images were obtained, and the training set and validation set were divided according to the principle of 9:1. To verify the generalization ability of the algorithm proposed by the present invention, the present invention added the cattle image data of other ranches as the test set. The test set contains 2482 cattle face pictures of 100 Holstein cows. Similarly, the test set was expanded, and finally 9928 test set pictures were obtained, and the accuracy of the algorithm was tested on the test set using the algorithm of the present invention.
[0046] This algorithm uses accuracy Acc (Accuracy), GFLOPs, and FPS as the performance evaluation indicators of the model. The calculation formula of Acc is as follows:
[0047]
[0048] TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative respectively. Acc refers to the proportion of all correctly predicted positive and negative samples in the total samples, where TP and TN represent true positive and true negative respectively; GFLOPs represents the computational complexity of the algorithm; FPS represents the number of pictures that the model can process per second, where N represents the total number of frames of images processed by the model in a period of time, and T represents the total time consumption of this period.
[0049] This method solves the problem of low accuracy of cattle face recognition caused by the complex and changeable breeding environment in the actual ranch and the non-planar characteristics of the cattle face shape, mainly including two major parts: the cattle face feature interaction stage and the identity recognition stage, as shown in the appendix Figure 1 is the overall architecture diagram of the cattle face recognition method of the present invention. The specific implementation process of the method is as follows:
[0050] A. Cattle face feature interaction stage:
[0051] Step 1: Considering the complex and changeable breeding environment of cattle farms and the non-planar characteristics of cattle face shapes in the real environment, if all training samples use the images finally extracted by the feature extraction network, due to the deepening of the number of network layers in the deep neural network, the original details in the cattle face images will become increasingly blurred or even lost after multiple convolutional operations, resulting in poor recognition effects. Considering that the shallow feature extraction layer has a smaller receptive field, it can enhance the expression ability of detailed features. Therefore, a multi-scale feature fusion module is designed to extract feature maps at different scales in different convolutional layers of the feature extraction network, and enable the extracted shallow features to interact with the deep features to achieve the purpose of accurately recognizing cattle faces. The feature interaction process is shown in Appendix Figure 2 as shown. To ensure the effectiveness of feature interaction, first unify the size of the cattle face images, and upsample the features extracted by the deep feature extraction layer to be the same size as the features extracted by the shallow feature extraction layer.
[0052] Step 2: To enable the model to learn the key features in the feature maps as efficiently as possible and reduce the interference of redundant information, a new attention mechanism M_CBAM is constructed at the output ends of the shallow network and the deep network respectively in the present invention. This attention mechanism designs a parallel structure of channel and spatial attention based on the traditional attention mechanism CBAM, and integrates adaptive weight coefficients, adjusts the weight coefficients of each attention branch according to the important features of the feature maps, and achieves automatic adjustment of the channel and spatial attention mechanisms with dynamic weighting under different cattle faces. The schematic diagram of the M_CBAM structure is shown in Appendix Figure 3 as shown. Considering that in rough set theory, knowledge is represented as a partitioning ability of the universe of discourse, and knowledge itself has granularity. By mapping features to a knowledge base, the importance of different local features for cattle face recognition can be represented by the coarseness or fineness of knowledge granularity. From the granularity principle, the coarser the knowledge granularity, the weaker the knowledge partitioning ability, indicating that the importance of this part of local features is lower, and thus the smaller the allocated weight, and vice versa. Therefore, applying this theory to the network, first map the intermediate features F 1 and F 2 to the knowledge base K=(V,X), where V represents the universe of discourse of the cattle face set, and X represents the set of cattle face attributes of each cow. The knowledge granularity of each part of the intermediate features is as follows:
[0053]
[0054] Define the knowledge partitioning ability as follows:
[0055]
[0056] The Sigmoid function normalizes this ability to obtain the importance:
[0057]
[0058] The definition of the adaptive weighted aggregation formula is as follows:
[0059]
[0060] In formula (4), x i (i = 1, 2,..., I) represents the elements of the I-dimensional feature vector T(x 1 , x 2 , …, x I ) transformed from the intermediate features. Since the domain of the cow face set is determined, the size of the knowledge granularity G of the intermediate features depends on the size of the feature vector T. The larger the modulus value of T, the richer the cow face attributes in this part, and the larger the G value. And it can be seen from formula (5) that the knowledge partitioning ability is inversely proportional to the knowledge granularity. Therefore, when the G value is larger, the N value is larger, indicating that the cow face information is processed at a more blurred level. In this case, the knowledge's ability to capture details is weak, so the ability to distinguish between different objects is reduced. Finally, this knowledge partitioning ability is normalized to the importance degree G. The larger the G value, the higher the importance degree of this part of the features. In formula (7), W represents the weighted modulus value of the image features after adaptive weighted aggregation.
[0061] B. Identity recognition stage:
[0062] Step 1: Considering that for traditional Triplet Loss, when the discrimination of training samples is low, the recognition accuracy of the trained model is low. Therefore, inspired by the clustering idea, the present invention designs a clustering triplet loss Cure-TripletLoss, specifically as Figure 4 shown. This loss first clusters the training samples to obtain different clusters, and assigns class labels to each cluster, thereby improving the discrimination of the samples. Then, combined with Triplet Loss as the overall loss, the recognition accuracy of the algorithm is improved. Its mathematical expression is:
[0063]
[0064] g.rep = p + α * (g.cen - p) (9)
[0065] In formula (8), g 1 and g 2 represent different classes, g.cen represents the center point of the new class. By selecting this point, the structure of the data cluster can be better represented, rather than relying on the centroid or mean; in formula (9), P represents the data item, α represents the shrinkage factor, and its size represents the degree of controlling the contraction of the data points inside the cluster towards the center. A suitable selection of the size of the shrinkage factor can ensure effective improvement of the clustering accuracy in the scenario of non-spherical clusters.
[0066] The above-mentioned Embodiment 1 is a specific application of the method of the present invention. In order to verify the effectiveness of the algorithm in Embodiment 1 of the present invention, ablation experiments are first carried out on a specific data set to verify the effect, and the experimental results are shown in Table 1.
[0067] Table 1 Ablation Experiment
[0068]
[0069]
[0070] As can be seen from Table 1, in terms of the backbone feature extraction network, MobileNetV3 has the advantages of smaller model parameters and computational complexity compared with Inception-ResNetV1. Therefore, using the MobileNetV3 backbone feature extraction network increases the overall FPS of the algorithm by approximately 3.38 frames / s; in terms of the fusion method, since the direct addition fusion method directly fuses the cow face features in equal proportion and does not highlight the local fine features of the cow face; while when using the multi-scale fusion method, by fusing the shallow local fine features and deep global features of the cow face, the model can have good detail feature expression ability. Therefore, the recognition accuracy is improved by 3.54% compared with the former; in terms of the attention mechanism, adding the M_CBAM attention mechanism can dynamically adjust the information importance for different cow face features, improve the extraction ability of the backbone network for important features, and increase the FPS of the algorithm by 0.91 frames / s on the premise that the recognition accuracy of the model is improved by 4.03%; in terms of the loss function, since the clustering triplet loss Cure-Triplet Loss first clusters the training samples to obtain different clusters and labels each cluster, improves the discrimination of the samples, and then combines Triplet Loss as the overall loss to assist the algorithm to converge. Therefore, compared with the traditional Triplet Loss, the recognition accuracy of the former algorithm is improved by 0.23% and the FPS is improved by 2.34 frames / s.
[0071] To demonstrate the effectiveness of the improved algorithm in complex scenarios, cow face images of different identities are selected to input into the network to extract feature vectors, and the Euclidean distances between the vectors are compared to determine whether they are the same cow. The test results are shown in the appendix Figure 5As shown in the figure, it can be seen that after the network adds multi-scale feature fusion, the Euclidean distance of the same bovine face feature vectors decreases, while the Euclidean distance of different bovine face feature vectors increases. After adding the M_CBAM attention mechanism, the Euclidean distance of the same bovine face image feature vectors decreases significantly, and the Euclidean distance of different bovine face image feature vectors increases significantly. After optimizing the loss function, the Euclidean distance of the same bovine face image feature vectors reaches the minimum, and the Euclidean distance of different bovine face image feature vectors reaches the maximum. At the same time, to better illustrate that the algorithm of the present invention can adaptively and selectively express according to important regions for bovine face features, the present invention conducts heatmap visualization on some bovine faces. As shown in Appendix Figure 6 It can be seen from the figure that compared with the other three algorithms, the dark colors of the heatmap of the algorithm of the present invention are concentrated in the facial area of the cow, verifying that the generalization ability of the bovine face recognition algorithm model is strong and the accuracy is relatively high. In summary, from the ablation experiment results, Appendix Figure 5 the Euclidean distance experiment of bovine face feature vectors and Appendix Figure 6 the visualization results of bovine face heatmaps can fully verify the effectiveness of the algorithm of the present invention.
[0072] To further evaluate the advantages and disadvantages of the model, the algorithm of the present invention is compared with FaceNet, MobileFaceNet, SK_ResNet and ECCSA-MFC methods. The comparison results are shown in Table 2.
[0073] Table 2 Comparison results of different models
[0074]
[0075] As can be seen from Table 2, the algorithm of Embodiment 1 of the present invention is significantly higher than other recognition algorithms in terms of recognition accuracy. The algorithm proposed by the present invention adds multi-scale fusion combined with a new attention mechanism, so the recognition accuracy is much higher than other face recognition algorithms. Since the M_CBAM attention mechanism is proposed in the improvement process of the algorithm of the present invention, the GFLOPs is slightly higher. The time complexity of the ECCSA-MFC algorithm is relatively small, but it cannot adapt to the interference of light changes, resulting in a low recognition accuracy and cannot be applied to the actual environment. Since SK_ResNet uses multiple SK_Bottlenecks to fuse multiple perceptual information, the algorithm complexity is lower than that of the algorithm of the present invention, but the recognition accuracy is slightly lower than that of the algorithm of the present invention. Compared with the relatively complex FaceNet model, the algorithm of the present invention has a recognition accuracy 1.59% higher than the FaceNet network while ensuring the light weight of the network. In summary, from the comparison results, the algorithm of the present invention has great advantages in both recognition accuracy and algorithm complexity, and can be better applied to actual complex scenarios.
[0076] Finally, the algorithm of Embodiment 1 of the present invention was deployed to the Jetson AGX Xavier embedded platform for algorithm application testing. The cow face was input into the back-end recognition algorithm for cow face recognition, and the recognition effect is as shown in the appendix Figure 7 As shown. By verifying the recognition results of the four groups of pictures with their ear tag numbers, it can be seen that the cow face recognition results are all correct, verifying the effectiveness of the transplantation of the algorithm model of the present invention to the embedded platform, and achieving a high recognition accuracy while ensuring the real-time requirement.
[0077] In addition to providing a cow face recognition method based on multi-scale feature interaction, the present invention also provides a cow face recognition system corresponding to the cow face recognition method, including:
[0078] Sampling module: used to unify the size of the cow face image, upsample the features extracted by the deep feature extraction network, and convert them into the same size as the features extracted by the shallow feature extraction network;
[0079] Construction module: used to construct the attention mechanism M_CBAM at the output ends of the deep feature extraction network and the shallow feature extraction network respectively, form a multi-scale feature fusion network, and realize the automatic adjustment of the channel and spatial attention mechanisms;
[0080] Convergence module: used to assist the convergence of the network by using the clustering triplet loss Cure-Triplet Loss.
[0081] Figure 8 is a schematic structural diagram of a computer device disclosed by the present invention. Refer to Figure 8 As shown, the computer device includes: an input device 63, an output device 64, a memory 62, and a processor 61; the memory 62 is used to store one or more programs; when the one or more programs are executed by the one or more processors 61, the one or more processors 61 implement the cow face recognition method provided in the above embodiment; wherein the input device 63, the output device 64, the memory 62, and the processor 61 can be connected by a bus or other means, Figure 8 Taking the connection by bus as an example.
[0082] The memory 62, as a computable device-readable and writable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions corresponding to the cow face recognition method described in the embodiments of the present application. The memory 62 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the device, etc. In addition, the memory 62 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 62 can further include a memory remotely set relative to the processor 61, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0083] The input device 63 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the device. The output device 64 can include display devices such as a display screen.
[0084] The processor 61 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 62.
[0085] The embodiments of the present application further provide a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute the cow face recognition method provided in the above embodiments. The storage medium is any of various types of memory devices or storage devices, including: installation media, such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memories or random access memories, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memories, such as flash memories, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium can also include other types of memories or combinations thereof. Additionally, the storage medium can be located in the first computer system in which the program is executed, or can be located in a different second computer system, and the second computer system is connected to the first computer system through a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The storage medium includes two or more storage media that can reside in different locations (such as in different computer systems connected through a network). The storage medium can store program instructions executable by one or more processors (such as specifically implemented as a computer program).
[0086] Finally, it should be noted that: Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but are mainly used to describe the features of specific embodiments of a particular invention. Certain features described in multiple embodiments in this specification can also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combinations. In addition, although the features may function in certain combinations as described above and are even initially claimed as such, one or more features from the claimed combination can in some cases be removed from the combination, and the claimed combination can be directed to a sub-combination or a variation of the sub-combination.
[0087] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0088] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the drawings are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0089] The foregoing is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.
Claims
1. A cow face recognition method based on multi-scale feature interaction, characterized in that: The steps include: Unify the size of the cow face image and upsample the features extracted by the deep feature extraction network to the same size as the features extracted by the shallow feature extraction network; The attention mechanism M_CBAM is constructed at the output end of the deep feature extraction network and the shallow feature extraction network respectively to form a multi-scale feature fusion network and realize automatic adjustment of the channel and spatial attention mechanism; Clustering triplet loss Cure-Triplet Loss is used to assist the network convergence.
2. The cow face recognition method according to claim 1, characterized in that: The features output by the third basic unit among the extracted features are taken as shallow features, and the features output at the end are taken as deep features; the above two features are respectively input into the M_CBAM attention mechanism, and the output features are concatenated.
3. The cow face recognition method according to claim 2, characterized in that: The splicing steps use the following calculation formula: Where: is the dimension of shallow features, is the dimension of the deep feature; H represents the height of the image; W represents the width of the image; C represents the number of channels of the image.
4. The cow face recognition method according to claim 1, characterized in that: The method of constructing the attention mechanism M_CBAM to form a multi-scale feature fusion network is to design a channel and spatial attention parallel structure on the traditional attention mechanism CBAM, and dynamically weighted fuse the feature vectors extracted by the two branches, introduce the knowledge granularity theory to calculate the dynamic weight coefficient, and then dynamically adjust the weight of the two-branch feature vector fusion.
5. The cow face recognition method according to claim 4, characterized in that: The specific calculation formula for constructing the attention mechanism M_CBAM is as follows: ξ=σ(N) (6) W=ξ1·W1+ξ2·W2 (7) In formula (4), G represents the I-dimensional feature vector (x1, x2, …, x i ) is mapped to the knowledge base of K = (V, X), where V represents the domain of the cow face set and X represents the cow face attributes of each cow; In formula (5), N represents the size of knowledge partitioning capability; In formula (6), ξ is the weight coefficient of the eigenvector; In formula (7), ξ1 and ξ2 are the weighting coefficients of the channel feature vector and the spatial feature vector respectively; W1 is the feature vector of the image after the channel attention mechanism, W2 is the feature vector of the image after the spatial attention mechanism, and W is the standard 128-dimensional vector after weighted fusion.
6. The cow face recognition method according to claim 1, characterized in that: The method of using clustering triplet loss Cure-TripletLoss to assist the convergence of the network includes the following steps: First, clustering is performed to obtain different clusters, and each cluster is labeled with a category, and then combined with Triplet Loss as the overall loss. The calculation formula is as follows: g.rep=p+α*(g.cen-p) (9) In formula (8), g1 and g2 represent different classes, g.rep and g.cen represent the center point and representative point of the new class g respectively; in formula (9), p represents the data item in g, and α represents the shrinkage factor.
7. The cow face recognition method according to any one of claims 1 to 6, characterized in that: Using accuracy Acc, GFLOPs and FPS as performance evaluation indicators, the Acc calculation formula is as follows: TP, TN, FP, and FN stand for true positive examples, true negative examples, false positive examples, and false negative examples, respectively. Acc refers to the proportion of all correctly predicted positive and negative samples to the total samples, where TP and TN stand for true positives and true negatives, respectively. GFLOPs represents the computational complexity of the algorithm. FPS represents the number of images that the model can process per second, where N represents the total number of image frames processed by the model in a period of time, and T represents the total time consumed in that period of time.
8. A cow face recognition system, characterized in that: include: Sampling module: used to unify the size of cow face images, upsample the features extracted by the deep feature extraction network, and convert them into the same size as the features extracted by the shallow feature extraction network; Construction module: used to construct the attention mechanism M_CBAM at the output of the deep feature extraction network and the shallow feature extraction network respectively, forming a multi-scale feature fusion network to realize automatic adjustment of channel and spatial attention mechanisms; Convergence module: used to assist the network convergence using clustering triplet loss Cure-Triplet Loss.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more computer programs, and the computer programs can execute the method according to any one of claims 1 to 7.
10. A device comprising a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cattle face recognition method, operation control device, electronic equipment and storage medium
CN114821658A
Cow face detection method for complex scene and detection system thereof
CN117152790A
Scene-adaptive double-branch cattle face efficient recognition method and related equipment
CN117809336A
System and method of machine learning using embedding networks
US20210110275A1
Cited By
A multi-modal information fusion-based cattle identity recognition method
CN122715212A