A sheep facial expression recognition method, system, device and medium
By introducing the SimAM attention mechanism, MobileViT block module and AA2_SPPF module in the YOLOv8n model, the SMEA-YOLOv8n model is formed, which solves the problem of inaccurate expression recognition of sheep faces in the prior art, and achieves higher accuracy in recognition of the faces and accuracy in judging the health status of sheep.
Patent Information
- Application Number
- CN202411645906.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The existing sheep face expression recognition methods cannot accurately extract the expression characteristics of the sheep face, which affects the judgment of the health status of the sheep.
By adding the SimAM attention mechanism and MobileViT block module to the neck network of the YOLOv8n model, and replacing the SPPF module of the backbone network as the AA2_SPPF module, the SMEA-YOLOv8n model is formed to extract the characteristics of the sheep face expression.
It improves the accuracy of sheep facial expression recognition, enhances the model's feature extraction and feature fusion ability in complex environments, can extract global and local features more accurately, and improves the accuracy of judging the healthy status of sheep.
Smart Images

Figure CN119360420B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image recognition, and in particular relates to a sheep facial expression recognition method, system, equipment and medium. Background Art
[0002] In animal pain management and welfare assessment, accurate assessment of pain is crucial to the health of sheep. This can provide timely warnings of sheep in abnormal conditions and reduce the losses caused by the spread of diseases due to delayed detection of problematic sheep. The facial expressions of sheep are considered to be one of the most specific indicators of pain. When sheep are sick, they will show facial pain, shortness of breath, loss of appetite and other characteristics. Therefore, the health of sheep can be assessed by their facial expressions.
[0003] McLennan et al. first proposed a standardized Sheep Pain Facial Expression Scale (SPFES), which scores the possible pain in each area separately according to the degree of ear rotation, nostril shape, eye contraction, etc., and adds the scores of the five areas of eye sockets, cheeks, ears, lips and jaws, and nose. If the score is greater than 1.5 points, the sheep is considered to be in pain. However, this traditional manual detection method is inefficient in large-scale house feeding, has a long existence time and low efficiency, and it is difficult to detect the problem of problematic sheep in time. In 2017, Lu et al. proposed a SPFES sheep facial expression recognition method based on machine learning, which describes the facial expression features of sheep through the histogram of oriented gradients (HOG) and uses support vector machine (SVM) to classify the pain level. However, this method only considers the front part of the sheep and does not consider the facial expressions of sheep in different postures. In 2020, Pessanha et al. extracted multimodal features and combined them with the appearance and geometric feature posture values of the region of interest described by the histogram of oriented gradients (HOG), and finally obtained a binary support vector machine classifier for whether the sheep was in pain. However, the feature extraction of sheep facial expressions in this method relied on manual work, and the ability to capture complex relationships in sheep facial expressions was weak. Compared with traditional machine learning methods, deep learning methods can automatically extract features through multi-level nonlinear transformations and learning, and can better handle complex expression recognition tasks. In order to improve the accuracy of facial pain classification of normal and abnormal sheep, Noor in 2019 achieved the classification of normal and abnormal sheep faces by transfer learning and fine-tuning the VGG16 and ResNet models, but this method did not consider eliminating features that are not related to the painful expressions of sheep faces in images under complex environments.
[0004] In summary, the existing sheep facial expression extraction methods cannot accurately extract the sheep's facial expressions, thus affecting the judgment of the health status of the sheep. Summary of the invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a sheep facial expression recognition method, comprising the following steps:
[0006] Obtain the sheep face image to be identified;
[0007] The SimAM attention mechanism is added after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, and the Mobi leViT block module is added after the third C2f module of the neck network. The SPPF module of the backbone network of the YOLOv8n model is replaced with the AA2_SPPF module to obtain the sheep facial expression recognition model SMEA-YOLOv8n. The sheep facial expression recognition model SMEA-YOLOv8n is trained using a sheep facial expression dataset containing normal sheep facial images and abnormal sheep facial images.
[0008] The sheep face image to be recognized is input into the trained SMEA-YOLOv8n, and the SimAM attention mechanism is used to remove interference factors irrelevant to the sheep facial expression in the sheep face image; the MobileViT block module is used to extract the global and local features of the sheep facial expression feature map after removing the interference factors; the AA2_SPPF module is used to extract the fine-grained features of the sheep facial expression to obtain the sheep facial expression image.
[0009] Preferably, the Mobile block module includes a convolution-based sheep facial expression local feature extraction module, a Transformer-based sheep facial expression global feature extraction module and a feature fusion module; the Transformer-based sheep facial expression global feature extraction module includes an Unfold unit, a Transformer unit and a Fold unit.
[0010] Preferably, the method of using the MobileViT block module to extract the global features and local features of the sheep facial expression feature map after removing interference factors comprises the following steps:
[0011] The sheep facial expression local feature extraction module based on convolution first extracts local features from the sheep facial expression feature map after removing interference factors through an n×n convolution kernel, and then adjusts the number of channels through a 1×1 convolution kernel;
[0012] The Unfold unit expands the feature map after adjusting the number of channels into N non-overlapping flat blocks. The Transformer unit encodes local and global information for the flat blocks. The Fold unit merges the encoded information to obtain a merged feature map.
[0013] The feature fusion module adjusts the number of channels of the merged feature map through a 1×1 convolution kernel, concatenates it with the sheep face expression feature map input into the MobileViT block module, and then outputs it after feature fusion through the convolution layer.
[0014] Preferably, the AA2_SPPF module is obtained by adding two adaptive average pooling layers Adaptive AvgPool2d to the SPPF module of the backbone network of the YOLOv8n model.
[0015] Preferably, the use of the AA2_SPPF module to extract the fine-grained features of the sheep facial expression is specifically as follows: using Adaptive AvgPool2d to calculate the average value of all pixels in the sheep facial expression feature map and perform pooling, removing noise and extreme values of the sheep facial expression feature map, adding global background information and edge information, and extracting the fine-grained features of the sheep facial expression.
[0016] Preferably, it also includes replacing the original CIoU loss function of YOLOv8n with the EfficiCIoU loss function, and using the EfficiCIoU loss function to train SMEA-YOLOv8n.
[0017] Preferably, the sheep facial expression recognition model SMEA-YOLOv8n is trained by using a sheep facial expression data set including normal sheep facial images and abnormal sheep facial images, comprising the following steps:
[0018] The sheep facial expression dataset is expanded and enhanced at a ratio of 1:4 to obtain an enhanced dataset.
[0019] Use the Label Img tool to label the images in the enhanced dataset;
[0020] The labeled data set is divided into training set, validation set and test set in a ratio of 7:2:1;
[0021] SMEA-YOLOv8n is trained using images in the training set, and the model is validated using images in the validation set after training.
[0022] The present invention also provides a sheep facial expression image recognition system, comprising:
[0023] An image acquisition module, used for acquiring a sheep face image to be identified;
[0024] A model building module is used to add SimAM attention mechanisms after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, add a Mobi leViT block module after the third C2f module of the neck network, and replace the SPPF module of the backbone network of the YOLOv8n model with the AA2_SPPF module to obtain the sheep facial expression recognition model SMEA-YOLOv8n; the sheep facial expression recognition model SMEA-YOLOv8n is trained using a sheep facial expression dataset containing normal sheep facial images and abnormal sheep facial images;
[0025] The sheep face expression recognition module is used to input the sheep face image to be recognized into the trained SMEA-YOLOv8n, and use the SimAM attention mechanism to remove interference factors irrelevant to the sheep face expression in the sheep face image; use the Mobi leViT block module to extract the global and local features of the sheep face expression feature map after removing the interference factors; use the AA2_SPPF module to extract the fine-grained features of the sheep face expression to obtain the sheep face expression image.
[0026] The present invention also provides a computer device, comprising a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the sheep facial expression recognition method.
[0027] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the sheep facial expression recognition method.
[0028] The sheep facial expression recognition method provided by the present invention has the following beneficial effects:
[0029] The present invention adds SimAM attention mechanisms after the second C2f module and the fourth C2f module of the neck network respectively, and adds a MobileViT block module after the third C2f module of the neck network. The SimAM attention mechanism estimates the importance of a single neuron through spatial inhibition between neurons, automatically allocates corresponding attention to spatial features and channel features, and eliminates interference factors irrelevant to sheep facial expressions in sheep face images, thereby enhancing the feature extraction capability of the model; the MobileViT block module can better understand the scene of the entire sheep facial expression image, provide richer contextual information, and further enhance the feature extraction and feature fusion capabilities of the model in complex environments such as poor lighting, thereby extracting more accurate global features and local features; finally, the improved AA2_SPPF module is used to obtain the global perspective information of the sheep facial expression and reduce the influence of different scales, so as to achieve the effect of adding some global background information and edge information, thereby helping the network to better extract small and fine-grained sheep facial expression key information capabilities, thereby improving the accuracy of sheep facial expression recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 Flow chart of a sheep facial expression recognition method according to an embodiment of the present invention;
[0032] Figure 2 This is the YOLOv8n model structure diagram;
[0033] Figure 3 SMEA-YOLOv8n model structure diagram for sheep facial expression recognition;
[0034] Figure 4 This is the structure diagram of SimAM’s attention mechanism;
[0035] Figure 5 It is the MobileViTblock module;
[0036] Figure 6 This is a comparison chart before and after the SPPF structure improvement; Figure 6 (a) is the SPPF structure diagram before improvement; Figure 6 (b) is the improved AA2_SPPF structure diagram;
[0037] Figure 7This is a comparison chart of mAP@0.5 after improvement for different sheep facial expressions;
[0038] Figure 8 PR curves before and after the improvement of the YOLOv8n model; Figure 8 (a) is the PR curve before improvement; Figure 8 (b) is the improved PR curve;
[0039] Fig. 9 This is the actual effect diagram of sheep facial expression recognition before and after the improvement of the YOLOv8n model; Fig. 9 (a) shows the recognition effect of the original benchmark model YOLOv8n. Fig. 9 (b) shows the recognition effect of the SMEA-YOLOv8n model of the present invention;
[0040] Fig.10 Comparison chart of mAP@0.5, Precision, and Recall after improvement of different YOLOv8n models. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.
[0042] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the technical solutions of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0043] In addition, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. In the description of the present invention, it should be noted that, unless otherwise clearly specified or limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. In the description of the present invention, unless otherwise specified, "plurality" means two or more, which will not be described in detail here.
[0044] Example
[0045] The present invention provides a method for recognizing sheep facial expressions, specifically: Figure 1 As shown, the following steps are included:
[0046] Step 1: Get the sheep face image to be recognized.
[0047] Step 2: Add SimAM attention mechanisms after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, add a MobileViTblock module after the third C2f module of the neck network, replace the SPPF module of the backbone network of the YOLOv8n model with the AA2_SPPF module, and obtain the sheep facial expression recognition model SMEA-YOLOv8n; train the sheep facial expression recognition model SMEA-YOLOv8n using a sheep facial expression dataset containing normal sheep facial images and abnormal sheep facial images.
[0048] As a single-stage, regressive target detection algorithm, the YOLOv8 model only needs to perform a forward calculation of the convolutional neural network once to complete the detection and category judgment of the target. The detection speed is faster than the two-stage algorithm such as Mask R-CNN that removes the region proposal network. In order to optimize the expression recognition of sheep faces, the YOLOv8n model framework is selected as the benchmark model from many aspects. The YOLOv8n model can be divided into five models: n, s, m, l, and x according to the scale and complexity of its network. In order to facilitate the deployment of the sheep face expression recognition model on mobile devices with limited resources, the YOLOv8n model with the smallest network depth and width is selected as the benchmark model for sheep face facial expression recognition. The network structure consists of four parts: the input part, the backbone network, the neck network, and the head network, such as Figure 2As shown in the figure. The input end of YOLOv8n is mainly responsible for performing image preprocessing operations on the sheep facial expression image to be detected, and then sending it to the backbone network. The backbone network is responsible for extracting feature information from the input sheep facial expression image. The neck network is responsible for integrating the features of the sheep facial expression extracted from the backbone network and performing feature fusion of sheep facial expression feature maps of different scales. Through the bidirectional feature transmission and fusion of top-down FPN and bottom-up PAN, deeper semantic information and positioning information of sheep facial expressions are obtained, and the sheep facial expression features are balanced and enriched. The head network further classifies and regresses the features after the neck network features are fused, so as to output the final expression detection results.
[0049] In order to effectively remove features irrelevant to pain in the facial features of the sheep face and weaken information irrelevant to the facial features of the sheep face, the present invention obtains an improved SMEA-YOLOv8n model by integrating the SimAM attention mechanism, the MobileViTblock module, the EfficiCIoU-Loss module and the AA2_SPPF module (Adaptive AvgPool2d-SPPF) on the basis of YOLOv8n.
[0050] SMEA-YOLOv8n model Figure 3 As shown in the figure, first, in order to enhance the feature fusion ability of the network in the process of sheep facial expression recognition and improve the sensitivity of small targets, the SimAM attention mechanism is added after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, and the MobileViTblock module is added after the third C2f module of the neck network; secondly, in order to improve the training effect and convergence speed of sheep facial expression recognition, the CIoU loss function is replaced by the EfficiCIoU loss function to improve the robustness of the algorithm; finally, in order to better obtain the global perspective information of sheep facial expressions and reduce the impact of different scales, the SPPF module of the backbone network of the YOLOv8n model is replaced by the AA2_SPPF module. The details are as follows:
[0051] (1) SimAM Attention Mechanism
[0052] Under adverse environmental conditions such as poor lighting conditions and occlusion, the extraction of key feature information of sheep facial expressions may be disturbed, resulting in low accuracy in expression recognition. To address the above problems, the present invention introduces a parameter-free SimAM (Simple Parameter-free Attention Module) attention mechanism module based on the baseline model to strengthen the focus on key areas of sheep facial expressions, effectively eliminate interference factors irrelevant to sheep facial expressions, and thus improve the network's feature extraction and feature fusion capabilities. The SimAM attention mechanism with complete 3D weights can estimate the importance of a single neuron through spatial inhibition between neurons without increasing parameters, automatically assign corresponding attention to spatial features and channel features, and enhance the model's feature fusion capabilities. Its principle diagram is shown in the figure. Figure 4 The core idea of SimAM is to identify and emphasize the feature information that contributes more to the final predicted structure by constructing an optimized energy function. This energy function takes into account the similarities and differences between feature maps, thereby dynamically adjusting the weight of each channel so that the model can better focus on the important parts while suppressing the insignificant parts. Formula (1) is the energy function e t :
[0053]
[0054] Among them, w t is the weight, b t is the bias, y is the binary label value, x i are other neurons; i is the index in the spatial dimension; M is the number of neurons; t is the target neuron of the input feature; λ is the weight constant, indicating whether the neuron is an important neuron;
[0055] According to formula (1), the weight w is obtained t and bias b t :
[0056]
[0057] Among them, t is the target neuron of the input feature, μ t and are the mean and variance of all neurons in the channel except t, respectively, and λ is the weight constant.
[0058] According to equations (2) and (3), the minimum energy function is calculated
[0059]
[0060] Among them, formula (4) can be obtained that the energy level is inversely proportional to the difference between neurons, that is, the importance of neurons is
[0061] Combination Figure 5 The fused output image is
[0062]
[0063] Among them, E represents the minimum energy function set of all neurons; X represents the input of the input feature map.
[0064] (2) MobileViT Attention Mechanism
[0065] In order to capture the characteristic information of the sheep facial expression of a small target in a global range, the present invention adds a Mobile block module after the third C2f module of the neck network of YOLOv8 (i.e., the second C2f module connected to the Detect part in the neck network of YOLOv8), so as to better understand the scene of the entire sheep facial expression image, provide richer contextual information, and further enhance the feature extraction and feature fusion capabilities of the model in complex environments such as poor lighting.
[0066] In order to meet the needs of lightweight, low-latency and accurate models for mobile vision tasks, Mehta et al. combined the attention mechanism of MobileNetV2 and Vision Transformer in convolutional neural networks (CNNs) and proposed a lightweight backbone network MobileViT through fewer channels and shallower networks. The network structure of MobileViT mainly includes Conv module, MV2 (MobileNetV2) module, MobileViT block module, global pooling layer and fully connected layer. Among them, the key module of the network, MobileViT block module, replaces the local processing in the standard convolution by introducing deeper global processing, while possessing the local characteristics and global learning ability of convolution. This structure can model the global and local information of sheep facial expressions with fewer parameters, which helps to achieve a lightweight model.
[0067] The MobileViTblock module consists of three parts: Figure 5As shown in the figure: the local feature extraction module of sheep face expression based on convolution, the global feature extraction module of sheep face expression based on Transformer, and the feature fusion module. The input feature map of sheep face expression first extracts local features through a convolution kernel of size n×n, and then adjusts the number of channels through a 1×1 convolution kernel; next, the feature map is processed by the Unfold, Transformer, and Fold units. The Unfold operation expands the feature map into N non-overlapping flat blocks, the Transformer encodes local and global information, and the Fold unit merges the encoded information; finally, the number of channels is adjusted through a 1×1 convolution kernel and spliced with the original input feature map, and then output after feature fusion through the convolution layer, which improves feature modeling and expression, thereby improving model performance.
[0068] (3)AA2_SPPF module
[0069] In order to reduce the computational complexity of the model and speed up the processing speed of the model, YOLOv8 continues to use the SPPF structure of YOLOv5 after version 6.0, using a 5×5 convolution kernel to perform three maximum pooling operations to replace the four convolution kernels of different sizes in the original SPP structure. The SPPF (Spatial Pyramid Pooling Fusion) structure is as follows Figure 6 As shown in (a), the feature extraction and receptive field processing of different scale feature information of sheep facial expressions are mainly performed through spatial pyramid pooling, but this structure only focuses on the edge information of sheep facial expressions and ignores the background information of different sheep facial expression images. Since the change of sheep facial expression is usually a collection of small changes in multiple facial regions, average pooling can capture these subtle changes and form stable global features. Therefore, the present invention improves the SPPF module in the backbone network of YOLOv8 to AA2_SPPF, whose structure is as follows Figure 6 As shown in (b), two AdaptiveAvgPool2d are used on the original basis to perform pooling by calculating the average value of all pixels in the area to smoothly extract features and reduce the influence of noise and extreme values, thereby achieving the effect of adding some global background information and edge information, helping the network to better extract small and fine-grained sheep facial expression feature information.
[0070] (4) EfficiCIoU loss function
[0071] In order to accurately capture the feature information of the sheep's facial expressions that are missed due to partial occlusion of the sheep, improve the positioning accuracy of the bounding box, speed up the convergence of the model, and enhance the learning ability of the image, the EfficiCIoU loss function is replaced by the original CIoU loss function. The loss function is the optimization target in the model training process and directly affects the performance and ability of the model. The loss function of the YOLOv8n model mainly includes classification loss and bounding box regression loss. Among them, bounding box regression is the key step to determine the target positioning performance. The loss function of bounding box regression is mainly divided into l-based n The loss of the norm and the loss based on the IoU series. The traditional IoU loss function only considers the overlap difference between the predicted box and the true box, and cannot handle the situation where the two do not overlap, resulting in the inability to calculate the gradient for optimization. To solve this problem, researchers have proposed a variety of IoU-based loss functions, such as GIoU, DIoU, CIoU, and EIoU. These loss functions not only consider the overlap, but also increase the spatial error and improve the positioning accuracy. The bounding box regression loss of the YOLOv8n model consists of Distribution FocalLoss and CIoU_Loss. CIoU (Complete Intersection over Union) is one of the current best bounding box regression loss functions, which comprehensively considers the three key geometric factors of overlapping area, center point distance, and aspect ratio. CIoU not only relies on the overlap measurement of the predicted box and the real box, but also comprehensively considers the relationship between their intersection and union, so as to more comprehensively evaluate the matching degree of the box; CIoU loss improves the accuracy of the center point positioning of the target detection model by considering the distance between the center point of the predicted box and the real box; by considering the difference between the aspect ratio of the predicted box and the real box, CIoU can more accurately adjust the shape of the predicted box and improve the model's ability to understand the target morphology. The calculation formula of the CIoU loss function is as follows:
[0072]
[0073] In the formula, b represents the center point coordinate of the target prediction box, b gt represents the coordinates of the center point of the label box, c represents the straight-line distance from the lower left corner to the upper right corner of the outer frame composed of the two target detection frames, α is the weight coefficient, v is the parameter for measuring the consistency of the aspect ratio, αv represents the aspect ratio influencing factor, and w gt and h gt Represent the width and height of the real box respectively.
[0074] However, in the bounding box regression process of CIoU, the width and height of the predicted box cannot be simply increased or decreased at the same time, because this does not conform to the change pattern and confidence of the actual width and height. Therefore, when the loss function converges to the ratio between the predicted box and the real box, it sometimes hinders the effective optimization of the model. To solve this problem, EIoU_Loss decomposes the aspect ratio influencing factor αv based on CIoU_Loss to more accurately calculate the width and height of the predicted box and the real box, thereby improving the shortcomings of CIoU_Loss. The calculation formula is as follows:
[0075]
[0076] In the formula, C w and C h Represent the width and height of the smallest bounding box covering the predicted box and the true box, respectively.
[0077] In order to further improve the adjustment speed and regression accuracy of the prediction box, EfficiCIoU_Loss combines the advantages of CIoU_Loss and EIoU_Loss loss functions: first, CIoU_Loss is used to adjust the aspect ratio of the prediction box until it converges to an appropriate range; then EIoU_Loss is used to fine-tune each edge until the width and height values reach the correct ratio. The calculation formula of EfficiCIoU_Loss is as follows:
[0078]
[0079] Step 3: Train SMEA-YOLOv8n, including the following steps:
[0080] (1) Obtain a sheep facial expression dataset containing normal sheep face images and abnormal sheep face images.
[0081] In 2020, NOOR et al. collected a total of 2,350 high-resolution sheep face images that met the SPFES standards from websites such as Pixabay and ImageNet, and then divided the SPFES standards into 1,407 normal sheep face images and 943 abnormal sheep face images to form a data set. The SPFES standard is shown in Table 1. The scores of the eye, ear, and nose areas are added together. If it is greater than 1.5 points, the sheep is considered to be in pain; if it is less than 1.5 points, the sheep is considered to be normal.
[0082] Table 1 Sheep Facial Expression Pain Scale
[0083]
[0084] First, the unclear, repeated and blurred sheep facial expression photos in the NOOR data set of the present invention were manually screened to determine the number of data sets, and then the pictures were renamed, and finally 938 normal sheep facial expression images and 154 abnormal sheep facial expression images were sorted out. In this study, the newly added 751 sheep facial expression images were scored and classified according to the SPFES standard. The scoring results are shown in Table 2. There are 62 sheep facial expression images with a score greater than 1.5 points, and 689 sheep facial expression images with a score less than 1.5 points. The sheep facial expression data set of this study has a total of 1843 images, of which there are 1627 normal sheep facial expression images and 216 abnormal sheep facial expression images.
[0085] Table 2 Scoring results of newly added sheep facial expression images
[0086] Rating score / points 0 1 2 3 4 5 6 7 8 9 10 Quantity / sheet 596 93 12 5 23 8 14 0 0 0 0
[0087] The sheep facial expression dataset of the present invention is divided into two parts: a public sheep facial expression dataset and a self-built sheep facial expression dataset. The public sheep facial expression dataset includes a public sheep facial expression dataset published by NOOR and others, and 426 sheep face images newly searched on public dataset websites such as Pixabay and VEER. Among them, the self-built sheep facial expression dataset comes from a field collection on a ranch in Inner Mongolia Autonomous Region in 2021. By collecting video data of sheep in the sheep house on site, using AfterEffects software to clip the collected video, and then using the OpenCV library in Python to perform original pixel frame capture of the collected video, a total of 325 clear sheep facial expression images were captured. The sheep in the dataset of the present invention include sheep of different breeds and goats, and include sheep of different ages.
[0088] (2) Data preprocessing
[0089] Considering that there are too few valid images of sheep facial expressions at present, in order to speed up the training speed of the sheep facial expression recognition model and improve the quality of the model's recognition of sheep facial expressions, the data set is expanded and enhanced at a ratio of 1:4. Random rotation, adding noise, changing brightness, adding mirrors and other transformations are used to increase the diversity of sheep facial expression images, thereby reducing the overfitting problem and increasing the generalization ability of the model. The data is enhanced to 9215 sheep facial expression images. Some images are enhanced as follows Figure 2 As shown. In order to ensure the accuracy of sheep facial expression recognition, the LabelImg tool is used to label the images in the dataset of the present invention, wherein normal sheep facial expressions are marked as normal and abnormal sheep facial expressions are marked as abnormal. In order to effectively evaluate the performance and generalization ability of the expression recognition model, the dataset of the present invention is divided into a training set, a validation set, and a test set in a ratio of 7:2:1.
[0090] (3) Model training
[0091] The images in the training set are input into SMEA-YOLOv8n, and SMEA-YOLOv8n is trained. The hardware environment of the present invention is configured as an operating system of Linux Ubuntu 18.04LTS, a 12vCPU Intel(R)Xeon(R)Platinum8255C CPU@2.50GHz, and an NVIDIA GeForce RTX 3080GPU with 10GB video memory. The accelerated computing architecture is CUDA11.0, and the deep learning framework used is PyTorch 1.7.0. The deployment environment is Python 3.8. The hyperparameters set the batch size to 8, the training cycle (epochs) to 150, the initial learning rate to 0.01, the momentum to 0.937, the weight decay to 0.0005, and the optimizer selects the stochastic gradient descent SGD algorithm. After training, the performance of SMEA-YOLOv8n is verified using the validation set, and the generalization performance and effect of the model are evaluated using the test set.
[0092] Step 4: Input the sheep face image to be recognized into the trained SMEA-YOLOv8n, and use the SimAM attention mechanism to remove interference factors irrelevant to the sheep facial expression in the sheep face image; use the MobileViT block module to extract the global and local features of the sheep facial expression feature map after removing the interference factors; use the AA2_SPPF module to extract the fine-grained features of the sheep facial expression to obtain the sheep facial expression image.
[0093] Step 5: Effect verification
[0094] In order to accurately and objectively evaluate the performance and efficiency of the sheep facial expression recognition model, the present invention adopts precision P (Precision), recall R (Recall), mean average precision mAP (MeanAverage Precision), parameter amount (Params), and computational effort (GFLOPs) as evaluation indicators.
[0095] The precision P represents the accuracy of the model, which measures the accuracy of the model's prediction of the positive class. The calculation formula is:
[0096]
[0097] Among them, TP stands for true positive examples, which refers to the number of samples correctly predicted by the model as positive; FP stands for false positive examples, which refers to the number of samples that the model incorrectly predicts as positive from negative classes; FN stands for false negative examples, which refers to the number of samples that the model incorrectly predicts as negative from positive classes.
[0098] Recall rate R (Recall) indicates the ratio of the number of samples correctly predicted as positive to the number of samples actually positive, which measures the recognition ability of the model for all samples actually positive. The calculation formula is:
[0099]
[0100] Among them, TP stands for true positive examples, which refers to the number of samples correctly predicted by the model as positive; FN stands for false negative examples, which refers to the number of samples that the model incorrectly predicts as negative.
[0101] The mean average precision (mAP) is used to evaluate the accuracy and comprehensiveness of the model in different categories. mAP can be used to evaluate the average detection accuracy of the model through different IoU thresholds (such as 0.5 and 0.5-0.95); the mean average precision (AP) is the area of the PR curve composed of precision and recall as the horizontal axis and vertical axis respectively, and then the AP values of all categories are averaged to obtain the mean average precision (mAP) of each type of target. The calculation formula is:
[0102]
[0103] The number of parameters (Params) is the total number of training parameters in the model, reflecting the complexity and scale of the model; a larger number of parameters means that the model is more complex and requires a larger storage capacity to store these parameters.
[0104] The amount of computation (GFLOPs) refers to the number of floating-point operations performed by the model per second, reflecting the complexity and resource consumption of the model; a larger amount of computation means that the model is more complex and requires more resources.
[0105] In order to verify that the SimAM attention mechanism can enhance the ability of extracting feature information of sheep facial expressions, the present invention conducts a comparative invention of the attention mechanism, and further analyzes the performance impact of different attention mechanisms on sheep facial expression recognition in complex environments such as poor lighting. The present invention adopts YOLOv8n as the benchmark model. First, the parameter-free SimAM attention mechanism is integrated into the SPPF module in the Backbone network, the C2f module connected to the Detect part in the Neck network, and the C2f module in the Neck network, respectively, denoted as After_SPPF, Concat_Detect, and Neck_C2f. The invention results are shown in Table 3. As can be seen from Table 3, the model after adding SimAM attention to the SPPF module in the Backbone network has the worst effect on sheep facial expression recognition, with mAP and Precision respectively 0.4% and 3.5% lower than the original Baseline; the model after integrating SimAM into the C2f module connected to the Detect part in the Neck network has the best effect on sheep facial expression recognition, with mAP reaching 90.3%, an increase of 2.3% compared with the baseline, and Recall reaching 82.5%, and GFLOPs has not changed, without increasing the complexity of the model. At the same time, in order to ensure a fair comparison of the impact of different attention modules on the model, the invention selects the three attention modules coordAtt (CA), CBAM, and CReToNeXt and integrates them into the same position in the backbone network and neck network of YOLOv8n respectively, and ensures that other structures of the network do not change, so as to achieve the reliability of the invention results. The invention results are shown in Table 4. Compared with the coordAtt (CA), CBAM, and CReToNeXt attention modules, the integration of the SimAM attention mechanism improves the mAP by 1.8%, 7.5%, and 1.2%, respectively, and the computational amount GFLOPs of each sheep face image is smaller than that of other attention modules, which further verifies that the SimAM attention mechanism can improve the network's ability to focus on the features of small targets, thereby accelerating the accuracy of sheep facial expression recognition.
[0106] Table 3 Comparison of SimAM in different locations of YOLOv8n model
[0107] Improved Modules mAP@0.5 / % Precision / % Recall / % GFLOPs baseline 88 83.2 81.9 8.1 After_SPPF 87.6 79.7 88.6 8.1 Concat_Detect 90.3 82.9 82.5 8.1 Neck_C2f 88.7 82.5 89.4 8.1
[0108] Table 4 Comparison of models with other attention mechanisms
[0109] Improved Modules mAP@0.5 / % Precision / % Recall / % GFLOPs baseline 88 83.2 81.9 8.1 CA 88.5 83.9 84.4 8.2 CBAM 82.8 82.1 82.7 8.9 CReToNeXt 89.1 84.2 87.5 9.2 SimAM 90.3 82.9 82.5 8.1
[0110] In order to better learn contextual information in complex environments such as poor lighting, the present invention combines the MobileViT block module with the Transformer structure based on the baseline network plus the SimAM attention mechanism, and captures the global and local information of the sheep's facial expression with fewer parameters, thereby helping to achieve a lightweight model. In order to verify the effect of integrating the MobileViT attention mechanism, the present invention replaces the SimAM module in the C2f module connected to the Detect part with the MobileViTblock module on the basis of the YOLOv8_SimAM model, and obtains C2f_1, C2f_2, and C2f_3. The results of the invention are shown in Table 5. C2f_2 has the best expression recognition effect. Although the calculation amount GFLOPs has increased, the mAP, Precision, and Recall reach 90.6%, 83.9%, and 85.8%, respectively, which are 0.3%, 1%, and 3.3% higher than YOLOv8_SimAM.
[0111] Table 5 Comparison of different positions of MobileViTblock module added to YOLOv8n model
[0112] Improved Modules mAP@0.5 / % Precision / % Recall / % GFLOPs YOLOv8_SimAM 90.3 82.9 82.5 8.1 C2f_1 88.7 82.4 88.6 11.1 C2f_2 90.6 83.9 85.8 11.1 C2f_3 89.5 81.6 87 11.1
[0113] In order to verify the effectiveness of the EfficiCIoU loss function in accelerating the convergence speed of the model, the present invention conducts a loss function improvement comparison invention. Based on the original benchmark model, the SimAM module and the MobileViT block module are integrated, and the EfficiCIoU loss function is compared with CIoU, EIoU, XIoU, WIoU, and SIoU respectively. It can be seen from Table 6 that the model adopts the EfficiCIoU loss function in mAP, Precision, and Recall compared with other loss functions such as CIoU. Compared with the CIoU loss function, the mAP is improved by 1.5%, the Precision remains basically stable, and the Recall is improved by 3.7%. The sheep face can be more accurately identified and a high-confidence sheep face expression result is given.
[0114] Table 6 Comparison of loss functions
[0115] Improved Modules mAP@0.5 / % Precision / % Recall / % CIo 90.6 83.9 85.8 EIU 88.7 82.6 85.4 XI X 89.8 83.2 84.1 WU 85.3 79.1 84.8 ikB 85.8 80.3 79.4 EfficiCIo 92.1 83.5 89.5
[0116] In order to further verify the effectiveness of the AA2_SPPF improved module, the present invention carried out an SPPF improvement comparison invention, and compared it with the AMSPPF module which respectively integrated the global average pooling layer and the global maximum pooling layer. The results are shown in Table 7 below. The AA2_SPPF module is 4.6%, 7.6% and 2.5% higher than the AMSPPF module in mAP, Precision and Recall respectively; the AA2_SPPF module is 0.4%, 2.5% and 1.5% higher than the original SPPF module in mAP, Precision and Recall respectively, which further illustrates that the present invention can help the network better extract small and fine-grained sheep facial expression feature information.
[0117] Table 7 Performance comparison of different SPPF improved modules
[0118] Module mAP@0.5 / % Precision / % Recall / % SPPF 92.1 83.5 89.5 AMSPPF 87.9 78.4 88.5 AA2_SPPF 92.5 86 91
[0119] In order to better verify the impact of the improved algorithm on the overall model, the present invention performs ablation and successively improves the YOLOv8n model as follows: the SimAM attention mechanism is integrated with the three C2f modules in the neck network connected to the Detect part (i.e., the second, third, and fourth C2f modules in the neck network of YOLOv8); the SimAM attention mechanism in the third C2f module in the neck network is replaced by the MobileViTblock module; the loss function is replaced by EfficiCIoU; the SPPF module in the backbone network is improved to AA2_SPPF. The ablation results are shown in Table 8. Indicates that the corresponding modules above are not added to the YOLOv8n model. For example, the first row indicates that no modules are added to the YOLOv8n benchmark model. The √ in the table indicates that the corresponding modules above are added to the YOLOv8n model. For example, the first column of the second row indicates that SimAM is added first. It can be seen from Table 8 that the mean average accuracy of the SMEA-YOLOv8n model of the present invention for sheep facial expression recognition reaches 92.5%, which is 4.5% higher than that of YOLOv8n, and the recognition effect is greatly improved. In addition, the recall rate and precision rate are increased by 9.1% and 2.8% respectively, reducing the false detection of sheep facial expressions. Table 9, Figure 7The model's improvement in the recognition of facial expressions of sheep faces is more comprehensive. After the neck network is fused with the SimAM module, the mAP of normal sheep facial expressions is increased by 2.5%, and the mAP of abnormal sheep facial expressions is increased by 2.1%. After adding the MobileViT block module, the mAP of normal sheep facial expressions is increased by 3.3%, and the mAP of abnormal sheep facial expressions is increased by 3.3%. After using EfficiCIoU for the bounding box regression loss function, the mAP of normal sheep facial expressions is increased by 3.6%, and the mAP of abnormal sheep facial expressions is increased by 4.7%. For the backbone network, the SPPF is improved to AA2_SPPF, which increases the mAP of normal sheep facial expressions by 3.7%, and the mAP of abnormal sheep facial expressions by 5.3%.
[0120] Table 8 Ablation invention
[0121]
[0122] Table 9 Comparison of sheep facial expression recognition
[0123]
[0124] In order to more intuitively demonstrate the improvement effect of the model of the present invention, the PR curves of the YOLOv8n model before and after improvement are compared, as shown in the figure: Figure 8 As shown. Figure 8 It can be seen that with the increase of the number of iterations of the two algorithms, the final results of Precision, Recall, and mAP of the improved model SMEA-YOLOv8n are better than those of the baseline YOLOv8n, which shows that the improved model can effectively improve the accuracy of sheep facial expression recognition and reduce the false detection rate. In order to verify the actual effect of the improved algorithm of the present invention, YOLOv8n and SMEA-YOLOv8n are visually compared. The sheep facial expression recognition results are shown in Figure 2. Fig. 9 As shown, Fig. 9 (a) shows the recognition effect of the original benchmark model YOLOv8n. Fig. 9 Figure (b) shows the recognition effect of the SMEA-YOLOv8n model of the present invention. Fig. 9 It can be seen that the YOLOv8n model has problems with incorrect recognition of sheep facial expression categories and low accuracy, while SMEA-YOLOv8n can recognize sheep facial expressions more accurately.
[0125] In addition, the present invention also compares the mAP@0.5 of different YOLOv8n models after improvement, as shown in the following figure. Fig.10 As shown, from Fig.10It can be seen that after the neck network is fused with the SimAM module (YOLOv8n+A), the mAP for normal sheep facial expressions is increased from 0.94 to 0.965, and the mAP for abnormal sheep facial expressions is increased from 0.82 to 0.841, compared with the baseline model YOLOv8n. This is because in the process of expression recognition, the extraction of sheep facial expression feature information is easily affected by the environment such as occlusion. The SimAM attention mechanism can efficiently focus on important expression feature areas, thereby improving the accuracy of expression recognition. After adding the MobileViTblock module (YOLOv8n+B), compared with YOLOv8n+A, the mAP for normal sheep facial expressions is increased from 0.965 to 0.973, and the mAP for abnormal sheep facial expressions is basically unchanged. This is because the MobileViT block module can extract the global information of sheep facial expressions through the self-attention mechanism in the Transformer and fuse the local information obtained by the convolution block to improve the recognition accuracy of the model. After replacing the loss function with EfficiCIoU (YOLOv8n+C), compared with YOLOv8n+B, the mAP for normal sheep facial expressions increased from 0.973 to 0.976, and the mAP for abnormal sheep facial expressions increased from 0.839 to 0.867. This is because the EfficiCIoU loss function combines the shortcomings of CIoU and EIoU and improves on them, further improving the adjustment speed and regression accuracy of the prediction box, thereby improving the recognition ability. After the backbone network improved SPPF to AA2_SPPF (YOLOv8n+D), compared with YOLOv8n+C, the mAP for normal sheep facial expressions increased from 0.976 to 0.977, and the mAP for abnormal sheep facial expressions increased from 0.867 to 0.873. This is because in AA2_SPPF, by adding two global average pooling layers, the key small and fine-grained sheep facial expression information can be extracted, thereby improving the accuracy.
[0126] The accuracy of the SMEA-YOLOv8n model increased from 0.913 to 0.951 in the normal (painless) sheep facial expression category, and from 0.75 to 0.769 in the abnormal (painful) sheep facial expression category, indicating that the model more accurately recognizes the normal facial expressions of sheep faces. The overall mAP increased from 0.88 to 0.925, indicating that the model has good recognition ability for the two categories of sheep facial expressions. In general, the SMEA-YOLOv8n model has significantly improved in sheep facial expression recognition. The improvement of the model in precision, mAP, and Recall shows that it can more accurately and comprehensively recognize the facial expressions of sheep, which can timely warn sheep in abnormal states, reduce more diseases caused by delayed detection of problematic sheep, and is conducive to welfare-oriented smart farming.
[0127] In summary, the present invention integrates the SimAM attention mechanism and the MobileViTblock module into the C2f module of the neck network. The SimAM attention mechanism estimates the importance of a single neuron through spatial inhibition between neurons, and automatically allocates corresponding attention to spatial features and channel features, thereby enhancing the feature extraction ability of the model; the MobileViT block module can better understand the scene of the entire sheep facial expression image, provide richer contextual information, and further enhance the feature extraction and feature fusion capabilities of the model in complex environments such as poor lighting, thereby extracting more accurate global features and local features; secondly, the loss function is replaced by the EfficiCIoU loss function, and the aspect ratio of the prediction box is adjusted by CIoU_Loss to converge to an appropriate range and each edge is finely adjusted by EIoU_Loss until the width and height values reach the correct ratio, thereby further improving the adjustment speed and regression accuracy of the prediction box, thereby improving the training effect of sheep facial expression recognition. Finally, through the improved AA2_SPPF module, the global perspective information of sheep facial expressions is obtained and the influence of different scales is reduced, so as to achieve the effect of adding some global background information and edge information, thereby helping the network to better extract small and fine-grained key information of sheep facial expressions. The experimental results show that compared with the baseline model YOLOv8n, the improved model of the present invention improves mAP, Precision, and Recal l by 4.5%, 2.8%, and 9.1%, respectively. Among them, the mAP of normal sheep facial expressions is improved by 3.7%, and the mAP of abnormal (painful) sheep facial expressions is improved by 5.3%.
[0128] The present invention also provides a sheep face expression image recognition system, comprising an image acquisition module, a model construction module and a sheep face expression recognition module. The image acquisition module is used to acquire a sheep face image to be recognized. The model construction module is used to add SimAM attention mechanisms after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, respectively, and add a MobileViTblock module after the third C2f module of the neck network, and replace the SPPF module of the backbone network of the YOLOv8n model with an AA2_SPPF module to obtain a sheep face expression recognition model SMEA-YOLOv8n; the sheep face expression recognition module is used to input the sheep face image to be recognized into SMEA-YOLOv8n, and use the SimAM attention mechanism to eliminate interference factors that are not related to the sheep face expression in the sheep face image; use the MobileViTblock module to extract the global features and local features of the sheep face expression feature map after removing the interference factors; use the AA2_SPPF module to extract the fine-grained features of the sheep face expression, and obtain the sheep face expression image.
[0129] The present invention also provides a computer device, comprising a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute a sheep facial expression recognition method.
[0130] The present invention also provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded by a processor to execute the sheep facial expression recognition method.
[0131] The above embodiments are only preferred specific implementation methods of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of the technical solutions that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belong to the protection scope of the present invention.
Claims
1. A sheep facial expression recognition method, characterized in that: The steps include: Obtain the sheep face image to be identified; The SimAM attention mechanism is added after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, and the MobileViT block module is added after the third C2f module of the neck network. The SPPF module of the backbone network of the YOLOv8n model is replaced with the AA2_SPPF module to obtain the sheep facial expression recognition model SMEA-YOLOv8n. The sheep facial expression recognition model SMEA-YOLOv8n is trained using a sheep facial expression dataset containing normal sheep facial images and abnormal sheep facial images. The sheep face image to be recognized is input into the trained SMEA-YOLOv8n, and the SimAM attention mechanism is used to remove interference factors irrelevant to the sheep face expression in the sheep face image; the MobileViT block module is used to extract the global and local features of the sheep face expression feature map after removing the interference factors; Use AA2_SPPF module to extract fine-grained features of sheep facial expressions and obtain sheep facial expression images; The AA2_SPPF module is obtained by adding two adaptive average pooling layers Adaptive AvgPool2d to the SPPF module of the backbone network of the YOLOv8n model; The AA2_SPPF module is used to extract the fine-grained features of the sheep's facial expressions, specifically: AdaptiveAvgPool2d is used to calculate the average value of all pixels in the sheep's facial expression feature map and perform pooling, the noise and extreme values of the sheep's facial expression feature map are removed, the global background information and edge information are added, and the fine-grained features of the sheep's facial expressions are extracted.
2. The sheep facial expression recognition method according to claim 1, characterized in that: The MobileViT block module includes a convolution-based sheep facial expression local feature extraction module, a Transformer-based sheep facial expression global feature extraction module and a feature fusion module; the Transformer-based sheep facial expression global feature extraction module includes an Unfold unit, a Transformer unit and a Fold unit.
3. The sheep facial expression recognition method according to claim 2, characterized in that: The method of using the MobileViT block module to extract the global features and local features of the sheep facial expression feature map after removing interference factors includes the following steps: The sheep facial expression local feature extraction module based on convolution first extracts local features from the sheep facial expression feature map after removing interference factors through an n×n convolution kernel, and then adjusts the number of channels through a 1×1 convolution kernel; The Unfold unit expands the feature map after adjusting the number of channels into N non-overlapping flat blocks. The Transformer unit encodes local and global information for the flat blocks. The Fold unit merges the encoded information to obtain a merged feature map. The feature fusion module adjusts the number of channels of the merged feature map through a 1×1 convolution kernel, and concatenates it with the sheep face expression feature map input into the MobileViTblock module, and then outputs it after feature fusion through the convolution layer.
4. The sheep facial expression recognition method according to claim 1, characterized in that: It also includes replacing the original CIoU loss function of YOLOv8n with the EfficiCIoU loss function, and using the EfficiCIoU loss function to train SMEA-YOLOv8n.
5. The sheep facial expression recognition method according to claim 4, characterized in that: The method of training the sheep facial expression recognition model SMEA-YOLOv8n by using a sheep facial expression data set including normal sheep facial images and abnormal sheep facial images comprises the following steps: The sheep facial expression dataset is expanded and enhanced at a ratio of 1:4 to obtain an enhanced dataset. Use the LabelImg tool to label the images in the augmented dataset; The labeled data set is divided into training set, validation set and test set in a ratio of 7:2:1; SMEA-YOLOv8n is trained using images in the training set, and the model is validated using images in the validation set after training.
6. A sheep facial expression image recognition system, characterized in that: include: An image acquisition module, used for acquiring a sheep face image to be identified; A model building module is used to add SimAM attention mechanisms after the second C2f module and the fourth C2f module of the neck network of the YOLOv8n model, add a MobileViT block module after the third C2f module of the neck network, replace the SPPF module of the backbone network of the YOLOv8n model with the AA2_SPPF module, and obtain the sheep facial expression recognition model SMEA-YOLOv8n; the sheep facial expression recognition model SMEA-YOLOv8n is trained by a sheep facial expression dataset containing normal sheep facial images and abnormal sheep facial images; wherein the AA2_SPPF module is obtained by adding two adaptive average pooling layers Adaptive AvgPool2d to the SPPF module of the backbone network of the YOLOv8n model; A sheep facial expression recognition module is used to input a sheep facial image to be recognized into the trained SMEA-YOLOv8n, and use the SimAM attention mechanism to remove interference factors that are not related to the sheep facial expression in the sheep facial image; use the MobileViT block module to extract global features and local features of the sheep facial expression feature map after removing the interference factors; use the AA2_SPPF module to extract fine-grained features of the sheep facial expression to obtain a sheep facial expression image; wherein, the use of the AA2_SPPF module to extract the fine-grained features of the sheep facial expression is specifically as follows: using Adaptive AvgPool2d to calculate the average value of all pixels in the sheep facial expression feature map and perform pooling, removing noise and extreme values of the sheep facial expression feature map, adding global background information and edge information, and extracting fine-grained features of the sheep facial expression.
7. A computer device, characterized in that: It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the sheep facial expression recognition method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the sheep facial expression recognition method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Driver safety belt detection method based on deep learning
CN117333852A
Sheep face image detection method based on YOLOv8-Swin-Transform-BiFPN neural network
CN118397657A