Holstein cow behavior identification method based on CAMILLA-YOLOv8n algorithm
By using the CAMILLA-YOLOv8n algorithm, combined with Coordinate Attention, MLLAttention, and Shape-IoU loss function, the efficiency and accuracy issues of Holstein cow behavior recognition were solved, achieving efficient recognition in complex environments.
Patent Information
- Application Number
- CN202411251433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-07
- Publication Date
- 2026-03-10
AI Technical Summary
Existing Holstein cow behavior recognition technologies are inefficient and inaccurate when identifying multiple behaviors, especially in complex environments where rapid and accurate identification is difficult. Furthermore, traditional methods may affect animal health and increase operating costs.
The CAMILLA-YOLOv8n algorithm is adopted. By integrating the Coordinate Attention mechanism into the C2f module, introducing the MLLA Attention mechanism into the P3, P4, and P5 layers of the Neck, improving the SPPF module to SPPF-GPE, and introducing the Shape-IoU loss function, the model's ability to recognize the behavior of Holstein cows is enhanced.
It improves the average accuracy and recognition speed of Holstein cow behavior recognition, meets the needs of practical application scenarios, and achieves accurate and rapid recognition in complex environments.
Smart Images

Figure CN121640560A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and Holstein cow behavior recognition, and particularly relates to a method for Holstein cow behavior recognition based on the CAMILLA-YOLOv8n algorithm. Background Technology
[0002] In recent years, with the rapid development of computer vision and deep learning technologies, the livestock industry has adopted deep learning to solve problems such as individual animal identification, animal movement tracking, body part recognition, and species classification. There is growing interest in using these models to examine the relationship between livestock behavior and related health issues. In animal husbandry, particularly in the field of Holstein dairy cow behavior recognition, the application of this technology has improved the production and management capabilities of Holstein dairy farming, enhanced disease prevention and control, and reduced operating costs.
[0003] As a major livestock country, China has long ranked among the world's top in terms of total livestock output, and the proportion of its livestock industry continues to rise. Against this backdrop, farm modernization has become an inevitable choice for livestock development. In livestock farming, studying animal behavior is crucial for understanding how animals interpret and respond to their environment. This allows us to use effective technologies to improve their health and welfare on farms. The daily behaviors of Holstein dairy cows, such as standing, eating, and lying down, as well as abnormal behaviors such as estrus, licking, and fighting, are all closely related to their physiological health. For example, normal standing, eating, and lying down behaviors are generally considered indicators of comfort in Holstein dairy cows, while abnormal fighting and licking behaviors may indicate health problems or environmental discomfort. Estrus behavior, as a significant external manifestation, is often difficult to detect in a timely manner due to its low frequency, short duration, subtlety, and lack of fixed location. If Holstein cows in estrus are not detected and signaled in time, farms must wait for the next estrus cycle, resulting in lost breeding opportunities, increased non-pregnant days, longer calving intervals, lower conception rates, and reduced milk production, thus impacting the profitability of Holstein dairy farming. Therefore, utilizing intelligent technologies for real-time monitoring and analysis of Holstein cow behavior in practical applications, and efficiently and accurately identifying such behavior, is crucial for timely understanding of Holstein cow health and improving farm economic efficiency.
[0004] Currently, methods for monitoring Holstein dairy cow behavior mainly fall into three categories: traditional worker observation, contact sensor detection, and non-contact image recognition detection. Traditionally, monitoring Holstein dairy cow behavior relies primarily on worker observation and recording, which is not only time-consuming and inefficient but also highly subjective. In densely populated Holstein dairy farms, real-time observation of every activity of each cow is impractical, making continuous, 24 / 7 large-scale monitoring difficult. Sensor-based monitoring of Holstein dairy cow behavior also has drawbacks. Contact sensor detection typically requires animals to wear sensors in specific locations, such as collars or ankles, to collect movement and physiological data for behavior identification. However, prolonged sensor use can cause discomfort, leading to stress responses and interfering with normal behavior. Furthermore, contact devices are often limited in function, mostly monitoring only one physiological indicator or movement behavior. Monitoring multiple behaviors requires multiple sensors with different functions, which may not be practical in actual farming.
[0005] Despite significant advancements in current technology, the field of Holstein cow behavior recognition still faces numerous challenges. Fuentes et al. proposed a spatiotemporal information-based recognition method using YOLOv3 for frame-level detection, context-based feature extraction, and combining 3D-CNN and optical flow to capture temporal features. Tested on 15 different cow behavior datasets recorded on Korean farms, the method achieved a mean accuracy (mAP) of 85.6%, effectively identifying individual cows, groups, and specific movements. However, the system exhibited some difficulty in recognizing smaller, specific movements and in handling changes in background and lighting conditions (especially at night). Wang et al. used an improved YOLOv8n model, E-YOLO, which utilizes normalized Wasserstein distance loss, a context-enhanced module, and a triple attention module to detect estrus behavior in cows. This model achieved an average accuracy of 93.90% for estrus detection, an average accuracy of 95.70% for mounting (APmounting), and an F1 score of 93.74%, demonstrating high efficiency and accuracy in real-world dairy farm environments. However, a limitation of the E-YOLO model is that it only focuses on detecting estrus and does not address other important cow behavior detection problems. Wang et al. proposed an efficient 3D CNN (E3D) algorithm based on a SandGlass-3D module combining 3D convolution and Dwise (depthwise separable convolution) specifically for recognizing basic cow movement behaviors. This model combines the SandGlass-3D module with an efficient channel attention (ECA) mechanism to effectively process spatiotemporal information in videos, quickly and accurately recognizing cow behaviors in natural environments with an accuracy of 98.17%. However, the model still faces challenges in terms of speed and efficiency on mobile devices with limited hardware resources. Yu et al. proposed a DRN-YOLO method based on DenseResNet for real-time monitoring of cow feeding behavior. By replacing the CSPDarknet backbone network with a self-designed DRNet backbone network, and based on the YOLOv4 algorithm, they utilized multi-scale and spatial pyramid pooling (SPP) structures to enhance the interaction of scale semantic features. Compared to YOLOv4, DRN-YOLO improved accuracy, recall, and mAP by 1.70%, 1.82%, and 0.97%, respectively. However, it faces challenges in recognizing low accuracy and insufficient feature extraction in complex farming environments. Bai et al. proposed an improved YOLOv3-based model, GC_Res2 YOLOv3, for recognizing dairy cow behavior. It combines the Res2 network structure and the Global Context Block to enhance the model's multi-scale perception capabilities and localization accuracy in dense scenes.The improved model achieved accuracies of 90.6%, 91.7%, 80.7%, and 98.5% in detecting four behaviors in cows: standing, lying down, walking, and mating, respectively. The overall average accuracy was 90.4%. However, despite demonstrating good generalization ability across various scenarios, it still struggled to distinguish between walking and standing behaviors in cows. Wang et al. proposed a lightweight cow mounting behavior recognition system based on an improved YOLOv5s. This system utilizes the concept of EfficientNetV2 to design a lightweight background network and introduces an attention mechanism, inverted residual structure, and depthwise separable convolutions. The model achieved an inference speed of up to 333.3 fps, with an inference time of 4.1 ms per image, and a mAP of 87.7%, which is 2.1% higher than the mAP of YOLOv5s. However, the lightweight network may result in a decrease in feature extraction capabilities, potentially affecting the model's accuracy. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and propose a method for recognizing the behavior of Holstein dairy cows based on the CAMILLA-YOLOv8n algorithm.
[0007] The problem solved by this invention is achieved by the following technical solution:
[0008] A method for recognizing the behavior of Holstein dairy cows based on the CAMILLA-YOLOv8n algorithm, characterized by the following steps:
[0009] Two Tiandy cameras (TD-H234S) were installed on the Holstein dairy farm and mounted on the support wall of the double-sloped cowshed beams for real-time data acquisition and processing. The cameras were fixed at a height of about 4.5m and tilted downwards at an angle of about 30°, allowing them to capture the entire area where the cows were active. Video data of the daily behavior of Holstein dairy cows was collected over a period of about 110 days.
[0010] We observed video data of daily behavior of Holstein dairy cows and selected six behaviors, including feeding, standing, lying down, licking, estrus, and fighting, as well as the empty state of the cow bed, as research subjects. These behaviors are closely related to the health evaluation of Holstein dairy cows.
[0011] Based on meticulous expert classification, 60 independent video clips containing Holstein cow behavior were manually selected. Each clip contained one or more complete instances of Holstein cow behavior to capture behavioral patterns of Holstein cows at different times and under different environments. Each clip ranged in length from 10 to 25 seconds, with a video resolution of 1280 pixels × 720 pixels and a frame rate of 30 fps. According to the specific behavioral characteristics of the Holstein cows, the CVAT data annotation tool was used to complete the annotation, and the results were exported in YOLO dataset format. The final dataset was labeled into 7 classes. To avoid excessive similarity between adjacent frames, differences between consecutive frames were removed, and redundant images were deleted. There was no data overlap between different sample sets. A total of 2418 images were collected, with a total of 23073 bounding boxes annotated. Of these, 1935 images were used for training and 483 for validation, representing 80% and 20% respectively for model training and validation.
[0012] Considering the challenges of varying scales, uneven lighting, and excessive occlusion in real-world Holstein dairy farm environments, this study employs various techniques to enrich the multi-scale behavioral data and background information of Holstein dairy cows in a labeled dataset of daily behavior. These techniques include Mosaic data augmentation, rotation transformation, irregular cropping, brightness variation, noise addition, horizontal and vertical flipping, and random occlusion transformation. Compared to single-image augmentation methods, Mosaic-enhanced images contain richer scene and multi-scale information, increasing the diversity of the dataset. Image flipping simulates Holstein dairy cow behavior samples from different perspectives. Image rotation transformation can simulate Holstein dairy cow behavior in different directions. Furthermore, adjusting image brightness can simulate different lighting conditions from daylight to night. Adding noise is used to simulate image quality under low light and inclement weather conditions. Finally, random cropping enhances the model's ability to recognize local behavioral features of Holstein dairy cows, which is crucial for understanding complex behavioral patterns. Training the model using the data-augmented dataset indirectly increases the number of samples, accelerates model convergence, and improves the model's generalization ability.
[0013] To address the challenges of increased background complexity and a larger number of Holstein cows resulting from Mosaic fusion of datasets, the CAMLLA-YOLOv8n model integrates a Coordinate Attention mechanism into the C2f module, forming the C2f-CA module. This module more accurately identifies and understands the spatial relationships between different Holstein cow locations, while also improving the model's sensitivity to key regions and its ability to filter background interference.
[0014] To address the challenge of recognizing multi-scale and multi-behavioral features of Holstein dairy cows, CAMLLA-YOLOv8n introduces the MLLA attention mechanism into the P3, P4, and P5 layers of the YOLOv8n model Neck, thus solving the challenge of recognizing Holstein dairy cow behavior due to large scale variations.
[0015] To address multi-scale Holstein cow behavior data, CAMLLA-YOLOv8n innovatively improves the SPPF module to form the SPPF-GPE module. By combining global average pooling and global max pooling, it optimizes small target recognition and enhances the model's ability to capture environmental changes and key aspects of Holstein cow behavior.
[0016] To address the limitations of traditional IoU loss in Holstein cow detection, CAMLLA-YOLOv8n introduces Shape-IoU, which focuses on the shape and scale features of the Bounding Box, improving the matching degree between the Prediction Box and the Ground Truth Box, and enhancing detection accuracy.
[0017] Advantages and positive effects of the present invention:
[0018] 1. This invention involves installing two Tiandy cameras (TD-H234S) on the support wall of the double-sloped cowshed beams at a Holstein dairy farm for real-time data acquisition and processing. The data collected video data of daily behavior of Holstein dairy cows over a period of approximately 110 days was used to construct a Holstein dairy cow behavior dataset.
[0019] 2. This invention improves the YOLOv8n model by: integrating a Coordinate Attention mechanism into the C2f module to form a C2f-CA module, which more accurately identifies and understands the spatial relationships between different Holstein cow locations, while improving the model's sensitivity to key regions and its ability to filter background interference; introducing an MLLA attention mechanism into the P3, P4, and P5 layers of the YOLOv8n model's Neck layer to address the challenge of Holstein cow behavior recognition due to large scale variations; improving the SPPF module to form an SPPF-GPE module, which optimizes small target recognition by combining global average pooling and global max pooling; and introducing Shape-IoU to focus on the shape and scale features of the Bounding Box, improving the matching degree between the Prediction Box and the Ground Truth Box, and enhancing detection accuracy. The CAMLLA-YOLOv8n model of this invention can effectively improve the average accuracy of Holstein cow behavior recognition in complex environments, meeting the needs for accurate and rapid recognition of Holstein cow behavior in practical application scenarios. Attached Figure Description
[0020] Figure 1 (a) is a monitoring video scene of the Holstein dairy cow behavior dataset collected in this invention. Figure 1 (b) Example of data segmentation for Holstein dairy cow behavior.
[0021] Figure 2 This invention provides the category distribution and annotation information for the Holstein cow behavior dataset.
[0022] Figure 3 These are sample images from a challenging dataset in the Holstein cow behavior dataset constructed for this invention.
[0023] Figure 4 This invention provides an enhanced visualization of the Holstein cow behavior dataset constructed for this purpose.
[0024] Figure 5 This is a diagram of the CAMLLA-YOLOv8n Holstein cow behavior recognition network architecture of the present invention.
[0025] Figure 6 The diagram shows the unmodified YOLOv8n network module structure for reference in this invention.
[0026] Figure 7 This is a diagram of the CA attention module of the present invention.
[0027] Figure 8 This invention provides a detailed architecture of CAMLLA-YOLOv8n and the MLLAttention attention mechanism.
[0028] Figure 9 This is a structural diagram of the SPPF-GPE of the present invention.
[0029] Figure 10 This invention provides a comprehensive performance comparison of the seven YOLO detection algorithms.
[0030] Figure 11 This is a comparative analysis chart of the accuracy, recall, and average accuracy of the seven YOLO detection algorithms of this invention.
[0031] Figure 12 The training and validation loss diagrams for the seven YOLO detection algorithms of this invention are shown.
[0032] Figure 13 Visualization of the ablation experimental results of different optimization modules of the CAMLLA-YOLOv8n algorithm of this invention in terms of Precision, Recall, mAP50 and mAP@0.5-0.95.
[0033] Figure 14This is a heatmap comparison of YOLOv8n and CAMLLA-YOLOv8n of this invention. The comparison is shown in three scenarios.
[0034] Figure 15 Visualization of the detection results of CAMLLA-YOLOv8n in this invention 1.
[0035] Figure 16 Visualization of the CAMLLA-YOLOv8n detection results of the present invention 2. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings. The described embodiments are merely some, not all, embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention.
[0037] Two Tiandy cameras (TD-H234S) were installed on the Holstein dairy farm, mounted on the support wall of the double-sloped cowshed beams, for real-time data acquisition and processing. The cameras were fixed at a height of about 4.5m and tilted downwards at an angle of about 30°, allowing them to capture the entire area where the Holstein dairy cows were active. Video data of the daily behavior of the Holstein dairy cows was collected over a period of about 110 days.
[0038] We observed daily behavioral video data of Holstein dairy cows, selecting six behaviors—feeding, standing, lying down, licking, estrus, and fighting—as well as the empty state of the cow bed as research subjects. These behaviors are closely related to the health evaluation of Holstein dairy cows.
[0039] Based on meticulous expert categorization, 60 independent video clips containing Holstein cow behaviors were manually selected. Each clip included one or more complete instances of Holstein cow behavior to capture behavioral patterns of Holstein cows at different times and under different environments. Each clip ranged in length from 10 to 25 seconds, with a video resolution of 1280 pixels × 720 pixels and a frame rate of 30 fps. Based on the defined behavioral characteristics of the Holstein cows, the CVAT data annotation tool was used, and the annotation results were exported in YOLO dataset format.
[0040] like Figure 1 As shown, Figure 1 (a) is a monitoring video scene of the Holstein dairy cow behavior dataset collected in this invention. Figure 1 (b) Example of data segmentation for Holstein dairy cow behavior.
[0041] like Figure 2 As shown,Figure 2 This presents the category distribution and annotation information for the Holstein cow behavior dataset of this invention. Figure 2 (a) shows the distribution of samples in each category in the constructed Holstein cow behavior dataset, with lying down, standing, and feeding behaviors being the most common, while estrus, licking, and fighting behaviors are less common. This distribution is consistent with the common behavioral patterns of Holstein cows in a farm environment and reflects the activity patterns of Holstein cows in a real farm setting. Figure 2 (b) shows the size and number of target boxes, which helps to optimize the anchor box size of the model and adjust the detection threshold of the model based on the target size. Figure 2 (c) describes the position of the target center point relative to the whole image, reflecting the spatial distribution characteristics of Holstein cows' positional preferences and behaviors in the image. Figure 2 (d) shows the aspect ratio distribution of the Holstein cow target box relative to the whole image, reflecting the multi-scale information of the Holstein cow.
[0042] The final dataset was labeled with 7 classes. To avoid excessive similarity between adjacent frames, differences between consecutive frames were removed and redundant images were deleted. There was no data overlap between different sample sets. A total of 2,418 images were collected, with a total of 23,073 bounding boxes labeled. Among them, there were 1,935 images in the training set and 483 images in the validation set. The proportions used for model training and model validation were 80% and 20%, respectively.
[0043] like Figure 3 As shown, Figure 3 The sample images in the Holstein dairy cow behavior dataset constructed for this invention are challenging examples, demonstrating problems such as inconsistent scale, uneven lighting, and excessive occlusion that occur in actual farm environments.
[0044] like Figure 4 As shown, Figure 4 This document presents a data augmentation demonstration of the Holstein cow behavior dataset constructed in this invention. The Mosaic data augmentation method involves randomly reading four images from the training set, performing random rotation, stitching, scaling, and translation operations, then applying color changes across the H, S, and V color gamuts, and finally stitching them together into a single image as training data. Figure 4 (a) is the original image of the sample. Figure 4 (b) The effect of Mosaic image enhancement. Figure 4(ci) represents the enhancement effects applied to a single image, including rotation, irregular cropping, brightness variation, noise addition, horizontal flipping, vertical flipping, and random occlusion. Compared to single-image enhancement methods, Mosaic-enhanced images contain richer scene and multi-scale information, increasing the diversity of the dataset. Image flipping simulates Holstein cow behavior samples from different perspectives. Image rotation can simulate the behavior of Holstein cows in different directions. Furthermore, adjusting image brightness can simulate different lighting conditions from daylight to night. Adding noise is used to simulate image quality under low light and inclement weather conditions. Finally, random cropping enhances the model's ability to recognize local behavioral features of Holstein cows, which is key to understanding complex behavioral patterns. Training the model using the augmented dataset indirectly increases the number of samples, accelerates model convergence, and improves the model's generalization ability.
[0045] like Figure 5 As shown, Figure 5 This is a diagram of the CAMLLA-YOLOv8n Holstein cow behavior recognition network architecture of this invention. To accurately identify cow behavior, this invention selects YOLOv8n as the basic model for cow behavior recognition. Improvements to the YOLOv8n model include: integrating a Coordinate Attention mechanism into the C2f module to form a C2f-CA module, which more accurately identifies and understands the spatial relationships between different Holstein cow locations, while improving the model's sensitivity to key regions and its ability to filter background interference; introducing an MLLA attention mechanism into the P3, P4, and P5 layers of the YOLOv8n model's Neck to address the challenge of Holstein cow behavior recognition due to large scale variations; improving the SPPF module to form the SPPF-GPE module, which optimizes small target recognition by combining global average pooling and global max pooling; and introducing Shape-IoU, focusing on the shape and scale features of the BoundingBox to improve the matching degree between the Prediction Box and the Ground Truth Box, thereby enhancing detection accuracy.
[0046] like Figure 6 As shown, YOLOv8n consists of four parts: Input, Backbone, Neck, and Head. The structural diagram of each module in the unmodified YOLOv8n network is shown below. Figure 6As shown. The Input layer, as the initial layer of the neural network, is responsible for receiving and processing the input image. The Backbone consists of a Convolutional Block, C2f, and SPPF (Spatial Pyramid Pooling Fast), responsible for learning complex feature representations from the input image. The Convolutional Block consists of a Convolutional Layer, an activation function SiLU, and a normalization layer BatchNorm2d. SPPF is used for feature extraction, and the C2f module, although lightweight, can obtain rich gradient information. YOLOv8n uses VFL Loss as the classification loss and DFL Loss+CIoULoss as the loss function for bounding box regression. For sample matching, it uses the Task-Aligned Assigner matching method. The Neck connects the Backbone and Head, and its design is crucial to the performance of the detection algorithm. The Neck adopts a FeaturePyramid Network (FPN) and Path Aggregation Network (PAN) structure, combining features at different scales to enhance the network's feature fusion capability. The Head contains three detection branches at different scales, obtaining the optimal detection box through non-maximum suppression (NMS). Head adopts the Anchor-Free approach, directly predicting the center point and size of the target object, no longer relying on preset anchors, thus improving the model's generalization ability on targets of various scales, shapes, and proportions.
[0047] like Figure 7 As shown, Figure 7 This is a diagram of the CA attention module of this invention. CAMLLA-YOLOv8n integrates the Coordinate Attention mechanism into the C2f module of the YOLOv8n network, forming a C2f-CA module. The C2f-CA module weights feature channels and utilizes the learned inter-channel dependencies to better understand the relationships between Holstein cows in different locations, enhancing the network's sensitivity to Holstein cow location information and effectively filtering out irrelevant background interference. The C2f-CA module significantly improves the model's ability to represent key behavioral regions of Holstein cows, guiding the model to focus on the true Region of Interest (ROI) and significantly improving the accuracy and robustness of behavior recognition.
[0048] like Figure 8 As shown, Figure 8This paper details the architecture of CAMLLA-YOLOv8n and its MLLAttention mechanism. CAMLLA-YOLOv8n improves upon the YOLOv8n model by enhancing the neck and introducing MLLAttention. YOLOv8n's head adopts an anchor-free approach, directly predicting the center point and size of the target object without relying on pre-defined anchors. P5, a small feature map, is used to detect large targets. The MLLAttention module is first added to the deepest P5 layer. In this layer, the P5 feature map, after passing through a series of convolutional layers, directly enters the MLLAttention module for processing. After MLLAttention processing in the P5 layer, the feature map is upsampled to increase its resolution and then concatenated with the original feature map from the P4 layer. The fused feature map from the P4 layer is then convolved again before being fed into the MLLAttention module. P3, a large feature map, is used to detect small targets. After MLLAttention processing in the P4 layer, the feature map is upsampled and fused with the original feature map from the P3 layer before entering the MLLAttention module in the P3 layer.
[0049] By applying MLLA attention at multiple levels (P3, P4, P5), CAMLLA-YOLOv8n can not only enhance the attention to the behavioral characteristics of Holstein cows at different scales, but also improve the accuracy of behavior recognition.
[0050] Scale adaptability of behavioral features: MLLAttention ensures the model can adapt to visual information at various scales through fine-tuning of feature maps from different levels. At the P5 level, by strengthening the focus on larger areas in the scene, the model can effectively identify behaviors requiring a wide field of view, such as fighting and mating. At the P3 level, by enhancing the capture of local details, it improves the ability to recognize detailed activities occurring within a smaller field of view, such as licking.
[0051] Spatial Context Enhancement for Behavior Recognition: Behavior recognition in Holstein cows relies not only on their posture and activities but also on their interactions with their surrounding environment. MLLAttention enhances the model's understanding of the environmental context by integrating spatial information from different scales, thereby improving the accuracy of recognizing behaviors such as standing and lying down that are related to changes in environmental position.
[0052] Temporal Feature Analysis of Dynamic Behavior: When identifying dynamic and complex behaviors such as estrus or fighting, MLLAttention focuses on changes in consecutive frames over time, capturing temporal information in the video sequence, which is crucial for understanding the behavior of Holstein cows within the video.
[0053] like Figure 9 As shown, Figure 9 This is a structural diagram of the SPPF-GPE of the present invention. In the face of the multi-scale challenges of Holstein cow behavior data, the wide distribution of target size across different pixel regions (small targets are 20-50 pixels wide and 30-100 pixels high; medium-sized targets are 150-300 pixels wide and 200-500 pixels high; and large targets are 300-600 pixels wide and 400-700 pixels high) significantly affects the model's recognition efficiency. In particular, the lack of detail and larger receptive field of small targets in deep feature maps increases the difficulty of recognition.
[0054] To address this issue, CAMLLA-YOLOv8n innovatively improves upon the YOLOv8n model by constructing SPPF-GPE (Global Pooling Enhanced SPPF), introducing global average pooling and global max pooling layers. GAP (Global Average Pooling) provides a global feature representation by calculating the average value across the entire feature map, which helps capture background and environmental information. GMP (Global Max Pooling) extracts the maximum value from each feature map, highlighting the most salient features and associating them with key parts of the Holstein cow or behavior, thus emphasizing the main object or behavior in the image, particularly suitable for detecting dynamic and complex behaviors such as fighting and mating. The fusion of global information from GAP and GMP yields broader background information and highlights key visual features. Furthermore, multi-scale pooling captures image details from coarse to fine, which is crucial for maintaining accuracy in behavior recognition under varying viewpoints and distances.
[0055] To verify the effectiveness of the proposed CAMLLA-YOLOv8n model, network performance was evaluated using six metrics: Precision, Recall, mAP@0.5, mAP@0.5-0.95, Params, and FLOPs. The calculations for Precision and Recall are shown below.
[0056]
[0057]
[0058] In this context, TP, FP, FN, and TN represent the number of actual positive class samples predicted as positive, the number of actual negative class samples predicted as positive, the number of actual positive class samples predicted as negative, and the number of actual negative class samples predicted as negative, respectively. Precision represents the proportion of true positive samples among those predicted as positive. Recall represents the proportion of all actual positive samples correctly identified as positive. mAP@0.5 represents the average AP value calculated for all images of each class when IoU is set to 0.5; mAP@0.5-0.95 represents the average mAP at different IoU thresholds (from 0.5 to 0.95, with a step size of 0.05). A higher mAP value indicates better detection performance of the object detection model. mAP is a key indicator for evaluating the performance of object detection algorithms, calculated by combining precision and recall through the area under the PR curve. Params, used to measure model memory usage, is the sum of all trainable parameters in the network. FLOPs represent the number of floating-point operations performed during a single inference operation, and are an important indicator of model complexity and computational efficiency.
[0059] To ensure fairness, all network models in the experiment were implemented and executed within the PyTorch framework. The hardware configuration used in this study was as follows: Ubuntu 20.04 operating system, Intel(R) Xeon(R) Platinum8352V CPU @ 2.10GHz, NVIDIA GeForce RTX 4090 GPU with 24GIB of VRAM, 64GB DDR5 RAM (32×2GB), and 4TB SSD. The deep learning framework used was PyTorch 2.2.2, CUDA version 12.2, cuDNN version 8.7, and Python version 3.9.19.
[0060] In deep learning model training, the setting of hyperparameters has a significant impact on the model's learning efficiency and final performance. During model training, the experimental parameters were set as follows: input image size was set to 640×640 pixels, batch size was 64, and to accelerate convergence, the initial learning rate was set to 0.01, the decay factor to 0.0005, the momentum factor to 0.937, and the batch size to 50 epochs. The Adam optimizer was used to iteratively optimize the network parameters, and other parameters were the default parameters from the YOLOv8 official documentation. After confirming model convergence, the model parameters were saved, and the model was evaluated.
[0061] To evaluate the performance of the CAMLLA-YOLOv8n model in Holstein cow behavior recognition, the same dataset as the CAMLLA-YOLOv8n model was used to evaluate the performance of six models, including YOLOv3, YOLOv5n, YOLOv5s, YOLOv7tiny, YOLOv8n, and YOLOv8s. As shown in Table 1, under the same experimental conditions, compared with the YOLOv8n model, the proposed CAMLLA-YOLOv8n Holstein cow behavior recognition algorithm achieved improvements of 2.187%, 1.619%, 1.839%, and 1.777% in Precision, Recall, mAP50, and mAP@0.5-0.95, respectively. Regarding Params and FLOPs, the Params of the CAMLLA-YOLOv8n model are 0.242M larger than those of YOLOv8n, remaining essentially unchanged. The CAMLLA-YOLOv8n model shows a slight increase in FLOPs, which improves the accuracy of Holstein cow behavior recognition to some extent. Although CAMLLA-YOLOv8n is the second most lightweight model after YOLOv5n, YOLOv5n's detection accuracy lags far behind CAMLLA-YOLOv8n. Table 1. Comparison of the overall performance of seven detection algorithms
[0062] like Figure 10 As shown, Figure 10 This paper compares the overall performance of seven YOLO detection algorithms presented in this invention. The distance of each curve from each axis represents the algorithm's performance on the corresponding metric. The larger the area enclosed by the curve, the better the algorithm's overall performance. For each curve's intersection with each axis, the closer it is to the outermost edge, the better the performance of that metric; the closer it is to the innermost edge, the worse the performance. YOLOv5n (orange curve) is close to the outermost edge in both Params and FLOPs, indicating that the YOLOv5n model is relatively lightweight in terms of Params and FLOPs. However, the CAMLLA-YOLOv8n model has Params and FLOPs second only to YOLOv5n and YOLOv8n, maintaining high performance despite lower resource consumption. In summary, the area enclosed by CAMLLA-YOLOv8n (red curve) in the figure is the largest, indicating that the CAMLLA-YOLOv8n model has achieved the highest performance in the four key metrics of Precision, Recall, mAP@0.5, and mAP@0.5-0.95, making it the best performing model overall among all current models.
[0063] like Figure 11 As shown, Figure 11This chart compares and analyzes the accuracy, recall, and mean accuracy of seven YOLO detection algorithms presented in this invention, showing the dynamic changes in four key performance indicators (Precision, Recall, mAP@0.5, and mAP@0.5-0.95) during training for the seven different YOLO models. In the initial stage of model training, all model performance indicators improve rapidly, and gradually stabilize as training progresses. In particular, the CAMLLA-YOLOv8n model exhibits high stability and excellent performance in the four key indicators of Precision, Recall, mAP50, and mAP@0.5-0.95. Its Precision and Recall both stabilize after the 30th epoch and show a slight performance improvement trend in the later stages, with fluctuations in Precision and Recall remaining around 94%.
[0064] like Figure 12 As shown, Figure 12 This diagram illustrates the training and validation loss of seven YOLO detection algorithms presented in this invention, depicting the loss changes of seven different YOLO models during the training and validation phases, including Box Loss, DFL Loss, and CLS Loss. In the initial stage of model training, due to the high learning rate, the loss curve rapidly decreases within the first 10 epochs, indicating enhanced model adaptability to the training data. The CAMLLA-YOLOv8n model performs particularly well in validation loss; as training progresses, its loss curve gradually stabilizes and essentially converges after approximately 30 epochs, with the loss value fluctuating around 0.3. This demonstrates that CAMLLA-YOLOv8n ensures high-accuracy detection while also possessing strong generalization ability, effectively suppressing overfitting.
[0065] The proposed CAMLLA-YOLOv8n model is based on YOLOv8n. It improves upon YOLOv8n by introducing CA attention into the C2f module, MLLA attention mechanism into the P3, P4, and P5 layers of the Neck layer, and innovatively improving the SPPF module to form the SPPF-GPE module, replacing the IoU loss function with the Shape-IoU loss function. To evaluate the impact of each optimized module of CAMLLA-YOLOv8n on model performance, ablation experiments were conducted using a variable-controlled method. Training and validation were performed on the same dataset and parameters, and the experimental results are shown in Table 2. Here, YOL0v8n represents the original network, and "√" indicates that this module was used. Table 2. Ablation Experiment Results of Different Optimization Modules
[0066] The increased background complexity resulting from Mosaic fusion and the growing number of Holstein cows necessitate a more accurate identification and understanding of the spatial relationships between different Holstein cow locations, enhancing the model's sensitivity to key regions and its ability to filter background interference. Firstly, by integrating a Coordinate Attention mechanism into the C2f module to form the C2f-CA module, the model's mAP50 and mAP@0.5-0.95 improved by 0.462% and 0.113%, respectively. Considering the challenges posed by large scale variations in Holstein cow behavior recognition, compared to YOLOv8n+CA, an MLLA attention mechanism was introduced in the P3, P4, and P5 layers of the Neck, resulting in improvements in precision and recall of 0.851% and 0.302%, respectively. Secondly, an innovative SPPF-GPE module was developed to improve the SPPF module. By combining global average pooling and global max pooling, small target recognition was optimized, improving the model's precision and recall by 1.189% and 0.684%, respectively. Finally, Shape-IoU loss optimization is introduced, focusing on the shape and scale features of the bounding box, to improve the matching degree between the prediction box and the ground truth box. CAMLLA-YOLOv8n improves the precision and recall by 0.851% and 0.302%, respectively.
[0067] like Figure 13 As shown, Figure 13 This document visualizes the ablation experimental results of different optimization modules of the CAMLLA-YOLOv8n algorithm in terms of Precision, Recall, mAP50, and mAP@0.5-0.95. The comprehensive experimental results show that each optimization module in the CAMLLA-YOLOv8n model improves the recognition accuracy to varying degrees, demonstrating the effectiveness of each optimization operation. Compared to the YOLOv8n model, although the Params and FLOPs of the CAMLLA-YOLOv8n model are slightly increased, it achieves significant improvements of 2.187%, 1.619%, 1.839%, and 1.777% in the four key performance indicators of Precision, Recall, mAP50, and mAP@0.5-0.95, respectively.
[0068] like Figure 14 As shown, Figure 14 This is a heatmap comparison between YOLOv8n and CAMLLA-YOLOv8n of this invention. The comparison is shown in three scenarios. The first row is the original image, the second row is the YOLOv8n heatmap, and the third row is the optimized CAMLLA-YOLOv8n heatmap.
[0069] To explore the improvement effect of CAMLLA-YOLOv8n and visually demonstrate the comparison before and after model optimization, this invention uses GradCAM to perform a visualization analysis of the key region responses of the YOLOv8n and optimized CAMLLA-YOLOv8n models. For example... Figure 14 As shown, the blue areas represent regions with lower confidence, while the red areas represent regions with higher confidence. It can be seen that compared to YOLOv8n, CAMLLA-YOLOv8n can more effectively focus on the heatmap values of key feature regions and dense areas in the same scenario, covering a wider range of target areas. Simultaneously, while focusing on a larger range of features, CAMLLA-YOLOv8n effectively suppresses interference from irrelevant information, improving the accuracy of Holstein cow behavior recognition. YOLOv8n, on the other hand, while focusing on a larger range of features, also picks up some irrelevant information.
[0070] like Figure 15 As shown, Figure 15 Visualization of the detection results of CAMLLA-YOLOv8n in this invention 1.
[0071] like Figure 16 As shown, Figure 16 Visualization of the CAMLLA-YOLOv8n detection results of this invention 2.
[0072] To further verify the effectiveness of the CAMLLA-YOLOv8n model in the Holstein cow behavior recognition task, detailed experiments were conducted. The results show that the average inference time for processing a single image using the CAMLLA-YOLOv8n model is 35.5 ms, which meets the requirements for real-time detection of video streams at 30 fps / s. With a confidence threshold of 0.65 and an IoU threshold of 0.5, the recognition results of CAMLLA-YOLOv8n on the dataset images were visualized. Figure 15 and Figure 16 As shown, CAMLLA-YOLOv8n performs well in various challenging scenarios, accurately identifying Holstein cow behavior in different environments, including background complexity caused by Mosaic fusion, increased Holstein cow numbers, multi-scale and multi-behavioral features of Holstein cows, and small target issues, effectively completing the identification task of this study.
Claims
1. A Holstein cow behavior recognition method based on a CAMILLA-YOLOv8n algorithm, characterized in that, Comprising the following steps: Step 1: Two Tiandy cameras (TD-H234S) were installed on-site at a Holstein dairy farm, mounted on the gable barn beam support wall for real-time data collection and processing. The cameras were fixed at a height of approximately 4.5 meters and tilted downward at an angle of about 30°, allowing them to capture the entire area of the dairy cow activity. Video data of Holstein dairy cow daily behavior spanning approximately 110 days was collected; Step 2: Observing the Holstein dairy cow daily behavior video data in Step 1, six behaviors of Holstein dairy cows were selected as research objects, including eating, standing, lying, licking, estrus, and fighting, which are closely related to Holstein dairy cow health evaluation; Step 3: Based on expert classification, 60 independent video clips containing Holstein dairy cow behavior were manually selected, each clip containing one or more complete instances of Holstein dairy cow behavior to capture the behavior patterns of Holstein dairy cows at different time periods and in different environments. The length of each clip varied from 10 to 25 seconds, with a video resolution of 1280 pixels x 720 pixels and a frame rate of 30 fps / s. According to the specific behavior characteristics of Holstein dairy cows divided in Step 2, the labeling results were exported in YOLO dataset format using the CVAT data labeling tool; the final dataset was labeled as 7 classes. To avoid excessive similarity between adjacent frames, the differences between consecutive frames were removed and redundant images were deleted, there was no data overlap between different sample sets, finally a total of 2418 images were collected, a total of 23073 boxes were labeled, of which 1935 were in the training set and 483 were in the validation set, the proportions for model training and model validation were 80% and 20% respectively. Step 4: Considering the problems of different scales of Holstein dairy cows, uneven lighting, excessive occlusion, etc. in the actual breeding environment, the labeled Holstein dairy cow daily behavior dataset images in Step 3 were enhanced using Mosaic data enhancement, rotation transformation, irregular cutting, brightness change, noise addition, left-right flipping, up-down flipping, and random occlusion transformation to enrich Holstein dairy cow multi-scale behavior data and background information. Compared to single image enhancement methods, Mosaic enhanced images contain more rich scene and multi-scale information, increasing the diversity of the dataset. Image flipping simulates Holstein dairy cow behavior samples under different perspectives. Image rotation transformation can simulate Holstein dairy cow behavior in different directions. In addition, adjusting the brightness of the image can simulate different lighting conditions from daylight to night. The method of adding noise is used to simulate image quality under low light and adverse weather conditions. Finally, random cropping enhances the model's ability to recognize local behavior features of Holstein dairy cows, which is crucial for understanding complex behavior patterns. Training the model using the data-enhanced dataset indirectly increases the number of samples, speeds up model convergence, and improves the model's generalization ability. The improved YOLOv8n model includes: fusing the Coordinate Attention mechanism in the C2f module to form the C2f-CA module, which more accurately identifies and understands the spatial relationship between different Holstein cow positions, while improving the sensitivity of the model to key areas and the filtering ability of background interference; introducing the MLLAttention mechanism in the P3, P4, and P5 layers of the YOLOv8n model Neck to address the challenges of Holstein cow behavior recognition due to large scale changes; improving the SPPF module to form the SPPF-GPE module, which optimizes small target recognition by combining global average pooling and global maximum pooling processing; introducing Shape-IoU, which focuses on the shape and scale features of BoundingBox, to improve the matching degree between PredictionBox and GroundTruthBox and enhance detection accuracy.
2. The Holstein cow behavior recognition method based on the CAMILLA-YOLOv8n algorithm according to claim 1, characterized in that, The steps include: In the above step 5, CAMLLA-YOLOv8n is a C2f-CA module formed by fusing the Coordinate Attention mechanism in the C2f module of the YOLOv8n network. The C2f-CA module performs weighted processing on the feature channels, uses the learned inter-channel dependency relationship, and the model can better understand the relationship between different positions of Holstein cows, enhancing the network's sensitivity to Holstein cow position information and effectively filtering out irrelevant background interference. The C2f-CA module significantly improves the model's ability to represent key areas of Holstein cow behavior, guiding the model to focus on real ROIs and significantly improving the accuracy and robustness of behavior recognition. The core idea of the CA attention mechanism is to embed position information into channel attention, allowing lightweight networks to pay attention to a larger area so that the model can better understand the relationship between different positions while avoiding excessive computational overhead. The CA attention mechanism first uses two one-dimensional global pooling operations to aggregate input features in the vertical and horizontal directions into two independent direction-aware feature maps. These two direction-specific feature maps are encoded into two attention maps, each capturing the long-range dependencies of the input feature map along one spatial direction. To alleviate the loss of position information caused by 2D global pooling, the channel attention is decomposed into two parallel 1D feature encoding processes, effectively integrating spatial coordinate information into the generated attention map. The CA attention mechanism first performs 2 global average pooling operations on the input feature map in the width and height directions, respectively, to obtain 2 feature maps that capture global features in the width and height directions. The above process is shown in equation (1): where z c represents global features; x c (i,j) is the input of the original features. In detail, first, a pooling kernel with size (H, 1) or (1, W) is used to encode each channel along the horizontal and vertical coordinates, respectively, to aggregate the features along two spatial directions, obtaining a pair of direction-aware feature maps. Therefore, the output of the c-th channel with height h can be represented as formula (2): Similarly, the output of the c-th channel with a width of w can be represented as equation (3): Then, after channel-level concatenation and convolution smoothing, a feature map with a global receptive field and more accurate information representation is generated, as shown in equation (4): f = δ(F1([z h ,z w ]) (4) Wherein, F1 is the feature map after batch normalization processing; f is the feature map obtained after Sigmoid activation function; delta is a nonlinear activation function; z h and z w are the aggregated feature maps. Then, through batch regularization and nonlinear mapping transformation, a tensor with the same number of channels as the input data is obtained, as shown in formula (5) and formula (6). g h = δ(F h (f h )) (5) g w = δ(F w (f w )) (6) where f h and f w are two independent tensors that f is split along the spatial dimensions; F h and F w are the transformations of f h and f w to tensors with the same number of channels as the input; g h and g w are the attention weights of the input feature map in the height and width directions respectively after the above calculation; δ is the Sigmoid function. Finally, the final feature map with attention weights in the width and height directions is obtained by multiplication weighting calculation on the original feature map, the position information is saved in the generated feature map with attention weights, and the last two attention maps are then multiplied to the input feature map to enhance the representation ability of the feature map. The final output of CoordinateAttention is shown in equation (7): where y c (i,j) denotes the final output feature; x c (i,j) denotes the original feature; and denote the features that have been aggregated to recover the original number of channels in the height and width dimensions, respectively.
3. The method for Holstein cow behavior recognition based on the CAMILLA-YOLOv8n algorithm according to claim 1, characterized in that, The steps include: In step 5 above, CAMLLA-YOLOv8n improves the Neck and introduces MLLAttention attention based on the YOLOv8n model. The Head of YOLOv8n adopts the Anchor-Free idea, directly predicting the center point and size of the target object, no longer relying on the preset Anchor. P5 is a small feature map used to detect large targets. The MLLAttention module is first added to the deepest P5 layer. After the P5 feature map passes through a series of convolution layers, it is directly processed in the MLLAttention module. After the MLLAttention processing in the P5 layer, the feature map is increased in resolution through upsampling, and then concatenated with the original feature map of the P4 layer. The fused feature map of the P4 layer is processed again through convolution and then sent to the MLLAttention module. P3 is a large feature map used to detect small targets. After the MLLAttention processing in the P4 layer, the feature map is upsampled and fused with the original feature map of the P3 layer, and then processed into the MLLAttention module of the P3 layer. By applying MLLAttention at multiple levels (P3, P4, P5), CAMLLA-YOLOv8n not only enhances attention on Holstein cow behavior characteristics at different scales, but also improves behavior recognition accuracy. Scale adaptability of behavior characteristics: MLLAttention ensures that the model can adapt to visual information of various scales by fine-tuning the feature maps from different levels. In the P5 layer, the model can effectively identify behaviors that require a wide field of view, such as fighting and estrus, by strengthening attention on larger areas in the scene. In the P3 layer, the model improves the ability to capture local details, such as licking, which occurs within a smaller field of view. Spatial context enhancement of behavior recognition: Holstein cow behavior recognition not only depends on the posture and activity of Holstein cows, but also closely related to their interaction with the surrounding environment. MLLAttention enhances the model's understanding of environmental context by integrating spatial information from different scales, thereby improving the recognition accuracy of behaviors related to changes in environmental position, such as standing and lying. Temporal feature analysis of dynamic behavior: When identifying dynamic and complex behaviors such as estrus or fighting, MLLAttention focuses on the changes in consecutive frames over time, capturing time dimension information in video sequences, which is crucial for understanding Holstein cow behavior within the video.
4. The method for Holstein cow behavior recognition based on the CAMILLA-YOLOv8n algorithm according to claim 1, characterized in that, The following steps are included: In step 5 above, in the face of the multi-scale challenge of Holstein cow behavior data, the wide distribution of target sizes in different pixel areas (small target width between 20-50 pixels, height between 30-100 pixels, medium-sized target width between 150-300 pixels, height between 200-500 pixels, and large-sized target width between 300-600 pixels, height between 400-700 pixels) significantly affects the recognition efficiency of the model. In particular, due to the lack of details in small targets in deep feature maps and large receptive fields, the recognition difficulty is increased. To address this issue, CAMLLA-YOLOv8n innovatively improves the SPPF module of YOLOv8n to form SPPF-GPE (Global Pooling Enhanced SPPF) based on the YOLOv8n model. Global average pooling and global maximum pooling layers are introduced, where GAP (Global Average Pooling) provides a form of global feature representation by calculating the average value of the entire feature map, which helps to capture background and environmental information. GMP (Global Max Pooling) extracts the maximum value in each feature map, emphasizing the most significant features associated with key parts of cows or behaviors, highlighting the main objects or behaviors in the image, especially suitable for detecting dynamic and complex behaviors such as Fighting and Mating. After the fusion of global information through GAP and GMP, more extensive background information and key visual features are obtained. In addition, multi-scale pooling can capture image details from coarse to fine, which is crucial for maintaining the accuracy of behavior recognition under varying angles and distances.
5. The method for Holstein cow behavior recognition based on the CAMILLA-YOLOv8n algorithm according to claim 1, characterized in that, The following steps are included: In step 5 above, as an important component of the detector positioning branch, Bounding Box Regression Loss plays a significant role in the target detection task. Existing bounding box regression methods such as IoU, GIoU, CIoU, and SIoU all add new geometric constraints to IoU. Typically, only the geometric relationship between GroundTruthBox and PredictionBox is considered, and the loss is calculated using the relative position and relative shape between the bounding boxes. However, the inherent properties of the box itself, such as shape and scale, are ignored, which makes the network sensitive to the position offset of the target. This leads to similar features between positive and negative samples, making it difficult for the network to converge. In order to overcome the shortcomings of existing research, based on the reality that the shape and scale factors of the bounding box itself will affect the regression results, this paper uses the bounding box regression method Shape-IoU which focuses on the shape and scale of the bounding box itself. It can calculate the loss by focusing on the shape and scale of the bounding box itself, so that the bounding box regression is more accurate, effectively improves the detection accuracy and is better than the existing method. CAMLLA-YOLOv8n is based on YOLOv8n model, using Shape-IoU to replace the original CIoU loss, which can calculate the loss by focusing on the shape and scale of the bounding box itself, so that the bounding box regression is more accurate.