Intelligent Monitoring Method for Sow Farrowing Behavior Based on Improved YOLOv8 Model
By improving the YOLOv8 model and adding the DLKA module and Adaptive Threshold Focal Loss, the problems of time-consuming manual observation and poor sensor performance in monitoring sow farrowing behavior have been solved. This has enabled accurate monitoring of sow posture changes and piglet birth conditions, thus improving the intelligent and refined management of the pig industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTH CHINA AGRICULTURAL UNIVERSITY
- Filing Date
- 2025-01-07
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for monitoring sow farrowing behavior suffer from problems such as time-consuming, labor-intensive, and inefficient manual observation, stress response caused by wearable devices, and poor sensor monitoring performance.
An improved YOLOv8 model was adopted. By adding the DLKA module to the head part of the model and introducing Adaptive Threshold Focal Loss, a sow farrowing behavior monitoring model was constructed, and iterative optimization training was carried out to improve the monitoring accuracy.
It enables automatic, precise, and real-time monitoring of sow farrowing behavior, reduces labor costs, improves production management efficiency, provides comprehensive management data, and promotes the intelligent and refined development of the pig farming industry.
Smart Images

Figure CN119992648B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically, to an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model. Background Technology
[0002] my country is a major livestock and poultry farming country, and pig farming is an important component of its livestock industry, providing a significant driving force for the development of animal husbandry. The farrowing process is a crucial stage in pig production, and its management quality directly affects piglet survival rates and sow health. In farrowing management, real-time recording and analysis of the frequency and duration of sow posture changes are of great significance for assessing farrowing progress, determining sow comfort, and intervening early in cases of abnormal farrowing.
[0003] Sows may exhibit various postural changes during farrowing, such as standing, sitting, lying on their side, and prone. The duration and frequency of these different postures reflect the sow's response to pain, anxiety, and fatigue. Studies have shown that sows typically maintain a lying-on or prone position for extended periods to facilitate parturition and lactation, while frequent standing or sitting may signal pain or discomfort. Therefore, accurately collecting and analyzing postural data during farrowing can help identify potential farrowing difficulties and provide farmers with a scientific management basis.
[0004] In addition, recording the birth time, order, and health status of piglets is equally crucial. This data reflects the sow's farrowing progress and the overall health of the piglets. By recording the birth time of each piglet, it's possible to determine whether farrowing was successful and if there were any abnormal intervals, thus allowing for early detection and intervention of potential problems such as dystocia or fetal distress. The initial health status of piglets after birth also provides important information for subsequent feeding and management, ensuring that weaker piglets receive timely care and support.
[0005] Currently, monitoring sow farrowing behavior in production mainly relies on manual observation. However, this method increases the time spent between humans and animals, posing a risk of zoonotic diseases. Furthermore, it is influenced by the subjective experience of the farm workers, making it time-consuming and labor-intensive, and unable to meet the needs of modern large-scale pig farming. Alternatively, wearable devices, photocells, and ultrasonic sensors can be used to monitor sow farrowing behavior. While this reduces labor, wearable devices may cause stress in the pigs, have limited power supply time, and are prone to falling off. Photocells and ultrasonic sensors are easily affected by the surrounding environment, resulting in relatively low sensitivity and unsatisfactory monitoring results. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, such as the time-consuming, labor-intensive, and inefficient nature of manual observation and the poor performance of sensor monitoring, this invention provides an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model. This method can accurately capture changes in the sow's posture during farrowing and accurately identify and record the birth of piglets, effectively promoting the pig industry towards intelligence and precision, and providing strong support for the high-quality development of my country's animal husbandry.
[0007] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0008] A method for intelligent monitoring of sow farrowing behavior based on an improved YOLOv8 model includes the following steps:
[0009] S1: Obtain and preprocess the sow farrowing behavior dataset, and divide the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows throughout the farrowing process;
[0010] S2: The YOLOv8 model is selected as the base model and improved. DLKA modules are added in front of the three detectors in the Head part of the YOLOv8 model, and Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detectors to build a sow farrowing behavior monitoring model.
[0011] S3: Use the training set to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the loss function is less than or equal to a preset threshold, and then complete the training to obtain the trained sow farrowing behavior monitoring model.
[0012] S4: Use the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluate the performance of the trained sow farrowing behavior monitoring model.
[0013] Preferably, step S1 includes:
[0014] Acquire daily lighting videos and nighttime infrared videos of several sows during farrowing, extract frames according to video time sequence, and filter out blurry and highly similar video frames to obtain the filtered dataset;
[0015] The preprocessing process involves labeling each frame of the filtered dataset and performing data augmentation.
[0016] The preprocessed dataset is divided into training and test sets according to a preset ratio.
[0017] Preferably, the filtering of blurry and highly similar video frames includes the following steps:
[0018] S1.1: Convert the images of each frame obtained by video frame extraction into floating-point images;
[0019] S1.2: All floating-point images are processed using a Laplace filter to highlight high-frequency components in each frame; the calculation formula for the Laplace filter is as follows:
[0020]
[0021] Where L(x,y) is the pixel value output by the Laplacian filter at image position (x,y), f(i,j) is the input pixel value at position (i,j), G(i,j) is the Gaussian kernel value at position (i,j), w is the width of the image, and h is the height of the image; a higher L(x,y) value indicates a clearer image.
[0022] S1.3: Divide all images into several groups at a certain frame interval, filter out the k frames with the lowest L(x,y) values in each group, and obtain the dataset after filtering the blurred frames, where k is a positive integer less than the frame interval.
[0023] S1.4: Using the SSIM algorithm, calculate the similarity between each frame image in the dataset after filtering blurred frames pairwise, and filter one of the two images with a similarity greater than or equal to a preset similarity threshold to obtain the filtered dataset.
[0024] Preferably, the label includes any one of the following:
[0025] Standing: The sow's limbs are straight to support her body, and her abdomen is off the ground;
[0026] Sitting posture: The sow sits on the ground with her forelegs extended for support, her hind legs bent, her hindquarters on the ground, and her abdomen off the ground;
[0027] Side-lying position: The sow lies completely on her side on the ground, with her shoulder, ribs and leg on one side in contact with the ground, her limbs relaxed, her abdomen completely in contact with the ground, and her teats exposed;
[0028] Lying prone: The sow's abdomen is on the ground, her forelegs and hind legs are naturally bent or straightened on both sides of her body, and her head can be raised or lowered.
[0029] Birth: Piglets are still attached to the sow's vulva, with part of their head or hind limbs exposed.
[0030] Piglets: refers to piglets that have completely detached from the sow's vulva.
[0031] Preferably, the data augmentation process includes any one or more of the following:
[0032] Angles are selected that are uniformly distributed within the range of [-180°, 180°], and each frame of the image is rotated randomly.
[0033] Randomly flip each frame of the image horizontally;
[0034] Randomly change the brightness, contrast, or saturation of each frame of the image;
[0035] Gaussian noise is randomly added to each frame of the image.
[0036] Preferably, in step S2, the DLKA module includes, in sequence, a first 2D convolutional layer, a GELU activation layer, a Deform-DW Conv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer;
[0037] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual summation connection; the output of the GELU activation layer is multiplied by the output of the second 2D convolutional layer in a weighted manner;
[0038] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence.
[0039] In the Deform-DW Conv2D submodule, the input features are processed by a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position.
[0040] Based on the output features of the Deform-DW Conv2D submodule, an inflation value is added and used as the input feature of the Deform-DW-D Conv2D submodule.
[0041] Preferably, in step S2, the Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle imbalanced datasets and difficult samples. The calculation formula for Adaptive Threshold Focal Loss is as follows:
[0042]
[0043] Among them, L ATFL The function value of Adaptive Threshold Focal Loss; p represents the predicted value for the next epoch. t This represents the average predicted probability value for the current epoch; λ is a hyperparameter.
[0044] Preferably, in step S4, real-time monitoring of the sow's farrowing behavior includes:
[0045] The trained sow farrowing behavior monitoring model was used to analyze video data of sows in the test set throughout the farrowing process to determine the sow's current posture and farrowing status in real time.
[0046] Based on the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of piglets are recorded.
[0047] Preferably, the preset judgment strategy includes:
[0048] When videos from the test set are input into the trained sow farrowing behavior monitoring model, recording is triggered only if the following conditions are met: the duration of the previous posture is recorded, and the frequency of each posture change so far is calculated:
[0049]
[0050] Where p is the sow's previous stable posture; c is the sow's current posture; q is the number of consecutive frames that maintain the same posture; t is a preset frame threshold; P q The probability that the current posture is detected by the trained sow farrowing behavior monitoring model; P is a preset probability threshold.
[0051] Preferably, in step S4, the metrics used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1 score, and mean average recognition rate (mAP).
[0052] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0053] This invention provides an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model. First, a sow farrowing behavior dataset is acquired and preprocessed, then divided into a training set and a test set. Next, the YOLOv8 model is selected as the base model and improved by adding DLKA modules before the three detectors in the Head part of the YOLOv8 model and introducing Adaptive Threshold Focal Loss as the classification loss to construct a sow farrowing behavior monitoring model. Then, the training set is used to iteratively optimize and train the sow farrowing behavior monitoring model until the loss function value is less than or equal to a preset threshold, completing the training and obtaining a well-trained sow farrowing behavior monitoring model. Finally, the trained sow farrowing behavior monitoring model is used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model is evaluated.
[0054] The present invention has the following advantages:
[0055] 1) This invention uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss function, which can effectively improve the accuracy of monitoring sow farrowing behavior;
[0056] 2) This invention utilizes computer vision technology, which can not only automatically, accurately, and in real time capture changes in the posture of sows, but also identify and record the birth of piglets, thereby reducing labor and improving the efficiency of production management.
[0057] 3) This invention can achieve 24-hour uninterrupted monitoring. Based on the data obtained from real-time monitoring by this invention, not only can the dynamic pattern of the sow's farrowing posture change be plotted, but the timeline and health indicators of piglets' birth can also be presented intuitively, providing farmers with comprehensive management basis.
[0058] 4) This invention only requires the collection of video data on the farrowing behavior of sows during monitoring. The non-contact monitoring will not cause stress to the pigs, which not only improves the farrowing experience of sows and the survival rate of piglets, but also promotes the pig farming industry towards intelligence and precision, providing strong support for the high-quality development of my country's animal husbandry. Attached Figure Description
[0059] Figure 1 This is a flowchart of an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model, as provided in Example 1.
[0060] Figure 2 This is a structural diagram of the YOLOv8 model provided in Example 1.
[0061] Figure 3 This is a structural diagram of the YOLOv8 model after adding the DLKA module, as shown in Example 1.
[0062] Figure 4 This is a structural diagram of the DLKA module provided in Example 1.
[0063] Figures 5-8 The images shown are of the sows in Stand, Sit, Sternal, and Lateral states under normal lighting conditions provided in Example 2.
[0064] Figures 9-12 The images provided in Example 2 show the detection results of images of sows in Stand, Sit, Sternal, and Lateral states under normal lighting conditions.
[0065] Figures 13-15 The images are infrared images of sows in Stand, Sit, and Lateral states collected at night, as provided in Example 2.
[0066] Figures 16-18 The results are the detection results of infrared images of sows in Stand, Sit, and Lateral states collected at night, as provided in Example 2.
[0067] Figure 19 The image shows piglets in their Birth state under normal lighting conditions, as provided in Example 2.
[0068] Figure 20 The image shows a sow in Piglet mode under normal lighting conditions, as provided in Example 2.
[0069] Figure 21 The image provided in Example 2 is an infrared image of a piglet in its Birth state, captured at night.
[0070] Figure 22 The image provided in Example 2 is an infrared image of a sow in Piglet mode acquired at night.
[0071] Figure 23 This is the result of recording the piglet birth information provided in Example 2.
[0072] Figure 24 This is the result of recording the sow posture change information provided in Example 2. Detailed Implementation
[0073] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0074] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;
[0075] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.
[0076] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0077] Example 1
[0078] like Figure 1 As shown, this embodiment provides an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model, including the following steps:
[0079] S1: Obtain and preprocess the sow farrowing behavior dataset, and divide the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows throughout the farrowing process;
[0080] S2: The YOLOv8 model is selected as the base model and improved. DLKA modules are added in front of the three detectors in the Head part of the YOLOv8 model, and Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detectors to build a sow farrowing behavior monitoring model.
[0081] S3: Use the training set to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the loss function is less than or equal to a preset threshold, and then complete the training to obtain the trained sow farrowing behavior monitoring model.
[0082] S4: Use the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluate the performance of the trained sow farrowing behavior monitoring model.
[0083] In the specific implementation process, the sow farrowing behavior dataset is first acquired and preprocessed. In this embodiment, the sow farrowing behavior dataset includes video data of several sows before, during and after farrowing. After preprocessing, the preprocessed sow farrowing behavior dataset is divided into a training set and a test set.
[0084] Next, the YOLOv8 model was selected as the base model, and improvements were made to the YOLOv8 model.
[0085] like Figure 2As shown, the YOLOv8 series of neural network models includes YOLOv8n, YOLOv8s, YOLOv8m, and YOLOv8l models. The YOLOv8 model architecture is based on a convolutional neural network (CNN) with various convolutional layers in the backbone. The backbone consists of a series of feature extraction layers, which are responsible for extracting features from the input image. These features are then fed into the Neck section for processing, and finally into three sets of detection heads in the Head section. The detection heads use a set of convolutional layers and fully connected layers for object detection and classification. Figure 2 In the diagram, the Conv module consists of a 3×3 Conv2d, BatchNorm, and SiLU activation function; the C2f module consists of 3 ConvModules and n Bottlenecks (with residual connections); and the Bottleneck consists of 2 ConvModules.
[0086] The improved YOLOv8 object detection network model is as follows: Figure 3 As shown, DLKA (Deformable Large Kernel Attention) modules are added before the three detection heads in the Head part of the YOLOv8 model. The addition of DLKA modules enables the network to extract more detailed features, thereby improving the accuracy of target detection.
[0087] like Figure 4 As shown, the DLKA module includes the following components connected in sequence: a first 2D convolutional layer, a GELU activation layer, a Deform-DWConv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer.
[0088] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual summation connection; the output of the GELU activation layer is multiplied by the output of the second 2D convolutional layer in a weighted manner;
[0089] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence.
[0090] In the Deform-DW Conv2D submodule, the input features are processed by a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position. That is, an offset is first generated through a standard convolution, and this offset is continuously adjusted in the later learning. With the offset, the convolutional kernel can adaptively adjust its position, so that the convolution operation can more accurately select important local regions in the image for feature extraction, rather than the fixed grid position in traditional convolution.
[0091] Based on the output features of the Deform-DW Conv2D submodule, an inflation value is added and used as the input feature of the Deform-DW-D Conv2D submodule;
[0092] The DLKA module adaptively adjusts the shape of the convolutional kernels based on the local structure of the feature map. Deformable convolution allows each kernel to shift, enabling it to sample different local regions at each location in the feature map. Then, based on an attention mechanism, it assigns weights to different regions in the feature map, highlighting key areas and weakening unimportant information. The input and output of DLKA are represented by the following mathematical formulas:
[0093] x out =Conv2D(x1*D-LKA-Attn(x1))+x in
[0094] x1 = GELU(Conv2D(x in ))
[0095] D-LKA-Attn(x1)=Conv2D(DeformDConv2D(DeformConv2D(x1)))
[0096] Where, x in For the input of the DLKA module; x out x1 is the output of the DLKA module; x1 is the intermediate feature; D-LKA-Attn is the deformable LKA attention; Conv2D is the 2D convolution; GeLU is the GeLU activation function; DeformConv2D is the computation process of the Deform-DWConv2D submodule; DeformDConv2D is the computation process of the Deform-DW-D Conv2D submodule; DeformDConv2D is an extension of DeformConv2D, which adds dilation to the convolution.
[0097] After improving the YOLOv8 object detection network model, we introduced Adaptive Threshold FocalLoss to replace the original classification loss of the three detection heads, while keeping other loss functions unchanged, and constructed a sow farrowing behavior monitoring model. This loss function is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle imbalanced datasets and difficult samples.
[0098] Then, the training set is used to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the loss function is less than or equal to the preset threshold, and the training is completed to obtain the trained sow farrowing behavior monitoring model.
[0099] Finally, the trained sow farrowing behavior monitoring model was used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model was evaluated.
[0100] This method uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss function, which can effectively improve the accuracy of monitoring sow farrowing behavior. This method can accurately capture the changes in sow posture during farrowing and accurately identify and record the birth of piglets, effectively promoting the pig industry towards intelligence and precision, and providing strong support for the high-quality development of my country's animal husbandry.
[0101] Example 2
[0102] This embodiment provides an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model, including the following steps:
[0103] S1: Obtain and preprocess the sow farrowing behavior dataset, and divide the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows throughout the farrowing process;
[0104] S2: The YOLOv8 model is selected as the base model and improved. DLKA modules are added in front of the three detectors in the Head part of the YOLOv8 model, and Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detectors to build a sow farrowing behavior monitoring model.
[0105] S3: Use the training set to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the loss function is less than or equal to a preset threshold, and then complete the training to obtain the trained sow farrowing behavior monitoring model.
[0106] S4: Use the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluate the performance of the trained sow farrowing behavior monitoring model.
[0107] Step S1 includes:
[0108] Acquire daily lighting videos and nighttime infrared videos of several sows during farrowing, extract frames according to video time sequence, and filter out blurry and highly similar video frames to obtain the filtered dataset;
[0109] The preprocessing process involves labeling each frame of the filtered dataset and performing data augmentation.
[0110] The preprocessed dataset is divided into training and testing sets according to a preset ratio;
[0111] The filtering of blurry and highly similar video frames includes the following steps:
[0112] S1.1: Convert the images of each frame obtained by video frame extraction into floating-point images;
[0113] S1.2: All floating-point images are processed using a Laplace filter to highlight high-frequency components in each frame; the calculation formula for the Laplace filter is as follows:
[0114]
[0115] Where L(x,y) is the pixel value output by the Laplacian filter at image position (x,y), f(i,j) is the input pixel value at position (i,j), G(i,j) is the Gaussian kernel value at position (i,j), w is the width of the image, and h is the height of the image; a higher L(x,y) value indicates a clearer image.
[0116] S1.3: Divide all images into several groups at a certain frame interval, filter out the k frames with the lowest L(x,y) values in each group, and obtain the dataset after filtering the blurred frames, where k is a positive integer less than the frame interval.
[0117] S1.4: Using the SSIM algorithm, calculate the similarity between each frame image in the dataset after filtering the blurred frames pairwise, and filter one of the two images with a similarity greater than or equal to a preset similarity threshold to obtain the filtered dataset.
[0118] The label includes any of the following:
[0119] Standing: The sow's limbs are straight to support her body, and her abdomen is off the ground;
[0120] Sitting posture: The sow sits on the ground with her forelegs extended for support, her hind legs bent, her hindquarters on the ground, and her abdomen off the ground;
[0121] Side-lying position: The sow lies completely on her side on the ground, with her shoulder, ribs and leg on one side in contact with the ground, her limbs relaxed, her abdomen completely in contact with the ground, and her teats exposed;
[0122] Lying prone: The sow's abdomen is on the ground, her forelegs and hind legs are naturally bent or straightened on both sides of her body, and her head can be raised or lowered.
[0123] Birth: Piglets are still attached to the sow's vulva, with part of their head or hind limbs exposed.
[0124] Piglets: refers to piglets that have completely detached from the sow's vulva;
[0125] The data augmentation process includes any one or more of the following:
[0126] Angles are selected that are uniformly distributed within the range of [-180°, 180°], and each frame of the image is rotated randomly.
[0127] Randomly flip each frame of the image horizontally;
[0128] Randomly change the brightness, contrast, or saturation of each frame of the image;
[0129] Gaussian noise is randomly added to each frame of the image;
[0130] In step S2, the DLKA module includes the following components connected in sequence: a first 2D convolutional layer, a GELU activation layer, a Deform-DWConv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer.
[0131] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual summation connection; the output of the GELU activation layer is multiplied by the output of the second 2D convolutional layer in a weighted manner;
[0132] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence.
[0133] In the Deform-DW Conv2D submodule, the input features are processed by a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position.
[0134] Based on the output features of the Deform-DW Conv2D submodule, an inflation value is added and used as the input feature of the Deform-DW-D Conv2D submodule;
[0135] In step S2, the Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle imbalanced datasets and difficult samples. The calculation formula for Adaptive Threshold Focal Loss is as follows:
[0136]
[0137] Among them, L ATFL The function value of Adaptive Threshold Focal Loss; p represents the predicted value for the next epoch. t This represents the average predicted probability value for the current epoch; λ is a hyperparameter.
[0138] In step S4, real-time monitoring of the sow's farrowing behavior includes:
[0139] The trained sow farrowing behavior monitoring model was used to analyze video data of sows in the test set throughout the farrowing process to determine the sow's current posture and farrowing status in real time.
[0140] Based on the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of piglets are recorded.
[0141] The preset judgment strategy includes:
[0142] When videos from the test set are input into the trained sow farrowing behavior monitoring model, recording is triggered only if the following conditions are met: the duration of the previous posture is recorded, and the frequency of each posture change so far is calculated:
[0143]
[0144] Where p is the sow's previous stable posture; c is the sow's current posture; q is the number of consecutive frames that maintain the same posture; t is a preset frame threshold; P q The probability that the trained sow farrowing behavior monitoring model detects the current posture; P is a preset probability threshold.
[0145] In step S4, the metrics used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1 score, and mean average recognition rate (mAP).
[0146] In the specific implementation process, the sow farrowing behavior dataset is first obtained and preprocessed. In this embodiment, the sow farrowing behavior dataset includes video data of several sows before, during and after farrowing.
[0147] Next, daily lighting videos and nighttime infrared videos of several sows during farrowing were extracted in chronological order, and blurry and highly similar video frames were filtered out. This included the following steps:
[0148] Each frame obtained by video frame extraction is converted into a 64-bit floating-point image; a Laplace filter is used to process all floating-point images to highlight the high-frequency components in each frame; the calculation formula for the Laplace filter is as follows:
[0149]
[0150] Where L(x,y) is the pixel value output by the Laplacian filter at image position (x,y), f(i,j) is the input pixel value at position (i,j), G(i,j) is the Gaussian kernel value at position (i,j), w is the width of the image, and h is the height of the image; a higher value of L(x,y) indicates more details and a clearer image.
[0151] Divide all images into several groups at 10-frame intervals, filter out the two frames with the lowest L(x,y) values in each 10-frame group, and obtain the dataset after filtering out the blurred frames.
[0152] Using the SSIM (Structure Similarity Index Measure) algorithm, the similarity between each frame image in the dataset after filtering blurred frames is calculated pairwise, and one of the two images with a similarity greater than or equal to a preset similarity threshold (0.9 in this embodiment) is filtered to obtain the filtered dataset.
[0153] After screening, 14,000 video frames were obtained, with a frame size of 1920×1080. Each frame was labeled and categorized into six classes based on the sow's posture and the piglet's birth: Standing, Sitting, Lateral, Sternal, Birth, and Piglet. The definitions of each category are as follows:
[0154] Standing: The sow's limbs are straight to support her body, and her abdomen is off the ground;
[0155] Sitting posture: The sow sits on the ground with her forelegs extended for support, her hind legs bent, her hindquarters on the ground, and her abdomen off the ground;
[0156] Side-lying position: The sow lies completely on her side on the ground, with her shoulder, ribs and leg on one side in contact with the ground, her limbs relaxed, her abdomen completely in contact with the ground, and her teats exposed;
[0157] Lying prone: The sow's abdomen is on the ground, her forelegs and hind legs are naturally bent or straightened on both sides of her body, and her head can be raised or lowered.
[0158] Birth: Piglets are still attached to the sow's vulva, with part of their head or hind limbs exposed;
[0159] Piglets: refers to piglets that have completely detached from the sow's vulva;
[0160] After classifying the images, the Stand class has 2000 images, the Sit class has 2000 images, the Sternal class has 1200 images, the Lateral class has 2800 images, the Birth class has 3000 images, and the Piglet class has 3000 images. These images are then divided into training and testing sets in a 9:1 ratio. The training set is further divided into training and validation datasets in an 8:2 ratio. The number of samples in each class after the division is shown in Table 1.
[0161] Table 1. Distribution of samples by category in the dataset
[0162]
[0163]
[0164] As can be seen from the sample quantity distribution of the four categories in Table 1, the number of samples in the Sternal class is significantly less than that in the Lateral class. In order to reduce the impact of class imbalance on the accuracy of model training, this embodiment replaces the Cross Entropy Loss used by the YOLOv8 object detection network with the Adaptive Threshold Focal Loss.
[0165] Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle imbalanced datasets and difficult samples. The formula for calculating Adaptive Threshold Focal Loss is as follows:
[0166]
[0167] Among them, L ATFL The function value of Adaptive Threshold Focal Loss; p represents the predicted value for the next epoch. t This represents the average predicted probability value for the current epoch; λ is a hyperparameter.
[0168] After labeling is completed, this embodiment also performs data augmentation processing to increase the diversity of image samples:
[0169] 1) Select angles that are uniformly distributed in the range of [-180°, 180°], and randomly rotate each frame of the image with a probability of 0.25;
[0170] 2) Randomly flip each frame of the image horizontally with a probability of 0.5;
[0171] 3) Randomly change the brightness, contrast, or saturation of each frame of the image with a probability of 0.5;
[0172] 4) Gaussian noise is randomly added to each frame of the image with a probability of 0.25;
[0173] After the above four steps, the aspect ratio of the initial images is maintained, all images are scaled to obtain the expanded dataset; the improved network model is trained using the expanded dataset, and then the improved model is evaluated using the test set.
[0174] Next, the YOLOv8 model was selected as the base model, and the YOLOv8 model was improved. In this embodiment, DLKA modules were added in front of the three detection heads in the Head part of the YOLOv8 model. The newly added DLKA modules enable the network to extract more detailed features, thereby improving the target detection accuracy.
[0175] This embodiment uses a 1920×1080 image as an example to describe in detail the process from image input to model output. The network first scales the 1920×1080 image down to 640×640, then inputs a 640×640 image into the improved network structure and outputs the detection result. The process is as follows:
[0176] Layer 1: The input image is passed through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1; this layer has 3 input channels (corresponding to the RGB channels of the image) and 64 output channels; the output feature map size is 320×320;
[0177] Layer 2: The output of the first convolutional layer is passed through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1; this layer has 64 input channels and 128 output channels; the size of the output feature map is 160×160;
[0178] Layer 3: The feature map output from the previous layer is passed through the C2f module, which consists of three convolutional layers with 128 input channels and 128 output channels. The first convolutional layer in the C2f module has 128 output channels, and the third convolutional layer has 128 output channels. The size of the output feature map is 160×160.
[0179] Layer 4: The output of the previous layer passes through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and padding of 1; this layer has 128 input channels and 256 output channels; the size of the output feature map is 80×80;
[0180] Layer 5: The feature map obtained from the 4th convolutional layer is input into the C2f module. The C2f module consists of two convolutional layers with 256 input channels and 256 output channels. The size of the output feature map of this layer is 80×80.
[0181] Layer 6: The feature map output from the previous layer is passed through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1; this layer has 256 input channels and 512 output channels; the size of the output feature map is 40×40.
[0182] Layer 7: The feature map output from layer 6 is input into the C2f module, which also consists of three convolutional layers and has 512 input channels and 512 output channels; the output feature map of the C3 module is 40×40 in size.
[0183] Layer 8: The feature map output from the previous layer is input into a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1; this layer has 512 input channels and 512 output channels; the size of the output feature map is 20×20;
[0184] Layer 9: The feature map output from layer 8 is passed through the C2f module; this layer consists of a convolutional layer with 512 input channels and 512 output channels, and the size of the output feature map is 20×20.
[0185] Layer 10: Input a 512-dimensional vector into the SPPF module, which consists of one 512 input channels and 512 output channels; the output feature map is 20×20 in size.
[0186] Then, the feature map output by the C2f module in the BackBone is concatenated with the upsampled feature map and fed into a new C2f module. The three new C2f modules input the feature map into the DLKA module. The DLKA module consists of two standard convolutional modules and two deformable convolutional modules. The standard convolutional modules are input into the deformable convolutional modules, and the deformable convolutional modules are input into the convolutional modules. The input features of the two standard convolutional modules are 256, 512, and 512, respectively. The input size is the same as the feature map output by the C2f module, which is 80×80, 40×40, and 20×20, respectively. Finally, the output is sent to the detection head for localization and classification.
[0187] The above network structure model uses a series of convolutional layers and C2f and DLKA modules to extract features from the input image at different spatial scales, and then performs classification and localization through global average pooling layers and fully connected layers;
[0188] Then, the training set is used to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the loss function is less than or equal to the preset threshold, and the training is completed to obtain the trained sow farrowing behavior monitoring model.
[0189] Finally, the trained sow farrowing behavior monitoring model was used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model was evaluated.
[0190] For real-time monitoring, this embodiment uses a trained sow farrowing behavior monitoring model to detect video data of sows in the test set throughout the farrowing process, and to determine the sow's current posture and farrowing status in real time.
[0191] like Figures 5-8 The images shown are of sows in different poses under normal lighting conditions. Figure 5 The sow is in a stand-up state. Figure 6 The sow is in a sit state. Figure 7 The sow is in a sternal state. Figure 8 The sow is in a latent state; Figures 9-12 They are respectively Figures 5-8 The corresponding test results;
[0192] like Figures 13-15 The images shown are infrared images of sows in different poses collected at night. Figure 13 The sow is in a stand-up state. Figure 14 The sow is in a sit state. Figure 15 The sow is in a latent state; Figures 16-18 They are respectively Figures 13-15 The corresponding test results;
[0193] like Figures 19-20 The image shown is of a piglet being born under normal lighting conditions. Figure 19 The piglets are in the Birth stage. Figure 20 The sow is in Piglet state;
[0194] like Figures 21-22 The image shown is an infrared image of piglets born at night. Figure 21 The piglets are in the Birth stage. Figure 22 The sow is in Piglet state;
[0195] Based on a preset judgment strategy and combined with the detection results of multiple frames, the duration of each pose, the frequency of pose changes, and the birth information of the piglets are recorded, such as... Figures 23-24 The image shows the recording results based on a comprehensive assessment of multiple frames. Figure 23 For recording the birth information of piglets, Figure 24 To record information on changes in the sow's posture;
[0196] This embodiment uses a test video of a sow's farrowing process as an example to illustrate the strategy for judging recorded postures: When the video is input into the trained sow farrowing behavior monitoring model, recording is triggered only if the following conditions are met: the duration of the previous posture is recorded, and the frequency of each posture change so far is calculated:
[0197]
[0198] Where p is the sow's previous stable posture; c is the sow's current posture; q is the number of consecutive frames that maintain the same posture; t is a preset frame threshold; P q The probability that the trained sow farrowing behavior monitoring model detects the current posture; P is a preset probability threshold.
[0199] The pseudocode representation of the above judgment process is as follows:
[0200]
[0201] For example, when a video stream is input, the model detects that the previous state of the sow in the video frame is Lateral, the current state is Sit, and the recording time is End Time. The model calculates the duration and posture change frequency of Lateral; when Sit satisfies q≥t, P q When P ≥ P, record the start time and set the previous state p to c;
[0202] For evaluating model performance, in this embodiment, the metrics used to assess the performance of the trained sow farrowing behavior monitoring model include: accuracy, precision, recall, F1 score, and mean average accuracy (mAP). This embodiment primarily uses mAP for evaluation, and its calculation formula is as follows:
[0203]
[0204] Among them, E AP,IOU E represents the model detection accuracy under the confidence level IoU condition. map This represents the average recognition rate of the multi-objective model.
[0205] To verify the effectiveness of the improvements to the YOLOv8 model, this embodiment uses the YOLOv8 object detection model as the base model and constructs multiple different models, each trained for 300 epochs. The trained models are then used to predict on the test set, and the mAP of each model is shown in Table 2.
[0206] Table 2 Comparison of mAP for different models
[0207] Model Stand Sit Sternal Lateral Birth Piglet mAP YOLOv8 base model 92.1 92.7 90.7 95.3 97.1 90.2 92.7 YOLOv8+DLKA 93.0 94.1 91.8 96.0 98.4 93.1 93.725 YOLOv8+AdLoss 92.7 93.2 91.3 95.6 97.4 90.7 93.2 YOLOv8+DLKA+AdLoss 93.2 94.4 92.1 96.3 98.7 93.4 94
[0208] Note: AdLoss indicates the introduction of Adaptive Threshold Focal Loss as the classification loss;
[0209] As shown in Table 2, the YOLOv8+DLKA+AdLoss model proposed in this embodiment has the best performance;
[0210] This method uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss function, which can effectively improve the accuracy of monitoring sow farrowing behavior. This method can accurately capture the changes in sow posture during farrowing and accurately identify and record the birth of piglets, effectively promoting the pig industry towards intelligence and precision, and providing strong support for the high-quality development of my country's animal husbandry.
[0211] The same or similar labels correspond to the same or similar parts;
[0212] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.
[0213] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for intelligent monitoring of sow farrowing behavior based on an improved YOLOv8 model, characterized in that, Includes the following steps: S1: Obtain and preprocess the sow farrowing behavior dataset, and divide the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows throughout the farrowing process; S2: The YOLOv8 model is selected as the base model and improved. DLKA modules are added in front of the three detectors in the Head part of the YOLOv8 model, and Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detectors to build a sow farrowing behavior monitoring model. The DLKA module comprises, in sequence: a first 2D convolutional layer, a GELU activation layer, a Deform-DW Conv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer; The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual summation connection; the output of the GELU activation layer is multiplied by the output of the second 2D convolutional layer in a weighted manner; The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence. In the Deform-DW Conv2D submodule, the input features are processed by a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position. Based on the output features of the Deform-DW Conv2D submodule, an inflation value is added and used as the input feature of the Deform-DW-DConv2D submodule; The Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle imbalanced datasets and difficult samples. The calculation formula for Adaptive Threshold Focal Loss is as follows: in, The function value of Adaptive Threshold Focal Loss; This represents the predicted value for the next epoch. This represents the average predicted probability value for the current epoch. For hyperparameters; S3: Use the training set to iteratively optimize and train the sow farrowing behavior monitoring model until the value of the AdaptiveThreshold Focal Loss is less than or equal to the preset threshold, and then complete the training to obtain the trained sow farrowing behavior monitoring model. S4: Use the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluate the performance of the trained sow farrowing behavior monitoring model.
2. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 1, characterized in that, Step S1 includes: Acquire daily lighting videos and nighttime infrared videos of several sows during farrowing, extract frames according to video time sequence, and filter out blurry and highly similar video frames to obtain the filtered dataset; The preprocessing process involves labeling each frame of the filtered dataset and performing data augmentation. The preprocessed dataset is divided into training and test sets according to a preset ratio.
3. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 2, characterized in that, The filtering of blurry and highly similar video frames includes the following steps: S1.1: Convert the images of each frame obtained by video frame extraction into floating-point images; S1.2: All floating-point images are processed using a Laplace filter to highlight high-frequency components in each frame; the calculation formula for the Laplace filter is as follows: in, Image location The pixel value output after passing through the Laplace filter. It is a location The input pixel value at that location, It is a location Gaussian kernel value at the location, It is the width of the image. It is the height of the image; A higher value indicates a clearer image; S1.3: Divide all images into several groups at certain frame intervals, and filter out images in each group. The k frames with the lowest values are used to obtain the dataset after filtering out the blurred frames, where k is a positive integer less than the frame number interval; S1.4: Using the SSIM algorithm, calculate the similarity between each frame image in the dataset after filtering blurred frames, and filter one of the two images with a similarity greater than or equal to a preset similarity threshold to obtain the filtered dataset.
4. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 2, characterized in that, The label includes any of the following: Standing: The sow's limbs are straight to support her body, and her abdomen is off the ground; Sitting posture: The sow sits on the ground with her forelegs extended for support, her hind legs bent, her hindquarters on the ground, and her abdomen off the ground; Side-lying position: The sow lies completely on her side on the ground, with her shoulder, ribs and leg on one side in contact with the ground, her limbs relaxed, her abdomen completely in contact with the ground, and her teats exposed; Lying prone: The sow's abdomen is on the ground, her forelegs and hind legs are naturally bent or straightened on both sides of her body, and her head can be raised or lowered. Birth: Piglets are still attached to the sow's vulva, with part of their head or hind limbs exposed. Piglets: refers to piglets that have completely detached from the sow's vulva.
5. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 2, characterized in that, The data augmentation process includes any one or more of the following: Take in [-180] ° 180 ° The angles are evenly distributed within the interval, and each frame of the image is randomly rotated; Randomly flip each frame of the image horizontally; Randomly change the brightness, contrast, or saturation of each frame of the image; Gaussian noise is randomly added to each frame of the image.
6. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 1, characterized in that, In step S4, real-time monitoring of the sow's farrowing behavior includes: The trained sow farrowing behavior monitoring model was used to analyze video data of sows in the test set throughout the farrowing process to determine the sow's current posture and farrowing status in real time. Based on the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of piglets are recorded.
7. The intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model according to claim 6, characterized in that, The preset judgment strategy includes: When videos from the test set are input into the trained sow farrowing behavior monitoring model, recording is triggered only if the following conditions are met: the duration of the previous posture is recorded, and the frequency of each posture change so far is calculated: in, To get the sow into a stable position; Current posture of the sow; The number of frames that continuously maintain the same posture; This is a preset frame rate threshold; The probability of the current posture being detected by the trained sow farrowing behavior monitoring model; This is a preset probability threshold.
8. A method for intelligent monitoring of sow farrowing behavior based on an improved YOLOv8 model according to any one of claims 1 to 7, characterized in that, In step S4, the metrics used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1 score, and mean average recognition rate (mAP).