Sow delivery behavior intelligent monitoring method based on improved YOLOv8 model
By improving the YOLOv8 model, combining the DLKA module and Adaptive Threshold Focal Loss, accurate monitoring of sow birth behavior and accurate recording of piglet births are achieved, the problem of inefficient monitoring in the existing technology is solved, and the intelligent and refined development of the pig farming industry has been promoted.
Patent Information
- Application Number
- CN202510024118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-07
AI Technical Summary
When monitoring the delivery behavior of sows, the manual observation method is time-consuming and labor-intensive and inefficient, while the sensor monitoring effect is poor and cannot meet the needs of modern large-scale pig farming.
Using the improved YOLOv8 model, by adding DLKA module in front of the three detection heads in the head part of the YOLOv8 model and introducing Adaptive Threshold Focal Loss, a sow birth behavior monitoring model is constructed to achieve accurate capture of sow birth behavior and accurate identification and recording of piglet birth conditions.
The accuracy of sow childbirth behavior monitoring has been improved, automatic, accurate and real-time posture change capture and piglet birth record have been achieved, labor burden has been reduced, production management efficiency has been improved, and pig farming has been promoted to develop towards intelligence and refinement.
Smart Images

Figure CN119992648A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more specifically, to an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model. Background Art
[0002] my country is a major livestock and poultry breeding country. Pig farming is an important part of my country's breeding industry, providing an important driving force for the development of my country's animal husbandry. The sow's farrowing process is a key stage in pig production, and its management quality is directly related to the survival rate of piglets and the health level of sows. In farrowing management, real-time recording and analysis of the frequency and duration of sow posture changes are of great significance for evaluating farrowing progress, judging the comfort of sows, and early intervention of farrowing abnormalities.
[0003] Sows may show a variety of posture changes during the birthing process, such as standing, sitting, lying on the side and lying prone. The duration and switching frequency of different postures reflect the sow's response to pain, anxiety and fatigue. Studies have shown that sows usually remain in a side-lying or prone position for a long time to complete birthing and lactation smoothly, while frequent standing or sitting may be a signal of pain or discomfort. Therefore, accurate statistics and analysis of posture data during birthing can help identify potential birthing difficulties and provide farmers with a scientific management basis.
[0004] In addition, it is also very important to record the birth time, order and health status of the piglets. These data can reflect the progress of the sow's delivery and the overall health level of the piglets. By recording the birth time of each piglet, it is possible to determine whether the delivery is smooth and whether the interval is abnormal, so as to detect and respond to possible problems such as dystocia or fetal distress early. The initial health status of the piglets after birth also provides an important reference for subsequent feeding and management, ensuring that weak piglets can receive timely care and support.
[0005] At present, the monitoring of sow farrowing behavior in production is mainly through manual observation. However, manual observation increases the time that humans and animals spend together, has the risk of zoonotic diseases, and is also affected by the subjective experience of breeders. It is time-consuming and labor-intensive and cannot meet the needs of modern large-scale pig farming. In addition, wearable devices, photocells, ultrasonic and other sensor technologies can also be used to monitor sow farrowing behavior. Although labor is reduced, wearable devices may cause stress reactions in pigs, the power supply time is limited and they are easy to fall off. Photocells and ultrasonic waves are easily affected by the surrounding environment and have relatively low sensitivity, so the monitoring effect is still not ideal. Summary of the invention
[0006] In order to overcome the defects of the manual observation method in the above-mentioned prior art, which is time-consuming, labor-intensive, inefficient, and has poor sensor monitoring effect, the present invention provides an intelligent monitoring method for sow delivery behavior based on an improved YOLOv8 model, which can accurately capture the posture changes of sows during delivery, and accurately identify and record the birth of piglets, effectively promoting the pig farming industry to move towards intelligence and refinement, and providing strong support for the high-quality development of my country's animal husbandry.
[0007] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0008] An intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model comprises the following steps:
[0009] S1: Obtaining a sow farrowing behavior dataset and preprocessing it, dividing the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows in the entire farrowing process;
[0010] S2: The YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved. The DLKA module is added before the three detection heads of the Head part of the YOLOv8 model, and the Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detection heads to build a sow farrowing behavior monitoring model;
[0011] S3: using the training set to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to a preset threshold, and obtaining a trained sow farrowing behavior monitoring model;
[0012] S4: Using the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluating the performance of the trained sow farrowing behavior monitoring model.
[0013] Preferably, the step S1 comprises:
[0014] Obtain daily illumination videos and night infrared videos of the farrowing process of several sows, extract frames according to the video time sequence, and filter out blurred and highly similar video frames to obtain a filtered data set;
[0015] Label each frame image in the filtered data set and perform data enhancement to complete preprocessing;
[0016] The preprocessed data set is divided into training set and test set according to the preset ratio.
[0017] Preferably, filtering blurred and highly similar video frames comprises the following steps:
[0018] S1.1: Convert each frame image obtained by video frame extraction into a floating point image;
[0019] S1.2: All floating point images are processed using a Laplace filter to highlight the high frequency components in each frame of the image; the calculation formula of the Laplace filter is as follows:
[0020]
[0021] Where L(x,y) is the pixel value output by the Laplacian filter at the image position (x,y), f(i,j) is the input pixel value at the position (i,j), G(i,j) is the Gaussian kernel value at the position (i,j), w is the width of the image, and h is the height of the image; the higher the L(x,y) value, the clearer the image;
[0022] S1.3: All images are continuously divided into several groups at a certain frame interval, and the k frames with the lowest L(x, y) values in each group of images are filtered out to obtain a data set after filtering the blurred frames, where k is a positive integer less than the frame interval;
[0023] S1.4: Using the SSIM algorithm, the similarities between the frames in the data set after filtering the blurred frames are calculated pairwise, and one of the two images whose similarity is greater than or equal to a preset similarity threshold is filtered to obtain the filtered data set.
[0024] Preferably, the label includes any one of the following:
[0025] Standing: The sow stretches her limbs to support her body, with her abdomen off the ground;
[0026] Sitting: The sow's front legs are straightened for support, and the hind legs are bent to sit on the ground, with the back half of the body on the ground and the abdomen off the ground;
[0027] Side-lying: The sow lies completely on the ground on her side, with the shoulder, ribs and legs on one side touching the ground, the limbs relaxed, the abdomen completely touching the ground, and the teats exposed;
[0028] Lying prone: The sow's belly is close to the ground, the front and hind legs are naturally bent or straightened on both sides of the body, and the head can be raised or lowered;
[0029] Birth: The piglet has not separated from the sow's vulva, with part of the head or hind limbs exposed;
[0030] Piglet: refers to a piglet that has completely separated from the sow's vulva.
[0031] Preferably, the data enhancement processing includes any one or more of the following:
[0032] Take angles evenly distributed in the interval [-180°, 180°] and randomly rotate each frame image;
[0033] Randomly flip each frame image horizontally;
[0034] Randomly change the brightness, contrast or saturation of each frame;
[0035] Gaussian noise is randomly added to each frame image.
[0036] Preferably, in step S2, the DLKA module includes: a first 2D convolutional layer, a GELU activation layer, a Deform-DW Conv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer and a third 2D convolutional layer connected in sequence;
[0037] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual sum connection; the output of the GELU activation layer is weighted multiplied with the output of the second 2D convolutional layer;
[0038] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence;
[0039] In the Deform-DW Conv2D submodule, the input features are passed through a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position;
[0040] Based on the output features of the Deform-DW Conv2D submodule, the expansion value is added as the input feature of the Deform-DW-D Conv2D submodule.
[0041] Preferably, in step S2, the Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively process unbalanced data sets and difficult samples. The calculation formula of AdaptiveThreshold Focal Loss is as follows:
[0042]
[0043] Among them, L ATFL is the function value of Adaptive Threshold Focal Loss; Represents the predicted value of the next epoch, p t Represents the average predicted probability value of the current epoch; λ is a hyperparameter.
[0044] Preferably, in step S4, real-time monitoring of the sow's farrowing behavior includes:
[0045] Use the trained sow farrowing behavior monitoring model to detect the video data of the sows in the test set throughout the farrowing process, and judge the sow’s current posture and farrowing status in real time;
[0046] According to the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of the piglets are recorded.
[0047] Preferably, the preset judgment strategy includes:
[0048] When the video in the test set is input into the trained sow farrowing behavior monitoring model, the recording is triggered only when the following conditions are met, and the duration of the last posture is recorded, and the frequency of each posture change so far is calculated:
[0049]
[0050] Among them, p is the previous stable posture of the sow; c is the current posture of the sow; q is the number of frames continuously detected to maintain the same posture; t is the preset frame number threshold; P q is the probability of the current posture detected by the trained sow farrowing behavior monitoring model; P is the preset probability threshold.
[0051] Preferably, in step S4, the indicators used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1-Score and mean average recognition rate (mAP).
[0052] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0053] The present invention provides an intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model. First, a sow farrowing behavior data set is obtained and preprocessed, and the preprocessed sow farrowing behavior data set is divided into a training set and a test set; then, the YOLOv8 model is selected as a basic model, and the YOLOv8 model is improved, DLKA modules are respectively added in front of three detection heads of the Head part of the YOLOv8 model, and Adaptive Threshold Focal Loss is introduced as a classification loss to construct a sow farrowing behavior monitoring model; then, the sow farrowing behavior monitoring model is iteratively optimized and trained using the training set, and the training is completed when the value of the loss function is less than or equal to a preset threshold, so as to obtain a trained sow farrowing behavior monitoring model; finally, the trained sow farrowing behavior monitoring model is used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model is evaluated;
[0054] The present invention has the following advantages:
[0055] 1) The present invention uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss loss function, which can effectively improve the accuracy of sow farrowing behavior monitoring;
[0056] 2) The present invention utilizes computer vision technology to not only automatically, accurately and in real time capture the posture changes of sows, but also identify and record the birth of piglets, thereby reducing labor and improving the efficiency of production management;
[0057] 3) The present invention can realize 24-hour uninterrupted monitoring. Based on the data obtained by real-time monitoring of the present invention, not only the dynamic pattern of changes in the sow's delivery posture can be drawn, but also the timeline and health indicators of piglet birth can be intuitively presented, providing a comprehensive management basis for breeders;
[0058] 4) The present invention only needs to collect video data of sows’ delivery behavior during monitoring. The contactless monitoring will not cause stress response in pigs. It not only improves the sows’ delivery experience and the survival rate of piglets, but also promotes the pig farming industry to move towards intelligence and refinement, providing strong support for the high-quality development of my country’s animal husbandry. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a flow chart of a method for intelligently monitoring sow farrowing behavior based on an improved YOLOv8 model provided in Example 1.
[0060] Figure 2 This is a diagram of the YOLOv8 model structure provided in Example 1.
[0061] Figure 3 This is a structural diagram of the YOLOv8 model after adding the DLKA module provided in Example 1.
[0062] Figure 4 This is a structural diagram of the DLKA module provided in Example 1.
[0063] Figures 5 to 8 They are respectively images of a sow in the Stand state, Sit state, Sternal state and Lateral state under daily lighting conditions provided in Example 2.
[0064] Figures 9 to 12 They are the detection results of the images of the sow in the Stand state, Sit state, Sternal state and Lateral state under the daily lighting conditions provided in Example 2.
[0065] Figures 13-15 They are infrared images of the sow in the Stand state, Sit state and Lateral state collected at night provided in Example 2 respectively.
[0066] Figures 16 to 18 They are the detection results of the infrared images of the sow in the Stand state, Sit state and Lateral state collected at night provided in Example 2.
[0067] Fig.19 This is an image of a piglet in the Birth state under daily lighting conditions provided in Example 2.
[0068] Fig. 20 This is an image of a sow in the Piglet state under daily lighting conditions provided in Example 2.
[0069] Fig.21 This is the infrared image of the piglet in the Birth state collected at night provided in Example 2.
[0070] Fig. 22 This is the infrared image of a sow in the Piglet state collected at night provided in Example 2.
[0071] Fig.23 This is the record result of the piglet birth information provided in Example 2.
[0072] Fig.24 This is the result of recording the sow posture change information provided in Example 2. DETAILED DESCRIPTION
[0073] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0074] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0075] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0076] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0077] Example 1
[0078] like Figure 1 As shown, this embodiment provides a method for intelligent monitoring of sow farrowing behavior based on an improved YOLOv8 model, comprising the following steps:
[0079] S1: Obtaining a sow farrowing behavior dataset and preprocessing it, dividing the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows in the entire farrowing process;
[0080] S2: The YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved. The DLKA module is added before the three detection heads of the Head part of the YOLOv8 model, and the Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detection heads to build a sow farrowing behavior monitoring model;
[0081] S3: using the training set to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to a preset threshold, and obtaining a trained sow farrowing behavior monitoring model;
[0082] S4: Using the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluating the performance of the trained sow farrowing behavior monitoring model.
[0083] In the specific implementation process, firstly, a sow farrowing behavior dataset is obtained and preprocessed. In this embodiment, the sow farrowing behavior dataset includes video data of several sows before, during and after farrowing. After the preprocessing is completed, the preprocessed sow farrowing behavior dataset is divided into a training set and a test set.
[0084] Then the YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved;
[0085] like Figure 2As shown in the figure, the YOLOv8 series of neural network models include YOLOv8n, YOLOv8s, YOLOv8m, and YOLOv8l models; the YOLOv8 model architecture is based on a convolutional neural network (CNN) with various convolutional layer backbones; the backbone part (BackBone) consists of a series of feature extraction layers, which are responsible for extracting features from the input image; these features are then sent to the Neck part for processing, and finally sent to the three groups of detection heads in the Head part, which use a group of convolutional layers and fully connected layers for target detection and classification; Figure 2 In the figure, the Conv module is composed of 3×3 Conv2d, BatchNorm and SiLU activation functions, the C2f module is composed of 3 ConvModules and n Bottlenecks (with residual connections); Bottleneck is composed of 2 ConvModules;
[0086] The improved YOLOv8 target detection network model is as follows Figure 3 As shown in the figure, DLKA (Deformable Large Kernel Attention) modules are added before the three detection heads of the Head part of the YOLOv8 model. The newly added DLKA module enables the network to extract more detailed features, thereby improving the target detection accuracy;
[0087] like Figure 4 As shown, the DLKA module includes: a first 2D convolutional layer, a GELU activation layer, a Deform-DWConv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer connected in sequence;
[0088] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual sum connection; the output of the GELU activation layer is weighted multiplied with the output of the second 2D convolutional layer;
[0089] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence;
[0090] In the Deform-DW Conv2D submodule, the input features are passed through a 3×3 convolution layer to calculate the offset field, and the offset field is used to guide the deformable convolution layer to adjust its sampling position; that is, an offset is first generated through a standard convolution learning, and this offset is continuously adjusted in the later learning. With the offset, the convolution kernel can adaptively adjust its position, so that the convolution operation can more accurately select important local areas in the image for feature extraction, rather than the fixed grid position in the traditional convolution;
[0091] Based on the output features of the Deform-DW Conv2D submodule, the expansion value is added as the input features of the Deform-DW-D Conv2D submodule;
[0092] The DLKA module adaptively adjusts the shape of the convolution kernel according to the local structure of the feature map. The deformable convolution allows each convolution kernel to shift, so that it can sample different local areas at each position of the feature map, and then assign weights to different areas in the feature map according to the attention mechanism, highlighting the key areas and weakening unimportant information. The input and output of DLKA are expressed by the following mathematical formula:
[0093] x out =Conv2D(x 1 *D-LKA-Attn(x 1 ))+x in
[0094] x 1 =GELU(Conv2D(x in ))
[0095] D-LKA-Attn(x 1 )=Conv2D(DeformDConv2D(DeformConv2D(x 1 )))
[0096] Among them, x in is the input of the DLKA module; x out is the output of the DLKA module; x 1 is the intermediate feature; D-LKA-Attn is the deformable LKA attention; Conv2D is the 2D convolution; GeLU is the GeLU activation function; DeformConv2D is the calculation process of the Deform-DWConv2D submodule; DeformDConv2D is the calculation process of the Deform-DW-D Conv2D submodule; DeformDConv2D is an extension of DeformConv2D, which is a convolution with the dilation technology added;
[0097] After improving the YOLOv8 target detection network model, Adaptive Threshold FocalLoss was introduced to replace the original classification loss of the three detection heads, and other loss functions remained unchanged to build a sow farrowing behavior monitoring model; this loss function is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle unbalanced data sets and difficult samples;
[0098] Then, the training set is used to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to the preset threshold, and the trained sow farrowing behavior monitoring model is obtained;
[0099] Finally, the trained sow farrowing behavior monitoring model was used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model was evaluated;
[0100] This method uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss loss function, which can effectively improve the accuracy of sow farrowing behavior monitoring; this method can accurately capture the posture changes of sows during farrowing, and accurately identify and record the birth of piglets, effectively promoting the pig farming industry to move towards intelligence and refinement, and providing strong support for the high-quality development of my country's animal husbandry.
[0101] Example 2
[0102] This embodiment provides a method for intelligently monitoring sow farrowing behavior based on an improved YOLOv8 model, comprising the following steps:
[0103] S1: Obtaining a sow farrowing behavior dataset and preprocessing it, dividing the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows in the entire farrowing process;
[0104] S2: The YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved. The DLKA module is added before the three detection heads of the Head part of the YOLOv8 model, and the Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detection heads to build a sow farrowing behavior monitoring model;
[0105] S3: using the training set to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to a preset threshold, and obtaining a trained sow farrowing behavior monitoring model;
[0106] S4: using the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluating the performance of the trained sow farrowing behavior monitoring model;
[0107] The step S1 comprises:
[0108] Obtain daily illumination videos and night infrared videos of the farrowing process of several sows, extract frames according to the video time sequence, and filter out blurred and highly similar video frames to obtain a filtered data set;
[0109] Label each frame image in the filtered data set and perform data enhancement to complete preprocessing;
[0110] Divide the preprocessed data set into a training set and a test set according to a preset ratio;
[0111] The filtering of blurred and highly similar video frames comprises the following steps:
[0112] S1.1: Convert each frame image obtained by video frame extraction into a floating point image;
[0113] S1.2: All floating point images are processed using a Laplace filter to highlight the high frequency components in each frame of the image; the calculation formula of the Laplace filter is as follows:
[0114]
[0115] Where L(x,y) is the pixel value output by the Laplacian filter at the image position (x,y), f(i,j) is the input pixel value at the position (i,j), G(i,j) is the Gaussian kernel value at the position (i,j), w is the width of the image, and h is the height of the image; the higher the L(x,y) value, the clearer the image;
[0116] S1.3: All images are continuously divided into several groups at a certain frame interval, and the k frames with the lowest L(x, y) values in each group of images are filtered out to obtain a data set after filtering the blurred frames, where k is a positive integer less than the frame interval;
[0117] S1.4: Using the SSIM algorithm, the similarity between the frames in the data set after filtering the blurred frames is calculated pairwise, and one of the two images whose similarity is greater than or equal to a preset similarity threshold is filtered to obtain the filtered data set;
[0118] The tag includes any of the following:
[0119] Standing: The sow stretches her limbs to support her body, with her abdomen off the ground;
[0120] Sitting: The sow's front legs are straightened for support, and the hind legs are bent to sit on the ground, with the back half of the body on the ground and the abdomen off the ground;
[0121] Side-lying: The sow lies completely on the ground on her side, with the shoulder, ribs and legs on one side touching the ground, the limbs relaxed, the abdomen completely touching the ground, and the teats exposed;
[0122] Lying prone: The sow's belly is close to the ground, the front and hind legs are naturally bent or straightened on both sides of the body, and the head can be raised or lowered;
[0123] Birth: The piglet has not separated from the sow's vulva, with part of the head or hind limbs exposed;
[0124] Piglet: refers to the piglet that has completely separated from the sow's vulva;
[0125] The data enhancement processing includes any one or more of the following:
[0126] Take angles evenly distributed in the interval [-180°, 180°] and randomly rotate each frame image;
[0127] Randomly flip each frame image horizontally;
[0128] Randomly change the brightness, contrast or saturation of each frame;
[0129] Gaussian noise is randomly added to each frame image;
[0130] In step S2, the DLKA module includes: a first 2D convolutional layer, a GELU activation layer, a Deform-DWConv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer connected in sequence;
[0131] The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual sum connection; the output of the GELU activation layer is weighted multiplied with the output of the second 2D convolutional layer;
[0132] The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence;
[0133] In the Deform-DW Conv2D submodule, the input features are passed through a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position;
[0134] Based on the output features of the Deform-DW Conv2D submodule, the expansion value is added as the input features of the Deform-DW-D Conv2D submodule;
[0135] In step S2, the Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively process unbalanced data sets and difficult samples. The calculation formula of Adaptive Threshold Focal Loss is as follows:
[0136]
[0137] Among them, L ATFL is the function value of Adaptive Threshold Focal Loss; Represents the predicted value of the next epoch, p t Represents the average predicted probability value of the current epoch; λ is a hyperparameter;
[0138] In step S4, real-time monitoring of the sow's farrowing behavior includes:
[0139] Use the trained sow farrowing behavior monitoring model to detect the video data of the sows in the test set throughout the farrowing process, and judge the sow’s current posture and farrowing status in real time;
[0140] According to the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of the piglets are recorded;
[0141] The preset judgment strategy includes:
[0142] When the video in the test set is input into the trained sow farrowing behavior monitoring model, the recording is triggered only when the following conditions are met, and the duration of the last posture is recorded, and the frequency of each posture change so far is calculated:
[0143]
[0144] Among them, p is the previous stable posture of the sow; c is the current posture of the sow; q is the number of frames continuously detected to maintain the same posture; t is the preset frame number threshold; P q is the probability of the current posture detected by the trained sow farrowing behavior monitoring model; P is the preset probability threshold;
[0145] In step S4, the indicators used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1-Score and average recognition rate (mAP).
[0146] In the specific implementation process, firstly, a sow farrowing behavior dataset is obtained and preprocessed. In this embodiment, the sow farrowing behavior dataset includes video data of several sows before, during and after farrowing;
[0147] Then, the daily illumination videos and night infrared videos of several sows’ farrowing process are extracted according to the video time sequence, and the blurred and highly similar video frames are filtered out, including the following steps:
[0148] Each frame image obtained by video frame extraction is converted into a 64-bit floating-point image; all floating-point images are processed using a Laplace filter to highlight the high-frequency components in each frame image; the calculation formula of the Laplace filter is as follows:
[0149]
[0150] Where L(x,y) is the pixel value output by the Laplacian filter at the image position (x,y), f(i,j) is the input pixel value at the position (i,j), G(i,j) is the Gaussian kernel value at the position (i,j), w is the width of the image, and h is the height of the image; the higher the L(x,y) value, the more details and the clearer the image;
[0151] All images are continuously divided into several groups at intervals of 10 frames, and the two frames with the lowest L(x,y) values in every 10 frames are filtered out to obtain the data set after filtering the blurred frames;
[0152] Using the SSIM (Structure Similarity Index Measure) algorithm, the similarity between the frames in the data set after filtering the blurred frames is calculated pairwise, and one of the two images whose similarity is greater than or equal to a preset similarity threshold (0.9 in this embodiment) is filtered to obtain a filtered data set;
[0153] After screening, 14,000 video frames were obtained with a frame size of 1920×1080. Each frame was labeled and divided into six categories according to the posture of the sow and the birth of the piglets. The categories include: Stand, Sit, Lateral, Sternal, Birth and Piglet. The definition of each category is as follows:
[0154] Standing: The sow stretches her limbs to support her body, with her abdomen off the ground;
[0155] Sitting: The sow's front legs are straightened for support, and the hind legs are bent to sit on the ground, with the back half of the body on the ground and the abdomen off the ground;
[0156] Side-lying: The sow lies completely on the ground on her side, with the shoulder, ribs and legs on one side touching the ground, the limbs relaxed, the abdomen completely touching the ground, and the teats exposed;
[0157] Lying prone: The sow's belly is close to the ground, the front and hind legs are naturally bent or straightened on both sides of the body, and the head can be raised or lowered;
[0158] Birth: The piglet has not separated from the sow's vulva, with part of the head or hind limbs exposed;
[0159] Piglet: refers to the piglet that has completely separated from the sow's vulva;
[0160] After the classification, the Stand class has 2000 pictures, the Sit class has 2000 pictures, the Sternal class has 1200 pictures, the Lateral class has 2800 pictures, the Birth class has 3000 pictures, and the Piglet class has 3000 pictures; these pictures are divided into training set and test set at a ratio of 9:1, and the training set is further divided into training data set and verification data set at a ratio of 8:2. The number of samples in each category after the division is shown in Table 1;
[0161] Table 1 Distribution of samples in each category of the dataset
[0162]
[0163]
[0164] From the sample number distribution of the four categories in Table 1, it can be seen that the number of samples in the Sternal class is significantly less than that in the Lateral class. In order to reduce the impact of class sample imbalance on the accuracy of model training, this embodiment replaces the Cross Entropy Loss used in the YOLOv8 target detection network with Adaptive Threshold Focal Loss.
[0165] Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively handle unbalanced data sets and difficult samples. The calculation formula of Adaptive Threshold Focal Loss is as follows:
[0166]
[0167] Among them, L ATFL is the function value of Adaptive Threshold Focal Loss; Represents the predicted value of the next epoch, p t Represents the average predicted probability value of the current epoch; λ is a hyperparameter;
[0168] After labeling is completed, this embodiment also performs data enhancement processing to increase the diversity of image samples:
[0169] 1) Take an angle uniformly distributed in the interval [-180°, 180°] and randomly rotate each frame image with a probability of 0.25;
[0170] 2) Randomly flip each frame horizontally with a probability of 0.5;
[0171] 3) Randomly change the brightness, contrast or saturation of each frame image with a probability of 0.5;
[0172] 4) Randomly add Gaussian noise to each frame image with a probability of 0.25;
[0173] After the above four steps, the aspect ratio of the initial image is maintained, and all images are scaled to obtain an expanded data set. The improved network model is trained using the expanded data set, and the improved model is evaluated using the test set.
[0174] Then, the YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved. In this embodiment, DLKA modules are added before the three detection heads of the Head part of the YOLOv8 model. The newly added DLKA modules enable the network to extract more detailed features, thereby improving the target detection accuracy.
[0175] This embodiment takes an image of 1920×1080 as an example to describe in detail the process from image input model to result output; the network will first scale 1920×1080 to 640×640, then input an image of 640×640 into the improved network structure, and output the detection result. The process is as follows:
[0176] Layer 1: The input image passes through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1. This layer has 3 input channels (corresponding to the RGB channels of the image) and 64 output channels. The size of the output feature map is 320×320.
[0177] Layer 2: The output of the first convolutional layer is passed through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1. This layer has 64 input channels and 128 output channels. The size of the output feature map is 160×160.
[0178] Layer 3: The feature map output by the previous layer is passed through the C2f module, which consists of three convolutional layers with 128 input channels and 128 output channels. The first convolutional layer in the C2f module has 128 output channels, the third convolutional layer has 128 output channels, and the size of the output feature map is 160×160.
[0179] Layer 4: The output of the previous layer passes through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1; this layer has 128 input channels and 256 output channels; the size of the output feature map is 80×80;
[0180] Layer 5: The feature map obtained by the 4th convolutional layer is input to the C2f module, which consists of two convolutional layers with 256 input channels and 256 output channels. The size of the feature map output by this layer is 80×80.
[0181] Layer 6: The feature map output by the previous layer is passed through a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1. This layer has 256 input channels and 512 output channels. The size of the output feature map is 40×40.
[0182] Layer 7: The feature map output from layer 6 is input to the C2f module, which also consists of three convolutional layers and has 512 input channels and 512 output channels; the size of the output feature map of the C3 module is 40×40;
[0183] Layer 8: The feature map output by the previous layer is input into a 2D convolutional layer with a kernel size of 3×3, a stride of 2, and a padding of 1. This layer has 512 input channels and 512 output channels. The size of the output feature map is 20×20.
[0184] Layer 9: Pass the feature map output from layer 8 through the C2f module; this layer consists of a convolutional layer with 512 input channels and 512 output channels, and the size of the output feature map is 20×20;
[0185] Layer 10: The 512-dimensional vector is input into the SPPF module, which consists of a 512 input channels and 512 output channels; the size of the output feature map is 20×20;
[0186] After that, the feature map output by the C2f module in BackBone will be concatenated with the upsampled output feature map and sent to the new C2f. The three new C2fs input the feature map to the DLKA module. The DLKA module consists of two standard convolution modules and two deformable convolutions. The standard convolution is input to the deformable convolution, and the deformable convolution is input to the convolution. The input features of the two standard convolutions are 256, 512, and 512 respectively. The input size is the same as the feature map output by C2f, which are 80×80, 40×40, and 20×20 respectively. Finally, it is output to the detection head for positioning and classification.
[0187] The above network structure model uses a series of convolutional layers and C2f and DLKA modules to extract features from the input image at different spatial scales, and then classifies and locates them through global average pooling layers and fully connected layers;
[0188] Then, the training set is used to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to the preset threshold, and the trained sow farrowing behavior monitoring model is obtained;
[0189] Finally, the trained sow farrowing behavior monitoring model was used to monitor the farrowing behavior of sows in the test set in real time, and the performance of the trained sow farrowing behavior monitoring model was evaluated;
[0190] For real-time monitoring, this embodiment uses the trained sow farrowing behavior monitoring model to detect the video data of the sows in the test set during the entire farrowing process, and determines the current posture and farrowing status of the sows in real time;
[0191] like Figures 5 to 8 As shown in the figure, they are images of sows in different postures under daily lighting. Figure 5 The sow is in the Stand state. Figure 6 The sow is in Sit state. Figure 7 The sow is in the Sternal state. Figure 8 The sow is in the Lateral state; Figures 9 to 12 They are Figures 5 to 8 The corresponding test results;
[0192] like Figures 13-15 As shown in the figure, they are infrared images of sows in different postures collected at night. Fig.13 The sow is in the Stand state. Fig.14 The sow is in Sit state. Fig.15 The sow is in the Lateral state; Figures 16 to 18 They are Figures 13-15The corresponding test results;
[0193] like Figures 19-20 The following is an image of piglets being born under normal lighting conditions. Fig.19 The piglets are in the Birth state. Fig. 20 The sow is in the Piglet state;
[0194] like Figures 21-22 The following is an infrared image of piglets being born, taken at night. Fig.21 The piglets are in the Birth state. Fig. 22 The sow is in the Piglet state;
[0195] According to the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of the piglets are recorded, such as Figures 23-24 As shown, it is the record result of comprehensive multi-frame judgment. Fig.23 To record the birth information of piglets, Fig.24 To record the sow’s posture changes;
[0196] This embodiment takes a test video of a sow's delivery process as an example to illustrate the judgment strategy for recording postures: when the video inputs the trained sow delivery behavior monitoring model, the recording is triggered only when the following conditions are met, and the duration of the previous posture is recorded, and the frequency of each posture change so far is calculated:
[0197]
[0198] Among them, p is the previous stable posture of the sow; c is the current posture of the sow; q is the number of frames continuously detected to maintain the same posture; t is the preset frame number threshold; P q is the probability of the current posture detected by the trained sow farrowing behavior monitoring model; P is the preset probability threshold;
[0199] The pseudo code of the above judgment process is expressed as:
[0200]
[0201] For example, when the video stream is input, the model detects that the previous state of the sow in the video frame is Lateral, the current state is Sit, and the recording time is End Time. The duration of Lateral and the frequency of posture change are calculated; when the Sit class satisfies q≥t,P q ≥P, record the Start Time and set the previous state p to c;
[0202] For the evaluation of model performance, in this embodiment, the indicators used to evaluate the performance of the trained sow farrowing behavior monitoring model include: accuracy, precision, recall, F1-Score and average recognition rate mAP. In this embodiment, the evaluation is mainly performed by the mAP indicator, and its calculation formula is as follows:
[0203]
[0204] Among them, E AP,IOU Indicates the model detection accuracy under the confidence IoU condition, E map is the average recognition rate of the multi-target model;
[0205] In order to verify the effectiveness of the improvement of the YOLOv8 model, this example uses the YOLOv8 target detection model as the basic model to build multiple different models, and each model is trained for 300 epochs; the trained models are used to predict the test set, and the mAP of each model is shown in Table 2:
[0206] Table 2 Comparison of mAP of different models
[0207] Model Stand Sit Sternal Lateral Birth Piglet mAP YOLOv8 basic model 92.1 92.7 90.7 95.3 97.1 90.2 92.7 YOLOv8+DLKA 93.0 94.1 91.8 96.0 98.4 93.1 93.725 YOLOv8+AdLoss 92.7 93.2 91.3 95.6 97.4 90.7 93.2 YOLOv8+DLKA+AdLoss 93.2 94.4 92.1 96.3 98.7 93.4 94
[0208] Note: AdLoss means the introduction of Adaptive Threshold Focal Loss as classification loss;
[0209] As can be seen from Table 2, the YOLOv8+DLKA+AdLoss model proposed in this embodiment has the best performance;
[0210] This method uses the DLKA module to improve the network structure of the traditional YOLOv8 model and introduces the AdaptiveThreshold Focal Loss loss function, which can effectively improve the accuracy of sow farrowing behavior monitoring; this method can accurately capture the posture changes of sows during farrowing, and accurately identify and record the birth of piglets, effectively promoting the pig farming industry to move towards intelligence and refinement, and providing strong support for the high-quality development of my country's animal husbandry.
[0211] The same or similar reference numerals correspond to the same or similar components;
[0212] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;
[0213] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. An intelligent monitoring method for sow farrowing behavior based on an improved YOLOv8 model, characterized in that: The following steps are involved: S1: Obtaining a sow farrowing behavior dataset and preprocessing it, dividing the preprocessed sow farrowing behavior dataset into a training set and a test set; the sow farrowing behavior dataset includes video data of several sows in the entire farrowing process; S2: The YOLOv8 model is selected as the basic model, and the YOLOv8 model is improved. The DLKA module is added before the three detection heads of the Head part of the YOLOv8 model, and the Adaptive Threshold Focal Loss is introduced to replace the original classification loss of the three detection heads to build a sow farrowing behavior monitoring model; S3: using the training set to iteratively optimize the sow farrowing behavior monitoring model until the training is completed when the value of the loss function is less than or equal to a preset threshold, and obtaining a trained sow farrowing behavior monitoring model; S4: Using the trained sow farrowing behavior monitoring model to monitor the farrowing behavior of sows in the test set in real time, and evaluating the performance of the trained sow farrowing behavior monitoring model.
2. According to claim 1, a sow farrowing behavior intelligent monitoring method based on an improved YOLOv8 model is characterized in that: The step S1 comprises: Obtain daily illumination videos and night infrared videos of the farrowing process of several sows, extract frames according to the video time sequence, and filter out blurred and highly similar video frames to obtain a filtered data set; Label each frame image in the filtered data set and perform data enhancement to complete preprocessing; The preprocessed data set is divided into training set and test set according to the preset ratio.
3. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 2 is characterized in that: The filtering of blurred and highly similar video frames comprises the following steps: S1.1: Convert each frame image obtained by video frame extraction into a floating point image; S1.2: All floating point images are processed using a Laplace filter to highlight the high frequency components in each frame of the image; the calculation formula of the Laplace filter is as follows: Where L(x,y) is the pixel value output by the Laplacian filter at the image position (x,y), f(i,j) is the input pixel value at the position (i,j), G(i,j) is the Gaussian kernel value at the position (i,j), w is the width of the image, and h is the height of the image; the higher the L(x,y) value, the clearer the image; S1.3: All images are continuously divided into several groups at a certain frame interval, and the k frames with the lowest L(x, y) values in each group of images are filtered out to obtain a data set after filtering the blurred frames, where k is a positive integer less than the frame interval; S1.4: Using the SSIM algorithm, the similarities between the frames in the data set after filtering the blurred frames are calculated pairwise, and one of the two images whose similarity is greater than or equal to a preset similarity threshold is filtered to obtain the filtered data set.
4. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 2 is characterized in that: The tag includes any of the following: Standing: The sow stretches her limbs to support her body, with her abdomen off the ground; Sitting: The sow's front legs are straightened for support, and the hind legs are bent to sit on the ground, with the back half of the body on the ground and the abdomen off the ground; Side-lying: The sow lies completely on the ground on her side, with the shoulder, ribs and legs on one side touching the ground, the limbs relaxed, the abdomen completely touching the ground, and the teats exposed; Lying prone: The sow's belly is close to the ground, the front and hind legs are naturally bent or straightened on both sides of the body, and the head can be raised or lowered; Birth: The piglet has not separated from the sow's vulva, with part of the head or hind limbs exposed; Piglet: refers to a piglet that has completely separated from the sow's vulva.
5. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 2 is characterized in that: The data enhancement processing includes any one or more of the following: Take angles evenly distributed in the interval [-180°, 180°] and randomly rotate each frame image; Randomly flip each frame image horizontally; Randomly change the brightness, contrast or saturation of each frame; Gaussian noise is randomly added to each frame image.
6. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 1, characterized in that: In step S2, the DLKA module includes: a first 2D convolutional layer, a GELU activation layer, a Deform-DW Conv2D submodule, a Deform-DW-D Conv2D submodule, a second 2D convolutional layer, and a third 2D convolutional layer connected in sequence; The input of the first 2D convolutional layer and the output of the third 2D convolutional layer form a residual sum connection; the output of the GELU activation layer is weighted multiplied with the output of the second 2D convolutional layer; The Deform-DW Conv2D submodule and the Deform-DW-D Conv2D submodule have the same structure, both including a 3×3 convolutional layer and a deformable convolutional layer connected in sequence; In the Deform-DW Conv2D submodule, the input features are passed through a 3×3 convolutional layer to calculate the offset field, and the offset field is used to guide the deformable convolutional layer to adjust its sampling position; Based on the output features of the Deform-DW Conv2D submodule, the expansion value is added as the input feature of the Deform-DW-DConv2D submodule.
7. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 1, characterized in that: In step S2, the Adaptive Threshold Focal Loss is based on Focal Loss and introduces an adaptive threshold mechanism to more effectively process unbalanced data sets and difficult samples. The calculation formula of Adaptive Threshold Focal Loss is as follows: Among them, L ATFL is the function value of Adaptive Threshold Focal Loss; Represents the predicted value of the next epoch, p t Represents the average predicted probability value of the current epoch; λ is a hyperparameter.
8. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 1, characterized in that: In step S4, real-time monitoring of the sow's farrowing behavior includes: Use the trained sow farrowing behavior monitoring model to detect the video data of the sows in the test set throughout the farrowing process, and judge the sow’s current posture and farrowing status in real time; According to the preset judgment strategy and combined with the detection results of multiple frames, the duration of each posture, the frequency of posture changes, and the birth information of the piglets are recorded.
9. The intelligent monitoring method for sow farrowing behavior based on the improved YOLOv8 model according to claim 8, characterized in that: The preset judgment strategy includes: When the video in the test set is input into the trained sow farrowing behavior monitoring model, the recording is triggered only when the following conditions are met, and the duration of the last posture is recorded, and the frequency of each posture change so far is calculated: Among them, p is the previous stable posture of the sow; c is the current posture of the sow; q is the number of frames continuously detected to maintain the same posture; t is the preset frame number threshold; P q is the probability of the current posture detected by the trained sow farrowing behavior monitoring model; P is the preset probability threshold.
10. A method for intelligently monitoring sow farrowing behavior based on an improved YOLOv8 model according to any one of claims 1 to 9, characterized in that: In step S4, the indicators used to evaluate the performance of the trained sow farrowing behavior monitoring model include any one or more of the following: accuracy, precision, recall, F1-Score and average recognition rate (mAP).
Citation Information
Patent Citations
Method, device and apparatus for identifying animal delivery
CN109460713A
Sow delivery state monitoring method and device
CN117095327A
Dermatoscope image-oriented deep learning vitiligo identification method
CN118279667A
Attitude estimation method based on improved YOLOV8 algorithm
CN118609205A
Early fire smoke detection method and system based on improved YOLOv9 algorithm
CN118799805A
Cited By
Sow delivery prediction method using vulva physiological time sequence characterization and prototype correction
CN121661056A
Sow delivery detection system based on space-time motion positioning and detection method thereof
CN122049996A