Lightweight piglet multi-behavior identification method based on improved YOLOv8

By improving YOLOv8's lightweight network architecture, including the Ghost module and the C2f-Faster-EMA module, the problem of no image preprocessing and high complexity of TSM algorithms in the prior art is solved, and efficient and reliable piglet multi-behavior recognition is achieved.

CN119942649APending Publication Date: 2025-05-06VEGETABLE RES INST GUANGDONG ACAD OF AGRI SERVICES +1

Patent Information

Application Number
CN202510095025.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art does not perform image preprocessing when constructing piglet data sets, which affects model performance. The TSM algorithm requires more training data when identifying piglet population behavior, resulting in high model complexity.

Method used

The lightweight network architecture of improved YOLOv8 is adopted, including the Ghost module and the C2f-Faster-EMA module, and the Ghost module reduces the parameter amount and calculation complexity, the FasterNet module improves the feature extraction speed, and the EMA module enhances the multi-scale feature extraction capability.

Benefits of technology

It effectively reduces the computational complexity and parameter quantity of the model, improves the identification efficiency and reliability, and realizes efficient real-time monitoring and identification in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942649A_ABST
    Figure CN119942649A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight piglet multi-behavior identification method based on improved YOLOv8. The lightweight piglet multi-behavior identification method comprises the following steps: dividing a data set into a training set, a verification set and a test set; a PBR-YOLO network architecture is designed, a PBR-YOLO lightweight network architecture is constructed, and the PBR-YOLO firstly uses a Ghost module network to replace an original backbone network; a C2f-Faster-EMA module and an ELMD (Extreme Local Mode Decomposition) module; the Ghost module network firstly generates an essential feature map through standard convolution, and then carries out linear transformation on the essential feature map to generate a Ghost feature map; a required parameter compression ratio is obtained by adjusting the parameter quantity of the essential feature map and the parameter quantity of the Ghost feature map in the Ghost module; then, a Faster Net module is introduced, and the Faster Net module is combined with EMA to form a C2f-Faster-EMA module; and the piglet behavior can be automatically detected in a practical application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a lightweight piglet multi-behavior recognition method based on improved YOLOv8. Background Art

[0002] Animal husbandry is an important pillar industry of China's national economy and plays a vital role in the national economy. With the increasing demand for meat products from consumers, the scale of livestock and poultry farming is expanding. The health of pigs directly affects the sustainable development and economic benefits of pig farming. Therefore, monitoring pig behavior and evaluating pig health levels are crucial to improving farming efficiency. In order to better manage the pig herd, farmers need to effectively identify and analyze the behavior of piglets, thereby optimizing management strategies and improving the production and economic benefits of the industry.

[0003] For example, the patent document with patent application number 202410424837.2 and publication date 2024.07.12 discloses an integrated detection system for lactating sow piglet behavior based on a patrol robot, including a patrol robot hardware system and a sow piglet behavior detection system; the patrol robot hardware system: using a track-type patrol robot to collect sow piglet behavior images and videos for storage, analysis and transmission; the sow piglet behavior detection system: analyzing sow static images and piglet dynamic short videos, detecting sow postures through YOLOv8 on the core computing unit of the patrol robot, detecting piglet group behavior through a lightweight TSM algorithm, and transmitting the results to the database for storage. Compared with traditional manual inspections, patrol robots and computer vision avoid manual intervention in lactating sows, minimize the risk of zoonotic diseases, improve patrol efficiency, and reduce the labor input cost of farms. At the same time, the platform-based information management system facilitates the formation of effective production and breeding experience, and has the characteristics of automation and intelligence.

[0004] In the above literature, no image preprocessing was performed when constructing the piglet dataset. Preprocessing is an important step in piglet behavior recognition and will directly affect the performance of the model. The TSM algorithm is used to detect the behavior of piglet groups. Since the TSM algorithm may require more training data to learn the behavior patterns in time series, especially in terms of the diversity and complexity of piglet behavior, the TSM algorithm is more complex in structure than the PBR-YOLO algorithm. Summary of the invention

[0005] The present invention provides a lightweight piglet multi-behavior recognition method based on improved YOLOv8, which effectively reduces the computational complexity and parameter amount of the model, thereby improving the recognition efficiency and reliability.

[0006] To achieve the above purpose, the technical solution of the present invention is: a lightweight piglet multi-behavior recognition method based on improved YOLOv8, which is implemented by a patrol robot, and a collection module is set on the patrol robot. The specific steps include: The S1 acquisition module obtains images of piglets in different behavioral states, pre-processes the images, and establishes and determines the piglet's position and behavior data to form a data set; presets the inspection route; S2 builds the PBR-YOLO lightweight network architecture, which includes the Ghost module and the C2f-Faster-EMA module; the C2f-Faster-EMA module includes the FasterNet module and the EMA module; the Ghost module is preset to have a speed-up multiple s, and the parameter compression ratio required by the feature map is preset to be equal to s; S3 presets the total number of feature maps output by the Ghost module as n, and determines the number of Ghost feature maps generated in the Ghost module according to the parameter compression ratio s and the total number of feature maps output by the Ghost module. The training set first generates an essential feature map through standard convolution in the Ghost module, and then linearly transforms the essential feature map to generate a Ghost feature map for output; The S4FasterNet module determines whether to perform convolution processing on the Ghost feature map generated in the input channel according to the Ghost feature map information in the input channel, and then forms an output image after adjusting the weights of two or more scale features and processing through the EMA module, recognizes the output image and determines the behavior and position information of the piglet; S5 The inspection robot moves along the preset route and identifies the piglet behavior and location information according to steps S3-S4.

[0007] The above settings first determine the speed-up multiple s required for the Ghost module, and then determine the number of Ghost feature maps generated in the Ghost module according to the speed-up multiple, so that the number of generated Ghost feature maps can be adjusted. This can greatly improve the speed under the same number of features obtained by ordinary convolution operations, enable rapid recognition, and will not make the entire system too complicated. In addition, by integrating the FasterNet module, the input channels with more input feature information can be convolved to further improve the model processing speed, and then the output image is processed by multi-scale convolution through the EMA module, so as to improve the multi-scale feature extraction capability and enhance the detection accuracy, and realize efficient real-time monitoring of piglet behavior in a resource-constrained environment. It can also realize the association between multiple channels and improve the reliability of detection data. Then, the inspection robot moves along the preset route and uses the above setting model for detection and recognition to ensure the reliability of recognition.

[0008] Furthermore, in step S1, the acquisition module includes a camera, which extracts an image every 15 frames from the video captured by the camera and saves it in PNG format; the image preprocessing process includes screening out images with different backgrounds, lighting and behavior modes, removing low-quality images with blur or noise exceeding a preset value, and then using LabelImg software to annotate the piglet behavior in the image, adding a bounding box and label for each behavior, and the labeled images are the data set.

[0009] With the above settings, an image is extracted every 15 frames from the video captured by the camera. Saving the image in PNG format ensures image quality and filters out images with different backgrounds, lighting, and behavior patterns, which helps the model learn to recognize piglet behaviors in different environments. It removes low-quality images that are blurry or noisy, and adds bounding boxes and labels to each behavior instance to ensure that each behavior type is accurately labeled.

[0010] Furthermore, in step S1, the data set is divided into a training set, a validation set and a test set, wherein the ratio of the training set, the validation set and the test set is set at 8:1:1.

[0011] In the above settings, the ratio of 8:1:1 balances the data requirements of the three stages of training, validation, and testing, ensuring that each stage has enough data while avoiding the problem of uneven data distribution.

[0012] Furthermore, the parameter calculation formula of standard convolution is The parameters of the Ghost module are divided into two parts, including the parameters of the essential feature map and the parameters of the Ghost feature map:

[0013] Parameters of the essential characteristic graph The number of channels of the output feature map is , the number of channels of the input feature map is The convolution kernel size is , Ghost feature maps are generated through low-computation operations. Assuming that the total number of feature maps generated by each essential feature map is The number of Ghost feature maps generated by each essential feature map is , the number of parameters of the convolution kernel is Then the parameter quantity of Ghost feature map is: , The total number of parameters of the Ghost module

[0014] The above settings combine the parameters obtained above to obtain the parameter compression ratio for: Simplifying, we can get the following expression: Referring to the calculation of the acceleration ratio, when When The impact on the total value is very small, and the parameter compression ratio is: Therefore, when When the Ghost module is used to replace the standard convolution, the compression ratio is approximately equal to the theoretical acceleration ratio, which means that the calculation speed of the model can be theoretically increased by s times.

[0015] GhostNet is used to replace the original backbone network, which effectively reduces the number of model parameters and computational complexity. At the same time, because the Ghost module optimizes the calculation method and reduces the generation of redundant features, the network can focus more on extracting key information and maintain efficient feature extraction capabilities.

[0016] Furthermore, in step S4, the FasterNet module includes a PConv 3x3 convolution layer, two 1x1 convolution layers, normalization and ReLU activation functions, and the FasterNet module performs convolution processing on the feature information contained in the input channel.

[0017] The above settings introduce the FasterNet module. By adopting an efficient network structure design, the module greatly reduces the computational complexity and memory access requirements while maintaining efficient feature extraction. The PConv 3x3 convolution layer uses partial convolution to extract features from the input data; two 1x1 convolution layers are used to further reduce the dimension and generate output features respectively; batch normalization and activation function ReLU are applied after the convolution layer to enhance the nonlinear expression ability of the network.

[0018] Furthermore, in step S4, the output image is formed after the EMA module adjusts the weights of more than two scale features, including: grouping the channels input to the EMA module, and convolving the groups through three parallel convolution branches, the three parallel convolutions include two 1×1 convolution branches to process the interaction between channels and capture local feature relationships, and the 3×3 convolution branches focus on extracting multi-scale spatial structure information, and capture pixel-level relationships across spatial dimensions through matrix dot product operations; the outputs of each branch are multiplied by the corresponding weight values ​​and then added to obtain the output image.

[0019] With the above settings, the EMA module evenly distributes the semantic features in the feature maps by performing multi-scale processing on the feature maps, and captures pixel-level dependencies through a cross-space learning mechanism.

[0020] Furthermore, step S2 also includes an ELMD module, and step S4 also includes: the output image is independently processed through a 1×1 (Conv_GN) convolution layer, and then the output image enters two parallel 3×3 convolution layers, and the ELMD detection head introduces group normalization processing; then the normalized image is activated through a SiLU activation function, and then the activated image is resized and output.

[0021] The above settings are processed independently by a 1×1 (Conv_GN) convolutional layer, which is mainly used to reduce the channel dimension and unify the feature dimension. Then, the feature map enters two parallel 3×3 convolutional layers (Conv_GN), which share convolution parameters and are used to further extract multi-scale features. The convolutional layer normalizes the features within each channel group, eliminating the dependence on the batch size. The activation function has advantages in enhancing gradient flow and improving training efficiency, which helps the model better handle complex feature relationships.

[0022] Furthermore, the normalization process includes: the channels of the input image are divided into G groups, each group contains channels, C is the total number of channels; For each group of channels, GN calculates its mean separately and variance , and normalized, its expression is as follows: ; Indicates the position on channel c in the nth batch characteristic information, and are the mean and variance of the g-th group, Is a preset constant.

[0023] The above settings can rarely normalize the input image features, making it easier to perform corresponding processing later.

[0024] Further, step S4 includes: the FasterNet module first identifies the data points in the window and determines the valid data points, and the PConv convolution layer performs a conventional convolution operation on the valid data points.

[0025] The above settings can ensure that convolution operations are performed only on valid data points, thereby ensuring both increased speed and data reliability.

[0026] Furthermore, the FaterNet module first identifies the data points in the window and determines the valid data points, specifically including: the PConv convolution layer defines a binary mask to distinguish the validity of the data points. The valid data points are marked as 1 in the mask, and the missing data points are marked as 0. The convolution operation is performed on both the original data and the mask.

[0027] With the above settings, the convolution operation is performed on both the original data and the mask, ensuring that the convolution kernel is only applied to locations where the mask value is 1. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of the present invention.

[0029] Figure 2 This is a structural block diagram of the inspection robot control module in the present invention.

[0030] Figure 3 Schematic diagram of the PBR-YOLO network structure in the present invention.

[0031] Figure 4 It is a schematic diagram of the Ghost module in the present invention.

[0032] Figure 5 Schematic diagram of characteristic diagrams of different channels in the present invention.

[0033] Figure 6 Schematic diagram of the FasterNet module and its key components in the present invention.

[0034] Figure 7 Schematic diagram of FasterNet, EMA module and its key components in the present invention.

[0035] Figure 8 It is a schematic diagram of the ELMD detection head in the present invention. DETAILED DESCRIPTION

[0036] Embodiment 1.

[0037] like Figure 1-8 As shown in the figure, a lightweight piglet multi-behavior recognition method based on improved YOLOv8 is used to identify piglet behaviors. It is implemented through a patrol robot and a collection module. The specific steps include: The S1 acquisition module obtains images of piglets in different behavioral states, pre-processes the images, and establishes and determines the piglet's position and behavior data to form a data set; presets the inspection route; S2 builds the PBR-YOLO lightweight network architecture, which includes the Ghost module and the C2f-Faster-EMA module; the C2f-Faster-EMA module includes the FasterNet module and the EMA module; the Ghost module is preset to have a speed-up multiple s, and the parameter compression ratio required by the feature map is preset to be equal to s; S3 presets the total number of feature maps output by the Ghost module as n, and determines the number of Ghost feature maps generated in the Ghost module according to the parameter compression ratio s and the total number of feature maps output by the Ghost module. The training set first generates an essential feature map through standard convolution in the Ghost module, and then linearly transforms the essential feature map to generate a Ghost feature map for output; The S4FasterNet module determines whether to perform convolution processing on the Ghost feature map generated in the input channel according to the Ghost feature map information in the input channel, and then forms an output image after adjusting the weights of more than two scale features and processing through the EMA module, recognizes the output image and determines the behavior and position information of the piglet. During the inspection, S5 moves along the preset route and identifies the piglet behavior and location information according to steps S3-S4.

[0038] like Figure 2 As shown in the figure, the inspection robot regularly inspects the farm area according to the preset route to ensure that all piglet activity areas are covered. At the same time, the inspection robot is equipped with a high-definition camera, which can adjust the exposure and resolution of the camera according to the farm environment. For example, in a relatively dark farm environment, the exposure and resolution are increased to ensure that the image clarity and frame rate meet the detection requirements, thereby capturing high-quality images and videos under different lighting, weather and environmental conditions; The acquisition module includes an image processing module and an image recognition module. The image preprocessing module preprocesses the collected piglet images to improve the recognition efficiency and accuracy of the model. The preprocessing steps include image cropping, scaling, rotation correction, grayscale, noise removal, etc. These operations help reduce the complexity of image data and highlight the key features of piglet behavior to provide clearer input data; The image recognition module is equipped with a PBR-YOLO lightweight network model to identify the collected piglet behavior images and obtain the detection results. The detection results include the location of the piglets in the pig house and the behavior information of the piglets.

[0039] The PBR-YOLO lightweight network model uses GhostNet to replace the original backbone network of YOLOv8, which simplifies the architecture and improves the detection speed. At the same time, the FasterNet module with the efficient multi-scale attention (EMA) mechanism is integrated into the C2f module to enhance the feature extraction capability. In addition, the efficient and lightweight multi-path detection head (ELMD) is used to reduce the computational complexity through parallel structure and shared convolution parameters.

[0040] Control module: responsible for fully controlling the behavior of the inspection robot, ensuring that the robot can complete the inspection task according to the predetermined route through precise path planning and dynamic navigation. Through the precise scheduling of the control module, the robot can complete the inspection task autonomously and efficiently, and ensure the stable operation of the system.

[0041] By applying the PBR-YOLO model to inspection robots, farmers can effectively identify and analyze the behavior of piglets, thereby optimizing management strategies and improving the production and economic benefits of the industry. Figure 3 As shown, in step S1, an image is extracted from the video every 15 frames to ensure that the diversity of piglet behaviors is captured, and the image is saved in PNG format to ensure image quality, and images with different backgrounds, lighting, and behavior modes are screened out, and low-quality images with blur or noise greater than a preset value are removed to ensure the clarity and diversity of the data set; the piglet behaviors in the image are annotated using LabelImg software, and a bounding box and label are added to each behavior to ensure that each behavior type is accurately labeled. In this embodiment, the LabelImg software is used for annotation as a prior art and is not described here.

[0042] In order to ensure the clarity and diversity of sample categories in the dataset, we manually selected piglet pictures to cover various situations and behavioral scenarios. In the end, we obtained 1,200 pictures containing various piglet behaviors. The dataset is divided into training set, validation set and test set. The three sets are set in a ratio of 8:1:1. The training set has 960 images, and the validation set has 120 images and the test set has 120 images. Each picture contains multiple instances of piglet behavior, and there are a total of 10,640 behavioral instances in the entire dataset; the annotated data is enhanced, and the collected pictures are randomly rotated, randomly cropped, color jittered, Gaussian noised, horizontally flipped, and vertically flipped; the YOLOv8 network model is trained based on the piglet dataset to identify various behavioral states of piglets and record relevant behavioral information; the trained target detection model adopts the improved PBR-YOLO model; like Figure 4As shown in the figure, the Ghost module reduces model parameters and computational costs by replacing a part of the standard convolution operation with a less computationally intensive linear transformation. In the Ghost module, the standard convolution operation generates an essential feature map, and the number of feature maps generated is half of the standard convolution output. A linear transformation is performed on the essential feature map to generate a redundant Ghost feature map.

[0043] In this embodiment, the linear transformation function is a linear function, as shown below: Y=WX+B, where B is the bias value, W is the weight matrix of the linear transformation, X is the essential feature map, and Y is the Ghost output feature map. The parameter calculation formula for the standard convolution is , the parameters of the Ghost module are divided into two parts, including the parameters of the essential feature map and the parameters of the Ghost feature map: ; Parameters of the essential characteristic graph The number of channels of the output feature map is , the number of channels of the input feature map is The convolution kernel size is Ghost feature maps are generated through low-computation operations. Assume that the total number of feature maps generated by each essential feature map is The number of Ghost feature maps generated by each essential feature map is , the number of parameters of the convolution kernel is Then the parameter quantity of Ghost feature map is:

[0044] The total number of parameters of the Ghost module: Therefore, the parameter compression ratio is obtained by combining the parameters obtained above. for: Simplifying, we can get the following expression: , refer to the calculation of the acceleration ratio, when When The impact on the total value is very small, and the parameter compression ratio is: , therefore, when When the Ghost module is used to replace the standard convolution, the compression ratio is approximately equal to the theoretical acceleration ratio, which means that the calculation speed of the model can be theoretically increased by s times.

[0045] The Ghost feature map generation complexity verification process is as follows: Generate each essential feature map using a cheap linear transformation Ghost feature map, where the cheap transformation convolution kernel size is , according to the standard convolution complexity calculation formula is as follows: ; in and are the height and width of the essential feature map, c is the number of input channels, k is the convolution kernel value, and d is the convolution kernel value of the Ghost feature module. Since the complexity of the Ghost feature map is the sum of the complexity of the standard convolution and the complexity of the Ghost feature map, the calculation complexity expression is as follows:

[0046] Combining the above formulas, we can get the total computational complexity expression of the Ghost module as follows:

[0047] Therefore, by substituting the FLOPs derived previously into the speedup definition formula, we can obtain:

[0048] After simplification, the expression is as follows:

[0049] when hour, , that is, the convolution kernel of the cheap operation is small, and s-1 in s-1+c is smaller than c, so s-1+c is close to c, so the formula can be further simplified as:

[0050] From the above formula, we can see that the speedup ratio of the Ghost module is close to , that is, through the Ghost module, we can get about The calculation speed is accelerated by times, which can verify that the calculation speed of the model is increased by s times mainly because the Ghost module needs to be accelerated by s times. After the speed is increased by s times, the feature output of the Ghost module needs to be n / s*(s-1).

[0051] like Figure 5 As shown in the figure, PConv convolves some input channels through dynamic channel selection, reducing the computation of redundant features while maintaining efficient feature extraction capabilities. By using the channel attention mechanism FasterNet module, the model can automatically adjust and optimize the use of channels according to different levels of feature information during training or inference.

[0052] like Figure 6As shown in the figure, the FasterNet module includes a PConv 3x3 convolution layer, two 1x1 convolution layers, normalization and ReLU activation functions, which greatly reduces the computational complexity and memory access requirements while maintaining efficient feature extraction. The FasterNet module is introduced, and by adopting an efficient network structure design, the module greatly reduces the computational complexity and memory access requirements while maintaining efficient feature extraction. Among them, PConv 3x3 convolution layer: uses partial convolution to extract features of input data; two 1x1 convolution layers: used to further reduce the dimension and generate output features respectively; normalization and activation function ReLU: after the convolution layer, normalization and ReLU activation function are applied to enhance the nonlinear expression ability of the network; FasterNet module performs convolution processing on the feature information in the input channel.

[0053] Step S4 includes: the FasterNet module first identifies the data points in the window and determines the valid data points, and the PConv convolution layer performs a conventional convolution operation on the valid data points. The FasterNet module first identifies the data points in the window and determines the valid data points. Specifically, the PConv convolution layer defines a binary mask to distinguish the validity of the data points. The valid data points are marked as 1 in the mask, and the missing data points are marked as 0. The convolution operation is performed on both the original data and the mask.

[0054] Assume that the input and output feature maps have the same dimensions and the same number of channels, and the convolution is performed on only c of the input. p The convolution operation is performed on each channel.

[0055] Then the FLOPs expression of PConv is as follows:

[0056] Among them, when When the number of channels c is one-fourth of the standard convolution, the floating-point operations of PConv are reduced to 1 / 16 of the standard convolution, and its expression is as follows:

[0057] Therefore, it can be concluded that the computational complexity of partial convolution is only 1 / 16 of that of standard convolution.

[0058] like Figure 7As shown in the figure, the EMA module is integrated into the FasterNet block to optimize the overall network structure. The efficient lightweight multi-path detection head (ELMD) is used to ensure efficient and accurate target detection and behavior recognition while maintaining low computational overhead through parallel structure and shared convolution parameters. ELMD is a lightweight multi-path detection head for detection tasks proposed by the present invention, which aims to solve the problems of high computational complexity, large number of parameters, and insufficient fusion of multi-scale features in the traditional YOLOv8 detection head. By designing parallel convolution paths and shared convolution mechanisms, ELMD effectively reduces computational overhead and enhances the fusion capability of multi-scale features, which is particularly suitable for resource-constrained edge devices and mobile platforms.

[0059] In step S4, the output image is formed after the EMA module adjusts the weights of more than two scale features, including: grouping the channels input to the EMA module, and convolving the groups through three parallel convolution branches, the three parallel convolutions include two 1×1 convolution branches to process the interaction between channels and capture local feature relationships, and the 3×3 convolution branches focus on extracting multi-scale spatial structure information, and capture pixel-level relationships across spatial dimensions through matrix dot product operations; the outputs of each branch are multiplied by the corresponding weight values ​​and then added to obtain the output image.

[0060] The ELMD module processing process in step S4 module also includes: independently processing the output image through a 1×1 convolution layer, and then the output image enters two parallel 3×3 convolution layers, and the ELMD detection head introduces group normalization processing; then the normalized image is activated through the SiLU activation function, and then the activated image is resized and output.

[0061] The normalization process includes: the channels of the input image are divided into G groups, each group contains channels, C is the total number of channels; For each group of channels, GN calculates its mean and variance , and normalized, its expression is as follows:

[0062] Indicates the position on channel c in the nth batch characteristic information, and are the mean and variance of the g-th group, Is a preset constant.

[0063] SiLU activation function, in order to enhance the nonlinear expression ability of the model, the ELMD detection head uses the SiLU activation function, and its expression is as follows:

[0064] Where x is the input image and e is a natural constant. The SiLU function has advantages in enhancing gradient flow and improving training efficiency, which helps the model better handle complex feature relationships.

[0065] Finally, it is necessary to adjust the scale and output. In ELMD, a scale adjustment layer is designed to address the problem that different detection heads process different target scale ranges. After processing multi-scale features, the scale adjustment layer will adjust the scale of the feature map according to the needs of different detection heads. This enables ELMD to better adapt to changes in the size of input data, thereby avoiding unnecessary complex calculations and further improving the flexibility and adaptability of the model; after processing these features, ELMD passes the adjusted feature maps to the regression head and classification head respectively to perform the bounding box regression and category classification tasks of the target. Since each part shares convolution features, ELMD effectively reduces redundant convolution operations, thereby reducing computational overhead; the model is deployed to actual application scenarios, and the piglet behavior is automatically detected by the inspection robot, and the farmer monitors the piglet behavior in real time.

[0066] This paper proposes a lightweight architecture based on GhostNet of the PBR-YOLO model, a C2f module optimized by FasterNet Block, and an efficient multi-path detection head (ELMD), which effectively reduces the computational complexity and parameter quantity of the model, and finally obtains an efficient and lightweight piglet multi-behavior recognition system. Farmers can monitor the behavior of piglets. This invention not only reduces the consumption of computing resources, but also improves the reliability and real-time performance of detection. It is suitable for intelligent monitoring and refined management in intensive pig farms, and can effectively reduce labor costs and improve breeding production efficiency. This embodiment experiments with different models. The comparative experimental results presented in Table 1 show that the FLOPs (floating point operations) and number of parameters of the SSD, FasterR-CNN, and RT-DETR models are significantly higher than those of other models, resulting in larger weight files of these models.

[0067] Specifically, the FLOPs of the SSD model is 207.5 G, the number of parameters is 24.68M, and the model size is 98.8 MB; the FLOPs of the Faster R-CNN model is 370.2 G, the number of parameters is 137.15M, and the model size is 113.7 MB; the FLOPs of the RT-DETR model is 125.7 G, the number of parameters is 41.95M, and the model size is 86.1 MB. These characteristics make it difficult for these three models to run efficiently in a resource-constrained environment. The model proposed in the present invention achieved 82.7% accuracy, 70.0% recall, 78.5% mAP (mean average precision), 6.6 milliseconds single-frame inference time, and 2.7 MB model size on the test set, which is significantly better than other models in terms of the number of parameters, floating-point operations (FLOPs), and model size. Compared with the original YOLOv8n model, the proposed model improves the accuracy by 3.1%, reduces the number of parameters by 1.78M, reduces the FLOPs by 4.2G, and reduces the model size by 3.6 MB, making it more applicable on resource-constrained devices. Compared with other YOLO models (including YOLOv7-tiny, YOLOv5n, and YOLOv10), the proposed model performs better in terms of parameter efficiency and model size. Specifically, compared with the YOLOv5n model with the least number of parameters, the proposed model improves the average accuracy by 12.9%, reduces the number of parameters by 30.5%, and reduces the model size by 30.7%.

[0068]

[0069] Finally, the S5 inspection robot moves along the preset route and identifies the piglet behavior and position information according to steps S3-S4. If the piglet behavior is identified as being immobile, the current position information is sent to the control module to control the inspection robot and then move it to a position close to the current position. The piglet behavior can then be confirmed at close range, such as whether it is sick or in other states.

[0070] The working principle of the present invention is as follows: first, by determining the speed-up multiple s required for the Ghost module, and then determining the number of Ghost feature maps generated in the Ghost module according to the speed-up multiple, the number of Ghost feature maps generated can be adjusted, so that the speed can be greatly improved relative to the same number of features obtained under ordinary convolution operations, and rapid recognition can be achieved without making the entire system too complicated. In addition, by integrating the FasterNet module, the input channel with more input feature information can be convolved to further improve the model processing speed, and then the output image is subjected to multi-scale convolution processing through the EMA module, so as to improve the multi-scale feature extraction capability and enhance the detection accuracy, thereby realizing efficient real-time monitoring of piglet behavior in a resource-constrained environment, and also realizing the association between multiple channels to improve the reliability of detection data, and then the inspection robot moves along the preset route and adopts the above-set model for detection and recognition to ensure the reliability of recognition.

Claims

1. A lightweight piglet multi-behavior recognition method based on improved YOLOv8, implemented by a patrol robot and a collection module, characterized in that: The specific steps include: S1 acquisition module obtains images of piglets in different behavioral states, pre-processes the images and establishes and determines the piglet position and behavior data to form a data set; presets the inspection route; S2 builds the PBR-YOLO lightweight network architecture, which includes the Ghost module and the C2f-Faster-EMA module; the C2f-Faster-EMA module includes the FasterNet module and the EMA module; the Ghost module is preset to have a speed-up multiple s, and the parameter compression ratio required by the feature map is preset to be equal to s; S3 presets the total number of feature maps output by the Ghost module as n, and determines the number of Ghost feature maps generated in the Ghost module according to the parameter compression ratio s and the total number of feature maps output by the Ghost module. ; The training set first generates an essential feature map through standard convolution in the Ghost module, and then the essential feature map is linearly transformed to generate a Ghost feature map for output; The S4FasterNet module determines whether to perform convolution processing on the Ghost feature map generated in the input channel according to the Ghost feature map information in the input channel, and then forms an output image after adjusting the weights of two or more scale features and processing through the EMA module, recognizes the output image and determines the behavior and position information of the piglet; S5 The inspection robot moves along the preset route and identifies the piglet behavior and location information according to steps S3-S4.

2. According to claim 1, a lightweight piglet multi-behavior recognition method based on improved YOLOv8 is characterized in that: In step S1, the acquisition module includes a camera, which extracts an image every 15 frames from the video captured by the camera and saves it in PNG format; the image preprocessing process includes screening out images with different backgrounds, lighting and behavior modes, removing low-quality images with blur or noise exceeding a preset value, and then using LabelImg software to annotate the piglet behavior in the image, adding a bounding box and label for each behavior, and the labeled images are the data set.

3. According to claim 1, a lightweight piglet multi-behavior recognition method based on improved YOLOv8 is characterized in that: In step S1, the data set is divided into a training set, a validation set, and a test set, wherein the ratio of the training set, the validation set, and the test set is set at 8:1:

1.

4. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 1 is characterized in that: The parameter calculation formula for standard convolution is The parameters of the module are divided into two parts, including the parameters of the essential feature map and the parameters of the Ghost feature map: Parameters of the essential characteristic graph The number of channels of the output feature map is , the number of channels of the input feature map is , the convolution kernel size is , Ghost feature maps are generated through low-computation operations. Assuming that the total number of feature maps generated by each essential feature map is The number of Ghost feature maps generated by each essential feature map is The parameters of the convolution kernel are Then the parameter quantity of Ghost feature map is: The total number of parameters of the Ghost module: .

5. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 1 is characterized in that: In step S4, the FasterNet module includes a PConv 3x3 convolution layer, two 1x1 convolution layers, normalization and ReLU activation functions. The FasterNet module performs convolution processing on the feature information contained in the input channel.

6. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 5 is characterized in that: In step S4, the output image is formed after the EMA module adjusts the weights of more than two scale features, including: grouping the channels input to the EMA module, and convolving the groups through three parallel convolution branches, the three parallel convolutions include two 1×1 convolution branches to process the interaction between channels and capture local feature relationships, and the 3×3 convolution branches focus on extracting multi-scale spatial structure information, and capture pixel-level relationships across spatial dimensions through matrix dot product operations; the outputs of each branch are multiplied by the corresponding weight values ​​and then added to obtain the output image.

7. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 1 is characterized in that: Step S2 also includes an ELMD module, and step S4 also includes: the output image is independently processed through a 1×1 convolution layer, and then the output image enters two parallel 3×3 convolution layers, and the ELMD detection head introduces group normalization processing; then the normalized image is activated through a SiLU activation function, and then the activated image is resized and output.

8. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 5 is characterized in that: The normalization process includes: the channels of the input image are divided into G groups, each group contains channels, C is the total number of channels; For each group of channels, GN calculates its mean and variance And normalized, the expression is as follows: Indicates the position on channel c in the nth batch characteristic information, and are the mean and variance of the g-th group, Is a preset constant.

9. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 1 is characterized in that: Step S4 includes: the FasterNet module first identifies the data points in the window and determines the valid data points, and the PConv convolution layer performs a conventional convolution operation on the valid data points.

10. The lightweight piglet multi-behavior recognition method based on improved YOLOv8 according to claim 9, characterized in that: The FaterNet module first identifies the data points in the window and determines the valid data points, specifically including: The PConv convolutional layer defines a binary mask to distinguish the validity of the data points. The valid data points are marked as 1 in the mask, and the missing data points are marked as 0. The convolution operation is performed on both the original data and the mask.

Citation Information

Patent Citations

  • Integrated detection system for behaviors of sows and piglets in lactation period based on inspection robot

    CN118334704A

Cited By

  • Lightweight denoising method based on convolutional neural network

    CN121053033A