Pig growth state evaluation method based on deep learning and visual large model
By combining YOLOv5-SNMS and ByteTrack algorithms for pig tracing, and using SegmentAnything model for instance segmentation, the problems of low accuracy and manual label dependence of pig growth health assessment methods in the prior art in complex environments are solved, and efficient and accurate pig growth status evaluation is achieved.
Patent Information
- Application Number
- CN202510249879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing pig growth health assessment method has low instance segmentation accuracy in complex breeding environments, lacks correlation information between video stream frames, and relies on a large amount of manual annotation, which is inefficient.
The improved YOLOv5-SNMS model is combined with the ByteTrack algorithm to conduct pig tracking, and the tracking results are used as a prompt for the visual big model SegmentAnything to perform instance segmentation. At the same time, a pig growth health assessment model based on binarized image feature extraction and linear regression networks was designed.
It improves the robustness of pig instance segmentation in complex scenarios, accurately identify pig growth abnormalities, evaluate pig growth status, reduces the dependence of manual labeling, and improves evaluation efficiency.
Smart Images

Figure CN120148070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, object detection, tracking, segmentation, and pig farming technology. More specifically, it relates to a method for evaluating the growth status of pigs based on deep learning and large vision models. Background Art
[0002] With the rapid development of artificial intelligence technology, the livestock industry is gradually transforming towards large-scale, precise, and intelligent directions. By accurately evaluating the growth and health status of pigs, the breeding efficiency can be significantly improved, the labor cost can be reduced, and the productivity can be enhanced.
[0003] The health status of pigs is a crucial indicator in breeding management. In particular, problems such as being thin, malnourished, disabled, or having growth deformities directly affect the growth and development of pigs, production efficiency, and pork quality. Traditional manual measurement methods are not only inefficient but also easily cause stress to pigs, affecting their health and welfare. With the expansion of the scale of the farm, it has become particularly important to establish a scientific and efficient method for evaluating pig growth to better achieve precise management of pig growth and health.
[0004] In recent years, deep learning technology has made significant progress in the field of computer vision, especially showing powerful performance in tasks such as face recognition and object detection. Subsequently, scholars applied instance segmentation in computer vision to the evaluation of pig growth and health. Although existing methods have been applied to a certain extent, most methods have three major problems: First, in the complex breeding environment of the pig farm, when pigs are occluded or adhered, the accuracy of instance segmentation is relatively low; second, the simple instance segmentation technology lacks the correlation information between video stream frames and cannot distinguish different pigs in the video stream; third, the training of existing instance segmentation models relies on a large amount of manual annotation, with low annotation efficiency and poor economy, which limits their wide application in actual breeding scenarios.
[0005] To address the deficiencies in the prior art, the present invention proposes a method for assessing the growth and health of pigs based on the large vision model (SegmentAnything). This method uses the YOLOv5-SNMS model improved from YOLOv5's post-processing, combines it with the ByteTrack algorithm to track pigs. Then, the pig tracking results are used as prompts for the large vision model (SegmentAnything), and the SegmentAnything model is used to perform instance segmentation on pigs to obtain binary masks of pigs. The method of using the pig tracking bounding box as a prompt for SegmentAnything for instance segmentation can effectively improve the robustness of the instance segmentation model in complex scenarios such as pig occlusion and adhesion. The present invention also independently designs a pig growth and health assessment model based on binary image feature extraction and linear regression network. This model uses the binary mask of pigs as input, can accurately detect the postural problems of pigs (malnutrition, growth deformities), analyze the health status of pigs, optimize the breeding process, and improve production efficiency.
[0006] The present invention improves the YOLOv5 model, combines it with the ByteTrack algorithm, uses the pig tracking results as prompts for the large vision model (SegmentAnything), and uses the SegmentAnything model to perform instance segmentation on pigs. It also independently designs a pig growth and health assessment model with the binary mask of pigs as input, thus constructing a method for assessing the growth and health of pigs based on Segment Anything. Compared with previous methods, the present invention is the first to use tracking information as a prompt and the large vision model Segment Anything for instance segmentation, which can track and segment pigs in real time in a video stream and can effectively handle complex scenarios such as pig adhesion and occlusion in pig pens. The pig mask is input into the independently designed pig growth and health assessment model, which can accurately identify the growth abnormalities of pigs, evaluate the growth status of pigs, and provide an innovative solution for the growth and health assessment of group-raised pigs. Summary of the Invention
[0007] The present invention aims to provide a method for assessing the growth and health of pigs based on Segment Anything. This technology improves the YOLOv5 model, combines it with the ByteTrack algorithm to track pigs. Then, the pig tracking results are used as prompts for the large vision model (SegmentAnything), and the SegmentAnything model is used to perform instance segmentation on pigs to obtain binary masks of pigs. This technology also independently designs a pig growth and health assessment model based on binary image feature extraction and linear regression network. This model uses the binary mask of pigs as input and can accurately detect the postural problems of pigs and analyze the health status of pigs.
[0008] A method for evaluating the growth and health of pigs based on the Segment Anything model, which is characterized by innovatively combining target tracking, instance segmentation with the evaluation of pig growth status, effectively solving the problem of monitoring the health status of individual pigs in complex breeding scenarios, including the following steps:
[0009] Step S1: Use an online camera device to obtain videos of group pigs in the group pigsty environment at different time periods and different densities.
[0010] Step S2: Use video frame extraction technology to obtain and organize pig activity pictures. Label the pigs in the pictures according to the order of the video stream picture frames, construct a pig tracking dataset, and divide the dataset into a training set, a validation set, and a test set.
[0011] Step S3: Design a YOLOv5-SNMS-Byte model based on the improved post-processing technology of YOLOv5 for complex scenarios such as pig occlusion and adhesion, so as to track pigs, and use the pig tracking box and pig ID information as prompts for SAM to construct a pig instance segmentation model; among them, the pig instance segmentation model is constructed based on SegmentAnything, and the generated pig tracking box information is used as a prompt and input into the SegmentAnything model. The SegmentAnything model outputs the corresponding pig mask for each frame according to the assigned pig ID, and performs binary processing on the pig mask. Finally, the filled color of the output mask is black.
[0012] Step S4: Use the pig tracking dataset to train the YOLOv5-SNMS-Byte model to obtain the best weight parameters of the YOLOv5-SNMS-Byte model.
[0013] Step S5: Design a pig growth status evaluation model based on binary mask feature extraction and linear regression network.
[0014] Step S6: Based on the video collected in S1, construct a pig growth status dataset, and train the pig growth status evaluation model in S5 to obtain the best weight parameters of the pig growth status evaluation model.
[0015] Step S7: Input the video into the pig instance segmentation model constructed in Step S3, and input the mask of each pig frame by frame and ID into the pig growth status evaluation model of the best weight model constructed in Step S6 to output the growth status classification results (normal, deformed, malnourished) of each pig.
[0016] Furthermore, in step S1, a Hikvision network camera (model: DS-IPC-T13HV3-IA) is used for real-time video recording. The resolution of this camera is set to 1920×1080, and the video frame rate is 24fps. The camera is controlled by a Python script to record the activities of pigs during the following time periods: 5:00 - 9:00, 11:00 - 15:00, 17:00 - 21:00, and 23:00 - 3:00 the next day. The recorded videos will be organized and saved as pig activity data videos for subsequent analysis.
[0017] Furthermore, in step S2, a pig tracking dataset is constructed. The videos collected in S1 are frame-sliced every 10 frames to obtain a pig detection picture set. LabelImg is used to annotate the pigs in the pictures, and the annotated files have the suffix txt, with the file name being the same as the picture name. A pig tracking dataset is obtained. The pig tracking dataset is divided into a training set, a test set, and a validation set in the ratio of 7:2:1.
[0018] Furthermore, in step S3, an improved version of YOLOv5-Byte is used to detect and track pigs. The improved YOLOv5-Byte replaces the original non-maximum suppression (NMS) in YOLOv5 with soft non-maximum suppression (soft-NMS). Soft-NMS reduces the situation of mistakenly deleting correct bounding boxes by gradually decreasing the scores of overlapping boxes instead of directly deleting them, especially suitable for scenarios with dense targets or blurred boundaries. It specifically includes the following steps:
[0019] S3-1: Input the video stream frame by frame into the feature extraction network of YOLOv5 to extract multi-scale feature maps; through several convolutional layers, pooling layers, and feature fusion operations in the network, generate high-dimensional feature maps;
[0020] S3-2: Based on the feature maps, the network performs regression operations on predefined candidate boxes to adjust the positions and sizes of the candidate boxes to more accurately cover the target areas; during this process, detection candidate boxes of pigs and their corresponding classification scores are generated, representing the confidence and class probabilities of the targets in the candidate boxes;
[0021] S3-3: When screening the pig detection candidate boxes, replace the traditional non-maximum suppression NMS with Soft-NMS; the method is as follows: for each candidate box, sort the candidate boxes in descending order of the candidate box scores; traverse each candidate box and dynamically adjust the scores of other candidate boxes; filter the candidate boxes with scores lower than the threshold, and retain the detection boxes with higher scores and better spatial distributions; finally, the candidate box screening will retain the box with the highest classification score to generate the final pig detection box; the information contained in the detection box includes the target category, classification probability, and the specific position coordinates of the bounding box;
[0022] Among them, soft-NMS is different from traditional NMS. Soft-NMS considers both the degree of overlap of prediction boxes and scores. When two boxes with high scores overlap significantly, Soft-NMS does not directly delete the box with a slightly lower score but attenuates its score. For each detected target i, the intersection over union IoU(i,j) with other targets j is calculated. Then, the attenuation exponent σ(i,j) is calculated using the intersection over union, where λ is a hyperparameter used to control the intensity of attenuation. Subsequently, the confidence score score(i) of target i is adjusted according to the attenuation factor. Then, all targets are re-sorted according to the adjusted score score'(i). The above process is repeated until the confidence score is lower than a certain threshold or the maximum detection number is reached. Under Soft-NMS, when two overlapping boxes are detected, the box with the highest score is retained, and the score of the other box is attenuated. If the attenuated score is still relatively high, the box still has a high priority in the next round of comparison, thus retaining more valuable information. This improves the robustness in complex scenarios such as pig occlusion and adhesion. The formula for σ(i,j) and the formula for score′(i) are defined in formulas (1) and (2) respectively
[0023] σ(i,j) = e -λ×IoU(i,j) (1)
[0024]
[0025] S3-4. The pig detection boxes obtained in step S3-3 are passed as input to the ByteTrack algorithm to achieve cross-frame association and tracking of each target. The process is as follows: First, the detections are divided into high-score detections (High) and low-score detections (Low). Subsequently, it uses a Kalman filter to predict the new position of each track in the Tracker within the current frame. The first association is achieved by matching the high-score detections (High) with 20 track sensors in the Tracker. Then, the unmatched detections and tracks are respectively assigned to Dremain and Tremain. The second association is performed between the low-score detections (Low) and the unmatched tracks in Tremain. After these two association steps, any remaining unmatched detections are regarded as background and thus deleted. As for the unmatched tracks (Tre-remain) after the second association, they are retained for a predefined number of frames, usually 30 frames, and then discarded. Finally, the unmatched high-score detections in Dremain are initialized as new tracks. This two-layer association process not only optimizes the matching of tracks and detections but also achieves excellent target tracking performance
[0026] Combining soft-NMS and ByteTrack significantly improves the detection robustness in complex occlusion scenarios. Traditional NMS adopts a relatively extreme suppression strategy for adjacent bounding boxes with an Intersection over Union (IoU) exceeding the threshold, which is prone to misdeleting valid detection bounding boxes due to the dense adhesion of pigs, resulting in missed detections. Soft-NMS, through a linear attenuation formula, attenuates the scores of low-score bounding boxes according to the IoU and hyperparameters instead of directly eliminating them, which not only retains the potential valid detections of occluded targets but also avoids the over-suppression of high-score bounding boxes. This mechanism enables partially occluded pigs to participate in trajectory matching in the subsequent double-layer matching of ByteTrack (high-score detection - initial trajectory association, low-score detection - secondary association of residual trajectories) even if their scores temporarily decrease, effectively alleviating the ID jump problem caused by occlusion. At the same time, the detection bounding boxes that still maintain high scores after attenuation are prioritized to participate in the next round of matching after reordering, forming a "detection - tracking" collaborative optimization - more clues of occluded targets are retained in the detection stage, and the trajectory continuity is maintained through Kalman prediction and multi-level association in the tracking stage, ultimately achieving more stable object detection and ID tracking in scenarios with dense pig herds and frequent occlusions, providing better tracking information for subsequent instance segmentation based on tracking bounding box prompts.
[0027] S3-5. Use the pig tracking bounding boxes obtained by the ByteTrack algorithm as prompt information and input them into the vision large model Segment-Anything. Use the tracking bounding boxes generated by the pig tracking algorithm as input prompts. Each box is represented by the upper-left coordinate (x1, y1) and the lower-right coordinate (x2, y2), containing the Region of Interest (ROI) of the pig and the corresponding ID information. The prompt decoder converts these prompt information (x1, y1, x2, y2) and ID into an embedded vector representation and further enhances its representation ability using the Transformer module. At the same time, the image is input into the image encoder based on Swin Transformer to extract local and global features; the mask decoder extracts context information from the embedded vector representation of the prompt information and the image features, and uses the cross-attention mechanism to combine the image features extracted by the image encoder with the prompt features output by the prompt encoder, gradually generating a target mask that matches the input prompt. The specific process is as follows:
[0028] S3-5-1. Use the tracking bounding boxes generated by the pig tracking algorithm as input prompts. Each box is represented by the upper-left coordinate (x1, y1) and the lower-right coordinate (x2, y2), containing the ROI of the pig and the corresponding ID information. These information are used as prompts and input into the prompt encoder of SegmentAnything. The image is input into the image encoder based on Swin Transformer;
[0029] S3-5-2. The image encoder performs hierarchical processing on the input image, uses an encoder based on Swin Transformer to extract multi-scale features, and finally outputs the multi-scale features of the image F = {F1, F2, F3, F4} (corresponding to different resolutions respectively).
[0030] S3-5-3. The prompt encoder enhances the feature representation of the anchor box prompt (x1, y1, x2, y2). w = x 2 -x 1 ,h = y 2 -y 1 ,and maps the geometric features to a high-dimensional space through sine positional encoding to obtain the embedded vector representation PE of the anchor box. pos = MLP(concat(center x , center y , w, h)). Assign a learnable embedded vector (similar to word embedding) to each unique ID, map the ID to a semantic space of a fixed dimension through a Lookup Table to obtain the embedded vector representation Embed of the ID. ID = EmbeddingLayer(ID). After concatenating the positional encoding and the ID embedding, fuse them through a lightweight Transformer to obtain the tracking information embedded vector representation TrackingEmbed = Transformer(concat(PEpos, EmbedID)) of the concatenation of the two.
[0031] S3-5-4. In the mask decoder, use the tracking information embedded vector representation TrackingEmbed as the Query, and the image multi-scale features F = {F1, F2, F3, F4} as the Key, perform cross-attention on different resolution feature maps respectively, and generate the final mask prediction result M = CrossAttention(TrackingEmbed, F) through multi-scale fusion.
[0032] Among them, in terms of weight allocation in the S3-5-4 mask decoder, a learnable gating mechanism (GatingNetwork) is introduced. This mechanism fuses the tracking embedding and the image features through a multi-layer perceptron (MLP) and adaptively adjusts the fusion weights of the two. The dynamic weight allocation mechanism can adaptively adjust the contribution ratio of the tracking embedding and the image features according to the features and context information of the target, thereby further improving the effect of feature fusion and enhancing the adaptability and accuracy of mask prediction.
[0033] Further, in step S4, the YOLOv5-SNMS-Byte model is trained using a pig tracking dataset to obtain the optimal weight parameters of the YOLOv5-SNMS-Byte model. The specific process is as follows:
[0034] S4-1. The pictures are preprocessed through the input end. The preprocessing includes Mosaic data augmentation and adaptive picture scaling;
[0035] S4-2. The preprocessed pictures enter the backbone network composed of modules such as Focus and CSP in YOLOv5-SNMS, and three feature maps of different sizes are obtained through the channel adjustment module;
[0036] S4-3. The three feature maps of different sizes obtained in step S4-2 are input into the Neck structure composed of FPN and PAN to obtain feature maps of three scales to detect targets of small, medium, and large sizes respectively;
[0037] S4-4. The three feature maps of the scales obtained in step S4-3 are input into the detection head to perform object detection on the feature maps of each scale, including predicting information such as the position, category, and confidence of the bounding box;
[0038] S4-5. The predicted bounding boxes obtained in step S4-4 are input into ByteTrack for detection box matching, and IDs are assigned to the detection boxes.
[0039] S4-6. When training YOLOv5-SNMS-ByteTrack, the detection part mainly evaluates the detection accuracy of the model through indicators such as mAP, Precision, Recall, and F1 Score on the validation set; the tracking part evaluates the tracking accuracy of the model through indicators such as MOTA, IDF1, and the number of identity switches, and further optimizes the overall system performance by means of debugging matching thresholds and Kalman filter parameters, etc.
[0040] Further, the pig growth status evaluation model based on the binary image feature extraction network and the linear regression network in step S5 has the pig binary mask obtained in S3 as the input and the pig growth status classification result as the output. This model extracts features from the pig mask generated in S3 and uses a deep regression prediction network to calculate the growth status classification result. OpenCV is used to analyze the binary pig mask to extract key features including the total number of black pixels, mask boundary, minimum bounding rectangle, body shape ratio, and morphological symmetry, etc. The specific steps are as follows:
[0041] S5-1. Calculate the total number of black pixels in the pig mask to reflect the overall size A of the pig, so as to judge whether the pig is too thin in body condition;
[0042] S5-2. Calculate the boundary of the pig mask and obtain the perimeter P reflecting the shape complexity, thereby determining whether the pig has physical disabilities;
[0043] S5-3. Calculate the minimum bounding rectangle of the pig mask and calculate the aspect ratio AR, thereby determining whether the pig is too thin;
[0044] S5-4. Calculate the body type ratio C of the pig mask, thereby determining whether the pig has growth deformities;
[0045] S5-5. Calculate the deviation degree CD between the centroid of the pig mask and the theoretical axis of symmetry, analyze the morphological symmetry, thereby determining whether the pig has growth deformities;
[0046] Subsequently, these features are input into a deep regression prediction network, which contains multiple linear layers and uses the ReLU activation function to enhance the stability and convergence of the model. At the same time, a Dropout mechanism is introduced to prevent overfitting. Finally, the predicted values output by the network pass through the softmax layer, and the growth state of the pig is finally output.
[0047] Furthermore, in step S6, the pig growth state evaluation model is trained, and the methods for dataset construction and training are as follows:
[0048] S6-1. First, prepare multiple pig images, including pigs with normal growth conditions and pigs with disabilities, deformities, and malnutrition, with a ratio of 1:1. Label them according to the steps of preparing the pig tracking dataset in S2, and then use the labeling results as prompts for Segment Anything for instance segmentation. Save the generated binary masks;
[0049] S6-2. According to whether the growth condition of the pig is normal or not, store its masks in three folders: "normal", "deformed", and "malnourished", with a ratio of 1:1:1 for the three masks. Divide the dataset into a training set, a test set, and a validation set at a ratio of 7:2:1;
[0050] S6-3. Input the binary mask of the pig obtained in S6-1 into the binary feature network of the pig growth health evaluation model to obtain the features of the binary mask of the pig;
[0051] S6-4. Input the features of the binary mask of the pig obtained in step S6-3 into the linear regression layer, and finally output the classification information on whether the pig has growth abnormalities;
[0052] S6-5. In the training stage, first match the predicted category with the true category to determine the attribution of positive and negative samples, and then adjust the weight parameters by calculating the loss function. At the end of each iteration, use the validation set to calculate the accuracy and average precision to continuously optimize the model parameters.
[0053] Furthermore, the maximum number of training iterations (epoch) for the linear regression part of the pig growth and health assessment model is set to 50, the size of the input binary mask image is 320×320, and the number of images input to the model each time (batch) is 8. The model learning rate is 0.0001.
[0054] Furthermore, in the training stage of the model, the weight parameters are adjusted through the cross-entropy loss and mean squared error loss functions. Where y is the true label (0 or 1), y' is the probability predicted by the model (between 0 and 1), m is the number of samples, y i is the true value of the i-th sample, y i * is the predicted value of the i-th sample, and θ is the parameter of the model (including weights and biases). Its calculation formula is as shown in equations (3)-(4):
[0055] L(y,y')=-(ylog(y')+(1-y)log(1-y')) (3)
[0056]
[0057] Advantages of the present invention:
[0058] 1. The present invention first introduces the SegmentAnything model into the field of group breeding pig instance segmentation, constructs a pig instance segmentation method based on SegmentAnything supplemented with tracking prompts, not only improves the accuracy of pig instance segmentation, but also solves the pain point that existing methods require a large amount of manual annotation. It provides a new solution for group breeding pig instance segmentation;
[0059] 2. The present invention independently designs a pig growth and health assessment model based on binary image feature extraction and deep regression network, which can give full play to the advantages of pig masks in pig body posture analysis, accurately identify the growth abnormalities of pigs, evaluate the growth status of pigs, and provide an innovative solution for the growth and health assessment of group breeding pigs. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is the implementation roadmap of the pig growth and health assessment system of the present invention;
[0061] Figure 2 is a schematic diagram of a pig photographed by the camera of the present invention;
[0062] Figure 3 is the pig growth and health assessment flowchart of the present invention;
[0063] Figure 4 is the pig instance segmentation flowchart based on SegmentAnything of the present invention;
[0064] Figure 5 It is the network structure diagram of the pig growth and health model of the present invention. Specific implementation manners
[0065] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0066] As Figure 1 shown, the present invention provides a complete route for implementing a pig growth and health assessment system based on Segment Anything. This process covers steps such as data collection, target tracking, target segmentation, feature modeling, supervised learning, and model evaluation. Based on deep learning and computer vision technologies, it realizes the intelligent assessment of the growth status of group pigs. In this example, ordinary planar images are taken by a pigsty monitoring camera with a resolution of 1080p. The schematic diagram of the camera is as Figure 2 shown, installed on one side of the pigsty close to the walkway, 2.5 m above the ground, to ensure that the pigs in the pigsty are completely photographed.
[0067] As Figure 3 shown, the present invention provides a pig growth and health assessment method based on Segment Anything. This method improves the YOLOv5 model and combines it with the ByteTrack algorithm to track pigs. Then, the pig tracking result is used as a prompt for the visual large model (Segment Anything), and the Segment Anything model is used to perform instance segmentation on pigs to obtain a binary mask of the pigs. The present invention also independently designs a pig growth and health assessment model based on binary image feature extraction and a linear regression network. This model uses the binary mask of the pigs as input and can accurately detect the physical problems of pigs (such as being thinner, malnourished, disabled, and growth deformities, etc.), so as to analyze the health status of pigs.
[0068] As Figure 3As shown in the blue part in [description of the figure], for the pig tracking part of this method, an improved version of the YOLOv5-Byte algorithm is adopted. In the improved YOLOv5-Byte, the non-maximum suppression (NMS) in the original YOLOv5 is replaced with soft non-maximum suppression (soft-NMS). Soft-NMS takes into account both the degree of overlap of the prediction boxes and the scores. When two high-score boxes overlap significantly, Soft-NMS does not directly delete the box with a slightly lower score, but attenuates its score. For each detected target i, the intersection over union IoU(i,j) with other targets j is calculated. Then the attenuation exponent σ(i,j) is calculated using the intersection over union, where λ is a hyperparameter used to control the intensity of attenuation. The confidence score score(i) of target i is then adjusted according to the attenuation factor. Then all targets are re-sorted according to the adjusted score score′(i). The above process is repeated until the confidence score is lower than a certain threshold or the maximum detection number is reached. Under Soft-NMS, when two overlapping boxes are detected, the box with the highest score is retained, and the score of the other box is attenuated. If the attenuated score is still relatively high, the box still has a higher priority in the next round of comparison, thus retaining more valuable information. This improves the robustness in complex scenarios such as pig occlusion and adhesion.
[0069] The matching algorithm selected by the improved YOLOv5-Byte is the ByteTrack algorithm. The ByteTrack algorithm first divides the detection boxes into high-score detection boxes (High) and low-score detection boxes (Low). Subsequently, it uses a Kalman filter to predict the new positions of each track in the Tracker within the current frame. The first association is achieved by matching the high-score detections (High) with 20 track sensors in the Tracker. Then the unmatched detections and tracks are respectively assigned to Dremain and Tremain. The second association is performed between the low-score detections (Low) and the unmatched tracks in Tremain. After these two association steps, any remaining unmatched detections are regarded as background and thus deleted. As for the unmatched tracks (Tre-remain) after the second association, they are retained for a predefined number of frames, usually 30 frames, and then discarded. Finally, the unmatched high-score detections in Dremain are initialized as new tracks. This two-layer association process not only optimizes the matching of tracks and detections but also significantly improves the overall tracking performance.
[0070] Combining soft-NMS and ByteTrack significantly improves the detection robustness in complex occlusion scenarios. Traditional NMS adopts a relatively extreme suppression strategy for adjacent bounding boxes with an Intersection over Union (IoU) exceeding the threshold, which is prone to misdeleting valid detection boxes due to the dense adhesion of pigs, resulting in missed detections. Soft-NMS, through a linear attenuation formula, attenuates the scores of low-score bounding boxes based on the IoU and hyperparameters instead of directly eliminating them, which not only retains the potential valid detections of occluded targets but also avoids over-suppression of high-score bounding boxes. This mechanism enables partially occluded pigs to participate in trajectory matching in the subsequent double-layer matching of ByteTrack (high-score detection - initial trajectory association, low-score detection - secondary association of residual trajectories) even if their scores temporarily decrease, effectively alleviating the ID jump problem caused by occlusion. At the same time, the detection bounding boxes that still maintain relatively high scores after attenuation are prioritized to participate in the next round of matching after reordering, forming a "detection - tracking" collaborative optimization - more clues of occluded targets are retained in the detection stage, and the trajectory continuity is maintained through Kalman prediction and multi-level association in the tracking stage, ultimately achieving more stable object detection and ID tracking in scenarios with dense pig populations and frequent occlusions, providing better tracking information for subsequent instance segmentation based on tracking box prompts.
[0071] As Figure 3 shown in the yellow part of Figure 4 it, the pig instance segmentation method adopts an instance segmentation method based on SegmentAnything supplemented with tracking information prompts. The specific implementation details are as
[0072] shown. First, the tracking bounding boxes generated by the pig tracking algorithm are used as input prompts. The bounding box coordinates (x1, y1, x2, y2) are normalized to proportional values relative to the image size to eliminate resolution differences. The geometric attributes of the bounding box (center point coordinates, width, height, area) are calculated to enhance the spatial representation, and then the geometric features are mapped to a high-dimensional space through sine position encoding (or learnable position encoding). For the pig ID, a learnable embedding vector (similar to word embedding) is assigned to each ID, and the ID is mapped to a semantic space of a fixed dimension through a Lookup Table. Finally, after concatenating the position encoding and the ID embedding, they are fused through a fully connected layer to obtain the embedding vector representation of the tracking prompt information.
[0073] Meanwhile, the image is input into an image encoder based on Swin Transformer to extract local and global features.
[0074] In the mask generation stage, the embedded vector representation of the tracking prompt information (Tracking Embedding) is used as the Query, and the multi-scale feature maps (Fi) output by the image encoder are used as the Key and Value to achieve feature alignment through cross-attention. At different image scales, through a hierarchical interaction method, cross-attention operations are performed separately on feature maps of different resolutions to generate the final mask prediction result.
[0075] In terms of weight assignment for mask generation, a learnable gating mechanism (Gating Network) is introduced to adaptively adjust the fusion weight of the tracking embedding and the image features, thereby optimizing the quality of the generated mask.
[0076] As Figure 3 shown in the green part of Figure 5 it, the pig growth health assessment model is designed based on a binary image feature extraction network and a linear regression network. Its specific implementation details are as
[0077] shown. Its input is the binary mask of pigs obtained in S3, and the output is the classification result of the pig growth state. This model extracts features from the pig masks generated in S3 and uses a deep regression prediction network to predict the growth state of pigs. OpenCV is used to analyze the binary pig masks to extract key features including the total number of black pixels, mask boundary, minimum bounding rectangle, body shape ratio, and morphological symmetry. Subsequently, these features are input into the deep regression prediction network, which contains multiple linear layers and uses the ReLU activation function to enhance the stability and convergence of the model, and at the same time introduces a Dropout mechanism to prevent overfitting. Finally, the predicted values output by the network pass through the softmax layer to output the classification result of the pig growth state.
[0077] The present invention is based on a pig tracking dataset and a pig growth state assessment dataset from a pigsty video stream. Among them, the pig tracking dataset is a self-built dataset based on the LabelImg software. It contains 980 picture sets of pigsty videos with a resolution of 1920×1080, under different lighting conditions and different densities. The pig growth state assessment dataset contains 480 pictures of pigs in different growth conditions, which are divided into three categories: normal, deformed, and malnourished, with a ratio of 1:1:1.
[0078] The present invention uses the Pytorch 1.9.0 deep learning framework and Python 3.8. The hardware environment is an Intel Core i7-13700K processor, an NVIDIA GeForce RTX 4090 24GB graphics card, and the operating system is Windows 10. After the environment is deployed, the pig tracking model is trained. During the experiment, the SGD optimizer is used, with an initial learning rate of 0.0001, an SGD optimizer momentum of 0.937, a weight decay of 0.0005, 100 iterations, and a batch-size of 12. After the pig tracking model is trained, the pig health assessment model is trained. During the training process, the SGD optimizer is used, with an initial learning rate of 0.0001, an SGD optimizer momentum of 0.937, a weight decay of 0.0005, 50 iterations, and a batch-size of 8.
[0079] The above is an example of the best implementation mode of the present invention. The parts not described in detail are the common general knowledge of those of ordinary skill in the art. The protection scope of the present invention is subject to the content of the claims. Any equivalent transformation based on the technical inspiration of the present invention is also within the protection scope of the present invention.
Claims
1. A method for evaluating the growth status of pigs based on deep learning and visual large models, characterized in that: Combining object tracking, instance segmentation and pig growth status assessment to achieve individual pig health status monitoring in complex breeding scenarios includes the following steps: Step S1, using online video equipment to obtain group pig activity videos at different time periods and different densities in a group pig house environment; Step S2, using video frame extraction technology to obtain and organize pig activity pictures; labeling the pigs in the pictures in the order of the video stream picture frames, constructing a pig tracking data set, and dividing the data set into a training set, a validation set, and a test set; Step S3, designing a YOLOv5-SNMS-Byte model based on the improved post-processing technology of YOLOv5 for complex scenes of pig occlusion and adhesion, so as to track the pigs, use the pig tracking frame and pig ID information as SAM prompts, build a pig instance segmentation model, and generate a binary mask matching the input prompt; Step S4, using the pig tracking dataset to train the YOLOv5-SNMS-Byte model to obtain the optimal weight parameters of the YOLOv5-SNMS-Byte model; Step S5, designing a pig growth status assessment model based on binary mask feature extraction and linear regression network; Step S6, based on the video collected in S1, construct a pig growth status data set, and train the pig growth status assessment model in S5 to obtain the optimal weight parameters of the pig growth status assessment model; Step S7, input the video into the pig instance segmentation model constructed in step S3, input the frame-by-frame and ID-by-ID masks of the pigs output by the model into the pig growth status assessment model of the optimal weight model constructed in step S6, and output the growth status classification results of each pig, including normal, deformed, and malnourished.
2. The method for evaluating pig growth status based on deep learning and visual large model according to claim 1, characterized in that: In step S1, a Hikvision DS-IPC-T13HV3-IA online camera is installed in a group pig house environment, and a python script is used to record the pig activities in four time periods of 5:00-9:00, 11:00-15:00, 17:00-21:00, and 23:00-3:00 every day in real time, and the records are organized and saved as pig activity data videos.
3. The method for evaluating pig growth status based on deep learning and visual large model according to claim 1, characterized in that: In the step S2, representative pig activity data videos are screened, and a pig detection picture set is obtained using a video frame extraction script; the pigs in the picture are labeled using LabelImg, and the labeled file has a txt suffix, and the file name is consistent with the picture name, thereby obtaining a pig tracking data set; The pig tracking dataset is divided into training set, test set and validation set in the ratio of 7:2:
1.
4. The method for evaluating pig growth status based on deep learning and visual large model according to claim 1, characterized in that: The specific steps in step S3 include: S3-1, input the video stream frame by frame into the feature extraction network of YOLOv5 to extract multi-scale feature maps; generate high-dimensional feature maps through several convolutional layers, pooling layers and feature fusion operations in the network; S3-2. Based on the feature map, the network performs regression operations on the predefined candidate boxes, adjusts the position and size of the candidate boxes, and makes them cover the target area more accurately. In this process, the pig detection candidate boxes and their corresponding classification scores are generated, indicating the confidence and category probability of the target in the candidate boxes. S3-3, when screening candidate frames for pig detection, replace the traditional non-maximum suppression NMS with Soft-NMS; the method is: for each candidate frame, sort the candidate frame from high to low according to the candidate frame score; traverse each candidate frame and dynamically adjust the scores of other candidate frames; filter the candidate frames with scores below the threshold, and retain the detection frames with higher scores and better spatial distribution; finally, the candidate frame screening will retain the frame with the highest classification score to generate the final pig detection frame; the detection frame contains information such as target category, classification probability, and specific location coordinates of the bounding box; S3-4, passing the pig detection frame obtained in step S3-3 as input to the ByteTrack algorithm to achieve cross-frame association and tracking of each target; the process is: first, the detection is divided into high-score detection High and low-score detection Low; then, the Kalman filter is used to predict the new position of each track in the Tracker in the current frame; the first association is achieved by matching the high-score detection High with the 20 track sensors in the Tracker; then the unmatched detections and tracks are assigned to the detection frame set Dremain with higher confidence and the detection frame set Tremain with lower confidence respectively; a second association is performed between the low-score detection Low and the unmatched tracks in Tremain; after these two association steps, any remaining unmatched detections are regarded as background and are therefore deleted; as for the unmatched tracks Tre-remain after the second association, they are retained for a predefined number of frames, usually 30 frames, and then discarded; finally, the unmatched high-score detections in Dremain are initialized as new tracks; S3-5. The pig tracking bounding box obtained by the ByteTrack algorithm is used as the prompt information and input into the visual large model Segment-Anything. The tracking box generated by the pig tracking algorithm is used as the input prompt. Each box is represented by the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2), containing the pig's ROI and corresponding ID information. The prompt decoder converts these prompt information (x1, y1, x2, y2) and ID into embedded vector representations, and uses the Transformer module to further enhance its representation capabilities. At the same time, the image is input into the image encoder based on the Swin Transformer to extract local and global features; the mask decoder extracts contextual information from the embedded vector representation of the prompt information and the image features, and uses the cross-attention mechanism to combine the image features extracted by the image encoder with the prompt features output by the prompt encoder to gradually generate a target mask that matches the input prompt.
5. The method for evaluating pig growth status based on deep learning and visual large model according to claim 1, characterized in that: In the step S5, the pig growth status assessment model based on binary mask feature extraction and linear regression network has the pig binary mask obtained in S3 as input and the pig growth status classification result as output; the pig growth status assessment model extracts features from the pig mask generated in S3, and calculates the growth status classification result using a deep regression classification network; OpenCV is used to analyze the binarized pig mask to extract key features including the total number of black pixels, mask boundary, minimum circumscribed rectangle, body size ratio and morphological symmetry; then, these features are input into the deep regression classification network, which includes multiple linear layers and uses a ReLU activation function to enhance the stability and convergence of the model, while introducing a Dropout mechanism to prevent overfitting; finally, the predicted value output by the network passes through a softmax layer to output the pig growth status category.
6. The method for evaluating pig growth status based on deep learning and visual large model according to claim 1, characterized in that: The specific steps in step S6 include: S6-1. First, prepare multiple pig images, including pigs with normal growth and pigs with disabilities, deformities, and malnutrition, with a ratio of 1:
1. Label them according to the steps of preparing the pig tracking dataset in S2, and then use the labeling results as prompts for the visual large model SegmentAnything to perform instance segmentation; save the generated binary mask; S6-2. According to whether the pigs are growing normally or not, their masks are stored in three folders: "normal", "deformed" and "malnutrition". The ratio of the three masks is 1:1:
1. The data set is divided into training set, test set and validation set in a ratio of 7:2:
1. S6-3, inputting the pig binary mask obtained in S6-1 into the binary feature network of the pig growth and health assessment model to obtain the features of the pig binary mask; S6-4, inputting the features of the pig binary mask obtained in step S6-3 into the linear regression layer, and finally outputting classification information of whether the pig has growth abnormality; S6-5. During the training phase, the predicted category is first matched with the true category to determine the attribution of positive and negative samples. Then, the weight parameters are adjusted by calculating the loss function. At the end of each iteration, the validation set is used to calculate the accuracy and average precision to continuously optimize the model parameters.
Citation Information
Patent Citations
Weaned piglet target tracking method based on deep learning
CN111709287A
Pig farm pig instance segmentation method based on deep learning
CN114332096A
Live pig multi-target tracking and behavior recognition method, computer equipment and storage medium
CN115830078A
Dairy cow body size measuring method based on deep learning algorithm
CN115965680A
Group health preserving pig weight dynamic monitoring method based on multi-target tracking and segmentation
CN117475368A
Cited By
RGB-D video salient target detection method based on depth-guided adaptive query
CN121640350A
Piglet weight evaluation method and system based on time sequence characteristics and attention mechanism
CN122336863A