Meat duck ingestion and drinking behavior recognition algorithm based on MOT and instance segmentation

By combining MOT and instance segmentation algorithms with TAM and Mask R-CNN networks, the problem of confusion in the recognition of meat duck behavior in a stacked cage environment is solved, achieving high-precision monitoring of feeding and drinking behavior, and supporting real-time analysis of health status and early warning of anomalies.

CN120877330APending Publication Date: 2025-10-31JIANGSU ACAD OF AGRI SCI

Patent Information

Application Number
CN202511017357.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In a stacked cage environment, traditional single-target tracking and simple segmentation methods are difficult to accurately distinguish between the feeding and drinking behaviors of meat ducks, leading to confusion in behavior recognition and affecting health monitoring and disease prevention and control.

Method used

We employ recognition algorithms based on MOT and instance segmentation, combining a TAM network structure for multi-target tracking and a Mask R-CNN network structure for behavior recognition and detection. We optimize mask generation using SAM and XMem models to achieve pixel-level target segmentation and behavior recognition.

Benefits of technology

It improves the accuracy and speed of identifying the feeding and drinking behavior of meat ducks, reduces costs, enhances monitoring efficiency, and supports real-time health status monitoring and early warning of abnormal behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877330A_ABST
    Figure CN120877330A_ABST
Patent Text Reader

Abstract

The invention provides a meat duck ingestion and drinking behavior recognition algorithm based on MOT and instance segmentation, and the method comprises the steps: collecting and marking an image of the ingestion and drinking behavior of a meat duck, and building a caged meat duck ingestion and drinking behavior detection data set through image data preprocessing; meat duck group multi-target segmentation and tracking based on target perception TAM; constructing and training a meat duck ingestion and drinking behavior recognition and detection network based on a Mask R-CNN algorithm; and classifying and judging the behaviors of the meat ducks, and counting the total frame number of the water drinking behaviors and the total frame number of the feeding behaviors of the target meat ducks in the video detection period to obtain the water drinking duration and the feeding duration. According to the invention, high-precision identification and time duration statistics of meat duck ingestion and water drinking behaviors can be realized in a stacked cage culture environment, the meat duck health state perception and fine management level of a farm can be improved, abnormal behaviors can be early warned in time, scientific feeding and risk control can be assisted, and the method has good practicability and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent aquaculture, specifically to the application of image segmentation technology and neural network technology, and particularly to an algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation. Background Technology

[0002] As poultry farming continues to advance towards intelligent and intensive operations, the stacked cage rearing model has been widely adopted in my country's duck farming sector due to its advantages of high-density and high-efficiency rearing. This model not only effectively increases yield per unit area but also demonstrates significant advantages in reducing labor costs, and has gradually become the mainstream trend in modern duck farming. In the closed, stacked cage rearing environment, how to achieve real-time identification of ducks' feeding and drinking behaviors based on sensor data, and subsequently conduct health status monitoring and abnormal behavior early warning, has become a key technical problem that urgently needs to be solved in the field of intelligent farming.

[0003] Chinese patent application CN119578884A discloses a big data-based poultry farming monitoring and management system. The method includes: real-time monitoring of the poultry house environment; immediately triggering a behavior monitoring module when an anomaly is detected and generating a specific environmental anomaly score; effectively extracting individual behavioral characteristics of each poultry by combining visual monitoring and infrared thermal imaging, generating individualized behavioral data streams, identifying potential individual health risks, accumulating long-term behavioral and environmental data of each poultry using a health record database, identifying individual behavioral patterns and health change trends using a time series analysis model, predicting possible health risks based on historical data and models, effectively improving the foresight of management and the efficiency of disease prevention and control; and automatically identifying the most influential abnormal characteristics through comprehensive risk scoring and contribution analysis, pushing them to management personnel, updating the health records of each poultry in real time, and improving the overall operational efficiency of the farm.

[0004] However, current technology still faces many challenges. In actual stacked cage environments, due to the confined space, dense flocks of ducks, and severe shading between individuals, coupled with frequent changes in ambient light and reflections, traditional single-target tracking and simple segmentation methods struggle to consistently and stably track each duck, often resulting in target identification confusion. For example, when multiple ducks gather around a waterer or feed trough, target overlap frequently occurs, making it difficult for the system to accurately distinguish between simultaneous drinking and feeding actions, leading to behavioral recognition confusion. Furthermore, the details of ducks' feeding and drinking actions are subtle and their transitions are rapid, making it difficult for traditional bounding box-based detection methods to capture the start and end points of actions, resulting in missed or incorrect judgments in behavioral recognition statistics. Due to these recognition errors, farm managers cannot monitor the health and feeding status of ducks in real time, making it difficult to promptly detect early warning signals such as abnormal drinking or reduced feed intake, delaying disease prevention and control and feeding adjustments, which directly affects the health and output of ducks, and may even lead to the spread of potential diseases and economic losses. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, this invention provides a feeding and drinking behavior recognition algorithm for meat ducks based on MOT and instance segmentation. The algorithm aims to achieve accurate and efficient recognition of the feeding and drinking behavior of meat ducks in a stacked cage environment, thereby helping farms reduce costs and improve monitoring efficiency while ensuring detection accuracy.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] An algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation, the technical steps of which are as follows:

[0008] Step S1000: Collect and label images of feeding and drinking behavior of meat ducks. After image data preprocessing, establish a dataset for detecting feeding and drinking behavior of caged meat ducks.

[0009] Step S2000: Multi-target segmentation and tracking of duck flocks based on target-aware TAM;

[0010] Step S3000: Construct and train a duck feeding and drinking behavior recognition and detection network based on the Mask R-CNN algorithm;

[0011] Step S4000: Classify and determine the behavior of the meat ducks, count the total number of frames of drinking behavior and feeding behavior generated by the target meat ducks during the video detection period, and obtain the drinking time and feeding time.

[0012] Furthermore, the specific implementation method of step S1000 includes the following steps:

[0013] Step S1100: Acquisition and annotation of images of feeding and drinking behavior of caged meat ducks;

[0014] Step S1200: Perform image preprocessing on the labeled images to establish a dataset for detecting the feeding and drinking behavior of caged ducks.

[0015] Furthermore, the images of the feeding and drinking behavior of the caged ducks are acquired through an image acquisition terminal.

[0016] Furthermore, the annotation of the images of caged ducks feeding and drinking behavior is performed using the SAM-Tool semi-automatic annotation tool to annotate the collected images.

[0017] Furthermore, the image annotation information includes: a category determination of the feeding and drinking behavior of ducks in the image, and the annotation information is stored in XML format.

[0018] Furthermore, the image enhancement processing method includes batch processing operations based on OpenCV to increase exposure, enhance contrast, and enhance saturation.

[0019] Furthermore, the specific implementation method of step S2000 includes the following steps:

[0020] Step S2100: Initial mask generation for meat ducks based on SAM and XMem, dynamic temporal tracking, and frame-level mask optimization.

[0021] Furthermore, the method for generating the initial duck mask, dynamic temporal tracking, and frame-level mask optimization includes: using the SAM model to annotate the duck image in the first frame of the preprocessed video to generate an initial target duck mask, and introducing a weak cueing method to continuously optimize the initial mask by modifying the segmentation results multiple times; using the XMem model to predict the target duck mask for subsequent frames, using the target duck mask of the first frame for tracking, and generating corresponding masks in subsequent video frames; and using the SAM model to perform frame optimization on the duck mask video predicted by XMem.

[0022] Furthermore, the specific implementation method of step S3000 includes the following steps:

[0023] Step S3100: Design of Mask R-CNN feature extraction and mask generation;

[0024] Step S3200: Preprocessing and segmentation of the image dataset of feeding and drinking behavior of caged meat ducks;

[0025] Step S3300: Setting training parameters for the caged duck feeding and drinking behavior recognition and detection model;

[0026] Step S3400: Training and optimization of the model for recognizing and detecting the feeding and drinking behavior of caged meat ducks.

[0027] Furthermore, the Mask R-CNN is the selected base network, employing a combination of a ResNet-101 deep residual network and an FPN feature pyramid network as the feature extraction network. This combined structure can extract features with different scales and semantic information. The Region Candidate Network is responsible for generating candidate target regions (RoIs), using anchor boxes and candidate box regression to provide potential target location information. The RoI classification network receives RoIs as input and uses the RoIAlign operation to accurately pool each RoI in the feature map, avoiding the quantization problem introduced by RoIPooling. RoIAlign uses bilinear interpolation to accurately calculate the four coordinate positions of each feature map, thereby improving segmentation accuracy. The mask generation network performs further convolution and upsampling operations on the RoIs to generate mask predictions for each RoI. These mask predictions correspond to the pixel-level segmentation results of the target duck in the image, accurately representing the shape and boundaries of the target.

[0028] Furthermore, the preprocessing method for the image dataset of feeding and drinking behavior of caged ducks is as follows: convert all XML format files to TXT format.

[0029] Furthermore, the method for dividing the image dataset of caged ducks' feeding and drinking behavior is as follows: 90% of the data is used as the training and validation set, and 10% of the data is used as the test set, and the training and validation set is divided into a 90% training set and a 10% validation set.

[0030] Furthermore, the training parameters of the caged duck feeding and drinking behavior recognition and detection model include: an initial learning rate of 0.001, a weight coefficient of 0.0005, a training threshold of 0.9, an input image size of 960×720 pixels, a training period of 300 epochs, and a batch size of 16.

[0031] Furthermore, the training and optimization method for the caged duck feeding and drinking behavior recognition and detection model is as follows: a virtual environment for training the model is built on a GPU server, the training set is input into the Mask R-CNN network structure for target detection model training, and after training, an inference model for caged duck feeding and drinking behavior recognition and detection is obtained; a validation set is input into the inference model for caged duck feeding and drinking behavior recognition and detection for validation, and the inference model is optimized based on the validation results, finally obtaining the best-performing inference model for caged duck feeding and drinking behavior recognition and detection.

[0032] Furthermore, the specific implementation method of step S4000 includes the following steps:

[0033] Step S4100: Classify and determine the behavior of the meat ducks. The behavior determination includes drinking behavior and feeding behavior. If the confidence level of the meat duck mask classification in a certain frame is higher than 90%, it is determined that the target in that frame image has performed the corresponding behavior.

[0034] Step S4200: Calculate the drinking behavior and total number of frames and feeding behavior of the target duck during the video detection period, and then calculate the drinking time and feeding time.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] This invention presents an algorithm for recognizing the feeding and drinking behavior of broiler ducks based on MOT and instance segmentation. On one hand, it employs a TAM network structure to achieve multi-target tracking and recognition of broiler ducks. TAM combines the advantages of SAM and XMem, and continuously improves the segmentation effect through an interactive process, making it suitable for multi-target tracking and recognition tasks of broiler ducks, improving the accuracy and speed of recognition, and providing strong support for broiler duck behavior recognition and behavior duration calculation tasks. On the other hand, it selects a Mask R-CNN network structure for target duck behavior recognition and detection. Based on Faster R-CNN, a target detection model with a branch for predicting masks is added, thereby achieving pixel-level target segmentation while detecting targets. By separating target broiler ducks from the background using Mask R-CNN, mask information for each target broiler duck is obtained, and further, the behavior of the broiler ducks is recognized, thus efficiently detecting the feeding and drinking behavior of caged broiler ducks, which can be widely used in the field of broiler duck farming. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the principle of an algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation, according to the present invention.

[0039] Figure 2 This is a Mask R-CNN network architecture diagram of a meat duck feeding and drinking behavior recognition algorithm based on MOT and instance segmentation according to the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] This invention employs a deep learning-based object detection algorithm to automatically identify the feeding and drinking behaviors of caged ducks. Mask R-CNN is a deep learning model widely used for object detection and instance segmentation. It extends Faster R-CNN by adding a branch for predicting the mask, thus achieving pixel-level object segmentation simultaneously with object detection. By using Mask R-CNN to separate the target ducks from the background, obtaining the mask information for each duck, and further recognizing the ducks' behaviors, this invention achieves extremely high real-time performance.

[0042] Based on the above description, this invention provides an algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation. Please refer to [link to relevant documentation]. Figure 1 As shown, it includes the following steps:

[0043] Step S1000: Collect and label images of feeding and drinking behavior of meat ducks. After image data preprocessing, establish a dataset for detecting feeding and drinking behavior of caged meat ducks.

[0044] Specifically, this step involves collecting and labeling images of feeding and drinking behaviors of meat ducks, preprocessing the image data, and constructing an image dataset of feeding and drinking behaviors of caged meat ducks for behavior detection.

[0045] Further, step S1000 includes:

[0046] Step S1100: Collection and annotation of images of feeding and drinking behavior of caged meat ducks.

[0047] In large-scale duck farming bases, autonomous inspection platforms are used to collect real-time image data of the feeding and drinking behaviors of caged ducks. The data collection process covers different time periods and lighting conditions to ensure the diversity and authenticity of the behavioral images.

[0048] After the data collection was completed, the SAM-Tool semi-automatic annotation tool was used to label the collected dataset with behavioral categories. Based on the behavioral characteristics of the ducks, the data was divided into "feeding behavior" and "drinking behavior" categories. A set of annotation information matching images with corresponding labels was generated. All annotation information was organized and saved as a standardized txt file for easy use in subsequent model training and detection tasks.

[0049] Step S1200: Perform image preprocessing on the labeled images of caged ducks feeding and drinking behavior to establish a dataset for detecting caged ducks feeding and drinking behavior.

[0050] After the images of caged ducks feeding and drinking behavior were labeled, issues such as uneven lighting and strong background interference in the layered cage environment were identified. Therefore, the labeled images underwent uniform preprocessing and enhancement to improve their readability and model recognizability.

[0051] Specifically, the data augmentation methods include, but are not limited to: increasing image brightness to alleviate low-light interference, enhancing image contrast to highlight target boundaries, and increasing image saturation to enhance the color features of duck feathers. These methods can effectively improve the feature difference between the foreground and background in the image, making it easier for subsequent segmentation models to identify the outlines of individual ducks, significantly improving the detection accuracy of the start and end times of feeding and drinking behaviors, and reducing misjudgments caused by blurred edges or reflective interference.

[0052] For example, taking a large-scale duck farm as an example, after the data images of caged ducks' feeding and drinking behavior were labeled, it was found that the lighting in the caged environment was weak in the morning and evening, and some areas were shaded, resulting in overall dark images. The background environment was similar in color to the ducks' feathers, and the ducks' edges were blurred. For example, near the feed trough, the strong reflection from the metal bars of the cages caused light spots and reflective stripes in the images, severely interfering with target recognition. An image enhancement preprocessing method was adopted, specifically including increasing image brightness, enhancing contrast to highlight the boundary between the ducks and the background, and appropriately increasing saturation to enhance the color features of the ducks' feathers. Through this batch processing operation, the visual interference caused by uneven lighting and reflections was effectively reduced, enabling the subsequent segmentation model to more clearly distinguish individual ducks from the background, thereby improving the accuracy of determining the start and end times of feeding and drinking behaviors and significantly reducing misjudgments caused by blurred target edges.

[0053] Step S2000: Multi-target segmentation and tracking of duck flocks based on target-aware TAM.

[0054] Specifically, this step employs a multi-target segmentation and tracking-based technical framework—the Target-Aware Module (TAM)—to perform frame-by-frame tracking and target mask generation of the feeding and drinking behaviors of a flock of meat ducks in a caged environment. This ensures accurate identification of the behavioral state and boundary information of each duck in dense scenes. The TAM framework is an integrated framework that coordinates the functions of the SegmentAnything Model (SAM) and the Memory-Augmented Video Object Segmentation (XMem) model, combining the advantages of both to achieve accurate multi-target mask generation for the flock of meat ducks.

[0055] Specifically, SAM is responsible for generating the initial mask for the first frame of the duck image in the video, providing a high-quality foundation for target segmentation; XMem, based on a temporal memory mechanism, combines the mask information from the previous frame to achieve cross-frame tracking and dynamic prediction of the target mask; TAM integrates the two, effectively improving the continuity and accuracy of segmentation and tracking. This integrated framework ensures that the mask of individual ducks remains intact and clear even in complex and heavily occluded cage environments, providing a solid foundation for the accurate identification of subsequent feeding and drinking behaviors.

[0056] Further, step S2000 includes:

[0057] Step S2100: Initial mask generation for meat ducks based on SAM and XMem, dynamic temporal tracking, and frame-level mask optimization.

[0058] Specifically, the SAM model is used to annotate the first frame of the duck image in the preprocessed video to generate an initial target duck mask. The initial mask is continuously optimized by modifying the segmentation results multiple times. Subsequently, the XMem model is used to predict the target duck mask for subsequent frames, and corresponding masks are generated in subsequent video frames, effectively solving the problem of multiple ducks moving and occluding in dense environments. In response to mask breakage and deformation caused by changes in lighting or occlusion during the XMem prediction process, the SAM model is further used to perform fine-grained frame optimization on the mask results of each frame, correcting mask defects and morphological abnormalities, and ensuring the integrity and continuity of the target mask.

[0059] For example, taking a large-scale duck farm as an example, in real-world scenarios, due to the frequent movement of multiple ducks in their cages, the target mask initially generated by the SAM model inevitably suffers from missing edges or incomplete outlines. To address this issue, through multiple rounds of adjustments and refinement of the segmentation results, the mask boundary of each duck is gradually improved, ensuring accurate coverage of its morphological features. For instance, when several ducks simultaneously approach a waterer area, due to severe occlusion, the XMem model, combined with the mask information from the previous frame, can accurately track and predict the dynamic changes in the boundaries of each duck, effectively distinguishing overlapping areas. To address mask breakage and deformation caused by changes in lighting or occlusion during XMem prediction, the SAM model is further used to perform frame-level optimization of the predicted mask, correcting these issues. This collaborative processing workflow ensures that the target mask remains clear and complete in the densely populated, complexly maneuvering environment of stacked cages, providing accurate basic data support for subsequent precise recognition of feeding and drinking behaviors.

[0060] Step S3000: Construct and train a detection network for recognizing and detecting the feeding and drinking behavior of meat ducks based on the Mask R-CNN algorithm.

[0061] Specifically, this step aims to construct a high-precision detection network model suitable for recognizing the feeding and drinking behavior of caged ducks. Mask R-CNN is selected as the basic detection network, combined with a ResNet-101 deep residual network and a Feature Pyramid Network (FPN) to construct a multi-scale, multi-semantic-level feature extraction module, effectively improving the expressive power of target features. A Region Proposal Network (RPN) is responsible for generating candidate target regions, and combined with anchor boxes and candidate box regression strategies, accurate localization of the duck's target location information is achieved. Through training and optimization of the overall network structure, a detection and inference model with accurate target recognition and behavior judgment capabilities is constructed, enabling effective monitoring and analysis of the feeding and drinking behavior of ducks in a tiered cage environment.

[0062] Further, please refer to Figure 2 As shown, Figure 2 This is a Mask R-CNN network architecture diagram of a meat duck feeding and drinking behavior recognition algorithm based on MOT and instance segmentation according to the present invention. The specific method includes the following steps:

[0063] Step S3100: Design of Mask R-CNN feature extraction and mask generation.

[0064] This step refines the design of the core components of the Mask R-CNN network. A feature extraction network module is constructed using a combination of a ResNet-101 deep residual network and an FPN feature pyramid network, capable of extracting duck target features with different scales and semantic information. The Region Candidate Network (RPN) is responsible for generating candidate target regions (RoIs) and providing potential target location information based on anchor boxes and candidate box regression mechanisms. Subsequently, the RoI classification network receives RoIs as input and uses the RoIAlign operation to accurately pool each RoI in the feature map, avoiding quantization errors caused by RoI Pooling. RoIAlign uses bilinear interpolation to accurately calculate the four coordinate positions of each feature map, thereby improving the spatial accuracy of target segmentation. Finally, the mask generation network (RoIs) performs further convolution and upsampling operations to generate mask predictions for each RoI. These mask predictions correspond to the pixel-level segmentation results of the target duck in the image, accurately representing the morphology and boundaries of the duck in the image.

[0065] For example, taking a large-scale duck farm as an example, in a real-world scenario, when ducks densely surround waterers to drink, traditional methods struggle to accurately segment the outline of each duck due to their similar colors and overlapping postures, leading to confusion in action recognition. This step, through precise region candidate generation and RoIAlign fine-grained pooling, effectively distinguishes the boundaries of adjacent ducks, ensuring accurate depiction of each duck's shape even in complex environments such as dim lighting, shadows, or glare from fences. The system can accurately capture the specific time point and start and end positions of each duck's feeding or drinking, avoiding recognition errors caused by overlapping targets and ensuring that the details of the behavior monitoring data truly reflect the actual situation on site.

[0066] Step S3200: Preprocessing and segmentation of the image dataset of feeding and drinking behavior of caged ducks.

[0067] Specifically, the image dataset of caged ducks' feeding and drinking behavior obtained in step S1000 is preprocessed. First, all labeled files are converted from the original XML format to a lightweight TXT format. Then, based on the collected image dataset of caged ducks' feeding and drinking behavior, a stratified random sampling method is used to divide the dataset into a training and validation set and a test set, with 90% of the data used as the training and validation set and 10% as the test set. The training and validation set is then further divided into a 90% training set and a 10% validation set, thus completing all preprocessing before image input.

[0068] Specifically, this step converts XML format to the lightweight TXT format, reducing file parsing and loading time and facilitating automated batch data processing. Simultaneously, in terms of data partitioning, the diverse duck images collected are allocated to the training, validation, and test sets according to the differences in day and night lighting and the proportion of different cage environments in actual farms, ensuring the model can cover various real-world scenarios. This effectively prevents recognition errors due to insufficient data in specific environments, ensuring the training process fully utilizes diverse samples and improves the behavior recognition system's accuracy in judging feeding and drinking actions in real-world farming environments.

[0069] Step S3300: Setting the training parameters for the caged duck feeding and drinking behavior recognition and detection model.

[0070] In the specific implementation process, the training parameters of the constructed model for recognizing and detecting the feeding and drinking behavior of caged ducks were scientifically configured. These training parameters included: an initial learning rate of 0.001, a weight coefficient of 0.0005, a training threshold of 0.9, an input image size of 960×720 pixels, a training period of 300 epochs, and a batch size of 16. All these parameters were determined through extensive experimental comparison and optimization, taking into account the feature distribution of the feeding and drinking behavior images of caged ducks and the model training requirements, to ensure the stability and convergence speed of the training process and improve the model's recognition accuracy and generalization ability.

[0071] Step S3400: Training and optimization of the model for recognizing and detecting the feeding and drinking behavior of caged meat ducks.

[0072] In the specific implementation process, a virtual runtime environment suitable for the Mask R-CNN object detection model is built in a server environment configured with GPU computing resources. The pre-processed training set of images depicting the feeding and drinking behavior of caged ducks is input into the Mask R-CNN network model to train the object detection model, ultimately obtaining a model suitable for recognizing and detecting the feeding and drinking behavior of caged ducks. After training, considering environmental variables such as complex background interference, multi-target occlusion, changes in day and night lighting, and different cage layouts in actual cage farming scenarios, the model is dynamically optimized and adjusted based on the recognition results fed back from the validation set. This includes, but is not limited to, fine-tuning the learning rate, adjusting the feature extraction layer structure, and balancing the weights of the loss function, continuously improving the model's recognition accuracy and robustness in diverse actual farming environments. Through this training and optimization method, a high-performance, highly adaptable object detection model is finally obtained, meeting the high-precision, real-time recognition requirements for the feeding and drinking behavior of ducks in actual farming environments.

[0073] Step S4000: Classify and determine the behavior of the meat ducks, count the total number of frames of drinking behavior and feeding behavior generated by the target meat ducks during the video detection period, and obtain the drinking time and feeding time.

[0074] This step aims to quantify the duration of feeding and drinking behaviors of caged ducks in monitoring videos. Specifically, it involves counting the number of frames of feeding and drinking behaviors identified in the video detection period for the target ducks, which serve as the basic quantitative basis for their feeding and drinking durations, respectively.

[0075] Specifically, this step precisely quantifies the actual duration of each duck's drinking and feeding behavior by counting the number of consecutive frames of drinking and feeding behavior in the video. This helps farm managers understand the ducks' health status and feeding patterns in real time. In tiered cage environments, this video-based automated behavior duration statistics replace manual observation, reducing labor costs and subjective errors. It also helps to promptly detect abnormal behaviors, such as reduced drinking or abnormal feeding, promoting scientific feeding and refined management.

[0076] Further, step S4000 includes:

[0077] Step S4100: Classify and determine the behavior of the meat ducks. The behavior determination includes drinking behavior and feeding behavior. If the confidence level of the meat duck mask in a certain frame is higher than 90% for the corresponding behavior, it is determined that the target in that frame image has performed the corresponding behavior.

[0078] Specifically, this step sets a confidence threshold of 90% for duck behavior to effectively filter out misjudgments caused by changes in lighting, occlusion, or blurred movements, ensuring that only target duck behaviors highly defined by the model are identified. For example, in a stacked cage environment, when a duck approaches a waterer but its movement is not obvious, only frames with a confidence level reaching the set standard are counted as drinking behavior. This avoids misidentification due to incomplete movements or background interference, thus ensuring accurate correspondence in behavioral statistics and providing reliable data support for subsequent scientific analysis of duck health and feeding / drinking habits.

[0079] Step S4200: Calculate the drinking behavior and total number of frames and feeding behavior of the target duck during the video detection period, and then calculate the drinking time and feeding time.

[0080] The specific formulas for the duration of drinking and feeding are as follows:

[0081]

[0082] Among them, T d The target drinking time for ducks, expressed in seconds (s); F d The total number of frames showing drinking behavior in the target ducks, expressed in frames; Te F represents the feeding time of the target meat ducks, expressed in seconds (s); e The total number of frames generated for feeding behavior in the target duck, expressed in frames; F ps To detect the video frame rate, the unit is frames per second (fps), which is 15 fps in this study.

[0083] For example, in the morning monitoring video of a large-scale duck farm, the system analyzed the behavior of a duck numbered A123. After monitoring the entire 30-minute video, the system identified 450 frames of the duck drinking and 600 frames of feeding. Combined with a video frame rate of 15fps, the calculated drinking time was 30 seconds and the feeding time was 40 seconds. Further detailed analysis showed that the duck frequently went to the waterer during the morning peak hours, but the stays were often short, and the drinking actions were intermittent, potentially indicating insufficient water intake. Meanwhile, its feeding time was concentrated between 8:00 and 9:00 AM, showing a relatively regular feeding habit. Based on this specific data, the farm management optimized the layout of the waterers and increased the frequency of water quality monitoring, promptly identifying and addressing equipment malfunctions and drinking environment issues, thereby effectively ensuring the ducks' drinking needs and preventing potential growth and development problems due to insufficient water access. The entire process was meticulously recorded, detailing the specific duration of the ducks' drinking and feeding, providing data support for the scientific management of the farm.

[0084] The parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.

[0085] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation, characterized in that, include: Images of ducks' feeding and drinking behavior were collected and labeled. After image data preprocessing, a dataset for detecting ducks' feeding and drinking behavior was established. Multi-target segmentation and tracking of meat duck flocks based on target-aware TAM; Construct and train a network for recognizing and detecting feeding and drinking behavior in meat ducks based on the Mask R-CNN algorithm; The behavior of meat ducks is classified and judged. The total number of frames of drinking behavior and feeding behavior generated by the target meat ducks during the video detection period is counted to obtain the drinking time and feeding time.

2. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 1, characterized in that, The specific implementation method for establishing a dataset for detecting the feeding and drinking behavior of caged meat ducks includes: Collection and annotation of images of feeding and drinking behavior of caged meat ducks; Image preprocessing was performed on the labeled images to establish a dataset for detecting the feeding and drinking behavior of caged ducks; The images of the feeding and drinking behavior of the caged ducks were collected using an image acquisition terminal. The annotation of the images of caged ducks feeding and drinking behavior was performed using the SAM-Tool semi-automatic annotation tool to annotate the collected images. The image preprocessing refers to image enhancement processing, including batch processing operations such as increasing exposure, enhancing contrast, and enhancing saturation; The dataset for detecting the feeding and drinking behavior of caged ducks refers to the data on the feeding and drinking behavior of ducks with labeled information.

3. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 1, characterized in that, The specific implementation method for multi-target segmentation and tracking of the meat duck flock includes the following steps: The target-aware TAM network combines the SAM large segmentation model and the XMem advanced VOS model. Initial mask generation, dynamic temporal tracking and frame-level mask optimization for meat ducks based on SAM and XMem.

4. The algorithm for recognizing the feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 3, characterized in that, The specific implementation method for generating the initial mask for the meat duck, dynamic temporal tracking, and frame-level mask optimization includes the following steps: The SAM model was used to annotate the duck image in the first frame of the preprocessed video to generate an initial target duck mask. A weak cueing method was introduced, and the initial mask was continuously optimized by modifying the segmentation results multiple times. The XMem model was used to predict the target duck mask for subsequent frames. The target duck mask in the first frame was used for tracking, and corresponding masks were generated in subsequent video frames. The SAM model was used to optimize the duck mask video predicted by XMem.

5. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 1, characterized in that: The specific implementation method for constructing and training a duck feeding and drinking behavior recognition and detection network based on the Mask R-CNN algorithm includes the following steps: The duck feeding and drinking behavior recognition and detection network includes a feature extraction network, a region candidate network, an RoI classification network, and a mask generation network; Design of Mask R-CNN feature extraction and mask generation; Preprocessing and segmentation of image datasets of feeding and drinking behavior of caged meat ducks; Training parameter settings for a model to identify and detect feeding and drinking behavior in caged meat ducks; Training and optimization of a model for recognizing and detecting feeding and drinking behavior in caged meat ducks.

6. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 5, characterized in that: The Mask R-CNN is a selected base network that uses a combination of a ResNet-101 deep residual network and an FPN feature pyramid network as the feature extraction network. This combined structure can extract features with different scales and semantic information. The region candidate network is responsible for generating candidate target regions (RoIs) and uses anchor boxes and candidate box regression to provide potential target location information. The RoI classification network takes RoIs as input and uses the RoIAlign operation to accurately pool each RoI in the feature map, avoiding the quantization problem introduced by RoIPooling. RoIAlign uses bilinear interpolation to accurately calculate the four coordinate positions of each feature map, thereby improving the segmentation accuracy. The mask generation network further convolves and upsamples the RoIs to generate mask predictions for each RoI. These mask predictions correspond to the pixel-level segmentation results of the target duck in the image, accurately representing the shape and boundaries of the target.

7. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 5, characterized in that: The preprocessing method for the image dataset of feeding and drinking behavior of caged ducks is as follows: convert all XML format files to TXT format; the partitioning method for the image dataset of feeding and drinking behavior of caged ducks is as follows: use 90% of the data as the training and validation set and 10% of the data as the test set, and further divide the training and validation set into a 90% training set and a 10% validation set.

8. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 5, characterized in that: The training parameters of the caged duck feeding and drinking behavior recognition and detection model include: an initial learning rate of 0.001, a weight coefficient of 0.0005, a training threshold of 0.9, and an input image size of 960×720 pixels.

9. The algorithm for recognizing feeding and drinking behavior of meat ducks based on MOT and instance segmentation according to claim 1, characterized in that: The specific implementation methods for the drinking time and feeding time include the following steps: The behavior of meat ducks is classified and determined, including drinking behavior and feeding behavior. If the confidence level of the meat duck mask in a certain frame is higher than 90%, it is determined that the target in that frame has performed the corresponding behavior. The drinking behavior and total number of frames and the feeding behavior and total number of frames generated by the target meat ducks during the video detection period are counted to determine the drinking time and feeding time.

Citation Information

Patent Citations

  • Poultry breeding monitoring management system based on big data

    CN119578884A

Cited By

  • Laying duck state monitoring method and system based on multi-modal data

    CN121145159A