A Multimodal Sensing Cascade Monitoring Method and System for Intensive Farming of Muscovy Ducks

By employing a multimodal sensing cascade monitoring method, utilizing a dual-track tracking system guided by head anchor points and multidimensional analysis, the problems of unstable identity tracking and difficulty in gait quantification of Muscovy ducks in intensive farming environments were solved, achieving efficient and reliable health monitoring.

CN122223064BActive Publication Date: 2026-07-31SICHUAN AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN AGRI UNIV
Filing Date
2026-05-15
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In intensive farming environments, existing technologies suffer from unreliable identification tracking of ducks, unreliable static physical signs, and difficulties in gait quantification, resulting in poor health monitoring outcomes.

Method used

A multimodal sensing cascaded monitoring method is adopted, using a dual-track tracking system guided by head anchor points, combined with visual, infrared speckle flow and spatial location information, to maintain identity and conduct detailed static physical characteristics, and to conduct detailed gait analysis through spatiotemporal dynamics analysis, so as to achieve continuous monitoring of Muscovy ducks.

Benefits of technology

It significantly improves the continuity and robustness of identity tracking, enables highly reliable diagnosis of stationary individuals and objective dynamic quantification of gait, and optimizes the computing power efficiency ratio in monitoring tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223064B_ABST
    Figure CN122223064B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal perception cascade monitoring method and system for intensively farmed Muscovy ducks, belonging to the field of Muscovy duck farming monitoring. It addresses the problem of unstable identity tracking in high-density farming environments caused by occlusion, complex textures, and discontinuous movement in existing technologies. This invention adopts a three-level linkage architecture: First, a dual-track tracking model based on head anchor point detection and body center guidance is constructed. Continuous identity tracking in occluded scenarios is achieved through feature vector maintenance and asynchronous Kalman filtering, and task routing is performed based on group activity comparison. Second, local regions of interest are cropped for stationary abnormal individuals, and infrared speckle micro-motion analysis and key point geometric pose are integrated. Finally, the spatiotemporal trajectory of the feet is extracted for moving individuals, and gait anomalies are quantitatively evaluated using dynamic indicators such as frequency domain energy, step height variation coefficient, and landing impact acceleration. This invention significantly improves the robustness of identity tracking and the anomaly detection rate, making it suitable for large-scale, high-density farming scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of monitoring in Muscovy duck farming, specifically to a multimodal sensing cascade monitoring method and system for intensively farmed Muscovy ducks. Background Technology

[0002] In large-scale Muscovy duck farming, timely health warnings and development monitoring are crucial for ensuring profitability. Currently, the industry mainly relies on manual inspections, but in high-density environments, abnormal individuals are easily obscured by the flock, and inspections are conducted randomly, making continuous tracking around the clock difficult. This leads to delayed disease detection and increases the risk of disease transmission.

[0003] Existing non-contact automated monitoring technologies mainly fall into three categories:

[0004] Global initial screening technology: an individual localization and counting scheme based on CBAM-YOLOv7 and ByteTrack / DeepSORT. This method enhances target feature extraction through an attention mechanism, enabling real-time population statistics and activity assessment in dense scenes. However, Muscovy ducks exhibit strong gregariousness, with individuals frequently burrowing and pushing each other, often obscuring or temporarily eliminating previously stable head features. Existing tracking algorithms heavily rely on the continuous visibility of visual features; once the head is lost, multi-target tracking systems struggle to maintain continuous identification, easily leading to ID drift or re-identification errors, resulting in significant deviations in activity statistics.

[0005] Static vital sign analysis techniques: physiological signal extraction schemes based on Euler video magnification (EVM) or infrared thermal imaging (IRT). These techniques calculate respiratory rate, prone position, and neck retraction posture by analyzing subtle pixel fluctuations or surface temperature distribution in RGB video. However, in actual Muscovy duck farming environments, the complex patterns on their skin, under natural or non-uniform lighting, result in low signal-to-noise ratios when passively extracting sub-pixel-level minute displacements. They are also susceptible to background noise and image artifacts, making it difficult to reliably obtain respiratory rhythms and distinguish with high confidence between deep sleep, pathological neck retraction, and biological death.

[0006] Dynamic gait analysis technology is divided into two categories: contact and passive. Contact methods (accelerometers, RFID tags, ground pressure sensors) can directly collect gait frequency, ground impact, and force differences between the left and right limbs. However, sensors are prone to detachment, contamination, or signal interruption, making long-term continuous sampling difficult. Passive methods (ToF cameras, binocular cameras) extract the trajectory of the feet and the center of the body through 3D reconstruction. However, under conditions of frequent duck occlusion, ground reflection, or water surface interference, the trajectory of key points is often broken, making it difficult to form a complete dynamic feature chain in the gait process. This results in insufficient stability for gait anomaly identification and continuous quantitative analysis.

[0007] In summary, existing technologies generally lack robustness under conditions of dense occlusion, complex textures, and discontinuous motion, and there is an urgent need for a closed-loop monitoring solution that can continuously perform identity maintenance, anomaly screening, and dynamic diagnosis. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multimodal sensing cascade monitoring method and system for intensively farmed Muscovy ducks, so as to solve the problems of unstable identity tracking, unreliable static physical characteristics judgment, and difficulty in gait quantification caused by occlusion, complex textures and discontinuous movement in high-density farming environments.

[0009] The present invention achieves the above objectives by adopting the following technical solution: Firstly, the present invention provides a multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks, comprising the following steps:

[0010] Step S1: Global initial screening and identity maintenance;

[0011] The dual-track tracking system based on head anchor point guidance can detect and track individual ducks in the breeding scene in real time, maintain the identity ID and movement parameters of each individual in real time, and generate task routing decisions based on the comparison of group activity.

[0012] Step S2: Detailed examination of static individual physical signs;

[0013] When the task routing decision determines that an individual is a static abnormal risk individual, the multimodal vital signs detailed investigation process is triggered. Combining visual, infrared speckle flow and spatial location information, the vital signs, body posture and spatial distribution of the static individual are diagnosed in parallel, and the static health status judgment result is output.

[0014] Step S3: Detailed examination of dynamic individual gait;

[0015] When the task routing decision determines that an individual is a group-moving individual and meets the preset physical scale threshold, the spatiotemporal dynamics gait analysis process is triggered. The temporal signals of key points of the individual's feet and head during continuous walking are extracted. The degree of gait abnormality is quantitatively evaluated by at least one of frequency domain energy analysis, step height variation coefficient analysis and landing impact acceleration analysis, and the dynamic gait determination result is output.

[0016] Furthermore, step S1 specifically includes:

[0017] S101, RGB video acquisition and feature annotation;

[0018] High-resolution RGB video streams from the breeding environment were collected. Head bounding box coordinates were marked for the visible duck head area in the video frames. The coordinates of the individual body center point were marked based on the head position and the unobstructed body parts, forming a spatial topological association dataset of head bounding box coordinates and body center point coordinates.

[0019] S102, Two-stage progressive training strategy and tracking feature modeling;

[0020] A deep learning model is constructed, and the first and second training stages are executed sequentially. The head and body feature vector of each individual is calculated and maintained in real time. The head and body feature vector is equal to the coordinates of the body center point minus the coordinates of the head box center. The magnitude of the head and body feature vector represents the body proportion feature, and the phase represents the pointing angle of the head and body relative to the image coordinate system.

[0021] S103, Improved asynchronous tracing logic that integrates spatial topological associations;

[0022] The head trajectory and body center trajectory are maintained synchronously. When the head is lost, the Kalman filter predictor takes over and maintains the continuity of identity with the body center point. When the head reappears, reverse addressing and reconnection are performed based on the head and body feature vectors. At the same time, the pixel displacement of the center point is extracted in real time as the individual's motion parameters.

[0023] S104. Task routing and offloading based on group activity comparison;

[0024] The movement of a single ID is compared with the average movement of the entire flock in real time. If an ID remains stationary for more than a first preset time and the cumulative displacement parameter is close to zero, it is identified as an abnormally stationary individual and sent to step S2. If an ID moves with the flock and its physical distance is less than the second preset threshold and its confidence level of the body center point is higher than the third preset threshold, it is identified as a moving individual to be investigated in detail and sent to step S3.

[0025] Furthermore, the two-stage progressive training strategy in step S102 specifically includes:

[0026] Phase 1: Head detection training;

[0027] Using the spatial topology association dataset, the backbone network and the first prediction head of the model are trained on the basis of the YOLOv11-m pre-trained weights. During the training process, the Mosaic data augmentation strategy is introduced to randomly stitch together images containing dead ducks, paralyzed individuals and sparsely distributed samples, so that the model can still establish a stable feature mapping when the head features are damaged or the environmental distribution is sparse.

[0028] Phase Two: Condition-Guided Association Training;

[0029] Based on the head model weights obtained from the first stage of training, while keeping the RGB backbone feature extraction path unchanged, a lightweight feature fusion module is introduced before the pose prediction head. The dynamic heatmap generated by the head center coordinates is used as an auxiliary prior. Only the parameters of the fusion module and the pose prediction head are optimized so that the model learns the conditional mapping relationship between the head position and the body center.

[0030] Furthermore, the asynchronous tracing logic in step S103 specifically includes:

[0031] During operation, the tracker synchronously maintains the head trajectory output by the head detection model and the body center trajectory predicted by the head anchor guidance center perception model.

[0032] When it is detected that the head of an ID is lost due to occlusion, the ID is not retrieved. Instead, the Kalman filter predictor takes over the motion estimation of the individual and automatically and asynchronously switches the tracking weight to the body center point to maintain the continuity of identity through the center point.

[0033] When the obscured head reappears, its currently generated head-body vector is extracted and matched with the historical average head-body feature vector in the ID file. If the error is within a preset threshold, reverse addressing and reconnection are performed to retrieve the original ID.

[0034] Furthermore, step S2 specifically includes:

[0035] S201: High-precision data filtering and key point calibration;

[0036] The original video stream was acquired using an RGB camera and then processed by frame extraction. Individuals in the duck flock that were in the foreground and whose limb outlines were not obscured were selected from the video frames as training samples. Eight biomarkers were then labeled on the selected high-quality individuals. The eight biomarkers specifically include: beak tip, top of head, neck point, body center, left hip point, right hip point, left foot, and right foot.

[0037] S202: Skeleton-aware model training;

[0038] The training process employs a progressive two-stage evolution strategy of feature-space;

[0039] The first stage focuses on basic anatomical feature extraction: dynamic local region of interest cropping and coordinate reprojection are performed on the original image based on the labeled body center point coordinates to generate a local aligned close-up set, enabling the model to learn the geometric topological constraints between key points;

[0040] The second stage is panoramic space adaptive training: the model is placed in the original panoramic image containing background interference, and a dynamic Gaussian heat map generated by the labeled body center point is introduced into the network as an auxiliary prior. By guiding the feature to focus on the target area, the model is guided to only look at the ducks in the foreground without obstruction. Random Gaussian displacement jitter and guidance signal random discarding mechanism are introduced in the training to enable the model to make coordinate corrections in combination with anatomical knowledge.

[0041] During the training phase, the guidance signal originates from manually labeled ground truth data; in the subsequent edge-side inference phase, the guidance signal is output in real time by the perception model in step S1.

[0042] S203: Multi-dimensional parallel diagnostic logic and comprehensive status determination;

[0043] Upon receiving the static anomaly risk signal from step S1, a local region of interest (ROI) image is dynamically cropped with the body center point of the ID as the center. The trained skeleton perception model is then used to perform inference on this ROI image, selectively outputting the coordinates and depth parameters of four core points: the tip of the mouth, the top of the head, the neck point, and the body center. Parallel determination is then performed according to at least one of the following dimensions:

[0044] Spatial distribution deviation judgment: Calculate the pixel distance between the ID and the geometric center of the current duck flock in the 2D coordinate system. If an individual is in a remote area outside the preset range for a long time, it is judged as an out-of-group state.

[0045] Micro-motion monitoring of vital signs: IR infrared speckle flow is invoked in the back region where the body center is located, and the signal within a preset time period is extracted using a sub-pixel displacement tracking algorithm. The presence of fluctuations in vital signs is confirmed by fast Fourier transform energy spectrum.

[0046] Determining extreme paralysis: The physical depth corresponding to the tip of the mouth is retrieved, and the physical height difference between it and the ground reference system is calculated. If the height difference is less than the set value, and the four points of the tip of the mouth, the top of the head, the neck, and the center of the body basically present a horizontal straight line trajectory in the image, then it is determined to be extreme paralysis.

[0047] Neck extension posture analysis: Calculate the ratio of the pixel displacement between the top of the head and the neck to the pixel displacement between the neck and the center of the body. If the ratio is lower than the preset ratio threshold and the static duration exceeds the preset static duration, it is judged as pathological neck retraction.

[0048] Furthermore, the spatiotemporal dynamic gait analysis in step S3 includes:

[0049] S301, Time-domain dynamic signal extraction;

[0050] The discrete coordinate sequences of the left foot, right foot, and the highest point of the head output by the skeleton perception model are converted into continuous dynamic signals with pathological diagnostic value. The continuous dynamic signals include at least the longitudinal displacement time sequence of the foot and the longitudinal displacement time sequence of the head.

[0051] S302, Step Height Signal Reconstruction and Stability Analysis;

[0052] Based on the aforementioned longitudinal foot displacement time-series signal, the relative height signal H(t), reflecting the periodic change in stride height, is reconstructed using the baseline zeroing method. Then, the set of all leg-lifting peaks in the relative height signal H(t) is extracted, and the coefficient of variation of the set of peaks is calculated. As a quantitative indicator of step height instability, among which This represents the maximum value of the foot's longitudinal coordinate within a preset time window. This represents the original vertical pixel coordinates of the foot keypoint at time t. The coefficient of variation representing the height of the leg-lift peak. This represents the set of peak points of the leg-raising wave. The standard deviation of the set P representing the peak heights. This represents the average value of the set P of wave crest heights;

[0053] S303, Instantaneous impact acceleration analysis upon landing;

[0054] Perform a second-order differential processing on the relative height signal H(t) to calculate the instantaneous acceleration of the foot upon landing. And extract its peak value as a quantitative indicator of ground impact intensity;

[0055] S304, Frequency Domain Energy Distribution and Gait Symmetry Analysis;

[0056] A fast Fourier transform is performed on the longitudinal displacement time-series signal of the foot to convert the time-domain signal into a frequency-domain signal. The integral energy intensity below the set frequency band is extracted, and the energy values ​​of the left and right limbs are calculated separately. Then, the frequency band energy asymmetry is calculated according to the following formula:

[0057] ;

[0058] in, This indicates the degree of energy asymmetry within the set frequency band. This represents the integral value of the power spectral density within a set frequency band after the left foot longitudinal displacement time-series signal undergoes a Fast Fourier Transform. This represents the integral value of the power spectral density within a set frequency band after the right foot longitudinal displacement time-series signal has undergone a fast Fourier transform.

[0059] S305: Analysis of Whole-Body Postural Compensation Characteristics

[0060] Based on the time-series signal of the longitudinal displacement of the top of the head, its standard deviation is calculated as a quantitative indicator to measure the vertical posture stability of an individual. The formula for calculating the standard deviation is as follows:

[0061] ;

[0062] in, This represents the standard deviation of the longitudinal displacement time-series signal at the top of the head. This represents the vertical pixel coordinates of the key point on the top of the head in the i-th frame image. This represents the average value of the Y-coordinate at the top of the head within the sampling window, and N represents the total number of frames within the sampling window.

[0063] Furthermore, step S3 also includes a multi-dimensional gait anomaly auxiliary judgment and early warning mechanism:

[0064] The logic controller performs the following decision criteria:

[0065] Maximum landing acceleration determination: The peak value of the instantaneous acceleration a(t) of the foot upon landing exceeds the preset acceleration threshold;

[0066] Step height instability determination: coefficient of variation of the peak of unilateral leg lift Exceeding the preset coefficient of variation threshold;

[0067] Frequency band power asymmetry determination: Frequency band energy asymmetry of both limbs Exceeding the preset asymmetry threshold;

[0068] Determination of center of gravity trajectory dispersion: Standard deviation of head undulation of the target individual Exceeding the preset dispersion threshold;

[0069] When an individual meets any two or more of the above four criteria for abnormality, the individual is deemed to have a risk of gait abnormality, triggering an early warning.

[0070] Furthermore, the method also includes an automatic audit termination mechanism:

[0071] If the target individual fails to trigger the static anomaly judgment criteria in step S2 or the gait anomaly judgment criteria in step S3 within the set continuous detailed investigation period, the detailed investigation task for that individual will automatically end, the current detailed investigation computing power will be released, and the system will return to the global initial screening monitoring state.

[0072] Secondly, the present invention also provides a multimodal sensing cascade monitoring system for intensively farmed Muscovy ducks, used to execute the above-mentioned method. The system includes: a multimodal data acquisition module comprising an RGB camera and an infrared speckle emitter; an edge computing terminal with a built-in deep learning inference engine and Kalman filter tracker, used to perform global initial screening and identity maintenance, static physical characteristic detailed examination, and dynamic gait detailed examination; and a cloud-based management platform used to receive and display the abnormal individual IDs and their tracking trajectories reported by the edge computing terminal. The edge computing terminal is configured to execute a funnel-shaped hierarchical linkage architecture: continuously running global initial screening and identity maintenance at a first computing power cost, and only triggering static physical characteristic detailed examination or dynamic gait detailed examination at a second computing power cost when a suspected abnormal target is identified, wherein the second computing power cost is higher than the first computing power cost.

[0073] Compared with the prior art, the present invention has the following beneficial effects:

[0074] This invention significantly enhances the continuity and robustness of identity tracking in extremely crowded environments. It employs an asymmetric tracking strategy of "head anchoring leading to center of gravity prediction," utilizing the relatively low occlusion rate of head features in dense environments as an initial entry point, and combining this with the body's center of gravity extracted from the feature map as the core tracking reference. Even when extreme occlusion causes a short-term visual loss of head features, the system can still maintain the continuity of identity identification using the motion vector of the center of gravity, effectively solving the problem of frequent "ID loss" or identifier drift caused by pushing, shoving, and crawling.

[0075] This system achieves highly reliable, in-depth diagnosis of the health status of stationary individuals. For seemingly abnormal, immobile individuals, the system achieves accurate classification of Muscovy ducks in a stationary state by cropping local regions of interest and fusing multimodal features such as active infrared speckle displacement, three-point line geometric posture, and group deviation. The introduction of active speckle sensing technology significantly improves the signal-to-noise ratio of physiological micro-movements, enabling the system not only to accurately capture faint respiratory rhythms but also to objectively distinguish between deep sleep, pathological neck retraction, and biological death by combining the geometric proportions of the neck and back, greatly reducing missed detections and false alarms.

[0076] This invention fills the gap in objective dynamic quantification criteria for non-contact gait monitoring. By constructing an integrated three-dimensional dynamic gait audit model, it achieves a technological leap from spatial geometric reconstruction to physical and mechanical assessment. Through quadratic difference processing of the foot spatial coordinate sequence, the system can quantify and extract key dynamic indicators such as the instantaneous impact acceleration upon landing, reflecting the force characteristics of the affected limb. This provides a physical criterion with clear mechanical significance for limp identification and can capture microscopic anomalies in the transition between gait symmetry and landing weight-bearing in individuals with early or mild limpness.

[0077] The computing power efficiency ratio in large-scale real-time monitoring tasks has been optimized. Thanks to the funnel-shaped hierarchical linkage architecture, the system has established an efficient "on-demand wake-up" mechanism. The global L1-level module maintains population counting and initial activity screening at extremely low computing power cost, and only triggers high-precision L2-level vital sign discrimination or L3-level dynamic auditing in a local area when a specific suspected target is identified. This flexible computing power distribution logic significantly reduces the computing load on the backend server while ensuring monitoring accuracy, enabling real-time concurrent monitoring of thousands of ducks throughout their entire life cycle on low-power edge devices. Attached Figure Description

[0078] Figure 1 This is a flowchart of a multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks provided in an embodiment of the present invention;

[0079] Figure 2 This is a superimposed comparison spectrum of the power spectral density (PSD) of the two legs of normal healthy ducks provided in an embodiment of the present invention;

[0080] Figure 3 This is a superimposed comparison spectrum of the power spectral density (PSD) of the two legs of a suspected lame group of Muscovy ducks provided in an embodiment of the present invention;

[0081] Figure 4 This is a bar chart comparing the coefficients of variation of the left and right leg lift peaks provided in an embodiment of the present invention. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0083] This invention provides a multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks, such as... Figure 1 As shown, it specifically includes:

[0084] Phase 1: Initial Global Screening and Identity Maintenance

[0085] The core task at this stage is to establish a stable identity ID for each duck in a high-density farming scenario, track its movement in real time, and route tasks based on the comparison of group activity levels.

[0086] Step 1: RGB Video Acquisition and Feature Annotation

[0087] In actual deployment at the farm, the Orbbec Gemini 335L camera was fixed directly above the duck house, with the lens pointing vertically downwards, covering a breeding area of ​​approximately 20 square meters. The system continuously acquired high-resolution RGB video streams (1920×1080, 20 frames / second) and extracted keyframes for training.

[0088] The annotation process utilizes the AnyLabeling manual annotation tool. Annotators perform two core tasks on the extracted video frames:

[0089] (1) Use a rectangle to define all visible duck head areas in the image and obtain the true coordinates of the head frame (coordinates of the top left and bottom right corners of the rectangle).

[0090] (2) Based on the position and direction of the head and the unobstructed body parts in the picture, mark the coordinates of the center point of each individual's body (single point coordinates).

[0091] The final output is a spatial topological association dataset containing "head bounding box coordinates" and "body center point coordinates". This dataset includes a high proportion of extreme condition samples such as dense occlusion, pathological neck retraction, and paralyzed postures, which are used to support the progressive training of perception models such as YOLOv11 and enhance the model's localization robustness in complex environments.

[0092] In this embodiment, a total of 1,000 frames of images were labeled, with an average of about 58 animals per frame, resulting in a total of about 58,000 head bounding boxes and corresponding body center points.

[0093] Step 2: Two-stage progressive training strategy and tracking feature modeling

[0094] This step is performed on a high-performance cloud server to build a deep learning model with spatial guidance capabilities.

[0095] This embodiment adopts a progressive logic of "first anchoring head features, then inferring the body's center of mass." The necessity of this approach lies in the fact that in high-density farming environments, the torso of Muscovy ducks is prone to frequent and extensive overlapping occlusion, and their movement patterns are extremely variable, resulting in poor robustness of directly regressing to the center of mass. In contrast, Muscovy duck head features have higher visual stability and saliency, a relatively low occlusion probability, and are easier to learn. By first establishing a stable head representation and using it as a localization guide, the system can utilize the inherent spatial constraints between the head and body to help the model "fill in" the body's center of mass position in an interference environment, thereby significantly improving the system's localization stability under high-density occlusion conditions. Therefore, this embodiment performs head detection first.

[0096] Phase 1: Head Detection Training

[0097] Using the head bounding box coordinate data obtained from step 1, the backbone network (CSPDarknet) and the first prediction head of the model are specifically trained based on the YOLOv11-m pre-trained weights. Training parameters: input image size 640×640, batch size=32, initial learning rate 0.01, SGD optimizer, momentum 0.937, weight decay 0.0005, and a total of 200 epochs of training.

[0098] To improve the model's robustness in real-world environments, a Mosaic data augmentation strategy was introduced during training. This strategy involves randomly stitching together four images—one containing a dead duck, one containing a limp individual, and one containing sparsely distributed samples—into a single composite image. For example, an image containing a dead duck (with its neck retracted and stationary), an image containing a limp individual, and two images of healthy individuals are stitched together into a single image. This allows the model to establish stable feature mapping relationships even when head features are impaired (e.g., dead ducks are limp, or there is occlusion between individuals) or the environment is sparsely distributed. After this training phase, the model achieved a recall rate of 96.5% for duck heads and a localization accuracy (mAP@0.5) of 97.2%.

[0099] Phase Two: Condition-Guided Association Training

[0100] This stage uses the head model weights obtained from the first stage of training as the initialization basis, keeping the RGB backbone feature extraction path unchanged, and introduces a lightweight feature fusion module before the pose prediction head. This module uses a dynamic heatmap generated from the head center coordinates as an auxiliary prior, which is applied only to the pose prediction branch to enhance the localization ability of the body center point. Specifically, the head center coordinates are first mapped to a Gaussian heatmap of the same size as the feature map (Gaussian kernel standard deviation σ = 15 pixels). This heatmap is then concatenated with the feature map output by the backbone network and input into the pose prediction head.

[0101] During training, the parameters of the backbone network and detection branches were frozen, and only the parameters of the fusion module and pose prediction head were optimized, enabling the model to learn the conditional mapping relationship between the head position and the body center. Training parameters: initial learning rate 0.001, Adam optimizer, training for 100 epochs. Experimental results show that in severely occluded scenes (body visible area <40%), the accuracy of body center point localization was greatly improved, and the localization error (pixel Euclidean distance) decreased from 12.3 pixels to 5.8 pixels.

[0102] During model inference, the system automatically calculates and maintains the feature vector V in real time, which is defined as the vector difference between the coordinates of the body center point and the coordinates of the head bounding box center:

[0103] V = Coordinates of body center point - Coordinates of head frame center point

[0104] In this system, the modulus of V represents the individual's body proportions (head-to-body distance), while the phase represents the pointing angle of the head and body relative to the image coordinate system (e.g., the head is in front of the left or behind the right of the body). This vector serves as dynamic reference data, providing a reliable tracking foundation for subsequent identity maintenance and reconnection mechanisms. The system maintains a 30-frame historical feature vector queue for each active ID and calculates its moving average as the individual's identity fingerprint.

[0105] Step 3: Improved asynchronous tracing logic that integrates spatial topological associations

[0106] The core algorithm in this step is based on ByteTrack multi-target tracking and Kalman filtering, with modifications to the underlying logic, and is executed by an edge computing terminal. During runtime, the tracker simultaneously maintains two trajectory sets: the head trajectory output by the head detection model and the body center trajectory predicted by the head anchor guidance center perception model.

[0107] The state vector of the Kalman filter is set to [x, y, vx, vy], where (x, y) are the coordinates of the body center point and (vx, vy) are the corresponding velocities. The process noise covariance matrix Q and the observation noise covariance matrix R are set according to the movement characteristics of the Muscovy duck: the diagonal elements of Q are [0.1, 0.1, 0.5, 0.5], and the diagonal elements of R are [1.0, 1.0].

[0108] Asynchronous switching mechanism: When the tracker detects that the head of an ID is lost due to occlusion (detection confidence is below 0.3 for more than 3 consecutive frames), the system does not reclaim the ID. Instead, the Kalman filter predictor takes over the motion estimation of that individual, and the tracking weights are automatically and asynchronously switched to the body center point. The continuity of the identity is maintained through continuous observation of the body center point. At this time, the coordinate update of the ID depends on the Kalman prediction and the observation of the body center point (if any), and no longer depends on the head bounding box.

[0109] Reverse addressing reconnection: When the occluded head reappears (detection confidence recovers to above 0.5), the system extracts the currently generated head-body vector V_current and matches it with the historical average feature vector V_history in the ID file. The matching error is a weighted sum of cosine similarity and relative magnitude error: if the relative magnitude error is <15% and the angle error is <20°, the match is considered successful, and reverse addressing reconnection is performed to forcibly retrieve the old ID. This mechanism effectively suppresses ID switching in dense environments. In a test scenario with 1000 ducks and a density of 15 ducks / square meter, the ID switching frequency of this solution is an average of 2.3 times per minute, while the traditional ByteTrack solution is 28.6 times per minute, a 12-fold improvement.

[0110] Motion parameter extraction: While maintaining tracking, the system extracts the pixel coordinates of the body center point in real time and calculates the Euclidean distance in a sampling period of 30 frames (about 1.5 seconds). The real-time motion parameters of each individual are obtained by accumulating the effective pixel displacement (removing jumps caused by tracking loss).

[0111] The formula for calculating exercise volume is:

[0112] ;

[0113] Where T is the total number of frames within the sampling period (T=30 in this embodiment). Let be the pixel coordinates of the body's center point in the i-th frame. This formula sums the Euclidean distances between adjacent frames to obtain the cumulative pixel displacement of the individual within the sampling period.

[0114] Individuals with a cumulative displacement of less than 5 pixels are considered to be in a stationary state.

[0115] Step 4: Task routing and distribution based on group activity comparison

[0116] While performing tracking tasks, the edge computing terminal aggregates and analyzes the movement data of all IDs within the current frame in real time. The system makes differentiated task flow decisions by comparing the activity performance of individual IDs with the average movement of the entire flock of ducks in real time.

[0117] (1) Static anomaly risk assessment and routing

[0118] If the system detects that the duck flock is generally in an active state (average group movement > 50 pixels / 30 frames), but a certain ID remains stationary for more than 30 minutes and its effective displacement accumulation parameter is close to zero (< 5 pixels / 30 frames), the system determines that the target poses an abnormal risk. At this time, the edge terminal will immediately capture a local region of interest image of the area where the ID is located (size is 1 / 4 of the original image, approximately 480×480 pixels), and extract the center point coordinates predicted by the head anchor guidance center perception model as a guidance signal, which will be sent to stage two for depth determination.

[0119] (2) Individual screening and routing during herd movement

[0120] For IDs that move smoothly with the duck flock, the system initiates a physical scale-based gating screening mechanism. The edge terminal retrieves the physical distance d (in meters) of the individual's head center point from the Gemini 335L depth camera. Only when d < 1.5 meters and the confidence level of the body center point is higher than 0.7 will the system generate a Gaussian heatmap attention signal with a radius of 50 pixels centered on the center point, and route the panoramic image to stage three for refined feature extraction.

[0121] Through the aforementioned dynamic routing mechanism, the system ultimately achieves automated focusing and monitoring of key individuals within a wide-angle panoramic view, avoiding the waste of computing power by running high-consumption models on the entire map.

[0122] Phase Two: Detailed Examination of Static Individual Vital Signs Based on Multimodal Guidance

[0123] This phase targets individuals identified as having static abnormal risk in Phase 1, and performs high-precision multimodal vital sign diagnosis.

[0124] Step 5: High-precision data filtering and key point calibration

[0125] The system uses an Orbbec Gemini 335L camera to acquire raw RGB video streams and performs frame extraction. Annotators use the AnyLabeling tool to perform rigorous sample selection: only individuals in the foreground of the duck flock whose entire body outline is completely unobstructed are selected as training samples. This embodiment yielded 3000 high-quality samples.

[0126] The labelers labeled eight biomarkers for each selected high-quality individual, including: the tip of the beak (T), the top of the head (H), the neck point (N), the body center (S) located between the two wings on the back, the left hip point (L-Hip), the right hip point (R-Hip), the left foot (L-Foot), and the right foot (R-Foot).

[0127] The core of this filtering and labeling logic lies in enabling the model to focus on the skeleton features under ideal conditions, ensuring that the model is not disturbed by overlapping background pixels during the learning phase. This not only provides an accurate geometric benchmark for subsequent diagnosis but also effectively avoids wasting computational resources on invalid pixels during the inference phase by ignoring low-quality targets.

[0128] Step 6: Training the skeleton-aware model

[0129] The training environment for this step is a cloud server, and its goal is to build a basic model that is applicable to both Stage 2 and Stage 3—the "skeleton perception model". This model is customized based on the official YOLO11-pose architecture, with the number of output channels modified to 8 (corresponding to 8 key points). It strictly defines the topological structure of 8 core biological key points, including the body center of gravity (S), the top of the head (H), the tip of the mouth (T), the neck (N), and the two hip points and the soles of the feet on both sides of the lower limbs.

[0130] During model training, the system employs a heatmap guidance strategy: it directly generates a dynamic Gaussian heatmap (Gaussian kernel standard deviation σ = 15 pixels) using the ground truth coordinates of the labeled body center (S) point, and introduces this heatmap as an auxiliary prior into the network. By guiding features to focus on the target region, the model is guided to see only ducks with unobstructed foreground. The model learns this guidance signal, establishing a strong correlation between the guided region and the anatomical topology of the duck. Experiments show that after adding heatmap guidance, the mean keypoint localization error (OKS) in occluded scenes decreased from 0.32 to 0.19.

[0131] The training process employs a progressive two-stage evolutionary strategy of "feature-space". The first stage focuses on basic anatomical feature extraction. The system performs dynamic local region of interest (ROI) cropping and coordinate reprojection on the original image using the labeled centroid (S) as the anchor point, generating a high signal-to-noise ratio (SNR) locally aligned close-up set. By shielding spatial location interference, the model can deeply learn the geometric and topological constraints between keypoints, effectively compensating for the technical limitations of the original monitoring's low resolution. Experimental data show that in this stage, the model achieves a keypoint mean accuracy (mAP50) of 99.45% and a high-precision regression level (mAP50-95) of 90.59% within the local perceptual field, establishing a solid recognition benchmark for subsequent analysis.

[0132] After mastering the structural features, the model enters the second stage of panoramic adaptive training. Here, the model is placed in an original image containing complex background interference. Although the system still provides a dynamic heatmap generated from the target center as an auxiliary prior, random Gaussian displacement jitter and a random discarding mechanism for guidance signals are artificially introduced. This "guidance degradation" strategy simulates the positioning deviation or depth information omission that may occur in head-guided localization models during real-world edge-side inference, forcing the model to autonomously perform secondary verification and coordinate repair based on visual features, incorporating the anatomical knowledge accumulated in the previous stage. Test results show that even under panoramic conditions and guidance drift pressure, the model's keypoint capture accuracy remains stable at 86.33%, demonstrating extremely strong guidance tolerance and independent recognition capabilities.

[0133] The skeleton perception model obtained in this step exhibits excellent autonomous spatial localization robustness. In a low-power environment at the edge, the model can utilize the guidance signal output from the pre-stage as a "selection gate," and output a highly stable keypoint sequence through autonomous topological regression within the gate. This mechanism ensures that even when the guidance signal is "inaccurate or incomplete," the system can still provide a high-purity underlying coordinate data stream for subsequent life / death determination and gait energy spectrum analysis.

[0134] It should be clarified that during the training phase, the guidance signal originates from manually labeled ground truth data; however, in the subsequent edge-side inference phase, this guidance signal is switched to real-time output by the "perception model" (i.e., the head anchor guidance center perception model) from Phase 1. This step ultimately produces a high-performance skeleton extraction model with a strong ability to capture foreground targets and support spatial guidance.

[0135] Step 7: Multi-dimensional parallel diagnostic logic and comprehensive status determination

[0136] This step is executed by the edge computing terminal, which uses the skeleton perception model trained in step 6 to perform a detailed examination of the "static ID". After receiving the risk signal from phase one, the edge terminal dynamically crops a 640×640 pixel local region of interest image with the body center (S) of the ID as the origin. Subsequently, the skeleton perception model guides the focus of features on the target region, performs inference on the local region of interest, and selectively outputs only the coordinates and depth parameters of four core points: the tip of the mouth (T), the top of the head (H), the neck point (N), and the body center (S) (the limb endpoints are not required for static diagnosis).

[0137] The execution entity is determined in parallel according to the following four dimensions:

[0138] (1) Spatial distribution deviation judgment: The system calculates the pixel distance between the ID and the geometric center of the current duck flock (the average position of the body center points of all active individuals) in the 2D coordinate system. In this embodiment, the normal activity range threshold is 300 pixels. If the distance between the individual's center point and the geometric center is >300 pixels and the duration exceeds 1 hour, it is judged as "out-of-group state".

[0139] Specifically, the system obtains the set of two-dimensional coordinates of all tracked duck head IDs in the current frame in real time. And calculate the geometric center of the population. :

[0140] ;

[0141] ;

[0142] Subsequently, the system quantifies and calculates the Euclidean distance between each ID coordinate and the geometric center. As the instantaneous spatial deviation of this individual:

[0143] ;

[0144] Further analysis of the average deviation distance of all individuals in the current image. and deviation from standard deviation The determination is made using dynamically defined spatial constraint thresholds:

[0145] ;

[0146] The coefficient 2.4 in the formula is the optimal hyperparameter determined through repeated manual verification in multiple experiments. Experimental results show that this coefficient can effectively filter out normal spatial disturbances caused by feeding or wandering, ensuring that the system only performs outlier state locking on extremely deviating individuals located on the outer edge of the group (approximately 1.6%), thereby greatly reducing false alarms while maintaining a high recall rate.

[0147] If the real-time distance of the target individual If the judgment condition is met and the state is maintained for more than 1 hour on the timeline, the system will immediately determine that the ID is in an outlier state.

[0148] (2) Monitoring of subtle movements of vital signs: IR speckle flow (i.e., active projection of an infrared speckle emitter to analyze subtle displacements of the speckle pattern) is invoked in the back region (40×40 pixel window) where the body center (S) is located. Signals within 30 seconds are extracted using a subpixel displacement tracking algorithm (LK optical flow method), and the presence of vital signs is confirmed by Fast Fourier Transform (FFT) energy spectrum analysis. The frequency corresponding to the respiratory rhythm is typically between 0.2 and 0.5 Hz (12 to 30 breaths / minute). If a significant energy peak (greater than 3 times the background noise) is detected in the 0.2 to 0.5 Hz frequency band, vital signs are considered present; otherwise, no vital signs are considered present.

[0149] In one embodiment of the present invention, addressing the perception challenge of passive visual optical flow "slipping" due to the lack of texture in the white feathers of Muscovy ducks, the present invention employs an active sensing principle, utilizing hardware to project high-frequency infrared speckle onto the duck's back to artificially construct sub-pixel-level feature point clusters. During the measurement phase, the system core utilizes the Lucas-Kanade sparse optical flow algorithm to perform high-frequency tracking on the speckle clusters, measuring their displacement vector within a 30-second sampling period. The raw signal undergoes detrending preprocessing to eliminate center-of-gravity drift interference, and is then mapped to the frequency domain via a fast Fourier transform to generate an energy spectrum. The system uses the dominant frequency... Calculate the respiratory rate per minute (BPM): .

[0150] Experiments showed that its respiratory energy peaks were precisely distributed within the physiological range of 0.56 Hz to 2.33 Hz. If the energy peaks met the survival energy level threshold criteria:

[0151] ;

[0152] ;

[0153] Indicates the peak energy level. Indicates the survival energy level.

[0154] This allows the system to determine if the target has vital signs, thus accurately distinguishing between deep sleep and biological death in a high-density environment.

[0155] (3) Judgment of beak tip touching the ground and paralysis: The system retrieves the physical depth corresponding to the beak tip (T) point (through depth data from the Gemini 335L camera) and calculates the physical height difference between it and the ground reference system. The ground reference system is obtained by calibrating the depth values ​​of multiple points on the ground of the breeding farm. If the height difference |Z_Beak - Z_Ground| < 50mm, where Z_Beak represents the physical depth value of the beak tip point and Z_Ground represents the depth value of the ground reference system, and the four points B, H, N, and S basically present a horizontal straight line trajectory in the image (the average vertical deviation of the fitted straight line of the four points is < 5 pixels), then it is judged as extreme paralysis (possibly death or serious illness).

[0156] In one embodiment of the present invention, ordinary RGB vision is limited by perspective deviation caused by the installation angle, making it difficult to accurately measure the true vertical height of an individual. This solution uses a 3D depth camera to directly acquire sub-pixel level point cloud data of the target area, providing a physical bottom surface reference for accurate height auditing. The system uses the RANSAC algorithm to perform spatial modeling on the point cloud. This process, through random sampling and iterative consensus comparison in cluttered environmental data, can spontaneously and effectively remove outlier interference points such as duck bodies or bedding, thereby accurately solving the plane equation representing the actual bottom surface of the farm:

[0157] ;

[0158] In the formula, A, B, and C represent the three components of the plane normal vector, and D represents the plane constant term.

[0159] During the judgment process, the system deeply integrates with the hardware-level D2C acceleration module of the depth camera to extract the sub-pixel coordinates of the mouth tip from the skeleton model. Real-time mapping to physical depth value And based on the camera intrinsic parameter matrix Perform a backprojection transformation to reconstruct 2D pixels into 3D physical space coordinates. The system then quantifies the vertical physical height of that point from the ground. :

[0160] ;

[0161] The core logic of the judgment focuses on the loss of the body's ability to support itself. If the height of the tip of the mouth (T) remains below the critical threshold of 50mm for more than one hour, and the tip of the mouth (T), the top of the head (H), the neck (N), and the center of the body (S) exhibit significant horizontal collinearity in 3D space (i.e., the line connecting these four points is essentially parallel to the ground), the system determines that the individual is in an extreme state of paralysis. This indicator uses depth data to reconstruct physical spatial pose, achieving high-reliability identification between pathological paralysis and normal head-tucking sleeping postures.

[0162] (4) Neck extension posture analysis: The system calculates the ratio of the pixel displacement between the top of the head (H) and the neck point (N) to the pixel displacement between the neck point (N) and the body center (S), focusing on measuring the morphological contraction factor that reflects the neck extension ratio:

[0163] ;

[0164] Experimental comparative data show that the true value of the morphological contraction factor in healthy individuals is approximately 0.9405, while in abnormal individuals exhibiting obvious pathological neck retraction, this value drops significantly to around 0.6933. The core judgment logic focuses on the continuous auditing of this morphological feature over time. If the system detects that the target's morphological contraction factor remains in a low range for more than 1 hour, it is immediately quantified as a pathological neck retraction state. This indicator, through dynamic quantification of anatomical proportions, can accurately identify early functional impairments caused by lack of feeding, thus effectively eliminating visual interference from normal head-retracting sleeping postures or momentary head-down preening at the algorithm level.

[0165] In this embodiment, the judgment results of the four dimensions are combined: if any two or more dimensions are abnormal, the system determines that the individual has a static health risk, generates a red warning, and notifies the management personnel through the cloud management platform.

[0166] Phase 3: Detailed Investigation of Dynamic Individual Gait Based on Spatiotemporal Dynamics Analysis

[0167] This phase involves performing refined gait quantitative analysis on the moving individuals selected in Phase 1 who require detailed investigation (with clear foreground, no severe occlusion, and appropriate distance).

[0168] Step 8: Extraction of time-domain dynamic signals

[0169] This step is executed by the edge computing terminal, utilizing the discrete coordinate sequences of the left hip point, right hip point, left foot, right foot, and the highest point of the head output by the skeleton perception model, and converting them into continuous dynamic signals with pathological diagnostic value. Specifically, the Y coordinates (vertical pixel coordinates, top left corner of the image origin, downward is positive) of the left and right feet and the Y coordinate of the head are extracted from each frame of the image to obtain the left and right foot Y(t) sequences and the head Y_head(t) sequence. The sampling window is uniformly set to 75 frames, based on a sampling rate of 20 FPS, the duration is approximately 3.75 seconds, which can ensure the complete capture of 2 to 3 gait cycles. By performing height trajectory reconstruction on the original coordinates, the system generates a relative height fluctuation curve with the ground as the zero point. Through extensive data comparison of the two sets of samples, five core biomechanical indicators were finally identified. The comparative mean values ​​collected in the experiment are shown in Table 1 below:

[0170] Table 1 Biomechanical Indicators

[0171]

[0172] Step 9: Step Height Signal Reconstruction and Stability Analysis

[0173] To quantify the spatial displacement stability of ducks during continuous walking, a step height signal instability assessment model was constructed. The system first extracts the longitudinal coordinates of key foot points and then reconstructs the relative height signal reflecting the periodic changes in step height using the baseline zeroing method. The reconstruction formula is as follows:

[0174] ;

[0175] In the formula, Let H(t) be the maximum value of the longitudinal coordinate of the foot within the sampling window (i.e., the lowest point of the foot, corresponding to the ground baseline), and let Y(t) be the longitudinal coordinate of the foot at time t. When the foot is lifted, Y(t) decreases (because the Y value decreases in the image when moving upwards), therefore H(t) is positive and the larger it is, the higher the foot is lifted.

[0176] After acquiring continuous step height signals H(t), the system accurately captures the set P of all leg-lifting peaks in the signal sequence (i.e., the maximum height of a single step, obtained by finding local maxima) and calculates the coefficient of variation of this peak set. This coefficient is used as the core quantitative indicator for measuring step height instability, and its calculation formula is defined as follows:

[0177] ;

[0178] Combination Figure 4The bar chart data shows that the coefficient of variation objectively captures the characteristic differences in gait rhythm among individuals in different health states. Experimental statistical results indicate that the mean coefficient of variation for the normal healthy group was 0.1951, remaining within a low and stable range, reflecting good consistency in the height of foot lift on both sides of the legs during walking in healthy ducks. In contrast, the mean coefficient of variation for the suspected lame group showed a significant upward trend, reaching an experimental mean of 0.3834. This change in data dimension captures the phenomenon of uneven stride height and increased spatial displacement fluctuations among individuals during walking. Through modeling and analysis of the indicators, the system can effectively eliminate the absolute height interference caused by the difference in duck body size, and provide a quantitative physical basis for the auxiliary identification of gait abnormalities from a statistical perspective.

[0179] Step 10: Instantaneous impact acceleration analysis upon landing

[0180] This step quantifies the impact intensity of an individual's limbs upon contact with the ground by calculating the instantaneous acceleration signal a(t). The system performs quadratic difference processing on the reconstructed step height relative height signal H(t), and its calculation formula is defined as:

[0181] ;

[0182] For a discrete sampling sequence, the acceleration is approximated as:

[0183] , ;

[0184] in Inter-frame time interval (in this embodiment) Second).

[0185] Based on the principles of mechanics, this index reflects the dynamic load characteristics of a unilateral limb at the moment of impact. Experimental statistics show that the average maximum ground acceleration of individuals in the normal healthy group was 71.00 (unit: pixels / frame²), with a smooth and evenly distributed pulse waveform; while the average acceleration of individuals in the suspected lame group significantly increased to 114.00, and the pulse trajectory exhibited obvious abnormal jumps. This surge in numerical dimensions objectively reflects the differences in the mechanical impact generated at the moment of ground contact between individuals.

[0186] Step 11: Frequency Domain Energy Distribution and Gait Symmetry Analysis

[0187] After acquiring the foot spatial displacement sequence of the target individual, the system first uses the Fast Fourier Transform (FFT) algorithm to convert the time-domain signal into a frequency-domain signal. The basic transformation formula is as follows:

[0188] ;

[0189] Through this transformation, the system obtains the power distribution characteristics of foot motion on the frequency axis. Figure 2 and Figure 3 The system-generated superimposed power spectral density (PSD) graphs of the two feet are presented as a comparison. The waveform distribution of these graphs clearly shows that the power spectral density curves of the two feet of healthy individuals highly overlap on the frequency horizontal axis, while the diseased individuals exhibit significant peak misalignment and energy level differences.

[0190] To further quantify the mechanical differences exhibited in the spectrum, the system extracts the energy intensity by performing a definite integral on the area below the 0.5Hz to 4.0Hz frequency band (which covers the main frequency components of the Muscovy duck's gait), using the following formula:

[0191] ;

[0192] The difference in force applied from both sides is measured using the energy asymmetry formula, which is:

[0193] ;

[0194] Analysis of individuals with right-leg lameness revealed a significant feature mapping between the atlas and the data in Table 2. The energy integral for the right foot decreased to 37.84, while the energy value for the left foot reached 620.99, resulting in a jump in asymmetry to 0.885. Coupled with a clear lateral misalignment of the red and blue peaks (dominant frequency shift), these indicators demonstrate that the system can accurately pinpoint the affected limb and quantify the degree of lameness.

[0195] Table 2. Analysis data of individuals with right leg lameness.

[0196]

[0197] Step 12: Analysis of whole-body postural compensation characteristics

[0198] The system extracts the longitudinal displacement sequence of the top of the head and calculates the standard deviation (SD) to assess the displacement dispersion of the target individual along the vertical axis. The calculation formula is as follows:

[0199] ;

[0200] Experimental sampling results showed that the mean SD value of head undulation in the normal healthy group was 18.89 pixels, while the value of this indicator in the suspected lame group increased significantly to 32.73 pixels. This significant change in this data dimension indicates that there is a clear mapping relationship between the degree of disorder in vertical head displacement and gait health (when one side of the body is lame, the duck's head will shake up and down more to maintain balance).

[0201] Through the By modeling the indicator in real time, the system can capture whole-body postural instability caused by abnormal walking function. The introduction of this indicator provides a quantitative physical criterion for locking pathological gait from a whole-body dynamics perspective, significantly improving the overall accuracy of the system's judgment.

[0202] Step 13: Multidimensional Gait Anomaly Assisted Judgment and Early Warning Mechanism

[0203] Through the dynamic characteristic modeling in steps 8 to 12, the system has acquired sufficient quantitative data features to effectively distinguish between normal and lame individuals. Based on statistical comparisons of multiple pairs of lame and normal duck data, this step is executed by the logic controller using the following judgment criteria (each threshold was determined through extensive experimental statistics):

[0204] Maximum landing acceleration determination: The peak value of the instantaneous acceleration a(t) of the foot landing exceeds the preset acceleration threshold (100 in this embodiment, which can be finely adjusted according to the deployment environment).

[0205] Step height instability determination: coefficient of variation of the peak of unilateral leg lift Exceeding the preset coefficient of variation threshold (0.25 in this embodiment);

[0206] Frequency band power asymmetry determination: Frequency band energy asymmetry of both limbs Exceeding the preset asymmetry threshold (0.70 in this embodiment);

[0207] Determination of center of gravity trajectory dispersion: Standard deviation of head undulation of the target individual The dispersion threshold is exceeded (30 pixels in this embodiment).

[0208] The system executes a multi-dimensional feature association early warning logic: If a target individual meets any two or more of the above four criteria simultaneously, the system immediately determines that the individual has a gait abnormality risk. Once the early warning is triggered, the logic controller will immediately mark the target in red in the real-time visual stream and lock its movement trajectory, while only reporting the ID number of the abnormal individual to the cloud management platform in real time.

[0209] For duck individuals in the gait detection phase (foreground and unobstructed), the system has an "automatic audit termination mechanism": if the target individual fails to trigger two or more of the above abnormal judgment criteria within a continuous 10-second detailed inspection cycle, the logic controller will determine that the individual's current gait is normal, automatically end the current gait detection task, release the current detailed inspection computing power, and return to the global initial screening monitoring state.

[0210] The following example, using a specific aquaculture scenario, illustrates the complete workflow of this invention.

[0211] Deployment Environment: A Muscovy duck farm with a duck house area of ​​100 square meters and a stocking density of 15 ducks / square meter, totaling approximately 1500 Muscovy ducks. Six Orbbec Gemini 335L cameras are evenly installed above the duck house, each camera covering an area of ​​approximately 16 square meters. An edge computing terminal is set up, and each camera is connected to the edge terminal via a GigE interface.

[0212] Initial screening phase: After system startup, each camera acquires a video stream in real time at 20 FPS. The global initial screening module processes each frame and assigns a temporary ID to each duck entering the field of view. Within the first 5 minutes, the system establishes a mapping relationship between the ID and the feature vector V. The average movement of the flock is approximately 80 pixels / 30 frames.

[0213] Static Anomaly Detection: After 2 hours of operation, the system detected a duck with ID=237 whose movement was only 2 pixels / 30 frames for 35 consecutive minutes, while other ducks were moving normally. The system extracted a region of interest from ID=237 and sent it for detailed investigation in Phase 2. Detailed investigation results: The distance between the body center and the geometric center of the flock was 520 pixels (>300), the neck extension ratio Ratio was 0.19 (<0.25), and the stillness duration was 35 minutes (>20 minutes), but a weak respiratory rate (0.3Hz) was detected in the trachea area. The overall assessment was "pathological neck retraction and separation from the flock," generating a yellow alert. After receiving notification, the duck was isolated, and an examination revealed an empty crop. It recovered after treatment.

[0214] Gait analysis: Another mallard with ID=512 moved with the flock and was 1.2 meters away from the camera. A detailed examination of the three gait phases was conducted during the trigger phase. System analysis of gait within 75 frames: right leg. (Greater than 0.25), peak landing acceleration of the right leg was 122 (greater than 100), energy asymmetry between the left and right legs was 0.82 (greater than 0.70), and the SD of the top of the head was 36.5 pixels (greater than 30). Three of the four items exceeded the standard, triggering a red alert. The system locked ID=512 in the image with a red box and reported it. The breeders found an external injury on the sole of its right leg and treated it in time, preventing the infection from worsening.

[0215] Computing performance: In the above scenario, the global initial screening module consumes approximately 25% of the GPU. When three individuals are simultaneously undergoing static detailed examination and two are undergoing gait detailed examination, the peak GPU consumption rises to 55%, with an average consumption of approximately 38%. Compared to traditional solutions (which always run the full set of pose estimation, consuming over 85% of the GPU), this invention significantly reduces the computational load and can run stably on edge devices.

[0216] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks, characterized in that, Includes the following steps: Step S1: Global initial screening and identity maintenance; A dual-track tracking system based on head anchor points is used to detect and track individual ducks in a breeding scenario in real time. It maintains the individual's identification ID and activity parameters in real time, and generates task routing decisions based on group activity comparisons. Specifically, this includes: S101, RGB video acquisition and feature annotation; High-resolution RGB video streams from the breeding environment were collected. Head bounding box coordinates were marked for the visible duck head area in the video frames. The coordinates of the individual body center point were marked based on the head position and the unobstructed body parts, forming a spatial topological association dataset of head bounding box coordinates and body center point coordinates. S102, Two-stage progressive training strategy and tracking feature modeling; A deep learning model is constructed, and the first and second training stages are executed sequentially. The head and body feature vector of each individual is calculated and maintained in real time. The head and body feature vector is equal to the coordinates of the body center point minus the coordinates of the head box center. The magnitude of the head and body feature vector represents the body proportion feature, and the phase represents the pointing angle of the head and body relative to the image coordinate system. S103, Improved asynchronous tracing logic that integrates spatial topological associations; The head trajectory and body center trajectory are maintained synchronously. When the head is lost, the Kalman filter predictor takes over and maintains the continuity of identity with the body center point. When the head reappears, reverse addressing and reconnection are performed based on the head and body feature vectors. At the same time, the pixel displacement of the center point is extracted in real time and accumulated as the individual's motion parameters. S104. Task routing and offloading based on group activity comparison; The movement volume of a single ID is compared with the average movement volume of the entire flock in real time. If an ID remains stationary for more than the first preset time and its movement volume parameter is less than the preset movement volume threshold, it is identified as an individual with abnormal stationary risk and sent to step S2. If an ID moves with the flock and its physical distance is less than the second preset threshold and its confidence level of the body center point is higher than the third preset threshold, it is identified as a moving individual to be investigated in detail and sent to step S3. Step S2: Detailed examination of static individual physical signs; When the task routing decision determines that an individual is a static abnormal risk individual, the multimodal vital signs detailed investigation process is triggered. Combining visual, infrared speckle flow and spatial location information, the vital signs, body posture and spatial distribution of the individual are diagnosed in parallel, and the static health status judgment result is output. Step S3: Detailed examination of dynamic individual gait; When the task routing decision determines that an individual is a group-moving individual and meets the preset physical scale threshold, the spatiotemporal dynamics gait analysis process is triggered. The temporal signals of key points of the individual's feet and head during continuous walking are extracted. The degree of gait abnormality is quantitatively evaluated by at least one of frequency domain energy analysis, step height variation coefficient analysis and landing impact acceleration analysis, and the dynamic gait determination result is output.

2. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 1, characterized in that, The two-stage progressive training strategy in step S102 specifically includes: Phase 1: Head detection training; Using the spatial topology association dataset, the backbone network and the first prediction head of the model are trained on the basis of the YOLOv11-m pre-trained weights. During the training process, the Mosaic data augmentation strategy is introduced to randomly stitch together images containing dead ducks, paralyzed individuals and sparsely distributed samples, so that the model can still establish a stable feature mapping when the head features are damaged or the environmental distribution is sparse. Phase Two: Condition-Guided Association Training; Using the head model weights obtained from the first stage of training as the initialization basis, and keeping the RGB backbone feature extraction path unchanged, a lightweight feature fusion module is introduced before the pose prediction head. The lightweight feature fusion module is used to concatenate the dynamic heatmap generated by the head center coordinates with the feature map output by the backbone network. Only the parameters of the lightweight feature fusion module and the pose prediction head are optimized so that the model learns the conditional mapping relationship between the head position and the body center.

3. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 1, characterized in that, The asynchronous tracing logic in step S103 specifically includes: During operation, the tracker synchronously maintains the head trajectory output by the head detection model and the body center trajectory predicted by the head anchor guidance center perception model. When it is detected that the head of an ID is lost due to occlusion, the ID is not retrieved. Instead, the Kalman filter predictor takes over the motion estimation of the individual and automatically and asynchronously switches the tracking weight to the body center point to maintain the continuity of identity through the center point. When the obscured head reappears, its currently generated head-body vector is extracted and matched with the historical average head-body feature vector in the ID file. If the error is within a preset threshold, reverse addressing and reconnection are performed to retrieve the original ID.

4. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 1, characterized in that, Step S2 specifically includes: S201: High-precision data filtering and key point calibration; The original video stream was acquired using an RGB camera and then processed by frame extraction. Individuals in the duck flock that were in the foreground and whose limb outlines were not obscured were selected from the video frames as training samples. Eight biomarkers were then labeled on the selected high-quality individuals. The eight biomarkers specifically include: beak tip, top of head, neck point, body center, left hip point, right hip point, left foot, and right foot. S202: Skeleton-aware model training; The training process employs a progressive two-stage evolution strategy of feature-space; The first stage focuses on basic anatomical feature extraction: dynamic local region of interest cropping and coordinate reprojection are performed on the original image based on the labeled body center point coordinates to generate a local aligned close-up set, enabling the model to learn the geometric topological constraints between key points; The second stage is panoramic space adaptive training: the model is placed in the original panoramic image containing background interference, and a dynamic Gaussian heat map generated by the labeled body center point is introduced into the network as an auxiliary prior. By guiding the feature to focus on the target area, the model is guided to only look at the ducks in the foreground without obstruction. Random Gaussian displacement jitter and guidance signal random discarding mechanism are introduced in the training to enable the model to make coordinate corrections in combination with anatomical knowledge. During the training phase, the guidance signal originates from manually labeled ground truth data; in the subsequent edge-side inference phase, the guidance signal is output in real time by the perception model in step S1. S203: Multi-dimensional parallel diagnostic logic and comprehensive status determination; Upon receiving the static anomaly risk signal from step S1, a local region of interest (ROI) image is dynamically cropped with the body center point of the ID as the center. The trained skeleton perception model is then used to perform inference on this ROI image, selectively outputting the coordinates and depth parameters of four core points: the tip of the mouth, the top of the head, the neck point, and the body center. Parallel determination is then performed according to at least one of the following dimensions: Spatial distribution deviation judgment: Calculate the pixel distance between the ID and the geometric center of the current duck flock in the 2D coordinate system. If an individual is in a remote area outside the preset range for a long time, it is judged as an out-of-group state. Micro-motion monitoring of vital signs: IR infrared speckle flow is invoked in the back region where the body center is located, and the signal within a preset time period is extracted using a sub-pixel displacement tracking algorithm. The presence of fluctuations in vital signs is confirmed by fast Fourier transform energy spectrum. Determination of mouth tip touching the ground and paralysis: retrieve the physical depth corresponding to the mouth tip point, calculate the physical height difference between it and the ground reference system. If the height difference is less than the preset height threshold, and the four points of mouth tip, top of head, neck point and body center present a horizontal straight line trajectory in the image, it is determined to be extreme paralysis. Neck extension posture analysis: Calculate the ratio of the pixel displacement between the top of the head and the neck to the pixel displacement between the neck and the center of the body. If the ratio is lower than the preset ratio threshold and the static duration exceeds the preset static duration, it is judged as pathological neck retraction.

5. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 1, characterized in that, The spatiotemporal dynamic gait analysis in step S3 includes: S301, Time-domain dynamic signal extraction; The discrete coordinate sequences of the left foot, right foot, and the highest point of the head output by the skeleton perception model are converted into continuous dynamic signals with pathological diagnostic value. The continuous dynamic signals include at least the longitudinal displacement time sequence of the foot and the longitudinal displacement time sequence of the head. S302, Step Height Signal Reconstruction and Stability Analysis; Based on the aforementioned longitudinal foot displacement time-series signal, the relative height signal H(t), reflecting the periodic change in stride height, is reconstructed using the baseline zeroing method. Then, the set of all leg-lifting peaks in the relative height signal H(t) is extracted, and the coefficient of variation of the set of peaks is calculated. As a quantitative indicator of step height instability, among which This represents the maximum value of the foot's longitudinal coordinate within a preset time window. This represents the original vertical pixel coordinates of the foot keypoint at time t. The coefficient of variation representing the height of the leg-lift peak. This represents the set of peak points of the leg-raising wave. The standard deviation of the set P representing the peak heights. This represents the average value of the set P of wave crest heights; S303, Instantaneous impact acceleration analysis upon landing; Perform a second-order differential processing on the relative height signal H(t) to calculate the instantaneous acceleration of the foot upon landing. And extract its peak value as a quantitative indicator of ground impact intensity; S304, Frequency Domain Energy Distribution and Gait Symmetry Analysis; A fast Fourier transform is performed on the longitudinal displacement time-series signal of the foot to convert the time-domain signal into a frequency-domain signal. The integral energy intensity below the set frequency band is extracted, and the energy values ​​of the left and right limbs are calculated separately. Then, the frequency band energy asymmetry is calculated according to the following formula: ; in, This indicates the degree of energy asymmetry within the set frequency band. This represents the integral value of the power spectral density within a set frequency band after the left foot longitudinal displacement time-series signal undergoes a Fast Fourier Transform. This represents the integral value of the power spectral density within a set frequency band after the right foot longitudinal displacement time-series signal has undergone a fast Fourier transform. S305: Analysis of Whole-Body Postural Compensation Characteristics Based on the time-series signal of the longitudinal displacement of the top of the head, its standard deviation is calculated as a quantitative indicator to measure the vertical posture stability of an individual. The formula for calculating the standard deviation is as follows: ; in, This represents the standard deviation of the longitudinal displacement time-series signal at the top of the head. This represents the vertical pixel coordinates of the key point on the top of the head in the i-th frame image. This represents the average value of the Y-coordinate at the top of the head within the sampling window, and N represents the total number of frames within the sampling window.

6. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 5, characterized in that, Step S3 also includes a multi-dimensional gait anomaly auxiliary judgment and early warning mechanism: The logic controller performs the following decision criteria: Maximum landing acceleration determination: The peak value of the instantaneous acceleration a(t) of the foot upon landing exceeds the preset acceleration threshold; Step height instability determination: coefficient of variation of the peak of unilateral leg lift Exceeding the preset coefficient of variation threshold; Frequency band power asymmetry determination: Frequency band energy asymmetry of both limbs Exceeding the preset asymmetry threshold; Determination of center of gravity trajectory dispersion: Standard deviation of head undulation of the target individual Exceeding the preset dispersion threshold; When an individual meets any two or more of the four criteria for abnormality, the individual is deemed to have a risk of gait abnormality, triggering an alert.

7. The multimodal sensing cascade monitoring method for intensively farmed Muscovy ducks according to claim 1, characterized in that, This method also includes an automatic audit termination mechanism: If the target individual fails to trigger the static anomaly judgment criteria in step S2 or the gait anomaly judgment criteria in step S3 within the set continuous detailed investigation period, the detailed investigation task for that individual will automatically end, the current detailed investigation computing power will be released, and the system will return to the global initial screening monitoring state.

8. A multimodal sensing cascade monitoring system for intensively farmed Muscovy ducks, characterized in that, The system is used to perform the method according to any one of claims 1 to 7, the system comprising: A multimodal data acquisition module, comprising an RGB camera and an infrared speckle emitter; The edge computing terminal has a built-in deep learning inference engine and Kalman filter tracker for performing global initial screening and identity maintenance, static vital sign detailed examination and dynamic gait detailed examination; The cloud-based management platform is used to receive and display the abnormal individual IDs and their tracking trajectories reported by the edge computing terminals; Furthermore, the edge computing terminal is configured to execute a funnel-shaped hierarchical linkage architecture: it continuously runs global initial screening and identity maintenance at a first computing power cost, and only when a suspected abnormal target is identified will it trigger a static vital signs detailed investigation or a dynamic gait detailed investigation at a second computing power cost, where the second computing power cost is higher than the first computing power cost.