Animal health monitoring and early warning method and system based on intelligent agent
By combining the agent approach with the YOLOv8-DeepSORT algorithm and a veterinary knowledge base, we have achieved 24/7 health monitoring and precise early warning for small-scale animal groups. This solves the problems of insufficient monitoring coverage, waste of computing resources, and delayed early warning in existing technologies, and improves the intelligence and efficiency of breeding management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Existing animal health monitoring methods cannot achieve continuous 24/7 coverage, are difficult to accurately identify behavioral details of small groups of animals, waste computational resources, have lagging early warning mechanisms, and lack professional veterinary knowledge support, leading to misjudgments or omissions.
An agent-based approach is adopted, using the YOLOv8-DeepSORT algorithm for individual identification and tracking, combined with a behavior recognition model and veterinary knowledge base for health assessment, generating early warning signals, and sending early warning information through video clips and keyframe identifiers.
It enables efficient and accurate behavioral identification and health monitoring of small-scale animal groups, optimizes computing efficiency, improves the level of precision and intelligence in breeding management, and enhances breeding efficiency and product quality.
Smart Images

Figure CN121817099A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to animal health monitoring methods, and more specifically to an animal health monitoring and early warning method and system based on intelligent agents. Background Technology
[0002] In modern livestock farming, animal health directly impacts farming efficiency and product quality. Traditionally, animal health monitoring relies primarily on manual inspections by farmers, a method with several limitations: First, manual inspections cannot provide uninterrupted 24 / 7 coverage, especially during critical periods like nighttime, where abnormalities are easily missed; second, for small groups of animals, manual inspections are inefficient, making it difficult to accurately record the behavioral details and duration of each animal; third, the accuracy of health assessments largely depends on the experience of farmers, and varying levels of expertise can lead to misjudgments or missed diagnoses, particularly in the difficulty of promptly detecting behavioral changes caused by early-stage diseases; finally, the early warning mechanism is often delayed, and the process from discovering an anomaly to notifying management can easily result in missing the optimal intervention window.
[0003] While some existing animal behavior monitoring solutions attempt to address these issues using video analytics, they still face challenges: most can only identify single behavioral patterns and lack precise statistics on the duration of behavior; algorithm design does not adequately consider the needs of small groups of animals, resulting in insufficient individual tracking accuracy; they fail to integrate professional veterinary knowledge to build a decision-making system, making it difficult to provide accurate health warnings directly; warning methods are limited, lacking the ability to push text, voice, and abnormal images together, hindering rapid response; furthermore, repetitive feature extraction occurs during target detection and behavior recognition, wasting computational resources and reducing the overall efficiency of the system.
[0004] Therefore, it is necessary to design a new method that can accurately identify animal behavior, automatically collect behavioral data, and combine professional veterinary knowledge to achieve intelligent early warning, optimize computing efficiency, effectively overcome the shortcomings of traditional manual monitoring methods and existing technologies, improve the level of precision and intelligence in breeding management, and thus enhance breeding efficiency and product quality. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide an animal health monitoring and early warning method and system based on intelligent agents.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an animal health monitoring and early warning method based on intelligent agents, comprising:
[0007] Acquire video of animal groups recorded by cameras;
[0008] An object detection model is used to identify and continuously track each frame of the animal group video, generating the position and movement trajectory information of each animal; the object detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism for individual identification and tracking.
[0009] A behavior recognition model is used to analyze the animal's behavior patterns based on the location and movement trajectory information of each animal. Specific behaviors are identified by extracting posture feature points, and relevant data is recorded.
[0010] Based on the relevant data and combined with the professional knowledge in the veterinary knowledge base, a large model inference engine is used to comprehensively assess the health status of animals. When abnormalities are found, corresponding warning signals are issued according to the severity.
[0011] Select relevant video clips and keyframes according to the warning level corresponding to the warning signal, add labels and explanatory text to obtain the warning information;
[0012] Send the aforementioned warning information.
[0013] This invention also provides an agent-based animal health monitoring and early warning system, comprising:
[0014] The video acquisition unit is used to acquire videos of animal groups recorded by a camera.
[0015] The detection and tracking unit is used to perform individual identification and continuous tracking of each frame of the animal group video using a target detection model, and to generate the position and movement trajectory information of each animal; wherein, the target detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism for individual identification and tracking;
[0016] The behavior recognition and statistics unit is used to analyze the behavior patterns of each animal based on its position and movement trajectory information using a behavior recognition model, identify specific behaviors by extracting posture feature points, and record relevant data.
[0017] The comprehensive assessment unit is used to comprehensively assess the health status of animals based on the relevant data and professional knowledge in the veterinary knowledge base, using a large model inference engine. When an abnormality is detected, it issues a corresponding warning signal according to the severity.
[0018] The early warning information generation unit is used to select relevant video clips and keyframes according to the early warning level corresponding to the early warning signal, and add identifiers and explanatory text to obtain early warning information;
[0019] A sending unit is used to send the warning information.
[0020] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0021] The advantages of this invention compared to existing technologies are as follows: This invention acquires video of animal groups via cameras and utilizes the YOLOv8-DeepSORT algorithm combined with a feature reuse mechanism to efficiently and accurately identify and continuously track each animal, generating location and movement trajectory information. Subsequently, a behavior recognition model is used to analyze this information to extract posture feature points, thereby accurately identifying specific behaviors and automatically statistically analyzing behavioral data. Combining the professional knowledge of a veterinary knowledge base, a large-scale model inference engine is used to comprehensively assess the animals' health status and issue corresponding warning signals when abnormalities are detected. Simultaneously, relevant video clips and keyframes are selected, labeled, and explanatory text is added to form warning information for transmission. This method not only optimizes computational efficiency and achieves accurate behavior identification and automatic data statistics, but also effectively overcomes the limitations of traditional manual monitoring methods, significantly improving the refinement and intelligence of aquaculture management, thereby increasing aquaculture efficiency and product quality.
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating the agent-based animal health monitoring and early warning method provided in an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram illustrating the individual standing time provided in an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of individual behavior classification provided in an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of the reasoning process provided for embodiments of the present invention;
[0028] Figure 5 A schematic block diagram of an agent-based animal health monitoring and early warning system provided in an embodiment of the present invention;
[0029] Figure 6 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0032] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0033] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0034] Please see Figure 1 , Figure 1 This is a flowchart illustrating the agent-based animal health monitoring and early warning method provided in this embodiment of the invention. This agent-based animal health monitoring and early warning method is applied in a server. By combining advanced target detection and behavior recognition models (such as the YOLOv8-DeepSORT algorithm and feature reuse mechanism), it achieves accurate identification and continuous tracking of individuals in animal group videos, extracts multi-dimensional correlation features to analyze animal behavior patterns, and then uses a large model inference engine and the professional knowledge of a veterinary knowledge base for in-depth evaluation, automatically filtering out abnormal situations and issuing corresponding early warning signals. This method not only automatically collects behavioral data and optimizes computational efficiency, but also overcomes the inefficiency and subjectivity of traditional manual monitoring methods, improving the precision and intelligence of breeding management, thereby helping to improve breeding efficiency and product quality. Furthermore, through automatic annotation and incremental training mechanisms, the model performance is continuously optimized to ensure long-term effective monitoring and early warning capabilities.
[0035] Figure 1 This is a flowchart illustrating the agent-based animal health monitoring and early warning method provided in an embodiment of the present invention. Figure 1As shown, the method includes the following steps S110 to S170.
[0036] S110. Acquire video of an animal group recorded by a camera.
[0037] In this embodiment, animal group video refers to video data obtained by continuously recording a small group of no more than 20 animals 24 hours a day from a fixed perspective using high-definition network cameras. These cameras are deployed in locations that can cover the entire monitoring area to ensure no blind spots and have night vision enhancement capabilities (such as infrared illumination technology) to ensure video clarity in low-light environments at night. The acquired video data is stored in H.265 encoding format and transmitted to the analysis layer in real time, ensuring real-time data transmission and low latency.
[0038] Specifically, videos of animal groups have the following characteristics:
[0039] Fixed viewing angle coverage: The camera installation location is carefully selected to ensure full coverage of the activity area of the small groups of animals that need to be monitored, avoiding blind spots or dead angles, thereby achieving effective monitoring of all individuals.
[0040] 24-hour continuous recording: The system supports recording around the clock, which provides rich raw data support for accurate behavioral pattern analysis and health status assessment, and helps to capture abnormal behavior at any time.
[0041] High-quality images: Thanks to the high-definition camera and night vision capabilities, video clarity is maintained even in low-light conditions, which is crucial for accurate target detection, tracking, and behavior recognition.
[0042] Standardized data format: Video data is stored in the H.265 encoding format. This efficient compression standard not only reduces the storage space requirement, but also improves the data transmission efficiency, which is beneficial for subsequent data processing and analysis.
[0043] Timestamps: To facilitate data analysis, video data is automatically stamped with timestamps, which allows each video segment to be precisely mapped to a specific point in time, making it particularly important for time-series analysis of events.
[0044] In summary, animal group videos, as one of the core inputs of this system, directly affect the accuracy of target detection, behavior recognition, and ultimately, health warnings. In-depth analysis of this video data enables refined management of the health status of small groups of animals, improving breeding efficiency and product quality.
[0045] S120. An object detection model is used to identify and continuously track each frame of the animal group video, generating the position and movement trajectory information of each animal; wherein, the object detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism to identify and track individuals.
[0046] In this embodiment, the object detection and tracking algorithm mainly refers to the YOLOv8-DeepSORT algorithm.
[0047] Motion trajectory information refers to the record of each animal's positional changes between consecutive frames, obtained by analyzing each frame of an image using the YOLOv8-DeepSORT algorithm combined with a feature reuse mechanism. This information includes not only the exact location of each animal at a given moment (represented by a target bounding box), but also their movement path over time (i.e., motion trajectory). This provides fundamental data support for subsequent behavior recognition and health status assessment.
[0048] In one embodiment, step S120 described above may include steps S121 to S124.
[0049] S121. The animal group video is preprocessed to obtain the preprocessing result.
[0050] In this embodiment, preprocessing results refer to a series of operations performed on the original video data to improve the performance of the object detection model, including but not limited to: adjusting the image size to a uniform size, normalizing the pixel value range, and removing noise. These steps help improve the accuracy and efficiency of subsequent processing steps.
[0051] S122. The total number of animal individuals in each frame of the preprocessed image is estimated in real time using a lightweight counting algorithm.
[0052] This step aims to quickly estimate the number of animals present in the current frame, allowing for dynamic adjustment of the YOLOv8 network parameters to accommodate the monitoring needs of different small groups. This lightweight counting algorithm provides a relatively accurate estimate of the total number of animals without affecting the overall system performance.
[0053] S123. When the total number of animal individuals is not greater than a set threshold, the anchor scale distribution in the object detection network is dynamically adjusted based on the total number of animal individuals, and an integrated network is constructed on the basis of the object detection network to realize single forward inference and form an object detection model.
[0054] In this embodiment, the object detection network refers to the YOLOv8 network.
[0055] In this step, if the total number of detected animals does not exceed a set threshold (e.g., 20), the scale distribution of anchors in the YOLOv8 network is adjusted according to the actual number of animals, thereby optimizing the target detection performance. Simultaneously, an integrated network with pose estimation functionality is built on top of YOLOv8, enabling the simultaneous acquisition of target bounding boxes, confidence scores, appearance features, and key pose points in a single forward inference.
[0056] S124. Input the preprocessing results into the target detection model to obtain the target bounding box, confidence level, appearance features and key pose points, and form the position and movement trajectory information of each animal.
[0057] In this embodiment, the target bounding box refers to the rectangular region surrounding each detected animal individual, used to accurately pinpoint its location in the image.
[0058] Confidence level: Represents the level of confidence of the model in a certain detection result. It is usually a value between 0 and 1. The higher the value, the more reliable the detection result is.
[0059] Appearance characteristics: Feature vectors that describe the appearance characteristics of an individual animal can be used to distinguish different individuals or species and are one of the key bases for individual tracking.
[0060] Key pose points: Coordinate points marked on specific parts of an animal's body (such as the head, torso, limbs, etc.) to capture changes in the animal's posture, thereby supporting behavior recognition.
[0061] In summary, step S120, through frame-by-frame processing of animal group videos and utilizing advanced target detection and tracking technologies, achieved precise localization of individual animals within a small group and tracking of their movement trajectories, laying a solid foundation for subsequent behavioral analysis and health management.
[0062] In one embodiment, step S124 described above may include steps S1241 to S1245.
[0063] S1241. Input the preprocessing results into the target detection model to extract appearance features, position temporal features, and pose similarity to obtain multi-dimensional associated features.
[0064] This step first inputs the preprocessed video frames into the object detection model (YOLOv8-DeepSORT collaborative optimization algorithm). This model not only identifies the location of each individual animal (represented by the target bounding box), but also extracts features from multiple dimensions:
[0065] Physical characteristics: Describes the physical features of an individual animal, such as color and texture.
[0066] Location temporal features: Based on the positional changes of an individual animal in the previous few frames of images, predict its current position and possible future position.
[0067] Posture similarity: Capturing changes in an animal's posture by using key posture points (such as key points on the head, torso, and limbs).
[0068] These features together constitute a multi-dimensional correlation feature containing a variety of information, providing data support for subsequent target tracking.
[0069] S1242. Construct a three-dimensional association matrix from the multi-dimensional association features, and use a lightweight Hungarian matching algorithm to associate the target with the historical tracking ID to obtain the association result.
[0070] After acquiring multi-dimensional correlation features, the system constructs a three-dimensional correlation matrix. This matrix comprehensively considers information from multiple dimensions, including appearance features, temporal location features, and pose similarity. Then, a lightweight Hungarian matching algorithm is used to determine the best match between the target in the current frame and targets with assigned tracking IDs in previous frames. This method effectively reduces computational complexity while improving matching accuracy.
[0071] S1243. Determine whether the target is occluded based on the association result.
[0072] Based on the matching results of step S1242, the system can evaluate the state of each target. If a target fails to match any historical tracking ID in the current frame, or if the matching quality drops significantly, it may mean that the target is occluded by other objects.
[0073] S1244. If the target is occluded, the position of the target is predicted by fitting the historical motion trajectory, and the tracking ID is re-matched based on the predicted position in the next frame image until a match is successful.
[0074] Once the target is confirmed to be occluded, the system uses previously recorded historical motion trajectory data to fit and predict the target's likely location in the next frame. Based on this predicted location, it attempts to re-match the tracking ID in subsequent frames. This mechanism helps maintain a continuous tracking link, ensuring a high tracking success rate even under brief occlusion.
[0075] S1245. Generate a sequence of target images with tracking IDs, including tracking stability data and multi-dimensional correlation features, to determine the location and movement trajectory information of each animal.
[0076] If the target is not obscured, then proceed to step S1245.
[0077] Regardless of whether the target is occluded, a sequence of target images with a tracking ID is ultimately generated. In addition, tracking stability data (such as tracking success rate, occlusion recovery time, etc.) and all multi-dimensional correlation features used for tracking decisions are included. This information is crucial for determining the precise location and movement trajectory of each animal and provides fundamental data support for subsequent behavioral analysis.
[0078] In summary, step S124 achieves efficient and accurate tracking of small groups of individual animals through a series of carefully designed sub-steps. This not only improves the robustness of the system but also effectively addresses challenges such as occlusion, ensuring high-precision positioning and trajectory tracking under long-term monitoring.
[0079] In this embodiment, the input video frames are preprocessed (including noise reduction and normalization) before entering the real-time animal count monitoring area, which is a key basis for subsequent decisions. The system uses a population density perception mechanism to determine whether the current number of animals is less than or equal to 20: if the condition is met, a processing path optimized for small groups is initiated; otherwise, it switches to the conventional large group counting mode. When a small group is confirmed, the system dynamically adjusts the anchor scale and proportion in the YOLOv8 network to adapt to changes in target size under different densities, thereby improving detection accuracy and avoiding the problem of missed detection of small targets or false detection of large targets caused by fixed anchors. At the same time, a "cross-frame feature consistency constraint" is introduced to enhance feature stability in low-light environments (such as nighttime infrared).
[0080] Next, the system invokes the YOLOv8 detection-pose integrated feature extraction network, synchronously outputting multi-task feature results through a single forward inference, including target bounding boxes, appearance features, 14 key pose points (such as head, torso, limbs, etc.), and temporal position features, and applies cross-frame consistency constraints to ensure feature continuity. This step realizes the core concept of "feature extraction once, reuse for multiple tasks," significantly reducing redundant computation overhead and improving overall efficiency.
[0081] The process then splits into two branches: Branch 1 is for tracking association calculation, which is used to match the current frame with the historical tracking ID; Branch 2 is for feature reuse and transfer, which directly encapsulates the extracted pose features and passes them to the downstream behavior recognition module, avoiding repeated calls to the model for pose estimation and achieving efficient utilization of computing resources.
[0082] In branch 1, the system first extracts multi-dimensional association features, which integrates appearance features, location and temporal features, and pose similarity information to construct a three-dimensional association matrix. Then, a lightweight Hungarian matching algorithm is used to optimally match the target in the current frame with historical tracking IDs. This algorithm is optimized for small group scenarios, reducing the complexity from O(n log n) to O(n log n) / n. 3Reduced to O(n) 2 This significantly improves matching efficiency. After matching, the system determines whether the target is occluded. If not, it directly associates and updates the tracking status; if occluded, it activates the occlusion recovery prediction mechanism: by fitting historical motion trajectories, it predicts the possible reappearance location of the target, and in the next frame, it re-matches the tracking ID based on the predicted location until tracking is successfully restored. This ensures a high tracking success rate (≥95%) and fast recovery capability (≤0.5 seconds) in frequent occlusion scenarios.
[0083] Finally, the system encapsulates and outputs a complete result including a target image sequence with a unique ID, tracking stability data (such as the number of consecutive frames and the time of loss), and reused pose feature points. On the one hand, it feeds back to the self-learning module as data support for sample selection and model optimization, and on the other hand, it pushes it to the behavior recognition and data statistics module for subsequent behavior analysis and health assessment, forming a complete "collection-analysis-decision-feedback" closed loop.
[0084] In summary, in the scenario of monitoring small groups of animals, through multiple technological innovations such as group density perception, dynamic parameter adjustment, multi-task feature fusion, lightweight matching algorithm and occlusion recovery mechanism, efficient, stable and low-latency individual tracking and feature reuse are achieved, providing a solid data foundation for subsequent intelligent decision-making.
[0085] This embodiment focuses on the intelligent monitoring of small groups (≤20 animals), innovatively employing a "small group adaptive YOLOv8-DeepSORT collaborative optimization algorithm" and introducing a mechanism of "one-time feature extraction, multi-task reuse". This method not only solves the problem of repetitive computation caused by traditional module separation, but also optimizes the core algorithm for the needs of small group monitoring.
[0086] Firstly, regarding the optimization of feature extraction through multi-task fusion, a "detection-pose integrated feature extraction network" was constructed based on YOLOv8. Through a single forward inference, this network can simultaneously output the target bounding box, confidence score, appearance features, and 14 key pose feature points (e.g., 2 points for the pig's head, 4 points for the torso, and 8 points for the limbs). This avoids the repeated use of YOLOv8 for feature extraction in the behavior recognition stage, thereby improving computational efficiency by more than 40% and meeting the low resource consumption requirements in small-group monitoring scenarios.
[0087] Secondly, considering the characteristics of animal groups at different densities, the system implements dynamic adjustment of anchors for group density perception. By dynamically adjusting the size and proportion of anchors based on real-time statistics of the number of animals in the monitoring area (ensuring no more than 20 animals), the system prevents missed detection of small targets or false detection of large targets due to fixed group density. Simultaneously, a "cross-frame feature consistency constraint" is introduced to improve the stability of feature extraction in low-light environments (such as nighttime infrared), resulting in an initial detection mAP exceeding 92%.
[0088] To optimize the tracking and association algorithm, the Hungarian matching algorithm in DeepSORT was improved, and a multi-dimensional association matrix of "appearance features + location temporal features + pose similarity" was established. A lightweight matching strategy was designed for small groups of no more than 20 animals, reducing the computational complexity of matching from O(n log n) to O(n log n). 3 Reduced to O(n) 2 This significantly improves processing speed. In addition, the added "occlusion recovery prediction mechanism" can predict the possible location of the target after occlusion by fitting historical motion trajectories, which improves the tracking success rate by 8%-10% and shortens the occlusion recovery time to no more than 0.5 seconds. It is very suitable for dealing with scenarios where individuals in small groups are prone to occlusion.
[0089] Finally, regarding the optimization of feature transfer, the system uses the extracted appearance features for binding to individual tracking IDs and directly encapsulates pose features for synchronous output to the behavior recognition stage, achieving feature reuse without redundancy. This optimization method ultimately achieves a tracking frame rate of ≥25fps and a tracking success rate of ≥95% (initial state) in small groups (≤20 individuals), while significantly reducing the overall consumption of computing resources.
[0090] In summary, the method of this embodiment, through the above-mentioned technological innovations, especially in multi-task fusion feature extraction, dynamic adjustment of group density perception anchors, optimization of tracking association algorithms, and optimization of feature flow, effectively improves the performance and efficiency of small group animal monitoring systems, and provides reliable data support for subsequent intelligent decision-making.
[0091] S130. Using a behavior recognition model, analyze the animal's behavior pattern based on the position and movement trajectory information of each animal, identify specific behaviors by extracting posture feature points, and record relevant data.
[0092] In this embodiment, relevant data refers to various data related to individual animal behavior, including but not limited to behavior type, start time, end time, and duration. This data is crucial for understanding the animal's health status.
[0093] In one embodiment, step S130 may include steps S131 to S136.
[0094] S131. The position and movement trajectory information of each animal are normalized in size, and the dynamic correlation and difference in several consecutive frames are calculated to construct a posture change vector describing the behavior pattern.
[0095] In this embodiment, the posture change vector refers to a series of values generated by analyzing the changes in animal posture feature points in consecutive frames, used to represent the animal's behavioral pattern over a period of time. For example, when walking, the limb feature points will exhibit periodic displacement, while the distance between the head and the feeding trough will change in a specific way when feeding.
[0096] S132. Transform the attitude change vector into semantically meaningful behavioral features to obtain the current attitude change vector;
[0097] In this embodiment, the current posture change vector refers to the abstract numerical posture change vector that has been transformed into interpretable behavioral features using a preset algorithm or model. For example, specific actions such as "walking" and "feeding" can be identified by the system through this transformation.
[0098] S133. Compare the current posture change vector with a preset template library to initially determine the behavior category and obtain the comparison result.
[0099] In this embodiment, the comparison result refers to matching the transformed behavioral features with standard templates in the database to identify the most likely corresponding behavioral category. The comparison result is one or more potential behavioral categories and their corresponding confidence scores.
[0100] S134. Output the individual's behavior label within a specific time window based on the comparison result, and remove behaviors whose duration does not meet the requirements to obtain the final behavior judgment result.
[0101] In this embodiment, the final behavior judgment result refers to the final behavior category being filtered and confirmed based on the comparison results of the previous step, combined with certain filtering conditions (such as behavior duration). For example, behaviors lasting less than 0.2 seconds may be considered a misjudgment and rejected.
[0102] S135. Based on the final behavior judgment result, record the start time, end time and duration of each behavior to obtain individual behavior data.
[0103] In this embodiment, individual behavior data refers to the detailed records of each identified behavior, including start time, end time, and duration. This information is crucial for subsequent statistical analysis.
[0104] S136. Summarize all individual behavior data, statistically analyze the distribution ratio of various behaviors and the frequency of abnormal behavior by period, and generate a structured report based on the statistical data to obtain relevant data.
[0105] Finally, the behavioral data collected from each individual will be summarized and analyzed to determine the distribution of different behavioral types and the frequency of any abnormal behaviors. These analytical results will be compiled into easy-to-understand and use structured reports to provide decision support for aquaculture managers.
[0106] Through this series of steps, the method of this embodiment enables precise monitoring and analysis of animal behavior within small groups, which helps to promptly identify potential health problems and thus take appropriate preventive measures.
[0107] In this embodiment, behavior recognition employs a "YOLO behavior classification algorithm with temporal consistency constraints on posture," and optimizes the process by leveraging the feature reuse mechanism of the preceding modules. It is specifically designed for the precise identification needs of small groups (≤20 individuals). Its core innovation lies in directly reusing the posture feature points output by the integrated target detection-tracking-feature reuse module. This eliminates the need to repeatedly call YOLOv8 for feature extraction, completing identification solely through "temporal feature fusion-behavior matching-filtering," significantly improving efficiency. Specifically, this method introduces a "Lightweight Temporal Feature Fusion Unit" (LTFU) to dynamically correlate and calculate the differences between reused posture feature points across 10 consecutive frames (equivalent to a 0.4-second time window), constructing a "posture change vector." For example, during walking, limb feature points exhibit periodic displacement vectors; while during feeding, the distance between the head and the feeding trough undergoes specific changes. Based on a pre-defined "posture change vector - behavior category" mapping library, the system uses cosine similarity matching to determine behavior and incorporates a "behavior duration filtering mechanism" to eliminate misidentified behaviors shorter than 0.2 seconds, such as brief posture adjustments. This ensures an initial behavior recognition accuracy of at least 85%. Furthermore, in terms of data statistics, the system accurately records the start and end times and duration of various behaviors for each individual animal within a small group, constructing an individual behavior time-series database. Simultaneously, it statistically analyzes the distribution percentage of various behaviors within the group and the frequency of abnormal behaviors according to pre-defined periods (e.g., hours, days), ultimately generating a statistical report including text descriptions (e.g., "Individual ID1: 08:00-08:30 feeding, lasting 30 minutes") and charts (e.g., individual behavior time-series curves, group behavior distribution pie charts). Concurrently, the system also calculates the accuracy of behavior recognition, which can be obtained through manual sampling verification or stability tracking, forming recognition quality feedback data. This data is input into the self-learning module to further optimize and improve system performance. This entire process not only improves the accuracy of behavior identification, but also provides breeding managers with detailed behavior analysis reports, which helps to identify potential health problems in a timely manner and take corresponding measures.
[0108] Among these, reusing posture feature points mainly refers to a mechanism designed in animal health monitoring and early warning systems to improve computational efficiency and reduce resource consumption. Specifically, it includes the following aspects:
[0109] Single-step feature extraction supports multiple tasks: In the integrated module of object detection-tracking-feature reuse, a "detection-pose integrated feature extraction network" is adopted. This network can simultaneously output the target bounding box, confidence score, appearance features, and key pose feature points (e.g., 2 points for the head, 4 points for the torso, and 8 points for the limbs of a pig) in a single forward inference process. This means that it is no longer necessary to repeatedly extract feature points for each frame of video data, thus saving a significant amount of computational resources.
[0110] Feature transfer optimization: The extracted pose feature points are directly used in the behavior recognition process, achieving efficient reuse of feature points. That is, in the behavior recognition and data statistics module, there is no need to call the YOLOv8 model again to extract features; instead, the previously obtained pose feature points are directly used to complete behavior recognition through methods such as temporal feature fusion. This not only improves processing speed but also reduces resource waste caused by redundant computation.
[0111] Improving computational efficiency: Through this "one-time feature extraction, multi-task reuse" mechanism, the computational efficiency of the entire system is significantly improved. The paper mentions that this method can improve computational efficiency by more than 40%, making it very suitable for monitoring small groups (no more than 20 animals), as small groups typically require lower resource consumption and higher real-time performance.
[0112] In summary, "reusing pose feature points" refers to a technical strategy used in video analysis where features are extracted from the same frame in a single operation and shared among multiple tasks to improve computational efficiency and reduce resource consumption. This mechanism is one of the keys to achieving efficient system operation.
[0113] Specifically, the system receives target data with unique IDs and reused pose feature points from the preceding module. This data already includes the location, tracking ID, and 14 key pose feature points (such as head, torso, and limbs) for each individual animal, eliminating the need for repeated extraction and thus avoiding computational redundancy. The system then enters a lightweight preprocessing stage, performing only basic operations such as size standardization on the input data. Since the pose features have already been extracted and reused in the preceding module, there is no need to call the model again for feature extraction, significantly reducing computational overhead.
[0114] Next, the system feeds 10 consecutive frames (approximately 0.4 seconds) of reused pose feature points into a lightweight temporal feature fusion unit (LTFU). This unit captures pose change trends through dynamic correlation and difference calculation, constructing a "pose change vector" with behavioral semantics. For example, walking behavior corresponds to periodic displacement vectors of the limbs, while feeding behavior is reflected in the change vector of the distance between the head and the feeding trough. This process realizes the transformation from static pose to dynamic behavior patterns, improving the accuracy of behavior recognition.
[0115] Subsequently, the system uses a cosine similarity matching algorithm to compare the constructed posture change vector with a pre-defined "posture change vector-behavior category" mapping library to initially determine the current behavior category, such as standing, lying down, side-lying, walking, or foraging. This matching mechanism, based on similarity calculation in a high-dimensional vector space, possesses good robustness and generalization ability.
[0116] After the initial classification is completed, the system introduces a behavior duration filtering mechanism to post-process the recognition results. This mechanism eliminates transient behaviors with a duration shorter than a set threshold (such as 0.2 seconds), such as falsely identified behaviors like briefly getting up to adjust posture, effectively reducing the false alarm rate and improving the reliability of the recognition results.
[0117] The filtered results are output as precise individual ID-behavior category mappings, i.e., the specific behavioral labels for each animal within a specific time window. Based on this, the system enters the individual behavior statistics stage, recording the start time, end time, and duration of each behavior, and constructing an individual behavior time-series database to provide data support for subsequent long-term analysis.
[0118] Next, the system performs group behavior statistics, calculating key indicators such as the distribution ratio of various behaviors in the group and the frequency of abnormal behavior according to preset periods (such as hours or days), so as to comprehensively reflect the overall activity patterns and potential health risks of the group.
[0119] Finally, the system generates a structured statistical report, including text descriptions (such as "Individual ID1: Feeding from 08:00 to 08:30, lasting 30 minutes") and visual charts (such as individual behavior time-series curves and group behavior distribution pie charts), intuitively presenting the behavioral data. This report is pushed to the agent's decision-making and early warning module as input for health status assessment. Simultaneously, the system statistically analyzes the accuracy of behavior recognition (which can be verified through manual sampling or by tracking stability), generating recognition quality feedback data. This feedback is then fed back to the self-learning and reinforcement learning optimization module to drive continuous model iteration and performance improvement.
[0120] The steps in this embodiment range from feature reuse to temporal modeling, then to behavior determination and statistical analysis, ultimately achieving closed-loop feedback optimization. This fully demonstrates the system's efficiency, accuracy, and intelligence in small group scenarios.
[0121] S140. Based on the relevant data and combined with the professional knowledge in the veterinary knowledge base, a large model inference engine is used to comprehensively assess the health status of the animal. When an abnormality is found, a corresponding warning signal is issued according to the severity.
[0122] In this embodiment, the warning signal refers to different levels of attention or action instructions generated based on the analysis of animal behavior data, which are used to indicate the urgency of potential health problems.
[0123] Step S140 focuses on using expertise from the veterinary knowledge base and a large-scale model inference engine to comprehensively assess animal health and issue early warning signals based on the severity of abnormalities. This process involves multiple sub-steps, including data format conversion, structuring, verification and cleaning, initial disease association analysis, deep inference, and early warning generation, aiming to ensure the accuracy of the assessment and the timeliness of the warnings.
[0124] In one embodiment, step S140 described above may include steps S141 to S145.
[0125] S141. Perform format conversion and structuring processing on the relevant data to obtain the processing result.
[0126] In this embodiment, the processing result refers to the specific measures taken in response to the warning signal and their effect feedback, including whether the health problem has been resolved or whether further observation is needed.
[0127] First, the system standardizes the received data (such as individual / group behavior statistics reports output by the behavior recognition and data statistics module). This includes converting raw text descriptions, time-series data, chart metadata, and other information into a data format conforming to the JSON-LD 1.1 standard, defining the semantic mapping of each field using `@context`, and using `@type` to identify data entity types, ensuring consistency and traceability in data parsing. The result of this step is a structured, machine-readable behavior statistics report.
[0128] S142. Apply predefined data validation rules to filter out data with missing key information or incorrect format in the processing results, and add timestamps and source identifiers to the cleaned data to obtain the processed data.
[0129] In this embodiment, the processed data refers to a set of structured information that has been cleaned, verified, and analyzed to support decision-making or further processing.
[0130] Next, predefined data validation rules are applied to filter data entries that are missing key information or have incorrect formats. After filtering, a timestamp and source identifier are added to each valid data entry to ensure data integrity and credibility. The resulting "processed data" not only contains accurate behavioral statistics but also includes timestamps and data source descriptions, enhancing data reliability and transparency.
[0131] S143. Using structured rule expressions in the knowledge base, and through precise matching and fuzzy matching techniques, perform an initial disease association analysis on the processed data to determine the probability of meeting the requirements, a list of suspected associated diseases, and their confidence scores, so as to obtain preliminary matching results.
[0132] In this embodiment, the preliminary matching result refers to the list of diseases related to the target animal's behavioral pattern and their correlation scores, which are selected from the veterinary knowledge base based on a rule-based matching algorithm.
[0133] Based on the processed data, preliminary disease association analysis is conducted using predefined structured rule expressions from a veterinary knowledge base, employing both exact and fuzzy matching techniques. This step aims to determine the probability of meeting the requirements, a list of potentially associated diseases, and their confidence scores—the preliminary matching results. This process considers not only the direct association between individual behavioral characteristics and diseases but also complex scenarios involving multiple behavioral combinations, thereby more comprehensively capturing potential health problems.
[0134] S144. The preliminary matching results are combined with structured behavioral time-series data and auxiliary features as input. The fine-tuned lightweight large model is used for deep reasoning. By introducing a thinking chain mechanism, the degree of behavioral abnormality, the degree of disease symptom fit, and the weight of environmental influence factors are analyzed step by step to generate multi-dimensional evaluation results. The multi-dimensional evaluation results include health status assessment, warning level, ranking of associated diseases, and confidence level.
[0135] In this embodiment, the multi-dimensional assessment result refers to the comprehensive evaluation of the animal's health status by combining multiple factors (such as the degree of behavioral abnormality, environmental influence, etc.) in the deep reasoning process of the large model, and giving the corresponding health status classification and warning level.
[0136] The initial matching results, along with structured behavioral time-series data and auxiliary features (such as breeding environment parameters, season, animal breed / age, etc.), are used as input to perform deep reasoning using a fine-tuned lightweight large model. By introducing a thought chain mechanism, factors such as the degree of behavioral abnormality, the degree of disease symptom fit, and the weight of environmental influences are analyzed step by step to generate multi-dimensional assessment results, including health status assessment, warning level, ranking of associated diseases, and confidence level. This method can effectively improve the accuracy and reliability of health assessment.
[0137] S145. Based on the multi-dimensional evaluation results and the corresponding warning levels, when an abnormal situation is detected, a corresponding warning signal is automatically issued according to the severity.
[0138] Finally, based on the multi-dimensional assessment results and corresponding warning levels, when an anomaly is detected, an appropriate warning signal is automatically issued according to its severity. These warning signals may include different levels of text warnings (general attention, requiring verification, and emergency response), accompanied by relevant visual evidence (image or video clip links) to help managers make quick response decisions.
[0139] In summary, step S140, through a series of carefully designed sub-steps, achieves fully automated management of the entire process from data collection to health assessment and early warning issuance, greatly improving the efficiency and accuracy of health management for small groups of animals.
[0140] In this embodiment, step S140 relies on professional veterinary knowledge and intelligent reasoning capabilities to achieve accurate assessment and differentiated early warning of animal health status. It integrates three core components: "veterinary knowledge base + large model reasoning engine + intelligent agent interface". The core algorithm is "multimodal data-driven health reasoning matching algorithm". Through the full-link design of "knowledge construction - data adaptation - two-stage reasoning - early warning generation - image evidence", it takes into account the professionalism of decision-making, reasoning efficiency and readability of early warning.
[0141] To ensure the professionalism, usability, and dynamic adaptability of the knowledge base, a complete end-to-end solution of "knowledge extraction - fusion - structured storage - iterative optimization" was adopted. First, in the knowledge sourcing and precise extraction stage, the resources covered three types of core authoritative resources, ensuring the comprehensiveness and reliability of the knowledge. These resources included authoritative veterinary textbooks (such as *Animal Infectious Diseases* and *Veterinary Clinical Diagnosis*), large-scale livestock clinical diagnosis and treatment cases from the past decade (including over 1000 typical disease-behavior correlation cases), and practical experience collected from over 50 senior livestock veterinarians through semi-structured interviews, which were then transformed into structured text. In the extraction stage, NLP technology fine-tuned with BERT was used for entity recognition and relation extraction, automatically extracting core entities and their relationships such as "disease name - typical behavioral characteristics - disease cycle - treatment measures - susceptible animal species," and supplemented by manual review and correction to avoid any possible extraction errors.
[0142] The next stage is knowledge fusion and structured storage. Based on the Neo4j graph database, a detailed structured knowledge graph was constructed, clarifying the node types and their relationships. Core nodes include disease nodes, behavior nodes, animal breed nodes, and environmental factor nodes; while relationships cover terms such as "associated behavior," "susceptible breed," and "triggering factors." Furthermore, key "animal behavior-disease association rules" were converted into computable structured expressions. For example: "Pig AND lying down behavior duration > 6 hours AND feeding behavior frequency → associated diseases: swine fever (confidence 0.85) / streptococcal infection (confidence 0.72)." This expression not only supports fuzzy matching of rules but also prioritizes rules based on the severity of the disease and the association confidence level.
[0143] Finally, a manually reviewed, iterative model was adopted for the dynamic iteration mechanism of the knowledge base to ensure the timeliness and accuracy of the knowledge. Managers regularly review the behavior-disease association rules in the knowledge base based on the actual disease occurrence and early warning response feedback in real-world farming scenarios. If any rules are found to be misjudged, the association confidence of the relevant rules can be directly updated or new constraints can be added (such as adding a "season = winter" restriction to enhance the effectiveness of the swine fever association rules). This approach allows the knowledge base to flexibly adapt to changes in farming scenarios, ensuring its long-term effectiveness and practicality. This end-to-end approach not only improves the quality of the knowledge base but also enhances its value in practical applications.
[0144] To build a complete data interaction mechanism and ensure the standardization and reliability of data input to the inference engine, the process involves four steps:
[0145] Data reception scope: Receive "individual / group behavior statistics reports" (including text descriptions, time-series data, and chart metadata) output by the behavior recognition and data statistics module through a standardized intelligent agent interface.
[0146] Format Conversion and Structuring: Data standardization and conversion are strictly performed in accordance with the JSON-LD 1.1 standard. Field semantics are defined using `@context`, and `@type`, field definitions, and data type constraints are clearly defined to ensure data parsing and traceability. Simultaneously, chart metadata is converted into a structured array. Specific standard JSON-LD format and chart data examples are as follows: ① Standard JSON-LD format example: {
[0147] @context:{
[0148] "@vocab":"http: / / example.org / animal-health-monitoring / vocab#",
[0149] "Data type":"DataType",
[0150] "Individual ID":"IndividualID",
[0151] "Statistical Period"
[0152] "BehaviorList",
[0153] "BehaviorType":"BehaviorType",
[0154] Duration": "Duration",
[0155] "Occurrence Period"
[0156] Chart Metadata: "ChartMetadata"
[0157] },
[0158] "@type":"BehaviorStatisticalReport",
[0159] Data type: "Behavioral statistics report",
[0160] "Individual ID":"Pig_01",
[0161] "Statistical period":{
[0162] "@type":"TimePeriod",
[0163] "Start Time":"2025-08-01T00:00:00",
[0164] End Time: "2025-08-01T24:00:00"
[0165] },
[0166] "List of Behaviors":[
[0167] {
[0168] "@type":"Behavior",
[0169] "Behavior type": "Lying down"
[0170] Duration: 7.5 hours
[0171] "Time period of occurrence":[
[0172] {
[0173] "@type":"TimePeriod",
[0174] Start Time: "2025-08-01T02:30:00",
[0175] End Time: "2025-08-01T10:00:00"
[0176] } ]
[0178] },
[0179] {
[0180] "@type":"Behavior",
[0181] "Behavior type": "feeding"
[0182] Duration: "0h"
[0183] "Time period of occurrence":[]
[0184] } ]
[0186] }
[0187] like Figure 2 and Figure 3 As shown, the JSON-LD defines the semantic mapping of each field through @context and clarifies the data entity type through @type, ensuring that large models can accurately parse data relationships; the chart metadata clearly presents core information such as time series and proportion through a structured array, directly supporting multi-dimensional reasoning and analysis of large models.
[0188] Introduce validation rules to filter low-quality data: Filtering conditions: missing core fields (such as individual ID, duration of behavior), incorrect data format (such as duration being non-numerical); Data traceability: add timestamps and source identifiers to the cleaned data to form a traceable structured data set.
[0189] To balance inference efficiency and evaluation accuracy, a two-stage strategy of "rapid filtering through rule matching + deep inference using a large model" was adopted. In the first stage, rapid filtering through rule matching, structured rule expressions from the knowledge base were used to perform initial screening through precise matching and fuzzy matching. For cases with a strong correlation between behavior and disease (e.g., poultry exhibiting lying-on-the-side behavior accompanied by convulsions directly corresponding to Newcastle disease), precise matching was used; while for more complex scenarios involving multiple behavioral combinations (e.g., pigs lying on the ground for more than 5 hours, exhibiting low activity and no feeding behavior, potentially pointing to swine fever or influenza), fuzzy matching was applied for analysis. The output of this stage includes a list of highly probable associated diseases (confidence ≥70%), a list of suspected associated diseases (confidence 30%-70%), and results with no associated diseases. Each result was labeled with its matching basis (e.g., behavioral field and rule ID) to ensure the traceability of the entire process.
[0190] The second stage, deep inference using a large model, employed a finely tuned lightweight Llama 3-8B model to optimize the input-output flow and improve inference efficiency. In this stage, preliminary matching results, structured behavioral time-series data, and auxiliary features (such as breeding environment parameters, season, animal breed / age, etc.) were first encapsulated into parseable prompt word templates as input. Next, a thought chain mechanism was introduced during inference to guide the model in a deeper analysis across multiple dimensions, from the degree of behavioral abnormality to the fit of disease symptoms and the weight of environmental influences, ultimately ranking the likelihood of disease occurrence. Finally, at the output layer, the model generates a multi-dimensional assessment report, covering health status (normal / suspected abnormal / severe abnormal), warning level (levels 1 to 3, representing general concern, need for verification, and emergency treatment, respectively), associated diseases ranked by probability and their confidence levels, and key influencing factors (e.g., summer season increases the probability of heat stress-related diseases). This two-stage strategy not only efficiently screens potential diseases but also provides more accurate health assessments and warning information, offering strong support for practical operation.
[0191] S150. Select relevant video clips and keyframes according to the warning level corresponding to the warning signal, and add labels and explanatory text to obtain warning information.
[0192] In this embodiment, the early warning information refers to a notification generated based on detected abnormal animal behavior, which includes keyframe image links, video clip links, and explanatory text. It aims to be quickly conveyed to management personnel through multiple channels so that appropriate measures can be taken in a timely manner.
[0193] In one embodiment, step S150 described above may include steps S151 to S156.
[0194] S151. Determine the individual ID, statistical period, and abnormal behavior based on the warning signal.
[0195] First, the individual ID of the affected animal, the time period (statistical period) during which the abnormal behavior occurred, and the specific type of abnormal behavior are identified based on the warning signal. This step is based on the analysis results of animal behavior from the previous modules.
[0196] S152. Select several images as keyframes at key moments of abnormal behavior, and extract video clips, adding time watermarks and individual IDs.
[0197] Next, for the confirmed abnormal behavior, the system selects images from several key moments as keyframes and extracts corresponding video clips from the original video. These moments typically include the start of the abnormal behavior, the most obvious characteristics, and typical moments during the ongoing process. Simultaneously, time watermarks and individual IDs are added to these keyframes and video clips to ensure traceability and clarity.
[0198] S153, Save keyframes and video clips and generate a unique access link.
[0199] Selected keyframes and video clips are saved to a local server, and a unique access link is generated for each file. This link will be used to construct subsequent early warning information, allowing administrators to directly click the link to view relevant video data.
[0200] S154. Add behavioral feature identifiers to keyframes and write explanatory text for each warning level, including anomaly description, risky disease, and recommended measures.
[0201] Key behavioral features are marked on the keyframes, and detailed explanatory text is written for different warning levels. This text typically includes a description of the abnormal behavior, possible associated diseases, and recommended measures. For example, for a Level 3 warning (emergency response), in addition to describing the abnormal behavior in detail, the possible disease names and confidence levels are clearly stated, and emergency treatment recommendations are given.
[0202] S155. Based on the warning level, embed the corresponding keyframe link or video clip link and explanatory text into the warning information.
[0203] Depending on the warning level, relevant keyframe links or video clip links, along with previously written explanatory text, are embedded in the warning information. This way, the warning information received by management personnel not only includes textual explanations but also allows them to quickly view relevant video footage by clicking on links, thus gaining a more intuitive understanding of the specific situation.
[0204] S156. Push early warning information through multiple channels and record push logs to track feedback.
[0205] Finally, warning information is pushed to managers through multiple channels (such as mobile apps, SMS, and voice calls), and a log of the push is recorded, including the push time, reception status, and feedback from managers. This not only ensures that warning information is delivered to relevant personnel in a timely and accurate manner, but also facilitates subsequent tracking and management of the progress of warning events.
[0206] Through the steps S151 to S156 described above, the system can effectively transform detected anomalies into highly actionable early warning information, helping farms to better manage and protect animal health.
[0207] In this embodiment, as Figure 4 As shown, differentiated early warning content is generated based on the reasoning results, and image evidence is extracted to help managers respond quickly. First, personalized early warning information is generated for different levels: Level 1 warning (general attention) aims to remind managers of minor abnormal behavior, such as an animal with individual ID Pig_01 lying on the ground for 4.5 hours on August 1, 2025, slightly exceeding the normal threshold of 3 hours, suggesting increased patrols and observation; Level 2 warning (verification required) emphasizes situations requiring immediate action, such as the same ID lying on the ground for 6.2 hours without eating on the same day, with associated diseases including swine fever and streptococcal infection, suggesting checking body temperature and mental state and isolating for observation; Level 3 warning (emergency response) describes the most serious situation, such as the individual lying on the ground with convulsions and no eating or drinking behavior within the same period, suspected of acute swine fever, requiring immediate emergency measures such as isolation, disinfection, and notification of a veterinarian.
[0208] To provide more intuitive evidence, the system precisely extracts keyframes and video clips based on the time period of the abnormal behavior. The selected keyframes typically cover the start time of the abnormal behavior, the moment when the characteristics are most obvious, and several instants in the typical process, with 3-5 images chosen for each abnormal situation. The video clips are extracted from 10 seconds before and after the abnormal period, using MP4 format and H.265 encoding, and are accompanied by a time watermark and individual ID to ensure traceability. These image files will be stored on a local server, and a unique access link will be generated with embedded warning text for easy viewing.
[0209] S160. Send the warning information.
[0210] Subsequently, in step S160, the system sends these warning messages to ensure that they are received by the target recipient in a timely manner.
[0211] Furthermore, the information push module employs a multi-channel collaborative push strategy to ensure that early warning information reaches management personnel quickly. The push content differs depending on the warning level: Level 1 warnings only send text messages; Level 2 warnings include text and links to images of abnormal behavior; and Level 3 warnings add a voice broadcast function with accompanying video clips of abnormal behavior. Simultaneously, the system regularly pushes periodic evaluation reports from the self-learning module, allowing management personnel to understand the model optimization progress. The push methods integrate mobile terminal apps, SMS, and voice calls, with voice broadcasts utilizing TTS technology and allowing users to customize the broadcast frequency. Throughout the process, the system backend records detailed push logs, including push time, reception status, and feedback results, thus forming a complete closed loop for early warning handling.
[0212] S170. Automatically filter out useful information based on newly collected data samples, use the information for automatic labeling and incremental training, and optimize the target detection model and behavior recognition model.
[0213] In one embodiment, step S170 described above may include steps S171 to S175.
[0214] S171. Based on image quality, target integrity, and recognition confidence criteria, select data samples with unique identifiers and behavioral classifications to obtain a sample set.
[0215] First, based on criteria such as image quality, target integrity, and recognition confidence, data samples with unique identifiers (e.g., individual IDs) and behavioral classifications are selected from the newly collected data. This step aims to ensure that the sample set used for subsequent processing is of high quality and high relevance, thereby improving the effectiveness of model training. For example, only frames with image sharpness exceeding a preset threshold, target integrity of no less than 90%, and recognition confidence within a specific range are selected for inclusion in the sample set.
[0216] S172. Use an object detection algorithm to pre-label the sample set, including bounding boxes, key points, and behavior labels, and calculate the overall confidence score.
[0217] Next, YOLOv8 is used to pre-annotate the selected sample set, which includes the generation of bounding boxes, keypoints, and behavior labels. This process provides preliminary structured information for each sample. Afterward, the system calculates the overall confidence score as a standard for evaluating the accuracy of the pre-annotation results. The data generated in this stage will serve as the basis for further optimization.
[0218] S173. Dynamically adjust the number of frozen layers based on the sample size, and introduce dynamic L2 regularization.
[0219] The core of this step is to dynamically adjust the number of frozen layers based on the sample size. For small sample sizes, most parameters of the backbone network can be frozen, with only the output layers fine-tuned. As the sample size increases, more feature fusion layer parameters are gradually unfrozen to better adapt to new data. Furthermore, introducing dynamic L2 regularization helps prevent overfitting and ensures the model's generalization ability.
[0220] S174. Design a reward function that combines tracking success rate, recognition accuracy and feature consistency, and optimize the parameters of the target detection model and behavior recognition model through reinforcement learning to obtain the updated target detection model and behavior recognition model.
[0221] We design a reward function that combines tracking success rate, recognition accuracy, and feature consistency to guide the reinforcement learning process. This reward function drives the optimization of model parameters by quantifying the model's performance in these three aspects. This approach can effectively improve the performance of object detection and behavior recognition models, enabling them to more accurately capture animal behavioral features.
[0222] S175. Conduct a comprehensive evaluation of the updated target detection model and behavior recognition model, and check whether the updated target detection model and behavior recognition model meet the predetermined performance indicators.
[0223] Finally, a comprehensive evaluation is performed on the updated object detection and behavior recognition models. This includes checking whether they meet the predetermined performance metrics, such as higher tracking success rates and behavior recognition accuracy. If the requirements are met, these models can be deployed in practical applications; otherwise, it may be necessary to return to the previous steps for further optimization.
[0224] In summary, step S170, through a series of carefully designed technical means, achieves full automation of the process from data screening and automatic labeling to model optimization, which greatly improves the adaptability and accuracy of the model and provides a solid guarantee for the efficient operation of the animal health monitoring and early warning system.
[0225] Step S170 of this embodiment improves the identification accuracy of specific small groups (≤20 animals). Its core algorithm innovation lies in the "fusion algorithm of small-sample incremental learning and reinforcement learning." This algorithm is specifically designed for the characteristics of small groups with limited data and easily captured individual differences. Through a complete closed-loop process of "sample selection-labeling-training-feedback-deployment," it achieves precise adaptation to the monitored small groups within just two days. Firstly, the core positioning of this module is to utilize 24-hour uninterrupted data collection and a closed-loop learning mechanism to quickly adjust the model to adapt to the individual characteristics of different animals, such as coat color, body size, and posture habits, thereby improving the accuracy of the target detection and behavior recognition modules. The optimization objective focuses on two key indicators: individual tracking success rate and behavior recognition accuracy, aiming to achieve both indicators reaching or exceeding 98% after the two-day learning cycle.
[0226] To ensure the quality and diversity of training samples, a "confidence-oriented dynamic sample selection algorithm" was employed. This process consisted of three stages: First, two types of input data were received in real time—"sequences of consecutive frames with IDs" output from the target detection and tracking module and "recognition results and accuracy feedback" output from the behavior recognition module. Next, low-quality samples were filtered based on four core selection criteria: image sharpness, target integrity, recognition confidence, and coverage of multiple poses / behaviors by a single ID sample. Finally, sample accumulation was completed in two phases over a two-day timeframe. The first phase focused on collecting basic samples, with at least 500 frames collected for each individual; the second phase concentrated on supplementing pose samples with low recognition accuracy to improve the overall sample library.
[0227] Subsequently, the process proceeds to automatic annotation and manual verification, employing an efficient model of "model pre-annotation + minimal manual verification" to balance efficiency and accuracy. First, pre-annotation is performed on a pre-trained YOLOv8 model capable of object detection, pose recognition, and preliminary behavior classification. The pre-annotation process includes standardization preprocessing, multi-branch annotation, calculation of overall annotation confidence, result standardization, and preliminary filtering, ultimately generating a candidate annotation set. Then, every 12 hours, 10% of the candidate samples are manually reviewed, focusing on low-confidence samples and correcting potential errors to ensure high-quality final training data. Next comes the freeze-fine-tuning adaptive incremental training phase, a novel method based on the YOLOv8 incremental training framework designed to balance model fit and overfitting risk. The number of freeze layers is flexibly adjusted based on the number of samples per ID, and dynamic L2 regularization is introduced to ensure generalization ability and avoid overfitting. Furthermore, to continuously optimize model performance, a reinforcement learning feedback mechanism is designed to drive continuous model improvement through rewards and penalties. The reward function comprehensively considers multiple dimensions, including individual tracking success rate, behavior recognition accuracy, and feature consistency of consecutive frames with the same ID, and includes reasonable constraints to ensure that the model conforms to the normal behavioral timeline patterns in actual aquaculture scenarios. Finally, after the above series of steps, the optimal model is selected based on metrics such as mAP, F1 score, and tracking success rate, and deployed to the object detection and behavior recognition module using a "hot update" method. This not only ensures the stable operation of the system but also establishes a continuous optimization mechanism, so that when new anomalies occur, a new round of small-batch incremental learning can be automatically triggered, ensuring that the model adapts to the dynamically changing aquaculture environment in the long term. This comprehensive and meticulous design enables this module to significantly improve the system's ability and accuracy in recognizing specific small groups in a short period of time.
[0228] Specifically, through an automated, closed-loop incremental learning mechanism, the target detection and behavior recognition model for a specific small group (≤20 animals) can be rapidly optimized within 2 days. The entire process uses "continuous frame sequences with IDs" and "behavior recognition results and accuracy feedback" as input data sources, constructing a complete closed loop from sample collection, automatic annotation, model training to performance evaluation and deployment.
[0229] First, the system receives video frame sequences with IDs from the target detection module and recognition results and accuracy feedback from the behavior recognition module, serving as the foundational data for subsequent processing. Then, the system initiates a "confidence-oriented dynamic sample selection rule," using multiple dimensions such as image clarity, target completeness, recognition confidence interval ([50%, 90%]), and individual pose / behavior diversity to perform high-quality screening of the raw data, ensuring high availability and representativeness of the training data. Based on this, a two-stage sample collection period of 48 hours begins: the first stage (0-24 hours) focuses on collecting basic samples, ensuring each individual ID has at least 500 frames of valid data; the second stage (24-48 hours) focuses on supplementing samples, prioritizing the completion of poses or behavior types with lower recognition accuracy in the early stages, thereby improving the coverage and balance of the sample library.
[0230] After data collection, the system automatically pre-annotates the sample set using a pre-trained YOLOv8 model, including bounding boxes, keypoints, and behavior labels, and calculates the overall annotation confidence score. Subsequently, the pre-annotation results are filtered, removing low-quality samples with a confidence score below 50%, and a stratified candidate annotation set is generated based on confidence score, providing efficient support for manual verification. Every 12 hours, the system automatically selects 10% of the candidate samples for manual review, focusing on correcting annotation errors in low-confidence samples, ultimately generating high-quality, effective training data.
[0231] Next, the system enters the "freeze-fine-tuning adaptive incremental training" phase. The number of frozen layers in the model is dynamically adjusted based on the sample size of each individual: when there are fewer than 300 frames, 90% of the backbone network parameters are frozen, and only the output layer is fine-tuned; when the sample size reaches or exceeds 300 frames, 20% of the feature fusion layer parameters are moderately unfrozen to enhance the model's ability to learn individual characteristics. Simultaneously, a "dynamic L2 regularization" mechanism is introduced, adjusting the regularization coefficients in real time based on sample diversity to prevent overfitting and ensure generalization ability.
[0232] During training, the system further introduces a "reinforcement learning feedback mechanism," constructing a multi-index weighted reward function that comprehensively considers "individual tracking success rate," "behavior recognition accuracy," and "consistency of features in consecutive frames with the same ID," with weights of 0.4:0.4:0.2 respectively. When the tracking or recognition performance of an individual ID significantly improves (e.g., tracking success rate increases by ≥5%, recognition accuracy increases by ≥3%), a positive reward is given, increasing the weight of its corresponding feature branch; conversely, if tracking loss or recognition error occurs, a negative penalty is applied, triggering parameter fine-tuning. Furthermore, the system incorporates "normal behavioral temporal patterns" from the veterinary knowledge base as a rationality constraint, prioritizing samples that do not conform to common sense for inclusion in the next round of learning, ensuring that the model's reasoning conforms to biological behavioral logic.
[0233] After the training is completed, the system evaluates the performance of the updated model. If the model performance improvement exceeds 3% and conforms to the temporal pattern of the veterinary knowledge base, it is determined to be qualified, and the new model is deployed to the target detection and behavior recognition module using the "hot update" method to achieve seamless replacement and ensure the stable operation of the system. If the performance fails to meet the standard or violates the knowledge constraints, the corresponding samples are included in the next round of learning and returned to the sample collection link for re-iteration, forming a continuous optimization loop.
[0234] Finally, the system continuously monitors the recognition quality of the previous module. Once new abnormal behaviors or recognition deviations are detected, a new round of small-batch incremental learning is triggered to achieve dynamic adaptation within 2 days, ensuring that the model maintains high precision and strong robustness in the long term and fully meets the actual needs of large individual differences and fast environmental dynamic changes in small-group farming scenarios. The entire process realizes an integrated intelligent optimization from data-driven to knowledge-constrained and from automatic learning to continuous evolution, which is the key technical path for the method in this embodiment to achieve efficient model adaptation under small-sample conditions.
[0235] In this embodiment, advanced technical means are used to achieve real-time monitoring, analysis, and early warning of the health status of small groups of animals with no more than 20 individuals, so as to improve the intelligence and refinement level of farming management. The technical route of the entire system can be summarized as: data collection, integrated feature extraction, tracking + behavior recognition (feature reuse), knowledge matching, decision-making early warning, information push, and self-learning iteration, forming a complete closed-loop work process.
[0236] First, the video acquisition module uses a fixed-angle high-definition camera to continuously collect video data of the animal group for 24 hours and transmits it to the analysis layer for subsequent processing in real time. A data marking interface is reserved during this process for convenient later data processing and analysis. Secondly, in the integrated module of target detection-tracking-feature reuse, the initial basic model is loaded and the YOLOv8-DeepSORT collaborative optimization algorithm is used to simultaneously output the target position, appearance features, and pose feature points with a single feature extraction and complete individual ID binding. This not only improves efficiency but also avoids repeated calculations and ensures the effective use of resources.
[0237] In the behavior recognition and data statistics module, the pose feature points from the previous step are directly reused, and the "YOLO behavior classification algorithm based on pose temporal consistency constraints" is used to judge the behavior type of animals, record the corresponding temporal data, and generate a statistical report including text and charts. These data are crucial for evaluating the performance of the model. To maintain the efficiency and accuracy of the model, the self-learning and reinforcement learning optimization module performs incremental updates every 24 hours based on the newly input data and hot-deploys new model versions to ensure the continuous optimization of the system.
[0238] After the intelligent agent interface transforms the behavioral statistics report into structured data, it inputs it into the large-scale model inference engine, which, combined with the veterinary knowledge base, performs an accurate assessment of health status. Based on the inference results, if an anomaly is detected, an early warning mechanism is triggered, including the extraction and display of abnormal images. Finally, the information push module pushes early warning information and self-learning progress reports to management personnel through various means such as an app, SMS, or voice calls, ensuring timely response and handling.
[0239] The innovation of this embodiment lies in its algorithms and technical solutions designed specifically for small-group characteristics, such as the YOLOv8-DeepSORT collaborative optimization algorithm and the integrated detection-pose feature extraction and reuse mechanism. These significantly improve the overall efficiency and accuracy of the system. Furthermore, by introducing self-learning and reinforcement learning mechanisms, rapid adaptation and accurate identification can be achieved even with limited data, ensuring long-term identification accuracy and system stability. This allows the system to not only meet the monitoring needs of specific small groups but also possess strong adaptability and convenient deployment, making it widely applicable in various small-group farming scenarios. It effectively reduces the risk of disease transmission and farming losses, improving the overall efficiency of farming management.
[0240] Therefore, compared with existing technologies, the method of this embodiment has significant advantages, such as improved monitoring comprehensiveness and accuracy, optimized computational efficiency, upgraded decision-making professionalism, high response efficiency, strong adaptability and convenient deployment, and guaranteed long-term stability. These features enable the system not only to meet the monitoring needs of specific small groups, but also to have broad application prospects, effectively reducing the risk of disease transmission and aquaculture losses, and improving the overall efficiency of aquaculture management.
[0241] The aforementioned agent-based animal health monitoring and early warning method acquires video of animal groups via cameras and utilizes the YOLOv8-DeepSORT algorithm combined with a feature reuse mechanism to efficiently and accurately identify and continuously track each animal, generating location and movement trajectory information. Subsequently, a behavior recognition model analyzes this information to extract posture feature points, thereby accurately identifying specific behaviors and automatically compiling behavioral data. Combining the professional knowledge of a veterinary knowledge base, a large-scale model inference engine is used to comprehensively assess the animal's health status and issue corresponding level early warning signals when abnormalities are detected. Simultaneously, relevant video clips and keyframes are selected, labeled, and explanatory text is added to form early warning information for transmission. This method not only optimizes computational efficiency and achieves accurate behavior identification and automatic data statistics but also effectively overcomes the limitations of traditional manual monitoring methods, significantly improving the precision and intelligence of aquaculture management, thereby increasing aquaculture efficiency and product quality.
[0242] Figure 5This is a schematic block diagram of an animal health monitoring and early warning system 300 based on an intelligent agent, provided in an embodiment of the present invention. Figure 5 As shown, corresponding to the above-described agent-based animal health monitoring and early warning method, the present invention also provides an agent-based animal health monitoring and early warning system 300. This agent-based animal health monitoring and early warning system 300 includes a unit for executing the above-described agent-based animal health monitoring and early warning method, and the system can be configured in a server. Specifically, please refer to... Figure 5 The intelligent agent-based animal health monitoring and early warning system 300 includes a video acquisition unit 301, a detection and tracking unit 302, a behavior recognition and statistics unit 303, a comprehensive evaluation unit 304, an early warning information generation unit 305, and a sending unit 306.
[0243] The video acquisition unit 301 is used to acquire videos of animal groups recorded by a camera; the detection and tracking unit 302 is used to perform individual identification and continuous tracking on each frame of the animal group video using a target detection model, generating the position and movement trajectory information of each animal; wherein, the target detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism for individual identification and tracking; the behavior recognition and statistics unit 303 is used to analyze the animal's behavior patterns based on the position and movement trajectory information of each animal using a behavior recognition model, identify specific behaviors by extracting posture feature points, and record relevant data; the comprehensive evaluation unit 304 is used to perform a comprehensive evaluation of the animal's health status based on the relevant data and combined with professional knowledge in the veterinary knowledge base, using a large model inference engine, and issuing corresponding warning signals according to the severity of any abnormalities detected; the warning information generation unit 305 is used to select relevant video segments and keyframes according to the warning level corresponding to the warning signal, add labels and explanatory text to obtain warning information; and the sending unit 306 is used to send the warning information.
[0244] In one embodiment, the system further includes:
[0245] The model update unit 307 is used to automatically filter out useful information based on newly collected data samples, and use the information for automatic annotation and incremental training to optimize the object detection model and behavior recognition model.
[0246] In one embodiment, the detection and tracking unit 302 includes:
[0247] The preprocessing subunit is used to preprocess the animal group video to obtain preprocessing results; the quantity estimation subunit is used to estimate the total number of animal individuals in each frame of the preprocessed image in real time using a lightweight counting algorithm; the construction subunit is used to dynamically adjust the anchor scale distribution in the object detection network based on the total number of animal individuals when the total number of animal individuals is not greater than a set quantity threshold, and to construct an integrated network on the basis of the object detection network to achieve single forward inference and form an object detection model; the detection processing subunit is used to input the preprocessing results into the object detection model to obtain the target bounding box, confidence, appearance features and key pose points, forming the position and motion trajectory information of each animal.
[0248] In one embodiment, the detection processing subunit includes:
[0249] The system comprises the following modules: an association module, which inputs the preprocessed results into the target detection model to extract appearance features, temporal location features, and pose similarity to obtain multi-dimensional association features; an ID association module, which constructs a three-dimensional association matrix from the multi-dimensional association features and uses a lightweight Hungarian matching algorithm to associate the target with historical tracking IDs to obtain association results; an occlusion judgment module, which determines whether the target is occluded based on the association results; a matching module, which, if the target is occluded, predicts the target's location by fitting historical motion trajectories and re-matches the tracking ID based on the predicted location in the next frame image until a match is successful; and a result generation module, which generates a target image sequence with tracking IDs, results containing tracking stability data and multi-dimensional association features to determine the location and motion trajectory information of each animal.
[0250] In one embodiment, the behavior recognition and statistics unit 303 includes:
[0251] The system comprises the following subunits: a vector construction subunit, which normalizes the size of the position and trajectory information of each animal and calculates the dynamic correlations and differences in several consecutive frames to construct a posture change vector describing the behavior pattern; a vector transformation subunit, which transforms the posture change vector into semantically meaningful behavioral features to obtain the current posture change vector; a comparison subunit, which compares the current posture change vector with a preset template library to preliminarily determine the behavior category to obtain a comparison result; a elimination subunit, which outputs the individual's behavior label within a specific time window based on the comparison result and eliminates behaviors whose duration does not meet the requirements to obtain the final behavior judgment result; a recording subunit, which records the start time, end time, and duration of each behavior based on the final behavior judgment result to obtain individual behavior data; and a summarization subunit, which summarizes all individual behavior data, statistically analyzes the distribution ratio of various behaviors and the frequency of abnormal behavior occurrences by period, and generates a structured report based on the statistical data to obtain relevant data.
[0252] In one embodiment, the comprehensive evaluation unit 304 includes:
[0253] The data processing subunit performs format conversion and structure processing on the relevant data to obtain the processing results. The data processing subunit applies predefined data validation rules to filter out data with missing key information or incorrect formats from the processing results, and adds timestamps and source identifiers to the cleaned data to obtain the processed data. The multiple matching subunit utilizes structured rule expressions in the knowledge base, employing precise matching and fuzzy matching techniques to perform initial disease association analysis on the processed data, determining the probability of meeting requirements, a list of suspected associated diseases, and their confidence scores to obtain preliminary matching results. The reasoning subunit combines the preliminary matching results with structured behavioral time-series data and auxiliary features as input, uses a fine-tuned lightweight large model for deep reasoning, and introduces a thought chain mechanism to progressively analyze the degree of behavioral abnormality, disease symptom fit, and environmental influence weighting factors to generate multi-dimensional evaluation results. These multi-dimensional evaluation results include health status assessment, warning level, associated disease ranking, and confidence level. The warning signal generation subunit automatically issues corresponding warning signals based on the severity of detected abnormalities, according to the multi-dimensional evaluation results and the corresponding warning level.
[0254] In one embodiment, the early warning information generation unit 305 includes:
[0255] The system comprises the following sub-units: a determination sub-unit, used to determine the individual ID, statistical period, and abnormal behavior based on the warning signal; an addition sub-unit, used to select several images as keyframes at key moments of abnormal behavior, extract video clips, and add time watermarks and individual IDs; a saving sub-unit, used to save keyframes and video clips and generate unique access links; a text addition sub-unit, used to add behavioral feature identifiers to keyframes and write explanatory text for each warning level, including abnormal descriptions, risk diseases, and recommended measures; a connection sub-unit, used to embed corresponding keyframe links or video clip links and explanatory text into the warning information according to the warning level; and a push sub-unit, used to push warning information through multiple channels and record push logs to track feedback.
[0256] In one embodiment, the model update unit 307 includes:
[0257] The sample set generation subunit is used to filter data samples with unique identifiers and behavior classifications based on image quality, target integrity, and recognition confidence criteria to obtain a sample set. The computation subunit is used to pre-annotate the sample set using YOLOv8, including bounding boxes, keypoints, and behavior labels, and to calculate the overall confidence score. The adjustment subunit is used to dynamically adjust the number of frozen layers based on the sample size, introducing dynamic L2 regularization. The design subunit is used to design a reward function that combines tracking success rate, recognition accuracy, and feature consistency, and to optimize the parameters of the target detection model and behavior recognition model through reinforcement learning to obtain updated target detection and behavior recognition models. The comprehensive evaluation subunit is used to comprehensively evaluate the updated target detection and behavior recognition models to check whether they meet predetermined performance indicators.
[0258] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned intelligent agent-based animal health monitoring and early warning system 300 and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0259] The aforementioned agent-based animal health monitoring and early warning system 300 can be implemented as a computer program, which can, for example... Figure 6 It runs on the computer device shown.
[0260] Please see Figure 6 , Figure 6 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0261] See Figure 6The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0262] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an agent-based animal health monitoring and early warning method.
[0263] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0264] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an agent-based animal health monitoring and early warning method.
[0265] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The processor 502 is used to run the computer program 5032 stored in the memory to implement all steps of the agent-based animal health monitoring and early warning method.
[0266] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0267] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0268] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform all the steps of the agent-based animal health monitoring and early warning method. The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0269] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0270] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0271] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the system of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0272] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0273] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An animal health monitoring and early warning method based on intelligent agents, characterized in that, include: Acquire video of animal groups recorded by cameras; An object detection model is used to identify and continuously track each frame of the animal group video, generating the position and movement trajectory information of each animal; the object detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism for individual identification and tracking. A behavior recognition model is used to analyze the animal's behavior patterns based on the location and movement trajectory information of each animal. Specific behaviors are identified by extracting posture feature points, and relevant data is recorded. Based on the relevant data and combined with the professional knowledge in the veterinary knowledge base, a large model inference engine is used to comprehensively assess the health status of animals. When abnormalities are found, corresponding warning signals are issued according to the severity. Select relevant video clips and keyframes according to the warning level corresponding to the warning signal, add labels and explanatory text to obtain the warning information; Send the aforementioned warning information.
2. The animal health monitoring and early warning method based on intelligent agents according to claim 1, characterized in that, After sending the warning information, the method further includes: Useful information is automatically filtered out from newly collected data samples, and this information is used for automatic annotation and incremental training to optimize the object detection model and behavior recognition model.
3. The animal health monitoring and early warning method based on intelligent agents according to claim 1, characterized in that, The method employs a target detection model to perform individual identification and continuous tracking on each frame of the animal group video, generating location and movement trajectory information for each animal, including: The animal group video is preprocessed to obtain the preprocessing result; The total number of animal individuals in each frame of the preprocessed image is estimated in real time using a lightweight counting algorithm. When the total number of animal individuals is not greater than a set threshold, the anchor scale distribution in the object detection network is dynamically adjusted based on the total number of animal individuals, and an integrated network is constructed on the basis of the object detection network to realize single forward inference and form an object detection model. The preprocessing results are input into the target detection model to obtain the target bounding box, confidence level, appearance features and key pose points, forming the position and movement trajectory information of each animal.
4. The animal health monitoring and early warning method based on intelligent agents according to claim 3, characterized in that, The preprocessing results are input into the target detection model to obtain the target bounding box, confidence score, appearance features, and key pose points, forming the position and movement trajectory information of each animal, including: The preprocessing results are input into the target detection model to extract appearance features, temporal location features, and pose similarity to obtain multi-dimensional association features; The multi-dimensional association features are used to construct a three-dimensional association matrix, and a lightweight Hungarian matching algorithm is used to associate the target with historical tracking IDs to obtain the association results. Based on the association results, determine whether the target is occluded; If the target is occluded, the position of the target is predicted by fitting the historical motion trajectory, and the tracking ID is re-matched based on the predicted position in the next frame image until a match is successful; Generate a sequence of target images with tracking IDs, along with results containing tracking stability data and multi-dimensional correlation features, to determine the location and movement trajectory information of each animal.
5. The animal health monitoring and early warning method based on intelligent agents according to claim 1, characterized in that, The behavior recognition model analyzes the animal's behavior patterns based on the position and movement trajectory information of each animal, identifies specific behaviors by extracting posture feature points, and records relevant data, including: The position and movement trajectory information of each animal are normalized in size, and the dynamic correlation and difference in several consecutive frames are calculated to construct a posture change vector describing the behavior pattern. The attitude change vector is transformed into semantically meaningful behavioral features to obtain the current attitude change vector; The current posture change vector is compared with a preset template library to preliminarily determine the behavior category and obtain the comparison result; Based on the comparison results, the individual's behavior labels within a specific time window are output, and behaviors that do not meet the duration requirements are removed to obtain the final behavior judgment result; Based on the final behavior judgment result, the start time, end time and duration of each behavior are recorded to obtain individual behavior data; All individual behavioral data are aggregated, and the distribution ratio of various behaviors and the frequency of abnormal behaviors are statistically analyzed on a periodic basis. Based on the statistical data, a structured report is generated to obtain relevant data.
6. The animal health monitoring and early warning method based on intelligent agents according to claim 1, characterized in that, Based on the relevant data and combined with professional knowledge from the veterinary knowledge base, a large-scale model inference engine is used to comprehensively assess the animal's health status. When abnormalities are detected, corresponding warning signals are issued according to the severity, including: The relevant data is then subjected to format conversion and structuring processing to obtain the processing result; The predefined data validation rules are applied to filter out data with missing key information or incorrect formatting in the processing results, and timestamps and source identifiers are added to the cleaned data to obtain the processed data; Using structured rule expressions in the knowledge base, and through precise matching and fuzzy matching techniques, an initial disease association analysis is performed on the processed data to determine the probability of meeting the requirements, a list of suspected associated diseases and their confidence scores, so as to obtain preliminary matching results. The preliminary matching results are combined with structured behavioral time-series data and auxiliary features as input. A fine-tuned lightweight large model is used for deep reasoning. By introducing a thinking chain mechanism, the degree of behavioral abnormality, the degree of disease symptom fit, and the weight of environmental influence factors are analyzed step by step to generate multi-dimensional evaluation results. The multi-dimensional evaluation results include health status assessment, warning level, ranking of associated diseases, and confidence level. Based on the multi-dimensional assessment results and the corresponding warning levels, when an anomaly is detected, a corresponding warning signal will be automatically issued according to the severity.
7. The animal health monitoring and early warning method based on intelligent agents according to claim 1, characterized in that, The step of selecting relevant video clips and keyframes based on the warning level corresponding to the warning signal, adding identifiers and explanatory text to obtain warning information includes: The individual ID, statistical period, and abnormal behavior are determined based on the warning signal. Select several images from key moments of abnormal behavior as keyframes, and extract video clips, adding time watermarks and individual IDs; Save keyframes and video clips and generate unique access links; Add behavioral feature identifiers to keyframes and write explanatory text for each warning level that includes an anomaly description, risky disease, and recommended measures; Based on the warning level, embed corresponding keyframe links or video clip links and explanatory text into the warning information; Early warning information is pushed out through multiple channels, and push logs are recorded to track feedback.
8. The animal health monitoring and early warning method based on intelligent agents according to claim 2, characterized in that, The process of automatically filtering useful information from newly collected data samples, using this information for automatic annotation and incremental training, and optimizing the object detection model and behavior recognition model includes: Data samples with unique identifiers and behavioral classifications were selected based on image quality, target integrity, and recognition confidence criteria to obtain a sample set; The sample set is pre-labeled using an object detection algorithm, including bounding boxes, key points, and behavior labels, and the overall confidence score is calculated. The number of frozen layers is dynamically adjusted based on the sample size, and dynamic L2 regularization is introduced. Design a reward function that combines tracking success rate, recognition accuracy, and feature consistency. Optimize the parameters of the object detection model and behavior recognition model through reinforcement learning to obtain updated object detection and behavior recognition models. A comprehensive evaluation of the updated object detection model and behavior recognition model is conducted to check whether the updated object detection model and behavior recognition model meet the predetermined performance indicators.
9. An animal health monitoring and early warning system based on intelligent agents, characterized in that, include: The video acquisition unit is used to acquire videos of animal groups recorded by a camera. The detection and tracking unit is used to perform individual identification and continuous tracking of each frame of the animal group video using a target detection model, and to generate the position and movement trajectory information of each animal; wherein, the target detection model uses an object detection and tracking algorithm combined with a feature reuse mechanism for individual identification and tracking; The behavior recognition and statistics unit is used to analyze the behavior patterns of each animal based on its position and movement trajectory information using a behavior recognition model, identify specific behaviors by extracting posture feature points, and record relevant data. The comprehensive assessment unit is used to comprehensively assess the health status of animals based on the relevant data and professional knowledge in the veterinary knowledge base, using a large model inference engine. When an abnormality is detected, it issues a corresponding warning signal according to the severity. The early warning information generation unit is used to select relevant video clips and keyframes according to the early warning level corresponding to the early warning signal, and add identifiers and explanatory text to obtain early warning information; A sending unit is used to send the warning information.
10. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 8.