Zero-perception dairy cow behavior observation system based on AI vision
The AI-based vision-based zero-perception dairy cow behavior observation system utilizes visual Transformer and deep forest classifier to achieve non-perception dairy cow behavior monitoring. This solves the shortcomings of traditional methods, such as sensor wearing and manual inspection, and provides high-precision, low-cost automated behavior recognition and health early warning. It adapts to complex environments and improves the efficiency of dairy cow breeding management.
Patent Information
- Application Number
- CN202510982319.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing dairy cow behavior monitoring technologies suffer from problems such as stress response caused by sensor wearing, easy equipment damage, limited data, and high manpower costs for manual video inspections, making it difficult to achieve large-scale, long-term continuous monitoring. Furthermore, existing visual methods have low recognition accuracy in environments with multiple cows, occlusion, significant changes in lighting, or complex behavioral patterns, and cannot meet the requirements for fully automated and interference-free behavior observation.
A zero-perception dairy cow behavior observation system based on AI vision is adopted. It utilizes the global spatial temporal feature modeling capability of the visual Transformer and the small sample adaptive advantage of the deep forest classifier. Video image data is collected through monitoring cameras, and after image preprocessing, features are extracted using the visual Transformer model. Combined with the deep forest classifier, multi-granularity window and cascade classification are performed to generate individual and group behavior information. Anomaly warning is achieved through model adaptive training.
It enables automatic long-term, full-scenario observation and health early warning of dairy cow behavior without the need for wearing sensors. It features high recognition accuracy, strong environmental adaptability, flexible system expansion and low operation and maintenance costs. It supports lightweight deployment and real-time inference of edge devices, reduces labor and equipment maintenance costs, and improves dairy cow welfare and breeding efficiency.
Smart Images

Figure CN120877374A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of animal behavior monitoring technology, and in particular to a zero-perception dairy cow behavior observation system based on AI vision. Background Technology
[0002] With the intelligent development of animal husbandry, dairy cow behavior monitoring technology plays a crucial role in ensuring dairy cow health, improving breeding efficiency, and achieving precise management. The industry generally adopts behavior observation methods based on sensor-equipped devices (such as collars and leg bands) and manual video inspections. Sensor solutions can achieve a rough identification of dairy cow behaviors such as movement, rumination, and feeding, but they are prone to triggering stress responses, affecting the natural behavioral expressions of dairy cows, and suffer from problems such as difficult maintenance, easy equipment damage, and limited data. Manual video inspection methods require a large investment of manpower, making it difficult to achieve large-scale, long-term continuous monitoring, and are prone to subjective errors and omissions, limiting both real-time performance and accuracy.
[0003] In recent years, computer vision and artificial intelligence technologies have been initially applied in the field of animal behavior recognition, such as behavior classification based on convolutional neural networks and motion analysis based on traditional object detection models. However, existing vision methods generally suffer from technical shortcomings such as strong dependence on local features, insufficient sensitivity to global behavioral changes, and limited temporal modeling capabilities. In actual farming environments with many cattle, occlusion, significant changes in lighting, or complex behavioral patterns, the recognition accuracy of existing models decreases, failing to meet the needs of large-scale, fully automated, and interference-free behavioral observation. Traditional deep learning models typically require a large amount of manually labeled data and high-performance computing hardware, making them difficult to adapt to the deployment and maintenance requirements of farms.
[0004] Therefore, how to provide a zero-perception dairy cow behavior observation system based on AI vision is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a zero-perception dairy cow behavior observation system based on AI vision. This invention fully utilizes the global spatial-temporal feature modeling capabilities of the visual Transformer and the small-sample adaptive advantages of the deep forest classifier. It describes the entire process from data acquisition, feature extraction, intelligent classification, anomaly detection, and model adaptive updating, achieving accurate identification and dynamic management of individual and group dairy cow behavior. Without requiring sensors or interfering with the natural behavior of the cows, it automatically completes long-term, full-scenario behavior observation and health warnings, possessing advantages such as high recognition accuracy, strong environmental adaptability, flexible system expansion, and low operation and maintenance costs.
[0006] The zero-perception dairy cow behavior observation system based on AI vision according to an embodiment of the present invention includes the following modules:
[0007] The data acquisition module is used to collect continuous video image data in the dairy farm through monitoring cameras to obtain raw image data;
[0008] The image preprocessing module is used to preprocess the raw image data and generate a standardized sequence of image frames;
[0009] The Visual Transformer feature extraction module is used to input standardized image frame sequences into the Visual Transformer model and output spatiotemporal feature vectors.
[0010] The Deep Forest Behavior Classification module receives spatiotemporal feature vectors and performs multi-granularity window and cascade classification, outputting behavior category labels.
[0011] The behavior analysis module is used to generate individual behavior trajectories and group behavior distributions, enabling visualization and statistical analysis.
[0012] The model adaptive training module is used to periodically collect new data and incrementally train the deep forest classifier.
[0013] The anomaly warning module is used to identify and record abnormal behavior and output anomaly warning information.
[0014] Optionally, modules can be integrated using the following methods:
[0015] S1. Install fixed monitoring cameras in the dairy farm and collect continuous video stream image data of dairy cows in the cowshed environment according to the preset acquisition frequency to obtain raw image data. Preprocess the raw image data to obtain a standardized image frame sequence.
[0016] S2. Input the standardized image frame sequence into the visual Transformer model, and use image block segmentation, linear embedding and position coding. Extract features from each frame image through a multi-layer Transformer encoder to obtain the spatial feature vector of each frame image. Then, obtain the spatiotemporal feature vector reflecting continuous behavior by fusing temporal information.
[0017] S3. Input the spatiotemporal feature vector into the deep forest classifier, and use the multi-granularity scanning structure and cascade structure to classify the spatiotemporal feature vector for behavior, and output the cow behavior category label corresponding to each frame.
[0018] S4. Based on the dairy cow behavior category tags, generate individual dairy cow behavior trajectory information and group behavior distribution information, visualize them, and perform real-time statistical analysis on the behavior data. If the behavior characteristics of an individual or group deviate from the historical threshold for a long period of time, output abnormal warning information.
[0019] S5. Deploy the visual Transformer model on edge computing devices and periodically collect new data. Combine cow behavior category labels and behavior trajectory information with group behavior distribution information to incrementally train the deep forest classifier and achieve adaptive behavior recognition in different scenarios and environments.
[0020] Optionally, the raw image data specifically includes a sequence of unprocessed color images from various monitoring cameras within the dairy farm, collected continuously.
[0021] Optionally, the preprocessing of the original image data specifically includes size normalization, illumination adjustment, and background noise removal.
[0022] Optionally, S2 specifically includes:
[0023] S21. For each frame of standardized image obtained, based on the motion trajectory and behavioral heat distribution of individual cows in the current frame and several historical frames, the spatiotemporal dynamic patch partitioning method is adopted to adaptively divide each frame of image into several image blocks of variable size and position. Among them, the active motion area and the behavioral change area are divided into high-density, small-size patches, and the static motion or background area is divided into low-density, large-size patches.
[0024] S22. Flatten each obtained Patch and convert it into a one-dimensional vector. Then, obtain the Patch embedding vector through linear transformation. All Patch embedding vectors are arranged in order to form the Patch feature sequence of the current frame, ensuring that the feature expression of each Patch has dynamic structural adaptability.
[0025] S23. Generate a behavior-aware position code for each Patch. The behavior-aware position code includes not only the two-dimensional spatial coordinate information of the Patch in the whole frame image, but also the average motion direction, average speed and behavior frequency of the Patch in the current frame and the previous several historical frames.
[0026] S24. Add the obtained Patch embedding vector to the generated behavior-aware location encoding vector element by element to obtain the behavior-space composite input feature of each Patch. Arrange the behavior-space composite input features of all Patches in sequence to form the Patch input feature sequence.
[0027] S25. Generate a category label vector based on the Patch input feature sequence. The category label vector is a learnable fixed-length vector, initially obtained by random initialization, and serves as an expression of global behavioral features. Insert the category label vector into the starting position of the Patch input feature sequence. The concatenated complete input sequence is then input into a multi-layer stacked Transformer encoder. Each layer of the Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization structure.
[0028] S26. In the multi-head self-attention mechanism of each Transformer encoder layer, a cross-frame adaptive attention mechanism is adopted. Specifically, in the self-attention weight calculation process, the correlation between all patches in the current frame is calculated, and the correlation between the current frame and related patches in adjacent historical frames is calculated simultaneously to realize the global interaction of behavior evolution information in the temporal dimension. After the forward calculation of all Transformer encoder layers, the output is a feature sequence containing all patches and category label vectors.
[0029] S27. Apply global average pooling to the output feature sequence, and calculate the average value of all feature vectors in the same frame in each corresponding dimension to obtain the spatial feature vector of each frame image.
[0030] S28. Input the spatial feature vectors of multiple consecutive frames of images into the temporal feature fusion function in chronological order. The temporal feature fusion function integrates the spatial feature vectors of multiple consecutive frames in chronological order and uses a recurrent neural network to model the temporal dependency between patches to generate a spatiotemporal feature vector that reflects the continuous dynamic behavior of the cow.
[0031] Optionally, S3 specifically includes:
[0032] S31. For the spatiotemporal feature vector of each obtained image frame, based on the variation amplitude and behavioral distribution characteristics of the spatiotemporal feature vector, an adaptive multi-granularity window mechanism is adopted to segment the spatiotemporal feature vector according to dynamically determined window lengths and sliding steps, generating multiple sets of granular feature subsets, denoted as {F1, F2, ..., F...}. M}, where M is the number of adaptive granular windows, F m This is the feature subset extracted at the m-th granularity.
[0033] S32. Extract the feature subsets F at each granularity respectively. m Given multiple completely random trees and multiple random forests, an adaptive tree number allocation and soft labeling mechanism is used for each category y. Specifically, the probability distribution on the behavior label space is applied to the behavior category to which the sample belongs. As a supervisory signal, the probability output of category y at granularity m is P. m (y|F m );
[0034] S33. Output the class probabilities {P1, P2, ..., P} under all granularity windows. M The features are concatenated to form a multi-granularity fused feature vector F. mg ;
[0035] S34. The obtained multi-granularity fusion feature vector F mg The input is fed into the first layer of a cascaded forest structure. Each layer of the forest consists of several completely random trees and random forests. The output is a class probability vector P. (l) (y|F mg ), where l is the number of layers in the cascaded structure;
[0036] S35. After training and inference at each layer of the cascaded structure, the maximum probability confidence of all behavior categories is calculated. If the maximum probability confidence of a certain behavior category is lower than the preset threshold θ, a new layer of cascaded forest is added to the behavior category, expanding the number of layers to l+1.
[0037] S36. Concatenate the class probability vector output by each layer with the original multi-granularity fusion feature vector as the input of the next layer. Repeat steps S34 and S35 until the maximum probability confidence of all classes is not lower than the threshold θ or the maximum number of layers is reached.
[0038] S37. When there are new behavioral samples or environmental changes, a multi-factor driven online incremental forest update mechanism is adopted, specifically including: for new sample X new and its tag y new Calculate the characteristic response distance d of the new sample for each tree in the existing forest. j If the characteristic response distance is d j If the dynamic threshold γ is exceeded, the j-th tree undergoes structural expansion or node re-splitting, and the tree weights are adjusted in real time based on the distribution of new sample labels and soft labels. A time decay factor α and an environmental adaptation factor β are introduced to jointly regulate the effective weights of each tree in the forest. The forest weight ω... t Used to indicate the effectiveness of the forest in classification decisions at the current moment;
[0039] S38. Finally, the behavior category with the highest probability in the category probability vector output by the last cascaded structure is taken as the cow behavior category label Y of the current frame image. * .
[0040] Optionally, S4 specifically includes:
[0041] S41. Associate the cow behavior category labels of each obtained image frame with the individual cow's identity and time sequence to construct a behavior label sequence for each cow, denoted as {y}. i,1 ,y i,2 ,...,y i,T}, where i is the cow number, T is the frame number within the observation period, and y i,t Label the behavior category of the i-th cow in frame t;
[0042] S42. Based on the behavior tag sequence, extract the behavior duration, behavior switching frequency, and behavior pattern features for each cow according to the temporal relationship. Specifically, this includes counting the length of the segment where the same behavior tag appears consecutively and the number of switches between different behavior tags to obtain the behavior duration sequence D. i,k and switching frequency S i D i,k S represents the duration of the k-th type of behavior of the i-th cow. i Total number of behavior switches;
[0043] S43. Group all individual cows' behavior tag sequences, duration sequences, and switching frequencies according to spatial location and number, calculate the individual distribution density of each behavior category within the same time period, and generate a group behavior distribution information matrix G. t,k G t,k Let the number of cows exhibiting the k-th type of behavior in frame t be ;
[0044] S44. Based on the behavior tag sequence, behavior duration, switching frequency and group behavior distribution information, use visualization methods to generate individual behavior trajectory maps, group behavior heat maps and statistical analysis curves, and output them to the behavior monitoring interface.
[0045] S45. Perform real-time statistical analysis on the behavioral tag sequence, duration, and switching frequency of each cow, and identify abnormal behavioral characteristics based on historical averages and set thresholds;
[0046] S46. When the anomaly detection formula satisfies E i,k When the value is 1, the system automatically outputs abnormal warning information, records the abnormal time, cow number, abnormal behavior type and deviation amount, and provides visual prompts on the monitoring interface or sends warning information through the data interface, thus completing the intelligent warning function for abnormal behavior.
[0047] Optionally, S5 specifically includes:
[0048] S51. Deploy the trained visual Transformer model on an edge computing device with inference capabilities to achieve local real-time feature extraction and spatiotemporal feature vector generation of the acquired standardized image frame sequence.
[0049] S52. After the visual Transformer model is deployed, new image data is collected periodically or at set time intervals, and the new data is input into the visual Transformer model to obtain the corresponding spatiotemporal feature vectors.
[0050] S53. Synchronously collect and organize the cow behavior category labels, behavior trajectory information and group behavior distribution information corresponding to the new data to form an incremental sample set;
[0051] S54. Input the incremental sample set into the deep forest classifier, perform incremental training on the deep forest classifier, and update the parameters of the deep forest classifier in real time.
[0052] S55. After incremental training is completed, the updated deep forest classifier is used to classify the collected image data into behaviors, and the cow behavior category labels and related behavior analysis results are output.
[0053] S56. Continuously loop through data acquisition, incremental model training, and behavior classification to achieve adaptive recognition of cow behavior and dynamic optimization of model performance in different scenarios and environments.
[0054] The beneficial effects of this invention are:
[0055] This invention deeply integrates a visual Transformer model with a deep forest classifier to construct a zero-perception behavior observation method and system for dairy farming scenarios, improving the accuracy and efficiency of dairy cow behavior recognition. It overcomes the technical bottlenecks of traditional methods relying on wearable sensors or manual inspections, achieving continuous and automated individual and group behavior recognition and health risk warning for large-scale dairy cow herds under completely interference-free conditions. Through multi-granularity adaptive windows, a class confidence-driven cascaded layer growth strategy, a soft-label mechanism, and an online incremental forest update mechanism, it solves the technical challenges of existing visual algorithms, such as insufficient sensitivity to behavioral changes in complex scenarios, weak small-sample learning ability, and poor model generalization adaptability.
[0056] Using the method of this invention, the system automatically adapts to different lighting, shading, and layout environments in a dairy farm, without relying on large-scale labeled data and high-performance hardware. It supports lightweight deployment and real-time inference of edge devices, and continuously optimizes the model through periodic collection of new data, ensuring stable and reliable long-term monitoring results. This significantly reduces labor and equipment maintenance costs, provides farm managers with multi-dimensional, visualized behavioral analysis and anomaly warnings, improves dairy cow welfare and farming efficiency, and has broad application value and good economic benefits. Attached Figure Description
[0057] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0058] Figure 1 This is a schematic diagram of the structure of the zero-perception dairy cow behavior observation system based on AI vision proposed in this invention;
[0059] Figure 2 This is a flowchart illustrating the zero-perception dairy cow behavior observation method based on AI vision proposed in this invention. Detailed Implementation
[0060] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0061] refer to Figure 1 The AI-based vision-based zero-perception dairy cow behavior observation system includes the following modules:
[0062] The data acquisition module is used to collect continuous video image data in the dairy farm through monitoring cameras to obtain raw image data;
[0063] The image preprocessing module is used to preprocess the raw image data and generate a standardized sequence of image frames;
[0064] The Visual Transformer feature extraction module is used to input standardized image frame sequences into the Visual Transformer model and output spatiotemporal feature vectors.
[0065] The Deep Forest Behavior Classification module receives spatiotemporal feature vectors and performs multi-granularity window and cascade classification, outputting behavior category labels.
[0066] The behavior analysis module is used to generate individual behavior trajectories and group behavior distributions, enabling visualization and statistical analysis.
[0067] The model adaptive training module is used to periodically collect new data and incrementally train the deep forest classifier.
[0068] The anomaly warning module is used to identify and record abnormal behavior and output anomaly warning information.
[0069] refer to Figure 2 The zero-perception dairy cow behavior observation method based on AI vision includes the following steps:
[0070] S1. Install fixed monitoring cameras in the dairy farm and collect continuous video stream image data of dairy cows in the cowshed environment according to the preset acquisition frequency to obtain raw image data. Preprocess the raw image data to obtain a standardized image frame sequence.
[0071] S2. Input the standardized image frame sequence into the visual Transformer model, and use image block segmentation, linear embedding and position coding. Extract features from each frame image through a multi-layer Transformer encoder to obtain the spatial feature vector of each frame image. Then, obtain the spatiotemporal feature vector reflecting continuous behavior by fusing temporal information.
[0072] S3. Input the spatiotemporal feature vector into the deep forest classifier, and use the multi-granularity scanning structure and cascade structure to classify the spatiotemporal feature vector for behavior, and output the cow behavior category label corresponding to each frame.
[0073] S4. Based on the dairy cow behavior category tags, generate individual dairy cow behavior trajectory information and group behavior distribution information, visualize them, and perform real-time statistical analysis on the behavior data. If the behavior characteristics of an individual or group deviate from the historical threshold for a long period of time, output abnormal warning information.
[0074] S5. Deploy the visual Transformer model on edge computing devices and periodically collect new data. Combine cow behavior category labels and behavior trajectory information with group behavior distribution information to incrementally train the deep forest classifier and achieve adaptive behavior recognition in different scenarios and environments.
[0075] In this embodiment, the original image data specifically includes a sequence of unprocessed color images output from various monitoring cameras within the dairy farm.
[0076] In this embodiment, the preprocessing of the original image data specifically includes size normalization, lighting condition adjustment, and background noise removal.
[0077] In this embodiment, S2 specifically includes:
[0078] S21. For each frame of standardized image obtained, based on the motion trajectory and behavioral heat distribution of individual cows in the current frame and several historical frames, the spatiotemporal dynamic patch partitioning method is adopted to adaptively divide each frame of image into several image blocks of variable size and position. Among them, the active motion area and the behavioral change area are divided into high-density, small-size patches, and the static motion or background area is divided into low-density, large-size patches.
[0079] S22. Flatten each obtained Patch and convert it into a one-dimensional vector. Then, obtain the Patch embedding vector through linear transformation. All Patch embedding vectors are arranged in order to form the Patch feature sequence of the current frame, ensuring that the feature expression of each Patch has dynamic structural adaptability.
[0080] S23. Generate a behavior-aware position code for each Patch. The behavior-aware position code includes not only the two-dimensional spatial coordinate information of the Patch in the whole frame image, but also the average motion direction, average speed and behavior frequency of the Patch in the current frame and the previous several historical frames.
[0081] S24. Add the obtained Patch embedding vector to the generated behavior-aware location encoding vector element by element to obtain the behavior-space composite input feature of each Patch. Arrange the behavior-space composite input features of all Patches in sequence to form the Patch input feature sequence.
[0082] S25. Generate a category label vector based on the Patch input feature sequence. The category label vector is a learnable fixed-length vector, initially obtained by random initialization, and serves as an expression of global behavioral features. Insert the category label vector into the starting position of the Patch input feature sequence. The concatenated complete input sequence is then input into a multi-layer stacked Transformer encoder. Each layer of the Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization structure.
[0083] S26. In the multi-head self-attention mechanism of each Transformer encoder layer, a cross-frame adaptive attention mechanism is adopted. Specifically, in the self-attention weight calculation process, the correlation between all patches in the current frame is calculated, and the correlation between the current frame and related patches in adjacent historical frames is calculated simultaneously to realize the global interaction of behavior evolution information in the temporal dimension. After the forward calculation of all Transformer encoder layers, the output is a feature sequence containing all patches and category label vectors.
[0084] S27. Apply global average pooling to the output feature sequence, and calculate the average value of all feature vectors in the same frame in each corresponding dimension to obtain the spatial feature vector of each frame image.
[0085] S28. Input the spatial feature vectors of multiple consecutive frames of images into the temporal feature fusion function in chronological order. The temporal feature fusion function integrates the spatial feature vectors of multiple consecutive frames in chronological order and uses a recurrent neural network to model the temporal dependency between patches to generate a spatiotemporal feature vector that reflects the continuous dynamic behavior of the cow.
[0086] In this embodiment, S3 specifically includes:
[0087] S31. For the spatiotemporal feature vector of each obtained image frame, based on the variation amplitude and behavioral distribution characteristics of the spatiotemporal feature vector, an adaptive multi-granularity window mechanism is adopted to segment the spatiotemporal feature vector according to dynamically determined window lengths and sliding steps, generating multiple sets of granular feature subsets, denoted as {F1, F2, ..., F...}. M}, where M is the number of adaptive granular windows, F m This is the feature subset extracted at the m-th granularity.
[0088] S32. Extract the feature subsets F at each granularity respectively. m Given multiple completely random trees and multiple random forests, an adaptive tree number allocation and soft labeling mechanism is used for each category y. Specifically, the probability distribution on the behavior label space is applied to the behavior category to which the sample belongs. As a supervisory signal, the probability output of category y at granularity m is P. m (y|F m ):
[0089]
[0090] Where, N m,y The number of trees assigned to category y at granularity m, where K is the total number of behavior categories. Let P be the soft-label distribution of the samples in class k. j,y (k|F m Let be the predicted probability of the j-th tree for category k;
[0091] The probability output is P m (y|F m This study enhances the intelligent adaptability and precise discrimination capabilities of deep forest classifiers in dairy cow behavior recognition. For different categories of behavioral samples, it dynamically allocates varying numbers of decision tree resources based on the actual classification difficulty and distribution. The model effectively focuses on rare or easily confused categories, reducing recognition errors caused by class imbalance. By introducing soft-label probability distributions, it fully considers the transitional and ambiguous nature of samples between various behavioral states, improving the model's ability to recognize samples with ambiguous behavioral boundaries and composite behaviors. It achieves the fusion and discrimination of multi-granularity and multi-class information, enhancing sensitivity to abnormal and rare behaviors in complex scenarios. This provides a solid mathematical foundation for high accuracy and robustness in behavior classification, greatly expanding the system's application capabilities in large-scale, dynamically changing environments, and promoting the automation, intelligence, and scientification of dairy cow behavior monitoring.
[0092] S33. Output the class probabilities {P1, P2, ..., P} under all granularity windows.M The features are concatenated to form a multi-granularity fused feature vector F. mg ;
[0093] S34. The obtained multi-granularity fusion feature vector F mg The input is fed into the first layer of a cascaded forest structure. Each layer of the forest consists of several completely random trees and random forests. The output is a class probability vector P. (l) (y|F mg ), where l is the number of layers in the cascaded structure;
[0094] S35. After training and inference at each layer of the cascaded structure, the maximum probability confidence of all behavior categories is calculated. If the maximum probability confidence of a certain behavior category is lower than the preset threshold θ, a new layer of cascaded forest is added to the behavior category, expanding the number of layers to l+1.
[0095] S36. Concatenate the class probability vector output by each layer with the original multi-granularity fusion feature vector as the input of the next layer. Repeat steps S34 and S35 until the maximum probability confidence of all classes is not lower than the threshold θ or the maximum number of layers is reached.
[0096] S37. When there are new behavioral samples or environmental changes, a multi-factor driven online incremental forest update mechanism is adopted, specifically including: for new sample X new and its tag y new Calculate the characteristic response distance d of the new sample for each tree in the existing forest. j If the characteristic response distance is d j If the dynamic threshold γ is exceeded, the j-th tree undergoes structural expansion or node re-splitting, and the tree weights are adjusted in real time based on the distribution of new sample labels and soft labels. A time decay factor α and an environmental adaptation factor β are introduced to jointly regulate the effective weights of each tree in the forest. The forest weight ω... t Used to represent the effectiveness of the forest in classification decisions at the current moment:
[0097] ω t =α·ω t-1 +β·(1-α)·δ t ;
[0098] Where, ω t-1 The forest weights at the previous time step, δ t The new weights generated by the current incremental learning are α, β, and γ, where α is the time decay factor, β is the environmental adaptation factor, and γ is the structural change threshold.
[0099] The multi-factor driven online incremental forest regeneration mechanism proposed in this invention enhances the long-term adaptability and decision-making effectiveness of the classification model in the dynamic environment of actual pastures by integrating two adjustment factors: time decay and environmental adaptation. Based on the response distance between new behavioral samples and the decision trees in the existing forest, the structure of the trees is flexibly adjusted to achieve rapid adaptation to new behaviors or environmental changes. The effective weights of each tree in classification decisions are dynamically adjusted to ensure that the model always focuses on the latest and most relevant behavioral feature data. The time decay factor ensures a smooth transition of the model to historical data, preventing overfitting to old patterns. The environmental adaptation factor improves the model's sensitivity to external changes and sudden events, enhancing its self-learning and self-evolution capabilities. Under conditions of continuous monitoring and continuous input of new data, it maintains a high level of behavioral recognition accuracy and robustness, providing a solid technical guarantee for smart farming and dairy cow health management.
[0100] S38. Finally, the behavior category with the highest probability in the category probability vector output by the last cascaded structure is taken as the cow behavior category label Y of the current frame image. * .
[0101] In this embodiment, S4 specifically includes:
[0102] S41. Associate the cow behavior category labels of each obtained image frame with the individual cow's identity and time sequence to construct a behavior label sequence for each cow, denoted as {y}. i,1 ,y i,2 ,...,y i,T}, where i is the cow number, T is the frame number within the observation period, and y i,t Label the behavior category of the i-th cow in frame t;
[0103] S42. Based on the behavior tag sequence, extract the behavior duration, behavior switching frequency, and behavior pattern features for each cow according to the temporal relationship. Specifically, this includes counting the length of the segment where the same behavior tag appears consecutively and the number of switches between different behavior tags to obtain the behavior duration sequence D. i,k and switching frequency S i D i,k S represents the duration of the k-th type of behavior of the i-th cow. i Total number of behavior switches;
[0104] S43. Group all individual cows' behavior tag sequences, duration sequences, and switching frequencies according to spatial location and number, calculate the individual distribution density of each behavior category within the same time period, and generate a group behavior distribution information matrix G. t,k G t,k Let the number of cows exhibiting the k-th type of behavior in frame t be ;
[0105] S44. Based on the behavior tag sequence, behavior duration, switching frequency and group behavior distribution information, use visualization methods to generate individual behavior trajectory maps, group behavior heat maps and statistical analysis curves, and output them to the behavior monitoring interface.
[0106] S45. Perform real-time statistical analysis on the behavioral tag sequence, duration, and switching frequency of each cow, and identify abnormal behavioral characteristics based on historical averages and set thresholds:
[0107]
[0108] Among them, E i,k As an anomaly marker, Let θ be the historical mean duration of the k-th type of behavior. D The threshold for the duration of the behavior. θ represents the historical average switching frequency. S For switching frequency thresholds;
[0109] This invention achieves automatic identification of abnormal behavior in dairy cows by using dual threshold discrimination based on behavior duration and switching frequency. It utilizes historical behavior averages as a benchmark, comparing the duration and switching frequency of an individual cow's behavior with its historical normal levels in real time. Significant deviations from preset thresholds are automatically flagged as abnormal, avoiding the problems of strong subjectivity, low efficiency, and missed or false judgments inherent in traditional manual interpretation. This improves the system's sensitivity and accuracy in detecting abnormal behavior characteristics, adapts to individual differences in different dairy cows and behavior types, and achieves individualized and dynamic intelligent early warning. It enhances the farm's ability to automatically monitor key events such as health risks, stress responses, and group anomalies, providing technical support for managers to take timely intervention measures, ensure cow welfare, and improve farming efficiency.
[0110] S46. When the anomaly detection formula satisfies E i,k When the value is 1, the system automatically outputs abnormal warning information, records the abnormal time, cow number, abnormal behavior type and deviation amount, and provides visual prompts on the monitoring interface or sends warning information through the data interface, thus completing the intelligent warning function for abnormal behavior.
[0111] In this embodiment, S5 specifically includes:
[0112] S51. Deploy the trained visual Transformer model on an edge computing device with inference capabilities to achieve local real-time feature extraction and spatiotemporal feature vector generation of the acquired standardized image frame sequence.
[0113] S52. After the visual Transformer model is deployed, new image data is collected periodically or at set time intervals, and the new data is input into the visual Transformer model to obtain the corresponding spatiotemporal feature vectors.
[0114] S53. Synchronously collect and organize the cow behavior category labels, behavior trajectory information and group behavior distribution information corresponding to the new data to form an incremental sample set;
[0115] S54. Input the incremental sample set into the deep forest classifier, perform incremental training on the deep forest classifier, and update the parameters of the deep forest classifier in real time.
[0116] S55. After incremental training is completed, the updated deep forest classifier is used to classify the collected image data into behaviors, and the cow behavior category labels and related behavior analysis results are output.
[0117] S56. Continuously loop through data acquisition, incremental model training, and behavior classification to achieve adaptive recognition of cow behavior and dynamic optimization of model performance in different scenarios and environments.
[0118] Example 1:
[0119] To verify the feasibility of this invention in practice, it was applied to a modern dairy farm with approximately 900 head of cattle annually. The farm's management team had been facing the practical problems of difficulty in identifying abnormal behaviors in dairy cows, high labor intensity of manual inspections, and low efficiency in collecting behavioral data. In key behavioral monitoring scenarios such as estrus and early stages of disease, traditional methods relying on manual experience and sensor collars were prone to misjudgments or missed detections, affecting scientific feeding and health management.
[0120] A total of 120 dairy cows were selected from four main barns to construct and deploy the intelligent observation system for dairy cow behavior based on the fusion of visual Transformer and deep forest proposed in this invention. Five high-definition cameras were installed in each barn, and the video signals were directly transmitted to a small edge computing server within the barn. The system automatically captured one high-definition image frame every two seconds. After data acquisition, size normalization, illumination correction, and noise removal were automatically performed, and all standardized image frames entered the AI visual analysis process.
[0121] The visual Transformer model automatically extracts key behavioral features from each frame and multiple consecutive frames of dairy cow images through dynamic patch partitioning and temporal feature fusion. The deep forest classifier uses multi-granularity windows and a cascading mechanism to intelligently distinguish between various behaviors such as feeding, drinking, resting, rumination, active movement, and abnormal movement. The system supports continuous accumulation of behavioral data and incremental model training, enabling adaptive optimization of the model under changes in barn environment, lighting, and season.
[0122] During three consecutive months of operation, the system automatically processed over 3 million images and recorded the behavior categories, durations, and switching frequencies of each cow. Comparison with data from manual inspections during the same period showed improved system accuracy and faster alerts and responses to abnormal events. During the peak estrus season in spring, the system automatically detected estrus behaviors in cows 25 times, all before manual inspections, with an average lead time of 6.1 hours; it also detected 20 instances of abnormal lying down, successfully assisting staff in timely intervention and reducing health risks. The mean square error for the average duration of feeding behavior was only 3.4 minutes², and the error for behavior switching frequency was less than 0.2 times / day, superior to manual observation.
[0123] Table 1 Comparison of recognition performance between the system and manual inspection under different behavior types
[0124]
[0125] As shown in Table 1, the system of this invention demonstrates performance advantages in various behavioral recognition tasks, including dairy cow feeding, drinking, rumination, lying down, active movement, and abnormal behavior. The system's automatic recognition accuracy is higher than manual inspection, reaching a maximum of 98.0% (lying down) and a minimum of 92.7% (abnormal behavior), while the accuracy of manual recording is generally lower, with a minimum of only 81.2%. The system's false negative rate is significantly lower than manual detection; in complex or abnormal behavior recognition, the system's false negative rate is only 2.6%, better than the 12.0% of manual inspection. Regarding response speed, the system relies on real-time edge computing, significantly reducing the average response time. The response time for most behaviors is no more than 0.8 hours, while manual inspection generally requires more than 2 hours, with even longer delays for some abnormal behaviors. This indicates that the system of this invention improves the scientific rigor, real-time performance, and accuracy of behavior recognition in large-scale farming scenarios, reduces the burden on management personnel, and provides technical support for dairy cow health protection and improved farming efficiency.
[0126] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A zero-perception dairy cow behavior observation system based on AI vision, characterized in that, Includes the following modules: The data acquisition module is used to collect continuous video image data in the dairy farm through monitoring cameras to obtain raw image data; The image preprocessing module is used to preprocess the raw image data and generate a standardized sequence of image frames; The Visual Transformer feature extraction module is used to input standardized image frame sequences into the Visual Transformer model and output spatiotemporal feature vectors. The Deep Forest Behavior Classification module receives spatiotemporal feature vectors and performs multi-granularity window and cascade classification, outputting behavior category labels. The behavior analysis module is used to generate individual behavior trajectories and group behavior distributions, enabling visualization and statistical analysis. The model adaptive training module is used to periodically collect new data and incrementally train the deep forest classifier. The anomaly warning module is used to identify and record abnormal behavior and output anomaly warning information.
2. A zero-perception dairy cow behavior observation method based on AI vision, applied to the zero-perception dairy cow behavior observation system based on AI vision as described in claim 1, characterized in that, Includes the following steps: S1. Install fixed monitoring cameras in the dairy farm and collect continuous video stream image data of dairy cows in the cowshed environment according to the preset acquisition frequency to obtain raw image data. Preprocess the raw image data to obtain a standardized image frame sequence. S2. Input the standardized image frame sequence into the visual Transformer model, and use image block segmentation, linear embedding and position coding. Extract features from each frame image through a multi-layer Transformer encoder to obtain the spatial feature vector of each frame image. Then, obtain the spatiotemporal feature vector reflecting continuous behavior by fusing temporal information. S3. Input the spatiotemporal feature vector into the deep forest classifier, and use the multi-granularity scanning structure and cascade structure to classify the spatiotemporal feature vector for behavior, and output the cow behavior category label corresponding to each frame. S4. Based on the dairy cow behavior category tags, generate individual dairy cow behavior trajectory information and group behavior distribution information, visualize them, and perform real-time statistical analysis on the behavior data. If the behavior characteristics of an individual or group deviate from the historical threshold for a long period of time, output abnormal warning information. S5. Deploy the visual Transformer model on edge computing devices and periodically collect new data. Combine cow behavior category labels and behavior trajectory information with group behavior distribution information to incrementally train the deep forest classifier and achieve adaptive behavior recognition in different scenarios and environments.
3. The zero-perception dairy cow behavior observation method based on AI vision according to claim 1, characterized in that, The raw image data specifically includes a sequence of unprocessed color images from multiple monitoring cameras continuously collected within the dairy farm.
4. The zero-perception dairy cow behavior observation method based on AI vision according to claim 1, characterized in that, The preprocessing of the original image data specifically includes size normalization, lighting condition adjustment, and background noise removal.
5. The zero-perception dairy cow behavior observation method based on AI vision according to claim 2, characterized in that, S2 specifically includes: S21. For each frame of standardized image obtained, based on the motion trajectory and behavioral heat distribution of individual cows in the current frame and several historical frames, the spatiotemporal dynamic patch partitioning method is adopted to adaptively divide each frame of image into several image blocks of variable size and position. Among them, the active motion area and the behavioral change area are divided into high-density, small-size patches, and the static motion or background area is divided into low-density, large-size patches. S22. Flatten each obtained Patch and convert it into a one-dimensional vector. Then, obtain the Patch embedding vector through linear transformation. All Patch embedding vectors are arranged in order to form the Patch feature sequence of the current frame, ensuring that the feature expression of each Patch has dynamic structural adaptability. S23. Generate a behavior-aware position code for each Patch. The behavior-aware position code includes not only the two-dimensional spatial coordinate information of the Patch in the whole frame image, but also the average motion direction, average speed and behavior frequency of the Patch in the current frame and the previous several historical frames. S24. Add the obtained Patch embedding vector to the generated behavior-aware location encoding vector element by element to obtain the behavior-space composite input feature of each Patch. Arrange the behavior-space composite input features of all Patches in sequence to form the Patch input feature sequence. S25. Generate a category label vector based on the Patch input feature sequence. The category label vector is a learnable fixed-length vector, initially obtained by random initialization, and serves as an expression of global behavioral features. Insert the category label vector into the starting position of the Patch input feature sequence. The concatenated complete input sequence is then input into a multi-layer stacked Transformer encoder. Each layer of the Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network, residual connections, and a layer normalization structure. S26. In the multi-head self-attention mechanism of each Transformer encoder layer, a cross-frame adaptive attention mechanism is adopted. Specifically, in the self-attention weight calculation process, the correlation between all patches in the current frame is calculated, and the correlation between the current frame and related patches in adjacent historical frames is calculated simultaneously to realize the global interaction of behavior evolution information in the temporal dimension. After the forward calculation of all Transformer encoder layers, the output is a feature sequence containing all patches and category label vectors. S27. Apply global average pooling to the output feature sequence, and calculate the average value of all feature vectors in the same frame in each corresponding dimension to obtain the spatial feature vector of each frame image. S28. Input the spatial feature vectors of multiple consecutive frames of images into the temporal feature fusion function in chronological order. The temporal feature fusion function integrates the spatial feature vectors of multiple consecutive frames in chronological order and uses a recurrent neural network to model the temporal dependency between patches to generate a spatiotemporal feature vector that reflects the continuous dynamic behavior of the cow.
6. The zero-perception dairy cow behavior observation method based on AI vision according to claim 2, characterized in that, S3 specifically includes: S31. For the spatiotemporal feature vector of each obtained image frame, based on the variation amplitude and behavioral distribution characteristics of the spatiotemporal feature vector, an adaptive multi-granularity window mechanism is adopted to segment the spatiotemporal feature vector according to dynamically determined window lengths and sliding steps, generating multiple sets of granular feature subsets, denoted as {F1, F2, ..., F...}. M }, where M is the number of adaptive granular windows, F m This is the feature subset extracted at the m-th granularity. S32. Extract the feature subsets F at each granularity respectively. m Given multiple completely random trees and multiple random forests, an adaptive tree number allocation and soft labeling mechanism is used for each category y. Specifically, the probability distribution on the behavior label space is applied to the behavior category to which the sample belongs. As a supervisory signal, the probability output of category y at granularity m is P. m (y|F m ); S33. Output the class probabilities {P1, P2, ..., P} under all granularity windows. M The features are concatenated to form a multi-granularity fused feature vector F. mg ; S34. The obtained multi-granularity fusion feature vector F mg The input is fed into the first layer of a cascaded forest structure. Each layer of the forest consists of several completely random trees and random forests. The output is a class probability vector P. (l) (y|F mg ), where l is the number of layers in the cascaded structure; S35. After training and inference at each layer of the cascaded structure, the maximum probability confidence of all behavior categories is calculated. If the maximum probability confidence of a certain behavior category is lower than the preset threshold θ, a new layer of cascaded forest is added to the behavior category, expanding the number of layers to l+1. S36. Concatenate the class probability vector output by each layer with the original multi-granularity fusion feature vector as the input of the next layer. Repeat steps S34 and S35 until the maximum probability confidence of all classes is not lower than the threshold θ or the maximum number of layers is reached. S37. When there are new behavioral samples or environmental changes, a multi-factor driven online incremental forest update mechanism is adopted, specifically including: for new sample X new and its tag y new Calculate the characteristic response distance d of the new sample for each tree in the existing forest. j If the characteristic response distance is d j If the dynamic threshold γ is exceeded, the j-th tree undergoes structural expansion or node re-splitting, and the tree weights are adjusted in real time based on the distribution of new sample labels and soft labels. A time decay factor α and an environmental adaptation factor β are introduced to jointly regulate the effective weights of each tree in the forest. The forest weight ω... t Used to indicate the effectiveness of the forest in classification decisions at the current moment; S38. Finally, the behavior category with the highest probability in the category probability vector output by the last cascaded structure is taken as the cow behavior category label Y of the current frame image. * .
7. The zero-perception dairy cow behavior observation method based on AI vision according to claim 2, characterized in that, S4 specifically includes: S41. Associate the cow behavior category labels of each obtained image frame with the individual cow's identity and time sequence to construct a behavior label sequence for each cow, denoted as {y}. i,1 ,y i,2 ,...,y i,T }, where i is the cow number, T is the frame number within the observation period, and y i,t Label the behavior category of the i-th cow in frame t; S42. Based on the behavior tag sequence, extract the behavior duration, behavior switching frequency, and behavior pattern features for each cow according to the temporal relationship. Specifically, this includes counting the length of the segment where the same behavior tag appears consecutively and the number of switches between different behavior tags to obtain the behavior duration sequence D. i,k and switching frequency S i D i,k S represents the duration of the k-th type of behavior of the i-th cow. i Total number of behavior switches; S43. Group all individual cows' behavior tag sequences, duration sequences, and switching frequencies according to spatial location and number, calculate the individual distribution density of each behavior category within the same time period, and generate a group behavior distribution information matrix G. t,k G t,k Let the number of cows exhibiting the k-th type of behavior in frame t be denoted by . S44. Based on the behavior tag sequence, behavior duration, switching frequency and group behavior distribution information, use visualization methods to generate individual behavior trajectory maps, group behavior heat maps and statistical analysis curves, and output them to the behavior monitoring interface. S45. Perform real-time statistical analysis on the behavioral tag sequence, duration, and switching frequency of each cow, and identify abnormal behavioral characteristics based on historical averages and set thresholds; S46. When the anomaly detection formula satisfies E i,k When the value is 1, the system automatically outputs abnormal warning information, records the abnormal time, cow number, abnormal behavior type and deviation amount, and provides visual prompts on the monitoring interface or sends warning information through the data interface, thus completing the intelligent warning function for abnormal behavior.
8. The zero-perception dairy cow behavior observation method based on AI vision according to claim 2, characterized in that, S5 specifically includes: S51. Deploy the trained visual Transformer model on an edge computing device with inference capabilities to achieve local real-time feature extraction and spatiotemporal feature vector generation of the acquired standardized image frame sequence. S52. After the visual Transformer model is deployed, new image data is collected periodically or at set time intervals, and the new data is input into the visual Transformer model to obtain the corresponding spatiotemporal feature vectors. S53. Synchronously collect and organize the cow behavior category labels, behavior trajectory information and group behavior distribution information corresponding to the new data to form an incremental sample set; S54. Input the incremental sample set into the deep forest classifier, perform incremental training on the deep forest classifier, and update the parameters of the deep forest classifier in real time. S55. After incremental training is completed, the updated deep forest classifier is used to classify the collected image data into behaviors, and the cow behavior category labels and related behavior analysis results are output. S56. Continuously loop through data acquisition, incremental model training, and behavior classification to achieve adaptive recognition of cow behavior and dynamic optimization of model performance in different scenarios and environments.
Citation Information
Cited By
Live pig breeding whole-process intelligent management system
CN121744070A