Cattle behavior recognition method based on measurement and time-space relation reasoning
By explicitly modeling the spatiotemporal relationship between cattle behavior and the environment using the STPN-STRFN-STEN network, the problems of low accuracy and poor robustness in cattle behavior recognition in existing methods are solved, achieving higher recognition accuracy and model stability.
Patent Information
- Application Number
- CN202511058797.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for identifying cattle behavior neglect explicit modeling of the cattle's behavior and its surrounding environment, resulting in low accuracy and poor robustness in complex environments, especially in terms of insufficient feature extraction and inadequate model generalization ability when dealing with complex behavioral features.
A cattle behavior recognition method based on metric and spatiotemporal relationship reasoning is adopted. The method is constructed by STPN-STRFN-STEN network and uses the relationship network to explicitly model the cattle subject and the environment. By combining the contrastive loss function and the triple loss function, the spatiotemporal relationship between cattle behavior and the environment is explicitly modeled, thereby improving the robustness and accuracy of the model.
Explicitly modeling the spatiotemporal relationship between cattle behavior and the environment improves the accuracy and interpretability of cattle behavior recognition models, effectively captures subtle differences between different behavioral states, and enhances the robustness and recognition accuracy of the models.
Smart Images

Figure CN120954090A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cattle behavior recognition technology, and in particular to a cattle behavior recognition method based on metric and spatiotemporal relationship reasoning. Background Technology
[0002] The behavioral patterns of cattle have attracted widespread attention. In-depth research on the behavior of cattle under different environmental and management conditions can promote the development of farm management towards a more intelligent and sustainable direction.
[0003] Cattle detection refers to the extraction of key biological characteristics in complex farm environments; the core is to distinguish cattle from other objects and handle interference such as occlusion and changes in lighting.
[0004] Cow behavior recognition is diverse and complex, with subtle differences between behaviors, especially in crowded or dynamically changing farm environments. Based on an improved YOLO v3 model, by optimizing anchor points, introducing a DenseBlock structure, and improving the bounding box loss function, the detection performance of estrus behavior in dairy cows has been significantly enhanced. However, lameness behavior in dairy cows involves complex background environments, leading to low accuracy and poor robustness in traditional RGB image-based lameness detection methods. Existing methods still face challenges in handling complex behavioral features, including insufficient feature extraction and inadequate model generalization ability. Although advanced model architectures can effectively model the spatiotemporal relationship of cow behavior, the inherent "black box" nature of deep learning means that its modeling process is often implicit. This implicit modeling approach may lead to insufficient generalization ability in extreme scenarios, thus affecting the reliability of the recognition results.
[0005] In the field of cattle behavior recognition, existing methods often neglect explicit modeling of the cattle behavior subject and its surrounding environment. Summary of the Invention
[0006] To address the shortcomings of existing methods, this invention solves the problem that existing methods neglect explicit modeling of the cattle's behavior and its surrounding environment.
[0007] The technical solution adopted in this invention is: a method for recognizing cattle behavior based on metric and spatiotemporal relationship reasoning, comprising the following steps: Step 1: Obtain a dataset of cattle behavior; In a preferred embodiment of the present invention, the construction of the dataset includes: performing video segmentation to locate the cows in each frame of the image; correcting the detected bounding boxes; and adding behavioral labels to each cow using a video annotation tool.
[0008] In a preferred embodiment of the present invention, cattle behavior includes: feeding, walking, running, standing rumination, lying down rumination, standing rest, lying down rest, drinking water, grooming, other, and hiding behaviors.
[0009] Step 2: Obtain the high and low video frame sequences; In a preferred embodiment of the present invention, the low frame rate is 8 and the high frame rate is 16.
[0010] Step 3: Input the low-frequency video frame sequence into the Slow branch of the STPN network, and input the high-frequency video frame sequence into the Fast branch of the STPN network; output the Slow feature map and the Fast feature map. Step 4: Input the Slow branch feature map and Fast branch feature map into the STRFN network; As a preferred embodiment of the present invention, the STRFN network includes: First, feature map K of the Slow branch and feature map J of the Fast branch are used to obtain feature map KT(N,8,C3,W2,H2) and feature map JT(N,16,C4,W3,H3) using ROI Align. Secondly, KT and JT are flattened into feature maps KT1(N,8,C3) for N targets to be identified. W2 H2) and feature map JT1(N,16,C4) W3 H3); Secondly, in the upper branch, KT1 is transformed using the first MLP, and a contrastive loss function is used. L constractive Make the feature vectors extracted by the first MLP have similar behaviors close to each other; Next, KT1 and JT1 are copied into feature maps KT2(N,8,C3) with the same shape as KT and JT, respectively. W2 H2, W1, H1) and feature map JT2(N, 16, C4) W3 H3, W1, H1); Secondly, KT2, KJ and JT2, KJ are merged into feature map KT3(N,8,C3) along the channel dimension. W2 H2+C1+2C2,W1,H1) and feature map JT3(N,16,C4) W3 H3+C2+(C1) / 2,W1,H1); Finally, KT3 and JT3 each use 1 1. Convolutional fusion yields feature map KT4(N,8,C5,W1,H1) and feature map JT4(N,16,C6,W1,H1).
[0011] Step 5: Input the feature map output by the STRFN network into the STEN network to complete the construction of the STPN-STRFN-STEN network.
[0012] In a preferred embodiment of the present invention, the second MLP of the STEN network adopts a triplet loss function.
[0013] As a preferred embodiment of the present invention, the STPN-STRFN-STEN network employs a total loss function. ;in, α and β For hyperparameters; Indicates batch size; L cls For binary cross-entropy loss; F a For anchor point features, F p For positive samples and F n This is a negative sample.
[0014] In a preferred embodiment of the present invention, the STPN-STRFN-STEN network is evaluated using AP and mAP evaluation metrics.
[0015] As a preferred embodiment of the present invention, a cattle behavior recognition system based on metric and spatiotemporal relation reasoning includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement the cattle behavior recognition method based on metric and spatiotemporal relation reasoning.
[0016] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code implements a method for recognizing cattle behavior based on metric and spatiotemporal relationship reasoning when executed by a processor.
[0017] The beneficial effects of this invention are: 1. This invention utilizes a relational network to explicitly model the relationship between the global feature map extracted from the backbone network and the cow subject extracted using ROI; in the branch processing spatial features, contrastive loss is used to project the extracted cow subject features in the spatial dimension into a nonlinear space to initially distinguish the state of the cow subject in various behaviors; by fusing the relationship features established between the cow subject features in the temporal and spatial dimensions and the global features of the environment through the relational network, the robustness of the cow behavior recognition model is further improved. 2. This invention is used to automatically mine the relationship network between cattle behavior subjects and the environment from cattle video data; it explicitly models the relationship between cattle behavior and the surrounding environment, improving the accuracy and interpretability of the cattle behavior recognition model; 3. In the spatial dimension, the main features of cattle (such as posture) can directly reflect their local behavioral information; by adopting a metric learning method based on contrastive loss, the extracted spatial features are mapped to a nonlinear feature space; it can effectively capture the subtle differences between different behavioral states, and distinguish similar but different behavioral patterns through distance measurement in the feature space, thereby improving the accuracy of behavior recognition. 4. By using a relational network, deep integration of the main features of cattle and the global features of the environment in the temporal and spatial dimensions can be achieved; by establishing the correlation between features, the interaction information between cattle behavior and the environment can be effectively captured, thereby significantly improving the robustness of the cattle behavior recognition model. Attached Figure Description
[0018] Figure 1 This is a flowchart of the cattle behavior recognition method based on measurement and spatiotemporal relationship reasoning of the present invention; Figure 2 These are example images of cattle behavior videos from this invention and the number of annotations for their corresponding behaviors; Figure 3 This is a diagram illustrating the reasoning behind the relationship between cattle and the environment in this invention; Figure 4 The present invention utilizes a contrastive loss function to measure the spatial features of the bovine subject. Figure 5 This is a visualization comparison of the feature dimensionality reduction results before and after using contrastive learning according to the present invention; Figure 6 This is a visualization result of a video sample of cows drinking water using GradCAM, as described in this invention. Figure 7 This is a visual sample of cattle performing grooming behavior according to the present invention. Detailed Implementation
[0019] The present invention will be further described below with reference to the accompanying drawings and embodiments. The drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0020] like Figure 1 As shown, the cattle behavior recognition method based on metric and spatiotemporal relation reasoning includes the following steps: Step 1: Obtain a dataset of cattle behavior; The dataset is a cattle behavior dataset collected, annotated, and published by [CVB: A Video Dataset of Cattle Visual Behaviors]; experimental data were collected under natural lighting conditions at a resolution of 1920×1080 pixels and a frame rate of 30 frames per second (FPS); a complete video surveillance dataset was constructed by deploying high-resolution cameras at the four corners of an experimental site containing eight Angus cattle; The process of building a dataset includes the following steps: First, the acquired video was segmented into segments with fixed time spans. Then, a pre-trained object detection and tracking model was applied to locate individual cattle in each frame. After domain experts corrected the detection bounding boxes, a video annotation tool (CVAT) was used to add behavioral labels to each cow in each frame. The final dataset contained 502 15-second video segments, which were labeled with 11 visually identifiable cattle behaviors in a pasture environment. Cattle behaviors include: grazing, walking, running, ruminating, ruminating, resting, drinking, grooming, other (behaviors not belonging to any of the above categories), and hiding. The dataset is labeled with the following format: video ID (video_id), time (second), bounding box coordinates (x1, y1, x2, y2), i.e., top left and bottom right coordinates, cattle ID (cattle_id), and behavior category (behavior).
[0021] Given the scarcity of samples and inconsistent data quality for some behavioral categories, this study processed the original dataset as follows to meet research needs: "ruminating-standing" and "resting-standing" were merged into the "standing" category, and "ruminating-lying" and "resting-lying" were merged into the "lying" category. Meanwhile, the "others" and "hidden" categories were removed. It is important to note that among the merged behavioral categories, only "standing" and "lying" are mutually exclusive; all other behavioral categories remain mutually exclusive. The remaining behavioral categories can be used for pairwise comparisons, effectively evaluating the model's spatiotemporal feature extraction capabilities.
[0022] In the data preprocessing stage, the video data was segmented in seconds, resulting in 7,530 valid video segments and 30,282 corresponding action bounding boxes. To ensure the reliability of the model evaluation, the dataset was divided into three parts: training set, validation set, and test set, with proportions of 70%, 10%, and 20%, respectively.
[0023] Considering that the behavior recognition task is based on bounding boxes provided by the dataset, the model needs to rely on these bounding boxes to extract the spatiotemporal features of the cattle. For each cow, there are 30 frames of bounding box data per second, where the spatiotemporal information of any one frame can characterize the spatiotemporal changes of that cow within that second. Therefore, using the bounding boxes from the first frame of each second as the model's input data improves computational efficiency while ensuring feature representativeness. The examples of each behavior class and the number of their corresponding bounding boxes are as follows: Figure 2 As shown.
[0024] Step 2: Obtain the high and low video frame sequences; Steps two, three, and five refer to the patented method for pig behavior recognition based on attention mechanism spatiotemporal perception and enhancement network; low frame rate = 8, high frame rate = 16. Step 3: Construct a spatial temporal perception network (STPN). Input the low video frame sequence into the Slow branch and the high video frame sequence into the Fast branch. Introduce an FL-SAM module after each residual block in the Slow branch. Introduce a KMFEM module after each residual block in the Fast branch to output Slow feature maps and Fast feature maps. The spatiotemporal awareness network includes: a spatial attention residual network and an inter-frame difference residual network; the spatial attention residual network consists of five residual blocks, with FL-SAM modules inserted between adjacent residual blocks; In the inter-frame differential residual network, it consists of five residual blocks, with KMFEM modules inserted between adjacent residual blocks; This invention improves the pig behavior recognition model framework based on a spatiotemporal perception and enhancement network with an attention mechanism, where the gray part represents the original model structure. The original model, based on a spatiotemporal perception and enhancement network with an attention mechanism, achieves spatiotemporal feature modeling of the cow's behavioral subject and its related interaction areas in video data. Building on this, this invention further introduces a relational network and metric learning method, and designs a new spatiotemporal relationship feature fusion module. This improvement effectively enhances the model's performance in recognizing cow behavior, especially in feature extraction and classification accuracy in complex scenes. The pig behavior recognition model based on attention-based spatiotemporal perception and augmentation networks has several limitations. First, the spatiotemporal features extracted using attention mechanisms and frame difference methods only achieve implicit modeling between the actor and the environment. In extreme cases, key features may exceed the bounding box after scaling. Second, in complex scenes (such as crowded environments), interference between multiple moving actors may lead to inaccurate feature extraction. These problems make it difficult for the model to capture the essential behavioral features in cattle video data, ultimately affecting the accuracy of behavior recognition. Therefore, establishing an explicit modeling mechanism between the actor and its surrounding environment has become a critical issue that urgently needs to be addressed. Step 4: Input the Slow branch feature map and Fast branch feature map into the spatial-temporal relationship fusion network (STRFN); The feature map K of the Slow branch and the feature map J of the Fast branch; wherein the size of feature map K is (B,8,C1,W1,H1) and the size of feature map J is (B,16,C2,W1,H1). The two branches are spliced and fused together in the low frame rate branch; the feature maps obtained from the two branches are cropped through ROI to obtain the spatial features and motion changes of relevant positions of each cow during its movement. First, feature map K and feature map J are used to obtain feature map KT and feature map JT, with KT having a size of (N,8,C3,W2,H2) and JT having a size of (N,16,C4,W3,H3). Secondly, KT and JT are flattened into feature maps KT1 and JT1 for N target recognition, where KT1 has a size of (N, 8, C3). W2 H2), the size of JT1 is (N, 16, C4) W3 H3); Secondly, in the upper branch of the spatiotemporal relationship fusion network, the first MLP is used to transform KT1, and a contrastive loss function is used. L constractive This makes the feature vectors extracted by the first MLP have similar behaviors close to each other and different categories separate, thus initially increasing the discriminative power of the identified targets; Next, KT1 and JT1 are copied into feature maps KT2 and JT2 with the same shape as KT and JT, respectively, where the size of KT2 is (N, 8, C3). W2 H2, W1, H1), the size of JT2 is (N, 16, C4) W3 H3, W1, H1); Secondly, in terms of channel dimension, KT2, KJ and JT2, KJ are merged into feature maps KT3 and JT3 respectively, where the size of KT3 is (N, 8, C3). W2 The values of H2+C1+2C2,W1,H1) and JT3 are (N,16,C4) W3 H3+C2+(C1) / 2,W1,H1); Finally, KT3 and JT3 each use 1 1. Convolutional fusion is used to identify the relationship between the target and global features, resulting in relational feature maps KT4(N,8,C5,W1,H1) and JT4(N,16,C6,W1,H1); Input the feature maps KT4 and JT4 into the spatiotemporal augmentation network.
[0025] like Figure 3 The diagram illustrates the relationship between cattle and their environment, where B represents the batch, N represents the number of cattle in each batch, C represents the number of channels, W represents the width of the feature map, and H represents the height of the feature map.
[0026] Relational reasoning plays a crucial role in behavior recognition tasks. This invention proposes a Relation Network (RN) for explicit modeling of a cow and its surrounding environment. The internal algorithm is as follows: Figure 3 As shown; by using relational networks, we can explicitly and automatically mine the connections between each cow and its environment in video clips during model training.
[0027] Extracting bovine behavioral features with good representational capabilities is crucial for establishing inferences about inter-behavioral relationships. Before applying a relational network, this invention introduces a contrastive loss function in the low-frame-rate branch for extracting spatial features. L constractive By projecting the feature vectors of each cow obtained through ROI cropping onto a high-dimensional space, a highly discriminative representation vector is obtained. This method ensures that the main features of cows exhibiting the same behavior are highly similar in the feature space, while maintaining significant differences between different behavioral categories. The specific implementation method is as follows: Figure 4 As shown; to verify the effectiveness of metric learning, the t-SNE algorithm is used to reduce the dimensionality of the metric vectors extracted by the trained model and visualize them; as shown. Figure 5 As shown, a contrastive loss function is introduced. L constractive After adopting the metric learning method, the spatial features of cattle subjects showed better inter-class discriminative power and intra-class clustering, effectively improving the representational power of the features.
[0028] Given that the spatiotemporal features of cattle movement significantly influence action classification, this invention employs a spatiotemporal relationship fusion network to explicitly model the correlation between the cattle's main features and environmental features in both spatial and temporal dimensions. In the low frame rate branch, a 1×1 convolutional layer fuses the spatial features of the cattle's main body after high-dimensional spatial measurement with the global spatiotemporal feature map extracted by the spatiotemporal perception network. This method effectively captures the contextual relationship between the target and global spatiotemporal features, significantly enhancing the ability to discriminate target features. In the high frame rate branch, a 1×1 convolutional layer is used to model the dynamic changes of the cattle's main body in the temporal dimension and its other correlations with the environment in the spatial dimension. This design not only enhances the temporal consistency of target features but also improves the stability of the detection process. By applying 1×1 convolutional operations in both the temporal and spatial dimensions, the model can simultaneously capture the dynamic change characteristics and static contextual information of the target, comprehensively improving the robustness and recognition accuracy of the model.
[0029] Step 5: Construct the spatial temporal enhancement network (STEN). In the first branch of the STEN, the CL-SAM module performs residual connections on the feature map KT4 and then inflates it. In the second branch, the feature map JT4 is Maxpooled and then fed into the LSTM network. The output feature maps from both branches are inflated and concatenated and then fed into the second MLP network for classification. It also includes: introducing triplet loss in the second MLP; Triplet Loss is used to optimize the feature space distribution and improve the stability of model training; it makes the feature representations of positive samples of the same class (Fp) closer to each other, while the feature representations of negative samples of different classes (Fn) are further apart.
[0030] Specifically, since Triplet Loss requires anchor vectors as a reference, anchor feature vectors are defined. F a Considering that the mean of each category of behavioral features has good stability, using the category mean as an anchor can effectively reduce the impact of sample fluctuations on the training process, while enhancing the compactness of intra-class features and the separation of inter-class features, thereby improving the classification performance of the model. Based on this F a Defined as the average value of the behavioral features of each category after the spatiotemporal augmentation network fuses the two branches; Formula (2) gives the anchor point features. F a Positive samples F p and negative samples Fn The constructed triplet loss function improves the model's discriminative ability by optimizing the feature space distribution; The formula for the triplet loss function is: (2) in, d 2() sets the metric function to the cosine function; margin This is a hyperparameter used to control the minimum difference between positive and negative samples, initially set to 0.1, which increases with training. epoch Exponential growth; margin The formula is: (3) Finally, considering that Triplet Loss primarily focuses on learning discriminative features and cannot directly optimize classification tasks, a binary cross-entropy loss (BCE Loss, Classifier) is introduced to ensure that the model can accurately predict the sample class, denoted as . L cls .
[0031] Simultaneously, a contrastive loss function, denoted as [function name missing], is used to extract high-discrimination spatial features of cattle subjects. L constractive Together they constitute the entire model (i.e. Figure 1 The loss function of the STPN-STRFN-STEN network composed of A, B, and C is given by the following formula: (4) in, α and β The hyperparameters for model training are all set to 0.5; This represents the batch size, the set of all triples extracted from the training data of B.
[0032] This joint loss function design ensures both the discriminative power of features and the accuracy of classification, thereby achieving a comprehensive improvement in model performance.
[0033] It also includes: evaluating the model of this invention using AP and mAP evaluation metrics; AP and mAP, based on the Precision-Recall curve, can simultaneously reflect the model's precision and recall, avoiding the limitations of a single metric. Furthermore, in imbalanced datasets, mAP can average the performance across multiple classes, making it suitable for multi-class classification tasks. Higher AP and mAP indicate better model performance in classification tasks and more accurate prediction of sample classes.
[0034] Experimental procedure: The experiment was conducted on an Ubuntu 18.04.6 system running Python 3.8, implemented using PyTorch 2.2.2 + CUDA 12.1. The server configuration included a CPU with 48 cores (Xeon E5-2678 v3) and four NVIDIA TeslaV100-PCIE-32 GB GPUs. Stochastic gradient descent (SGD) was used for backpropagation with a momentum of 0.9 and a batch size of 16. The AdamW optimizer was employed with an initial learning rate of 0.001 and weight decay of 0.02. Cosine annealing learning rate decay was applied, where the learning rate annealed from its initial value to its minimum value according to a cosine function during each epoch from epoch 0 to epoch 200, and then restarted in the next epoch.
[0035] This invention's improved model was compared with STPEN (Spatial-Temporal Perception and Enhancement Network) and several existing action recognition models (including TSN, TSM, I3D, and SqueezeNet). STPEN employs a dual-branch network architecture, performing spatial feature extraction and motion modeling separately, and finally fusing the features from the spatial and temporal branches for action classification. TSN divides the video into fixed-length segments, extracts features from each segment and performs temporal modeling, and finally fuses the features from all segments to complete the classification task. TSM captures temporal information changes by shifting input features along the temporal dimension. The I3D (Inflated 3D ConvNet) model uses an inflated weight method to extend the pre-trained 2D convolutional network weights to a 3D convolutional network, thereby more effectively processing spatiotemporal information in the video. SqueezeNet significantly reduces the number of network parameters while maintaining high classification performance through its innovative Fire module and 1x1 convolutional design.
[0036] As shown in Table 1, compared with the I3D, TSM, TSN, and STPEN models, the mAP of our proposed model is 6.79 percentage points higher than the I3D model, 11.08 percentage points higher than the TSM model, 6.96 percentage points higher than the TSN model, 4.92 percentage points higher than the SqueezeNet model, and 5.19 percentage points higher than the STPEN model.
[0037] Table 1. Comparison of mAP and AP between the model of this invention and existing models.
[0038] This invention uses a relational network to explicitly model the spatiotemporal relationships of cattle behavior in videos. Through feature measurement based on a contrastive loss function, it effectively distinguishes the features of each behavior in the spatial dimension of the cattle subject. Compared with traditional action recognition models, this model exhibits stronger behavior representation capabilities and achieves higher recognition accuracy. Notably, this improvement balances other behaviors well, with particularly significant enhancements in the classification of grooming and running behaviors. Experimental results show that, compared with methods that only extract traditional spatiotemporal features, combining a relational network to explicitly establish connections between extracted features achieves higher accuracy. The visualization results of video samples of cattle drinking water using GradCAN are shown below. Figure 6 As shown.
[0039] This invention conducted ablation experiments on the low frame rate branch using feature measure (MF) based on contrastive loss function and spatial relation network (SRN), and the high frame rate branch using temporal relation network (TRN). The experiments used a controlled variable method, keeping other parameters such as learning rate, optimizer, and training epochs consistent except for the relation network and feature measure in the spatial and temporal dimensions. The results are shown in Table 2.
[0040] Table 2 Comparison of mAP and corresponding behavioral AP of the models obtained from ablation experiments
[0041] Experimental results show that SRN significantly outperforms TRN in behaviors closely related to spatial features of cattle and their environment, such as Drinking, Grooming, Lying, and Staming. However, it lags behind in Running, a behavior more closely related to temporal features. The results indicate that when relational networks help the network extract features in both temporal and spatial dimensions, SRN focuses on extracting spatial relational features of cattle behavior, while TRN focuses on extracting temporal relational features. Furthermore, adding MF (Multi-dimensional Feature Fusion) to initially distinguish the spatial features of the cattle behavior subject before using SRN for explicit relational modeling of cattle behavior can comprehensively improve model performance. Finally, the complete model, by combining MF, SRN, and TRN, achieves comprehensive modeling of spatiotemporal features and subject relationships, thus achieving optimal performance in various behavior recognition tasks. This result verifies the importance of multi-dimensional feature fusion and explicit relational modeling in behavior recognition tasks.
[0042] This invention aims to accurately model the spatiotemporal interaction between cattle and their surrounding environment. It fully utilizes video data of cattle behavior to analyze the spatiotemporal interaction characteristics between the cattle and their environment, achieving accurate classification of various behaviors. In the algorithm implementation, the relationship network employs 1×1 convolution operations in both spatial and temporal dimensions, effectively fusing the association information between the cattle's behavior and the global spatiotemporal feature map. To further improve the accuracy of the relationship network modeling, this invention innovatively introduces a feature measurement method based on contrastive loss. This method performs clustering optimization of different categories of cattle behavior features in a high-dimensional space, significantly improving the model's classification performance without increasing the number of model parameters. Experimental results show that this algorithm can not only accurately capture the spatiotemporal features of cattle behavior but also effectively distinguish similar behavior categories, providing a reliable solution for cattle behavior recognition in complex scenarios.
[0043] Traditional CNN models, such as VGG16, ResNet, and VisionTransformer (ViT), have been widely used in animal behavior recognition to extract animal behavioral features. However, traditional CNN models have limitations when processing video data: First, they typically employ local spatiotemporal convolution operations, which can only implicitly extract the relationship between the behavioral subject and other spatial features in the video; second, because they mainly focus on local spatial features, they are difficult to effectively model spatiotemporal dependencies over long periods. 3D CNNs, by performing convolution operations simultaneously in the spatiotemporal dimensions, can more effectively capture spatiotemporal features in videos; however, this method is computationally expensive, often requiring a large amount of computing resources, which may lead to excessive memory consumption during training and significantly increase time and computational overhead. On the other hand, Vision... Transformer (ViT) captures long-range dependencies globally through its self-attention mechanism, and compared to traditional CNNs, it can better model global contextual information between video frames. However, ViT has limitations: first, it requires large-scale training data to fully realize its performance advantages; second, it is prone to overfitting on small datasets; and finally, due to its reliance on the self-attention mechanism, the model has high complexity, resulting in long training and inference times, which poses a challenge to real-time behavior recognition tasks and requires further model optimization to reduce computational latency.
[0044] Although the contrastive loss-based measurement method can effectively distinguish the spatiotemporal pose of cattle and perform well in modeling the relationship between the actor and the environment, measuring in high-dimensional space may make it difficult for the model to accurately capture the interaction relationship of the actor in complex environment. Figure 7 Two visualizations showcasing the model's recognition of grooming behavior are presented: Figure 7In (a), the model successfully focused attention on the abdominal region, which is closely associated with grooming behavior; however, in Figure 7 In (b), when the cattle were facing away from the camera and performing grooming, the model incorrectly focused its attention on the limbs. This phenomenon indicates that the model still has significant shortcomings in modeling the interaction between the cattle and their own bodies, and there is an urgent need to develop new methods that can explicitly model this spatiotemporal relationship. In addition, the experimental results show that compared with traditional models, this method has achieved a significant improvement in the recognition accuracy of grooming and running behaviors. However, compared with other behavior categories, the recognition accuracy of these two behaviors is still far behind, which may be related to the insufficient number of training samples. Therefore, developing augmentation methods suitable for small sample data to further improve the recognition rate has become the focus of current research.
[0045] This invention considers applying an attention mechanism to the main features of cattle extracted from the ROI, and establishing a correlation model between cattle and themselves by dividing different body regions. This may provide an effective solution to the current problem. For the small sample size problem, transfer learning and meta-learning methods have shown great potential and can be a key direction for future research. In subsequent work, we will conduct in-depth research on the above issues in order to further improve the accuracy and robustness of cattle behavior recognition.
[0046] This invention proposes a feature learning algorithm based on metric learning and spatiotemporal relationship reasoning, aiming to accurately model the spatiotemporal interaction between cattle and their surrounding environment, thereby improving the accuracy of cattle behavior recognition. The algorithm consists of three core modules: a spatial relationship network, a temporal relationship network, and a feature measurement module based on contrastive loss. The spatial and temporal relationship networks explicitly model the spatiotemporal feature association between the cattle's behavior and its surrounding environment, enhancing the model's interpretability and significantly improving recognition accuracy. Furthermore, the feature measurement module based on contrastive loss further improves the accuracy of the relationship network modeling by clustering and optimizing the features of the cattle's main body. Experimental results show that the proposed algorithm achieves a recognition accuracy of 87.19% on a cattle video dataset containing seven behaviors, representing improvements of 5.19% and 4.92% compared to the basic model and the best traditional model, SqueezeNet, respectively. These results confirm that the algorithm can effectively capture the spatiotemporal interaction features between cattle and their environment in cattle videos, achieving high-precision behavior recognition. This research provides technical support for promoting the intelligent and sustainable development of farm management and also provides a scientific basis for animal health management.
[0047] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning, characterized in that, Includes the following steps: Step 1: Obtain a dataset of cattle behavior; Step 2: Obtain the high and low video frame sequences; Step 3: Input the low-frequency video frame sequence into the Slow branch of the STPN network and the high-frequency video frame sequence into the Fast branch of the STPN network; output the Slow feature map and the Fast feature map. Step 4: Input the Slow branch feature map and Fast branch feature map into the STRFN network; Step 5: Input the feature map output by the STRFN network into the STEN network to complete the construction of the STPN-STRFN-STEN network.
2. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, The STRFN network includes: First, feature map K of the Slow branch and feature map J of the Fast branch are used to obtain feature map KT(N,8,C3,W2,H2) and feature map JT(N,16,C4,W3,H3) using ROI Align. Secondly, KT and JT are flattened into feature maps KT1(N,8,C3) for N targets to be identified. W2 H2) and feature map JT1(N,16,C4) W3 H3); Secondly, in the upper branch of the STRFN network, the first MLP is used to transform KT1, and the contrastive loss function is used. L constractive Make the feature vectors extracted by the first MLP have similar behaviors close to each other; Next, KT1 and JT1 are copied into feature maps KT2(N,8,C3) with the same shape as KT and JT, respectively. W2 H2, W1, H1) and feature map JT2(N, 16, C4) W3 H3, W1, H1); Secondly, KT2, KJ and JT2, KJ are merged into feature map KT3(N,8,C3) along the channel dimension. W2 H2+C1+2C2,W1,H1) and feature map JT3(N,16,C4) W3 H3+C2+(C1) / 2,W1,H1); Finally, KT3 and JT3 each use 1 1. Convolutional fusion yields feature map KT4(N,8,C5,W1,H1) and feature map JT4(N,16,C6,W1,H1).
3. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, The second MLP of the STEN network uses a triplet loss function.
4. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, The STPN-STRFN-STEN network uses a total loss function. ;in, α and β For hyperparameters; Indicates batch size; L cls For binary cross-entropy loss; F a For anchor point features, F p For positive samples and F n This is a negative sample.
5. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, The STPN-STRFN-STEN network was evaluated using AP and mAP metrics.
6. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, The construction of the dataset includes: performing video segmentation to locate the cows in each frame of the image; correcting the detected bounding boxes; and adding behavioral labels to each cow using video annotation tools.
7. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, Cattle behavior includes: feeding, walking, running, standing rumination, lying down rumination, standing rest, lying down rest, drinking, grooming, other, and hiding behaviors.
8. The method for recognizing cattle behavior based on metric and spatiotemporal relation reasoning according to claim 1, characterized in that, Low frame rate = 8, high frame rate = 16.
9. A cattle behavior recognition system based on metric and spatiotemporal relation reasoning, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing instructions to implement the cattle behavior recognition method based on metric and spatiotemporal relation reasoning as described in any one of claims 1-8.
10. A computer-readable medium storing computer program code, characterized in that, The computer program code, when executed by a processor, implements the cattle behavior recognition method based on metric and spatiotemporal relation reasoning as described in any one of claims 1-8.