Chicken behavior recognition and positioning method based on spatio-temporal feature learning

Through the improved ActionFormer model and spatiotemporal feature learning method, the shortcomings of chicken behavior recognition methods in the prior art in the temporal resolution are solved, and accurate identification of the behaviors occurring in chicken videos and precise positioning of the time intervals are achieved.

CN120220185APending Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510295940.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing video-based chicken behavior recognition methods have insufficient time resolution, and cannot effectively deal with different behavioral events that occur in video clips when chickens are active.

Method used

A method of chicken behavior recognition and positioning based on spatiotemporal feature learning is proposed. Through the improved ActionFormer model, the time-sensitive pre-training module TSP, the time attention module TAM, the improved transformer block and three sets of cascaded feature pyramids are integrated, and combined with the weighted regression loss function, the chicken behavior recognition and positioning model CBLFormer is formed.

Benefits of technology

It effectively improves the time resolution of the video-based behavior recognition method, can identify the behaviors that occur one after another in the video and locate the corresponding time intervals of the behaviors, and is suitable for continuous monitoring of multiple behavior events of chickens in multi-objective scenarios in the chicken space-time behavior detection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220185A_ABST
    Figure CN120220185A_ABST
Patent Text Reader

Abstract

The invention discloses a chicken behavior recognition and positioning method based on spatio-temporal feature learning, and aims to recognize successive behaviors of chickens in a video and position a time interval corresponding to each behavior. The backbone of the CBLFormer completes hierarchical construction of feature maps based on 1D convolution and improved transformer block, a time attention module TAM is integrated to capture a fine-grained time sequence dependency relationship, the check part carries out feature fusion and interaction based on three groups of cascaded feature pyramids, and the head part adopts a transformer structure to carry out fine-grained modeling on the feature maps with different time resolutions from the check. The method provided by the invention can effectively improve the time resolution of a video behavior recognition model, and can also be integrated into a chicken space-time behavior detection system for continuous monitoring of multiple behavior events of chickens in a multi-target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chicken behavior recognition and localization methods, and specifically to a chicken behavior recognition and localization method based on spatio-temporal feature learning for recognizing the behaviors that occur successively in chickens in a video and localizing the time intervals in which each behavior occurs. Background Art

[0002] Timely acquisition of poultry behavior information can provide support for management and decision-making in the poultry breeding process, improve poultry welfare and prevent disease transmission. Currently, it mainly relies on farm workers to observe poultry behavior information and judge the health status of poultry based on this. This method has high labor intensity, strong subjectivity and low efficiency, and it is difficult to meet the development needs of modern poultry farming. With the aggravation of the aging population trend and the increase in labor costs, there is an urgent need for an automated and intelligent method to replace manual inspection.

[0003] Computer vision technology is considered an effective means for detecting livestock and poultry behaviors due to its advantages of high efficiency, low cost, and non-stress. Currently, the methods for realizing livestock and poultry behavior recognition based on computer vision technology are mainly divided into two categories: image-based methods and video-based methods. For image-based behavior recognition methods, detectors of the YOLO series and the R-CNN series are usually used for supervised training on image datasets to enable the detectors to have the ability to recognize different behaviors of livestock and poultry, and then the performance of the detectors is evaluated on new datasets. Such methods can achieve the recognition of most behaviors of livestock, such as basic behaviors of animals such as pigs, cows, and sheep, such as eating, drinking, standing, lying down, estrus, etc., and specific behaviors such as rumination of cows and sheep. This method is also applicable to the recognition of basic behaviors of poultry, including resting, eating, drinking, feather pecking, standing, fighting, exploring, mating, etc. Image-based behavior recognition methods usually comprehensively apply technologies such as data augmentation, attention mechanism, loss function improvement, and feature fusion strategy optimization to improve the recognition rate and stability of the model in complex breeding scenarios. However, the limitation of such methods is that the expression of livestock and poultry behaviors is a dynamic process. Defining livestock and poultry behaviors based on single-frame images is likely to lead to ambiguous and ambiguous behavior semantics due to the high similarity of the spatial features of certain behaviors at specific moments, reducing the reliability of behavior recognition results.

[0004] In contrast, video-based livestock and poultry behavior recognition methods can effectively combine the temporal and spatial characteristics of behaviors, thereby solving the possible semantic ambiguity problems between different behaviors and improving the accuracy of behavior recognition models. According to the ways in which the model processes the spatial and temporal characteristics in the video, such methods can be subdivided into two types. The first method is to first use a 2D feature extractor to extract the spatial features of each frame in the video clip to form a feature sequence that can represent the livestock and poultry behaviors, and then input the feature sequence into a long short-term memory neural network for temporal modeling. The second method is to use a neural network model to simultaneously model the spatio-temporal characteristics of the video clip, and then realize the recognition of livestock and poultry behaviors in the classification layer of the model. It should be noted that both types of methods require that only one behavior event is included in the video clip, which can usually be achieved by using a tracker to clip out the video clip of a single behavior event from a video with multiple concurrent behavior events. That is to say, video-based livestock and poultry behavior recognition methods usually rely on video clips with a fixed time window length. However, when livestock and poultry (especially poultry) are in a relatively active state, behavior transitions may occur within this time window, and the existing models can only predict a single behavior within the same time window and cannot effectively capture such behavior transitions, that is, the video-based behavior recognition model has deficiencies in temporal resolution. The deficiency in temporal resolution further restricts the application of the behavior recognition model integrated into a spatio-temporal behavior detection system for continuous behavior monitoring.

[0005] Obtaining chicken behavior information is of great significance for evaluating the health status and production performance of chickens. However, the existing video-based behavior recognition methods have insufficient temporal resolution, manifested as predicting a single behavior event for a single video clip and being unable to effectively handle the situation where different behavior events occur successively in the video clip caused by the active state of chickens. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, the present invention proposes a chicken behavior recognition and localization method based on spatio-temporal feature learning for recognizing the behaviors that occur successively in the video of chickens and localizing the time intervals corresponding to the behaviors.

[0007] A chicken behavior recognition and localization method based on spatio-temporal feature learning includes the following steps:

[0008] 1) Obtain chicken flock video data and make a pre-training data set and a chicken behavior recognition and localization data set;

[0009] 2) Perform data augmentation on the chicken behavior recognition and localization data set obtained in step 1) to obtain an enhanced behavior recognition and localization data set;

[0010] 3) Based on the ActionFormer model, add the time-sensitive pre-training module TSP to the data input end of the ActionFormer model, integrate the time attention module TAM in the backbone of the ActionFormer model, and improve the transformer block in the backbone to an improved transformer block. Replace the feature pyramid structure of the neck of the ActionFormer model with three cascaded feature pyramids, improve the 1D convolutional-based head of the ActionFormer model to a linear transformer-based head, and improve the regression loss function of the ActionFormer model to a weighted regression loss function to form an improved model;

[0011] 4) Use the pre-training dataset described in step 1) to fine-tune the time-sensitive pre-training module TSP in the improved model obtained in step 3) to obtain a time-sensitive pre-training module TSP with inference ability, and form a chicken behavior recognition and localization model CBLFormer;

[0012] 5) Use the enhanced behavior recognition and localization dataset described in step 2) to train and optimize the chicken behavior recognition and localization model CBLFormer in step 4). The time-sensitive pre-training module TSP with inference ability in step 4) is used to extract the video of the enhanced behavior recognition and localization dataset into a feature sequence to obtain a chicken behavior recognition and localization model CBLFormer with inference ability;

[0013] 6) Use the chicken behavior recognition and localization model CBLFormer with inference ability to recognize chicken videos, and obtain the chicken behavior categories that occur successively in the video and the time intervals when each chicken behavior occurs.

[0014] In step 1), the method for making the pre-training dataset includes: intercepting video segments of a single chicken from the obtained chicken flock videos. The duration of the video segments is 1s to 100s. Each video segment only contains the same behavior repeatedly occurring for the same chicken, and assign a class label to each video segment to constitute the pre-training dataset.

[0015] In step 1), the method for making the chicken behavior recognition and localization dataset includes: intercepting video segments of a single chicken from the obtained chicken flock videos, limiting the duration of the video segments to 1s to 100s. Each video segment contains two different behaviors that occur successively for the same chicken, and the expression time of each behavior must exceed 0.5s. Mark the information of each video segment, including the video file name, video duration, resolution, time interval, and class label, to constitute the chicken behavior recognition and localization dataset.

[0016] In step 2), the data augmentation method for the chicken behavior recognition and localization dataset includes: offsetting the endpoints of the time interval corresponding to the chicken behavior by 1%-5% of the entire interval length, and restricting that the right endpoint of the time interval corresponding to the current behavior in each video and the left endpoint of the time interval corresponding to the next behavior cannot overlap, so as to prevent confusion between behaviors.

[0017] In step 3), the principle of the temporal attention module TAM is as follows: The feature sequence X with size output from the previous feature extraction block is subjected to max pooling and average pooling along the time step T direction, where D is the dimension of the feature sequence, to obtain a feature vector P with size (1, T) max and P avg . The feature vectors P max and P avg are added bitwise and passed through a Sigmoid function to obtain a weight vector W0, and its calculation process is shown in Equation (1). The feature sequence X is weighted using the weight vector W0 to obtain a weighted feature sequence X' with size , and its calculation method is shown in Equation (2).

[0018] W0 = Sigmoid(P max + P avg ) (1)

[0019] X' = X · W0 (2)

[0020] Through the integration of the temporal attention module TAM, the model can dynamically adjust the degree of attention to different time steps during the feature extraction process, and effectively capture the time steps that help distinguish the boundaries between different chicken behaviors.

[0021] In step 3), the improved transformer block is improved from the original transformer block. The improvement method is: introducing a hybrid position encoding module before the first Norm layer of the transformer block, and introducing a multi-scale dilated convolutional layer after the first Norm layer. Among them, the multi-scale dilated convolutional layer adds the outputs of the convolution with a dilation rate of 1, the convolution with a dilation rate of 2, and the convolution with a dilation rate of 4 bitwise. The hybrid position encoding automatically adjusts the ratio of the absolute position encoding and the relative position encoding through the learnable parameter α, realizing the dynamic fusion of the global and local position perception capabilities, and its calculation method is shown in Equation (4).

[0022] PE mixed = α · PE abs + (1 - α) · PE rel (4)

[0023] Among them, PE abs and PE rel represent the absolute position encoding and the relative position encoding respectively. α is a learnable parameter used to dynamically adjust the weighted ratio of the two, and α ∈ [0, 1].

[0024] The introduction of the hybrid position encoding module in the Improved transformer block can enhance the model's ability to perceive the positions of different behavior sequences of chickens, and the introduction of the multi-scale dilated convolutional layer can further enhance the model's ability to extract sequence information of different scales.

[0025] In step 3), each of the three cascaded feature pyramids contains an encoder-decoder structure. In the decoder, channel alignment and feature integration are completed based on 1×1 convolutional operations to achieve the hierarchical fusion of multi-scale features.

[0026] In step 3), the calculation method of the weighted regression loss function is shown in Equation (3).

[0027]

[0028] Among them, M is the number of positive samples, IoU(b i , g i ) is the intersection over union of the predicted time interval b i and the ground-truth time interval g i , ρ(b i , g i ) is the Euclidean distance between the center points of the predicted time interval b i and the ground-truth time interval g i , c i is the length of the minimum closed interval of the predicted time interval b i and the ground-truth time interval g i , α d is the weight coefficient between IoU and the center deviation, and L reg is the weighted regression loss.

[0029] The introduction of the weighted regression loss function can enhance the model's comprehensive optimization ability for the time interval overlap quality and the interval center positioning accuracy of chicken behaviors during the training process.

[0030] Further preferably, the key of a chicken behavior recognition and positioning method based on spatio-temporal feature learning is a chicken behavior recognition and positioning model CBLFormer that fuses spatio-temporal features, which mainly includes the following steps:

[0031] A chicken behavior recognition and localization model CBLFormer based on spatio-temporal feature learning. First, the TSP module is used to extract features from the videos in the chicken behavior recognition and localization dataset, obtaining a video feature sequence of size (D, T), where D is the feature dimension and T is the number of features, i.e., the number of time steps. Then, the video feature sequence is fed into the backbone of CBLFormer, and masked 1D convolution and improved transformer blocks are used to extract semantic information at different levels in the feature sequence layer by layer. The temporal attention module TAM (Temporal Attention Module) is integrated to capture fine-grained temporal dependencies, thereby obtaining five feature maps with different temporal resolutions. The five feature maps with different temporal resolutions are fed into the neck for feature fusion and interaction based on three groups of cascaded feature pyramids. The five fused feature maps are fed into the head based on a linear transformer, where the classification head is responsible for predicting the behavior category of the chicken at each time step, and the regression head is responsible for predicting the distances from each time step to the start and end points of the behavior interval, finally realizing the recognition of chicken behavior categories and the localization of the time intervals corresponding to the categories.

[0032] The principle of behavior recognition and localization of the CBLFormer model: The chicken video is extracted with features by the pre-trained TSP model to form a feature sequence X = {x1, x2, …, x T}}, and the task of behavior recognition and localization is to obtain the corresponding y i for each x i (1 ≤ i ≤ t), y i = (s i , e i , a i ), where s i is the start point of the behavior, e i is the end point of the behavior, and a i is the category of the behavior. The above x i → y i problem can be transformed into x i → y′ i , y′ i = where is the distance from the current time step to the start point of the behavior, is the distance from the current time step to the end point of the behavior, and p(a i ) is the probability that the current time step belongs to each behavior category. The decoding method for the behavior category is shown in Equation (5), and the decoding methods for the start and end points of the behavior interval are shown in Equations (6 - 7).

[0033] a i = argmax(p(ai )) (5)

[0034]

[0035] Specifically, CBLFormer is improved based on the temporal action localization model ActionFormer. The main improvement points include: 1) integrating the temporal attention module TAM in the backbone part of the model; 2) improving the transformer block used in the backbone part of the model to an improved transformer block; 3) replacing the feature pyramid structure FPN in the neck part of the model with three cascaded feature pyramids; 4) improving the 1D convolutional-based head of the model to a linear transformer-based head; 5) improving the regression loss function of the model to a weighted regression loss function.

[0036] The TSP (Temporally-Sensitive Pretraining) module in CBLFormer is used for video feature extraction, aiming to provide high-quality and temporally sensitive feature representations for subsequent chicken behavior recognition and localization tasks. Fine-tuning the TSP module using the pre-trained dataset enhances the model's ability to distinguish different chicken behavior categories, enabling the model to better adapt to the target feature space. The key parameters during fine-tuning of TSP include clip length, stride, and fps, where the clip length is set to 4, the stride is set to 4, and the fps is set to 25. This parameter setting can effectively capture the short-term and long-term dependencies of the video, providing high-quality video feature representations for subsequent behavior recognition and localization tasks. When extracting video features using the fine-tuned TSP module, the clip length, stride, and fps parameters are kept consistent with those during fine-tuning of TSP. The duration of each clip is 0.16s, that is, the time resolution of the CBLFormer model is 0.16s. This resolution can make full use of the temporal sensitivity of TSP to provide strong support for the accurate recognition of chicken behavior categories and the precise localization of time boundaries.

[0037] The principle of the temporal attention module TAM used in CBLFormer is as follows: Max-pooling and average-pooling the feature sequence X with size output from the previous feature extraction block along the time step T to obtain a feature vector P with size (1, T) max and P avg , and the feature vector P max and P avgThe weight vector W0 is obtained by adding the two bits and passing through a Sigmoid function. The calculation process is shown in formula (1). The weight vector W is used to weight the feature sequence X, and the size is The weighted feature sequence X′ is calculated as shown in formula (2).

[0038] W0=Sigmoid(P max +P avg ) (1)

[0039] X′=X·W0 (2)

[0040] The principle of the improved transformer block used in CBLFormer is as follows: a multi-scale dilated convolution operation is introduced after the Norm layer in the original transformer block, with dilation rates of 1, 2, and 4, respectively, to capture contextual information of different temporal resolutions from the feature sequence. The outputs of the convolution with dilation rate of 1, the convolution with dilation rate of 2, and the convolution with dilation rate of 4 are added and sent to the local multi-head convolutional attention layer (MHCA) for self-attention calculation to model the long-range dependency between feature sequences. After the Norm layer and the MLP layer, the calculation result of the entire improved transformer block is output. In addition, hybrid position encoding is used in the improved transformer block to enhance the model's position perception ability for the temporal behaviors of different chickens. The hybrid position encoding automatically adjusts the ratio of absolute position encoding and relative position encoding through the learnable parameter α, realizing the dynamic fusion of global and local position perception capabilities. The calculation method is shown in formula (4).

[0041] PE mixed =α·PE abs +(1-α)·PE rel (4)

[0042] Among them, PE abs and PE rel They represent absolute position encoding and relative position encoding respectively, α is a learnable parameter used to dynamically adjust the weighted ratio of the two, and α∈[0,1].

[0043] For the three cascaded feature pyramid structures used in the neck part of CBLFormer, each group of feature pyramids contains an encoder-decoder structure. In the decoder, channel alignment and feature integration are completed based on 1×1 convolution operations to achieve the layer-by-layer fusion of multi-scale features. This design enables high-level features to guide the learning of low-level features at an earlier stage, thereby effectively improving the feature representation ability of CBLFormer in the chicken behavior recognition and localization tasks.

[0044] For the structure of the linear-transformer-based head in CBLFormer, the classification head and the regression head adopt the same structure to model the long-range dependencies between feature sequences. The classification head predicts the classification scores for each time step through a 1×1 masked 1D convolutional layer, and the regression head predicts the offset values from the start and end points of the chicken behavior time interval for each time step based on the masked 1D convolutional layer. Compared with the original head that only relies on masked 1D convolution, the improved head can perform more fine-grained modeling on the five feature sequences with different time resolutions from the neck, providing more accurate outputs for the chicken behavior classification task and the behavior interval localization task.

[0045] The total loss function of the CBLFormer model consists of a classification loss function and a regression loss function. The classification loss function uses Focal loss, which effectively improves the learning ability for difficult-to-distinguish chicken behavior categories by assigning higher weights to the feature sequences with classification difficulties. The regression loss function uses DIoU loss to optimize the loss between the predicted offset values and the ground truth offset values for each time step. To further improve the localization accuracy of the CBLFormer model for the time intervals of complex temporal actions, the present invention improves the regression loss function DIoU loss to WDIoU loss. WDIoU loss introduces a dynamic weight balance mechanism between IoU and the center deviation on the basis of DIoU loss, which is used to enhance the comprehensive optimization ability of the model for the time interval overlap quality and the interval center localization accuracy of the target behavior during the training process. The calculation method of WDIoU loss is shown in Equation (3).

[0046]

[0047] Where M is the number of positive samples, IoU(b i ,g i ) is the intersection over union of the predicted time interval b i and the ground truth time interval g i of sample i, and ρ(b i ,gi ) is the Euclidean distance of the center point for the predicted time interval b i and the true time interval g i denoted as c i is the Euclidean distance of the center point for the predicted time interval b i and the true time interval g i and the length of the minimum closed interval of g is denoted as α d is the weight coefficient between IoU and the center deviation, denoted as L reg is the weighted regression loss. By setting the value of α d to 0.6, the CBLFormer model can achieve optimal performance on the behavior recognition and localization dataset in the present invention.

[0048] Compared with the prior art, the present invention has the following advantages:

[0049] The present invention proposes a chicken behavior recognition and localization model CBLFormer based on spatio-temporal feature learning, which fully utilizes the local feature extraction ability of the convolutional operator and the long-range modeling ability of the self-attention mechanism to extract semantic information at different levels in the video feature sequence layer by layer. The time attention module TAM is used to enhance the model's ability to capture fine-grained temporal dependence relationships. The feature fusion and interaction ability of the model is enhanced by using a feature pyramid based on three cascades. The weighted regression loss function is used to enhance the model's comprehensive optimization ability for the time interval overlap quality and the interval center localization accuracy of the target behavior during the training process. Finally, the recognition of the behaviors that occur successively by chickens in the video and the localization of the time interval corresponding to each chicken behavior are realized. The spatio-temporal feature learning-based chicken behavior recognition and localization method proposed by the present invention can effectively improve the time resolution of video-based behavior recognition methods, and can also be integrated into a chicken spatio-temporal behavior detection system for continuous monitoring of multiple behavior events of chickens in a multi-target scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the overall architecture diagram of the chicken behavior recognition and model CBLFormer in the specific embodiment of the present invention;

[0051] Figure 2 is the architecture diagram of the time attention module of the chicken behavior recognition and model CBLFormer in the specific embodiment of the present invention;

[0052] Figure 3 is the architecture diagram of the improved Transformer block in the chicken behavior recognition and model CBLFormer in the specific embodiment of the present invention;

[0053] Figure 4 is the architecture diagram of the improved Head in the chicken behavior recognition and model CBLFormer in the specific embodiment of the present invention;

[0054] Figure 5 It is a feature visualization diagram of the CBLFormer model in the specific implementation manner of the present invention. Specific implementation manner

[0055] The specific implementation manner of the present invention will be described in detail below in conjunction with the accompanying drawings.

[0056] The present invention proposes a chicken behavior recognition and localization model CBLFormer based on spatio-temporal feature learning, as Figure 1 shown. First, use the time-sensitive pre-training module TSP to extract features from the videos in the chicken behavior recognition and localization dataset, obtaining a video feature sequence of size (D, T), where D is the feature dimension and T is the number of features, i.e., the number of time steps. Then, send the video feature sequence into the backbone of CBLFormer, and use masked 1D convolution and improved transformer blocks to extract semantic information at different levels in the feature sequence layer by layer, and integrate the temporal attention module TAM (Temporal Attention Module) to capture fine-grained temporal dependence relationships, thereby obtaining five feature maps with different temporal resolutions. Send the five feature maps with different temporal resolutions into the neck for feature fusion and interaction based on three groups of cascaded feature pyramids. Send the fused five feature maps into the head based on the linear transformer, where the classification head is responsible for predicting the behavior category of the chicken at each time step, and the regression head is responsible for predicting the distances from each time step to the start and end points of the behavior interval, finally realizing the recognition of chicken behavior categories and the localization of the time intervals corresponding to the categories.

[0057] The principle of behavior recognition and localization of the CBLFormer model: The chicken video is extracted with features by the pre-trained TSP model to form a feature sequence X = {x1, x2, …, x T}, and the task of behavior recognition and localization is to obtain y i corresponding to each x i (1 ≤ i ≤ T), y i = (s i , e i , a i ), where s i is the start point of the behavior, e i is the end point of the behavior, and a i is the category of the behavior. The above x i → y i problem can be transformed into x i → y′ i , y′ i = where is the distance from the current time step to the start point of the behavior, is the distance from the current time step to the end point of the behavior, p(a i ) is the probability that the current time step belongs to each behavior category. The decoding method of the behavior category is shown in Equation (5), and the decoding methods of the start and end points of the behavior interval are shown in Equations (6 - 7).

[0058] a i = argmax(p(a i )) (5)

[0059]

[0060] Specifically, CBLFormer is improved based on the temporal action localization model ActionFormer. The main improvement points include: 1) integrating the temporal attention module TAM in the backbone part of the model; 2) improving the transformer block used in the backbone part of the model to the improved transformer block; 3) replacing the feature pyramid structure FPN in the neck part of the model with three cascaded feature pyramids; 4) improving the 1D convolutional-based head of the model to a linear transformer-based head; 5) improving the regression loss function of the model to the weighted regression loss function.

[0061] The TSP (Temporally-Sensitive Pretraining) module in CBLFormer is used for video feature extraction, aiming to provide high-quality and temporally sensitive feature representations for subsequent chicken behavior recognition and localization tasks. Fine-tuning the TSP module using the pre-training dataset enhances the model's ability to distinguish different chicken behavior categories, enabling the model to better adapt to the target feature space. The key parameters during fine-tuning of TSP include clip length, stride, and fps, where the clip length is set to 4, the stride is set to 4, and the fps is set to 25. This parameter setting can effectively capture the short-term and long-term dependencies of the video, providing high-quality video feature representations for subsequent behavior recognition and localization tasks. When extracting video features using the fine-tuned TSP module, the clip length, stride, and fps parameters are kept consistent with those during fine-tuning of TSP. The duration represented by each clip is 0.16s, that is, the time resolution of the CBLFormer model is 0.16s. This resolution can make full use of the temporal sensitivity of TSP, providing strong support for the accurate recognition of chicken behavior categories and the precise localization of time boundaries.

[0062] The principle of the temporal attention module TAM used in CBLFormer is asFigure 2 As shown. Max-pooling and average-pooling are performed on the feature sequence X of size along the time step T, where D is the dimension of the feature sequence, to obtain a feature vector P of size (1, T) max and P avg . The feature vectors P max and P avg are added bitwise and passed through a Sigmoid function to obtain the weight vector W0, and its calculation process is shown in Equation (1). The feature sequence X is weighted using the weight vector W0 to obtain a weighted feature sequence X′ of size , and its calculation method is shown in Equation (2).

[0063] W0 = Sigmoid(P max + P avg ) (1)

[0064] X′ = X · W0 (2)

[0065] The principle of the improved transformer block used in CBLFormer is as Figure 3 shown. After the Norm layer in the original transformer block, multi-scale dilated convolution operations are introduced, with dilation rates of 1, 2, and 4 respectively, to capture context information with different temporal resolutions from the feature sequence. The outputs of the convolution with dilation rate 1, the convolution with dilation rate 2, and the convolution with dilation rate 4 are added bitwise and then fed into the Local Multi-Head Convolutional Attention (MHCA) layer for self-attention calculation to model the long-range dependencies between feature sequences. After passing through the Norm layer and the MLP layer, the calculation result of the entire improved transformer block is output. In addition, hybrid position encoding is adopted in the Improved transformer block to enhance the model's position perception ability for different chicken temporal behaviors. The hybrid position encoding automatically adjusts the ratio of absolute position encoding and relative position encoding through the learnable parameter α, realizing the dynamic fusion of global and local position perception abilities, and its calculation method is shown in Equation (4).

[0066] PE mixed = α · PE abs + (1 - α) · PE rel (4)

[0067] where, Pe abs and PE relThey represent absolute position encoding and relative position encoding respectively. α is a learnable parameter used to dynamically adjust the weighted ratio of the two, and α ∈ [0, 1].

[0068] For the three cascaded feature pyramid structures used in the neck of CBLFormer, as Figure 1 shown, each group of feature pyramids contains an encoder-decoder structure. In the decoder, channel alignment and feature integration are completed based on 1×1 convolutional operations to achieve the hierarchical fusion of multi-scale features. This design enables high-level features to guide the learning of low-level features at an earlier stage, thus effectively enhancing the feature representation ability of CBLFormer in the chicken behavior recognition and localization tasks.

[0069] The structure of the head based on the linear transformer in CBLFormer is as Figure 4 shown. The classification head and the regression head adopt the same structure to model the long-range dependencies between feature sequences. The classification head predicts the classification scores for each time step through a 1×1 masked 1D convolutional layer, and the regression head predicts the offset values from the start and end points of the chicken behavior time interval for each time step based on the masked 1D convolutional layer. Compared with the original head that only relies on the masked 1D convolution, the improved head can perform more fine-grained modeling on the five feature sequences with different time resolutions from the neck, providing more accurate outputs for the chicken behavior classification task and the behavior interval localization task.

[0070] The total loss function of the CBLFormer model consists of a classification loss function and a regression loss function. The classification loss function uses Focal loss, which effectively enhances the learning ability for difficult-to-distinguish chicken behavior categories by assigning higher weights to the feature sequences with classification difficulties. The regression loss function uses DIoU loss to optimize the loss between the predicted offset values and the true offset values for each time step. To further improve the localization accuracy of the time interval of complex temporal actions by the CBLFormer model, the present invention improves the regression loss function DIoU loss to WDIoU loss. WDIoU loss introduces a dynamic weight balancing mechanism between IoU and the center deviation on the basis of DIoU loss, which is used to enhance the comprehensive optimization ability of the model for the time interval overlap quality and the interval center localization accuracy of the target behavior during the training process. The calculation method of WDIoU loss is shown in Equation (3).

[0071]

[0072] where M is the number of positive samples, IoU(b i ,g i ) is the intersection over union of the predicted time interval b i and the ground truth time interval g i , ρ(b i ,g i ) is the Euclidean distance between the centers of the predicted time interval b i and the ground truth time interval g i , c i is the length of the minimum closed interval of the predicted time interval b i and the ground truth time interval g i , α d is the weight coefficient between IoU and the center deviation, and L reg is the weighted regression loss. By setting the value of α d to 0.6, the CBLFormer model can achieve optimal performance on the action recognition and localization dataset in the present invention.

[0073] After training the CBLFormer model using the action recognition and localization dataset, the present invention visualized the output results of the model at the head. By performing the Softmax operation on the classification output, the action class scores at each time step were calculated, and the weighted boundary offset values (onset and offset) of the regression head output were obtained by weighting the boundary offsets with the action class scores, as Figure 5 shown. The score curve clearly shows the conversion of action classes. The peak of the classification score at the 51st time step is significant, corresponding to the time node when the chicken behavior changes from resting to activity, indicating that the model has high temporal sensitivity and action class discrimination ability in the classification task. From the predicted histogram of the weighted boundary offsets, it can be seen that near the action conversion node, the predicted values of the weighted onset and offset show significant peaks, and the offsets far from the action conversion node gradually decrease until they increase again near the start of the resting behavior and near the end of the activity behavior. The model successfully distinguishes the central region and the boundary region of different behaviors. This visualization result proves that the CBLFormer model can effectively identify chicken behaviors and locate the behavior boundaries, providing strong support for the interpretability of the model decision-making process.

Claims

1. A chicken behavior recognition and positioning method based on spatiotemporal feature learning, characterized in that: The following steps are involved: 1) Obtain chicken video data and create pre-training datasets and chicken behavior recognition and positioning datasets; 2) performing data enhancement on the chicken behavior recognition and positioning data set obtained in step 1) to obtain an enhanced behavior recognition and positioning data set; 3) Adopting the ActionFormer model, adding a time-sensitive pre-training module to the data input of the ActionFormer model, adding a temporal attention module to the backbone of the ActionFormer model, and improving the transformer block in the backbone of the ActionFormer model to an improved transformer block, replacing the feature pyramid structure of the neck of the ActionFormer model with three sets of cascaded feature pyramids, improving the head of the ActionFormer model to a head based on a linear transformer, and improving the regression loss function of the ActionFormer model to a weighted regression loss function, thereby forming an improved model; 4) using the pre-training data set described in step 1) to fine-tune the time-sensitive pre-training module in the improved model obtained in step 3) to obtain a time-sensitive pre-training module with reasoning ability, and obtain a chicken behavior recognition and positioning model CBLFormer; 5) using the enhanced behavior recognition and positioning data set described in step 2) to train and optimize the chicken behavior recognition and positioning model CBLFormer of step 4), wherein the time-sensitive pre-training module with reasoning ability of step 4) is used to extract the video of the enhanced behavior recognition and positioning data set into a feature sequence to obtain the chicken behavior recognition and positioning model CBLFormer with reasoning ability; 6) The chicken behavior recognition and positioning model CBLFormer with reasoning ability is used to identify the chicken video, and obtain the chicken behaviors that occur successively in the video and the time interval of each chicken behavior.

2. The chicken behavior recognition and positioning method based on spatiotemporal feature learning according to claim 1 is characterized in that: In step 1), the method for preparing the pre-training data set includes: Video clips of individual chickens are captured from the obtained chicken flock video data. The duration of the video clips is 1s to 100s. Each video clip only contains the same behavior repeated by the same chicken. A category label is assigned to each video clip to form a pre-training dataset.

3. The chicken behavior recognition and positioning method based on spatiotemporal feature learning according to claim 1, characterized in that: In step 1), the method for preparing the chicken behavior recognition and positioning data set includes: Video clips of individual chickens are captured from the obtained chicken flock video data, and the video clip length is limited to 1s to 100s. Each video clip contains multiple different behaviors of the same chicken, and the expression time of each behavior must exceed 0.5s. The information of each video clip is annotated, including the video file name, video length, resolution, time interval and category label, to form a chicken behavior recognition and positioning dataset.

4. The chicken behavior recognition and positioning method based on spatiotemporal feature learning according to claim 1 is characterized in that: In step 3), the temporal attention module processes data through the following process: The feature sequence X input to the temporal attention module is pooled by maximum and average pooling along the time step T direction to obtain the feature vector P max and the eigenvector P avg , the feature vector P max and the eigenvector P avg The weight vector W0 is obtained by adding the features and passing it through a Sigmoid function. The calculation process is shown in formula (1). The weight vector W0 is used to weight the feature sequence X to obtain the weighted feature sequence X ′ , its calculation method is shown in formula (2); W0=Sigmoid(P max +P avg ) (1) X ′ =X·W0 (2)。 5. The chicken behavior recognition and positioning method based on spatiotemporal feature learning according to claim 1 is characterized in that: In step 3), the transformer block in the backbone of the ActionFormer model is improved to an improved transformer block, including: A hybrid position encoding module is introduced before the first Norm layer of the transformer block, and a multi-scale dilated convolutional layer is introduced after the first Norm layer of the transformer block.

6. The chicken behavior recognition and positioning method based on spatiotemporal feature learning according to claim 1 is characterized in that: In step 3), the calculation method of the weighted regression loss function is shown in formula (3): Among them, M is the number of positive samples, IoU(b i ,g i ) is the prediction time interval b of sample i i and the real time interval g i The intersection-over-combination ratio, ρ(b i ,g i ) is the prediction time interval b i and the real time interval g i Euclidean distance of the center point, c i is the prediction time interval b i and the real time interval g i The minimum closure interval length, α d is the weight coefficient between IoU and center deviation, L reg is the weighted regression loss.