A method for beef cattle behavior recognition based on video spatiotemporal features
Through the spatiotemporal aggregation and excitation model of the dual-branch spectrum channel, the long and short spatiotemporal characteristics in the video are extracted and spatial characteristics are enhanced, which solves the problem of insufficient recognition accuracy of beef cattle behavior in the prior art, and achieves higher recognition accuracy and effect.
Patent Information
- Application Number
- CN202310469187.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-04-27
AI Technical Summary
The existing cattle behavior recognition method based on deep learning cannot fully utilize the temporal characteristics of the video, and it is difficult to distinguish behaviors with highly similar motion characteristics. The three-dimensional convolutional neural network is only suitable for the recognition of videos of a single target behavior, and it is difficult to meet the needs of the real breeding process.
The spatiotemporal aggregation and excitation model of the two-branch spectrum channel is adopted to construct a beef cattle behavior recognition model through the video sampling module, feature extraction network and prediction module. The model extracts the long and short space-time characteristics of the video through the motion excitation module, the motion aggregation module and the channel spectrum attention module, and enhances the spatial characteristics through the spatial feature extraction branches to ensure that the model captures the feature information of the beef cattle behavior video more comprehensively.
The accuracy of beef cattle behavior recognition was improved, and the F1 score of the DB-TEAF method on the beef cattle behavior data set reached 86.35%, and the accuracy rate reached 92.61%, which significantly improved the recognition accuracy in the prior art.
Smart Images

Figure CN116612409B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of video image understanding, relates to beef cattle behavior, and specifically relates to a beef cattle behavior recognition method based on video spatiotemporal features. Background Art
[0002] The behavior of beef cattle can reflect their health status and thus the quality of beef. Beef cattle behavior recognition is an important technical means in the modern beef cattle breeding industry. It refers to directly identifying the behavior of beef cattle through video images, thereby obtaining the behavioral status of beef cattle without contact. Its goal is to provide support for beef cattle health management.
[0003] In recent years, with the development of image processing technology and deep learning methods, animal behavior recognition methods based on deep learning have made significant progress. The current animal behavior recognition methods can be regarded as detecting or classifying animal targets in videos or images. Compared with traditional single computer vision methods, they effectively improve the detection efficiency and accuracy. For example, Wang Shaohua and He Dongjian proposed a method for recognizing the estrus behavior of dairy cows based on the improved YOLO v3 model in the paper "Research on the recognition of estrus behavior of dairy cows based on the improved YOLO v3 model". The size of the anchor box of the YOLO v3 model was re-clustered and optimized, and dense modules were introduced. The bounding box loss function with the intersection-over-union ratio and the center distance of the two boxes was used as the measurement method to improve the recognition accuracy of the model and realize the recognition of the cow's climbing behavior in complex breeding links. For example, Ma et al. proposed a three-dimensional convolutional neural network based on the improved Rexnet in the paper "Basic motion behavior recognition of single dairy cow based on improvedRexnet 3D network." By fusing the temporal information of the image sequence, the recognition of basic motion behaviors of dairy cows such as lying, standing, and walking was realized. For example, Qiao et al. proposed a three-dimensional convolutional neural network based on long short-term memory in the paper “C3D-ConvLSTM based cow behavior classification using video data for precision livestock farming”. First, the features of the video sequence were extracted through the three-dimensional convolutional neural network, and then the long short-term memory module was added to the three-dimensional convolutional neural network to further extract the spatiotemporal features of the image sequence, thereby realizing the recognition of individual cow behaviors such as eating, exploring, grooming, walking, and standing.
[0004] However, there are still some problems with the current cattle behavior recognition method based on deep learning:
[0005] (A) Using object detection methods to recognize cattle behavior cannot fully utilize the temporal characteristics of the video and cannot distinguish behaviors with highly similar motion features.
[0006] (B) The method of using three-dimensional convolutional neural network to realize cattle behavior recognition is still in the research stage. It is currently only applicable to the recognition of a single target behavior video and is difficult to meet the needs of real breeding links. Summary of the invention
[0007] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a beef cattle behavior recognition method based on video spatiotemporal features, so as to solve the technical problem that the accuracy of the beef cattle behavior recognition method in the prior art needs to be further improved.
[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions to achieve the above problems:
[0009] A method for identifying beef cattle behavior based on video spatiotemporal features, the method comprising the following steps:
[0010] Step S1, constructing a beef cattle behavior dataset:
[0011] By collecting real farm monitoring videos and cropping them to obtain beef cattle behavior video clips, the beef cattle behavior video clips are labeled according to the beef cattle behavior to form a beef cattle behavior dataset.
[0012] The beef cattle behaviors include mounting behavior, running behavior, fighting behavior and normal behavior.
[0013] Step S2, constructing a beef cattle behavior recognition model:
[0014] A dual-branch spectrum channel spatiotemporal aggregation and excitation model is adopted to construct a beef cattle behavior recognition model, which includes a video sampling module, a feature extraction network and a prediction module.
[0015] The method for constructing the beef cattle behavior recognition model includes the following methods:
[0016] Step S201, constructing a video sampling module.
[0017] Step S202, constructing a feature extraction network.
[0018] Step S203, constructing a prediction module.
[0019] Step S3, beef cattle behavior recognition model training:
[0020] Step S4, beef cattle behavior identification:
[0021] The beef cattle behavior video is input into the beef cattle behavior recognition model trained in step S3 to recognize the beef cattle behavior.
[0022] The present invention also has the following technical features:
[0023] Specifically, in step 201, the specific process of constructing the video sampling module is as follows: split the cattle behavior video into video frames, input the video sampling module, first sample 30 frames at equal intervals, and form a video frame sequence V o = {I(x,y,1),(x,y,2),…,(x,y,T), calculate the video frame sequence V o The sparse regular operator of the RGB difference between two adjacent frames is used, and the four groups of adjacent frames with the largest sparse regular operator are selected to form a video frame sequence M, and the video frame sequence M is input into the subsequent feature extraction network.
[0024] Where:
[0025] V o represents a sequence of video frames;
[0026] I(x,y,t) represents each acquired video frame;
[0027] t represents the frame number, t = 1, 2, ..., T;
[0028] x is the width of the input video frame;
[0029] y is the height of the input video frame.
[0030] Specifically, in step S202, the specific process of constructing the feature extraction network is as follows: the feature extraction network includes two branches, namely, a motion excitation and aggregation branch and a spatial feature extraction branch; the two branches extract temporal features and spatial features from the input video frame sequence M to obtain temporal feature maps Y respectively. t and spatial feature map Y c .
[0031] More specifically, the specific process of step S202 is:
[0032] Step S20201, the video frame sequence M is input into the feature extraction network, and after a convolution operation, a feature map F is output, and the size of F is [N, T, C, H, W].
[0033] Where:
[0034] N represents the batch size;
[0035] T represents the time dimension;
[0036] C represents the number of channels;
[0037] H represents the feature map height;
[0038] W represents the feature map width.
[0039] Step S20202, constructing a motion excitation and aggregation branch, wherein the motion excitation and aggregation branch includes a motion excitation and aggregation module and a channel spectrum attention module.
[0040] The motion excitation and aggregation module includes a motion excitation module and a motion aggregation module; the feature map F is input to the motion excitation and aggregation branch, and the motion excitation module extracts the motion difference of adjacent frames to obtain the feature map F. o ; The motion aggregation module performs feature fusion on all channels, extracts long-term features, and obtains the feature map g.
[0041] After compression by the channel spectrum attention module according to different weights, the feature map H is obtained; after multiple operations, the output time feature map Y of the motion excitation and aggregation branch is obtained t .
[0042] Step S20203, construct a spatial feature enhancement branch containing 3 convolution modules; used to enhance the spatial features in the feature map F; after each convolution operation, the size of the output feature map is transformed to half of the input feature map, the input of the spatial feature extraction branch is the feature map F in step S20201, and the output is the spatial feature map Y c .
[0043] Preferably, in step S20201, H=H input / 2, W = W input / 2, C = 64, H input =256 and W input =455,H input is the height of the input video frame, W input The width of the input video frame.
[0044] Specifically, in step S203, the specific process of constructing the prediction module is as follows: splicing the time feature map Y t and spatial feature map Y c The feature extraction network output Y is obtained. The feature extraction network output Y undergoes channel transformation, full connection operation and maximum pooling operation, and finally outputs the predicted label L of the video pre .
[0045] Specifically, in step S3, the specific process of the beef cattle behavior recognition model training is as follows: when training the beef cattle behavior recognition model, a focal loss function is used to calculate the loss L of the model prediction probability and the real label of the video. f , back propagation loss L f , repeat the iteration until the number of iterations reaches the preset initial value to complete the training.
[0046] L f =-αt (1-p t ) γ log(p t )
[0047] Where:
[0048] α t Indicates the weights of different categories;
[0049] p t Represents the class probability predicted by the model;
[0050] γ is used to suppress the loss contribution of simple samples and promote the loss contribution of difficult samples.
[0051] Compared with the prior art, the present invention has the following technical effects:
[0052] (I) The present invention can extract frames with significant motion differences from a video frame sequence to ensure that the motion information of beef cattle is fully retained; furthermore, the long- and short-term spatiotemporal features in the video are extracted through a motion excitation module and a motion aggregation module, and the representation weights of different channels of the feature map for motion are adjusted through a spectral channel attention mechanism, and a spatial feature extraction branch is constructed to enhance the spatial features. The model captures the feature information in the beef cattle behavior video more comprehensively and has a higher accuracy in identifying the beef cattle behavior.
[0053] (II) The DB-TEAF method of the present invention achieved an F1 score of 86.35% and an accuracy of 92.61% on the constructed beef cattle behavior dataset, and the results of beef cattle behavior video recognition were more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is the structural diagram of the beef cattle behavior recognition model.
[0055] Figure 2 It is a structural diagram of motion excitation and aggregation branches.
[0056] Figure 3 3 is a comparison chart of the beef cattle behavior video recognition results of the method of the present invention and other existing methods in the embodiment.
[0057] Figure 4 (a) Visualized heat map of None (blank) control.
[0058] Figure 4(b) is a visualized heat map of the BAM method.
[0059] Figure 4(c) is a visualized heat map of the SE method.
[0060] Figure 4(d) is a visualized heat map of the SRM method.
[0061] Figure 4(e) is a visualized heat map of the ECA method.
[0062] FIG4( f ) is a visualized heat map of the DB-TEAF method of the present invention.
[0063] Figure 5 It is the behavior recognition result of each category of the method of the present invention and the model trained by using other loss functions in the embodiment of the present invention.
[0064] The specific contents of the present invention are further explained in detail below in conjunction with embodiments. DETAILED DESCRIPTION
[0065] It should be noted that, unless otherwise specified, all algorithms in the present invention adopt algorithms known in the prior art.
[0066] In order to cope with the problems of complex background and unbalanced cattle behavior categories in real breeding scenes, the present invention provides a cattle behavior recognition method based on video spatiotemporal features (DB-TEAF method), which includes:
[0067] Firstly, a beef cattle behavior dataset consisting of real farm surveillance videos was constructed and the dataset was divided into training set, validation set and test set.
[0068] Then, a dual-branch spectral channel spatiotemporal aggregation and excitation model is constructed to realize the recognition of beef cattle behavior videos. The model includes a video sampling module, a feature extraction network and a prediction module. The video sampling module is used to obtain the most motion-representative video frames from the video and input them into the feature extraction network. The feature extraction network consists of a motion excitation and aggregation branch and a spatial feature extraction branch, which respectively extract spatiotemporal features and spatial features from the image sequence.
[0069] Finally, the features of the two branches are concatenated, channel compressed, fully connected, and pooled by the prediction module to obtain the predicted label of the video.
[0070] The present invention can extract temporal and spatial feature information representative of motion in a video, and the prediction of beef cattle behavior videos is more accurate.
[0071] In accordance with the above technical scheme, specific embodiments of the present invention are given below. It should be noted that the present invention is not limited to the following specific embodiments, and all equivalent changes made on the basis of the technical scheme of this application fall within the protection scope of the present invention.
[0072] Example:
[0073] This embodiment provides a method for identifying cattle behavior based on video spatiotemporal features, the method comprising the following steps:
[0074] Step S1, constructing a beef cattle behavior dataset:
[0075] By collecting real farm monitoring videos and cropping them to obtain beef cattle behavior video clips, the beef cattle behavior video clips are labeled according to the beef cattle behavior to form a beef cattle behavior dataset.
[0076] Beef cattle behaviors include mounting behavior, running behavior, fighting behavior and normal behavior.
[0077] In this embodiment, the beef cattle behavior data set includes 1285 videos, including 771 training sets, 257 validation sets, and 257 test sets. The duration of each video is between 3 seconds and 15 seconds. In addition, the categories in the data set are unbalanced, with 176 climbing behaviors, 90 fighting and running behaviors, and 931 normal behaviors. The distribution of data in different categories is more in line with the actual situation.
[0078] Step S2, constructing a beef cattle behavior recognition model:
[0079] like Figure 1 As shown, a dual-branch spectral channel spatiotemporal aggregation and excitation model is used to construct a beef cattle behavior recognition model, which includes a video sampling module, a feature extraction network and a prediction module.
[0080] The construction of the beef cattle behavior recognition model includes the following methods:
[0081] Step S201, constructing a video sampling module:
[0082] The beef cattle behavior video is split into video frames and input into the video sampling module. First, 30 frames are sampled at equal intervals to form a video frame sequence V o = {I(x,y,1),(x,y,2),…,(x,y,T), calculate the video frame sequence V o The sparse regular operator of the RGB difference between two adjacent frames is used, and the four groups of adjacent frames with the largest sparse regular operator are selected to form a video frame sequence M, and the video frame sequence M is input into the subsequent feature extraction network.
[0083] Where:
[0084] V o represents a sequence of video frames;
[0085] I(,y,t) represents each acquired video frame;
[0086] t represents the frame number, t = 1, 2, ..., T;
[0087] x is the width of the input video frame;
[0088] y is the height of the input video frame.
[0089] In this embodiment, preferably, x=455, y=256, and T=30.
[0090] Preferably, in this embodiment, the size of M is [8, 256, 455].
[0091] Step S202, constructing a feature extraction network:
[0092] The feature extraction network includes two branches, namely, a motion excitation and aggregation branch and a spatial feature extraction branch; the two branches extract temporal features and spatial features from the input video frame sequence M to obtain temporal feature maps Y respectively. t and spatial feature map Y c .
[0093] The specific process of step S202 is:
[0094] Step S20201: The video frame sequence M is input into the feature extraction network, and a convolution operation is performed to output a feature map F. The size of F is [N, T, C, H, W], where N represents the batch size, T represents the time dimension, C represents the number of channels, H represents the feature map height, W represents the feature map width, and H = H input / 2, W = W input / 2, C = 64, H input =256 and F input =455 is the height and width of the input video frame.
[0095] Preferably, in this embodiment, in step S20201, the convolution kernel size of the convolution operation is 7×7, the step size is 2, and the padding is 3.
[0096] Step S20202, as Figure 2 As shown in FIG. 1 , a motion excitation and aggregation branch is constructed, which includes a motion excitation and aggregation module and a channel spectrum attention module, wherein the motion excitation and aggregation module includes a motion excitation module and a motion aggregation module. The feature map of the input motion excitation and aggregation branch is F, and the motion excitation module extracts the features of the motion differences of adjacent frames to obtain the feature map F. o ; The motion aggregation module performs feature fusion on all channels, extracts long-term features, and obtains feature map G; the channel spectrum attention module compresses them according to different weights to obtain feature map H. After multiple operations, the output time feature map Y of the motion excitation and aggregation branch is obtained. t .
[0097] In this embodiment, preferably, in step S20202, the motion excitation module reduces the number of channels of F by a 1×1 convolution operation with a step size of 1 to obtain F r , divide it by time dimension, and get F r (t), adjacent F M (t) Calculate the motion difference F M (t) = conv trans (Fr (t+1))-F r (t),1≤t≤T-1,F r The size is F M The size of (t) is Use global pooling to calculate the motion difference F of all adjacent frames M .
[0098] F M =Pool([F M (1),F M (2),…,F M (T)])
[0099] Perform a 1×1 convolution operation and convert F M The number of channels of is expanded to the same as X, and the motion weights are calculated using the Sigmoid activation function to obtain the matrix M:
[0100] M=2(conv exp ( M )-1,M∈R N×T×C×1×1
[0101] Where:
[0102] Sigmoid represents the activation function;
[0103] conv exp Indicates channel expansion.
[0104] R represents the set of real numbers.
[0105] Use residual connection F and M to obtain F0:
[0106] F o =+F⊙M
[0107] Where:
[0108] + indicates corresponding addition of matrix elements;
[0109] ⊙ represents channel multiplication;
[0110] F o The dimensions are the same as F.
[0111] In this embodiment, preferably, in step S20202, the motion aggregation module aggregates F along the channel dimension. o Split into 4 groups of fragments, the last 3 fragments use 1 channel convolution and 1 convolution operation of size 3×3 to calculate the channel and spatial features respectively, and use residual connection to establish the hierarchical relationship between two adjacent groups of fragments:
[0112]
[0113]
[0114]
[0115] In the formula The dimensions are [N,T,C / 4,H,W].
[0116] Splice the four segments to get the output feature map G of the motion aggregation module:
[0117] G=[G 1 ; G 2 ; G 3 ; G 4 ],∈R N×T×C×H×W
[0118] Where:
[0119] R represents the set of real numbers.
[0120] In this embodiment, preferably, in step S20202, the channel spectrum attention module takes the feature map G as input. j (j∈N×T) is split into multiple sequences For each sequence Perform a two-dimensional discrete cosine transform to obtain the result of channel attention:
[0121]
[0122] Where:
[0123] [u i ,v i ]for The two-dimensional frequency components of
[0124] H and w with G j same.
[0125] Splicing Freq i , the attention weight vector is calculated by full connection and sigmoid activation
[0126] att=sigmoid(fc(cat([Freq0,Freq1,…,Freq n-1 ]))).
[0127] Where:
[0128] fc represents the fully connected layer;
[0129] cat represents a concatenation operation.
[0130] Each feature map G j Channels are compressed according to their corresponding attention weights:
[0131]
[0132] Fusion of all H j Get the output H of the channel spectral attention module.
[0133] Step S20203, refer to the structure in Table 1, and construct a spatial feature enhancement branch containing 3 convolution modules. It is used to enhance the spatial features in the feature map F, where N represents the batch size, Conv2d represents the two-dimensional convolution operation, BatchNorm2d represents the two-dimensional normalization layer, Relu represents the linear rectification function, and Maxpooling represents the maximum pooling operation. After each convolution operation, the size of the output feature map is transformed to half of the input feature map. The input of the spatial feature extraction branch is the feature map F in step S20201, and the output is the spatial feature map Y c .
[0134] Table 1. Spatial feature extraction branch structure
[0135]
[0136]
[0137] Preferably, in step S20203 of this embodiment, the convolution module includes a convolution layer with a convolution kernel size of 3×3 and a step size of 2, a normalization layer and a maximum pooling operation.
[0138] Step S203, constructing a prediction module:
[0139] Splicing time feature map Y t and spatial feature map Y c The feature extraction network output Y is obtained. The feature extraction network output Y undergoes channel transformation, full connection operation and maximum pooling operation, and finally outputs the predicted label L of the video pre .
[0140] Step S3, beef cattle behavior recognition model training:
[0141] When training the beef cattle behavior recognition model, the focal loss function is used to calculate the loss L between the model prediction probability and the video true label. f : Back propagation loss L f , repeat the iteration until the number of iterations reaches the preset initial value to complete the training.
[0142] L f =-α t (1- t ) γ log(p t )
[0143] Where:
[0144] α t Indicates the weights of different categories;
[0145] p t Represents the class probability predicted by the model;
[0146] γ is used to suppress the loss contribution of simple samples and promote the loss contribution of difficult samples.
[0147] Step S4, beef cattle behavior identification:
[0148] The beef cattle behavior video is input into the beef cattle behavior recognition model trained in step S3 to recognize the beef cattle behavior.
[0149] In the present invention, the test beef cattle behavior video is input into the beef cattle behavior recognition model trained in step S3, and the model outputs the model prediction label L pre , and use precision (Precision), recall (Recall), F1 score (F1-score) and accuracy (Accuracy) for evaluation.
[0150] Comparative Example 1:
[0151] This comparison compares the proposed method with 6 latest behavior recognition methods, including C3D (3D convolutional neural network), TSN (time segmentation network), I3D (inflated 3D convolutional network), TSM (time offset module), TEA (motion excitation and aggregation network), and TAM (time adaptive module). There are 4 evaluation indicators, namely precision, recall, F1 score, and accuracy.
[0152] Experimental verification: The results are shown in Table 1. The four evaluation indicators of the method (DB-TEAF) in the embodiment of the present invention are all optimal results and achieve the highest F1 score (86.35%) and accuracy (92.61%). Compared with the second best (TAM), DB-TEAF improves the F1 score by 7.48% and the accuracy by 2.34%. Figure 3 The results of the method in this embodiment and the three best methods for beef cattle climbing, fighting, running, and normal behavior video recognition are shown in Figure 1, where the red label indicates that the prediction category is wrong. Figure 3 From the first row, we can see that due to the high diversity of the training data of normal behavior, each method can well identify the videos of normal behavior of beef cattle. Figure 3It can be seen from the second, third and fourth rows that the method proposed in this embodiment focuses on the long and short spatiotemporal features of different channels, can accurately focus on the motion behavior in the video and accurately identify it, and has stronger motion capture capabilities and spatiotemporal feature acquisition capabilities than C3D, TEA and TAM.
[0153] Table 1 Beef cattle behavior recognition results of different methods
[0154]
[0155] Comparative Example 2:
[0156] This comparative example uses a sparse time sampling method and a dense sampling method to replace the video sampling module in step S201, and the other steps are the same as those in the embodiment.
[0157] Experimental verification: The results are shown in Table 2. Compared with the sparse time sampling method, the method of this embodiment has an accuracy improvement of 10.12%, a recall rate improvement of 2.14%, an F1 score improvement of 7.94%, and an accuracy improvement of 5.06%. The improvement in the evaluation indicators of the dense sampling method is more obvious. The results show that the video sampling method proposed in this embodiment can extract frames with significant motion in the video sequence, thereby improving the recognition effect of the beef cattle behavior recognition model.
[0158] Table 2 Beef cattle behavior recognition results of models with different video sampling modules
[0159]
[0160] Comparative Example 3:
[0161] This comparative example gives 5 beef cattle behavior recognition models combined with different attention modules. The difference between these methods and the embodiment is that the attention module in step S202 is removed or replaced in this comparative example, and the other steps are the same as the embodiment. Among them, the attention mechanisms used in 4 comparative examples are BAM (bottleneck attention module), SE (shrinkage and excitation module), SRM (style recalibration module), and ECA (efficient channel attention module). There are 4 evaluation indicators, namely precision, recall, F1 score, and accuracy.
[0162] Experimental verification: The results are shown in Table 3. The four evaluation indicators of this case method (DB-TEAF) are all optimal and achieve the highest F1 score (86.35%) and accuracy (92.61%). Figure 4(a) to Figure 4(f)This is a heat map of the method of this embodiment and other comparative methods. The closer the color of the image area is to dark red, the higher the model's attention to this area. The method of this embodiment can focus more on the beef cattle where the behavior occurs, which is beneficial for the model to extract features with better motion representation capabilities, thereby improving the accuracy of beef cattle behavior recognition.
[0163] Table 3. Beef cattle behavior recognition results using different attention mechanism modules
[0164]
[0165]
[0166] Comparative Example 4:
[0167] This comparative example provides a beef cattle behavior recognition model. The difference between this method and the embodiment is that this comparative example removes the spatial feature enhancement branch of step S202 in the embodiment, forming a different feature extraction network structure, and the other steps are the same as the embodiment.
[0168] Experimental verification: The results are shown in Table 4. Compared with the comparative example 4 without the spatial feature enhancement branch, the method of this embodiment with the spatial feature enhancement branch has an accuracy improvement of 8.79%, a recall rate improvement of 6.63%, an F1 score improvement of 7.32%, and an accuracy improvement of 4.67%. The results show that the spatial feature enhancement branch can improve the feature extraction network's ability to capture the behavioral characteristics of beef cattle, thereby enabling the beef cattle behavior recognition model to achieve better results.
[0169] Table 4 Beef cattle behavior recognition results under different feature extraction network structures
[0170]
[0171] Comparative Example 5:
[0172] In this comparative example, when training the model in step S3, a cross entropy loss function is used to calculate the loss of the model prediction probability and the true label of the video, and the other steps are the same as those in the embodiment.
[0173] Experimental verification: The results are shown in Table 5. Compared with the conventional model trained with the cross entropy loss function, the model trained with the focal loss function in this embodiment has an accuracy improvement of 2.36%, a recall rate improvement of 12.91%, an F1 score improvement of 10.3%, and an accuracy improvement of 5.06%. Figure 5As shown, the accuracy of the method in this embodiment is improved for the climbing, running, fighting or normal behavior of beef cattle, and the improvement is greater for the fighting and running behaviors with a relatively small number in the data set. The results show that the focus loss function training model used in this embodiment meets the needs of actual scenarios and further improves the accuracy of the beef cattle behavior recognition model.
[0174] Table 5 Beef cattle behavior recognition results of models trained with different loss functions
[0175]
Claims
1. A method for identifying cattle behavior based on spatiotemporal features of video, characterized in that: The method comprises the following steps: Step S1, constructing a beef cattle behavior dataset: By collecting real farm monitoring videos and cutting them to obtain beef cattle behavior video clips, the beef cattle behavior video clips are labeled according to the beef cattle behavior to form a beef cattle behavior dataset; The beef cattle behaviors include mounting behavior, running behavior, fighting behavior and normal behavior; Step S2, constructing a beef cattle behavior recognition model: A dual-branch spectral channel spatiotemporal aggregation and excitation model is used to construct a beef cattle behavior recognition model, which includes a video sampling module, a feature extraction network and a prediction module; The method for constructing the beef cattle behavior recognition model includes the following methods: Step S201, constructing a video sampling module; Step S202, constructing a feature extraction network; The feature extraction network includes two branches, namely, a motion excitation and aggregation branch and a spatial feature extraction branch; the two branches extract temporal features and spatial features from the input video frame sequence M to obtain temporal feature maps respectively. and spatial feature map ; The specific process of constructing the feature extraction network is as follows: Step S20201: The video frame sequence M is input into the feature extraction network, and a feature map is output after a convolution operation. , The size is ; Where: Indicates the batch size; Represents the time dimension; Indicates the number of channels; Indicates the feature map height; Indicates the feature map width; Step S20202, constructing a motion excitation and aggregation branch, wherein the motion excitation and aggregation branch includes a motion excitation and aggregation module and a channel spectrum attention module; The motion excitation and aggregation module includes a motion excitation module and a motion aggregation module; the feature map of the input motion excitation and aggregation branch is The motion excitation module extracts the motion differences of adjacent frames and obtains the feature map ; The motion aggregation module performs feature fusion on all channels, extracts long-term features, and obtains feature maps ; The channel spectrum attention module compresses the feature map according to different weights. ; After multiple operations, the output time characteristic graph of motion excitation and aggregation branch is obtained ; Step S20203, construct a spatial feature enhancement branch containing three convolution modules; used to enhance the feature map After each convolution operation, the size of the output feature map is transformed to half of the input feature map, and the input of the spatial feature extraction branch is the feature map in step S20201 , the output is a spatial feature map ; Step S203, constructing a prediction module; The specific process of constructing the prediction module is as follows: splicing the time feature map and spatial feature map Get the feature extraction network output , the feature extraction network output After channel transformation, full connection operation and maximum pooling operation, the predicted label of the video is finally output ; Step S3, beef cattle behavior recognition model training: Step S4, beef cattle behavior identification: The beef cattle behavior video is input into the beef cattle behavior recognition model trained in step S3 to recognize the beef cattle behavior.
2. The method for identifying cattle behavior based on video spatiotemporal features according to claim 1, characterized in that: In step 201, the specific process of constructing the video sampling module is as follows: split the cattle behavior video into video frames, input the video sampling module, first sample 30 frames at equal intervals, and form a video frame sequence , calculate the video frame sequence The sparse regular operator of the RGB difference between two adjacent frames is used, and the four groups of adjacent frames with the largest sparse regular operator are selected to form a video frame sequence M, and the video frame sequence M is input into the subsequent feature extraction network; Where: represents a sequence of video frames; Represents each video frame obtained; Indicates the frame number. ; is the width of the input video frame; The height of the input video frame.
3. The method for identifying cattle behavior based on video spatiotemporal features as claimed in claim 1, characterized in that: In step S20201, , , , =256 and =455, is the height of the input video frame, The width of the input video frame.
4. The method for identifying cattle behavior based on video spatiotemporal features according to claim 1, characterized in that: In step S3, the specific process of the beef cattle behavior recognition model training is as follows: when training the beef cattle behavior recognition model, a focal loss function is used to calculate the loss of the model prediction probability and the real label of the video. , back propagation loss , repeat the iteration until the number of iterations reaches the preset initial value to complete the training; Where: Indicates the weights of different categories; Represents the class probability predicted by the model; It is used to suppress the loss contribution of simple samples and promote the loss contribution of difficult samples.