Three-flow convolutional network-based giraffe behavior identification method
By using a three-stream convolutional network for giraffe behavior recognition, combined with multi-target tracking and global feature modeling, the problem of local detail and global context modeling in existing giraffe behavior recognition technologies is solved, achieving high-precision individual behavior recognition and micro-motion recognition.
Patent Information
- Application Number
- CN202511046319.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to simultaneously model both local detailed movements and global contextual behaviors of giraffes in complex scenarios, and lack individual ID tracking capabilities, resulting in insufficient accuracy and poor robustness in behavior recognition. This is especially true in wild environments where multiple giraffes coexist and pose occlusion is frequent, where micro-movement recognition accuracy is inadequate.
A three-stream convolutional network is used for video target detection and pose estimation. A multi-target tracking algorithm is combined for cross-frame tracking to generate different types of video segments. Multi-scale behavioral features are extracted through a global feature modeling network, and adaptive feature fusion is performed using a gated attention fusion module. Finally, the behavior category recognition result is output through a probabilistic model.
It achieves efficient and accurate recognition of individual giraffe behaviors in complex scenarios, improves the accuracy of micro-movement recognition, especially the recognition of stereotypical behaviors such as licking trees, and has the ability to perceive multiple spatiotemporal scales and continuously recognize individuals.
Smart Images

Figure CN120976970A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic animal behavior recognition, and in particular to a giraffe behavior recognition method based on a three-stream convolutional network. BACKGROUND
[0002] As an important herbivorous animal, the stereotyped behavior of giraffes reflects their health and welfare level. The automatic recognition of their daily behavior is of great significance in wild animal ecological monitoring, ethological research and zoo intelligent management. Current mainstream behavior recognition methods mostly rely on single-stream convolutional neural networks or RNN structures, which are difficult to model local detailed actions and global contextual behaviors at the same time. Moreover, most of these methods do not have individual ID tracking capabilities, leading to chaotic behavior attribution and inability to continuously track behavior changes. In particular, in the wild environment where multiple giraffes coexist and posture occlusion is frequent, the existing algorithms have insufficient recognition accuracy for micro-motions such as licking trees and other stereotyped behaviors, and have poor robustness. Therefore, there is an urgent need for a behavior recognition framework that has multi-temporal and spatial scale perception capabilities, can combine local behavior modeling of key parts, and supports continuous identification of individuals. SUMMARY
[0003] To solve the above problems, the present application proposes a giraffe behavior recognition method based on a three-stream convolutional network, aiming to achieve efficient and accurate recognition of giraffe individual daily behavior in complex scenarios.
[0004] To solve the above technical problems, the technical solution of the present application is as follows:
[0005] A giraffe behavior recognition method based on a three-stream convolutional network, comprising the following steps:
[0006] Collecting videos containing giraffes in outdoor scenes;
[0007] Using a convolutional neural network to perform real-time target detection based on giraffe individuals for each frame of image in the video, outputting the corresponding bounding boxes of all giraffes, each bounding box being labeled with a unique ID, and performing posture estimation within the range of each bounding box to extract the key feature points of the current giraffe and the coordinates corresponding to the key feature points;
[0008] Based on the associated bounding boxes, key feature points and coordinates, using a multi-target tracking algorithm combined with skeleton posture similarity to process the video, and tracking each bounding box across frames;
[0009] For each tracked bounding box in the video, a video subset containing three different types of sub-clips is generated, wherein the first sub-clip only includes the pose changes of the current bounding box, the second sub-clip is the video clip after frame rate reduction of the first sub-clip, and the third sub-clip only includes the region expanded from the mouth key point extracted from the key feature points;
[0010] The first sub-clip, the second sub-clip and the third sub-clip are respectively input into different global feature modeling networks to obtain corresponding self-attention spatio-temporal feature maps, which are dimensionally reduced to form three-dimensional behavior feature vectors;
[0011] The three-dimensional behavior feature vectors are input into a gated attention fusion module to dynamically adjust the weight proportion of each behavior feature vector, adaptively fuse the features for different behavior types, and output a fused feature vector;
[0012] The fused feature vector is input into a fully connected layer, and a behavior class recognition result corresponding to the current video subset is output through a probability model.
[0013] The beneficial effects of the present application are as follows: by using a convolutional neural network to perform target detection and pose estimation on each frame of image in the video, the bounding box and key feature point coordinates of each giraffe are obtained, then a multi-target tracking algorithm is used in combination with skeleton pose similarity to realize cross-frame tracking and ID maintenance of individuals, for each tracked giraffe individual, a regular frame rate whole body video clip, a frame reduction clip, and a local action clip centered on the mouth are extracted, the three types of clips are respectively input into a global feature modeling network to extract multi-scale behavior feature vectors, adaptive feature fusion is realized through a gated attention fusion module, and finally a typical behavior is output through a fully connected layer and a probability model, realizing high-precision giraffe behavior recognition. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 The flowchart of the giraffe behavior recognition method based on a three-stream convolutional network disclosed in the embodiments of the present application is shown. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present application clearer and more explicit, the content of the present application will be further described in detail below in combination with the drawings and specific embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that only the parts related to the present application are shown in the drawings for convenience of description, but not all the contents.
[0016] The embodiments of the present application propose a giraffe behavior recognition method based on a three-stream convolutional network, as shown in Figure 1As shown, comprising the following steps:
[0017] Step 1, collect the video containing giraffes in outdoor scenes.
[0018] In this embodiment 1, first determine the acquired video needs to be set in the area where giraffes frequently move, ensure that the video stream contains a large number of giraffe individuals, ensure sufficient data quantity, and serve as the basis for subsequent optimization, and the acquired video needs to be determined in advance The frame rate serves as the basis for subsequent sub-fragment parameter setting, see the following steps.
[0019] Step 2, using a convolutional neural network to perform real-time target detection on each frame of image in the video based on giraffe individuals, output all giraffe corresponding bounding boxes, each bounding box is marked with a unique ID, and pose estimation is performed within the range of each bounding box, key feature points of the current giraffe are extracted, and coordinates corresponding to the key feature points are extracted.
[0020] In an example, the extraction process of the coordinates corresponding to the key feature points includes:
[0021] For each frame of image I t (a frame of image of the video at time t), the YOLOvX-Pose network is used to detect the position of the giraffe individual, and the corresponding bounding box is generated;
[0022] Get the bounding box of each giraffe individual Wherein, are the upper left corner coordinates of the bounding box, are the width and height of the bounding box, respectively;
[0023] Extract the key feature points of the current giraffe from the bounding box, and the set of key feature points is represented as represents the K key feature point coordinate set of the i giraffe in the t frame, wherein, represents the coordinate of the k key feature point of the i giraffe in the t frame, and the key feature point at least includes the mouth, head, shoulder, knee and tail.
[0024] The YOLOvX-Pose network used in step 2 of the present scheme can perform real-time target detection on each frame of video, outputting the position of the bounding box of each giraffe. After marking each bounding box with a unique ID, each giraffe is assigned a unique bounding box as the basis for subsequent cross-frame tracking. Meanwhile, the network includes a pose estimation branch that can extract multiple key feature points including the mouth, head, shoulder, knee, and tail, as well as their coordinates, as the data basis for subsequent step 3. At the same time, the network structure of the YOLOvX-Pose network also supports multi-target simultaneous detection, with high precision and fast inference ability, suitable for complex backgrounds of dynamic outdoor shooting videos. Each frame of the input video is detected, and the bounding box and key feature point coordinates of each giraffe are output.
[0025] Step 3: Based on the associated bounding box, key feature points, and coordinates, a multi-target tracking algorithm is used in combination with skeleton pose similarity to process the video and track each bounding box across frames.
[0026] In an example, tracking each bounding box across frames includes:
[0027] Taking giraffe individuals as the detection target, the OC-SORT algorithm is used to associate across frames based on target appearance features. On the basis of Kalman filtering and Hungarian matching, a skeleton vector sequence is constructed. For each giraffe i, the skeleton structure is defined as:
[0028]
[0029] In the formula, is the skeleton vector of the kth key feature point and the lth key feature point of the ith giraffe, (k, l) ∈ ε, and ε is the connection relationship between key feature points. It can be understood that is defined as and is the difference between the two key feature point coordinates k and l.
[0030] The matching similarity of the skeleton structure is defined as the weighted cosine angle similarity:
[0031]
[0032] In the formula, is the skeleton structure vector set of the ith giraffe in the tth frame, represents the angle between the corresponding skeleton vectors of the ith individual and the jth individual in the tth frame and the t-1th frame;
[0033] The matching similarity is used as the weighting factor for unique ID matching.
[0034] The step 3 of the scheme is based on the target detection and key feature point recognition of step 2, and further introduces the OC-SORT algorithm to realize the cross-frame ID tracking of giraffe individuals. Specifically, the method data correlates the target detection result and the target appearance feature, and adds the skeleton structure similarity calculation based on the key points in the matching process. By analyzing the angle similarity of the skeleton posture vector of each giraffe, the matching accuracy under occlusion and rapid posture change is improved. This strategy effectively prevents ID switching when multiple giraffes interact, and by constructing the skeleton structure vector and using the similarity of key feature points to participate in the matching weight calculation, the stable ID of the target individual between different frames is realized.
[0035] Step 4, for the historical trajectory of each tracked bounding box in the video, a video subset containing three different types of sub-clips (clip clips) is generated, wherein the first sub-clip only includes the posture change of the current bounding box, the second sub-clip is the video clip after reducing the frame rate of the first sub-clip, and the third sub-clip only includes the region expanded from the mouth key point extracted from the key feature points.
[0036] Specifically, the first sub-clip (whole body video clip with regular frame rate) includes a video clip of a preset length (such as 5 seconds) extracted from the video at a regular frame rate (such as 25 fps), and the posture change within the current bounding box is retained, which is used to separate the motion dynamic features of the giraffe in a short time.
[0037] The second sub-clip (reduced frame clip) is a video clip after reducing the frame rate of the first sub-clip by a preset ratio, and the video length is the same as the first sub-clip. Taking the frame rate of the first sub-clip as 25fps as an example, the frame rate is reduced to 1 / 5, i.e. 5fps, as the second sub-clip.
[0038] The third sub-clip (mouth-centered local action clip) is based on the key feature points and the corresponding coordinates, and the mouth key point is extracted from the key feature points. In a preset manner, a new mouth local area video is expanded, and the frame rate and video length are the same as the second sub-clip. In an example, based on the coordinates of the mouth key point detected in step 2, the mouth key point is taken as the center, and the same Euclidean distance d is extended upward, downward, left and right to crop, forming a new mouth local area video. The frame rate is the same as the second sub-clip, and is named as the third sub-clip, which is particularly used to extract the mouth local micro-motion features such as tree licking and chewing, and belongs to the key step of giraffe behavior recognition in the scheme.
[0039] Step 5, input the first sub-clip, the second sub-clip and the third sub-clip into different global feature modeling networks respectively to obtain corresponding self-attention spatio-temporal feature maps. After dimension reduction processing, three-dimensional behavior feature vectors are formed.
[0040] In an example, the method for calculating the behavior feature vector comprises:
[0041] The first sub-fragment, the second sub-fragment and the third sub-fragment are respectively input into different ViT networks, each ViT network divides the input fragment into a time dimension * a spatial fragment, and extracts a self-attention spatio-temporal feature map corresponding to different fragments D is an embedding dimension, and after dimension reduction processing of the self-attention spatio-temporal feature map, a behavior feature vector of three dimensions is formed
[0042] Step 6: input the behavior feature vector of three dimensions into a gated attention fusion module, dynamically adjust the weight proportion of each behavior feature vector, adaptively fuse features for different behavior types, and output a fused feature vector.
[0043] In an example, the calculation process of the fused feature vector comprises:
[0044] The behavior feature vector of three dimensions is input into a gated attention fusion module, and a gating coefficient is defined as α A , α B , α C satisfy the following conditions:
[0045] α A +α B +α C = 1;
[0046]
[0047] In the formula, σ(g) represents a sigmoid function, W i and b j are the jth branch parameters that can be learned in the weight calculation of the GAF module;
[0048] The weight proportion of each behavior feature vector is dynamically adjusted, adaptive feature fusion is performed for different behavior types, and the fused feature vector is output:
[0049]
[0050] and are behavior feature vectors obtained after dimension reduction processing of the self-attention spatio-temporal feature maps F A , F B and F C .
[0051] In step 6 of the scheme, the Gated Attention Fusion (GAF) learns the optimal feature weight combination mode corresponding to each behavior, dynamically adjusts the weight proportion of the 3-flow behavior feature vector through the gating mechanism, thereby realizing adaptive feature fusion for different behavior types, and obtaining a fusion feature vector.
[0052] In step 7, the fusion feature vector is sent to the full connection layer, and the behavior category recognition result corresponding to the current video subset is output through the probability model.
[0053] The prediction process of the behavior category recognition result includes:
[0054] The fusion feature vector is input into the neural network structure composed of the full connection layer and the Softmax classifier. In the neural network structure composed of the full connection layer and the Softmax classifier, the output behavior prediction vector y is set as The element y k represents the probability of belonging to the following 5 types of behaviors: y1 is walking, y2 is standing, y3 is tree licking, y4 is eating, and y5 is running. The final output behavior recognition result is represented as c = argmax (y k ).
[0055] In summary, the recognition scheme has the following beneficial effects:
[0056] (1) Unlike the traditional SORT method which only relies on the target center displacement to evaluate the next frame position of the target, the inter-frame multi-target tracking method based on key feature point detection is adopted, so that giraffe tracking can still be performed using Kalman filtering in the case of occlusion or non-linear motion.
[0057] (2) The traditional multi-flow behavior recognition network often inputs two spatial domain identical video features into the network without specially extracting the features of the specific behavior expression region, which makes it difficult to extract the space-time features of the giraffe tree licking micro-behavior. The introduction of the feature extraction flow of the giraffe mouth region in the three-flow ViT network of the present application enables the network to more meticulously capture the giraffe tree licking stereotyped behavior, and reduces the frame rate and the giraffe mouth small spatial domain without significantly increasing the overall network operation amount.
[0058] (3) The introduced gated attention fusion module enables the network to dynamically and adaptively adjust the weight of the three-stream network to extract features, and compared with the static fusion strategy, the method improves the recognition accuracy of multiple types of behaviors, especially the "lick tree" type micro-behavior. The above embodiments are only for illustrating the technical concept and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and to implement it, and cannot limit the protection scope of the present application. Any equivalent changes or modifications made according to the essence of the present application shall be covered within the protection scope of the present application.
Claims
1. A giraffe behavior recognition method based on a three-stream convolutional network, characterized in that, Includes the following steps: Collect videos featuring giraffes in outdoor settings; A convolutional neural network is used to perform real-time target detection based on individual giraffes in each frame of the video, outputting bounding boxes corresponding to all giraffes. Each bounding box is labeled with a unique ID, and pose estimation is performed within the range of each bounding box location to extract the key feature points of the current giraffe and the coordinates corresponding to the key feature points. Based on the associated bounding boxes, key feature points, and coordinates, a multi-target tracking algorithm combined with skeleton pose similarity is used to process the video, and cross-frame tracking is performed on each bounding box; For the historical trajectory of each tracked bounding box in the video, a video subset containing three different types of sub-segments is generated. The first sub-segment includes only the pose change of the current bounding box, the second sub-segment is a video segment with reduced frame rate of the first sub-segment, and the third sub-segment includes only the region expanded from the mouth key points extracted separately from the key feature points. The first sub-segment, the second sub-segment, and the third sub-segment are respectively input into different global feature modeling networks to obtain corresponding self-attention spatiotemporal feature maps. After dimensionality reduction processing, the self-attention spatiotemporal feature maps form three-dimensional behavioral feature vectors. The three-dimensional behavioral feature vectors are input into the gated attention fusion module, the weight ratio of each behavioral feature vector is dynamically adjusted, adaptive feature fusion is performed for different behavior types, and the fused feature vector is output. The fused feature vector is fed into a fully connected layer, and the behavior category recognition result corresponding to the current video subset is output through a probability model.
2. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 1, characterized in that, The first sub-segment includes a video segment of a preset length extracted from the video at a normal frame rate, while retaining the pose changes within the current bounding box.
3. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 2, characterized in that, The second sub-segment is a video segment whose frame rate is reduced by a preset ratio after the first sub-segment is completed, and the video length is the same as that of the first sub-segment.
4. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 1, characterized in that, The third sub-segment extracts key mouth points from the key feature points based on the key feature points and the corresponding coordinates, and expands them into a new local mouth region video according to a preset method. The frame rate and video length are the same as those of the second sub-segment.
5. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 1, characterized in that, The process of extracting the coordinates of the key feature points includes: For each frame image I t The YOLOvX-Pose network was used to detect the location of individual giraffes and generate corresponding bounding boxes. Obtain the bounding box of each individual giraffe. in, These are the coordinates of the top left corner of the bounding box. These are the width and height of the bounding box, respectively; Extract key feature points of the current giraffe from the bounding box, where the set of key feature points is represented as follows: in, This represents the coordinates of the k-th key feature point of the i-th giraffe in frame t. Key feature points include at least the mouth, head, shoulders, knees, and tail.
6. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 5, characterized in that, The cross-frame tracking for each of the bounding boxes includes: Using individual giraffes as detection targets, the OC-SORT algorithm is employed for cross-frame association based on target appearance features. Building upon Kalman filtering and Hungarian matching, a skeleton vector sequence is constructed. For each giraffe i, the skeleton structure is defined as follows: In the formula, Let be the skeleton vector of the k-th key feature point and the l-th key feature point of the i-th giraffe, (k,l)∈ε, where ε is the connection relationship between the key feature points; The matching similarity of the skeleton structure is defined as the weighted cosine similarity: In the formula, Let be the set of skeletal structure vectors of the i-th giraffe in frame t. This represents the angle between the skeleton vectors of the i-th individual and the j-th individual in frame t and frame t-1. The matching similarity is used as a weighting factor for unique ID matching.
7. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 1, characterized in that, The method for calculating the behavioral feature vector includes: The first sub-segment, the second sub-segment, and the third sub-segment are each input into different ViT networks. Each ViT network divides the input segment into a temporal dimension multiplied by a spatial segment, and extracts the self-attention spatiotemporal feature maps corresponding to different segments. D represents the embedding dimension. After dimensionality reduction processing of the self-attention spatiotemporal feature map, a three-dimensional behavioral feature vector is formed.
8. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 7, characterized in that, The calculation process of the fused feature vector includes: The three-dimensional behavioral feature vectors are input into the gated attention fusion module, and the gate coefficient is defined as α. A α B α C It meets the following conditions: α A +α B +α C =1; In the formula, σ(g) represents the sigmoid function, and W i and b i All of these are learnable branch parameters calculated in the GAF module; The weight ratio of each behavior feature vector is dynamically adjusted, and adaptive feature fusion is performed for different behavior types to output the fused feature vector: and These are the self-attention spatiotemporal feature maps F A F B and F C The behavioral feature vector obtained after dimensionality reduction.
9. The giraffe behavior recognition method based on a three-stream convolutional network as described in claim 8, characterized in that, The prediction process for the behavior category identification results includes: fuse feature vectors In a neural network structure consisting of a fully connected layer and a Softmax classifier, the output behavior prediction vector is set. element y k Let y1 represent the probability of belonging to the following 5 behavior categories: y1 walking, y2 standing, y3 licking a tree, y4 eating, and y5 running. The final output of the behavior recognition result is represented as c = arg max(y k ).