Wild animal fine-grained behavior recognition method based on text supplement state
By combining the skeleton encoder and text encoder in the behavior recognition network, and using three-level text description and contrastive learning, we solve the problems of data subjectivity and high computational resource consumption in fine-grained behavior recognition of wild animals, and achieve more accurate and efficient behavior recognition.
Patent Information
- Application Number
- CN202510733693.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies for fine-grained behavior recognition of wild animals have problems such as strong data subjectivity, device implantation affecting behavior, large consumption of computing resources, limited sample size, and poor model generalization performance, making it difficult to accurately capture the subtle behavioral characteristics of wild golden monkeys.
A method based on text supplementation state is adopted. By combining the skeleton encoder and the text encoder, the action recognition network is trained using three-level text description and contrastive learning, the domain prior knowledge is integrated, and the fine-grained semantic information is distilled to the skeleton encoder to achieve the retention of multi-granularity text description.
It improves the accuracy and robustness of behavior recognition, reduces computational costs, enhances the ability to recognize fine-grained behaviors, reduces dependence on text encoders, and achieves more efficient behavior recognition.
Smart Images

Figure CN120599700A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of identifying biological behaviors in videos, and in particular relates to a method for identifying fine-grained behaviors of wild animals based on text supplementation status. Background Art
[0002] As an endangered primate species, fine-grained behavioral recognition of wild golden snub-nosed monkeys is crucial for studying their social structure, ecological adaptability, and developing scientific conservation strategies. Traditional manual observation methods are not only inefficient but also inherently flawed, with data subjectivity. In complex outdoor environments, factors such as varying lighting, obstruction by vegetation, and the uncertainty of animal activity make visual-based behavioral recognition a significant technical challenge. Accurately capturing subtle behavioral characteristics of golden snub-nosed monkeys, particularly under non-cooperative sampling conditions, is a key technical challenge that needs to be addressed.
[0003] Existing behavior recognition technologies fall into three main categories. The first category consists of traditional methods based on manual observation. While simple to implement, these methods rely heavily on the observer's expertise, suffer from inherent flaws such as high subjectivity and poor data consistency, and are difficult to achieve long-term continuous monitoring. The second category consists of contact-based monitoring technologies based on physical sensors, including inertial measurement units and GPS trackers. While these methods can obtain accurate motion parameters, the implantation of these devices can cause physical and psychological stress on the animals, altering their natural behavior patterns. Furthermore, these technologies present practical challenges such as difficult maintenance and limited battery life. The third category consists of non-contact monitoring technologies based on computer vision. For example, methods that combine object detection algorithms with 3D pose estimation achieve automated recognition, but they still face numerous technical challenges when applied to wild animals such as golden snub-nosed monkeys. Static appearance features are poorly adaptable to complex outdoor environments. While methods based on time series modeling can capture dynamic behavioral features, they suffer from technical bottlenecks such as high computational resource consumption and difficulty modeling long-term dependencies. Furthermore, these deep learning methods generally require large amounts of labeled data for training. However, the difficulty in collecting behavioral data on golden snub-nosed monkeys results in a limited sample size, making the models prone to overfitting and generalization performance difficult to guarantee. Especially when dealing with fine-grained behavior recognition tasks, existing methods are obviously insufficient in their ability to distinguish similar behaviors, making it difficult to meet the actual needs of scientific research and conservation work.
[0004] To address the above technical bottlenecks, a golden monkey behavior recognition method is needed that can make up for the shortcomings of visual data in fine-grained distinction, alleviate the difficulty of small sample learning, and improve the interpretability of the model. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to address the deficiencies in the above-mentioned prior art and provide a method for fine-grained behavior recognition of wild animals based on text supplementation status. The method has a simple structure and a reasonable design. The behavior recognition network includes a skeleton encoder and a text encoder, so that the behavior recognition network integrates domain prior knowledge on the basis of data-driven learning; the behavior recognition network is trained based on contrastive learning, and the fine-grained semantic information of the three-level text description is "distilled" to the skeleton encoder, so that the skeleton encoder still retains multi-granularity text description information during single-modal reasoning.
[0006] To solve the above technical problems, the present invention adopts a technical solution: a method for fine-grained wildlife behavior recognition based on text supplementation status, characterized by comprising the following steps:
[0007] Step 1: Data collection: Obtain the golden monkey video dataset D and divide the video dataset D into a training set and a test set;
[0008] Step 2: Divide the body of the golden monkey into K parts according to the partitioning strategy, where K is a positive integer not less than 2;
[0009] Step 3: Generate text description: Create a three-level text description T for the behavior of the golden monkey, including a global text description, a part text description, and a synonym extended text description;
[0010] Step 4: Build an action recognition network: The action recognition network includes a skeleton encoder and a text encoder; train and update the skeleton encoder;
[0011] Step 5: Train the action recognition network based on contrastive learning: Input the training set and text description into the skeleton encoder and text encoder respectively to obtain the training skeleton features and text features, calculate the global contrast loss and part contrast loss, and optimize and update the skeleton encoder and text encoder;
[0012] Step 6: Test the behavior recognition network:
[0013] The test set is input into the trained skeleton encoder for feature extraction and behavior classification;
[0014] Step 7: Apply the behavior recognition network:
[0015] Videos of golden snub-nosed monkeys are collected and fed into a behavior recognition network, which then outputs fine-grained behavior categories and confidence levels of the monkeys.
[0016] The above-mentioned method for identifying fine-grained behavior of wild animals based on text supplementation status is characterized in that the fine-grained behavior categories include basic behavior categories and sub-category behavior categories.
[0017] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplementation status is characterized in that the specific method of training the behavior recognition network based on contrastive learning in step 5 is:
[0018] Step 501: construct global positive and negative sample pairs and calculate the global contrast loss;
[0019] Step 502: construct positive and negative sample pairs of parts and calculate the part contrast loss;
[0020] Step 503: weighted fusion of global contrast loss and local contrast loss to obtain total loss;
[0021] Step 504: Calculate the gradient based on the total loss, and optimize and update the skeleton encoder and the text encoder simultaneously through back propagation;
[0022] Step 505: Iterate and update until a stop condition is reached, otherwise repeat steps 501 to 504.
[0023] The above-mentioned method for fine-grained wildlife behavior recognition based on text supplementation status is characterized in that the specific steps of constructing global positive and negative sample pairs in step 501 are:
[0024] Step 5011: obtain a training skeleton sequence of the golden monkey based on the training set, input the training skeleton sequence to a skeleton encoder, the skeleton encoder outputs training skeleton features, and the training skeleton features are subjected to global average pooling to obtain global skeleton features;
[0025] Step 5012: The global text description is input into a text encoder for semantic encoding to obtain global text features.
[0026] Step 5013: The global skeleton features and global text features of the same behavior constitute a global positive sample pair, and the global skeleton features and global text features of different behaviors constitute a global negative sample pair.
[0027] The above-mentioned method for fine-grained wildlife behavior recognition based on text supplementation status is characterized in that the specific steps of constructing the positive and negative sample pairs of parts in step 502 are:
[0028] Step 5021: A training skeleton sequence of the golden monkey is obtained based on the training set, and the training skeleton sequence is input into a skeleton encoder. The skeleton encoder outputs training skeleton features, and the training skeleton features are subjected to local average pooling to obtain part skeleton features of K parts.
[0029] Step 5022: Input the part text description into the text encoder to obtain part text features of K parts;
[0030] Step 5023: The part skeleton features and part text features of the same behavior constitute a part positive sample pair, and the part skeleton features and part text features of different behaviors constitute a part negative sample pair, and the part contrast loss is calculated.
[0031] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplementation status is characterized in that the part contrast loss is obtained by weighted calculation of the contrast loss of each sub-part.
[0032] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplementation status is characterized in that: in the iterative training of the behavior recognition network based on contrastive learning, the synonym extended text description is alternately input into the text encoder with the corresponding global text description in a random replacement manner to obtain global text features.
[0033] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplement status is characterized in that: the specific method of training and updating the skeleton encoder in step four includes: obtaining a training skeleton sequence of golden monkeys based on the training set, inputting the training skeleton sequence into the skeleton encoder, the skeleton encoder outputting training skeleton features, predicting the behavior probability distribution through the classifier, calculating the gradient based on the cross-entropy loss function, and updating the parameters of the skeleton encoder.
[0034] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplement status is characterized in that: the specific method for testing the behavior recognition network in step six is: obtaining a test skeleton sequence of the golden monkey based on the test set, inputting the test skeleton sequence into the skeleton encoder trained in step five, the skeleton encoder outputs the test skeleton features, and the test skeleton features are passed through the classifier to obtain the predicted behavior probability; taking the category with the highest probability as the prediction result, calculating the accuracy, and obtaining the accuracy of the behavior recognition network.
[0035] The above-mentioned method for fine-grained behavior recognition of wild animals based on text supplementation status is characterized in that: the skeleton encoder is composed of GC-MTC modules, each GC-MTC module contains a graph convolution GC layer and a multi-scale temporal convolution MTC module, and the multi-scale temporal convolution MTC module includes: 1×1 convolution branch, temporal convolution branch with an expansion rate of 1, temporal convolution branch with an expansion rate of 2, and MaxPool branch.
[0036] Compared with the prior art, the present invention has the following advantages:
[0037] 1. The present invention has a simple structure, reasonable design, and is easy to implement and operate.
[0038] 2. The present invention adopts a three-level text description including global text description, part text description and synonym extended text description, which work together to construct a fine-grained semantic representation.
[0039] 3. The behavior recognition network of the present invention includes a skeleton encoder and a text encoder, which enables the behavior recognition network to integrate domain prior knowledge on the basis of data-driven learning, thereby enhancing the ability to characterize the multi-level characteristics of golden monkey behavior and improving the behavior recognition network's robust understanding of professional terms and their variant expressions, ultimately achieving more accurate fine-grained behavior recognition.
[0040] 4. The present invention trains an action recognition network based on contrastive learning, and "distills" the fine-grained semantic information of the three-level text description to the skeleton encoder, so that the skeleton encoder still retains multi-granularity text description information during single-modal reasoning.
[0041] 5. The testing process completely removes the text encoder and runs the skeleton encoder independently, making the computational cost of the inference phase comparable to that of traditional single-modal behavior recognition methods, eliminating the real-time dependence on the text encoder while retaining the representation optimization effect obtained by multimodal contrastive learning in the training phase.
[0042] In summary, the present invention has a simple structure and a reasonable design. The behavior recognition network includes a skeleton encoder and a text encoder, so that the behavior recognition network can integrate domain prior knowledge on the basis of data-driven learning; the behavior recognition network is trained based on contrastive learning, and the fine-grained semantic information of the three-level text description is "distilled" to the skeleton encoder, so that the skeleton encoder still retains multi-granularity text description information during single-modal reasoning.
[0043] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Flow chart of the method of the present invention.
[0045] Figure 2 Schematic diagram of the body division strategy of the present invention.
[0046] Figure 3 A schematic diagram of the global text description of the present invention.
[0047] Figure 4 This is a flow chart of the present invention's contrastive learning-based behavior recognition network training.
[0048] Figure 5 This is a training flowchart of the behavior recognition network of the present invention.
[0049] Figure 6 This is a test flow chart of the behavior recognition network of the present invention. DETAILED DESCRIPTION
[0050] The method of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments of the present invention.
[0051] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0054] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used herein to describe the spatial positional relationship of a device or feature to other devices or features as shown in the figures. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figures. For example, if the device in the drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.
[0055] like Figure 1 As shown, the present invention provides a method for fine-grained wildlife behavior recognition based on text supplementation status, comprising the following steps:
[0056] Step 1: Data collection: Obtain the golden monkey video dataset D and divide the video dataset D into a training set and a test set.
[0057] A golden monkey video dataset D is obtained. The video dataset D includes multiple golden monkey video samples. The multiple golden monkey video samples are divided into a training set and a test set according to a preset ratio, for example, 8:2.
[0058] Using a computer vision-based skeleton keypoint detection algorithm, such as the OpenPose algorithm, each frame of the golden monkey video sample is processed to identify pre-defined joints on the monkey's body, resulting in a joint point set V. Based on prior knowledge of the monkey's skeletal structure, after determining the keypoint set V, keypoint pairs with skeletal connections are recorded. These keypoint pairs form a skeletal edge set E. For example, if a skeletal connection exists between joint points A and B, then (A, B) will be recorded in the skeletal edge set E.
[0059] Get the three-dimensional coordinate information of the joint points, arrange the coordinate information in chronological order and batches, and get the golden monkey skeleton sequence S, S∈R B×3×N×Ts , B represents the batch size, that is, the number of video samples input for each training, 3 represents the dimension of the coordinates (x, y, z) of the joint points in the image coordinate system, N represents the number of joint points in each frame, and Ts represents the sequence length, that is, the number of time steps, which represents the number of consecutive frames.
[0060] Step 2: Divide the golden snub-nosed monkey's body into K parts according to a partitioning strategy, where K is a positive integer not less than 2. It should be noted that the partitioning strategy can be binary, quartered, or sextupled. The binary partitioning strategy divides the golden snub-nosed monkey's body into upper and lower limbs; the quartered partitioning strategy divides the golden snub-nosed monkey's body into upper and lower limbs, head, and trunk; and the sextupled partitioning strategy further focuses on the two sides of the golden snub-nosed monkey's body, dividing it into upper and lower limbs, head, trunk, left side, and right side.
[0061] like Figure 2 As shown, in this embodiment, the quartering method is adopted, that is, K=4.
[0062] Step 3: Generate text description: Create a three-level text description T for the behavior of the golden monkey, including a global text description, a part text description, and a synonym extended text description.
[0063] In practice, a large language model (LLM) was used to construct a three-level text description T for the behavior of golden snub-nosed monkeys. A byte-pair encoding tokenizer based on the CLIP model was used to tokenize the text description and convert it into word vectors.
[0064] Taking the "feeding" behavior of golden monkeys as an example, explain the third-level text description.
[0065] Among them, the global text description refers to the basic behavior category, that is, the label, which is generated by life science experts based on behavioral observations. It only contains basic semantic information and is used to define the macro category of behavior; the global text description is "feeding".
[0066] like Figure 3 As shown in the figure, the labels of each column of golden monkey video frames from left to right are: eating, grooming, walking, climbing trees and mating.
[0067] Part text descriptions refer to the subcategories of behavior for each part, such as "left forelimb grasping fruit from a branch," "right forelimb assisting in securing and adjusting the fruit's position," "mouth opening and closing to chew," and "head tilting downward to approach food." Based on ethological knowledge, local changes in golden snub-nosed monkey body parts are described and independent text descriptions are generated for multi-part comparative learning, enhancing the skeleton encoder's sensitivity to local features.
[0068] The synonym-extended text description enriches the semantic representation by adding temporal, spatial and dynamic details (such as "the left forelimb quickly grabs the fruit 1.2 meters above the ground at a 45-degree angle for 0.8 seconds", "the right forelimb applies slight force to fix the fruit and adjusts it to the front of the mouth, with a contact time of 0.5 seconds", "the lower jaw opens and closes quickly three times to complete chewing", "the head is tilted down 20 degrees to accurately locate the fruit, and the neck muscles remain tense"), providing the text encoder with contextually relevant spatiotemporal information, thereby improving the model's recognition accuracy of complex feeding behaviors.
[0069] The present invention adopts a large language model to generate a three-level text description including global text description, part text description and synonym extended text description. The global text description defines the overall behavior pattern, the part text description refines the local action features, and the synonym extended text description covers the semantic variant expression and is used to randomly replace the global text description. The three work together to construct a fine-grained semantic representation.
[0070] Step 4: Build an action recognition network: The action recognition network includes a skeleton encoder and a text encoder; train and update the skeleton encoder E s .
[0071] Skeleton Encoder E s The system consists of multiple GC-MTC modules, each of which includes a graph convolution (GC) layer and a multi-scale temporal convolution (MTC) module. The multi-scale temporal convolution (MTC) module consists of four parallel branches: a 1×1 convolution branch, a temporal convolution branch with a dilation rate of 1, a temporal convolution branch with a dilation rate of 2, and a MaxPool branch. These four branches, in order, preserve the original temporal features, extract macroscopic features of motion patterns, capture high-frequency actions in fine-grained behaviors, and identify low-frequency, large-scale motion.
[0072] Skeleton Encoder E s The feature processing flow can be divided into two core stages: graph structure modeling and multi-scale temporal feature extraction.
[0073] In the graph convolution GC layer, the golden monkey skeleton sequence S is input, and the skeleton is modeled as a graph G = [V, E]. The spatial features of the joint points are aggregated through the graph convolution operation to capture the spatial dependencies of the skeleton and output a feature matrix that integrates the spatial structure information.
[0074] The multi-scale temporal convolution (MTC) module takes as input the output features of the graph convolution (GC) layer. For each branch, the features are first processed in the time dimension to facilitate the temporal convolution operation. Then, after the four branches are independently operated, the output feature dimensions remain consistent. The outputs of the four branches are concatenated in the channel dimension to obtain a matrix that integrates the multi-scale temporal features, which is the skeleton feature.
[0075] Train and update the skeleton encoder E s The specific method includes: obtaining a training skeleton sequence S of the golden monkey based on the training set, inputting the training skeleton sequence S into the skeleton encoder E s , skeleton encoder E s Output training skeleton feature E s (S), predicting the behavior probability distribution through the classifier, based on the cross entropy loss function L cls (E s (S)) Calculate the gradient and update the skeleton encoder E s The parameters are updated iteratively until the iteration stopping condition is met.
[0076] Cross Entropy Loss Function Among them, C represents the type of behavior category, B represents the batch size, that is, the number of samples input for each training, and y ij Indicates the true label of the jth sample belonging to the i-th behavior, S j Represents the skeleton encoder E s The global skeleton feature of the j-th sample output, P i (S j ) represents the predicted probability that the j-th sample belongs to the i-th behavior, 1≤j≤B.
[0077] Skeleton features offer significant advantages in capturing posture changes and modeling temporal dynamics. However, they are unable to represent biological attributes and, therefore, cannot capture key discriminative information such as age and gender. Text features, on the other hand, can proactively inject prior knowledge into the recognition model, enabling the definition of social relationship constraints based on knowledge from animal behavior and life sciences. However, text descriptions cannot provide detailed information about posture changes and lack spatial positioning capabilities. Therefore, the behavior recognition network includes a text encoder, enabling it to integrate domain prior knowledge based on data-driven learning. This not only enhances the ability to represent the multi-level characteristics of golden snub-nosed monkey behavior, but also improves the network's robust understanding of specialized terminology and its variant expressions, ultimately achieving more accurate and fine-grained behavior recognition.
[0078] Step 5: Training the behavior recognition network based on contrastive learning: Figure 4 and Figure 5 As shown, the training set and text description are input into the skeleton encoder and text encoder respectively to obtain the training skeleton features and text features, calculate the global contrast loss and part contrast loss, and optimize and update the skeleton encoder E s and text encoder E t The specific method is:
[0079] Step 501: Construct global positive and negative sample pairs and calculate the global contrast loss:
[0080] Step 5011: Get the training skeleton sequence S of the golden monkey based on the training set, and input the training skeleton sequence S to the skeleton encoder E. s , skeleton encoder E s Output training skeleton feature E s (S), training skeleton features E s (S) Obtain global skeleton features through global average pooling;
[0081] Step 5012: The global text description is input into the text encoder for semantic encoding to obtain global text features.
[0082] The text encoder is based on the Transformer. It uses a byte-pair encoding tokenizer based on the CLIP model to tokenize the global text description and convert it into word vectors. The word vectors are then fed into the text encoder, where they are semantically encoded through multiple layers of Transformer blocks. Finally, the output word vector sequences are aggregated to generate fixed-dimensional global text features.
[0083] In the iterative training of the action recognition network based on contrastive learning, the synonym extended text description is input into the text encoder alternately with the corresponding global text description in a random replacement manner to obtain the global text features.
[0084] Step 5013: The global skeleton features and global text features of the same behavior constitute a global positive sample pair (S i ,T i ), S i represents the global skeleton features of the i-th behavior in the sample, T i It represents the global text features belonging to the i-th behavior in the global text description. The global skeleton features and global text features of different behaviors constitute the global negative sample pairs (S i , T u ), calculate the global contrast loss L com 1≤i≤C, 1≤u≤C. j represents the jth behavior category, and C represents the total number of behavior types.
[0085] For example, S i Skeleton sequence features representing “feeding” behavior, T i The global text description of "eating" is represented by the two, which describe the same behavior and constitute a global positive sample pair.
[0086] S i Skeleton sequence features representing “feeding” behavior, T u The global text description of "hug" is represented. The two descriptions do not belong to the same behavior and constitute a global negative sample pair.
[0087] The global positive sample pair and the global negative sample pair constitute the global positive and negative sample pair. The constructed global positive and negative sample pairs are input into the global contrast loss function layer to calculate the global contrast loss L com , E S,T represents the expectation of the skeleton-text pair, KL(·) represents the KL divergence calculation, which measures the difference between two probability distributions. represents the predicted similarity distribution from skeleton to text, sim(S j , T j ) represents the global skeleton feature S i With the global text feature T i The cosine similarity of τ is used to adjust the sharpness of the similarity distribution. i , T u ) represents the global skeleton feature S i With the global text feature T u The cosine similarity of y S2T represents the ground-truth similarity label from skeleton to text, represents the predicted similarity distribution from text to skeleton, sim(T i , S i ) represents the global text feature T i With the global skeleton feature S iThe cosine similarity, sim(T u , S i ) represents the global text feature T u With the global skeleton feature S i The cosine similarity of y T2S Represents the ground-truth similarity labels from text to skeleton.
[0088] By minimizing the global contrast loss L com , so that the similarity of the positive sample pairs approaches 1, the distance between the skeleton features and text features in the positive sample pairs is shortened, the similarity of the negative sample pairs approaches 0, the distance between the skeleton features and text features in the negative sample pairs is extended, and the global skeleton feature S is forced to i With the global text feature T u Semantic alignment in vector space.
[0089] Step 502: Construct positive and negative sample pairs for each part and calculate the part contrast loss:
[0090] Step 5021: Get the training skeleton sequence S of the golden monkey based on the training set, and input the training skeleton sequence S to the skeleton encoder E. s , skeleton encoder E s Output training skeleton feature E s (S), training skeleton features E s (S) Obtain the part skeleton features of K parts through local average pooling.
[0091] In step 5022, the part text descriptions are input into a text encoder to obtain part text features for K parts. A byte-pair encoding tokenizer based on the CLIP model is used to tokenize the part text descriptions and convert them into word vectors. The word vectors are then input into the text encoder, where they are semantically encoded through multiple layers of Transformer blocks. Finally, the output word vector sequences are aggregated to obtain fixed-dimensional part text features.
[0092] Step 5023: The part skeleton features and part text features of the same behavior constitute the part positive sample pairs, and the part skeleton features and part text features of different behaviors constitute the part negative sample pairs. Calculate the part contrast loss L multi .
[0093] For example, K = 4, 1 represents the forelimbs, 2 represents the hind limbs, 3 represents the head, and 4 represents the trunk. i1 represents the key features of the forelimbs in the “feeding” behavior extracted by the encoder, T i1 The encoding features representing "the left forelimb grabs the fruit on the branch" and "the right forelimb assists in fixing and adjusting the position of the fruit" in the text description of the part belong to the description of the same behavior and the same part, are completely matched, and constitute a positive sample pair of the part.
[0094] T u1 The encoding feature of "grasping the hair and combing it with the forelimbs" in the body part description belongs to the grooming behavior, which is consistent with S i1 Descriptions that do not belong to the same behavior constitute part negative sample pairs.
[0095] Part contrast loss L multi It is calculated by weighting the contrast loss of each sub-part.
[0096] The part positive sample pair and the part negative sample pair constitute the part positive and negative sample pair, and the constructed part positive and negative sample pair is input into the part contrast loss function layer. Where K represents the total number of divided parts, L k represents the part contrast loss of the k-th part. 1≤k≤K.
[0097] in S ik represents the skeleton feature of the kth part of the sample belonging to the i-th behavior, T ik The text skeleton feature representing the k-th part of the text description of the i-th behavior in the text description, represents the similarity distribution of part-level predictions from skeleton to text, (S ik , T ik ) represents a positive sample pair, (S ik , T uk ) represents a negative sample pair, sim(S ik , T ik ) represents the part skeleton feature S ik and part text feature T ik The cosine similarity of Represents the similarity distribution of part-level prediction from text to skeleton, sim(T ik , S ik ) represents the part text feature T ik and part skeleton features S ik The cosine similarity of .
[0098] Step 503: Weighted fusion of global contrast loss and part contrast loss to obtain the total loss L total , L total =L com +L multi .
[0099] Step 504: Calculate the gradient based on the total loss and optimize and update the skeleton encoder E through back propagation. s and text encoder E t ;
[0100] Step 505: Iterate and update until a stop condition is reached, otherwise repeat steps 501 to 504.
[0101] Based on the skeleton encoder E s Output skeleton feature E s (S) and text encoder E t Output text features E t (T) Calculate the global contrast loss L com Compared with the part loss L multi , global contrast loss L com Compared with the part loss L multi Used to align skeleton features E s (S) and text features E t (T) is distributed in the vector space, achieving cross-modal semantic alignment. During training, through contrastive learning, the fine-grained semantic information of the three-level text description is "distilled" into the skeleton encoder, allowing the skeleton encoder to retain multi-granular text description information during unimodal reasoning.
[0102] Step 6: Test the behavior recognition network:
[0103] Input the test set into the trained skeleton encoder E s , perform feature extraction and behavior classification.
[0104] like Figure 6 As shown in the figure, the specific method of testing the behavior recognition network is as follows: based on the test set, the test skeleton sequence S of the golden monkey is obtained, and the test skeleton sequence S is input into the skeleton encoder E trained in step 5. s , skeleton encoder E s Output test skeleton feature E s (S), test skeleton feature E s (S) The predicted behavior probability is obtained through the classifier; the category with the highest probability is taken as the prediction result, the accuracy is calculated, and the accuracy of the behavior recognition network is obtained.
[0105] During the training phase, the skeleton encoder learns to encode the semantics of text descriptions into the skeleton feature space through contrastive learning. A global contrastive loss is used for cross-modal alignment, prompting the skeleton encoder to learn discriminative features that embody multi-granular semantics, effectively enhancing the semantic expressiveness of the skeleton features. During testing, the text encoder is completely stripped away, running the skeleton encoder independently. This brings the computational cost of the inference phase to parity with traditional single-modality-based action recognition methods. By leveraging feature decoupling and a pre-stored knowledge base, comprehensive output from coarse-grained classification to fine-grained state supplementation is achieved, eliminating the real-time dependence on the text encoder while retaining the representation optimization effects achieved through multimodal contrastive learning during the training phase.
[0106] Step 7: Apply the behavior recognition network:
[0107] Videos of golden snub-nosed monkeys are collected and fed into a behavior recognition network, which then outputs fine-grained behavior categories and confidence levels of the monkeys.
[0108] Fine-grained behavior categories include basic behavior categories and part behavior categories.
[0109] For example, a video of a golden monkey is fed into a behavior recognition network, and the behavior recognition network outputs:
[0110] Basic behavior categories: hug;
[0111] Part behavior categories: Forelimbs: grabbing fruits, pushing aside branches and leaves; Head: lowering the head to get closer to food, turning towards food; Trunk: leaning forward to get closer to food.
[0112] The above description is merely an embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for fine-grained wildlife behavior recognition based on text supplementation status, characterized by: The following steps are involved: Step 1: Data collection: Obtain the golden monkey video dataset D and divide the video dataset D into a training set and a test set; Step 2: Divide the body of the golden monkey into K parts according to the partitioning strategy, where K is a positive integer not less than 2; Step 3: Generate a three-level text description: Create a three-level text description T for the behavior of the golden monkey, including a global text description, a part text description, and a synonym extended text description; Step 4: Build an action recognition network: The action recognition network includes a skeleton encoder and a text encoder; train and update the skeleton encoder; Step 5: Train the action recognition network based on contrastive learning: Input the training set and text description into the skeleton encoder and text encoder respectively to obtain training skeleton features and text features. The trained skeleton features and text features are input into the action recognition network to output fine-grained action categories. The global contrast loss and part contrast loss are calculated to optimize and update the skeleton encoder and text encoder. Step 6: Test the behavior recognition network: The test set is input into the trained skeleton encoder for feature extraction and behavior classification; Step 7: Apply the behavior recognition network: Videos of golden snub-nosed monkeys are collected and fed into a behavior recognition network, which then outputs fine-grained behavior categories and confidence levels of the monkeys.
2. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 1, characterized in that: Fine-grained behavior categories include basic behavior categories and part behavior categories.
3. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 1, characterized in that: The specific method of training the behavior recognition network based on contrastive learning in step 5 is: Step 501: construct global positive and negative sample pairs and calculate the global contrast loss; Step 502: construct positive and negative sample pairs of parts and calculate the part contrast loss; Step 503: weighted fusion of global contrast loss and local contrast loss to obtain total loss; Step 504: Calculate the gradient based on the total loss, and optimize and update the skeleton encoder and the text encoder simultaneously through back propagation; Step 505: Iterate and update until a stop condition is reached, otherwise repeat steps 501 to 504.
4. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 3 is characterized by: The specific steps of constructing global positive and negative sample pairs in step 501 are: Step 5011: obtain a training skeleton sequence of the golden monkey based on the training set, input the training skeleton sequence to a skeleton encoder, the skeleton encoder outputs training skeleton features, and the training skeleton features are subjected to global average pooling to obtain global skeleton features; Step 5012: The global text description is input into a text encoder for semantic encoding to obtain global text features. Step 5013: The global skeleton features and global text features of the same behavior constitute a global positive sample pair, and the global skeleton features and global text features of different behaviors constitute a global negative sample pair.
5. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 3 is characterized in that: The specific steps of constructing the positive and negative sample pairs of parts in step 502 are: Step 5021: A training skeleton sequence of the golden monkey is obtained based on the training set, and the training skeleton sequence is input into a skeleton encoder. The skeleton encoder outputs training skeleton features, and the training skeleton features are subjected to local average pooling to obtain part skeleton features of K parts. Step 5022: Input the part text description into the text encoder to obtain part text features of K parts; Step 5023: The part skeleton features and part text features of the same behavior constitute a part positive sample pair, and the part skeleton features and part text features of different behaviors constitute a part negative sample pair, and the part contrast loss is calculated.
6. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 5, characterized in that: The part contrast loss is calculated by weighting the contrast losses of each sub-part.
7. A method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 3 or 4, characterized in that: In the iterative training of the action recognition network based on contrastive learning, the synonym extended text description is input into the text encoder alternately with the corresponding global text description in a random replacement manner to obtain the global text features.
8. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 1, characterized in that: The specific method for training and updating the skeleton encoder in step 4 includes: obtaining a training skeleton sequence of the golden monkey based on the training set, inputting the training skeleton sequence into the skeleton encoder, the skeleton encoder outputting the training skeleton features, predicting the behavior probability distribution through the classifier, calculating the gradient based on the cross entropy loss function, and updating the parameters of the skeleton encoder.
9. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 1, characterized in that: The specific method for testing the behavior recognition network in step six is as follows: obtain the test skeleton sequence of the golden monkey based on the test set, input the test skeleton sequence into the skeleton encoder trained in step five, the skeleton encoder outputs the test skeleton features, and the test skeleton features are passed through the classifier to obtain the predicted behavior probability; take the category with the highest probability as the prediction result, calculate the accuracy, and obtain the accuracy of the behavior recognition network.
10. The method for fine-grained wildlife behavior recognition based on text supplementation status according to claim 1, characterized in that: The skeleton encoder consists of GC-MTC modules. Each GC-MTC module contains a graph convolution GC layer and a multi-scale temporal convolution MTC module. The multi-scale temporal convolution MTC module includes: 1×1 convolution branch, temporal convolution branch with dilation rate 1, temporal convolution branch with dilation rate 2, and MaxPool branch.