Vision-text cross-modal panda behavior recognition method based on attention mechanism
By using Transformer networks and cross-modal attention mechanisms, combined with symmetric cross-entropy loss to optimize parameters, the problem of insufficient interaction depth between visual and textual features in giant panda behavior recognition was solved, thereby improving the accuracy and robustness of giant panda behavior recognition.
Patent Information
- Application Number
- CN202511541994.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing cross-modal learning methods for giant panda behavior recognition suffer from problems such as insufficient capture of video temporal information, insufficient depth of interaction between visual and textual features, and inability of preprocessing results to be effectively fed back to the attention mechanism, resulting in insufficient accuracy and robustness in behavior recognition.
We employ a cross-modal attention mechanism based on Transformer networks, and construct paired samples of video frame sequences and text descriptions by combining visual-text feature alignment and bidirectional matching learning with symmetric cross-entropy loss to optimize parameters. We introduce a behavior label-driven semantic anchor design, optimize feature fusion and preprocessing strategies, and enhance the model's ability to focus on key behavioral features.
It improves the accuracy of giant panda behavior recognition, solves the problem of insufficient interaction depth between visual and text features, enhances the integrity and robustness of the model's dynamic representation of behavior, and adapts to visual data scenarios in complex field environments.
Smart Images

Figure CN120997886A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of attention mechanism technology, and in particular to a visual-text cross-modal giant panda behavior recognition method based on attention mechanism. Background Technology
[0002] The current giant panda behavior recognition process mainly consists of the following steps: First, data collection involves capturing images or videos of giant pandas eating, resting, and other behaviors in captivity or the wild using high-definition infrared cameras, sensors, and other equipment, while simultaneously recording environmental information. Second, data preprocessing involves denoising and standardizing the raw data, and having professionals label the behavior categories to construct a structured dataset. Next, model building and training are performed, using algorithms such as convolutional neural networks and long short-term memory networks. After dividing the dataset, the model parameters are iteratively optimized, and the structure is adjusted to improve classification ability. Subsequently, the model is evaluated using a test set, and optimization is performed based on indicators such as accuracy. Supplementary data or improved methods are used for behaviors with high recognition errors. Finally, the model is deployed and dynamically updated, used for real-time monitoring, generating behavior reports and early warnings, and the model is regularly updated with new data to adapt to changes in behavior patterns.
[0003] For example, Chinese invention patent CN111967358B provides a neural network gait recognition method based on an attention mechanism, including the following steps: segmenting the training set and test set from the benchmark dataset; pre-training the network using a gait extraction model without an embedded attention mechanism to make the model adaptable to human gait; embedding temporal and spatial attention mechanism modules into the network and loading the pre-trained network model parameters; and retraining the attention mechanism-based gait recognition feature extraction model using the dataset.
[0004] For example, Chinese invention patent CN112257572A discloses a behavior recognition method based on a self-attention mechanism. This method employs a keyframe target position prediction module and a continuous frame action category prediction module based on a multi-angle attention mechanism. While performing continuous frame action detection, it can also achieve target localization. Replacing the 3D convolutional network with the keyframe target position prediction and continuous frame action category prediction modules based on a multi-angle attention mechanism improves the model's parallel computing capabilities on GPUs. Furthermore, the multi-angle attention mechanism-based keyframe target position prediction and continuous frame action category prediction modules avoid the compatibility issues that arise when converting or deploying models under different deep learning frameworks due to 3D convolution.
[0005] The current technology has the following technical problems: Existing cross-modal learning methods, which jointly align image and text features and directly apply CLIP (Contrastive Language-Image Pre-trained Model) to video behavior recognition, still face multiple challenges: First, the capture of temporal information in videos is insufficient, making it difficult to fully represent the dynamic process of behavior; second, the interaction depth and correlation between visual and text features are insufficient, failing to achieve effective semantic fusion; furthermore, the current behavior recognition preprocessing stage lacks a unified standard, and the preprocessing results cannot be effectively fed back to the attention mechanism, further affecting the model's ability to focus on key behavioral features. Summary of the Invention
[0006] To address the technical problems of insufficient interaction depth and correlation between visual and textual features in existing technologies, and the inability to effectively feed preprocessing results back to the attention mechanism, this invention provides a visual-text cross-modal giant panda behavior recognition method based on an attention mechanism. The technical solution is as follows: A visual-text cross-modal giant panda behavior recognition method based on an attention mechanism is provided. The method includes: Step 1: Inputting a multimodal dataset into an initial behavior recognition model, extracting feature data of giant panda behavior through a Transformer network, and introducing a cross-modal attention mechanism to align the feature data; Step 2: Based on the feature data of giant panda behavior, constructing paired samples of video frame sequences and text descriptions, and conducting bidirectional matching learning in a unified embedding space using a cross-modal representation network. An optimizer uses symmetric cross-entropy loss as the main loss function to update parameters, requiring paired samples to move closer together and unpaired samples to separate. Through validation and early stopping mechanisms, when the training end marker is detected, the obtained behavior recognition model parameters and inference process are fixed, thus marking the initial behavior recognition model as a behavior recognition model; Step 3: Obtaining target giant panda behavior data, optimizing the visual data in the target giant panda behavior data, inputting the optimized target giant panda behavior data into the behavior recognition model, and analyzing and identifying the target giant panda behavior category through the behavior recognition model.
[0007] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: (1) The present invention is a visual-text cross-modal giant panda behavior recognition method based on attention mechanism. First, the multimodal dataset is input into the initial model, and the giant panda behavior features are extracted through the Transformer network. A customized cross-modal attention mechanism is introduced to realize feature alignment, enhance the interaction depth between visual and text features, and break through the bottleneck of insufficient semantic fusion. Then, based on the features, video frame sequence-text description paired samples are constructed. Bidirectional matching learning is carried out in a unified embedding space through the cross-modal representation network. The parameters are optimized with symmetric cross-entropy loss. The model is solidified by combining verification and early stopping mechanisms to fully capture video temporal information and solve the problem of incomplete dynamic representation of behavior. Finally, the visual part of the target giant panda behavior data is optimized. The preprocessing strategy is adjusted according to the quality parameters so that the optimization results assist the attention mechanism to focus on key features, improve the current situation where there is no unified standard for preprocessing and no feedback, and finally accurately identify the target behavior category, providing reliable technical support for giant panda monitoring research.
[0008] (2) This invention uses a behavior-label-driven semantic anchor design, introducing category labels as explicit anchors to link category embedding, video global features, and text global features for three-way interaction, significantly optimizing cross-modal feature fusion and category discrimination. First, L2 norm normalization is performed on the three to ensure that features interact at a uniform scale, avoiding the impact of magnitude differences on fusion accuracy. Then, text features guide video feature attention weighting, generating accurate attention weights along the time dimension, effectively aggregating key temporal information of the video, and solving the problem of insufficient utilization of video dynamic features in traditional fusion. Second, the category embedding is extended to a homomorphic tensor, enabling category, video, and text to interact directly in the same spatial location, strengthening the semantic correlation among the three, and overcoming the problem of missing semantic anchors caused by relying solely on bidirectional video-text interaction. Finally, the original category score is obtained through element-level multiplication and word-dimensional averaging, providing a high-quality discrimination basis for subsequent loss calculation and inference, greatly improving the accuracy and reliability of behavior recognition.
[0009] (3) This invention achieves dynamic control of data cleaning by comparing the quality characterization coefficient with the threshold value in a hierarchical manner, avoiding the inadequacy of the unified process for adapting to different quality data, ensuring efficient processing of high-quality data and accurate optimization of low-quality data; it introduces a dual judgment of the quality characterization coefficient increment and normal data conditions, combined with the constraint of the number of executions, which maximizes the repair effect of low-quality data through parameter adjustment, avoids excessive optimization consumption, and ensures data input and triggers model initialization when the repair is ineffective, thus balancing recognition efficiency and accuracy; it also links the deviation of the quality characterization coefficient with the temperature coefficient of the visual Transformer encoder, so that the model attention weight dynamically focuses on key features according to the data quality, effectively making up for the feature defects of low-quality visual data, and ultimately improving the robustness and accuracy of behavior recognition, adapting to the complex and ever-changing visual data scenarios in the monitoring of wild giant panda behavior. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of a visual-text cross-modal giant panda behavior recognition method based on an attention mechanism provided in an embodiment of the present invention; Figure 2 This is a flowchart of the target giant panda data optimization process provided in an embodiment of the present invention; Figure 3 This is a flowchart of the determination and initialization behavior recognition model provided in the embodiments of the present invention. Detailed Implementation
[0012] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0013] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0014] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0015] This invention provides a visual-text cross-modal giant panda behavior recognition method based on an attention mechanism. For example... Figure 1 The flowchart shown is for a visual-text cross-modal giant panda behavior recognition method based on an attention mechanism. The processing flow of this method may include the following steps: Step 1: Input the multimodal dataset into the initial behavior recognition model, extract the feature data of giant panda behavior through the Transformer network, and introduce a cross-modal attention mechanism to align the feature data.
[0016] The training of behavior recognition models is based on multimodal datasets. The input during the training phase includes visual features of pre-processed panda behavior video frame sequences (such as frame-level feature sequences extracted through convolution and Transformer), corresponding behavior category labels, and associated text descriptions. The training framework is usually built on a pre-trained base model. For example, a video Transformer model (such as Video Swin Transformer) pre-trained on general video understanding tasks (such as action classification and temporal feature modeling) is selected as the basic backbone for visual feature learning. At the same time, a language model (such as a lightweight version of BERT) pre-trained on a large-scale text corpus is used as the basic module for text feature learning. By leveraging the general visual temporal dependency capture ability and text semantic understanding ability already learned by the base model, the training difficulty and data requirements of the model on panda-specific behavior data are reduced. During training, the model relies on the base model. On the one hand, it learns the mapping relationship between visual features and behavior category labels to optimize the feature extraction accuracy of the visual branch for dynamic changes in giant panda behavior (such as paw movements when eating and limb posture changes when climbing). On the other hand, it combines text descriptions and uses a cross-modal attention mechanism to allow the visual branch and text branch to learn interactively, strengthening the understanding of behavioral semantics (e.g., the text "giant panda eating bamboo" can help the model focus on the key area of "mouth contact with bamboo" in vision). At the same time, it uses loss functions (such as cross-entropy loss) to iteratively optimize the parameters of the base model and subsequent task adaptation layers, so that the model gradually transitions from "general feature learning" to "giant panda specific behavior discrimination", and finally masters the ability to accurately distinguish behavior categories from visual dynamic changes. The input to the model inference stage is a sequence of video frames depicting panda behaviors to be identified, along with text descriptions. After processing by the feature extraction module derived from the basic model, visual feature sequences and text feature sequences are formed respectively. The two types of features are then integrated through a cross-modal fusion layer. The output is the panda behavior category (such as eating, climbing, resting, etc.) corresponding to the video frame sequence, presented in the form of standardized labels, to achieve accurate identification and classification of panda behaviors.
[0017] The multimodal dataset contains video frame sequences of giant panda behavior, behavior category information, and text descriptions. The video frame sequences have been standardized and annotated. The behavior category information consists of pre-defined standardized labels, each corresponding to at least one type of giant panda behavior, such as eating, climbing, or resting. These labels serve as template input parameters to trigger the generation of text descriptions. The text descriptions are generated based on the pre-defined templates, and their content corresponds to the behavior category information. They map the giant panda behavior scenes and core content presented in the video frame sequences in the form of short natural language sentences. For example, when the behavior category is eating, a text description like "a video of a panda eating bamboo" is generated. This achieves the transformation from behavior category to natural language expression and establishes a connection between text modality, visual modality, and behavior category information.
[0018] Standardized annotation refers to the process of labeling and describing the behavior of giant pandas in videos, following unified and standardized rules and procedures. Specifically, a standardized definition system covering various giant panda behaviors (such as eating, climbing, and resting) is first established to ensure consistent understanding of behavior categories among different annotators. Then, according to this system, the video clips containing giant panda behavior are labeled frame-by-frame or segment-by-segment to identify the behavior categories exhibited by the pandas. This ensures consistency and accuracy in terminology and behavior judgment criteria, providing reliable and standardized foundational information for subsequent work such as constructing multimodal datasets and training behavior recognition models based on this labeled data.
[0019] Feature data of giant panda behavior was extracted using a Transformer network. This feature data included both video and textual features. The specific extraction process was as follows: S1. Obtain the video sequence to be processed. The video sequence consists of multiple frames of images. Represent the video sequence as follows: Each frame of the image has a preset height and width. Indicates the first A frame image, where ℝ represents the set of real numbers, and the tensor takes real values; 3 indicates RGB three channels. T For video frame rate, H and W These represent the height and width of the image, respectively.
[0020] S2. Encode the video sequence using a visual Transformer to obtain a frame-level feature sequence, specifically including: S21. First, extract each frame of the video sequence. F t According to the preset patch size P × P Divided into The initial features are obtained by linearly mapping the image patches using two-dimensional convolution: ; in The number of image patches, D For feature dimensions.
[0021] S22. Next, a learnable single [class] category label is added to the front of the linearly mapped image patch sequence. And add video location encoding to the sequence. : ; Subsequently, the feature sequence with added position encoding will be added. The input is fed into a Transformer encoder with a preset number of layers (e.g., L layers) for global feature modeling, while the output dimension remains unchanged. The processed features are: ; S23. Next, extract the features corresponding to the category labels as the visual features of the corresponding single-frame image: ; Finally, the visual features of all single-frame images are arranged in frame order to obtain the frame-level feature sequence of the video: ; S3. Modeling the frame-level feature sequence using a temporal Transformer to obtain temporally enhanced global video features, specifically including: S31. Input the frame-level feature sequence into the temporal Transformer, which consists of multiple layers of residual attention blocks and is used to capture inter-frame temporal dependencies.
[0022] S32. Perform multi-layer residual attention block processing on the frame-level feature sequence in sequence. The processing of each layer is as follows: capture the inter-frame temporal dependency through the multi-head self-attention mechanism in the current layer residual attention block, perform layer normalization on the features after capturing the temporal dependency, and then send the processed features into the feedforward fully connected layer for feature transformation.
[0023] Let the input be , No. The formula for calculating the multi-head self-attention residual connectivity is: ; The formula for calculating the residual connectivity of a feedforward network is: ; Wherein, MHA(·) represents multi-head self-attention mechanism, which can capture the temporal dependencies between frames in the frame-level feature sequence from the perspective of multiple attention heads at the same time (such as the relationship between the giant panda raising its head in the first frame and eating in the third frame), LN is layer normalization, and MLP(·) is feedforward fully connected layer.
[0024] go through L t After layer calculation, the temporally enhanced global features of the video are obtained: ; S33. After all the residual attention blocks in all layers have been processed, output the temporally enhanced global video features.
[0025] The text feature data extraction process is as follows: obtain the input text description, text description W In the form of a word sequence, ,in, For the number of words in the text, w i For the first One word.
[0026] First, each word is mapped to an embedding vector of fixed dimensions. e i For the first i Word embedding vectors of words: ; The embedding operation Embed(·) is a mapping operation that maps discrete symbols (such as words, characters, etc.) to a continuous vector space of fixed dimensions.
[0027] Add learnable text position encoding Then, the text feature sequence after adding positional encoding is: ; The position-encoded sequence z 0 input to L txt In a Transformer model composed of layered text Transformer residual attention blocks, let... Z 0= z 0, then similarly, the text Transformer has the first... The formula for calculating the multi-head self-attention residual of a layer is: ; The feedforward residual is: ; Where MHA(·) represents multi-head self-attention mechanism, LN represents layer normalization, and MLP(·) represents feedforward fully connected network.
[0028] After all residual attention blocks in all layers have been processed, the text word-level feature sequence is obtained: ; The feature of the first position in the text sequence is taken as the global semantic vector of the text: ; The cross-modal attention mechanism specifically refers to: acquiring category embeddings, video global features, and text global features, and performing L2 norm normalization on each of them; applying attention guidance to video features using text features, calculating the similarity matrix between text and video, obtaining attention weights along the time dimension using the Softmax function, and weighting and aggregating video features using the attention weights to generate text-guided video features; expanding the category embeddings into tensors that are homologous to the text-guided video features, enabling direct interaction between categories, text, and video in the same spatial location; performing element-wise multiplication on the interacting features, averaging along the word dimension, and obtaining the original category score that can be used for cross-entropy loss calculation and inference stages, without Softmax normalization.
[0029] This paper proposes a cross-modal attention mechanism to achieve efficient interaction between visual and textual modalities. The mechanism uses behavioral labels as category embeddings for TRIPLE, employing category embedding as a semantic anchor. First, text provides fine-grained guidance to the video, then categories interact directly with the aligned video-text representations, thus obtaining highly discriminative joint features through a concise computational path.
[0030] Let the learnable class embedding vector be... Video frame-level features are Text word-level features are ,in C For the number of categories, B For batch number, For the number of words in the text, T For video frame rate, D Let be the number of feature dimensions. First, perform L2 norm normalization on the three dimensions to make them mathematically lie on the same hypersphere: ; symbol This indicates forward / left assignment. This indicates L2 norm normalization.
[0031] Subsequently, text features are applied to video features with attention at the word level to calculate the similarity matrix between text and video. S : ; a ∈[1, C ], b ∈[1, B ], , t ∈[1, T ], d ∈[1, D These correspond to the category, batch, word, frame, and feature index, respectively. S abit Pointer matrix S Corresponding dimensions The corresponding index below abit The same applies to the elements thereafter.
[0032] The attention weight matrix is obtained by using Softmax along the video frame dimension. A : ; This formula involves a summation operation along the frame dimension. To avoid confusion, the summation frame index is re-denoted as... The temperature coefficient τ1 = 0.01 is used to sharpen the attention weight distribution.
[0033] The video features are weighted and aggregated using these weights to obtain the video feature matrix of text-guided attention (att). , By its components Combining based on indices (a, b, i, d): ; Next, embed the category. Expand (expend: exp) to be related to Isomorphic matrices : ; That is, a matrix Extend Dimensions, but without changing the original element values. By its components It is composed of indices (a, b, i, d).
[0034] In the expanded space, categories, text, and video interact directly at the same location. After element-wise multiplication, the average is taken along the word and feature dimensions to obtain the original category score before Softmax normalization. logits That is, the original category score matrix here. Z : ; Where ⊙ represents element-wise multiplication (Hadamard product).
[0035] This ternary interaction mechanism uses category embedding as its core anchor point. It first allows the text to "see" the video in a fine-grained manner, and then allows the category to directly interact with the aligned video-text representation. The calculation path is concise and the alignment effect is good.
[0036] Step 2: Based on the feature data of giant panda behavior, construct paired samples of video frame sequences and text descriptions. The cross-modal representation network conducts bidirectional matching learning in a unified embedding space, and the optimizer updates the parameters using symmetric cross-entropy loss as the main loss function. Paired samples are required to move closer to each other, and unpaired samples are required to separate from each other. Through validation and early stopping mechanisms, when the end of training is detected, the obtained behavior recognition model parameters and inference process are solidified, thereby marking the initial behavior recognition model as the behavior recognition model.
[0037] The optimizer updates parameters using symmetric cross-entropy loss as the main loss function. Specifically, the main loss function is guided by the bidirectional matching relationship between video and text, and text and video. It first calculates two types of unidirectional matching losses separately, and then combines them to form the final symmetric cross-entropy loss. The video-to-text matching loss measures the similarity between video features and text features for each video in the training batch using cosine similarity. The similarity distribution is scaled by a temperature coefficient to maximize the similarity between the video and text of the same category and minimize the similarity with text of different categories. The text-to-video matching loss uses the same logic to maximize the similarity between each text and videos of the same category and minimize the similarity with videos of different categories.
[0038] In one example implementation, a bidirectional matching learning objective is designed, and the model is trained using symmetric cross-entropy loss to improve the alignment capability of cross-modal representations. Specifically, the matching relationships between video to text and text to video are considered simultaneously. By maximizing the similarity of matched pairs and minimizing the similarity of unmatched pairs, consistency constraints in the cross-modal feature space are achieved.
[0039] Let the global features of the text be The global features of the video are Both have been processed using L2 normalization. Similarity function Using cosine similarity: ; For video-to-text matching loss, with a batch size of B In the case of, let In order to be with the first b For a set of video tags with the same category index, the loss can be defined as: ; Where τ2 is a temperature coefficient hyperparameter used to scale the smoothness of the similarity distribution; K , b , As a vector index, when used as a subscript, it represents the first position of the corresponding vector. K , b , Each element.
[0040] Similarly, the text-to-video matching loss can be defined as: ; The final symmetric cross-entropy loss is: ; This loss function guides video features and corresponding category features to be closer in the embedding space during training, while separating mismatched feature pairs, thereby enhancing the model's performance in cross-modal retrieval and classification tasks.
[0041] Step 3: Obtain the target giant panda's behavior data, optimize the visual data in the target giant panda's behavior data, input the optimized target giant panda's behavior data into the behavior recognition model, and analyze and identify the target giant panda's behavior category through the behavior recognition model.
[0042] In this invention, the target giant panda behavior data specifically refers to the dataset to be analyzed for behavior categories. This dataset includes video data recording the giant panda's activities, as well as textual descriptions added later to describe the video content. When processing this data, optimization is only performed on the video data. The core reason is that video data is often collected in complex environments such as the wild, and is easily affected by factors such as changes in lighting, vegetation obstruction, and camera shake, leading to quality issues such as blurriness and loss of detail, which in turn affects the model's accuracy in extracting behavioral features. Textual descriptions, on the other hand, are generated later through manual or automated annotation, resulting in stable content and no clarity-related quality fluctuations, requiring no additional optimization. Therefore, this invention prioritizes optimizing the video data to improve the quality of visual features, and then inputs the optimized complete target data into the behavior recognition model. This ensures that the model, relying on the synergy of high-quality visual data and stable textual descriptions, accurately analyzes and identifies the specific behavior category of the giant panda.
[0043] Figure 2 This is a flowchart of the target giant panda data optimization process provided in an embodiment of the present invention. Figure 1After the model training in steps one through three is completed, the visual data contained in the target giant panda behavior data to be input into the trained model needs to be optimized. The quality representation coefficient of the visual data is compared with the defined quality representation coefficient. If the quality representation coefficient of the visual data is greater than or equal to the defined quality representation coefficient, the current data cleaning process is executed to perform data optimization. The optimized target giant panda behavior data is then input into the behavior recognition model, which analyzes and identifies the target giant panda behavior category. If the quality representation coefficient of the visual data is less than the defined quality representation coefficient, the current data cleaning process is adjusted. The deviation of the quality representation coefficient is compared with a reference value. If the deviation is less than the reference value, the radius of the slight deblurring kernel in the data cleaning process is increased based on the deviation. If the deviation is greater than or equal to the reference value, the radius of the slight deblurring kernel and the global sharpness reconstruction weights are increased based on the deviation. After adjustment, data optimization is performed, and the quality representation coefficient of the visual data is updated.
[0044] Figure 3 This is a flowchart of the behavior recognition model initialization process provided in this embodiment of the invention. After the data optimization process is completed, the quality characterization coefficient of the visual data is updated, and it is determined whether to issue a quality warning for the visual data in the target giant panda behavior data. If the determination is yes, a quality warning is issued. If the determination is no, it is determined whether there are normal data conditions. If they exist, the target giant panda behavior data is input into the behavior recognition model. If they do not exist, the parameter adjustment strategy is continued. After normal data conditions are detected within the defined number of executions, the target giant panda behavior data is input into the behavior recognition model. If no normal data conditions are detected within the defined number of executions, the behavior recognition model is initialized, and the target giant panda behavior data is input into the behavior recognition model.
[0045] The visual data in the target giant panda behavior data is optimized by: obtaining the perceptual hash entropy value, the local contrast entropy, and the inter-frame optical flow gradient magnitude variance of the visual data.
[0046] Perceptual hash entropy refers to the information entropy calculated based on the perceptual hash features of visual data (such as images and video frames). It is used to quantify the richness and disorder of texture details in visual data. The higher the entropy value, the more complete the visual data details and the better the quality. Conversely, it indicates that details are lost or the degree of blurriness is relatively high. First, the visual data is converted into a grayscale image and scaled to a fixed size (such as 8×8 pixels) to reduce redundant information interference. Discrete cosine transform is performed on the scaled image (to extract low-frequency components, reflecting the overall structure and key details of the image). A threshold is set to binarize the low-frequency components, generating a perceptual hash sequence composed of 0s and 1s. Finally, the probability distribution of 0s and 1s in the hash sequence is statistically analyzed, and the perceptual hash entropy value is calculated based on the information entropy formula.
[0047] Local contrast entropy is an information entropy obtained by dividing visual data into local regions, calculating the contrast of each region, and statistically analyzing its probability distribution. It is used to characterize the complexity and discernibility of local brightness differences in visual data. A higher entropy value means richer local brightness levels and stronger detail differentiation. The visual data is divided into several non-overlapping local regions according to a preset size (such as 16×16 pixels); the contrast of each region is calculated (usually the difference between the maximum and minimum pixel values in the region, or the standard deviation of the region's grayscale); the distribution of contrast values of all regions is statistically analyzed to obtain the probability density function of contrast; the entropy value of this distribution is calculated by substituting it into the information entropy formula, which is the local contrast entropy.
[0048] Inter-frame optical flow gradient magnitude variance, for video-like visual data, describes the dispersion of the optical flow field gradient magnitude between two adjacent frames. It is used to quantify the intensity of inter-frame motion (such as camera shake or rapid target movement). A larger variance indicates more unstable inter-frame motion, easily leading to blurring or ghosting; conversely, a smaller variance indicates smoother motion. First, for two adjacent video frames, a dense optical flow algorithm (such as the Farneback algorithm) is used to generate the optical flow field: by analyzing the positional change of each pixel between the two frames, the horizontal motion component (horizontal movement distance) and vertical motion component (vertical movement distance) of that pixel are determined, thus quantifying the motion state of all pixels. Next, the spatial gradient of the motion vector of each pixel is calculated: for the horizontal motion component of a single pixel, its difference from the horizontal motion components of adjacent pixels is analyzed to obtain... The horizontal gradient is calculated by analyzing the difference between the vertical motion component of the pixel and the vertical motion component of its neighboring pixels. This vertical gradient reflects the rate of change of each pixel's motion state in the surrounding area. Then, the horizontal and vertical gradients of the same pixel are calculated together. Specifically, the gradient values in the two directions are squared separately, the two squared values are added together, and the square root of the sum is taken. This process fuses the gradient information in the two directions into a single value, which is the optical flow gradient magnitude of the pixel. This value can intuitively reflect the strength of the motion change of a single pixel. Finally, the optical flow gradient magnitude values of all pixels in two frames are collected, and the overall dispersion of these values is statistically analyzed, i.e., variance processing is performed. The result is the inter-frame optical flow gradient magnitude variance.
[0049] The perceptual hash entropy value is compared with the minimum perceptual hash entropy value, the local contrast entropy is compared with the minimum local contrast entropy value, and the maximum value of the inter-frame optical flow gradient magnitude variance is compared with the inter-frame optical flow gradient magnitude variance. The results of each comparison are weighted and summarized to obtain the quality characterization coefficient of the visual data. The quality characterization coefficient of the visual data is used to quantitatively characterize the sharpness quality of the visual data.
[0050] In weighted summation, each comparison result is multiplied by its corresponding weight coefficient and then summed. The weight coefficient refers to the proportion of the parameter in the quality characterization coefficient, and its value ranges from 0 to 1. Multiple sets of comparative experiments are used to verify the rationality of the weights. Different weight combinations are adjusted to calculate the quality characterization coefficient. The accuracy of the model in recognizing behaviors such as eating, climbing, and resting of giant pandas is observed. The weight allocation scheme that makes the overall recognition accuracy of the model optimal and the robustness strongest is selected. Finally, the specific proportion of each parameter in the quality characterization coefficient is determined to ensure that the weight coefficient can objectively reflect the actual impact value of the parameter on data quality.
[0051] The determination of the minimum allowable value of perceptual hash entropy, the minimum allowable value of local contrast entropy, and the maximum value of the magnitude variance of inter-frame optical flow gradient requires relevant technical personnel to determine the values through both data statistics and experimental verification, taking into account the visual data characteristics and model requirements of the giant panda behavior recognition scenario. First, a large number of wild giant panda video samples (covering different lighting, occlusion, and shaking scenarios) were collected. The perceptual hash entropy, local contrast entropy, and inter-frame optical flow gradient magnitude variance of each sample were extracted, and the distribution range and data characteristics of the three types of parameters were statistically analyzed. Second, with the core criterion of the parameters corresponding to visual quality meeting the model's feature extraction requirements, effective samples that enable the model to accurately recognize giant panda behavior were selected. The lower limit range of the perceptual hash entropy and local contrast entropy, and the upper limit range of the inter-frame optical flow gradient magnitude variance in these samples were determined. For example, a perceptual hash entropy value that is too low will lead to the loss of texture details, so the lowest value of this entropy value in the effective samples was taken as the minimum allowable value. A local contrast entropy value that is too low means that the brightness and darkness of the picture are blurred, so the lowest value in the effective samples was taken as the minimum allowable value. An inter-frame optical flow gradient magnitude variance that is too high indicates violent movement and picture distortion, so the highest value in the effective samples was taken as the maximum allowable value. Finally, the rationality of the parameters was verified through multiple sets of comparative experiments, and the model was fine-tuned until the recognition accuracy and robustness of giant panda behavior were optimal within the parameter range. Finally, the allowable extreme values of the three types of parameters were determined.
[0052] When visual data experiences drastic inter-frame motion due to camera shake, rapid target movement, or other reasons, the variance of the inter-frame optical flow gradient magnitude increases. This drastic motion easily leads to image blurring and detail misalignment, thereby impairing the integrity of the texture details in the visual data and causing a decrease in perceptual hash entropy. Since perceptual hash entropy depends on the richness of texture details, the loss of details directly reduces the disorder of its 0 and 1 hash sequence. Simultaneously, image blurring weakens the brightness differences in local areas, reducing the brightness gradient between adjacent areas and causing a synchronous decrease in local contrast entropy. Conversely, if the variance of the inter-frame optical flow gradient magnitude is small (smooth inter-frame motion), image details are preserved, the perceptual hash entropy remains at a high level to reflect texture richness, and the local contrast entropy also maintains a high value due to clear brightness levels. In the calculation of the quality characterization coefficient, the perceptual hash entropy and local contrast entropy are both positive influencing factors. The higher the value of the two, the better the data detail and discrimination. The inter-frame optical flow gradient magnitude variance is a negative influencing factor. The higher the value of the two, the more serious the motion interference. The three factors are linked through the chain from motion stability to detail integrity to brightness and darkness discrimination, and together determine the final value of the quality characterization coefficient, thereby quantifying the overall quality of visual data.
[0053] The quality characterization coefficients of the visual data are compared with the defined quality characterization coefficients. The defined quality characterization coefficients, representing the minimum allowable value of the quality characterization coefficients, are formulated based on the adaptability of giant panda visual data quality to the performance of the behavior recognition model, and are determined through multi-stage data analysis and experimental verification. First, visual data samples of wild giant pandas under different collection scenarios (such as changes in lighting, vegetation obstruction, camera shake, etc.) are collected. The perceptual hash entropy, local contrast entropy, and inter-frame optical flow gradient magnitude variance of each sample are extracted. The quality characterization coefficients of each sample are calculated according to preset weights (determined based on the degree of influence of each parameter on the model's feature extraction), forming an initial coefficient dataset. Second, samples with different quality characterization coefficients are input into the behavior recognition model. The model's behavior recognition accuracy for each sample is statistically analyzed. Samples with recognition accuracy meeting a preset threshold (such as above 90%) are selected, and the lowest value of the quality characterization coefficient among these effective samples is determined as the defined quality characterization coefficient.
[0054] If the quality characterization coefficient of the visual data is greater than or equal to the defined quality characterization coefficient, it indicates that the texture detail integrity and brightness level differentiation of the visual data meet the standards, and the inter-frame motion stability is within a reasonable range, satisfying the basic requirements of the behavior recognition model for feature extraction. The model can effectively identify typical behaviors of giant pandas such as eating, climbing, resting, and activity based on this type of data. Then, the current data cleaning process is executed, mainly to optimize data redundancy and minor interference. Specifically, this includes: removing invalid frames without giant panda targets in the video, removing local noise pixels caused by brief lens obstruction, lightly calibrating the brightness and contrast of the image to balance the consistency of visual features, and retaining the temporal correspondence between the text description and the video frames. This ensures that the optimized data maintains the integrity of the original behavioral information while reducing the impact of irrelevant interference on the model's inference. The optimized target giant panda behavior data is then input into the behavior recognition model.
[0055] If the quality characterization coefficient of the visual data is less than the defined quality characterization coefficient, the current data cleaning process needs adjustment. This indicates that the overall quality of the current visual data does not meet the basic requirements of the behavior recognition model. The initial data cleaning process was designed based on a preset scenario where the quality characterization coefficient of the visual data is greater than or equal to the defined quality characterization coefficient. This process can only handle minor redundancy or interference and cannot specifically address the aforementioned low-quality issues (such as the inability to repair severe blur or suppress image misalignment caused by violent movements). Continuing with the original process will not only fail to improve data quality but may also lead to further loss of useful information due to ineffective processing, resulting in a significant decrease in the accuracy of subsequent model recognition. Therefore, the current data cleaning process needs to be adjusted.
[0056] Adjusting the current data cleaning process refers to executing parameter adjustment strategies. Specifically, these strategies involve differentiating the quality characterization coefficients of the defined quality characterization coefficients from those of the visual data (i.e., difference processing), comparing the result with the defined quality characterization coefficients (i.e., ratio processing), and labeling the result as the quality characterization coefficient deviation. This deviation indicates the degree of deviation between the defined quality characterization coefficients and the visual data's quality characterization coefficients. The greater the deviation, the lower the quality of the visual data is compared to the defined quality characterization coefficients.
[0057] The deviation of the quality characterization coefficient is compared with the reference value of the deviation of the quality characterization coefficient. The reference value of the deviation of the quality characterization coefficient is a critical threshold used to define the visual data quality as slightly substandard or severely substandard, and is formulated by relevant technical personnel.
[0058] If the deviation of the quality characterization coefficient is less than the reference value, indicating that the visual data quality is slightly substandard and the core problem is mostly slight blur (such as slight lens fogging or slight detail blurring caused by low light), then the radius of the slight deblurring kernel in the data cleaning process is adjusted based on the deviation of the quality characterization coefficient. This is done by multiplying the deviation of the quality characterization coefficient by the current radius of the slight deblurring kernel and adding the result to the current radius of the slight deblurring kernel, thus increasing the radius of the slight deblurring kernel. At this point, only the radius of the slight deblurring kernel is adjusted because the radius of the deblurring kernel directly determines the range of the blurred areas that the deblurring algorithm can repair. Slight blur corresponds to a small and shallow blur area. By increasing the radius of the deblurring kernel, the algorithm can more accurately cover the slightly blurred areas while avoiding image distortion caused by over-repair.
[0059] If the deviation of the quality characterization coefficient is greater than or equal to the reference value, indicating that the visual data quality is severely substandard, there is often a combination of moderate to severe blurring and weakened global texture (such as overall blurring caused by thick vegetation occlusion, or inter-frame detail misalignment caused by severe camera shake). Simply increasing the radius of the deblurring kernel cannot repair the loss of global texture. The deblurring kernel focuses on the edge recovery of local blurred areas, while global sharpness reconstruction can compensate for large-scale detail loss by adjusting the overall contrast of the image and strengthening the texture level. Based on the increase of the deviation of the quality characterization coefficient, the radius of the slight deblurring kernel in the data cleaning process is adjusted. At the same time, based on the increase of the deviation of the quality characterization coefficient, the weight of global sharpness reconstruction in the data cleaning process is adjusted. That is, the deviation of the quality characterization coefficient is multiplied by the weight of global sharpness reconstruction, and the result is added to the current weight of global sharpness reconstruction, thereby completing the adjustment of the current weight of global sharpness reconstruction. Through dual parameter adjustment, both the edge details of local blurred areas are repaired and the texture level of the global image is reconstructed, which can effectively improve the quality of severely substandard data and make it meet the basic requirements of model feature extraction (such as clearly presenting the paw movements of a giant panda when eating, or the limb posture when climbing). Ensuring robustness of model recognition: Avoiding the problem of limited quality improvement due to insufficient single repair methods, dual adjustment can make the visual features of severely substandard data more consistent with those of compliant data after cleaning, reducing the recognition error caused by fluctuations in input quality, and ensuring stable recognition accuracy for various behaviors of giant pandas.
[0060] After adjustment, data optimization is performed. Once data optimization is complete, the quality characterization coefficient of the visual data is updated to determine whether to issue a quality warning for the visual data in the target giant panda's behavioral data.
[0061] To determine whether to issue a quality warning for the visual data in the target giant panda behavior data, the specific determination process is as follows: the quality characterization coefficient of the updated visual data is compared with the quality characterization coefficient of the unupdated visual data, the result is divided by the quality characterization coefficient of the unupdated visual data, the result is marked as the quality characterization coefficient increment, and compared with the defined quality characterization coefficient increment.
[0062] The definition of the quality characterization coefficient increment refers to the minimum improvement that the quality characterization coefficient of visual data should achieve after the data cleaning process. To obtain the definition of the quality characterization coefficient increment, multiple sets of panda visual data samples with slightly and severely substandard quality are first collected. After processing with corresponding data cleaning and adjustment strategies, the quality characterization coefficients before and after cleaning for each set of samples are recorded, and the increment is calculated. Then, samples with different increments after cleaning are input into a behavior recognition model, and samples with recognition accuracy reaching a preset passing standard are selected. The minimum value of the quality characterization coefficient increment among these valid samples is extracted as the initial reference value. Subsequently, experiments are repeated under different field collection scenarios such as low light and high occlusion, and the initial reference value is weighted and calibrated based on the data proportion of each scenario. Finally, the stability of the calibrated definition value is verified by continuously inputting multiple sets of low-quality samples, ultimately determining the definition of the quality characterization coefficient increment that can accurately judge the effectiveness of cleaning and adjustment.
[0063] If the increment of the quality characterization coefficient is less than or equal to the increment of the defined quality characterization coefficient, it indicates that the cleaning and adjustment are ineffective. At this time, the data still has unresolved quality defects (such as blurring not being effectively eliminated or missing texture details). If it is directly input into the model, it will lead to insufficient accuracy in the model's extraction of giant panda behavioral features, which will cause problems such as decreased recognition accuracy, misjudgment (such as misjudging resting as eating) or missed judgment, affecting the reliability of subsequent giant panda behavior analysis results. Therefore, it is determined to issue a quality warning for the visual data in the target giant panda behavior data, that is, generate a quality defect analysis report. Combined with the previous parameter records (such as the variance of the magnitude of the inter-frame optical flow gradient and the local contrast entropy), the specific reasons for the insufficient increment are located, and a defect report is generated and uploaded to the technical personnel.
[0064] If the increment of the quality characterization coefficient is greater than the increment of the defined quality characterization coefficient, and there are normal data conditions, then the target giant panda behavior data will be input into the behavior recognition model; this means that the data simultaneously meets the two core conditions of effective cleaning and adjustment and quality compliance, and has the basic qualifications to be input into the behavior recognition model.
[0065] If the increment of the quality characterization coefficient is greater than the increment of the defined quality characterization coefficient, and there are no normal data conditions, the parameter adjustment strategy will continue to be executed. After normal data conditions are detected within the defined number of executions, the target giant panda behavior data will be input into the behavior recognition model. If no normal data conditions are detected within the defined number of executions, the behavior recognition model will be initialized, and the target giant panda behavior data will be input into the behavior recognition model.
[0066] Defining the number of executions refers to the maximum number of times the parameter adjustment strategy is allowed to be executed. The process of determining this number of executions must balance the data cleaning effect and processing efficiency, and revolve around two core principles: avoiding over-adjustment and ensuring quality compliance. For example, firstly, by experimentally analyzing the quality improvement patterns of different quality defect data after 1-5 adjustments, it is determined that most data can meet the quality standards after 3-4 adjustments, and further increasing the number of adjustments will only increase computational costs without significant effect. Secondly, considering the timeliness requirements of the behavior recognition model for data processing, the risk of process delays caused by excessive number of adjustments is eliminated. Finally, the minimum value between the maximum number of adjustments required for most data to meet the standards and the upper limit of the timeliness allowed is determined as the defined number of executions, ensuring that ineffective adjustments that waste resources are avoided while covering the quality repair needs of most data.
[0067] Normal data conditions refer to the updated visual data having a quality characterization coefficient greater than or equal to the defined quality characterization coefficient.
[0068] If, within the defined number of executions, normal data conditions are still not detected after multiple parameter adjustments (i.e., the updated visual data quality characterization coefficient consistently fails to meet the defined standard), considering that although the data is not fully up to standard, it has undergone multiple rounds of cleaning and optimization, and the degree of quality defects has been significantly reduced, the target giant panda behavior data can be input into the behavior recognition model. At the same time, the behavior recognition model is initialized by resetting the model's internal parameters to enhance its adaptability and robustness to data with slight quality defects. This approach avoids process blockages caused by excessively pursuing perfect data compliance and provides dual guarantees of data optimization and model adaptation through adaptive adjustments at the model end, minimizing the impact of low-quality data on the accuracy of giant panda behavior recognition and ensuring that the recognition process is both efficient and reliable.
[0069] The behavior recognition model is initialized as follows: the quality representation coefficients of the final visual data are obtained, and the deviation of the quality representation coefficients is updated; based on the updated deviation of the quality representation coefficients, the temperature coefficient of the multi-head self-attention mechanism of the visual Transformer encoder in the model is increased, thereby identifying the target giant panda behavior category through the behavior recognition model. The deviation of the quality representation coefficients is multiplied by the current temperature coefficient of the multi-head self-attention mechanism, and the result is accumulated with the current temperature coefficient of the multi-head self-attention mechanism to complete the increase and adjustment of the temperature coefficient of the multi-head self-attention mechanism.
[0070] The multi-head self-attention mechanism of the visual Transformer encoder is responsible for capturing key behavioral features of giant pandas in the image (such as limb movements and posture contours). The temperature coefficient directly affects the accuracy of attention weight allocation. When the data has slight quality defects (such as blurred details or weakened local features), the attention mechanism under the original temperature coefficient may over-focus on irrelevant background information (such as vegetation and lighting) or fail to accurately locate the behavioral feature areas of the giant panda. Increasing the temperature coefficient based on the updated deviation can reduce the concentration of attention weights, allowing the model to more flexibly extract behavioral feature details that are not clear but still exist in the image (such as the panda's paw movements and body posture trends in a blurred state), reducing feature misjudgments caused by missing data details. The advantage of this adjustment is that it eliminates the need for meaningless repeated data cleaning. Instead, through dynamic adaptation of model parameters and data quality, it maximizes the retention of effective behavioral feature information even when the data is not fully up to standard but has been optimized. This improves the model's robustness to low-quality data, ensures the accuracy of giant panda behavior category recognition, and avoids efficiency losses caused by continuous data processing. It achieves the goal of efficient recognition by actively adapting the model when data optimization is limited.
[0071] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0073] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An attention mechanism-based visual-text cross-modal giant panda behavior recognition method, characterized in that, The method comprises: Step 1, inputting a multi-modal data set into an initial behavior recognition model, extracting feature data of panda behavior through a Transformer network, and introducing a cross-modal attention mechanism to align the feature data; Step 2, based on the feature data of panda behavior, constructing a pair of sample of video frame sequence-text description, carrying out bidirectional matching learning in a unified embedding space by a cross-modal representation network, and updating parameters by an optimizer with a symmetric cross-entropy loss as the main loss function, requiring paired samples to approach each other and unpaired samples to separate from each other, and being constrained by verification and early stopping mechanism, when the training end flag is monitored, the obtained behavior recognition model parameters and inference process are solidified, so that the initial behavior recognition model is marked as a behavior recognition model; Step 3, obtaining target panda behavior data, and optimizing the visual data in the target panda behavior data, inputting the optimized target panda behavior data into the behavior recognition model, and identifying the target panda behavior category through the behavior recognition model.
2. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 1, characterized in that, The multi-modal data set comprises panda behavior video frame sequence, behavior category information and text description; The video frame sequence has completed standardized annotation processing; The behavior category information is a preset standardized label, the label corresponds to at least one behavior type of a panda, and is used as a template input parameter to trigger generation of the text description; The text description is generated based on a preset template, and its content corresponds to the behavior category information, and is mapped to the panda behavior scene and core content presented by the video frame sequence in the form of a natural language short sentence.
3. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 1, characterized in that, The feature data of panda behavior is extracted through the Transformer network, wherein the feature data comprises video feature data and text feature data, and the video feature data is extracted in the following specific process: S1. obtaining a video sequence to be processed, the video sequence being composed of multiple images, each image having a preset height and width; S2. encoding the video sequence by a visual Transformer to obtain a frame-level feature sequence, specifically including: S21. dividing each image in the video sequence into multiple image blocks according to a preset patch size, and linearly mapping the image blocks by two-dimensional convolution; S22. adding a learnable class label to the front end of the linearly mapped image block sequence, and adding position encoding to the sequence; S23. inputting the sequence with position encoding into a Transformer encoder with a preset number of layers for global feature modeling, extracting the feature corresponding to the class label as the visual feature of the corresponding single-frame image, arranging the visual features of all single-frame images in sequence according to frames to obtain the frame-level feature sequence of the video; S3. modeling the frame-level feature sequence by a time series Transformer to obtain a time series enhanced video global feature, specifically including: S31. inputting the frame-level feature sequence into the time series Transformer, the time series Transformer being composed of multiple residual attention blocks; S32. sequentially performing multi-layer residual attention block processing on the frame-level feature sequence, and each layer processing process comprises: capturing inter-frame time sequence dependency relationship through a multi-head self-attention mechanism in a current layer residual attention block, performing layer normalization processing on the feature after capturing the time sequence dependency relationship, and sending the processed feature to a feedforward fully connected layer for feature conversion; S33. outputting the time sequence enhanced video global feature after all layer residual attention block processing is completed.
4. The visual-text cross-modal giant panda behavior recognition method based on an attention mechanism according to claim 3, characterized in that, The text feature data, and the specific extraction process comprises: obtaining an input text description, wherein the text description is in the form of a word sequence; performing preprocessing on the word sequence, mapping each word to an embedded vector of a fixed dimension, and adding learnable position encoding to the embedded vector sequence; inputting the embedded vector sequence after adding the position encoding to a Transformer model; processing the sequence through the Transformer model, and each layer residual attention block processing process comprises: processing the input feature by using a multi-head self-attention mechanism, and sending the feature after layer normalization to a feedforward fully connected network for feature conversion; obtaining a text feature sequence after all layer residual attention block processing is completed, and extracting a feature at a first position of the sequence as a text global semantic vector.
5. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 1, characterized in that, The cross-modal attention mechanism, and specifically comprises: obtaining a category embedding, a video global feature and a text global feature, and performing L2 norm normalization processing on the three; implementing attention guidance on the video feature by using the text feature, calculating a similarity matrix of the text and the video, obtaining an attention weight by using a Softmax function along a time dimension, weighting and aggregating the video feature by using the attention weight, and generating a text guided video feature; extending the category embedding into a tensor in the same form as the text guided video feature, so that the category, the text and the video directly interact at the same spatial position; performing element-level multiplication operation on the interacted feature, taking an average value along a word dimension, and obtaining an original category score which can be used for cross-entropy loss calculation and reasoning stage and has not been subjected to Softmax normalization.
6. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 5, characterized in that, The parameter updating is implemented by the optimizer by using the symmetric cross-entropy loss as a main loss function, wherein the main loss function specifically comprises: taking a bidirectional matching relationship of video to text and text to video as an optimization guide, calculating two types of one-way matching losses respectively, and combining the two to form a final symmetric cross-entropy loss; wherein the matching loss of video to text measures the similarity of the video feature and the text feature by using cosine similarity for each video in a training batch, scales the similarity distribution by using a temperature coefficient, maximizes the similarity of the video and the text of the same category, and minimizes the similarity of the video and the text of different categories; the matching loss of text to video adopts the same logic, maximizes the similarity of the text and the video of the same category, and minimizes the similarity of the text and the video of different categories.
7. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 1, characterized in that, The data optimization processing on the visual data in the target giant panda behavior data specifically comprises: obtaining a perceptual hash entropy value of the visual data, a local contrast entropy of the visual data and a frame-to-frame optical flow gradient modulus variance of the visual data; The perception hash entropy value is compared with a minimum allowed perception hash entropy value, the local contrast entropy is compared with a minimum allowed local contrast entropy, the inter-frame optical flow gradient modulus length variance running maximum value is compared with the inter-frame optical flow gradient modulus length variance, and each comparison result is weighted and aggregated to obtain a quality representation coefficient of the visual data, which is used to quantitatively represent the clear quality of the visual data. The quality representation coefficient of the visual data is compared with a defined quality representation coefficient. If the quality representation coefficient of the visual data is greater than or equal to the defined quality representation coefficient, a current data cleaning process is executed to perform data optimization processing, and the target giant panda behavior data after the optimization processing is input into the behavior recognition model. If the quality representation coefficient of the visual data is less than the defined quality representation coefficient, the current data cleaning process is adjusted.
8. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 7, characterized in that, The adjustment of the current data cleaning process refers to the execution of a parameter adjustment strategy, and the parameter adjustment strategy specifically refers to: The defined quality representation coefficient is differentially processed with the quality representation coefficient of the visual data, and the processing result is compared with the defined quality representation coefficient, and the processing result is marked as a quality representation coefficient deviation; The quality representation coefficient deviation is compared with a quality representation coefficient deviation reference value; If the quality representation coefficient deviation is less than the quality representation coefficient deviation reference value, a slight deblurring kernel radius in the data cleaning process is adjusted based on the quality representation coefficient deviation increase; If the quality representation coefficient deviation is greater than or equal to the quality representation coefficient deviation reference value, the slight deblurring kernel radius in the data cleaning process is adjusted based on the quality representation coefficient deviation increase, and the global clarity reconstruction weight in the data cleaning process is adjusted based on the quality representation coefficient deviation increase; After the adjustment is completed, data optimization processing is performed, and after the data optimization processing is completed, the quality representation coefficient of the visual data is updated to determine whether to perform quality warning on the visual data in the target giant panda behavior data.
9. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 8, characterized in that, The determination of whether to perform quality warning on the visual data in the target giant panda behavior data specifically includes the following steps: The updated quality representation coefficient of the visual data is differentially processed with the quality representation coefficient of the visual data before the update, the processing result is divided by the quality representation coefficient of the visual data before the update, the processing result is marked as a quality representation coefficient increment, and the quality representation coefficient increment is compared with a defined quality representation coefficient increment; If the quality representation coefficient increment is less than or equal to the defined quality representation coefficient increment, it is determined to perform quality warning on the visual data in the target giant panda behavior data; If the quality representation coefficient increment is greater than the defined quality representation coefficient increment, and there is a data normal condition, the target giant panda behavior data is input into the behavior recognition model; If the quality characterization coefficient increment is greater than the defined quality characterization coefficient increment, and there is no data normal condition, the parameter adjustment strategy is continuously executed, and after the data normal condition is monitored within the defined execution times, the target giant panda behavior data is input to the behavior recognition model, and if the data normal condition is not monitored within the defined execution times, the behavior recognition model is initialized, and the target giant panda behavior data is input to the behavior recognition model; The data normal condition refers to that the quality characterization coefficient of the updated visual data is greater than or equal to the defined quality characterization coefficient.
10. The attention mechanism based visual-text cross-modal panda behavior recognition method according to claim 9, characterized in that, The initialization process of the behavior recognition model is as follows: Obtain the quality characterization coefficient of the final visual data, and update the quality characterization coefficient deviation; Based on the increased quality characterization coefficient deviation, adjust the multi-head self-attention mechanism temperature coefficient of the visual Transformer encoder in the model, so as to analyze and identify the target giant panda behavior category through the behavior recognition model.
Citation Information
Patent Citations
A neural network gait recognition method based on attention mechanism
CN111967358B
Behavior recognition method based on self-attention mechanism
CN112257572A
Multi-sensor fusion SLAM method and system based on data quality comprehensive evaluation
CN119557846A
Field panda behavior monitoring method and system based on artificial intelligence
CN119863824A
Liquid crystal display driving control method and system
CN120708558A
Cited By
Panda data detection method and system based on learnable motion saliency modulation
CN121747156A
A Data Detection Method and System for Giant Pandas Based on Learnable Motion Saliency Modulation
CN121747156B