An animal pose estimation method and system based on implicit encoding neural network representation

The GLFormer model, integrating CNN and Transformer architectures with implicit encoding, addresses the challenges of occlusions and limited data in animal pose estimation, enhancing accuracy and efficiency in complex scenes.

CN119964205BActive Publication Date: 2025-07-15HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437759.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-15
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing animal pose estimation methods are not effective in complex scenarios, especially when the animal part is blocked, which affects the accuracy of behavior prediction, and existing models are prone to overfitting on small-scale data sets.

Method used

Combining the advantages of convolutional neural network (CNN) and Transformer, the hidden encoding neural network representation is adopted. By combining the encoder, codebook and decoder, the global and local characteristics of animal poses are learned, the global and local attention mechanisms are used for feature extraction, and the dependence between key points is represented by hidden vectors to optimize the pose estimation model.

Benefits of technology

It improves the accuracy and completeness of animal pose estimation, optimizes resource usage during training, reduces computing resource consumption, and can accurately capture key points in complex environments, solving the problem of pose incompleteness caused by occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964205B_ABST
    Figure CN119964205B_ABST
Patent Text Reader

Abstract

The present invention discloses an animal pose estimation method and system based on implicit encoded neural network representation. The method includes: obtaining an image of an animal with at least one key point occluded; inputting the obtained image into a feature extraction model to output an extracted feature map; inputting the extracted feature map into a trained implicit encoded neural network to output an estimated pose of the animal with at least one key point occluded. Among them, the training of the implicit encoded neural network includes: training a combined encoder, a codebook and a decoder, where the combined encoder is used to convert an animal pose into multiple latent variable features, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vectors; training a classification head, which is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to animal pose estimation based on deep learning. Specifically, it is an animal pose estimation method and system based on the representation of a hidden-coded neural network. Background Art

[0002] Animal pose estimation is an important research direction in computer vision. It locates the main joint points or parts of an animal through a computer system, analyzes the pose and actions of the animal in an image or video, and can provide a detailed interpretation of the animal's behavior. In many applications, animal pose estimation is the basis for tasks such as behavior recognition and tracking, health status monitoring, and reproductive behavior. Compared with traditional methods of wearing tags and sensors, deep learning methods have achieved optimal performance and enabled low-cost, non-contact, and non-invasive detection. In the field of bioscience, some researchers have built their own networks by adapting to the characteristics of tasks or datasets and achieved good results. For example, Biderman et al. used convolutional neural network stacks to propose a new network architecture LightningPose for video animal keypoint prediction; Walter et al. proposed a high-real-time multi-object animal pose detection tool Trex. In addition, some animal pose estimations directly adopt human pose estimation methods. For example, Mathis et al. developed an open-source tool DeepLabCut for keypoint detection of mice and fruit flies based on the human pose estimation method DeeperCut. The unlabeled automated annotation system OpenMonkeyStudio for rhesus macaque poses proposed by Bala et al. uses the human pose estimation algorithm CPM.

[0003] The method and process of animal pose estimation based on deep learning are similar to those of human pose estimation. The mainstream methods are the top-down paradigm and training based on keypoint heatmaps. The top-down method first obtains the bounding box of the animal, and then detects the keypoints in each bounding box to obtain the complete pose. The heatmap-based method refers to converting each keypoint marked in the image into the form of a Gaussian heatmap for training. When the model makes predictions, it converts the image features into a series of tensors based on the Gaussian distribution, performs the argmax operation on the tensors, and maps them to the original image to obtain the coordinate positions of the keypoints. The animal pose estimation model is mainly divided into two parts: the backbone network for feature extraction and the prediction head. The backbone network can be built according to the task as described above, or some classic models based on convolutional neural networks (CNNs) such as ResNet, HRNet, LiteHRNet, etc. can be used. In recent years, with the popularity of ViT (Vision Transformer), models based on the Transformer architecture have also been used in animal pose estimation tasks, such as Swin-Transformer, HRFormer, ViTPose, etc. The prediction head generally uses the paradigm proposed by SimpleBaseline, maps the feature map output by the backbone network into a keypoint heatmap, and then decodes and restores it to the coordinates of the original image. Some works such as DARK modulation smooth the distribution of the keypoint heatmap, making the final result obtained by decoding the predicted heatmap more accurate.

[0004] Due to the greater richness of animal species, more diverse pose characteristics, as well as more crowded, occluded situations and complex environmental conditions, the previous methods are limited in animal pose estimation. On the currently commonly used small-scale animal pose datasets, although CNN-based methods are easier to train, their prediction performance is often limited by the local receptive field and lacks the modeling of long-range dependence relationships. The models based on Transformer that learn well-converged global dependence relationships are highly data-intensive. Under relatively limited animal pose data, the model assumptions are too complex, the feature dimensions are too many, and it is easy to cause overfitting and deteriorate the effect. Some methods have explored how to combine the advantages of CNN and Transformer. For example, HRFormer has achieved a good balance between performance and the number of parameters and computational complexity, but its self-attention mechanism using local windows lacks the modeling of long-range dependence, and the information interaction between windows also needs to be improved.

[0005] In addition, the existing mainstream animal pose estimation paradigm only estimates the positions of key points visible in the image features. When the image instance is only partially visible or severely occluded, the model can only give the pose composed of the key points of the visible part, and has limited ability to guess the position and pose of the entire animal body, which may affect subsequent prediction of animal actions and behaviors. Summary of the Invention

[0006] To overcome the deficiencies of the above-mentioned prior art, the present invention provides an animal pose estimation method and system based on implicit encoding neural network representation, which combines the advantages of convolutional neural network (CNN) and Transformer, and aims to solve the challenges in the prior art regarding complex scenarios through an effective global and local feature extraction mechanism.

[0007] According to one aspect of the specification of the present invention, there is provided an animal pose estimation method based on implicit encoding neural network representation, including:

[0008] Obtaining an image of an animal with at least one key point occluded;

[0009] Inputting the obtained image into a feature extraction model to output an extracted feature map;

[0010] Inputting the extracted feature map into a trained implicit encoding neural network to output an estimated pose of an animal with at least one key point occluded; wherein, the training of the implicit encoding neural network includes: training a combined encoder, a codebook, and a decoder, where the combined encoder is used to convert an animal pose into multiple latent variable features, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vectors; training a classification head, where the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.

[0011] As a further technical solution, training the combined encoder, the codebook, and the decoder includes:

[0012] Projecting a pose into multiple latent vector features, where each latent vector feature encodes a sub-structure of a pose;

[0013] Discretizing the latent vector features using a shared codebook such that a pose can be represented by several discrete vectors in the codebook;

[0014] Inputting the discrete vectors in the codebook into the decoder to restore the original pose;

[0015] Jointly training the combined encoder, the codebook, and the decoder by minimizing the reconstruction error.

[0016] As a further technical solution, training the classification head includes:

[0017] Use two basic residual convolutional blocks to transform the output feature map of the backbone network, and then flatten and expand the feature plane through a linear projection layer to change its dimension;

[0018] Reconstruct the expanded feature into a matrix, and use a multi-layer perception mechanism to process the feature and output the classification score of the latent variable;

[0019] Use the cross-entropy loss and the distance loss between the predicted value and the ground truth of the key point position to train the classification head.

[0020] As a further technical solution, input the obtained image into a feature extraction model, including:

[0021] Use the bottleneck structure of the deep residual network for preliminary feature extraction;

[0022] Based on the preliminary extracted features, combine the global and local attention mechanisms to perform feature extraction on the global and local windows;

[0023] Use a feed-forward network to enhance the interaction of the extracted global and local feature information.

[0024] As a further technical solution, the feed-forward network uses a two-branch convolutional structure, including depthwise separable convolutions with different kernel sizes.

[0025] As a further technical solution, the method further includes:

[0026] After obtaining the image of an animal with at least one key point occluded, perform strided convolution on the image to reduce the resolution;

[0027] After obtaining the key point feature information of the animal pose, convert the key point feature information to the resolution of the original image and output the final predicted pose.

[0028] According to one aspect of the specification of the present invention, there is provided an animal pose estimation system based on the representation of the latent coding neural network, obtained by using the animal pose estimation method based on the representation of the latent coding neural network.

[0029] As a further technical solution, the system includes:

[0030] An image input module for obtaining an image of an animal with at least one key point occluded;

[0031] A feature extraction module for inputting the obtained image into a feature extraction model and outputting the extracted feature map;

[0032] A pose estimation module is configured to input the extracted feature map into a trained latent encoding neural network and output the estimated pose of an animal with at least one key point occluded. Wherein, the training of the latent encoding neural network includes: training a combined encoder, a codebook and a decoder. The combined encoder is used to convert an animal pose into multiple latent variable features, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vector; training a classification head, and the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.

[0033] According to one aspect of the specification of the present invention, there is provided an animal pose estimation device based on a latent encoding neural network representation, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the above-mentioned animal pose estimation method based on a latent encoding neural network representation.

[0034] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the above-mentioned animal pose estimation method based on a latent encoding neural network representation.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. The present invention combines GLFormer with an occlusion key point prediction method based on latent encoding, which can achieve more effective animal pose estimation. GLFormer provides strong feature extraction capabilities, making the latent encoding method more accurate and reliable when extracting pose information. The combination of the two not only improves the overall accuracy and integrity of pose estimation, but also optimizes the resource usage during training and improves efficiency.

[0037] 2. As a new animal pose estimation model, GLFormer combines the advantages of convolutional neural network (CNN) and Transformer, and has many significant advantages: (1) GLFormer can more effectively extract long-range and short-range features through an optimized global and local information extraction mode. This structure enables the model to accurately capture the key points of animals when dealing with complex environments (such as crowded and occluded scenes), thus improving the accuracy of pose estimation. (2) The model pays attention to computational efficiency in design. By integrating the advantages of CNN and Transformer, it performs well on small-scale animal pose datasets. By reducing the number of parameters and dependence on training data, GLFormer reduces the consumption of computing resources.

[0038] 3. The occlusion key point prediction method based on latent encoding can be used as a new post - processing mode for animal pose estimation and has the following advantages: (1) Without relying on prior assumptions, this method can capture the dependency relationships between key points by learning latent vectors. This representation method enables the model to still make reliable pose inferences when facing key point occlusions, thus solving the common pose incompleteness problem in traditional Top - down methods. (2) Through the combination of an encoder, a codebook, and a decoder, the latent encoding method can effectively compress and represent pose information, enabling each pose to be characterized by a latent vector. This not only improves computational efficiency but also reduces the complexity and memory footprint of the model. Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following - described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic flowchart of an animal pose estimation method based on latent - encoding neural network representation provided by an embodiment of the present invention.

[0041] Figure 2 It is an overall architecture diagram of GLFormer provided by an embodiment of the present invention.

[0042] Figure 3 It is a schematic structural design diagram of GLFormerBlock provided by an embodiment of the present invention.

[0043] Figure 4 It is a schematic diagram of a cross - window interaction mode for enhancing local information aggregation provided by an embodiment of the present invention.

[0044] Figure 5 It is a schematic diagram for training an encoder, a codebook, and a decoder provided by an embodiment of the present invention.

[0045] Figure 6 It is a schematic diagram for training a classification head provided by an embodiment of the present invention, where the * sign indicates frozen parameters during module training.

[0046] Figure 7 It is a schematic structural diagram of an encoder network provided by an embodiment of the present invention.

[0047] Figure 8 It is a schematic workflow diagram of a classification head provided by an embodiment of the present invention.

[0048] Figure 9Schematic diagram of the combination of GLFormer and the method for predicting key points with latent coding provided by the embodiments of the present invention.

[0049] Figure 10 Schematic diagram of an animal pose estimation system based on the representation of a neural network with latent coding provided by the embodiments of the present invention. Detailed implementation manners

[0050] Existing backbone networks show limitations in animal pose estimation tasks. Models based on convolutional neural networks (CNNs) are easy to train and converge on small-scale datasets, but their prediction performance is often limited by the local receptive field and cannot fully capture global information in complex scenarios. This makes CNN models perform poorly when dealing with complex backgrounds such as crowding and occlusion. In addition, although Transformer-based network models have better global information perception capabilities and can model long-range dependencies in complex scenarios, due to the high number of parameters and complexity of the Transformer model, overfitting problems are likely to occur on small-scale animal pose datasets, resulting in insufficient model generalization ability and thus affecting prediction accuracy. Therefore, existing animal pose estimation model architectures are difficult to maintain efficient performance in complex scenarios and data-scarce situations.

[0051] Currently, the mainstream animal pose estimation methods mainly rely on estimating the positions of visible key points in the image. When part of the animal's body is occluded or only partially visible, existing methods usually can only infer the pose of the visible part, and it is difficult to effectively predict the positions of the key points of the occluded part. This limitation not only leads to incomplete pose estimation but also affects the accuracy of subsequent behavior prediction. Especially in scenarios such as animal groups, complex natural environments, or experimental scenarios with occlusion, this visibility limitation reduces the practicality of pose estimation.

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Additionally, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. This combination is not restricted by the order of steps and / or the structural composition mode, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0053] An embodiment of the present invention provides an animal pose estimation method based on implicit encoding neural network representation, as Figure 1 shown. First, obtain an image of an animal with at least one key point occluded; then, input the obtained image into a feature extraction model to output an extracted feature map; then, input the extracted feature map into a trained implicit encoding neural network to output an estimated pose of the animal with at least one key point occluded; wherein, the training of the implicit encoding neural network includes: training a combined encoder, a codebook, and a decoder, the combined encoder is used to convert an animal pose into multiple latent variable features, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vector; training a classification head, the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.

[0054] An embodiment of the present invention proposes a novel animal pose estimation backbone network model called GLFormer, which combines the advantages of convolutional neural networks (CNNs) and Transformers, aiming to address the challenges in complex scenarios in the prior art through an effective global and local feature extraction mechanism. The following will be described in detail from the overall architecture and key modules.

[0055] The overall architecture of GLFormer in the embodiment of the present invention is as Figure 2 shown. The task of animal pose estimation is to detect the positions of K key points (such as the neck, knees, and hooves, etc.) from an input image with a size of H×W×3. The current mainstream method relies on the generation of key point heatmaps, that is, estimating the H'×W' heatmap of each key point, and the heatmap represents the confidence of the key point in the feature map. GLFormer follows a heatmap-based pipeline. First, the input image passes through two strided convolutions to reduce the resolution, and then feature extraction is performed by the backbone network (GLFormer). The extracted feature map generates K key point heatmaps through the head of pose estimation. Finally, based on these heatmaps, the key point positions are inferred and converted to the resolution of the original image to output the final predicted pose. The architecture of GLFormer includes multiple branches to represent features of different resolutions. In the early stage of the network, the bottleneck structure BottleNeck of the deep residual network ResNet is used for feature extraction because convolution has a good effect on extracting local features in the early stage. Next, GLFormerBlock is used for feature extraction in the subsequent stage, and combines a global and local information fusion mechanism to ensure that the high-resolution branch retains more detailed information in the final stage, helping with accurate pose estimation.

[0056] As Figure 3As shown in the figure, the GLFormerBlock is divided into two stages, including the global and local attention module Globaland local attention (GL-Attn) and the convolutional-based feed-forward network module Feed-Forward Network (ConvFFN). The feature map output by the convolutional backbone in the first stage is input into GL-Attn for feature extraction of the global and local windows, and then the convolutional-based dual-branch FFN enhances the interaction of the global and local feature information. Finally, the size and number of channels of the feature map output by the Block are the same as those of the input. Specifically, it can be expressed as follows, where X in represents the input feature, and X out represents the output feature:

[0057]

[0058] In the embodiments of the present invention, animal pose estimation requires the model to fully extract low-frequency global information and fine-grained local information to make reliable predictions on animal key points. GL-Attn (Global and local attention) is the key module for model feature extraction, and its global and local branches contain standard self-attention structures. After separating the channels of the input feature, they are respectively input into the global and local branches. Among them, the channel separation rate γ is a hyperparameter that can adjust the channel ratio of the two input branches to obtain better performance. The output features of the two branches are concatenated and then channel-shuffled, which helps to improve the feature representation ability of the model.

[0059] The standard self-attention mechanism has excellent long-range information modeling ability, but this also brings a considerable amount of Flops (floating-point operations per second) and memory costs. To reduce the number of parameters and computational volume, as Figure 3 shown in the figure, we perform average pooling on K and V for downsampling, aggregate the global information in the horizontal and vertical directions respectively, and then perform standard self-attention processing on Q, K, and V. Specifically, after linearly projecting the feature map of the input global branch, the obtained Q, K, V ∈ R M ×N×C , where M is the number of heads, N is the sequence length, and C is the number of channels. Convert K and V into the form of (C × M, H, W), perform average pooling on K along the W axis, pool it into a feature vector of size (C × M, H, 1), and then convert it into the shape of (M, H, C). Perform average pooling on V along the H axis, and similarly obtain the shape of (M, W, C). Multiply the Q matrix by the K matrix to obtain an attention map of size (M, N, H). After performing the softmax normalization operation on the attention map, multiply it by the V matrix that has been average-pooled along the H axis, and finally the shape of the output feature map is (M, N, C). As shown in the following formula, where x inDenote the input feature, FC denote the fully connected layer, Q g , K g , V g respectively denote Q, K, V generated by the global branch, AvgPool denote the average pooling operation, x Global denote the output feature of the global branch.

[0060]

[0061] The mode of the global branch can effectively generate a global receptive field and capture low-frequency global information, but it has insufficient ability to model high-frequency fine-grained local features. Therefore, GLFormer adopts a local window-based self-attention mechanism to capture local features. As Figure 3 shown, for the feature map x in ∈R M×N×C of the input local branch, where M is the number of heads, N is the sequence length, and C is the number of channels. Convert it into the shape of x in ∈R M×H×W×C , and perform padding around the feature Figure Four to ensure that its height H and width W can be divisible by the window size S. For converting the overall feature map into window-based feature vectors, x in →{x1, x2……x w}, with the shape of x w ∈R M×S^2×C , and after linear layer projection and dimension rearrangement, the obtained Q, K, V ∈ R M×S×S×C . For V, use depthwise separable convolution with a kernel size of 3×3 to perform local information aggregation with shared weights, and process Q and K to generate context-aware weights. Specifically, after performing local information aggregation on Q and K using two 3×3 depthwise separable convolutions, calculate the attention weight map Attn win , and then obtain context-aware weights with values between (-1, 1) through a series of pointwise convolutional mappings and Swish, Tanh activation functions. Different from ordinary Softmax normalization, the normalization method of GLFormer introduces stronger non-linear fitting ability. Finally, calculate the Hadamard product of Attn win and V after local information aggregation to generate the output of the local branch. As shown in the following formula, where x win denote the input feature of the window, FC denote the fully connected layer, Q l , K l , V l respectively denote Q, K, V generated by the global branch, DwConv3 denote 3×3 depthwise separable convolution, PwConv denote pointwise convolution, d denote the number of channels processed by each head, x outRepresents the output features of each window.

[0062]

[0063] For the feature information generated by each window, it is aggregated into the size of the local branch original input feature map and finally fused into the output feature map:

[0064]

[0065] The self-attention based on local windows in the local branch lacks connections between windows, which limits its modeling ability. Embodiments of the present invention fully consider the spatial invariance of convolution, that is, the convolution kernel shares weight parameters at all positions on the feature map, and propose a cross-window interaction mechanism for enhancing local information aggregation, as Figure 4 shown. It uses a double-branch convolution structure in the feed-forward network, including depthwise separable convolutions with different kernel sizes, to increase the ability to model information exchange between local windows. Specifically, for the feature x in input to the feed-forward network, after point convolution mapping, BN layer, and GELU activation function, it is input to the double-branch module that realizes window information interaction. This module consists of depthwise separable convolutions of 3×3 and 5×5 in parallel. The feature information aggregated in different ways is added after branching, and then input to point convolution mapping, BN layer, and GELU activation function to output x out . The specific implementation method is as shown in the formula:

[0066]

[0067] In embodiments of the present invention, GLFormer sets some specific configurations according to the examples of HRFormer. The number of modules, blocks, and heads in each stage of GLFormer is the same as that of HRFormer. For the number of channels, GLFormer-S refers to the example configuration of HRFormer-S, but GLFormer-B changes the number of channels in each branch of HRFormer-B from a multiple of 78 to 80 so that the number of channels received by the operations after channel separation is an integer. According to the convention of Swin-Transformer and HRFormer, the window size in the local branch is defaulted to 7×7.

[0068] In the occlusion key point prediction based on implicit coding in embodiments of the present invention, its overall process is described as follows. The occlusion key point prediction based on implicit coding can learn the dependence relationship between the position coordinates of key points in the representation learning stage without providing any prior assumptions. Its main purpose is to learn the true overall pose and represent each pose sub-structure and their mutual relationships with a series of latent vectors.

[0069] The occlusion key point prediction based on implicit encoding in the embodiments of the present invention is divided into two stages.

[0070] In the first stage, a combined encoder, a codebook, and a decoder are trained. As Figure 5 shown, a pose is projected into M implicit vector features, and each feature vector encodes a sub-structure of a pose. Then, the implicit vectors are discretized through a shared codebook. The implementation method is to add an index to each implicit vector, indexing to the codebook entry with the closest vector distance in the feature space, so that a pose can be represented by M discrete vectors in the codebook. These discrete vectors can be regarded as "key point prototypes" or "pose sub-patterns" learned through training. The space represented by the codebook is large enough to accurately represent all pose sub-structures. Then, these codebook vectors are input into the decoder to recover the original pose. The main training objective of this stage is the codebook and the decoder, which are jointly trained by minimizing the reconstruction error.

[0071] In the second stage, the parameters of the backbone network, the codebook, and the decoder are frozen, and a classification head Class Head is trained. As Figure 6 shown, the role of the classification head is to map the continuous features extracted by the backbone network into discrete implicit vector categories. Once the classification head determines the category to which each discrete feature belongs, this category will be used as an index to look up the corresponding discrete vector in the codebook. These discrete vectors are combined and then the key point feature information of the final pose is recovered through the decoder. The features processed by the decoder can be directly used to generate the two-dimensional coordinates (x, y) of the key points, and these coordinates are the final results of pose estimation.

[0072] Specifically, the following process is included:

[0073] (1) Training the encoder and decoder

[0074] Represent the original pose as G ∈ R K×D , where K is the number of body joints and D is the dimension of each joint. Here, D = 2 for two-dimensional poses, representing the relative coordinates (x, y) on the picture. It is necessary to train a combined encoder f e ( ) to convert a pose into M implicit variable features:

[0075]

[0076] where each implicit variable feature t i ∈ R H approximately corresponds to a sub-structure of the pose, and this sub-structure represents the interdependence of several key points. This representation has a lot of redundancy because different implicit variables may have overlapping connections, and the redundant representation makes it robust to the occlusion of key points. Figure 7Shows the structure and working process of the encoder network.

[0077] First, the positions of each key point are fed into a fully connected layer to map and increase the feature dimension. Then, the features are fed into a series of MLP-Mixer modules to deeply fuse the features of different key points. Finally, M latent variable features are extracted by applying a linear projection to all the key point features. Using the codebook C = (c1, …, c V ) T ∈R V×N defines the potential embedding space, where V is the number of codebook vectors. Use the embedding space to quantize each latent variable feature t i through nearest neighbor search, as shown in the following formula:

[0078]

[0079] Use q(t i ) to represent the index of the corresponding codebook entry. Then, the discrete vector (c q(t1) , c q(t2) , …, c q(tM) ) is input into the decoder network to recover the original pose:

[0080]

[0081] The structure of the decoder network is the reverse order of the encoder network, but uses a shallower series of MLP-Mixer networks (only containing one block). The encoder network, codebook, and decoder network are jointly trained by minimizing the following loss on the training dataset. In the formula, sg is to stop the gradient backpropagation, and β is a hyperparameter.

[0082]

[0083] (2) Training the classification head

[0084] The structure of the classification head Class Head is as Figure 8 shown, consisting of a series of residual convolutional blocks and MLP-Mixer layers.

[0085] Using the trained codebook and decoder, animal pose estimation can be regarded as a classification task. First, two basic residual convolutional blocks are used to transform the output feature map of the backbone network, and then the features are flattened and the dimension is changed through a linear projection layer, where Conv represents the residual convolutional module and FC represents the linear layer:

[0086]

[0087] Reconstruct the unfolded features into a matrix X f ∈RM×N , where N is the size H×W of the feature map. Then, 4 MLP-Mixer layers are used to process the features, and the hidden variable classification scores are output, which are then transformed into a probability distribution through Softmax. In the formula, M represents the MLP-Mixer operation:

[0088]

[0089] The training loss function of the classification head consists of two parts, namely the cross-entropy loss of classification and the distance loss between the predicted value of the key point position and the ground truth, as follows:

[0090]

[0091] where is the predicted category, and L is the hidden vector category generated by providing the ground truth of the key point position to the encoder. refers to the key point coordinates predicted by the model, and G refers to the true coordinates of the key points.

[0092] In summary, the present invention proposes an animal pose estimation method that combines GLFormer and an occlusion key point prediction method based on hidden coding. The trained GLFormer can be combined with a classification head, a codebook, and a decoder to predict the pose of the occluded animal. Figure 9 That is, the final structure diagram of the model and the schematic diagrams of the input image and the output image.

[0093] As Figure 9 shown, in the inference and prediction stage, after inputting an image with the feature information of some key points of the animal artificially occluded into GLFormer to extract the feature information, through the post-processing method based on the hidden coding representation, the key point positions of the occluded parts can be accurately predicted and corrected, and at the same time, the coherence of the pose is optimized.

[0094] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with a processor function. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, the embodiments of the present invention provide an animal pose estimation system based on hidden coding neural network representation, which is used to execute an animal pose estimation method based on hidden coding neural network representation in the above method embodiments.

[0095] See Figure 10, the system includes: an image input module for acquiring an image of an animal with at least one key point occluded; a feature extraction module for inputting the acquired image into a feature extraction model and outputting an extracted feature map; a pose estimation module for inputting the extracted feature map into a trained latent coding neural network and outputting an estimated pose of an animal with at least one key point occluded; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook, and a decoder, the combined encoder for converting an animal pose into multiple latent variable features, the codebook for providing discrete vectors, and the decoder for restoring key point feature information of the animal pose according to the determined discrete vectors; training a classification head, the classification head for mapping continuous features extracted by a backbone network into discrete latent vector categories.

[0096] An animal pose estimation system based on latent coding neural network representation provided by an embodiment of the present invention adopts Figure 10 several modules therein, combines the advantages of convolutional neural networks and Transformers, and aims to solve the challenges in the prior art regarding complex scenarios through an effective global and local feature extraction mechanism.

[0097] It should be noted that the system embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting corresponding functional modules, and its principle is basically the same as that of the above system embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above system embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solutions, improve the modules in the above system embodiment to obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.

[0098] Based on the same inventive concept as the above embodiment, an embodiment of the present invention further provides an animal pose estimation device based on latent coding neural network representation, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the above-mentioned animal pose estimation method based on latent coding neural network representation.

[0099] In an embodiment of the present invention, the memory can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or can also be a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present invention can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0100] In an embodiment of the present invention, the processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0101] Based on the same inventive concept as the above embodiments, the embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions cause the computer to execute the steps of the method for animal pose estimation based on implicit coding neural network representation.

[0102] The terms "comprising" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. An animal pose estimation method based on the representation of a hidden-coded neural network, characterized in that Including: Obtaining an image of an animal with at least one key point occluded; Inputting the obtained image into a feature extraction model, and performing preliminary feature extraction using the bottleneck structure of a deep residual network; the preliminarily extracted features are input into a global and local attention module for feature extraction of global and local windows, and then the global and local feature information is enhanced and interacted by a convolutional-based dual-branch FFN, and finally a feature map output by the Block; Inputting the preliminarily extracted features into a global and local attention module for feature extraction of global and local windows, including: separating the channels of the input features and respectively inputting them into global and local branches, where the channel separation rate is used to adjust the channel ratio of the two input branches, and the output features of the two branches are concatenated and then shuffled in channels; After performing a linear projection on the feature map of the input global branch, we obtain Q, K, V ∈ R M×N×C , where M is the number of heads, N is the sequence length, and C is the number of channels; convert K and V into the form of (C×M, H, W), perform average pooling on K along the W axis to obtain a feature vector of size (C×M, H, 1), and then convert it into the shape of (M, H, C); perform average pooling on V along the H axis, and similarly obtain the shape of (M, W, C); multiply the Q matrix by the K matrix to obtain an attention map of size (M, N, H), perform a softmax normalization operation on the attention map, and then multiply it by the V matrix after average pooling along the H axis to output the feature map; For the feature map x of the input local branch in ∈R M×N×C , convert it into the shape of x in ∈R M×H×W×C , and pad around the feature map; for converting the overall feature map into window-based feature vectors, x in →{x1, x2……xw}, with the shape of x w ∈R M×S^2×C , through linear layer projection and dimension rearrangement, obtain Q, K, V∈R M×S×S×C ; for V, use depthwise separable convolution with a kernel size of 3×3 for shared-weight local information aggregation, then perform matrix multiplication to calculate the attention weight map, and then obtain context-aware weights through a series of pointwise convolutional mappings and Swish and Tanh activation functions, calculate the Hadamard product of the attention weight map and V after local information aggregation, and generate the output of the local branch; Inputting the extracted feature map into a trained latent coding neural network, and outputting the estimated pose of an animal with at least one key point occluded; wherein, the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, the combined encoder is used to convert an animal pose into multiple latent variable features, each latent variable feature corresponds to a sub-structure of the pose, and the sub-structure represents the interdependence of several key points, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vectors; training a classification head, the classification head is used to map the continuous features extracted by the backbone network to discrete latent vector categories.

2. The animal pose estimation method based on the implicit encoding neural network representation according to claim 1, wherein Training a combined encoder, a codebook and a decoder, including: Projecting a pose into multiple latent vector features, each latent vector feature encoding a sub-structure of the pose; Discretizing the latent vector features using a shared codebook, so that a pose can be represented by several discrete vectors in the codebook; Inputting the discrete vectors in the codebook into the decoder to restore the original pose; Co-training the combined encoder, the codebook and the decoder by minimizing the reconstruction error.

3. The animal pose estimation method based on the representation of the latent-coded neural network according to claim 1, wherein Training a classification head, including: Using two basic residual convolutional blocks to transform the feature map output by the backbone network, and then flattening and expanding the feature plane through a linear projection layer and changing its dimension; Reconstructing the expanded features into a matrix, and processing the features using a multi-layer perception mechanism to output latent variable classification scores; Training the classification head using cross-entropy loss and the distance loss between the predicted value and the true value of the key point position.

4. The animal pose estimation method based on the implicit coding neural network representation according to claim 1, wherein, The method further includes: After obtaining an image of an animal with at least one key point occluded, performing strided convolution on the image to reduce the resolution; After obtaining the key point feature information of the animal pose, converting the key point feature information to the resolution of the original image and outputting the final predicted pose.

5. An animal pose estimation system based on implicit encoding neural network representation, characterized in that, Implemented by using the animal pose estimation method based on latent coding neural network representation according to any one of claims 1 to 4, including: An image input module for obtaining an image of an animal with at least one key point occluded; A feature extraction module for inputting the obtained image into a feature extraction model and outputting the extracted feature map; The pose estimation module is configured to input the extracted feature map into the trained latent encoding neural network and output the estimated pose of at least one animal with key points occluded; wherein, the training of the latent encoding neural network includes: training a combined encoder, a codebook, and a decoder, the combined encoder is used to convert an animal pose into multiple latent variable features, the codebook is used to provide discrete vectors, and the decoder is used to restore the key point feature information of the animal pose according to the determined discrete vectors; training a classification head, the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.

6. An animal pose estimation device based on the representation of a hidden-coded neural network, characterized in that, It includes a memory and a processor, the memory stores program instructions executed by the processor, and the processor invokes the program instructions to execute the steps of an animal pose estimation method according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of an animal pose estimation method according to any one of claims 1 to 4.