Animal attitude estimation method and system based on implicit coding neural network representation
By combining the advantages of hidden coding neural networks and CNN/Transformer, the occlusion and overfitting problems in complex scenarios in animal pose estimation are solved, and more efficient and accurate pose estimation is achieved.
Patent Information
- Application Number
- CN202510437759.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The prior art is difficult to effectively deal with the occlusion problem in complex scenarios in animal pose estimation, and the Transformer-based model is prone to overfitting on small-scale data sets, resulting in limited prediction performance.
Using a method based on hidden encoding neural network representation, combining the advantages of convolutional neural network (CNN) and Transformer, the classification head is trained by combining encoder, codebook and decoder to realize global and local feature extraction, and solve the problems of occlusion and overfitting.
It improves the overall accuracy and integrity of animal pose estimation, optimizes the use of training resources, reduces computing resource consumption, and enhances the model's prediction ability in complex scenarios.
Smart Images

Figure CN119964205A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular to animal posture estimation based on deep learning, and specifically to an animal posture estimation method and system based on latent coding neural network representation. Background Art
[0002] Animal posture estimation is an important research direction in computer vision. It locates the main joints or parts of animals through computer systems, analyzes the posture and movement of animals in images or videos, and can provide a detailed interpretation of animal behavior. In many applications, animal posture estimation is the basis for tasks such as behavior recognition and tracking, health status monitoring, and reproductive behavior. Compared with the traditional method of wearing tags and sensors, deep learning methods have achieved optimal performance and realized low-cost, non-contact, and non-destructive detection. In the field of biological sciences, some researchers have built networks by adapting the characteristics of tasks or data sets and achieved good results. For example, Biderman et al. used convolutional neural network stacking to propose a new network architecture LightningPose for video animal key point prediction; Walter et al. proposed a high-real-time multi-target animal posture detection tool Trex. Other animal posture estimation directly uses human posture estimation methods. For example, Mathis et al. developed an open source tool DeepLabCut based on the human posture estimation method DeeperCut for key point detection of mice and fruit flies. Bala et al. proposed a markerless rhesus monkey posture automatic annotation system OpenMonkeyStudio that uses the human posture estimation algorithm CPM.
[0003] The method and process of animal posture estimation based on deep learning are similar to those of human posture estimation. The mainstream methods are top-down paradigm and keypoint heatmap-based training. The top-down method first obtains the bounding box of the animal, and then detects the key points in each bounding box to obtain the complete posture. The heatmap-based method refers to converting each key point annotated in the image into a Gaussian heatmap for training. When the model predicts, it converts the image features into a series of tensors based on Gaussian distribution, performs argmax operations on the tensors, and maps them to the original image to obtain the key point coordinates. The animal posture estimation model is mainly divided into two parts: the backbone network for feature extraction and the prediction head. The backbone network can be built according to the task as described above, or some classic convolutional neural network (CNN)-based models such as ResNet, HRNet, LiteHRNet, etc. can be used. In recent years, with the popularity of ViT (Vision Transformer), animal posture estimation tasks have also begun to use models based on Transformer architecture, such as Swin-Transformer, HRFormer, ViTPose, etc. The prediction head generally uses the paradigm proposed by SimpleBaseline to map the feature map output by the backbone network into a key point heat map, and then decode it to restore the coordinates of the original image. Some works such as DARK modulation smooth the distribution of key point heat maps, making the final result obtained by decoding the prediction heat map more accurate.
[0004] Due to the richer species of animals, more diverse posture features, more crowded, occluded and complex environmental conditions, the previous methods have limited their effectiveness in animal posture estimation. On the current small-scale animal posture datasets, although CNN-based methods are easier to train, their prediction performance is often limited by the local receptive field and lacks modeling of long-range dependencies. Transformer-based models are highly data-intensive in learning well-converged global dependencies. Under relatively limited animal posture data, the model assumptions are too complex and the feature dimensions are too many, which can easily cause overfitting and poor results. Some methods have explored how to combine the advantages of CNN and Transformer. For example, HRFormer has achieved a good balance between performance and the amount of parameters and computation, but it uses the self-attention mechanism of local windows, lacks modeling of long-range dependencies, and the information interaction between windows is also worth improving.
[0005] In addition, the existing mainstream animal posture estimation paradigm is to only estimate the position of key points of visible image features. When the image instance is only partially visible or severely occluded, the model can only give the posture composed of the key points of the visible part, and has limited ability to guess the position and posture of the animal's entire body, which may affect the subsequent prediction of animal movements and behaviors. Summary of the invention
[0006] In order to overcome the above-mentioned shortcomings of the prior art, the present invention provides an animal posture estimation method and system based on hidden coding neural network representation, which combines the advantages of convolutional neural network (CNN) and Transformer, and aims to solve the challenges of complex scenes in the prior art through effective global and local feature extraction mechanisms.
[0007] According to one aspect of the present invention, there is provided a method for estimating an animal posture based on a hidden coding neural network representation, comprising: Obtain an image of the animal with at least one key point occluded; Inputting the acquired image into a feature extraction model and outputting an extracted feature map; The extracted feature map is input into a trained latent coding neural network, and an estimated posture of an animal with at least one key point blocked is output; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, wherein the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; and training a classification head, wherein the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
[0008] As a further technical solution, training a combined encoder, codebook and decoder includes: Project a posture into multiple latent vector features, each of which encodes a substructure of the posture; The latent vector features are discretized using a shared codebook so that a posture can be represented by several discrete vectors in the codebook; Input the discrete vector in the codebook into the decoder to restore the original posture; The combined encoder, codebook, and decoder are jointly trained by minimizing the reconstruction error.
[0009] As a further technical solution, training the classification head includes: Use two basic residual convolution blocks to transform the backbone network output feature map, and then use a linear projection layer to flatten the feature and change its dimension; Reconstruct the expanded features into a matrix, process the features using a multi-layer perception mechanism, and output the latent variable classification score; The classification head is trained using cross entropy loss and the distance loss between the predicted keypoint locations and the true values.
[0010] As a further technical solution, the acquired image is input into a feature extraction model, including: Use the bottleneck structure of the deep residual network for preliminary feature extraction; Based on the preliminary extracted features, global and local window feature extraction is performed in combination with global and local attention mechanisms; A feed-forward network is used to enhance the interaction between the extracted global and local feature information.
[0011] As a further technical solution, the feedforward network uses a dual-branch convolution structure, including depth-wise separable convolutions with different kernel sizes.
[0012] As a further technical solution, the method further includes: After acquiring an image of the animal with at least one key point occluded, performing a strided convolution on the image to reduce the resolution; After obtaining the key point feature information of the animal posture, the key point feature information is converted into the resolution of the original image, and the final predicted posture is output.
[0013] According to one aspect of the present invention, there is provided an animal posture estimation system based on hidden coding neural network representation, which is obtained by using the animal posture estimation method based on hidden coding neural network representation.
[0014] As a further technical solution, the system includes: An image input module, for acquiring an image of an animal with at least one key point occluded; A feature extraction module, used for inputting the acquired image into a feature extraction model and outputting an extracted feature map; A posture estimation module is used to input the extracted feature map into a trained latent coding neural network, and output the estimated posture of an animal with at least one key point occluded; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; training a classification head, the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
[0015] According to one aspect of the invention specification, there is provided an animal posture estimation device based on hidden coding neural network representation, comprising a memory and a processor, wherein the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the animal posture estimation method based on hidden coding neural network representation.
[0016] According to one aspect of the present specification, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions enable the computer to execute the steps of the animal posture estimation method based on hidden coding neural network representation.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention combines GLFormer with the occlusion key point prediction method based on implicit coding to achieve more effective animal posture estimation. GLFormer provides powerful feature extraction capabilities, making the implicit coding method more accurate and reliable in extracting posture information. The combination of the two not only improves the overall accuracy and integrity of posture estimation, but also optimizes resource usage during training and improves efficiency.
[0018] 2. As a new animal posture estimation model, GLFormer combines the advantages of convolutional neural networks (CNN) and Transformer and has many significant advantages: (1) GLFormer can more effectively extract long-range and short-range features through optimized global and local information extraction modes. This structure enables the model to accurately capture the key points of animals when dealing with complex environments (such as crowded and occluded scenes), thereby improving the accuracy of posture estimation. (2) The model is designed with a focus on computational efficiency. By integrating the advantages of CNN and Transformer, it performs well on small-scale animal posture datasets. By reducing the number of parameters and dependence on training data, GLFormer reduces the consumption of computing resources.
[0019] 3. The occluded key point prediction method based on hidden coding can be used as a new post-processing mode for animal posture estimation, with the following advantages: (1) This method can capture the dependency between key points by learning hidden vectors without relying on prior assumptions. This representation method enables the model to perform reliable posture inference when facing key point occlusion, thereby solving the problem of incomplete posture commonly seen in previous top-down methods. (2) By combining the encoder, codebook and decoder, the hidden coding method can effectively compress and represent posture information, so that each posture is represented by a hidden vector. This not only improves computational efficiency, but also reduces model complexity and memory usage. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction is given below to the drawings used in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of an animal posture estimation method based on hidden coding neural network representation is provided in an embodiment of the present invention.
[0022] Figure 2 The overall architecture diagram of GLFormer provided in an embodiment of the present invention.
[0023] Figure 3 A schematic diagram of the GLFormerBlock structure design provided in an embodiment of the present invention.
[0024] Figure 4 A schematic diagram of a cross-window interaction mode for enhancing local information aggregation provided by an embodiment of the present invention.
[0025] Figure 5 A schematic diagram of a training combination encoder, codebook and decoder provided in an embodiment of the present invention.
[0026] Figure 6 A schematic diagram of a training classification head provided in an embodiment of the present invention, wherein the * indicates that parameters are frozen during module training.
[0027] Figure 7 A schematic diagram of the structure of an encoder network provided in an embodiment of the present invention.
[0028] Figure 8 A schematic diagram of the workflow of a classification head provided in an embodiment of the present invention.
[0029] Fig. 9 A schematic diagram of the combination of GLFormer and the implicit coding key point prediction method provided in an embodiment of the present invention.
[0030] Fig.10 A schematic diagram of an animal posture estimation system based on hidden coding neural network representation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] Existing backbone networks show limitations in animal posture estimation tasks. Models based on convolutional neural networks (CNNs) are easy to train and converge on small-scale datasets, but their prediction performance is often limited by local receptive fields and cannot fully capture global information in complex scenes. This makes CNN models ineffective when dealing with complex backgrounds such as crowding and occlusion. In addition, although Transformer-based network models have better global information perception capabilities and can model long-range dependencies in complex scenes, due to the high number of parameters and complexity of Transformer models, they are prone to overfitting problems on small-scale animal posture datasets, resulting in insufficient model generalization capabilities, which in turn affects prediction accuracy. Therefore, it is difficult for existing animal posture estimation model architectures to maintain efficient performance in complex scenes and data scarcity.
[0032] The current mainstream animal posture estimation methods are mainly based on the position estimation of visible key points in the image. When part of the animal's body is occluded or only partially visible, existing methods can usually only infer the posture of the visible part, while it is difficult to effectively predict the position of the key points of the occluded part. This limitation not only leads to incomplete posture estimation, but also affects the accuracy of subsequent behavior prediction. Especially in animal groups, complex natural environments, or experimental scenarios with occlusion, this visibility limitation reduces the practicality of posture estimation.
[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention are arbitrarily combined with each other to form a new technical solution. This combination is not restricted by the sequence of steps and / or the structural composition mode, but must be based on the ability of ordinary technicians in this field to achieve. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that this combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0034] The embodiment of the present invention provides an animal posture estimation method based on hidden coding neural network representation, such as Figure 1As shown, first, an image of an animal with at least one key point blocked is obtained; then, the obtained image is input into a feature extraction model to output an extracted feature map; then, the extracted feature map is input into a trained latent coding neural network to output an estimated posture of the animal with at least one key point blocked; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; training a classification head, the classification head is used to map the continuous features extracted by the backbone network to discrete latent vector categories.
[0035] The embodiment of the present invention proposes a new backbone network model for animal posture estimation, called GLFormer, which combines the advantages of convolutional neural network (CNN) and Transformer, and aims to solve the challenges of complex scenes in the prior art through effective global and local feature extraction mechanisms. The following is a detailed description of the overall architecture and key modules.
[0036] The overall architecture of GLFormer described in the embodiment of the present invention is as follows Figure 2 As shown in the figure. The task of animal pose estimation is to detect the positions of K key points (such as neck, knees, and hooves) from an input image of size H×W×3. The current mainstream method relies on the generation of key point heat maps, that is, estimating the H'×W' heat map of each key point, which represents the confidence of the key point in the feature map. GLFormer follows a heat map-based pipeline. First, the input image is reduced in resolution by two strided convolutions, and then the backbone network (GLFormer) performs feature extraction. The extracted feature map generates heat maps of K key points through the head of pose estimation. Finally, the key point positions are inferred based on these heat maps and converted to the resolution of the original image to output the final predicted pose. The architecture of GLFormer includes multiple branches to represent features of different resolutions. In the early stages of the network, the bottleneck structure BottleNeck of the deep residual network ResNet is used for feature extraction because convolution has a good effect on extracting local features in the early stages. Next, GLFormerBlock is used for feature extraction in subsequent stages and combines the fusion mechanism of global and local information to ensure that the high-resolution branch retains more detail information in the final stage, facilitating accurate pose estimation.
[0037] like Figure 3As shown in Figure 1, GLFormerBlock is divided into two stages, including the global and local attention module Globaland local attention (GL-Attn) and the convolution-based feed-forward network module Feed-Forward Network (ConvFFN). The feature map output by the convolution backbone in the first stage is input into GL-Attn for global and local window feature extraction, and then the convolution-based dual-branch FFN enhances the interaction of global and local feature information. Finally, the feature map size and number of channels output by the Block are consistent with the input. Specifically, it can be expressed as follows, where X in represents the input features, X out Represents the output features:
[0038] In an embodiment of the present invention, animal posture estimation requires the model to fully extract low-frequency global information and fine-grained local information to reliably predict the key points of the animal. GL-Attn (Global and local attention) is a key module for model feature extraction. The global and local branches it contains are standard self-attention structures. The channels of the input features are separated and input into the global and local branches respectively, where the channel separation rate γ is a hyperparameter that can adjust the channel ratio of the input two branches to obtain better performance. The output features of the two branches are spliced and then shuffled, which helps to improve the feature representation ability of the model.
[0039] The standard self-attention mechanism has superior long-range information modeling capabilities, but this also brings considerable Flops (floating point operations per second) and memory costs. In order to reduce parameters and computation, such as Figure 3 As shown in , we downsample K and V by average pooling, aggregate global information in the horizontal and vertical directions respectively, and then perform standard self-attention on Q, K, and V. Specifically, after linearly projecting the feature map of the input global branch, we obtain Q, K, V ∈ R M ×N×C , where M is the number of heads, N is the sequence length, and C is the number of channels. Convert K and V into the form of (C×M, H, W). K performs average pooling on the W axis, pooling into a feature vector of size (C×M, H, 1), and then converts it into a shape of (M, H, C). V performs average pooling on the H axis, and similarly obtains a shape of (M, W, C). Multiply Q by the K matrix to obtain an attention map of size (M, N, H). After performing softmax normalization on the attention map, multiply it with the V matrix after average pooling on the H axis, and the final output feature map shape is (M, N, C). As shown in the following formula, where x in represents input features, FC represents the fully connected layer, Qg , K g 、V g They represent Q, K, and V generated by the global branch, AvgPool represents the average pooling operation, and x Global Represents the global branch output feature.
[0040]
[0041] The global branch model can effectively generate a global receptive field and capture low-frequency global information, but it lacks the ability to model high-frequency fine-grained local features. To this end, GLFormer uses a local window-based self-attention mechanism to capture local features. Figure 3 As shown, for the feature map x of the input local branch in ∈R M×N×C , where M is the number of heads, N is the sequence length, and C is the number of channels. Convert it to x in ∈R M×H×W×C The shape and characteristics Figure 4 Padding is performed around to ensure that its height H and width W are divisible by the window size S. For converting the overall feature map into a window-based feature vector, x in →{x1,x2……x w}, with shape x w ∈R M×S^2×C , after linear layer projection and dimension rearrangement, we get Q, K, V∈R M×S×S×C For V, a depth-wise separable convolution with a kernel size of 3×3 is used to aggregate local information of shared weights, and Q and K are processed to generate context-aware weights. Specifically, after using two 3×3 depth-wise separable convolutions to aggregate local information of Q and K, a matrix multiplication is performed to calculate the attention weight map Attn win , and then through a series of point convolution mapping and Swish, Tanh activation functions to obtain context-aware weights, the value is between (-1,1). Different from the ordinary Softmax normalization, GLFormer's normalization method introduces stronger nonlinear fitting capabilities. Finally, Attn is calculated win The Hadamard product of V after local information aggregation generates the output of the local branch. As shown in the following formula, where x win represents the input features of the window, FC represents the fully connected layer, Q l , K l 、V l They represent Q, K, and V generated by the global branch, respectively. DwConv3 represents 3×3 depth-separable convolution. PwConv represents point convolution. d represents the number of channels processed by each head. out Represents the output features of each window.
[0042]
[0043] For the feature information generated by each window, it is aggregated into the size of the original input feature map of the local branch, and finally fused into the output feature map:
[0044] The local window-based self-attention in the local branch lacks connections between windows, which limits its modeling ability. The embodiment of the present invention fully considers the spatial invariance of convolution, that is, the convolution kernels share weight parameters at all positions on the feature map, and proposes a cross-window interaction mechanism to enhance local information aggregation, such as Figure 4 It uses a dual-branch convolution structure in the feedforward network, including depth-separable convolutions with different kernel sizes to increase the ability to model information exchange between local windows. Specifically, for the feature x input to the feedforward network in After point convolution mapping, BN layer and GELU activation function, it is input into the dual-branch module for realizing window information interaction. The module is composed of 3×3 and 5×5 depth-separable convolutions in parallel. The feature information aggregated in different ways is added after the branch, and then input into point convolution mapping, BN layer and GELU activation function to output x out The specific implementation method is shown in the formula:
[0045] In the embodiment of the present invention, GLFormer sets some specific configurations according to the instance of HRFormer. The number of modules, blocks and heads in each stage of GLFormer is the same as that of HRFormer. For the number of channels, GLFormer-S refers to the instance configuration of HRFormer-S, but GLFormer-B changes the number of channels in each branch of HRFormer-B from a multiple of 78 to 80, so that the number of channels received by the operation after channel separation is an integer. According to the conventions of Swin-Transformer and HRFormer, the window size in the local branch defaults to 7×7.
[0046] In the occlusion key point prediction based on implicit coding in the embodiment of the present invention, the overall process is described as follows. The occlusion key point prediction based on implicit coding can learn the dependency between the key point position coordinates in the representation learning stage without providing any prior assumptions. Its main purpose is to learn the real overall posture and use a series of potential vectors to represent each posture substructure and their relationship.
[0047] The occlusion key point prediction based on implicit coding in the embodiment of the present invention is divided into two stages.
[0048] The first stage trains a combined encoder, codebook, and decoder. Figure 5 As shown in the figure, a posture is projected into M latent vector features, and each feature vector encodes a substructure of the posture. The latent vectors are then discretized through a shared codebook. This is achieved by adding an index to each latent vector, indexing it to the codebook entry with the closest vector distance in the feature space, so that a posture can be represented by M discrete vectors in the codebook. These discrete vectors can be regarded as "key point prototypes" or "posture sub-patterns" learned through training. The space represented by the codebook is large enough to accurately represent all posture substructures. These codebook vectors are then input into the decoder to restore the original posture. The main training targets at this stage are the codebook and decoder, which are trained together by minimizing the reconstruction error.
[0049] In the second stage, the parameters of the backbone network, codebook and decoder are frozen, and the classification head is trained. Figure 6 As shown in the figure, the role of the classification head is to map the continuous features extracted by the backbone network into discrete latent vector categories. Once the classification head determines the category to which each discrete feature belongs, the category will be used as an index to find the corresponding discrete vector in the codebook, and these discrete vectors will be combined and then restored through the decoder to restore the key point feature information of the final posture. The features processed by the decoder can be directly used to generate the two-dimensional coordinates (x, y) of the key points, which are the final results of the posture estimation.
[0050] Specifically, the process includes the following: (1) Training encoder and decoder Denote the original pose as G∈R K×D , where K is the number of body joints, D is the dimension of each joint, and D = 2 is the two-dimensional pose, which represents the relative coordinates (x, y) on the image. It is necessary to train a combined encoder f e ( ) to convert a posture into M latent variable features:
[0051] Each latent variable feature t i ∈R H The approximation corresponds to a substructure of the pose that represents the interdependencies of several keypoints. This representation has a lot of redundancy because different latent variables may have overlapping connections, and the redundant representation makes it robust to occlusion of keypoints. Figure 7 The structure and workflow of the encoder network are shown.
[0052] First, the position of each key point is fed into a fully connected layer mapping to increase the feature dimension, and then the feature is fed into a series of MLP-Mixer modules to deeply fuse the features of different key points. Finally, M latent variable features are extracted by applying linear projection to all key point features. Use codebook C = (c1, …, c V ) T ∈R V×N Define a potential embedding space, where V is the number of codebook vectors. Use the embedding space to find the nearest neighbor for each latent variable feature t i Quantify as shown below:
[0053] Use q(t i ) to represent the index of the corresponding codebook entry. Then the discrete vector (c q(t1) , c q(t2) , …,c q(tM) ) Input the decoder network to restore the original pose:
[0054] The decoder network structure is the reverse order of the encoder network, but uses a shallower MLP-Mixer series network (containing only one block). The encoder network, codebook, and decoder network are trained together by minimizing the following loss on the training dataset. In the formula, sg is the stop gradient backpropagation, and β is a hyperparameter.
[0055]
[0056] (2) Training the classification head The structure of the Class Head is as follows Figure 8 As shown, it consists of a series of residual convolution blocks and MLP-Mixer layers.
[0057] Using the trained codebook and decoder, animal posture estimation can be regarded as a classification task. First, two basic residual convolution blocks are used to transform the output feature map of the backbone network, and then the features are flattened and their dimensions are changed through a linear projection layer, where Conv represents the residual convolution module and FC represents the linear layer:
[0058] Reconstruct the expanded features into matrix X f ∈R M×N , N is the feature map size H×W. Then use 4 MLP-Mixer layers to process the features, output the latent variable classification scores, and then convert them into probability distributions through Softmax. In the formula, M represents the MLP-Mixer operation:
[0059] The training loss function of the classification head consists of two parts, namely the cross entropy loss of classification and the distance loss between the predicted value and the true value of the key point position, as shown below:
[0060] in is the predicted category, and L is the latent vector category generated by providing the encoder with the true value of the key point position. refers to the key point coordinates predicted by the model, and G refers to the real coordinates of the key points.
[0061] In summary, the present invention proposes an animal posture estimation method combining GLFormer and an occlusion key point prediction method based on latent coding. The trained GLFormer can be combined with a classification head, a codebook and a decoder to predict the posture of the occluded animal. Fig. 9 That is the final structure diagram of the model and the input and output pictures.
[0062] like Fig. 9 As shown in the figure, in the inference and prediction stage, after the feature information of some key points of the animal is artificially occluded and input into GLFormer to extract the feature information, a post-processing method based on latent coding representation is used to accurately predict and correct the position of the key points of the occluded part, while optimizing the consistency of the posture.
[0063] The implementation basis of each embodiment of the present invention is to implement programmed processing through a device with a processor function. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this reality, on the basis of the above embodiments, an embodiment of the present invention provides an animal posture estimation system based on hidden coding neural network representation, which is used to execute an animal posture estimation method based on hidden coding neural network representation in the above method embodiment.
[0064] See also Fig.10The system includes: an image input module for acquiring an image of an animal with at least one key point blocked; a feature extraction module for inputting the acquired image into a feature extraction model and outputting an extracted feature map; a posture estimation module for inputting the extracted feature map into a trained latent coding neural network and outputting an estimated posture of the animal with at least one key point blocked; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; training a classification head, the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
[0065] An animal posture estimation system based on hidden coding neural network representation provided by an embodiment of the present invention adopts Fig.10 Several modules in it combine the advantages of convolutional neural networks and Transformer, aiming to solve the challenges of complex scenes in existing technologies through effective global and local feature extraction mechanisms.
[0066] It should be noted that the system embodiments provided by the present invention are not only used to implement the methods in the above-mentioned method embodiments, but also used to implement the methods in other method embodiments provided by the present invention. The only difference lies in the setting of corresponding functional modules, and the principles thereof are basically the same as the principles of the above-mentioned system embodiments provided by the present invention. As long as technical personnel in this field refer to the specific technical solutions in other method embodiments on the basis of the above-mentioned system embodiments, obtain corresponding technical means and technical solutions composed of these technical means by combining technical features, and on the premise of ensuring the practicality of the technical solutions, they will improve the modules in the above-mentioned system embodiments to obtain corresponding system class embodiments for implementing the methods in other method class embodiments.
[0067] Based on the same inventive concept as the above-mentioned embodiment, an embodiment of the present invention also provides an animal posture estimation device based on hidden coding neural network representation, including a memory and a processor, the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the animal posture estimation method based on hidden coding neural network representation.
[0068] In an embodiment of the present invention, the memory may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or a volatile memory (volatile memory), such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in an embodiment of the present invention may also be a circuit or any other device that can implement a storage function, for storing program instructions and / or data.
[0069] In the embodiments of the present invention, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0070] Based on the same inventive concept as the above-mentioned embodiment, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the steps of the animal posture estimation method based on hidden coding neural network representation.
[0071] The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for estimating animal posture based on hidden coding neural network representation, characterized in that: include: Obtain an image of the animal with at least one key point occluded; Inputting the acquired image into a feature extraction model and outputting an extracted feature map; The extracted feature map is input into a trained latent coding neural network, and an estimated posture of an animal with at least one key point blocked is output; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, wherein the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; and training a classification head, wherein the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
2. The method for estimating animal posture based on hidden coding neural network representation according to claim 1, characterized in that: Training the combined encoder, codebook, and decoder includes: Project a posture into multiple latent vector features, each of which encodes a substructure of the posture; The latent vector features are discretized using a shared codebook so that a posture can be represented by several discrete vectors in the codebook; Input the discrete vector in the codebook into the decoder to restore the original posture; The combined encoder, codebook, and decoder are jointly trained by minimizing the reconstruction error.
3. The method for estimating animal posture based on hidden coding neural network representation according to claim 1, characterized in that: Train the classification head, including: Use two basic residual convolution blocks to transform the backbone network output feature map, and then use a linear projection layer to flatten the feature and change its dimension; Reconstruct the expanded features into a matrix, process the features using a multi-layer perception mechanism, and output the latent variable classification score; The classification head is trained using cross entropy loss and the distance loss between the predicted keypoint locations and the true values.
4. The method for estimating animal posture based on hidden coding neural network representation according to claim 1, characterized in that: Inputting the acquired image into a feature extraction model comprises: Use the bottleneck structure of the deep residual network for preliminary feature extraction; Based on the preliminary extracted features, global and local window feature extraction is performed in combination with global and local attention mechanisms; A feed-forward network is used to enhance the interaction between the extracted global and local feature information.
5. The method for estimating animal posture based on hidden coding neural network representation according to claim 4, characterized in that: The feed-forward network uses a dual-branch convolutional structure, including depth-wise separable convolutions with different kernel sizes.
6. The method for estimating animal posture based on hidden coding neural network representation according to claim 1, characterized in that: The method further comprises: After acquiring an image of the animal with at least one key point occluded, performing a strided convolution on the image to reduce the resolution; After obtaining the key point feature information of the animal posture, the key point feature information is converted into the resolution of the original image, and the final predicted posture is output.
7. An animal posture estimation system based on hidden coding neural network representation, characterized in that: It is obtained by using an animal posture estimation method based on hidden coding neural network representation as described in any one of claims 1 to 6.
8. The animal posture estimation system based on hidden coding neural network representation according to claim 7, characterized in that: include: An image input module, for acquiring an image of an animal with at least one key point occluded; A feature extraction module, used for inputting the acquired image into a feature extraction model and outputting an extracted feature map; A posture estimation module is used to input the extracted feature map into a trained latent coding neural network, and output the estimated posture of an animal with at least one key point occluded; wherein the training of the latent coding neural network includes: training a combined encoder, a codebook and a decoder, the combined encoder is used to convert an animal posture into multiple latent variable features, the codebook is used to provide a discrete vector, and the decoder is used to restore the key point feature information of the animal posture according to the determined discrete vector; training a classification head, the classification head is used to map the continuous features extracted by the backbone network into discrete latent vector categories.
9. An animal posture estimation device based on hidden coding neural network representation, characterized in that: It includes a memory and a processor, wherein the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of an animal posture estimation method based on hidden coding neural network representation as described in any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the steps of an animal posture estimation method based on hidden coding neural network representation as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-view-angle information fusion human body posture estimation and spatial positioning method
CN113780205A
CNN and transformer combined SR-Fuse crack image segmentation method
CN117853500A
Lightweight remote sensing change detection method and system based on double-time-image remote sensing image
CN118334532A
Anti-shielding animal attitude estimation method based on online distillation
CN119296135A
Remote sensing image segmentation method and device, medium and equipment
CN119625302A
Cited By
Online adaptation method for compensating for domain offset in on-orbit non-cooperative spacecraft pose tracking
CN121095272A