A method and system for single-view animal 3D pose estimation

By improving the combination of the HRNet backbone network and the weakly supervised learning module, and using semantic prototypes and attention mechanisms to perform cross-scale feature fusion to generate pseudo-labeled data, the problem of high dependence on labeled data and human experience in 3D pose estimation of wild animals is solved, thereby improving the accuracy and performance of pose estimation.

CN120726133BActive Publication Date: 2025-11-18JIANGXI ACAD OF FORESTRY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511223747.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-18
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies for animal 3D pose estimation in wild scenarios suffer from problems such as high dependence on labeled data, high dependence on human experience, and low performance limits, making it difficult to accurately identify species, especially in complex backgrounds and natural animal movement situations.

Method used

We employ an improved HRNet backbone network for feature extraction, perform cross-scale feature fusion through semantic prototypes and attention mechanisms, and combine a weakly supervised learning module to generate pseudo-labeled data using a small amount of 3D labeled data to construct a closed-loop optimization mechanism. We also utilize similarity, geometric constraints, and ranking loss to improve the accuracy of pose estimation.

Benefits of technology

With limited 3D labeled data and a large amount of unlabeled 2D data, it improves the accuracy and performance ceiling of 3D pose estimation, and can adaptively adjust the spatial attention calculation granularity to achieve efficient animal 3D pose recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726133B_ABST
    Figure CN120726133B_ABST
Patent Text Reader

Abstract

The application discloses a kind of single-view animal 3D posture estimation method and system, comprising: backbone network receives 2D image data, the backbone network is trained, obtains trained backbone network;The output of backbone network is connected with 3D posture estimation network, constructs weakly supervised learning module, uses 3D labeled data to train weakly supervised learning module, forms initial model, uses initial model to predict 2D unlabeled data, obtains pseudo-labeled data set, and real 3D labeled data is merged to form training set, and weakly supervised learning module is trained;Real-time acquisition single-view animal image is input into trained weakly supervised learning module and carries out single-view animal 3D posture estimation;The application has the advantages that based on a large number of 2D unlabeled data, a small amount of 3D labeled data, under the premise of small dependence on labeled data, simultaneously not depending on artificial experience, effectively improve the performance upper limit on 3D posture estimation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D pose estimation, and specifically to a method and system for 3D pose estimation of a single-view animal. Background Technology

[0002] Images and videos automatically captured by cameras deployed in the wild need to be able to automatically identify the specific species of animals appearing within them with high accuracy. However, in wild settings, due to complex backgrounds and the animals' natural movements during filming, they often exhibit complex postures when entering and leaving the monitored area, making species identification difficult. With the aid of animal posture information, it is possible to more effectively and accurately identify animal species in images and videos captured by a single camera.

[0003] There are various existing methods for animal 3D pose estimation. The following evaluation of existing methods focuses on three dimensions: the degree of dependence on labeled data, the reliance on human experience, and the performance ceiling.

[0004] 1) Traditional methods based on model fitting

[0005] Overview: A 3D template model of an animal (e.g., a geometric model, a skeletal model) is pre-built. Then, an optimization algorithm is used to match the model with features such as contours and key points in a 2D image, thereby estimating 3D pose parameters such as rotation, translation, and joint angles.

[0006] Advantages: It has a clear physical meaning and can still function even when training data is insufficient, meaning it has little dependence on labeled data.

[0007] Disadvantages: It has poor generalization ability to complex postures and requires artificial design of features. It is difficult to handle animal species with large differences in appearance, that is, it relies on human experience and has a low performance ceiling.

[0008] 2) Traditional methods based on geometric features

[0009] Overview: Extract geometric features from two-dimensional images, such as distances, angles, and symmetry between joints, and then combine them with perspective projection principles (such as pinhole camera models) and prior knowledge (such as animal body proportions) to reconstruct three-dimensional structures.

[0010] Advantages: It does not require a large amount of labeled data, making it suitable for simple scenarios, i.e., it has little dependence on labeled data.

[0011] Disadvantages: It heavily relies on prior knowledge, and its accuracy drops significantly under complex perspectives or asymmetric postures, meaning it depends on human experience and has a low performance ceiling.

[0012] 3) Deep learning methods based on direct regression

[0013] Overview: Convolutional Neural Networks (CNNs) are used to directly regress the 3D joint coordinates of animals from images, typically using an "end-to-end" training method.

[0014] Advantages: It can automatically learn image features and has strong adaptability to complex poses, meaning it does not rely on human experience and has a high performance ceiling.

[0015] Disadvantages: It requires a large amount of 3D labeled data, while animal 3D datasets are relatively scarce, and the labeling cost is high, meaning it is highly dependent on labeled data.

[0016] 4) Deep learning methods based on 2D to 3D enhancement

[0017] Overview: First, predict the 2D key points of the animal in the image, and then use geometric transformations (such as triangulation and perspective projection) or deep learning models to elevate the 2D coordinates to 3D space.

[0018] Advantages: It can leverage abundant 2D annotation data to reduce dependence on 3D data, meaning it has a low dependence on annotation data.

[0019] Disadvantages: The mapping from 2D to 3D is ambiguous and requires additional constraints (such as prior knowledge of limb length), which means it relies on human experience and has a low performance ceiling.

[0020] 5) Deep learning methods based on parameterized models

[0021] Overview: The animal's 3D pose is represented as a parametric model (such as a skeletal model or a statistical shape model). The model's parameters (such as joint angles and body size) are predicted by a network, thereby generating the 3D pose.

[0022] Advantages: It reduces the difficulty of directly predicting 3D pose and can generate smooth and reasonable 3D poses, meaning it has less dependence on labeled data.

[0023] Disadvantages: Model building requires specialized biomechanical knowledge and has limited generalization ability for rare postures, i.e., it relies on human experience and has a low performance ceiling.

[0024] 6) Generative Adversarial Deep Learning Methods

[0025] Overview: Generate plausible 3D poses from single-view images by learning the latent distribution of animal poses using generative adversarial networks (GANs) or variational autoencoders (VAEs).

[0026] Advantages: It can generate diverse poses, alleviating the problem of insufficient data, that is, it does not rely on human experience and has little dependence on labeled data.

[0027] Disadvantages: Training is difficult, and the generated poses may lack geometric consistency, meaning the performance ceiling is low.

[0028] In summary, most existing methods suffer from problems such as reliance on human experience, low performance ceiling, and high dependence on labeled data. Summary of the Invention

[0029] The technical problem to be solved by this invention is how to effectively improve the performance ceiling of 3D pose estimation accuracy based on the existing large amount of unlabeled 2D data and small amount of labeled 3D data, while having little dependence on labeled data and without relying on human experience.

[0030] This invention solves the above-mentioned technical problems through the following technical means: a method for single-view animal 3D pose estimation, comprising:

[0031] S1. The backbone network receives 2D image data. The backbone network is based on the HRNet model and is improved. The feature maps output by the previous branch of the fusion layer (FuseLayer) or transition layer (TransitionLayer) of the HRNet model are first clustered to generate semantic prototypes, then an association matrix is ​​generated for mutual guidance between different resolutions, then an attention mechanism is formed by channel distribution entropy to weight the features of different resolutions, and finally the branches of different resolutions are fused through semantic association and spatial attention to obtain fused features. The backbone network is then trained to obtain the trained backbone network.

[0032] S2. Fix the backbone network parameters, connect the output of the backbone network to the 3D pose estimation network to construct a weakly supervised learning module, train the weakly supervised learning module using 3D labeled data to form an initial model, use the initial model to predict 2D unlabeled data to obtain a pseudo-labeled dataset, merge the pseudo-labeled dataset and the real 3D labeled data to form a training set, train the weakly supervised learning module, and obtain the trained weakly supervised learning module.

[0033] S3. Input the real-time acquired single-view animal images into the trained weakly supervised learning module to perform single-view animal 3D pose estimation.

[0034] Furthermore, the working process of the backbone network is as follows:

[0035] S11. Feature maps output by the network branch at the previous step of the fusion or transition layer in the HRNet model. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ;use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ;

[0036] S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ;

[0037] S13. Calculate the channel distribution entropy. ,in, Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. Features of each channel; design a dynamic convolution kernel generator to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, and process the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. The formula for the dynamic convolution kernel generator is expressed as follows: ,in, The learnable threshold, Indicates spatial location Kernel size at the location; Indicates if, Indicates other;

[0038] S14, Utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. Feature map With final attention weight Element-wise multiplication yields attention-enhanced features. Then , After upsampling and The first-scale fusion feature is obtained by splicing. ,Will , After downsampling and The splicing yields the third-scale fusion features. Fusion features at the second scale .

[0039] Furthermore, S12 includes:

[0040] The prototype feature maps at each scale are unified to an intermediate scale, and the corresponding feature map is: , , ,in, Indicates downsampling, Indicates upsampling, The length and width of the prototype feature map at an intermediate scale;

[0041] The guidance weights of the high-resolution prototype on the low-resolution prototype are calculated using the following formula:

[0042]

[0043] The guidance weights of the low-resolution prototype to the high-resolution prototype are calculated using the following formula:

[0044]

[0045] in, Representation of feature map The One channel, Representation of feature map The One channel, 1≤ ≤N, 1≤ ≤N, The length, width, and channel are respectively... , , Feature map ; Represents spatial dimension Find the average. The number of cluster centers. For normalized exponential functions, The first term represents the guiding weight of the high-resolution prototype on the low-resolution prototype. Line 1 Column elements, The first term represents the guiding weight of the low-resolution prototype on the high-resolution prototype. Line 1 Column elements.

[0046] Furthermore, the aforementioned utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. ,include:

[0047] Third-scale channel attention weights MLP stands for Multilayer Perceptron. express Activation function, first-scale channel attention weights Second-scale channel attention weights Generate the final attention weights. , k=1,2,3, ⊙ is the Hadamard product.

[0048] Furthermore, the process of training the backbone network is as follows:

[0049] A loss function is constructed, the parameters of the backbone network are adjusted, and the value of the loss function is calculated. Training stops when the value of the loss function is minimized, resulting in a trained backbone network. The loss function is a contrastive loss function, and its formula is as follows: , , , All are weighting coefficients; For similarity loss and ,in, For expectations; This is the pose geometric feature vector of the current frame. To and The corresponding positive sample feature vectors are generally feature vectors from adjacent frames, or intermediate sample feature vectors generated through pose interpolation. To and The corresponding negative sample feature vector; For temperature coefficient, Indicates the first One negative sample, The number of negative samples. For geometric constraint loss and , Let be the physiological constraint threshold for species s. Indicates two postures and Euclidean distance in 3D space Represents the L2 norm; For sorting loss and , , , These are the feature vectors of three adjacent frames.

[0050] Furthermore, the loss function also includes a temporal loss function, which uses the output of the backbone network of three consecutive frames as the input of the RAFT network, calculates the bidirectional reprojection loss, bidirectional optical flow inverse loss, optical flow difference L1 / L2 loss, and optical flow gradient smoothing loss, and then weights and fuses them as the temporal consistency loss.

[0051] Furthermore, the 3D pose estimation network is a stacked multi-layer convolutional network, and the output of the 3D pose estimation network is the 3D joint coordinates.

[0052] The present invention also provides a system for single-view animal 3D pose estimation, comprising:

[0053] The self-supervised training module is used for the backbone network to receive 2D image data. The backbone network is based on the HRNet model and is improved. The feature maps output by the network branches of the previous step of the fusion layer or transition layer of the HRNet model are first clustered to generate semantic prototypes, then an association matrix is ​​generated to guide each other between different resolutions, then an attention mechanism is formed by channel distribution entropy to weight the features of different resolutions, and finally the branches of different resolutions are fused through semantic association and spatial attention to obtain fused features, and the backbone network is trained to obtain the trained backbone network.

[0054] The weakly supervised training module is used to fix the parameters of the backbone network, connect the output of the backbone network to the 3D pose estimation network, construct the weakly supervised learning module, train the weakly supervised learning module using 3D labeled data to form an initial model, use the initial model to predict 2D unlabeled data to obtain a 3D pseudo-labeled dataset, merge the 3D pseudo-labeled dataset and the real 3D labeled data to form a training set, train the weakly supervised learning module, and obtain the trained weakly supervised learning module.

[0055] The pose estimation module is used to input real-time acquired single-view animal images into a trained weakly supervised learning module to perform single-view animal 3D pose estimation.

[0056] Furthermore, the working process of the backbone network is as follows:

[0057] S11. Feature maps output by the network branch at the previous step of the fusion or transition layer in the HRNet model. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ;use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ;

[0058] S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ;

[0059] S13. Calculate the channel distribution entropy. ,in, Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. Features of each channel; design a dynamic convolution kernel generator to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, and process the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. The formula for the dynamic convolution kernel generator is expressed as follows: ,in, The learnable threshold, Indicates spatial location Kernel size at the location; Indicates if, Indicates other;

[0060] S14, Utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. Feature map With final attention weight Element-wise multiplication yields attention-enhanced features. Then , After upsampling and The first-scale fusion feature is obtained by splicing. ,Will , After downsampling and The splicing yields the third-scale fusion features. Fusion features at the second scale .

[0061] Furthermore, S12 includes:

[0062] The prototype feature maps at each scale are unified to an intermediate scale, and the corresponding feature map is: , , ,in, Indicates downsampling, Indicates upsampling, The length and width of the prototype feature map at an intermediate scale;

[0063] The guidance weights of the high-resolution prototype on the low-resolution prototype are calculated using the following formula:

[0064]

[0065] The guidance weights of the low-resolution prototype to the high-resolution prototype are calculated using the following formula:

[0066]

[0067] in, Representation of feature map The One channel, Representation of feature map The One channel, 1≤ ≤N, 1≤ ≤N, The length, width, and channel are respectively... , , Feature map ; Represents spatial dimension Find the average. The number of cluster centers. For normalized exponential functions, The first term represents the guiding weight of the high-resolution prototype on the low-resolution prototype. Line 1 Column elements, The first term represents the guiding weight of the low-resolution prototype on the high-resolution prototype. Line 1 Column elements.

[0068] Furthermore, the aforementioned utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. ,include:

[0069] Third-scale channel attention weights MLP stands for Multilayer Perceptron. express Activation function, first-scale channel attention weights Second-scale channel attention weights Generate the final attention weights. , k=1,2,3, ⊙ is the Hadamard product.

[0070] Furthermore, the process of training the backbone network is as follows:

[0071] A loss function is constructed, the parameters of the backbone network are adjusted, and the value of the loss function is calculated. Training stops when the value of the loss function is minimized, resulting in a trained backbone network. The loss function is a contrastive loss function, and its formula is as follows: , , , All are weighting coefficients; For similarity loss and ,in, For expectations; This is the pose geometric feature vector of the current frame. To and The corresponding positive sample feature vectors are generally feature vectors from adjacent frames, or intermediate sample feature vectors generated through pose interpolation. To and The corresponding negative sample feature vector; For temperature coefficient, Indicates the first One negative sample, The number of negative samples. For geometric constraint loss and , Let be the physiological constraint threshold for species s. Indicates two postures and Euclidean distance in 3D space Represents the L2 norm; For sorting loss and , , , These are the feature vectors of three adjacent frames.

[0072] Furthermore, the loss function also includes a temporal loss function, which uses the output of the backbone network of three consecutive frames as the input of the RAFT network, calculates the bidirectional reprojection loss, bidirectional optical flow inverse loss, optical flow difference L1 / L2 loss, and optical flow gradient smoothing loss, and then weights and fuses them as the temporal consistency loss.

[0073] Furthermore, the 3D pose estimation network is a stacked multi-layer convolutional network, and the output of the 3D pose estimation network is the 3D joint coordinates.

[0074] The advantages of this invention are:

[0075] (1) Based on a small amount of 3D labeled data and the feature extraction capability of the trained backbone network, this invention performs feature extraction and pose estimation on a large amount of 2D unlabeled data to generate pseudo-3D labeled labels, thereby obtaining a large amount of 3D pseudo-labeled data. At the same time, based on the large amount of 3D pseudo-labeled data already generated, it is further used as training input to perform new pose estimation on a large amount of 2D unlabeled data, generating new pseudo-3D labeled labels. The weakly supervised learning module is iteratively optimized step by step, thereby effectively improving the performance ceiling of 3D pose estimation accuracy without relying on human experience, while having little dependence on labeled data.

[0076] (2) The backbone network of this invention establishes the association between features at different scales through semantic prototypes, allowing high-level semantics to guide low-level detail attention, and low-level details to correct the deviation of high-level semantics. The computational granularity of spatial attention is adaptively adjusted according to the complexity (entropy value) of the feature region, balancing accuracy and efficiency. The attention weights act on both the original scale features and the cross-scale associated features, forming a closed-loop optimization.

[0077] (3) Existing contrastive learning methods often use InfoNCE or SimCLR loss, which only focus on the surface similarity of image-level or overall features. This invention uses similarity loss. Geometric constraint loss And sorting loss The weighted summation of the three loss components constrains the pose features from three dimensions: surface similarity, physical plausibility, and temporal continuity, thereby further improving the accuracy of 3D pose estimation. In addition, a temporal loss function is set to ensure that the 3D pose changes in consecutive frames conform to the physical laws of animal movement and temporal continuity, compensating for the ambiguity of single-frame estimation.

[0078] (4) Currently, labeled data is scarce and costly, while unlabeled data is plentiful and readily available. Without this method, 3D pose estimation is only applicable to a small amount of labeled data, which can easily lead to insufficient performance. Without this method, it is impossible to make effective use of a large amount of unlabeled data at low cost, resulting in insufficient performance. Using this method, a large amount of unlabeled data can be used more fully and at low cost, thereby improving performance. Attached Figure Description

[0079] Figure 1 This is a flowchart of the backbone network training process in a single-view animal 3D pose estimation method disclosed in Embodiment 1 of the present invention.

[0080] Figure 2 This is a schematic diagram of the high-resolution feature fusion process in a single-view animal 3D pose estimation method disclosed in Embodiment 1 of the present invention;

[0081] Figure 3This is a schematic diagram of the low-resolution feature fusion process in a single-view animal 3D pose estimation method disclosed in Embodiment 1 of the present invention;

[0082] Figure 4 This is a schematic diagram of the training process of the weakly supervised learning module in a single-view animal 3D pose estimation method disclosed in Embodiment 1 of the present invention. Detailed Implementation

[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0084] Example 1

[0085] Embodiment 1 of the present invention provides a method for 3D pose estimation of an animal in a single view, comprising the following steps:

[0086] S1. The backbone network receives 2D image data. Based on HRNet, the backbone network is an improvement. The feature maps output from the previous branch of HRNet's FuseLayer or TransitionLayer are first clustered to generate semantic prototypes, then an association matrix is ​​generated for mutual guidance between different resolutions. Next, an attention mechanism based on channel distribution entropy is used to weight features at different resolutions. Finally, branches at different resolutions are fused through semantic association and spatial attention to obtain fused features. The backbone network is then trained to obtain the trained backbone network. The backbone network training process is as follows: Figure 1 As shown; the specific principle and process are as follows:

[0087] To effectively extract pose features at different scales from 2D images, the network design innovates upon HRNet. HRNet is an existing technology, as described in the paper "Wang J, Sun K, Cheng T, et al. Deep High-Resolution Representation Learning for Visual Recognition[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, PP(99):1-1.DOI:10.1109 / TPAMI.2020.2983686.". The HRNet model employs a pyramid and anti-pyramid structure with downsampling followed by upsampling, while maintaining a high-resolution main line. Feature fusion is performed between multiple branches of different resolutions (downsampling of high-resolution features supplements low-resolution branches, and upsampling of low-resolution features enhances high-resolution branches), ensuring that high-resolution features can always fuse with the semantic information of low-resolution branches, while low-resolution branches can also acquire high-resolution detail information. Among them, the FuseLayer of the HRNet model is mainly responsible for information fusion between feature maps of different resolutions. The conventional approach is to adjust feature maps of different resolutions to the same size through interpolation or downsampling before adding or concatenating them. This allows features from different resolutions to interact, ensuring that each feature map contains rich semantic and spatial information. However, this fusion method has limitations. For example, information loss or bias may occur during interpolation or downsampling, and simple addition or concatenation cannot fully exploit the complex relationships between features. The TransitionLayer in the HRNet model is used to generate new resolution branches. It typically uses convolutional operations, such as using a 3x3 convolution with a stride of 2 to downsample the output of the original branch, thus obtaining a new low-resolution branch. While this method increases the network's ability to capture information at different scales, it relies solely on the output of a single branch from the previous layer (usually the last branch) when generating new branches, failing to fully utilize information from other branches, resulting in insufficient information utilization.

[0088] This invention innovatively designs the FuseLayer and TransitionLayer of the HRNet model's backbone network, introducing a novel attention-based mechanism for multi-scale fusion. Existing attention mechanisms mostly focus on channel / spatial dependencies of single-scale features or static dimensionality decomposition, but neglect two core issues:

[0089] 1) Semantic association of cross-scale features in convolutional networks (mutual guidance between low-level detailed features and high-level semantic features).

[0090] 2) Differences in the dynamic complexity of feature regions (simple regions do not require fine-grained attention, while complex regions require more precise attention allocation).

[0091] To address this, this invention designs an attention mechanism from two novel perspectives: "cross-scale semantic collaboration" and "dynamic granularity adaptation," achieving more accurate feature weighted fusion of FuseLayer and TransitionLayer. The main improvements are threefold:

[0092] 1) Cross-scale semantic prototype association: Establish associations between features at different scales through semantic prototypes, allowing high-level semantics to guide low-level details, and low-level details to correct deviations in high-level semantics;

[0093] 2) Dynamic granularity attention calculation: Adaptively adjust the calculation granularity of spatial attention according to the complexity (entropy value) of the feature region to balance accuracy and efficiency;

[0094] 3) Two-way feedback mechanism: Attention weights are applied to both the original scale features and cross-scale related features simultaneously, forming a closed-loop optimization.

[0095] The comparison before and after the improvement is as follows:

[0096] Before the improvement, the HRNet model used simple upsampling for low-resolution branches when performing high-resolution FuseLayer operations, and simple downsampling with a step size of 3 for high-resolution branches when performing low-resolution FuseLayer operations.

[0097] like Figure 2 and Figure 3 As shown, the improved method first clusters different resolutions to generate semantic prototypes, then generates an association matrix for mutual guidance between different resolutions, then uses channel distribution entropy to form an attention mechanism to weight features at different resolutions, and finally fuses branches at different resolutions through semantic association and spatial attention. The specific process of the improved backbone network of this invention is as follows:

[0098] S11. Feature maps output from the previous branch of the HRNet's FuseLayer or TransitionLayer. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ;use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ;

[0099] Specifically, let the L1, L2, and L3 branches of the previous step network in the FuseLayer or TransitionLayer output feature maps. (High resolution, fine granularity) (Medium resolution, medium granularity) (Low resolution coarse granularity), where, express A 3D real vector space, , , .

[0100] For each scale feature, a "semantic prototype" is generated, representing the core semantic information of that scale. The channel dimension of the feature map is divided into multiple semantic groups, each represented by a learnable "semantic prototype." This clustering is "dynamic" because it adaptively discovers semantic relationships between feature channels through learnable parameters, rather than the static partitioning of traditional clustering algorithms.

[0101] right To perform dynamic clustering along the channel dimension, first initialize N cluster centers (N=8 in this example). , This represents the Nth cluster center. express The following are defined as 3D real vector spaces, with similar definitions for other real vector spaces, differing only in their dimensions; these will not be elaborated upon further. Each cluster center is a... These are dimensional vectors, and the cluster centers are updated through backpropagation during training to capture different semantic patterns.

[0102] For each channel Calculate its cosine similarity with all cluster centers. ,in Represents the vector dot product. Using the L2 norm, the final correlation matrix is ​​obtained. ,in This indicates the strength at which channel j belongs to prototype (cluster center) i;

[0103] The channel dimension is compressed to the prototype dimension through matrix multiplication, i.e. Intuitively explained for each position ( ) eigenvectors Mapped to Each dimension This represents the response strength of the i-th prototype at that location.

[0104] S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ;

[0105] Specifically, this step establishes dependencies between prototypes at different scales to achieve semantic guidance. By establishing semantic dependencies between features at different scales, it enables high-level semantics to guide low-level details, and low-level details to correct high-level semantics. This bidirectional association breaks the limitation of traditional attention mechanisms operating only at a single scale, allowing the network to utilize complementary information from multi-scale features simultaneously.

[0106] The prototype feature maps at various scales generated by S11 Unified to intermediate scale And perform cross-scale comparisons, i.e. , , ,in , Indicates downsampling, Indicates upsampling, This represents the length and width of the prototype feature map at an intermediate scale.

[0107] The guiding weight for measuring the high-resolution prototype over the low-resolution prototype. , Representation of feature map The One channel for high-resolution prototype index (1≤ ≤N), Representation of feature map The One channel for low-resolution prototype index (1≤ ≤N), The length, width, and channel are respectively... , , Feature map GlobalAvg() represents the spatial dimension. The average is calculated to obtain the global semantic representation of the prototype, then divided by... The goal is to scale to stabilize the gradient, ultimately obtaining .

[0108] Similarly, the guidance weights of the low-resolution prototype to the high-resolution prototype were calculated. .

[0109] S13. Adjusting the computational granularity of spatial attention based on the complexity of the feature region: The core idea is to adaptively adjust the granularity of attention computation according to the complexity of the feature region, significantly reducing computational overhead while maintaining accuracy. This design solves the "one-size-fits-all" problem of traditional attention mechanisms, avoiding wasting computational resources in simple regions. Specifically, for feature maps... Calculate each spatial location ( The channel distribution entropy, i.e. High entropy values ​​indicate uniform distribution among channels (such as edges and textured areas where multiple features coexist), while low entropy values ​​indicate concentrated distribution among channels (such as solid color backgrounds where a single feature dominates). Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. The characteristics of each channel are considered; a dynamic convolution kernel generator is designed to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, i.e. ,in, The learnable threshold, Indicates spatial location The size of the convolution kernel at a given location affects the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. .

[0110] S14. Combining semantic prototype association and dynamic spatial attention, the final weights are generated. The core objective is to fuse the cross-scale semantic association information generated in S12 with the dynamic granular spatial attention generated in S13 to form the final feature enhancement weights. This fusion not only achieves synergy between channel and spatial dimensions but also establishes a bidirectional feedback mechanism between cross-scale features. Specifically, utilizing... and Channel attention weights are generated at various scales, with low-resolution channel weights generated under high-resolution semantic guidance, i.e., the third-scale channel attention weights. MLP stands for Multilayer Perceptron. Similarly, high-resolution channel weights are generated under the guidance of low-resolution semantics, i.e., the channel attention weights at the first scale. The medium-resolution channel weights are generated from their own prototype, namely the channel attention weights at the second scale. , express Activation function.

[0111] Channel attention weights and spatial attention maps at various scales The final attention weights are obtained by performing element-wise multiplication. k=1,2,3, ⊙ is the Hadamard product (element-by-element multiplication), which achieves synergistic enhancement of channel and spatial dimensions.

[0112] Feature map With final attention weight Element-wise multiplication yields attention-enhanced features. ⊙ Finally, the low-resolution features are upsampled and then concatenated with the high-resolution features, that is, , After upsampling and The first-scale fusion feature is obtained by splicing. The high-resolution features are downsampled and then concatenated with the low-resolution features, that is... , After downsampling and The splicing yields the third-scale fusion features. The mid-resolution features remain unchanged, i.e., the fused features at the second scale. This creates cross-scale feedback, ultimately enabling innovative modifications to FuseLayer. Compared to the innovative modifications to TransitionLayer, TransitionLayer, due to the need to implement new resolution branches, requires an additional downsampling step after performing the aforementioned attention-weighted fusion on different resolution branches.

[0113] S2. Fix the backbone network parameters, connect the backbone network output to the 3D pose estimation network to construct a weakly supervised learning module. Train the weakly supervised learning module using 3D labeled data to form an initial model. Use the initial model to predict 2D unlabeled data to obtain a pseudo-labeled dataset. Merge the pseudo-labeled dataset and the real 3D labeled data to form a training set, and train the weakly supervised learning module to obtain a trained weakly supervised learning module; Figure 4 As shown, the process of training the backbone network is as follows: constructing a loss function, adjusting the parameters of the backbone network, calculating the value of the loss function, and stopping training when the value of the loss function is minimized to obtain the trained backbone network; wherein, the loss function is constructed by contrastive learning network and temporal consistency network. The specific process is as follows:

[0114] The core task of contrastive learning networks is to guide the model to learn the latent spatial structure of 3D poses by constructing a similarity metric for pose features, focusing on fine-grained distinction of pose geometric features. The input of a contrastive learning network is the output of a backbone network, which consists of stacked dilated convolutional layers and a ViT structure. The output is a 72-dimensional pose geometric feature vector f, containing three sub-vectors:

[0115] Joint coordinate distribution vector (32-dimensional): Encodes the relative positional distribution of key joints (such as head, trunk, and extremities) in normalized space;

[0116] Motion trend vector (20-dimensional): The joint displacement direction and amplitude characteristics of consecutive frames are compressed through principal component analysis (PCA);

[0117] Topological structure vector (20-dimensional): Based on the skeletal connection strength encoding of graph convolution output, it reflects the differences in posture structure among different species (such as the limb ratio between birds and mammals).

[0118] The three sub-vectors are orthogonalized (ensuring that the cosine similarity between vectors is <0.1) to avoid feature redundancy and improve the discrimination efficiency of contrastive learning.

[0119] Existing contrastive learning methods often employ InfoNCE or SimCLR losses, focusing only on surface similarity of image-level or overall features. This invention uses a similarity loss method. Geometric constraint loss And sorting loss Weighted summation, i.e., contrastive loss function α, β, and γ are weighting coefficients, and the three loss components constrain the pose features from three dimensions: surface similarity, physical rationality, and temporal continuity.

[0120] Similarity loss An improved InfoNCE loss is employed, focusing on the surface similarity measure of pose features. It not only treats temporally adjacent frames within the same video as positive sample pairs but also generates "intermediate samples" through pose interpolation (e.g., generating a pseudo-sample for frame t+1 by linear interpolating the poses of frames t and t+2), enhancing the continuity of positive samples. For negative sample pairs, a "hard example mining" mechanism (an optimization strategy for negative sample selection) is introduced, prioritizing samples with similar poses but belonging to different motion stages (e.g., subtle differences between standing and squatting). The formula is as follows: ,in, This is the pose geometric feature vector of the current frame. To and The corresponding positive sample feature vectors are generally feature vectors from adjacent frames, or intermediate sample feature vectors generated through pose interpolation. To and The corresponding negative sample feature vectors come from samples with "similar postures but different movement stages" selected by the "difficult example mining" mechanism (such as samples with subtle differences between standing and squatting). For temperature coefficient, Indicates the first One negative sample, The number of negative samples is adaptively adjusted according to the difficulty of the samples (difficult samples). = 0.1, easy example samples = 0.5), to avoid the similarity measurement bias caused by the difference in sample difficulty.

[0121] Geometric constraint loss Introduce the physical rationality constraint of 3D pose to ensure that the poses corresponding to similar features conform to the biological motion law. Traditional contrast loss only judges "similarity" through the cosine similarity of feature vectors, and easily ignores the physical constraints of poses (such as joint movement range, species physiological structure limitations). The geometric constraint loss explicitly associates "feature similarity" with "3D joint Euclidean distance difference". When two pose features are similar, force their corresponding 3D joint spatial distributions to conform to the biological motion law (such as the maximum telescopic distance of the legs of quadruped animals does not exceed the physiological constraint threshold of species s , specifically, when the cosine similarity of the feature vectors of two poses is higher than the threshold (such as 0.8), force the Euclidean distance difference between the joints of their corresponding 3D poses not to exceed the species-specific threshold (such as the maximum telescopic distance of the legs of quadruped animals). If it exceeds the threshold, impose a penalty. The formula is: , where is the physiological constraint threshold of species s (obtained by statistical analysis of a small amount of 3D annotation data), to ensure the consistency between the similarity of pose features and biological physical characteristics. represents two poses and the Euclidean distance in 3D space, represents the L2 norm.

[0122] Sorting loss Ensure that the pose features at different stages of the same action are arranged in chronological order through temporal constraints, and strengthen the motion continuity. Traditional contrast learning only focuses on the binary judgment of "similar / dissimilar", ignoring the "sequential relationship" of poses in time series (such as the stage progression of "starting→running→decelerating"). The sorting loss forces the feature vectors to be arranged in chronological order in the latent space through the triplet loss (t1 < t2 < t3), ensuring that the difference between the middle frame feature and the front and back frames conforms to the natural temporal logic of motion, and focusing on the continuity of animal motion (such as gait cycle, action transition), solving the ambiguity that "the same pose may belong to different action stages" (such as "squat" may be the transition of "standing→jumping" or "jumping→landing", and its stage attribute is clarified through temporal sorting). For continuous actions such as "starting - running - decelerating", use the triplet loss to force the feature vectors to satisfy the time progression relationship in the latent space. The formula is: , , , are the feature vectors of three adjacent frames, ensuring that the difference between the middle frame feature and the front and back frames conforms to the natural temporal logic of motion.

[0123] The temporal consistency constraint module of the temporal consistency network aims to ensure that the 3D pose changes of consecutive frames conform to the physical laws of animal movement and temporal continuity, thus compensating for the ambiguity of single-frame estimation. The model's input consists of three consecutive frames of images (…). The backbone network output, with a temporally consistent network structure of RAFT, outputs forward optical flow. ( arrive (pixel motion vector field), reverse optical flow ( arrive (pixel motion vector field), continuous optical flow ( arrive (Pixel motion vector field). RAFT refers to Recurrent All-Pairs Field Transforms, a deep neural network for optical flow estimation. It is an existing technology. The paper is from "Teed Z, Deng J. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow (Extended Abstract). [C] / / International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, 2021. DOI:10.24963 / IJCAI.2021 / 662.".

[0124] The temporal consistency loss is weighted by existing bidirectional reprojection loss, bidirectional optical flow reciprocity loss, optical flow difference L1 / L2 loss, and optical flow gradient smoothing loss.

[0125] Bidirectional reprojection loss is commonly used in fields such as visual SLAM and is considered an existing technology. It is mentioned in the paper "Campos C, Elvira R, Juan J. Gómez Rodríguez, et al. ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM[J]. 2020.DOI:10.1109 / TRO.2021.3075644." This paper proposes using bidirectional reprojection error as the objective function to optimize pose and employs the Huber function to provide robustness against pseudo-matches.

[0126] The bidirectional optical flow reciprocal loss is discussed in the paper "Zhu AZ, Liu W, Wang Z, et al. RobustnessMeets Deep Learning: An End-to-End Hybrid Pipeline for Unsupervised Learning of Egomotion[J].2018.DOI:10.48550 / arXiv.1812.08351.", which sets constraints and applies these constraints to disparity and optical flow to construct the relevant consistency loss function.

[0127] The L2-optical flow in the L1 / L2 loss of optical flow difference was first proposed and solved by Horn and Schunk in 1981. The relevant paper is "Horn BKP, Schunck BG. Determining Optical Flow[J]. Artificial Intelligence, 1981, 17(1-3):185-203.DOI:10.1016 / 0004-3702(81)90024-2." This paper proposed an optical flow calculation method based on the L2 norm, which is called the HS algorithm. The L1-optical flow was proposed relatively late, and there is no specific widely recognized foundational paper to define it. However, many papers on optical flow use L1-optical flow when it comes to robust optical flow calculation, such as when dealing with optical flow with large offsets. Furthermore, in the paper "Teed Z, Deng J. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow (Extended Abstract). [C] / / International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, 2021. DOI:10.24963 / IJCAI.2021 / 662.", the optical flow difference L1 loss was also used to calculate the distance between the predicted optical flow and the actual optical flow.

[0128] The optical flow gradient smoothing loss proposed in "Yin Z, Shi J. GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose[J]. IEEE, 2018.DOI:10.1109 / CVPR.2018.00212." is an adaptive geometric consistency loss that extends the smoothness loss of the two-dimensional optical flow field.

[0129] The backbone network of the weakly supervised part adopts the backbone network of the self-supervised part that has been trained. During weakly supervised learning, the parameters of this part are fixed and used only to extract high-dimensional features as input to the 3D pose estimation network. The 3D pose estimation network architecture is a stacked multi-layer convolutional network, and the output is the 3D joint coordinates (x, y, z). The 3D pose estimation loss adopts the MPJPE (Mean Per Joint Position Error) loss to calculate the Euclidean distance between the predicted 3D pose and the real annotation. MPJPE loss is an existing technology. The relevant paper is "Ionescu C, Papava D, Olaru V, et al. Human3.6M: Large Scale Datasets and Predictive Methods for 3D HumanSensing in Natural Environments[J].IEEE Transactions on Pattern Analysis & Machine Intelligence, 2014, 36(7):1325-1339.DOI:10.1109 / TPAMI.2013.248.".

[0130] The process of iterative pseudo-label-guided training is as follows:

[0131] 1) First step: Basic model initialization (relies on a small number of real annotations)

[0132] With fixed backbone network parameters, the output of the backbone network is connected to a 3D pose estimation network to construct a weakly supervised learning module. This module is then trained using labeled 3D data to form an initial model. Finally, the initial model is trained using a small amount of labeled 3D data based on cross-entropy.

[0133] 2) Second step: pseudo-label generation and filtering (reducing noise interference)

[0134] The initial model is used to predict 2D unlabeled data to obtain the predicted probability distribution of each sample. Based on a threshold (incrementing from 0.7 to 0.9), the highest probability is selected as the pseudo label to form a 3D pseudo-labeled dataset.

[0135] 3) The third step is joint training (integrating real labels and filtered pseudo labels).

[0136] The 3D pseudo-labeled dataset and the real 3D labeled data are merged to form a new training set. A new model is trained, and when calculating the loss, the real labeled samples are given a higher weight (fixed at 1), and the pseudo-labeled samples are given a lower weight (increasing from 0.7 to 0.9). The total loss is then calculated.

[0137] 4) Fourth step: Iterative optimization (gradually improving the quality of the model and pseudo-labels)

[0138] Repeat steps two and three until the model converges on the validation and test set metrics.

[0139] S3. Input the real-time acquired single-view animal images into the trained weakly supervised learning module to perform single-view animal 3D pose estimation.

[0140] Through the above technical solution, this invention uses a small amount of 3D labeled data and the feature extraction capability of a trained self-supervised learning module to perform feature extraction and pose estimation on a large amount of unlabeled 2D data, generating pseudo-3D labeled tags, thereby obtaining a large amount of 3D pseudo-labeled data. At the same time, based on the large amount of 3D pseudo-labeled data already generated, it is further used as training input to perform new pose estimation on a large amount of unlabeled 2D data, generating new pseudo-3D labeled tags. During the iteration process, the confidence level is sorted and unreliable pseudo-labels are filtered out, and the weakly supervised learning module is gradually iteratively optimized.

[0141] Example 2

[0142] Based on Embodiment 1, Embodiment 2 of the present invention also provides a system for single-view animal 3D pose estimation, comprising:

[0143] The self-supervised training module is used for the backbone network to receive 2D image data. The backbone network is based on the HRNet model and is improved. The feature maps output by the network branches of the previous step of the fusion layer or transition layer of the HRNet model are first clustered to generate semantic prototypes, then an association matrix is ​​generated to guide each other between different resolutions, then an attention mechanism is formed by channel distribution entropy to weight the features of different resolutions, and finally the branches of different resolutions are fused through semantic association and spatial attention to obtain fused features, and the backbone network is trained to obtain the trained backbone network.

[0144] The weakly supervised training module is used to fix the parameters of the backbone network, connect the output of the backbone network to the 3D pose estimation network, construct the weakly supervised learning module, train the weakly supervised learning module using 3D labeled data to form an initial model, use the initial model to predict 2D unlabeled data to obtain a 3D pseudo-labeled dataset, merge the 3D pseudo-labeled dataset and the real 3D labeled data to form a training set, train the weakly supervised learning module, and obtain the trained weakly supervised learning module.

[0145] The pose estimation module is used to input real-time acquired single-view animal images into a trained weakly supervised learning module to perform single-view animal 3D pose estimation.

[0146] Specifically, the working process of the backbone network is as follows:

[0147] S11. Feature maps output by the network branch at the previous step of the fusion or transition layer in the HRNet model. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ;use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ;

[0148] S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ;

[0149] S13. Calculate the channel distribution entropy. ,in, Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. Features of each channel; design a dynamic convolution kernel generator to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, and process the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. The formula for the dynamic convolution kernel generator is expressed as follows: ,in, The learnable threshold, Indicates spatial location Kernel size at the location; Indicates if, Indicates other;

[0150] S14, Utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. Feature map With final attention weight Element-wise multiplication yields attention-enhanced features. Then , After upsampling and The first-scale fusion feature is obtained by splicing. ,Will , After downsampling and The splicing yields the third-scale fusion features. Fusion features at the second scale .

[0151] More specifically, S12 includes:

[0152] The prototype feature maps at each scale are unified to an intermediate scale, and the corresponding feature map is: , , ,in, Indicates downsampling, Indicates upsampling, The length and width of the prototype feature map at an intermediate scale;

[0153] The guidance weights of the high-resolution prototype on the low-resolution prototype are calculated using the following formula:

[0154]

[0155] The guidance weights of the low-resolution prototype to the high-resolution prototype are calculated using the following formula:

[0156]

[0157] in, Representation of feature map The One channel, Representation of feature map The One channel, 1≤ ≤N, 1≤ ≤N, The length, width, and channel are respectively... , , Feature map ; Represents spatial dimension Find the average. The number of cluster centers. For normalized exponential functions, The first term represents the guiding weight of the high-resolution prototype on the low-resolution prototype. Line 1 Column elements, The first term represents the guiding weight of the low-resolution prototype on the high-resolution prototype. Line 1 Column elements.

[0158] More specifically, the utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. ,include:

[0159] Third-scale channel attention weights MLP stands for Multilayer Perceptron. express Activation function, first-scale channel attention weights Second-scale channel attention weights Generate the final attention weights. , k=1,2,3, ⊙ is the Hadamard product.

[0160] More specifically, the process of training the backbone network is as follows:

[0161] A loss function is constructed, the parameters of the backbone network are adjusted, and the value of the loss function is calculated. Training stops when the value of the loss function is minimized, resulting in a trained backbone network. The loss function is a contrastive loss function, and its formula is as follows: , , , All are weighting coefficients; For similarity loss and ,in, For expectations; This is the pose geometric feature vector of the current frame. To and The corresponding positive sample feature vectors are generally feature vectors from adjacent frames, or intermediate sample feature vectors generated through pose interpolation. To and The corresponding negative sample feature vector; For temperature coefficient, Indicates the first One negative sample, The number of negative samples. For geometric constraint loss and , Let be the physiological constraint threshold for species s. Indicates two postures and Euclidean distance in 3D space Represents the L2 norm; For sorting loss and , , , These are the feature vectors of three adjacent frames.

[0162] More specifically, the loss function is a temporal loss function, which uses the output of the backbone network of three consecutive frames of images as the input of the RAFT network, calculates the bidirectional reprojection loss, bidirectional optical flow inverse loss, optical flow difference L1 / L2 loss, and optical flow gradient smoothing loss, and then weights and fuses them as the temporal consistency loss.

[0163] Specifically, the 3D pose estimation network is a stacked multi-layer convolutional network, and the output of the 3D pose estimation network is the 3D joint coordinates.

[0164] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for 3D pose estimation of an animal from a single view, characterized in that, include: S1. The backbone network receives 2D image data. The backbone network is based on the HRNet model and is improved. The feature maps output by the network branch of the previous step of the fusion layer or transition layer of the HRNet model are first clustered to generate semantic prototypes, then an association matrix is ​​generated to guide the mutual guidance between different resolutions, then an attention mechanism is formed by channel distribution entropy to weight the features of different resolutions, and finally the branches of different resolutions are fused through semantic association and spatial attention to obtain fused features. The backbone network is then trained to obtain the trained backbone network. S2. Fix the backbone network parameters, connect the output of the backbone network to the 3D pose estimation network to build a weakly supervised learning module, train the weakly supervised learning module using 3D labeled data to form an initial model, use the initial model to predict 2D unlabeled data to obtain a 3D pseudo-labeled dataset, merge the 3D pseudo-labeled dataset and the real 3D labeled data to form a training set, train the weakly supervised learning module, and obtain the trained weakly supervised learning module. S3. Input the real-time acquired single-view animal images into the trained weakly supervised learning module to perform single-view animal 3D pose estimation.

2. The method for single-view animal 3D pose estimation according to claim 1, characterized in that, The working process of the backbone network is as follows: S11. Feature maps output by the network branch at the previous step of the fusion or transition layer in the HRNet model. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ;use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ; S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ; S13. Calculate the channel distribution entropy. ,in, Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. Features of each channel; design a dynamic convolution kernel generator to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, and process the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. The formula for the dynamic convolution kernel generator is expressed as follows: ,in, The learnable threshold, Indicates spatial location Kernel size at the location; Indicates if, Indicates other; S14, Utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. Feature map With final attention weight Element-wise multiplication yields attention-enhanced features Then , After upsampling and The first-scale fusion feature is obtained by splicing. ,Will , After downsampling and The splicing yields the third-scale fusion features. Fusion features at the second scale .

3. The method for single-view animal 3D pose estimation according to claim 2, characterized in that, S12 includes: The prototype feature maps at each scale are unified to an intermediate scale, and the corresponding feature map is: , , ,in, Indicates downsampling, Indicates upsampling, The length and width of the prototype feature map at an intermediate scale; The guidance weights of the high-resolution prototype on the low-resolution prototype are calculated using the following formula: The guidance weights of the low-resolution prototype to the high-resolution prototype are calculated using the following formula: in, Representation of feature map The One channel, Representation of feature map The One channel, 1≤ ≤N, 1≤ ≤N, The length, width, and channel are respectively... , , Feature map ; Represents spatial dimension Find the average. The number of cluster centers. For normalized exponential functions, The first term represents the guiding weight of the high-resolution prototype on the low-resolution prototype. Line number Column elements, The first term represents the guiding weight of the low-resolution prototype on the high-resolution prototype. Line number Column elements.

4. The method for single-view animal 3D pose estimation according to claim 3, characterized in that, The use of and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. ,include: Third-scale channel attention weights MLP stands for Multilayer Perceptron. express Activation function, first-scale channel attention weights Second-scale channel attention weights Generate the final attention weights. , k=1,2,3, ⊙ is the Hadamard product.

5. The method for single-view animal 3D pose estimation according to claim 4, characterized in that, The process of training the backbone network is as follows: A loss function is constructed, the parameters of the backbone network are adjusted, and the value of the loss function is calculated. Training stops when the value of the loss function is minimized, resulting in a trained backbone network. The loss function is a contrastive loss function, and its formula is as follows: , , , All are weighting coefficients; For similarity loss and ,in, For expectations; This is the pose geometric feature vector of the current frame. To and The corresponding positive sample feature vector, To and The corresponding negative sample feature vector; For temperature coefficient, Indicates the first One negative sample, The number of negative samples. For geometric constraint loss and , Let be the physiological constraint threshold for species s. Indicates two postures and Euclidean distance in 3D space Represents the L2 norm; For sorting loss and , , , These are the feature vectors of three adjacent frames.

6. The method for single-view animal 3D pose estimation according to claim 1, characterized in that, The 3D pose estimation network is a stacked multi-layer convolutional network, and the output of the 3D pose estimation network is the 3D joint coordinates.

7. A system for single-view animal 3D pose estimation, characterized in that, include: The self-supervised training module is used for the backbone network to receive 2D image data. The backbone network is based on the HRNet model and is improved. The feature maps output by the network branches of the previous step of the fusion layer or transition layer of the HRNet model are first clustered to generate semantic prototypes, then an association matrix is ​​generated to guide each other between different resolutions, then an attention mechanism is formed by channel distribution entropy to weight the features of different resolutions, and finally the branches of different resolutions are fused through semantic association and spatial attention to obtain fused features, and the backbone network is trained to obtain the trained backbone network. The weakly supervised training module is used to fix the parameters of the backbone network, connect the output of the backbone network to the 3D pose estimation network, construct the weakly supervised learning module, train the weakly supervised learning module using 3D labeled data to form an initial model, use the initial model to predict 2D unlabeled data to obtain a 3D pseudo-labeled dataset, merge the 3D pseudo-labeled dataset and the real 3D labeled data to form a training set, train the weakly supervised learning module, and obtain the trained weakly supervised learning module. The pose estimation module is used to input real-time acquired single-view animal images into a trained weakly supervised learning module to perform single-view animal 3D pose estimation.

8. The system for single-view animal 3D pose estimation according to claim 7, characterized in that, The working process of the backbone network is as follows: S11. Character map of the network branch output of the previous step in the fusion or transition layer of the HRNet model. Dynamic clustering along the channel dimension is performed on the feature maps. For each channel, its cosine similarity to all cluster centers is calculated, and each cosine similarity result is used as a matrix element to construct an association matrix. ; use Calculate the prototype feature map for each scale; where different scales correspond to different resolutions. Indicates the first One scale, ; S12. Unify the prototype feature maps at each scale to an intermediate scale, and calculate the guidance weights of the high-resolution prototype on the low-resolution prototype. And the guiding weights of low-resolution prototypes to high-resolution prototypes ; S13. Calculate the channel distribution entropy. ,in, Indicates channel dimension, Indicates the total number of channel dimensions. Indicates except Other channel dimensions besides; Representation of feature map Spatial location ( ) No. Features of each channel; design a dynamic convolution kernel generator to assign convolution kernels of different sizes to each spatial location based on the entropy value of the channel distribution entropy, and process the feature map. Perform granular adaptive convolution, and activate the output of the granular adaptive convolution using the Sigmoid function to obtain the final spatial attention map. The formula for the dynamic convolution kernel generator is expressed as follows: ,in, The learnable threshold, Indicates spatial location Kernel size at the location; Indicates if, Indicates other; S14, Utilization and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. Feature map With final attention weight Element-wise multiplication yields attention-enhanced features Then , After upsampling and The first-scale fusion feature is obtained by splicing. ,Will , After downsampling and The splicing yields the third-scale fusion features. Fusion features at the second scale .

9. A system for single-view animal 3D pose estimation according to claim 8, characterized in that, S12 includes: The prototype feature maps at each scale are unified to an intermediate scale, and the corresponding feature map is: , , ,in, Indicates downsampling, Indicates upsampling, The length and width of the prototype feature map at an intermediate scale; The guidance weights of the high-resolution prototype on the low-resolution prototype are calculated using the following formula: The guidance weights of the low-resolution prototype to the high-resolution prototype are calculated using the following formula: in, Representation of feature map The One channel, Representation of feature map The One channel, 1≤ ≤N, 1≤ ≤N, The length, width, and channel are respectively... , , Feature map ; Represents spatial dimension Find the average. The number of cluster centers. For normalized exponential functions, The first term represents the guiding weight of the high-resolution prototype on the low-resolution prototype. Line number Column elements, The first term represents the guiding weight of the low-resolution prototype on the high-resolution prototype. Line number Column elements.

10. A system for single-view animal 3D pose estimation according to claim 9, characterized in that, The use of and Generate channel attention weights at various scales, and then compare these channel attention weights with the spatial attention map. The final attention weights are obtained by performing element-wise multiplication. ,include: Third-scale channel attention weights MLP stands for Multilayer Perceptron. express Activation function, first-scale channel attention weights Second-scale channel attention weights Generate the final attention weights. , k=1,2,3, ⊙ is the Hadamard product.

Citation Information

Patent Citations

  • Human body posture estimation method based on improved HRNet network in operating room scene

    CN114373226A

  • Three-dimensional human body posture estimation method for virtual clothing show

    CN116030498A