Animal attitude estimation method and system based on super key point coding

By employing a super keypoint encoding method and utilizing a Bayesian prior framework and Transformer interaction mechanism, the problems of weak cross-species generalization ability and large computational redundancy in animal pose estimation are solved, achieving high-precision and efficient animal pose estimation, which is applicable to livestock monitoring and wildlife conservation.

CN121010987APending Publication Date: 2025-11-25SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511110378.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing animal pose estimation methods suffer from poor cross-species generalization ability, large computational redundancy, and lack of anatomical constraints, resulting in accuracy and efficiency bottlenecks in complex scenarios, especially in real-time performance in wildlife conservation and livestock monitoring.

Method used

A super-keypoint encoding-based method is adopted, which generates a high-level semantic super-keypoint set through a Bayesian prior framework. Combining density peak clustering and dynamic pruning techniques, the Transformer interaction mechanism is used to optimize the anatomical topological constraints of visual markers and super-keypoints, and outputs a high-precision keypoint heatmap.

Benefits of technology

It achieves high accuracy and efficiency in cross-species animal pose estimation, supports millisecond-level pose analysis in complex scenarios, and meets the real-time needs of livestock monitoring and wildlife protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010987A_ABST
    Figure CN121010987A_ABST
Patent Text Reader

Abstract

The invention relates to an animal attitude estimation method and system based on super key point coding, and the system comprises the steps: modeling cross-species homologous key point distribution characteristics through a Bayesian prior framework, and generating high-level semantic super key points through density peak clustering; then image blocks are coded into visual marks, and the visual marks and the super key point marks are jointly input into a dynamic pruning Transform architecture; pruning redundancy calculation through a norm-driven attention scoring mechanism, and jointly learning a visual constraint and anatomical structure relationship; and finally outputting a high-precision key point heat map. According to the system, a super key point-visual marker collaborative optimization architecture is constructed, the limitation of animal form diversity processing, redundancy calculation and cross-species generalization ability in the prior art is broken through, and an efficient attitude estimation technology suitable for scenes such as animal husbandry monitoring and wild animal protection is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to animal pose estimation technology in the field of computer vision, specifically disclosing an animal pose estimation system and method based on super keypoint encoding. Background Technology

[0002] Animal pose estimation, as a core task of computer vision, has significant application value in fields such as precision animal husbandry and wildlife conservation. Existing methods are mainly divided into three categories: human model transfer methods, which directly transfer human pose estimation models, but ignore the diversity of animal morphology, resulting in poor cross-species generalization ability; heatmap regression methods, which supervise keypoint localization through Gaussian heatmaps, but suffer from coordinate quantization errors and computational redundancy; and unsupervised methods, which utilize clustering or contrastive learning to reduce label dependence, but their accuracy drops significantly in complex scenes.

[0003] In the prior art, patent number CN202410101315A discloses a method for estimating the pose of wild animals in real-world scenarios, specifically including: S1, constructing a dataset of animal pose estimation images and processing the constructed dataset based on style transfer; S2, constructing a free simple baseline pose estimation model based on heatmap generation based on group whitening operation, generating heatmaps using the model, and training the model using the heatmaps; S3, correcting the model and designing coordinate representation methods and heatmap decoding methods; S4, adopting a lightweight pose estimation network decoding architecture to complete the estimation of wild animal poses in real-world scenarios. Although this method alleviates heatmap quantization errors through Taylor expansion coordinate decoding, it relies on fixed keypoint definitions and cannot model the semantic associations of joints between different species, failing to solve the cross-species generalization bottleneck caused by differences in animal anatomical structures; furthermore, lightweight deconvolution only reduces local computational load and does not optimize the processing of globally redundant visual features; in addition, although Taylor expansion improves sub-pixel accuracy, it does not establish anatomical constraint relationships between keypoints. Therefore, there is an urgent need for a new paradigm for animal pose estimation that integrates cross-species keypoint semantic modeling, adaptive computational optimization, and anatomical structure constraint learning to break through the accuracy and efficiency bottlenecks in complex scenarios. Summary of the Invention

[0004] The first objective of this invention is to provide an animal pose estimation method based on super keypoint encoding, aiming to solve the problems of poor cross-species generalization ability, large computational redundancy, and lack of anatomical constraints caused by the diversity of animal morphology in existing technologies. To solve the above technical problems, an animal pose estimation method based on super keypoint encoding is provided, comprising the following steps:

[0005] S1. Super keypoint generation: The statistical distribution features of animal keypoints are extracted through a Bayesian prior framework, and a set of super keypoints with high-level semantics is generated using the density peak clustering algorithm.

[0006] S2, Tokenization representation, divides the input image into visual blocks and encodes them into visual tags with location embeddings, while mapping superkeypoints to learnable tags;

[0007] S3, Dynamic Label Pruning: The dynamic norm pruning module removes redundant visual labels and retains labels with high attention weights.

[0008] S4, Transformer Interaction and Prediction: The combined input of pruned visual labels and super keypoint labels to the Transformer encoder learns cross-label interaction relationships through a self-attention mechanism and outputs a keypoint heatmap.

[0009] In one embodiment, a pre-trained feature extractor f is used on the labeled image set of the basic animal categories. θ Extract local features O from all keypoints r r And modeled as a Gaussian distribution O r ~N(ζr,σ 2 I), then calculate the distribution mean ζ of the keypoints r in category y. r :

[0010]

[0011] Where I' is the image. These are the coordinates of the key points. Based on ζ. r Perform density peak clustering to adaptively generate L super keypoint sets M = {μ1, μ2, ..., μ...} L The same supercritical point encompasses homologous anatomical sites across different species.

[0012] In one embodiment, images from a labeled image set of basic animal categories are used.

[0013] Divided into A patch, and for the patch p i Perform linear projection to obtain the corresponding encoding v i :

[0014]

[0015] And add a location code to it, updating its code as follows:

[0016] v i ←v i +pe i

[0017] M is then mapped to learnable labels {t1,…,t} L}, output the label sequence T = {[v1,…,v...} N ],[t1,…,tL ]}.

[0018] In one embodiment, the query vector q, key vector k, and value vector v are first calculated:

[0019]

[0020] And generate attention matrix A:

[0021]

[0022] For the input visual label sequence {v1,…,v N Calculate the attention score ij :

[0023]

[0024] A binary mask M is generated based on a dynamic threshold τ set according to the k-th largest score. ij :

[0025] M ij =I(score) ij ≥τ)

[0026] Then perform the pruning operation:

[0027] x pruned =M⊙x

[0028] In one embodiment, the Transformer interaction employs a multi-head self-attention mechanism as follows:

[0029] MSA(T)=[SA1(T);SA2(T);…;SA h (T)]W P

[0030] in d is the label dimension, d i =d.

[0031] In one embodiment, the density peak clustering does not preset the number of clusters, but adaptively identifies densely distributed key point regions based on k-nearest neighbor distance, and forces different semantic key points within the same species to be assigned to different super key points.

[0032] In one embodiment, the dynamic threshold τ is determined by selecting the k-th maximum score, satisfying τ = topk(score). ij Furthermore, the number of visual markers after pruning is 50%-70% of the original number.

[0033] In one embodiment, the keypoint prediction results are evaluated using the Target Keypoint Similarity (OKS) metric:

[0034]

[0035] Where d i To predict the Euclidean distance from the ground truth coordinates, s is the square root of the object's bounding box area, and k is... i This is the keypoint type normalization constant.

[0036] The second objective of this invention is to provide an animal pose estimation system based on super keypoint encoding, which aims to solve the problems of insufficient generalization ability and poor real-time performance of existing systems under complex animal morphologies, and to achieve cross-species anatomical constraint learning and efficient deployment of edge devices through a modular architecture.

[0037] To address the aforementioned technical issues, an animal pose estimation system based on super keypoint encoding is provided. The system implements the aforementioned animal pose estimation method based on super keypoint encoding, including a prior modeling module, a tokenization encoding module, a dynamic pruning module, a cross-tag interaction module, and a pose decoding module.

[0038] In one embodiment, the prior modeling module extracts animal keypoint features using a Bayesian distribution framework and performs density peak clustering to generate a set of super keypoints; the tokenization encoding module maps the input image into blocks as visual tags, while embedding the super keypoints as learnable tags; the dynamic pruning module calculates a binary mask based on attention scores to filter out low-weight visual tags; the cross-tag interaction module jointly optimizes the constraint relationship between visual tags and super keypoint tags through a multi-head self-attention mechanism; and the pose decoding module outputs keypoint heatmap coordinates based on the interacted super keypoint tags.

[0039] Implementing the embodiments of the present invention will have the following beneficial effects:

[0040] 1. The animal pose estimation method based on super keypoint encoding in this embodiment uses a Bayesian prior framework to model the cross-species keypoint distribution O. r ~N(ζ) r ,σ 2 I) Adaptively generate a high-level semantic superkeypoint set M = {μ1, μ2, ..., μ} using density peak clustering. L}; Joint dynamic norm pruning x pruned =M⊙x removes 50-70% of redundant visual labels and uses a multi-head self-attention mechanism MSA(T) = [SA1(T); SA2(T); ...; SA h (T)]W P Jointly learn the anatomical topological constraints of visual markers and super-keypoint markers, and finally base it on the OKS index. Outputs high-precision key point heatmaps. Compared with existing technologies, this method solves the problems of weak generalization ability, large computational redundancy, and lack of anatomical constraints caused by animal morphological diversity through a super key point-visual tag collaborative optimization architecture.

[0041] 2. The animal pose estimation system based on super keypoint encoding in this embodiment includes a prior modeling module (performing Bayesian distribution clustering to generate super keypoints), a tokenization encoding module (projecting image blocks into visual tags and adding positional encoding), a dynamic pruning module (generating binary masks based on attention scores), a cross-tag interaction module (multi-head self-attention learning anatomical topological constraints between visual tags and super keypoint tags), and a pose decoding module (parses super keypoint tags to output heatmap coordinates), constructing an end-to-end anatomical constraint-computation optimization joint pipeline. Compared to existing systems, this system overcomes the limitations of traditional animal pose estimation techniques in cross-species adaptability, real-time performance, and edge deployment, supporting millisecond-level pose analysis in complex scenarios and meeting the efficient real-time requirements of scenarios such as livestock monitoring and wildlife protection. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a flowchart of the animal pose estimation method based on super keypoint encoding as described in Embodiment 1 of the present invention;

[0044] Figure 2 This is the animal pose estimation system architecture based on super keypoint encoding as described in Embodiment 2 of the present invention;

[0045] Figure 3 This is a network structure diagram of the animal pose estimation method based on super keypoint encoding as described in Embodiment 1 of the present invention;

[0046] Figure 4 This is an implementation effect diagram of the animal pose estimation system based on super key point encoding as described in Embodiment 2 of the present invention;

[0047] Figure 5 This is an implementation effect diagram of the animal pose estimation method based on super keypoint encoding described in Embodiment 1 of the present invention; Detailed Implementation

[0048] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0049] It should be noted that when a component is said to be "fixed to" another component, it can be directly attached to the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0051] Example 1

[0052] Please refer to Figure 1 , 3 Embodiment 1 of the present invention provides an animal pose estimation method based on super keypoint encoding. This embodiment includes:

[0053] A pre-trained feature extractor f is used on the labeled image set of the input basic animal categories. θ Extract local features O from all keypoints r r And modeled as a Gaussian distribution O r ~N(ζ) r ,σ 2 I), then calculate the distribution mean ζ of the keypoints r in category y. r :

[0054]

[0055] Based on ζ r Perform density peak clustering to adaptively generate L super keypoint sets M = {μ1, μ2, ..., μ...} L Images from a labeled image set representing basic animal categories. Divided into The patch, through linear projection and by adding positional encoding, updates its encoding as follows:

[0056] v i ←vi +pe i

[0057] Construct a labeled sequence T = {[v1,…,v...} N ],[t1,…,t L ]}.

[0058] Then, the query vector q, key vector k, and value vector v are calculated:

[0059]

[0060] And generate attention matrix A:

[0061]

[0062] For the input visual label sequence {v1,…,v N Calculate the attention score ij :

[0063]

[0064] A binary mask M is generated based on a dynamic threshold τ set according to the k-th largest score. ij :

[0065] M ij =I(score) ij ≥τ)

[0066] Then perform the pruning operation:

[0067] x pruned =M⊙x

[0068] The input to the Transformer interaction uses a multi-head self-attention mechanism:

[0069] MSA(T)=[SA1(T);SA2(T);…;SA h (T)]W P ,

[0070]

[0071] The super keypoint prediction results are decoded into a keypoint heatmap and evaluated using the target keypoint similarity (OKS) metric.

[0072]

[0073] Implementing Embodiment 1 of the present invention, cross-species keypoint distribution O is modeled using a Bayesian prior framework. r ~N(ζ) r ,σ 2I) Adaptively generate a high-level semantic superkeypoint set M = {μ1, μ2, ..., μ} using density peak clustering. L}; Joint dynamic norm pruning x pruned =M⊙x removes 50-70% of redundant visual labels and uses a multi-head self-attention mechanism MSA(T) = [SA1(T); SA2(T); ...; SA h (T)]W P Jointly learn the anatomical topological constraints of visual markers and super-keypoint markers, and finally base it on the OKS index.

[0074] Outputs high-precision keypoint heatmaps. Compared to existing technologies, this method addresses the issues of weak generalization ability, large computational redundancy, and lack of anatomical constraints caused by animal morphological diversity through a super keypoint-visual tag collaborative optimization architecture.

[0075] Example 2

[0076] The animal pose estimation system based on super keypoint coding in this second embodiment protects a different subject from the animal pose estimation method based on super keypoint coding in the first embodiment. The specific differences are as follows: The animal pose estimation system based on super keypoint coding includes a prior modeling module, a tokenization coding module, a dynamic pruning module, a cross-tag interaction module, and a pose decoding module.

[0077] In one optional embodiment, the prior modeling module extracts animal keypoint features using a Bayesian distribution framework and performs density peak clustering to generate a set of super keypoints; the tokenization encoding module maps the input image into blocks as visual tags, while embedding the super keypoints as learnable tags; the dynamic pruning module calculates a binary mask based on attention scores to filter out low-weight visual tags; the cross-tag interaction module jointly optimizes the constraint relationship between visual tags and super keypoint tags through a multi-head self-attention mechanism; and the pose decoding module outputs keypoint heatmap coordinates based on the super keypoint tags after interaction.

[0078] Implementing Embodiment 2 of the present invention will have the following beneficial effects:

[0079] The system comprises a prior modeling module (performing Bayesian clustering to generate super keypoints), a tokenization and encoding module (projecting image blocks to visual markers and adding positional encodings), a dynamic pruning module (generating binary masks based on attention scores), a cross-marker interaction module (multi-head self-attention learning anatomical topological constraints between visual markers and super keypoint markers), and a pose decoding module (parses super keypoint markers to output heatmap coordinates), constructing an end-to-end anatomical constraint-computation optimization joint pipeline. Compared to existing systems, this system overcomes the limitations of traditional animal pose estimation techniques in cross-species adaptability, real-time performance, and edge deployment, supporting millisecond-level pose analysis in complex scenarios and meeting the high-efficiency, real-time requirements of scenarios such as livestock monitoring and wildlife conservation.

Claims

1. An animal pose estimation method based on super keypoint encoding, characterized in that... include: S1. Super keypoint generation: The statistical distribution features of animal keypoints are extracted through a Bayesian prior framework, and a set of super keypoints with high-level semantics is generated using the density peak clustering algorithm. S2, Tokenization representation, divides the input image into visual blocks and encodes them into visual tags with location embeddings, while mapping superkeypoints to learnable tags; S3, Dynamic Label Pruning: The dynamic norm pruning module removes redundant visual labels and retains labels with high attention weights. S4, Transformer Interaction and Prediction: The combined input of pruned visual labels and super keypoint labels to the Transformer encoder learns cross-label interaction relationships through a self-attention mechanism and outputs a keypoint heatmap.

2. The animal pose estimation method based on super keypoint encoding according to claim 1, characterized in that, The generation of super key points includes: A pre-trained feature extractor f is used on the labeled image set of basic animal categories. θ Extract local features O from all keypoints r r And modeled as a Gaussian distribution O r ~N(ζ) r ,σ 2 I), then calculate the distribution mean ζ of the keypoints r in category y. r : Where I' is the image. For keypoint coordinates, based on ζ r Perform density peak clustering to adaptively generate L super keypoint sets M = {μ1, μ2, ..., μ...} L The same supercritical point encompasses homologous anatomical sites across different species.

3. The animal pose estimation method based on super keypoint encoding according to claim 1, characterized in that, The tokenized representation includes: Images from the labeled image set of basic animal categories Divided into A patch, and for the patch p i Perform linear projection to obtain the corresponding encoding v i : And add a location code to it, updating its code as follows: V i ←v i +on i M is then mapped to learnable labels {t1,…,t} L }, output the label sequence T = {[v1,…,v...} N ],[t1,…,t L ]}.

4. The animal pose estimation method based on super keypoint encoding according to claim 1, characterized in that, The dynamic tagging and pruning includes: First, calculate the query vector q, the key vector k, and the value vector v: And generate attention matrix A: For the input visual label sequence {v1,…,v N Calculate the attention score ij : Where A ih Let H represent the attention weight between the i-th and h-th labels, where H is the total number of attention heads. norm ) hj Representing the normalized features of the h-th head and j-th dimension, a binary mask M is generated based on a dynamic threshold τ set according to the k-th largest score. ij : M ij =I(score ij ≥τ) x pruned =M⊙x Then perform the pruning operation X. pruned .

5. The animal pose estimation method based on super keypoint encoding according to claim 1, characterized in that, The Transformer interaction and prediction include: The Transformer interaction employs a multi-head self-attention mechanism as follows: MSA(T)=[SA1(T);SA2(T);…;SA h (T)]W P in d is the labeled dimension, d i =d.

6. The method according to claim 2, characterized in that: The density peak clustering does not preset the number of clusters, and adaptively identifies densely distributed key point regions based on k-nearest neighbor distance. Different semantic key points within the same species are forcibly assigned to different super key points.

7. The method according to claim 3, characterized in that: The dynamic threshold τ is determined by selecting the k-th maximum score, satisfying τ = topk(score). ij Furthermore, the number of visual markers after pruning is 50%-70% of the original number.

8. The method according to claim 1, characterized in that: Keypoint prediction results are evaluated using the Target Keypoint Similarity (OKS) metric: Where d i To predict the Euclidean distance from the ground truth coordinates, s is the square root of the object's bounding box area, and k is... i This is the keypoint type normalization constant.

9. An animal pose estimation system based on super keypoint encoding, used in the animal pose estimation method based on super keypoints as described in any one of claims 1-7, characterized in that, include: The prior modeling module extracts animal keypoint features through a Bayesian distribution framework and performs density peak clustering to generate a super keypoint set. The tokenization encoding module maps the input image into blocks as visual tags, and embeds super key points as learnable tags. A dynamic pruning module calculates a binary mask based on attention scores to filter out low-weight visual markers; A cross-tag interaction module, which jointly optimizes the constraint relationship between visual tags and super keypoint tags through a multi-head self-attention mechanism; The attitude decoding module outputs the key point heatmap coordinates based on the super key point markers after interaction.

Citation Information

Patent Citations

  • Wild animal attitude estimation method in real scene

    CN117912116A