Animal pose estimation method based on key point perception enhancement

By constructing the RCBPose network, introducing the RepLK module and the BRA mechanism, and using coordinate-separable convolution and a two-layer routed attention mechanism, the occlusion and background interference problems in animal pose estimation under complex backgrounds are solved, achieving higher robustness and accuracy.

CN121259706BActive Publication Date: 2026-08-25KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511306830.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-08-25
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing animal pose estimation methods are inaccurate in complex farm environments due to occlusion and background interference, resulting in inaccurate pose and key point information. In particular, they are difficult to accurately identify and track individual instances in multi-animal interaction scenarios.

Method used

An animal pose estimation network, RCBPose, is constructed using an improved Darknet-53 backbone network, CSConv module, Neck network, and Head part. The RepLK module and BRA mechanism are introduced, and the ability to perceive key points and learn global structure is enhanced through coordinate separable convolution and two-layer routing attention mechanism.

Benefits of technology

It significantly improves the robustness and accuracy of the model in handling complex backgrounds and local key points, and can better adapt to changes in the body size and posture of different animals, thereby improving the accuracy of multi-target animal posture estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121259706B_ABST
    Figure CN121259706B_ABST
Patent Text Reader

Abstract

The application relates to an animal posture estimation method based on key point perception enhancement and belongs to the field of computer vision. The method comprises the following steps: constructing and training an animal posture estimation network RCBPose; wherein the RCBPose comprises an improved Darknet-53 backbone network, a CSConv module, a Neck network and a Head part; the improved Darknet-53 backbone network is used for performing feature extraction on an image to be subjected to animal posture estimation, so as to obtain a high-dimensional feature map; the CSConv module is used for performing convolution processing on the high-dimensional feature map, so as to output a multi-scale feature map; the Neck network is used for performing feature fusion on the multi-scale feature map, so as to obtain fused features containing animal key point positions and global position information; and the Head part is used for identifying and predicting the fused features, so as to obtain spatial position information between animal key points and complete animal posture estimation. The application aims to solve the technical problem that, in the prior art, animal posture and key point information are inaccurate due to occlusion and background interference in a complex breeding farm environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an animal pose estimation method based on key point perception enhancement, belonging to the field of computer vision. Background Technology

[0002] Animal pose estimation is fundamental to behavior recognition research, as pose estimation techniques can accurately describe an animal's movement trajectory and posture changes. However, existing methods still face challenges in capturing global body structure information and the spatial relationships of local keypoints, such as complex background interference, animal body size diversity, and missing local keypoints, which affect the robustness and accuracy of the models.

[0003] In behavior recognition research, animal behavior is often manifested through the spatiotemporal dynamic changes of its keypoint movements. Using the keypoint coordinate sequence obtained from 2D pose estimation, a spatiotemporal skeleton map can be further constructed, providing a basis for identifying animal movement patterns and behavior categories. For example, by analyzing the relative position changes and movement trajectories of keypoints in consecutive frames, basic animal movements (such as walking, running, and standing) as well as more complex behavioral patterns (such as foraging and attacking) can be distinguished. Furthermore, 2D pose estimation can be combined with data from other modalities (such as RGB video, infrared images, or depth information) to form a multimodal fusion method, further improving the accuracy and robustness of behavior recognition. This multimodal fusion method can fully utilize the complementary advantages of different data sources when dealing with complex backgrounds, occlusion, and individual differences, achieving more accurate behavior detection and abnormal action recognition. Based on the number of target animals in the scene, pose estimation methods can be divided into single-target animal pose estimation and multi-target animal pose estimation.

[0004] Single-animal pose estimation aims to predict the pose of a single animal by detecting keypoints in videos or images. To classify the poses of individual cattle, Khin et al. proposed a cattle pose classification system based on DeepLabCut and SVM. First, the DeepLabCut network was used to extract keypoints from the video, and then an SVM classifier was used to classify the features, identifying six poses. This method achieved an overall classification accuracy of 88.75% on actual dairy farm video sequences. However, during the transition from standing to sitting and from sitting to stretching, some poses may appear simultaneously with other poses, leading to misclassification and increasing the difficulty of classification. Li et al. developed three deep cascaded convolutional neural networks, including a convolutional pose machine model, a stacked dense hourglass network, and a convolutional heatmap regression model, for robust cattle pose estimation. This method achieved an average accuracy of 90.39% on 16 keypoints, but could not accurately estimate similar poses. Zhang et al. proposed an animal pose estimation algorithm based on SHN, designing an efficient residual module and a lightweight dual-branch feature fusion module to reduce network computation and parameter size. Although this method has achieved good performance in animal pose estimation tasks, the accuracy of pose estimation needs to be improved in the case of animal occlusion and complex backgrounds. When dealing with multiple individual animals, the model performance needs to be validated and enhanced.

[0005] Multi-animal pose estimation (MAE) needs to handle more complex scenes because it requires identifying and locating each target animal and accurately distinguishing keypoints between different target animals. MAE can be divided into top-down and bottom-up approaches. The bottom-up approach first uses keypoint detection algorithms to identify keypoints for all animal targets in the image, and then combines these points to match the correct animal target. Therefore, the bottom-up approach focuses on keypoint algorithms and is more effective in scenes with multiple interacting animals. However, during the keypoint association stage, overlap between animals can lead to the inability to correctly identify and track individual instances. While the bottom-up approach has certain advantages in handling multi-animal interaction scenes, avoiding the initial detection step for each animal, the appearance differences between animals and small individuals often pose challenges to keypoint localization. Furthermore, when dealing with multiple individual animals, especially in complex environments, the overlap between individuals often makes it difficult for the bottom-up approach to accurately identify the animal's pose during the keypoint association stage. Summary of the Invention

[0006] The purpose of this invention is to provide an animal pose estimation method based on key point perception enhancement, which aims to solve the technical problem that existing methods are inaccurate in animal pose and key point information due to occlusion and background interference in complex farm environments.

[0007] To achieve the above objectives, the technical solution of the present invention is: an animal pose estimation method based on keypoint perception enhancement, the method comprising the following steps:

[0008] Step 1: Construct and train the animal pose estimation network RCBPose; wherein, the RCBPose includes an improved Darknet-53 backbone network, CSConv module, Neck network and Head part;

[0009] Step 2: Input the image of the animal pose estimation to be performed into the trained RCBPose, and extract features through the improved Darknet-53 backbone network to obtain a high-dimensional feature map containing animal pose information; wherein, the improved Darknet-53 backbone network is the original Darknet-53 backbone network with the introduction of RepLK module and BRA mechanism.

[0010] Step 3: The high-dimensional feature map containing animal pose information is convolved using the CSConv module to output a multi-scale feature map; wherein the convolution process is coordinate convolution and depthwise separable convolution in sequence.

[0011] Step 4: Perform feature fusion on the multi-scale feature map through the Neck network to obtain fused features that include animal key point locations and global location information;

[0012] Step 5: Identify and predict the fused features through the Head part to obtain the spatial location information between animal key points and complete animal posture estimation.

[0013] Optionally, the improved Darknet-53 backbone network adopts a C2f architecture.

[0014] Optionally, the CSConv module is located between the backbone network and the Neck, and is used to integrate several feature maps of different scales across layers.

[0015] Optionally, the Neck network uses a PANet structure to perform upsampling and downsampling operations on the multi-scale feature maps.

[0016] Optionally, the Head section adopts a decoupled structure design, separating the classification head from the detection head.

[0017] Optionally, the training animal pose estimation network RCBPose adopts a positive sample matching strategy and incorporates a focus loss function. It also introduces data augmentation operations in the first few preset training cycles and stops data augmentation operations in the last few preset training cycles.

[0018] Optionally, the coordinate convolution specifically includes:

[0019] The input high-dimensional feature map M containing animal posture information ),in , These are the width and height of the high-dimensional feature map, respectively. To determine the number of channels, firstly, two coordinate channels x and y are generated for M, representing the horizontal and vertical coordinates of each pixel, respectively. Then, the horizontal and vertical coordinates are normalized to the range [-1, 1]. Finally, they are concatenated with the input high-dimensional feature map M to form a new high-dimensional feature map M1. ).

[0020] Optionally, the depthwise separable convolution includes depthwise convolution and pointwise convolution, specifically:

[0021] After coordinate convolution processing, the output is a new high-dimensional feature map M1. The input channels are passed to depthwise separable convolution for processing. In the depthwise convolution stage, a convolution kernel is applied independently to each input channel. The expression for depthwise convolution is:

[0022]

[0023] In the formula, M2 is the output feature map after depthwise convolution; i and j represent the row and column indices of the output feature map, respectively, and k represents the number of channels in the feature map. This represents the depthwise convolution kernel, where m and n correspond to the offsets of the convolution kernel in the spatial dimension, respectively.

[0024] In the pointwise convolution stage, a 1×1 convolution kernel is used to convolve the output of the depthwise convolution to adjust the number of channels in the output feature map and combine information from different input channels. The expression is as follows:

[0025]

[0026] In the formula, M3 is the output feature map after point convolution, and W... 1,1,l,t The parameters of the point convolution kernel are denoted as , where the kernel size is 1×1, t is the number of channels in the feature map, and l represents the channel index of the input feature map.

[0027] Compared with the prior art, the beneficial effects of the present invention are:

[0028] (1) By introducing the RepLK (Revisiting Large Kernel) module, this invention significantly expands the receptive field of the model, thereby enhancing its implicit learning ability of the animal’s overall body structure, enabling it to better adapt to changes in the body size and posture of different animals.

[0029] (2) This invention designs a Coordinate Separable Convolution (CSConv), which enhances the ability to perceive key point coordinates by introducing coordinate information. At the same time, in order to deal with the problems of blurred local key points and interference from complex backgrounds, RCBPose incorporates a Bi-Level Routing Attention (BRA) mechanism, which significantly improves the model's ability to perceive local key points and blurred body parts by dynamically allocating attention perception. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the animal RGB images and their 16 annotation information in the dataset used in this invention;

[0031] Figure 2 This is a network structure diagram of the animal pose estimation method based on key point perception enhancement of the present invention;

[0032] Figure 3 This is a structural diagram of the RepLK module of this invention;

[0033] Figure 4 This is a structural diagram of the BRA attention mechanism of the present invention;

[0034] Figure 5 This is a structural diagram of the standard convolution and CSConv convolution of this invention. Detailed Implementation

[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0036] Example 1, as Figure 2 As shown, an animal pose estimation method based on keypoint perception enhancement includes the following steps:

[0037] Step 1: Construct and train the animal pose estimation network RCBPose; wherein, the RCBPose includes an improved Darknet-53 backbone network, CSConv module, Neck network and Head part;

[0038] Optionally, RCBPose in this embodiment can be built on top of YOLOv8;

[0039] Optionally, the training animal pose estimation network RCBPose adopts a positive sample matching strategy and incorporates a distribution focal loss function. During training, it uses the YOLOX training method, which includes advanced techniques for handling imbalanced datasets and improving model robustness. Data augmentation operations are introduced in the first few preset training cycles and stopped in the last few preset training cycles. In this embodiment, data augmentation operations are stopped in the last 10 training cycles.

[0040] It is understandable that this embodiment helps increase data diversity by introducing data augmentation operations in the early stages of training, but stopping this operation in the later stages of training allows the model to focus on learning finer details, reducing noise interference that may be introduced during the augmentation process, thereby further improving the accuracy of the model.

[0041] Step 2: Input the image of the animal pose estimation to be performed into the trained RCBPose, and extract features through the improved Darknet-53 backbone network to obtain a high-dimensional feature map containing animal pose information; wherein, the improved Darknet-53 backbone network is the original Darknet-53 backbone network with the introduction of RepLK module and BRA mechanism.

[0042] Optionally, the images for animal pose estimation in this embodiment can be selected from three representative publicly available animal datasets: NWAFU-Cattle, Horse-10, and Lamb.

[0043] The NWAFU-Cattle dataset contains 2134 high-quality images of dairy and beef cattle, each image containing at least one individual with partial body occlusion. This dataset not only provides accurate bounding boxes but also covers three typical behavioral poses (standing, walking, and lying down), and these images were collected from diverse environmental scenes.

[0044] The Horse-10 dataset is a benchmark dataset specifically designed for animal pose estimation tasks. It contains 8,114 frames of images covering 30 different breeds of horses. The dataset features detailed annotations for 22 key body parts in each image, and was collected from various environments including wild and farm settings, including individuals with different coat colors.

[0045] The Lamb dataset has certain limitations in terms of sample diversity and environmental coverage. It contains only 400 images of lambs collected in a farm environment. This environmental uniformity may affect the generalization ability of the model and limit its application in a wider range of scenarios.

[0046] Optionally, in this embodiment, to ensure consistency and comparability, 16 key body points are marked for each type of animal, such as... Figure 1 As shown.

[0047] Optionally, the improved Darknet-53 backbone network adopts a C2f architecture.

[0048] It is important to understand that Darknet-53 effectively solves the gradient vanishing problem during training by extensively using residual connections to build deep network structures, thus significantly improving the model's convergence performance.

[0049] Furthermore, the C2f structure achieves model lightweighting while ensuring a rich gradient information flow by adding residual connections. This improvement not only maintains the model's high performance but also reduces the computational load, making it more suitable for real-time applications.

[0050] Furthermore, the RepLK module significantly expands the receptive field using large-size convolutional kernels. Initial feature extraction is performed first through two parallel initial CSB blocks. Each CSB block contains a 2D convolutional layer (Conv2d), a batch normalization layer (BN), and a SiLU activation function. The 2D convolutional layer handles the convolution operation and extracts features from the input data; the batch normalization layer (BN) normalizes the convolutional data to improve model stability and convergence speed. SiLU, as the activation function, introduces non-linearity, enabling the model to learn more complex mapping relationships. The feature maps processed by the two CSB blocks are concatenated along the channel dimension to obtain a fused feature map, with the overall structure as shown below. Figure 3 As shown;

[0051] Furthermore, the Two-Level Route Attention (BRA) mechanism is a dynamic query and perceptual sparse attention mechanism that effectively filters out most irrelevant key-value pairs in coarse regions, and then applies fine-grained token-to-token attention within the joint routing region. The overall structure of BRA is as follows: Figure 4 As shown.

[0052] First, BRA will input the feature map ( (H and W represent the length and width, and C represents the number of channels) is divided into Several different regions, among which This represents the dimensionality factor of the partition, where each region contains a feature vector. ,but Become Q, K, V are obtained through linear mapping. The expression is:

[0053]

[0054] In the formula, W q W k W v These are the projection weights for query, key, and value, respectively.

[0055] Then, a directed graph is constructed using the adjacency matrix to find the regions that each given region should focus on and participate in. Simply put, the average of Q and K for each region is first calculated to obtain Q for each region. r K r ( Then, the transpose multiplication method is used to calculate the region Q. r and K r The adjacency matrix A of the correlation between them r =Q r (K T ) T The correlation graph is pruned by retaining only the top k connections of each region. Then, the directed path topkIndex(A) is used to prune the correlation graph. r Select the k most relevant region indices for each region to obtain the region index matrix I. r ( This leads to finer-grained token-to-token attention; each query token in a region only focuses on the token generated by I. r (i,1) , I r (i,2) , …, I r (i,k) All key-value pairs in the k routing regions of the index. To facilitate GPU computation, dense matrix multiplication is used to aggregate the key-value tensor, expressed as:

[0056]

[0057] In the formula, K g and V g It is a tensor of aggregated keys and values, containing only the tokens of the top k most relevant regions; gather(∙) represents collecting the corresponding key-value pairs by index; I r This represents the index of the top k related regions corresponding to each region token.

[0058] Finally, an attention operation is performed on the aggregated key-value pairs, expressed as:

[0059]

[0060] In the formula, O represents the final output result; Attention(Q, K) g V g ) represents the aggregated K g Vg Perform standard attention operations with query Q; LCE(V) is a general building block for multi-head context aggregation, used for context enhancement.

[0061] Specifically, taking the Horse-10 dataset as an example, the horse images in the dataset are input into the RCBPose network, where the improved Darknet-53 backbone network first extracts features. This backbone network introduces a C2f structure, a RepLK module, and a BRA mechanism on the basis of the original structure, thereby improving the feature representation capability and finally obtaining a high-dimensional feature map containing horse posture information.

[0062] Understandably, the improved Darknet-53 backbone network significantly enhances the robustness and accuracy of the model in handling complex backgrounds and local key points by using the RepLK module in series with the BRA mechanism.

[0063] Step 3: The high-dimensional feature map containing animal pose information is convolved using the CSConv module to output a multi-scale feature map; wherein the convolution process is coordinate convolution and depthwise separable convolution in sequence.

[0064] It's important to understand that the CSConv module not only preserves spatial location information but also significantly improves computational efficiency. Furthermore, by appending coordinate channels to the input feature map, it enhances standard convolution, enabling the network to perceive the location of each pixel, thereby improving performance on spatially relevant tasks. The overall structure of the CSConv module is as follows: Figure 5 As shown.

[0065] Optionally, the CSConv module is located between the backbone network and the Neck, and is used to integrate several feature maps of different scales across layers.

[0066] Optionally, the coordinate convolution specifically includes:

[0067] The input high-dimensional feature map M containing animal posture information ),in , These are the width and height of the high-dimensional feature map, respectively. To determine the number of channels, firstly, two coordinate channels x and y are generated for M, representing the horizontal and vertical coordinates of each pixel, respectively. Then, the horizontal and vertical coordinates are normalized to the range [-1, 1]. Finally, they are concatenated with the input high-dimensional feature map M to form a new high-dimensional feature map M1. ).

[0068] Optionally, the depthwise separable convolution includes depthwise convolution and pointwise convolution, specifically:

[0069] After coordinate convolution processing, the output is a new high-dimensional feature map M1. The input is passed to a depthwise separable convolution for processing. In this embodiment, the expression for the depthwise convolution is:

[0070]

[0071] In the formula, M2 is the output feature map after depthwise convolution; i and j represent the row and column indices of the output feature map, respectively, and k represents the number of channels in the feature map. This represents the depthwise convolution kernel, where m and n correspond to the offsets of the convolution kernel in the spatial dimension, respectively.

[0072] Understandably, depthwise separable convolution is an efficient convolutional operation that can reduce computational costs and model parameters while maintaining or even improving overall model performance. This is achieved by breaking down traditional convolutional operations into two smaller, more efficient operations: depthwise convolution (DW) and pointwise convolution (PW). In the depthwise convolution stage, a convolutional kernel is applied independently to each input channel. This means that if the input feature map has C+2 channels, there will be C+2 convolutional kernels. Depthwise convolution is performed only in the spatial dimension and does not require merging information across channels, thus significantly reducing computational costs compared to traditional convolution.

[0073] In the pointwise convolution stage, a 1×1 convolution kernel is used to convolve the output of the depthwise convolution to adjust the number of channels in the output feature map and combine information from different input channels. The expression is as follows:

[0074]

[0075] In the formula, M3 is the output feature map after point convolution, and W... 1,1,l,t The parameters of the point convolution kernel are denoted as , where the kernel size is 1×1, t is the number of channels in the feature map, and l represents the channel index of the input feature map.

[0076] Understandably, the kernel size used in the point convolution stage is 1×1, but it spans all input channels. Therefore, if the output of a depthwise convolution has k channels, the point convolution will apply k 1×1×l kernels, where k is the number of output channels of the point convolution layer.

[0077] Specifically, taking the Horse-10 dataset as an example, after feature extraction, the feature maps are fed into the CSConv module, which integrates feature maps of three different scales across layers. This module combines the spatial awareness of coordinate convolution with the efficient computational characteristics of depthwise separable convolution, enabling the model to more accurately capture the spatial relationships between key points and output multi-scale feature maps.

[0078] Step 4: Perform feature fusion on the multi-scale feature map through the Neck network to obtain fused features that include animal key point locations and global location information;

[0079] Optionally, the Neck network uses a PANet (Path Aggregation Network) structure to perform upsampling and downsampling operations on the multi-scale feature maps.

[0080] Specifically, PANet introduces a cross-layer fusion method to combine low-resolution shallow feature maps containing location information with high-resolution deep feature maps rich in semantic information. This feature fusion process significantly enhances the model's representational power by passing feature information along a specific path and propagating low-level semantic features upwards. These operations collectively improve the expressive power of multi-scale features, enabling the model to better capture the complexity and subtle differences in livestock postures in dynamic environments.

[0081] Specifically, taking the Horse-10 dataset as an example, the multi-scale feature map output by the CSConv module is further input into the Neck network to finally obtain a fused feature containing the location of horse body key points and global location information.

[0082] Step 5: Identify and predict the fused features through the Head part to obtain the spatial location information between animal key points and complete animal posture estimation.

[0083] Optionally, the Head section adopts a decoupled structure design, separating the classification head from the detection head.

[0084] Specifically, by separating the classification head from the detection head, the Head part facilitates more accurate feature extraction and pixel-level prediction, enabling the network to handle different scales and semantic information more flexibly. This decoupled structure ensures the model's effectiveness in handling multi-scale targets, particularly in complex scenes involving multiple overlapping and interacting animal poses, significantly improving segmentation capabilities.

[0085] Specifically, taking the Horse-10 dataset as an example, the fused features processed by the Neck network are finally input into the Head part. The classification head is responsible for identifying the type of key points (such as left eye, right ear, etc.), while the detection head focuses on the regression prediction of the specific location of the key points, thereby obtaining the precise spatial location information between the key points of the horse.

[0086] In summary, this invention designs an animal pose estimation network, RCBPose. To efficiently capture subtle changes in local keypoints of animals under different environments, this invention employs a coordinate-separable convolution, enhancing the model's ability to perceive the coordinates of keypoints on the animal's body, thereby overcoming the problem of keypoint localization caused by occlusion in complex backgrounds. Simultaneously, the introduction of a large convolutional kernel module significantly expands the global receptive field, improving the model's perception of changes in the overall joint structure of the animal and further uncovering the intrinsic positional relationships between keypoints. To efficiently filter irrelevant keypoint pairs and increase attention to key regions, this invention adopts a two-layer routing attention mechanism, BRA, which performs more refined calculations on the interrelationships of keypoints within different key regions, thereby achieving flexible allocation of computational resources and reducing the overall computational load.

[0087] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. An animal pose estimation method based on keypoint perception enhancement, characterized in that, The method includes the following steps: Step 1: Construct and train the animal pose estimation network RCBPose; wherein, the RCBPose includes an improved Darknet-53 backbone network, CSConv module, Neck network and Head part; Step 2: Input the image of the animal pose estimation to be performed into the trained RCBPose, and extract features through the improved Darknet-53 backbone network to obtain a high-dimensional feature map containing animal pose information; wherein, the improved Darknet-53 backbone network is the original Darknet-53 backbone network with the introduction of RepLK module and BRA mechanism. Step 3: The high-dimensional feature map containing animal pose information is convolved using the CSConv module to output a multi-scale feature map; wherein the convolution process is coordinate convolution and depthwise separable convolution in sequence. Step 4: Perform feature fusion on the multi-scale feature map through the Neck network to obtain fused features that include animal key point locations and global location information; Step 5: Identify and predict the fused features through the Head part to obtain the spatial location information between animal key points and complete animal posture estimation.

2. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The improved Darknet-53 backbone network adopts a C2f architecture.

3. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The CSConv module, located between the backbone network and Neck, is used to integrate several feature maps of different scales across layers.

4. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The Neck network uses a PANet structure to perform upsampling and downsampling operations on the multi-scale feature maps.

5. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The Head section adopts a decoupled structure design, separating the classification head from the detection head.

6. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The training animal pose estimation network RCBPose adopts a positive sample matching strategy and incorporates a focus loss function. It introduces data augmentation operations in the first few preset training cycles and stops data augmentation operations in the last few preset training cycles.

7. The animal pose estimation method based on keypoint perception enhancement according to claim 1, characterized in that, The coordinate convolution is specifically as follows: The input high-dimensional feature map M containing animal posture information ),in , These are the width and height of the high-dimensional feature map, respectively. To determine the number of channels, firstly, two coordinate channels x and y are generated for M, representing the horizontal and vertical coordinates of each pixel, respectively. Then, the horizontal and vertical coordinates are normalized to the range [-1, 1]. Finally, they are concatenated with the input high-dimensional feature map M to form a new high-dimensional feature map M1. ).

8. The animal pose estimation method based on key point perception enhancement according to claim 7, characterized in that, The depthwise separable convolution includes depthwise convolution and pointwise convolution, specifically: After coordinate convolution processing, the output is a new high-dimensional feature map M1. The input channels are passed to depthwise separable convolution for processing. In the depthwise convolution stage, a convolution kernel is applied independently to each input channel. The expression for depthwise convolution is: ; In the formula, M2 is the output feature map after depthwise convolution; i and j represent the row and column indices of the output feature map, respectively, and k represents the number of channels in the feature map. This represents the depthwise convolution kernel, where m and n correspond to the offsets of the convolution kernel in the spatial dimension, respectively. In the pointwise convolution stage, a 1×1 convolution kernel is used to convolve the output of the depthwise convolution to adjust the number of channels in the output feature map and combine information from different input channels. The expression is as follows: ; In the formula, M3 is the output feature map after point convolution, and W... 1,1,l,t The parameters of the point convolution kernel are denoted as , where the kernel size is 1×1, t is the number of channels in the feature map, and l represents the channel index of the input feature map.