Human body posture key point recognition method based on feature enhancement high resolution
By introducing a hierarchical residual connection structure and multi-scale convolutional attention module into the human pose estimation model, FE-HRNet is constructed, and the problems of robustness and computational complexity of human pose estimation in population-intensive scenarios are solved, and key point positioning with higher accuracy and lower complexity are achieved.
Patent Information
- Application Number
- CN202510468055.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
The existing top-down human posture estimation method is not robust enough in population-intensive scenarios, has high computational complexity, and is difficult to achieve high-precision key point positioning while maintaining moderate parameters.
The hierarchical residual connection structure and multi-scale convolutional attention module are used to enhance feature representation capabilities, build a FE-HRNet model, and improve the accuracy and robustness of key points through multi-scale feature fusion and adaptive attention mechanism.
Achieve higher detection accuracy and lower computational complexity in dense population environments, significantly improving the detection capability and positioning accuracy of the model in complex poses.
Smart Images

Figure CN120375423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image recognition technology, and specifically to a method for identifying key points of human body postures based on feature enhancement and high resolution. Background Art
[0002] Human Pose Estimation (HPE) technology aims to locate human body parts from input data such as images and videos and construct a human body representation. HPE technology can be divided into two categories: two-dimensional human pose estimation (2D HPE) and three-dimensional human pose estimation (3D HPE). 2D HPE aims to accurately locate and identify the positions of key joints of the human body from two-dimensional images or video sequences. This type of method mainly uses a convolutional neural network (CNN) to process the input image and predicts the positions and connection relationships of the joints by generating heatmaps or vector fields. The existing network architectures mainly include two methods: the top-down method and the bottom-up method. The former first detects the human body and then estimates the pose; the latter directly estimates all visible joints and then groups them. 3D HPE aims to reconstruct the three-dimensional skeletal structure and joint positions of the human body from two-dimensional images or video sequences and construct a complete human skeletal model in three-dimensional space. Different from 2D HPE based on a single-direction image, 3D HPE needs to combine multi-direction images for combined analysis. 2D HPE mainly includes two research directions: single-person pose estimation and multi-person pose estimation. Single-person pose estimation mainly focuses on the accurate recognition of a single human body pose in an image. With the wide application of CNN, researchers have designed more refined network structures to capture the diverse changes of human joints. However, when multiple human bodies appear in the image, the problem complexity increases significantly. It is not only necessary to identify the joint points of each individual but also to accurately associate these joint points with the corresponding human body. To solve the multi-person pose estimation problem, researchers have proposed two main methods: the bottom-up method and the top-down method. The bottom-up method first detects all possible key points in the image and then combines the key points of the same individual into a complete human body pose. This method can effectively handle human body overlaps, but it needs to detect and cluster key points for a large number of pixels, and the computational cost is relatively high. The top-down method first detects the human body region in the image through a human body detector and estimates the pose for each region. This method can handle the pose estimation problem at different scales well, but the performance of the human body detector directly affects the accuracy of subsequent pose estimation.
[0003] In the top-down approach, Xiao et al. incorporated deconvolution layers into the ResNet backbone to construct an architecture capable of effectively representing high-resolution heatmaps; on this basis, Wang et al. proposed the High-Resolution Network (HRNet), which learns high-resolution representations by parallel processing and multi-scale fusion of features at different resolutions, while maintaining high-resolution feature maps in the main branch all the time, further improving the pose estimation performance and laying a foundation for subsequent research; Huang et al. proposed an unbiased data processing method (UDP) to address the quantization error problem in human pose estimation tasks, which solved the bias problem in standard data processing through unbiased coordinate system transformation and unbiased keypoint format transformation. To reduce the computational complexity of the pose estimation model, Yu et al. proposed LiteHRNet based on HRNet, exploring the lightweight high-resolution network design for the first time and achieving efficient feature fusion through the conditional channel weighting mechanism, providing a new solution for keypoint localization; Li et al. further optimized the design of Dite-HRNet, enhancing the spatial feature expression through dynamic kernel aggregation and introducing adaptive context pooling to improve the keypoint localization accuracy; Zhang et al. proposed a lightweight high-resolution human pose estimation network HF-HRNet for mobile devices, which addressed the problem of slow inference speed of existing methods on devices such as NPUs by designing a hardware-friendly HUM module (including two core components, CAD and MSCA), significantly improving the model performance while ensuring the pose estimation accuracy. With the successful application of Transformer in computer vision, Transformer-based human pose estimation algorithms have been proposed one after another. TransPose uses a Transformer encoder to learn the dependencies between pixels; TokenPose uses each keypoint embedding as a token and learns constraints and appearance cues from the image at the same time; Hrformer replaces the residual modules in the HRNet model with Transformer modules to model the global context information, thus improving the accuracy of keypoint localization; ViTPose uses the encoder of the vision Transformer to capture the global information of the human body region and combines a lightweight decoder to accurately predict the keypoint positions, achieving excellent performance in human pose estimation tasks. In terms of optimizing the feature enhancement mechanism, ECANet proposed by Wang et al. achieved significant performance improvement while minimizing the number of parameters through the local cross-channel interaction strategy; Bao et al. organically combined ECANet with HRNet, showing excellent performance in complex scenarios such as fast movement.These works provide important inspirations for the design of efficient pose estimation networks; Yuan et al. proposed an efficient sparse attention mechanism NSA, which effectively solved the problem of low efficiency of traditional attention mechanisms in processing context by combining coarse-grained token compression, fine-grained token selection, and sliding window techniques.
[0004] During the top-down pose estimation process, multiple human instances within a single bounding box are predicted to handle overlaps, but the performance improvement is still limited. As the backbone network for basic feature extraction, it plays a key role in feature representation learning and overlap processing. However, the performance of most current top-down methods is limited by the capacity of traditional backbone networks. Existing studies generally improve accuracy by increasing the complexity of model parameters, but achieving high accuracy while maintaining a moderate number of parameters is still an important challenge, and the performance robustness in crowded human scenarios needs to be improved urgently. Summary of the Invention
[0005] The object of the present invention is to provide a method for identifying human pose key points based on feature-enhanced high resolution in view of the deficiencies of the prior art. This method has higher accuracy than existing algorithms in crowded environments, and the model parameter quantity and the required floating-point operation quantity of this method are superior to existing key point detection algorithms. This method shows stronger detection ability and higher positioning accuracy when dealing with complex poses.
[0006] The technical solution for achieving the object of the present invention is as follows:
[0007] A method for identifying human pose key points based on feature-enhanced high resolution, comprising the following steps:
[0008] 1) Identify the body pose: Based on the images extracted according to the teaching duration of knowledge points in the smart classroom scenario, use the YOLOv11 algorithm to detect and identify all learner bounding frames in the images. The target bounding box consists of four elements {x min , y min , w, h}, where x min , y min represent the x coordinate and y coordinate of the upper left corner of the bounding box respectively, and w and h represent the width and height of the bounding box respectively. Input the images extracted according to the teaching duration of knowledge points in the smart classroom scenario into the backbone network Feature-Enhanced High-Resolution Network (FE-HRNet for short) to extract the two-dimensional coordinate positions of 13 key points including the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, and left and right hips, and accurately identify the body pose;
[0009] 2) Enhance the feature representation ability of the learner's pose key points: Input the image features extracted in step 1) into the backbone network FE-HRNet, and adopt the adaptive fusion of multi-scale features to achieve dynamic feature enhancement in the spatial and channel dimensions, including:
[0010] 2-1) Construct a hierarchical class residual connection structure to achieve fine-grained multi-scale feature representation, effectively expanding the receptive field range of the network. A multi-scale feature extraction module is adopted, which significantly enhances the representation ability of the network while maintaining a similar computational complexity. A set of smaller convolutional kernels are used to replace the original 3×3 convolution, and a hierarchical residual connection is used to connect convolutional groups of different scales. In each hierarchical class residual connection structure, the input feature map first undergoes a convolution operation with a 1×1 convolutional kernel and is equally divided into s feature subsets, denoted as x i , where i ∈ {1, 2,... s}, and each feature subset x i has the same spatial dimension as the input feature map. Except for x1, each x i has a corresponding 3x3 convolution operation, denoted as K i , and the output of K i is represented by y i . The feature subset x i is added to the output of K i-1 (), and then fed into K i (). To reduce parameters while increasing s, the 3×3 convolution operation of x1 is omitted, and it is written as shown in formula (1):
[0011]
[0012] Through the above design, each 3×3 convolution K i ( ) in the hierarchical class residual connection structure can receive feature information from all feature partitions {x j , j ≤ i}. Each time the feature partition x j passes through a 3×3 convolution operator, the receptive field of its output features will be correspondingly expanded. As the features interact and fuse repeatedly between different branches, the receptive field of the convolutional kernel is also continuously expanding. By introducing the size control parameter s, s = 4, as Figure 2 shown, this structure can flexibly adjust the granularity of feature partitioning, enabling the network to perform feature learning and extraction within a wider receptive field range while maintaining a low computational complexity and storage overhead;
[0013] 2-2) Construct a multi-scale convolutional attention module, capture spatial context information at different scales through multi-branch depthwise separable convolutions, and combine the channel attention mechanism to adaptively enhance key features, significantly improving the localization ability of human key points. First is the 3×3 depthwise separable convolutional layer, which is used to aggregate local information of the input features. Second is the multi-branch depthwise strip convolutional structure. Each branch uses two orthogonal 1D depthwise separable convolutions to decompose the standard convolution, which are (1×7, 7×1), (1×11, 11×1), and (1×21, 21×1) respectively. This decomposition strategy not only significantly reduces the computational complexity but also efficiently captures multi-scale context information through approximate 2D convolution operations of 7×7, 11×11, and 21×21. Finally, a 1×1 convolutional layer is used to model the dependencies between different channels, and the output is directly used as the attention weight to perform element-wise multiplication with the input features. Through the adaptive fusion of multi-scale features, dynamic feature enhancement in both spatial and channel dimensions is achieved, which can effectively handle object and detail feature expressions at different scales. The formal representation is shown in Formulas (2) and (3):
[0014]
[0015] where Att and Out are the attention map and output respectively, Scale i , i ∈ {1, 2, 3} represents the i-th branch, DW-Conv represents depthwise convolution, is the multiplication operation of element matrices, and F represents the input feature;
[0016] 2-3) Embed the hierarchical residual structure and multi-scale convolution into the high-resolution network HRNet (High-Resolution Network). FE-HRNet constructs four parallel feature branches. The resolutions of the four parallel branches from top to bottom are sampled to 1, 1 / 2, 1 / 4, and 1 / 8 of the input image resolution in sequence, and the corresponding feature channel numbers are expanded to 32, 64, 128, and 256 in sequence. In the feature extraction process, the feature map resolution is first reduced to 1 / 4 of the input image resolution through two 3×3 convolutional layers with a stride of 1, and then feature extraction is carried out through four stages: In the first stage, a hierarchical residual connection structure is introduced, and fine-grained feature extraction is achieved through 4 improved residual units; the second, third, and fourth stages each contain 1, 4, and 3 multi-resolution modules respectively. Each module consists of two parts: parallel multi-resolution convolution and multi-resolution fusion, and contains 4 residual units. Among them, each unit contains 2 3×3 convolutions in different resolution branches. At the end of each stage, a multi-scale convolutional attention module is connected to extract more detailed category and spatial context information. Finally, multi-scale feature fusion is used to generate the key point heat map, and the heat map is shown in Formula (4):
[0017]
[0018] where f t-3 ,f t-2 ,f t-1 ,f t respectively represent the corresponding outputs of the original input sizes with resolutions of 1 / 8, 1 / 4, 1 / 2, and 1 in the fourth stage of the network. Conv is the convolution operation;
[0019] 3) Training: Training is carried out using 1 NVIDIA GeForce RTX 4090 graphics processing unit. The adaptive moment estimation optimizer, i.e., the Adam optimizer, and the mean square error (MSE) loss function are used. The initial learning rate is 0.0005, the batch size is 64, the number of training epochs is set to 210, the input size of the network is 256×192, and the model is implemented on the PyTorch framework. According to the optimized loss function, the model can accurately detect the key points of human postures, especially with better robustness and accuracy in dense scenes;
[0020] 4) Apply the trained FE-HRNet model to the human pose estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature person and 13 key points.
[0021] This technical solution is applied to the recognition of students' meta-actions in offline smart classrooms and has the following characteristics:
[0022] It provides a human pose estimation model based on feature-enhanced high resolution, aiming to solve the problem of student pose estimation in dense smart classroom scenarios, specifically manifested in:
[0023] 1) Innovatively embed the hierarchical residual connection structure into the initial stage of multi-scale feature extraction of HRNet. By designing multiple more fine-grained feature branches, the receptive field range of the network is expanded, the network's ability to understand the relationship of different-scale features in the image is enhanced, and the overall performance of the model is improved;
[0024] 2) Introduce the multi-scale convolutional attention module into HRNet, effectively enhancing the network's multi-scale feature encoding ability. This module can aggregate multi-scale context information and establish long-range spatial dependencies for human pose estimation, thereby improving the positioning accuracy of key point regions;
[0025] 3) FE-HRNet has good generalization ability and robustness in different scenarios. Especially the high-precision performance achieved on the SCP dataset for smart classrooms provides a solid foundation for the application of this model in actual teaching scenarios.
[0026] This method has higher accuracy than existing algorithms in a densely populated environment, and the model parameters and floating-point operation requirements of this method are superior to existing keypoint detection algorithms. This method shows stronger detection ability and higher positioning accuracy when dealing with complex postures. Description of the Drawings
[0027] Figure 1 It is the block diagram of the FE-HRNet structure in the embodiment;
[0028] Figure 2 It is the block diagram of the hierarchical residual structure in the embodiment;
[0029] Figure 3 It is the block diagram of the multi-scale convolutional attention structure in the embodiment;
[0030] Figure 4 It is the MAP convergence diagram of each model under the MS COCO 2017 dataset in the embodiment;
[0031] Figure 5 It is the MAP convergence diagram of each model under the Smart Classroom Pose dataset in the embodiment;
[0032] Figure 6 It is the detection result of each model on the student with ID 230***10 in the embodiment;
[0033] Figure 7 It is the detection result of each model on the student with ID 230***13 in the embodiment;
[0034] Figure 8 It is the detection result of each model on the student with ID 230***45 in the embodiment. Detailed Implementation Manner
[0035] The following further elaborates on the content of the present invention in conjunction with the drawings and embodiments, but does not limit the present invention.
[0036] Embodiment:
[0037] Refer to Figure 1 , a human pose keypoint recognition method based on feature-enhanced high resolution, includes the following steps:
[0038] 1) Identify the body posture: Based on the images extracted according to the teaching duration of knowledge points in the smart classroom scenario, use the YOLOv11 algorithm to detect and identify all the learner boundary frames in the images. The target bounding box consists of four elements {x min , y min , w, h}, where x min , y minrespectively represent the x - coordinate and y - coordinate of the upper - left corner of the bounding box, w and h respectively represent the width and height of the bounding box. Input the images extracted by the lecture duration of knowledge points in the intelligent classroom scenario into the backbone network FE - HRNet to extract the two - dimensional coordinate positions of 13 key points including the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, and left and right hips, and accurately identify the body posture;
[0039] 2) Enhance the feature representation ability of the learner's posture key points: Input the image features extracted in step 1) into the backbone network FE - HRNet, and adopt the adaptive fusion of multi - scale features to achieve dynamic feature enhancement in the spatial and channel dimensions, including:
[0040] 2 - 1) Construct a hierarchical class - residual connection structure to achieve fine - grained multi - scale feature representation, effectively expanding the network receptive field range. Adopt a multi - scale feature extraction module, which significantly enhances the network's representation ability while maintaining a similar computational complexity. Replace the original 3×3 convolution with a set of smaller convolution kernels, and connect different - scale convolution groups in a hierarchical residual connection manner. In each hierarchical class - residual connection structure, the input feature map first undergoes a convolution operation with a 1×1 convolution kernel, and is equally divided into s feature subsets, denoted as x i , where i ∈ {1, 2,... s}, and each feature subset x i has the same spatial dimension as the input feature map. Except for x1, each x i has a corresponding 3x3 convolution operation, denoted as K i , and use y i to represent the output of K i (). Add the feature subset x i to the output of K i-1 (), and then feed it into K i (). In order to reduce parameters while increasing s, the convolution operation with a 3×3 convolution kernel for x1 is omitted, and it is written as shown in formula (1):
[0041]
[0042] Through the above design, each 3×3 convolution K i () in the hierarchical class - residual connection structure can receive feature information from all feature partitions {x j , j ≤ i}. Each time the feature partition x j passes through a 3×3 convolution operator, the receptive field of its output feature will be correspondingly expanded. As the features interact and fuse repeatedly between different branches, the receptive field of the convolution kernel is also continuously expanding. By introducing the size control parameter s, s = 4, as Figure 2As shown, this structure can flexibly adjust the granularity of feature segmentation, enabling the network to perform feature learning and extraction within a wider receptive field range while maintaining low computational complexity and storage overhead;
[0043] 2-2) Construct a multi-scale convolutional attention module, which captures spatial context information at different scales through multi-branch depthwise separable convolutions and combines a channel attention mechanism to adaptively enhance key features, significantly improving the localization ability of human keypoints. First is a 3×3 depthwise separable convolution layer for aggregating local information of the input features. Second is a multi-branch depthwise strip convolution structure, where each branch uses two orthogonal 1D depthwise separable convolutions to decompose the standard convolution, namely (1×7, 7×1), (1×11, 11×1), and (1×21, 21×1). This decomposition strategy not only significantly reduces the computational complexity but also efficiently captures multi-scale context information through approximate 2D convolution operations of 7×7, 11×11, and 21×21. Finally, a 1×1 convolution layer is used to model the dependencies between different channels, and the output is directly used as the attention weight to perform element-wise multiplication with the input features, achieving dynamic feature enhancement in both spatial and channel dimensions through the adaptive fusion of multi-scale features, which can effectively handle object and detail feature expressions at different scales. The formal representation is shown in Formulas (2) and (3) as follows:
[0044]
[0045] where Att and Out are the attention map and the output respectively, Scale i , i ∈ {1, 2, 3} represents the i-th branch, DW-Conv represents depth convolution, is the multiplication operation of element matrices, and F represents the input feature;
[0046] (2-3) Incorporate the hierarchical residual structure and multi-scale convolution into the backbone network HRNet. FE-HRNet constructs four parallel feature branches. The resolutions of the four parallel branches from top to bottom are sampled to 1, 1 / 2, 1 / 4, and 1 / 8 of the input image resolution in sequence, and the corresponding number of feature channels is expanded to 32, 64, 128, and 256 in sequence. In the feature extraction process, the resolution of the feature map is first reduced to 1 / 4 of the input image resolution through two 3×3 convolutional layers with a stride of 1, and then feature extraction is carried out through four stages: In the first stage, the hierarchical residual connection structure of Res2Net is introduced, and fine-grained feature extraction is achieved through 4 improved residual units; the second, third, and fourth stages respectively contain 1, 4, and 3 multi-resolution modules. Each module consists of two parts: parallel multi-resolution convolution and multi-resolution fusion, and contains 4 residual units. Among them, each unit contains 2 3×3 convolutions in different resolution branches. At the end of each stage, a multi-scale convolutional attention module is connected to extract more detailed category and spatial context information. Finally, multi-scale feature fusion is used to generate the keypoint heatmap, and the heatmap is shown in formula (4):
[0047]
[0048] where f t-3 , f t-2 , f t-1 , f t represent the corresponding outputs of the network at the 4th stage with resolutions of 1 / 8, 1 / 4, 1 / 2, and 1 of the original input size respectively, and Conv represents the convolution operation;
[0049] (3) Training: Use 1 NVIDIA GeForce RTX 4090 graphics processing unit for training. Use the adaptive distance estimation optimizer, namely the Adam optimizer, and the mean square error MSE loss function. The initial learning rate is 0.0005, the batch size is 64, the number of training epochs is set to 210, the input size of the network is 256×192, and the model is implemented on the pytorch framework. According to the optimized loss function, the model can accurately detect the human body pose keypoints, especially with better robustness and accuracy in dense scenes;
[0050] 4) Apply the trained FE-HRNet model to the human pose estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature person and 13 key points. To verify the performance of the method in this example, it was compared with body pose estimation methods such as ResNet50, ResNet101, HRNet, ECA-HRNe, and NSA-HRNet. These methods were trained on the COCO2017 dataset. To ensure fairness in the comparison, the performance results presented in the experiment were all trained from scratch. Among them, Figure 4 For the comparison of the MAP convergence curves during the training process, Table 1 shows the experimental performance results of each model on the MS COCO 2017 dataset:
[0051] Table 1 Metrics of Each Model on the MS COCO 2017 Dataset
[0052]
[0053] , Figure 4 shows the comparison of the MAP convergence curves of different models during the training process. The horizontal axis represents the number of training epochs, and the vertical axis represents the MAP value. It can be observed from the figure that in the initial stage of training (before about 180 epochs), the convergence speed of HRNet is relatively slow, and the MAP value is lower than that of the ResNet50 baseline model. This is mainly because the parallel multi-branch architecture adopted by HRNet requires a longer parameter optimization stage. As the training progresses, HRNet gradually demonstrates its powerful feature expression ability and finally achieves better performance than ResNet-50 under the condition of similar parameter quantities. NSA-HRNet (orange curve) improves the convergence speed by introducing a sparse attention mechanism, but the method in this example exhibits more superior training characteristics: on the one hand, as Figure 3 shown, benefiting from the hierarchical residual structure of the Res2Net module and the multi-scale feature enhancement mechanism of the MSCA module, this model shows a faster convergence rate; on the other hand, throughout the training process, FE-HRNet always maintains a leading MAP level, and the final performance is significantly better than all the comparison methods. This result not only verifies the effectiveness of the proposed architecture but also indicates its better optimization efficiency and training stability. From the data comparison and quantitative analysis in Table 1, the following conclusions can be drawn:
[0054] (1) Compared with 75.92% of the baseline HRNet, the method in this example fully inherits the excellent network architecture of HRNet and further improves the performance through the multi-scale feature enhancement mechanism. In terms of the number of parameters, FE-HRNet is similar to HRNet and lower than ResNet-50 and ResNet-101, reflecting the high efficiency of the model design;
[0055] (2) The method in this example adopts a scratch training strategy and introduces a multi-scale feature enhancement mechanism. The model performance reaches 78.02% MAP on the MAP metric, significantly exceeding ResNet-50 and ResNet-101 (improving by 2.36% and 1.75% respectively), and at the same time significantly leading the baseline HRNet and ECA-HRNet (improving by 2.10% and 1.85% respectively). NSA-HRNet has made a breakthrough in dealing with long-range feature dependencies and achieved a MAP performance of 77.25%. However, in comparison, the method in this example still has a further improvement of 0.77% in accuracy and slightly lower computational complexity, demonstrating a better balance of efficiency and performance;
[0056] (3) The method in this example shows comprehensive advantages in feature representation ability and localization accuracy. Compared with ResNet-50 with a similar parameter scale, the method in this example improves by 2.36% and 1.81% respectively on the AP and AR metrics. Especially in terms of the key point recall rate, the method in this example reaches a level of 80.40%, significantly superior to all comparison methods, fully demonstrating the technical advancement of the model in the human key point detection task;
[0057] Evaluate the method in this example on the Smart Classroom Pose dataset. The performance of the model algorithm on the validation set is shown in Table 2, and the MAP convergence curve of the model is as Figure 5 shown:
[0058] Table 2 Metrics of each model on the Smart Classroom Pose dataset
[0059]
[0060] , Figure 5Shows the training convergence curves of different models on the SCP dataset. In the initial stage of training (rounds 0 - 25), all models showed a rapid upward trend, demonstrating strong learning ability. The method in this example (brown curve) showed a faster convergence speed and a higher performance level during this stage. In the middle stage of training (rounds 25 - 100), the performance growth of each model gradually slowed down and entered a stable improvement stage. The performance of NSA - HRNet (orange curve) and ECA - HRNet (purple curve) was similar and slightly higher than that of ResNet50 (blue curve) and HRNet (green curve). The method in this example continued to maintain a leading advantage with less fluctuation, showing better training stability. In the later stage of training (rounds 100 - 200), all models basically reached the convergence state and the performance tended to be stable. Among them, the method in this example finally converged to the highest mAP level. Although the other models reached relatively high accuracies, the gap was obvious. This fully verified the superiority of the method in this example in feature extraction and expression ability. This convergence curve demonstrated the advantage of the method in this example in terms of final performance, reflecting its better training efficiency and stability;
[0061] Table 2 shows the performance comparison results of each model on the SCP dataset. On this dataset, the method in this example achieved an average precision (AP) of 98.00%, which was 0.55% and 0.17% higher than ResNet - 50 and ResNet - 101 respectively. Compared with HRNet (97.74%) and ECA - HRNet (97.75%) with similar parameter scales, the method in this example achieved performance gains of 0.26% and 0.25% respectively, verifying the effectiveness of the proposed method in the smart classroom scenario. In terms of the AP50 index, the method in this example reached an excellent performance of 99.45%, exceeding all comparison methods, indicating its obvious advantage in large - scale key point localization. NSA - HRNet reached a high level of 98.96% in the AR index, but was still inferior to the method in this example in terms of comprehensive accuracy and computational complexity. Combining the experimental results of the COCO and SCP datasets, FE - HRNet had good generalization ability and robustness in different scenarios. Especially the high - precision performance achieved on the SCP dataset for smart classrooms provided a solid foundation for the application of this model in actual teaching scenarios. To verify the effectiveness of the method in this example in actual scenarios, the research selected three students (IDs are 230***10, 230***13, 230***45) in the smart classroom scenario as test samples, and used ResNet50, ResNet101, HRNet, ECA - HRNet, NSA - HRNet, and the model of the method in this example to conduct human pose key point detection experiments on them respectively, and analyzed and evaluated the performance of each model in the human key point detection task, where Figure 6 、 Figure 7 、 Figure 8The detection results of the above six models for three students are respectively shown. (a)-(e) correspond to the detection outputs of the above six models in sequence. Tables 3-5 show the corresponding Figure 6 , Figure 7 , Figure 8 keypoint detection situations. Among them, "√" indicates the successfully detected keypoints, and "—" indicates the undetected keypoints:
[0062] Table 3 Detection results of each model for students with different IDs
[0063]
[0064] . It can be observed from Figure 6 that obvious keypoint detection errors occur in the elbow area for ResNet50 and ResNet101. These errors are marked by red circles. HRNet also has positioning deviations at the upper limb joints, while ECA-HRNet fails to achieve effective detection on this sample. In contrast, NSAHRNet and FE-HRNet present the most complete detection results and high-precision keypoint localization. Figure 7 shows the detection results of ID 13. The performance of each model on this sample is relatively stable. Among them, the detection results of the method in this example and HRNet are particularly prominent, and the keypoints are more comprehensively covered. However, HRNet still has detection errors in the arm area (shown by red circles). For the sample with ID 45 with complex postures, in the complex posture situation under occlusion conditions, the detection effects of the ResNet series models decrease significantly, while the HRNet series shows stronger robustness. In particular, the method in this example maintains a high detection quality.
[0065] From the data analysis in Table 3, in the keypoint detection of the student with ID 10, the method in this example successfully detected 11 keypoints, significantly better than the 6-point and 7-point detection results of ResNet50 and ResNet101. In the keypoint detection of the student with ID 13, the performance of FE-HRNet in this example is further improved to 13 keypoints. The performance of other models is relatively stable but the detection accuracy is low. For the sample with ID 45 with complex postures, the detection performance of most models shows a significant decreasing trend. The number of detected points of ResNet50 and ResNet101 decreases to 5 points and 3 points respectively, while the method FE-HRNet in this example still maintains a high-precision detection result of 13 keypoints, fully demonstrating its excellent feature recognition stability and robustness under complex conditions.
[0066] The experimental results show that the method FE-HRNet in this example demonstrates significant advantages in key point detection performance in the smart classroom scenario. This method can not only accurately locate the key points of learners, but also maintain stable detection effects in the scenario of dense human distribution. In the scenario of dense students, the method in this example effectively reduces the probability of incorrect allocation of students' key points through an efficient feature extraction and multi-scale information fusion mechanism, and successfully realizes the distinction of different target key points. Even in the complex scenario where the body parts of students overlap, the model still maintains a high positioning accuracy, fully demonstrating the advantages of the method FE-HRNet in this example in terms of feature expression. The experimental results fully verify the superiority of the method FE-HRNet model in the human pose estimation task in the classroom scenario, especially showing stronger detection ability and higher positioning accuracy when dealing with complex poses. Among them, on the MS COCO 2017 dataset, GFLOPs and Params are 12.97G and 30.76M respectively, and the MAP of FE-HRNet is 78.02%. On the Smart Classroom Pose dataset, the MAP of FE-HRNet is 98.00%. The inventive method shows relatively excellent indicators in different environments.
Claims
1. A method for identifying human body pose key points based on feature enhancement and high resolution, characterized in that, It includes the following steps: 1) Identify body postures: Based on the images extracted according to the teaching duration of knowledge points in the intelligent classroom scenario, use the YOLOv11 algorithm to detect and identify all the learner boundary frames in the images. The target bounding box consists of four elements {x min , y min , w, h}, where x min , y min represent the x coordinate and y coordinate of the upper left corner of the bounding box respectively, and w and h represent the width and height of the bounding box. Input the images extracted according to the teaching duration of knowledge points in the intelligent classroom scenario into the backbone network FE-HRNet to extract the two-dimensional coordinate positions of 13 key points including the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, and left and right hips, and accurately identify the body postures; 2) Enhance the feature representation ability of the learner's pose key points: Input the image features extracted in step 1) into the backbone network FE-HRNet, and adopt the adaptive fusion of multi-scale features to achieve dynamic feature enhancement in the spatial and channel dimensions, including: 2-1) Construct a hierarchical class residual connection structure to achieve fine-grained multi-scale feature representation. Replace the original 3×3 convolution with a set of smaller convolution kernels, and connect different-scale convolution groups in a hierarchical residual connection manner. In each hierarchical class residual connection structure, the input feature map first undergoes a convolution operation with a 1×1 convolution kernel, and is equally divided into s feature subsets, denoted as x i , where i ∈ {1, 2,... s}, and each feature subset x i has the same spatial dimension as the input feature map. Except for x1, each x i has a corresponding 3x3 convolution operation, denoted as K i , and the output of K i is represented by y i . The feature subset x i is added to the output of K i-1 , and then fed into K i to reduce the parameters while increasing s. Omit the 3×3 convolution operation of x1, and write it as shown in formula (1): Each 3×3 convolution K in the hierarchical class residual connection structure i () can receive feature information from all feature partitions {x j , j ≤ i}, and each time the feature partition x j passes through a 3×3 convolution operator, the receptive field of its output feature will expand accordingly. As the features interact and fuse repeatedly between different branches, the receptive field of the convolution kernel is also continuously expanding. A size control parameter s is introduced, where s = 4; 2-2) Construct a multi-scale convolutional attention module. First is a 3×3 depthwise separable convolutional layer, which is used to aggregate the local information of the input features. Secondly is a multi-branch depthwise strip convolutional structure. Each branch respectively uses two orthogonal one-dimensional depthwise separable convolutions to decompose the standard convolution, which are (1×7, 7×1), (1×11, 11×1), and (1×21, 21×1). At the same time, through approximate 7×7, 11×11, and 21×21 two-dimensional convolution operations, efficient capture of multi-scale context information is achieved. Finally, a 1×1 convolutional layer is used to model the dependencies between different channels, and the output is directly used as the attention weight to perform element-wise multiplication with the input features. The formal representation is shown in formulas (2) and (3): where Att and Out are the attention map and the output respectively, Scale i , i ∈ {1, 2, 3} represents the i-th branch, DW-Conv represents depthwise convolution, is the multiplication operation of element matrices, and F represents the input feature; 2-3) Embed the hierarchical residual structure and multi-scale convolution into the backbone network HRNet. FE-HRNet constructs four parallel feature branches. The resolutions of the four parallel branches from top to bottom are sampled to 1, 1 / 2, 1 / 4, and 1 / 8 of the input image resolution in sequence, and the corresponding feature channel numbers are expanded to 32, 64, 128, and 256 in sequence. The feature extraction process first reduces the feature map resolution to 1 / 4 of the input image resolution through two 3×3 convolutional layers with a stride of 1, and then performs feature extraction through four stages: In the first stage, the hierarchical residual connection structure of Res2Net is introduced, and fine-grained feature extraction is achieved through 4 improved residual units; the second, third, and fourth stages respectively contain 1, 4, and 3 multi-resolution modules. Each module consists of two parts: parallel multi-resolution convolution and multi-resolution fusion, and contains 4 residual units. Among them, each unit contains 2 3×3 convolutions in different resolution branches. At the end of each stage, a multi-scale convolutional attention module is connected to extract more detailed category and spatial context information. Finally, multi-scale feature fusion is used to generate the key point heat map, and the heat map is shown in formula (4): where f t-3 , f t-2 , f t-1 , f t respectively represent the corresponding outputs of the original input sizes with resolutions of 1 / 8, 1 / 4, 1 / 2, and 1 in the fourth stage of the network, and Conv is the convolution operation; 3) Training: Use 1 NVIDIA GeForce RTX 4090 graphics processing unit for training. Use the adaptive distance estimation optimizer, that is, the Adam optimizer and the mean square error MSE loss function. The initial learning rate is 0.0005, the batch size is 64, the number of training epochs is set to 210, the input size of the network is 256×192, and the model is implemented on the pytorch framework. According to the optimized loss function, make the model accurately detect the human body pose key points; 4) Apply the trained FE-HRNet model to the human body pose estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature person and 13 key points.
Citation Information
Cited By
Multi-modal forged video detection method based on multi-head addition cross attention mechanism
CN120635786A
Unmanned aerial vehicle image target detection network based on double-branch attention
CN120876839A
Human body posture intelligent recognition system based on artificial intelligence
CN121033945A
An AI-based intelligent human posture recognition system
CN121033945B
Industrial part surface defect detection method and system with robustness
CN121258875A