Sparse attention feature enhancement-based shielded human body posture key point identification method
By adopting the sparse attention feature enhancement method in the recognition of key points of human posture, combined with PANet and NS-CBAM modules, the problem of low accuracy in occlusion key points recognition in smart classroom scenes is solved, achieving higher robustness and accuracy.
Patent Information
- Application Number
- CN202510283447.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing methods of identifying key points in human postures are less accurate when dealing with occlusion scenarios, especially in smart classroom scenarios. Due to dense seating and multimodal interference, traditional methods are difficult to effectively identify occlusion key points.
Using a method based on sparse attention feature enhancement, local target box and key point information are captured in the intermediate features, and feature representation is updated. Combined with the PANet neck network and NS-CBAM sparse cycle dual-channel attention mechanism, the robustness and accuracy of the model in occlusion scenarios are improved.
It effectively solves the problem of insufficient detection accuracy of human key points under complex occlusion, improves the robustness and accuracy of the model in practical applications, and realizes accurate identification of key points of body posture.
Smart Images

Figure CN120220183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image recognition technology, especially the technology of human body pose key point recognition in image recognition. Specifically, it is a method for occluded human body pose key point recognition based on sparse attention feature enhancement. Background Art
[0002] Human body pose estimation is a technology that infers postures and analyzes motion states based on the spatial structure relationship between human body key points in images or videos, and has wide applications in fields such as motion capture, motion analysis, augmented reality, and human-computer interaction. In the field of education, human body pose estimation technology plays an important role. It can quantitatively evaluate classroom participation and characterize the activity of learning behaviors, which is crucial for accurately identifying the body postures of students.
[0003] Currently, in the smart classroom scenario, students are usually in a dense classroom environment, which makes the occlusion problem an urgent challenge to be solved. On the one hand, due to the severe occlusion caused by dense seating, the mutual occlusion and overlapping phenomena among students will significantly reduce the accuracy of traditional human body pose estimation methods. On the other hand, the multi-modal interference caused by the dynamic changes of teaching equipment (such as the light and shadow changes of projectors, the complex background formed by foldable desks and chairs) further exacerbates the feature space aliasing of existing AI models in occluded classroom scenarios.
[0004] Existing methods for human body pose key point recognition still have many deficiencies when dealing with occlusion scenarios. Relevant researchers usually regard the detection of regular key points and occluded key points as a unified problem and solve it through a single network. However, due to the discreteness between occlusion features and key point features, a very small amount of occlusion data is difficult to guide the network to converge in the occlusion direction, resulting in poor performance of the network when dealing with occlusion problems. To solve problems such as incomplete high-level features of occluded key points, there is an urgent need for an occluded human body pose key point recognition method that can not only remove obstacle interference features but also use multi-dimensional features matching the true characteristics of key points to enhance the expression ability of key point features and construct a complete and high-precision high-order key point feature representation. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for occluded human body pose key point recognition based on sparse attention feature enhancement in view of the deficiencies of the prior art. This method captures local target box and key point information in intermediate features, and through updating the feature representation, it can effectively solve the problem of insufficient accuracy of human body key point detection in complex occlusion situations, improve the robustness and accuracy of the model in practical applications, and thus achieve accurate recognition of body pose key points.
[0006] The technical solution for achieving the purpose of the present invention is as follows:
[0007] A method for identifying occluded human pose key points based on sparse attention feature enhancement, comprising the following steps:
[0008] 1) Input the images extracted according to the teaching duration of knowledge points in the smart classroom scenario into the backbone network CSPDarknet53-SPPF, and extract multi-level key point features including the two-dimensional coordinate positions of 17 key points, the Euclidean distances between adjacent key points of the left and right eyes and the left and right shoulders, and the relative included angles between the shoulder-elbow and wrist, and hip-knee and ankle key points. The 17 key points are the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankle positions, as well as the starting coordinates, width, height, and category of the learner detection box;
[0009] 2) Enhance the global representation ability of the learner's pose key point features: Input the extracted features into the PANet neck network, and use the multi-scale feature fusion optimization model to adapt to the occlusion scenario, and further enhance the global representation ability of the learner's pose key point features, including:
[0010] 2-1) PANet, as a fusion of multi-scale features, associates the anchor box Anchor with the pose key points. One anchor box Anchor matches one target. Each anchor box Anchor includes the human body bounding box and 2D pose information. Each detection head contains two decoupled heads for bounding box localization and key point regression respectively. The target bounding box is determined by {C x , C y , W, H, b conf , c conf} 6 elements, and each key point is determined by {x, y, c} 3 elements. Among them, x, y, and c respectively represent the key point position and class confidence. For each anchor box Anchor, 51 elements related to 17 key points of the human body and 6 elements of the bounding box are associated to generate the required predicted key points [Pv] as shown in formula (1):
[0011]
[0012] In the formula, C x , C y are the horizontal and vertical coordinates of the center point of the bounding box respectively, W and H are the width and height of the bounding box respectively, b conf , c conf are the bounding box confidence and the predicted class confidence respectively, and K is the target key point feature constant;
[0013] 3) Construct a shared and separated dual-channel ShareSepHead detection head: The input of the separated dual-channel ShareSepHead detection head is the feature maps from different levels, and the output is the learner target and key point feature S out , including:
[0014] 3-1) The shared separation dual-channel embedding formula is as shown in formula (2):
[0015]
[0016] In formula (2), S b1 and S b2 represent the response inputs of the occluded features;
[0017] 3-2) ShareSepHead uses a shared feature extraction layer to process multi-task learning. Detection and pose achieve task adaption through gradient gating. A deformable convolution (Deformable Conv) is introduced in the pose branch to enhance the response of the occluded region F def The response of the occluded region F def The embedding formula is as shown in formula (3):
[0018]
[0019] In formula (3), F def represents the feature map after deformable convolution processing, w k represents the weight of the k-th sampling point, p represents the center position of the convolution kernel, and Δp k represents the offset predicted by the shallow features;
[0020] 4) Construct a sparse dual-loop attention mechanism NS-CBAM module: The input of the NS-CBAM module is the representation of the occluded region response, and the output is the response weight of the key point region. Specifically:
[0021] 4-1) NS-CBAM adopts two parts: sparse loop dual-channel attention and spatial attention, which cooperate to optimize the expression of feature key points. The channel attention module uses global average pooling and max pooling to generate a channel attention map to enhance the information of the key point detection channels. The spatial attention module uses a sparse dynamic hierarchical convolution kernel and dynamic weight parameters to highlight the responses of all key point regions in the image. The output of NS-CBAM is as shown in formula (4):
[0022] NS-CBAM Output = spatial att·(M c (F)) (4),
[0023] In formula (4), M c (F) represents the channel attention weight, spatial att represents the spatial attention weight, and x is the input feature map;
[0024] 4-2) The channel attention module uses global average pooling and global max pooling to generate a channel attention map to enhance the information of the key point detection channels. The output of the channel attention module Mc (F) As shown in formula (5):
[0025] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (5),
[0026] In formula (5), F is the input feature map, AvgPool and MaxPool respectively represent the global average pooling and max pooling operations, MLP represents the multi-layer perceptron, and σ represents the Sigmoid activation function;
[0027] 4-3) The spatial attention module uses convolutional operations to highlight the response of the key point regions in the image. The output M s (F) As shown in formula (6):
[0028] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (6),
[0029] In formula (6), f 7×7 represents a 7×7 convolutional operation, and [AvgPool(F); MaxPool(F)] represents concatenating the average pooling and max pooling results along the channel axis;
[0030] 4-4) The sparse dynamic hierarchical convolutional kernel in NS-CBAM is a key component of the spatial attention module. By alternately using convolutional kernels of different scales, multi-scale feature extraction and fusion of the feature map are achieved. This innovative design enhances the model's ability to capture features of different scales, and at the same time, through dynamic weight parameters further optimizes the occluded key point feature representation,
[0031] The output of the sparse dynamic hierarchical convolutional kernel is shown in formula (7):
[0032]
[0033] In formula (7), i represents the index of the convolutional layer, Conv2d represents the two-dimensional convolutional operation, and padding represents the padding size;
[0034] The hierarchical attention fusion output is shown in formula (8):
[0035] att i = σ(ns_conv i (combined))·α i (8),
[0036] In formula (8), σ represents the Sigmoid activation function, combined is the concatenation result of the average value and the maximum value of the input feature map, and α i represents the dynamic weight parameter;
[0037] The final spatial attention weight output is shown in formula (9):
[0038]
[0039] In formula (9), N is the total number of convolutional layers, and M s (F) represents the output weight of the spatial attention module;
[0040] 5) Design the loss function Aol (Anti-occlusion loss, abbreviated as Aol). The input of the loss function Aol is the key point region response weight output in step 4), and the output is the loss function of the dynamic key points, including:
[0041] 5-1) The output of the loss function Aol is shown in formula (10):
[0042]
[0043] In formula (10), λ is the balance factor, and σ i represents the normalization factor, IoU represents an index to measure the overlapping degree between the predicted region and the real region, and L XIoU represents the weight of the XIoU loss;
[0044] 5-2) The calculation of the loss function Aol is based on the improved OKS′, and the calculation of OKS′ is shown in formula (11):
[0045]
[0046] In formula (11), represents the key point weight, d i represents the Euclidean distance between the real position and the predicted position of the i-th key point, s represents the human body scale factor, and v i represents the visibility flag;
[0047] 6) Construct the YOLOXSP model: The YOLOXSP model architecture consists of a CSPDarknet53-SPPF backbone network, a PANet neck network, an NS-CBAM sparse cyclic dual-channel attention mechanism, and a new type of ShareSepHead detection head connected in sequence;
[0048] 7) The YOLOXSP model is trained using the loss function Aol. The training process is as follows: collect the human body pose image dataset in the smart classroom scenario and label the key points of the human body pose, ensuring that the dataset includes samples of occlusion scenarios. Configure the training environment, with the number of iteration rounds set to 300, Batch size set to 16, the input image size to 388×288, enable automatic mixed-precision training, and set the learning rate of the backbone network to 0.01. According to the optimized loss function, the model can accurately detect the key points of the human body pose, especially with better robustness and accuracy in occlusion scenarios.
[0049] 8) Apply the trained YOLOXSP model to the human body pose estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature (person) and 17 key points.
[0050] This technical solution is applied to the identification of key points of the body postures of students in a dense and occluded smart classroom, and has the following characteristics:
[0051] 1) The innovative occluded human body pose key point recognition network architecture YOLOXSP, optimized specifically for student pose recognition in the smart classroom scenario, effectively addresses the problem of low detection and recognition rates caused by crowded students or insignificant differences in the front and back backgrounds.
[0052] 2) Adopt the design of the ShareSepHead detection head, combined with the new NS-CBAM sparse cyclic dual-channel attention mechanism, which precisely improves the positioning accuracy of occluded key points. This design successfully solves the problems of missing and difficult recognition of learner pose key points in complex backgrounds and occlusion environments, and enhances the robustness of the model in practical applications.
[0053] 3) This technical solution shows high effectiveness and reliability in identifying key points of body postures in a dense and occluded smart classroom scenario, providing strong support for improving classroom participation assessment, learning state analysis, and teaching strategy optimization.
[0054] This technical solution is optimized specifically for the smart classroom scenario, effectively addressing the problem of low recognition rate caused by student congestion or background differences; through the design of the ShareSepHead detection head, it fully exploits the channel and spatial feature information, enhancing the model's ability to recognize occluded human body pose key points; in the dense and occluded scenario of the smart classroom, it has high effectiveness and reliability in recognizing the key points of body pose, providing a new technical means for the fields of smart classroom and pose estimation. This method captures local target boxes and key point information in intermediate features, and through updating the feature representation, it can effectively solve the problem of insufficient accuracy of human key point detection in complex occlusion situations, improve the robustness and accuracy of the model in practical applications, and thus achieve accurate recognition of the key points of body pose. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic flowchart of the method of the embodiment;
[0056] Figure 2 It is a schematic diagram of the YOLOXSP model framework in the embodiment;
[0057] Figure 3 It is a schematic diagram of the structure of the shared separation dual-channel ShareSepHead detection head in the embodiment;
[0058] Figure 4 It is a schematic diagram of the structure of the NS-CBAM sparse dual attention mechanism in the embodiment;
[0059] Figure 5 It is a schematic diagram of multi-person dense pose estimation in the smart classroom scenario with background in the embodiment;
[0060] Figure 6 It is a schematic diagram of multi-person dense pose estimation in the smart classroom scenario without background in the embodiment;
[0061] Figure 7 It is a schematic diagram of the comparative experiment of YOLOXSP with advanced models on the Ochuman dataset in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The following further elaborates on the content of the present invention in conjunction with the drawings and embodiments, but does not limit the present invention.
[0063] Embodiment:
[0064] Referring to Figure 1 、 Figure 2 A method for recognizing occluded human body pose key points based on sparse attention feature enhancement includes the following steps:
[0065] 1) Input the images extracted by the teaching duration according to knowledge points in the intelligent classroom scenario into the backbone network CSPDarknet53 - SPPF, and extract multi - level key - point features including the two - dimensional coordinate positions of 17 key points, the Euclidean distances between adjacent key points of the left and right eyes and the left and right shoulders, and the relative included angles between the key points of the shoulder - elbow and wrist, and hip - knee and ankle. The 17 key points are the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, as well as the starting coordinates, width, height, and category of the learner detection box;
[0066] 2) Strengthen the global representation ability of the learner's pose key - point features: Input the extracted features into the PANet neck network, and use the multi - scale feature fusion to optimize the model's adaptability to the occlusion scenario, including:
[0067] 2 - 1) PANet, as a network that fuses multi - scale features, associates the anchor box Anchor with the pose key points. One anchor box Anchor matches one target. Each anchor box Anchor includes the human body bounding box and 2D pose information. Each detection head contains two decoupled heads respectively for bounding box localization and key - point regression. The target bounding box is determined by 6 elements {C x ,C y ,W,H,b conf ,c conf}, and each key point is determined by 3 elements {x,y,c}. Among them, x, y, and c respectively represent the key - point position and class confidence. For each anchor box Anchor, 51 elements related to 17 key points of the human body and 6 elements of the bounding box are associated to generate the required predicted key points [Pv] as shown in formula (1):
[0068]
[0069] In the formula, C x ,C y are respectively the horizontal and vertical coordinates of the center point of the bounding box, W and H are respectively the width and height of the bounding box, b conf ,c conf are respectively the bounding - box confidence and the predicted - class confidence, and K is the target key - point feature constant;
[0070] 3) Construct a shared - separation dual - channel ShareSepHead detection head: As shown in Figure 3 , the input of the separation dual - channel ShareSepHead detection head is the feature maps from different levels, and the output is the learner target and key - point feature S out , including:
[0071] 3 - 1) The shared - separation dual - channel embedding formula is as shown in formula (2):
[0072]
[0073] In formula (2), S b1 and S b2 represent the response inputs of the occlusion features;
[0074] 3-2) ShareSepHead uses a shared feature extraction layer to process multi-task learning. Detection and pose achieve task adaptation through gradient gating, and deformable convolution Deformable Conv is introduced in the pose branch to enhance the response of the occlusion area F def , and the response of the occlusion area F def is embedded as shown in formula (3):
[0075]
[0076] In formula (3), F def represents the feature map after deformable convolution processing, w k represents the weight of the k-th sampling point, p represents the center position of the convolution kernel, and Δp k represents that the offset is predicted by the shallow feature;
[0077] 4) Construct a sparse double-loop attention mechanism NS-CBAM module: As Figure 4 shown, the input of the NS-CBAM module is the response representation of the occlusion area, and the output is the response weight of the key point area. Specifically:
[0078] 4-1) NS-CBAM uses sparse loop dual-channel attention and spatial attention. The two parts cooperate to optimize the expression of feature key points. The channel attention module uses global average pooling and max pooling to generate a channel attention map to enhance the information of the key point detection channel. The spatial attention module uses a sparse dynamic hierarchical convolution kernel and dynamic weight parameters to highlight the responses of all key point areas in the image. The output of NS-CBAM is shown in formula (4):
[0079] NS-CBAM Output = spatial att·(M c (F)) (4),
[0080] In formula (4), M c (F) represents the channel attention weight, spatial att represents the spatial attention weight, and x is the input feature map;
[0081] 4-2) The channel attention module uses global average pooling and global max pooling to generate a channel attention map to enhance the information of the key point detection channel. The output of the channel attention module M c (F) is shown in formula (5):
[0082] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (5),
[0083] In formula (5), F is the input feature map, AvgPool and MaxPool respectively represent the global average pooling and max pooling operations, MLP represents the multi-layer perceptron, and σ represents the Sigmoid activation function;
[0084] 4-3) The spatial attention module uses convolutional operations to highlight the responses in the key point regions of the image. The output M s (F) is as shown in formula (6):
[0085] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (6),
[0086] In formula (6), f 7×7 represents a 7×7 convolutional operation, and [AvgPool(F); MaxPool(F)] represents concatenating the average pooling and max pooling results along the channel axis;
[0087] 4-4) The sparse dynamic hierarchical convolutional kernel in NS-CBAM is a key component of the spatial attention module. By alternately using convolutional kernels of different scales, it realizes multi-scale feature extraction and fusion of the feature map. This innovative design enhances the model's ability to capture features of different scales. At the same time, through dynamic weight parameters it further optimizes the occluded key point feature representation,
[0088] The output of the sparse dynamic hierarchical convolutional kernel is as shown in formula (7):
[0089]
[0090] In formula (7), i represents the index of the convolutional layer, Conv2d represents the two-dimensional convolutional operation, and padding represents the padding size;
[0091] The hierarchical attention fusion output is as shown in formula (8):
[0092] att i = σ(ns_conv i (combined)) · α i (8),
[0093] In formula (8), σ represents the Sigmoid activation function, combined is the concatenation result of the average value and the maximum value of the input feature map, and α iRepresents the dynamic weight parameter;
[0094] The final spatial attention weight output is shown in formula (9):
[0095]
[0096] In formula (9), N is the total number of convolutional layers, M s (F) represents the output weight of the spatial attention module;
[0097] 5) Design the loss function Aol (Anti-occlusion loss, abbreviated as Aol). The input of the loss function Aol is the key point region response weight output in step 4), and the output is the loss function of the dynamic key points, including:
[0098] 5-1) The output of the loss function Aol is shown in formula (10):
[0099]
[0100] In formula (10), λ is the balance factor, σ i represents the normalization factor, IoU represents the index to measure the overlap degree between the predicted region and the real region, L XIoU represents the weight of the XIoU loss;
[0101] 5-2) The calculation of the loss function Aol is based on the improved OKS′, and the calculation of OKS′ is shown in formula (11):
[0102]
[0103] In formula (11), represents the key point weight, d i represents the Euclidean distance between the real position and the predicted position of the i-th key point, s represents the human body scale factor, v i represents the visibility flag;
[0104] 6) Construct the YOLOXSP model: The YOLOXSP model architecture consists of a CSPDarknet53-SPPF backbone network, a PANet neck network, an NS-CBAM sparse cyclic dual-channel attention mechanism, and a new type of ShareSepHead detection head connected in sequence;
[0105] 7) The YOLOXSP model is trained using the loss function Aol. The training process in this example is as follows: Collect the human pose image dataset in the smart classroom scenario and annotate the human pose key points, ensuring that the dataset contains samples of occlusion scenarios. Configure the training environment, with the number of iteration rounds set to 300, the Batch size set to 16, the input image size to 388×288, enable automatic mixed precision training, and set the learning rate of the backbone network to 0.01. According to the optimized loss function, the model can accurately detect the human pose key points, especially having better robustness and accuracy in occlusion scenarios.
[0106] Apply the trained YOLOXSP model to the human pose estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature (person) and 17 key points. Among them, Table 1 is the schematic diagram of the comparative experiment of YOLOXSP with advanced models on the Ochuman dataset in this example:
[0107] Table 1 Comparative experiment of YOLOXSP with advanced models on the Ochuman dataset
[0108]
[0109] As Figure 5 and Figure 6 shown, Figure 5 and Figure 6 show the estimation of multi-person dense poses by the YOLOXSP model in the smart classroom scenario with and without background. It can be seen from the figure that in the smart classroom scenario with background, the YOLOXSP model can accurately detect the key points of multiple people. Even in the case of dense people and occlusion, the model can better identify the key point positions of each person, indicating that the model has good robustness and accuracy in complex background environments. In the smart classroom scenario without background, the YOLOXSP model can still accurately detect the key points of multiple people. After removing the background, the model focuses more on the pose estimation of people, and the accuracy of key point detection is further improved, which shows that the model in this example performs better without background interference.
[0110] As Figure 7 shown, Figure 7Shows the comparison experiment results of the YOLOXSP model in this example with other advanced models on the Ochuman dataset. It can be seen from the figure that during the training process of the YOLOXSP model, the mean average precision (mAP) gradually converges and remains at a relatively high level, superior to other models such as MRSA, CID, DHRNet, DEKR, and SPM, etc. This indicates that the YOLOXSP model in this example performs excellently in the multi-person dense pose estimation task in the smart classroom scenario, especially having better robustness and accuracy in occlusion scenarios. Compared with existing advanced models, the YOLOXSP in this example has significant advantages in both the accuracy of key point detection and the generalization ability of the model.
Claims
1. A method for recognizing key points of occluded human body posture based on sparse attention feature enhancement, characterized in that: The steps include: 1) Input the images extracted according to the teaching time of knowledge points in the smart classroom scenario into the backbone network CSPDarknet53-SPPF to extract multi-level key point features including the two-dimensional coordinate positions of 17 key points, the linear Euclidean distance between adjacent key points of the left and right eyes and the left and right shoulders, and the relative angles between the key points of the shoulder, elbow and wrist, hip, knee and ankle. The 17 key points are the positions of the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, as well as the starting coordinates, width, height, and category of the learner detection frame; 2) Strengthen the global representation ability of the learner's posture key point features: The extracted features are input into the PANet neck network, and the multi-scale feature fusion optimization model's adaptability to occluded scenes is used, including: 2-1) PANet is used as a fusion multi-scale feature to associate anchor boxes with posture key points. One anchor box matches one target. Each anchor box includes a human body bounding box and 2D posture information. Each detection head contains two decoupling heads for bounding box positioning and key point regression respectively. The target bounding box is composed of {C x ,C y ,W,H,b conf ,c conf }6 elements, each key point is determined by 3 elements {x, y, c}, where x, y and c represent the key point position and category confidence respectively. For each anchor box Anchor, 51 elements of 17 key points of the human body and 6 elements of the bounding box are associated to generate the required predicted key point [Pv] as shown in formula (1): In the formula, C x , C y are the horizontal and vertical coordinates of the center point of the bounding box, W and H are the width and height of the bounding box, respectively. conf 、c conf They are the bounding box confidence and the predicted category confidence, respectively, and K is the target key point feature constant; 3) Construct a shared separation dual-channel ShareSepHead detection head: The separated dual-channel ShareSepHead detection head inputs feature maps from different levels and outputs learner targets and key point features S out ,include: 3-1) The shared separation dual-channel embedding formula is shown in formula (2): S out =BN(Conv shared (S b1 ))⊕Attn spatial (S b2 ) (2), In formula (2), S b1 and S b2 A response input representing an occlusion feature; 3-2) ShareSepHead uses a shared feature extraction layer to process multi-task learning. Detection and posture are adaptive to each other through gradient gating. Deformable Conv is introduced in the posture branch to enhance the response of the occluded area. def , occlusion area response F def The embedding formula is shown in formula (3): In formula (3), F def represents the feature map after deformed convolution, w k represents the weight of the kth sampling point, p represents the center position of the convolution kernel, Δp k It indicates that the offset is predicted by shallow features; 4) Construct a sparse dual-cycle attention mechanism NS-CBAM module: The input of the NS-CBAM module is the occluded area response representation, and the output is the key point area response weight, specifically: 4-1) NS-CBAM uses sparse cyclic dual-channel attention and spatial attention to collaboratively optimize the expression of feature key points. The channel attention module uses global average pooling and maximum pooling to generate channel attention maps. The spatial attention module uses sparse dynamic hierarchical convolution kernels and dynamic weight parameters to highlight the responses of all key point areas in the image. The output of NS-CBAM is shown in formula (4): NS-CBAM Output =spatial at·(M c (F)) (4), In formula (4), M c (F) represents the channel attention weight, spatial att represents the spatial attention weight, and x is the input feature map; 4-2) The channel attention module uses global average pooling and global maximum pooling to generate the channel attention map. The output M of the channel attention module is c (F) As shown in formula (5): M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (5), In formula (5), F is the input feature map, AvgPool and MaxPool represent global average pooling and maximum pooling operations respectively, MLP represents multi-layer perceptron, and σ represents Sigmoid activation function; 4-3) The spatial attention module uses convolution operation to highlight the response of the key point area in the image. The output M of the spatial attention module s (F) As shown in formula (6): M s (F)=σ(f 7×7 ([AvgPool(F);MaxPool(F)])) (6), In formula (6), f 7×7 Represents a 7×7 convolution operation, [AvgPool(F); MaxPool(F)] means concatenating the average pooling and maximum pooling results along the channel axis; 4-4) The sparse dynamic hierarchical convolution kernel in NS-CBAM is a key component of the spatial attention module. By alternately using convolution kernels of different scales, multi-scale feature extraction and fusion of feature maps are achieved. The output of the sparse dynamic layered convolution kernel is shown in formula (7): In formula (7), i represents the index of the convolution layer, Conv2d represents the two-dimensional convolution operation, and padding represents the padding size; The hierarchical attention fusion output is shown in formula (8): att i =σ(ns_conv i (combined))·α i (8), In formula (8), σ represents the Sigmoid activation function, combined is the concatenation of the average and maximum values of the input feature map, and α i Represents dynamic weight parameters; The final spatial attention weight output is shown in formula (9): In formula (9), N is the total number of convolutional layers, M s (F) represents the output weight of the spatial attention module; 5) Design loss function Aol: The input of loss function Aol is the key point regional response weight output in step 4), and the output is the loss function of dynamic key points, including: 5-1) The output of the loss function Aol is shown in formula (10): In formula (10), λ is the balance factor, σ i represents the normalization factor, IoU represents the index to measure the overlap between the predicted area and the true area, and L XIoU Represents the weight of XIoU loss; 5-2) The calculation of the loss function Aol is based on the improved OKS′. The calculation of OKS′ is shown in formula (11): In formula (11), represents the key point weight, d i represents the Euclidean distance between the true position and the predicted position of the i-th key point, s represents the human scale factor, and v i Indicates visibility flag; 6) Build the YOLOXSP model: The YOLOXSP model architecture consists of a sequentially connected CSPDarknet53-SPPF backbone network, a PANet neck network, a NS-CBAM sparse recurrent dual-channel attention mechanism, and a new ShareSepHead detection head; 7) The loss function Aol is used to train the YOLOXSP model. The training process is as follows: collect a human posture image dataset in the smart classroom scenario, annotate the human posture key points, ensure that the dataset contains occlusion scene samples, configure the training environment, set the number of iterations to 300, the batch size to 16, the input image size to 388×288, enable automatic mixed precision training, and set the backbone network learning rate to 0.01; based on the optimized loss function, the model accurately detects the human posture key points, especially in occlusion scenes, with better robustness and accuracy; 8) Apply the trained YOLOXSP model to the human posture estimation task in the smart classroom scenario, perform key point detection on the input image or video, and output the detection results, including the learner target box category feature person and 17 key points.
Citation Information
Cited By
Navigation logistics real-time dangerous goods identification system and method
CN120564137A
Human body posture intelligent recognition system based on artificial intelligence
CN121033945A
An AI-based intelligent human posture recognition system
CN121033945B