Human body posture feature extraction model, method and system and posture estimation network, method and system
By using the Swiftformer model and the inverse residual module, combined with the additive attention mechanism and coordinate classification method, the problems of high model complexity and insufficient feature extraction capabilities in the existing technology are solved, and efficient pose feature extraction and estimation are achieved.
Patent Information
- Application Number
- CN202510327525.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing human pose estimation method is sensitive to image resolution, the heat map method brings high overhead and quantization error, the coordinate regression method is large in calculation and complex in the model, and the existing model feature extraction ability is insufficient.
The Swiftformer model is used for feature extraction, combined with the inverse residual module and the pose estimation head, and through the additive attention mechanism and coordinate classification method, the model parameter quantity and calculation complexity are reduced, and the feature capture ability is improved.
In the case of reducing the amount of parameters and calculations, the ability to capture complex human posture characteristics is improved, the calculation complexity is reduced, and the accuracy and robustness of posture estimation are improved.
Smart Images

Figure CN120299061A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biometric technologies for identifying human bodies in images or videos, and particularly relates to a human body pose feature extraction model, method, system, and pose estimation network, method, and system. Background Art
[0002] Most of the existing human body pose estimation methods are based on Gaussian heatmaps and coordinate regression. The heatmap method renders each key point in the GroundTruth into a Gaussian heatmap, and uses the maximum value mapping as the pose estimation result. However, the heatmap method is sensitive to image resolution, and at the same time, the heatmap generation mapping brings high overhead and quantization errors. The coordinate regression method directly predicts the coordinate positions of human key points through the model, and realizes human body pose estimation through the coordinate positions of the key points. These two methods have high requirements for the feature extraction ability of human body poses. Currently, commonly used models, such as the baseline method SimpleBaseline-50 and SimpleBaseline-101 models that use heatmaps, and the PRTR method based on coordinate regression, etc., involve complex feature extraction model structures and large amounts of calculations. Therefore, there is an urgent need for a feature extraction technology with a low model complexity and a strong feature extraction ability. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the present invention proposes a human body pose feature extraction model, method, system, and pose estimation network, method, and system, which can improve the model's ability to capture complex human body pose features while introducing fewer parameters and calculations. The specific technical solutions are as follows:
[0004] In the first aspect, a human body pose feature extraction model is provided. In the first implementable manner of the first aspect, it includes:
[0005] A feature extraction module configured to extract features from a human body pose estimation picture using the Swiftformer model to obtain feature maps of different resolution sizes;
[0006] An inverted residual module configured to perform convolution and self-attention operations on each feature map to obtain a one-dimensional Gaussian feature vector corresponding to the human body pose estimation picture.
[0007] Combined with the first implementable manner of the first aspect, in the second implementable manner of the first aspect, the inverted residual module includes:
[0008] An upsampling MLP layer configured to perform upsampling processing on the feature map;
[0009] The self-attention layer is configured to capture local features of the feature map after dimensionality increase processing by using separable convolution, and perform dynamic modeling on the global features of the feature map by using window self-attention;
[0010] The dimensionality reduction MLP layer is configured to perform dimensionality reduction processing on the high-dimensional features captured by the self-attention layer to obtain a one-dimensional Gaussian feature vector corresponding to the human pose estimation image.
[0011] In a second aspect, a human pose estimation network is provided. In a first realizable manner of the second aspect, it includes:
[0012] The human pose feature extraction model as described in the first or second realizable manner of the first aspect, configured to perform feature extraction on the input human pose estimation image to obtain a corresponding one-dimensional Gaussian feature vector;
[0013] The pose estimation head is configured to estimate the human pose in the human pose estimation image according to the one-dimensional Gaussian feature vector.
[0014] Combined with the first realizable manner of the second aspect, in a second realizable manner of the second aspect, the pose estimation head estimating the human pose includes:
[0015] Using the coordinate classification method to perform vertical classification and horizontal classification on the one-dimensional Gaussian feature vector respectively;
[0016] Restoring the vertical classification result and the horizontal classification result to the height and width of the human pose estimation image respectively through the scaling factor to obtain the corresponding human key point positions.
[0017] In a third aspect, a human pose feature extraction method is provided, including:
[0018] Construct the human pose feature extraction model as described in the first or second realizable manner of the first aspect, and train the human pose feature extraction model by using the constructed data set;
[0019] Obtain the human pose image to be extracted, and extract a one-dimensional Gaussian feature vector from the human pose image by using the trained human pose feature extraction model.
[0020] In a fourth aspect, a human pose feature estimation method is provided, including:
[0021] Construct the human pose estimation network as described in the second realizable manner of the second aspect, and train the human pose estimation network by using the constructed data set;
[0022] Obtain the human pose image to be estimated, and estimate the human key point positions in the human pose image by using the trained human pose estimation network.
[0023] In a fifth aspect, a human body posture feature extraction system is provided, which is characterized by including:
[0024] A model training module, configured to construct a human body posture feature extraction model as described in the first or second implementation manner of the first aspect, and train the human body posture feature extraction model using the constructed data set;
[0025] A feature extraction module, configured to obtain a human body posture image to be extracted, and extract a one-dimensional Gaussian feature vector from the human body posture image using the trained human body posture feature extraction model.
[0026] In a sixth aspect, a human body posture feature estimation system is provided, including:
[0027] A network training module, configured to construct a human body posture estimation network as described in the second implementation manner of the second aspect, and train the human body posture estimation network using the constructed data set;
[0028] A posture estimation module, configured to obtain a human body posture image to be estimated, and estimate the positions of human body key points in the human body posture image using the trained human body posture estimation network.
[0029] Advantageous effects: By using the human body posture feature extraction model, method, system, and posture estimation network, method, system of the present invention. The human body posture feature extraction model can utilize the efficient additional attention mechanism of the Swiftformer model to accurately extract a human body posture feature map rich in context information. The Swiftformer model replaces the quadratic matrix multiplication operation with linear element multiplication, reducing the number of model parameters and reducing the algorithm complexity from quadratic to linear with respect to the number of tokens. And without sacrificing any accuracy, it replaces the key-value interaction with a linear layer, which can significantly reduce the computational complexity of the model.
[0030] Through the set inverted residual module, the detailed information of the key point coordinate information in the feature map can be further extracted and restored, and at the same time, the connection between different key points can be captured, further improving the coordinate extraction ability of the feature extraction model. And the inverted residual module is only composed of standard convolution and multi-head self-attention, reducing the complexity of the feature extraction model, and being able to improve the capture ability of the feature extraction model for complex human body posture features with fewer introduced parameters and calculations. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the specific implementation manners of the present invention, the drawings required for use in the specific implementation manners will be briefly introduced below. In all the drawings, the components or parts do not necessarily draw according to the actual scale.
[0032] Figure 1 It is an architecture diagram of a human body posture feature extraction model provided by an embodiment of the present invention;
[0033] Figure 2 It is a framework structure diagram of a human body posture estimation network provided by an embodiment of the present invention;
[0034] Figure 3 It is a flowchart of a human body posture feature extraction method provided by an embodiment of the present invention;
[0035] Figure 4 It is a flowchart of a human body posture feature estimation method provided by an embodiment of the present invention;
[0036] Figure 5 It is a system block diagram of a human body posture feature extraction system provided by an embodiment of the present invention;
[0037] Figure 6 It is a system block diagram of a human body posture feature estimation system provided by an embodiment of the present invention. Specific embodiments
[0038] Hereinafter, embodiments of the technical solution of the present invention will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and therefore are only examples and cannot be used to limit the protection scope of the present invention.
[0039] As Figure 1 shown in the architecture diagram of the human body posture feature extraction model, the extraction model includes:
[0040] A feature extraction module configured to extract features from a human body posture estimation picture using a Swiftformer model to obtain feature maps of different resolution sizes;
[0041] An inverted residual module configured to perform convolution and self-attention operations on each feature map to obtain a one-dimensional Gaussian feature vector corresponding to the human body posture estimation picture.
[0042] Specifically, the extraction model includes a feature extraction module and an inverted residual module. Among them, the feature extraction module uses the latest efficient additive attention model, that is, the Swiftformer model, to extract feature maps of different resolution sizes from the input human body posture estimation picture to obtain a multi-scale feature representation corresponding to the human body posture estimation picture. The inverted residual module (InvertedResidual Block) can further extract and restore the details of the key point coordinate information in the feature maps extracted by the feature extraction module, and at the same time capture the connections between different key points, so as to obtain a one-dimensional Gaussian feature vector corresponding to the human body posture estimation picture.
[0043] Compared with the traditional Transformer model, the Swiftformer model adopts an additive attention mechanism, effectively replacing the quadratic matrix multiplication operation with linear element multiplication, reducing the number of model parameters and reducing the algorithm complexity from quadratic to linear with respect to the number of tokens. At the same time, replacing the key-value interaction with a linear layer significantly reduces the computational complexity of the model without sacrificing any accuracy. Moreover, the efficient attention mechanism design can be used in all stages of the network to achieve more effective global information capture.
[0044] Specifically, the additive attention mechanism effectively utilizes the pairwise interaction between tokens through element multiplication, rather than using dot product operations to capture global context information. The traditional attention mechanism encodes the relevance scores of the context information of the input sequence based on the interaction of three attention components: Q (query), K (key), and V (value). The additive attention mechanism deletes the key-value interaction without degrading performance, and a linear projection layer alone is sufficient to satisfy the relationship between query and key.
[0045] The specific operations include: multiplying the query matrix by the learnable parameter vector π and performing a softmax operation to obtain the query attention weights. Then, based on the learned attention weights, it is merged with the query matrix to obtain a single global query vector. This process can be expressed by the following formula:
[0046]
[0047] where γ is the query attention weight, q is the global query vector, Q is the query matrix, n is the feature length, and d is the scaling factor.
[0048] On this basis, the global query vector is used to interact with the key matrix K element by element to capture context information, and this method is more computationally straightforward in the MHSA (multi-head attention mechanism). The output of the additive attention is specifically:
[0049] X = normalized(Q) + L(K * q);
[0050] where L represents linear transformation processing and normalized represents normalization processing.
[0051] The following will combine Figure 1 to elaborate on the specific operation steps of the Swiftformer model in detail.
[0052] In this embodiment, the Swiftformer model includes a patch embedding layer, and the output of the patch embedding layer is connected to four hierarchical steps of different scales in cascade. All the hierarchical steps are the same and are composed of a convolutional encoder (Conv Encoder) and a Swiftformer encoder (Swiftformer Encoder).
[0053] Specifically, after the human pose estimation picture of size H×W×C is input into the feature extraction module, it will first pass through a patch embedding layer with a 3×3 convolution of stride 2 to obtain a feature map of size, and the specific calculation formula of the patch embedding layer is as follows:
[0054] F1 = Conv 3×3,stride=2 (2).
[0055] The feature map output by the patch embedding layer is sent to the first-level hierarchical step. The spatial features of the feature map are extracted through the convolutional encoder, and the global information and the relationship between the local and the local in the spatial features output by the convolutional encoder are captured through the Swiftformer encoder using the additive attention mechanism. The specific calculation formula is as follows:
[0056] X1 = F1 + Conv 1×1 (Conv 1×1 (Dwconv 3×3 (F1)));
[0057] The specific calculation formulas corresponding to the other hierarchical steps are as follows:
[0058] X i = F i + Conv 1×1 (Conv 1×1 (Dwconv 3×3 (F i )));
[0059] where i = 2, 3, 4.
[0060] Adjacent hierarchical steps are connected by a downsampling layer, so as to reduce the spatial size by 2 times and increase the feature dimension. The specific calculation formula of the downsampling layer output is as follows:
[0061] F i = Downsample(X i-1 ).
[0062] The feature expression output by the entire feature extraction module is:
[0063]
[0064] Since the visibility of human body key points is not only affected by clothing, environment, perspective, etc., but also faces problems such as occlusion, intersection, and dislocation. Moreover, the human pose estimation method for coordinate classification also puts forward higher requirements for the quality of the feature map.
[0065] To this end, a reverse residual module is creatively introduced in the extraction model claimed in the present invention. The reverse residual module can further extract and recover the detailed information of the key point coordinates in the feature map, and at the same time capture the connections between different key points, further improving the coordinate extraction ability of the feature extraction model. And in this embodiment, the reverse residual module is only composed of standard convolution and multi-head self-attention, reducing the complexity of the feature extraction model, and being able to improve the capture ability of the feature extraction model for complex human pose features with less parameter quantity and calculation amount.
[0066] In this embodiment, optionally, the reverse residual module includes:
[0067] A dimensionality increasing MLP layer, configured to perform dimensionality increasing processing on the feature map;
[0068] A self-attention layer, configured to capture local features of the feature map after dimensionality increasing processing by using separable convolution, and perform dynamic modeling on the global features of the feature map by using window self-attention;
[0069] A dimensionality decreasing MLP layer, configured to perform dimensionality decreasing processing on the high-dimensional features captured by the self-attention layer to obtain a one-dimensional Gaussian feature vector corresponding to the human pose estimation picture.
[0070] Specifically, the reverse residual module includes a dimensionality increasing MLP layer, a self-attention layer, and a dimensionality decreasing MLP layer. Among them, the feature map output by the feature extraction module is first subjected to dimensionality increasing operation through the dimensionality increasing MLP layer. Then, the self-attention layer uses separable convolution to capture local features of the feature map after dimensionality increasing processing. After fusing the local features with the feature map, the window self-attention mechanism (EW-MHSA) is used to perform dynamic modeling on the global features of the feature map. Finally, the dimensionality decreasing MLP layer performs dimensionality decreasing processing on the extracted global features to obtain a one-dimensional Gaussian feature vector corresponding to the human pose estimation picture. The specific calculation process is as follows:
[0071]
[0072] ξ(·)=(X μ +DWCONV 3×3 (X μ ))(EW-MHSA);
[0073] X′ μ =MLP(ξ(·)).
[0074] Such asFigure 2 The framework structure diagram of the human pose estimation network shown, the estimation network includes:
[0075] The human pose feature extraction model as described above, configured to extract features from the input human pose estimation image to obtain a corresponding one-dimensional Gaussian feature vector;
[0076] The pose estimation head, configured to estimate the human pose in the human pose estimation image according to the one-dimensional Gaussian feature vector.
[0077] Specifically, the estimation network includes the above-mentioned human pose feature extraction model and the pose estimation head. Among them, through the human pose feature extraction model, the capture ability of the estimation network for complex human pose features can be improved while significantly reducing the number of parameters and the amount of computation, providing a basis for the subsequent pose estimation head to accurately estimate the human pose. The pose estimation head can accurately estimate the human pose according to the one-dimensional Gaussian feature vector extracted by the human pose feature extraction model, achieving a good balance between the network computational complexity and accuracy.
[0078] In this embodiment, optionally, the pose estimation head estimating the human pose includes:
[0079] Using the coordinate classification method to perform vertical classification and horizontal classification on the one-dimensional Gaussian feature vector respectively;
[0080] Restoring the vertical classification result and the horizontal classification result to the height and width of the human pose estimation image respectively through the scaling factor to obtain the corresponding human key point positions.
[0081] Specifically, when representing the key point positions as a two-dimensional Gaussian distribution based on the Gaussian heat map method, a large amount of error will be introduced, especially in the case of occlusion and complex poses. At the same time, there is redundant post-processing when restoring the heat map. The coordinate regression method has poor scalability.
[0082] Therefore, the pose estimation head can use the coordinate classification method to decompose the coordinate positioning problem into a classification task for one-dimensional vector representation. Specifically, the pose estimation head first performs decoupling operations on the one-dimensional Gaussian feature vector extracted by the human pose feature extraction model in the horizontal and vertical directions, and uses the coordinate classification results in the horizontal and vertical directions obtained by the classifier, which are respectively and Then, according to the set scaling factor, the vertical classification result and the horizontal classification result are restored to the horizontal coordinate and the vertical coordinate of the human pose estimation input image respectively to obtain the corresponding human key point positions. In this way, the intermediate supervision of the high-resolution heat map is avoided, and a large number of bins are used to reduce the quantization error to the sub-pixel level.
[0083] As Figure 3 shown in the flowchart of the human body pose feature extraction method, the extraction method includes:
[0084] Step 1: Construct the above-mentioned human body pose feature extraction model, and use the constructed dataset to train the human body pose feature extraction model;
[0085] Step 2: Obtain the human body pose image to be extracted, and use the trained human body pose feature extraction model to extract a one-dimensional Gaussian feature vector from the human body pose image.
[0086] Specifically, first, the above-mentioned human body pose feature extraction model can be constructed, and the constructed human body pose feature extraction model can be trained through a pre-constructed dataset to obtain a trained human body pose feature extraction model. Then, the human body pose image to be estimated can be obtained, and the human body pose image can be input into the trained human body pose feature extraction model. Through the human body pose feature extraction model, the one-dimensional Gaussian feature vector corresponding to the human body in the image can be obtained, providing a basis for the subsequent pose estimation head to accurately estimate the human body pose.
[0087] As Figure 4 shown in the flowchart of the human body pose feature estimation method, the estimation method includes:
[0088] Step D1: Construct the above-mentioned human body pose estimation network, and use the constructed dataset to train the human body pose estimation network;
[0089] Step D2: Obtain the human body pose image to be estimated, and use the trained human body pose estimation network to estimate the positions of the human body key points in the human body pose image.
[0090] Specifically, first, the above-mentioned human body pose estimation network can be constructed, and the constructed human body pose estimation network can be trained through a pre-constructed dataset to obtain a trained human body pose estimation network. Then, the human body pose image to be estimated can be obtained, and the human body pose image can be input into the trained human body pose estimation network. The one-dimensional Gaussian feature vector in the human body pose image is captured by the human body pose feature extraction model in the human body pose estimation network, and then the one-dimensional Gaussian feature vector is vertically classified and horizontally classified by the pose estimation head, and restored to the height and width of the human body pose estimation image according to the scaling factor to obtain the corresponding positions of the human body key points.
[0091] As Figure 5 shown in the system block diagram of the human body pose feature extraction system, the extraction system includes:
[0092] A model training module, configured to construct the above-mentioned human pose feature extraction model and train the human pose feature extraction model using the constructed data set;
[0093] A feature extraction module, configured to obtain a human pose image to be extracted and extract a one-dimensional Gaussian feature vector from the human pose image using the trained human pose feature extraction model.
[0094] Specifically, the extraction system includes a model training module and a feature extraction module. Among them, the model training module can construct the above-mentioned human pose feature extraction model and train the constructed human pose feature extraction model using a pre-constructed data set to obtain a trained human pose feature extraction model. The feature extraction module can obtain a human pose image to be estimated and input the human pose image into the trained human pose feature extraction model. Through the human pose feature extraction model in the human pose feature extraction model, a one-dimensional Gaussian feature vector corresponding to the human body in the image can be obtained, providing a basis for the subsequent pose estimation head to accurately estimate the human pose.
[0095] As Figure 6 shown in the system block diagram of the human pose feature estimation system, the estimation system includes:
[0096] A network training module, configured to construct the above-mentioned human pose estimation network and train the human pose estimation network using the constructed data set;
[0097] A pose estimation module, configured to obtain a human pose image to be estimated and estimate the positions of human key points in the human pose image using the trained human pose estimation network.
[0098] Specifically, the estimation system includes a network training module and a pose estimation module. Among them, the network training module can construct the above-mentioned human pose estimation network and train the constructed human pose estimation network using a pre-constructed data set to obtain a trained human pose estimation network. The pose estimation module can obtain a human pose image to be estimated and input the human pose image into the trained human pose estimation network. Through the human pose feature extraction model in the human pose estimation network, a one-dimensional Gaussian feature vector in the human pose image is captured, and then the one-dimensional Gaussian feature vector is vertically classified and horizontally classified by the pose estimation head, and restored to the height and width of the human pose estimation image according to the scaling factor to obtain the corresponding positions of human key points.
[0099] To verify the pose estimation effect of the present invention, compared with the baseline methods using heatmaps such as SimpleBaseline-50, SimpleBaseline-101, and SimpleBaseline-152, the present invention has an improvement of 6.5%, 5%, and 4.1% in AP respectively, while reducing the model GFlops (computational complexity) by 34.8%, 41.1%, and 67.5%. It can be seen that the model structure of the present invention is simple without complex component designs, significantly reducing the computational complexity. Among them, AP is the abbreviation of Average Precision, which is a common evaluation index for pose estimation performance.
[0100] Compared with HRNet-W32 of the multi-scale high-resolution heatmap fusion CNN method, the GFlops (computational complexity) is reduced by 29%, and at the same time, various indicators have certain improvements. Compared with its upgraded version HRNet-W48, the present invention has achieved competitive results with its 55% #Params (number of parameters) and 34.9% GFLOPs (computational complexity).
[0101] Compared with the above HRNet series models, the present invention uses a Transformer network with a full-stage additive attention mechanism to significantly reduce the high computational overhead of the HRNet method due to maintaining high-resolution feature maps and fusion. Moreover, the present invention adopts a method based on coordinate classification, without generating two-dimensional Gaussian heatmaps, avoiding redundant post-processing operations and quantization errors, improving the sensitivity to key points while significantly reducing the computational complexity. The model of the present invention is simpler than the heatmap method and has a higher recognition accuracy.
[0102] In addition, the model of the present invention was uploaded to the Codalab official website, and the experimental results of MS COCO test-dev were returned, further confirming the performance of the present invention. Compared with the integral regression method Integral Pose Regression, the present invention has an improvement of 9.4% in AP. Compared with the coordinate regression-based PRTR method, an improvement of 2.5 AP is achieved. Since the solution space of the coordinate classification method adopted by the present invention is smaller than that of coordinate regression, the model and loss function design are more easily adapted to the model. Compared with the heatmap-based CPN and RMPE methods, the present invention has an improvement of about 2.0 AP with its 20% GFLOPs (computational complexity).
[0103] Compared with the Transformer-based pose estimation methods TFPose and TokenPose-B, the present invention has achieved improvements of 2.0 and 0.2 AP respectively. At the same time, it leads the TokenPose-B method by about 0.8 AP in the AP50 and AP75 metrics, proving the effective adaptability of the human pose feature extraction model of the present invention to the coordinate feature extraction in the overall human pose estimation task. Compared with the latest knowledge distillation method DistilPose-L, the present invention has improved by 0.5 AP and 1.2 APM, reflecting the strong robustness of the present invention to the key points that are prone to change and occlusion. Among them, AP50 and AP75 are based on the key point similarity, taking different thresholds to measure whether the detection of each human body is accurate.
[0104] Experiments were also carried out on the multi-person pose estimation dataset MPII and compared with advanced methods in recent years. Compared with the model CPM, an improvement of 2.9% PCK@0.5 was achieved. Compared with the heatmap-based methods SimpleBaseline-50, SimpleBaseline-101, and SimpleBaseline-152, improvements of 2.1%, 1.3%, and 1.5% PCK@0.5 were achieved respectively while the number of parameters and the amount of calculation were significantly reduced. Among them, PCKh@0.5 is an evaluation metric for the MPII multi-person pose estimation dataset, representing the percentage of the correct key point metric relative to the head size.
[0105] Compared with the HRNet-W32 method that maintains high-resolution heatmaps to improve the capture of detailed features, there is an improvement of about 0.7 in the evaluation metrics of complex key points in the lower body such as Hip, Knee, and Ankle. This shows that the human pose feature extraction model of the present invention has shown good results in capturing the relationship of key points of the whole body context, and at the same time reflects the effectiveness of the combination of CNN and Transformer in the inverted residual module of the present invention in restoring edge detail information.
[0106] Compared with the TokenPose-L / D24 method that uses a knowledge distillation Transformer network, an improvement of about 0.8 was achieved in the metrics of complex key points such as Hip and Ankle while ensuring the accuracy of key points in the upper body such as Head and Shoulder. Compared with the latest coordinate classification method RTMPose-m model, an improvement of about 1.4 was also achieved in the PCK@0.5 metric, which also reflects the rationality of the network structure design of the present invention and the good balance between model complexity and accuracy.
[0107] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered within the scope of the claims and the description of the present invention.
Claims
1. A human body posture feature extraction model, characterized in that including: a feature extraction module configured to extract features from a human pose estimation image using a Swiftformer model to obtain feature maps of different resolution sizes; an inverted residual module configured to perform convolution and self-attention operations on each feature map to obtain a one-dimensional Gaussian feature vector corresponding to the human pose estimation image.
2. The human body posture feature extraction model according to claim 1, wherein The inverted residual module includes: an upsampling MLP layer configured to perform upsampling processing on the feature map; a self-attention layer configured to capture local features of the feature map after upsampling processing using separable convolution and perform dynamic modeling on the global features of the feature map using window self-attention; a downsampling MLP layer configured to perform downsampling processing on the high-dimensional features captured by the self-attention layer to obtain a one-dimensional Gaussian feature vector corresponding to the human pose estimation image.
3. A human body pose estimation network, characterized in that, including: a human pose feature extraction model as claimed in claim 1 or 2, configured to extract features from an input human pose estimation image to obtain a corresponding one-dimensional Gaussian feature vector; a pose estimation head configured to estimate the human pose in the human pose estimation image according to the one-dimensional Gaussian feature vector.
4. The human body pose estimation network according to claim 3, wherein The pose estimation head estimating the human pose includes: performing vertical classification and horizontal classification on the one-dimensional Gaussian feature vector respectively using a coordinate classification method; restoring the vertical classification result and the horizontal classification result to the height and width of the human pose estimation image respectively through a scaling factor to obtain corresponding human key point positions.
5. A method for extracting human body posture features, characterized in that, including: constructing a human pose feature extraction model as claimed in claim 1 or 2 and training the human pose feature extraction model using the constructed dataset; obtaining a human pose image to be extracted and extracting a one-dimensional Gaussian feature vector from the human pose image using the trained human pose feature extraction model.
6. A method for estimating human body posture features, characterized in that, including: constructing a human pose estimation network as claimed in claim 4 and training the human pose estimation network using the constructed dataset; obtaining a human pose image to be estimated and estimating the human key point positions in the human pose image using the trained human pose estimation network.
7. A human body posture feature extraction system, characterized in that including: a model training module configured to construct a human pose feature extraction model as claimed in claim 1 or 2 and train the human pose feature extraction model using the constructed dataset; a feature extraction module configured to obtain a human pose image to be extracted and extract a one-dimensional Gaussian feature vector from the human pose image using the trained human pose feature extraction model.
8. A human body posture feature estimation system, characterized in that, including: a network training module configured to construct a human pose estimation network as claimed in claim 4 and train the human pose estimation network using the constructed dataset; a pose estimation module configured to obtain a human pose image to be estimated and estimate the human key point positions in the human pose image using the trained human pose estimation network.
Citation Information
Patent Citations
Image processing method and device, computer equipment and storage medium
CN116977814A
Hand posture estimation method and system and electronic equipment
CN117690188A
Attitude estimation device and method
CN118447528A
Realization method of remote sensing image road extraction task based on Swift-SegEdgeNet
CN118736434A
Face detection method, face recognition method, behavior recognition method and system
CN118942137A