Light-weight and high-robustness 2D pedestrian attitude estimation method and system and storage medium
Through the HRMamba-EMA model, the feature extraction and fusion is performed using VSSBlock and EMA attention mechanisms, which solves the problems of high computational complexity and feature loss in complex scenarios, and realizes efficient pedestrian pose estimation, which is suitable for vehicle-mounted scenarios.
Patent Information
- Application Number
- CN202510428397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-25
AI Technical Summary
When handling complex scenarios, existing HRNet networks cannot fully capture global context information, resulting in high computational complexity, large memory consumption, and easy to cause details and local features to be lost during feature fusion.
The HRMamba-EMA model is adopted, and the branches with different resolutions are extracted features using VSSBlock, and a multi-scale fusion module based on EMA is introduced. Global information modeling is performed through the state space model (SSM), and feature fusion is performed in combination with the EMA attention mechanism.
It significantly reduces the amount of model parameters, improves the computing efficiency and feature extraction capabilities, enhances the integration ability of multi-scale information, improves the accuracy and computing efficiency of pedestrian pose estimation, and is suitable for on-board scenarios with resource-constrained.
Smart Images

Figure CN120375008A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology. Specifically, it relates to a lightweight and highly robust 2D human pose estimation method, system, and storage medium. Background Art
[0002] Human pose estimation is one of the core tasks in the field of computer vision, aiming to accurately locate human key points from images or videos and construct a human skeletal structure to describe the pose. With the rapid development of artificial intelligence technology, this task has shown important value in scenarios such as action recognition, human-computer interaction, virtual reality, motion analysis, and medical rehabilitation.
[0003] The top-down human pose estimation method is a two-stage method that combines human detection and key point localization. First, this method uses an object detection algorithm to identify and frame the positions of all people in the input image. Subsequently, for each detected human bounding box, single-person pose estimation is performed to predict the positions of the key points of the human body within the box. Since human detection is performed first and then the pose of each human body is estimated separately, it can handle occlusion problems more finely and avoid key point confusion in multi-person scenarios. In addition, this method has good adaptability to human bodies of different scales, can handle complex backgrounds and multi-person scenarios, and has high accuracy and robustness.
[0004] The HRNet network is widely used as a highly representative top-down human pose estimation method. It is a neural network architecture that focuses on high-resolution feature representation. Its core advantage lies in achieving multi-scale information fusion while maintaining high-resolution features by parallelly connecting high-resolution and low-resolution convolutional streams. This design not only avoids the loss of spatial information caused by resolution reduction but also enhances the network's sensitivity to position, thus significantly improving the accuracy of key point localization. However, the HRNet network relies on the capabilities of CNN. CNN expands the receptive field layer by layer through convolution, captures local features, and models long-range dependencies by stacking multiple convolutional layers. It has great limitations in remote modeling and performs poorly in capturing long-range dependency relationships. Especially when dealing with complex scenarios, it may not be able to fully capture global context information, thereby affecting the performance of the model. This not only increases the computational complexity but also leads to a sharp increase in the number of parameters. In addition, the multi-scale fusion part of the network is prone to losing details and local features after multiple upsamplings, further restricting its performance in the human pose estimation task. On the other hand, the network needs to process feature maps of multiple resolutions and perform multiple fusions between resolutions, resulting in a large amount of computation and memory consumption. Summary of the Invention
[0005] The technical problems to be solved by the present invention are:
[0006] The existing HRNet cannot fully capture global context information when dealing with complex scenarios, and it is prone to losing details and local features in feature fusion, with high computational complexity and large memory consumption.
[0007] The technical solution adopted by the present invention to solve the above technical problems:
[0008] The present invention provides a 2D pedestrian pose estimation method, including the following steps:
[0009] Step 1: Collect images through an in-vehicle camera and detect pedestrian targets in the images.
[0010] Step 2: Construct an HRMamba-EMA model based on the HRNet network. Each branch of the model is constructed with a VSSBlock to process branches with different resolutions to extract features, and a multi-scale fusion module based on EMA is introduced for multi-scale feature fusion.
[0011] Step 3: Perform pose estimation and key point detection on the obtained pedestrian target detection results based on the HRMamba-EMA model.
[0012] Further, in step 1, the detection of pedestrian targets in the images is specifically based on the YOLO object detection algorithm to detect pedestrian targets in the images.
[0013] Further, the functional implementation process of the HRMamba-EMA model in step 2 includes:
[0014] In the first stage, two 3×3 convolutional layers are used to downsample the input image, reducing the resolution to 1 / 4 of the original, and at the same time increasing the number of channels to 64; subsequently, four Bottleneck modules are used to extract features and increase the number of channels to 256.
[0015] In the second and fifth stages, each branch of the model contains four VSSBlock modules to process branches with different resolutions to extract features; in the fifth stage, a multi-scale fusion module based on EMA is added, and an EMA attention mechanism is introduced before each upsampling to fuse features of different scales. Finally, a heatmap of each key point of the pedestrian is output.
[0016] Further, the functional implementation process of the VSS block is as follows: After layer normalization, the input is split into two branches. In the first branch, the input first passes through a linear layer and then through the activation function Silu. In the second branch, the input is first processed through a linear layer, a depthwise separable convolution, and the activation function Silu, and then input into the SS2D module for further feature extraction. Subsequently, the features are normalized through layer normalization. Finally, the two paths are merged, the features are merged using a linear layer, and this result is combined with the residual connection to obtain an enhanced feature tensor.
[0017] Further, the functional implementation process of the SS2D module is as follows:
[0018] Scanning expansion: First, the input image is divided into 16 small images of equal size of 4×4, and then the small images are scanned one by one in the order from left to right, from top to bottom, from bottom to top, and from right to left.
[0019] Selective scanning of the S6 block: The S6 block is of the Mamba structure, with the state space model SSM as the main body, relying on a continuous system. The input sequence x(t) ∈ R is mapped to the output y(t) ∈ R through the intermediate implicit state h(t) ∈ R N by a linear time-invariant system, which is represented as a linear ordinary differential equation:
[0020] h′(t) = Ah(t) + Bx(t)
[0021] y(t) = Ch(t)
[0022] where A is the state transition matrix, A ∈ R N×N , B is the input projection matrix for mapping the input to the hidden state space, B ∈ R N×1 , C is the output projection matrix for calculating the final output, C ∈ R D×1 , and N is the state size;
[0023] Introduce a time scale parameter Δ, and use the discretization rule to convert A and B into continuous parameters and The zero-order hold is used as the discretization rule, specifically:
[0024]
[0025] After discretization, based on the SSM model, calculate through linear recursion:
[0026]
[0027] y(t) = Ch(t)
[0028] Or calculate through global convolution:
[0029]
[0030] Among them, is a convolution kernel, L is the sequence length, and y is the final output;
[0031] Scanning and merging: Re - combine the small images processed by block S6 in order to form an image with the same size as the input.
[0032] Furthermore, the functional implementation process of the EMA - based multi - scale fusion module is as follows:
[0033] Feature maps of different scales are respectively subjected to feature extraction by the EMA module and gradually up - sampled. At the same time, after the feature fusion of different scales, an EMA module is continuously added to further enhance the feature expression ability.
[0034] Furthermore, the functional implementation process of the EMA module is as follows:
[0035] First, divide the input feature map into G sub - features in the channel dimension. Subsequently, use three parallel paths to extract the attention weights of the grouped feature maps. Two of the paths are 1×1 branches, and the other path is a 3×3 branch. In the 1×1 branches, the channels are encoded along two spatial directions respectively, and in the 3×3 branch, a single 3×3 convolution is stacked to capture multi - scale feature representations;
[0036] The output of the 1x1 branch encodes global spatial information through two - dimensional global average pooling, and the output of the 3x3 branch is directly converted into the corresponding dimensional shape; then, the output is aggregated through matrix dot - product operation to generate the first spatial attention map; finally, the output feature maps within each group are aggregated through the Sigmoid function of the two generated spatial attention weight values to capture pixel - level pairing relationships.
[0037] The present invention also provides a 2D pedestrian pose estimation method system, which has program modules corresponding to the steps of the method described in any one of the above technical solutions, and executes the steps in the above - mentioned 2D pedestrian pose estimation method when running.
[0038] The present invention also provides a computer - readable storage medium. The computer - readable storage medium stores a computer program, and the computer program is configured to implement the steps in the 2D pedestrian pose estimation method described in any one of the above technical solutions when called by a processor.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] The method of the present invention is based on HRNet. With VSSBlock as the network backbone and adopting the Mamba structure, it uses the state space model (SSM) to directly perform global information modeling through the shared state update matrices A, B, and C, avoiding the redundancy of stacking multiple layers of convolutions. It can capture global features in a single layer and efficiently handle long-range dependencies. The computational complexity of Mamba is O(N), far lower than O(N2) of CNN. At the same time, due to the sharing of the state update matrix, the redundant parameters are greatly reduced, thereby improving the computational efficiency and feature extraction ability, and having significant advantages in global modeling and computational efficiency. In the multi-scale fusion part, the EMA attention mechanism is introduced to effectively enhance the feature expression ability, fuse features of different scales, and enhance the integration ability of multi-scale information.
[0041] Based on the human pose estimation model, the present invention significantly reduces the number of parameters of the model, speeds up the inference speed, can process more data or perform faster real-time inference under the same hardware conditions, reduces the storage requirements of the model, and makes it more suitable for resource-constrained vehicle-mounted scenarios. At the same time, the method improves the accuracy of pedestrian pose estimation of the model, making it more competitive in actual pedestrian pose estimation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the HRMamba-EMA network structure in an embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of the Bottleneck residual block in an embodiment of the present invention;
[0044] Figure 3 It is a schematic diagram of the VSSblock structure in an embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of the SS2D module structure in an embodiment of the present invention;
[0046] Figure 5 It is a schematic diagram of the EMA module structure in an embodiment of the present invention;
[0047] Figure 6 It is a schematic diagram of the multi-scale fusion module based on EMA in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are only a part of the embodiments or examples of the present invention, rather than all of them. All other embodiments or examples obtained by those of ordinary skill in the art based on the embodiments or examples in the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0049] To make the above objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings.
[0050] Specific Embodiment 1: As shown in Figures 1 to 6 The present invention provides a 2D pedestrian pose estimation method, including the following steps:
[0051] Step 1: Collect images through an in-vehicle camera and detect pedestrian targets in the images;
[0052] Step 2: Construct an HRMamba-EMA model based on the HRNet network. Each branch of the model is constructed with a VSSBlock to process branches with different resolutions to extract features, and a multi-scale fusion module based on EMA is introduced for multi-scale feature fusion;
[0053] Step 3: Perform pose estimation and key point detection on the obtained pedestrian target detection results based on the HRMamba-EMA model.
[0054] Specific Embodiment 2: The detection of pedestrian targets in the images in Step 1 is specifically to detect pedestrian targets in the images based on the YOLO object detection algorithm. Other parts of this embodiment are the same as those in Specific Embodiment 1.
[0055] Specific Embodiment 3: The HRMamba-EMA model in Step 2 is based on HRNet, combined with the Mamba and EMA attention mechanisms. As shown in Figure 1 The process of its function implementation is as follows:
[0056] In the first stage, two 3×3 convolutional layers (stride = 2) are used to downsample the input image, reducing the resolution to 1 / 4 of the original, and at the same time increasing the number of channels to 64; subsequently, four Bottleneck modules as shown in Figure 2 are used to extract features, and the number of channels is increased to 256;
[0057] In the second and fifth stages, VSSBlock is used to replace Bassicblock. Each branch of the model contains four VSSBlock modules, which are used to process branches with different resolutions to extract features. In the fifth stage, a multi-scale fusion module based on EMA is added to solve the problem of information loss during the upsampling and downsampling processes. The EMA attention mechanism is adopted before each upsampling operation, which can effectively fuse features of different scales, thereby enhancing the integration ability of multi-scale information.
[0058] Finally, the heatmap of each key point of the pedestrian is output. Other parts of this implementation scheme are the same as those of the second specific implementation scheme or the third specific implementation scheme.
[0059] Specific implementation scheme four: The functional implementation process of the VSSblock is as follows: As Figure 3 shown, the input is split into two branches after layer normalization. In the first branch, the input first passes through a linear layer and then through the activation function Silu. In the second branch, the input is first processed through a linear layer, a depthwise separable convolution, and the activation function Silu, and then input into the SS2D (2D selective scanning) module for further feature extraction. Subsequently, the features are normalized through layer normalization. Finally, the two paths are merged, the features are merged using a linear layer, and this result is combined with the residual connection to obtain the output result. Other parts of this implementation scheme are the same as those of the third specific implementation scheme.
[0060] Specific implementation scheme five: The functional implementation process of the SS2D module is as follows: As Figure 4 shown, it includes: scan expansion, selective scanning through the S6 block, and scan merging:
[0061] Scan expansion: First, the input image is divided into 16 small images of equal size in a 4×4 manner, and then the small images are scanned one by one in the order from left to right, from top to bottom, from bottom to top, and from right to left; to ensure that each element in the feature map integrates information from all other positions in different directions, thereby generating a global receptive field without increasing the linear computational complexity.
[0062] S6 block selective scanning: The S6 block is of the Mamba structure, with the state space model SSM as the main body, depending on the continuous system. The input sequence x(t) ∈ R is mapped to the output y(t) ∈ R through the intermediate implicit state h(t) ∈ R N by a linear time-invariant system, which is represented as a linear ordinary differential equation:
[0063] h′(t) = Ah(t) + Bx(t) (1)
[0064] y(t) = Ch(t)
[0065] where A is the state transition matrix that determines how the hidden state evolves, and A ∈ R N×N ; B is the input projection matrix used to map the input into the hidden state space, and B ∈ R N×1 , C is the output projection matrix used to calculate the final output, and C ∈ R D×1 , and N is the state size;
[0066] The SSM discretizes this continuous system to make it more suitable for deep learning scenarios. The SSM introduces a time scale parameter Δ and uses a fixed discretization rule to convert A and B into continuous parameters and The zero-order hold (ZOH) is used as the discretization rule, specifically:
[0067]
[0068] After discretization, the model based on the SSM is calculated through linear recursion:
[0069]
[0070] or through global convolution:
[0071]
[0072] where is a convolution kernel representing the impulse response of the entire system; L is the sequence length; and y is the final output.
[0073] The Mamba model is based on the SSM. When calculating the B and C matrices, a selection mechanism is added, and an additional linear layer is introduced to select the input control quantity and state quantity, strengthening the model's adaptability to different inputs and solving the problem that the output of the traditional SSM completely depends on the ordered data input. Once the data increases or decreases or the order changes, the SSM cannot handle it.
[0074] Scan merging: The scan merging part is the opposite of the scan expansion part. It recombines the small images processed by the S6 block into a large image with the same size as the input in order. Other aspects of this implementation are the same as those of the fourth specific implementation.
[0075] Specific implementation six: As Figure 6 shown, the functional implementation process of the EMA-based multi-scale fusion module is as follows:
[0076] The features at different scales are respectively subjected to feature extraction by the EMA module. In order to better retain and transmit information, reduce information loss, and improve the feature alignment effect, an upsampling method is adopted step by step. At the same time, after the features at different scales are fused, the EMA module is added continuously to further enhance the feature expression ability. Other parts of this implementation scheme are the same as those of the fifth specific implementation scheme.
[0077] Specific implementation scheme seven: The functional implementation process of the EMA module includes three parts: feature grouping, parallel sub-networks, and cross-space learning, as Figure 5 shown, specifically:
[0078] First, the input feature map is divided into G sub-features in the channel dimension. Subsequently, three parallel routes are used to extract the attention weights of the grouped feature maps. Two of the routes are 1×1 branches, and the other route is a 3×3 branch. In order to obtain the dependency relationships of all channels and reduce the computational amount, the channels are encoded along two spatial directions in the 1×1 branches respectively, and a single 3×3 convolution is stacked in the 3×3 branch to capture multi-scale feature representations; in this way, the EMA module can not only encode cross-channel information to adjust the importance of different channels, but also retain accurate spatial structure information in the channels.
[0079] In the cross-space learning part, the EMA module enriches feature aggregation by providing a cross-space information aggregation method in different spatial dimension directions. Specifically, the output of the 1x1 branch encodes global spatial information through two-dimensional global average pooling, and the output of the 3x3 branch is directly converted into the corresponding dimension shape; then, the outputs are aggregated through matrix dot product operations to generate the first spatial attention map; finally, the output feature maps within each group are aggregated through the Sigmoid function of the two generated spatial attention weight values to capture pixel-level pairing relationships and highlight the global context of all pixels. The final output of the EMA module has the same size as the input and can be effectively stacked into modern architectures.
[0080] The EMA module of this implementation scheme has strong feature extraction and feature expression capabilities, enabling the model to pay more attention to important features while suppressing irrelevant or noisy features. This mechanism can improve the feature quality and reduce information loss. Other parts of this implementation scheme are the same as those of the sixth specific implementation scheme.
[0081] A 2D pedestrian pose estimation method (algorithm) proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on the algorithm.
[0082] Based on the method proposed by the present invention, a 2D pedestrian pose estimation system is developed using a programming language. The system has program modules corresponding to the steps of the above technical solution and executes the steps in the above 2D pedestrian pose estimation method when running.
[0083] Store the computer program of the developed system (software) on a computer-readable storage medium. The computer program is configured to implement the steps of the above-mentioned 2D pedestrian pose estimation method when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.
[0084] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0085] The computational programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor and can implement these computational programs using high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disc, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0086] The beneficial effects of the present invention will be described below in conjunction with specific embodiments.
[0087] Embodiment 1
[0088] An ablation experiment was conducted on an open-source MRI dataset to verify the improvement effect of the Mamba structure and the multi-scale fusion module based on EMA on the model of the present invention in pedestrian pose estimation. The results of the ablation experiment under normal lighting are shown in Table 1.
[0089] Table 1
[0090]
[0091] In the MRI dataset, each image contains only one human body and is hardly affected by occlusion or blur, resulting in a relatively high overall accuracy for human keypoint detection. As shown in Table 1, under normal lighting conditions, compared with the baseline network, after only introducing Mamba, AP50, AP75, AP90, and mAP increased by 1.5%, 0.7%, 5.2%, and 1.8% respectively, while AR50, AR75, AR90, and mAR increased by 0.5%, 0.9%, 3.0%, and 1.1% respectively. After only introducing EMA, AP50, AP75, AP90, and mAP increased by 1.5%, 0.7%, 5.1%, and 1.8% respectively, while AR50, AR75, AR90, and mAR increased by 0.5%, 0.8%, 2.9%, and 1.1% respectively. When both Mamba and EMA are introduced simultaneously, the values of AP and AR are further improved. Compared with the baseline, AP50, AP75, AP90, and mAP increased by 1.5%, 0.7%, 6.2%, and 2.0% respectively, while AR50, AR75, AR90, and mAR increased by 0.5%, 0.9%, 3.5%, and 1.2% respectively. At the same time, the number of parameters decreased by 37% and the FLOPs decreased by 25%.
[0092] The results of HRMamba-EMA and other advanced human pose estimation models are shown in Table 2. The compared models include ResNet series (ResNet50, ResNet101, ResNet152), Hourglass, LiteHRNet-30, VGG16, Swin-T, AlexNet, ShuffleNetV2, MobileNetV2, HRFormer, and HRNet.
[0093] Table 2
[0094]
[0095]
[0096] As shown in Table 2, except for the Hourglass method, the input image size of other methods is 256x192. When IOU = 0.50 and 0.75, except for the Alexnet method, the AP and AR values of the remaining methods are relatively high, indicating that the current mainstream methods perform well under lower accuracy requirements. Comparing the average precision, only the mAP and mAR values of Resnet101, Resnet152, HRNet, and the HRMamba-EMA of the present invention exceed 90%, showing better performance under higher accuracy requirements, and the mAP and mAR values of HRMamba-EMA are the highest. It can be seen that the HRMamba-EMA of the present invention can achieve a relatively high pedestrian pose estimation accuracy while maintaining a relatively low computational complexity.
[0097] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the scope of protection of the present invention.
Claims
1. A 2D pedestrian pose estimation method, characterized in that, It includes the following steps: Step 1: Collect images through an in-vehicle camera and detect pedestrian targets in the images; Step 2: Construct an HRMamba-EMA model based on the HRNet network. Each branch of the model is constructed with a VSSBlock to process branches with different resolutions for feature extraction, and an EMA-based multi-scale fusion module is introduced for multi-scale feature fusion; Step 3: Based on the HRMamba-EMA model, perform pose estimation and key point detection on the obtained pedestrian target detection results.
2. The 2D pedestrian pose estimation method according to claim 1, characterized in that In Step 1, the detection of pedestrian targets in the images is specifically based on the YOLO object detection algorithm to detect pedestrian targets in the images.
3. The 2D pedestrian pose estimation method according to claim 2, characterized in that, The functional implementation process of the HRMamba-EMA model in Step 2 includes: In the first stage, two 3×3 convolutional layers are used to downsample the input image, reducing the resolution to 1 / 4 of the original, and at the same time increasing the number of channels to 64; subsequently, four Bottleneck modules are used to extract features and increase the number of channels to 256; In the second and fifth stages, each branch of the model contains four VSSBlock modules to process branches with different resolutions for feature extraction; in the fifth stage, an EMA-based multi-scale fusion module is added, and the EMA attention mechanism is introduced before each upsampling to fuse features of different scales. Finally, a heatmap of each key point of the pedestrian is output.
4. The 2D pedestrian pose estimation method according to claim 3, characterized in that, The functional implementation process of the VSSblock is as follows: The input is split into two branches after layer normalization. In the first branch, the input first passes through a linear layer and then through the activation function Silu. In the second branch, the input is first processed through a linear layer, a depthwise separable convolution, and the activation function Silu, and then input into the SS2D module for further feature extraction. Subsequently, the features are normalized through layer normalization. Finally, the two paths are merged, and the features are merged using a linear layer, and this result is combined with the residual connection to obtain an enhanced feature tensor.
5. The 2D pedestrian pose estimation method according to claim 4, wherein The functional implementation process of the SS2D module is as follows: Scan expansion: First, the input image is divided into 16 small images of equal size in a 4×4 manner, and then the small images are scanned one by one in the order from left to right, from top to bottom, from bottom to top, and from right to left; Selective Scanning of Block S6: Block S6 has a Mamba structure, with the state space model SSM as the main body, relying on continuous systems. The input sequence x(t) ∈ R is mapped to the output y(t) ∈ R through the intermediate implicit state h(t) ∈ R N by a linear time-invariant system, represented as a linear ordinary differential equation: h′(t) = Ah(t) + Bx(t) y(t) = Ch(t) where A is the state transition matrix, A ∈ R N×N , B is the input projection matrix for mapping the input to the hidden state space, B ∈ R N×1 , C is the output projection matrix for calculating the final output, C ∈ R D×1 , and N is the state size; Introduce a time scale parameter Δ and convert A and B into continuous parameters using the discretization rule and The zero-order hold is used as the discretization rule, specifically: After discretization, based on the SSM model, it is calculated through linear recursion: y(t) = Ch(t) Or calculated through global convolution: Among them, is a convolution kernel, L is the sequence length, and y is the final output; Scan merging: The small images processed by the S6 block are recombined in order into an image with the same size as the input.
6. The 2D pedestrian pose estimation method according to claim 5, wherein The functional implementation process of the EMA-based multi-scale fusion module is as follows: Features of different scales are respectively subjected to feature extraction by the EMA module and gradually upsampled. At the same time, after the feature fusion of different scales, the EMA module is continuously added to further enhance the feature expression ability.
7. The 2D pedestrian pose estimation method according to claim 6, wherein The functional implementation process of the EMA module is as follows: First, the input feature map is divided into G sub-features in the channel dimension. Subsequently, three parallel paths are used to extract the attention weights of the grouped feature maps, where two paths are 1×1 branches and the other path is a 3×3 branch. In the 1×1 branches, the channels are encoded along two spatial directions respectively, and in the 3×3 branch, a single 3×3 convolution is stacked to capture multi-scale feature representations; The output of the 1x1 branch encodes the global spatial information through two-dimensional global average pooling, and the output of the 3x3 branch is directly transformed into the corresponding dimensional shape; then, the outputs are aggregated through a matrix dot product operation to generate the first spatial attention map; finally, the output feature maps within each group are aggregated through the Sigmoid function of the two generated spatial attention weight values to capture the pixel-level pairing relationship.
8. A 2D pedestrian pose estimation method system, characterized in that, The system has program modules corresponding to the steps of the method described in any one of claims 1 to 7 above, and executes the steps in the above 2D pedestrian pose estimation method when running.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps in the 2D pedestrian pose estimation method described in any one of claims 1 to 7 when called by a processor.
Citation Information
Patent Citations
Multi-person attitude estimation method and system based on GAMHR-Net
CN115019338A
Human body posture estimation method based on scale feature and hierarchical feature fusion
CN117711023A