Methods, devices, equipment, and media for 3D human pose estimation based on spatiotemporal fusion of multimodal features.

By using a multimodal feature spatiotemporal fusion method, joint relationship modeling is established and long-term and short-term features are captured, which solves the jitter problem in 3D human pose estimation, realizes stable and coherent 3D human pose estimation, and improves accuracy.

CN122290219APending Publication Date: 2026-06-26PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIV SHENZHEN GRADUATE SCHOOL
Filing Date
2026-05-27
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing 3D human pose estimation methods are prone to pose jitter and inconsistency in continuous frame sequences, especially under occlusion conditions, where multimodal feature fusion is difficult to overcome temporal jitter and occlusion interference.

Method used

By extracting multimodal features from image sequences, a joint relationship model between two-dimensional human pose and other modal features is established. Long-term global features and short-term detail features in attention-enhanced feature sequences are captured, and three-dimensional human pose is estimated by combining long-term global features and short-term detail features.

Benefits of technology

It achieves stable and coherent 3D human pose estimation in consecutive frames, reduces pose jitter, and improves the accuracy of 3D human pose estimation in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290219A_ABST
    Figure CN122290219A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for 3D human pose estimation based on spatiotemporal fusion of multimodal features. The method includes extracting a multimodal feature sequence from an image sequence; performing joint relationship modeling on the multimodal features to obtain an attention-enhanced feature sequence; capturing long-term global features and short-term detail features from the attention-enhanced feature sequence; and estimating 3D human pose based on the long-term global features and short-term detail features. This application utilizes short-term detail features reflecting periodic or detailed changes in the sequence, as well as long-term global features reflecting overall trends and cyclical patterns, to perform 3D human pose estimation. This achieves stable and coherent 3D human pose estimation in continuous frames, reduces pose jitter caused by independent estimation in single frames, and effectively improves the accuracy of 3D human pose estimation in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus, device and medium for estimating three-dimensional human pose based on spatiotemporal fusion of multimodal features. Background Technology

[0002] 3D human pose estimation not only provides technical support for complex application scenarios but also demonstrates enormous potential in multiple fields. For example, in areas such as intelligent monitoring, virtual reality, augmented reality, and healthcare, accurate human pose estimation can help improve the intelligence level of systems and even enable real-time monitoring of individual safety. 3D human pose estimation aims to infer the positions of various joints of the human body in three-dimensional space from images or videos. The current mainstream implementation method is a lift-based two-stage approach. In the first stage, this type of method uses an existing 2D human pose estimator to estimate 2D human coordinates. In the second stage, using only the 2D human coordinates as input, the 2D human coordinates are lift-up to 3D human coordinates.

[0003] However, existing methods are generally limited to single-frame processing, which makes the estimated pose in continuous frame sequences prone to frequent jumps, seriously affecting the stability and coherence of 3D pose. Especially in the presence of occlusion, multimodal feature fusion at the single-frame level is difficult to effectively overcome temporal jitter and occlusion interference, making this problem more prominent.

[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide a method, device, equipment and medium for three-dimensional human pose estimation based on spatiotemporal fusion of multimodal features, which addresses the shortcomings of the existing technology.

[0006] To address the aforementioned technical problems, the first aspect of this application provides a three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features, wherein the three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features specifically includes: Multimodal feature extraction is performed on the image sequence to obtain a multimodal feature sequence, wherein the multimodal features in the multimodal feature sequence include two-dimensional human pose, depth features, and image features; Joint relationship modeling is performed on the multimodal features in the multimodal feature sequence to obtain attention-enhanced feature sequence. The joint relationship modeling is used to establish the interaction between the two-dimensional human posture and other modal features in the multimodal features to mine complementary information between the two-dimensional human posture and other modal features and to understand the human skeletal system. The long-term overall features and short-term detail features in the attention-enhanced feature sequence are captured. The long-term overall features are used to reflect the overall trend and cyclical pattern of the image sequence, while the short-term detail features are used to reflect the periodicity or short-term detail changes of the image sequence. High-frequency joint motion features are captured from the short-term detailed features, and three-dimensional human posture is estimated based on the long-term overall features and the high-frequency joint motion features.

[0007] The aforementioned three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features, wherein the step of modeling joint relationships among the multimodal features in the multimodal feature sequence to obtain an attention-enhanced feature sequence specifically includes: For each multimodal feature in the multimodal feature sequence, a multi-head attention mechanism is used to fuse the multimodal features to obtain a multimodal fused feature; The joint relationships of the multimodal fusion features are modeled using an empty attention mechanism to obtain attention-enhanced features.

[0008] The aforementioned 3D human pose estimation method based on multimodal feature spatiotemporal fusion, wherein capturing the long-term overall features and short-term detail features in the attention-enhanced feature sequence specifically includes: Obtain a preset number of moving average windows, where each moving average window has a different scale; Extract the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature component corresponding to each moving average window, and determine the long-term overall feature based on all overall feature components; The detail feature components corresponding to each moving average window are determined based on the overall feature components and the attention enhancement features, and short-term detail features are determined based on all detail feature components.

[0009] The aforementioned three-dimensional human pose estimation method based on multimodal feature spatiotemporal fusion, wherein the step of extracting the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature component corresponding to each moving average window specifically includes: For each moving average window, mirror padding is performed along the sequence length dimension of the attention-enhanced feature sequence to ensure that the sequence length dimension of the overall feature components is the same as the sequence length dimension of the attention-enhanced feature sequence. The moving average window is slid across the filled attention-enhanced feature sequence, and average pooling is performed on the attention-enhanced features covered by the moving average window to obtain the overall feature components corresponding to the moving average window.

[0010] The aforementioned 3D human pose estimation method based on multimodal feature spatiotemporal fusion, wherein determining the detail feature components corresponding to each moving average window based on the overall feature components corresponding to each moving average window and the attention-enhanced features specifically includes: For each moving average window, the attention enhancement feature and the overall feature component corresponding to the moving average window are subtracted element by element to obtain the detail feature component corresponding to the moving average window.

[0011] The aforementioned three-dimensional human pose estimation method based on multimodal feature spatiotemporal fusion, wherein capturing the high-frequency joint motion features in the short-term detail features specifically includes: The short-term detail features are extracted using multiple parallel high-frequency information extraction paths to obtain multiple joint motion high-frequency feature components. The high-frequency information extraction paths include downsampling, equal-length convolution, and upsampling operations. Multiple joint motion high-frequency feature components are spliced ​​together to obtain spliced ​​high-frequency features; The spliced ​​high-frequency features are convolved to obtain the high-frequency features of joint motion.

[0012] The aforementioned three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features, wherein the estimation of three-dimensional human pose based on the long-term overall features and high-frequency joint motion features specifically includes: Multilayer perceptrons are used to map and interact with long-term overall features to obtain intermediate overall features; The intermediate overall features are fused with the high-frequency features of joint movements to form a human posture sequence; The human posture sequence is subjected to a dimension reshaping operation to determine the three-dimensional human posture.

[0013] A second aspect of this application provides a three-dimensional human pose estimation device based on spatiotemporal fusion of multimodal features, wherein the three-dimensional human pose estimation device based on spatiotemporal fusion of multimodal features specifically includes: The acquisition module is used to extract multimodal features from the image sequence to obtain a multimodal feature sequence, wherein the multimodal features in the multimodal feature sequence include two-dimensional human pose, depth features, and image features; An enhancement module is used to perform joint relationship modeling on the multimodal features in the multimodal feature sequence to obtain an attention-enhanced feature sequence. The joint relationship modeling is used to establish the interaction between the two-dimensional human posture and other modal features in the multimodal features to mine complementary information between the two-dimensional human posture and other modal features and to understand the human skeletal system. The capture module is used to capture long-term overall features and short-term detail features in the attention-enhanced feature sequence. The long-term overall features are used to reflect the overall trend and cyclical pattern of the image sequence, and the short-term detail features are used to reflect the periodicity or short-term detail changes of the image sequence. The pose estimation module is used to capture the high-frequency joint motion features in the short-term detailed features and estimate the three-dimensional human pose based on the long-term overall features and the high-frequency joint motion features.

[0014] A third aspect of this application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the three-dimensional human pose estimation method based on multimodal feature spatiotemporal fusion as described above.

[0015] A fourth aspect of this application provides a terminal device, which includes: a processor and a memory; The memory stores a computer-readable program that can be executed by the processor; When the processor executes the computer-readable program, it implements the steps in the three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features as described above.

[0016] Beneficial Effects: This application provides a method, apparatus, device, and medium for 3D human pose estimation based on spatiotemporal fusion of multimodal features. The method includes extracting multimodal features from an image sequence to obtain a multimodal feature sequence; modeling joint relationships of the multimodal features in the multimodal feature sequence to obtain an attention-enhanced feature sequence; capturing long-term global features and short-term detail features in the attention-enhanced feature sequence; and estimating 3D human pose based on the long-term global features and short-term detail features. This application first determines the attention-enhanced feature sequence by modeling joint relationships of the multimodal features; then captures short-term detail features reflecting periodic or detailed changes in the sequence, as well as long-term global features reflecting overall trends and cyclical patterns; finally, it combines the long-term global features and short-term detail features to perform 3D human pose estimation, achieving stable and coherent 3D human pose estimation in continuous frames, reducing pose jitter caused by independent estimation in a single frame, and effectively improving the accuracy of 3D human pose estimation in complex scenes. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram of the principle of the three-dimensional human pose estimation system provided in the embodiments of this application.

[0019] Figure 2 A flowchart of a three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features provided in this application embodiment.

[0020] Figure 3 An example diagram showing the three-dimensional human pose estimation results provided in the embodiments of this application.

[0021] Figure 4 This is a schematic diagram of the principle of the three-dimensional human pose estimation device based on spatiotemporal fusion of multimodal features provided in the embodiments of this application.

[0022] Figure 5 A schematic block diagram of the terminal device provided in the embodiments of this application. Detailed Implementation

[0023] This application provides a method, apparatus, device, and medium for three-dimensional human pose estimation based on spatiotemporal fusion of multimodal features. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.

[0024] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0025] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0026] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.

[0027] Research has revealed that 3D human pose estimation not only provides technical support for complex application scenarios but also demonstrates enormous potential in multiple fields. For example, in areas such as intelligent surveillance, virtual reality, augmented reality, and healthcare, accurate human pose estimation can help improve the intelligence level of systems and even enable real-time monitoring of individual safety. 3D human pose estimation aims to infer the positions of various joints of the human body in three-dimensional space from images or videos. The current mainstream implementation method is a lift-based two-stage approach. This type of method uses an existing 2D human pose estimator to estimate 2D human coordinates in the first stage, and in the second stage, using only the 2D human coordinates as input, lifts them to 3D human coordinates.

[0028] However, existing methods are generally limited to single-frame processing, which makes the estimated pose in continuous frame sequences prone to frequent jumps, seriously affecting the stability and coherence of 3D pose. Especially in the presence of occlusion, multimodal feature fusion at the single-frame level is difficult to effectively overcome temporal jitter and occlusion interference, making this problem more prominent.

[0029] To address the aforementioned issues, this application embodiment involves extracting multimodal features from image sequences to obtain multimodal feature sequences; modeling joint relationships within the multimodal features of the multimodal feature sequences to obtain attention-enhanced feature sequences; capturing long-term global features and short-term detail features within the attention-enhanced feature sequences; and estimating 3D human pose based on the long-term global features and short-term detail features. This application first determines the attention-enhanced feature sequences by modeling joint relationships within the multimodal features; then captures short-term detail features reflecting periodic or detailed changes, as well as long-term global features reflecting overall trends and cyclical patterns within the sequence; finally, it combines the long-term global features and short-term detail features to perform 3D human pose estimation, achieving stable and coherent 3D human pose estimation across consecutive frames, reducing pose jitter caused by independent estimation in single frames, and effectively improving the accuracy of 3D human pose estimation in complex scenes.

[0030] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.

[0031] This embodiment provides a 3D human pose estimation method based on spatiotemporal fusion of multimodal features. The method utilizes a 3D human pose estimation system, which includes a feature extraction component and a 3D pose estimation component. The feature extraction component extracts multimodal features, and the 3D pose estimation component is connected to the feature extraction component to estimate the 3D human pose based on the multimodal feature sequence provided by the feature extraction component. Specifically, as shown... Figure 1 As shown, the feature extraction component includes several parallel feature extraction modules. Each feature extraction module is used to extract one modality of features, such as multimodal features including 2D human pose, depth features, and image features. The feature extraction component can include a feature extraction module for 2D human pose, a feature extraction module for depth features, and a feature extraction module for image features. The 3D pose estimation component includes an attention enhancement module, a temporal dual-branch module, and a pose estimation module. The attention enhancement module is used to determine attention enhancement features, which can include a multimodal feature fusion unit and a spatial attention unit. The temporal dual-branch module is used to capture long-term overall features and short-term detail features. The pose estimation module is used to estimate the 3D human pose.

[0032] like Figure 2 As shown in the embodiments of this application, the three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features specifically includes: S10. Perform multimodal feature extraction on the image sequence to obtain a multimodal feature sequence.

[0033] Specifically, an image sequence comprises multiple frames, such as monocular RGB images captured by a monocular camera for each frame. The image sequence can include all video frames from a video clip. However, since the image acquisition frequency is higher than the frequency of human posture changes, consecutive frames may contain a large amount of redundant information, and directly processing all video frames would increase the computational burden. Therefore, in practical applications, key video frames can be selected from the video clip according to a preset sampling interval, such as selecting one frame every 3 or 5 frames. This preserves sufficient temporal information while effectively reducing the data volume and improving subsequent processing efficiency.

[0034] A multimodal feature sequence is a sequence composed of multiple modal features extracted from each frame of an image sequence, arranged in chronological order. Multimodal features include, but are not limited to, 2D human pose, depth features, and image features. 2D human pose can be obtained by extracting joint coordinates from images using existing 2D pose estimation algorithms. Depth features can be acquired by depth sensors or obtained through monocular image depth estimation algorithms. Image features can be extracted from images using feature extraction models such as convolutional neural networks to extract discriminative visual features. By extracting multimodal features from each frame of the image sequence, a multimodal feature sequence can be constructed, providing a data foundation for subsequent joint relationship modeling and pose estimation.

[0035] S20. Perform joint relationship modeling on the multimodal features in the multimodal feature sequence to obtain the attention-enhanced feature sequence.

[0036] Specifically, attention-enhanced features are cross-modal joint representations obtained by modeling multimodal features through joint relationships. The joint relationship modeling aims to establish the interaction between two-dimensional human pose and other modal features (such as depth features, image features, etc.) to mine complementary information between two-dimensional human pose and other modal features and to understand the human skeletal system.

[0037] In one embodiment, modeling the joint relationships of the multimodal features in the multimodal feature sequence to obtain the attention-enhanced feature sequence specifically includes: For each multimodal feature in the multimodal feature sequence, a multi-head attention mechanism is used to fuse the multimodal features to obtain a multimodal fused feature; The joint relationships of the multimodal fusion features are modeled using an empty attention mechanism to obtain attention-enhanced features.

[0038] Specifically, multi-head attention mechanisms are used to establish interactions between 2D human pose and other modal features (such as depth features and image features) to uncover complementary information between them. The core of the multi-head attention mechanism lies in using multiple parallel attention heads to interact and fuse multimodal features from different subspaces. Each attention head independently calculates the association weights between different modal features, capturing local dependencies between modalities. The outputs of each attention head are then concatenated and linearly transformed to obtain preliminary multimodal fusion features. This process effectively integrates the advantageous information of different modal features, such as the skeletal structure information provided by 2D human pose, the spatial relationships reflected by depth features, and the texture details contained in image features, thus improving the representational power of multimodal fusion features.

[0039] The process of obtaining multimodal fusion features can be represented as follows: , , in, Represents multimodal features, This indicates the bulls' self-attention. Indicates intermediate features, This represents a multilayer perceptron. This represents the multimodal fusion feature.

[0040] Furthermore, the spatial attention mechanism is used to focus on the spatial structural relationships between human joints to understand the human skeletal system. When modeling joint relationships using this mechanism, the multimodal fusion features are first flattened. These flattened features are then input into a spatial attention unit. The spatial attention unit calculates the spatial association weights between different joint points, dynamically adjusting these weights to highlight key joint features and suppress interference from irrelevant background or redundant information. Specifically, the spatial attention unit performs convolution operations on the multimodal fusion features to generate a spatial attention map. Each element in this attention map corresponds to a weight at a different location in the original feature map; a higher weight indicates a greater contribution to joint relationship modeling at that location. Subsequently, the multimodal fusion features are multiplied element-wise with the spatial attention map to obtain features incorporating spatial attention information—i.e., attention-enhanced features. Through this process, the model can more accurately capture the spatial distribution patterns and interdependencies of human joints, laying a more reliable foundation for subsequent temporal feature capture.

[0041] S30. Capture the long-term overall features and short-term detail features in the attention-enhanced feature sequence.

[0042] Specifically, long-term global features are used to reflect the overall trend and cyclical pattern of an image sequence, while short-term detail features are used to reflect the periodicity or short-term detail changes of an image sequence. In other words, long-term global features are the global correlation features of an image sequence, while short-term detail features are the local detail features of an image sequence. By combining long-term global features and short-term detail features, a holistic view of an image sequence can be captured.

[0043] In one embodiment, capturing the long-term global features and short-term detail features in the attention-enhanced feature sequence specifically includes: S31. Obtain a preset number of moving average windows, wherein each moving average window has a different scale; S32. Extract the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature component corresponding to each moving average window, and determine the long-term overall feature based on all overall feature components. S33. Determine the detail feature components corresponding to each moving average window based on the overall feature components and the attention enhancement features, and determine the short-term detail features based on all detail feature components.

[0044] In step S31, the preset number can be flexibly set according to the length of the image sequence and the actual application scenario requirements. For example, it can be set to three moving average windows of different scales. Each moving average window has a different scale, that is, each moving average window corresponds to a different time span. For example, the three moving average windows of different scales correspond to the first time span, the second time span, and the third time span, respectively. The first time span is longer than the second time span, and the second time span is longer than the third time span.

[0045] In step S32, extracting the overall trend of the attention enhancement feature sequence according to each moving average window means sliding the moving average window over the attention enhancement feature sequence and extracting the overall trend within the coverage area of ​​the moving average window at each sliding step. For example, a moving average window with a first time span can capture the macroscopic pattern of human posture changes over time in the image sequence, such as the periodic swaying trend when walking; a moving average window with a second time span can reflect the overall changes in posture over a short period of time, such as the gradual increase or decrease in the amplitude of arm swing; a moving average window with a third time span can focus on the overall correlation of posture between adjacent frames.

[0046] In one specific embodiment, the step of extracting the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature components corresponding to each moving average window specifically includes: For each moving average window, mirror padding is performed along the sequence length dimension of the attention-enhanced feature sequence to ensure that the sequence length dimension of the overall feature components is the same as the sequence length dimension of the attention-enhanced feature sequence. The moving average window is slid across the filled attention-enhanced feature sequence, and average pooling is performed on the attention-enhanced features covered by the moving average window to obtain the overall feature components corresponding to the moving average window.

[0047] Specifically, the mirror fill operation avoids excessive weakening of edge frame features during pooling by adding symmetrical feature values ​​at the beginning and end of the sequence, ensuring that the overall feature component at each position can integrate information from the preceding and following windows. For example, when the moving average window spans 5 frames, mirror values ​​of 2 frames are added to the beginning and end of the sequence, so that the overall feature component of the first frame can contain the feature information of the following 2 frames, and the last frame can fuse the features of the previous 2 frames. After mirror fill, the moving average window slides across the attention-enhanced feature sequence with a fixed step size, and average pooling is performed on the attention-enhanced features covered by the moving average window. That is, the mean of all attention-enhanced features within the moving average window is calculated and used as the feature element of the overall feature component at the current position. Through average pooling operations with different moving average windows, overall feature components reflecting the overall trend of different time spans can be obtained. For example, a moving average window with a long time span corresponds to a slow change trend, while a moving average window with a short time span captures a fast change trend. The fixed step size is 1, meaning that each time step is slid by one time step, ensuring that the sequence length of the overall feature components is consistent with the original attention-enhanced feature sequence. This allows each overall feature component to fully cover the time dimension of the attention-enhanced feature sequence, laying the foundation for subsequent fusion of overall trends at different scales.

[0048] Furthermore, after obtaining the overall feature components corresponding to each moving average window, all overall feature components can be averaged to obtain long-term overall features. This effectively integrates the overall trends across different time spans, forming a comprehensive representation of the overall trend and cyclical patterns of the image sequence. Simultaneously, by averaging all overall feature components to obtain long-term overall features, trend information at different scales can be balanced, avoiding potential local biases that may exist in a single window, and enabling the long-term overall features to more comprehensively reflect the changing patterns of human posture over a longer period.

[0049] In step S33, determining the detail feature components based on the overall feature components and attention-enhanced features corresponding to each moving average window aims to capture local fluctuations by observing the difference between the attention-enhanced feature sequence and the overall trend. This is because the overall feature components reflect the overall trend at the scale of the moving average window, so the detail feature components obtained by subtracting the attention-enhanced features from the overall feature components can highlight short-term fluctuations and subtle changes in the attention-enhanced features that deviate from the overall trend. For example, when the overall feature components represent the average posture trend of human walking over a certain period of time, the detail feature components can reflect subtle adjustments or sudden movements of joints in each frame of the image, such as changes in the height of the feet or non-periodic swings of the arms.

[0050] Based on this, in one embodiment, determining the detail feature components corresponding to each moving average window based on the overall feature components corresponding to each moving average window and the attention enhancement features specifically includes: For each moving average window, the attention enhancement feature and the overall feature component corresponding to the moving average window are subtracted element by element to obtain the detail feature component corresponding to the moving average window.

[0051] Specifically, the element-wise subtraction operation involves subtracting each element in the attention-enhanced feature sequence from the corresponding global feature component element. The difference obtained is the detail feature component element at that position. For example, if the keypoint coordinate feature value of the attention-enhanced feature in a certain frame is ( , ), and the overall feature component elements corresponding to this frame are ( , Then the element of the detail feature component at that position is ( , By subtracting elements one by one, we can accurately separate out the local changes in attention enhancement features that exceed the overall trend. These local changes include short-term dynamic details of human posture, such as rapid joint rotations and subtle adjustments to body balance.

[0052] Furthermore, after obtaining the detail feature components corresponding to each moving average window, short-term detail features can be determined based on all detail feature components. For example, all detail feature components can be averaged to obtain short-term detail features. These short-term detail features and long-term overall features together constitute a comprehensive description of the attention-enhanced feature sequence, providing rich feature support for subsequent pose estimation tasks that includes both macro trends and micro details.

[0053] S40. Capture the high-frequency joint motion features in the short-term detailed features, and estimate the three-dimensional human posture based on the long-term overall features and the high-frequency joint motion features.

[0054] Specifically, the 3D human pose is estimated by fusing long-term overall features and short-term detail features. The estimated 3D human pose is a sequence of 3D human poses, which includes the 3D human pose corresponding to each frame in the image sequence. In estimating the 3D human pose based on the long-term overall features and short-term detail features, the long-term overall features provide the changing trends and cyclical patterns of the human pose over a longer time range, such as the periodic swaying pattern during walking, while the short-term detail features capture subtle adjustments and sudden movements at joints in each frame, such as non-periodic arm swings and changes in foot lift height. By fusing the long-term overall features and short-term detail features, the estimation of the 3D human pose considers both the overall motion trend and local dynamic details, thereby improving the accuracy and robustness of the pose estimation. For example, as... Figure 3 As shown, the three-dimensional human pose estimated in the embodiments of this application also has high accuracy in the case of self-occlusion.

[0055] Furthermore, high-frequency joint motion features are high-frequency information such as detailed jitter and short-term loops in joint motion within short-term detail features, including rapid joint rotation within a short period and subtle positional shifts between adjacent frames. These high-frequency joint motion features can be extracted using a pre-defined feature extraction module, such as one based on a convolutional neural network, which separates the high-frequency joint motion features from the short-term detail features. These high-frequency joint motion features reflect the dynamic changes in human posture in an instant, which is crucial for accurately capturing rapid movements or complex posture adjustments.

[0056] In one embodiment, capturing the high-frequency joint motion features in the short-term detail features specifically includes: The short-term detail features are extracted using multiple parallel high-frequency information extraction paths at multiple scales to obtain multiple joint motion high-frequency feature components. Multiple joint motion high-frequency feature components are spliced ​​together to obtain spliced ​​high-frequency features; The spliced ​​high-frequency features are convolved to obtain the high-frequency features of joint motion.

[0057] Specifically, the high-frequency information extraction path includes downsampling, equal-length convolution, and upsampling operations. The downsampling operation compresses the scale of the short-term detail features to extract large-scale motion features. The equal-length convolution operation performs convolution operations on the short-term detail features to extract high-frequency detail features. The upsampling operation restores the scale of the short-term detail features to ensure consistent output dimensions across multiple parallel high-frequency information extraction paths. Furthermore, in practice, different high-frequency information extraction paths can be configured with different downsampling factors and convolution kernel sizes to adapt to the joint motion feature extraction requirements of different frequency ranges. For example, one path can use 2x downsampling combined with a 3×3 convolution kernel to focus on capturing medium-scale high-frequency motion; another path can use 4x downsampling combined with a 5×5 convolution kernel to extract high-frequency changes at a larger scale. After extracting the multi-scale high-frequency feature components, the components are concatenated along the channel dimension to form a concatenated high-frequency feature containing multi-dimensional high-frequency information. Subsequently, channel fusion and feature dimensionality reduction are performed on the spliced ​​high-frequency features using a 1×1 convolution kernel, ultimately outputting high-frequency joint motion features with unified dimensions, providing accurate dynamic details for subsequent 3D human pose estimation.

[0058] Furthermore, when estimating the 3D human posture based on the long-term overall features and the high-frequency joint motion features, a multilayer perceptron can be used to map and interact with the long-term overall features to obtain intermediate overall features. Then, the intermediate overall features are fused with the high-frequency joint motion features (e.g., by pixel-by-pixel addition) to form a human posture sequence. Finally, a dimensionality reshaping operation is performed on this human posture sequence to determine the 3D human posture, such as using... Dimensional reshaping, etc.

[0059] In summary, this embodiment provides a 3D human pose estimation method based on spatiotemporal fusion of multimodal features. The method includes extracting multimodal features from an image sequence to obtain a multimodal feature sequence; modeling joint relationships within the multimodal features in the multimodal feature sequence to obtain an attention-enhanced feature sequence; capturing long-term global features and short-term detail features from the attention-enhanced feature sequence; and estimating 3D human pose based on the long-term global features and short-term detail features. This application first determines the attention-enhanced feature sequence by modeling joint relationships within the multimodal features; then captures short-term detail features reflecting periodic or detailed changes, as well as long-term global features reflecting overall trends and cyclical patterns; finally, it combines the long-term global features and short-term detail features to perform 3D human pose estimation, achieving stable and coherent 3D human pose estimation across consecutive frames, reducing pose jitter caused by independent estimation in single frames, and effectively improving the accuracy of 3D human pose estimation in complex scenes.

[0060] Based on the aforementioned 3D human pose estimation method based on multimodal feature spatiotemporal fusion, this embodiment provides a 3D human pose estimation device based on multimodal feature spatiotemporal fusion, such as... Figure 4 As shown, the three-dimensional human pose estimation device based on multimodal feature spatiotemporal fusion specifically includes: The acquisition module 100 is used to extract multimodal features from the image sequence to obtain a multimodal feature sequence, wherein the multimodal features in the multimodal feature sequence include two-dimensional human pose, depth features and image features; The enhancement module 200 is used to perform joint relationship modeling on the multimodal features in the multimodal feature sequence to obtain the attention enhancement feature sequence. The joint relationship modeling is used to establish the interaction between the two-dimensional human posture and other modal features in the multimodal features to mine the complementary information between the two-dimensional human posture and other modal features and to understand the human skeletal system. The capture module 300 is used to capture long-term overall features and short-term detail features in the attention-enhanced feature sequence. The long-term overall features are used to reflect the overall trend and cyclical pattern of the image sequence, and the short-term detail features are used to reflect the periodicity or short-term detail changes of the image sequence. The pose estimation module 400 is used to capture the high-frequency joint motion features in the short-term detail features and estimate the three-dimensional human pose based on the long-term overall features and the high-frequency joint motion features.

[0061] Based on the above-described three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features, this embodiment provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in the three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features as described in the above embodiment.

[0062] Based on the aforementioned 3D human pose estimation method based on multimodal feature spatiotemporal fusion, this application also provides a terminal device, such as... Figure 5 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical instructions in the memory 22 to execute the methods described in the above embodiments.

[0063] Furthermore, the logical instructions in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0064] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.

[0065] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, as well as transient storage media.

[0066] Furthermore, the specific process of loading and executing multiple instruction processors in the aforementioned storage medium and terminal device has been described in detail in the above method, and will not be repeated here.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A three-dimensional human pose estimation method based on multi-modal feature spatio-temporal fusion, characterized in that, The aforementioned three-dimensional human pose estimation method based on multimodal feature spatiotemporal fusion specifically includes: Multimodal feature extraction is performed on the image sequence to obtain a multimodal feature sequence, wherein the multimodal features in the multimodal feature sequence include two-dimensional human pose, depth features, and image features; Joint relationship modeling is performed on the multimodal features in the multimodal feature sequence to obtain attention-enhanced feature sequence. The joint relationship modeling is used to establish the interaction between the two-dimensional human posture and other modal features in the multimodal features to mine complementary information between the two-dimensional human posture and other modal features and to understand the human skeletal system. The long-term overall features and short-term detail features in the attention-enhanced feature sequence are captured. The long-term overall features are used to reflect the overall trend and cyclical pattern of the image sequence, while the short-term detail features are used to reflect the periodicity or short-term detail changes of the image sequence. High-frequency joint motion features are captured from the short-term detailed features, and three-dimensional human posture is estimated based on the long-term overall features and the high-frequency joint motion features.

2. The three-dimensional human pose estimation method based on multi-modal feature spatio-temporal fusion according to claim 1, characterized in that, The step of modeling the joint relationships of the multimodal features in the multimodal feature sequence to obtain the attention-enhanced feature sequence specifically includes: For each multimodal feature in the multimodal feature sequence, a multi-head attention mechanism is used to fuse the multimodal features to obtain a multimodal fused feature; The joint relationships of the multimodal fusion features are modeled using an empty attention mechanism to obtain attention-enhanced features.

3. The three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features according to claim 1, characterized in that, The capture of long-term overall features and short-term detailed features in the attention-enhanced feature sequence specifically includes: Obtain a preset number of moving average windows, where each moving average window has a different scale; Extract the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature component corresponding to each moving average window, and determine the long-term overall feature based on all overall feature components; The detail feature components corresponding to each moving average window are determined based on the overall feature components and the attention enhancement features, and short-term detail features are determined based on all detail feature components.

4. The three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features according to claim 3, characterized in that, The step of extracting the overall trend of the attention-enhanced feature sequence according to each moving average window to obtain the overall feature components corresponding to each moving average window specifically includes: For each moving average window, mirror padding is performed along the sequence length dimension of the attention-enhanced feature sequence to ensure that the sequence length dimension of the overall feature components is the same as the sequence length dimension of the attention-enhanced feature sequence. The moving average window is slid across the filled attention-enhanced feature sequence, and average pooling is performed on the attention-enhanced features covered by the moving average window to obtain the overall feature components corresponding to the moving average window.

5. The three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features according to claim 3 or 4, characterized in that, The step of determining the detailed feature components corresponding to each moving average window based on the overall feature components corresponding to each moving average window and the attention-enhanced features specifically includes: For each moving average window, the attention enhancement feature and the overall feature component corresponding to the moving average window are subtracted element by element to obtain the detail feature component corresponding to the moving average window.

6. The three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features according to claim 1, characterized in that, The capture of high-frequency joint motion features in the short-term detail features specifically includes: The short-term detail features are extracted using multiple parallel high-frequency information extraction paths to obtain multiple joint motion high-frequency feature components. The high-frequency information extraction paths include downsampling, equal-length convolution, and upsampling operations. Multiple joint motion high-frequency feature components are spliced ​​together to obtain spliced ​​high-frequency features; The spliced ​​high-frequency features are convolved to obtain the high-frequency features of joint motion.

7. The three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features according to claim 1, characterized in that, The estimation of three-dimensional human posture based on the long-term overall features and high-frequency joint motion features specifically includes: Multilayer perceptrons are used to map and interact with long-term overall features to obtain intermediate overall features; The intermediate overall features are fused with the high-frequency features of joint movements to form a human posture sequence; The human posture sequence is subjected to a dimension reshaping operation to determine the three-dimensional human posture.

8. A three-dimensional human pose estimation device based on spatiotemporal fusion of multimodal features, characterized in that, The aforementioned three-dimensional human pose estimation device based on multimodal feature spatiotemporal fusion specifically includes: The acquisition module is used to extract multimodal features from the image sequence to obtain a multimodal feature sequence, wherein the multimodal features in the multimodal feature sequence include two-dimensional human pose, depth features, and image features; An enhancement module is used to perform joint relationship modeling on the multimodal features in the multimodal feature sequence to obtain an attention-enhanced feature sequence. The joint relationship modeling is used to establish the interaction between the two-dimensional human posture and other modal features in the multimodal features to mine complementary information between the two-dimensional human posture and other modal features and to understand the human skeletal system. The capture module is used to capture long-term overall features and short-term detail features in the attention-enhanced feature sequence. The long-term overall features are used to reflect the overall trend and cyclical pattern of the image sequence, and the short-term detail features are used to reflect the periodicity or short-term detail changes of the image sequence. The pose estimation module is used to capture the high-frequency joint motion features in the short-term detailed features and estimate the three-dimensional human pose based on the long-term overall features and the high-frequency joint motion features.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the three-dimensional human pose estimation method based on multimodal feature spatiotemporal fusion as described in any one of claims 1-7.

10. A terminal device, characterized in that, include: Processor and memory; The memory stores a computer-readable program that can be executed by the processor; When the processor executes the computer-readable program, it implements the steps in the three-dimensional human pose estimation method based on spatiotemporal fusion of multimodal features as described in any one of claims 1-7.