Human posture estimation method based on multi-source data, terminal device and storage medium

CN122551402APending Publication Date: 2026-08-11SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]三维人体姿态估计广泛应用于应急救援、医疗康复等领域,现有实现方案主要分为纯毫米波雷达、纯惯性传感器、多模态融合三类,其中,纯毫米波雷达方案受点云稀疏、多径效应影响,易出现方向模糊、动态细节丢失,跨场景泛化能力差;纯惯性传感器方案需要在人体多处佩戴多个设备,佩戴负担大,还存在积分造成的位置漂移问题;传统雷达与惯性传感器融合方案多采用简单特征拼接,无法适配两类异构数据的特性,融合效率低、抗干扰能力弱

Benefits of technology

[0007]在本申请实施例中,通过从骨盆惯性数据中提取运动时序特征构建仿射变换矩阵,再对雷达点云进行视角校正,有效消除了因人体运动、转向等带来的雷达观测视角偏差,让雷达点云与人体运动姿态在空间上精准对齐,为后续特征融合与姿态估计提供了更准确、低畸变的输入数据,显著提升了后续姿态估计的鲁棒性与精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551402A_ABST
    Figure CN122551402A_ABST
Patent Text Reader

Abstract

This application relates to the field of radar target detection and attitude recognition technology, and provides a human attitude estimation method, terminal device, and storage medium based on multi-source data. The method includes: collecting radar point cloud data and pelvic inertial data during human attitude changes; performing viewpoint correction on the radar point cloud data based on the pelvic inertial data to obtain viewpoint-aligned radar point cloud data; performing feature interaction fusion on the viewpoint-aligned radar point cloud data and pelvic inertial data to obtain fused features; and estimating human attitude based on the fused features to obtain the estimated human attitude. This method utilizes pelvic inertial data to correct radar point cloud viewpoint deviations, and then achieves deep integration of multi-source information through feature interaction fusion, which can reduce environmental interference and significantly improve the accuracy and reliability of human attitude estimation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of radar target detection and attitude recognition technology, and in particular relates to human attitude estimation methods, terminal devices and storage media based on multi-source data. Background Technology

[0002] 3D human pose estimation is widely used in emergency rescue, medical rehabilitation, and other fields. Existing solutions mainly fall into three categories: pure millimeter-wave radar, pure inertial sensors, and multimodal fusion. Pure millimeter-wave radar solutions are prone to directional ambiguity and loss of dynamic details due to point cloud sparsity and multipath effects, and suffer from poor cross-scene generalization. Pure inertial sensor solutions require wearing multiple devices at various points on the body, resulting in a heavy wearing burden and positional drift caused by integration. Traditional radar and inertial sensor fusion solutions often employ simple feature stitching, which cannot adapt to the characteristics of two heterogeneous data types, leading to low fusion efficiency and weak anti-interference capabilities. To address these shortcomings, there is an urgent need in this field for a technical solution that offers low wearing burden, good fusion results, and enhanced pose estimation accuracy and robustness. Summary of the Invention

[0003] This application provides a human pose estimation method, terminal device, and storage medium based on multi-source data, which can reduce environmental interference and significantly improve the accuracy and reliability of human pose estimation results.

[0004] In a first aspect, embodiments of this application provide a method for estimating human pose from multi-source data, a terminal device, and a storage medium, including: Collect radar point cloud data and pelvic inertial data during human posture changes; The radar point cloud data is corrected by viewing angle based on pelvic inertial data to obtain radar point cloud data with viewing angle alignment. The radar point cloud data and pelvic inertial data after viewpoint alignment are fused by feature interaction to obtain fused features; Human pose estimation is performed based on fused features to obtain the target pose of the human body.

[0005] In this embodiment, two types of data are collected. Inertial data is used to correct the viewpoint of the radar point cloud. Then, the corrected radar features and inertial features are interactively fused, and finally, accurate human posture estimation is achieved based on the fused features. Specifically, the motion and posture information of pelvic inertial data (acceleration, angular velocity) provides a benchmark for viewpoint correction of the radar point cloud, eliminating viewpoint deviations caused by human turning, lateral movement, etc., and aligning the spatial coordinates of the radar data with the human motion state. This solves the problem of feature distortion under viewpoint shift in pure radar methods. Secondly, through the interactive fusion of radar spatial features and inertial temporal features, the radar's accurate capture capability of human contours and spatial positions is preserved, while the motion dynamic information of inertial data is used to supplement the speed and posture change details that are difficult for radar to obtain directly. This achieves information complementarity between the two modes. Therefore, the above method uses pelvic inertial data to correct the viewpoint deviation of the radar point cloud, and then achieves deep integration of multi-source information through feature interaction fusion, which can reduce environmental interference and significantly improve the accuracy and reliability of human posture estimation results.

[0006] In one possible implementation of the first aspect, the radar point cloud data is subjected to viewpoint correction based on pelvic inertial data to obtain viewpoint-aligned radar point cloud data, including: Temporal features were extracted from pelvic inertial data to obtain motion temporal features; Construct an affine transformation matrix based on the temporal characteristics of motion; The radar point cloud data is corrected by affine transformation matrix to obtain the radar point cloud data with viewpoint alignment.

[0007] In this embodiment, by extracting motion temporal features from pelvic inertial data to construct an affine transformation matrix, and then performing viewpoint correction on the radar point cloud, the radar observation viewpoint deviation caused by human movement and turning is effectively eliminated, allowing the radar point cloud and human motion posture to be accurately aligned in space. This provides more accurate and low-distortion input data for subsequent feature fusion and posture estimation, significantly improving the robustness and accuracy of subsequent posture estimation.

[0008] In one possible implementation of the first aspect, an affine transformation matrix is ​​constructed based on the motion temporal characteristics, including: The rotation vector corresponding to the human body rotation angle is obtained based on the motion time sequence characteristics; The rotation vector is normalized to obtain the normalized rotation vector; Construct the affine transformation matrix based on the normalized rotation vector.

[0009] In this embodiment, by extracting rotation vectors from motion temporal features and constructing an affine transformation matrix after normalization, the perspective changes caused by human motion can be accurately and stably described. This avoids the numerical deviations or transformation anomalies that are prone to occur when directly using the original rotation vectors to construct the matrix, ensuring the accuracy and reliability of subsequent radar point cloud perspective correction and providing a more robust spatial alignment basis for subsequent attitude estimation.

[0010] In one possible implementation of the first aspect, the radar point cloud data is subjected to viewpoint correction based on an affine transformation matrix to obtain viewpoint-aligned radar point cloud data, including: Generate a standardized sampling grid based on the affine transformation matrix; Affine transformation is performed on the radar point cloud data according to the standardized sampling grid to obtain the affine transformed radar point cloud data. Spatial transformation is performed on the radar point cloud data after affine transformation to obtain viewpoint aligned radar point cloud data.

[0011] In this embodiment, a standardized sampling grid is generated by an affine transformation matrix, and affine transformation and spatial alignment correction are performed on the radar point cloud data accordingly. This achieves accurate spatial alignment between the radar point cloud and the human motion posture, effectively eliminating feature deviations caused by viewpoint shift and motion distortion, ensuring the quality of input data for subsequent feature fusion and attitude estimation, and providing a low-distortion, high-consistency radar data foundation for subsequent steps.

[0012] In one possible implementation of the first aspect, the viewpoint-aligned radar point cloud data and pelvic inertial data are subjected to feature interaction fusion to obtain fused features, including: The radar point cloud data after viewpoint alignment is feature-encoded to obtain the first feature data; Feature encoding is performed on the pelvic inertial data to obtain the second feature data; The first feature data and the second feature data are subjected to cross-attention fusion processing to obtain cross-attention features; The cross-attention features are adaptively weighted to obtain the fused features.

[0013] In the embodiments of this application, effective features are extracted from the two types of data respectively, intermodal correlations are established by cross-attention, and the feature ratio is optimized by adaptive weighting to achieve deep fusion of spatial and motion information, improve feature expression ability and anti-interference ability, and ensure the pose estimation effect.

[0014] In one possible implementation of the first aspect, human pose estimation is performed based on fused features to obtain the target pose estimation of the human body, including: Temporal features are extracted from the fused features to obtain global temporal features; Human pose estimation is performed based on the last frame of global temporal features to obtain the initial estimated human pose. The initial estimated attitude is corrected to obtain the target estimated attitude.

[0015] In this embodiment, feature encoding is performed on the radar point cloud data and pelvic inertial data after viewpoint alignment, and then fusion features are obtained through cross-attention fusion and adaptive weighting. This achieves deep interaction and complementarity between radar spatial information and inertial motion information. Cross-attention is used to capture the correlation and dependence between the two modes, and adaptive weighting is used to dynamically adjust the feature contribution, thereby generating more comprehensive and robust fusion features, providing more reliable feature support for subsequent human pose estimation.

[0016] In one possible implementation of the first aspect, the initial estimated pose is corrected to obtain the target estimated pose, including: Extracting Doppler channel data from radar point cloud data; The Doppler channel data is pooled to obtain the Doppler feature sequence; The initial estimated attitude is corrected based on the Doppler feature sequence to obtain the target estimated attitude.

[0017] In this embodiment, Doppler channel data is extracted from radar point cloud and pooled to obtain Doppler feature sequence, which is then used to correct the initial estimated attitude. This effectively introduces dynamic information reflecting the local motion speed of the human body, making up for the detail deviations when estimating attitude based solely on temporal and spatial features. This achieves accurate correction of the initial attitude, resulting in higher accuracy and stronger robustness of the final output target estimated attitude, especially with a significant reduction in error in dynamic motion scenarios.

[0018] In one possible implementation of the first aspect, the initial estimated pose is corrected based on the Doppler feature sequence to obtain the target estimated pose, including: The last frame of global temporal features, the initial estimated pose features, and the Doppler feature sequence are concatenated to obtain the concatenated features. The residual correction data for the human skeleton coordinates are calculated based on the splicing characteristics; The initial estimated attitude is superimposed with the residual correction data to obtain the target estimated attitude.

[0019] In this embodiment, a comprehensive characterization is constructed by splicing time sequence, attitude and Doppler multi-class features. Based on this, the skeleton coordinate residual is solved and superimposed on the initial attitude to complete the correction. By fully combining global motion law, initial attitude information and local velocity features, the estimation deviation is accurately compensated, and the accuracy and reliability of the final attitude result are greatly improved.

[0020] In a second aspect, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a human pose estimation method based on multi-source data as described in any of the first aspects above.

[0021] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements a human pose estimation method based on multi-source data as described in any of the first aspects above.

[0022] Fourthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the human pose estimation method based on multi-source data as described in any of the first aspects above.

[0023] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating the human pose estimation method based on multi-source data provided in an embodiment of this application. Figure 2 This is a schematic diagram of a process for correcting radar point cloud data according to an embodiment of this application; Figure 3 This is a schematic diagram of the process for constructing an affine transformation matrix provided in an embodiment of this application; Figure 4 This is a schematic diagram of a process for correcting radar point cloud data provided in another embodiment of this application; Figure 5 This is a schematic diagram of the feature interaction fusion process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the target attitude estimation process provided in an embodiment of this application; Figure 7 This is a schematic diagram of the target attitude estimation process provided in another embodiment of this application; Figure 8 This is a schematic diagram of the target attitude estimation process provided in another embodiment of this application; Figure 9 This is a schematic diagram of the overall architecture of a human pose estimation method based on multi-source data provided in an embodiment of this application; Figure 10 This is a schematic diagram of the overall architecture of attitude fine estimation provided in an embodiment of this application; Figure 11 This is a schematic diagram showing the ablation comparison results of different modal inputs provided in the embodiments of this application; Figure 12 This is a schematic diagram of the ablation experiment results of the MIMAC-Net core components provided in the embodiments of this application; Figure 13 This is a schematic diagram showing the comparison results between the two evaluation protocols provided in the embodiments of this application and the mRI baseline method; Figure 14 This is a schematic diagram showing the comparison results between the mRI dataset provided in this application and existing methods; Figure 15 This is a visual comparison diagram of different methods provided in the embodiments of this application; Figure 16 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0027] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0028] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0029] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0032] 3D human pose estimation is widely used in emergency rescue, medical rehabilitation, and other fields. Existing solutions mainly fall into three categories: pure millimeter-wave radar, pure inertial sensors, and multimodal fusion. Pure millimeter-wave radar solutions are prone to directional ambiguity and loss of dynamic details due to point cloud sparsity and multipath effects, and suffer from poor cross-scene generalization. Pure inertial sensor solutions require wearing multiple devices at various points on the body, resulting in a heavy wearing burden and positional drift caused by integration. Traditional radar and inertial sensor fusion solutions often employ simple feature stitching, which cannot adapt to the characteristics of two heterogeneous data types, leading to low fusion efficiency and weak anti-interference capabilities. To address these shortcomings, there is an urgent need in this field for a technical solution that offers low wearing burden, good fusion results, and enhanced pose estimation accuracy and robustness.

[0033] To address the aforementioned technical challenges, this application proposes a 3D human pose estimation method based on the fusion of millimeter-wave radar and a single inertial sensor (MmWave-Inertial Mutual Attention Correction Network, MIMAC-Net). This invention aims to overcome point cloud sparsity and spatial ambiguity by leveraging the physical prior of a single pelvic inertial measurement unit (IMU) with minimal hardware invasiveness. Simultaneously, it designs a spatial alignment network and a gated cross-attention mechanism to achieve efficient and deep fusion of heterogeneous modalities. Furthermore, it combines Doppler enhancement with the efficient temporal modeling capabilities of state-space models (such as Mamba) to accurately capture high-frequency motion details, achieving high-precision, robust, and long-term non-intrusive 3D human pose estimation. This approach combines the advantages of multiple sensor data types, effectively improving the accuracy and anti-interference capability of pose estimation.

[0034] See Figure 1 This is a flowchart illustrating a human pose estimation method based on multi-source data provided in an embodiment of this application. It is intended as an example and not a limitation. The method may include the following steps: S101 collects radar point cloud data and pelvic inertial data during human posture changes.

[0035] In this embodiment of the application, when the human body makes various movements and the posture changes continuously, millimeter-wave radar point cloud data and pelvic inertial data are collected simultaneously. The two types of data are time-synchronized and the frame sequences correspond one-to-one, serving as the original input for the subsequent attitude estimation algorithm.

[0036] For example, a single inertial measurement unit (IMU) is worn in the pelvic region of the human body, and paired with an external millimeter-wave radar to form a synchronous acquisition system. Before acquisition, the device clock is synchronized to ensure that the timing of the two types of data is aligned. During the entire process of the subject completing rehabilitation and daily movements and posture changes such as walking, squatting, lunging, and limb extension, the millimeter-wave radar continuously receives echo signals and generates radar point cloud data representing the spatial contour of the human body frame by frame. The IMU in the pelvis synchronously acquires inertial information such as three-axis acceleration, three-axis angular velocity, and posture rotation angle, and outputs pelvic inertial data for the corresponding frame, ultimately forming dual-channel raw data with perfect time and movement matching.

[0037] For example, when a subject completes a set of continuous lunge movements, from the initial step and forward lean to the recovery of posture, the radar collects point cloud data corresponding to each frame of the human body shape in real time. The inertial measurement unit simultaneously records the changes in pelvic angle deflection, motion acceleration, and angular velocity. Each frame of radar point cloud is accompanied by pelvic inertial data with the same timestamp, providing a complete and matching data source for subsequent viewpoint correction, feature fusion, and three-dimensional posture estimation.

[0038] S102, perform perspective correction on radar point cloud data based on pelvic inertial data to obtain perspective-aligned radar point cloud data.

[0039] In this embodiment, the attitude and orientation information provided by pelvic inertial data is used to correct the viewing angle deviation of radar point cloud caused by changes in human orientation. The radar point cloud under different viewing angles is uniformly converted to the standard coordinate system, and finally radar point cloud data with regular viewing angle and consistent spatial reference is obtained, thus eliminating the problem of directional ambiguity.

[0040] In one embodiment, see Figure 2 This is a schematic diagram of a process for correcting radar point cloud data according to an embodiment of this application, as shown below. Figure 2 As shown, step S102 includes: S201, extract the temporal features from the pelvic inertial data to obtain the motion temporal features.

[0041] In this embodiment, temporal feature mining is performed on the synchronously collected pelvic inertial data. Based on the temporal model, the continuously changing motion patterns, posture trends and dynamic information in the data are extracted, and finally, the motion temporal features that can characterize the human motion state are obtained, providing effective feature support for subsequent viewpoint correction and feature fusion.

[0042] For example, a gated recurrent unit (GRU) is used to extract temporal features from pelvic inertial data. First, a single IMU data point from frame t is given. By iteratively calculating the hidden state from the previous moment, the continuous motion change trend between frames is extracted, and the temporal hidden features of the current frame are output. The calculation formula is:

[0043] in, Representing the hidden state output in the previous moment, GRU relies on its internal control structure to memorize and transmit historical motion information, capturing the temporal correlation of inertial data.

[0044] S202, construct the affine transformation matrix based on the motion time sequence characteristics.

[0045] In this embodiment, the rotation parameters required for radar point cloud are calculated using the pelvic motion time sequence features extracted above, and an affine transformation matrix is ​​further constructed. This matrix is ​​the core calculation basis for realizing radar point cloud view correction.

[0046] In one embodiment, see Figure 3 This is a schematic diagram of the process for constructing an affine transformation matrix according to an embodiment of this application, as shown below. Figure 3As shown, step S202 includes: S301, obtain the rotation vector corresponding to the human body rotation angle based on the motion timing characteristics.

[0047] In this embodiment, based on the motion timing features extracted from pelvic inertial data, a rotation vector representing the change in human orientation is solved through network computation. This vector is used to represent the rotation relationship of the human body relative to the radar coordinate system in the current frame and is the core parameter for the subsequent construction of the affine transformation matrix.

[0048] For example, the motion timing features output by the gated loop unit The regression yields a rotation vector containing sine and cosine components of the angle. The calculation formula is:

[0049] In the formula, FC is a fully connected layer. , These correspond to the original components of the target rotation angle. They also correspond to the original cosine and sine values ​​of the human body rotation angle, together forming a rotation vector describing the human body's posture deflection. To avoid the problem of angle periodicity, this scheme does not directly predict the angle value, but instead indirectly represents the rotation relationship through this vector.

[0050] S302, normalize the rotation vector to obtain the normalized rotation vector.

[0051] In this embodiment of the application, the rotation vector obtained above is... Normalization is performed to eliminate the influence of amplitude, resulting in a standard rotation vector that conforms to the unit vector constraint, thereby ensuring that subsequent angle calculations and matrix construction satisfy geometric rules.

[0052] For example, let the original rotation vector be... The normalization calculation is performed according to the vector magnitude, as shown in the following formula:

[0053] S303, constructing an affine transformation matrix based on the normalized rotation vector.

[0054] In this embodiment, an affine transformation matrix is ​​constructed using the sine and cosine values ​​of the rotation angle contained in the normalized rotation vector, according to the two-dimensional rotation transformation rules. This matrix is ​​used to realize the viewpoint rotation transformation of the radar point cloud. :

[0055] It contains only rotation components and does not introduce scaling or translation, describing the transformation of a radar feature map from an arbitrary orientation to the normal coordinate system's positive viewpoint.

[0056] In the above method, by extracting rotation vectors from motion temporal features and constructing affine transformation matrices after normalization, the perspective changes caused by human motion can be accurately and stably described. This avoids the numerical deviations or transformation anomalies that are prone to occur when directly using the original rotation vectors to construct the matrix, ensuring the accuracy and reliability of subsequent radar point cloud perspective correction and providing a more robust spatial alignment basis for subsequent attitude estimation.

[0057] S203, perform viewpoint correction on the radar point cloud data according to the affine transformation matrix to obtain viewpoint aligned radar point cloud data.

[0058] In this application, the affine transformation matrix constructed above is used to perform spatial rotation transformation and resampling operations on the original radar point cloud feature map, correcting the viewing angle deviation caused by the change of human body orientation, unifying all frames of radar data into the standard coordinate system, and finally obtaining radar point cloud data with consistent viewing angle reference and elimination of directional ambiguity.

[0059] In the above method, by extracting motion temporal features from pelvic inertial data to construct an affine transformation matrix, and then performing viewpoint correction on the radar point cloud, the radar observation viewpoint deviation caused by human movement and turning is effectively eliminated, so that the radar point cloud and human motion posture are accurately aligned in space, providing more accurate and low-distortion input data for subsequent feature fusion and attitude estimation, and significantly improving the robustness and accuracy of subsequent attitude estimation.

[0060] In another embodiment, see Figure 4 This is a flowchart illustrating the correction process for radar point cloud data provided in another embodiment of this application, such as... Figure 4 As shown, step S203 includes: S401 generates a standardized sampling grid based on the affine transformation matrix.

[0061] In this embodiment of the application, based on the constructed affine transformation matrix, a regular standard sampling grid is generated according to the coordinate transformation rules. This grid is the coordinate reference for subsequent resampling of radar feature maps and completion of viewpoint correction.

[0062] For example, the affine transformation matrix This is applied to perform grid sampling operations on radar feature maps. First, based on... Generate standard sampling grid For each target pixel in the grid The sampling position of the radar feature map is calculated by affine transformation. The calculation process is as follows:

[0063] S402, perform an affine transformation on the radar point cloud data according to the standardized sampling grid to obtain the affine transformed radar point cloud data.

[0064] In this embodiment, following the coordinate mapping rules defined by the standardized sampling grid, each coordinate position of the original radar point cloud is mapped one by one to the corresponding standard coordinate position of the grid, completing the overall coordinate rotation transformation. The entire radar point cloud data undergoes a viewpoint transformation synchronously with the coordinates, ultimately yielding affine transformed radar point cloud data with corrected viewpoint.

[0065] For example, when the subject turns to the right, the original radar point cloud is shifted to the right side of the screen. According to the coordinate mapping relationship of the standardized sampling grid, the offset point cloud coordinates are transformed one by one to the standard positive position. The entire radar point cloud completes the viewpoint rotation, resulting in radar point cloud data with a regular viewpoint.

[0066] S403 performs a spatial transformation on the affine-transformed radar point cloud data to obtain viewpoint-aligned radar point cloud data.

[0067] In this embodiment, the radar point cloud data after affine transformation is further spatially normalized to correct the subtle deviations generated during the transformation process, and finally obtain radar point cloud data with completely unified perspective and standardized spatial reference, thus completely eliminating the spatial ambiguity caused by human orientation.

[0068] For example, using a bilinear interpolation kernel For the original radar feature map Perform spatial transformation to obtain the aligned radar feature map. :

[0069] Where N represents the neighboring pixels around the sampling point, Represents the number of pixels in any channel c. .

[0070] In the above method, a standardized sampling grid is generated by an affine transformation matrix, and affine transformation and spatial alignment correction are performed on the radar point cloud data accordingly. This achieves accurate spatial alignment between the radar point cloud and the human motion posture, effectively eliminating feature deviations caused by viewpoint shift and motion distortion, ensuring the quality of input data for subsequent feature fusion and attitude estimation, and providing a low-distortion, high-consistency radar data foundation for subsequent steps.

[0071] S103, the radar point cloud data and pelvic inertial data after viewpoint alignment are fused by feature interaction to obtain fused features.

[0072] In this embodiment, the radar features after viewpoint alignment and the pelvic inertial features are fused across modes. By leveraging the complementary effects between the features, the effective information of the two types of data is integrated, and finally, a fused feature containing both spatial morphology and motion dynamic information is generated.

[0073] In one embodiment, see Figure 5 This is a schematic diagram of the feature interaction fusion process provided in an embodiment of this application, such as... Figure 5 As shown, step S103 includes: S501, the radar point cloud data after viewpoint alignment is feature encoded to obtain the first feature data.

[0074] In this embodiment, the radar point cloud data with the viewpoint alignment completed is feature encoded, and its core information such as spatial geometry and contour is extracted by a dedicated encoder and converted into a high-dimensional feature vector of the same dimension to obtain the first feature data corresponding to the radar.

[0075] For example, a radar spatial encoder is used to perform multi-layer convolution operations on the viewpoint-aligned radar feature map (i.e., viewpoint-aligned radar point cloud data). Perform feature extraction and dimension mapping to output radar feature vectors. This process compresses the spatial structure, intensity, and velocity information of sparse radar point clouds into regular high-dimensional features, adapting them for subsequent cross-modal fusion. For example, when a subject completes a squatting motion, the viewpoint-aligned radar point cloud fully presents the contour changes of the human body during the squatting process. This set of radar data is fed into a spatial encoder, where multiple layers of convolution refine the spatial features such as human shape and limb distribution, ultimately generating first feature data with unified dimensions. This accurately represents the current spatial form of the human body, preparing it for subsequent interactive fusion with inertial features.

[0076] S502, perform feature encoding on the pelvic inertial data to obtain the second feature data.

[0077] In this embodiment of the application, feature encoding is performed on the pelvic inertial data to mine temporal information such as motion trends and posture dynamics in the data, and it is transformed into a standardized feature vector to obtain the second feature data corresponding to the inertial data.

[0078] For example, a timing encoder is used to process pelvic inertial data. Perform feature extraction and dimensionality transformation to output second feature data. For example, when a subject walks, pelvic inertial data fluctuates continuously with each step. This data is input into a time encoder to extract dynamic information such as step frequency, pelvic swing amplitude, and rotation trend, generating second feature data for subsequent fusion calculations with radar features.

[0079] S503, perform cross-attention fusion processing on the first feature data and the second feature data to obtain cross-attention features.

[0080] In this embodiment of the application, in order to achieve complementarity of heterogeneous features and establish semantic associations between modalities, the present invention adopts a cross-attention mechanism, with the number of attention heads set to 4. As the query vector Q, IMU features Using key vector K and value vector V, we retrieve the inertial motion states most relevant to the current radar observation, thus obtaining attention features enhanced by inertial information. :

[0081] in, It is a learnable projection matrix.

[0082] S504 performs adaptive weighting on the cross-attention features to obtain the fused features.

[0083] In this embodiment, a learnable gating coefficient is introduced to adaptively adjust the injection strength of IMU information and prevent noise interference. The final fusion features Generated via gated residual join:

[0084] This gated residual structure gives the model the ability to "soft switch," preserving its spatial features when the radar signal is clear, and selectively fusing complementary information from the IMU when the radar features are interfered with by noise such as obstruction or multipath interference, thereby improving the stability of attitude estimation.

[0085] In the above method, effective features are extracted by encoding the two types of data respectively, intermodal correlation is established by cross attention, and feature proportion is optimized by adaptive weighting to achieve deep fusion of spatial and motion information, improve feature expression ability and anti-interference ability, and ensure pose estimation effect.

[0086] S104, perform human pose estimation based on fused features to obtain the target pose estimation of the human body.

[0087] In this embodiment of the application, the fusion feature that integrates spatial and motion information is used to calculate the posture parameters of various parts of the human body through a posture calculation model, and finally outputs an accurate estimated posture of the human target.

[0088] In the above method, pelvic inertial data is used to correct the radar point cloud view deviation, and then feature interaction fusion is used to achieve deep integration of multi-source information, which can reduce environmental interference and significantly improve the accuracy and reliability of human pose estimation results.

[0089] In one embodiment, see Figure 6 This is a schematic diagram of the target pose estimation process provided in an embodiment of this application, as shown below. Figure 6 As shown, step S104 includes: S601, extract time-series features from the fused features to obtain global time-series features.

[0090] In this embodiment, multi-frame fusion features are input into the Mamba state space model to mine the temporal correlation and dynamic evolution of actions between consecutive frames, and to extract global temporal features that carry overall motion information.

[0091] For example, multimodal features after cross-attention fusion The data is fed into the Mamba state-space model, where temporal modeling is performed with linear complexity. Key pose cues are efficiently filtered using discretized recursive equations, and the output is a latent state containing rich context (i.e., global temporal features).

[0092] S602, perform human pose estimation based on the last frame of global temporal features to obtain the initial estimated human pose.

[0093] In this embodiment, the feature data of the last frame in the global temporal features is extracted and sent to the pose estimation module for calculation to preliminarily calculate the initial estimated pose of the human body.

[0094] For example, take the state of the last frame. This serves as the global pose representation at the current moment. Next, the global pose features extracted by Mamba are... Input a fully connected layer to predict the initial coarse skeleton of the human body. (i.e., initial attitude estimation):

[0095] This allows for the rapid determination of the overall anatomical structure of the human body, which includes J joints.

[0096] For example, after a subject completes a series of turning movements, the features of the last frame of the time sequence are extracted. The model combines the spatial and motion information at the end of the movement to calculate the position of the human torso and limbs at this time, and obtain the corresponding initial estimated posture.

[0097] S603, correct the initial estimated attitude to obtain the target estimated attitude.

[0098] In the embodiments of this application, the initial pose result is optimized by combining temporal and feature information, and deviations and errors are corrected to finally obtain an accurate and reliable human target pose estimation.

[0099] In the above method, feature encoding is performed on the radar point cloud data and pelvic inertial data after viewpoint alignment, and then fusion features are obtained through cross-attention fusion and adaptive weighting. This achieves deep interaction and complementarity between radar spatial information and inertial motion information. It not only captures the correlation and dependence between the two modes by utilizing cross-attention, but also dynamically adjusts the feature contribution through adaptive weighting, thereby generating more comprehensive and robust fusion features, providing more reliable feature support for subsequent human pose estimation.

[0100] In another embodiment, see Figure 7 This is a schematic diagram of the target pose estimation process provided in another embodiment of this application, such as... Figure 7 As shown, step S603 includes: S701 extracts Doppler channel data from radar point cloud data.

[0101] In this embodiment, the coarse skeleton often lacks dynamic details at the limb extremities, therefore fine regression is introduced, and a Doppler enhancement mechanism is employed at this stage. Specifically, a lightweight dual-pooling side-branch encoder is designed to explicitly mine Doppler features. From the raw radar input... Extracting Doppler channel data separately .

[0102] For example, a single radar point cloud data point contains multi-channel information such as spatial coordinates, echo intensity, and Doppler velocity. An extraction operator is set, the original radar point cloud set is input, all point cloud units are traversed and Doppler dimension values ​​are filtered to obtain Doppler channel data.

[0103] S702 performs pooling processing on the Doppler channel data to obtain the Doppler feature sequence.

[0104] In this embodiment, the extracted Doppler channel data is subjected to pooling dimensionality reduction and feature aggregation to compress the data volume while retaining the core motion information, thereby generating a time-series Doppler feature sequence.

[0105] For example, for Doppler channel data After processing by independent convolutional layers, a dual pooling strategy can be used to apply max pooling and average pooling operators in parallel to capture local limb extreme velocities and overall motion trends, respectively. These are then concatenated to generate a Doppler feature sequence. :

[0106] S703, the initial estimated attitude is corrected based on the Doppler feature sequence to obtain the target estimated attitude.

[0107] In this embodiment, by combining the motion velocity information reflected by the Doppler feature sequence, the error of the initial posture is corrected, the posture parameters are optimized, and finally an accurate human target posture estimation is obtained.

[0108] In the above method, Doppler channel data is extracted from radar point cloud and pooled to obtain Doppler feature sequence, which is then used to correct the initial estimated attitude. This effectively introduces dynamic information reflecting the local motion speed of the human body, making up for the detail deviation when estimating attitude based solely on temporal and spatial features. This achieves accurate correction of the initial attitude, and the final output target estimated attitude has higher accuracy and stronger robustness, especially with a significant reduction in error in dynamic motion scenarios.

[0109] In another embodiment, see Figure 8 This is a schematic diagram of the target pose estimation process provided in another embodiment of this application, such as... Figure 8 As shown, step S703 includes: S801 concatenates the last frame of global temporal features, the initial estimated attitude features, and the Doppler feature sequence to obtain concatenated features.

[0110] In this embodiment, features of three different dimensions are spliced ​​and fused together according to their dimensions, integrating temporal information, posture information and motion speed information to form a complete spliced ​​feature, providing comprehensive data support for subsequent posture correction.

[0111] For example, three types of effective features are extracted: global temporal features (last frame features), coarse pose features output from the encoded initial pose, and features extracted from the T-th frame of the Doppler feature sequence. These features are then concatenated along the feature channel dimension without destroying the original information of each part, forming a multi-source information fusion concatenated feature F, whose formula is: F=

[0112] S802, calculate the residual correction data for the human skeleton coordinates based on the splicing features.

[0113] In this embodiment of the application, based on the splicing features that integrate multiple types of information, the deviation values ​​of the coordinates of each joint of the human skeleton from the actual state are calculated to obtain residual correction data for correcting posture. For example, the splicing feature F is input into a fine regression network to predict the coordinate residual correction amount (i.e., residual correction data). :

[0114] S803, the initial estimated attitude is superimposed with the residual correction data to obtain the target estimated attitude.

[0115] In this embodiment, the coordinate residual correction amount calculated by superimposing the initial attitude is used to compensate for the estimation deviation, ultimately obtaining a more accurate human target estimated attitude. The calculation formula is:

[0116] For example, if there is a slight deviation in the elbow coordinates in the initial posture, the coordinates are fine-tuned after the corresponding residual correction value is superimposed, the error is corrected, and the final human posture that fits the actual state is output.

[0117] In the above method, a comprehensive representation is constructed by splicing time sequence, attitude and Doppler multi-class features. Based on this, the skeleton coordinate residual is solved and superimposed on the initial attitude to complete the correction. By fully combining the global motion law, initial attitude information and local velocity features, the estimation deviation is accurately compensated, and the accuracy and reliability of the final attitude result are greatly improved.

[0118] See Figure 9 This is a schematic diagram of the overall architecture of a human pose estimation method based on multi-source data provided in an embodiment of this application, as shown below. Figure 9 As shown, the whole is divided into four core modules: 1. Input module: Radar point cloud, IMU (accelerometer + gyroscope) data; 2. IMU-guided spatial alignment network: performs viewpoint correction on radar point clouds and outputs aligned radar data; 3. Gated cross-attention fusion module: Aligns radar features and IMU features for cross-modal interactive fusion; 4. Mamba cascaded network: It is divided into two branches: coarse estimation and fine estimation, and outputs the final human target pose.

[0119] The attitude estimation steps based on the above structure include: Step 1: Input Data Collection The system receives two types of raw data: Radar point cloud: a sparse set of points that records the spatial position and reflection intensity of a human body; IMU data includes acceleration and gyroscope signals, reflecting the pelvic movement status.

[0120] Step 2: IMU-guided radar spatial alignment Pelvic IMU data is fed into GRU (Gated Cyclic Unit) + MLP (Multilayer Perceptron) to estimate affine transformation parameters. A standardized sampling grid is constructed based on this parameter; The original radar point cloud is subjected to grid sampling and affine transformation to obtain a viewpoint-aligned radar point cloud, thus eliminating viewpoint deviation.

[0121] Step 3: Gated Cross-Attention Feature Fusion Feature encoding: After alignment, the radar point cloud is fed into the radar encoder to extract spatial features. That is, the first feature data; IMU data is fed into the IMU encoder (GRU+MLP) to extract motion features; Cross-attention interaction: Using radar features as the query (Q) and IMU features as the key (K) and value (V), attention weights are calculated using Softmax to generate interaction features. That is, cross-attention features; Gated fusion: generating weights using the Sigmoid gating mechanism ( The interactive features are modulated, and then residually connected with the radar features and subjected to layer normalization (LN) to obtain the final fused features:

[0122] Step 4: Mamba Cascaded Network (Coarse Estimation + Fine Estimation) (1) Coarse estimation branch The fused features are fed into the Mamba state-space model to model long-term temporal motion patterns; The output is fed into the CoarseHead (coarse estimation head) to generate the initial estimated pose.

[0123] (2) Precise estimation branch After alignment, the radar data is sent to a dual Doppler pooling module to extract motion velocity features; The Doppler features, initial pose coding features and Mamba temporal features are concatenated and then fed into FineHead. Calculate the skeleton coordinate residual correction amount, superimpose it on the initial attitude, and obtain the final target estimated attitude.

[0124] See Figure 10 This is a schematic diagram of the overall architecture of attitude fine estimation provided in an embodiment of this application, as shown below. Figure 10 As shown, it includes a coarse estimation branch and a fine estimation branch. The coarse estimation branch generates an initial human skeleton (i.e., the initial estimated attitude coarse skeleton) based on the radar-inertial measurement unit (IMU) fusion features and the Mamba state space model. The fine estimation branch extracts motion features based on the aligned radar features and performs residual correction on the coarse skeleton, finally outputting a high-precision fine skeleton (i.e., the target estimated attitude).

[0125] 1) Coarse estimation branch: Generate the initial human skeleton Input features: Received Radar-IMU fusion features (a time-series feature sequence that integrates radar spatial information and IMU motion information).

[0126] Mamba time series modeling: First, the fused features are standardized using LayerNorm (layer normalization); Divided into two parallel paths: Path 1: Linear layer (fully connected layer) directly maps features; Path 2: Linear layer → Conv1D (one-dimensional convolution) → SelectiveSSM (selective state-space model), capturing temporal dependencies; The outputs of the two paths are first connected via residuals (⊕), and then multiplied and fused (…). ), and send it to the Linear layer; After another residual connection, the output characteristics of the Mamba module are obtained.

[0127] Coarse skeleton generation: Mamba output features are fed into CoarseHead (coarse estimation head), regression is performed to obtain the initial human joint coordinates, and the coarse skeleton (initial estimated pose) is output.

[0128] (ii) Precise estimation branch: Correction yields a high-precision skeleton. Input features: Receive aligned radar features (radar data with view correction completed, including Doppler channel information).

[0129] Dual Doppler pooling: First, two layers of Conv2D (two-dimensional convolution) are used to extract features of the radar in both spatial and channel dimensions. It is further divided into two parallel pooling paths: Path 1: AvgPool (average pooling) extracts the overall mean information of motion features; Path 2: MaxPool (maximum pooling), extracting local extremum information of motion features; The two pooling results are additively fused (⊕) to obtain Doppler features that integrate global and local motion information.

[0130] Feature stitching and residual correction: The Doppler features are concatenated with the Mamba output features of the coarse estimation branch (⊕) and then fed into the FineHead (fine estimation head). FineHead regression yields the residual corrections between the rough skeleton and the actual pose; The residual correction is superimposed on the coarse skeleton to obtain the final fine skeleton (target estimated pose).

[0131] To achieve high-precision 3D pose reconstruction and ensure that the generated skeleton conforms to human kinematic constraints, this invention designs a multi-task joint loss function, which consists of three parts: fine pose regression loss, coarse pose loss, and bone length consistency loss.

[0132] in, and The predicted 3D key points and the actual annotations were measured separately in the fine regression stage and the coarse regression stage. Error between:

[0133] The bone length consistency loss is defined as follows: it penalizes the difference between the predicted bone length and the actual bone length to prevent non-physical deformations such as limb stretching and contraction in the predicted skeleton.

[0134] in, This represents a predefined set of human skeleton connections. Represents the index of two connected joints. This represents the Euclidean distance.

[0135] To comprehensively evaluate the effectiveness and robustness of the proposed MIMAC-Net method for 3D pose estimation of complex movements, this invention conducted extensive experiments on the publicly available mRI dataset. mRI is currently the only large-scale human pose estimation dataset that simultaneously integrates three modalities, including a commercial millimeter-wave radar (TI IWR1443 Boost), six wearable IMUs (WitMotion BWT901CL), and two RGB-D cameras (Microsoft KinectV2) used to generate the gold standard. This dataset collected over 160,000 frames of synchronized data from 12 complex rehabilitation training movements (such as limb extension, lunge, squat, and walking) from 20 subjects. This invention uses only millimeter-wave radar point clouds and single pelvic IMU data as input to verify the multimodal fusion potential of the lightweight MIMAC-Net model.

[0136] This application employs a 5-fold cross-validation strategy, strictly adhering to the two data partitioning methods defined by mRI: S1 (random partitioning) and S2 (split across subjects). In S1, all subject sample frames are randomly shuffled and partitioned into training, validation, and test sets in a 7:1:2 ratio. S2 is more challenging, randomly selecting data from 80% of the subjects (16 individuals) for training, with 14 individuals serving as the training set, 2 as the validation set, and the remaining 20% ​​(4 individuals) as the test set.

[0137] To verify the key gain of introducing a single IMU as an auxiliary mode for millimeter-wave radar attitude estimation and to explore the performance of multimodal sensor fusion, this invention designed a set of modal comparison experiments. In the experiments, the performance differences between the complete MIMAC-Net model (millimeter-wave + single IMU) and the pure radar variant model were compared. The pure radar variant model is an ablation model of MIMAC-Net, retaining its backbone structure while removing IMU-related modules and components, such as the IMU-SAN module and IMU feature encoder. The cross-modal cross-attention module is replaced with a millimeter-wave radar self-attention module, relying solely on radar point cloud features for attitude regression.

[0138] See Figure 11 This is a schematic diagram illustrating the ablation comparison results of different modal inputs provided in the embodiments of this application, such as... Figure 11 As shown, the multimodal MIMAC-Net, with the introduction of a single IMU, exhibits significant performance advantages. Specifically, compared to the baseline model relying solely on radar data, MIMAC-Net's MPJPE drops dramatically from 119.55 mm to 96.17 mm, a reduction of 19.6%; the average position error per joint (PA-MPJPE) after Protodyakonov analysis alignment also decreases by 17.02 mm, representing a performance improvement of 16.5%. This significant performance gain powerfully demonstrates the crucial value of a single IMU in multimodal fusion attitude reconstruction, revealing the deep complementarity between radar and inertial sensors at the physical level. This low-cost modal enhancement strategy, incorporating a single IMU, successfully addresses the robustness limitations of pure radar solutions under varying viewpoints and dynamic blurring with minimal hardware cost, significantly improving the robustness and accuracy of attitude estimation while maintaining a lightweight system.

[0139] To systematically verify the necessity and effectiveness of each core component in the MIMAC-Net framework, this invention designed a series of ablation experiments. Using the full model that performs best under the cross-subject protocol (S2) as a baseline, the following five key components were removed or replaced one by one using the "controlled variable method": IMU-guided spatial alignment network, gated cross-attention fusion module, Doppler information flow, coarse-to-fine cascaded network, and Mamba model. All experiments were conducted under the same training strategy and hyperparameter settings, and employed a 5-fold cross-validation method.

[0140] See Figure 12 This is a schematic diagram of the ablation experiment results of the MIMAC-Net core components provided in the embodiments of this application, as shown below. Figure 12As shown, the complete MIMAC-Net outperforms all variants, demonstrating the unique value and indispensability of each component in the existing architecture. First, when the IMU-guided spatial alignment network is removed and replaced with self-supervised alignment relying on radar features, the MPJPE increases by 3.28 mm, indicating that sparse radar point clouds struggle to infer accurate global rotation. Introducing absolute orientation information from a single pelvic IMU as a physical constraint effectively eliminates spatial ambiguity in radar and enhances the geometric consistency of features. Second, when the gated cross-attention module is removed and replaced with simple feature stitching, performance degrades significantly, with the MPJPE increasing by 6.5 mm. The results show that simple linear stitching cannot effectively handle information differences between heterogeneous modalities. The cross-attention mechanism dynamically filters the most complementary motion features in the IMU sequence, and the gated network adaptively adjusts weights based on signal quality, significantly improving the robustness of pose reconstruction.

[0141] Furthermore, removing the explicit double-pooling Doppler branch and the Doppler feature injection of the cascade head resulted in a 4.91mm decrease in model performance, demonstrating that explicit Doppler features help capture details of rapidly micro-moving limb extremities and correct motion blur in dynamic movements. Next, removing the coarse-to-fine residual cascade structure and degenerating into a single-stage MLP regression network increased MPJPE by 5.65mm, validating that the residual cascade strategy decomposes the complex 3D pose regression task into progressively refined subtasks, reducing optimization difficulty and enhancing the accuracy and stability of prediction results. Finally, removing the Mamba module caused the largest performance degradation, with MPJPE surging by 10.73mm. This significant difference strongly demonstrates the superiority of the Mamba architecture; its selective state-space mechanism can more efficiently capture long-range spatiotemporal dependencies, ensuring the coherence of actions and the integrity of pose. In conclusion, the superior performance of MIMAC-Net stems from the organic collaboration of its various components, which together constitute a robust and efficient pose estimation system.

[0142] To comprehensively evaluate the performance and generalization ability of MIMAC-Net under standard evaluation settings, this invention conducts a rigorous quantitative comparison with the best results reported in the original paper on the mRI dataset. The mRI baseline method uses pure millimeter-wave radar as input and employs a CNN architecture for pose estimation. This invention strictly follows the experimental settings of the mRI paper, conducting tests under both randomized partitioning (S1) and cross-subject partitioning (S2), and reporting performance metrics for Scheme 1 (12 types of actions) and Scheme 2 (only 10 standard actions). The last two actions in Scheme 1 are free-form stretching and relaxation, and walking along a straight line. These two actions are designed to increase the diversity of the dataset, with the 11th action determined by each subject and the 12th action involving overall displacement.

[0143] See Figure 13 This is a schematic diagram showing the comparison results between the two evaluation protocols provided in the embodiments of this application and the mRI baseline method, as shown below. Figure 13 As shown, MIMAC-Net significantly outperforms baseline methods across all evaluation protocols and partitioning settings, demonstrating a substantial performance advantage. Specifically, in the most challenging S2 scenario, for Scheme 1, MIMAC-Net drastically reduces the MPJPE from 186.6 mm to 96.2 mm, a performance improvement of 48.5%; for Scheme 2, the MPJPE reaches an astonishing 76.28 mm. This significant performance improvement across scenarios is primarily attributed to the IMU-guided spatial alignment network and the Doppler-enhanced Mamba cascade network introduced in this invention. The former effectively addresses the challenges of perspective blurring and generalization caused by subject differences in the S2 protocol, while the latter significantly enhances the ability to capture long sequences and dynamic actions.

[0144] Of particular note is that the advantages of the MIMAC-Net of this invention are more pronounced under random partitioning (S1). For Scheme 1, the MPJPE is significantly reduced from 163.3 mm to 48.5 mm, a reduction of 70.3%, while in Scheme 2, an extremely low error of 28.71 mm is achieved. This indicates that when the training data is sufficiently rich and covers the body characteristics of the subjects, MIMAC-Net can approach the physical resolution limit of millimeter-wave radar, demonstrating great performance potential and application value.

[0145] To evaluate the human pose estimation performance of the proposed MIMAC-Net method, this invention compares it with five current state-of-the-art methods on the same mRI dataset, including mmHPE, mmBaT, GF-DecNet, PoseGraphNet, and CGAN, covering a variety of mainstream deep learning techniques such as CNN, RNN, GCN, Transformer, and Generative Adversarial Network (GAN).

[0146] See Figure 14 This is a schematic diagram showing the comparison results between the mRI dataset provided in this application and existing methods, as shown in the embodiment. Figure 14As shown, MIMAC-Net achieves the best pose reconstruction accuracy, with MPJPE and PA-MPJPE reaching 96.17mm and 86.42mm respectively, significantly outperforming all compared methods. The second-best method, mmHPE, although also employing a coarse-to-fine cascaded network, lacks an IMU-guided spatial alignment network and cross-attention mechanism, resulting in an increase of 5.89mm in MPJPE. Compared to the Transformer-based mmBaT method, MIMAC-Net significantly reduces MPJPE by 12.23mm. This demonstrates that when processing long-sequence sparse radar signals, MIMAC-Net's Mamba not only possesses stronger modeling capabilities than Transformer, but its linear complexity also results in a significantly lower parameter / computational cost (2.46M / 2.29G) compared to mmBaT (11.68M / 5.07G).

[0147] Furthermore, while both GF-DecNet and PoseGraphNet employ graph convolutional networks, they lack explicit Doppler feature injection, failing to capture dynamic limb end-effector details. PoseGraphNet's MPJPE increases by 12.97 mm, and GF-DecNet's MPJPE even reaches a worst of 153.10 mm, a significant increase of 56.93 mm, with its parameter count almost 50 times that of MIMAC-Net. Although CGAN boasts extremely low parameter / computational cost (0.02M / 0.01G), its overly simple generative adversarial network struggles to capture the complex dynamic spatiotemporal features of sparse point clouds, resulting in an MPJPE as high as 127.09 mm. In conclusion, MIMAC-Net achieves the optimal balance between performance and efficiency, breaking through the pose estimation accuracy bottleneck of the contrasting methods while maintaining a lightweight architecture.

[0148] In addition to its significant advantages in quantitative metrics, to more intuitively evaluate the reconstruction quality, a visual comparison was made between the 3D human poses generated by MIMAC-Net and contrasting methods and the ground truth. To verify the reconstruction effect on more challenging and complex movements, frames containing complex motion patterns such as limb crossing and rapid squatting were specifically selected; see [link to relevant documentation]. Figure 15 This is a visual comparison diagram of different methods provided in the embodiments of this application, such as... Figure 15 As shown, the comparison of skeleton reconstruction results of different methods under five typical movements (right forward lunge, squat, left lunge, left leg raise, and double arm raise) is presented. It can be seen that the skeleton reconstructed by MIMAC-Net (red) is closest to the actual value (blue) in all movement types. Especially in large-amplitude and self-occluded movements such as squats and lunges, pure radar comparison methods often fail to accurately capture the knee flexion angle and limb orientation, resulting in stiff or misaligned leg postures, while MIMAC-Net ensures a high degree of geometric consistency due to its spatial alignment network.

[0149] Furthermore, in rapid extremity movements such as raising the left leg and raising both arms, where the extremities are far from the body center and radar reflection points are sparse, methods like PoseGraphNet, GF-DecNet, and CGAN exhibit significant extremity localization drift or loss (e.g., insufficient leg raising height). In contrast, MIMAC-Net, with its Doppler-enhanced cascaded network, successfully captures the instantaneous velocity and limb details of the wrist and ankle, thus reconstructing a more natural and kinetically accurate skeletal pose. The visualized qualitative results are highly consistent with the quantitative data in Table 4, further validating MIMAC-Net's core advantage in handling complex dynamic movements.

[0150] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0151] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0152] Figure 16 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. For example... Figure 16 As shown, the terminal device 16 of this embodiment includes: at least one processor 160 ( Figure 16 (Only one is shown in the diagram) a processor, a memory 161, and a computer program 162 stored in the memory 161 and executable on at least one processor 160. When the processor 160 executes the computer program 162, it implements the steps in any of the above embodiments of the human pose estimation method based on multi-source data.

[0153] The terminal device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. This terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 16 This is merely an example of terminal device 16 and does not constitute a limitation on terminal device 16. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0154] The processor 160 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0155] In some embodiments, memory 161 may be an internal storage unit of terminal device 16, such as a hard disk or memory of terminal device 16. In other embodiments, memory 161 may be an external storage device of terminal device 16, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on terminal device 16. Furthermore, memory 161 may include both internal and external storage units of terminal device 16. Memory 161 is used to store operating system, applications, boot loader, data, and other programs, such as program code for computer programs. Memory 161 can also be used to temporarily store data that has been output or will be output.

[0156] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0157] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0160] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0161] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A human pose estimation method based on multi-source data, characterized in that, The method includes: Collect radar point cloud data and pelvic inertial data during human posture changes; The radar point cloud data is corrected by viewing angle based on the pelvic inertial data to obtain radar point cloud data with viewing angle alignment. The radar point cloud data and the pelvic inertial data after viewpoint alignment are subjected to feature interaction fusion to obtain fused features; Human pose estimation is performed based on the fused features to obtain the target estimated pose of the human body.

2. The human pose estimation method based on multi-source data as described in claim 1, characterized in that, The step of performing viewpoint correction on the radar point cloud data based on the pelvic inertial data to obtain viewpoint-aligned radar point cloud data includes: Temporal features are extracted from the pelvic inertial data to obtain motion temporal features; Construct an affine transformation matrix based on the temporal characteristics of motion; The radar point cloud data is subjected to viewpoint correction based on the affine transformation matrix to obtain viewpoint-aligned radar point cloud data.

3. The human pose estimation method based on multi-source data as described in claim 2, characterized in that, The construction of the affine transformation matrix based on the motion temporal characteristics includes: The rotation vector corresponding to the human body rotation angle is obtained based on the motion timing characteristics. The rotation vector is normalized to obtain a normalized rotation vector; The affine transformation matrix is ​​constructed based on the normalized rotation vector.

4. The human pose estimation method based on multi-source data as described in claim 2, characterized in that, The step of performing viewpoint correction on the radar point cloud data based on the affine transformation matrix to obtain viewpoint-aligned radar point cloud data includes: A standardized sampling grid is generated based on the affine transformation matrix; The radar point cloud data is subjected to an affine transformation based on the standardized sampling grid to obtain the affine transformed radar point cloud data. A spatial transformation is performed on the radar point cloud data after affine transformation to obtain the radar point cloud data after viewpoint alignment.

5. The human pose estimation method based on multi-source data as described in claim 1, characterized in that, The radar point cloud data and pelvic inertial data, after being aligned with the viewpoint, are subjected to feature interaction fusion to obtain fused features, including: The radar point cloud data after viewpoint alignment is feature-encoded to obtain the first feature data; The pelvic inertial data is feature-encoded to obtain second feature data; The first feature data and the second feature data are subjected to cross-attention fusion processing to obtain cross-attention features; The cross-attention features are adaptively weighted to obtain the fused features.

6. The human pose estimation method based on multi-source data as described in claim 1, characterized in that, The step of estimating human pose based on the fused features to obtain the target estimated pose of the human body includes: Temporal feature extraction is performed on the fused features to obtain global temporal features; Human pose estimation is performed based on the last frame of the global temporal features to obtain the initial estimated human pose. The initial estimated attitude is corrected to obtain the target estimated attitude.

7. The human pose estimation method based on multi-source data as described in claim 6, characterized in that, The step of correcting the initial estimated pose to obtain the target estimated pose includes: Extract Doppler channel data from the radar point cloud data; The Doppler channel data is pooled to obtain a Doppler feature sequence; The initial estimated pose is corrected based on the Doppler feature sequence to obtain the target estimated pose.

8. The human pose estimation method based on multi-source data as described in claim 7, characterized in that, The step of correcting the initial estimated pose based on the Doppler feature sequence to obtain the target estimated pose includes: The last frame of the global temporal features, the initial estimated pose features, and the Doppler feature sequence are concatenated to obtain the concatenated features. The residual correction data for the human skeleton coordinates are calculated based on the splicing features. The initial estimated attitude is superimposed with the residual correction data to obtain the target estimated attitude.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.