A spatio-temporal decoupled lightweight method and system for continuous human pose estimation of WiFi

CN122839331APending Publication Date: 2026-09-29KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611017144.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]现有WiFi人体姿态估计方法仍存在多方面明显缺陷:其一,连续姿态建模能力不足

Benefits of technology

时空显式解耦,保护时序因果结构。本发明摒弃传统二维卷积时空强耦合的处理方式,采用“先时序建模、后空间提取”的解耦策略,通过因果时序卷积严格保留信号的时序因果属性,更契合连续动作的动态演变规律,有效提升了连续姿态估计的精度与运动平滑性,减少逐帧预测抖动。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839331A_ABST
    Figure CN122839331A_ABST
Patent Text Reader

Abstract

This invention discloses a spatiotemporally decoupled, lightweight WiFi continuous human pose estimation method and system, belonging to the field of WiFi wireless sensing and human pose estimation technology. To address the problems of strong spatiotemporal feature coupling in existing WiFi pose estimation methods, which disrupts temporal causality, suffers from insufficient continuous motion modeling, high computational overhead, and lacks skeleton topological constraints, this invention adopts an encoder-decoder architecture: First, a grouped fusion temporal convolutional network is used to extract the temporal dynamic features of the CSI signal, strictly preserving the temporal causal attributes of the signal; then, an asymmetric convolutional network is used to extract the spatial frequency features between subcarriers; an axial self-attention mechanism is introduced to model the skeleton topological dependence of human key points; finally, a lightweight decoder combined with skeleton length constraint loss outputs continuous pose coordinates. This invention significantly improves the accuracy and smoothness of continuous pose estimation while greatly compressing the number of model parameters and floating-point computation, exhibiting excellent cross-scenario generalization performance, and can be widely applied to non-contact sensing scenarios such as smart healthcare and human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of WiFi wireless sensing and human pose estimation technology, specifically involving a spatiotemporally decoupled lightweight WiFi continuous human pose estimation method and system. Background Technology

[0002] Human pose estimation (HPE) aims to quantify the joint positions and skeletal morphology of the human body in three-dimensional space, and has significant application value in fields such as smart healthcare, human-computer interaction, and intelligent security. Current mainstream pose estimation solutions mostly rely on vision devices or wearable sensors: while vision-based solutions offer high accuracy, they are susceptible to lighting and occlusion, and pose privacy risks; wearable device solutions, while not having privacy issues, require users to wear invasive devices, limiting their application scenarios.

[0003] In contrast, sensing technologies based on WiFi Channel State Information (CSI) have gradually become a research hotspot in the field of human perception due to their advantages such as being contactless, low-cost, privacy-friendly, and unaffected by light occlusion. With the development of WiFi sensing technology, researchers have begun to extend it from coarse-grained tasks such as action recognition and breathing monitoring to fine-grained human pose estimation tasks, that is, solving the nonlinear mapping from high-dimensional CSI time-series signals to the coordinates of key points in human space.

[0004] Existing WiFi human pose estimation methods still suffer from several significant shortcomings: First, insufficient continuous pose modeling capabilities. Current research primarily evaluates discrete action categories, lacking systematic modeling of continuous action transition states. Frame-by-frame prediction results are prone to "jitter," making it difficult to adapt to the continuous and smooth motion of the human body in real-world scenarios. Second, strong spatiotemporal coupling disrupts the temporal causality of the signal. Most methods equate CSI data with two-dimensional static images, using standard two-dimensional convolution to simultaneously process the temporal and subcarrier dimensions. This indiscriminate mixing of two physically distinct dimensions disrupts the original signal's temporal causal structure, easily leading to dynamic information confusion and feature redundancy. Third, high model computational complexity hinders edge deployment. Existing solutions suffer from low serial computation efficiency in CNN-LSTM cascade architectures, while pure Transformer architectures involve computational overhead on the order of the quadratic sequence length. Models with large parameter counts and high floating-point computational demands cannot adapt to IoT edge devices with limited computing power. Fourth, lack of skeleton topological constraints. Most direct coordinate regression methods do not explicitly model the physical topological dependencies of human key points, and the output posture is prone to problems such as joint misalignment and abnormal bone length, which do not conform to the laws of human kinematics.

[0005] In summary, existing WiFi human pose estimation technologies cannot simultaneously meet the requirements of high accuracy, low computational overhead, and strong structural rationality in continuous scenarios, making it difficult to support large-scale deployment in real-world scenarios. Summary of the Invention

[0006] To achieve the above objectives, this invention adopts the following technical solution: Based on a spatiotemporally decoupled encoder-decoder architecture, a lightweight WiFi continuous human pose estimation method with spatiotemporal decoupling is proposed, comprising the following steps: Step S1, CSI data acquisition and preprocessing. Collect multi-link WiFi channel state information, extract signal amplitude features, and construct a standardized spatiotemporal input feature tensor through time alignment and sliding window processing.

[0007] Step S2: Extract temporal dynamic features using grouped fusion temporal convolution. A Grouped Fusion Temporal Convolutional Network (GF-TCN) is constructed. Grouped convolution reduces computational redundancy, and causal dilation convolution captures long-range temporal dependencies. Simultaneously, a channel fusion mechanism adaptively selects advantageous subcarriers, strictly preserving the temporal causal attributes of the signal throughout the entire process.

[0008] Step S3: Extract spatial frequency features using an asymmetric convolutional network. An asymmetric convolutional structure that slides only along the subcarrier dimension is employed to extract spatial correlations between subcarriers while maintaining a constant temporal dimension structure. Multi-level residual blocks progressively compress the subcarrier dimension, completing the mapping from frequency domain features to the semantic dimension of human key points.

[0009] Step S4: Axial self-attention modeling of skeleton topological constraints. An axial self-attention mechanism is introduced to decouple attention computation along two orthogonal directions: the first stage aggregates the temporal features of a single keypoint along the temporal dimension, and the second stage models the physical topological dependencies of multiple keypoints along the spatial dimension, achieving skeleton structure constraints with low computational overhead.

[0010] Step S5: The lightweight decoder regresses continuous pose coordinates. The high-dimensional encoded features are progressively mapped to keypoint spatial coordinates through a convolutional decoder, and the output is optimized by combining the bone length constraint loss, finally outputting continuous temporal human pose estimation results.

[0011] Further, step S1 is specifically implemented as follows: configure a WiFi acquisition system with multiple transmitting and receiving antennas to extract the effective subcarrier amplitude information of each communication link; stitch all links along the subcarrier dimension to form a panoramic amplitude feature; and generate a spatiotemporal input tensor with the dimension of subcarrier number × time step based on the set sampling rate and sliding window length, aligned with time.

[0012] Furthermore, the specific implementation of the grouped fusion temporal convolutional network in step S2 is as follows: the input features are uniformly divided into multiple independent subgroups along the subcarrier channel dimension, and each group independently performs one-dimensional temporal convolution. The number of parameters and computational cost of a single group convolution is only 1 / G of that of a standard one-dimensional convolution (G is the number of groups); causal dilation convolution is used within each group to expand the receptive field through an exponentially growing dilation factor, and the convolution operation only accesses current and historical information, strictly maintaining temporal causality; based on the grouped convolution, a progressive channel fusion mechanism is used to adaptively select advantageous subcarriers that are strongly correlated with human motion and filter out weakly correlated noise components.

[0013] Furthermore, the specific implementation of the asymmetric convolutional network in step S3 is as follows: an asymmetric convolutional kernel with a receptive field of 1×k is used, and one-dimensional sliding convolution is performed only in the subcarrier dimension while the temporal dimension remains unchanged; multi-level asymmetric residual blocks are constructed, and each residual block compresses the subcarrier dimension through downsampling convolution with a stride of (1,2) while expanding the number of feature channels; through multi-level mapping, the high-dimensional subcarrier features are gradually converted into semantic features corresponding to the number of human key points, with each output dimension corresponding to one human key point.

[0014] Furthermore, the axial self-attention mechanism in step S4 is specifically implemented as follows: The first stage is temporal dimension attention: the input features are reshaped into the form of "number of key points × number of channels × time step", and the temporal sequence of each key point is processed in parallel; query, key, and value matrices are generated through learnable projection, and temporal dependencies are calculated using grouped scaling dot product attention to highlight temporal features that contribute significantly to pose regression; the second stage is spatial node attention: the input features are reshaped into the form of "time step × number of channels × number of key points", and attention calculation is performed along the key point dimension to model the global topological dependency relationship between key points of the human skeleton; the two-stage calculation is cascaded and executed, and the final output is an encoded feature that simultaneously possesses temporal dynamic characteristics and spatial structural constraints.

[0015] Furthermore, the specific implementation of the decoder and loss function in step S5 is as follows: The decoder adopts a multi-layer convolutional structure: first, dimensionality reduction is achieved through 3×3 convolution, combined with batch normalization and activation functions to enhance nonlinear expression; then, 1×1 convolution is used to compress the channels to the coordinate dimension; finally, adaptive max pooling is used to aggregate temporal information and output keypoint coordinate tensors; the loss function uses the Smooth L1 norm as the main loss to measure the coordinate prediction error; an additional bone length constraint loss is introduced to penalize the deviation between the predicted bone length and the actual bone length, thereby enhancing the rationality of the human body structure in the output pose; the total loss is a weighted sum of the main loss and the bone constraint loss.

[0016] Furthermore, the method employs an allowable domain accuracy evaluation mechanism, using the percentage of correct key points (PCK) and the mean joint position error (MPJPE) as core evaluation indicators to comprehensively evaluate the attitude estimation accuracy and absolute positioning deviation under different error thresholds.

[0017] This invention also provides a spatiotemporally decoupled, lightweight WiFi continuous human posture system, comprising: The data preprocessing module is used to collect multi-link WiFi channel status information, extract amplitude features, perform time alignment and window slicing, and output a standardized spatiotemporal input tensor. The temporal feature extraction module, connected to the data preprocessing module, is used to extract the temporal dynamic features of CSI signals through a grouped fusion temporal convolutional network and adaptively select the dominant subcarriers. The spatial feature extraction module, connected to the temporal feature extraction module, is used to extract the spatial correlation between subcarriers through an asymmetric convolutional network, and to complete the progressive mapping from the subcarrier dimension to the key point dimension. The topology constraint module, connected to the spatial feature extraction module, is used to model the temporal dependency of key points and the topology constraint of the skeleton space through a two-stage axial self-attention mechanism to optimize feature representation. The pose decoding module, connected to the topology constraint module, is used to regress the spatial coordinates of key points through a lightweight convolutional decoder and output a continuous and smooth human pose estimation result by combining the skeletal constraint loss.

[0018] Compared with the prior art, the present invention has the following advantages: Spatiotemporal explicit decoupling preserves the temporal causal structure. This invention abandons the traditional two-dimensional convolutional approach with strong spatiotemporal coupling and adopts a decoupling strategy of "temporal modeling first, spatial extraction later". By strictly preserving the temporal causal attributes of the signal through causal temporal convolution, it better fits the dynamic evolution law of continuous actions, effectively improves the accuracy and motion smoothness of continuous pose estimation, and reduces frame-by-frame prediction jitter.

[0019] Extremely lightweight design, adapted for edge deployment. Through the combined optimization of grouped convolution, asymmetric convolution and pooling regression, the number of model parameters and floating-point operations are significantly reduced; on a self-built dataset, only 2.23M parameters and 0.07B floating-point operations are required, with computational overhead far lower than existing mainstream models, and can be directly deployed on IoT edge devices with limited computing power.

[0020] Explicit topological constraints ensure that pose results conform to kinematic laws. An axial self-attention mechanism is introduced to explicitly model the physical dependencies of key points in the human skeleton with low computational increment. Combined with bone length constraint loss, this effectively avoids problems such as joint misalignment and abnormal bone proportions, resulting in more reasonable pose output.

[0021] It exhibits excellent generalization performance and strong scenario adaptability. The spatiotemporally decoupled feature extraction method reduces the risk of overfitting the model to specific scenarios or subjects, maintaining stable high accuracy performance in heterogeneous scenarios across subjects and datasets, and possessing good practical application value. Attached Figure Description

[0022] Figure 1 This is a diagram of the overall WiFlow network architecture of the present invention; Figure 2 This is a topology diagram showing the experimental environment and equipment deployment for this invention. Figure 3 This is a visualization comparison chart of eight common daily action posture estimations according to the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the following embodiments are only for illustrating the present invention and are not intended to limit the scope of the present invention.

[0024] Example 1: Core Method Flow This embodiment provides a spatiotemporally decoupled, lightweight WiFi continuous human pose estimation method, the overall architecture of which is as follows: Figure 1 As shown, the specific implementation steps are as follows: S1 CSI Data Acquisition and Preprocessing A WiFi sensing and data acquisition system was constructed, consisting of one transmitter with three transmitting antennas and two receivers, each with three receiving antennas, forming a total of 18 physical communication links. Both transceivers were equipped with Intel 5300 network cards, operating in the 5GHz band with a bandwidth of 20MHz, and the CSI sampling rate was set to 600Hz.

[0025] In an orthogonal frequency division multiplexing (OFDM) system, the channel frequency response (CFR) of the k-th subcarrier at time t can be modeled as: in and These represent the amplitude response and phase response of the signal, respectively. The signal captured by the receiver is a vector superposition of multipath propagating signals, which can be decomposed into static components. Dynamic components caused by human movement : in A set of dynamic propagation paths, Let n be the complex attenuation factor of the nth path. λ represents the path propagation distance, and λ is the wireless carrier wavelength.

[0026] During the acquisition process, amplitude information of 30 effective subcarriers from each link was extracted, and phase data that is susceptible to noise contamination was discarded. The subcarriers of 18 links were stitched together along the dimensions to obtain panoramic amplitude features of 540 spatial frequency points.

[0027] Using a 30FPS visual annotation signal as the time reference, a sliding time window of length T=20 is set, and the CSI stream is aligned with the visual pose label frame by frame. Finally, a spatiotemporal input feature tensor with a dimension of 540×20 is generated, corresponding to the temporal input of the pose of a single frame.

[0028] To address the missing key points caused by occlusion in visual annotations, a temporal consistency cleaning mechanism using linear interpolation between consecutive frames is employed to repair the labels, ensuring the smoothness of motion in the pose sequence.

[0029] S2 grouping fusion temporal convolution extracts temporal dynamic features The spatiotemporal tensor generated in step S1 is input into a grouped fusion temporal convolutional network (GF-TCN) to prioritize feature modeling in the temporal dimension. The network consists of 4 GF-TCN layers, with the inflation factor increasing exponentially with the number of layers in the order of 1, 2, 4, and 8, gradually expanding the temporal receptive field.

[0030] Let the input time series feature matrix be Where L is the time length and N is the number of subcarrier channels. X is uniformly divided into G subgroups along the channel dimension. The one-dimensional grouped convolution of any channel n within the g-th group is calculated as follows: in Here, b is the activation function, l is the time axis position, and M is the kernel size. Let m be the weight of the g-th convolutional kernel. Through grouped calculation, the number of parameters and computational cost of a single-layer convolution is only 1 / G of that of a standard one-dimensional convolution.

[0031] Within each group, causal dilation convolution is used to expand the receptive field. The calculation method is as follows: Where d is the dilation factor and k is the kernel size. Ensure that the output at each moment depends only on the input at the current and historical moments, and strictly maintain the temporal causality of the signal.

[0032] Based on grouped convolution, the network adaptively evaluates the attitude discrimination capability of each subcarrier through a progressive channel fusion mechanism, retains the advantageous subcarriers with significant dynamic features, filters out noise components that are weakly correlated with human motion, and outputs time-series features with compressed dimensions.

[0033] S3 asymmetric convolutional network extracts spatial frequency features The temporally encoded features are input into an asymmetric convolutional network. This module performs convolution operations only along the subcarrier dimension, keeping the temporal dimension constant and avoiding disruption of the established temporal structure.

[0034] Let the input feature tensor be Where B is the batch size. Let T be the number of input channels, T be the timing length, and S be the number of subcarriers. The feature transform of the nth asymmetric residual block is: in It is an asymmetric convolution kernel that operates only along the subcarrier dimension. This represents the convolution operation. The SiLU activation function is used. A downsampling convolution with a stride of (1,2) is used to achieve progressive compression of the subcarrier dimension.

[0035] The module comprises four levels of asymmetric residual blocks, each consisting of three 1×3 asymmetric convolutional kernels, enhanced by residual connections to improve feature representation. Each residual block progressively compresses the subcarrier dimension through downsampling, while simultaneously expanding the number of feature channels from 8 to 64. After four levels of mapping, the subcarrier dimension is progressively compressed from 240 to 15, corresponding to 15 skeletal key points in the human body, completing the transformation from frequency domain physical features to key point semantic features, and outputting encoded features with a dimension of 64×20×15.

[0036] S4 Axial Self-Attention Modeling Skeleton Topological Constraints Input the encoded features output from step S3 into the axial self-attention module, set the number of groups to 8, and perform attention calculation in two stages.

[0037] The first stage is temporal attention: features are reshaped into the form (B×K)×C×T (K is the number of keypoints, C is the number of channels, and T is the temporal length), and attention calculations are performed in parallel on the temporal sequence of each keypoint. Query, key, and value matrices are generated through learnable linear projections. Divide the T-dimensional features into G groups, where the dimension of each group is d=T / G. Calculate the attention within each group using a scaled dot product. By adaptively weighting the feature contributions at different time steps through temporal attention, the temporal information most valuable for pose estimation is enhanced.

[0038] The second stage is spatial node attention: the features are reshaped into the form of (B×T)×C×K, attention calculations are performed on the K key points at each time step, the global topological dependencies between key points are modeled, and the physical structural constraints of the human skeleton are integrated into the feature representation.

[0039] After the two-stage calculation is completed, the output has optimization characteristics that simultaneously possess temporal dynamics and spatial structural constraints.

[0040] S5 lightweight decoder regresses continuous attitude coordinates The optimized features are input into the decoder. First, the number of channels is reduced from 64 to 32 through 3×3 convolution, and then batch normalization and SiLU activation function are used to refine the features. Then, the channels are compressed to 2 through 1×1 convolution, which correspond to the x and y two-dimensional coordinates of the key points respectively.

[0041] Subsequently, adaptive max pooling was used to aggregate the temporal dimension from 20 to 1, resulting in a 2×15×1 feature tensor. After tensor compression and transpose, a 15×2 keypoint coordinate matrix was output, corresponding to the human pose in a single frame.

[0042] During the training phase, a combined loss function is used to optimize the network. The main loss is measured using the Smooth L1 norm to measure the coordinate prediction error. in To predict coordinates, Where N is the true coordinates, N is the number of samples, and K is the number of keypoints.

[0043] Additional bone length constraint loss is introduced to enhance the structural rationality of the output pose: in Let h( be the set of skeletal edges) ) is the smoothing error function.

[0044] The total loss function is a weighted sum of the two: in In this embodiment, λ=0.2 is set as the skeleton constraint weight.

[0045] Example 2: Network Parameters and Experimental Verification This embodiment provides the core network parameter settings and verifies the performance of the invention through multiple sets of experiments. This invention is implemented in PyTorch and trained and tested on an NVIDIA GeForce RTX 4090D GPU environment, using the AdamW optimizer with weight decay. The initial learning rate is The training process lasts for 50 epochs with a batch size of 64.

[0046] 2.1 Network Core Module Parameter Settings The input / output dimensions and core parameters of each module in the WiFlow network are shown in the table below: Table 1 WiFlow Network Core Module Parameter Settings 2.2 Evaluation Indicators The experiment used the percentage of correct key points (PCK) and the mean joint position error (MPJPE) as the core evaluation indicators. PCK was defined as: in The normalization factor is the distance from the right shoulder to the left hip. I ( ) is a Boolean indicator function, and α is the error threshold.

[0047] MPJPE is defined as: This indicator directly calculates the average Euclidean distance between the predicted key points and the actual coordinates, quantifying the absolute positioning deviation.

[0048] 2.3 Performance Verification of Self-Built Dataset The self-built dataset contains 360,000 strictly synchronized CSI-pose samples across 8 categories of continuous daily actions from 5 volunteers. Two data partitioning methods were used for validation. Setting 1: Random partitioning (subject-related), with training, validation, and test sets divided at 70%:15%:15%; Setting 2: Cross-subject partitioning (subject-independent), using leave-one-out cross-validation; Performance comparison of setting 1 (random partitioning) The performance and computational efficiency of each model under random partitioning are compared in the table below: Table 2 Performance comparison under setting 1 (random partitioning) As shown in Table 2, the WiFlow model of this invention significantly outperforms the benchmark model in all accuracy metrics, achieving 99.48% PCK@50 and a low MPJPE of 0.007m. Meanwhile, it has only 2.23M parameters and only 0.07B of computation, resulting in significantly faster training speed and higher inference efficiency, achieving a dual improvement in accuracy and efficiency.

[0049] Performance comparison in setting 2 (across subject partitions) The leave-one-out method was used to verify the model's cross-user generalization ability, and the results are shown in the table below: Table 3 Performance comparison under setting 2 (leave-one-out validation across users) As shown in Table 3, the present invention maintains stable high performance in cross-user generalization scenarios, with an average PCK@20 of 87.26%; even on the most challenging Subject 3 data, it still significantly outperforms all benchmark models, verifying the strong generalization capability of the spatiotemporal decoupling architecture.

[0050] 2.4 Generalization Validation on Public Datasets The generalization performance was further validated on the publicly available large-scale WiFi sensing dataset MM-Fi, using a combination of Protocol 3 (containing 27 action classes) and the random partitioning Setting 1. The results are shown in the table below: Table 4 Performance comparison under the MM-Fi dataset As shown in Table 4, in cross-dataset and complex action scenarios, the PCK@50 of WiFlow of this invention reaches 88.46%, which is 2.72 percentage points higher than the second-best model, and the MPJPE is as low as 0.120m. At the same time, the number of parameters is only 1.06M and the computation is only 0.02B, which is far lower than all benchmark models, fully verifying its excellent cross-scenario generalization ability and lightweight advantages.

[0051] 2.5 Ablation Experiment Verification To verify the effectiveness of each core module, ablation experiments were conducted on a self-built dataset, and the results are shown in the table below: Table 5 Comparison of ablation experimental performance of core modules As shown in Table 5, the spatiotemporal decoupling architecture (GF-TCN + asymmetric convolution) contributed the most to the performance improvement. After replacing it with two-dimensional convolution, PCK@10 decreased by 7.81 percentage points. The group fusion strategy, causal dilation convolution and axial attention mechanism all made positive contributions to the performance, verifying the necessity and effectiveness of the design of each module.

[0052] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A spatiotemporally decoupled, lightweight WiFi continuous human pose estimation method, characterized in that, Includes the following steps: S1 collects multi-link WiFi channel status information, extracts signal amplitude features, and constructs a standardized spatiotemporal input feature tensor through time alignment and sliding window processing. S2 extracts temporal dynamic features through a grouped fusion temporal convolutional network, reduces computational redundancy by using grouped convolution, captures long-range temporal dependencies by combining causal dilated convolution, and adaptively selects advantageous subcarriers to preserve the temporal causal attributes of the signal throughout the process. S3 extracts spatial frequency features through an asymmetric convolutional network, performs convolution only along the subcarrier dimension while keeping the temporal dimension constant, and progressively compresses the subcarrier dimension through multi-level residual blocks to complete the mapping from frequency domain features to the semantic dimension of human key points. S4 introduces an axial self-attention mechanism, decoupling attention computation along two orthogonal directions of temporal and keypoint space, respectively modeling single keypoint temporal dependency and multi-keypoint topological constraint, and optimizing feature representation; S5 uses a lightweight convolutional decoder to map encoded features into keypoint spatial coordinates, and combines bone length constraint loss to optimize the output, thus obtaining continuous temporal human pose estimation results.

2. The spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of step S1 is as follows: configure a WiFi acquisition system with multiple transmitting and receiving antennas, extract the effective subcarrier amplitude information of each communication link, and discard the phase data; stitch all links along the subcarrier dimension to form a panoramic amplitude feature; Based on the set sampling rate and sliding window length, a spatiotemporal input tensor of subcarrier number × time step is generated by time alignment.

3. The spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of the grouped fusion temporal convolutional network in step S2 is as follows: The input features are uniformly divided into multiple independent subgroups along the subcarrier channel dimension. Each group performs a one-dimensional temporal convolution independently. The computational cost of a single group is 1 / G of the standard one-dimensional convolution, where G is the number of groups. Each group uses causal dilated convolution, with the dilation factor increasing exponentially with the number of network layers. The convolution operation only accesses information from the current and historical moments, strictly maintaining temporal causality. The subcarrier discrimination capability is adaptively evaluated through a progressive channel fusion mechanism to select advantageous subcarriers that are strongly correlated with human motion.

4. The spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of the asymmetric convolutional network in step S3 is as follows: An asymmetric convolution kernel with a receptive field of 1×k is used, and one-dimensional sliding convolution is performed only in the subcarrier dimension, while the structure remains constant in the temporal dimension. Construct multi-level asymmetric residual blocks. Each level of residual block compresses the subcarrier dimension through downsampling convolution with a stride of (1,2) while expanding the number of feature channels. The high-dimensional subcarrier features are transformed into semantic features corresponding to the number of human body key points through multi-level mapping, with each output dimension corresponding to one human body key point.

5. A spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of the axial self-attention mechanism in step S4 is as follows: The first stage is temporal dimension attention: the input features are reshaped into the form of "number of key points × number of channels × time step", the temporal sequence of each key point is processed in parallel, and the temporal dependency is calculated by grouping and scaling dot product attention, and high-value temporal features are highlighted by weighting. The second stage is spatial node attention: the input features are reshaped into the form of "time step × number of channels × number of key points", attention calculation is performed along the key point dimension, and the global topological dependency between key points of the human skeleton is modeled. The two-stage computation is executed in cascade, and the output has encoding characteristics that simultaneously possess temporal dynamics and spatial structural constraints.

6. The spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of the decoder in step S5 is as follows: first, dimensionality reduction is achieved through 3×3 convolution, combined with batch normalization and activation functions to enhance nonlinear expression; then, channels are compressed to the coordinate dimension through 1×1 convolution; finally, temporal information is aggregated through adaptive max pooling to output the key point coordinate tensor.

7. A spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The specific implementation of the loss function in step S5 is as follows: the Smooth L1 norm is used as the main loss to measure the coordinate prediction error, and an additional bone length constraint loss is introduced to penalize the length deviation between the predicted bone and the real bone. The total loss is the weighted sum of the main loss and the bone constraint loss.

8. A spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 1, characterized in that, The method uses the percentage of correct key points and the average joint position error as core evaluation indicators to comprehensively evaluate the attitude estimation accuracy and absolute positioning deviation under different thresholds.

9. A spatiotemporally decoupled lightweight WiFi continuous human pose estimation method according to claim 2, characterized in that, Step S1 also includes label preprocessing: for missing key points in visual annotation caused by occlusion, a temporal consistency cleaning mechanism of linear interpolation between consecutive frames is used to repair them, ensuring the smoothness of motion of the pose sequence.

10. A spatiotemporally decoupled, lightweight WiFi continuous human pose estimation system, characterized in that, include: The data preprocessing module is used to collect multi-link WiFi channel state information, extract amplitude features, perform time alignment and window slicing, and output a standardized spatiotemporal input tensor. The temporal feature extraction module, connected to the data preprocessing module, is used to extract the temporal dynamic features of CSI signals through a grouped fusion temporal convolutional network and adaptively select the dominant subcarriers. The spatial feature extraction module, connected to the temporal feature extraction module, is used to extract the spatial correlation between subcarriers through an asymmetric convolutional network, and to complete the progressive mapping from the subcarrier dimension to the key point dimension. The topology constraint module, connected to the spatial feature extraction module, is used to model the temporal dependency of key points and the topology constraint of the skeleton space through a two-stage axial self-attention mechanism to optimize feature representation. The pose decoding module, connected to the topology constraint module, is used to regress the spatial coordinates of key points through a lightweight convolutional decoder and output a continuous and smooth human pose estimation result by combining the skeletal constraint loss.