Human body posture sensing method and device based on channel state information

By preprocessing and feature fusion of channel state information, the problem of low efficiency in signal feature utilization in existing technologies is solved, and higher precision human posture perception is achieved.

CN120853259APending Publication Date: 2025-10-28XIDIAN UNIV HANGZHOU RES INST +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510967648.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing human posture perception technologies based on wireless communication signals have low efficiency in utilizing signal features, resulting in insufficient perception accuracy.

Method used

By employing a channel state information preprocessing method, spatiotemporal tensors, frequency tensors, and wavelet tensors are obtained respectively. Feature fusion and prediction are then performed using a pre-trained target network model to output human pose information.

Benefits of technology

It improves the expressive power and utilization efficiency of signal features, and enhances the accuracy and perception precision of human posture prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853259A_ABST
    Figure CN120853259A_ABST
Patent Text Reader

Abstract

The invention discloses a human body posture sensing method and device based on channel state information. The method comprises the following steps: firstly, acquiring the channel state information of a communication signal for human body posture sensing; secondly, preprocessing the channel state information to respectively obtain a time-space domain tensor, a frequency domain tensor and a wavelet domain tensor; and inputting the time-space domain tensor, the frequency domain tensor and the wavelet domain tensor into a pre-trained target network model, performing feature fusion on the time-space domain tensor, the frequency domain tensor and the wavelet domain tensor by the target network model, performing human body posture prediction based on the fused features, and outputting human body posture information. According to the method, the fusion features of the time-space domain tensor, the frequency domain tensor and the wavelet domain tensor are adopted for human body posture prediction, joint information of the time-space domain, the frequency domain and the wavelet domain of communication signals can be fully mined, the expression ability and the utilization efficiency of signal features are improved, and therefore the human body posture prediction accuracy of the model is improved, and the human body posture prediction efficiency is improved. And the sensing precision of the human body posture is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to a method and device for human posture perception based on channel state information. Background Technology

[0002] With the synergistic evolution of sensing and communication technologies and 5G-A technology, the field of wireless sensing is ushering in new development opportunities. Sensing and communication systems achieve a high degree of integration of communication and sensing functions by sharing hardware platforms and spectrum resources. The core advantage of wireless sensing technology lies in its non-contact nature, meaning that users do not need to wear or carry any additional devices. This characteristic makes wireless sensing technology more convenient and user-friendly compared to traditional sensor technologies. Wireless sensing technology also has significant advantages over traditional vision-based technologies, such as cameras. First, wireless sensing technology is not limited by lighting conditions and can operate stably in low-light or even no-light environments, while camera performance degrades significantly at night or in complex lighting conditions. Second, wireless communication signals have the ability to penetrate obstacles, making wireless sensing effective even in non-line-of-sight scenarios where the target is obscured, whereas camera-based sensing requires the target to be within the line of sight to capture an image and complete the sensing.

[0003] Currently, wireless sensing technology has been extensively studied in fields such as human posture perception, vital sign monitoring, user authentication, behavior analysis, and intelligent interaction. However, current human posture perception technologies based on wireless communication signals typically rely on time-domain or frequency-domain signal features for posture estimation, resulting in low efficiency in utilizing signal features and consequently low accuracy in perceiving human posture. Summary of the Invention

[0004] To address the aforementioned problems in the prior art, this invention provides a method and apparatus for human posture perception based on channel state information.

[0005] The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, the present invention provides a human posture perception method based on channel state information, comprising: Acquire channel state information of communication signals used for human posture perception; The channel state information is preprocessed to obtain the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor, respectively. The spatiotemporal tensor, frequency tensor, and wavelet tensor are input into a pre-trained target network model, which outputs human pose information. The target network model is used to fuse features of the spatiotemporal tensor, frequency tensor, and wavelet tensor, and to predict human pose based on the fused features, thus outputting human pose information.

[0006] Secondly, the present invention provides a human posture sensing device based on channel state information, comprising: The acquisition module is used to acquire channel state information of communication signals used for human posture perception. The preprocessing module is used to preprocess the channel state information to obtain the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor, respectively. The prediction module is used to input spatiotemporal domain tensors, frequency domain tensors, and wavelet domain tensors into a pre-trained target network model and output human pose information. The target network model is used to perform feature fusion on the spatiotemporal domain tensors, frequency domain tensors, and wavelet domain tensors, and to predict human pose based on the fused features, outputting human pose information.

[0007] This invention provides a method and apparatus for human posture perception based on channel state information. First, the channel state information of the communication signal used for human posture perception is acquired. Then, the channel state information is preprocessed to obtain spatiotemporal tensors, frequency domain tensors, and wavelet domain tensors. These tensors are then input into a pre-trained target network model. The target network model performs feature fusion on the spatiotemporal, frequency, and wavelet domain tensors and predicts human posture based on the fused features, outputting human posture information. This invention uses the fused features of spatiotemporal, frequency, and wavelet domain tensors for human posture prediction, which can fully exploit the joint information of the communication signal in the spatiotemporal, frequency, and wavelet domains, improving the expressive power and utilization efficiency of signal features, thereby improving the accuracy of the model's human posture prediction and enhancing the perception precision of human posture.

[0008] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a human posture perception method based on channel state information provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the prediction process of the target network model in a human posture perception method based on channel state information provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the loss change during the training process of a human posture perception method based on channel state information provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the test error change during the training process of a human posture perception method based on channel state information provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of the human posture perception device based on channel state information provided in an embodiment of the present invention. Detailed Implementation

[0010] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0011] This invention provides a method for human posture perception based on channel state information, see [link to relevant documentation]. Figure 1 , the method comprises the following steps: S10. Obtain channel state information of communication signals used for human posture perception.

[0012] For example, Channel State Information (CSI) can be extracted from the communication signals of multiple terminal devices communicating normally with a 5G base station. CSI includes IQ data, where I (In-phase) refers to the real part of the communication signal on the carrier phase reference, and Q (Quadrature) refers to the imaginary part of the communication signal on the carrier phase reference. For instance, in this embodiment, each frame of IQ data can be represented as a 6×4×30×20×2 tensor, representing the I and Q values ​​on 6 transmit antennas, 4 receive antennas, 30 subcarriers, and 20 consecutive data packets.

[0013] S20. Preprocess the channel state information to obtain the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor, respectively.

[0014] Specifically, since channel state information is not suitable as direct input to the target network model, further preprocessing of the channel state information is needed to make it easier for the target network model to learn features related to human posture. Amplitude and phase data can be separated from the channel state information, and the amplitude and phase data can be transformed to obtain the spatiotemporal tensor, frequency tensor, and wavelet tensor corresponding to the communication signal.

[0015] Optionally, in step S20, the channel state information is preprocessed to obtain spatiotemporal tensors, frequency tensors, and wavelet tensors, specifically including: S201. Separate amplitude data and phase data from channel state information.

[0016] For example, amplitude data separated from channel state information It can be represented as: (1) Phase data separated from channel state information It can be represented as: (2) in, Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the data packet index.

[0017] S202. The phase data is unwound and the linear trend term is removed to obtain the first preprocessed data. The amplitude data is filtered and exponentially smoothed to obtain the second preprocessed data. The first preprocessed data and the second preprocessed data are concatenated to obtain the spatiotemporal tensor.

[0018] For example, phase data can be processed along the subcarrier dimension (i.e. Phase unwrapping is performed on the dimension), and then the linear trend term is removed by least squares fitting to obtain the first preprocessed data.

[0019] Optionally, the first preprocessed data can be represented as: (3) in, This represents the first preprocessed data. This represents the phase data after unwinding. Indicates the slope. Represents the intercept parameter. Indicates a subcarrier.

[0020] Here, by performing least-squares fitting on the phase data after unwinding, the fitted line can be obtained. , Let be the slope of the fitted line. This is the intercept parameter of the fitted line. Since the phase data is unwrapped along the subcarrier dimension, in the schematic diagram of the fitted line, the horizontal axis represents the subcarrier. The vertical axis represents phase data.

[0021] In addition, the amplitude data can be processed by median filtering, as shown in equation (4), to suppress occasional isolated sharp noise.

[0022] (4) Next, the filtered amplitude data is processed along the data packet dimension (i.e. The second preprocessed data can be obtained by exponentially smoothing the data (dimensions) to enhance the continuity of time series.

[0023] Optionally, the second preprocessed data can be represented as: (5) in, and These represent the second preprocessed data corresponding to different packet indices. and These represent the amplitude data corresponding to different data packet indices after filtering. As a smoothing factor, Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the data packet index.

[0024] Here, smoothing factor Used to smooth amplitude data.

[0025] Finally, the first preprocessed data and the second preprocessed data By concatenating the components, we obtain the spacetime tensor, which can be represented as: .

[0026] S203. Perform Fast Fourier Transform on the channel state information to obtain the third preprocessed data, and use multiple main frequency data in the third preprocessed data to form a frequency domain tensor.

[0027] For example, each transmitting antenna and each receiving antenna can be paired, for instance, 6 transmitting antennas and 4 receiving antennas can form 24 antenna pairs. Performing a Fast Fourier Transform on the channel state information of each antenna pair yields third preprocessed data. Multiple key frequency data points are retained from this third preprocessed data; for example, 10 key frequency data points can be retained to form... A tensor of dimension will Let it be denoted as frequency domain tensor .

[0028] Alternatively, the frequency domain tensor can be represented as: (6) in, Represents the frequency domain tensor. The in-phase component in the channel state information. These are the orthogonal components in the channel state information. This indicates the clock speed data. This indicates the total number of clock frequency data. Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the packet index. This indicates the total number of data packet indexes. Represents the imaginary unit. Represents angular frequency. Indicates phase as The complex exponential signal.

[0029] S204. Perform continuous wavelet transform on the amplitude data to obtain multiple coefficient matrices; determine the wavelet domain tensor based on the multiple coefficient matrices.

[0030] For example, taking each frame of IQ data as including 6 transmit antennas, 4 receive antennas, 30 subcarriers, and 20 data packets, 24 antenna pairs can be formed. For each antenna pair, the corresponding amplitude data is first extracted. Then for each subcarrier corresponding to Continuous wavelet transform can be performed using three scale parameters (1, 2, and 4) to control the scaling of the wavelet. After discretization, the... Convolution with the wavelet function yields sub-coefficient matrices at different scales. Then, the multiple sub-coefficient matrices corresponding to each subcarrier are merged to obtain a single coefficient matrix for each subcarrier. .

[0031] Here, This represents 20 amplitude data points for a specific subcarrier of a specific antenna pair.

[0032] Alternatively, the coefficient matrix can be represented as: (7) in, This represents the amplitude data corresponding to each antenna pair consisting of a transmitting antenna and a receiving antenna. Indicates the scale parameter. Indicates time shift, Describing wavelet functions, Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the data packet index.

[0033] Here, It can be the wavelet function Mexican Hat.

[0034] Further, in step 204, the wavelet domain tensor is determined based on multiple coefficient matrices, specifically including: Stack the multiple coefficient matrices corresponding to each antenna pair to obtain a stacked matrix; calculate the mean of the stacked matrix along the subcarrier dimension to obtain the subtensor corresponding to each antenna pair; The sub-tensors corresponding to each antenna pair are merged to obtain the wavelet domain tensor.

[0035] Specifically, the coefficient matrix corresponding to all subcarriers of each antenna pair Stacking can yield a stacking matrix. Then, by averaging the stacked matrix along the subcarrier dimension, the subtensor corresponding to each antenna pair can be obtained. Finally, merge the subtensors corresponding to all antenna pairs. The wavelet domain tensor can be obtained. .

[0036] S30. Input the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor into the pre-trained target network model and output human posture information.

[0037] The target network model is used to fuse features of spatiotemporal tensors, frequency tensors, and wavelet tensors, and to predict human pose based on the fused features, outputting human pose information.

[0038] Specifically, refer to Figure 2 The target network model can extract features from the input spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor respectively, obtaining spatiotemporal domain features, frequency domain tensors, and wavelet domain tensors. Then, based on a multi-head cross-domain attention mechanism, the spatiotemporal domain features, frequency domain tensors, and wavelet domain tensors are fused to obtain a feature fusion tensor, realizing deep interaction between spatiotemporal domain features, frequency domain features, and wavelet domain features. Next, the feature fusion tensor is temporally encoded to obtain a temporal tensor. Finally, the temporal tensor is used for pose regression prediction to output human pose information, thereby completing human pose perception.

[0039] Optionally, the target network model may include a feature extraction network, a feature fusion network, an encoder, and a pose regression prediction network; In step S30, the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor are input into the pre-trained target network model to output human pose information, specifically including: S301. Input the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor into the feature extraction network, and output spatiotemporal domain features, frequency domain features, and wavelet domain features with consistent dimensions.

[0040] For example, refer to Figure 2 The feature extraction network can include a spatiotemporal domain feature extraction module, a frequency domain feature extraction module, and a wavelet domain feature extraction module.

[0041] For the input spatiotemporal tensor The spatiotemporal domain feature extraction module splits the spatiotemporal domain tensor into frames according to the data packet dimension, then merges the receiving antenna and transmitting antenna into the same dimension, and performs linear interpolation to adjust each frame of data. Dimension tensors. Simultaneously, the input channels of the residual network (e.g., the ResNet18 backbone network) are adjusted to 2, adapting to dual-channel input of amplitude and phase data for channel state information, and the output dimension of the final fully connected layer is adjusted to 256. Each frame, after being processed by the adjusted ResNet18 network, can output... The spatiotemporal features of the dimension, where B is the batch size, which represents the size of the sample set for a single model training run.

[0042] For the input frequency domain tensor The frequency domain feature extraction module can employ a 2-layer Conv2d-BN-ReLU-MaxPool structure to compress the frequency domain tensor to... Dimensions, then mapped to Frequency domain characteristics of dimensionality.

[0043] For the input wavelet domain tensor The wavelet domain feature extraction module can consist of two layers of 1D dilated convolutions, capable of capturing cross-scale temporal dynamic features, and ultimately outputting... Wavelet domain features of dimensionality.

[0044] S302. Input the spatiotemporal domain features, frequency domain features, and wavelet domain features into the feature fusion network to perform feature fusion and output the feature fusion tensor.

[0045] For example, refer to Figure 2 The feature fusion network can introduce a multi-head cross-domain attention mechanism. First, it fuses the spatiotemporal domain features and frequency domain features separately to obtain spatiotemporal-frequency domain cross-attention features. Then, it fuses the spatiotemporal domain features and frequency domain features to obtain spatiotemporal-wavelet domain cross-attention features. Finally, it fuses the spatiotemporal-frequency domain cross-attention features and the spatiotemporal-wavelet domain cross-attention features to output a feature fusion tensor.

[0046] S303. Input the feature fusion tensor into the encoder for temporal encoding and output a temporal tensor.

[0047] For example, refer to Figure 2 The encoder can consist of three Transformer encoder layers and pooling layers. After the feature fusion tensor is temporally encoded by the Transformer encoder, the pooling layers perform average pooling along the temporal dimension to output a temporal tensor, which can be represented as follows: .

[0048] Each Transformer encoder layer can include a self-attention layer and a feedforward neural network (FFN) layer. The self-attention layer can introduce residual connections and layer normalization to ensure gradient stability. The feedforward neural network layer receives the output of the self-attention layer and performs further nonlinear transformations on it to capture more complex features and representations.

[0049] In this embodiment, the self-attention layer can capture long-range dependencies, taking into account both short-term dynamics and global trends. Stacked Transformer encoder structures can improve the model's hierarchical feature representation capabilities, and average pooling can ensure a globally consistent output representation.

[0050] S304. Input the temporal tensor into the pose regression prediction network to predict the pose and output human pose information.

[0051] Optionally, the pose regression prediction network may include a global regression prediction network and multiple regional pose prediction networks, with the multiple regional pose prediction networks respectively used to predict pose information for different regions of the human body.

[0052] For example, refer to Figure 2 The global regression prediction network is used to perform global pose regression on the temporal tensor. Multiple regional pose prediction networks may include a regional pose prediction network corresponding to the torso, an arm, and a leg. In this embodiment, both the global regression prediction network and the multiple regional pose prediction networks can be multilayer perceptron networks (MLPs).

[0053] In step S304, the temporal tensor is input into the pose regression prediction network for pose prediction, and the human pose information is output, including: S3041. Input the temporal tensor into the global regression prediction network to perform global prediction and obtain the global prediction pose tensor.

[0054] For example, refer to Figure 2 Translate the time tensor By inputting the data into a global regression prediction network for global prediction, the global prediction pose tensor can be obtained. , is represented as: (8) in, This represents the prediction results of the global regression prediction network. This indicates the total number of joints, and 3 refers to the dimension of the three-dimensional coordinates of each joint.

[0055] Here, It can be 25 or 14, and this embodiment does not limit it.

[0056] S3042. According to the preset region division rules, the global prediction attitude tensor is divided into regions to obtain sub-global prediction attitude tensors corresponding to multiple regions.

[0057] For example, to improve local joint accuracy, a partitioned residual refinement is introduced. A preset region partitioning rule divides the human body into multiple regions, such as the torso region, arm region, and leg region. Based on the preset region partitioning rule, the global predicted pose tensor is divided into regions, resulting in a sub-global predicted pose tensor corresponding to each region.

[0058] S3043. After concatenating the temporal tensor with the sub-global prediction pose tensors corresponding to multiple regions, input them into the regional pose prediction networks corresponding to multiple regions for partitioned prediction to obtain the pose information of multiple regions.

[0059] For example, refer to Figure 2 We set up region pose prediction networks for the torso, arm, and leg regions respectively, and used temporal tensors... The tensor is concatenated with the subglobal predicted pose tensor corresponding to each region, and the concatenation result is input into the regional pose prediction network corresponding to each region to perform regional prediction and output the pose information of multiple regions.

[0060] S3044. Combine the posture information of multiple regions to output human posture information.

[0061] For example, refer to Figure 2 It stitches together pose information from multiple regions to output human pose information. This embodiment ensures overall posture consistency through global prediction, and then refines each region locally using region prediction, which can improve the position prediction accuracy of key joints. The hierarchical structure takes into account both the whole body and local areas, enhancing the model's expressive power.

[0062] During the training of the target network model, the three-dimensional real coordinates of the human joints of the target object and the channel state information can be collected. The collected channel state information is preprocessed according to step S20 to obtain the spatiotemporal domain tensor, frequency domain tensor and wavelet domain tensor. The obtained spatiotemporal domain tensor, frequency domain tensor and wavelet domain tensor are used to construct the training dataset. The collected three-dimensional real coordinates are used as label data, and the target loss function is set to train the target network model.

[0063] Optionally, the target loss function used to train the target network model can be expressed as: (9) in, This represents the weighted mean square error of the key points. Indicates loss of skeletal consistency. Represents the time-series smoothing loss. , , These represent the balance factors of the corresponding terms.

[0064] The objective loss function used in this embodiment forms a multi-dimensional and complementary loss system, which can significantly improve the model's performance and robustness in complex actions and dynamic scenarios.

[0065] Alternatively, the weighted mean square error of the key points can be expressed as: (10) in, This indicates the size of the sample set used in a single model training iteration. , This indicates the total number of joints. , This represents the weight values ​​for different key points. Represents the predicted coordinates of the joints. Represents the actual coordinates of the joints. This represents the Euclidean distance between the predicted coordinates and the actual coordinates of a key point.

[0066] In this embodiment, Based on the average Euclidean distance between the predicted and actual coordinates of joints quantized by MPJPE, different weights are assigned to different joints to emphasize the importance of key joints. By increasing the error penalty for joints in the core torso region, the model is guided to prioritize correcting the predictions of torso region joints, ensuring overall posture stability and reducing overall posture jitter.

[0067] Loss of skeletal consistency can be expressed as: (11) in, This represents a parent-child joint pair, which consists of two joints used to determine the length of a single bone. , This represents the set of parent-child pairs of key nodes. This represents the predicted coordinates of the parent node. Represents the predicted coordinates of the sub-joints. This indicates the calculation of the length of each bone. This indicates the calculation of the variance in length of each bone.

[0068] In this embodiment, the skeletal consistency loss, by penalizing skeletal length fluctuations, prevents physically unreasonable stretching and twisting of joint predictions, thus ensuring biomechanical properties.

[0069] The time series smoothing loss can be expressed as: (12) in, This indicates the predicted coordinates of the key point in the current iteration. This indicates the predicted coordinates of the same key point in the next iteration. This represents the distance between two consecutive predicted coordinates of a key point. This represents the smoothing weight.

[0070] In this embodiment, the temporal smoothing loss encourages the model to generate coherent and smooth action sequences by constraining the magnitude of positional changes of key points in consecutive frames, thereby reducing jitter and disjointed transitions.

[0071] This invention provides a human posture perception method based on channel state information. First, the channel state information of the communication signal used for human posture perception is acquired. Then, the channel state information is preprocessed to obtain spatiotemporal tensors, frequency domain tensors, and wavelet domain tensors. These tensors are then input into a pre-trained target network model. The target network model performs feature fusion on the spatiotemporal, frequency, and wavelet domain tensors and predicts human posture based on the fused features, outputting human posture information. This invention uses the fused features of spatiotemporal, frequency, and wavelet domain tensors for human posture prediction, which can fully exploit the joint information of the communication signal in the spatiotemporal, frequency, and wavelet domains, improving the expressive power and utilization efficiency of signal features, thereby improving the accuracy of the model's human posture prediction and enhancing the perception precision of human posture.

[0072] The following simulation experiment further illustrates the human posture perception method based on channel state information provided by this invention.

[0073] I. Experimental Conditions Model Training Environment: Training experiments were conducted on a server running Ubuntu 20.04, with software configured as Python 3.10, PyTorch 2.4.0, CUDA 11.8, and cuDNN 8. The server hardware consisted of an NVIDIA Tesla P100-16GB GPU with a floating-point performance of 5.18 TFLOPS, 16GB of GPU memory, and 732.16GB / s of GPU bandwidth, along with 16 PCIe channels and 15.75GB / s of PCIe bandwidth. It also featured 12 Intel(R) Xeon(R) Platinum 8260 CPUs @ 2.30GHz, 60GB of usable memory, a hard disk bandwidth of 276.40MB / s, and 200GB of usable space.

[0074] Data Acquisition: In an indoor office setting, posture data of a single person within a 3m x 3m area was collected, including the 3D true coordinates of human joints and IQ data. The acquisition equipment consisted of three UEs: a 5G base station with 4 receiving antennas (Rx) and a UE terminal with 2 transmitting antennas (Tx). Each frame of data was saved as... The tensor represents the IQ data of 20 data packets from 6 transmit antennas, 4 receive antennas, and 30 subcarriers. Simultaneously, a Kinect V2 was deployed to synchronously acquire the 3D real-world coordinates of human joints. A total of 40,668 frames of data were collected in the experiment, of which 32,640 frames were used for training and 8,028 frames were used for testing.

[0075] The evaluation indicators for the experimental results are: MPJPE: Quantizes the average Euclidean distance between the predicted and true coordinates of a joint, where K represents the number of test frames and N represents the number of joints.

[0076]

[0077] II. Experimental Results and Analysis exist Figure 3 In the first 20 epochs, the training loss decreased from approximately 1.25 to 0.20, then slowly decreased and stabilized, indicating that the model had high learning efficiency in the early stages and that there were no gradient explosion or vanishing problems. The training loss continued to decrease and stabilize without oscillation or rebound, indicating that the model complexity was matched with the data scale and that overfitting did not occur.

[0078] exist Figure 4 The test set error decreased synchronously, further validating the good generalization ability. The MPJPE metric is measured in meters. The initial MPJPE value was approximately 0.67, and eventually stabilized below 0.09, meaning the error was below 90 mm, a reduction of 86%. This indicates a significant improvement in the model's pose estimation ability. Furthermore, the MPJPE stabilized after approximately 40 epochs, synchronizing with the training loss curve, indicating the effectiveness of the training strategy.

[0079] Corresponding to the above-described method for human posture perception based on channel state information, this embodiment of the invention also provides a human posture perception device based on channel state information; such as Figure 5 As shown, the device may include: The acquisition module is used to acquire channel state information of communication signals used for human posture perception. The preprocessing module is used to preprocess the channel state information to obtain the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor, respectively. The prediction module is used to input spatiotemporal domain tensors, frequency domain tensors, and wavelet domain tensors into a pre-trained target network model and output human pose information. The target network model is used to perform feature fusion on the spatiotemporal domain tensors, frequency domain tensors, and wavelet domain tensors, and to predict human pose based on the fused features, outputting human pose information.

[0080] For details regarding the device, please refer to the steps of the human posture perception method based on channel state information provided in the first aspect; these will not be repeated here.

[0081] This invention provides a human posture perception device based on channel state information. First, it acquires the channel state information of the communication signal used for human posture perception. Then, it preprocesses the channel state information to obtain spatiotemporal tensors, frequency domain tensors, and wavelet domain tensors. Next, it inputs these tensors into a pre-trained target network model. The target network model fuses features from the spatiotemporal, frequency, and wavelet domain tensors and predicts human posture based on the fused features, outputting human posture information. This invention uses the fused features of spatiotemporal, frequency, and wavelet domain tensors for human posture prediction, which can fully exploit the joint information of the communication signal in the spatiotemporal, frequency, and wavelet domains, improving the expressive power and utilization efficiency of signal features, thereby improving the accuracy of the model's human posture prediction and enhancing the perception precision of human posture.

[0082] It should be noted that the device is basically similar to the method embodiment, so the description is relatively simple. For relevant parts, please refer to the description of the method embodiment.

[0083] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0084] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0085] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0086] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for human posture perception based on channel state information, characterized in that, include: Acquire channel state information of communication signals used for human posture perception; The channel state information is preprocessed to obtain spatiotemporal tensors, frequency tensors, and wavelet tensors, respectively. The spatiotemporal tensor, the frequency tensor, and the wavelet tensor are input into a pre-trained target network model to output human pose information. The target network model is used to perform feature fusion on the spatiotemporal tensor, the frequency tensor, and the wavelet tensor, and to predict human pose based on the fused features, thus outputting human pose information.

2. The human posture perception method based on channel state information according to claim 1, characterized in that, The preprocessing of the channel state information to obtain spatiotemporal tensors, frequency tensors, and wavelet tensors includes: Amplitude data and phase data are separated from the channel state information; The phase data is unwound and the linear trend term is removed to obtain the first preprocessed data. The amplitude data is filtered and exponentially smoothed to obtain the second preprocessed data. The first preprocessed data and the second preprocessed data are concatenated to obtain the spatiotemporal tensor. The channel state information is processed by Fast Fourier Transform to obtain third preprocessed data, and multiple main frequency data in the third preprocessed data are used to form a frequency domain tensor. The amplitude data is subjected to continuous wavelet transform to obtain multiple coefficient matrices; based on the multiple coefficient matrices, the wavelet domain tensor is determined.

3. The human posture perception method based on channel state information according to claim 2, characterized in that, The first preprocessed data is represented as follows: ; in, This represents the first preprocessed data. This represents the phase data after unwinding. Indicates the slope. Represents the intercept parameter. Indicates subcarrier; The second preprocessed data is represented as follows: ; in, and These represent the second preprocessed data corresponding to different packet indices. and These represent the amplitude data corresponding to different data packet indices after filtering. As a smoothing factor, Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the data packet index.

4. The human posture perception method based on channel state information according to claim 2, characterized in that, The frequency domain tensor is represented as: ; in, This represents the frequency domain tensor. The in-phase component in the channel state information. These are the orthogonal components in the channel state information. This refers to the main frequency data. This indicates the total number of main frequency data. Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the packet index. This indicates the total number of data packet indexes. Represents the imaginary unit. Represents angular frequency. Indicates phase as The complex exponential signal.

5. The human posture perception method based on channel state information according to claim 2, characterized in that, The coefficient matrix is ​​represented as follows: ; in, This represents the amplitude data corresponding to each antenna pair consisting of a transmitting antenna and a receiving antenna. Indicates the scale parameter. Indicates time shift, Describing the wavelet function, Indicates the transmitting antenna. Indicates the receiving antenna. Indicates subcarrier, Indicates the packet index; The determination of the wavelet domain tensor based on the plurality of coefficient matrices includes: Stack the multiple coefficient matrices corresponding to each antenna pair to obtain a stacked matrix; calculate the mean of the stacked matrix along the subcarrier dimension to obtain the sub-tensor corresponding to each antenna pair; The sub-tensors corresponding to each antenna pair are merged to obtain the wavelet domain tensor.

6. The human posture perception method based on channel state information according to claim 1, characterized in that, The target network model includes a feature extraction network, a feature fusion network, an encoder, and a pose regression prediction network; The step of inputting the spatiotemporal tensor, the frequency tensor, and the wavelet tensor into a pre-trained target network model and outputting human pose information includes: The spatiotemporal tensor, the frequency tensor, and the wavelet tensor are input into the feature extraction network, which outputs spatiotemporal features, frequency features, and wavelet features with consistent dimensions. The spatiotemporal domain features, the frequency domain features, and the wavelet domain features are input into the feature fusion network for feature fusion, and the feature fusion tensor is output. The feature fusion tensor is input into the encoder for temporal encoding, and a temporal tensor is output. The temporal tensor is input into the pose regression prediction network to predict pose and output human pose information.

7. The human posture perception method based on channel state information according to claim 6, characterized in that, The posture regression prediction network includes a global regression prediction network and multiple regional posture prediction networks, which are used to predict posture information for different regions of the human body. The step of inputting the temporal tensor into the pose regression prediction network for pose prediction and outputting human pose information includes: The temporal tensor is input into the global regression prediction network for global prediction to obtain the global prediction pose tensor. According to the preset region division rules, the global predicted attitude tensor is divided into regions to obtain sub-global predicted attitude tensors corresponding to multiple regions. After concatenating the temporal tensor with the subglobal prediction pose tensors corresponding to the multiple regions, the concatenation is input into the regional pose prediction network corresponding to the multiple regions for partitioned prediction, thereby obtaining the pose information of the multiple regions. The posture information of the multiple regions is spliced ​​together to output human posture information.

8. The human posture perception method based on channel state information according to claim 7, characterized in that, The target loss function used to train the target network model is expressed as follows: ; in, This represents the weighted mean square error of the key points. Indicates loss of skeletal consistency. Indicates the time-series smoothing loss. , , These represent the balance factors of the corresponding terms.

9. The human posture perception method based on channel state information according to claim 8, characterized in that, The weighted keypoint mean square error is expressed as: ; in, This indicates the size of the sample set used in a single model training iteration. N represents the total number of joints. , This represents the weight values ​​for different key points. Represents the predicted coordinates of the joints. Represents the actual coordinates of the joints. This represents the Euclidean distance between the predicted and actual coordinates of a key point. The loss of skeletal consistency is represented as: ; in, This represents a parent-child joint pair, which includes two joints used to determine the length of a single bone. ,in, This represents the set of parent-child pairs of key nodes. This represents the predicted coordinates of the parent node. Represents the predicted coordinates of the sub-joints. This indicates the calculation of the length of each bone. This indicates the calculation of the variance in length of each bone. The temporal smoothing loss is expressed as: ; in, This indicates the predicted coordinates of the key point in the current iteration. This indicates the predicted coordinates of the same joint point in the next iteration. This represents the distance between two adjacent predicted coordinates of the key point. This represents the smoothing weight.

10. A human posture sensing device based on channel state information, characterized in that, include: The acquisition module is used to acquire channel state information of communication signals used for human posture perception. The preprocessing module is used to preprocess the channel state information to obtain the spatiotemporal domain tensor, frequency domain tensor, and wavelet domain tensor, respectively. The prediction module is used to input the spatiotemporal domain tensor, the frequency domain tensor, and the wavelet domain tensor into a pre-trained target network model and output human pose information; wherein, the target network model is used to perform feature fusion on the spatiotemporal domain tensor, the frequency domain tensor, and the wavelet domain tensor, and perform human pose prediction based on the fused features, and output human pose information.

Citation Information

Cited By

  • Business expansion work order handwritten date identification method and system

    CN121033870A

  • Lightweight user behavior recognition method based on multi-level feature extraction algorithm

    CN121366448A