CSI space-time human body key point detection method combining speed modeling
By combining the CSI space-time human body key point detection method with velocity modeling, the time and space information are extracted using multi-layer space-time modeling module and self-attention mechanism, and the velocity branch is introduced, the problem of key point estimation jump in the existing CSI human body posture estimation method is solved, and a more stable and accurate human body posture estimation is achieved.
Patent Information
- Application Number
- CN202510470199.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-27
AI Technical Summary
The existing CSI human pose estimation method relies on single-frame data and is difficult to capture the timing correlation between continuous actions, resulting in obvious jumps in the skeleton sequence, reducing the stability and reliability of pose estimation.
The CSI space-time human body key point detection method combined with velocity modeling is adopted. By synchronously collecting video data and CSI data, a CSI human body key point timing detection network is established, and time and space information are extracted using multi-layer space-time modeling modules and self-attention mechanisms, and velocity branches are introduced to constrain the displacement and direction of key points.
A more accurate and smoother continuous CSI estimation is achieved, which significantly improves the stability and accuracy of estimation of key points in humans. The generated skeleton sequence is more in line with the natural movement trajectory of the human body, and the posture estimation is smoother and more stable.
Smart Images

Figure CN120220245A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a CSI spatio-temporal human key point detection method combined with speed modeling. Background Art
[0002] Most existing CSI human pose estimation methods mainly rely on single-frame data for independent speculation. This makes it difficult to effectively capture the temporal correlation between consecutive actions during continuous pose estimation. Since adjacent frame data is not fully combined, the generated skeleton sequence may exhibit obvious jumping phenomena and cannot reflect the natural transition of human actions. Especially in scenarios where real-time tracking and monitoring of human poses are required (such as monitoring continuous actions or long-term dynamic postures), this jumping phenomenon will greatly reduce the stability and reliability of pose estimation.
[0003] In addition, existing CSI human key point detection networks based on single-frame estimation have deficiencies in extracting temporal information in human pose estimation methods. They fail to fully exploit the correlation of CSI data in the time dimension and the pose association between frames. This leads to limited accuracy of human pose estimation and unsatisfactory performance when dealing with scenarios with strong continuity and dynamics.
[0004] Meanwhile, traditional CSI human key point detection methods focus on the extraction of static spatial structure features and ignore the temporal dynamic changes of the spatial relationship of the human skeleton during the movement process. As the movement progresses, the relative positions between human key points will continuously change. Therefore, the detection method needs to be able to dynamically learn the changing rules of the human structure in space to more accurately reflect the evolution of human poses. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects and deficiencies in the prior art and provide a CSI spatio-temporal human key point detection method combined with speed modeling. This detection method can achieve more accurate and stable continuous CSI estimation, significantly improving the stability and accuracy of human key point estimation. In addition, this detection method further standardizes the movement trajectory of the human body, making the generated skeleton sequence more conform to the trajectory of natural human movement, thereby achieving smoother and more stable pose estimation.
[0006] To achieve the above purpose, the present invention is realized through the following technical solutions: A CSI spatio-temporal human key point detection method combined with speed modeling, characterized by comprising the following steps:
[0007] First step, synchronously collect video data and CSI data, and perform timestamp alignment operation; annotate the human key point data of the video data, and obtain the CSI data corresponding to the human key point data;
[0008] In the second step, the human key point data and CSI data are divided into a data set and a training set. The sliding window method is used to extract the human key point data and CSI data in the data set and the training set, obtaining T-frame key point skeleton sequence data and corresponding T-frame CSI time series data;
[0009] In the third step, a CSI human key point time series detection network is established; the training set is used to train the CSI human key point time series detection network, obtaining a trained CSI human key point time series detection network;
[0010] The CSI human key point time series detection network adopts a multi-layer spatio-temporal modeling module. Each spatio-temporal modeling module is composed of the fusion of two-branch self-attention mechanisms of "time-space" and "space-time" to obtain human key point features, so as to capture the time correlation and spatial information between consecutive frames of the T-frame CSI time series data;
[0011] Each spatio-temporal modeling module also leads out a speed branch. The results of the speed branches of each spatio-temporal modeling module are added together to fuse and obtain global speed features, realizing the constraint of the displacement and direction of human key points between consecutive frames, making the generated key point skeleton sequence more in line with the natural movement trajectory of the human body, so as to achieve smoother and more stable pose estimation;
[0012] In the fourth step, the trained CSI human key point time series detection network is used to detect human key points, realizing human key point estimation and speed estimation, so as to achieve human pose estimation.
[0013] In the above solution, the present invention can solve the jump problem of key point estimation in the prior art. The present invention takes multi-frame CSI data as input and outputs corresponding multi-frame key points, realizing more accurate and smoother continuous CSI estimation, significantly improving the stability and accuracy of key point estimation. This improvement provides a more reliable solution for tasks such as continuous action monitoring and long-term pose tracking. In addition, the present invention adopts a multi-layer spatio-temporal modeling module. Each spatio-temporal modeling module is composed of the fusion of two-branch self-attention mechanisms of "time-space" and "space-time". Among them, the time self-attention mechanism captures the time correlation between consecutive frames of the CSI signal, extracts the dynamic features of human actions, and avoids the jump of key point estimation, while the spatial multi-head self-attention mechanism learns the spatial structure features of the human skeleton, ensuring that the spatial relationship between key points conforms to the human body structure law. The two branches are connected in different orders to learn time and spatial information with emphasis. At the same time, the spatio-temporal modeling module also leads out a speed branch. The results of the speed branches of each spatio-temporal modeling module are added together to fuse the global speed information. The speed estimation can constrain the displacement and direction of key points between consecutive frames, making the generated skeleton sequence more in line with the natural movement trajectory of the human body, and achieving smoother and more stable pose estimation.
[0014] Specifically, in the second step, the extraction of human key point data and CSI data from the dataset and the training set using the sliding window method to obtain T-frame key point skeleton sequence data and the corresponding T-frame CSI time series data means that:
[0015] The human key point data in the first step is G sample ∈R 17×2 , representing the coordinates of 17 human key points; the CSI data in the first step is a frame of CSI signal X sample ∈R 3×90×5 , representing 3 transmitting antennas, 90 representing 3 receiving antennas multiplied by 30 amplitude data subcarriers, and 5 CSI consecutive sampling data;
[0016] After extracting the data using the sliding window method, T-frame CSI time series data X C ∈R T×3×90×5 and T-frame key point skeleton sequence data G kp ∈R T×17×2 are obtained, where T is the length of the time series.
[0017] In the third step, the CSI human key point time series detection network includes a feature extraction module, several spatio-temporal modeling modules, a velocity decoder, and a key point decoder; several spatio-temporal modeling modules are cascaded and connected to the feature extraction module and the key point decoder respectively; the velocity branches led out by each spatio-temporal modeling module are added and fused and then connected to the velocity decoder.
[0018] The feature extraction module is composed of 3 convolutional modules using the Relu activation function; the T-frame CSI time series data is downsampled by the maximum pooling layer after passing through the first convolutional module, and then downsampled by the following two convolutional modules, and finally expanded by the fully connected layer to obtain the output of the feature extraction module:
[0019]
[0020] Among them, are the outputs after downsampling by each convolutional module for the T-frame CSI time series data X C respectively, and the output of the third convolutional module after downsampling is where J is the number of human key points, J = 17;
[0021] Flatten in the (H,W) dimension and expand the information of the last dimension using the fully connected layer to obtain the output of the feature extraction module:
[0022]
[0023] Among them, dim is the output dimension of the fully connected layer.
[0024] Each of the spatio-temporal modeling modules includes a first branch composed of a time module and a space module connected in sequence, and a second branch composed of a space module and a time module connected in sequence;
[0025] Add the position encoding to different dimensions of the output of the feature extraction module to obtain the input of the spatio-temporal modeling module:
[0026]
[0027] Input the input of the spatio-temporal modeling module into the first branch composed of a time module and a space module connected in sequence and the second branch composed of a space module and a time module connected in sequence respectively, perform self-attention mechanism fusion, and obtain the human key point feature fusion result;
[0028] Among them, F 0 ∈R T×J×dim , is a learnable spatial encoding parameter, is a learnable temporal encoding parameter.
[0029] The space module consists of a spatial multi-head self-attention mechanism, and extracts the spatial features of each time step in T time steps for the input of the space module t represents the t-th time step, t ∈ 1, …, T,;
[0030] Use the self-attention mechanism to obtain 3 vectors in the multi-head attention mechanism
[0031]
[0032] Among them are learnable projection matrices respectively, i represents the i-th spatio-temporal modeling module, i ∈ 1, …, N, t represents the t-th time step, i ∈ 1, …, T,; h represents the h-th number of heads, h ∈ 1, …, H,;
[0033] Finally, obtain the output of the spatial multi-head attention:
[0034]
[0035] Among them, is the projection parameter matrix, d k is the dimension of K s , i ∈ 1, …, N,
[0036] After using the same spatial multi-head attention mechanism for each of the T time steps, the results of the T time steps are stacked and reshaped to return to the original dimension (T, J, dim) and input into a multi-layer perceptron. Then, through residual connection and layer normalization, the output of the final spatial module is obtained.
[0037]
[0038] The calculation process of the entire spatial module is denoted by S i indicating that i represents the i-th spatio-temporal modeling module;
[0039] The time module consists of a time multi-head self-attention mechanism, and the input of the time module is Flatten the number of human key points J and the dimension where dim is located into C flatten dimensions, obtaining
[0040]
[0041] Use the self-attention mechanism to obtain three vectors in the multi-head attention mechanism:
[0042]
[0043] where are learnable projection matrices respectively. i represents the i-th spatio-temporal modeling module, i ∈ 1, …, N; h represents the h-th head number, h ∈ 1, …, H;
[0044] Finally, the output of the time multi-head attention is obtained:
[0045]
[0046] where, is the projection parameter matrix, d k is the dimension of the K T matrix, i ∈ 1, …, N
[0047] After reshaping the output TMHSA of the time multi-head attention back to the input shape (T, J, dim), it is input into a multi-layer perceptron, and then through residual connection and layer normalization, the output of the final time module is obtained
[0048]
[0049] where i ∈ 1, …, N. The entire calculation process of the time module is denoted by T i indicating that i represents the i-th spatio-temporal modeling module.
[0050] Calculate the learnable weight parameters of the first branch and the second branch in the $i$-th spatio-temporal module The results of the two weight parameters are added up to 1, and the calculation formula is as follows:
[0051]
[0052] Among them, $W$ is the learnable parameter matrix, concat represents concatenating the results of the two branches, and the softmax function converts the two weight parameters into a probability distribution so that the sum of the two weights is 1;
[0053] Multiply the weight parameters element-wise with the outputs of the first branch and the second branch to obtain the final branch fusion result $F$ i , $F$ i will also be used as the feature input for the next spatio-temporal modeling module:
[0054]
[0055] Among them, represents the feature output of the first branch composed of the sequential connection of the time module and the space module in the $i$-th spatio-temporal modeling module; represents the feature output of the second branch composed of the sequential connection of the space module and the time module in the $i$-th spatio-temporal modeling module; $F$ i-1 is the feature output of the $(i - 1)$-th spatio-temporal modeling module and also the feature input of the $i$-th spatio-temporal modeling module.
[0056] In each spatio-temporal modeling module, a first velocity branch and a second velocity branch are introduced:
[0057]
[0058] Among them, represents the spatial module feature output of the first branch composed of the sequential connection of the time module and the space module in the $i$-th spatio-temporal modeling module; represents the time module feature output of the second branch composed of the sequential connection of the space module and the time module in the $i$-th spatio-temporal modeling module; $F$ i-1 is the feature output of the $(i - 1)$-th spatio-temporal modeling module and also the feature input of the $i$-th spatio-temporal modeling module;
[0059] Calculate the learnable weight parameters of the first velocity branch and the second velocity branch in the $i$-th spatio-temporal module The results of the two weight parameters are added up to 1, and the calculation formula is as follows;
[0060]
[0061] W M W is a learnable parameter matrix; concat represents concatenating the results of the two velocity branches; the softmax function converts the two weight parameters into a probability distribution such that the sum of the two weights is 1;
[0062] The weight parameters are weighted and fused with the outputs of the first velocity branch and the second velocity branch to finally obtain the velocity feature of the i-th spatio-temporal modeling module:
[0063]
[0064] where, V i ∈R T×J×dim .
[0065] The CSI human key-point time-series detection network is composed of N spatio-temporal modeling modules in N cascades, and N velocity features are obtained. The velocity features are input to the velocity decoder and added together to fuse the velocity information of different time and space scales:
[0066]
[0067] V sum is input to the Transformer Encoder module, and the result is output:
[0068] V feature = TransformerEncoderLayer(V sum )[0,:,:]
[0069] V feature ∈R J×dim ;
[0070] V feature is flattened and input to two fully connected layers for size transformation, and finally the output shape is transformed to the representation shape of the velocity to obtain the velocity estimation result of the human key points:
[0071]
[0072] where, O v1 , O v2 are the intermediate results of the velocity decoder.
[0073] The CSI human key-point time-series detection network is composed of N spatio-temporal modeling modules in N cascades, and the key-point feature is the output F of the last spatio-temporal module N ∈R T×J×dim ;
[0074] Flatten the number J of human body key points and the dimension where dim is located; Flatten F N Input it into two fully connected layers of the key point decoder for size transformation, and finally transform it into the human body key point estimation result:
[0075]
[0076] Among them, O k1 , O k2 is the intermediate result of the key point decoder;
[0077] Calculate the loss function including the key point loss and the velocity loss:
[0078]
[0079] Among them, the true velocity information annotation G speed = G kp [-1,:,:] - G kp [0,:,:], G kp ∈R T×17×2 , is the T-frame key point skeleton sequence data G kp ∈R T×17×2 ; α represents the weight of the velocity information in the loss formula;
[0080] Judge whether the CSI human body key point time series detection network is trained completed according to the loss function.
[0081] The advantages of the CSI spatio-temporal human body key point detection method combining velocity modeling in the present invention are as follows:
[0082] 1. CSI human body key point time series detection network:
[0083] The CSI human body key point time series detection network in the present invention can process multi-frame CSI data, output corresponding multi-frame key points, and achieve more accurate and more stable continuous key point estimation. By learning the key point velocity information, this network effectively reduces the jump phenomenon of key point estimation and improves the stability and accuracy of pose estimation.
[0084] 2. Spatio-temporal modeling module for joint estimation of velocity and key points:
[0085] The spatio-temporal modeling module introduced in the present invention includes two-branch self-attention mechanisms of "time-space" and "space-time", which can simultaneously extract the time and space information in CSI data. The key point features are fused by learnable weights from the final outputs of the two branches; the velocity features are simultaneously derived from the results of the time module of the "time-space" branch and the space module of the "space-time" branch, and are fused by learnable weights. The results of the velocity branches of all spatio-temporal modeling modules are added together to obtain the global velocity feature. The final global velocity feature consists of one layer of Transformer and two layers of fully connected layers, which are used to obtain the final key point estimation and velocity estimation to achieve human pose estimation.
[0086] Through velocity estimation, the present invention can obtain the absolute velocity and direction information of the key point sequence, further standardize the movement trajectory of the human body, make the generated skeleton sequence more conform to the natural movement trajectory of the human body, and thus achieve smoother and more stable pose estimation.
[0087] 3. Key point estimation method with velocity modeling:
[0088] By calculating the difference between the last frame and the first frame of the key point sequence, the absolute velocity and direction information of the key points are obtained, and the trajectory stability of the continuous key point detection task is increased. This method is not only applicable to the key point estimation of CSI data, but also can be applied to the key point estimation of continuous video frames, 2D to 3D key point estimation, and other temporal key point detection tasks, with wide applicability and scalability.
[0089] Compared with the prior art, the present invention has the following advantages and beneficial effects: The CSI spatio-temporal human key point detection method combined with velocity modeling in the present invention can achieve more accurate and smoother continuous CSI estimation, significantly improving the stability and accuracy of human key point estimation. In addition, this detection method further standardizes the movement trajectory of the human body, makes the generated skeleton sequence more conform to the natural movement trajectory of the human body, and thus achieves smoother and more stable pose estimation. Brief Description of the Drawings
[0090] Figure 1 is the structural diagram of the CSI human key point temporal detection network of the present invention;
[0091] Figure 2 is the structural diagram of the feature extraction module in the CSI human key point temporal detection network of the present invention;
[0092] Figure 3 is the structural diagram of a single spatio-temporal modeling module in the CSI human key point temporal detection network of the present invention; Detailed Description of the Invention
[0093] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0094] Embodiment
[0095] As Figures 1 to 3 shown, the CSI spatio-temporal human key point detection method combining speed modeling of the present invention includes the following steps:
[0096] In the first step, video data and CSI data are synchronously collected and timestamp alignment operation is performed; the human key point data of the video data is labeled, and the CSI data corresponding to the human key point data is obtained.
[0097] In this embodiment, 3 seconds of data is collected for each action, and sampling is performed at a rate of 30 frames per second. Therefore, each action contains 90 consecutive time frames. The present invention uses 17 key points of the COCO dataset for labeling. The representation form of human key points can be extended to 18 points and 25 points of OpenPose, as well as three-dimensional single-person pose estimation networks and their representation forms. The present invention estimates the human skeleton information from the video frames aligned with the WiFi-CSI time through the AlphaPose network and provides it to the CSI network for learning.
[0098] In the second step, the human key point data and CSI data are divided into a dataset and a training set, and the sliding window method is used to extract the human key point data and CSI data of the dataset and the training set, obtaining T-frame key point skeleton sequence data and corresponding T-frame CSI time series data.
[0099] This embodiment has 20 participants. Each participant performs 17 actions in 4 rooms, and each action is repeated 20 times, including normal behaviors such as standing and sitting, as well as exercise behaviors such as jumping and push-ups, and abnormal behaviors such as falling. When dividing the dataset, the 20 testers are randomly divided into a dataset and a training set according to a ratio of 4:1.
[0100] In the third step, a CSI human key point time series detection network is established; the training set is used to train the CSI human key point time series detection network, obtaining a trained CSI human key point time series detection network;
[0101] The CSI human key point time series detection network adopts a multi-layer spatio-temporal modeling module. Each spatio-temporal modeling module is composed of the fusion of two-branch self-attention mechanisms of "time-space" and "space-time" to obtain human key point features, so as to capture the time correlation and spatial information between consecutive frames of the T-frame CSI time series data;
[0102] Each spatio-temporal modeling module also leads to a velocity branch. The results of the velocity branches of each spatio-temporal modeling module are added together to fuse and obtain the global velocity feature, which realizes the constraint of the displacement and direction of the human body key points between consecutive frames, making the generated key point skeleton sequence more conform to the natural movement trajectory of the human body, so as to achieve smoother and more stable pose estimation;
[0103] In the fourth step, the trained CSI human body key point time series detection network is used to detect the human body key points, realizing the estimation of the human body key points and the velocity, so as to realize the human body pose estimation.
[0104] In the second step, the extraction of the human body key point data and the CSI data of the data set and the training set by using the sliding window method to obtain the T-frame key point skeleton sequence data and the corresponding T-frame CSI time series data means:
[0105] The human body key point data in the first step is G sample ∈R 17×2 , representing the coordinates of 17 human body key points; the CSI data in the first step is a frame of CSI signal X sample ∈R 3×90×5 , representing 3 transmitting antennas, 90 representing the subcarriers composed of 3 receiving antennas multiplied by 30 amplitude data, and 5 CSI continuous sampling data;
[0106] After extracting the data by using the sliding window method, the T-frame CSI time series data X C ∈R T×3×90×5 and the T-frame key point skeleton sequence data G kp ∈R T×17×2 are obtained, where T is the time series length.
[0107] In the third step, the CSI human body key point time series detection network includes a feature extraction module, several spatio-temporal modeling modules, a velocity decoder, and a key point decoder; several spatio-temporal modeling modules are cascaded and connected to the feature extraction module and the key point decoder respectively; the velocity branches led out by each spatio-temporal modeling module are added and fused and then connected to the velocity decoder.
[0108] Specifically:
[0109] (1) The feature extraction module of the present invention is composed of 3 convolutional modules using the Relu activation function; the T-frame CSI time series data is downsampled by the maximum pooling layer after passing through the first convolutional module, and then downsampled by the latter two convolutional modules, and finally expanded by the fully connected layer to obtain the output of the feature extraction module:
[0110]
[0111] Among them, are the T-frame CSI time series data XC The output after downsampling by each convolutional module layer, and the output of downsampling by the third convolutional module layer is where J is the number of human key points, J = 17;
[0112] Take Flatten in the (H, W) dimension and use a fully connected layer to expand the information of the last dimension to obtain the output of the feature extraction module:
[0113]
[0114] where, dim is the output dimension of the fully connected layer.
[0115] (2) Each of the spatio-temporal modeling modules includes a first branch composed of a time module and a space module connected in sequence, and a second branch composed of a space module and a time module connected in sequence;
[0116] Add positional encoding to different dimensions of the output of the feature extraction module to obtain the input of the spatio-temporal modeling module:
[0117]
[0118] Input the input of the spatio-temporal modeling module into the first branch composed of a time module and a space module connected in sequence and the second branch composed of a space module and a time module connected in sequence respectively, and perform self-attention mechanism fusion to obtain the human key point feature fusion result;
[0119] where, F 0 ∈R T×J×dim , is a learnable spatial encoding parameter, is a learnable temporal encoding parameter.
[0120] (3) The space module consists of a spatial multi-head self-attention mechanism, and extracts the spatial features of each time step in T time steps from the input of the space module where t represents the t-th time step, t ∈ 1,..., T; t represents the t-th time step, t ∈ 1,..., T,;
[0121] Use the self-attention mechanism to obtain 3 vectors in the multi-head attention mechanism
[0122]
[0123] where are learnable projection matrices, where \(i\) represents the \(i\)-th spatio-temporal modeling module, \(i\in\{1,\ldots,N\}\), \(t\) represents the \(t\)-th time step, \(t\in\{1,\ldots,T\}\), and \(h\) represents the \(h\)-th number of heads, \(h\in\{1,\ldots,H\}\).
[0124] Finally, the output of the spatial multi-head attention is obtained:
[0125]
[0126] Among them, is the projection parameter matrix, \(d\) k is the dimension of \(K\) s matrix, \(i\in\{1,\ldots,N\}\).
[0127] After using the same spatial multi-head attention mechanism for \(T\) time steps, the results of the \(T\) time steps are stacked and reshaped to the original dimension \((T, J, \text{dim})\) and then input into a multi-layer perceptron. After that, through residual connection and layer normalization, the output of the final spatial module is obtained.
[0128] The calculation process of the entire spatial module is denoted by \(S\) i where \(i\) represents the \(i\)-th spatio-temporal modeling module.
[0129] (4) The time module consists of a time multi-head self-attention mechanism. The input of the time module is Flatten the number of human key points \(J\) and the dimension where \(\text{dim}\) is located into \(C\) flatten dimensions, obtaining
[0130]
[0131] Use the self-attention mechanism to obtain three vectors in the multi-head attention mechanism:
[0132]
[0133] Among them are learnable projection matrices, where \(i\) represents the \(i\)-th spatio-temporal modeling module, \(i\in\{1,\ldots,N\}\), and \(h\) represents the \(h\)-th number of heads, \(h\in\{1,\ldots,H\}\).
[0134] Finally, the output of the time multi-head attention is obtained:
[0135]
[0136] Among them, is the projection parameter matrix, \(d\) k is the dimension of the \(K\) T matrix, \(i\in\{1,\ldots,N\}\).
[0137] After reshaping the output TMHSA of the temporal multi-head attention back to the input shape (T, J, dim), it is input into a multi-layer perceptron, and then through residual connection and layer normalization, the output of the final temporal module is obtained.
[0138]
[0139] where i ∈ 1, …, N, and the whole process calculated by the temporal module is denoted by T i indicating that i represents the i-th spatio-temporal modeling module;
[0140] (5) Calculate the learnable weight parameters of the first branch and the second branch in the i-th spatio-temporal module as The results of the two weight parameters are added up to 1, and the calculation formula is as follows:
[0141]
[0142] where W is a learnable parameter matrix, and concat represents concatenating the results of the two branches; the softmax function converts the two weight parameters into a probability distribution, making the sum of the two weights equal to 1;
[0143] Multiply the weight parameters element-wise with the outputs of the first branch and the second branch to obtain the final branch fusion result F i , F i will also be used as the feature input for the next spatio-temporal modeling module:
[0144]
[0145] where, represents the feature output of the first branch composed of the sequential connection of the temporal module and the spatial module in the i-th spatio-temporal modeling module; represents the feature output of the second branch composed of the sequential connection of the spatial module and the temporal module in the i-th spatio-temporal modeling module; F i-1 is the feature output of the (i - 1)-th spatio-temporal modeling module and also the feature input of the i-th spatio-temporal modeling module.
[0146] (6) Introduce the first velocity branch and the second velocity branch in each spatio-temporal modeling module:
[0147]
[0148] where, represents the spatial module feature output of the first branch composed of the sequential connection of the temporal module and the spatial module in the i-th spatio-temporal modeling module; Represents the output of the temporal module feature of the second branch formed by sequentially connecting the spatial module and the temporal module in the $i$-th spatio-temporal modeling module; $F$ i-1 Is the feature output of the $(i - 1)$-th spatio-temporal modeling module and also the feature input of the $i$-th spatio-temporal modeling module;
[0149] Calculate the learnable weight parameters of the first velocity branch and the second velocity branch in the $i$-th spatio-temporal module The results of the two weight parameters are added up to 1, and the calculation formula is as follows:
[0150]
[0151] $W$ M $W$ is a learnable parameter matrix; concat represents concatenating the results of the two velocity branches; the softmax function Converts the two weight parameters into a probability distribution so that the sum of the two weights is 1;
[0152] Weight the weight parameters With the outputs of the first velocity branch and the second velocity branch for weighted fusion, and finally obtain the velocity feature of the $i$-th spatio-temporal modeling module:
[0153]
[0154] Where, $V$ i $\in \mathbb{R}$ T×J×dim .
[0155] (7) The CSI human key point temporal detection network is composed of $N$ spatio-temporal modeling modules in $N$-cascade, and a total of $N$ velocity features are obtained. The velocity features are input to the velocity decoder and added together to fuse the velocity information of different time and space scales:
[0156]
[0157] Input $V_{sum}$ into the Transformer Encoder module and output the result:
[0158] $V$ feature $=$ TransformerEncoderLayer($V$ sum )[0, :, :]
[0159] $V$ feature $\in \mathbb{R}$ J×dim ;
[0160] Flatten $V$ feature And input it into two fully connected layers for size transformation, and finally output the shape transformation to the representation shape of the velocity to obtain the velocity estimation result of the human key points:
[0161]
[0162] Among them, O v1 , O v2 is the intermediate result of the speed decoder.
[0163] (8) The CSI human key-point time-series detection network is composed of N spatio-temporal modeling modules in N cascades, and the key-point feature is the output F of the last spatio-temporal module N ∈R T×J×dim ;
[0164] Flatten the number J of human key points and the dimension where dim is located; input F N into two fully connected layers of the key-point decoder for size transformation, and finally transform it into the human key-point estimation result:
[0165]
[0166] Among them, O k1 , O k2 is the intermediate result of the key-point decoder;
[0167] Calculate the loss function including the key-point loss and the speed loss:
[0168]
[0169] Among them, the true speed information annotation G speed corresponding to the T-frame CSI time-series data = G kp [-1,:,:] - G kp [0,:,:], G kp ∈R T×17×2 , is the T-frame key-point skeleton sequence data G kp ∈R T×17×2 ; α represents the weight of the speed information in the loss formula;
[0170] Judge whether the CSI human key-point time-series detection network is trained completed according to the loss function.
[0171] The CSI human key point temporal detection network of the present invention mainly consists of a feature extraction module, a spatio-temporal modeling module, a speed decoder, and a key point decoder. The feature extraction module reduces the dimension of the T-frame CSI time series data through three convolutional layers, extracts the human key point features, and reduces the data dimension, laying a foundation for subsequent processing. The spatio-temporal modeling module is the core part, which is composed of multiple cascaded spatio-temporal modeling modules. Each module contains two branches, namely "time-space" and "space-time", which respectively use the multi-head self-attention mechanism to extract features in the time dimension and the space dimension, and fuse the results of the two branches through learnable weights. At the same time, a speed branch is introduced to learn the motion information of the key points, enhancing the ability to capture dynamic changes. The speed decoder and the key point decoder are responsible for converting the output of the spatio-temporal modeling module into the final human key point position and speed information. Among them, the speed decoder processes the speed features through the Transformer Encoder, and the key point decoder maps the features to the human key point position through the fully connected layer. The loss function combines the key point loss and the speed loss, and uses the mean square error (MSE) to measure the difference between the predicted value and the true value, so as to optimize the network parameters and improve the network performance.
[0172] The advantages of the CSI spatio-temporal human key point detection method combining speed modeling of the present invention are as follows:
[0173] 1. CSI human key point temporal detection network:
[0174] The CSI human key point temporal detection network of the present invention can process multi-frame CSI data, output corresponding multi-frame key points, and achieve more accurate and smoother continuous key point estimation. By learning the key point speed information, the network effectively reduces the jump phenomenon of key point estimation, improving the stability and accuracy of pose estimation.
[0175] 2. Spatio-temporal modeling module for joint estimation of speed and key points:
[0176] The spatio-temporal modeling module introduced in the present invention contains self-attention mechanisms for the "time-space" and "space-time" branches, which can simultaneously extract the time and space information in the CSI data. The key point features are fused by the final outputs of the two branches through learnable weights; the speed features are simultaneously extracted from the results of the time module of the "time-space" branch and the space module of the "space-time" branch, and fused through learnable weights. The speed branch results of all spatio-temporal modeling modules are added together to obtain the global speed feature. The final global speed feature consists of one layer of Transformer and two layers of fully connected layers, which are used to obtain the final key point estimation and speed estimation to achieve human pose estimation.
[0177] Through speed estimation, the present invention can obtain the absolute speed and direction information of the key point sequence, further standardize the human motion trajectory, make the generated skeleton sequence more conform to the natural motion trajectory of the human body, and thus achieve smoother and more stable pose estimation.
[0178] 3. Key point estimation method with speed modeling:
[0179] By calculating the difference between the last frame and the first frame of the key point sequence, the absolute speed and direction information of the key points are obtained, and the trajectory stability of the continuous key point detection task is increased. This method is not only applicable to the key point estimation of CSI data, but also can be applied to the key point estimation of continuous video frames, 2D to 3D key point estimation, and other temporal key point detection tasks, with wide applicability and scalability.
[0180] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A CSI spatiotemporal human key point detection method combined with velocity modeling, characterized by: The following steps are involved: The first step is to synchronously collect video data and CSI data and perform timestamp alignment operations; annotate the human key point data of the video data and obtain the CSI data corresponding to the human key point data; The second step is to divide the human key point data and CSI data into a data set and a training set, and use the sliding window method to extract the human key point data and CSI data of the data set and the training set to obtain T-frame key point skeleton sequence data and the corresponding T-frame CSI time series data; The third step is to establish a CSI human key point timing detection network; use the training set to train the CSI human key point timing detection network to obtain a trained CSI human key point timing detection network; The CSI human key point temporal detection network adopts a multi-layer spatiotemporal modeling module. Each spatiotemporal modeling module is formed by the fusion of the "time-space" and "space-time" two-branch self-attention mechanism to obtain the human key point features to capture the temporal correlation and spatial information between consecutive frames of the T-frame CSI time series data. Each spatiotemporal modeling module also leads to a velocity branch. The velocity branch results of each spatiotemporal modeling module are added and fused to obtain the global velocity feature, so as to constrain the displacement and direction of the key points of the human body between consecutive frames, so that the generated key point skeleton sequence is more consistent with the natural motion trajectory of the human body, so as to achieve smoother and more stable posture estimation; The fourth step is to use the trained CSI human key point timing detection network to detect the human key points, realize human key point estimation and speed estimation, and realize human posture estimation.
2. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 1 is characterized in that: In the second step, the human body key point data and CSI data of the data set and the training set are extracted by the sliding window method to obtain T-frame key point skeleton sequence data and corresponding T-frame CSI time series data: The human key point data in the first step is G sample ∈R 17×2 , represents the coordinates of 17 key points of the human body; the CSI data of the first step is a frame of CSI signal X sample ∈R 3×90×5 , represents 3 transmitting antennas, 90 represents 3 receiving antennas multiplied by 30 subcarriers of amplitude data, and 5 CSI continuous sampling data; After extracting data using the sliding window method, we get T frame CSI time series data X C ∈R T×3×90×5 and T frame key point skeleton sequence data G kp ∈R T×17×2 , where T is the length of the time series.
3. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 1, characterized in that: In the third step, the CSI human key point timing detection network includes a feature extraction module, several spatiotemporal modeling modules, a speed decoder and a key point decoder; A number of spatiotemporal modeling modules are cascaded and connected to the feature extraction module and the key point decoder respectively; the speed branches derived from each spatiotemporal modeling module are added and fused and then connected to the speed decoder.
4. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 3 is characterized by: The feature extraction module consists of three layers of convolution modules using the Relu activation function. The T-frame CSI time series data is downsampled using the maximum pooling layer after passing through the first layer of convolution modules, and then downsampled through the next two layers of convolution modules. Finally, the fully connected layer is expanded to obtain the feature extraction module output: in, are T frames of CSI time series data X C After downsampling the output of each convolution module, the output of the third convolution module downsampling is Where J is the number of key points of the human body, J = 17; Will Flatten in the (H, W) dimension and use the fully connected layer to expand the information of the last dimension, and get the output of the feature extraction module: in, dim is the output dimension of the fully connected layer.
5. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 4 is characterized in that: Each of the spatiotemporal modeling modules includes a first branch consisting of a time module and a space module connected in sequence, and a second branch consisting of a space module and a time module connected in sequence; Add position encoding to feature extraction module output The different dimensions of , get the input of the spatiotemporal modeling module: The input of the spatiotemporal modeling module is respectively input into the first branch consisting of the sequential connection of the time module and the space module and the second branch consisting of the sequential connection of the space module and the time module, and the self-attention mechanism is fused to obtain the fusion result of the key point features of the human body; Among them, F 0 ∈R T×J×dim , is the learnable spatial encoding parameter, is a learnable temporal encoding parameter.
6. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 5, characterized in that: The spatial module is composed of a spatial multi-head self-attention mechanism. Extract the spatial features of each of the T time steps t represents the t-th time step t∈1,…,T,; Use the self-attention mechanism to obtain the three vectors in the multi-head attention mechanism in are the learnable projection matrices, i represents the i-th spatiotemporal modeling module, t∈1,…,N, and t represents the t-th time step t∈1,…,T,; Finally, the output of spatial multi-head attention is obtained: in, is the projection parameter matrix, d k It's K s The dimension of , i∈1,…,N, After using the same spatial multi-head attention mechanism for T time steps, the results of T time steps are stacked and transformed back to the initial dimension (T, J, dim) and input into the multi-layer perceptron. Then, the output of the final spatial module is obtained through residual connection and layer normalization. The calculation process of the entire space module is represented by S i Indicates that i represents the i-th spatiotemporal modeling module; The time module is composed of a time multi-head self-attention mechanism. The input of the time module is Flatten the number of human key points J and the dimension where dim is located into C flatten Dimension, get Use the self-attention mechanism to obtain the three vectors in the multi-head attention mechanism: in They are respectively the learnable projection matrices, i represents the i-th spatiotemporal modeling module, i∈1,…,N,; h represents the h-th head number, h∈1,…,H,; Finally, the output of temporal multi-head attention is obtained: in, is the projection parameter matrix, d k It's K T The dimension of the matrix, i∈1,…,N, The output of the temporal multi-head attention, TMHSA, is transformed back to the input shape (T, J, dim) and then input into the multi-layer perceptron. Then, the output of the final temporal module is obtained through residual connection and layer normalization. Where i∈1,…,N, the entire process of time module calculation is T i Indicates that i represents the i-th spatiotemporal modeling module.
7. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 6, characterized in that: Calculate the learnable weight parameters of the first branch and the second branch in the i-th spatiotemporal module The sum of the two weight parameters is 1, and the calculation formula is as follows: Among them, W is the learnable parameter matrix, concat represents the concatenation of the results of the two branches, and the softmax function The two weight parameters are converted into probability distributions so that the sum of the two weights is 1; The weight parameter Perform element-by-element point multiplication with the output of the first branch and the second branch to obtain the final branch fusion result F i , F i It will also be used as the feature input for the next spatiotemporal modeling module: in, Represents the feature output of the first branch consisting of the sequential connection of the time module and the space module in the i-th spatiotemporal modeling module; represents the feature output of the second branch consisting of the sequential connection of the spatial module and the temporal module in the i-th spatiotemporal modeling module; F i-1 It is the feature output of the i-1th spatiotemporal modeling module and also the feature input of the i-th spatiotemporal modeling module.
8. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 6, characterized in that: In each spatiotemporal modeling module, the first velocity branch and the second velocity branch are derived: in, Represents the spatial module feature output of the first branch consisting of the sequential connection of the time module and the space module in the i-th spatiotemporal modeling module; Represents the temporal module feature output of the second branch consisting of the sequential connection of the spatial module and the temporal module in the i-th spatiotemporal modeling module; F i-1 It is the feature output of the i-1th spatiotemporal modeling module and also the feature input of the i-th spatiotemporal modeling module; Calculate the learnable weight parameters of the first speed branch and the second speed branch in the i-th spatiotemporal module The sum of the two weight parameters is 1, and the calculation formula is as follows: W M W is the learnable parameter matrix; concat represents the concatenation of the results of the two speed branches; the softmax function converts The two weight parameters are converted into probability distributions so that the sum of the two weights is 1; The weight parameter The outputs of the first velocity branch and the second velocity branch are weightedly fused to finally obtain the velocity features of the i-th spatiotemporal modeling module: Among them, V i ∈R T×J×dim .
9. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 8, characterized in that: The CSI human key point timing detection network uses N spatiotemporal modeling modules in N cascades to obtain a total of N speed features. The speed features are input into the speed decoder and added to fuse speed information of different time and space scales: In sum ∈R T×J×dim ; V sum Input into the Transformer Encoder module and output the result: V feature =TransformerEncoderLayer(V sum )[0,:,:] In feature ∈R J×dim ; V feature The flattened input is sent to two fully connected layers for size transformation, and the final output shape is transformed into the shape representing the speed, and the speed estimation results of the key points of the human body are obtained: Q speed ∈R J×2 ; Among them, O v1 ,O v2 is the intermediate result of the velocity decoder.
10. The CSI spatiotemporal human key point detection method combined with velocity modeling according to claim 9, characterized in that: The CSI human key point temporal detection network is composed of N spatiotemporal modeling modules in N cascades, and the key point feature is the output F of the last spatiotemporal module. N ∈R T×J×dim ; Flatten the number of human key points J and the dimension where dim is located; N The two fully connected layers input to the key point decoder are resized and finally transformed into the human key point estimation result: THE kp ∈R T×J×2 ; Among them, O k1 , O k2 is the intermediate result of the key point decoder; Calculate the loss function including key point loss and speed loss: Among them, the real speed information corresponding to the T frame CSI time series data is labeled G speed =G kp [-1,:,:]-G kp [0,:,:],G kp ∈R T×17×2 , is the T-frame key point skeleton sequence data G kp ∈R T×17×2 ; α represents the weight of speed information in the loss formula; The loss function is used to determine whether the CSI human key point timing detection network has been trained.