Human body action recognition method based on decoupling and structured modeling of CSI (Channel State Information)

By combining modal decomposition and structured modeling techniques with capsule networks to decouple action and environmental components, the accuracy and robustness of CSI-based human action recognition are improved, and the consistency and generalization problems of action recognition in complex environments are solved.

CN120929964APending Publication Date: 2025-11-11HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511090257.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing CSI-based human motion recognition technologies struggle to effectively decouple small-scale fading components in complex multipath propagation environments, resulting in poor motion recognition consistency. Furthermore, data-driven methods lack generalization ability across different locations and scenarios.

Method used

By employing modal decomposition and structured modeling techniques, a human motion recognition model is constructed through capsule networks to achieve dual decoupling of motion and environment components. Furthermore, motion features are optimized through master classification, enhancement, and path reconstruction to improve recognition accuracy and robustness.

Benefits of technology

It effectively decouples environmental and action features, improving the accuracy of location-independent human action recognition and its cross-location generalization ability, achieving a recognition accuracy of 97.17%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929964A_ABST
    Figure CN120929964A_ABST
Patent Text Reader

Abstract

The invention discloses a human body action recognition method based on decoupling and structured modeling of CSI (Channel State Information). The method comprises the following steps: 1, collecting CSI action sample data; 2, preprocessing the collected CSI data; 3, completely integrated empirical mode decomposition, adaptive noise and Hilbert transform are introduced to carry out action-environment component dual decoupling on the preprocessed CSI data; 4, constructing a human body action recognition model based on the capsule network; and 5, generating unified action characterization insensitive to position change from the decoupled action signals by using capsule network vectorization coding and a dynamic routing mechanism, and constructing an enhancement path and a reconstruction path to respectively improve inter-class discrimination and structural consistency. According to the method, position-independent human body action recognition is realized based on modal decomposition and structured modeling technologies, action recognition under different position conditions can be effectively adapted, and high accuracy is achieved during new position testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, specifically to a human motion recognition method based on CSI decoupling and structured modeling. Background Technology

[0002] With the rapid development of IoT technology, human motion recognition has been widely applied in smart homes, health monitoring, virtual reality, and other fields. Compared to traditional technologies such as video surveillance, wearable sensors, and radar, Wi-Fi-based human motion recognition technology has attracted researchers' attention due to its advantages such as low cost, no need for additional equipment, and privacy protection. Compared to Signal Strength Indicator (RSSI), Channel State Information (CSI) can provide more stable and richer sensing information and is widely used in the field of Wi-Fi-based human motion recognition.

[0003] However, CSI (Channel Frequency Response) reflects the channel characteristics of a signal during propagation, influenced by multiple paths including direct path and reflections, scattering, and diffraction from environmental objects such as walls, furniture, and human bodies. HAR (Hybrid AR) based on CSI relies on changes in channel characteristics caused by human actions for identification. However, CSI includes not only large-scale fading components (low-frequency gradual decay components) caused by factors such as propagation path distance and environmental layout, but also small-scale fading components (rapid fluctuation components) caused by multipath effects. When the subject's position and scene remain constant, small-scale fading not only severely interferes with human action recognition, but also causes both large-scale and small-scale fading to change simultaneously when the subject is in different positions or scenes due to the combined effects of positional and spatial changes. Even when performing the same action, the channel characteristics will exhibit significantly different dynamic features and temporal patterns, resulting in poor consistency in action representation.

[0004] Currently, researchers mainly explore from two levels: physical model-driven and data-driven. Although significant progress has been made in the field of cross-location human motion recognition, the following challenges remain: (1) In complex indoor multipath propagation environments, CSI signals contain small-scale channel fading caused by human motion and small-scale channel fading caused by the environment, which overlap in time-frequency characteristics. Existing physical model-driven methods based on DFS have demonstrated their effectiveness in decoupling motion and environmental components, but further decoupling of small-scale fading components is needed. (2) Human motion is composed of the coordinated movement of various joint units, resulting in an organic connection between the corresponding global and local features. Due to the lack of a clear mathematical mapping method to fully characterize this structural relationship, existing data-driven methods can effectively extract local, global, or distributed features, but it is difficult to maintain a stable structural relationship between global and local features under different location conditions, which limits the generalization ability and recognition robustness of the model in new locations or new scenarios. Summary of the Invention

[0005] To address the shortcomings of existing methods, this invention proposes a human motion recognition method based on CSI decoupling and structured modeling. This method aims to achieve position-independent human motion recognition based on modal decomposition and structured modeling techniques, thereby effectively adapting to motion recognition under different positional conditions and improving the accuracy of human motion recognition in new positions.

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The human action recognition method based on CSI decoupling and structured modeling of the present invention is characterized by the following steps: Step 1: Collect CSI samples and preprocess them to obtain the preprocessed CSI sample set. ,in, Indicates the first N is a preprocessed CSI sample; N represents the total number of CSI samples. Step 2, for the first A preprocessed CSI sample By performing a dual decoupling of action and environment components, the first... Instantaneous amplitude - instantaneous phase - instantaneous frequency spectrum ; Step 3: Construct a human action recognition model based on capsule networks, consisting of an encoder and a decoder, and then... After processing, the estimated reconstructed tensor is obtained. This allows us to construct the classification loss for the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path ; Step 3.1: The encoder consists of a multi-scale feature extraction module, a capsuled vector coding layer, and a fully connected dynamic routing mechanism layer, and performs... The process is performed to obtain the high-level semantic capsule set of the i-th CSI sample. ,in, express The A high-level semantic capsule, ; The number of high-level semantic capsules; Step 3.2: The decoder consists of a main classification path, an enhancement path, and a reconstruction path, and performs... This process is used to construct the classification loss for the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path ; Step 4: Construct the total loss of the human motion recognition model using equation (7). It is used to train and optimize the human action recognition model until the loss tends to stabilize or the maximum number of training iterations is reached, thereby obtaining the trained human action recognition model to classify CSI action samples. (7) In equation (7), , , These are the weight coefficients for the main classification path, the enhancement path, and the reconstruction path, respectively.

[0007] The human action recognition method based on CSI decoupling and structured modeling described in this invention is characterized in that step 1 is performed as follows: Step 1.1: Select a rectangular area indoors, and use a router as the WiFi signal transmitting device on the outside of the rectangular area, and use two network cards as receiving devices, denoted as ... , ; Step 1.2: Divide the rectangular area evenly into smaller regions. Set the center point of each smaller region as a calibration point. Randomly select several non-calibrated locations within the rectangular area as test points. Repeat this process several times at both the calibration points and the test points. Human-like movements, and through , Collect CSI samples; Step 1.3, , Two CSI sample data points for each action at each location were merged to obtain the final CSI dataset. ,in, Represents the CSI sample set. This represents a set of action category labels, allowing The Middle Each CSI sample is denoted as ,and The number of subcarriers is , , middle Action category labels are denoted as ; Step 1.4, obtain middle After taking the amplitude values ​​of the first subcarrier data, transient outlier removal and smoothing are performed to obtain the second... One preprocessed CSI sample, denoted as... .

[0008] Furthermore, step 2 is performed as follows: Step 2.1: Use the empirical mode decomposition method to analyze the signal. In The subcarrier data is decomposed into multiple scales to obtain The intrinsic mode functions at different frequency scales and a residual signal; Step 2.2, Removal The higher-order intrinsic mode functions corresponding to the noise component and the lower-order intrinsic mode functions corresponding to the environmental component are identified in the intrinsic mode functions. This allows for signal reconstruction of the remaining intrinsic mode functions, yielding the 1st... A preliminary decoupling action reconstruction sample ; Step 2.3: Construct using Hilbert transform middle The complex analytical representation of each subcarrier data is used, and the instantaneous amplitude, instantaneous phase, and instantaneous frequency of each subcarrier data are calculated based on the complex analytical representation. These values ​​are then stacked along the channel dimension to obtain... Instantaneous amplitude-instantaneous phase-instantaneous frequency spectrum .

[0009] Furthermore, step 3.1 is performed as follows: Step 3.1.1: The multi-scale feature extraction module... Multi-scale fusion feature extraction is performed to obtain... Multi-scale fusion features ; Step 3.1.2, the capsuled vector coding layer pair Processing is performed to obtain Normalized primary capsule set ; Step 3.1.3, the fully connected dynamic routing mechanism layer... Perform adaptive aggregation to obtain High-level semantic capsule collection .

[0010] Furthermore, step 3.1.1 is performed as follows: Step 3.1.1.1: Use convolutional layers to... Perform channel compression to obtain the first Low-order structural features ; Step 3.1.1.2: Use standard convolutional layers to... Short-range detail extraction is performed to obtain the first... Local detailed features ; Step 3.1.1.3: Utilize the expansion rate dilated convolution pairs Large-scale structure capture was performed to obtain the first Large-scale structural features ; Step 3.1.1.4, will and After splicing, we get Multi-scale fusion features .

[0011] Furthermore, step 3.1.2 is performed as follows: Step 3.1.2.1, for Perform depthwise separable convolution to obtain The local dynamic structure is obtained, and after mapping the channel dimension of the local dynamic structure, the first... Structural enhancement features ; Step 3.1.2.2, will The channel dimension is reconstructed into vector form to obtain the first... A collection of primary capsules ,in, express The One primary capsule, , This represents the number of primary capsules. Steps 3.1.2.3: [The remaining text appears to be incomplete and requires further context.] Each primary capsule in the set is normalized and compressed to obtain a normalized set of primary capsules. ,in, express The A normalized primary capsule.

[0012] Furthermore, step 3.1.3 is performed as follows: Step 3.1.3.1, will The k-th affine transformation matrix of the j-th normalized primary capsule of the i-th CSI sample Multiply to obtain the j-th normalized primary capsule of the i-th CSI sample. Projection of a high-level semantic capsule ; Step 3.1.3.2: Calculate the soft routing coefficients of the j-th normalized primary capsule for the i-th CSI sample using equation (1). ; , (1) In equation (1), This represents the j-th normalized primary capsule of the i-th CSI sample and the... Attention scores among high-level semantic capsules Let represent the query vector for the i-th CSI sample. This represents the j-th normalized primary capsule of the i-th CSI sample. The key vectors of a high-level semantic capsule; Step 3.1.3.3, using (2) to... and Perform weighted fusion to obtain the first The first CSI sample High-level semantic capsule features ; (2) Step 3.1.3.4, through function pairs Compress to obtain The A high-level semantic capsule .

[0013] Furthermore, step 3.2 is performed as follows: Step 3.2.1, Calculation of the main classification path vector norm Thus, the classification loss can be constructed using equation (3). : (3) In equation (3), Let represent the value of the k-th bit in the one-hot encoded vector of the action category of the i-th CSI sample, that is, when the true action category of the i-th CSI sample is the k-th class, let Otherwise, let ; , These represent two threshold parameters, one for correctly identifying the action category and the other for incorrectly identifying it. This represents the scaling factor when an action category is incorrectly identified. Step 3.2.2: The enhanced path employs a multi-branch embedding mechanism based on spline mapping. Perform high-order nonlinear modeling to obtain the action class prediction probability of the i-th CSI sample. Thus constructing cross-entropy loss ; Step 3.2.3, the reconstructed path pair Processing is performed to obtain the same as Alignment estimation reconstruction tensor Thus, the MSE loss for reconstructed path optimization is constructed. .

[0014] Furthermore, step 3.2.2 is performed as follows: Step 3.2.2.1, using M Groups of sub-paths with different structural dimensions based on spline mapping respectively After processing, M sub-path feature representations are obtained. ,in, This represents the feature representation of the m-th sub-path; Step 3.2.2.2: Use a linear transformation to be learned to... Mapping to M paths, pay attention to the score. Then use the softmax function to... The process is performed to obtain the fusion weights. ,in, This indicates the score for the m-th path. This represents the m-th fusion weight; Step 3.2.2.3: Obtain the multi-granularity structure embedding representation using equation (4). ; (4) Step 3.2.2.4, for Perform channel-level modulation to obtain the first modulation features ; Step 3.2.2.5, will Flattened for the first Global feature representation The input is then processed through a fully connected layer and a softmax function to perform action classification prediction, thus obtaining the predicted action class probability for the i-th CSI sample. ; Step 3.2.2.6: Calculate the cross-entropy loss using equation (5). : (5) In equation (5), express The Middle The predicted probability of a certain action category.

[0015] Furthermore, step 3.2.3 is performed as follows: Step 3.2.3.1, for Perform a linear transformation operation to obtain the latent features of the i-th CSI sample. ; Step 3.2.3.2, will Reconstructed with Using tensors of the same dimension, we obtain the initial structure mapping for the i-th CSI sample. ; Step 3.2.3.3: Using local convolution and sub-pixel upsampling... Reduced to structure tensor And then Perform a cropping operation on the edge area to obtain the result. Alignment estimation reconstruction tensor ; Step 3.2.3.4: Construct the mean squared error loss for reconstructed path optimization using equation (6). : (6).

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention combines mode decomposition and instantaneous feature extraction techniques. By adaptively decomposing the CSI signal using mode decomposition and selecting intermediate frequency intrinsic mode functions for reconstruction, it achieves decoupling of large-scale fading features of the channel caused by the environment from action features. By using Hilbert transform to construct a complex analytical representation of the action signal and extracting its instantaneous features, it further achieves decoupling of small-scale fading features dominated by the environment from action features, thereby effectively improving the accuracy of location-independent human action recognition.

[0017] 2. This invention designs a human action recognition model based on capsule networks. The main classification path vectorizes the spatial relationships of encoded local dynamic units and generates robust unified action representations through dynamic routing. Two auxiliary paths, enhancement and reconstruction, are introduced to improve inter-class discriminability and structural consistency, respectively. The three branches collaboratively optimize structured modeling, discriminative enhancement, and structural consistency, thereby improving cross-location generalization ability. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the implementation of the present invention; Figure 2 This is a diagram of the encoder module of the present invention; Figure 3 This is a diagram of the decoder module of the present invention; Figure 4 This is a floor plan showing the location of the lobby during the data collection process; Figure 5 This is a confusion matrix representing the accuracy of cross-location human motion recognition for the six categories of this invention; Figure 6 This is a graph showing the effect of the number of training calibration points on the accuracy of human motion classification in this invention. Detailed Implementation

[0019] In this embodiment, a human motion recognition method based on CSI decoupling and structured modeling is described, with the system structure as follows: Figure 1 As shown, the procedure is as follows: Step 1: Collect CSI samples and preprocess them to obtain the preprocessed CSI sample set. ,in, Indicates the first One preprocessed CSI sample; Step 1.1: Select a rectangular area indoors, and use a router as the WiFi signal transmitter outside the rectangular area. Use two network cards as receivers, denoted as ... , ; In this embodiment, the lobby of the experimental building is selected as the experimental site. A router is used as the WiFi signal transmitting device in front of the rectangular area, and network cards are used as receiving devices in the back and one side of the rectangular area.

[0020] Step 1.2: Divide the rectangular area evenly into smaller regions. Set the center point of each smaller region as the calibration point. Randomly select several non-calibrated locations within the rectangular area as test points. Repeat this process several times at both the calibration point and the test points. Human-like movements, and through , Collect CSI samples; In this embodiment, 16 location calibration points with a spacing of 1.6m were set in a rectangular area, and 6 random locations were set as test points. The experimenter repeated 6 types of actions (clapping, side kicking, waving, squatting, front kicking, and chest expansion) 10 times at each location.

[0021] Step 1.3, , Two CSI sample data points for each action at each location were merged to obtain the final CSI dataset. , Represents the CSI sample set. This represents a set of action category labels, allowing The Middle Each CSI sample is denoted as ,and The number of subcarriers is , N is the total number of CSI samples. middle Action category labels are denoted as In this embodiment, , .

[0022] Step 1.4, obtain middle After taking the amplitude values ​​of the first subcarrier data, transient outlier removal and smoothing are performed to obtain the second... One preprocessed CSI sample, denoted as... ; In this embodiment, a Hampel filter based on the statistical characteristics of a sliding window is used to identify and remove transient outliers from the amplitude values ​​of 180 subcarriers; a Savitzky-Golay filter is combined for smooth reconstruction, which reduces high-frequency noise while preserving edge trends.

[0023] Step 2, for the first A preprocessed CSI sample By performing a dual decoupling of action and environment components, the first... Instantaneous amplitude - instantaneous phase - instantaneous frequency spectrum ; Step 2.1: Use the empirical mode decomposition method to analyze the signal. In The subcarrier data is decomposed into multiple scales to obtain The intrinsic mode functions at different frequency scales and a residual signal are used in this embodiment. In this embodiment, a fully integrated empirical mode decomposition and adaptive noise method are used to decompose the 180 subcarrier data at multiple scales.

[0024] Step 2.2, Removal The higher-order intrinsic mode functions corresponding to the noise component and the lower-order intrinsic mode functions corresponding to the environmental component are identified in the intrinsic mode functions. This allows for signal reconstruction of the remaining intrinsic mode functions, yielding the 1st... A preliminary decoupling action reconstruction sample In this embodiment, the intrinsic mode function with a center frequency of 1~60Hz is selected for signal reconstruction, which initially achieves the decoupling of the large-scale fading characteristics and action characteristics of the channel caused by the environment in the CSI signal.

[0025] Step 2.3: Construct using Hilbert transform middle The complex analytical representation of each subcarrier data is used, and the instantaneous amplitude, instantaneous phase, and instantaneous frequency of each subcarrier data are calculated based on the complex analytical representation. These values ​​are then stacked along the channel dimension to obtain... Instantaneous amplitude-instantaneous phase-instantaneous frequency spectrum In this embodiment, from The instantaneous amplitude-instantaneous phase-instantaneous frequency spectrum is extracted to further decouple the small-scale fading characteristics of multipath in the environmental channel from the action characteristics.

[0026] Step 3: Construct a human action recognition model based on capsule networks, consisting of an encoder and a decoder. The encoder comprises a multi-scale feature extraction module, a capsuleized vector encoding layer, and a fully connected dynamic routing mechanism layer, as shown below. Figure 2 As shown; the decoder consists of a main classification path, an enhancement path, and a reconstruction path, as follows: Figure 3 As shown; and for The process is performed to obtain the classification loss of the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path .

[0027] Step 3.1, The input is processed in the encoder to obtain the high-level semantic capsule set of the i-th CSI sample. ,in, express The A high-level semantic capsule, ; C represents the number of high-level semantic capsules; in this embodiment, C = 6, corresponding to 6 types of human actions. Step 3.1.1: The multi-scale feature extraction module... Multi-scale fusion feature extraction is performed to obtain... Multi-scale fusion features ; Step 3.1.1.1: Use convolutional layers to... Perform channel compression to obtain the first Low-order structural features In this embodiment, the kernel size of the convolutional layer is... ; Step 3.1.1.2: Use standard convolutional layers to... Short-range detail extraction is performed to obtain the first... Local detailed features In this embodiment, the kernel size of the standard convolutional layer is [size missing]. ; Step 3.1.1.3: Utilize the expansion rate dilated convolution pairs Large-scale structure capture was performed to obtain the first Large-scale structural features In this embodiment, the dilation rate of the void convolution is... The kernel size is ; Step 3.1.1.4, will and After splicing, we get Multi-scale fusion features .

[0028] Step 3.1.2, Capsule-based vector coding layer Processing is performed to obtain Normalized primary capsule set ; Step 3.1.2.1, for Perform depthwise separable convolution to obtain The local dynamic structure is obtained, and after mapping the channel dimension of the local dynamic structure, the first... Structural enhancement features ; Step 3.1.2.2, will The channel dimension is reconstructed into vector form to obtain the first... A collection of primary capsules ,in, express The One primary capsule, , This represents the number of primary capsules. Steps 3.1.2.3: [The remaining text appears to be incomplete and requires further context.] Each primary capsule in the set is normalized and compressed to obtain a normalized set of primary capsules. ,in, express The A normalized primary capsule.

[0029] Step 3.1.3, Fully Connected Dynamic Routing Mechanism Layer Perform adaptive aggregation to obtain High-level semantic capsule collection ; Step 3.1.3.1, will The k-th affine transformation matrix of the j-th normalized primary capsule of the i-th CSI sample Multiply to obtain the j-th normalized primary capsule of the i-th CSI sample. Projection of a high-level semantic capsule ; Step 3.1.3.2: Calculate the soft routing coefficients of the j-th normalized primary capsule for the i-th CSI sample using equation (1). ; , (1) In equation (1), This represents the j-th normalized primary capsule of the i-th CSI sample and the... Attention scores among high-level semantic capsules Let represent the query vector for the i-th CSI sample. This represents the j-th normalized primary capsule of the i-th CSI sample. The key vector of a high-level semantic capsule.

[0030] Step 3.1.3.3, using (2) and Perform weighted fusion to obtain the first The first CSI sample High-level semantic capsule features ; (2) Step 3.1.3.4, through function pairs Compress to obtain The A high-level semantic capsule .

[0031] Step 3.2, Decoder pair The process is performed to obtain the classification loss of the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path ; Step 3.2.1: Calculation of the main classification path vector norm Thus, the classification loss can be constructed using equation (3). : (3) In equation (3), Let represent the value of the k-th bit in the one-hot encoded vector of the action category of the i-th CSI sample, that is, when the true action category of the i-th CSI sample is the k-th class, let Otherwise, let ; , These represent two threshold parameters, one for correctly identifying the action category and the other for incorrectly identifying it. This represents the scaling factor when an action category is incorrectly identified. Step 3.2.2: The enhanced path employs a multi-branch embedding mechanism based on spline mapping. Perform high-order nonlinear modeling to obtain the action class prediction probability of the i-th CSI sample. Thus constructing cross-entropy loss ; Step 3.2.2.1, using MGroups of sub-paths with different structural dimensions based on spline mapping respectively After processing, M sub-path feature representations are obtained. ,in, This represents the feature representation of the m-th sub-path; Step 3.2.2.2: Use a linear transformation to be learned to... Mapping to M paths, pay attention to the score. Then use the softmax function to... The process is performed to obtain the fusion weights. ,in, This indicates the score for the m-th path. This represents the m-th fusion weight. In this embodiment, three sub-paths with different structural dimensions are set, with structural dimensions of 64, 128, and 256 respectively, to capture multi-granularity structural features.

[0032] Step 3.2.2.3: Obtain the multi-granularity structure embedding representation using equation (4). ; (4) Step 3.2.2.4, for Perform channel-level modulation to obtain the first modulation features In this embodiment, the Feature-wise Linear Modulation mechanism is used to embed multi-granularity structures. Channel-level modulation is performed to obtain modulation characteristics. ; Step 3.2.2.5, will Flattened for the first Global feature representation The input is then processed through a fully connected layer and a softmax function to perform action classification prediction, thus obtaining the predicted action class probability for the i-th CSI sample. .

[0033] Step 3.2.2.6: Calculate the cross-entropy loss using equation (5). : (5) In equation (5), express The Middle The predicted probability of a certain action category.

[0034] Step 3.2.3: Reconstruct the path pair Processing is performed to obtain the same as Alignment estimation reconstruction tensor Thus, the mean squared error loss for reconstructed path optimization is constructed. ; Step 3.2.3.1, for Perform a linear transformation operation to obtain the latent features of the i-th CSI sample. ; Step 3.2.3.2, will Reconstructed with Using tensors of the same dimension, we obtain the initial structure mapping for the i-th CSI sample. .

[0035] Step 3.2.3.3: Using local convolution and sub-pixel upsampling... Reduced to structure tensor And then Perform a cropping operation on the edge area to obtain the result. Alignment estimation reconstruction tensor ; Step 3.2.3.4: Construct the MSE loss for reconstructed path optimization using equation (6). : (6) Step 4: Construct the total loss of the human motion recognition model using equation (7). It is used to train and optimize the human action recognition model until the loss tends to stabilize or the maximum number of training iterations is reached, thereby obtaining the trained human action recognition model to classify CSI action samples. (7) In equation (7), , , These are the weight coefficients for the main classification path, the enhancement path, and the reconstruction path, respectively.

[0036] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0037] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0038] In this embodiment, the optimal model is used to classify and predict the CSI data under six actions obtained at six random test points.

[0039] The application effects of this invention will be described in detail below with reference to experiments.

[0040] Experimental conditions: To verify the performance of the proposed system, data was collected and tested in the lobby of an experimental building with an area of ​​7m × 7m. For example... Figure 4 As shown, Using a TL-WDR6500 router, , An Intel 5300 network card was used. All experiments were conducted on a computer equipped with an Intel i9-9700K CPU and an NVIDIA GeForce 3080 GPU.

[0041] Experiment 1: Cross-location recognition experiment The model was trained using data from all calibration points, and tested using data from test points. The experimental results are as follows: Figure 5 As shown. Figure 5 The confusion matrix of the recognition accuracy of the six action categories is presented. The results show that the recognition accuracy of the six action categories all exceed 96%, and the average recognition accuracy reaches 97.17%, indicating that the proposed method has good cross-location robustness.

[0042] Experiment 2: The impact of different numbers of training locations on recognition accuracy Experiments were conducted using samples of 4, 6, 8, 10, and 12 calibration points, respectively. Multiple sampling averages were performed for each training location configuration. The experimental results are shown below. Figure 6 As shown in the figure, the results indicate that even with only four training locations, the model's recognition accuracy did not decrease significantly and still reached over 90%. This demonstrates that the proposed method can effectively overcome the impact of reducing the number of training locations and possesses strong spatial generalization ability.

Claims

1. A human motion recognition method based on CSI decoupling and structured modeling, characterized in that, The procedure is as follows: Step 1: Collect CSI samples and preprocess them to obtain the preprocessed CSI sample set. ,in, Indicates the first N is a preprocessed CSI sample; N represents the total number of CSI samples. Step 2, for the first A preprocessed CSI sample By performing a dual decoupling of action and environment components, the first... Instantaneous amplitude - instantaneous phase - instantaneous frequency spectrum ; Step 3: Construct a human action recognition model based on capsule networks, consisting of an encoder and a decoder, and then... After processing, the estimated reconstructed tensor is obtained. This allows us to construct the classification loss for the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path ; Step 3.1: The encoder consists of a multi-scale feature extraction module, a capsuled vector coding layer, and a fully connected dynamic routing mechanism layer, and performs... The process is performed to obtain the high-level semantic capsule set of the i-th CSI sample. ,in, express The A high-level semantic capsule, ; The number of high-level semantic capsules; Step 3.2: The decoder consists of a main classification path, an enhancement path, and a reconstruction path, and performs... This process is used to construct the classification loss for the main classification path. Enhance the cross-entropy loss of the path. Mean squared error loss of the reconstruction path ; Step 4: Construct the total loss of the human motion recognition model using equation (7). It is used to train and optimize the human action recognition model until the loss tends to stabilize or the maximum number of training iterations is reached, thereby obtaining the trained human action recognition model to classify CSI action samples. (7) In equation (7), , , These are the weight coefficients for the main classification path, the enhancement path, and the reconstruction path, respectively.

2. The human motion recognition method based on CSI decoupling and structured modeling according to claim 1, characterized in that, Step 1 is performed as follows: Step 1.1: Select a rectangular area indoors, and use a router as the WiFi signal transmitting device on the outside of the rectangular area, and use two network cards as receiving devices, denoted as ... , ; Step 1.2: Divide the rectangular area evenly into smaller regions. Set the center point of each smaller region as a calibration point. Randomly select several non-calibrated locations within the rectangular area as test points. Repeat this process several times at both the calibration points and the test points. Human-like movements, and through , Collect CSI samples; Step 1.3, , Two CSI sample data points for each action at each location were merged to obtain the final CSI dataset. ,in, Represents the CSI sample set. This represents a set of action category labels, allowing The Middle Each CSI sample is denoted as ,and The number of subcarriers is , , middle Action category labels are denoted as ; Step 1.4, obtain middle After taking the amplitude values ​​of the first subcarrier data, transient outlier removal and smoothing are performed to obtain the second... One preprocessed CSI sample, denoted as... .

3. The human motion recognition method based on CSI decoupling and structured modeling according to claim 1, characterized in that, Step 2 is performed as follows: Step 2.1: Use the empirical mode decomposition method to analyze the signal. In The subcarrier data is decomposed into multiple scales to obtain The intrinsic mode functions at different frequency scales and a residual signal; Step 2.2, Removal The higher-order intrinsic mode functions corresponding to the noise component and the lower-order intrinsic mode functions corresponding to the environmental component are identified in the intrinsic mode functions. This allows for signal reconstruction of the remaining intrinsic mode functions, yielding the 1st... A preliminary decoupling action reconstruction sample ; Step 2.3: Construct using Hilbert transform middle The complex analytical representation of each subcarrier data is used, and the instantaneous amplitude, instantaneous phase, and instantaneous frequency of each subcarrier data are calculated based on the complex analytical representation. These values ​​are then stacked along the channel dimension to obtain... Instantaneous amplitude-instantaneous phase-instantaneous frequency spectrum .

4. The human motion recognition method based on CSI decoupling and structured modeling according to claim 1, characterized in that, Step 3.1 is performed as follows: Step 3.1.1: The multi-scale feature extraction module... Multi-scale fusion feature extraction is performed to obtain... Multi-scale fusion features ; Step 3.1.2, the capsuled vector coding layer pair Processing is performed to obtain Normalized primary capsule set ; Step 3.1.3, the fully connected dynamic routing mechanism layer... Perform adaptive aggregation to obtain High-level semantic capsule collection .

5. The human motion recognition method based on CSI decoupling and structured modeling according to claim 4, characterized in that, Step 3.1.1 is performed as follows: Step 3.1.1.1: Use convolutional layers to... Perform channel compression to obtain the first Low-order structural features ; Step 3.1.1.2: Use standard convolutional layers to... Short-range detail extraction is performed to obtain the first... Local detailed features ; Step 3.1.1.3: Utilize the expansion rate dilated convolution pairs Large-scale structure capture was performed to obtain the first Large-scale structural features ; Step 3.1.1.4, will and After splicing, we get Multi-scale fusion features .

6. The human motion recognition method based on CSI decoupling and structured modeling according to claim 4, characterized in that, Step 3.1.2 is performed as follows: Step 3.1.2.1, for Perform depthwise separable convolution to obtain The local dynamic structure is obtained, and after mapping the channel dimension of the local dynamic structure, the first... Structural enhancement features ; Step 3.1.2.2, will The channel dimension is reconstructed into vector form to obtain the first... A collection of primary capsules ,in, express The One primary capsule, , This represents the number of primary capsules. Steps 3.1.2.3: [The remaining text appears to be incomplete and requires further context.] Each primary capsule in the set is normalized and compressed to obtain a normalized set of primary capsules. ,in, express The A normalized primary capsule.

7. The human motion recognition method based on CSI decoupling and structured modeling according to claim 4, characterized in that, Step 3.1.3 is performed as follows: Step 3.1.3.1, will The k-th affine transformation matrix of the j-th normalized primary capsule of the i-th CSI sample Multiply to obtain the j-th normalized primary capsule of the i-th CSI sample. Projection of a high-level semantic capsule ; Step 3.1.3.2: Calculate the soft routing coefficients of the j-th normalized primary capsule for the i-th CSI sample using equation (1). ; , (1) In equation (1), This represents the j-th normalized primary capsule of the i-th CSI sample and the... Attention scores among high-level semantic capsules Let represent the query vector for the i-th CSI sample. This represents the j-th normalized primary capsule of the i-th CSI sample. The key vectors of a high-level semantic capsule; Step 3.1.3.3, using (2) to... and Perform weighted fusion to obtain the first The first CSI sample High-level semantic capsule features ; (2) Step 3.1.3.4, through function pairs Compress to obtain The A high-level semantic capsule .

8. The human motion recognition method based on CSI decoupling and structured modeling according to claim 1, characterized in that, Step 3.2 is performed as follows: Step 3.2.1, Calculation of the main classification path vector norm Thus, the classification loss can be constructed using equation (3). : (3) In equation (3), Let represent the value of the k-th bit in the one-hot encoded vector of the action category of the i-th CSI sample, that is, when the true action category of the i-th CSI sample is the k-th class, let Otherwise, let ; , These represent two threshold parameters, one for correctly identifying the action category and the other for incorrectly identifying it. This represents the scaling factor when an action category is incorrectly identified. Step 3.2.2: The enhanced path employs a multi-branch embedding mechanism based on spline mapping. Perform high-order nonlinear modeling to obtain the action class prediction probability of the i-th CSI sample. Thus constructing cross-entropy loss ; Step 3.2.3, the reconstructed path pair Processing is performed to obtain the same as Alignment estimation reconstruction tensor Thus, the MSE loss for reconstructed path optimization is constructed. .

9. A human motion recognition method based on CSI decoupling and structured modeling according to claim 8, characterized in that, Step 3.2.2 is performed as follows: Step 3.2.2.1, using M Groups of sub-paths with different structural dimensions based on spline mapping respectively After processing, M sub-path feature representations are obtained. ,in, This represents the feature representation of the m-th sub-path; Step 3.2.2.2: Use a linear transformation to be learned to... Mapping to M paths, pay attention to the score. Then use the softmax function to... The process is performed to obtain the fusion weights. ,in, This indicates the score for the m-th path. This represents the m-th fusion weight; Step 3.2.2.3: Obtain the multi-granularity structure embedding representation using equation (4). ; (4) Step 3.2.2.4, for Perform channel-level modulation to obtain the first modulation features ; Step 3.2.2.5, will Flattened for the first Global feature representation The input is then processed through a fully connected layer and a softmax function to perform action classification prediction, thus obtaining the predicted action class probability for the i-th CSI sample. ; Step 3.2.2.6: Calculate the cross-entropy loss using equation (5). : (5) In equation (5), express The Middle The predicted probability of a certain action category.

10. A human motion recognition method based on CSI decoupling and structured modeling according to claim 8, characterized in that, Step 3.2.3 is performed as follows: Step 3.2.3.1, for Perform a linear transformation operation to obtain the latent features of the i-th CSI sample. ; Step 3.2.3.2, will Reconstructed with Using tensors of the same dimension, we obtain the initial structure mapping for the i-th CSI sample. ; Step 3.2.3.3: Using local convolution and sub-pixel upsampling... Reduced to structure tensor And then Perform a cropping operation on the edge area to obtain the result. Alignment estimation reconstruction tensor ; Step 3.2.3.4: Construct the mean squared error loss for reconstructed path optimization using equation (6). : (6)。