A CSI-based location-independent human activity recognition method

By using a CSI-based multivariate temporal neural network, combined with contrastive learning and incremental learning, the problem of insufficient position generalization ability in existing technologies is solved. This enables position-independent action recognition and continuous learning of new action categories with limited samples, thereby improving the model's generalization and learning capabilities.

CN116645729BActive Publication Date: 2026-02-03HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310657374.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-05
Publication Date
2026-02-03
Estimated Expiration
2043-06-05

AI Technical Summary

Technical Problem

Existing CSI-based human motion recognition systems have shortcomings in position generalization ability, require a large number of samples for training and have difficulty continuously learning new types of actions. Existing methods often forget old tasks after learning new tasks, and cannot achieve continuous and effective position-independent recognition.

Method used

We employ a CSI-based multivariate temporal graph neural network, combining contrastive learning and incremental learning. By constructing a feature extraction network for the multivariate temporal graph neural network, we capture spatiotemporal dependencies using a hybrid jump propagation layer and an extended perception layer, construct positive samples to enhance the model's generalization ability, and achieve position-independent action recognition and continuous learning of new action categories.

Benefits of technology

It enables the recognition of actions at any indoor location with limited samples of specific locations and categories, and can recognize new categories of actions without retraining the model, thus improving the model's generalization and continuous learning capabilities and reducing the number of samples required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645729B_ABST
    Figure CN116645729B_ABST
Patent Text Reader

Abstract

The application discloses a CSI-based position-independent human activity continuous learning recognition method, which comprises the following steps: 1, collecting CSI action sample data; 2, pre-processing the CSI action sample data; 3, constructing positive samples by randomly scaling the pre-processed samples in the time dimension; 4, constructing a multivariate time graph neural network and extracting CSI action sample features; 5, calculating the similarity between the sample feature values and the positive samples and the feature values of the remaining samples, obtaining a comparison loss, and optimizing the feature extraction network; 6, freezing the feature extraction network, sending the features obtained from the input samples into a classifier for training to obtain a classification model. When the application continuously learns new action categories, the user does not need to retrain the feature extraction network, and the new and old action recognition in any position in the room can be realized by providing limited position new category samples to train the classifier, and the practicability is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of wireless communication, and particularly relates to a position-independent action recognition technology based on a comparative incremental learning graph neural network. BACKGROUND

[0002] In recent years, human action recognition plays an important role in the fields of safety monitoring, entertainment, smart home, etc. Different technologies are adopted in human action recognition systems, such as wearable sensors, radars, computer vision, etc. Compared with other human action recognition technologies, WiFi-based technology has the advantages of low cost, no need to wear devices, and no privacy issues.

[0003] WiFi signals include received signal strength information (RSSI) and channel state information (CSI). RSSI is widely used in WiFi-based human action recognition, and it is the aggregate signal strength of multipath, has the advantages of simplicity and low hardware requirements, but the perception accuracy based on RSSI is low. The CSI signal can extract more rich multipath information from subcarriers, and CSI is based on the physical layer, describes the amplitude and phase characteristics of the channel, and can better reflect the fine-grained channel information.

[0004] When the position of a human changes when collecting actions, the collected WiFi signals are not the same due to the different multipath superposition signals generated at the receiver. Therefore, a problem of system generalization performance is faced in human action recognition based on CSI, that is, the position generalization ability. In practical applications, the position and environment of a human when performing actions are not fixed, and there is a demand for recognizing new category actions. A direct solution is to collect activity samples at all positions for training, and when a new category action appears, new samples need to be collected and the model needs to be retrained with the old samples. However, this requires a lot of time and effort, and the cost is too high. Therefore, a system with strong generalization ability is needed, which only needs limited position samples to recognize actions at any position, and only a small amount of samples can recognize new category actions.

[0005] Currently, the technology to solve this problem is embodied in two aspects. On the one hand, researchers reduce the number of samples required by extracting motion signals separately to reduce the impact of position changes. For example, researchers have developed a cross-domain gesture recognition system Widar3.0 by extracting an environment-independent motion signal BVP from the signal; there are also gesture recognition systems that separate gesture signals from background information using rank reduction and sparse decomposition algorithms; there are also methods that introduce the concept of multi-view in gesture recognition to obtain features independent of target position and direction, thereby achieving one-to-one mapping of gesture and signal features. However, these methods have limitations, including being susceptible to factors such as occlusion and direction, and having a small identifiable area range. The other aspect is to reduce the number of samples required for the target position through transfer learning or meta-learning. First, researchers train the model with limited position data samples, then use transfer learning to migrate the model to the target position, and only use a small number of samples to achieve recognition of the target position action; or collect limited action categories to train the model, and then migrate to the target domain to identify new class actions. However, after learning new activities, the new model's recognition accuracy for previously learned activities drops significantly, which means that the model forgets the old task in order to learn the new task. It is only effective in the current scenario and cannot achieve continuous learning. SUMMARY

[0006] The present application is to solve the above-mentioned deficiencies in the prior art, and proposes a CSI-based position-independent human activity recognition method, which can recognize actions in any position in the room with limited position and limited category action samples as the training set, and only a small number of samples are required for new category actions without retraining the entire network, thereby reducing the number of samples required for recognizing actions in any position in the room and new category actions.

[0007] To achieve the above-mentioned application purposes, the present application adopts the following technical solutions:

[0008] The CSI-based position-independent human activity continuous learning recognition method of the present application is characterized by the following steps:

[0009] Step 1, collection and preprocessing of CSI data:

[0010] Step 1.1, select a rectangular area in the indoor space; use a router as the WIFI signal transmitting device outside the rectangular area, denoted as AP, and the transmitting device AP has TX root transmitting antennas, and use a network card as the receiving device outside the rectangular area, denoted as RP, and the receiving device RP has RX root receiving antennas;

[0011] Step 1.2: Divide the rectangular area into m blocks evenly, and set the center point of each block as a calibration point, thus obtaining a total of m calibration points; randomly select n points from the non-calibrated points in the rectangular area as test points; select m1 positions from the m positions as training positions, and use the remaining m2 positions and n arbitrary positions as test positions, where m = m1 + m2;

[0012] Step 1.3: Perform a type of human action at the j-th training position, and during the execution of the human action, use the receiving device RP to collect CSI signals on different channels transmitted by the transmitting device AP u times, thereby obtaining CSI data for m1 calibration points; wherein, the CSI data of the i-th action performed at the j-th point with dimensions TX×RX×Δ×T is taken as a single sample, denoted as . Where T is the sequence length of a single sample, Δ is the number of subcarriers, j∈[1,m1], i∈[1,a];

[0013] Step 1.4: Select a single sample The CSI data of the sub-th subcarrier between the t-th transmit antenna and the r-th receive antenna at time t is denoted as CSI. tx,rx,sub,t The CSI data at time point t on the sub-th subcarrier between the first transmitting antenna and the rx-th receiving antenna. 1,rx,sub,t Divide by the CSI data at time point t of the sub-th subcarrier between the second transmit antenna and the rx-th receive antenna. 2,rx,sub,t Then, the CSI quotient data of a single sample after removing carrier frequency offset and sampling frequency offset is obtained; the amplitude and phase of a single sample are obtained from the CSI quotient data; tx∈[1,TX]; rx∈[1,RX]; sub∈[1,Δ]; t∈[1,T];

[0014] The residual error of the phase on each subcarrier in the CSI quotient data of the single sample is removed by linear transformation, so that the phase after error removal is concatenated with the amplitude of the single sample in the dimension of subcarrier and a single CSI sample with dimension RX×2Δ×T is formed.

[0015] Based on the different numbers of the receiving antennas, a single CSI sample is divided into RX CSI samples, and the k-th CSI sample performing the i-th action at the j-th point is denoted as .

[0016] Step 1.5: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Randomly generate an index between 1 and T, and select the first Q indices, sort them in ascending order to obtain the sorted index sequence. Then, use the sorted index sequence to analyze the k-th CSI sample. Perform Q samplings to obtain the k-th sample.

[0017] Step 1.6: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Sampling is performed at T / Q intervals to obtain the k-th CSI sample.

[0018] Step 1.7, and After splicing, the k-th sample is obtained.

[0019] Step 2: Establish a feature extraction network based on a multivariate temporal graph neural network, including: a graph learning module, a temporal convolution module, and a graph convolution module.

[0020] Step 2.1: Construct the adjacency matrix A using the graph learning module;

[0021] Step 2.2: Use a convolutional layer with a 1×1 kernel to concatenate the k-th sample. Projecting onto the latent space yields the k-th hidden state.

[0022] Step 2.3: The temporal convolution module includes two dilated perceptron layers, one of which is followed by a tangent hyperbolic activation function layer; the other dilated perceptron layer is followed by a sigmoid activation function layer; the dilated perceptron layer contains several one-dimensional dilated convolutional layers with different kernel sizes.

[0023] The kth hidden state The input is processed by the temporal convolution module, and after passing through two dilated perceptron layers and their corresponding activation function layers, the tangent hyperbolic activation feature vector is obtained. and gate vector Activate the eigenvectors of the tangent hyperbola and gate vector Multiply them to obtain the k-th time gait feature.

[0024] Step 2.4: The graph convolution module consists of two hybrid skip propagation layers, each containing a graph convolutional layer and a multilayer perceptron. One of the graph convolutional layers in the hybrid skip propagation layer uses the adjacency matrix A to perform feature aggregation on the input features, while the other graph convolutional layer uses the transpose A of the adjacency matrix A. T Perform feature aggregation on the input features;

[0025] The input graph convolution module processes the data through two hybrid skip propagation layers, resulting in two graph convolution features. and Convolution features of two graphs and The k-th spatiotemporal fusion feature is obtained by adding them together.

[0026] Step 3: Iteratively train the network using supervised contrastive loss:

[0027] Step 3.1: Integrate spatiotemporal features according to The corresponding original spatiotemporal fusion features and The corresponding data augmentation spatiotemporal fusion features are split, and the set of spatiotemporal fusion features of all categories at all training locations after splitting is denoted as X;

[0028] Step 3.2: Use equation (1) to calculate the loss function L. sup :

[0029]

[0030] In equation (1), |X| represents the total number of features of set X, x q P(q) represents the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature in set X; P(q) is the sum of the original spatiotemporal fusion features or data-augmented spatiotemporal fusion features excluding the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature x. q In addition to and with x q A set of original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of the same category, where |P(q)| represents the total number of features in set P(q), and y p P(q) represents the p-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to P(q); O(q) represents the feature excluding the q-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature y. q The set of all original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of all other categories of the same type, |O(q)| represents the total number of features in the set |O(q)|, z oO(q) represents the o-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to O(q); τ is the temperature coefficient;

[0031] Step 3.3: Train the feature extraction network using gradient descent and calculate the loss function L. sup To update the network parameters until the loss function L is reached. sup The process continues until convergence, thus obtaining a well-trained feature extraction network.

[0032] Step 4: Construct a classifier consisting of two fully connected layers and a ReLU layer in between. When training the classifier on the CSI dataset obtained based on the action categories at m1 training locations, freeze the network parameters of the trained feature extraction network and calculate the cross-entropy loss function to update the classifier parameters until the cross-entropy loss function converges, thereby obtaining the human action classification model. After preprocessing the action sample data at (m2+n) test locations, input the CSI data into the human action classification model to obtain the corresponding classification results.

[0033] Step 5: When learning a new type of action b, repeat the process of step 1 to collect and process CSI data. Then, update the classifier together with the CSI data of the previous old type a according to the process of step 4. This will result in a human action classification model that can identify new and old actions and is independent of location. This will enable the classification and recognition of CSI data corresponding to the new and old type action sample data at (m2+n) test locations.

[0034] The present invention provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the human activity recognition method, and the processor is configured to execute the program stored in the memory.

[0035] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the human activity recognition method.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] 1. This invention proposes a novel structural solution for position-independent action recognition. By using contrastive learning to enhance the generalization ability of the model and combining it with incremental learning, an effective learning mechanism is formed, which makes full use of perceptual information and provides a powerful condition for realizing position-independent action recognition and continuously learning new types of actions.

[0038] 2. This invention proposes to construct positive samples from both temporal and spatial perspectives: temporally, the samples are randomly scaled in the time dimension to construct positive samples that are different but of the same category; spatially, labels are used to classify actions of the same category at different locations as positive samples. Thus, the model can fully extract the common features of similar actions in the action samples, while highlighting the differences between categories. This improves the model's ability to generalize to different locations and its ability to continuously learn new types of actions. As a result, the feature extraction module does not need to be retrained, and the entire model can have excellent performance when recognizing new and old actions at any location.

[0039] 3. This invention uses a multivariate time-plot neural network to fully utilize the potential spatial dependencies between variable pairs in a single sample. It uses a hybrid jump propagation layer and an extended perception layer to capture the spatiotemporal dependencies in the time series, thereby fully extracting the features of multivariate time series such as CSI action data, effectively fusing spatiotemporal features, and providing effective support for subsequent comparative learning of features. Attached Figure Description

[0040] Figure 1 This is a system structure diagram of the present invention;

[0041] Figure 2 This is a model structure diagram of the present invention;

[0042] Figure 3 This is a structural diagram of the temporal convolution module of the present invention;

[0043] Figure 4 This is a structural diagram of the graph convolution module of the present invention. Detailed Implementation

[0044] In this embodiment, as Figure 1 As shown, a location-independent human activity continuous learning recognition method based on CSI is performed according to the following steps:

[0045] Step 1: Collection and preprocessing of CSI data:

[0046] Step 1.1: Select a rectangular area in the indoor space; use a router as the WIFI signal transmitting device outside the rectangular area, denoted as AP, and the transmitting device AP has TX transmitting antennas; use a network card as the receiving device outside the rectangular area, denoted as RP, and the receiving device RP has RX receiving antennas.

[0047] In this embodiment, the AP uses a TL-WDR6500 router with two transmitting antennas; the RP uses an Intel 5300 network card with three receiving antennas.

[0048] Step 1.2: Divide the rectangular area into m blocks evenly, and set the center point of each block as a calibration point, thus obtaining a total of m calibration points; randomly select n points from the non-calibrated points in the rectangular area as test points; select m1 positions from the m positions as training positions, and use the remaining m2 positions and n arbitrary positions as test positions, where m = m1 + m2.

[0049] In this embodiment, the rectangular area is evenly divided into 16 blocks, resulting in 16 calibration points; 6 points are randomly selected from the non-calibration point locations as test points; 12 calibration points are selected from the 16 locations for training, and the remaining 4 locations and 6 other locations are randomly distributed as test locations.

[0050] Step 1.3: Perform a type of human action at the j-th training position; and during the execution of the human action, use the receiving device RP to collect the CSI signals on different channels sent by the transmitting device AP u times, thereby obtaining CSI data for m1 calibration points; among them, the CSI data of the i-th action performed at the j-th point with dimensions TX×RX×Δ×T is taken as a single sample, denoted as . Where T is the sequence length of a single sample, Δ is the number of subcarriers, j∈[1,m1], i∈[1,a];

[0051] In this embodiment, a total of 6 actions were performed, with each action being sampled 10 times consecutively. The sampling rate was 500 packets per second, and the sampling lasted for 5 seconds, resulting in a T size of approximately 2500. The number of subcarriers Δ was 30, and the operation was performed on a computer equipped with an Intel i9-9700K CPU and an NVIDIA GeForce 3080 GPU.

[0052] Step 1.4: Select a single sample The CSI data of the sub-th subcarrier between the t-th transmit antenna and the r-th receive antenna at time t is denoted as CSI. tx,rx,sub,t The CSI data at time point t on the sub-th subcarrier between the first transmitting antenna and the rx-th receiving antenna. 1,rx,sub,t Divide by the CSI data at time point t of the sub-th subcarrier between the second transmit antenna and the rx-th receive antenna. 2,rx,sub,t Then, the CSI quotient data of a single sample after removing carrier frequency offset and sampling frequency offset is obtained; the amplitude and phase of a single sample are obtained from the CSI quotient data; tx∈[1,TX]; rx∈[1,RX]; sub∈[1,Δ]; t∈[1,T];

[0053] By removing the residual phase error on each subcarrier in the CSI quotient data of a single sample through linear transformation, the phase after error removal is concatenated with the amplitude of the single sample in the subcarrier dimension to form a single CSI sample with dimension RX×2Δ×T.

[0054] Based on the different numbers of the receiving antennas, a single CSI sample is divided into RX CSI samples, and the k-th CSI sample performing the i-th action at the j-th point is denoted as .

[0055] Step 1.5: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Randomly generate an index between 1 and T, and select the first Q indices, sort them in ascending order to obtain the sorted index sequence. Then, use the sorted index sequence to analyze the k-th CSI sample. Perform Q samplings to obtain the k-th sample.

[0056] Step 1.6: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Sampling is performed at T / Q intervals to obtain the k-th CSI sample.

[0057] Step 1.7, and After splicing, the k-th sample is obtained.

[0058] Step 2, as follows Figure 2 As shown, a feature extraction network based on a multivariate temporal graph neural network is established, including: a graph learning module, a temporal convolution module, and a graph convolution module.

[0059] Step 2.1: Construct the adjacency matrix A using the graph learning module;

[0060] In this embodiment, the number of nodes corresponding to the number of subcarriers of the sample is first input, and the node embedding is initialized. The embedding of the i-th node is denoted as E. i Let E be the embedding of the j-th node. j In equation (1), θ i Let θ be the network parameter of the i-th node, and θ in equation (2) j Let M be the network parameter of the j-th node; α is a hyperparameter controlling the saturation rate of the activation function, set to 0.5; the i-th hidden feature obtained by equation (1) is denoted as M. i The j-th hidden feature obtained through equation (2) is denoted as M. j ; The adjacency matrix A is obtained through equation (3). i,j A i,jThis refers to the relationship between corresponding nodes i and j in the adjacency matrix;

[0061] M i =tanh(αE) i θ i (1)

[0062] M j =tanh(αE) j θ j (2)

[0063]

[0064] Retain the k largest values ​​of the current node in the adjacency matrix, set other values ​​to zero, and reduce the number of neighboring nodes; then set the current neighbor node A... i,j diagonal A j,i Setting the parameters to zero reduces the computational cost of subsequent graph convolutions while allowing the model parameters to adapt to changes with new training data, thus learning the final graph structure adjacency matrix A.

[0065] Step 2.2: Use a convolutional layer with a 1×1 kernel to concatenate the k-th sample. Projecting onto the latent space yields the k-th hidden state.

[0066] Step 2.3, as follows Figure 3 As shown, the temporal convolution module includes two dilated perceptron layers. One dilated perceptron layer is followed by a tangent hyperbolic activation function layer, and the other dilated perceptron layer is followed by a sigmoid activation function layer. The dilated perceptron layer contains several one-dimensional dilated convolutional layers with different kernel sizes. In this embodiment, there are 4 one-dimensional dilated convolutional layers with kernel sizes of 1×2, 1×3, 1×6, and 1×7.

[0067] The kth hidden state The input is processed by the temporal convolution module, and after passing through two dilated perceptron layers and their corresponding activation function layers, the tangent hyperbolic activation feature vector is obtained. and gate vector Activate the eigenvectors of the tangent hyperbola and gate vector Multiply them to obtain the k-th time gait feature.

[0068] Step 2.4, as follows Figure 4As shown, the graph convolutional module consists of two hybrid skip propagation layers, each containing a graph convolutional layer and a multilayer perceptron. One of the graph convolutional layers in the hybrid skip propagation layer uses the adjacency matrix A to aggregate the input features, while the other graph convolutional layer uses the transpose A of the adjacency matrix A. T Perform feature aggregation on the input features;

[0069] The input graph convolution module processes the data through two hybrid skip propagation layers, resulting in two graph convolution features. and Convolution features of two graphs and The k-th spatiotemporal fusion feature is obtained by adding them together.

[0070] Step 3: Iteratively train the network using supervised contrastive loss:

[0071] Step 3.1: Integrate spatiotemporal features according to The corresponding original spatiotemporal fusion features and The corresponding data augmentation spatiotemporal fusion features are split, and the set of spatiotemporal fusion features of all categories at all training locations after splitting is denoted as X;

[0072] Step 3.2: Use equation (4) to calculate the loss function L. sup :

[0073]

[0074] In equation (4), |X| represents the total number of features of set X, x q P(q) represents the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature in set X; P(q) is the sum of the original spatiotemporal fusion features or data-augmented spatiotemporal fusion features excluding the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature x. q In addition to and with x q A set of original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of the same category, where |P(q)| represents the total number of features in set P(q), and y p P(q) represents the p-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to P(q); O(q) represents the feature excluding the q-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature y. q The set of all original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of all other categories of the same type, |O(q)| represents the total number of features in the set |O(q)|, z oThis represents the o-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to O(q); τ is the temperature coefficient; in this embodiment, τ is taken as 0.1.

[0075] Step 3.3: Train the feature extraction network using gradient descent and calculate the loss function L. sup To update the network parameters until the loss function L sup The process continues until convergence, thus obtaining a well-trained feature extraction network.

[0076] Step 4: Construct a classifier consisting of two fully connected layers and a ReLU layer in between. When training the classifier on the CSI dataset obtained based on the action categories at m1 training locations, freeze the network parameters of the trained feature extraction network and calculate the cross-entropy loss function to update the classifier parameters until the cross-entropy loss function converges, thus obtaining the human action classification model. After preprocessing the action sample data at (m2+n) test locations, input the CSI data into the human action classification model to obtain the corresponding classification results.

[0077] Step 5: When learning a new type of action b, repeat the process in Step 1 to collect and process CSI data. Then, update the classifier together with the CSI data of the previous type a old action according to the process in Step 4. This will result in a human action classification model that can identify both new and old actions and is independent of location. After preprocessing the sample data of new and old action types at (m2+n) test locations, classify and identify the human action classification model corresponding to the corresponding CSI data to obtain the corresponding classification results.

[0078] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0079] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

Claims

1. A location-independent, continuous learning method for recognizing human activities based on CSI, characterized in that, The procedure is as follows: Step 1: Collection and preprocessing of CSI data: Step 1.1: Select a rectangular area in the indoor space; use a router as the WIFI signal transmitting device outside the rectangular area, denoted as AP, and the transmitting device AP has TX transmitting antennas; use a network card as the receiving device outside the rectangular area, denoted as RP, and the receiving device RP has RX receiving antennas. Step 1.2: Divide the rectangular area into m blocks evenly, and set the center point of each block as a calibration point, thus obtaining a total of m calibration points; randomly select n points from the non-calibrated points in the rectangular area as test points; select m1 positions from the m positions as training positions, and use the remaining m2 positions and n arbitrary positions as test positions, where m = m1 + m2; Step 1.3: Perform a type of human action at the j-th training position, and during the execution of the human action, use the receiving device RP to collect CSI signals on different channels transmitted by the transmitting device AP u times, thereby obtaining CSI data for m1 calibration points; wherein, the CSI data of the i-th action performed at the j-th point with dimensions TX×RX×Δ×T is taken as a single sample, denoted as . Where T is the sequence length of a single sample, Δ is the number of subcarriers, j∈[1, m1], i∈[1, a]; Step 1.4: Select a single sample The CSI data of the sub-th subcarrier between the t-th transmit antenna and the r-th receive antenna at time point t is denoted as CSI. tx,rx,sub,t The CSI data at time point t on the sub-th subcarrier between the first transmitting antenna and the rx-th receiving antenna. 1,rx,sub,t Divide by the CSI data at time point t of the sub-th subcarrier between the second transmit antenna and the rx-th receive antenna. 2,rx,sub,t Then, the CSI quotient data of a single sample after removing carrier frequency offset and sampling frequency offset is obtained; the amplitude and phase of a single sample are obtained from the CSI quotient data; tx∈[1,TX]; rx∈[1,RX]; sub∈[1,Δ]; t∈[1,T]; The residual error of the phase on each subcarrier in the CSI quotient data of the single sample is removed by linear transformation, so that the phase after error removal is concatenated with the amplitude of the single sample in the dimension of subcarrier and a single CSI sample with dimension RX×2Δ×T is formed. Based on the different numbers of the receiving antennas, a single CSI sample is divided into RX CSI samples, and the k-th CSI sample performing the i-th action at the j-th point is denoted as . Step 1.5: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Randomly generate an index between 1 and T, and select the first Q indices, sort them in ascending order to obtain the sorted index sequence. Then, use the sorted index sequence to analyze the k-th CSI sample. Perform Q samplings to obtain the k-th sample. Step 1.6: Perform the i-th action on the k-th CSI sample at the j-th point after denoising. Sampling is performed at T / Q intervals to obtain the k-th CSI sample. Step 1.7, and After splicing, the k-th sample is obtained. Step 2: Establish a feature extraction network based on a multivariate temporal graph neural network, including: a graph learning module, a temporal convolution module, and a graph convolution module. Step 2.1: Construct the adjacency matrix A using the graph learning module; Step 2.2: Use a convolutional layer with a 1×1 kernel to concatenate the k-th sample. Projecting onto the latent space yields the k-th hidden state. Step 2.3: The temporal convolution module includes two dilated perceptron layers, one of which is followed by a tangent hyperbolic activation function layer; the other dilated perceptron layer is followed by a sigmoid activation function layer; the dilated perceptron layer contains several one-dimensional dilated convolutional layers with different kernel sizes. The kth hidden state The input is processed by the temporal convolution module, and after passing through two dilated perceptron layers and their corresponding activation function layers, the tangent hyperbolic activation feature vector is obtained. and gate vector Activate the eigenvectors of the tangent hyperbola and gate vector Multiply them to obtain the k-th time gait feature. Step 2.4: The graph convolution module consists of two hybrid skip propagation layers, each containing a graph convolutional layer and a multilayer perceptron. One of the graph convolutional layers in the hybrid skip propagation layer uses the adjacency matrix A to perform feature aggregation on the input features, while the other graph convolutional layer uses the transpose A of the adjacency matrix A. T Perform feature aggregation on the input features; The input graph convolution module processes the data through two hybrid skip propagation layers, resulting in two graph convolution features. and Convolution features of two graphs and The k-th spatiotemporal fusion feature is obtained by adding them together. Step 3: Iteratively train the network using supervised contrastive loss: Step 3.1: Integrate spatiotemporal features according to The corresponding original spatiotemporal fusion features and The corresponding data augmentation spatiotemporal fusion features are split, and the set of spatiotemporal fusion features of all categories at all training locations after splitting is denoted as X; Step 3.2: Use equation (1) to calculate the loss function L. sup : In equation (1), |X| represents the total number of features of set X, x q P(q) represents the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature in set X; P(q) is the sum of the original spatiotemporal fusion features or data-augmented spatiotemporal fusion features excluding the q-th original spatiotemporal fusion feature or data-augmented spatiotemporal fusion feature x. q In addition to and with x q A set of original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of the same category, where |P(q)| represents the total number of features in set P(q), and y p P(q) represents the p-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to P(q); O(q) represents the feature excluding the q-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature y. q The set of all original spatiotemporal fusion features or data-augmented spatiotemporal fusion features of all other categories of the same type, |O(q)| represents the total number of features in the set |O(q)|, z o O(q) represents the o-th original spatiotemporal fusion feature or data-enhanced spatiotemporal fusion feature in the feature set corresponding to O(q); τ is the temperature coefficient; Step 3.3: Train the feature extraction network using gradient descent and calculate the loss function L. sup To update the network parameters until the loss function L is reached. sup The process continues until convergence, thus obtaining a well-trained feature extraction network. Step 4: Construct a classifier consisting of two fully connected layers and a ReLU layer in between. When training the classifier on the CSI dataset obtained based on the action categories at m1 training locations, freeze the network parameters of the trained feature extraction network and calculate the cross-entropy loss function to update the classifier parameters until the cross-entropy loss function converges, thereby obtaining the human action classification model. After preprocessing the action sample data at (m2+n) test locations, input the CSI data into the human action classification model to obtain the corresponding classification results. Step 5: When learning a new type of action b, repeat the process of step 1 to collect and process CSI data. Then, update the classifier together with the CSI data of the previous old type a according to the process of step 4. This will result in a human action classification model that can identify new and old actions and is independent of location. This will enable the classification and recognition of CSI data corresponding to the new and old type action sample data at (m2+n) test locations.

2. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the human activity continuous learning and recognition method of claim 1, and the processor is configured to execute the program stored in the memory.

3. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when run by the processor, executes the steps of the continuous learning and recognition method for human activity as described in claim 1.

Citation Information

Patent Citations

  • Gesture recognition method based on small samples

    CN114818864A

  • Gesture recognition and position classification combined deep learning method

    CN115331311A