Human motion recognition method based on infrared and WiFi

By combining the feature extraction and fusion method of infrared video signals and WiFi signals, the problem of insufficient human body movement recognition performance in the absence of light environment is solved, and higher recognition accuracy and stability are achieved.

CN117058759BActive Publication Date: 2025-05-13NORTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311030061.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-16
Publication Date
2025-05-13
Estimated Expiration
2043-08-16

AI Technical Summary

Technical Problem

In the absence of light environment, it is difficult for the prior art to effectively recognize human movements. Infrared video provides low-resolution images while WiFi signals are rough and depend on the surrounding environment, resulting in poor HAR performance.

Method used

The human body movement recognition method combined with infrared and WiFi is adopted, and the characteristics of infrared video signals and WiFi signals are extracted respectively through the dual-stream convolutional residual network and the bidirectional long and short-term memory network, and feature fusion and reconstruction are combined with deep neural network and deep typical correlation analysis network, and finally action classification is performed through the support vector machine.

Benefits of technology

It effectively solves the problem of insufficient performance of human body movement recognition in the unilluminated environment, improves the accuracy of human body movement recognition based on infrared data, and significantly improves the accuracy of human body movement recognition in the unilluminated environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058759B_ABST
    Figure CN117058759B_ABST
Patent Text Reader

Abstract

The present application relates to a human motion recognition method combining infrared and WiFi. The method combines infrared and WiFi data that are not affected by light, fully considers the heterogeneous differences between infrared video and WiFi signals, and designs a dual-branch network to extract discriminant features from different modalities respectively; in order to make full use of the complementarity of multimodal information, a feature fusion method based on subspace projection is adopted to fuse the features of different modalities and perform discriminant analysis; the method nonlinearly projects the multimodal features into a common space, and performs the final action classification through a support vector machine. The method of the present application effectively solves the problem of human motion recognition under insufficient light conditions, optimizes the performance of human motion recognition based on infrared data, and further improves the accuracy of human motion recognition in a non-lighting environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of motion recognition, and in particular to a method for human motion recognition combining infrared and WiFi. Background Art

[0002] Human action recognition (HAR) is an important research topic in the fields of intelligent transportation, intelligent interactive systems, medical monitoring, etc. Most HAR methods focus on applications in visible light environments. However, in some real-world scenarios, such as nighttime video surveillance systems and public safety applications, it is difficult to collect visible light images or videos in dim or dark scenes.

[0003] Infrared (IR) videos based on thermal imaging are widely considered to be an effective candidate for achieving HAR tasks in the absence of light. However, IR videos usually provide low-resolution images without color and texture information, which may lead to poor HAR performance. Therefore, HAR in lightless environments remains a huge challenge faced today. Meanwhile, WiFi signals are considered to be another reliable option for achieving HAR in environments with poor lighting conditions and have attracted widespread attention in recent years. For HAR tasks, WiFi has the following advantages over other wireless signals: 1) the target does not require wearable devices; 2) the popularity of smart home wireless communication services means that WiFi signals can be easily collected at a very cheap price. However, in practice, since WiFi signals are rough and highly dependent on the surrounding environment, the current HAR performance using only a single WiFi signal has been unsatisfactory. Summary of the invention

[0004] In order to overcome at least one deficiency in the prior art, the present application provides a human motion recognition method combining infrared and WiFi.

[0005] In a first aspect, a method for human motion recognition combining infrared and WiFi is provided, comprising:

[0006] Get the infrared video signal and WiFi signal of the human action to be recognized:

[0007] A two-stream convolutional residual network is used to extract infrared features of infrared video signals, and a bidirectional long short-term memory network is used to extract WiFi features of WiFi signals;

[0008] A deep neural network is used to extract the nonlinear features of infrared features and WiFi features, which are recorded as the first modal features and the second modal features.

[0009] The first modal feature and the second modal feature are input into a deep canonical correlation analysis network, and a reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determines the reconstructed first modal feature and the reconstructed second modal feature when the sum of the reconstruction errors of the two modalities is minimized;

[0010] Calculate the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determine the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum;

[0011] Mapping the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to the common space to obtain a mapping matrix;

[0012] The mapping matrix is ​​input into the linear machine learning classifier SVM to obtain the recognition result of the human action to be recognized.

[0013] In one embodiment, a two-stream convolutional residual network is used to extract infrared features of infrared video signals, and a bidirectional long short-term memory network is used to extract WiFi features of WiFi signals, including:

[0014] Multiple single-frame images of infrared video signals and superimposed optical flow frames are input into a two-stream convolutional residual network to obtain the infrared features of the infrared video signals;

[0015] The CSI information of the WiFi signal is preprocessed, and the preprocessed CSI information is input into a bidirectional long short-term memory network to obtain the WiFi features of the WiFi signal.

[0016] In one embodiment, the specific implementation function of the deep canonical correlation analysis network adopts the following formula:

[0017]

[0018] Constraints:

[0019]

[0020]

[0021]

[0022] Where N is the number of samples, i is the sample number, f(X) is the first modal feature, X is the infrared feature, g(Y) is the second modal feature, Y is the WiFi feature, and U is the CCA direction of all output units of the deep neural network, U = [u1, ...u l ...,u L ],u l is the CCA direction of the lth output unit of the deep neural network, L is the number of output units in the deep neural network, V is the CCA direction of all output units of the deep neural network, V = [v1, ...v m …, v L ],v m is the CCA direction of the mth output unit of the deep neural network, the directions of U and V are different, λ is the trade-off parameter, f(x i ) is the infrared feature x corresponding to the i-th sample i The corresponding first modal feature, p is the first reconstruction nonlinear function, g(y i ) is the WiFi feature y corresponding to the i-th sample i The corresponding second modal feature, q is the second reconstruction nonlinear function; I is the unit matrix, r x and r y is the regularization parameter.

[0023] In one embodiment, the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum are mapped to the common space to obtain a mapping matrix, using the following formula:

[0024]

[0025] Among them, d is the mapping matrix, is the parameter vector of the deep canonical correlation analysis network corresponding to the jth modal feature when the correlation is maximum, z jk is the kth sample of the jth mode, n j is the number of samples of the jth mode, D jk is the mapping value of the kth sample of the jth mode.

[0026] In a second aspect, a human motion recognition device combining infrared and WiFi is provided, comprising:

[0027] Signal acquisition module, used to obtain infrared video signals and WiFi signals of human movements to be identified:

[0028] A first feature extraction module is used to extract infrared features of infrared video signals using a two-stream convolutional residual network and to extract WiFi features of WiFi signals using a bidirectional long short-term memory network;

[0029] A second feature extraction module is used to respectively extract nonlinear features of infrared features and WiFi features using a deep neural network, which are recorded as first modal features and second modal features;

[0030] A feature reconstruction module, used for inputting the first modal feature and the second modal feature into a deep canonical correlation analysis network, wherein a reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determines the reconstructed first modal feature and the reconstructed second modal feature when the sum of the reconstruction errors of the two modalities is minimized;

[0031] A network parameter determination module, used to calculate the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determine the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum;

[0032] A mapping module, used to map the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to a common space to obtain a mapping matrix;

[0033] The recognition module is used to input the mapping matrix into the linear machine learning classifier SVM to obtain the recognition result of the human action to be recognized.

[0034] In one embodiment, the first feature extraction module is further used to:

[0035] Multiple single-frame images of infrared video signals and superimposed optical flow frames are input into a two-stream convolutional residual network to obtain the infrared features of the infrared video signals;

[0036] The CSI information of the WiFi signal is preprocessed, and the preprocessed CSI information is input into a bidirectional long short-term memory network to obtain the WiFi features of the WiFi signal.

[0037] In one embodiment, the specific implementation function of the deep canonical correlation analysis network adopts the following formula:

[0038]

[0039] Constraints:

[0040]

[0041]

[0042]

[0043] Where N is the number of samples, i is the sample number, f(X) is the first modal feature, X is the infrared feature, g(Y) is the second modal feature, Y is the WiFi feature, and U is the CCA direction of all output units of the deep neural network, U = [u1, ...u l ...,u L ],u l is the CCA direction of the lth output unit of the deep neural network, L is the number of output units in the deep neural network, V is the CCA direction of all output units of the deep neural network, V = [v1, ...v m ..., v L ],v m is the CCA direction of the mth output unit of the deep neural network, the directions of U and V are different, λ is the trade-off parameter, f(x i ) is the infrared feature x corresponding to the i-th sample i The corresponding first modal feature, p is the first reconstruction nonlinear function, g(y i ) is the WiFi feature y corresponding to the i-th sample i The corresponding second modal feature, q is the second reconstruction nonlinear function; I is the unit matrix, r x and r y is the regularization parameter.

[0044] In one embodiment, the specific implementation function of the mapping module is expressed by the following formula:

[0045]

[0046] Among them, d is the mapping matrix, is the parameter vector of the deep canonical correlation analysis network corresponding to the jth modal feature when the correlation is maximum, z jk is the kth sample of the jth mode, n j is the number of samples of the jth mode, D jk is the mapping value of the kth sample of the jth mode.

[0047] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned human motion recognition method combining infrared and WiFi is implemented.

[0048] In a fourth aspect, a computer program product is provided, including a computer program / instruction. When the computer program / instruction is executed by a processor, the above-mentioned human motion recognition method combining infrared and WiFi is implemented.

[0049] Compared with the prior art, the present application has the following beneficial effects: the method of the present application combines infrared and WiFi data that are not affected by light, fully considers the heterogeneous differences between infrared video and WiFi signals, and designs a dual-branch network to extract discriminant features from different modalities respectively; in order to make full use of the complementarity of multimodal information, a feature fusion method based on subspace projection is adopted to fuse the features of different modalities and perform discriminant analysis; the method nonlinearly projects the multimodal features into a common space, and performs the final action classification through a support vector machine. The method of the present application effectively solves the problem of human action recognition under insufficient light conditions, optimizes the performance of human action recognition based on infrared data, and further improves the accuracy of human action recognition in a non-lighting environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The present application may be better understood by referring to the following description given in conjunction with the accompanying drawings, which together with the following detailed description are included in this specification and form a part of this specification. In the drawings:

[0051] Figure 1 A flowchart of a method for human motion recognition combining infrared and WiFi according to an embodiment of the present application is shown;

[0052] Figure 2 A structural block diagram of a human motion recognition device combining infrared and WiFi according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0053] The exemplary embodiments of the present application will be described below in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many implementation-specific decisions can be made in the process of developing any such actual embodiments in order to achieve the specific goals of the developer, and these decisions may vary from embodiment to embodiment.

[0054] It is also necessary to explain here that, in order to avoid obscuring the present application due to unnecessary details, only the device structure closely related to the scheme according to the present application is shown in the drawings, while other details that are not closely related to the present application are omitted.

[0055] It should be understood that the present application is not limited to the described implementation forms due to the following description with reference to the accompanying drawings. In this article, where feasible, the embodiments can be combined with each other, features between different embodiments can be replaced or borrowed, and one or more features can be omitted in one embodiment.

[0056] This application fully considers the characteristics of infrared video (optical flow information) and WiFi signal (time domain and frequency domain information), adopts different feature extraction strategies, and obtains typical features. In addition, infrared features and WiFi features are very different and cannot be directly used for discriminant analysis. Generally speaking, feature fusion technology includes early fusion and late fusion. The former is feature-based and usually simply connects the representation of multimodal features after feature extraction. The latter fuses the prediction scores obtained in the last layer of each single-modal network. The prediction score only represents the abstract semantic prediction of the network and lacks the fusion of medium and low-level features between different modes. Therefore, this application adopts an unsupervised feature fusion mechanism that can explore nonlinear relationships between data. It adds an autoencoder to generate reconstructed features to optimize the learned feature representation; it maps the optimized feature representation to a common space and obtains better expression capabilities through learning.

[0057] The present application embodiment provides a method for human motion recognition combining infrared and WiFi. Figure 1 A flowchart of a method for human motion recognition combining infrared and WiFi according to an embodiment of the present application is shown. Figure 1 , methods include:

[0058] Step S1, obtaining infrared video signals and WiFi signals of human motion to be identified. Here, an infrared camera and a WiFi acquisition device can be used to synchronously acquire infrared video signals and WiFi signals of human motion.

[0059] Step S2, using a two-stream convolutional residual network to extract infrared features of infrared video signals, and using a bidirectional long short-term memory network to extract WiFi features of WiFi signals.

[0060] Specifically, multiple single-frame images of the infrared video signal and the superimposed optical flow frame are input into a two-stream convolutional residual network to obtain the infrared features of the infrared video signal, where the superimposed optical flow frame is obtained based on the infrared video signal; the CSI information of the WiFi signal is preprocessed, and the preprocessed CSI information is input into a bidirectional long short-term memory network to obtain the WiFi features of the WiFi signal; here, the preprocessing may be to remove outliers.

[0061] In step S3, a deep neural network is used to extract nonlinear features of infrared features and WiFi features respectively, which are recorded as first modal features and second modal features.

[0062] Here, the infrared signature X, X = [x1, ..., x N ] and WiFi features Y, Y = [y1, ..., y N ] as the features of the two modes, N is the number of samples, and the observation results are formed into pairs, expressed as (x1, y1), ..., (xN ,y N ), x N is the infrared feature corresponding to the Nth sample, y N is the WiFi feature corresponding to the Nth sample. Let f represent the multi-dimensional mapping implemented by the deep neural network (DNN), a depth of K f The multidimensional mapping f of the DNN is realized in the form of A nested mapping where W j is the weight parameter of the jth layer, j = 1, ..., K f , f j is the mapping of the jth layer. The activation function can include Sigmoid, Tanh, ReLU, etc. The set of learnable parameters in DNN is represented as W f , for example in the case of DNN

[0063] The infrared features and WiFi features are respectively input into the deep neural network to obtain nonlinear features, which are recorded as the first modal features f(X) and the second modal features g(Y).

[0064] Step S4, inputting the first modal feature and the second modal feature into the deep canonical correlation analysis network, the reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determines the reconstructed first modal feature and the reconstructed second modal feature when the sum of the reconstruction errors of the two modalities is minimized.

[0065] Here, the deep canonical correlation analysis network is improved by adding a separation autoencoder (SplitAE). The deep canonical correlation analysis network includes two independent reconstruction networks p and q for reconstructing the first modal features and the second modal features. The deep canonical correlation analysis network and the separation autoencoder are used to construct the sum of the reconstruction errors of the two modalities, and determine the reconstructed first modal features and the reconstructed second modal features when the sum of the reconstruction errors of the two modalities is minimized. The specific implementation function of the deep canonical correlation analysis network adopts the following formula:

[0066]

[0067] Constraints:

[0068]

[0069]

[0070]

[0071] Where N is the number of samples, i is the sample number, f(X) is the first modal feature, X is the infrared feature, g(Y) is the second modal feature, Y is the WiFi feature, and U is the CCA direction of all output units of the deep neural network, U = [u1, ...u l ...,u L ],u l is the CCA direction of the lth output unit of the deep neural network, L is the number of output units in the deep neural network, V is the CCA direction of all output units of the deep neural network, V = [v1, ...v m …, v L ],v m is the CCA direction of the mth output unit of the deep neural network, the directions of U and V are different, λ is the trade-off parameter, f(x i ) is the infrared feature x corresponding to the i-th sample i The corresponding first modal feature, p is the first reconstruction nonlinear function, g(y i ) is the WiFi feature y corresponding to the i-th sample i The corresponding second modal feature, q is the second reconstruction nonlinear function; I is the unit matrix, r x and r y is the regularization parameter.

[0072] Step S5, calculate the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determine the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature.

[0073] The above process can be expressed by the following formula:

[0074]

[0075] in, is the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, is the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum, that is, the two optimal solutions. θ1 is the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature, and θ2 is the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature.

[0076] Step S6, mapping the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to the common space, to obtain a mapping matrix. Here, the common space is a set of coordinate systems for data representation.

[0077] Specifically, the following formula can be used:

[0078]

[0079] Among them, d is the mapping matrix, is the parameter vector of the deep canonical correlation analysis network corresponding to the jth modal feature when the correlation is maximum, z jk is the kth sample of the jth mode, n j is the number of samples of the jth mode, D jk is the mapping value of the kth sample of the jth mode.

[0080] Step S7, input the mapping matrix into a linear machine learning classifier SVM to obtain a recognition result of the human action to be recognized.

[0081] Based on the same inventive concept as the human motion recognition method combining infrared and WiFi, this embodiment also provides a corresponding human motion recognition device combining infrared and WiFi. Figure 2 FIG. 1 shows a structural block diagram of a human motion recognition device combining infrared and WiFi according to an embodiment of the present application, see Figure 2 , the device comprises:

[0082] The signal acquisition module 21 is used to acquire the infrared video signal and WiFi signal of the human body action to be identified:

[0083] A first feature extraction module 22, for extracting infrared features of infrared video signals using a two-stream convolutional residual network, and extracting WiFi features of WiFi signals using a bidirectional long short-term memory network;

[0084] A second feature extraction module 23 is used to respectively extract nonlinear features of infrared features and WiFi features using a deep neural network, which are recorded as first modal features and second modal features;

[0085] A feature reconstruction module 24 is used to input the first modal feature and the second modal feature into a deep canonical correlation analysis network, and a reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determines the reconstructed first modal feature and the reconstructed second modal feature when the sum of the reconstruction errors of the two modalities is minimized;

[0086] A network parameter determination module 25, used to calculate the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determine the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum;

[0087] A mapping module 26, used to map the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to a common space to obtain a mapping matrix;

[0088] The recognition module 27 is used to input the mapping matrix into a linear machine learning classifier SVM to obtain a recognition result of the human body action to be recognized.

[0089] The human motion recognition device combining infrared and WiFi of this embodiment has the same inventive concept as the human motion recognition method combining infrared and WiFi mentioned above. Therefore, the specific implementation method of the device can be seen in the embodiment part of the human motion recognition method combining infrared and WiFi mentioned above, and its technical effect corresponds to the technical effect of the above method, which will not be repeated here.

[0090] In order to further verify the effectiveness of the method of the present application, a comparative experiment was conducted on the performance of the method of the present application and the existing methods in human action recognition.

[0091] The data of infrared and WiFi signals are collected using sensor devices. A human action dataset (IR-WI10) is constructed, including infrared video and CSI of WiFi. The dataset is obtained in indoor home scenes without considering lighting conditions. Ten coarse-grained daily actions are designed: answering the phone, drinking water, turning a book, squatting, sitting, stretching, sweeping the floor, taking off a coat, putting on a coat, and wiping a chair. The duration of each action is 5s, and each action is performed 30 times by 10 different volunteers in a fixed position. During the acquisition process, when collecting WiFi signals, the infrared video of the corresponding action is recorded according to the agreed synchronization trigger clock. This dataset is collected from 10 subjects and contains 10 action categories.

[0092] IR-WI10 dataset: It contains infrared video information and its corresponding WiFi signal CSI feature data; IR10 (infrared video) dataset: It only contains infrared video information in the IR-WI10 dataset; WI10 (WiFi) dataset: It only contains WiFi signal CSI information in the IR-WI10 dataset.

[0093] Table 1 compares the detailed parameter information of the infrared dataset (IR10) collected in this experiment with several reference datasets. It can be seen that most of the parameters of the IR10 dataset are similar to the standards of other datasets. It is worth noting that the IR10 dataset has a larger number of samples and a longer video length than other infrared datasets. Table 2 shows the specific information of the collected IR-WI10 dataset.

[0094] Table 1 Parameter comparison of some infrared action datasets and IR10

[0095] Dataset InfAR IITR-IAR IR10 Total number of videos 600 1470 3000 Action Category 12 21 10 Resolution 293*256 1024*768 320*240 Frame rate 25 30 25 Average length (seconds) 4 2~8 5~10

[0096] Table 2 IR-WI10 data and specific information

[0097] Dataset Number of samples Sample category Sample feature dimension IR10 (infrared video) 3000 10 4096 WI10(WiFi) 3000 10 512 TR-WI10 3000 10 (IR4096+WiFi512)

[0098] According to the characteristics of the two data sources, considering the heterogeneity gap between infrared video and WiFi signals, this application designs a dual-branch network to extract the identification features of different data sources respectively. The network can effectively extract the optical flow information of infrared video and the time-frequency information of WiFi signals. The infrared and WiFi features of 3000*4096 and 3000*512 dimensions are extracted using a dual-stream convolutional network and a bidirectional LSTM network respectively.

[0099] In order to verify the hypothesis that WiFi signals can improve the performance of human action recognition based on infrared (IR) video, a series of comparative experiments were conducted. First, action recognition was achieved using infrared and WiFi single data samples, and the corresponding features were extracted. Then, the two modal features were fused for classification. Table 3 shows the results of the feature fusion experiment. It can be seen that the accuracy of single-modal feature recognition is low, and the accuracy based on CSI signals is even lower. The main reason is the noise sensitivity of WiFi signals, while infrared data often carries more identification information than CSI signals. After using CSI signals, the infrared recognition accuracy of the model increased from 81% to 87.2%. It further proves that multimodal feature fusion can improve the performance of HAR.

[0100] Table 3 Results of feature fusion experiments

[0101] Modal characteristics Feature fusion (yes / no) Accuracy (%) WiFi no 39.8 IR no 81.00 lR+WiFi Yes, SVM 83.80 IR+WiFi Yes, Softmax 84.00 IR+WiFi Yes, CCA-Based 87.2

[0102] In this experiment, the accuracy (ACC) performance indicator is used to evaluate the effectiveness of the proposed method. ACC can be described as: the total number of samples collected in the experiment is recorded as N, for one action video sample i, its true action is y i , the action of the sample is predicted as y through the feature fusion mechanism of this application i ′, the accuracy calculation formula is:

[0103]

[0104] Where δ(y i ,y i ′) is a function used to evaluate the actual action y i and the predicted action y i ′. If y i andi ′ refers to the same action, then δ(y i , y′ i )=1, if it is not the same action, δ(y i , y′ i )=0.

[0105] Since the feature extraction methods of IR and WiFi are completely different, they cannot be directly used for discriminant analysis. There is a heterogeneity gap between the two features. In order to eliminate this difference and make full use of the information complementarity of multi-source data, a feature fusion method based on subspace projection is adopted. The features of different data sources are fused and discriminant analysis is performed. The model adds an autoencoder to dynamically adjust the relationship between the input features and the reconstructed features, nonlinearly project the multi-source features into the common space, and output the optimal projection vector representation. Finally, the support vector machine classifier is used to realize the classification of the common space mapping representation.

[0106] In order to verify the effect of the multimodal feature fusion mechanism, several CCA-based methods were used for comparative experiments. The experimental conclusions are shown in Table 4. In the CCA algorithm, PCA is first used for dimensionality reduction, but the results of the first dimensionality reduction are mixed together and difficult to distinguish. In order to reduce the nonlinear dimension, t-distributed random neighbor embedding (T-SNE) is used. By using T-SNE to represent high-dimensional data sets in two-dimensional or three-dimensional space, data visualization and secondary dimensionality reduction are achieved. By comparing the feature representation images of DCCA and Conv-DCCA algorithms after T-SNE operation, it can be seen that DCCA has better performance. This application also performs T-SNE dimensionality reduction operation, and its results are better than other multimodal methods. In addition, experiments are also conducted using multi-view algorithms such as MCCA and GMCCA. The final experimental results show that the multimodal data features extracted by this application for HAR are most suitable for the proposed multimodal feature fusion method.

[0107] Table 4 Recognition effect based on CCA algorithm

[0108] Multi-view algorithm MCCA GMCCA GMLDA GMMFA MvD Accuracy (%) 25.70 33.70 30.20 33.00 64.00 Multimodal Algorithms Conv-DCCA CCA TKMv DCCA This application Accuracy (%) 79.00 87.20 90.50 91.80 92.40

[0109] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned human motion recognition method combining infrared and WiFi is implemented.

[0110] An embodiment of the present application provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the above-mentioned human motion recognition method combining infrared and WiFi is implemented.

[0111] In summary, this application has the following technical effects:

[0112] This application combines infrared and WiFi data that are not affected by light, fully considering the heterogeneous differences between infrared video and WiFi signals, and designs a dual-branch network to extract discriminant features from different modalities respectively; in order to fully utilize the complementarity of multimodal information, a feature fusion method based on subspace projection is adopted to fuse the features of different modalities and perform discriminant analysis; this method nonlinearly projects multimodal features into a common space, and performs final action classification through a support vector machine. The method of this application effectively solves the problem of human action recognition under insufficient light conditions, optimizes the performance of human action recognition based on infrared data, and further improves the accuracy of human action recognition in a non-lighting environment.

[0113] The above are only various implementations of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A human motion recognition method combining infrared and WiFi, characterized in that: include: Get the infrared video signal and WiFi signal of the human action to be recognized: A two-stream convolutional residual network is used to extract infrared features of the infrared video signal, and a bidirectional long short-term memory network is used to extract WiFi features of the WiFi signal; A deep neural network is used to extract nonlinear features of the infrared feature and the WiFi feature respectively, which are recorded as a first modal feature and a second modal feature; Inputting the first modal feature and the second modal feature into a deep canonical correlation analysis network, wherein a reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determining the reconstructed first modal feature and the reconstructed second modal feature when the sum of reconstruction errors of the two modalities is minimized; Calculating the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determining the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum; Mapping the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to a common space to obtain a mapping matrix; The mapping matrix is ​​input into a linear machine learning classifier SVM to obtain a recognition result of the human body action to be recognized.

2. The method according to claim 1, characterized in that in, A two-stream convolutional residual network is used to extract infrared features of the infrared video signal, and a bidirectional long short-term memory network is used to extract WiFi features of the WiFi signal, including: Inputting multiple single-frame images and superimposed optical flow frames of the infrared video signal into the dual-stream convolutional residual network to obtain infrared features of the infrared video signal; The CSI information of the WiFi signal is preprocessed, and the preprocessed CSI information is input into the bidirectional long short-term memory network to obtain the WiFi features of the WiFi signal.

3. The method according to claim 1, characterized in that The specific implementation function of the deep canonical correlation analysis network adopts the following formula: Constraints: Where N is the number of samples, i is the sample number, f(X) is the first modal feature, X is the infrared feature, g(Y) is the second modal feature, Y is the WiFi feature, and U is the CCA direction of all output units of the deep neural network, U = [u1,…u l …,u L ],u l is the CCA direction of the lth output unit of the deep neural network, L is the number of output units in the deep neural network, V is the CCA direction of all output units of the deep neural network, V=[v1,…v m …,v L ],v m is the CCA direction of the mth output unit of the deep neural network, the directions of U and V are different, λ is the trade-off parameter, f(x i ) is the infrared feature x corresponding to the i-th sample i The corresponding first modal feature, p is the first reconstruction nonlinear function, g(y i ) is the WiFi feature y corresponding to the i-th sample i The corresponding second modal feature, q is the second reconstruction nonlinear function; I is the unit matrix, r x and r y is the regularization parameter.

4. The method according to claim 1, characterized in that in, The parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum are mapped to the common space to obtain a mapping matrix using the following formula: Where d is the mapping matrix, is the parameter vector of the deep canonical correlation analysis network corresponding to the jth modal feature when the correlation is maximum, z jk is the kth sample of the jth mode, n j is the number of samples of the jth mode, D jk is the mapping value of the kth sample of the jth mode.

5. A human motion recognition device combining infrared and WiFi, characterized in that: include: Signal acquisition module, used to obtain infrared video signals and WiFi signals of human movements to be identified: A first feature extraction module, configured to extract infrared features of the infrared video signal using a two-stream convolutional residual network, and extract WiFi features of the WiFi signal using a bidirectional long short-term memory network; A second feature extraction module, used to respectively extract nonlinear features of the infrared feature and the WiFi feature using a deep neural network, which are recorded as a first modal feature and a second modal feature; A feature reconstruction module, used for inputting the first modal feature and the second modal feature into a deep canonical correlation analysis network, wherein a reconstruction network in the deep canonical correlation analysis network reconstructs the first modal feature and the second modal feature, and determines the reconstructed first modal feature and the reconstructed second modal feature when the sum of the reconstruction errors of the two modalities is minimized; A network parameter determination module, used to calculate the correlation between the reconstructed first modal feature and the reconstructed second modal feature, and determine the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum; A mapping module, used for mapping the parameter vector of the deep canonical correlation analysis network corresponding to the first modal feature when the correlation is maximum, and the parameter vector of the deep canonical correlation analysis network corresponding to the second modal feature when the correlation is maximum to a common space to obtain a mapping matrix; The recognition module is used to input the mapping matrix into a linear machine learning classifier SVM to obtain a recognition result of the human body action to be recognized.

6. The device according to claim 5, characterized in that The first feature extraction module is further used for: Inputting multiple single-frame images and superimposed optical flow frames of the infrared video signal into the dual-stream convolutional residual network to obtain infrared features of the infrared video signal; The CSI information of the WiFi signal is preprocessed, and the preprocessed CSI information is input into the bidirectional long short-term memory network to obtain the WiFi features of the WiFi signal.

7. The device according to claim 5, characterized in that The specific implementation function of the deep canonical correlation analysis network adopts the following formula: Constraints: Where N is the number of samples, i is the sample number, f(X) is the first modal feature, X is the infrared feature, g(Y) is the second modal feature, Y is the WiFi feature, and U is the CCA direction of all output units of the deep neural network, U = [u1,…u l …,u L ],u l is the CCA direction of the lth output unit of the deep neural network, L is the number of output units in the deep neural network, V is the CCA direction of all output units of the deep neural network, V=[v1,…v m …,v L ],v m is the CCA direction of the mth output unit of the deep neural network, the directions of U and V are different, λ is the trade-off parameter, f(x i ) is the infrared feature x corresponding to the i-th sample i The corresponding first modal feature, p is the first reconstruction nonlinear function, g(y i ) is the WiFi feature y corresponding to the i-th sample i The corresponding second modal feature, q is the second reconstruction nonlinear function; I is the unit matrix, r x and r y is the regularization parameter.

8. The device according to claim 5, characterized in that The specific implementation function of the mapping module is expressed by the following formula: Where d is the mapping matrix, is the parameter vector of the deep canonical correlation analysis network corresponding to the jth modal feature when the correlation is maximum, z jk is the kth sample of the jth mode, n j is the number of samples of the jth mode, D jk is the mapping value of the kth sample of the jth mode.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for human motion recognition combining infrared and WiFi as described in any one of claims 1 to 4 is implemented.

10. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the human motion recognition method combining infrared and WiFi as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Human body action recognition method and device based on multi-modal feature fusion

    CN111898442A

  • Visual language recognition method and device, electronic equipment and storage medium

    CN114581812A