Multi-path gait recognition system and method based on WiFi-video cross-modal fusion
Through the multi-path gait recognition system with cross-modal fusion of WiFi-video, video generation of ideal channel state information combined with WiFi measurement actual information, it solves the problems of large sample demand and many devices in a multi-path environment, and achieves high-precision gait recognition.
Patent Information
- Application Number
- CN202311017066.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-08-14
AI Technical Summary
The existing WiFi-based gait recognition system has the problem of large sample size requirements and special placement of multiple receivers in multi-path indoor environments, and the multi-modal fusion system is sensitive to environmental noise and is difficult to scale to multi-path recognition.
A multi-path gait recognition system based on WiFi-video cross-modal fusion is adopted to generate ideal channel state information through the video feature extraction unit, and combined with the WiFi information extraction unit to measure the actual channel state information, and use neural network matching to achieve gait recognition to reduce device and sample requirements.
It realizes accurate identification of wireless gait in multi-path indoor environment, reduces equipment deployment costs and sample requirements, and improves identification accuracy.
Smart Images

Figure CN117115907B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless sensing technology, and in particular to a multi-path gait recognition system and method based on WiFi-video cross-modal fusion. Background Art
[0002] Currently, gait recognition systems can be implemented using a variety of sensors, with WiFi-based systems becoming a research hotspot. These systems offer advantages such as requiring no custom equipment, being insensitive to ambient brightness, and elegantly protecting user privacy. Existing research on WiFi-based gait recognition systems can be categorized into three types based on their application scenarios: single-path, multi-path, and multimodal fusion.
[0003] Existing single-path based systems extract reliable gait features from CSI (Channel State Information). Due to the diversity of gait features on different paths, these systems are generally applicable to single-path scenarios. For example, WiWho, disclosed in reference [2], is the first WiFi-based gait recognition system used to identify a person from a small group of people in a device-free manner. WifiU, disclosed in reference [1], uses principal component analysis (PCA) and spectrum enhancement technology to generate CSI spectrograms and estimates the speed of the torso and legs respectively for user identification. In WiFi-ID, disclosed in reference [3], continuous wavelet transform (CWT) and related feature selection algorithms are used to extract gait features in the time domain and frequency domain, and sparse approximation-based classification (SAC) is selected as the classifier. WiAU [4] uses residual network (ResNet) technology with two loss functions to verify legitimate users and identify illegal users. These gait recognition systems require less than 60 samples per person. For a walking path length of 5-10m, the sampling time is less than ten minutes, which is acceptable in practical use. The typical path used for recognition is parallel or perpendicular to the line of sight (LoS) between the WiFi transmitter (TX) and receiver (RX), thus requiring only a pair of transceivers. However, these systems require the subject to walk along a specific straight path. This requirement hinders the practical use of these systems in general indoor environments with multiple walking paths.
[0004] In view of the shortcomings of existing single-path based systems, in order to support general indoor scenes, existing multi-path based systems are able to identify people walking on multiple paths. At present, some solutions have been proposed to solve this problem. For example, the AGait method disclosed in reference [5] uses an attention-based recurrent neural network (RNN) encoder-decoder framework to achieve eight-direction gait recognition. However, this method requires a large number of gait samples, that is, each person must walk in the monitoring area for up to 60 minutes (an average of 966 samples per person) to train the model. In addition, it also uses two separate, specially placed RXs to monitor the torso and legs respectively. WiDIGR disclosed in reference [6] realizes direction-independent gait recognition by placing a TX and two RXs to form orthogonal Fresnel zones. The subject walks in the area from the same starting point but in different directions, and the signals of the two RXs are combined into a unified spectrum map about the walking direction. Afterwards, features are extracted using manual and Gabor filters for classification. For each walking direction, WiDIGR requires a similar number of samples as the single-path based system. Wi-PIGR, disclosed in reference [7], is a spectrum graph integration technology based on WiDIGR, using convolutional neural networks (CNN) and long short-term memory (LSTM) for feature extraction to achieve higher recognition accuracy. However, this solution requires more samples, four to five times more than a single-path-based system.
[0005] In summary, current multi-path based solutions currently have at least two problems: the first is the large number of samples, and the second is the dedicated placement of multiple RXs.
[0006] Existing systems based on multimodal fusion,
[0007] WiFi-based gait recognition has the inherent limitation of being sensitive to environmental noise. By combining other methods, WiFi-based gait recognition can achieve better results.
[0008] XModal-ID[8] proposed a cross-modal approach for WiFi-based human recognition, where a given video clip contains walking candidates. XModal-ID generates an ideal CSI through simulation and then matches it with the collected CSI to achieve subject recognition. However, XModal-ID cannot be applied to multi-path gait recognition because it only considers paths perpendicular to the WiFi LoS, cannot be extended to other paths due to the path-sensitive nature of WiFi, and does not consider the impact of the environment on the extracted gait features.
[0009] Reference [1] W. Wang, A. X. Liu, and M. Shahzad, “Gait recognition using wifi signals,” in Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing, 2016, pp. 363–373.
[0010] Reference [2] Y. Zeng, P. H. Pathak, and P. Mohapatra, “Wiwho: Wifi-based person identification in smart spaces,” in 2016 15th ACM / IEEE International Conference on Information Processing in Sensor Networks (IPSN). IEEE, 2016, pp. 1–12.
[0011] Reference [3] J. Zhang, B. Wei, W. Hu, and S. S. Kanhere, “Wifi-id: Human identification using wifi signal,” in 2016 International Conference on Distributed Computing in Sensor Systems (DCOSS). IEEE, 2016, pp. 75–82.
[0012] Reference [4] C. Lin, J. Hu, Y. Sun, F. Ma, L. Wang, and G. Wu, “Wiau: An accurate device-free authentication system with resnet,” in 2018 15th Annual IEEE International Conference on Sensing, Communication, and Networking (SECON). IEEE, 2018, pp. 1–9.
[0013] Reference [5] Y. Xu, W. Yang, M. Chen, S. Chen, and L. Huang, “Attention-based gait recognition and walking direction estimation in wi-fi networks,” IEEE Transactions on Mobile Computing, vol. 21, no. 2, pp. 465–479, 2022.
[0014] Reference [6] L. Zhang, C. Wang, M. Ma, and D. Zhang, “Widigr: Direction independent gait recognition system using commercial wi-fi devices,” IEEE Internet of Things Journal, vol. 7, no. 2, pp. 1178–1191, 2019.
[0015] Reference [7] L. Zhang, C. Wang, and D. Zhang, “Wi-pigr: Path independent gait recognition with commodity wi-fi,” IEEE Transactions on Mobile Computing, 2021.
[0016] Reference [8] B. Korany, C. R. Karanam, H. Cai, and Y. Mostofi, “Xmodal-id: Using wifi for through-wall person identification from candidate video footage,” in The 25th Annual International Conference on Mobile Computing and Networking, 2019, pp. 1–15.
[0017] References[9]Z.Song,H.Zhou,S.Wang,J.Fan,K.Guo,W.zhou,X.Wang,and
[0018] References
[10] H.Cai, B.Korany, CRKaranam, and Y.Mostofi, “Teaching rf tosense without rf training measurements,” Proceedings of the ACM onInteractive, Mobile, Wearable and Ubiquitous Technologies, vol.4, no.4, pp.1–22, 2020.
[0019] References
[11] K. Qian, C. Wu, Y. Zhang, G. Zhang, Z. Yang, and Y. Liu, “Widar2.0: Passive human tracking with a single wi-fi link,” in Proceedings of the16thAnnual International Conference on Mobile Systems, Applications, and Services, 2018, pp.350–361.
[0020] In view of this, the present invention is proposed. Summary of the Invention
[0021] The purpose of the present invention is to provide a multi-path gait recognition system and method based on WiFi-video cross-modal fusion, which can combine WiFi and video to achieve accurate wireless multi-path gait recognition, thereby solving the above-mentioned technical problems existing in the prior art.
[0022] The purpose of the present invention is achieved through the following technical solutions:
[0023] A multi-path gait recognition system based on WiFi-video cross-modal fusion, comprising:
[0024] Video-based gait feature extraction unit, gait feature storage and matching unit, WiFi-based information extraction unit and gait recognition unit; wherein,
[0025] The video-based gait feature extraction unit is communicatively connected to the gait feature storage and matching unit, and is capable of generating path-related ideal channel state information from an acquired video containing at least one walking subject, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit;
[0026] The gait feature storage and matching unit is communicatively connected to the gait recognition unit, and can match the path-related ideal gait features input by the video-based gait feature extraction unit to obtain path-related ideal channel state information gait features, and output the matched path-related ideal channel state information gait features to the gait recognition unit;
[0027] The WiFi-based information extraction unit is communicatively connected to the gait recognition unit and can measure actual channel state information from an ambient WiFi signal containing at least one walking subject acquired in the monitoring area, extract actual channel state information gait features from the actual channel state information, and estimate a walking path, and output the obtained actual channel state information gait features and walking path to the gait recognition unit respectively;
[0028] The gait recognition unit can identify the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
[0029] A multi-path gait recognition method based on WiFi-video cross-modal fusion, using the multi-path gait recognition system based on WiFi-video cross-modal fusion described in the present invention, includes the following steps:
[0030] generating path-related ideal channel state information from an acquired video containing at least one walking subject by a video-based gait feature extraction unit of the system, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit;
[0031] Matching the path-related ideal gait features input by the video-based gait feature extraction unit with the gait feature storage and matching unit to obtain path-related ideal channel state information gait features, and outputting the matched path-related ideal channel state information gait features to the gait recognition unit;
[0032] The WiFi-based information extraction unit of the system measures actual channel state information from ambient WiFi signals containing at least one walking subject acquired in the monitoring area, extracts actual channel state information gait features and estimates a walking path from the actual channel state information, and outputs the obtained actual channel state information gait features and walking path to the gait recognition unit respectively;
[0033] The gait recognition unit identifies the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
[0034] Compared with the prior art, the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the present invention has the following beneficial effects:
[0035] A video-based gait feature extraction unit simulates Wi-Fi signals based on a video of a walking subject, thereby deriving ideal channel state information for the walking subject walking on different paths. A gait feature storage and matching unit then derives gait features of the ideal channel state information associated with the path matched by the ideal channel state information. A Wi-Fi-based information extraction unit then collects and measures actual channel state information of actual Wi-Fi signals within the monitoring area to estimate the walking subject's path. The gait recognition unit then matches this information with the ideal channel state information to achieve multi-path walking subject recognition. This combination of video and Wi-Fi information allows the generated ideal CSI to be applied to different paths, enabling accurate wireless gait recognition with less equipment and fewer samples. The system and method of the present invention enable walking path recognition based on a pair of Wi-Fi transceivers, resulting in a low-cost system deployment, i.e., eliminating the need to deploy multiple specially placed Wi-Fi receivers. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 A schematic diagram of the structure of a multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0038] Figure 2 Schematic diagram of an application scenario of a multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0039] Figure 3 Schematic diagram of walking trajectories extracted based on WiFi for a multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0040] Figure 4 This paper evaluates the accuracy of the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the embodiment of the present invention.
[0041] Figure 5 Comparison of the number of recommended training set samples for the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the embodiment of the present invention.
[0042] Figure 6 This is the effect of feature alignment of the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the embodiment of the present invention.
[0043] Figure 7 The accuracy differences of different paths in the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the embodiment of the present invention are shown.
[0044] Figure 8 This is the walking example fusion effect of the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0045] Figure 9 This paper investigates the impact of the number of training samples on the path recognition and gait recognition performance of the multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0046] Figure 10 A schematic diagram of a path recognition scenario of a multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention.
[0047] Figure 11 A schematic diagram of the path recognition accuracy of a multi-path gait recognition system based on WiFi-video cross-modal fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the specific content of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments, and do not constitute a limitation of the present invention. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0049] First, the following terms may be used in this article:
[0050] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y includes both “X” or “Y” and “X and Y”.
[0051] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.
[0052] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.
[0053] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," and "fixed" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this document based on specific circumstances.
[0054] The terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings and are only for the convenience and simplification of description, and do not explicitly or implicitly indicate that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be understood as a limitation to this document.
[0055] The multi-path gait recognition system based on WiFi-video cross-modal fusion provided by the present invention is described in detail below. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professionals in this field. If specific conditions are not specified in the embodiments of the present invention, they are carried out in accordance with conventional conditions in the field or conditions recommended by the manufacturer. The reagents or instruments used in the embodiments of the present invention, for which the manufacturer is not specified, are all conventional products that can be purchased commercially.
[0056] like Figure 1 As shown, an embodiment of the present invention provides a multi-path gait recognition system based on WiFi-video cross-modal fusion, including:
[0057] Video-based gait feature extraction unit, gait feature storage and matching unit, WiFi-based information extraction unit and gait recognition unit; wherein,
[0058] The video-based gait feature extraction unit is communicatively connected to the gait feature storage and matching unit, and is capable of generating path-related ideal channel state information from an acquired video containing at least one walking subject, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit;
[0059] The gait feature storage and matching unit is communicatively connected to the gait recognition unit, and can match the path-related ideal gait features input by the video-based gait feature extraction unit to obtain path-related ideal channel state information gait features, and output the matched path-related ideal channel state information gait features to the gait recognition unit;
[0060] The WiFi-based information extraction unit is communicatively connected to the gait recognition unit and can measure actual channel state information from an ambient WiFi signal containing at least one walking subject acquired in the monitoring area, extract actual channel state information gait features from the actual channel state information, and estimate a walking path, and output the obtained actual channel state information gait features and walking path to the gait recognition unit respectively;
[0061] The gait recognition unit can identify the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
[0062] Preferably, in the above system, the video-based gait feature extraction unit includes:
[0063] Pedestrian detection module, 3D Mesh generation module, channel state information simulation module, spectrum graph generation module and original feature extraction module; Among them,
[0064] The pedestrian detection module is in communication with the 3D Mesh generation module and can detect the boundaries of pedestrians in a video frame containing at least one walking subject video using a Mask-RCNN algorithm and capture a corresponding image of the video frame;
[0065] The 3D mesh generation module is communicatively connected to the channel state information simulation module and can generate a 3D mesh of the human body surface from the corresponding images of the video frames captured by the pedestrian detection module using a human body mesh recovery algorithm, and use the 3D meshes generated from the corresponding images of all captured video frames to form a 3D motion model that describes the body surface movement of the subject when walking; specifically, the 3D meshes generated from the corresponding images of all captured video frames are used to form a sequence, and the 3D motion model that describes the body surface movement of the subject when walking is used as the sequence;
[0066] The channel state information simulation module is in communication with the spectrum diagram generation module and can obtain the visible point of the 3D motion model relative to the WiFi device from the 3D motion model generated by the 3D mesh generation module, and simulate the propagation of the WiFi signal based on the visible point to generate ideal channel state information;
[0067] The spectrum graph generation module is in communication with the original feature extraction module and can convert the ideal channel state information generated by the channel state information simulation module into ideal spectrum graphs of different paths using short-time Fourier transform;
[0068] The original feature extraction module can obtain the minimum gait cycle, minimum stride, minimum average energy, minimum trunk frequency point, minimum leg speed and minimum trunk speed, maximum gait cycle, maximum stride, maximum average energy, maximum trunk frequency point, maximum leg speed and maximum trunk speed, average gait cycle, average stride, average average energy, average trunk frequency point, average leg speed and average trunk speed and FD features from the ideal spectrum of different paths obtained by the spectrum generation module and form a 21-dimensional feature vector, combine the corresponding subject, path information and feature vector, and store them in the feature map as the ideal gait feature.
[0069] Preferably, in the above system, the channel state information simulation module simulates and generates ideal channel state information by simulating the propagation of WiFi signals according to the viewpoint in the following manner, including:
[0070] Extracting points that reflect electromagnetic signals to the receiver from the 3D grid using a hidden point removal algorithm;
[0071] The method in XModal-ID is used to generate simulated channel state information of simulated WiFi signals for multiple paths by simulating WiFi devices placed at different locations.
[0072] Preferably, in the above system, the WiFi-based information extraction unit includes:
[0073] Channel state information acquisition module, actual channel state information gait feature extraction module and walking path estimation module; wherein,
[0074] The channel state information acquisition module is communicatively connected to the actual channel state information gait feature extraction module and the walking path estimation module, respectively, and can measure the actual channel state information from the acquired ambient WiFi signal containing at least one walking subject;
[0075] The actual channel state information gait feature extraction module can sequentially perform data preprocessing, Butterworth bandpass filtering, and principal component analysis on the actual channel state information measured by the channel state information acquisition module to obtain an enhanced spectrum, and extract the actual channel state information gait feature from the enhanced spectrum;
[0076] The walking path estimation module can obtain a preliminary walking trajectory from the actual channel state information measured by the channel state information acquisition module using the Widar2.0 algorithm, extract a trajectory trend from the preliminary walking trajectory, and estimate the walking path by performing path classification on the extracted trajectory trend using the nearest node algorithm.
[0077] Preferably, in the above system, the gait recognition unit includes:
[0078] At least one feature alignment network, at least one similarity evaluation network and a recognition result output module; wherein,
[0079] Each of the feature alignment networks is communicatively connected to an output terminal of the gait feature storage and matching unit, and is capable of learning an offset of a real gait feature affected by environmental noise, and fusing the offset affected by environmental noise with an ideal channel state information gait feature output by the gait feature storage and matching unit to implement feature alignment to obtain a semi-ideal channel state information gait feature;
[0080] Each of the similarity evaluation networks is communicatively connected to the feature alignment network and the WiFi-based information extraction unit, respectively, and is capable of determining the similarity between the semi-ideal channel state information gait features output by the feature alignment network and the actual channel state information gait features output by the WiFi-based information extraction unit;
[0081] The recognition result output module is in communication with the similarity evaluation network and can obtain a gait recognition result after sorting the similarities of all candidates output by the similarity evaluation network.
[0082] Preferably, in the above system, each of the feature alignment networks adopts a single hidden layer feedback neural network with 21 units, which takes the ideal channel state information gait features of each subject (i.e., the pedestrian in the video) as input and the subject's corresponding actual channel state information gait features as supervision data, and can output the feature vector difference between the ideal channel state information gait features and the actual channel state information gait features;
[0083] Each of the similarity evaluation networks adopts a single hidden layer feedback neural network with 30 units, which takes the actual channel state information gait features and the semi-ideal channel state information gait features output by the feature alignment network as input, and can output the similarity between the actual channel state information gait features and the semi-ideal channel state information gait features.
[0084] Preferably, in the above system, each of the feature alignment networks corresponds to a judgment path;
[0085] Each similarity evaluation network corresponds to a judgment path.
[0086] An embodiment of the present invention further provides a multi-path gait recognition method based on WiFi-video cross-modal fusion, which uses the multi-path gait recognition system based on WiFi-video cross-modal fusion, and includes the following steps:
[0087] generating path-related ideal channel state information from an acquired video containing at least one walking subject by a video-based gait feature extraction unit of the system, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit;
[0088] Matching the path-related ideal gait features input by the video-based gait feature extraction unit with the gait feature storage and matching unit to obtain path-related ideal channel state information gait features, and outputting the matched path-related ideal channel state information gait features to the gait recognition unit;
[0089] The WiFi-based information extraction unit of the system measures actual channel state information from ambient WiFi signals containing at least one walking subject acquired in the monitoring area, extracts actual channel state information gait features and estimates a walking path from the actual channel state information, and outputs the obtained actual channel state information gait features and walking path to the gait recognition unit respectively;
[0090] The gait recognition unit identifies the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
[0091] In summary, the system and method of the embodiments of the present invention combine video with WiFi information to simulate WiFi signals based on the video of the pedestrian, thereby deriving ideal channel state information for the pedestrian walking on different paths. This information is then matched with the ideal channel state information gait features associated with the paths corresponding to the ideal channel state information. The walking path of the pedestrian is estimated by collecting and measuring the actual channel state information of WiFi signals in the monitoring area. Multi-path pedestrian identification is achieved by matching the gait recognition unit with the ideal channel state information. This combination of video and WiFi information allows the generated ideal CSI to be applied to different paths, enabling accurate wireless gait identification with less equipment and fewer samples.
[0092] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the multi-path gait recognition system and method based on cross-modal fusion of WiFi and video provided by the embodiments of the present invention are described in detail below with reference to specific embodiments.
[0093] Example 1
[0094] like Figure 1As shown, an embodiment of the present invention provides a multi-path gait recognition system based on cross-modal fusion of WiFi and video. The system includes the following units: a video-based gait feature extraction unit that can generate path-related ideal CSI from the video, extract gait features, and store them in a gait feature storage and matching unit for matching; a WiFi-based information extraction unit that estimates the walking path and extracts gait features from the measured actual CSI; and a gait recognition unit that can match gait features from different modalities based on a neural network method to identify the walking subject.
[0095] The following describes in detail each component of the multi-path gait recognition system of this embodiment.
[0096] (1) Video-based gait feature extraction unit, including the following modules:
[0097] 11) Pedestrian detection module:
[0098] For each detection process, a short video of a walking subject (generally referring to a pedestrian) walking is required. The video should be in a side view to obtain overall body movement information. Please note that it is not necessary to record the video in the application scenario. Due to the complex recording environment, the effect of directly using the video frame in the subsequent steps is poor. Therefore, the present invention uses the Mask-RCNN algorithm to detect the boundaries of pedestrians in the video frame and capture the corresponding picture of the video frame. In this way, the pedestrian detection module can intercept a series of video frame pictures containing pedestrians.
[0099] 12) 3D Mesh Generation Module:
[0100] For each video frame obtained by the pedestrian detection module, the 3D Mesh Generation module generates a 3D mesh of the human surface using the Human Mesh Recovery (HMR) algorithm, which has the advantage of inferring 3D pose and shape parameters directly from image pixels. The sequence of meshes generated across all frames forms a 3D motion model that roughly describes the movement of the subject's body surface as they walk.
[0101] 3) Channel State Information (CSI) simulation module
[0102] After obtaining the viewpoints of the 3D mesh through the 3D Mesh Generation Module, the Channel State Information (CSI) Simulation Module simulates the propagation of WiFi signals to generate ideal CSI. This process is completed by obtaining visible points and simulating WiFi signals.
[0103] First, since the electromagnetic signal from the WiFi device will be reflected back from the surface of the object, only some points in the 3D grid can reflect the electromagnetic signal to the RX. The present invention uses the Hidden Point Removal (HPR) algorithm to extract these points, which are represented as visible points. In addition, to better match the temporal resolution of WiFi sensing, the video-based point cloud is upsampled from 60Hz to 200Hz, as can be seen in reference
[10] . Note that oversampling will produce multiple continuous identical signals during the signal simulation process, which is not conducive to subsequent spectrum generation and feature extraction. Therefore, the present invention does not upsample to 1kHz as the WiFi sensing rate.
[0104] Then, the method in the XModal-ID scheme disclosed in reference [8] is used to generate simulated CSI. The difference is that the present invention generates signals for multiple paths by simulating WiFi devices placed in different locations. Figure 1 As shown, it simulates Figure 2 The positions of TX and RX corresponding to the three judgment paths.
[0105] 13) Spectrum graph generation module
[0106] Since the velocity information of different parts of the human body is mixed in the CSI, time-frequency analysis techniques are essential to find some periodicity in different frequency components. Similar to the WifiU scheme disclosed in reference [1], the simulated CSI is converted into a spectrogram using short-time Fourier transform (STFT).
[0107] For the 200 Hz sampling rate of the analog CSI, the present invention sets the FFT size to 80 samples and the sliding window step size to 1 sample. Therefore, the resulting frequency resolution is 2.50 Hz and the time resolution is 5 ms, which is suitable for tracking human walking.
[0108] In addition, due to the hidden point algorithm mentioned above, the number of points calculated at adjacent moments is different, which means that the signal strength of the simulated CSI varies greatly at different moments. Therefore, the present invention eliminates this difference by applying column normalization to the spectrum graph.
[0109] Since the frequencies of high energy under different paths vary greatly, instead of selecting a single classifier for all judgment paths, different classifiers are trained for different paths. In this way, their performance can be guaranteed with a limited number of measured CSI samples.
[0110] 14) Feature extraction module
[0111] After obtaining the ideal spectrograms of different paths through the spectrogram generation module, the following gait features can be extracted.
[0112] (141) Frequency distribution (FD): The frequency distribution is obtained by accumulating the frequency over time and represents the distribution of different frequency components in the signal. In other words, FD roughly describes the speed ratio of different body parts when a person walks [1]. Since the frequency generated by human motion is mainly concentrated in the range of 10-70 Hz, this paper divides the FD from 1-80 Hz into 8 parts and takes the average value to obtain the FD feature.
[0113] (142) Energy distribution (ED): Energy distribution is obtained by frequency accumulation, which includes the sum of the energy of the human body movement cycle and body parts.
[0114] (43) Speed curve: The speed curves of the trunk, legs, and trunk outline can be obtained by using the method in the WifiU solution disclosed in reference [1]. Please note that due to the well-known multipath effect of wireless signals, the calculated speed curves will be different on different walking paths.
[0115] (144) Gait cycle: For the energy distribution obtained, the present invention uses the autocorrelation-based method proposed in the WifiU scheme disclosed in reference [1] to obtain the gait cycle. This function carries information about the gait cycle and the walking cycle.
[0116] (145) Stride length: Stride length is obtained by multiplying the average value of gait period and trunk velocity.
[0117] (146) Average energy: The average value in the energy distribution contains the sum of the speed information of the whole body when a person walks, which is affected by body shape, walking speed, etc.
[0118] (147) Dominant frequency point: The point that accounts for 90% of the calculated frequency distribution represents the dominant frequency of the entire walking process.
[0119] In summary, a 21-dimensional feature vector was obtained, including the gait period, stride length, average energy, trunk frequency points, minimum, maximum, average, and variance values of leg and trunk speeds, as well as FD features. The corresponding walking subject and path information were then combined with the feature vector and stored in a gait feature database for use in phase similarity measurement.
[0120] (2) WIFI-based information extraction unit, including the following modules:
[0121] (21) The channel state information acquisition module can measure the actual channel state information from the acquired ambient WiFi signal containing at least one walking subject; the system collects the measured CSI on the RX side. For a pair of transmit and receive antennas, CSI values are obtained from the 30 OFDM subcarriers used by 802.11n. Therefore, in a scenario using 1 TX antenna and 3 RX antennas, 1×3×30=90 CSI values will be obtained for each received 802.11n frame. In the recognition stage, for the continuously measured CSI, the present invention uses a dynamic threshold algorithm to detect the start of walking, and then intercepts the CSI lasting 3 seconds, which represents an instance of stable walking, and then further estimates the walking path and uses the following two parts to extract gait features.
[0122] Path determination
[0123] The path determination part estimates the accurate walking path of the walking subject based on the measured actual CSI.
[0124] Basic idea: There are some WiFi-based human tracking methods (such as the Widar2.0 algorithm disclosed in reference
[11] ), but their accuracy is not perfect due to the inherent noise of WiFi. Therefore, the basic idea of path determination in this paper includes two parts, namely, walking trajectory estimation based on existing methods to obtain an inaccurate trajectory, namely the initial walking trajectory, and then path recognition to determine whether the trajectory belongs to one of the judgment paths.
[0125] (22) Walking trajectory estimation module: The rough walking trajectory is obtained as the initial walking trajectory using the method of the Widar2.0 scheme because it has the advantages of good performance (i.e., an average accuracy of 0.75m within a 6×5m tracking area) and easy deployment (i.e., only requires a pair of transceivers).
[0126] The Widar2.0 algorithm extracts channel parameters such as angle of arrival (AoA), time of flight (ToF), and Doppler shift frequency (DFS) from multiple subcarrier signals in the measured actual CSI. It then uses ToF and DFS to estimate the distance to the WiFi device and finally combines AoA to derive the position of the mobile walking subject.
[0127] The CSI sampling rate is 1kHz. In the Wider2.0 method, the window size of the channel parameter estimation is 100 samples. Therefore, the present invention can obtain an inaccurate walking trajectory of the walking subject, which is a position sequence with a time granularity of 0.1s.
[0128] Figure 3 An example of estimated path trajectories is drawn when a walking subject follows Figure 2When walking along the three judgment paths in [1]. By comparing the estimated trajectories with the ground truth, we observe two things. On the one hand, due to the inherent noise in CSI, the estimated trajectories deviate from the ground truth and cannot be directly used for path determination. On the other hand, due to the diversity of the selected judgment paths, there is still a chance for accurate path determination.
[0129] (221) Path Recognition Currently, inaccurate walking trajectories are obtained, and then accurate path recognition is achieved by extracting trajectory trends and using a classifier.
[0130] (222) Trajectory trend: The following data are selected to express the trajectory trend. First, linear parameters are obtained by linearly fitting the trajectory. Then, the trajectory is divided into three segments, and the center points of these segments are taken. Finally, the sum of all points in the middle segment represents stable walking. For the selected data, they are reshaped into a one-dimensional vector. This trend expression method is selected to achieve a good compromise between computational complexity and recognition accuracy.
[0131] (223) Path classification: The KNN algorithm was selected for path classification because it is simple and efficient. After automatic hyperparameter optimization of KNN by MATLAB, Manhattan distance was used to measure the distance between different trends, and the number of neighbors was set to 10. According to the path labels of adjacent trends, the probability of each path can be obtained. Experiments show that the accuracy of path recognition is proportional to the diversity between the judgment paths. When the distance between two paths exceeds 2m, the accuracy will exceed 90%, while when the distance is less than 1.5m, the accuracy is less than 80%. Therefore, multiple judgment paths with a center point distance of 2m were selected to balance the versatility and WiFi tracking ability. Tested in the corresponding dataset, the path recognition accuracy was as high as 86.9%.
[0132] (23) Actual channel state information gait feature extraction module
[0133] Since the influence of environmental noise cannot be ignored, in order to reduce the influence of noise and obtain highly reliable gait features from the measured CSI, the actual channel state information gait feature extraction module of the present invention reduces the influence of noise through the following process.
[0134] (231) Data preprocessing: The noise in the raw CSI typically contains interference from nearby devices, as well as noise generated by TX / RX internal state transitions (such as transmit power adaptation, internal CSI reference level changes, and imperfect clock synchronization), see references [1],
[22] . Due to carrier frequency offset (CFO), the phase noise in the CSI is too large, and only the amplitude is used to identify human gait. Bandpass filters and principal component analysis (PCA) are selected to denoise and compress multiple OFDM subcarriers while preserving information about human motion.
[0135] (232) Butterworth Bandpass Filter: The measured CSI may contain noise in various frequency bands, such as high-frequency noise caused by internal high-level pulses and low-frequency noise caused by nearby electronic devices. The frequencies generated by human motion are mainly concentrated in the range of 10-70 Hz, corresponding to a typical indoor walking speed of 0.3 to 2 meters per minute. Therefore, the present invention uses a Butterworth bandpass filter with a cutoff frequency of 5-80 Hz to eliminate noise in irrelevant frequency bands.
[0136] (233) PCA selection: All OFDM subcarriers have strong linear correlation, indicating redundancy in CSI. Therefore, we use PCA to automatically discover the correlation between CSI subcarriers and reorganize them into components that primarily represent human activity. To further remove noise components, we use the PCA selection algorithm proposed by IMFi in reference [9] to calculate the low-frequency energy ratio and then select the better components by setting a threshold.
[0137] (232) Spectrogram Generation: Similar to Sec. IV-D, the present invention converts the preprocessed CSI into a spectrogram using STFT. The sampling rate for the WiFi device is 1 kHz. Then, to maintain a temporal resolution similar to that of the ideal spectrogram, the present invention sets the FFT and sliding window step sizes to 256 samples and 8 samples, respectively. Thus, it achieves an appropriate frequency resolution of 3.9 Hz and a temporal resolution of 8 ms.
[0138] (233) Spectrogram enhancement: Although noise is filtered in the preprocessing, background noise with the same frequency as human activity is still not processed. Therefore, the present invention uses spectrogram enhancement technology to further remove noise. The processing method of WifiU in reference [1] can be adopted, using the frequency domain denoising method and a two-dimensional Gaussian low-pass filter on the spectrogram.
[0139] (233) Feature extraction: To compare the ideal and measured CSI, the same feature extraction method as in Sec. IV-E is used.
[0140] (3) Gait recognition unit, including the following modules:
[0141] The gait features of the ideal and measured actual CSI have been obtained through the processing of the above units. However, classification based on direct comparison is still inefficient. The reason is that the ideal CSI is a simulated approximation that does not take into account environmental influences, such as the long-term impact of room layout and the short-term impact of changes in channel conditions. Therefore, the present invention adopts a neural network-based method to solve environmental problems from two aspects. First, feature alignment is performed to obtain "semi-ideal" CSI gait features that include long-term environmental effects. Then, the similarity measurement between the NN-based semi-ideal features and the measured features is used to deal with short-term effects.
[0142] (31) Feature alignment network, which can achieve feature alignment related to the environment
[0143] For a given path, due to the influence of the long-term specific room environment, the measured features of different subjects will share the same offsets affected by environmental noise. Therefore, feature alignment is performed by learning these common offsets and fusing them with the ideal features to obtain semi-ideal CSI gait features.
[0144] For each path, the present invention uses a separate neural network for accent retrieval. The network architecture is a single-hidden-layer feedback neural network with 21 units. During the training phase, the network is fed with each subject's ideal characteristics as input, along with their corresponding measured characteristics as supervision data. This way, the network learns the specific accent of a room, specifically the difference between the ideal and measured characteristics. The present invention prefers to use data from multiple subjects to train the network, to avoid being constrained by specific individual characteristics.
[0145] During the inference phase, the ideal CSI gait features of each subject are processed by the network. The output is then the subject-dependent semi-ideal CSI gait features that contain long-term specific room accents.
[0146] (32) Similarity Evaluation Network
[0147] Even if a semi-ideal CSI gait feature is obtained, it is still unacceptable to directly compare it with the measured actual CSI gait feature due to short-term environmental effects. Therefore, the present invention uses another NN to perform similarity evaluation between features.
[0148] The similarity evaluation network is also a single-hidden-layer feedback neural network with 30 units. It takes as input the 21-dimensional difference between the measured and candidate semi-ideal features and outputs the similarity between them. Similar to feature alignment, the similarity evaluation network is trained separately for each judgment path.
[0149] During the training phase, an equal number of positive and negative samples are prepared to form the training set. Positive samples are a pair of half ideal and half measured features belonging to the same subject, with a similarity of 1. In contrast, negative samples contain features from different subjects, and their similarity is 0.
[0150] Once a measured feature vector of a judgment path is obtained in the recognition phase, a trained network is selected for that path and the similarity of each possible candidate to the semi-ideal feature is evaluated. Then, the gait recognition result can be obtained by ranking the similarities of all candidates.
[0151] (33) Recognition result output module, which can output multi-path gait recognition results
[0152] Because feature alignment and similarity assessment networks are trained separately for each path, if the path determination component can accurately identify the path, the corresponding network can be selected to complete the gait recognition task. However, due to inaccurate WiFi positioning results, path determination can only output the probability of each path. Therefore, path probability is treated as a weight for the similarity assessment results. By summing up the weighted similarities of different paths, the overall similarity of gait recognition can be obtained.
[0153] Once a pedestrian is identified as walking along a judgment path, a gait recognition result is obtained. Since the pedestrian naturally walks within the target room, a series of walking instances are obtained, each of which occurs at different times along different judgment paths. The gait recognition results from these different instances are superimposed on the system of the present invention, thereby achieving better recognition performance.
[0154] The system of the embodiment of the present invention has at least the following advantages:
[0155] By combining WiFi and video, multi-path gait recognition is achieved through cross-modal fusion, and the problem of multi-modal feature differences is optimized. The human positioning task based on low WiFi samples is realized. Video is well utilized to assist wireless devices in implementing the recognition task of the multi-path gait perception system, which will greatly improve the usability of the actual system, such as easy deployment and low sampling.
[0156] Example 2
[0157] (1) Experimental setup
[0158] (11) Experimental scenario: Select a 10×8 meter conference room for evaluation. Figure 4 As shown in the figure, three judgment paths were selected to represent the common paths people take in and out of the conference room. These paths are about 5 meters long, corresponding to 6-10 steps of normal human walking.
[0159] (12) Subject Information: A person's gait depends on their height, weight, and age. In the experiment of this embodiment, 18 volunteers were selected as subjects, including 15 males and 3 females, aged from 22 to 27 years old, with heights ranging from 1.60 to 1.85 meters and weights ranging from 50 to 80 kg. In the experiment, each subject was required to walk back and forth on the judgment path 15 times, and the corresponding sampling time for each subject was 4 to 8 minutes.
[0160] (13) Video data acquisition: The camera is set to 1280×720 resolution and 60 frames per second for video recording. The camera records people Figure 4 The subject walks on path 3 in the image and is 7 meters away from path 3. 15 video instances are collected for each subject, for a total of 270 instances.
[0161] (14) WiFi data collection: To collect WiFi data, two industrial computers (Zhanmei HT770) were used as WiFi transceivers, equipped with Intel 5300 network cards and running Ubuntu 12.04LTS. The TX has one antenna and broadcasts data packets into the air. The RX has three antennas, forming a unified linear array. We used the Linux 802.11nCSITool
[24] to collect CSI measurements. The WiFi device was set to monitor mode at 5.18GHz. The packet transmission rate was set to 1kHz. To better obtain human body information, the WiFi TX and RX were placed on a table at a height of 1m and the distance between them was 5m. 90 WiFi instances were collected for each volunteer, for a total of 1620 instances.
[0162] (15) Training set: The experimental group size was set to 3–6 people based on the number of common family members. This group was randomly selected from 18 subjects to avoid the influence of special subjects on the experimental results. Since only WiFi data was used in the recognition phase, 20% of the WiFi data was randomly divided as test data, and the rest was used as training data. We performed this experiment more than 100 times and calculated the average performance.
[0163] (2) Performance comparison
[0164] (21) Accuracy evaluation: The subject recognition accuracy of the system of the present invention is evaluated. When the subject walks naturally in the room, the system of the present invention is expected to achieve better performance by accumulating the results of different walking instances. Therefore, the recognition results are collected after the subject walks through one or three random path instances, and are denoted as WiVi-1 and WiVi-3, respectively. The system of the present invention is compared with the most advanced gait recognition systems with the same device deployment as the experiment of this embodiment (i.e., only one pair of transceivers), including WifiU[1], WiWho[2] and WiFi-ID[3], as well as WiDIGR[6], which is a typical multi-path gait recognition system with two RXs forming orthogonal Fresnel zones. Figure 4 The accuracy achieved under different group sizes is plotted. First, the performance of WiVi-1 to WiVi-3 (WiVi-1 and WiVi-3 are both systems of the present invention) shows that the more walking examples collected, the better the results achieved. Second, WiVi is significantly better than WifiU, WiWho, and WiFi-ID. Even for WiVi-1, the average improvements are 105.6%, 118.4%, and 117.5%, respectively. The reason is that these baseline algorithms are only designed for one specific path, which is parallel or horizontal to the LoS of the WiFi transceiver. Third, compared with WiDIGR, the system of the present invention achieves similar performance with fewer WiFi devices. The gap between WiVi-1 and WiDIGR is less than 3.14% in all cases, and WiVi-3 even achieves an average improvement of 9.1%.
[0165] (22) Sample size comparison: The recommended sample size is compared with other gait recognition systems. In addition to the four systems used for accuracy evaluation, Wi-PIGR[7] and AGait[5] are also examined here. Figure 5 As shown, the system requires 60 samples, which is similar to single-path systems (WifiU, WiWho, and WiFi-ID), but fewer than multi-path systems (WiDIGR and WiPIGR). Compared to AGait, these numbers are reduced by 57.1%, 70.0%, and 93.7%, respectively. Since other multi-path solutions require more WiFi devices than we do, we can conclude that the system achieves multi-path identification capabilities with fewer samples and fewer devices.
[0166] (3) Secondary assessment
[0167] The effects of different parameters of the system of the present invention were further evaluated.
[0168] (31) Distance between judgment paths: As mentioned above, it is necessary to ensure the diversity of the selected judgment paths. Here, the path determination accuracy of different adjacent distances between paths is evaluated. The example layout of the path is as follows Figure 10 The distance or angle between adjacent paths is 0.5 m in the perpendicular and parallel scenarios, and 15° in the scattered scenario. Figure 11 Figure 1 shows the path determination accuracy between path 1 and the other paths. As expected, when the two paths are too close, i.e., when the adjacent distance is ≤ 1.5m or the angle is ≤ 30°, the results are poor. However, when the paths are well separated (i.e., when the adjacent distance is ≥ 2m or the angle is ≥ 45°), the path determination results are above 90%. Note that most of the paths in the parallel scenario are far away from the WiFi transceiver, so the path determination accuracy is poor in this case. Therefore, the experimental results prove our discriminatory path selection principle.
[0169] (32) Validity of feature alignment: For all judgment paths, Figure 6 The top-1 and top-3 recognition accuracies are plotted with and without feature alignment. Even with a group size of 6, our system achieves 75.0% top-1 and 94.8% top-3 accuracy. Furthermore, feature alignment improves top-1 accuracy by an average of 3.1% and a maximum of 4.9%, demonstrating the effectiveness of context-dependent feature alignment.
[0170] (33) Determine the balance between paths: Figure 7 The recognition accuracy of different paths with different group sizes is plotted. The figure shows that the results are well balanced, with the largest gap less than 5.1%.
[0171] (34) The impact of sample size: For a walking instance, Figure 9 The accuracy of path determination and gait recognition is plotted for different numbers of training samples. Both accuracies increase with increasing samples and stabilize when the number is ≥ 60. Therefore, we choose 60 as the recommended sample size for our system. Furthermore, the results for path determination are much more stable, ranging from 82.9% to 86.9%, indicating that our proposed path determination method is less dependent on training data.
[0172] (35) Fusion of multiple walking instances: Figure 8 We have demonstrated the effectiveness of fusion across multiple walking instances. We further evaluated the fusion results for different numbers of instances in detail. As expected, accuracy increases with the number of instances, and this increase stabilizes when the number of instances exceeds three. Therefore, we prefer to use the fusion results for the three most recent walking instances in our system.
[0173] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0174] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.
Claims
1. A multi-path gait recognition system based on WiFi-video cross-modal fusion, characterized by: include: Video-based gait feature extraction unit, gait feature storage and matching unit, WiFi-based information extraction unit and gait recognition unit; wherein, The video-based gait feature extraction unit is communicatively connected to the gait feature storage and matching unit, and is capable of generating path-related ideal channel state information from an acquired video containing at least one walking subject, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit; The gait feature storage and matching unit is communicatively connected to the gait recognition unit, and can match the path-related ideal gait features input by the video-based gait feature extraction unit to obtain path-related ideal channel state information gait features, and output the matched path-related ideal channel state information gait features to the gait recognition unit; The WiFi-based information extraction unit is communicatively connected to the gait recognition unit and can measure actual channel state information from an ambient WiFi signal containing at least one walking subject acquired in the monitoring area, extract actual channel state information gait features from the actual channel state information, and estimate a walking path, and output the obtained actual channel state information gait features and walking path to the gait recognition unit respectively; The gait recognition unit can identify the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
2. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to claim 1 is characterized in that: The video-based gait feature extraction unit includes: Pedestrian detection module, 3D Mesh generation module, channel state information simulation module, spectrum graph generation module and original feature extraction module; Among them, The pedestrian detection module is in communication with the 3D Mesh generation module and can detect the boundaries of pedestrians in a video frame containing at least one walking subject video using a Mask-RCNN algorithm and capture a corresponding image of the video frame; The 3D mesh generation module is in communication with the channel state information simulation module and can generate a 3D mesh of the human body surface from the corresponding images of the video frames captured by the pedestrian detection module using a human body mesh recovery algorithm, and use the 3D meshes generated from the corresponding images of all captured video frames to form a 3D motion model that describes the body surface movement of the subject when walking; The channel state information simulation module is in communication with the spectrum diagram generation module and can obtain the visible point of the 3D motion model relative to the WiFi device from the 3D motion model formed by the 3D mesh generation module, and simulate the propagation of the WiFi signal based on the visible point to simulate and generate ideal channel state information; The spectrum graph generation module is in communication with the original feature extraction module and can convert the ideal channel state information generated by the channel state information simulation module into ideal spectrum graphs of different paths using short-time Fourier transform; The original feature extraction module can obtain the minimum gait cycle, minimum stride, minimum average energy, minimum trunk frequency point, minimum leg speed and minimum trunk speed, maximum gait cycle, maximum stride, maximum average energy, maximum trunk frequency point, maximum leg speed and maximum trunk speed, average gait cycle, average stride, average average energy, average trunk frequency point, average leg speed and average trunk speed and FD features from the ideal spectrum of different paths obtained by the spectrum generation module and form a 21-dimensional feature vector, combine the corresponding subject, path information and feature vector, and store them in the feature map as the ideal gait feature.
3. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to claim 2 is characterized in that: The channel state information simulation module simulates and generates ideal channel state information by simulating the propagation of WiFi signals according to the viewpoint in the following manner, including: Extracting points that reflect electromagnetic signals to the receiver from the 3D grid using a hidden point removal algorithm; The method in XModal-ID is used to generate simulated channel state information of simulated WiFi signals for multiple paths by simulating WiFi devices placed at different locations.
4. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to any one of claims 1 to 3, characterized in that: The WiFi-based information extraction unit includes: Channel state information acquisition module, actual channel state information gait feature extraction module and walking path estimation module; wherein, The channel state information acquisition module is communicatively connected to the actual channel state information gait feature extraction module and the walking path estimation module, respectively, and can measure the actual channel state information from the acquired ambient WiFi signal containing at least one walking subject; The actual channel state information gait feature extraction module can sequentially perform data preprocessing, Butterworth bandpass filtering, and principal component analysis on the actual channel state information measured by the channel state information acquisition module to obtain an enhanced spectrum, and extract the actual channel state information gait feature from the enhanced spectrum; The walking path estimation module can obtain a preliminary walking trajectory from the actual channel state information measured by the channel state information acquisition module using the Widar2.0 algorithm, extract a trajectory trend from the preliminary walking trajectory, and estimate the walking path by performing path classification on the extracted trajectory trend using the nearest node algorithm.
5. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to any one of claims 1 to 3, characterized in that: The gait recognition unit comprises: At least one feature alignment network, at least one similarity evaluation network and a recognition result output module; wherein, Each feature alignment network is communicatively connected to an output terminal of the gait feature storage and matching unit, and is capable of learning an offset of a real gait feature affected by environmental noise, and fusing the offset affected by environmental noise with an ideal channel state information gait feature output by the gait feature storage and matching unit to achieve feature alignment to obtain a semi-ideal channel state information gait feature; Each similarity evaluation network is respectively in communication with the feature alignment network and the WiFi-based information extraction unit, and is capable of determining the similarity between the semi-ideal channel state information gait features output by the feature alignment network and the actual channel state information gait features output by the WiFi-based information extraction unit; The recognition result output module is connected to each similarity evaluation network for communication, and can sort the similarities of all candidates output by each similarity evaluation network to obtain a gait recognition result.
6. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to claim 5 is characterized in that: Each feature alignment network uses a single hidden layer feedback neural network with 21 units. It takes the ideal channel state information gait features of each subject as input and the corresponding actual channel state information gait features of the subject as supervision data, and can output the feature vector difference between the ideal channel state information gait features and the actual channel state information gait features. Each similarity evaluation network adopts a single hidden layer feedback neural network with 30 units, which takes the actual channel state information gait features and the semi-ideal channel state information gait features output by the feature alignment network as input, and can output the similarity between the actual channel state information gait features and the semi-ideal channel state information gait features.
7. The multi-path gait recognition system based on WiFi-video cross-modal fusion according to claim 5 is characterized in that: Each feature alignment network corresponds to a judgment path; Each similarity evaluation network corresponds to a judgment path.
8. A multi-path gait recognition method based on WiFi-video cross-modal fusion, characterized in that: The multi-path gait recognition system based on WiFi-video cross-modal fusion according to any one of claims 1 to 7 comprises the following steps: generating path-related ideal channel state information from an acquired video containing at least one walking subject by a video-based gait feature extraction unit of the system, extracting path-related ideal gait features from the path-related ideal channel state information, and storing the extracted path-related ideal gait features in the gait feature storage and matching unit; Matching the path-related ideal gait features input by the video-based gait feature extraction unit with the gait feature storage and matching unit to obtain path-related ideal channel state information gait features, and outputting the matched path-related ideal channel state information gait features to the gait recognition unit; The WiFi-based information extraction unit of the system measures actual channel state information from ambient WiFi signals containing at least one walking subject acquired in the monitoring area, extracts actual channel state information gait features and estimates a walking path from the actual channel state information, and outputs the obtained actual channel state information gait features and walking path to the gait recognition unit respectively; The gait recognition unit identifies the corresponding walking subject based on a neural network matching method using the path-related ideal channel state information gait features output by the gait feature storage and matching unit, the actual channel state information gait features output by the WiFi-based information extraction unit, and the walking path.
Citation Information
Patent Citations
Skin type finger gesture recognition method based on smart watch
CN110069199A
Multi-mode pedestrian identity recognition method and system based on pedestrian appearance and gait information
CN111860291A