Multimodal gait user recognition method and system based on wifi acoustics
By combining WiFi signals and footstep sound signals collected by a microphone array, a multimodal gait recognition system was constructed, which solved the adaptability and robustness problems of WiFi gait recognition in complex scenarios and achieved high-precision user identification.
Patent Information
- Application Number
- CN202411176712.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-08-26
AI Technical Summary
Existing WiFi gait recognition technologies suffer from poor adaptability in complex scenarios, high deployment difficulty, environmental interference, and poor robustness.
A multimodal gait recognition method based on WiFi acoustics is adopted. Footstep sound signals are collected by a microphone array to calculate the footstep sound AoA sequence. Combined with the PLCR features of WiFi signals, a geometric model is constructed, multimodal fusion is performed, a velocity-angle matrix is constructed, and a lightweight neural network is trained for gait recognition.
It enables non-invasive extraction of user gait features without relying on wearable sensors or cameras in a single-link environment, improving the robustness and accuracy of recognition, reducing the difficulty of device deployment, and solving recognition problems in complex scenarios.
Smart Images

Figure CN119138884B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of user identity recognition under wireless sensing technology, and in particular to a multi-modal gait user recognition method and system based on WiFi acoustics. BACKGROUND
[0002] Wireless sensing technology is widely used in the field of smart Internet of Things. It transmits and receives WiFi, RFID and sound signals through devices with signal transceiving functions such as routers, RFID readers and smart speakers, extracts fine-grained features associated with human behavior through signal processing technology, and realizes the perception and recognition of human behavior or location without the need for users to carry sensor devices.
[0003] Gait recognition technology is a classification technology based on pattern recognition. The main idea is that the potential biological characteristics of a user are closely coupled with the corresponding gait when walking. Based on the physiological structure and movement habits of the human body, the gait of the same user always tends to be consistent, while the gaits of different users are different, which brings the possibility of user recognition.
[0004] With the rapid development of Internet of Things technology, intelligent applications based on wireless sensing technology have been widely deployed, and mature communication protocols and signal processing technologies provide reliable support for these services. WiFi sensing, as one of the core technologies of wireless sensing, is commonly used in human activity recognition (HAR) and positioning and tracking tasks. Its principle is that when WiFi signals encounter obstacles, the signals will undergo changes such as reflection and diffraction. Therefore, the position and posture changes caused by human activities will also affect the signal patterns, such as amplitude, phase, etc. Because different types and frequencies of movements have different effects on signals, by analyzing the characteristics and change patterns of signals, human movements can be effectively recognized and perceived. In recent years, researchers have often used channel state information (CSI) to analyze these changes. For example, the existing patent application document with publication number CN114757237A, "Speed-independent gait recognition method based on WiFi signal", includes: extracting the CSI amplitude of WiFi that changes over time during the walking process of the personnel; preprocessing the CSI amplitude; determining whether a person is walking in the environment and extracting the walking activity segment; converting the walking activity segment into a time-frequency graph of the same size; building a speed-independent gait recognition model based on DANN, which includes a feature extractor, an identity recognizer, and a speed recognizer. The feature extractor is used to extract potential features from the input time-frequency graph, the identity recognizer is used to predict the identity of the measured target using the features extracted by the feature extractor, and the speed recognizer is used to predict the speed of the measured target using the features extracted by the feature extractor; training the speed-independent gait recognition model and outputting the identity of the measured target. And the existing patent application document with publication number CN114783054A, "Gait recognition method based on wireless and video feature fusion", includes: using a video acquisition device to obtain pedestrian gait recognition video data, using a segmentation network to obtain high-quality pedestrian contour graph sequences from the video frames, performing standardization and cropping operations on the contour graphs to process them into a unified format, and then using a time-space-based deep neural network to obtain video features of the pedestrians; using a common commercial wireless signal device, one end sends physical layer channel state information (CSI) data, the other end receives, and performs denoising and normalization preprocessing on the CSI data, and uses a multi-scale convolutional neural network to extract wireless features of the pedestrians from the preprocessed CSI data; finally, the extracted wireless and video features are fused to perform identity prediction. In the foregoing prior art, CSI provides fine-grained physical layer information in signal propagation, such as propagation distance, power attenuation, and scattering, which reflects the combined effects of the environment and human dynamics. When the user moves, the activities and position changes of the limbs reflect their walking habits, such as stride, step frequency, and arm swinging habits, etc.These features provide the potential to identify different gaits, allowing the new generation of smart homes to provide more personalized and differentiated services according to different user identities.
[0005] Existing WiFi-based user identification work usually relies on extracting the short-time Fourier transform (STFT) spectrum from the CSI, and achieving user identification through the spectral features of different users walking. However, using only WiFi features can only obtain the PLCR, but not the real speed of the human body; at the same time, the spectral features are easily affected by many factors such as walking trajectory, direction, etc., which means that the same user walking different trajectories may produce different features, resulting in incorrect identification of these trajectories as different users, which cannot meet the increasingly diverse and complex scene requirements. Subsequent work introduces an additional orthogonal WiFi link to obtain another direction of the spectrum, and maps it from the orthogonal direction to the user's moving direction to solve the problem of trajectory dependence of the features, but the strict requirements for the number of devices and deployment location make it difficult to deploy the system in the home.
[0006] The widespread use of sound devices has also shown the potential of acoustic perception. When a user is walking, they will continuously produce footstep sounds, which can also be used as observation features for user identification. In existing acoustic perception work, a ring array (hexagonal microphone) is usually used to receive sound signals, which has 6 microphones at 6 different positions on the six corners. The typical size of these microphone arrays is usually small to fit common smart speaker products, and the spacing between different microphones and the length of the array are usually centimeters, which can be ignored relative to the distance between the array and the sound source. Therefore, according to the far-field effect, we can regard the sound wave as a parallel wave, which ignores the amplitude difference between the sound signals received by each microphone, and only considers the time delay difference between each received signal. This time delay difference provides the condition for calculating the sound source position. On the other hand, the sound signal can also obtain the corresponding energy spectrum, and through the footstep sound spectrum, features such as step frequency and intensity can be obtained for identity differentiation. However, pure audio is easily affected by other factors such as ground material, shoes worn, or environmental noise, and has poor robustness.
[0007] In view of the difficulty of WiFi gait features to overcome walking trajectory and direction dependence, and the vulnerability of sound features to interference, this method aims to combine the advantages of both, and proposes a multi-modal gait recognition system based on WiFi acoustics, which only uses a single device link to achieve trajectory and direction-independent user gait recognition.
[0008] The prior art document "Research on Passive Indoor Personnel Positioning Technology Based on WiFi" estimates the angle of arrival of the signal, which is easily affected by environmental noise. As the distance between the target and the receiver increases, a small error in the estimation of the angle of arrival of the signal will result in a large loss in the final positioning accuracy, thereby severely reducing the performance of the system. The system designs a double-window AOA estimation algorithm based on time and space dimensions, which effectively guarantees the time complexity of the algorithm while reducing the environmental noise in the AOA estimation. By subtracting a small constant value alpha from the CSI amplitude of all sensors and adding a large constant value beta to the CSI amplitude of the reference sensor, the influence of DFS ambiguity can be eliminated. The MUSIC algorithm is used to obtain the signal arrival angle estimation value. Due to the existence of the Doppler effect, it can be known that when the person to be measured moves in the indoor environment, the propagation path length of the human body reflection signal will change accordingly, which will cause the frequency of the received signal to shift to different degrees compared with the frequency of the signal transmitted by the transmitting end. When the person to be measured moves towards the signal receiving end, the WiFi receiving signal frequency will be higher than the original signal frequency sent by the transmitting end. Correspondingly, when the person to be measured moves away from the signal receiving end, the WiFi receiving signal frequency will be lower than the original signal frequency sent by the transmitting end. The aforementioned comparative document records that the signal arrival angle estimation algorithm in the system does not depend on the Doppler frequency shift estimation value, but the comparative document also mentions that d represents the distance from the target to the receiver, and the change range thereof can be obtained by the Doppler frequency shift estimation value. If the change range is related to the sampling probability and the angle solution, the error of the Doppler frequency shift may be transferred to the angle estimation, which may cause cumulative error.
[0009] In summary, the prior art has the technical problems of poor adaptability to complex scenes, high deployment difficulty, environmental interference, and poor robustness. SUMMARY
[0010] The technical problem to be solved by the present application is how to solve the technical problems of poor adaptability to complex scenes, high deployment difficulty, environmental interference, and poor robustness in the prior art.
[0011] The present application solves the above technical problems by adopting the following technical solutions: a multi-modal gait user recognition method based on WiFi acoustics comprises:
[0012] S1, moving the identified target within the range of the preset indoor device, collecting WiFi signals by using the preset WiFi device to extract CSI features, calculating the reflection path length change rate PLCR according to the CSI features, and constructing a Fresnel zone elliptical model to obtain the PLCR features;
[0013] S2, set and utilize the microphone array to collect the footstep sound of the identified target, calculate the footstep sound AoA sequence according to the signal of the footstep sound, construct a ray model, model according to the preset signal propagation mode and mathematical geometric theory, and calculate the theoretical sample offset value n corresponding to all angles theta sim , match the actual measured sample offset n shift with the theoretical sample offset value n sim , obtain the closest angle value, extract the signal angle of arrival AoA of each footstep sound from the footstep sound signal through sound source positioning, process to obtain a continuous angle sequence distributed along the time domain, and convert the original gait feature into an AoA angle feature;
[0014] S3, confirm the user movement trajectory through multi-modal fusion; simultaneously, construct a geometric model by combining the PLCR feature and the AoA angle feature, so as to obtain the trajectory position sequence of the user movement; construct a speed-angle matrix, and perform a conversion operation on the speed-angle matrix and a frequency-angle matrix according to the relationship between the speed and the frequency, so as to obtain the movement direction angle at each time from the trajectory position sequence; take the movement direction angle, the speed-angle matrix and the frequency-angle matrix as indexes, match the speed in the PLCR speed spectrum, so as to map and convert the PLCR speed spectrum into an independent gait speed spectrum;
[0015] S4, obtain the independent speed spectrum of different identified targets, construct a data-driven user identification model, construct and train a lightweight neural network, obtain a suitable gait recognition model, and perform gait recognition to determine the identity of the identified target.
[0016] The microphone and the WiFi transceiver are placed in the same position in the application, which is used to build a smart speaker or a smart voice assistant prototype, and the WiFi transceiver and the microphone have the functions of transmitting and receiving WiFi signals and collecting footstep sounds, respectively. WiFi and microphone are used to collect WiFi signals caused by user walking and footstep sounds when walking, respectively. A corresponding feature extraction and fusion algorithm is designed to extract human movement speed, angle, gait spectrum and other information. Finally, the gait spectrum is mapped to the acoustic angle to obtain a trajectory direction-independent gait feature, and a more robust user identification is realized.
[0017] The application can extract the gait feature of the user without relying on wearable sensors or camera devices, avoid the cumbersome device deployment requirement, and effectively solve the privacy and lighting condition problems based on camera recognition. The application can realize accurate identification effect under a single link, and improve the feasibility of deploying and applying non-intrusive identity recognition technology in real scenes.
[0018] In a more specific technical solution, S1 includes:
[0019] S11, using the preset CSI extraction tool, extracting the CSI from the WiFi signal using the following logic:
[0020] H(f, t) = H s (f, t) + H d (f, t) = H s (f, t) + A(f, t)e -j2πf(t)
[0021] In the formula, H s (f, t) and H d (f, t) are static signal components and dynamic signal components containing user movement, respectively; S12, the reflection path length at different times is recorded as d(t), and the Doppler frequency shift DFS is determined using the following logic:
[0022]
[0023]
[0024] Where f D (t) is the Doppler frequency shift DFS, λ is the signal wavelength, and τ is the signal flight time;
[0025] S13, using the following logic, determine the relationship between the reflection path length change rate PLCR and the Doppler frequency shift DFS:
[0026]
[0027] S14, using the following logic, integrate the reflection path length change rate PLCR in time to obtain the path length d(t):
[0028]
[0029] In a more specific technical solution, S2, according to the preset geometric relationship, using the following logic, the time difference Δt received by different microphones m i , m j in the microphone array:
[0030]
[0031] Using the following logic, express the sample offset n shift according to the time difference Δt:
[0032] n shift = Δt × fs
[0033] According to the sample offset n shift , using the following logic, the angle θ is obtained:
[0034]
[0035] where L is the distance between two microphones, c is the sound speed, and fs is the acoustic sampling rate. sound
[0036] Specifically, in the theoretical value solving process of the present application, the configuration and size of the microphone array are known, and the sound speed of sound propagation and the microphone signal sampling rate are also known. Therefore, for any sound signal coming from an angle in a two-dimensional plane, we can calculate the distance difference of the signal to the six microphones, i.e. the time difference, by using a geometric formula; and according to the signal sampling rate, we can further obtain the sample offset between the six signals.
[0037] In the actual value solving process of the present application, in the actual footstep sound collection, we can solve the sample point offset between two signal sequences by cross-correlation.
[0038] According to the sample offset n shift , the cross-correlation relationship between the microphone signals is expressed by using the following cross-correlation function logic, and the discrete point sequence C[n shift ] is obtained by processing the microphone cross-correlation relationship:
[0039] C[n shift ] = ∑(m i [n]·m j [n+n shift ]),
[0040] The discrete point sequence C[n shift ] is used to represent the hysteresis of the signal sequence [n+n shift ] relative to the signal sequence [n].
[0041] The peak value of the discrete point sequence C[n shift ] is calculated by using the following logic to find the sample offset value of the signal sequence [n+n shift ] relative to the signal sequence [n]:
[0042]
[0043] Specifically, in the present application, a hexagonal microphone array is used to collect footstep sound signals. Since the positions of each microphone in the hexagonal microphone array are different, the time of the same sound reaching the six microphones is also different. Therefore, the six-channel signals obtained will have a slight offset, i.e. the sample offset in the present application.
[0044] In a more specific technical solution, S2 includes:
[0045] S21, set the value range of the simulation angle θ', and bring the simulation angle into the preset geometric model to obtain the simulation sample offset angle matrix n by using the following logic:sim :
[0046]
[0047] where L is the distance between two microphones, c is the sound speed, and fs is the acoustic sampling rate. sound
[0048] A reference microphone is preset, and time delays Δt of signals received by the rest of the microphones relative to the signal received by the reference microphone are calculated.
[0049] S22, by similarity calculation, obtaining the applicable theoretical sample offset that is most matched with the actual sample offset, and taking the simulation disclosure θ' corresponding to the applicable theoretical sample offset as the angle of arrival of the sound signal:
[0050]
[0051] For the step sound signal generated at each step, the applicable theoretical sample offset is matched The step point quantity alignment angle sequence is obtained:
[0052] θ seq ={θ1,θ2,…,θ k}.
[0053] Specifically, since the traditional algorithm can only obtain an integer n shift (because the position index of the peak value can only take an integer), but in actual situations, the sample offset usually contains a decimal, which will cause errors in the final angle solution. In view of the foregoing situation, the present application adopts a theoretical and actual sample offset matching manner, that is, by modeling through signal propagation mode and mathematical geometric theory, the theoretical sample offset value n sim corresponding to all possible angles θ is calculated, and then the actual measured sample offset n shift is matched with the theoretical value, so as to find the closest angle value, which can more accurately solve the sample offset.
[0054] In a more specific technical solution, in S21, each microphone is taken as a reference microphone respectively, so as to expand the dimension of the simulation sample offset angle matrix n sim , according to the projection theorem, the theoretical sample offset is obtained by using the following logic:
[0055]
[0056] where m i ,m j are different microphones i,j∈{1,2,3,4,5,6}, Δd is the additional flight distance of the signal, is an angle vector, is a distance vector between two microphones.
[0057] The application can realize trajectory-independent gait feature extraction with limited equipment, and improve the user recognition accuracy in actual scenarios. The footstep sound sequence of a user walking can be collected by a microphone array. During walking, the footstep sound is discrete, that is, there is a short silent segment between adjacent two steps. In contrast to the traditional sound source positioning, which usually targets continuous sound sources. The application scenario of the application is a continuous moving sound source, which is a footstep, and is different from the traditional research scenario. Therefore, the moving process is regarded as a combination of multiple static sound sources, and only the sound points are calculated, and the silent part between the steps is abandoned. Therefore, the user's step point is determined by a threshold detection or correlation matching method. In this way, not only can the calculation interference caused by invalid information of the silent segment be avoided, but also the calculation amount can be reduced, and the operation efficiency can be improved.
[0058] In a more specific technical solution, S3 comprises:
[0059] S31, converting the obtained PLCR velocity spectrum to obtain f mat (t) frequency spectrum; since the dimension of f mat is velocity-angle-time, the time step and the corresponding angle can be determined through the obtained β(t), and the corresponding velocity is matched;
[0060] S32, combining the velocity values of each time step together, the entire PLCR velocity spectrum to trajectory-independent and direction-independent independent velocity spectrum is realized.
[0061] In a more specific technical solution, S31 comprises:
[0062] S311, combining the WiFi Fresnel ellipse and the angle sequence ray to solve the user position; wherein the entire perception area is divided into not less than 2 grids, each grid coordinate is marked as (x, y), and the Fresnel ellipse equation is constructed by using the reflection path length d(t):
[0063]
[0064] According to the moving direction angle θ, the angle sequence ray equation is constructed:
[0065] y t = tanθ seq (t)·x t
[0066] S312, convert the Fresnel ellipse equation and the angle sequence ray equation to obtain the error calculation equation:
[0067]
[0068] S313. Substitute the coordinates of each grid into the error calculation equation, perform maximum likelihood estimation, and obtain the user position (x) at each time step. t ,y t ):
[0069]
[0070] S314. In the process of solving gait features, polar coordinates v are used. ρ v θ The reflection path length variation rate (PLCR) represents the rate of change of the reflection path length.
[0071] S315. Obtain and construct the velocity-angle matrix v based on the velocity and angle characteristics. mat According to the velocity-frequency formula, the velocity-angle matrix v mat This is transformed into a frequency matrix, so that the PLCR velocity spectrum can be mapped using the following logic:
[0072]
[0073] The user's location trajectory is obtained through processing, and the displacement direction angle β(t) is calculated based on two adjacent coordinate points.
[0074] Specifically, in the theoretical calculation of this invention, the AOA of [0, 360]° is substituted to calculate the theoretical sample offset corresponding to each angle; the actual sample offset obtained by cross-correlation is matched with the theoretical value, and the AOA angle corresponding to the theoretical value when the two are closest is the signal arrival angle to be sought.
[0075] In a more specific technical solution, S314, v in polar coordinates ρ Represents the speed value, v θ Represents the direction of velocity:
[0076]
[0077] In a more specific technical solution, in S315, δ x and δ y The orientation coefficient is determined by the location of the WiFi device.
[0078]
[0079] In the formula, Indicates the user's location. and These indicate the locations of the transmitter and receiver, respectively.
[0080] In more specific technical solutions, a multimodal gait user recognition system based on WiFi acoustics includes:
[0081] A PLCR feature processing module is configured to make the identified target move within a preset indoor device range, collect WiFi signals by using a preset WiFi device, extract a CSI feature, calculate a reflection path length change rate PLCR according to the CSI feature, construct an elliptical model of a Fresnel zone according to the PLCR, and process the PLCR feature;
[0082] An AoA angle feature processing module is configured to set and collect footstep sounds of the identified target by using a microphone array, calculate a footstep sound AoA sequence according to signals of the footstep sounds, construct a ray model, model according to a preset signal propagation mode and a mathematical geometric theory, and calculate theoretical sample offset values n corresponding to all angles θ sim The actual measured sample offset n shift is matched with the theoretical sample offset value n sim to obtain the closest angle value, the signal arrival angle AoA of each footstep sound is extracted from the footstep sound signals by sound source positioning, a continuous angle sequence distributed along a time domain is processed to convert the original gait feature into an AoA angle feature;
[0083] A multi-modal fusion module is configured to confirm a user movement trajectory by multi-modal fusion, construct a geometric model by combining the PLCR feature and the AoA angle feature, calculate a trajectory position sequence of the user movement, construct a speed-angle matrix, and perform a conversion operation of the speed-angle matrix and a frequency-angle matrix according to a relationship between speed and frequency to calculate a movement direction angle at each time from the trajectory position sequence, and use the movement direction angle, the speed-angle matrix and the frequency-angle matrix as indexes to match a speed in a PLCR speed spectrum to map and convert the PLCR speed spectrum into an independent gait speed spectrum, the multi-modal fusion module is connected with the PLCR feature processing module and the AoA angle feature processing module.
[0084] A gait recognition module is configured to obtain independent speed spectrums of different identified targets, construct a data-driven user recognition model, construct and train a lightweight neural network, obtain a suitable gait recognition model, and perform gait recognition to determine the identity of the identified target, and the gait recognition module is connected with the multi-modal fusion module.
[0085] Compared with the prior art, the present application has the following advantages:
[0086] The microphone and the WiFi transceiver are placed in the same position to build a smart speaker or a smart voice assistant prototype, and the WiFi signal and the microphone sound collecting function are simultaneously provided, the WiFi and the microphone are used to collect the WiFi signal caused by the user walking and the footstep sound when walking, a corresponding feature extraction and fusion algorithm is designed, and information such as human moving speed, angle, gait spectrum and the like is extracted, finally, the gait spectrum is mapped to the acoustic angle to obtain the gait feature irrelevant to the trajectory direction, and higher robustness of user identification is realized.
[0087] The gait feature of the user can be extracted in a non-invasive manner without relying on wearable sensors or camera devices, the cumbersome device deployment requirement is avoided, the privacy and lighting condition problem based on camera recognition is effectively solved, accurate recognition effect can be realized under a single link, and the feasibility of deploying and applying the non-invasive identity recognition technology in a real scene is improved.
[0088] The present application adopts a mode of matching the theoretical and actual sample point offset, that is, the theoretical sample offset value n corresponding to all possible angles theta is calculated through signal propagation mode and mathematical geometry theory sim , and the actual measured sample offset n shift is matched with the theoretical value, so that the closest angle value is found, and the sample point offset solving can be more accurate.
[0089] The present application can realize the trajectory and direction independent gait feature extraction with limited devices, and improve the user identification accuracy in the actual scene. The footstep sound sequence of the user walking can be collected through the microphone array, and the footstep sound is discrete during walking, that is, there is a short silent segment between adjacent two steps. Compared with the traditional sound source positioning which usually aims at continuous sound source. The application scenario of the present application is a continuous moving sound source, that is, footstep, which is different from the traditional research scenario. Therefore, the moving process is regarded as a combination of multiple static sound sources, only the steps with sound are calculated, and the silent part between the steps is abandoned. Therefore, the steps of the user are determined through the threshold detection or correlation matching method. In this way, not only the calculation interference caused by the invalid information of the silent segment can be avoided, but also the calculation amount can be reduced, and the operation efficiency can be improved.
[0090] The present application solves the technical problems of poor adaptability to complex scenes, large deployment difficulty, environmental interference and poor robustness in the prior art. DETAILED DESCRIPTION
[0091] Figure 1 The basic steps of the WiFi acoustic based multi-modal gait user identification method of embodiment 1 of the present application are shown in the figure.
[0092] Figure 2 This is a schematic diagram of data flow processing for the multimodal gait user recognition method based on WiFi acoustics according to Embodiment 1 of the present invention;
[0093] Figure 3 This is a schematic diagram illustrating the specific steps of acoustic signal feature extraction in Embodiment 1 of the present invention;
[0094] Figure 4 This is a schematic diagram of footstep detection by cross-correlation matching in Embodiment 1 of the present invention;
[0095] Figure 5 This is a schematic diagram showing how different trajectories can lead to different state velocity spectra in the conventional method of Embodiment 1 of the present invention;
[0096] Figure 6 This is a schematic diagram illustrating the time delay of the microphone array receiving sound signals in Embodiment 1 of the present invention;
[0097] Figure 7 This is a schematic diagram of the trajectory calculated jointly by Fresnel elliptical clusters and angular rays in Embodiment 1 of the present invention;
[0098] Figure 8 A schematic diagram illustrating the specific steps for deriving the independent velocity spectrum gait features of Embodiment 1 of the present invention;
[0099] Figure 9 This is a diagram of the gait recognition neural network architecture of Embodiment 1 of the present invention;
[0100] Figure 10 This is a schematic diagram of the data stream processing for deriving independent velocity spectrum gait features in Embodiment 1 of the present invention;
[0101] Figure 11 This is a schematic diagram of the user identification confusion matrix in Embodiment 1 of the present invention. Detailed Implementation
[0102] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0103] Example 1
[0104] like Figure 1 and Figure 2 As shown, the multimodal gait user recognition method based on WiFi acoustics provided by this invention includes the following basic steps:
[0105] S1. WiFi signal feature extraction;
[0106] In the embodiment, 1 WiFi transmitter and 1 receiver are deployed to build a single-link WiFi signal collection device for collecting WiFi signals when the human body moves; after collecting the signal data packet, the CSI information can be extracted from the original signal through the CSI tool; when the human body moves, the relative position and distance of the signal source change constantly, resulting in DFS, according to the communication principle, the PLCR can be calculated from the frequency information, which describes the activity when the human body moves, and provides an information basis for subsequent calculation of user gait features.
[0107] In the embodiment, the CSI can be conveniently obtained from the WiFi signal through the CSI Tool toolbox:
[0108] H(f, t) = H s (f, t) + H d (f, t) = H s (f, t) + A(f, t)e -j2πf(t) ,
[0109] Where H s (f, t) and H d (f, t) are static signal components and dynamic signal components containing user movement conditions, respectively. The Doppler frequency offset is generated due to the change of the reflection signal path, we denote the reflection path length at different times as d(t), then the Doppler frequency shift can be denoted as follows:
[0110]
[0111] Where f D (t) is DFS, λ is the signal wavelength, and τ is the signal flight time. And the PLCR is the signal path change rate, that is Therefore, according to the above formula, if we denote the PLCR as r, it can be expressed as Therefore, the relationship between DFS and PLCR is as follows:
[0112]
[0113] Through the above derivation, it can be known that the path length d is the integral of the PLCR over time. In the background art and scheme, it is mentioned that the user's moving track needs to be calculated, and in the WiFi feature, the Fresnel zone needs to be constructed, which is a series of confocal ellipses. According to the Fresnel ellipse and the signal propagation principle, the signal propagation path length needs to be used as the parameter of the ellipse. The path length d(t) can be expressed as follows:
[0114]
[0115] Fresnel zones refer to a plurality of equidistant ellipsoidal regions existing between a signal transmitting end and a receiving end. The signal transmitting end and the receiving end are located at two foci of the ellipses respectively. When the target enters different Fresnel zones, the path lengths of the reflected signal and the direct signal will be different, resulting in changes in the superposition effect of the two at the receiving end. When the object crosses multiple Fresnel zones, the received signal strength will periodically increase and decrease, and this change can depict the user's activity of crossing the Fresnel zone.
[0116] S2, acoustic signal feature extraction;
[0117] In this embodiment, a small microphone array is deployed to simulate a smart voice assistant for collecting footstep sounds when the user walks; through sound source positioning technology, the signal angle of arrival AoA of each footstep sound is extracted from the sound signal, and finally a continuous angle sequence distributed along the time domain is obtained. The angle information can be used to convert the original gait features into independent features irrelevant to the trajectory and direction.
[0118] The principle of acoustic positioning is based on the analysis of sound signal propagation, including signal flight time, angle of arrival, etc. The sound signal emitted from the transmitting end (sound source) will form an acoustic angle of arrival on the antenna array of the receiving end (such as a microphone), and the angle value reflects the azimuth angle of the sound source; based on the speed-distance calculation formula, the distance of the sound source can be further obtained, and the angle and the distance can be used to solve the specific position of the sound source.
[0119] In this embodiment, since the positions of the microphones on the array are different, the flight distances and times of the signals emitted by the same sound source to different microphones are also different. These signals are presented in the form of one-dimensional time series data, and the data at each time point is a sample. The time delay Δt caused by the additional flight distance Δd causes a slight sample deviation of the sound signals captured by each microphone, and we call this feature sample shift.
[0120] In this embodiment, according to the geometric relationship, the time difference Δt at which different microphones m i and m j receive sound can be represented by the following formula:
[0121]
[0122] The sample shift n shift can be represented by the following formula:
[0123] n shift = Δt × fs,
[0124] The angle θ can be solved by the following formula:
[0125]
[0126] where L is the distance between two microphones, c sound is the sound speed (340m / s in air), and fs is the acoustic sampling rate, so that the sample offset n shift can be obtained as soon as θ is obtained.
[0127] In the embodiment, for the actually collected sound signal, a generalized cross-correlation algorithm can be used, which is widely used in signal time delay solving problems. Specifically, in sound source positioning, the target signal received by each microphone on the array comes from the same sound source, so the signal amplitude and other mode characteristics of the signals received by each microphone are very similar, and the signals have strong correlation. The aforementioned cross-correlation method can calculate the similarity between two sequences and judge the time step difference between two signals. According to this phenomenon, the cross-correlation function peak of each pair of microphone signals can be calculated to calculate the sample offset n shift . For a pair of microphones m i and m j , the signal of m i is composed of n samples, and the corresponding signal received by m j has a certain time delay and sample offset, so it is represented as n+n shift , and the cross-correlation formula between the two signals is as follows:
[0128] C[n shift ] = ∑(m i [n]·m j [n+n shift ]),
[0129] In the embodiment, the aforementioned function function can obtain a set of discrete point sequences C[n shift ] to represent the hysteresis of the signal sequence [n+n shift ] relative to the signal sequence [n].
[0130] Therefore, the index corresponding to the peak value of C[n shift ] can be found, that is, the sample offset value of the signal sequence [n+n shift ] relative to the signal sequence [n]:
[0131]
[0132] As Figure 3 shown, in the embodiment, the operation of the acoustic signal feature extraction further includes the following specific steps:
[0133] S21, simulate the sample offset;
[0134] In the embodiment, the simulation angle θ' ∈ [1, 360] (covering all possible angles in two-dimensional space) is set, and the simulation angle is brought into the previous geometric model to calculate the simulation sample offset angle matrix n sim :
[0135]
[0136] where L is the distance between two microphones, c sound is the sound speed, and fs is the acoustic sampling rate, which are all constants that can be obtained in advance. Since the time delay Δt is required, a reference microphone needs to be determined first, and then the signals received by other microphones relative to the signal received by the reference microphone are calculated.
[0137] As shown in Figure 4 and Figure 5 , in the embodiment, the hexagonal ring microphone in Figure 4 and Figure 5 is taken as an example, n sim should be a matrix of A x 6 dimensions, where A represents the number of angles θ', and 6 represents that each microphone is calculated relative to the reference microphone once, including the reference microphone itself.
[0138] In the embodiment, in order to further improve the resolution and accuracy, each microphone is taken as a reference microphone respectively, so that the dimension of n sim is expanded to A x 6 x 6. The specific calculation method uses the projection theorem and is expressed as follows:
[0139]
[0140] where m i ,m j are different microphones i, j ∈ {1, 2, 3, 4, 5, 6}, Δd is the additional flight distance of the signal, is the angle vector, is the distance vector between two microphones. In this way, each angle θ' has a corresponding set of 6 x 6 theoretical sample offsets
[0141] S22, similarity measure;
[0142] In the embodiment, the actual sample offset n shift and the theoretically calculated theoretical sample offset n sim are obtained by cross-correlation, and then the theoretical sample offset value that best matches the actual sample offset can be found by similarity calculation, and the θ' corresponding to the sample offset with the highest similarity is the sound signal arrival angle to be solved:
[0143]
[0144] In this embodiment, taking footstep detection as an example, the signal generated by each step can be used to obtain a... Ultimately, through this matching, an angle sequence θ aligned with the number of steps can be obtained. seq ={θ1,θ2,…,θ k}
[0145] S3, multimodal fusion computing;
[0146] like Figure 6 As shown, in this embodiment, WiFi and acoustic signal features are combined for computation. PLCR features and AoA angle features are obtained in the WiFi and acoustic feature calculation parts, respectively. By combining these two features to construct a geometric model, the user's trajectory position sequence can be obtained. A velocity-angle matrix is constructed, and based on the relationship between velocity and frequency, the velocity-angle matrix and frequency-angle matrix are transformed to obtain the movement direction angle at each moment from the user's trajectory position. Using the angle and the aforementioned matrix as an index, the velocity in the PLCR velocity spectrum is matched, transforming the PLCR velocity spectrum, which originally represented the frequency of WiFi signal path changes, into an independent gait velocity spectrum independent of trajectory and direction.
[0147] The path length change rate (PLCR) describes the rate at which the propagation path length of a WiFi signal changes after reflection from a target. This rate of change is one of the causes of Doppler frequency shift (DFS). In indoor positioning scenarios, when a person moves relative to the signal transmission link, it causes a change in the signal transmission path length, resulting in Doppler frequency shift and PLCR. By analyzing the characteristics of these two factors, information related to the target's direction of motion and velocity can be obtained.
[0148] like Figure 7 As shown, in the process of solving the user's movement trajectory position in this embodiment, the user's position is first solved by combining the WiFi Fresnel ellipse with the angle sequence rays. Specifically, the entire sensing area is divided into several grids, each with grid coordinates (x, y). Then, the Fresnel ellipse equation is constructed using d(t):
[0149]
[0150] Construct an angle sequence of rays using θ as a parameter:
[0151] y t =tanθ seq (t)·x t
[0152] Transform the two equations into the form of error calculation equation:
[0153]
[0154] Finally, substitute each grid coordinate in the area into the error calculation equation, and obtain the user's position (x t ,y t ) at each time through maximum likelihood estimation:
[0155]
[0156] In the gait feature solving process of the embodiment, PLCR is expressed in polar coordinates v ρ and v θ for convenience of calculation, which respectively represent the speed value and the speed direction:
[0157]
[0158] In the embodiment, a speed-angle matrix v mat is established, which is composed of speed features with a granularity of [0 m / s:0.02 m / s:2 m / s] and angle features with a granularity of [1°:1°:360°], and the dimension of the entire matrix is 101*360. According to the speed frequency formula, the v mat matrix is converted into a frequency matrix to realize the mapping of the PLCR speed spectrum:
[0159]
[0160] wherein δ x and δ y are direction coefficients determined by the WiFi device position:
[0161]
[0162] wherein, user location, and transmitter and receiver positions, respectively.
[0163] In the prior art literature in the background art, due to the existence of phase ambiguity with an amplitude of 2, the search space of AOA is defined in [-90°, 90°]. This means that the person being positioned can only move on one side of the receiving antenna array, which limits the deployment and user movement space in practical applications, while the search space of the present application is [0°, 360°], covering the entire two-dimensional plane.
[0164] Similarly, under a single WiFi link, the positioning performance of the technical solution recorded in the comparative document is lower than that of the present application.
[0165] In this embodiment, we can get a matrix at each time, so the whole f mat is a tensor distributed along the time step, with dimensions of 101*360*T (speed-angle-time), T represents the time step of the whole user movement process, which is equal to the length of PLCR and AoA.
[0166] From the foregoing, the user position trajectory is obtained, and the direction angle of displacement is calculated according to two adjacent coordinate points. In order to distinguish from the AoA in the foregoing, we will record this direction angle as β(t). In addition, the PLCR spectrum has been determined according to the WiFi stage.
[0167] As Figure 8 shown, in this embodiment, the independent speed spectrum gait feature derivation process in the multi-modal fusion calculation process includes the following specific steps:
[0168] S31, the obtained PLCR speed spectrum is converted to obtain the f mat (t) frequency spectrum;
[0169] S32, since the dimension of f mat is speed-angle-time, we can determine the time step and the corresponding angle through β(t) obtained by us, and then match the corresponding speed;
[0170] S33, combine the speed value of each time step together, which realizes the whole PLCR speed spectrum to trajectory, direction-independent independent speed spectrum.
[0171] S4, user gait recognition;
[0172] In this embodiment, by obtaining the speed spectrum of different users, a data-driven user recognition model is constructed, a lightweight neural network is adopted, and through training the model, the user identity is recognized according to the input speed spectrum.
[0173] As Figures 9 to 11As shown, in this embodiment, a lightweight neural network is built for training and recognition. Different volunteers (users) are used as labels, and their respective gait speed curves are used as training samples for the neural network to train the model. The trained model can recognize users according to the speed features calculated from the signal. Since the speed feature is one-dimensional time series information, and it is a typical classification task, we use a combination of convolutional neural network (CNN) and recurrent neural network (RNN). The network includes but is not limited to: 2 convolutional layers for feature extraction, each layer adds a batch normalization layer, a ReLU activation layer and a max pooling layer. In addition, it also includes an LSTM layer to capture time information. After LSTM, a fully connected layer is used to learn the spatio-temporal features extracted by CNN and LSTM, and finally link a softmax layer to calculate and output the class probability, output the user label with the maximum probability.
[0174] Embodiment 2
[0175] In this embodiment, the following experiments are performed to verify the performance of the proposed method, including different gait categories and user identities, which proves the feasibility and robustness of the method.
[0176] In this embodiment, in an indoor scene, two mini computers equipped with Intel 5300 network cards and three antennas are deployed to simulate home WiFi devices, with a cost similar to commercial WiFi routers. Two WiFi devices are used as transmitters and receivers for transmitting and receiving WiFi signals, with a signal frequency of 5.32 GHz and a bandwidth of 20 MHz. In terms of acoustic signals, a Seeed Respeaker circular array with six small microphones is deployed. The Seeed Respeaker circular array can be set up, for example, as a hexagonal microphone array with microphones located at the six corners, and the array base has a side length of 4.75 cm, which conforms to the structure of most commercial smart speakers, such as Alibaba Tmall Genius and Amazon Echo. The microphone sampling rate is set to 48 kHz, which covers the frequency range of possible footstep sounds, and has a good trade-off in terms of computational efficiency and operation time overhead. One of the WiFi devices is deployed with the microphone array to form a "smart voice assistant" prototype. The microphone array is operated through Raspberry Pi, and both WiFi devices are connected to a notebook computer equipped with an Intel i7-11800H CPU and 16G RAM through the SSH protocol of MobaXTerm, and are controlled through corresponding instructions to transmit and receive signals. Finally, Matlab R2023a is used for code writing and algorithm implementation.
[0177] In this embodiment, 12 volunteers of different heights, weights, genders, and ages were recruited to collect data and evaluate the system. Each volunteer was required to walk normally along several different trajectories, repeating the process a certain number of times. The collected data was divided into training and testing sets to train and test the user recognition network model.
[0178] In summary, this invention places the microphone and WiFi transceiver in the same location to build a prototype of a smart speaker or smart voice assistant. It has both WiFi signal transmission and reception and microphone sound recording functions. The WiFi and microphone are used to collect WiFi signals and footsteps generated when the user walks, respectively. Corresponding feature extraction and fusion algorithms are designed to extract information such as human movement speed, angle, and gait spectrum. Finally, the gait spectrum is mapped to the acoustic angle to obtain trajectory orientation-independent gait features, thereby achieving more robust user recognition.
[0179] This invention can non-invasively extract a user's gait features without relying on wearable sensors or camera devices, avoiding cumbersome equipment deployment requirements and effectively solving privacy and lighting conditions issues related to camera-based recognition. It can achieve accurate recognition results in a single-link environment, improving the feasibility of deploying and applying non-invasive identity recognition technology in real-world scenarios.
[0180] This invention employs a method of matching theoretical and actual sample point offsets, namely, by modeling using signal propagation modes and mathematical geometry, calculating the theoretical sample offset value n corresponding to all possible angles θ. sim Then offset the actual measured sample by n shift By matching the theoretical value, the closest angle value can be found, enabling more accurate solution for sample point offset.
[0181] This invention enables trajectory- and direction-independent gait feature extraction using limited equipment, improving user recognition accuracy in real-world scenarios. A microphone array can capture the sound sequence of a user's footsteps. Footsteps are discrete, meaning there are brief silent intervals between adjacent steps. Traditional sound source localization typically targets continuous sound sources. However, this invention addresses the continuous movement of footsteps, a different scenario from traditional research. Therefore, this invention treats movement as a combination of multiple static sound sources, calculating only the steps with sound, ignoring the silent intervals between steps. Thus, threshold detection or correlation matching is first used to determine the user's step points. This not only avoids computational interference from invalid information in silent intervals but also reduces computational load and improves efficiency.
[0182] The present application solves the technical problems of poor adaptability to complex scenes, great deployment difficulty, environmental interference and poor robustness in the prior art.
[0183] The above examples are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-modal gait user recognition method based on WiFi acoustics, characterized in that, The method comprises: S1, moving the identified target within the preset indoor device range, collecting WiFi signals by using the preset WiFi device to extract CSI features, calculating the reflection path length change rate PLCR according to the CSI features, constructing a Fresnel zone ellipse model, and processing to obtain the PLCR features; S2, set and use a microphone array to collect the footstep sound of the identified target, calculate a footstep sound AoA sequence according to a signal of the footstep sound, construct a ray model, model according to a preset signal propagation mode and mathematical geometric theory, and solve a theoretical sample offset value n corresponding to all angles θ sim , match the actually measured sample offset n shift with the theoretical sample offset value n sim , obtain the closest angle value, extract a signal arrival angle AoA of each footstep sound from the footstep sound signal through sound source positioning, process to obtain a continuous angle sequence distributed along a time domain, and convert original gait features into AoA angle features; S3, confirming the user movement trajectory through multi-modal fusion; combining the PLCR features and the AoA angle features to construct a geometric model, and calculating the trajectory position sequence of the user movement; constructing a speed-angle matrix, and performing a conversion operation on the speed-angle matrix and a frequency-angle matrix according to the relationship between the speed and the frequency, to calculate the movement direction angle at each time from the trajectory position sequence; taking the movement direction angle, the speed-angle matrix, and the frequency-angle matrix as indexes, matching the speed in the PLCR speed spectrum, and mapping and converting the PLCR speed spectrum into an independent gait speed spectrum; S4, obtaining the independent gait speed spectrum of different identified targets, constructing a data-driven user identification model, constructing and training a lightweight neural network, obtaining a suitable gait recognition model, and performing gait recognition to determine the identity of the identified target.
2. The WiFi acoustic based multi-modal gait user recognition method of claim 1, wherein, The S1 comprises: S11, using a preset CSI extraction tool to extract the CSI from the WiFi signals by using the following logic: H(f, t) = H s (f, t) + H d (f, t) = H s (f, t) + A(f, t)e -j2πf(t) where H s (f,t) and H d (f,t) are the static signal component and the dynamic signal component containing the user movement, respectively. S12, recording the reflection path length at different times as d(t), and determining the Doppler frequency shift DFS by using the following logic: where f D (t) is the Doppler shift DFS, λ is the signal wavelength, and τ is the signal time of flight. S13, determining the relationship between the reflection path length change rate PLCR and the Doppler frequency shift DFS by using the following logic: S14, integrating the reflection path length change rate PLCR in time to obtain the path length d(t) by using the following logic: 3.The WiFi acoustic based multi-modal gait user recognition method of claim 1, wherein, In the S2, according to a preset geometric relationship, the following logic is used to express the time difference At at which different microphones m i , m j receive the sound With the following logic, the sample offset n is expressed from the time difference At shift : n shift = Δt x fs According to the sample offset n shift The angle θ is processed with the following logic: In the formula, L is the distance between the two microphones, and c sound Where f is the speed of sound, and fs is the acoustic sampling rate; According to the sample offset n shift , the cross-correlation relationship between the microphone signals is expressed by using the following cross-correlation function logic, and a discrete point sequence C[n shift ] is obtained by processing according to the microphone cross-correlation relationship. C[n shift ] = ∑(m i [n] · m j [n + n shift ]), Using the discrete point sequence C[n shift ], the hysteresis of the signal sequence [n+n shift ] with respect to the signal sequence [n] is represented. Using the following logic, the peak value of the discrete point sequence C[n shift ] is calculated to find the sample offset value of the signal sequence [n+n shift ] relative to the signal sequence [n]:
4. The WiFi acoustic based multi-modal gait user recognition method of claim 1, wherein, The S2 comprises: S21, set the value range of the simulation angle θ', bring the simulation angle into the preset geometric model to obtain the simulation sample offset angle matrix n by using the following logic sim : In the formula, L is the distance between the two microphones, and c sound Where f is the speed of sound, and fs is the acoustic sampling rate; A preset reference microphone, calculating the time delay Δt of the signals received by the remaining microphones relative to the signals received by the reference microphone; S22, obtaining the most suitable theoretical sample offset that matches the actual sample offset by similarity calculation, and taking the simulation depth θ' corresponding to the suitable theoretical sample offset as the sound signal angle of arrival: for each of the step sound signals generated for each step, matching the applicable theoretical sample offset obtaining a sequence of step point number alignment angles θ seq = {θ1, θ2,..., θ k}.
5. The WiFi acoustic based multi-modal gait user recognition method of claim 4, wherein, In the S21, each of the microphones is taken as the reference microphone to extend the dimension of the emulated sample offset angle matrix n sim According to the projection theorem, the logic is used to process the theoretical sample offset where m i ,m j are different microphones i,j e {1,2,3,4,5,6}, Ad is the additional flight distance of the signal, is the angle vector, is the distance vector between two microphones.
6. The WiFi acoustic based multi-modal gait user recognition method of claim 1, wherein, The S3 comprises: S31, the obtained PLCR velocity spectrum is converted to obtain f mat (t) frequency spectrum; since the dimension of f mat is velocity-angle-time, through the obtained β(t), the time step and the corresponding angle can be determined, and then the corresponding velocity is matched. S32, combining the speed values at each time step together to realize the conversion of the entire PLCR speed spectrum into a trajectory-independent gait speed spectrum.
7. The WiFi acoustic based multi-modal gait user recognition method of claim 6, wherein, The S31 comprises: S311, combining the WiFi Fresnel ellipse and the angle sequence ray to solve the user position; wherein the entire perception area is divided into no less than two grids, each grid is marked as (x, y), and the Fresnel ellipse equation is constructed by using the reflection path length d(t): According to the movement direction angle θ, the angle sequence ray equation is constructed: y t = tan θ seq (t) · x t S312, converting and processing the Fresnel ellipse equation and the angle sequence ray equation to obtain an error calculation equation: S313, substitute the coordinates of each grid into the error calculation equation, perform maximum likelihood estimation, and calculate the user position (x t ,y t ) at each time point S314, in the gait feature solving process, polar coordinates v ρ , θ representing the reflection path length change rate PLCR; S315. Obtain and construct the velocity-angle matrix v based on the velocity characteristics and angle characteristics. mat According to the velocity-frequency formula, the velocity-angle matrix v mat The frequency matrix is transformed into a frequency matrix, and the PLCR velocity spectrum is mapped using the following logic: The user position trajectory is obtained by processing, and the displacement direction angle β(t) is calculated according to two adjacent coordinate points.
8. The WiFi acoustic based multi-modal gait user recognition method of claim 7, wherein, In the S314, the v in the polar coordinates ρ representing a speed value, v θ representing a speed direction: 9.The WiFi acoustic based multi-modal gait user recognition method of claim 7, wherein, In the S315, δ x and δ y are direction coefficients determined by the WiFi device position: wherein denotes the user position, and denotes the transmitter and receiver position, respectively.
10. A multi-modal gait user recognition system based on WiFi acoustics, characterized in that, The system comprises: The PLCR feature processing module is configured to make the identified target move within a preset indoor device range, collect WiFi signals by using a preset WiFi device, extract a CSI feature, calculate a reflection path length change rate PLCR according to the CSI feature, construct an elliptical model of a Fresnel zone according to the PLCR, and process the PLCR feature; The AoA angle feature processing module is configured to set and collect the footstep sound of the identified target by using the microphone array, calculate the footstep sound AoA sequence according to the signal of the footstep sound, construct a ray model, model according to a preset signal propagation mode and mathematical geometric theory, and solve the theoretical sample offset value n corresponding to all angles θ sim The actual measured sample offset n shift is matched with the theoretical sample offset value n sim to obtain the closest angle value, the signal arrival angle AoA of each footstep sound is extracted from the footstep sound signal by sound source positioning, a continuous angle sequence distributed along the time domain is obtained by processing, and the original gait feature is converted into the AoA angle feature. The multi-modal fusion module is configured to confirm a user movement trajectory by multi-modal fusion, construct a geometric model by combining the PLCR feature and the AoA angle feature, and obtain a trajectory position sequence of the user movement; construct a speed-angle matrix, perform a conversion operation on the speed-angle matrix and a frequency-angle matrix according to a relationship between speed and frequency, and obtain a movement direction angle at each time from the trajectory position sequence; and use the movement direction angle, the speed-angle matrix, and the frequency-angle matrix as indexes to match a speed in a PLCR speed spectrum, so as to map and convert the PLCR speed spectrum into an independent gait speed spectrum. The multi-modal fusion module is connected with the PLCR feature processing module and the AoA angle feature processing module. The gait recognition module is configured to obtain the independent gait speed spectrum of the identified target, construct a data-driven user recognition model, construct and train a lightweight neural network, obtain a suitable gait recognition model, perform gait recognition, and determine the identity of the identified target. The gait recognition module is connected with the multi-modal fusion module.
Citation Information
Patent Citations
Speed-independent gait recognition method based on WiFi signal
CN114757237A
Gait recognition method based on wireless and video feature fusion
CN114783054A
Multi-mode indoor positioning and tracking method based on Wi-Fi and acoustics
CN116430308A
System for improving sleep through feedback
CN118044225A