A multi-modal indoor positioning and tracking method based on wi-fi and acoustics

By combining Wi-Fi and acoustic multimodal positioning methods, and utilizing CSI and microphone array technology, high-precision indoor positioning and tracking is achieved without the need for cumbersome equipment deployment. This solves the problems of cumbersome equipment deployment and low accuracy in existing technologies, and is adaptable to various scenarios and user mobility.

CN116430308BActive Publication Date: 2026-03-17TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310303313.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2026-03-17
Estimated Expiration
2043-03-27

AI Technical Summary

Technical Problem

Existing Wi-Fi and acoustic positioning technologies suffer from problems such as cumbersome equipment deployment, low accuracy, and strict environmental requirements in indoor positioning and tracking, especially in single-link systems where it is difficult to achieve high-precision tracking without wearable devices.

Method used

Combining Wi-Fi and acoustic multimodal localization methods, a Fresnel ellipse model is constructed by extracting PLCR through CSI data collected from commercial Wi-Fi devices, and an angular sequence ray model is constructed by collecting footstep sounds from a microphone array. Finally, user location tracking is achieved through maximum likelihood estimation.

Benefits of technology

It achieves high-precision indoor positioning and tracking without the need for cumbersome equipment deployment and manual calibration, adapting to different scenarios and user movement, thus improving the feasibility and accuracy of the positioning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116430308B_ABST
    Figure CN116430308B_ABST
Patent Text Reader

Abstract

This invention discloses an indoor wireless positioning method based on WiFi and acoustic multimodal combination, comprising the following steps: A user moves freely within the sensing scene; a commercial WiFi device collects the user's CSI (Continuous Sensor Indicator) during movement, extracts the PLCR (Pulse Path Length) from the CSI, calculates the path length of the reflected signal, and constructs a Fresnel zone elliptical model; a microphone array collects the user's footsteps during movement, and the sample offset of different microphone pairs is solved by the peak value of the cross-correlation function of each pair of microphone signals; an angle is calculated based on the sample offset to construct a ray model; the Fresnel elliptical model and the ray model are combined to calculate the error, and the user's positioning and tracking are completed through maximum likelihood optimal estimation. Compared with traditional WiFi positioning methods, this invention eliminates the additional phase calibration cost and improves tracking accuracy; compared with traditional acoustic tracking systems, tracking is possible even if the user does not actively issue voice commands, achieving true passive tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of indoor tracking and positioning technology, specifically relating to a multimodal indoor positioning and tracking method based on Wi-Fi and acoustics. Background Technology

[0002] With the rapid development of IoT technology, intelligent applications based on wireless sensing technology are widely used, and increasingly mature communication protocols provide a solid foundation for wireless sensing services. Wi-Fi sensing is an important technology in wireless sensing, commonly used in Human Activity Recognition (HAR) tasks. When a Wi-Fi signal encounters an obstacle, it undergoes changes in propagation methods such as reflection and diffraction. Similarly, when a person is active, changes in body position and posture cause corresponding changes in the reflected Wi-Fi signal. Different types, amplitudes, and frequencies of movement affect the signal to varying degrees; therefore, by analyzing the characteristics and patterns of the signal, human movement can be perceived and recognized. In recent years, Channel State Information (CSI) has been commonly used for analysis. It is a fine-grained feature at the physical layer that reflects the combined effects of propagation distance, power attenuation, and scattering during signal propagation between the transmitter and receiver. CSI carries richer environmental and human information and has stronger anti-interference capabilities, making it suitable for indoor positioning tasks.

[0003] In existing acoustic localization methods, two common microphone arrays are used to receive sound: circular arrays (hexagonal microphones) and linear arrays, with six and four independent microphones respectively. Due to the different array arrangements, the timing of sound reception varies among the microphones. This reception delay can be used to calculate the difference in propagation distance for each microphone based on the principles of sound propagation. This step typically utilizes cross-correlation algorithms, followed by calculating the angle of arrival using geometric relationships to ultimately determine the direction of the sound source. However, current acoustic localization methods often impose strict limitations on the sound source, such as requiring a high signal-to-noise ratio and a stationary sound source. Further research is needed for the localization and tracking of moving targets.

[0004] In Wi-Fi-based indoor positioning technology, traditional methods employ fingerprinting, where fingerprints represent the signal characteristics (such as signal amplitude, phase difference, and intensity) corresponding to different locations within the room. This method requires pre-collecting fingerprint features from various indoor locations to build a fingerprint database. When the user is within the sensing range, the user's signal fingerprint features are extracted from the signal transmitter and receiver. Then, similarity measurement methods such as nearest neighbor and probabilistic methods are used to match these fingerprint features with the fingerprint database to find the user's current location corresponding to that fingerprint. However, building the fingerprint database is a labor-intensive process, requiring significant time and effort. Subsequent research focused on fine-grained derivation of Wi-Fi signal propagation models, using Channel State Information (CSI) as the primary feature. Relevant variable parameters closely coupled with human movement were extracted, such as Angle of Arrival (AoA), Time of Flight (ToF), and Doppler frequency offset. Based on these variables, kinematic information such as the azimuth angle and speed of human movement were derived to achieve positioning and tracking. However, these methods typically require extracting the aforementioned parameters from multiple transceiver links (a path between a set of transmitters and receivers is counted as one link), which increases the difficulty of deployment and practical application in real-world scenarios, resulting in poor feasibility. Researchers subsequently proposed the Widar2.0 method, which can extract multi-dimensional parameters from a single link, achieving decimeter-level single-link indoor positioning; however, it requires rigorous calibration, otherwise the accuracy is not ideal. Therefore, this invention combines acoustic positioning and WiFi positioning technologies, proposing an indoor wireless positioning method based on a combination of WiFi and acoustic multimodal technologies to achieve indoor positioning and tracking under a single link. Summary of the Invention

[0005] The purpose of this invention is to provide an indoor wireless positioning method and system based on WiFi and acoustic multimodal combination. This invention can realize real-time tracking and positioning of users without wearable devices, without the need for cumbersome equipment deployment and manual calibration. It can achieve high-precision positioning performance under a single link, thereby improving the feasibility of deploying and applying wireless positioning systems in real-world scenarios.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] The purpose of this invention is to provide an indoor wireless positioning method based on WiFi and acoustic multimodal combination, comprising the following steps:

[0008] S1. Users can move freely within the perceived scene area;

[0009] S2. Collect the CSI of the user when moving using commercial WiFi equipment, extract the PLCR from the CSI, calculate the path length of the reflected signal, and then construct the Fresnel zone ellipse model based on the obtained path length of the reflected signal.

[0010] S3. Use a microphone array to collect the user's footsteps when moving. Solve the sample offset of different microphone pairs by using the peak value of the cross-correlation function of the two microphone signals. Then solve the angle based on the sample offset and construct an angle sequence ray model.

[0011] S4. By combining the Fresnel ellipse model and the angle sequence ray model, the error is calculated, and then the user's location and tracking are completed through maximum likelihood optimal estimation.

[0012] Preferably, in step S2, the CSI of the user during movement is collected using a commercial WiFi device, the PLCR is extracted from the CSI, the path length of the reflected signal is calculated, and then a Fresnel zone ellipse model is constructed based on the obtained path length of the reflected signal. The specific steps are as follows:

[0013] A21. The CSI (Cost Indicator) of a user during movement is collected using a commercial WiFi device with a single-link Wi-Fi signal. The CSI expression is shown in the following formula:

[0014]

[0015] Among them, H S (f,t) and H d (f,t) represent the static and dynamic components of CSI, respectively; The amplitude and phase of the dynamic signal are represented by d(t); the path length of the dynamic signal is represented by d(t).

[0016] A22. Extract PLCR from CSI, using f respectively. D R(t) and R(t) represent Doppler frequency offset (DFS) and PLCR, respectively. D R(t) and R(t) are obtained by the following formula:

[0017] R(t)=L′ d (t)

[0018]

[0019] Among them, L′ d (t) represents the relationship between L d The differentiation operation of R(t) is performed because R(t) describes the rate of change of the signal propagation path, and the rate is obtained by differentiating the displacement distance; f represents the communication frequency; c light Represents the speed of light (3 × 10⁻⁶) 8 m / s);

[0020] A23. Calculate the path length L of the reflected signal. d (t0): The initial location of the user at time t0. Then the initial signal path length L at time t0 d (t0) is represented by the following formula:

[0021]

[0022] As the user moves and their location changes, the signal reflection path length also changes. Therefore, the reflected signal path length L at time t is... d (t) represents the initial length L. d (t0) and t 0~ The sum of the changes in the path length of the reflected signal within time t, where the change is PLCRR(t) over time t. 0~ The integral over time period t, L d (t) is specifically represented as follows:

[0023]

[0024] A24. Based on the obtained reflected signal path length, construct the Fresnel zone elliptical model: using the reflected signal path length L... d Using the link length between the transceivers (t) as a parameter, a Fresnel elliptic model is constructed, specifically represented as follows:

[0025] .

[0026] Preferably, in step S3, the sample offset of different microphone pairs is solved by the peak value of the cross-correlation function of the pairwise microphone signals, and then the angle is solved based on the sample offset to construct an angle sequence ray model. The specific steps are as follows:

[0027] A31. Calculate the measured sample offset n of the footstep sound signal reaching each microphone on the microphone array. shift The specific steps are as follows:

[0028] A311. Determine the reference microphone and set the reference microphone as m. i and m j Let L represent the distance between the two, then the reference microphone m i and m j The time difference Δt between receiving the sound is:

[0029] Among them, c sound Speed ​​of sound;

[0030] A312. Calculate the measured sample offset n based on the peak value of the cross-correlation function of the pairwise microphone signals. shiftFor a pair of microphones m i and m j Let m i The signal consists of n samples, while m samples produce the sample shift. j The sample points are represented as n+n shift A set of discrete point sequences is obtained through the cross-correlation function. This represents n+n shift The peak value of the sequence is n, relative to the lag of n. shift Specifically, it is expressed as follows:

[0031] C[n shift ]=Σ(m i [n]·m j [n+n shift ])

[0032]

[0033] A32. Constructing the theoretical sample offset n using geometric methods virtual The vector connecting the microphone pairs (denoted as) The incident ray vector of the sound signal received by the microphone. Perform projection and solve:

[0034] ;

[0035] A33. Match the measured sample offset of the obtained footsteps with the theoretical sample offset matrix, and measure n. shift and n virtual The similarity is used to find the theoretical sample offset value that best matches the measured sample offset, which is the angle of arrival AoAθ of the sound ray to the microphone.

[0036]

[0037] A34. Construct an angle sequence ray model using θ as a parameter, as shown below:

[0038] y t =tanθ·x t .

[0039] Preferably, in step S4, the Fresnel ellipse model and the angular sequence ray model are combined to calculate the error, and then the user's location tracking is completed through maximum likelihood optimal estimation. The specific steps are as follows:

[0040] By combining the constructed Fresnel ellipse model and the angle sequence ray model, and transforming the form of the two equations, we obtain the equation for calculating the error:

[0041]

[0042] Finally, the user's location (x) is obtained through maximum likelihood estimation. t ,y t ):

[0043]

[0044] The present invention also aims to provide an indoor wireless positioning system based on a combination of WiFi and acoustic multimodal technologies, comprising a WiFi signal processing unit, an acoustic signal processing unit, and a multimodal fusion positioning unit, wherein the output terminals of the WiFi signal processing unit and the acoustic signal processing unit are both connected to the input terminal of the multimodal fusion positioning unit.

[0045] The WiFi signal processing unit is used to collect CSI of human movement through single-link WiFi signal, then extract PLCR in the time domain from CSI, and then calculate the time-varying signal reflection path length of PLCR through gradient integration, and construct Fresnel ellipse model based on the obtained signal reflection path length and link deployment location parameters.

[0046] The acoustic signal processing unit is used to estimate the angle of arrival (AoA) of the sound signal of the user's footsteps and to construct an angle sequence ray model based on the continuous footstep angle sequence.

[0047] The multimodal fusion positioning unit is used to combine the Fresnel ellipse model and the angle sequence ray model to calculate the error. Through maximum likelihood optimal estimation, it completes the user's positioning and tracking, and connects all the coordinates of the entire movement process to output the user's movement trajectory.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] (1) This invention is the first to combine WiFi with acoustics to achieve multimodal positioning and tracking. Compared with the manual calibration and dense equipment deployment of traditional positioning methods, this invention does not require cumbersome pre-setting and can achieve real-time positioning and tracking of users without wearable devices under single-link deployment, which is more in line with the needs and feasibility of real indoor scenarios. Compared with traditional WiFi positioning methods, this invention eliminates the additional phase calibration cost and improves tracking accuracy; compared with traditional acoustic tracking systems, this invention can track users even if they do not actively issue voice commands, achieving true passive tracking, and there are no additional restrictions on the placement of the microphone (such as the requirement to be close to the ground or against a wall).

[0050] (2) This invention can be applied to different scenarios and situations, such as microphone height, device position, different trajectories, different moving speeds, different shoe types, etc., and can achieve fast, accurate and high-performance positioning effect. Attached Figure Description

[0051] Figure 1 This is a flowchart of an indoor wireless positioning method based on WiFi and acoustic multimodal combination proposed in this invention;

[0052] Figure 2 A schematic diagram illustrating the time delay of a microphone array receiving an audio signal;

[0053] Figure 3 This is a schematic diagram illustrating the solution for the theoretical sample offset.

[0054] Figure 4 This is a schematic diagram of the combined model of Fresnel elliptical clusters and angular rays;

[0055] Figure 5 This is a diagram illustrating the effect of passive positioning and tracking. Detailed Implementation

[0056] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0057] The technical terms and their explanations used in this invention are as follows:

[0058] Wireless sensing technology is a widely used technology in the field of smart Internet of Things. It uses professional transceiver devices to capture wireless signals such as Wi-Fi, RFID, and sound. By analyzing the signal characteristics, it can realize the sensing of human activities without the need for users to carry mobile phones, watches, sensors, or other devices.

[0059] Indoor positioning and tracking technology: a key part of the wireless sensing field. Its basic idea is to extract the characteristics of the signal propagation mode changes caused by the target's motion from the radio frequency signal (such as the signal angle of arrival, signal flight time, Doppler frequency offset, etc.), and then analyze the user's motion pattern to complete the tracking and positioning of the target.

[0060] A Fresnel zone refers to a series of concentric, equifocal ellipsoidal regions between the transmitter and receiver, with the transmitter and receiver located at the two foci of the ellipse. When a target is located in different Fresnel zones, the propagation path lengths of the reflected and direct signals differ, resulting in varying superposition effects between the transmitter and receiver. As an object passes through several Fresnel zones, the signal strength exhibits periodic changes of strengthening and weakening, which can be used for localization.

[0061] The path length change rate (PLCR) describes the rate at which the length of the Wi-Fi signal propagation path changes after reflection from a target. It is the cause of Doppler frequency shift (DFS). In indoor positioning scenarios, movement of a person relative to the Wi-Fi link causes changes in the signal transmission path, resulting in Doppler frequency shift and PLCR. Both reflect the changes in the signal path caused by the target relative to the link, allowing us to further deduce information such as the target's direction of motion and velocity.

[0062] Acoustic localization principle: It is based on the sound emitted by the target to be located. The signal emitted from the transmitting end (sound source) will form an acoustic angle of arrival on the antenna array of the receiving end (microphone array). The value of this angle can be calculated according to the direction of the signal incident, thereby helping to find the location of the sound source.

[0063] Sound signal sample shift: For microphone arrays used for sound reception, when a sound source emits sound, the distance between the microphones and the sound source varies, resulting in differences in the time it takes for the sound signal to be received. A visual example is the binaural effect in humans; the sound from one direction arrives at the left and right ears with different delays, helping us to identify the direction of the sound source. Since a sound signal is composed of several discrete samples, the time delay of the signal received by each microphone creates a sample shift. Based on the speed of sound propagation and the sample shift, geometric calculations can be used to obtain the angle of arrival of the sound signal, thereby enabling the localization of the sound source.

[0064] Example 1

[0065] Reference Figure 1 An indoor wireless positioning method based on WiFi and acoustic multimodal combination includes the following steps:

[0066] S1. Users can move freely within the perceived scene area;

[0067] S2. Collect CSI signals when users move using commercial WiFi devices, extract PLCR from CSI signals, calculate the path length of reflected signals, and construct a Fresnel zone ellipse model.

[0068] In this embodiment, two commercial devices equipped with Intel 5300 network cards (one transmitter and one receiver) were deployed in a 4×4m open indoor environment to collect Wi-Fi information. Both devices were pre-installed with the Ubuntu 14.04.3 operating system. The transmitter was equipped with a single antenna, while the receiver had three antennas arranged linearly, with an antenna spacing of 2.5cm. The signal frequency range was 5.31–5.33 GHz. For acoustic signals, a Seeed Respeaker circular array (hexagonal microphones) with six microphones was used, conforming to the structure of most commercial smart speakers, such as Google Home and Amazon Echo. The microphones and Wi-Fi transmitter were deployed in the same location to simulate a "voice assistant," supporting both acoustic and Wi-Fi functions. The microphone array was a regular hexagon, with microphones located at the six corners, each with a side length of 4.75cm. The microphone sampling rate was set to 48kHz, covering the entire possible frequency range of footsteps. The microphone array and Wi-Fi transceiver are both connected to a laptop equipped with an Intel i7-11800H CPU and 16GB of RAM via SSH protocol. Data transmission and reception can be controlled remotely via commands from the host computer. Finally, code is written and executed in Matlab.

[0069] Specifically, based on a 4×4m positioning area in the real environment, a coordinate system is determined with the lower left corner of this square area as the origin. The transmitter and microphone are deployed at position (0,0), and the receiver is deployed at position (4,0). Here, the coordinates of the transmitter and receiver are respectively... and This is indicated so that subsequent formulas can be defined.

[0070] The reflection path length variation rate (PLCR) in the time domain is extracted from the CSI. Then, the time-varying signal reflection path length is calculated by gradient integration. Based on the reflected signal path length and the location parameters of the link deployment, a Fresnel zone ellipse (model) is constructed.

[0071] The CSI (Continuous Sensor Identity) of a user during movement is collected using a commercial WiFi device with a single-link Wi-Fi signal. The CSI expression is shown in the following formula:

[0072]

[0073] Among them, H S (f,t) and H d (f,t) represent the static and dynamic components of CSI, respectively; d(t) represents the amplitude and phase of the dynamic signal; d(t) represents the path length of the dynamic signal.

[0074] Human movement causes changes in the propagation path of dynamic signals. For ease of writing, the length of the signal reflection path is denoted as L. d (t).

[0075] When a human body moves within the perception range, it traverses several Fresnel zones, resulting in changes in Doppler frequency offset (DFS) and path length change rate (PLCR). These two parameters describe the rates of change of signal frequency and signal path length, respectively, and can be calculated and converted between each other. Here, they are represented by f. D Let f(t) and R(t) represent DFS and PLCR, respectively. D R(t) and R(t) are calculated using equations (2) and (3) respectively:

[0076] R(t)=L′ d (t)

[0077]

[0078] Among them, L′ d (t) represents the relationship between L d The differentiation operation of R(t) is performed because R(t) describes the rate of change of the signal propagation path, and the rate can be obtained by differentiating the displacement distance; f represents the communication frequency; c light Represents the speed of light (3 × 10⁻⁶) 8 m / s).

[0079] The reflected signal travels from the transmitter, is reflected by the user, and is then received by the receiver. Therefore, the path length of the reflected signal is the sum of the distances from the user to the transmitter and receiver. Given the user's initial location at time t0. The initial signal path length L at this moment d (t0), specifically represented as shown in equation (4):

[0080]

[0081] As the user moves and their location changes, the signal reflection path length also changes. Therefore, the reflected signal path length L at time t is... d (t) represents the initial length L. d (t0) and t 0~ The sum of the changes in the path length of the reflected signal within time t is specifically expressed as shown in equation (5):

[0082]

[0083] Wherein, the change is PLCRR(t) at t 0~ The integral over the time interval t yields the path length of the reflected signal required to construct the Fresnel elliptical cluster.

[0084] S3. Use a microphone array to collect the user's footsteps when moving. Solve the sample offset of different microphone pairs by using the peak value of the cross-correlation function of the two microphone signals. Then, solve the angle of arrival AoA of the sound signal of the user's footsteps when walking based on the sample offset, and construct an angle sequence ray model.

[0085] Specifically, in this embodiment, we first detect footsteps and denoise the detected valid footstep sounds. The measured sample offset is then obtained by calculating the peak value of the cross-correlation function. A hexagonal microphone (a regular hexagon with six microphones located at each vertex) is used to collect the user's footstep sounds. When calculating the sound signal sample offset, we first need to determine the reference microphone and calculate the offset of the other microphones relative to it. Here, the reference microphone is set to m. i and m j The distance between the two is denoted by L.

[0086] Since the distance between the microphones is only a few centimeters, which is very small relative to the sensing area, the propagation path of the sound signal can be regarded as a series of parallel rays. Therefore, the reference microphone m... i and m j The time difference (delay value) Δt for receiving sound can be written as:

[0087]

[0088] Among them, c sound The speed of sound is 340 m / s; for the microphone sampling rate f s m i and m j The measured sample offset n between shift The relationship between AoAθ is expressed as shown in equation (7):

[0089]

[0090] The `round(·)` operator represents rounding, and its specific principle is as follows: Figure 2 As shown.

[0091] Therefore, to obtain θ, we first need to obtain the measured value of the sample offset. Specifically, since the target signal received by each microphone on the array comes from the same sound source in sound source localization, there is a strong correlation between the signals of each channel. Based on this principle, we calculate the measured sample offset n based on the peak value of the cross-correlation function of the pairwise microphone signals. shift Then, Δt is obtained, and θ is solved. For a pair of microphones m i and m j Let m iThe signal consists of n samples, while m samples produce the sample shift. j The sample points are represented as n+n shift A discrete sequence of points can be obtained through the cross-correlation function. This represents n+n shift The peak value of the sequence is n, relative to the lag of n. shift Specifically, it is expressed as follows:

[0092] C[n shift ]=∑(m i [n]·m j [n+n shift ])

[0093]

[0094] For the six microphones in the hexagonal array, the sample offset is calculated for each pair, resulting in a 6×6 sample offset matrix. This operation is performed on the k detected footsteps, ultimately yielding a complete sample offset matrix (6×6×k) corresponding to the footstep sounds of a human walking.

[0095] Next, a theoretical sample offset matrix is ​​constructed, which corresponds one-to-one with the theoretical sample offset of each possible acoustic signal incident angle. The calculated footstep sample offsets are then matched with this matrix to obtain the relative angle between each footstep and the microphone. A ray model is then constructed based on the slope calculated from the angle using the ray model.

[0096] Specifically, refer to Figure 3 Sample offset is caused by the additional propagation distance and time delay of the sound signal to reach different microphones. Therefore, we can construct the theoretical sample offset n geometrically. virtual The vector connecting the microphone pairs (denoted as) The incident ray vector of the sound signal received by the microphone. The projection and solution are represented as follows:

[0097]

[0098] A32, By measuring n shift and n virtual Based on the similarity, find the theoretical sample offset value that best matches the measured sample offset:

[0099]

[0100] S4. By combining the Fresnel ellipse model and the angle sequence ray model, the error is calculated, and then the user's location and tracking are completed through maximum likelihood optimal estimation.

[0101] With the reflected signal path length L d (t) Using the link length between the transceivers as a parameter, construct a Fresnel ellipse model:

[0102]

[0103] An angle sequence ray is constructed using θ as a parameter, as shown below:

[0104] y t =tanθ·x t .

[0105] Then, the two equations are transformed. A schematic diagram of the joint model of Fresnel elliptical groups and angular rays can be found here. Figure 4 The equation for calculating the error is obtained, as follows:

[0106]

[0107] Finally, the user's location (x) is obtained through maximum likelihood estimation. t ,y t ):

[0108]

[0109] The final passive positioning and tracking results are shown below. Figure 5 .

[0110] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concept, should be covered within the scope of protection of the present invention.

Claims

1. A method for indoor wireless positioning based on WiFi and acoustic multi-modal combination, characterized in that, The method comprises the following steps: S1, the user freely moves in the perception scene range; S2, the CSI when the user moves is collected by using a commercial WiFi device, the PLCR is extracted from the CSI, the reflection signal path length is calculated, and the Fresnel zone ellipse model is constructed according to the obtained reflection signal path length; S3, the footstep sound when the user moves is collected by using a microphone array, the sample offset of different microphone pairs is solved by using the cross-correlation function peak value of two microphones, the angle is solved according to the sample offset, and the angle sequence ray model is constructed; the specific steps are as follows: A31, the measured sample offset of the footstep sound signal to each microphone of the microphone array is obtained The specific steps are as follows: A311, determine the reference microphone, and set the reference microphone as and The distance between the two is represented by The reference microphone and The time difference between the two receiving sound t is: wherein is the speed of sound; A312. Calculate the measured sample offset based on the peak value of the cross-correlation function of the pairwise microphone signals. For a pair of microphones and ,set up The signal consists of n samples, and the sample shift is caused by... The sample points are represented as A set of discrete point sequences is obtained through the cross-correlation function. , in order to indicate The peak value of the sequence is relative to the lag of n. Specifically, it is expressed as follows: ; A32. Theoretical sample offset is constructed geometrically , through the vector of the connection between the microphones, denoted , the incident ray vector of the sound signal received by the microphones Project and solve: ; A33. Match the measured sample offset of the obtained footsteps with the theoretical sample offset matrix, and measure... and The similarity is used to find the theoretical sample offset value that best matches the measured sample offset. That is, the angle of arrival AoA of the sound ray at the microphone: ; A34、to The angle sequence ray model is constructed as a parameter, and is specifically represented as follows: ; S4, the Fresnel ellipse model and the angle sequence ray model are combined, the error is calculated, and the positioning and tracking of the user are completed through maximum likelihood optimal estimation; the specific steps are as follows: The Fresnel ellipse model and the angle sequence ray model are combined, the forms of the two equations are converted, and the error calculation equation is obtained: Finally, the user's location is obtained by maximum likelihood estimation : 。 2.The indoor wireless positioning method based on WiFi and acoustic multi-modal combination according to claim 1, characterized in that, In step S2, the CSI when the user moves is collected by using a commercial WiFi device, the PLCR is extracted from the CSI, the reflection signal path length is calculated, and the Fresnel zone ellipse model is constructed according to the obtained reflection signal path length; the specific steps are as follows: A21, the CSI when the user moves is collected by using a single-link Wi-Fi signal by using a commercial WiFi device, and the CSI expression is as follows: wherein, and denote the static and dynamic components of the CSI, respectively; denote the amplitude and phase of the dynamic signal; denote the dynamic signal path length; A22. Extracting the PLCR from the CSI, respectively using and denote the Doppler frequency offset and the PLCR, and are calculated by wherein, denotes a derivation operation on denotes a derivation operation on describes the rate of change of the signal propagation path, and the rate is derived from the displacement distance; denotes the communication frequency; denotes the speed of light; A23, calculating the reflection signal path length : given the initial position of the user at the time instant then the initial signal path length at the time instant is represented as follows: As the user moves, the location constantly changes, and the signal reflection path length also changes accordingly, and the reflection signal path length at time t is the initial length and t 0~ The sum of the change amounts of the reflection signal path length within time t, wherein the change amount is PLCR At t 0~ The integral within the time period t, is specifically represented as follows: ; A24. Construct a Fresnel zone ellipse model based on the path length of the reflected signal: path length of the reflected signal A25. Construct a Fresnel ellipse model with the length of the link between the transceivers as a parameter, specifically as follows: 。 3. A system for operating a WiFi and acoustic multi-modal combined indoor wireless positioning method according to claim 1 or 2, characterized in that, The device comprises a WiFi signal processing unit, an acoustic signal processing unit and a multi-modal fusion positioning unit, and the output ends of the WiFi signal processing unit and the acoustic signal processing unit are connected with the input end of the multi-modal fusion positioning unit, wherein, The WiFi signal processing unit is used for collecting the CSI when the human body moves by using a single-link Wi-Fi signal, then extracting the PLCR in the time domain from the CSI, calculating the signal reflection path length based on the time change of the PLCR by using gradient integration, and constructing the Fresnel zone ellipse model according to the obtained signal reflection path length; The acoustic signal processing unit is used for estimating the angle of arrival (AoA) of the sound signal of the footstep sound when the user walks, and constructing the angle sequence ray model according to the continuous angle sequence of the footstep sound; The multi-modal fusion positioning unit is used for combining the Fresnel ellipse model and the angle sequence ray model, calculating the error, completing the positioning and tracking of the user through maximum likelihood optimal estimation, and outputting the moving track of the user by connecting all the coordinates in the whole moving process.

Citation Information

Patent Citations

  • Wireless indoor positioning and sensing method and system and storage medium

    CN110049550A

  • Indoor target passive tracking method based on WiFi multi-dimensional parameter characteristics

    CN110809240A