A Bluetooth fingerprint library updating method based on voiceprint quality estimation

By introducing voiceprint quality estimation into the Bluetooth fingerprint library, using Bluetooth path loss model and acoustic signal positioning technology, the Bluetooth fingerprint library is updated in real time, solving the problem of insufficient accuracy and stability in the dynamic environment of the traditional Bluetooth fingerprint library, and realizing adaptive maintenance and efficient positioning services.

CN120264221BActive Publication Date: 2025-08-22HUZHOU INST OF ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510729979.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-22
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing Bluetooth fingerprint library update method is difficult to maintain high accuracy and stability when facing dynamic changes in the indoor environment. The traditional regular update method is time-consuming and labor-intensive and cannot cope with real-time environmental changes, resulting in a decrease in positioning accuracy.

Method used

Using a method based on voiceprint quality estimation, a base station group with Bluetooth broadcasting and acoustic signal transmission functions is deployed, and a Bluetooth path loss model is used for preliminary positioning, and acoustic signals are received synchronously for quality evaluation, effective sound signals are selected for TDOA positioning calculation, and the Bluetooth signal strength is reversed to update the fingerprint library in real time.

Benefits of technology

Effectively respond to dynamic changes in the indoor environment, realize adaptive maintenance of fingerprint data, improve positioning accuracy and stability, reduce manual maintenance costs, and adapt to real-time update requirements in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264221B_ABST
    Figure CN120264221B_ABST
Patent Text Reader

Abstract

The present application relates to the field of Bluetooth positioning technology, and specifically discloses a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation. The method deploys a base station group with both Bluetooth broadcasting and acoustic signal transmission functions in the target area, uses the Bluetooth path loss model for preliminary positioning, and estimates the position based on the received signal strength of each location node to establish a Bluetooth fingerprint library. At the same time, the method synchronously receives the acoustic signals of each location, and performs a quality assessment on the received acoustic signals. The effective acoustic signals are screened out for TDOA positioning calculation, the acoustic signal positioning coordinates of each location are determined, and the acoustic signal positioning accuracy is judged based on the position quality assessment method. Then, the Bluetooth strength of the current location is inferred based on the accurate acoustic signal location information to update the Bluetooth fingerprint database in real time. This method uses the propagation stability of the acoustic signal to compensate for the directional sensitivity of the Bluetooth signal, can effectively cope with the dynamic changes of the indoor environment, and realize the adaptive maintenance of fingerprint data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Bluetooth positioning technology, and more specifically, to a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation. Background Art

[0002] With the rapid development of network communication technologies and the widespread adoption of smart mobile devices, location-based services have provided significant convenience in outdoor scenarios. However, in indoor scenarios, such as finding a car in an underground parking lot, quickly locating trapped individuals, and finding people in large transportation hubs, providing high-precision location-based services remains challenging due to the limitations of indoor positioning technology. To address these challenges, a variety of technologies and solutions have been proposed, including Wi-Fi, Bluetooth, and ultra-wideband (UWB). While each of these technologies has its advantages, they struggle to strike a balance between cost, accuracy, coverage, and smartphone compatibility. For example, while UWB-based positioning systems can provide centimeter-level positioning accuracy, their receiver and transmitter costs are relatively high, and they lack smartphone compatibility, significantly limiting their practical applications.

[0003] Among existing technologies, Bluetooth fingerprint-based positioning methods are favored for their low cost, ease of deployment, and high device compatibility. They achieve relatively accurate positioning without the need for additional hardware, particularly in environments with high mobility. The flexible deployment capabilities of Bluetooth Low Energy (BLE) beacons allow for easy installation indoors and utilize spread spectrum communication technology to transmit positioning data. They maintain stable signal transmission even in obstructed walls and complex environments, making them a simple and cost-effective indoor positioning solution.

[0004] However, the main drawback of Bluetooth fingerprint-based positioning methods lies in the construction and updating of the fingerprint library. Dynamic environmental changes, such as the movement of people, furniture, and wireless interference, can cause significant fluctuations in the Bluetooth signal reception strength (RSSI). Consequently, the initially constructed fingerprint library may not accurately reflect the real-time signal environment. Furthermore, the polarization effect of the device antenna and slight changes in its position can cause variations in signal strength, further impacting the stability of the fingerprint library. While regularly updating the fingerprint library can improve positioning accuracy, this process is cumbersome and time-consuming, and the maintenance cost is very high, especially when deployed over a large area.

[0005] Traditional methods for updating Bluetooth fingerprint libraries primarily involve manual re-collection and periodic batch updates. However, manual re-collection requires technicians to manually measure and record new RSSI values ​​at different locations. While this method ensures a certain level of accuracy, it is time-consuming and labor-intensive, and the maintenance cost is extremely high, especially in environments with frequent changes. Periodic batch updates, on the other hand, involve updating the entire fingerprint library at set intervals. While this reduces some labor costs, it cannot adapt to immediate environmental changes, resulting in the possibility that the updated fingerprint library may still be inaccurate. Furthermore, these methods cannot effectively handle signal instability caused by real-time signal fluctuations and multipath effects. Excessively high update frequencies also consume computing and storage resources, making it difficult to maintain good positioning results in dynamic and complex environments.

[0006] Therefore, a Bluetooth fingerprint library updating method based on voiceprint quality estimation is expected. Summary of the Invention

[0007] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation, which deploys a base station group with both Bluetooth broadcasting and acoustic signal transmission functions in the target area, uses the Bluetooth path loss model for preliminary positioning, and performs position estimation based on the received signal strength of each location node to establish a Bluetooth fingerprint library; at the same time, it synchronously receives the acoustic signals of each location, and performs quality evaluation on the received acoustic signals, screens out effective acoustic signals for TDOA positioning calculation, determines the acoustic signal positioning coordinates of each location, and judges the acoustic signal positioning accuracy based on the position quality evaluation method, and then reversely infers the Bluetooth strength of the current location based on the accurate acoustic signal location information to update the Bluetooth fingerprint database in real time. This method uses the stability of acoustic signal propagation to compensate for the directional sensitivity of the Bluetooth signal, can effectively cope with the dynamic changes of the indoor environment, and realize the adaptive maintenance of fingerprint data.

[0008] Accordingly, according to one aspect of the present application, a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation is provided, which includes:

[0009] Construct a Bluetooth fingerprint library, where each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position};

[0010] receiving acoustic signals at the respective positions;

[0011] determining whether the acoustic signal at each position meets a signal quality standard, and if the acoustic signal at each position meets the signal quality standard, performing position estimation based on the acoustic signal at each position to obtain acoustic signal location coordinates at each position;

[0012] The deviation of the hyperbola intersection of the acoustic signal positioning coordinates is calculated based on a parameterized method to generate an acoustic signal position calculation quality assessment value;

[0013] Based on the comparison between the acoustic signal position calculation quality evaluation value of each position and the sound positioning quality threshold, it is determined whether the RSSI value of each position needs to be updated.

[0014] Accordingly, according to another aspect of the present application, a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation is provided, which includes:

[0015] Construct a Bluetooth fingerprint library, where each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position};

[0016] receiving acoustic signals at the respective positions;

[0017] Determining whether a signal quality standard is met based on the RSSI values ​​of the respective positions and the acoustic signals at the respective positions, and in response to the signal quality standard being met, performing position estimation based on the acoustic signals at the respective positions to obtain acoustic signal positioning coordinates of the respective positions;

[0018] The deviation of the hyperbola intersection of the acoustic signal positioning coordinates is calculated based on a parameterized method to generate an acoustic signal position calculation quality assessment value;

[0019] Based on the comparison between the acoustic signal position calculation quality evaluation value of each position and the sound positioning quality threshold, it is determined whether the RSSI value of each position needs to be updated.

[0020] Compared with the existing technology, the Bluetooth fingerprint library update method based on voiceprint quality estimation provided by this application deploys a base station group with both Bluetooth broadcasting and acoustic signal transmission functions in the target area, uses the Bluetooth path loss model for preliminary positioning, and estimates the position based on the received signal strength of each location node to establish a Bluetooth fingerprint library; at the same time, it synchronously receives the acoustic signals of each location, and performs quality assessment on the received acoustic signals, screens out the effective acoustic signals for TDOA positioning calculation, determines the acoustic signal positioning coordinates of each location, and judges the acoustic signal positioning accuracy based on the position quality assessment method, and then reversely infers the Bluetooth strength of the current location based on the accurate acoustic signal location information to update the Bluetooth fingerprint database in real time. This method uses the stability of acoustic signal propagation to compensate for the directional sensitivity of the Bluetooth signal, can effectively cope with the dynamic changes of the indoor environment, and realize the adaptive maintenance of fingerprint data. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 Schematic diagram of the deployment of a broadcast positioning system based on Bluetooth and acoustic signal fusion according to an embodiment of the present application.

[0023] Figure 2 Schematic diagram of polar coordinates comparing the directionality of acoustic signals and Bluetooth RSSI signals according to an embodiment of the present application.

[0024] Figure 3 Flowchart of a method for updating a Bluetooth fingerprint library based on voiceprint quality estimation according to an embodiment of the present application.

[0025] Figure 4 This is a flowchart of step S1 in the Bluetooth fingerprint library updating method based on voiceprint quality estimation according to an embodiment of the present application.

[0026] Figure 5 Schematic diagram of TDOA positioning estimation results with no error and with noise error.

[0027] Figure 6 This is a flowchart of determining whether a signal quality standard is met based on the RSSI value of each position and the acoustic signal of each position in the Bluetooth fingerprint library update method based on voiceprint quality estimation according to Example 2 of the present application.

[0028] Figure 7 This is a flowchart of step S33 in the Bluetooth fingerprint library updating method based on voiceprint quality estimation according to Example 2 of the present application. DETAILED DESCRIPTION

[0029] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0030] Example 1

[0031] In response to the shortcomings of existing Bluetooth fingerprint library update methods, this application proposes a Bluetooth fingerprint library update method based on voiceprint quality estimation. It deploys a base station group with both Bluetooth broadcasting and acoustic signal transmission functions in the target area, uses the Bluetooth path loss model for preliminary positioning, and estimates the position based on the received signal strength of each location node to establish a Bluetooth fingerprint library. At the same time, it synchronously receives the acoustic signals of each location, and performs quality assessment on the received acoustic signals. It selects effective acoustic signals for TDOA positioning calculation, determines the acoustic signal positioning coordinates of each location, and judges the acoustic signal positioning accuracy based on the position quality assessment method. Then, based on the accurate acoustic signal position information, it reversely infers the Bluetooth strength of the current location to update the Bluetooth fingerprint database in real time. This method uses the stability of acoustic signal propagation to compensate for the directional sensitivity of the Bluetooth signal, can effectively cope with the dynamic changes of the indoor environment, and realize the adaptive maintenance of fingerprint data.

[0032] Specifically, Figure 1 Schematic diagram of the deployment of a broadcast positioning system based on Bluetooth and acoustic signal fusion according to an embodiment of the present application. Figure 1 As shown in the figure, the positioning system based on Bluetooth and acoustic signal fusion broadcasting mainly consists of an intelligent positioning terminal, a multimodal positioning base station with Bluetooth and acoustic signal broadcasting functions, a central computing server, etc. N multimodal positioning base stations are deployed at the four corners of the ceiling edge to emit acoustic signals with a certain delay for distinguishing the signal source. At the same time, the built-in Bluetooth beacon module broadcasts BLE messages. The intelligent positioning terminal receives the acoustic signal from the acoustic signal base station and receives the BLE message at the same time. The Bluetooth beacon module communicates with the central positioning server that performs positioning. Figure 2 As shown in the figure, due to the polarization effect of the Bluetooth onboard antenna, the RSSI value centered on the traditional Bluetooth beacon is not uniformly distributed. If the polarization directions of the transmitting and receiving antennas do not match, the signal strength will decrease. At the same time, when the angle between the antenna and the receiving device varies significantly, the received signal strength will fluctuate, making it difficult to ensure the accuracy of the Bluetooth fingerprint library, which in turn affects the accuracy and stability of the overall positioning system. However, the sound pressure level distribution based on acoustic signal positioning has stable directionality during propagation in indoor environments, which can be used to compensate for the inaccuracy of fingerprint library data caused by the rapid fluctuation of Bluetooth RSSI with the environment.

[0033] Figure 3 Flowchart of the Bluetooth fingerprint library update method based on voiceprint quality estimation according to an embodiment of the present application. Figure 3As shown, according to an embodiment of the present application, a Bluetooth fingerprint library updating method based on voiceprint quality estimation includes the following steps: S1, constructing a Bluetooth fingerprint library, wherein each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position}; S2, receiving the acoustic signal of each position; S3, judging whether the acoustic signal of each position meets the signal quality standard, and after the acoustic signal of each position meets the signal quality standard, performing position estimation based on the acoustic signal of each position to obtain the acoustic signal positioning coordinates of each position; S4, calculating the hyperbola intersection deviation of the acoustic signal positioning coordinates based on a parameterized method to generate an acoustic signal position calculation quality evaluation value; S5, determining whether the RSSI value of each position needs to be updated based on a comparison between the acoustic signal position calculation quality evaluation value of each position and the acoustic positioning quality threshold.

[0034] In the aforementioned Bluetooth fingerprint library update method based on voiceprint quality estimation, step S1 constructs a Bluetooth fingerprint library. Each Bluetooth fingerprint in the library consists of {RSSI value for each location, two-dimensional coordinates for each location}. It should be understood that the Bluetooth fingerprint library is the core data foundation of the indoor positioning system. Its essence is to achieve positioning inference by recording the mapping relationship between signal characteristics (RSSI values) and physical coordinates at different locations. This application uses the Bluetooth Path Loss Model (FSPL) for initial library construction. This aims to use a mathematical model to predict signal attenuation patterns, reduce manual sampling costs, and provide a benchmark framework for subsequent dynamic updates of the Bluetooth fingerprint database.

[0035] Specifically, traditional Bluetooth fingerprint libraries rely on static RSSI data collection in fixed environments, but in actual dynamic scenarios, factors such as personnel flow, obstacle movement, and electromagnetic interference can cause significant fluctuations in signal strength (RSSI). If the uncalibrated raw RSSI value is used directly to build the fingerprint library, a systematic deviation will be introduced. For example, when the polarization direction of the device antenna does not match the environment, the RSSI may attenuate by 3-5dB; when there are temporary obstacles in the signal path (such as human body obstruction), the instantaneous RSSI fluctuation can reach more than 10dB. These non-steady-state noises will cause the fingerprint library to mismatch with the actual signal environment, resulting in positioning drift errors. Therefore, this application adopts the classic Freese free space path loss model to establish a physical mapping relationship between signal strength and distance to compensate for device hardware differences and environmental baseline noise. For example, by using the reference distance The calibration signal strength at ) and path loss exponent , establish the mathematical relationship between signal strength and distance:

[0036]

[0037] in, is the current distance The received signal strength value at is the reference distance The received signal strength value at It is the path loss index, which is used to reflect the attenuation characteristics of the signal in a specific environment.

[0038] Figure 4 FIG1 is a flow chart of step S1 in the method for updating the Bluetooth fingerprint library based on voiceprint quality estimation according to an embodiment of the present application. Figure 4 As shown, step S1 includes: S11, constructing a fingerprint matrix for each position; S12, calculating the distance between the RSSI vector of each reference position in the fingerprint matrix and the observed RSSI vector to obtain a distance vector; S13, selecting k reference positions corresponding to the smallest distances from the distance vector; S14, calculating weights for the k smallest distances, and performing weighted averaging on the k reference positions based on the weights to obtain the two-dimensional coordinates of the observed position.

[0039] Specifically, first, the target area is divided into K grid nodes, each location node corresponds to a unique two-dimensional coordinate. After deployment, the base station group periodically broadcasts Bluetooth signals, and the smart terminal receives RSSI measurement values ​​from L beacon base stations at each location node. , recorded as the reference position RSSI vector of each location node, where is the index of the corresponding position node, The beacon index of the transmitted signal. The reference position RSSI vectors of K position nodes are combined to form a fingerprint matrix , where each row corresponds to the signal characteristics of a location node:

[0040]

[0041] Next, the fingerprint matrix is ​​associated with the two-dimensional coordinates of each location node and stored to form an initial Bluetooth fingerprint library. Based on this, whenever a new Bluetooth signal is received, the RSSI vector of each reference location in the fingerprint matrix is ​​compared with the newly received RSSI observation vector. A similarity calculation method is used to evaluate the degree of match between the new signal and each fingerprint in the library. This allows the location of the new signal to be estimated based on the coordinates corresponding to similar fingerprints.

[0042] Specifically, first, the distance between the RSSI vector of each reference position in the fingerprint matrix and the observed RSSI vector is calculated to obtain a distance vector. It should be understood that the observed RSSI vector is a set of RSSI values ​​scanned and recorded in real time by the smart terminal (such as a mobile phone) carried by the user at the current position, reflecting the actual signal strength distribution in the current environment. In order to find the position most similar to the currently observed Bluetooth signal in the fingerprint library, so as to estimate the current position of the user, this application adopts Euclidean distance as a metric, and calculates the distance between the RSSI vector of each reference position in the fingerprint matrix and the observed RSSI vector to obtain a distance vector containing K distance values, where each distance value corresponds to the degree of difference in signal strength between a reference position in the fingerprint matrix and the current position of the user.

[0043] Next, the k reference positions corresponding to the smallest distances are selected from the distance vector. It should be understood that, considering that similar positions usually have similar signal characteristics, this application narrows the positioning search range and improves positioning accuracy by selecting the k reference points with the smallest distances. In addition, through multi-reference point collaborative decision-making, the local characteristics of the signal distribution around the current observation point can be captured, avoiding the problem of a single nearest neighbor being sensitive to noise, reducing positioning jumps caused by single-point signal fluctuations, and improving the robustness of position estimation.

[0044] Then, weights are calculated for the k smallest distances, and based on the weights, a weighted average is performed on the k reference positions to obtain the two-dimensional coordinates of the observation position. It should be understood that the k reference points are considered to be candidate positions that best match the signal characteristics of the user's current position, but the credibility of each reference point is different. The closer the reference point is to the observation point and the more similar the signal environment is, the greater the contribution of its position information to the estimation of the user's current position. Therefore, the present application designs a dynamic weight function that comprehensively considers the Euclidean distance between the reference point and the observation point and the similarity of the signal strength, and assigns higher weights to reference points with small distances and similar signal characteristics.

[0045] In particular, considering that some base station signals may not be detected due to obstruction or excessive distance in the distance calculation, not every message contains the RSSI values ​​of all beacons. If equal weights are directly assigned to all RSSI values, then the beacon fingerprint with fewer RSSI values ​​(i.e., the lack of RSSI values) may obtain the same weight in the distance calculation as the fingerprint with more RSSI values ​​and closer distance, which will lead to inaccurate results. Therefore, this application adds a constant term To adjust the weight distribution strategy, the beacon with missing RSSI value is given a lower weight to reduce its negative impact on the overall positioning result. The weighting function is expressed as:

[0046]

[0047] in, is the candidate reference point position, is the corresponding distance estimate, Indicates data integrity. Indicates the total number of beacons, Indicates the number of valid RSSI values.

[0048] In the above weighted function, the inverse distance term is introduced to strengthen the influence of the nearest neighbor, so that the distance The smaller the candidate reference point position, the greater the weight. At the same time, by adding the constant This is used to adjust for weight deviations caused by missing RSSI values, ensuring that reference points with higher signal integrity receive greater weight in positioning decisions. Compensation terms are also used to preserve the value of reference points with partially missing data, preventing them from being completely ignored due to incomplete data and maintaining global stability. Finally, the weights of the k candidate reference point locations are normalized, and a weighted average is calculated for these k reference locations to obtain the two-dimensional coordinates of the observed location.

[0049] In the above-mentioned Bluetooth fingerprint library update method based on voiceprint quality estimation, step S2 receives acoustic signals from each location. It should be understood that Bluetooth signals are susceptible to multipath effects and environmental interference, resulting in significant fluctuations in RSSI values. However, acoustic signals have a stable sound pressure level distribution and directionality when propagating indoors, which can compensate for the shortcomings of Bluetooth signals. Therefore, the present application synchronously receives acoustic signals from each location. That is, the multimodal base station simultaneously transmits an acoustic signal with a specific delay (such as a linear frequency modulation pulse) while transmitting the Bluetooth signal, thereby providing an auxiliary data source independent of Bluetooth for high-precision positioning.

[0050] In the above-mentioned Bluetooth fingerprint library update method based on voiceprint quality estimation, the step S3 determines whether the acoustic signal at each position meets the signal quality standard, and after the acoustic signal at each position meets the signal quality standard, the position estimation is performed based on the acoustic signal at each position to obtain the acoustic signal positioning coordinates at each position. It should be understood that since the acoustic signal in the indoor environment may be affected by multipath reflection, noise interference or nonlinear distortion of the equipment, the quality of the acoustic signal received at different positions in the indoor area is different. Therefore, it is necessary to perform a quality evaluation between the acoustic signal received by the smart terminal and the reference signal to determine whether the positioning signal quality standard is met, to ensure that only high-confidence data enters the positioning process. In a specific example of the present application, the acoustic signal at each position is subjected to channel impulse response analysis, signal peak distribution analysis, signal reverberation time analysis and signal frequency distortion analysis to determine whether the acoustic signal at each position meets the signal quality standard. The specific evaluation process is as follows:

[0051] Channel Impulse Response Analysis: The Channel Impulse Response (CIR) can be used to obtain the multipath effect and delay spread of the signal during transmission to further evaluate the signal quality. In an indoor environment, the signal is , the output signal received after propagation through the indoor channel is , the two satisfy the following relationship:

[0052]

[0053] in, is the received output signal, is the impulse response of the channel, is the length of the channel impulse response. In order to simplify the convolution solution process, the input signal can be Convert to a Toeplitz (T) matrix , the output signal and channel impulse response Expressed in vector form:

[0054]

[0055] in, is the length of the output signal.

[0056] Then, construct the Toeplitz matrix of the input signal :

[0057]

[0058] at this time It can be rewritten into matrix form:

[0059]

[0060] Directly solve the above formula, the matrix may appear Therefore, the present application adopts the least square method to solve the channel impulse response of the received signal. :

[0061]

[0062] in, Representation matrix The transposed matrix of .

[0063] Signal peak distribution analysis: When receiving acoustic signals, the non-line-of-sight (NLOS) scenario usually causes the signal to be affected by a strong multipath effect, which is manifested as multiple peaks in the signal peak distribution on the time axis. The peak distribution of the received signal can be analyzed to determine whether it is in an NLOS environment. Extract the peak distribution of the signal and record the delay of multiple main peaks and its corresponding amplitude In the NLOS scenario, the signal will have multiple paths, resulting in peaks that are densely distributed in time. Therefore, the time difference between adjacent peaks can be calculated to determine whether there is a strong multipath effect. For all extracted peaks, the time difference between each two adjacent peaks is calculated. :

[0064]

[0065] Set the time difference threshold , used to determine the significance of multipath effect. If the time difference between adjacent peaks is , then the signal is judged to be affected by strong multipath effects and is considered to be in an NLOS scenario. At the same time, the relative size of the calculated peak amplitude is used to determine the NLOS scenario. Generally, in a Line-Of-Sight (LOS) scenario, the amplitude of the direct path signal is large, and the amplitude of the subsequent reflected path gradually decreases. In an NLOS scenario, the amplitude of the subsequent path may approach or even exceed the first peak. Calculate the first peak With subsequent peak The relative amplitude ratio :

[0066]

[0067] Set an amplitude ratio threshold here , if a subsequent peak satisfies , we can determine that there is a strong multipath effect. and amplitude ratio If a signal meets any or all of the above conditions, it is judged to be in an NLOS scenario, indicating that its quality may be low and unsuitable for high-precision positioning.

[0068] Signal reverberation time analysis: In acoustic analysis, reverberation time (RT) is a key indicator for evaluating environmental acoustic characteristics and signal quality. Reverberation time RT60 is the time required for the acoustic signal to decay from its initial intensity to 60dB. Reverberation time RT60 can be obtained from the energy decay curve. Usually, the impulse response can be used to The energy decay is used to estimate RT60:

[0069]

[0070] in, Indicates the signal at time The energy in the environment is attenuated. The larger the value, the more sound waves are reflected in the environment and the stronger the reverberation is, which will have a greater impact on the signal transmission quality. Here we set an RT60 threshold , when the RT60 value exceeds the threshold When the environmental reverberation is too severe, the signal quality cannot meet the requirements of high-precision positioning of acoustic signals. , it can judge the impact of environmental reverberation and ensure that the signal quality is within an acceptable range to support high-precision acoustic signal positioning.

[0071] Signal frequency distortion analysis: This is to evaluate the quality of acoustic signals based on the frequency distortion of the received signal. The signal fidelity is mainly determined by comparing the difference in frequency response between the received signal and the reference signal. If the signal undergoes frequency distortion during transmission due to multipath effects, noise interference, or nonlinear effects of the equipment, the amplitude and phase of the signal may deviate, resulting in a decrease in quality. and are the spectrum of the reference signal and the received signal respectively, then the frequency response function for:

[0072]

[0073] in, Amplitude Indicates the gain change of the frequency component, angle Indicates phase change. Amplitude distortion is mainly reflected in Amplitude If the amplitude deviation is too large in some frequency bands, it means that the strength of the received signal in these frequency bands has been distorted. An amplitude deviation threshold is defined. , that is, when Deviation 1 exceeds At the same time, the phase response is detected, and the phase distortion is mainly reflected in In the case of deviation from the reference phase, a large phase offset will cause signal distortion and affect positioning accuracy. A phase deviation threshold is defined. , that is, when The deviation in a certain frequency band exceeds ° is considered as phase distortion.

[0074] After the acoustic signals at each location meet the signal quality standard, the screened high-quality acoustic signals are further converted into precise location coordinates, thereby providing reliable data support for updating the Bluetooth fingerprint library. In a specific example of the present application, position estimation is performed based on the acoustic signals at each location to obtain the acoustic signal positioning coordinates at each location, including: performing position estimation optimization on the time difference of arrival of the acoustic signals based on a parameterized method to obtain the acoustic signal positioning coordinates at each location.

[0075] Specifically, first, position estimation is performed based on the time difference of arrival of acoustic signals to obtain the initial acoustic signal positioning coordinates for each location. It should be understood that since the signal propagation speed is known (for example, the speed of sound in air is approximately 340 meters per second), by comparing the time differences of arrival of signals from different base stations, the distance difference of the location relative to each base station can be calculated. A hyperbolic solution is then performed based on this distance difference information to determine the initial acoustic signal positioning coordinates for that location. The specific steps are as follows:

[0076] (1) Calculate the arrival time: The acoustic signal from the base station is at time The sound signal is sent, and the arrival time of the positioning terminal receiving the signal is , the propagation time between the two Expressed as:

[0077]

[0078] Assuming that sound waves travel in straight lines, the distance estimation can be simplified to:

[0079]

[0080] in, represents the speed of sound, To locate the terminal, is the known location of the acoustic signal base station.

[0081] (2) Calculate the arrival time difference: To eliminate the time difference between base station sound signals and the arrival time difference between base station sound signals and base station sound signals and simplifies the positioning problem by calculating the arrival time difference between different base stations. The specific formula is:

[0082]

[0083] in, and From the base station and base stations The time when the signal reaches the intelligent positioning terminal. Base stations can be calculated and Distance difference to the positioning terminal:

[0084]

[0085] TDOA between multiple base stations can establish multiple such equations, each equation representing the distance difference between the positioning terminal and two base stations.

[0086] (3) Solving the transmitter position: Assuming there are N base stations, multiple arrival time difference equations can be formed. By combining these equations, mathematical optimization algorithms (such as least squares method, nonlinear least squares method, etc.) can be used to solve the position coordinates of the positioning terminal, that is, to solve the initial acoustic signal positioning coordinates of each position.

[0087] Position estimation is achieved through the above-mentioned acoustic signal-based TDOA method, and its positioning accuracy is high under noise-free conditions. However, the performance of this positioning method is easily affected by adverse environmental factors such as indoor non-line-of-sight (NLOS) and multipath propagation, resulting in deviations in the received timestamp and positioning inaccuracy. Figure 5 Schematic diagram of TDOA positioning estimation results with no error and with noise error. Figure 5 As shown, the sample geometry for N=4 is described, where For acoustic signal base station, and are the actual and estimated positioning terminal locations, respectively. From a geometric perspective, the noise-free TDOA equation defines a hyperbola, on which the positioning terminal should be located in two-dimensional (2-D) space. Using at least two hyperbolas, the position of the positioning terminal can be determined by the intersection of the lines. It can be observed that due to multipath propagation, erroneous timestamps are collected at the corresponding smart terminal, which can cause the hyperbolas to be severely stretched, causing them to no longer intersect at a precise point, significantly degrading positioning performance.

[0088] To address TDOA positioning under adverse conditions, this application further optimizes the time difference of arrival of acoustic signals based on a parameterized approach to obtain the acoustic signal positioning coordinates for each location. This involves reshaping the relationship between the transmitter position and the TDOA measurement, and then optimizing the newly introduced parameters (each associated with a corresponding hyperbola). This involves searching for a point on the TDOA-defined hyperbola that is closest to all other hyperbolas. The specific steps are as follows:

[0089] (1) Constructing a hyperbolic model: In this scenario, multiple base stations The acoustic signal is sent and reaches the positioning terminal. The TDOA value between each base station is calculated by the arrival time of the signal received by the positioning terminal, and the position of the terminal is then located. The TDOA hyperbola equation can be written as follows to describe the position of the smart terminal The geometric relationship between each base station:

[0090]

[0091] in, 、 Represents base stations and The midpoint coordinates of represents the semi-major axis of the hyperbola, represents the half-segment axis, is the rotation angle of the hyperbola. These hyperbolas define the time difference between each pair of base stations and can be used to construct multiple hyperbola intersections to infer the location of the smart terminal.

[0092] (2) Parameterized hyperbola equation: The positioning terminal calculates the arrival time difference by receiving the signal from the base station , and based on these TDOA measurements, find the point on the hyperbola closest to all base stations. To achieve this goal, the optimal point can be found by parameterizing the hyperbola equation:

[0093]

[0094] (3) Optimizing parameters: The goal is to find points on each hyperbola that minimize the Euclidean distance between them. Under the parameterized model, the minimization problem can be defined as follows:

[0095]

[0096] This nonlinear least squares problem can be solved by an optimization algorithm (such as gradient descent or least squares method) to obtain a set of parameters , thus obtaining the optimal point on each hyperbola.

[0097] (4) Calculate the final position estimate of the intelligent terminal: Find the optimal point on the hyperbola After that, the least squares cost function is further applied to determine the final position of the smart terminal , the optimization goal at this time is to minimize the error between all hyperbola intersections and obtain the final position through the least squares method. The estimate of can be solved by the following minimization problem:

[0098]

[0099] in, is the set of points calculated in the previous step according to the parameterized hyperbola equation, This algorithm ensures positioning accuracy by minimizing the error between the intersection points of the hyperbolas. Even in the presence of unfavorable environmental factors such as multipath propagation, TDOA positioning performance can be improved through optimization algorithms.

[0100] In the above-mentioned Bluetooth fingerprint library update method based on voiceprint quality estimation, the step S4 calculates the hyperbola intersection deviation of the acoustic signal positioning coordinates based on a parameterized method to generate an acoustic signal position calculation quality evaluation value. It should be understood that in actual scenarios, the position estimation of the smart terminal is updated in real time, and the measurement results will be different each time. In order to avoid low-quality acoustic signal positioning data from contaminating the Bluetooth fingerprint library, this application evaluates the quality of the obtained positioning solution by optimizing the minimum value of the objective function, and optimizes the final value of the objective function. The smaller the value, the smaller the deviation of each point on the hyperbola, which indicates that the calculated positioning point is more consistent and accurate. Specifically, the acoustic signal position calculation quality evaluation value of each position is calculated using the following formula:

[0101]

[0102] in, is the number of acoustic signal base stations, The hyperbola intersection parameters optimized for the parametric method, 、 is the hyperbolic index, 、 are coordinate components, is the square of the two-norm, A quality assessment value is calculated for the acoustic signal position.

[0103] It should be understood that the essence of the Q value is the sum of the squared distances between all hyperbola parameter points, reflecting the consistency of the hyperbola intersection points. By adjusting the parameter t, all hyperbolas are made to intersect at the same point as much as possible, thereby minimizing Q. When Q = 0, all hyperbolas intersect at one point, resulting in extremely high positioning accuracy (in an ideal noise-free scenario). When Q > 0, multipath or noise interference is present.

[0104] In the above-mentioned Bluetooth fingerprint library update method based on voiceprint quality estimation, the step S5 determines whether the RSSI value of each position needs to be updated based on the comparison between the acoustic signal position calculation quality evaluation value of each position and the acoustic positioning quality threshold. It should be understood that in actual positioning, the optimal objective function value can be used to determine whether the RSSI value of each position needs to be updated. To construct the normalized quality score of the position. According to the system accuracy requirements and test results, set a sound positioning quality threshold . Set thresholds according to different application scenarios .like , then the calculation quality of the position meets the requirements, and the RSSI fingerprint library can be updated based on the acoustic signal positioning coordinates of the position; if Q≥Qth, then the calculation quality of the position does not meet the requirements. At this time, the RSSI fingerprint library cannot be updated directly based on the acoustic signal positioning coordinates of the position. The positioning result needs to be discarded and the acoustic signal position measurement must be performed again to improve the data quality.

[0105] Obtaining accurate acoustic signal positioning coordinates Then, based on the acoustic positioning coordinates of the current position The RSSI value at the current location is recalculated with the base station position to update the Bluetooth positioning fingerprint library, that is, the RSSI value of the current location is calculated based on the above-mentioned Free Space Path Loss Model according to the actual distance between the current location and each Bluetooth base station. and compare it with the RSSI value in the fingerprint library In the actual Bluetooth positioning fingerprinting process, RSSI signals are often affected by various factors such as environmental noise, physical obstacles such as walls, and interference between devices, resulting in random fluctuations in RSSI values. Furthermore, RSSI values ​​can fluctuate rapidly over time. Therefore, it is necessary to establish a model based on an adaptive Kalman filter (AKF) in state space to smooth the RSSI value and accurately integrate the current location measurement value with historical data. The specific steps are as follows:

[0106] Define the state space model: The state transition equation is set to describe the dynamic change of RSSI value over time, that is, the slow change trend of RSSI over time and the fluctuation caused by noise. The change of RSSI is set to be relatively stable, but a certain amount of random fluctuation is allowed:

[0107]

[0108] in, It's time The actual RSSI value, is the process noise, and its variance is .

[0109] The observation equation is set to describe each measured RSSI value, taking into account measurement errors and environmental noise:

[0110]

[0111] in, is the measured RSSI value, is the measurement noise, whose variance is .

[0112] Adaptive noise adjustment: In order to more accurately handle the rapid fluctuations of RSSI values ​​and environmental noise, the process noise can be adaptively adjusted and measurement noise , making the filter more flexible. Then, calculate the residuals (i.e. the difference between the measured and predicted values):

[0113]

[0114] The square of the residual is used to dynamically adjust the noise covariance to adapt to the rapid fluctuations of RSSI and random noise:

[0115]

[0116]

[0117] in, and is a smoothing factor that ensures that the noise estimate is moderately smooth and not overly sensitive.

[0118] Adaptive Kalman filter: First, the current RSSI value is predicted based on the estimated value of the previous time step And predict the current state covariance based on the covariance estimate of the previous step :

[0119]

[0120]

[0121] Second, the Kalman gain is calculated based on the predicted covariance and the covariance adjusted for measurement noise to determine the weights of historical and current measurements:

[0122]

[0123] The Kalman gain is used to fuse the historical prediction and the current measured RSSI value to obtain the updated RSSI value. :

[0124]

[0125] Updated covariance Reflects the uncertainty of the filter's estimate of the current RSSI, and the next prediction will be based on this covariance:

[0126]

[0127] The final result The fusion of the current location and historical measurements represents the best estimate of the RSSI signal in its current state and is stored in the Bluetooth fingerprint library, replacing the original historical value. This updated RSSI value is more stable and better adapts to environmental changes and rapid RSSI fluctuations. This allows the Bluetooth fingerprint library to be updated in real time, ensuring that the RSSI value more accurately reflects the signal propagation characteristics of the environment, particularly addressing issues such as multipath and signal attenuation in complex indoor environments. By updating the fingerprint library in real time, the Bluetooth positioning system can maintain high positioning accuracy in diverse environments, particularly in complex indoor environments, providing more reliable positioning services.

[0128] Example 2

[0129] In particular, traditional acoustic signal quality assessment methods based on channel impulse response, peak distribution, reverberation time, and frequency distortion analysis can effectively detect multipath effects, environmental reverberation, and signal distortion. However, these methods require fixed thresholds (such as time difference thresholds, amplitude ratio thresholds, and reverberation time thresholds) for each analysis. In complex real-world environments, these thresholds are difficult to dynamically adjust, leading to increased false positives. For example, sudden noise in a shopping mall can instantly change the reverberation time, but fixed thresholds cannot adaptively filter out such interference. Furthermore, traditional methods focus only on local features in the time domain (peak time difference) and frequency domain (spectral distortion), failing to fully utilize the signal's joint time-frequency characteristics (such as the time-frequency distribution pattern of transient noise). This makes it difficult to capture the global characteristics of complex interference, and they are highly sensitive to non-stationary signals (such as instantaneous multipath changes caused by human movement). They rely on multiple measurements and averaging, resulting in limited real-time performance. Based on this, this application proposes an optimized joint quality assessment method for acoustic and Bluetooth signals. By introducing a deep learning algorithm to extract and classify the joint time-frequency feature patterns of acoustic and Bluetooth signals, and combining their characteristics, a dynamic assessment of their signal quality is achieved. This method not only fully considers the time-frequency distribution characteristics of the acoustic signal, but also comprehensively utilizes information such as the RSSI value of the Bluetooth signal. Leveraging the powerful feature learning and generalization capabilities of the deep learning model, it overcomes the problem of traditional methods where fixed thresholds cannot adapt to complex environmental interference. By jointly processing acoustic and Bluetooth signals, changes in signal quality can be more accurately captured, and multipath effects, noise interference, and non-stationary signal changes under different environmental conditions can be addressed in real time, significantly improving the accuracy and real-time performance of signal quality assessment.

[0130] In Example 2, a Bluetooth fingerprint library update method based on voiceprint quality estimation includes: constructing a Bluetooth fingerprint library, wherein each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position}; receiving the acoustic signal of each position; judging whether the signal quality standard is met based on the RSSI value of each position and the acoustic signal of each position, and in response to meeting the signal quality standard, performing position estimation based on the acoustic signal of each position to obtain the acoustic signal positioning coordinates of each position; calculating the hyperbola intersection deviation of the acoustic signal positioning coordinates based on a parameterized method to generate an acoustic signal position calculation quality evaluation value; and determining whether the RSSI value of each position needs to be updated based on a comparison between the acoustic signal position calculation quality evaluation value of each position and the acoustic positioning quality threshold.

[0131] Figure 6 This is a flow chart of determining whether the signal quality standard is met based on the RSSI value of each position and the acoustic signal of each position in the Bluetooth fingerprint library update method based on voiceprint quality estimation according to Example 2 of the present application. Figure 6As shown, based on the RSSI values ​​of the various positions and the acoustic signals of the various positions, judging whether the signal quality standard is met, includes: S31, converting the acoustic signal into a time-frequency diagram through wavelet analysis to obtain the acoustic signal time-frequency diagram; S32, performing joint time-frequency feature extraction based on void convolution coding on the acoustic signal time-frequency diagram to obtain an acoustic signal time-frequency feature coding diagram; S33, performing feature enhancement on the acoustic signal time-frequency feature coding diagram to obtain an acoustic signal time-frequency feature enhancement coding diagram; S34, normalizing the Bluetooth signal RSSI values ​​at the various positions to obtain a normalized Bluetooth signal RSSI feature vector; S35, fusing the acoustic signal time-frequency feature enhancement coding diagram with the Bluetooth signal RSSI feature vector to form a joint time-frequency feature coding diagram; S36, judging whether the joint signal meets the signal quality standard based on the joint time-frequency feature coding diagram.

[0132] Specifically, step S31 utilizes wavelet analysis to convert the acoustic signal into a time-frequency diagram to obtain a time-frequency diagram of the acoustic signal. It should be understood that acoustic signals contain rich characteristic information in both the time and frequency domains. Traditional analysis methods using either the time or frequency domain alone cannot simultaneously capture the local temporal and frequency characteristics of the signal. However, wavelet transform, through multi-scale decomposition, can accurately characterize the transient characteristics of the signal (such as impulse noise and multipath reflections) in the time and frequency domains. Converting the acoustic signal into a time-frequency diagram intuitively presents the energy distribution of the signal at different times and frequencies, thereby clearly distinguishing between the effective components (such as direct sound waves) and interference (such as multipath reflections and burst noise) in the signal. For example, direct sound waves appear as continuous high-energy bands in a time-frequency diagram, while transient noise appears as isolated spikes.

[0133] Specifically, step S32 performs time-frequency feature extraction on the acoustic signal time-frequency graph based on dilated convolutional coding to obtain an acoustic signal time-frequency feature coding graph. That is, in order to further extract representative time-frequency features from the acoustic signal time-frequency graph, thereby providing more discriminative feature information for subsequent signal quality assessment, the present application employs a dilated convolutional neural network to extract time-frequency features from the acoustic signal time-frequency graph. It should be understood that traditional convolutional neural networks (CNNs) have a limited receptive field and are difficult to capture long-range time-frequency dependencies (such as periodic interference). However, by introducing interval sampling, dilated convolution can expand the receptive field without increasing the number of parameters, effectively extracting time-frequency features in a wider range of contexts. This allows the capture of long-range dependencies across time and frequency (such as periodic reverberation) while retaining local details (such as signal mutation edges), and encodes these into an acoustic signal time-frequency feature coding graph, providing high-information-density input for subsequent signal quality assessment.

[0134] Specifically, step S33 performs feature enhancement on the acoustic signal time-frequency feature coding map to obtain an acoustic signal time-frequency feature enhanced coding map. It should be understood that, considering that the acoustic signal time-frequency feature coding map may contain certain redundant information (such as the random pattern of background noise), it is difficult to directly use it for accurate signal quality assessment. Therefore, the present application further performs feature enhancement on the acoustic signal time-frequency feature coding map to jointly model the spatial distribution characteristics and semantic association information of the acoustic signal time-frequency feature coding map. Based on cross-channel feature dependencies, the expressiveness of key features is enhanced while suppressing the interference of non-key information, thereby obtaining a more discriminative acoustic signal time-frequency feature enhanced coding map.

[0135] Figure 7 This is a flowchart of step S33 in the Bluetooth fingerprint library update method based on voiceprint quality estimation according to embodiment 2 of the present application. Figure 7 As shown, the step S33 includes: S331, extracting the channel feature vector of the (i, j)th pixel position from the acoustic signal time-frequency feature coding map as the acoustic signal time-frequency channel feature vector to be enhanced; S332, based on the acoustic signal time-frequency feature coding map, performing information compensation coding based on spatial-semantic dual-domain context guidance on the acoustic signal time-frequency channel feature vector to be enhanced to obtain an acoustic signal time-frequency channel enhancement component implicit coding vector; S333, fusing the acoustic signal time-frequency channel enhancement component implicit coding vector and the acoustic signal time-frequency channel feature vector to be enhanced to obtain an enhanced acoustic signal time-frequency channel feature vector, wherein the enhanced acoustic signal time-frequency channel feature vector is the channel feature vector of the pixel position (i, j) of the acoustic signal time-frequency feature enhancement coding map.

[0136] In a specific example of the present application, step S331 can be expressed as follows:

[0137]

[0138]

[0139] in, 、 and Respectively represent the height, width and number of channels of the time-frequency feature coding diagram of the acoustic signal, Represents the time-frequency feature coding diagram of the acoustic signal, Represents the channel feature vector of the (i, j)th pixel position in the time-frequency feature coding map of the acoustic signal, Represents the time-frequency channel feature vector of the acoustic signal to be enhanced.

[0140] It should be understood that each pixel position of the acoustic signal time-frequency feature coding map corresponds to a multidimensional channel feature vector, which carries the local features such as the signal energy distribution and frequency domain pattern at that time-space point. However, due to sensor noise, multipath effects or environmental interference, local features may have missing or distorted information, making it difficult to accurately express the complex time-frequency characteristics of the acoustic signal. Therefore, the present application extracts the channel feature vector of the (i, j)th pixel position in the acoustic signal time-frequency feature coding map as the channel feature vector to be enhanced in the acoustic signal time-frequency, so as to focus on the refined processing of local features, thereby enhancing the features of the entire acoustic signal time-frequency feature coding map from point to surface.

[0141] In a specific example of the present application, the step S332 includes: first, performing n random scans on the acoustic signal time-frequency feature coding map to obtain n channel feature vectors as a sparse set of acoustic signal time-frequency reference feature vectors. ,in, 、 、 and They represent the first, second, and third eigenvectors in the sparse set of acoustic signal time-frequency reference feature vectors. and A sound signal time-frequency reference feature vector. It should be understood that the fixed receptive field of the traditional convolutional network limits its ability to capture global contextual information, especially when facing a complex and changeable sound environment, and the global calculation of feature associations at all positions will bring high computational complexity. Therefore, the present application generates a sparse set of sound signal time-frequency reference feature vectors by randomly scanning the sound signal time-frequency feature coding map, which can introduce spatial diversity and avoid overfitting local noise. At the same time, it constructs a sound signal time-frequency feature context information library at a lower cost, providing rich context information guidance for subsequent feature enhancement processing.

[0142] Next, the Poincare distance between the acoustic signal time-frequency channel feature vector to be enhanced and each acoustic signal time-frequency reference feature vector in the sparse set of acoustic signal time-frequency reference feature vectors is calculated to obtain a spatial modulation matrix of the acoustic signal time-frequency channel feature vector to be enhanced composed of multiple Poincare distances, which can be expressed as follows:

[0143]

[0144]

[0145] in, represents the square of the norm of the vector, represents the inverse hyperbolic cosine function, express and The Poincare distance between express and The Poincare distance between Represents the spatial modulation matrix of the time-frequency channel features to be enhanced of the acoustic signal.

[0146] It should be understood that the traditional Euclidean distance is difficult to effectively represent hierarchical relationships in high-dimensional feature space, while the Poincare distance is applicable to hyperbolic space and can better model tree-like or hierarchical structures (such as the multipath reflection path of the acoustic signal). This application constructs a spatial modulation matrix of the acoustic signal time-frequency channel feature to be enhanced by quantifying the Poincare distance between the acoustic signal time-frequency channel feature vector to be enhanced and each acoustic signal time-frequency reference feature vector. This can more finely characterize the spatial correlation and hierarchical structure differences between the acoustic signal time-frequency channel feature to be enhanced and each reference feature. For example, the distance between the direct sound wave feature and the multipath reflection feature is large, while the distance between different reflection paths is small, thereby providing accurate spatial correlation information guidance for subsequent feature enhancement processing.

[0147] Then, the implicit semantic association between the acoustic signal time-frequency channel feature vector to be enhanced and each acoustic signal time-frequency reference feature vector in the sparse set of acoustic signal time-frequency reference feature vectors is calculated to obtain a set of acoustic signal time-frequency channel feature semantic association coding matrices to be enhanced, which can be expressed as follows:

[0148]

[0149] in, represents the weight matrix, represents the transpose of a vector, represents vector multiplication, represents the normalized exponential function, express and The semantic correlation coding matrix of the time-frequency channel features to be enhanced of the acoustic signal.

[0150] Here, since the surface similarity of the time-frequency features of the acoustic signal (such as energy distribution) is difficult to reflect the deep semantic association (such as signal category, interference type), in order to establish the implicit semantic association between the time-frequency channel feature vector of the acoustic signal to be enhanced and the time-frequency reference feature vectors of each acoustic signal from a higher level, this application introduces a cross-attention mechanism. By calculating the attention weights between the time-frequency channel feature vector of the acoustic signal to be enhanced and the time-frequency reference feature vectors of each acoustic signal in the sparse set, the degree of association between the two at the semantic level is revealed, thereby providing rich contextual semantic information guidance for subsequent feature enhancement.

[0151] Secondly, the information compensation coding vector between each acoustic signal time-frequency reference feature vector in the sparse set of acoustic signal time-frequency reference feature vectors and the acoustic signal time-frequency channel feature vector to be enhanced is calculated to obtain a set of acoustic signal time-frequency channel information compensation coding vectors to be enhanced, which is expressed as: ,in, express Relative to Here, by calculating the complementary information between each acoustic signal time-frequency reference feature vector and the acoustic signal time-frequency channel feature vector to be enhanced, a set of acoustic signal time-frequency channel information compensation coding vectors to be enhanced is obtained, which helps the subsequent feature enhancement processing to focus on the feature components that are missing or weakened at the position to be enhanced. For example, when the frequency band energy is missing due to occlusion, complementary information can be extracted from the reference features to ensure the targeted feature enhancement process and avoid the information redundancy and mismatch problems caused by the direct use of reference features.

[0152] Furthermore, based on the set of the acoustic signal time-frequency channel feature spatial modulation matrix and the acoustic signal time-frequency channel feature semantic association coding matrix, the set of the acoustic signal time-frequency channel information compensation coding vectors to be enhanced is subjected to explicit modeling modulation aggregation coding to obtain the acoustic signal time-frequency channel enhancement component implicit coding vector. In a preferred example of the present application, first, based on the set of the acoustic signal time-frequency channel feature spatial modulation matrix and the acoustic signal time-frequency channel feature semantic association coding matrix, each acoustic signal time-frequency channel information compensation coding vector in the set of the acoustic signal time-frequency channel information compensation coding vectors to be enhanced is subjected to dual-domain nested constraint optimization to obtain a set of optimized acoustic signal time-frequency channel information compensation coding vectors to be enhanced, which is expressed as follows:

[0153]

[0154]

[0155] in, express The corresponding semantically nested optimized acoustic signal time-frequency channel information compensation coding vector to be enhanced, express The corresponding optimized acoustic signal time-frequency channel information to be enhanced is compensated by the coding vector.

[0156] It should be understood that the acoustic signal time-frequency channel information compensation coding vector to be enhanced contains compensation information of each reference feature relative to the channel feature to be enhanced, but directly using this information may lead to information redundancy or inconsistency. Therefore, the present application combines the set of the acoustic signal time-frequency channel feature spatial modulation matrix to be enhanced and the acoustic signal time-frequency channel feature semantic association coding matrix to be enhanced (i.e., utilizing the spatial correlation and semantic correlation between each reference feature and the channel feature to be enhanced) to perform dual-domain nested constraint optimization on the corresponding acoustic signal time-frequency channel information compensation coding vector to be enhanced. Here, the semantic association coding matrix of the acoustic signal time-frequency channel features to be enhanced can be regarded as a canonical field, and the spatial modulation matrix of the acoustic signal time-frequency channel features to be enhanced can be regarded as a covariant field. Therefore, the present application utilizes the coupling effect of the canonical field and the covariant field to force the acoustic signal time-frequency channel information compensation coding vector to maintain a smooth transition of local neighborhood features under the covariance guidance of the acoustic signal time-frequency channel feature spatial modulation matrix to be enhanced, and at the same time strengthens the consistency of semantic association under the canonical constraint of the acoustic signal time-frequency channel feature semantic association coding matrix, and finally realizes the optimized expression of the acoustic signal time-frequency channel information compensation coding vector to be enhanced, thereby enhancing the robustness and accuracy of the acoustic signal time-frequency features.

[0157] Then, each acoustic signal time-frequency channel feature semantic association coding matrix in the set of the acoustic signal time-frequency channel feature semantic association coding matrices to be enhanced is used as a primary mask modulation unit, and the acoustic signal time-frequency channel feature spatial modulation matrix to be enhanced is used as a secondary mask modulation unit. The set of the optimized acoustic signal time-frequency channel information compensation coding vectors to be enhanced is subjected to explicit modeling modulation aggregation coding to obtain an implicit coding vector of the acoustic signal time-frequency channel enhancement component to be enhanced, which is expressed as follows:

[0158]

[0159] in, Represents the implicit coding vector of the enhanced component of the time-frequency channel feature to be enhanced of the acoustic signal.

[0160] Here, in order to balance the contribution of spatial and semantic information, the present application explicitly controls the information fusion process through the hierarchical modulation of the first-level semantic mask and the second-level spatial mask. Specifically, the semantic association coding matrix of the acoustic signal time-frequency channel feature to be enhanced is used as a first-level mask modulation unit to preliminarily modulate the acoustic signal time-frequency channel information compensation coding vector to be enhanced, aiming to enhance the contribution of the high-correlation compensation vector through the semantic mask, and the spatial modulation matrix of the acoustic signal time-frequency channel feature to be enhanced is used as a second-level mask modulation unit, which is further weighted by the spatial mask to highlight the compensation effect of spatially adjacent features, thereby effectively strengthening the aggregation of compensation information that is beneficial to the acoustic signal time-frequency channel feature vector to be enhanced, and suppressing the interference of irrelevant or redundant information.

[0161] In a specific example of the present application, step S333 can be expressed as follows:

[0162]

[0163] in, and Represents different fusion weight parameters, Represents the time-frequency channel feature vector of the enhanced acoustic signal.

[0164] Here, by fusing the implicit coding vector of the acoustic signal time-frequency channel enhancement component to be enhanced and the acoustic signal time-frequency channel feature vector to be enhanced, the feature expression of the channel to be enhanced can be targetedly enhanced on the basis of retaining the effective information of the original features, so that the generated enhanced acoustic signal time-frequency channel feature vector can more accurately reflect the physical characteristics and semantic information of the target sound field environment, thereby improving the accuracy and robustness of acoustic signal processing.

[0165] Specifically, in step S34, the Bluetooth signal RSSI values ​​at the various positions are normalized to obtain a normalized Bluetooth signal RSSI feature vector. In a specific example of the present application, the Bluetooth signal RSSI values ​​at the various positions are first collected. These RSSI values ​​are measurements of the signal strength between the Bluetooth receiver and the device. Then, the RSSI values ​​are normalized using a minimum-maximum normalization or standardization method to convert the RSSI values ​​at different positions or times into feature vectors of a unified scale. Specifically, the normalized Bluetooth signal RSSI feature vector can eliminate the amplitude difference in signal strength, so that the Bluetooth signal characteristics at different positions or environments can be directly compared. In this way, by normalizing the RSSI signal, the stability and comparability of the signal are guaranteed, which is helpful for subsequent feature fusion and quality assessment.

[0166] Specifically, step S35 fuses the acoustic signal time-frequency feature enhancement coding map with the RSSI value information of the Bluetooth signal to form a joint time-frequency feature coding map. In a specific example of the present application, first, the acoustic signal time-frequency feature enhancement coding map represents the multidimensional features of the acoustic signal in the time-frequency domain, including the spectral characteristics and time variation information of the acoustic signal; while the RSSI value information of the Bluetooth signal represents the strength information of the Bluetooth signal. In order to achieve a joint feature representation of the acoustic signal and the Bluetooth signal, the step fuses these two different types of signal features. Specifically, the acoustic signal time-frequency feature enhancement coding map and the Bluetooth signal RSSI feature vector can be spliced ​​in the feature dimension by feature splicing to form a new joint feature vector. This joint feature vector contains the time-frequency information and strength information of both the acoustic signal and the Bluetooth signal, and can provide comprehensive information for subsequent signal quality assessment. In some applications, if the weights of the two signals are different, they can also be fused by weighted averaging to make the influence of one signal in the joint feature more significant. In this way, by fusing the features of the acoustic signal and the Bluetooth signal to form a joint time-frequency feature coding map, the system can simultaneously utilize the information of the two signals, thereby improving the accuracy and reliability of signal quality assessment.

[0167] Specifically, the step S36 determines whether the joint signal meets the signal quality standard based on the joint time-frequency feature coding map. In a specific example of the present application, first, the joint time-frequency feature coding map represents the fused time-frequency features of the acoustic signal and the RSSI features of the Bluetooth signal, and contains the multi-dimensional features of the acoustic signal and the Bluetooth signal in the time-frequency domain and the strength information of the Bluetooth signal. In order to determine whether the joint signal meets the signal quality standard, the step introduces an automated quality assessment method based on a deep learning model. Specifically, the joint time-frequency feature coding map is input into a pre-trained deep learning classifier for quality assessment.

[0168] In a specific example of this application, a deep learning model is first trained on a large number of signal samples of known quality to learn the characteristic patterns of signals of different quality levels. These training samples include the joint time-frequency features of acoustic and Bluetooth signals, along with their corresponding quality labels (e.g., high or low quality). During the training process, the model automatically learns the complex nonlinear relationship between signal characteristics and quality criteria.

[0169] When the joint time-frequency feature coding map is input into the deep learning classifier, the classifier will classify it according to the learned feature patterns to determine whether the joint signal is a high-quality signal or a low-quality signal. Specifically, the classifier outputs a probability value, which represents the probability that the signal meets the predetermined quality standard. When the output value exceeds the set threshold, the joint signal is judged to be "high quality", otherwise it is judged to be "low quality". This automated process based on supervised learning helps to avoid the subjectivity and environmental dependence of traditional methods that rely on fixed thresholds. By adopting a deep learning model, the system can automatically extract complex patterns from the time-frequency features of the joint signal, evaluate the signal quality in real time, and has strong adaptability, which can accurately evaluate the signal in different environments. In addition, the deep learning model can also continuously optimize the quality assessment criteria according to changes in the environment, further improving the accuracy and reliability of the assessment.

[0170] In summary, the Bluetooth fingerprint library update method based on voiceprint quality estimation described in this application is explained. It deploys a base station group with both Bluetooth broadcasting and acoustic signal transmission functions in the target area, uses the Bluetooth path loss model for preliminary positioning, and estimates the position based on the received signal strength of each location node to establish a Bluetooth fingerprint library; at the same time, it synchronously receives the acoustic signals of each location, and performs quality evaluation on the received acoustic signals, screens out effective acoustic signals for TDOA positioning calculation, determines the acoustic signal positioning coordinates of each location, and judges the acoustic signal positioning accuracy based on the position quality evaluation method, and then reversely infers the Bluetooth strength of the current location based on the accurate acoustic signal location information to update the Bluetooth fingerprint database in real time. This method uses the stability of acoustic signal propagation to compensate for the directional sensitivity of the Bluetooth signal, can effectively cope with the dynamic changes of the indoor environment, and realize the adaptive maintenance of fingerprint data.

Claims

1. A Bluetooth fingerprint library update method based on voiceprint quality estimation, characterized in that: include: Construct a Bluetooth fingerprint library, where each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position}; receiving acoustic signals at the respective positions; determining whether the acoustic signal at each position meets a signal quality standard, and if the acoustic signal at each position meets the signal quality standard, performing position estimation based on the acoustic signal at each position to obtain acoustic signal location coordinates at each position; The deviation of the hyperbola intersection of the acoustic signal positioning coordinates is calculated based on a parameterized method to generate an acoustic signal position calculation quality assessment value; Based on the comparison between the acoustic signal position calculation quality evaluation value of each position and the acoustic positioning quality threshold, determine whether the RSSI value of each position needs to be updated, including: after determining that the acoustic signal position calculation quality evaluation value is less than the acoustic positioning quality threshold, reversely infer the Bluetooth strength of the current position based on the accurate acoustic signal position information, that is, obtain the distance between the current position and the base station based on the acoustic signal positioning coordinates and the base station position, and then calculate the RSSI value of the current position based on the path loss model.

2. The method for updating the Bluetooth fingerprint database based on voiceprint quality estimation according to claim 1, characterized in that: Build a Bluetooth fingerprint library, including: Construct the fingerprint matrix of each position; Calculating the distance between each reference position RSSI vector in the fingerprint matrix and the observed RSSI vector to obtain a distance vector; Selecting k reference positions corresponding to the smallest distances from the distance vector; Weights are calculated for the k smallest distances, and based on the weights, weighted averaging is performed on the k reference positions to obtain the two-dimensional coordinates of the observation position.

3. The method for updating the Bluetooth fingerprint database based on voiceprint quality estimation according to claim 1, characterized in that: Determining whether the acoustic signals at each position meet the signal quality standard includes: Channel impulse response analysis, signal peak distribution analysis, signal reverberation time analysis, and signal frequency distortion analysis are performed on the acoustic signals at the various positions to determine whether the acoustic signals at the various positions meet signal quality standards.

4. The method for updating the Bluetooth fingerprint database based on voiceprint quality estimation according to claim 1, characterized in that: After the acoustic signals at the respective positions meet the signal quality standard, position estimation is performed based on the acoustic signals at the respective positions to obtain the acoustic signal positioning coordinates at the respective positions, including: The position estimation and optimization of the time difference of arrival of the acoustic signals are performed based on a parameterized method to obtain the acoustic signal positioning coordinates of the various positions.

5. The method for updating the Bluetooth fingerprint database based on voiceprint quality estimation according to claim 1, characterized in that: The hyperbola intersection deviation of the acoustic signal positioning coordinates is calculated based on a parameterized method to generate an acoustic signal position calculation quality evaluation value, including: calculating the acoustic signal position calculation quality evaluation value of each position using the following formula, wherein the formula is: in, is the number of acoustic signal base stations, The hyperbola intersection parameters optimized for the parametric method, 、 is the hyperbolic index, 、 are coordinate components, is the square of the two-norm, A quality assessment value is calculated for the acoustic signal position.

6. A Bluetooth fingerprint library update method based on voiceprint quality estimation, characterized in that: include: Construct a Bluetooth fingerprint library, where each Bluetooth fingerprint in the Bluetooth fingerprint library is {RSSI value of each position, two-dimensional coordinates of each position}; receiving acoustic signals at the respective positions; Determining whether a signal quality standard is met based on the RSSI values ​​of the respective positions and the acoustic signals at the respective positions, and in response to the signal quality standard being met, performing position estimation based on the acoustic signals at the respective positions to obtain acoustic signal positioning coordinates of the respective positions; The deviation of the hyperbola intersection of the acoustic signal positioning coordinates is calculated based on a parameterized method to generate an acoustic signal position calculation quality assessment value; Based on the comparison between the acoustic signal position calculation quality evaluation value of each position and the acoustic positioning quality threshold, determine whether the RSSI value of each position needs to be updated, including: after determining that the acoustic signal position calculation quality evaluation value is less than the acoustic positioning quality threshold, reversely infer the Bluetooth strength of the current position based on the accurate acoustic signal position information, that is, obtain the distance between the current position and the base station based on the acoustic signal positioning coordinates and the base station position, and then calculate the RSSI value of the current position based on the path loss model.

7. The method for updating a Bluetooth fingerprint database based on voiceprint quality estimation according to claim 6, characterized in that: Determining whether a signal quality standard is met based on the RSSI values ​​at the respective positions and the acoustic signals at the respective positions includes: Converting the acoustic signal into a time-frequency graph through wavelet analysis to obtain an acoustic signal time-frequency graph; Extracting the acoustic signal time-frequency features based on dilated convolution coding on the acoustic signal time-frequency graph to obtain an acoustic signal time-frequency feature coding graph; Performing feature enhancement on the acoustic signal time-frequency feature coding map to obtain an acoustic signal time-frequency feature enhanced coding map; Normalizing the RSSI values ​​of the Bluetooth signals at the respective positions to obtain normalized RSSI feature vectors of the Bluetooth signals; The acoustic signal time-frequency feature enhancement coding map is combined with the Bluetooth signal RSSI feature vector to form a joint time-frequency feature coding map; Based on the joint time-frequency feature coding map, it is determined whether the joint signal meets the signal quality standard.

8. The method for updating a Bluetooth fingerprint database based on voiceprint quality estimation according to claim 7, characterized in that: The joint time-frequency feature coding map of the acoustic signal and the Bluetooth signal is enhanced to obtain a joint time-frequency feature enhanced coding map of the acoustic signal and the Bluetooth signal, including: Extracting the channel feature vector at the (i, j)th pixel position from the joint time-frequency feature coding map as the joint time-frequency channel feature vector to be enhanced for the acoustic signal and the Bluetooth signal; Based on the joint time-frequency feature coding map, performing information compensation coding based on spatial-semantic dual-domain context guidance on the joint time-frequency channel feature vector of the acoustic signal and Bluetooth signal to be enhanced, so as to obtain an implicit coding vector of the enhanced component of the joint time-frequency channel to be enhanced of the acoustic signal and Bluetooth signal; The implicit coding vector of the enhanced component of the joint time-frequency channel to be enhanced of the acoustic signal and Bluetooth signal and the joint time-frequency channel feature vector to be enhanced of the acoustic signal and Bluetooth signal are fused to obtain the enhanced joint time-frequency channel feature vector of the acoustic signal and Bluetooth signal, wherein the enhanced joint time-frequency channel feature vector of the acoustic signal and Bluetooth signal is the channel feature vector of the pixel position (i, j) of the joint time-frequency feature enhancement coding image.

9. The method for updating a Bluetooth fingerprint database based on voiceprint quality estimation according to claim 8, characterized in that: Based on the joint time-frequency feature coding map of the acoustic signal and the Bluetooth signal, information compensation coding based on spatial-semantic dual-domain context guidance is performed on the feature vector of the channel to be enhanced of the joint time-frequency feature to obtain an implicit coding vector of an enhanced component of the channel to be enhanced of the joint time-frequency feature, including: Randomly scan the joint time-frequency feature coding map to extract feature vectors of different channels; Calculate the Poincare distance between eigenvectors to measure the similarity of each vector and generate the eigenspace scheduling matrix; Based on the semantic association of feature vectors, information compensation coding is used to optimize features; The optimized feature vector is subjected to explicit modeling modulation and aggregate coding to obtain the final enhanced channel feature vector.

10. The method for updating a Bluetooth fingerprint database based on voiceprint quality estimation according to claim 9, characterized in that: Based on the set of the spatial modulation matrix of the joint time-frequency channel feature to be enhanced of the acoustic signal and the Bluetooth signal and the set of the joint time-frequency feature semantic association coding matrix, explicit modeling modulation aggregation coding is performed on the set of the joint time-frequency channel information compensation coding vectors to be enhanced to obtain an implicit coding vector of the joint time-frequency channel enhancement component, including: Based on the above-mentioned spatial modulation matrix and semantic association coding matrix, a dual-domain nested constraint optimization is applied to each information compensation coding vector to generate an optimized information compensation coding vector set; The joint semantic association coding matrix is ​​used as the first-level mask modulation unit and the spatial modulation matrix is ​​used as the second-level mask modulation unit. The optimized coding vector set is explicitly modeled and aggregated, and finally the implicit coding vector of the enhanced component of the joint time-frequency channel to be enhanced is obtained.

11. The method for updating a Bluetooth fingerprint database based on voiceprint quality estimation according to claim 10, characterized in that: Determining whether the combined signal meets a signal quality standard based on a combined time-frequency feature enhancement coding map of the acoustic signal and the Bluetooth signal includes: The joint time-frequency feature enhancement coding map of the acoustic signal and the Bluetooth signal is input into a signal quality module based on a classifier, and the classifier is used to analyze and determine whether the joint signal meets a preset signal quality standard.

Citation Information

Patent Citations

  • Selectively performing positioning procedure at access terminal based on behavior model

    CN103797332A

  • Indoor parking lot positioning method based on Bluetooth AOD and inertial navigation

    CN116634561A