A data-driven audio ranging method and system
By introducing short-time Fourier transform and convolutional neural networks into the audio ranging method, the problem of insufficient ranging accuracy and reliability in the existing technology is solved, and higher accuracy and robustness are achieved, and it is suitable for audio ranging applications in complex environments.
Patent Information
- Application Number
- CN202410953552.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-16
AI Technical Summary
The existing audio ranging method based on cross-correlation functions faces problems such as environmental noise interference, signal attenuation and frequency offset, hardware differences, multipath effect, and dynamic changes in signal and noise in practical applications, resulting in insufficient ranging accuracy and reliability.
The data-driven audio ranging method is used to convert the time domain signal into a time-frequency image through short-time Fourier transform (STFT), and feature extraction and arrival-time estimation are performed using a convolutional neural network (CNN).
It significantly improves the accuracy and robustness of audio ranging, reduces the impact of ambient noise and multipath effects, enhances the anti-interference ability of the system, and maintains efficient performance under different environments and equipment.
Smart Images

Figure CN118865960B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the field of indoor positioning technology, and particularly relates to a data-driven audio ranging method and system. Background Art
[0002] In the field of indoor positioning, most of the detection methods based on sound signals use cross-correlation functions to determine the time of arrival of signals. Although this method has good effects under ideal conditions, there are multiple technical problems in practical applications, which affect its high precision and stability. The following are the main technical problems existing in the prior art:
[0003] 1. Environmental noise interference:
[0004] In the actual application environment, the sound signal will be interfered by environmental noise during the propagation process, and these noises will be superimposed on the target signal, resulting in a decrease in signal quality. The presence of noise makes the similarity curve of the cross-correlation function unclear, thus affecting the accuracy of the time of arrival.
[0005] 2. Signal attenuation and frequency offset:
[0006] When the sound signal propagates in the air, due to the influence of factors such as distance, air damping, and reflection, the signal intensity at different frequency points will attenuate to varying degrees. At the same time, the signal also shows a frequency offset. These changes make the actually collected signal significantly different from the ideal signal, resulting in difficulty in accurately matching and positioning based on the cross-correlation function.
[0007] 3. Influence of hardware differences:
[0008] Different speaker and microphone devices will introduce their respective nonlinear and response characteristics during the signal transmission and reception processes. These hardware characteristics cause signal distortion during the transmission and reception processes, further increasing the difficulty and error of signal detection.
[0009] 4. Multipath effect:
[0010] In an indoor environment, the sound signal will be reflected on objects such as walls, floors, and furniture to form multipath signals. These multipath signals are superimposed in the microphone, resulting in changes in the time delay and form of the signal. It is difficult for the cross-correlation function to distinguish between the direct path signal and the reflected path signal, thus affecting the positioning accuracy.
[0011] 5. Dynamic changes of signal and noise:
[0012] The noise and signal characteristics in the environment vary with time, especially in dynamic environments (such as crowd activities, equipment operation, etc.). These dynamic changes make the difference between the pre-designed ideal signal and the actual collected signal more complex, further increasing the challenges to the stability and reliability of the cross-correlation function-based method.
[0013] Existing sound signal detection methods based on the cross-correlation function face multiple technical problems in practical applications, such as environmental noise interference, signal attenuation and frequency shift, hardware difference effects, multipath effects, and the dynamic changes of signals and noise. These problems make it very challenging to detect sound signals with high precision and stability, and more advanced algorithms and technologies are urgently needed to improve the detection accuracy and reliability. Summary of the Invention
[0014] In view of the problems existing in the prior art, the present invention provides a data-driven audio ranging method.
[0015] The present invention is implemented as follows. A data-driven audio ranging method includes:
[0016] S1: Preprocess the original signal, and convert the filtered time-domain signal into a time-frequency image through the short-time Fourier transform (STFT).
[0017] S2: Use a convolutional neural network model to obtain the arrival time estimate.
[0018] Further, the specific steps of S1 include:
[0019] Convert the time-domain signal into a time-frequency image through the short-time Fourier transform (STFT), and this image reflects the change of the energy of each frequency point of the signal over time;
[0020] Assume that the length of the sliding window is WL and the moving step size of the window is SL, then the time resolution is (SL / Fs) s, and the frequency resolution is (Fs / WL) Hz; the detailed steps of STFT are as follows:
[0021] First, start sliding the window from the starting point of the received sound signal data. At this time, the window function is centered at t = τ0, and the signal is processed with the window function:
[0022] y(t) = x(t) · w(t - τ0)
[0023] Then, perform Fourier transform on the data to obtain the PSD matrix of the sound data of the first window,
[0024] Represents the vector of the received signal at the time delay (0, τ0].
[0025]
[0026] where \(x(t)\) represents the received signal for the voice data segment; \(w\) is the Hamming window function; \(f\) m depends on the sampling rate of the smartphone (Frequency of Sampling, \(F_s\)), ranging from 0 Hz to \(F_s / 2\) Hz. \(f\) m and \(\tau_0\) are defined as follows:
[0027]
[0028] \(\tau_0=(WL / 2) / F_s\)
[0029] Finally, the calculation method of the PSD matrix representing the voice data segment of the \(n\)th window is as follows:
[0030]
[0031] In the formula, is the vector of the received signal at the time delay (\(\tau\) n-1 , \(\tau\) n ), where \(x(t\) n ) and \(\tau\) n are defined as follows:
[0032] \(x(t\) n ) = \(R[(n - 1)\times SL:WL+(n - 1)\times SL]\)
[0033] \(\tau\) n = \([WL / 2+(n - 1)\times SL] / F_s\)
[0034] When performing the short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, and the size of the window determines the frequency resolution of the input image; preferably, the step size of the window sliding is 1 ms, equivalent to a resolution of 0.34 m, and for the convenience of performing the fast Fourier transform, the size of the window is selected as 512.
[0035] Furthermore, the specific steps of \(S2\) include:
[0036] The first convolutional layer of the audio detection model performs a convolution operation on the time-frequency diagram with a kernel size of 3*3, stride and padding of 1. The output result does not change the size of the image, but the number of channels changes from 1 to 8, expecting to learn different features in the time-frequency diagram. And in this process, a rectified linear unit function is used for non-linear activation to improve the non-linear fitting ability of the model. Since the features of the Chirp signal are relatively obvious in the time-frequency diagram and the computing power in the practical process needs to be considered, the number of channels is reduced from 8 to 1 through the second convolution operation to control the amount of data entering the fully connected layer. Of course, after this process, ReLU is still used for non-linear activation. Subsequently, the two-dimensional image is converted into a one-dimensional vector, and the coefficients are regularized through dropout. During training, such regularization operations can accelerate the training speed of the model and prevent the model from overfitting to certain features. Of course, in actual regression, it will not be used. Finally, there are two consecutive fully connected layers. By fitting the learned features, the arrival time of the Chirp in the sound signal is finally obtained.
[0037] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:
[0038] First, the CC-DD algorithm proposed by the present invention has very stable performance, and the maximum CDF-95% is only 1.551 meters. Looking at the error results at different speeds as a whole, from the maximum CDF-68%, the CC-DD method optimizes 33% compared with CCFM, and from the maximum CDF-95%, it optimizes 80%. The dynamic ranging performance has been greatly improved.
[0039] The data-driven audio ranging method proposed by the present invention solves a number of key problems in existing audio ranging technologies. Traditional audio ranging methods rely on simple signal processing and models, and cannot effectively cope with signal noise and multipath interference in complex environments, resulting in insufficient ranging accuracy and reliability. By introducing the short-time Fourier transform (STFT) and convolutional neural network (CNN), the present invention can extract more signal features from time-frequency images, improve the ability to process signals in complex environments, and significantly improve the ranging accuracy and robustness.
[0040] First, the present invention converts the time-domain signal into a time-frequency image through the short-time Fourier transform (STFT), so that the features of the signal in time and frequency are fully displayed. Compared with traditional methods, STFT provides higher time and frequency resolutions, can capture the changing features of the signal more accurately, and reduces the influence of environmental noise and multipath effects on the ranging result. The preferred sliding window and stride settings further improve the quality of the time-frequency image, providing a reliable basis for subsequent feature extraction and analysis.
[0041] Secondly, the present invention uses a Convolutional Neural Network (CNN) for feature extraction and time-of-arrival estimation. The CNN can automatically learn and extract important features in the time-frequency image. Through multiple layers of convolution and non-linear activation, it captures the obvious features of the Chirp signal in the time-frequency diagram. Compared with traditional manual feature extraction methods, the CNN has stronger generalization ability and adaptability, and can maintain high-efficiency feature extraction performance in a variety of complex environments. In addition, the design of using dropout and fully connected layers enhances the regularization effect of the model, prevents overfitting, and improves the training speed and ranging accuracy of the model.
[0042] Finally, the present invention has high robustness and flexibility in system design. Through reasonable configuration of convolutional layers and fully connected layers, the system can process audio signals in various environments and has strong anti-interference ability. Especially in application scenarios such as indoor positioning and smart home, the method of the present invention can accurately calculate the distance between the device and the user and provide high-precision positioning services. Overall, while solving the problems of the existing technology, the present invention has achieved significant technological progress and provided new ideas and solutions for the application and development of audio ranging technology.
[0043] Second, the technical solution of the present invention solves a series of technical problems of the existing technology and has achieved significant technological progress. The following is a detailed elaboration of these technical problems and technological progress:
[0044] Technical problems solved by the existing technology:
[0045] 1. Accuracy problem of traditional ranging methods: Traditional audio ranging methods are interfered by various factors, such as environmental noise, multipath effect, etc., resulting in low ranging accuracy. The present invention improves the accuracy and stability of ranging through data-driven methods and deep learning models.
[0046] 2. Complexity of signal processing: Traditional audio signal processing methods involve complex mathematical operations and signal processing algorithms. The present invention simplifies the signal processing process and improves the processing efficiency by using the Short-Time Fourier Transform (STFT) and Convolutional Neural Network (CNN).
[0047] 3. Poor environmental adaptability: Traditional audio ranging methods perform quite differently in different environments. The present invention enables the system to better adapt to different environmental conditions and improves the robustness of the ranging method through the powerful learning ability of the deep learning model.
[0048] Significant technological progress:
[0049] 1. Improved ranging accuracy: By combining STFT and CNN, the present invention can more accurately capture the features in the audio signal, thereby significantly improving the ranging accuracy. This is crucial for applications that require precise distance measurement (such as autonomous driving, robot navigation, etc.).
[0050] 2. Enhanced environmental adaptability: Since the deep learning model has powerful feature learning and abstraction capabilities, the present invention can maintain high ranging performance under various environmental conditions. This makes the method more widely applicable in practical applications.
[0051] 3. Simplify the signal processing process: By using STFT to convert the time domain signal into a time-frequency image and combining it with CNN for feature extraction and arrival time estimation, the present invention simplifies the traditional complex signal processing process and improves processing efficiency.
[0052] 4. Promote the application of deep learning in the field of audio processing: This invention demonstrates the effectiveness of deep learning in audio processing tasks such as audio ranging, and is expected to promote the application and development of deep learning in a wider range of audio processing fields.
[0053] In summary, the technical solution of the present invention solves multiple technical problems in traditional audio ranging methods by combining STFT and deep learning models, and has achieved significant technical progress, providing new ideas and methods for the development of audio ranging technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a schematic diagram of a data-driven audio ranging algorithm provided by an embodiment of the present invention;
[0055] Figure 2 This is a corridor scene of a static ranging comparison experiment provided by an embodiment of the present invention: (a) no one; (b) someone;
[0056] Figure 3 is the slow ranging result in the dynamic ranging experiment provided by the embodiment of the present invention;
[0057] Figure 4 is the ranging result at normal speed in the dynamic ranging experiment provided by the embodiment of the present invention;
[0058] Figure 5 It is a fast ranging result in the dynamic ranging experiment provided by the embodiment of the present invention;
[0059] Figure 6 is the error cumulative distribution of slow dynamic ranging provided by an embodiment of the present invention;
[0060] Figure 7 is the error cumulative distribution of normal speed dynamic ranging provided by an embodiment of the present invention;
[0061] Figure 8 is the error cumulative distribution of fast dynamic ranging provided by the embodiments of the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] The following are 4 specific application embodiments for further illustrating the application of the data-driven audio ranging method of the present invention:
[0064] Embodiment 1: Audio ranging in smart home
[0065] In a smart home environment, audio ranging technology can be used to achieve more intelligent home control. For example, when a user issues a voice command in the living room, the smart home system needs to accurately determine the user's location in order to precisely control the corresponding device.
[0066] 1. The user issues a voice command, such as "Turn on the living room lights".
[0067] 2. The microphone array in the smart home system receives the voice signal.
[0068] 3. The system preprocesses the received original voice signal, including steps such as filtering and denoising.
[0069] 4. Use the short-time Fourier transform (STFT) to convert the preprocessed time-domain signal into a time-frequency image. In this step, the window size is set to 512, and the step size of window sliding is 1 ms to balance the time resolution and frequency resolution.
[0070] 5. Input the generated time-frequency image into a pre-trained convolutional neural network model for arrival time estimation. The model extracts features in the image through multiple layers of convolution and non-linear activation functions and outputs the arrival time estimation.
[0071] 6. According to the arrival time estimation, the smart home system determines the approximate location of the user.
[0072] 7. The system automatically turns on the living room light closest to the user according to the user's location to provide a comfortable lighting environment.
[0073] Through this embodiment, it can be seen that the application of the present invention in the smart home improves the intelligence level and user experience of the home system.
[0074] Embodiment 2: Obstacle distance detection in autonomous vehicles
[0075] In autonomous vehicles, accurately determining the distance to obstacles (such as pedestrians, other vehicles, etc.) is crucial for ensuring driving safety. Audio ranging technology can be used as an auxiliary means to improve the accuracy and reliability of obstacle detection.
[0076] 1. The audio sensor of the autonomous vehicle captures the sound signals of the obstacles ahead (such as vehicle engine sounds, pedestrian voices, etc.).
[0077] 2. The vehicle control system preprocesses the received sound signals, including steps such as amplification and filtering, to improve the signal quality.
[0078] 3. The short-time Fourier transform is used to convert the preprocessed time-domain signal into a time-frequency image to better extract signal features.
[0079] 4. The time-frequency image is input into a pre-trained convolutional neural network model for time-of-arrival estimation. This step can help the vehicle control system accurately determine the source distance of the sound signal.
[0080] 5. Based on the time-of-arrival estimation, the vehicle control system calculates the specific distance to the obstacle and makes a comprehensive judgment in combination with other sensor data (such as cameras, radars, etc.).
[0081] 6. According to the judgment result, the vehicle control system takes corresponding obstacle avoidance measures, such as decelerating and steering, to ensure driving safety.
[0082] Through this embodiment, the application potential of the present invention in the field of autonomous driving can be seen, providing a new technical means to improve driving safety and comfort.
[0083] Embodiment 3: Indoor positioning system
[0084] In indoor environments such as large shopping malls and museums, traditional GPS positioning is difficult to achieve accurate positioning due to signal problems. Therefore, a high-precision indoor positioning system is needed to help users find the target location.
[0085] System deployment:
[0086] 1) Signal preprocessing module:
[0087] Multiple microphones installed at different positions indoors are used to capture the sound signals emitted by users.
[0088] The original sound signals are preprocessed, and the filtered time-domain signals are converted into time-frequency images through the short-time Fourier transform (STFT). The sliding window length is set to 512, and the window moving step size is 1 ms to obtain high-resolution time-frequency images.
[0089] 2) Convolutional Neural Network Module:
[0090] Use the preprocessed time-frequency image as input and process it through a convolutional neural network model.
[0091] The first convolutional layer uses a 3x3 convolutional kernel, performs convolution operations with a stride and padding of 1, changes the number of channels from 1 to 8, and performs non-linear activation through ReLU.
[0092] The second convolutional layer reduces the 8 channels to 1 channel and performs non-linear activation through ReLU.
[0093] The conversion module converts the two-dimensional image into a one-dimensional vector and performs regularization through dropout.
[0094] The fully connected layer module contains two fully connected layers, fits the learned features, and finally estimates the arrival time of Chirp in the sound signal.
[0095] By deploying a data-driven audio ranging system, the indoor positioning system can accurately calculate the distances between the user and each microphone, and determine the specific location of the user through a multi-point positioning algorithm. This system shows high precision and high reliability in practical applications, significantly improving the user's positioning experience.
[0096] Example 4: Smart Home System
[0097] In a smart home system, voice control devices need to accurately identify the user's location in order to execute corresponding instructions (such as adjusting lights, temperature, etc.) according to the user's location.
[0098] 1) Signal Preprocessing Module:
[0099] Install multiple voice control devices in the home environment to capture the user's voice signal.
[0100] Preprocess the original voice signal, convert the filtered time-domain signal into a time-frequency image through short-time Fourier transform (STFT). Set the sliding window length to 512 and the moving step to 1 ms to obtain a high-precision time-frequency image.
[0101] 2) Convolutional Neural Network Module:
[0102] Use the preprocessed time-frequency image as input and process it through a convolutional neural network model.
[0103] The first convolutional layer uses a 3x3 convolutional kernel, performs convolution operations with a stride and padding of 1, changes the number of channels from 1 to 8, and performs non-linear activation through ReLU.
[0104] The second convolutional layer reduces the number of channels from 8 to 1 and performs non-linear activation through ReLU.
[0105] The conversion module converts the two-dimensional image into a one-dimensional vector and regularizes it through dropout.
[0106] The fully-connected layer module contains two fully-connected layers, fits the learned features, and finally estimates the arrival time of Chirp in the speech signal.
[0107] The data-driven audio ranging system can accurately calculate the distance between the user and each voice control device, achieving precise positioning in the smart home environment. The system can automatically adjust the status of home devices according to the user's location, improving the user's comfort and convenience. This system shows efficient and intelligent characteristics in practical applications, significantly enhancing the user experience of smart homes.
[0108] As Figure 1 shown, an embodiment of the present invention provides a data-driven audio ranging method, which includes:
[0109] S1: Preprocess the original signal and convert the filtered time-domain signal into a time-frequency image through short-time Fourier transform (STFT).
[0110] S2: Use a convolutional neural network model to obtain the arrival time estimate.
[0111] The specific steps of S1 include:
[0112] Convert the time-domain signal into a time-frequency image through short-time Fourier transform (STFT), and this image reflects the change of the energy of each frequency point as the signal changes over time;
[0113] Assume that the length of the sliding window is WL and the moving step size of the window is SL, then the time resolution is (SL / Fs) s, and the frequency resolution is (Fs / WL) Hz; the detailed steps of STFT are as follows:
[0114] First, start sliding the window from the starting point of the received sound signal data. At this time, the window function is centered at t = τ0, and the signal is processed with the window function:
[0115] y(t) = x(t) · w(t - τ0)
[0116] Then, perform Fourier transform on the data to obtain the PSD matrix of the sound data of the first window,
[0117] representing the vector of the received signal at the time delay (0, τ0].
[0118]
[0119] where \(x(t)\) represents the received signal of the voice data segment; \(w\) is the Hamming window function; \(f\) m depends on the sampling rate of the smartphone (Frequency of Sampling, \(F_s\)), ranging from 0 Hz to \(F_s / 2\) Hz. \(f\) m and \(\tau_0\) are defined as follows:
[0120]
[0121] \(\tau_0=(WL / 2) / F_s\)
[0122] Finally, the calculation method of the PSD matrix representing the voice data segment of the \(n\)th window is as follows:
[0123]
[0124] In the formula, is the vector of the received signal at the time delay (\(\tau\) n-1 , \(\tau\) n ), where \(x(t\) n ) and \(\tau\) n are defined as follows:
[0125] \(x(t\) n ) = \(R[(n - 1)\times SL:WL+(n - 1)\times SL]\)
[0126] \(\tau\) n = \([WL / 2+(n - 1)\times SL] / F_s\)
[0127] When performing the short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, and the size of the window determines the frequency resolution of the input image; preferably, the step size of the window sliding is 1 ms, which is equivalent to a resolution of 0.34 m, and for the convenience of performing the fast Fourier transform, the size of the window is selected as 512.
[0128] The specific content of S2 includes:
[0129] The first layer of convolution of the audio detection model convolves the time-frequency map with a kernel of size 3*3 with parameters of stride and padding of 1. The output result does not change the size of the image, but the number of channels changes from 1 to 8, expecting to learn different features in the time-frequency map; and in this process, the rectified linear unit function is used for non-linear activation to improve the non-linear fitting ability of the model; because the features of the Chirp signal are more obvious in the time-frequency map and the computing power in the practical process needs to be considered, the number of channels is reduced from 8 to 1 through the second convolution operation to control the amount of data entering the fully connected layer; of course, after this process, ReLU is still used for non-linear activation. Subsequently, the two-dimensional image is converted into a one-dimensional vector, and the coefficients are regularized through dropout; during training, such regularization operations can accelerate the training speed of the model and prevent the model from overfitting to certain features; of course, during actual regression, it will not be used; finally, two fully connected layers are followed, and by fitting the learned features, the arrival time of the Chirp in the sound signal is finally obtained.
[0130] Evidence related to the technical effects obtained in the embodiments of the present invention.
[0131] (1) Static ranging experiments at different distances
[0132] The present invention selects a typical narrow indoor environment, which has relatively good sealing and less interference, and is an ideal experimental scenario, and is completely different from the locations in the training set, which is conducive to verifying the effectiveness of training the model in similar scenarios.
[0133] In addition, the present invention selects four different terminal devices for experiments. Among them, three are mobile phones issued on the market, and another uses a work card as the terminal. The three mobile phones are Xiaomi 5s Plus, Huawei Nova 7, and Vivo X80 Pro respectively. These three mobile phones come from different manufacturers and have different ideas in microphone selection. And the release times of the three mobile phones are also different. Among them, Xiaomi 5s Plus was released in 2016, Huawei Nova 7 was released in 2020, and Vivo X80 Pro was released in 2022, so the usage times of the devices are also different. Different mobile phone manufacturers and different usage times greatly enrich the diversity of device terminals. At the same time, these three terminals are not included in the training data set and can be used to verify the ranging effect of the trained model on brand-new device terminals. The work card is the above-mentioned data collection device and is used for comparison in the experiment.
[0134] The specific process of the experiment is as follows, as Figure 2(a) As shown. In the experiment of the present invention, 4 speaker devices were used, 3 of which were placed at one end of the scenario as the base stations to be range-measured. Since the time synchronization between the mobile phone device and the speaker terminal is not stable, it is difficult to directly obtain the range measurement. In the present invention, the remaining one base station is placed together with the terminal. At such a suitable short distance, through the traditional audio signal detection based on the cross-correlation function model, it can be basically ensured that the arrival time of this base station is accurate.
[0135] In this way, the distances from the three base stations at the far end to the terminal can be indirectly estimated. In this experiment, several different distances were selected for evaluation. At each point, audio was collected for a certain duration through the above-mentioned different terminals respectively, and the terminal remained stationary during the collection process. And as Figure 2 (b) shown, during the test, human occlusion was also carried out to evaluate the ranging effect under non-ideal conditions.
[0136] The audio base station devices were placed at one end of the corridor, and marks were made at 5 meters, 10 meters, 20 meters, and 30 meters respectively. Data was collected using the above four different device terminals. In the corridor environment, at these four different distances, the statistical results of the static ranging error are shown in Table 1. There are three comparison algorithms, namely CCFM-0.3, CCFM-0.5, and CD-DD. Among them, CCFM-0.3 represents the detection algorithm using the traditional cross-correlation function model, and the scaling factor λ = 0.3 is selected when finally calculating the ATD of the signal. Correspondingly, CCFM-0.5 is also a detection algorithm based on the cross-correlation function model, with the scaling factor λ = 0.5, while CD-DD represents the data-driven audio signal detection algorithm proposed by the present invention. There are three statistical indicators for the error results, namely the root mean square error RMSE (Root Mean Square Error), the standard deviation STD (Standard Deviation), and the maximum value MAX (Maximum).
[0137] Table 1 Statistical results of static ranging errors at different distances (unit: meter)
[0138]
[0139] From the results in Table 1, at a distance of 5 meters, in such a corridor environment at a relatively short distance, which is equivalent to the ideal situation in the laboratory, the effects of the two algorithms CCFM-0.3 and CCFM-0.5 are very good, and basically each index is better than that of the CD-DD method. This is because the good conditions make the signal received by the microphone very close to the ideal signal, and the similarity curve calculated by the cross-correlation function only has an obvious single peak, and the signal-to-noise ratio is very high. However, observing the performance of the Xiaomi 5sPlus at a distance of 5 meters is relatively poor compared to other terminals. This is due to the difference caused by the hardware microphone performance. This is also a problem that traditional detection algorithms based on cross-correlation functions are difficult to solve.
[0140] At other distances, as the distance increases, the attenuation and interference of the sound signal in the propagation path will increase, and the corresponding detection effect will decrease. Even when it reaches a distance of 10 meters, there are some mutations in the effect of the Vivo X80Pro smartphone. The root mean square error results of CCFM-0.3 and CCFM-0.5 change from about 0.070 meters at 5 meters to about 0.400 meters at 10 meters, and the maximum error is also close to 1.400 meters. At this distance, the maximum error of CCFM-0.3 of the Xiaomi 5sPlus even reaches 2.165 meters. For the subsequent 20 meters and 30 meters, the overall performance stabilizes. This is because the interference in the corridor scene is relatively small and the space is relatively narrow, and the energy attenuation of the sound signal is much better than that in an open environment, which makes the signal-to-noise ratio of the data collected in this environment relatively high. The CD-DD method proposed in the present invention has very stable performance at different distances, which benefits from the data augmentation during model input, enabling the convolutional network model to have stable detection capabilities at different distances. The root mean square error of the CD-DD method is stable at about 0.2 meters. The smallest root mean square error is 0.118 meters of the Huawei Nova7 at 10 meters, and the largest is 0.263 meters of the Xiaomi 5sPlus at 10 meters. The standard deviation is stable at about 0.1 meters. The maximum error is 1.105 meters of the Vivo X80Pro at 20 meters. Overall, the CD-DD method proposed in the present invention maintains a relatively stable ranging effect for different mobile phones at different distances. Although it does not have as good ranging performance as the detection algorithm based on the cross-correlation function model under ideal conditions, the difference is not very large. Looking only at the statistical results of the CD-DD method, it can be found that there is no difference in the detection results of the three smartphones and the result of the work permit, which indicates that the CD-DD method has more advantages in terms of device differences.
[0141] (2) Comparison and analysis of dynamic ranging effects at different speeds
[0142] After comparing the static ranging performance, the dynamic ranging effect cannot be ignored. After all, in actual use, the tracked and located device is in motion most of the time. The test scenarios, equipment and methods are basically the same as those of static tests. In the dynamic ranging comparison experiment, the present invention designs experiments from several aspects such as different moving speeds, whether there is human body occlusion and different scenes. In the dynamic ranging comparison experiment, each experiment starts from about 30 meters away from the base station, approaches the base station equipment at different speeds, and when the distance is only about 5 meters, it moves away from the base station and returns to the starting point, and then immediately approaches and returns, for a total of two round trips. During the movement, it is a uniform linear motion, and each time it passes a ground marking point, the arrival timestamp is recorded, so that the true value of the ranging at the marking point can be confirmed according to the time. Three forms of expression are given for each group of experiments. First, the graph of the distance measurement results over time vividly shows the changes in the results; second, the cumulative distribution graph of the distance measurement errors under different conditions vividly shows the distribution of the overall error; third, the quantitative statistical table of the error results, the statistical indicators are one standard deviation and two standard deviations, abbreviated as CDF-68% and CDF-95%.
[0143] The dynamic collection of audio signals will be affected by the Doppler effect, and this effect will be greater as the relative speed of movement increases. Therefore, in this set of experiments, different terminals use different movement speeds in the process of moving away from and approaching the base station. There are three movement speeds in the experiment, namely slow, normal and fast. Normal speed is the daily walking speed of people, while slow is relatively slower, and fast means that the steps are more frequent and the speed is slightly faster. However, these three speeds are all within the range of pedestrians' usual movement speeds, and the fast one does not mean running. This set of experiments was still conducted in a corridor environment. At different speeds, the dynamic ranging results are as follows: Figure 3 , Figure 4 and Figure 5 .
[0144] First Look Figure 3Ranging results of different devices under medium and slow movement. The figure shows the change of ranging results of the devices over time. The blue curve represents the result of CCFM-0.3, the red line represents the ranging result of CCFM-0.5, and the yellow curve is the result of the data-driven audio signal detection algorithm CD-DD proposed by the present invention. Theoretically, the experimental results should be smooth curves. However, overall, the actual results of all devices have many spikes, and even some algorithms have many jumps. Generally speaking, the ranging result of Huawei Nova7 should be the best, with only a few jumps. In the result of the work ID card, the result of CCFM-0.3 is significantly abnormal, and interference in front of the direct path is often detected. The CCFM-0.5 result of Xiaomi 5sPlus is abnormal when ranging near the base station is far away at about 100 seconds, and the abnormality seems to be very regular, seemingly having a complementary relationship with the true ranging. It can basically be determined that the signal of the same echo path is detected, and this echo path is very likely to come from the wall at the end of the corridor.
[0145] Looking again Figure 4 Ranging results under normal speed movement and Figure 5 Ranging results under fast movement. From the perspective of time, as the walking speed increases, the time taken to walk the same test path gradually decreases. From the perspective of distance, the situation where CCFM-0.5 frequently detects echoes seems to decrease, while there are still many cases where CCFM-0.3 detects interference in front of the direct path. In comparison, the detection effect of the CD-DD method proposed by the present invention is generally better. As time goes by, its distance observations change relatively smoothly. However, due to the large amount of information that needs to be expressed in the figure, only the general change of the distance observations can be seen, and the specific values of the jumps cannot be visually seen.
[0146] Table 2 Statistical results of dynamic ranging errors at different speeds (unit: meter)
[0147]
[0148]
[0149] Therefore, in addition to the figure of the actual detection results over time, this experiment also statistically analyzed the cumulative distribution of errors in the results, such as Figure 6 , Figure 7 and Figure 8As shown. Different detection algorithms are distinguished by colors, while different devices are distinguished by different line types. In this way, it can be clearly compared that the CD-DD method proposed by the present invention is superior to the traditional detection algorithm based on the cross-correlation function model at any speed, because in the three figures, the blue curves, that is, the error cumulative distributions of CD-DD, are superior to the other two detection algorithms. At different speeds, the statistical results of dynamic ranging errors are shown in Table 2. Generally speaking, the ranging error at slow speed is better than that at normal speed, and the normal speed is better than that at fast speed. This is exactly because the above-mentioned Doppler frequency shift increases with the increase of the relative motion speed. At slow speed, for CCFM, the maximum of CDF-68% reaches 1.000 m on several terminal devices, and CDF-95% even reaches 7.900 m; while for the CC-DD algorithm, the result of CDF-68% is only 0.650 m, and at the same time, the maximum of CDF-95% is only 1.180 m. In the error results at normal speed and fast speed, the performance of the CC-DD algorithm proposed by the present invention is also very stable, and the maximum of CDF-95% is only 1.551 m. Looking at the error results at different speeds as a whole, in terms of the maximum CDF-68%, the CC-DD method optimizes 33% compared with CCFM, and in terms of the maximum CDF-95%, it optimizes 80%, and the dynamic ranging performance has been greatly improved.
[0150] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or field programmable gate arrays and programmable logic devices, can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.
[0151] The above is only the specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A data-driven audio ranging method, characterized in that: The method includes: S1: Preprocess the original signal and convert the filtered time domain signal into a time-frequency image through short-time Fourier transform (STFT); S2: Use the convolutional neural network model to obtain the arrival time estimate; The S1 specifically includes: The time domain signal is converted into a time-frequency image through short-time Fourier transform (STFT), which reflects the change of energy of each frequency point of the signal over time; When performing short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, while the size of the window determines the frequency resolution of the input image; 1ms is the step size of the window sliding, which is equivalent to a resolution of 0.34m, and the window size is selected as 512; The detailed steps of the short-time Fourier transform (STFT) include: The sliding window starts from the starting point of the received sound signal data. At this time, the window function is centered on time and the signal is processed by the window function. Perform Fourier transform on the data to obtain the PSD matrix of the sound data of the first window, which represents the vector of the received signal at the time delay; The S2 specifically includes: The first convolution layer of the audio detection model uses a kernel of size 3x3 with a stride and padding of 1 to perform convolution operations with the time-frequency map. The output does not change the image size, but the number of channels increases from 1 to 8. Use linear rectification function for nonlinear activation; The second convolution operation reduces the 8 channels to 1 channel, and the linear rectification function ReLU is used for nonlinear activation. Subsequently, the two-dimensional image is converted into a one-dimensional vector, and the coefficients are regularized by random dropout; Finally, the learned features are fitted through two fully connected layers to obtain the arrival time estimation of the chirp in the sound signal.
2. A data-driven audio ranging system, characterized in that: The system includes: A signal preprocessing module is used to preprocess the original signal and convert the filtered time domain signal into a time-frequency image through short-time Fourier transform (STFT); A convolutional neural network module to process the time-frequency images and estimate the arrival time; The signal preprocessing module specifically includes: It is used to convert the time domain signal into a time-frequency image through short-time Fourier transform (STFT), which reflects the change of energy of each frequency point of the signal over time; It is used to set the length of the sliding window to L and the moving step of the window to S, so that the time resolution is seconds, the frequency resolution is Hz; Fs is the sampling frequency, the step size of the sliding window is 1ms, and the window size is 512; The detailed steps of the short-time Fourier transform (STFT) include: Slide the window from the starting point of the received sound signal data and perform window function processing with t as the center; Perform Fourier transform on the data to obtain the power spectrum density (PSD) matrix of the sound data of the first window; Among them, the signal segment is processed by the Hamming window function and then Fourier transformed; The convolutional neural network module specifically includes: The first convolution layer uses a convolution kernel of size 3x3 and performs convolution operation with the time frequency image with parameters of stride and padding of 1. The output channel is changed from 1 to 8, and nonlinear activation is performed through the linear rectification function ReLU; The second convolution layer uses the convolution kernel to reduce 8 channels to 1 channel and performs nonlinear activation through the linear rectification function ReLU; The conversion module converts the two-dimensional image into a one-dimensional vector and regularizes the coefficients by random dropout; The fully connected layer module contains two fully connected layers, which are used to fit the learned features and finally estimate the arrival time of the chirp in the sound signal.
Citation Information
Patent Citations
Method for region positioning of indoor sound source based on convolutional neural network
CN109001679A