Audio signal direct vision state identification method and system based on energy envelope skewness

Through the audio signal processing method based on energy envelope skewness, the LoS and NLoS states of the signal are accurately identified, which solves the problem that indoor positioning is susceptible to interference from non-line-of-sight signals, and improves positioning stability and recognition accuracy.

CN119946558AActive Publication Date: 2025-05-06SHENZHEN CANQOON TECHNOLOGY CO LTD

Patent Information

Application Number
CN202411955018.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Indoor positioning is susceptible to interference from non-line-of-sight signals, resulting in a decrease in positioning stability. It is difficult for the prior art to accurately judge the LoS and NLoS states of the signal.

Method used

The audio signal direct-view state recognition method based on the skewness of the energy envelope is adopted. The energy density map is obtained through the audio signal processing module, the statistics module sorts and normalizes the energy, the calculation module fits the envelope and calculates the skewness, and the judgment module estimates the LoS/NLoS state of the signal.

Benefits of technology

High accuracy classification of LoS states of single Chirp signals is achieved, which reduces sensitivity to equipment and environment differences, and improves the robustness and recognition accuracy of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946558A_ABST
    Figure CN119946558A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of indoor positioning, and discloses an audio signal direct vision state identification method and system based on energy envelope skewness, and the method comprises the steps: solving an energy density map of an audio signal through an audio signal processing module; sorting and normalizing the energy of each frequency point in the single-frame EDM through a statistical module to obtain an energy distribution histogram; fitting an envelope line of the energy histogram by adopting a kernel density estimation method through an operation module, and calculating the skewness of the envelope line; and estimating the LoS / NLoS state of the signal through a discrimination module. According to the invention, high-accuracy, low-cost and low-complexity signal state identification can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the field of indoor positioning technology, and in particular relates to a method and system for identifying the direct-view state of an audio signal based on the skewness of an energy envelope. Background Art

[0002] As one of the important links of ubiquitous Beidou, the importance of indoor positioning is becoming more and more significant in the context of national strategy and the needs of the times. Thanks to the advantages of high adaptability of audio signal terminals and long propagation distance of a single base station, audio-based indoor positioning technology has developed rapidly. Common audio positioning systems are mostly based on the time of arrival (ToA) of the detection signal to estimate the distance from the terminal to the base station. However, the indoor topology varies, and the movement of pedestrians is random and changeable. The terminal is easily interfered by non-line-of-sight (NLoS) signals in complex channel environments, which increases the ToA detection error and leads to a decrease in positioning stability. To this end, judging the LoS and NLoS states of the signal is a prerequisite for accurate positioning ( Figure 1 ).

[0003] In view of the above analysis, the technical problems that need to be solved urgently in the prior art are:

[0004] Indoor positioning is susceptible to non-line-of-sight signal interference, so determining the LoS and NLoS states of the signal is a prerequisite for accurate positioning. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention provides a method and system for identifying the direct-view state of an audio signal based on the skewness of an energy envelope.

[0006] The present invention is implemented as follows: a method for identifying a direct-view state of an audio signal based on the skewness of an energy envelope, characterized in that the method for identifying a direct-view state of an audio signal based on the skewness of an energy envelope specifically comprises:

[0007] S1: Obtaining the energy density map (EDM) of the audio signal through the audio signal processing module;

[0008] S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram;

[0009] S3: Through the operation module, the envelope of the energy histogram is fitted using the kernel density estimation method, and the envelope skewness is calculated;

[0010] S4: Estimate the LoS / NLoS state of the signal through the discrimination module.

[0011] Further, in S1, when the terminal starts the microphone sensor, firstly, a bandpass filter is applied to the collected original audio time domain data to filter out the interference of non-target frequency band environmental noise; secondly, the filtered audio signal is subjected to short-time Fourier transform to obtain EDM, as follows:

[0012] (1) When the terminal starts the microphone sensor, it first applies a 12th-order Butterworth bandpass filter to the original audio time domain data with a sampling rate of 48kHz to filter out the interference of environmental noise in non-target frequency bands:

[0013]

[0014] Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, is the filtered time domain audio stream data;

[0015] (2) The audio stream is continuously monitored, and short-time Fourier transform (STFT) calculation is performed on each frame of data with a time unit of 100 ms. STFT is used in conjunction with a window function, and the Hanning window parameters with a window length of l and an overlap rate of k are selected.

[0016] Furthermore, in S2, the energy of each frequency point in a single frame of EDM is quickly sorted in ascending order, and the energy is normalized with reference to the maximum frequency point energy of the LoS signal at a distance of 1m from the base station:

[0017]

[0018] Among them, E i,j Indicates the EDM frequency energy value corresponding to the image subscript index i and j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value;

[0019] Taking 5 normalized energy units as statistical intervals, the energy values ​​of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1,x2,...,x n}, where n represents the total number of intervals, and a histogram is drawn.

[0020] Furthermore, in S3, a Gaussian kernel function is first selected to weight each frequency point of the energy histogram:

[0021]

[0022] Among them, h is the bandwidth parameter, which is given by the empirical rule. The kernel density estimation function is:

[0023]

[0024] The skewness of the energy histogram envelope is calculated as follows:

[0025]

[0026] Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Respectively represent the energy value probability mean and standard deviation of the envelope.

[0027] Further, in S4, the skewness of the energy histogram envelope of all collected audio data at different distances is calculated and the average is taken. If the skewness of the envelope is positive, the terminal is in the LoS state of the base station; if the skewness of the envelope is negative, the terminal is in the NLoS state of the base station.

[0028] Another object of the present invention is to provide an audio signal direct view state recognition system based on energy envelope skewness, the system specifically comprising:

[0029] An audio signal processing module, used for obtaining an energy density diagram of an audio signal;

[0030] The statistical module is used to sort and normalize the energy of each frequency point in a single frame of EDM to obtain an energy distribution histogram;

[0031] A calculation module is used to fit the envelope of the energy histogram and calculate the skewness of the envelope;

[0032] The discrimination module is used to estimate the LoS / NLoS state of the signal.

[0033] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0034] First, the present invention adopts the EDM energy histogram envelope skewness method, which has high accuracy and coverage for the classification of the LoS state of a single Chirp signal and is insensitive to differences in equipment and environment.

[0035] The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: NLoS, as one of the important pain points faced by the indoor positioning industry in the process of technology implementation, often consumes a lot of manpower and material resources in the process of environmental survey, base station layout plan optimization and site implementation. In order to reduce the impact of NLoS on the final position estimation results, construction personnel often have to modify the existing base station locations, and through additional weak current construction, or even additional encrypted base stations to ensure that the terminal can always receive a good signal. By accurately judging the LoS and NLoS states of the signals, it is possible to downgrade signals of poor quality in the positioning algorithm, greatly reducing their impact on the output of the positioning results, so that the base station location design, weak current construction, base station installation, and later operation and maintenance in the project process are no longer subject to the influence of environmental conditions, achieving cost reduction and efficiency improvement.

[0036] The technical solution of the present invention fills the technical gap in the industry at home and abroad: Since the indoor positioning industry is still dominated by radio frequency technologies such as Bluetooth and ultra-wideband (UWB) in the current projects implemented, a large number of discussions on NLoS are also based on radio frequency signals. However, the core patents or authorizations of these radio frequency technologies are in the hands of foreign technology companies or alliances. As a high-privacy, high-precision positioning signal source, the research on the application of audio in the field of indoor positioning is still a "blue ocean". Based on the audio Chirp signal, the present invention identifies the direct-view state of the audio signal in a simple and efficient way, which increases the robustness of the positioning algorithm in the face of poor signals, improves the usability of audio positioning technology in the scene application process, and effectively increases the voice of domestic technology in the competition in the indoor positioning industry.

[0037] The technical solution of the present invention solves a technical problem that people have been eager to solve but have never succeeded in solving: the discussion on the NLoS problem in the industry has a long history. In recent years, with the expansion of deep learning technology and the surge in the number of IoT devices connected to the network, more and more results have emerged that use big data and network models to identify signal states. However, for consumer-grade terminals, especially embedded devices, in order to ensure low power consumption and low computing memory overhead, the model has to be quantized or pruned before it can run normally, which introduces the technical problems of low precision and low recall.

[0038] The present invention summarizes and generalizes the characteristics of a large number of audio data signals in different scenarios, and proposes a signal direct-viewing state recognition method based on the skewness of the energy envelope. This method achieves efficient judgment of the signal state by the terminal with low computational complexity and algorithm operating conditions.

[0039] Radio frequency signals account for the vast majority of discussions and research on NLoS because they have a larger bandwidth, their signal encoding is more free, and the signals can carry more information and are easier to quantify. In fact, as one of the carriers of traditional multimedia, audio can have extremely rich time-frequency characteristics through appropriate signal processing methods. Moreover, as a mechanical wave, the physical phenomena such as signal diffraction, reflection, refraction, and diffraction generated by audio signals in the face of complex environmental topology are stable and evaluable. The energy density diagram of audio signals can reliably reflect the signal's obstruction status.

[0040] Second, the existing audio signal direct-view state recognition technology usually relies on complex multi-signal coordination or multi-base station auxiliary algorithms, which makes the system more sensitive to environmental noise interference, especially when the signal spectrum is wide or the background noise is strong, the recognition accuracy is significantly reduced. The present invention can effectively filter out non-target frequency band noise and extract signal features by introducing the analysis method of energy density map (EDM) and envelope skewness, thereby improving the robustness and recognition accuracy of the system in complex environments.

[0041] Traditional methods rely on multi-base station collaboration to enhance the recognition accuracy of the signal direct line of sight state, which not only increases the complexity and cost of hardware deployment, but also limits the flexibility of application scenarios. This paper uses the energy characteristics of single-base station audio signals and combines the kernel density estimation method to construct a single-base station distance perception network (DPNet) based on deep learning, which significantly reduces hardware requirements and achieves high-precision LoS / NLoS state discrimination in a single-base station scenario.

[0042] Existing methods usually use high-dimensional features and complex models for recognition, and feature extraction takes a long time and cannot meet the needs of real-time applications. The present invention uses lightweight algorithms such as fast sorting, normalization, and kernel density estimation to achieve efficient signal feature extraction and state discrimination. Combined with an optimized modular architecture, the system can quickly calculate and complete state recognition based on real-time audio data collection, significantly improving real-time performance and operating efficiency.

[0043] The technical solution proposed in the present invention can adapt to a variety of complex application scenarios, such as indoor positioning, smart home, vehicle communication and wireless channel optimization, which not only simplifies the deployment process, but also improves the recognition accuracy and system reliability. By innovatively combining the energy envelope skewness and deep learning technology, the present invention has made significant technological progress in the field of state recognition of audio signals, providing reliable support for the intelligent upgrading of related industries, and has broad market prospects and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of LoS and NLoS states of audio signals in an environment provided by an embodiment of the present invention;

[0045] Figure 2 is a flow chart of a method for identifying a direct-viewing state of an audio signal based on energy envelope skewness provided by an embodiment of the present invention;

[0046] Figure 3 It is a comparison of LoS and NLoS of a single frame EDM at different distances in a corridor scenario provided by an embodiment of the present invention;

[0047] Figure 4 is a normalized energy distribution histogram provided by an embodiment of the present invention;

[0048] Figure 5 It is the single-frame EDM energy histogram envelope at different distances in the corridor scene provided by an embodiment of the present invention, (left) LoS condition; (right) NLoS condition;

[0049] Figure 6 It is a module diagram of an audio signal direct-viewing state recognition system based on energy envelope skewness provided by an embodiment of the present invention;

[0050] Figure 7 This is the LoS / soft-NLoS state classification confusion matrix provided by an embodiment of the present invention, (upper left) overall sample; (upper right) Nova8 Pro sample; (lower left) corridor sample; (lower right) conference room sample;

[0051] Figure 8 It is the signal state recognition accuracy and recall rate of three different test devices provided in the embodiment of the present invention;

[0052] Fig. 9 These are the signal state recognition accuracy and recall rate for two different positioning scenarios provided by the embodiments of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] Example 1: LoS / NLoS status identification of indoor audio signals

[0055] When using audio signals for target detection in an indoor environment, the audio signals may be in a non-line-of-sight (NLoS) state during propagation due to obstacles such as furniture and walls. This embodiment identifies the LoS / NLoS state of the audio signal by using an energy envelope skewness method.

[0056] 1) Signal acquisition: The audio signal is collected through a microphone array installed in the room. The signal is a continuous speech signal over a period of time.

[0057] 2) Energy density map calculation: The collected audio signal is input into the audio signal processing module, and the spectrum is calculated using short-time Fourier transform (STFT) to generate the corresponding energy density map (EDM).

[0058] 3) Energy histogram construction: sort the energy of all frequency points within a certain time frame in EDM, and generate an energy distribution histogram through normalization.

[0059] 4) Kernel density estimation and skewness calculation: Use the kernel density estimation method to fit the energy distribution histogram, generate a continuous envelope, and calculate the skewness value of the envelope.

[0060] 5) State determination: Compare the skewness value with the preset threshold range. If the skewness value is small, it is determined to be LoS (line of sight); if the skewness value is large, it is determined to be NLoS (non-line of sight).

[0061] 6) Result output: The system output signal is currently in LoS or NLoS state.

[0062] The test results show that in indoor environments, the average skewness value of LoS signals is 0.8, while the average skewness value of NLoS signals is 2.5, and the recognition accuracy rate reaches 95%.

[0063] Example 2: Signal blocking status detection in outdoor environment

[0064] In unmanned driving or robot navigation scenarios, audio signals are used to detect whether obstacles block the sensor's line of sight (LoS / NLoS status), thereby optimizing path planning.

[0065] Implementation steps:

[0066] 1) Signal acquisition: The audio sensor array on the unmanned vehicle emits a broadband signal and collects the audio signals reflected from the environment.

[0067] 2) Energy density map calculation: The received reflected signal is preprocessed and the energy density map is generated using short-time Fourier transform.

[0068] 3) Energy distribution histogram generation: Extract the energy data of the frequency points from the single-frame energy density map, normalize them after sorting, and construct the energy distribution histogram.

[0069] 4) Envelope fitting and skewness calculation: The kernel density estimation method is used to fit the histogram, generate the envelope and calculate its skewness.

[0070] 5) State determination: Compare the skewness value with the threshold. If the skewness value is low, the audio signal is judged to be in LoS state (no obstacle); if the skewness value is high, it is judged to be in NLoS state (with obstacle).

[0071] 6) Path adjustment: If the NLoS state is detected, the unmanned vehicle automatically adjusts the navigation path to avoid obstacles.

[0072] Through testing in open areas and obstacle areas, the unmanned vehicle was able to successfully distinguish between LoS and NLoS states, with a LoS state recognition rate of 98% and a NLoS state recognition rate of 96%. This method effectively improves the path planning efficiency and safety of the unmanned vehicle.

[0073] These two embodiments respectively demonstrate the practical application of audio signal LoS / NLoS state recognition in indoor and outdoor scenarios, reflecting the versatility and efficiency of the method.

[0074] like Figure 2 As shown, an embodiment of the present invention provides a method for identifying a direct-viewing state of an audio signal based on energy envelope skewness, the method specifically comprising:

[0075] S1: Obtaining the energy density map (EDM) of the audio signal through the audio signal processing module;

[0076] S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram;

[0077] S3: Through the operation module, the envelope of the energy histogram is fitted using the kernel density estimation (KDE) method, and the envelope skewness is calculated;

[0078] S4: Estimate the LoS / NLoS state of the signal through the discrimination module.

[0079] First, the audio signal is preprocessed by the audio signal processing module and decomposed into multiple frequency components to generate the corresponding energy density map (EDM). EDM is a two-dimensional representation with time as the horizontal axis and frequency as the vertical axis. The value of each point in the graph represents the energy density at the corresponding time and frequency. The core of this step is to extract the spectral information of the signal through Fourier transform or short-time Fourier transform, and calculate the energy of each frequency point to provide a data basis for subsequent analysis.

[0080] In the second step, the statistical module sorts the energy data of all frequency points in a single frame of EDM by size, normalizes it, and generates an energy distribution histogram. The purpose of normalization is to eliminate the influence of different signal strengths so that the histogram reflects the distribution characteristics of energy between frequency points. This process converts the original frequency energy distribution into a standardized frequency distribution model by counting the proportion of the energy value of each frequency point to the total energy, which is convenient for subsequent calculation and analysis.

[0081] The operation module fits the energy distribution histogram through the Kernel Density Estimation (KDE) method to generate a smooth envelope. Kernel density estimation is a non-parametric statistical method that can generate a continuous probability density function (envelope) based on the discrete data points of the histogram. Subsequently, the skewness of the envelope is calculated, that is, the degree of deviation of the envelope from the symmetry of its mean. The positive and negative and absolute size of the skewness value can reflect the imbalance and deviation direction of the energy distribution in the signal, providing a feature quantity for state discrimination.

[0082] Finally, the discrimination module estimates the LoS (line-of-sight) or NLoS (non-line-of-sight) state of the signal based on the skewness value of the envelope. Generally, the energy distribution of the LoS signal is more concentrated and the skewness is smaller; while the energy distribution of the NLoS signal is more dispersed due to the multipath effect, and the skewness value is significantly increased. The discrimination module uses a preset classification model or threshold to classify the direct line of sight state of the signal by comparing the skewness feature value and outputs the discrimination result.

[0083] Through the above steps, this method realizes the accurate identification of the LoS / NLoS state of the audio signal. The core of this method is to capture the statistical characteristics of the signal's direct-view state through the skewness characteristics of the energy envelope, which has high discrimination accuracy and applicability.

[0084] In S1, after the terminal starts the microphone sensor, a 12th-order Butterworth bandpass filter (BPF) is first applied to the original audio time domain data with a sampling rate of 48 kHz to filter out the interference of environmental noise in non-target frequency bands.

[0085]

[0086] Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, It is the filtered time domain audio stream data.

[0087] The audio stream is continuously monitored and STFT calculation is performed on each frame of data with a time unit of 100ms.

[0088] STFT is usually used in conjunction with a window function to reduce frequency leakage caused by non-integer period sampling. To ensure that EDM has sufficient time resolution and frequency resolution, after repeated tests, the Hanning window parameters of 512 window length and 87.5% overlap were finally selected.

[0089] After STFT operation, the audio signal within the 100ms window length obtains an EDM with a pixel size of 33×68, with a time resolution of approximately 1.3ms / pixel and a frequency resolution of approximately 93.75Hz / pixel.

[0090] In a corridor 1.7m wide and 35m long, the front and back EDM of a single audio signal at different distances are compared. Taking Huawei Nova8 Pro mobile phone as an example, 60s of audio data are collected at each distance. Some of the results are shown below: Figure 3 shown.

[0091] On the one hand, from the perspective of the image: in horizontal comparison, as the physical distance increases, although the direct path of the signal gradually dims due to energy attenuation, it is always the most prominent part of the EDM; in vertical comparison, the signal in the LoS state has a concentrated direct path, while the signal in the NLoS state is more dispersed and the brightness is significantly reduced.

[0092] On the other hand, under the premise that the environment does not change, the direct path energy in the LoS state accounts for a high proportion, but due to the superposition of reverberation energy and direct path energy, it is difficult to divide the boundary between the two based on the threshold.

[0093] In S2, considering that different energy distributions of EDM actually reflect the numerical fluctuation of the direct path energy ratio, the energy of each frequency point in a single frame of EDM is sorted, normalized and a histogram is drawn.

[0094] The energy of each frequency point in a single frame of EDM is quickly sorted in ascending order, and the energy is normalized with reference to the maximum frequency point energy of the LoS signal at a distance of 1m from the base station:

[0095]

[0096] Among them, E i,j Indicates the EDM frequency energy value corresponding to the image subscript index i and j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value.

[0097] Taking 5 normalized energy units as statistical intervals, the energy values ​​of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1,x2,...,x n}, where n represents the total number of intervals. Draw a histogram, such as Figure 4 shown.

[0098] The S3, kernel density estimation is often used to measure the probability density curve of a set of discrete data. First, select the Gaussian kernel function (GaussianKernel) to weight each frequency point of the energy histogram:

[0099]

[0100] Where h is the bandwidth parameter, given by the rule of thumb. Then the kernel density estimation function is:

[0101]

[0102] Figure 5 The results of fitting the energy histogram envelope at different distances from the terminal to the base station using the KDE method in the LoS and NLoS states are shown. Figure 5 , the overall frequency energy in both states decreases approximately linearly as the physical distance between the terminal and the base station increases; and the non-audio signal parts are concentrated in the value range of 120 to 160. Comparing the LoS state on the left and the NLoS state on the right, the normalized energy range of the audio signal in the former is correspondingly greater than that in the latter by about 20 energy units; more importantly, the center of gravity of the energy distribution of the two has shifted significantly - the center of gravity of the former tends to be centered and to the right, while the center of gravity of the latter tends to be to the left.

[0103] Based on the standard normal distribution, the skewness statistic measures the symmetry of the data distribution: when the skewness is positive, the data center of gravity is on the left side of the peak; when the skewness is negative, the data center of gravity is on the right side of the peak. The skewness calculation of the energy histogram envelope is as follows:

[0104]

[0105] Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Respectively represent the energy value probability mean and standard deviation of the envelope.

[0106] In S4, the skewness of the energy histogram envelope of all collected audio data at different distances is calculated and the average is taken. The results are shown in Table 1.

[0107] Table 1. The skewness of the single-frame EDM energy histogram envelope at different distances when soft occlusion exists in the corridor scene

[0108]

[0109] As can be seen from Table 1, when facing the base station (LoS), the calculated skewness is positive, while when facing away from the base station (NLoS), the calculated skewness is negative; the calculation results are consistent with the image representation. Therefore, the envelope skewness, as a measure to describe the asymmetry of energy distribution, reflects the facing away state of the received signal with high reliability, that is, if the envelope skewness is positive, the terminal is in the LoS state of the base station; if the envelope skewness is negative, the terminal is in the NLoS state of the base station.

[0110] like Figure 6 As shown, an embodiment of the present invention provides an audio signal direct viewing state recognition system based on energy envelope skewness, specifically comprising:

[0111] An audio signal processing module, used for obtaining an energy density diagram of an audio signal;

[0112] The statistical module is used to sort and normalize the energy of each frequency point in a single frame of EDM to obtain an energy distribution histogram;

[0113] A calculation module is used to fit the envelope of the energy histogram and calculate the skewness of the envelope;

[0114] The discrimination module is used to estimate the LoS / NLoS state of the signal.

[0115] The LoS state judgment of the signal based on the skewness of the EDM energy histogram envelope belongs to a binary classification problem, so five general indicators, namely accuracy, precision, specificity, recall, and F1 score, are used to evaluate the performance of the classification method:

[0116]

[0117] Among them, TP, TN, FP, and FN are the number of samples that correctly classify EDM as LoS, correctly classify EDM as NLoS, misclassify EDM as LoS, and misclassify EDM as NLoS. The result of the judgment of the LoS / soft-NLoS state of the signal mainly affects the rationality of sending EDM to DPNet to perceive the distance. Therefore, it is better to misjudge LoS as NLoS than to misjudge NLoS as LoS. Therefore, under the premise of ensuring a high TP value, the smaller the FP value, the better, that is, the accuracy and F1 score can better express the value of binary classification.

[0118] We used three new test devices, Huawei Nova8 Pro, Huawei Mate30, and Oppo Reno5, to collect 16,920 frames of corridor and conference room data ranging from 1 to 35 meters. Figure 7 The confusion matrix of the overall sample and some single sample classification results is shown. Figure 8The signal state recognition accuracy and recall rate of three different test devices are shown. Fig. 9 The signal state recognition accuracy and recall rate of two different positioning scenarios are shown, and Table 2 gives the detailed classification performance of all samples. Observation results show that the precision and recall rate of the overall samples are approximately 0.98 and 0.96, respectively, that is, 2 frames of misjudgment and 4 frames of missed judgment occurred in the EDM of 100 frames of LoS. In terms of comprehensive performance, the classification effects of three different devices and two different indoor scenes are compared. The Nova8 Pro with better performance in the former is only about 1% to 2% better than the Reno5 with poorer performance, while there is almost no difference in the latter. The data comparison results show that the EDM energy histogram envelope skewness method has high accuracy and coverage for the classification of the LoS state of a single Chirp signal, and is insensitive to differences in equipment and environment.

[0119] Table 2 LoS / soft-NLoS state classification performance of different sample sources for single Chirp signal

[0120]

[0121] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0122] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for identifying the direct viewing state of an audio signal based on the skewness of an energy envelope, characterized in that: The method specifically includes: S1: Obtaining an energy density diagram of the audio signal through an audio signal processing module; S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram; S3: Through the operation module, the envelope of the energy histogram is fitted using the kernel density estimation method, and the envelope skewness is calculated; S4: Estimate the LoS / NLoS state of the signal through the discrimination module.

2. The method for identifying the direct-view state of an audio signal based on the energy envelope skewness according to claim 1, characterized in that: In S1, when the terminal starts the microphone sensor, a bandpass filter is first applied to the collected original audio time domain data to filter out the interference of non-target frequency band environmental noise; secondly, the filtered audio signal is subjected to short-time Fourier transform to obtain EDM, as follows: (1) When the terminal starts the microphone sensor, it first applies a 12th-order Butterworth bandpass filter to the original audio time domain data with a sampling rate of 48kHz to filter out the interference of environmental noise in non-target frequency bands: Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, is the filtered time domain audio stream data; (2) The audio stream is continuously monitored and STFT calculation is performed on each frame of data with a time unit of 100 ms. STFT is used in conjunction with a window function, and the Hanning window parameters with a window length of l and an overlap rate of k are selected.

3. The method for identifying the direct-view state of an audio signal based on the energy envelope skewness according to claim 1, characterized in that: S2, quickly sorts the energy of each frequency point in a single frame of EDM in ascending order, and normalizes the energy with reference to the maximum frequency point energy of the LoS signal at a distance of 1m from the base station: Among them, E i,j Indicates the EDM frequency energy value corresponding to the image subscript index i and j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value; Taking 5 normalized energy units as statistical intervals, the energy values ​​of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1,x2,...,x n }, where n represents the total number of intervals, and a histogram is drawn.

4. The method for identifying the direct-view state of an audio signal based on the energy envelope skewness according to claim 1, characterized in that: In S3, a Gaussian kernel function is first selected to weight each frequency point of the energy histogram: Among them, h is the bandwidth parameter, which is given by the empirical rule. The kernel density estimation function is: The skewness of the energy histogram envelope is calculated as follows: Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Respectively represent the energy value probability mean and standard deviation of the envelope.

5. The method for identifying the direct viewing state of an audio signal based on the energy envelope skewness according to claim 1, characterized in that: The S4 calculates the skewness of the energy histogram envelope of all collected audio data at different distances and takes the average. If the skewness of the envelope is positive, the terminal is in the LoS state of the base station; if the skewness of the envelope is negative, the terminal is in the NLoS state of the base station.

6. A system for identifying the direct-view state of an audio signal based on the energy envelope skewness as claimed in claims 1-5, characterized in that: The system specifically includes: A data preprocessing module for collecting audio signals and converting the time-domain audio signals into energy density maps (EDMs); The normalization module is used to sort the energy values ​​of the frequency points in the EDM and perform normalization according to the energy values ​​of the reference frequency points; The operation module uses the kernel density estimation method to fit the envelope of the energy histogram and calculates the skewness of the envelope; The discrimination module determines the signal’s direct-view state (LoS / NLoS) by analyzing the envelope skewness.

7. The audio signal direct view state recognition system according to claim 6, characterized in that: The data preprocessing module comprises: The filter unit is used to perform a 12th-order Butterworth bandpass filter on the collected time-domain audio signal to filter out environmental noise interference in non-target frequency bands; A short-time Fourier transform unit is used to perform a short-time Fourier transform (STFT) on the filtered audio signal to generate an energy density map of each frame time window; The frequency range of the filter is dynamically adjusted according to the characteristics of the target audio signal.

8. The audio signal direct view state recognition system according to claim 6, characterized in that: The normalization module is used to: Quickly sort the frequency energy in EDM from small to large; The maximum frequency energy value at 1 meter away from the base station in the line of sight (LoS) state is used as the normalization reference value, and the energy values ​​of all frequency points are normalized to the [0,1] interval; The distribution frequency of the frequency point energy after statistical normalization is calculated, and the statistical interval is divided according to the set energy unit to draw an energy histogram.

9. The audio signal direct view state recognition system according to claim 6, characterized in that: The discrimination module comprises: The kernel density estimation unit smoothes the energy histogram based on the Gaussian kernel function and fits the energy envelope; A skewness calculation unit, which determines the symmetry of the energy distribution by calculating the skewness of the envelope; The state determination unit determines the direct view state of the signal based on the comparison between the envelope skewness value and the preset threshold, where: If the skewness is positive, it is judged to be in LoS state; The skewness is negative, which is considered to be NLoS state.

Citation Information

Patent Citations

  • Detecting the location of a phone using RF wireless and ultrasonic signals

    CN107850667A

  • Chirp signal detection method under multipath and non-line-of-sight indoor environment

    CN115954015A

  • Indoor audio fingerprint positioning method and system, medium, equipment and terminal

    CN116164751A

  • Truck type identification system and identification method based on Bluetooth network

    CN117392857A

  • Feature preprocessing and extracting method for multi-channel audio positioning

    CN117630818A

Cited By

  • Micro-seismic activity monitoring method and system based on energy release intensity K line

    CN122488207A