A method and system for identifying the direct-view state of an audio signal based on energy envelope skewness
The direct-looking state of the audio signal is identified through the energy envelope skewness method, which solves the problem of non-visual interference in indoor positioning, and realizes high-precision and low-complexity signal state judgment, which is suitable for indoor positioning, smart home and on-board communication scenarios.
Patent Information
- Application Number
- CN202411955018.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Indoor positioning is susceptible to interference from non-line-of-sight signals, resulting in a decrease in positioning stability, and it is difficult for the prior art to accurately judge the direct-looking state of the signal.
The audio signal direct-looking state recognition method based on the energy envelope skewness is adopted, and the LoS/NLoS state is identified through audio signal processing, energy density graph calculation, kernel density estimation and skewness calculation.
It improves the accuracy and robustness of indoor positioning, reduces hardware requirements and computing complexity, adapts to a variety of complex application scenarios, and improves recognition accuracy and real-timeness.
Smart Images

Figure CN119946558B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of indoor positioning technology, and in particular relates to a method and system for identifying the direct-view state of an audio signal based on energy envelope skewness. Background Art
[0002] As one of the important links of ubiquitous Beidou, the importance of indoor positioning is becoming increasingly significant in the context of national strategy and the needs of the times. Thanks to the advantages of high adaptability of audio signal terminals and long propagation distance of a single base station, audio-based indoor positioning technology has developed rapidly. Common audio positioning systems are mostly based on the time of arrival (ToA) of the detection signal to estimate the distance from the terminal to the base station. However, the indoor topology is different, and the movement of pedestrians is random and changeable. The terminal is easily interfered by the non-line-of-sight (NLoS) signal in a complex channel environment, which increases the ToA detection error and leads to a decrease in positioning stability. To this end, judging the LoS and NLoS status of the signal is a prerequisite for accurate positioning ( Figure 1 ).
[0003] In view of the above analysis, the technical problems that need to be solved urgently in the existing technology are:
[0004] Indoor positioning is susceptible to interference from non-line-of-sight signals, so determining the LoS and NLoS states of the signal is a prerequisite for accurate positioning. Summary of the Invention
[0005] In view of the problems existing in the prior art, the present invention provides a method and system for identifying the direct-view state of an audio signal based on the skewness of an energy envelope.
[0006] The present invention is achieved by providing a method for identifying a direct view state of an audio signal based on the skewness of an energy envelope, characterized in that the method for identifying a direct view state of an audio signal based on the skewness of an energy envelope specifically comprises:
[0007] S1: Obtain the energy density map (EDM) of the audio signal through the audio signal processing module;
[0008] S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram;
[0009] S3: Using the calculation module, the kernel density estimation method is used to fit the envelope of the energy histogram and the skewness of the envelope is calculated;
[0010] S4: Estimate the LoS / NLoS state of the signal through the discrimination module.
[0011] Furthermore, in S1, when the terminal starts the microphone sensor, a bandpass filter is first applied to the collected raw audio time domain data to filter out the interference of non-target frequency band ambient noise; secondly, the filtered audio signal is subjected to short-time Fourier transform to obtain EDM, as follows:
[0012] (1) When the terminal starts the microphone sensor, it first applies a 12th-order Butterworth bandpass filter to the original audio time domain data with a sampling rate of 48kHz to filter out the interference of environmental noise in non-target frequency bands:
[0013]
[0014] Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, is the filtered time domain audio stream data;
[0015] (2) The audio stream is continuously monitored, and a short-time Fourier transform (STFT) is performed on each frame of data with a time unit of 100ms. The STFT is used in conjunction with a window function, and the Hanning window parameters with a window length of l and an overlap rate of k are selected.
[0016] Furthermore, in S2, the energy of each frequency point in a single EDM frame is quickly sorted in ascending order, and the energy is normalized with reference to the maximum frequency point energy of the LoS signal at a distance of 1 m from the base station:
[0017]
[0018] Among them, E i,j Indicates the EDM frequency energy value of the corresponding image subscript index i, j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value;
[0019] Taking 5 normalized energy units as statistical intervals, the energy values of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1, x2, ..., x n}, where n represents the total number of intervals, and a histogram is drawn.
[0020] Furthermore, in step S3, a Gaussian kernel function is first selected to weight each frequency point of the energy histogram:
[0021]
[0022] Where h is the bandwidth parameter, given by the empirical rule, the kernel density estimation function is:
[0023]
[0024] The skewness of the energy histogram envelope is calculated as follows:
[0025]
[0026] Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Represent the energy value probability mean and standard deviation of the envelope respectively.
[0027] Furthermore, in S4, the skewness of the energy histogram envelope of all collected audio data at different distances is calculated and the average is taken. If the skewness of the envelope is positive, the terminal is in the LoS state of the base station; if the skewness of the envelope is negative, the terminal is in the NLoS state of the base station.
[0028] Another object of the present invention is to provide an audio signal direct view state recognition system based on energy envelope skewness, the system specifically comprising:
[0029] An audio signal processing module, used to obtain an energy density map of the audio signal;
[0030] The statistical module is used to sort and normalize the energy of each frequency point in a single frame of EDM to obtain an energy distribution histogram;
[0031] The calculation module is used to fit the envelope of the energy histogram and calculate the skewness of the envelope;
[0032] The discrimination module is used to estimate the LoS / NLoS state of the signal.
[0033] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0034] First, the present invention adopts the EDM energy histogram envelope skewness method, which has high accuracy and coverage for the classification of the LoS state of a single chirp signal and is insensitive to the differences in equipment and environment.
[0035] The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: NLoS is one of the important pain points faced by the indoor positioning industry in the process of technology implementation, and often consumes a lot of manpower and material resources in the process of environmental survey, base station layout plan optimization and site implementation. In order to reduce the impact of NLoS on the final position estimation result, construction personnel often have to modify the existing base station site, and carry out additional weak current construction, or even additional encrypted base stations to ensure that the terminal can always receive a good signal. By accurately judging the LoS and NLoS status of the signal, it is possible to downgrade the weight of the signal of poor quality in the positioning algorithm, greatly reducing its impact on the output of the positioning result, so that the base station site design, weak current construction, base station installation, post-operation and maintenance and other links in the project process are no longer subject to the influence of environmental conditions, thereby achieving cost reduction and efficiency improvement.
[0036] The technical solution of the present invention fills the technical gaps in the industry at home and abroad: Since the current indoor positioning industry still mainly uses radio frequency technologies such as Bluetooth and ultra-wideband (UWB) in implemented projects, a large number of discussions on NLoS are also based on radio frequency signals. However, the core patents or authorizations of these radio frequency technologies are all in the hands of foreign technology companies or alliances. As a high-privacy, high-precision positioning signal source, the research on its application in the field of indoor positioning is still a "blue ocean". Based on audio Chirp signals, the present invention uses a simple and efficient method to identify the direct line of sight state of audio signals, thereby increasing the robustness of the positioning algorithm in the face of poor signals, improving the usability of audio positioning technology in scene applications, and effectively increasing the voice of domestic technologies in the competition in the indoor positioning industry.
[0037] The technical solution of this invention solves a long-standing technical challenge that has eluded successful solutions: the NLoS problem, a long-standing topic within the industry. In recent years, with the expansion of deep learning technology and the surge in the number of IoT devices connected to the internet, a growing number of research projects have emerged using big data and network models to identify signal states. However, for consumer-grade devices, especially embedded ones, models must be quantized or pruned to ensure low power consumption and minimal computational memory overhead. This, in turn, introduces the technical challenges of low precision and recall.
[0038] By summarizing and generalizing the signal characteristics of a large number of audio data in different scenarios, the present invention proposes a signal direct-view state recognition method based on the skewness of the energy envelope. This method realizes efficient judgment of the signal state by the terminal with low computational complexity and algorithm operating conditions.
[0039] RF signals dominate discussions and research on NLoS because they possess greater bandwidth, offer greater freedom in signal encoding, and can carry more information that is easily quantified. In reality, audio, as a traditional multimedia medium, can possess rich time-frequency characteristics through appropriate signal processing. Furthermore, as mechanical waves, audio signals exhibit stable and measurable physical phenomena such as diffraction, reflection, refraction, and reflection in complex environmental topologies. Energy density diagrams of audio signals can reliably reflect signal obstruction.
[0040] Second, existing audio signal direct-view state recognition technologies typically rely on complex multi-signal collaboration or multi-base station-assisted algorithms, making the system more sensitive to environmental noise interference. This significantly reduces recognition accuracy, especially in situations with wide signal spectra or strong background noise. This new method, by introducing an energy density map (EDM) and envelope skewness analysis method, can effectively filter out non-target frequency band noise and extract signal features, improving the system's robustness and recognition accuracy in complex environments.
[0041] Traditional methods rely on multi-base station collaboration to enhance the accuracy of line-of-sight recognition, which not only increases the complexity and cost of hardware deployment but also limits the flexibility of application scenarios. This paper utilizes the energy characteristics of single-base station audio signals, combined with kernel density estimation methods, to construct a deep learning-based single-base station distance perception network (DPNet). This significantly reduces hardware requirements while achieving high-precision LoS / NLoS state discrimination in single-base station scenarios.
[0042] Existing methods typically use high-dimensional features and complex models for recognition, resulting in time-consuming feature extraction and failing to meet the demands of real-time applications. This invention utilizes lightweight algorithms such as quick sorting, normalization, and kernel density estimation to achieve efficient signal feature extraction and state discrimination. Combined with an optimized modular architecture, the system rapidly calculates and completes state recognition based on real-time audio data acquisition, significantly improving real-time performance and operational efficiency.
[0043] The technical solution proposed in this invention can adapt to a variety of complex application scenarios, such as indoor positioning, smart homes, in-vehicle communications, and wireless channel optimization. It not only simplifies the deployment process but also improves recognition accuracy and system reliability. By innovatively combining energy envelope skewness with deep learning technology, this invention has achieved significant technological advancement in the field of audio signal state recognition, providing reliable support for the intelligent upgrade of related industries and possessing broad market prospects and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 2 is a schematic diagram of LoS and NLoS states of audio signals in an environment provided by an embodiment of the present invention;
[0045] Figure 2 This is a flow chart of a method for identifying a direct view state of an audio signal based on energy envelope skewness provided by an embodiment of the present invention;
[0046] Figure 3 This is a comparison of LoS and NLoS of a single-frame EDM at different distances in a corridor scenario provided by an embodiment of the present invention;
[0047] Figure 4 is a normalized energy distribution histogram provided by an embodiment of the present invention;
[0048] Figure 5 The single-frame EDM energy histogram envelopes at different distances in the corridor scenario provided by an embodiment of the present invention are (left) LoS condition; (right) NLoS condition;
[0049] Figure 6 This is a module diagram of an audio signal direct view state recognition system based on energy envelope skewness provided by an embodiment of the present invention;
[0050] Figure 7 This is the LoS / soft-NLoS state classification confusion matrix provided by an embodiment of the present invention, (upper left) overall sample; (upper right) Nova8 Pro sample; (lower left) corridor sample; (lower right) conference room sample;
[0051] Figure 8 The signal state recognition accuracy and recall rate of three different test devices provided in the embodiment of the present invention are:
[0052] Figure 9 These are the signal state recognition accuracy and recall rate for two different positioning scenarios provided by the embodiments of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0054] Example 1: LoS / NLoS status identification of indoor audio signals
[0055] When using audio signals for target detection in indoor environments, obstacles such as furniture and walls may cause non-line-of-sight (NLoS) states during audio signal propagation. This embodiment uses the energy envelope skewness method to identify the LoS / NLoS states of audio signals.
[0056] 1) Signal acquisition: The audio signal is collected through a microphone array installed in the room. The signal is a continuous speech signal over a period of time.
[0057] 2) Energy density map calculation: The collected audio signal is input into the audio signal processing module, and the spectrum is calculated using short-time Fourier transform (STFT) to generate the corresponding energy density map (EDM).
[0058] 3) Energy histogram construction: sort the energy of all frequency points within a certain time frame in the EDM and generate an energy distribution histogram through normalization.
[0059] 4) Kernel density estimation and skewness calculation: Use the kernel density estimation method to fit the energy distribution histogram, generate a continuous envelope, and calculate the skewness value of the envelope.
[0060] 5) Status determination: Compare the skewness value with the preset threshold range. If the skewness value is small, it is determined to be LoS (line of sight); if the skewness value is large, it is determined to be NLoS (non-line of sight).
[0061] 6) Result output: The system output signal is currently in LoS or NLoS state.
[0062] The test results show that in indoor environments, the average skewness value of LoS signals is 0.8, while the average skewness value of NLoS signals is 2.5, and the recognition accuracy rate reaches 95%.
[0063] Example 2: Signal blocking status detection in outdoor environment
[0064] In unmanned driving or robot navigation scenarios, audio signals are used to detect whether obstacles block the sensor's line of sight (LoS / NLoS state), thereby optimizing path planning.
[0065] Implementation steps:
[0066] 1) Signal acquisition: The audio sensor array on the unmanned vehicle emits a broadband signal and collects the audio signals reflected from the environment.
[0067] 2) Energy density map calculation: The received reflected signal is preprocessed and the energy density map is generated using short-time Fourier transform.
[0068] 3) Energy distribution histogram generation: Extract the energy data of the frequency points from the single-frame energy density map, sort and normalize them, and construct the energy distribution histogram.
[0069] 4) Envelope fitting and skewness calculation: The kernel density estimation method is used to fit the histogram, generate the envelope and calculate its skewness.
[0070] 5) State determination: The skewness value is compared with the threshold. If the skewness value is low, the audio signal is judged to be in the LoS state (no obstacle); if the skewness value is high, it is judged to be in the NLoS state (there is an obstacle).
[0071] 6) Path adjustment: If the NLoS state is detected, the unmanned vehicle automatically adjusts the navigation path to avoid obstacles.
[0072] Through testing in both open and obstructed areas, the autonomous vehicle was able to successfully distinguish between LoS and NLoS states, with a LoS recognition rate of 98% and a NLoS recognition rate of 96%. This method effectively improves the efficiency and safety of the autonomous vehicle's path planning.
[0073] These two examples demonstrate the practical application of audio signal LoS / NLoS state recognition in indoor and outdoor scenarios, respectively, demonstrating the versatility and efficiency of the method.
[0074] like Figure 2 As shown, an embodiment of the present invention provides a method for identifying the direct view state of an audio signal based on the skewness of the energy envelope, the method specifically comprising:
[0075] S1: Obtain the energy density map (EDM) of the audio signal through the audio signal processing module;
[0076] S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram;
[0077] S3: Using the calculation module, the kernel density estimation (KDE) method is used to fit the envelope of the energy histogram and calculate the envelope skewness;
[0078] S4: Estimate the LoS / NLoS state of the signal through the discrimination module.
[0079] First, the audio signal is preprocessed by the audio signal processing module and decomposed into multiple frequency components to generate a corresponding energy density map (EDM). The EDM is a two-dimensional representation with time on the horizontal axis and frequency on the vertical axis. The value of each point in the map represents the energy density at the corresponding time and frequency. The core of this step is to extract the signal's spectral information through Fourier transform or short-time Fourier transform and calculate the energy of each frequency point, providing a data foundation for subsequent analysis.
[0080] In the second step, the statistics module sorts the energy data for all frequencies in a single EDM frame by magnitude and normalizes it to generate an energy distribution histogram. Normalization eliminates the influence of varying signal strengths, allowing the histogram to reflect the energy distribution characteristics across frequencies. This process calculates the proportion of each frequency's energy value to the overall energy, converting the original frequency energy distribution into a standardized frequency distribution model for subsequent calculations and analysis.
[0081] The computation module uses the kernel density estimation (KDE) method to fit the energy distribution histogram and generate a smooth envelope. Kernel density estimation is a nonparametric statistical method that generates a continuous probability density function (envelope) from the discrete data points of the histogram. The skewness of the envelope is then calculated, which is the degree to which the envelope deviates from its symmetry with respect to its mean. The positive and negative sign and absolute magnitude of the skewness value can reflect the imbalance and direction of the energy distribution in the signal, providing a characteristic quantity for state discrimination.
[0082] Finally, the discrimination module estimates the signal's LoS (line-of-sight) or NLoS (non-line-of-sight) status based on the envelope's skewness. Typically, LoS signals have a more concentrated energy distribution and a smaller skewness; whereas NLoS signals, due to multipath effects, have a more dispersed energy distribution and a significantly larger skewness. Using a pre-set classification model or threshold, the discrimination module compares the skewness feature values to classify the signal's line-of-sight status and outputs the discrimination result.
[0083] Through the above steps, this method can accurately identify the LoS / NLoS state of the audio signal. Its core lies in capturing the statistical characteristics of the signal's direct-view state through the skewness characteristics of the energy envelope, which has high discrimination accuracy and applicability.
[0084] In the aforementioned S1, after the terminal starts the microphone sensor, a 12th-order Butterworth bandpass filter (BPF) is first applied to the original audio time domain data with a sampling rate of 48 kHz to filter out interference from ambient noise in non-target frequency bands.
[0085]
[0086] Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, It is the filtered time-domain audio stream data.
[0087] The audio stream is continuously monitored, and STFT calculation is performed on each frame of data with a time unit of 100ms.
[0088] STFT is often used in conjunction with a window function to mitigate frequency leakage caused by non-integer period sampling. To ensure sufficient time and frequency resolution for EDM, a Hanning window with a window length of 512 and an overlap ratio of 87.5% was selected after repeated testing.
[0089] The audio signal within the 100ms window length is processed by STFT to obtain an EDM with a pixel size of 33×68, with a time resolution of approximately 1.3ms / pixel and a frequency resolution of approximately 93.75Hz / pixel.
[0090] In a corridor 1.7m wide and 35m long, the forward and backward EDM of a single audio signal at different distances are compared. Taking the Huawei Nova8 Pro mobile phone as an example, 60s of audio data are collected at each distance. Some of the results are shown below. Figure 3 shown.
[0091] On the one hand, from the perspective of images: horizontal comparison, as the physical distance increases, although the signal direct path gradually dims due to energy attenuation, it is always the most prominent part of the EDM; vertical comparison, the signal in the LoS state has a concentrated direct path, while the signal in the NLoS state is more dispersed and the brightness is significantly reduced.
[0092] On the other hand, under the premise that the environment does not change, the direct path energy in the LoS state accounts for a high proportion, but since the reverberation energy and the direct path energy are superimposed on each other, it is difficult to divide the boundary between the two based on the threshold.
[0093] In S2, considering that different energy distributions of EDM actually reflect the numerical fluctuation of the direct path energy ratio, the energy of each frequency point in a single frame of EDM is sorted, normalized, and a histogram is drawn.
[0094] The energy of each frequency point in a single EDM frame is quickly sorted in ascending order, and the energy is normalized with reference to the maximum frequency point energy of the LoS signal at 1m away from the base station:
[0095]
[0096] Among them, E i,j Indicates the EDM frequency energy value of the corresponding image subscript index i, j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value.
[0097] Taking 5 normalized energy units as statistical intervals, the energy values of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1, x2, ..., x n}, where n represents the total number of intervals. Draw a histogram, such as Figure 4 shown.
[0098] The kernel density estimation in S3 is often used to calculate the probability density curve of a set of discrete data. First, a Gaussian kernel function is selected to weight each frequency point of the energy histogram:
[0099]
[0100] Where h is the bandwidth parameter, which is given by empirical rules. Then the kernel density estimation function is:
[0101]
[0102] Figure 5 The results of fitting the energy histogram envelope at different distances from the terminal to the base station using the KDE method are shown in the LoS and NLoS states. Figure 5 In both states, the overall frequency energy decreases roughly linearly with increasing physical distance from the terminal to the base station. The non-audio signal portion is concentrated in the 120-160 range. Comparing the LoS state (left) with the NLoS state (right), the normalized audio signal energy range in the LoS state (left) is approximately 20 energy units greater than that in the NLoS state (right). More importantly, the center of gravity of the energy distribution is significantly shifted between the two states—the former is centered and to the right, while the latter is shifted to the left.
[0103] Based on the standard normal distribution, the skewness statistic measures the symmetry of the data distribution: a positive skewness indicates that the data center of gravity is to the left of the peak; a negative skewness indicates that the data center of gravity is to the right of the peak. The skewness of the energy histogram envelope is calculated as follows:
[0104]
[0105] Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Represent the energy value probability mean and standard deviation of the envelope respectively.
[0106] In the step S4, the skewness of the energy histogram envelope of all collected audio data at different distances is calculated and the average is taken. The results are shown in Table 1.
[0107] Table 1. The skewness of the single-frame EDM energy histogram envelope at different distances depending on whether there is soft occlusion in the corridor scene.
[0108]
[0109] As shown in Table 1, the calculated skewness is always positive when facing the base station (LoS), while it is always negative when facing away from the base station (NLoS). The calculated results are consistent with the image representation. Therefore, the envelope skewness, as a measure of energy distribution asymmetry, reliably reflects the facing away state of the received signal. That is, a positive envelope skewness indicates that the terminal is in the LoS state with the base station, while a negative envelope skewness indicates that the terminal is in the NLoS state with the base station.
[0110] like Figure 6 As shown, an embodiment of the present invention provides an audio signal direct view state recognition system based on energy envelope skewness, specifically comprising:
[0111] An audio signal processing module, used to obtain an energy density map of the audio signal;
[0112] The statistical module is used to sort and normalize the energy of each frequency point in a single frame of EDM to obtain an energy distribution histogram;
[0113] The calculation module is used to fit the envelope of the energy histogram and calculate the skewness of the envelope;
[0114] The discrimination module is used to estimate the LoS / NLoS state of the signal.
[0115] Determining the LoS status of a signal based on the skewness of the EDM energy histogram envelope is a binary classification problem. Therefore, five general metrics, including accuracy, precision, specificity, recall, and F1 score, are used to evaluate the performance of the classification method:
[0116]
[0117] TP, TN, FP, and FN represent the number of EDM samples correctly classified as LoS, correctly classified as NLoS, incorrectly classified as LoS, and incorrectly classified as NLoS, respectively. The determination of the LoS / soft-NLoS state of the signal significantly impacts the rationality of feeding the EDM into the DPNet for distance perception. Therefore, it is better to misclassify LoS as NLoS than to misclassify NLoS as LoS. Therefore, while maintaining a high TP value, a smaller FP value is preferred. This means that precision and F1 score better reflect the value of binary classification.
[0118] We used three new test devices, Huawei Nova8 Pro, Huawei Mate30, and Oppo Reno5, to collect 16,920 frames of corridor and conference room data ranging from 1 to 35 meters. Figure 7 The confusion matrix of the overall sample and some individual sample classification results is shown. Figure 8The signal state recognition accuracy and recall rate of three different test devices are demonstrated. Figure 9 The signal state recognition accuracy and recall rate for two different positioning scenarios are demonstrated, and Table 2 gives the detailed classification performance of all samples. Observation results show that the precision and recall rate of the overall sample are approximately 0.98 and 0.96, respectively, which means that 2 frames are misclassified and 4 frames are missed in 100 frames of EDM LoS. In terms of comprehensive performance, the classification effects of three different devices and two different indoor scenarios are compared. The Nova8 Pro, which performs better in the former, only outperforms the Reno5, which performs worse, by about 1% to 2%, while there is almost no difference in the latter. The data comparison results show that the EDM energy histogram envelope skewness method has high accuracy and coverage for the classification of the LoS state of a single chirp signal, and is insensitive to differences in equipment and environment.
[0119] Table 2 LoS / soft-NLoS state classification performance of different sample sources for single Chirp signal
[0120]
[0121] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0122] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A method for identifying the direct view state of an audio signal based on the skewness of the energy envelope, characterized in that: The method specifically includes: S1: Obtain the energy density map of the audio signal through the audio signal processing module; S2: Through the statistical module, the energy of each frequency point in a single frame of EDM is sorted and normalized to obtain an energy distribution histogram; S3: Using the calculation module, the kernel density estimation method is used to fit the envelope of the energy histogram and the skewness of the envelope is calculated; S4: Estimate the LoS / NLoS state of the signal through the discrimination module; In S1, after the terminal starts the microphone sensor, it first applies a bandpass filter to the collected raw audio time domain data to filter out the interference of non-target frequency band ambient noise; then, the filtered audio signal is subjected to a short-time Fourier transform to obtain an EDM, as follows: (1) When the terminal starts the microphone sensor, it first applies a 12th-order Butterworth bandpass filter to the original audio time domain data with a sampling rate of 48kHz to filter out the interference of environmental noise in non-target frequency bands: Among them, s(t) is the original time domain audio stream data, f BPF [·] represents a Butterworth bandpass filter, is the filtered time domain audio stream data; (2) Continuously monitor the audio stream and perform STFT calculation on each frame of data with a time unit of 100ms. STFT is used in conjunction with a window function, and the Hanning window parameters are selected with a window length of l and an overlap rate of k; In step S2, the energy of each frequency point in a single EDM frame is quickly sorted in ascending order, and the energy is normalized with reference to the maximum frequency point energy of the LoS signal at a distance of 1 meter from the base station: Among them, E i,j Indicates the EDM frequency energy value of the corresponding image subscript index i, j, E max Indicates the maximum EDM frequency energy value at a distance of 1m in the LoS state. Indicates the normalized EDM frequency energy value; Taking 5 normalized energy units as statistical intervals, the energy values of the frequency points falling in the corresponding intervals are counted to obtain the data frequency of each interval {x1, x2, ..., x n }, where n represents the total number of intervals, and a histogram is drawn; In S3, a Gaussian kernel function is first selected to weight each frequency point of the energy histogram: Where h is the bandwidth parameter, given by the empirical rule, the kernel density estimation function is: The skewness of the energy histogram envelope is calculated as follows: Among them, e represents the normalized energy value probability corresponding to each point fitted by the envelope, and σ e Represent the energy value probability mean and standard deviation of the envelope respectively.
2. The method for identifying the direct view state of an audio signal based on the energy envelope skewness according to claim 1, wherein: In the step S4, the skewness of the energy histogram envelope of all collected audio data at different distances is calculated and the average is taken. If the skewness of the envelope is positive, the terminal is in the LoS state of the base station; if the skewness of the envelope is negative, the terminal is in the NLoS state of the base station.
3. An audio signal direct view state recognition system based on the audio signal direct view state recognition method based on energy envelope skewness according to any one of claims 1-2, characterized in that: The system specifically includes: a data preprocessing module for collecting audio signals and converting the time-domain audio signals into energy density maps (EDMs); Normalization module, used to sort the energy values of the EDM frequency points and normalize them according to the energy values of the reference frequency points; The calculation module uses the kernel density estimation method to fit the envelope of the energy histogram and calculate the skewness of the envelope; The discrimination module determines the direct viewing state of the signal by analyzing the skewness of the envelope.
4. The audio signal direct view state recognition system according to claim 3, wherein: The data preprocessing module includes: The filter unit is used to perform a 12th-order Butterworth bandpass filter on the collected time-domain audio signal to filter out environmental noise interference in non-target frequency bands; A short-time Fourier transform unit is used to perform a short-time Fourier transform (STFT) on the filtered audio signal to generate an energy density map of each frame time window; The frequency range of the filter is dynamically adjusted according to the characteristics of the target audio signal.
5. The audio signal direct view state recognition system according to claim 3, wherein: The normalization module is used to: Quickly sort the frequency energy in EDM from small to large; The maximum frequency energy value at 1 meter from the base station in the line-of-sight (LoS) state is used as the normalization reference value, and the energy values of all frequency points are normalized to the range of [0, 1]. The distribution frequency of the frequency point energy after statistical normalization is calculated, and the statistical interval is divided according to the set energy unit to draw an energy histogram.
6. The audio signal direct view state recognition system according to claim 3, wherein: The discrimination module includes: The kernel density estimation unit smoothes the energy histogram based on the Gaussian kernel function and fits the energy envelope; a skewness calculation unit, which determines the symmetry of the energy distribution by calculating the skewness of the envelope; The state determination unit determines the direct view state of the signal based on the comparison between the envelope skewness value and the preset threshold, where: If the skewness is positive, it is determined to be in LoS state; If the skewness is negative, it is determined to be in NLoS state.
Citation Information
Patent Citations
Chirp signal detection method under multipath and non-line-of-sight indoor environment
CN115954015A
Feature preprocessing and extracting method for multi-channel audio positioning
CN117630818A