Sound directional perception and enhancement method and system
Through the multi-sensing fusion processing of microwave transceiver and microphone, the voice capture and enhancement problems in complex noise environments are solved, efficient directional speech recognition and enhancement are achieved, and noise interference is reduced.
Patent Information
- Application Number
- CN202110962510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-08-20
AI Technical Summary
The prior art is difficult to achieve effective voice capture and enhancement in complex noise environments, and the signal processing process is complex and time-consuming. The beamforming microphone array requires huge equipment and low directional accuracy.
The micro vibration information of the sound source is collected in a directional manner through a microwave transceiver, and the ambient sound signal is collected synchronously with the microphone for fusion processing, including high-pass filtering, time-frequency analysis and matching filtering, realizing directional perception and enhancement of sound.
It realizes speech recognition and enhancement in a strong background noise and clutter aliasing environment, reduces interference from other sound sources and ambient noise, and improves the accuracy of directional perception.
Smart Images

Figure CN115881109B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of audio processing, specifically a method and system for sound directional perception and enhancement Background Art
[0002] Existing speech enhancement methods often use signal processing algorithms to post-process speech signals to reduce ambient noise and enhance the speech signal. This method generally has a good suppression effect on ambient noise, primarily Gaussian white noise. However, the sound perception and enhancement effect are poor in complex noise environments and those with mixed noise. Furthermore, the signal processing process is complex and time-consuming. Microphone array technology based on beamforming can achieve directional speech perception, but this method often requires a large microphone array and, due to its low directional accuracy, is inevitably affected by interference from ambient noise and other nearby sound sources. Summary of the Invention
[0003] Aiming at the problem that existing sound perception methods are difficult to achieve voice capture and enhancement in situations such as strong background noise and clutter interference, the present invention proposes a sound directional perception and enhancement method and system, which uses microwave transceivers and microphones for multi-sensor fusion to achieve directional perception of sound source targets and sound enhancement.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a method for directional sound perception and enhancement. The method uses a microwave transceiver to directionally collect micro-vibration information of a sound source to be measured, while simultaneously collecting ambient sound signals through a microphone. The sound signal extracted from the vibration information is fused with the ambient sound signal to achieve directional sound perception and enhancement.
[0006] The extraction refers to: performing high-pass filtering on the vibration signal of the sound source target to invert the sound signal of the directional sound source.
[0007] The fusion refers to: performing time-frequency analysis on the directional sound signal and the ambient sound signal respectively to obtain their time-frequency distribution characteristics and dividing them into time-frequency units with a length of L and a width of W respectively, processing the baseband signal of the microwave transceiver within each time-frequency unit, extracting the vibration information of the target sound source, performing matched filtering on the ambient sound signal based on it, and then converting the filtering result from the time-frequency domain to the time domain signal.
[0008] The microwave transceiver includes: a continuous wave microwave signal source, a power divider, a power amplifier, a mixer, a low-pass filter, a conditioning circuit, a transceiver antenna and a processor, wherein: the continuous wave microwave signal source generates a continuous wave microwave signal, which is output to the power amplifier through the power divider to be emitted by the transmitting antenna, and output to the mixer to be mixed with the reflected signal received by the receiving antenna for processing, the mixer is connected to the low-pass filter and the mixed signal is input to the low-pass filter, the low-pass filter is connected to the conditioning circuit, the conditioning circuit is connected to the processor, and the processor collects and processes the baseband signal output by the conditioning circuit into directional vibration information of the sound source target.
[0009] Technical Effects
[0010] The present invention comprehensively solves the problem that existing sound perception methods are difficult to achieve speech recognition and enhancement under strong background noise interference. It uses the vibration information of the sound source target obtained by directional measurement of a microwave transceiver to perform matched filtering on the sound signal collected by the microphone to achieve directional perception and enhancement of sound. Compared with existing conventional technical means, the technical details of the present invention are significantly improved as follows: method innovation, and the use of microwave transceivers and microphones for multi-sensor fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 Flowchart of the present invention;
[0012] Figure 2 Schematic diagram of time-frequency characteristics of an embodiment;
[0013] Figure 3 Schematic diagram of time-frequency blocks in an embodiment;
[0014] Figure 4 Schematic diagram of the embodiment effect;
[0015] Figure 5 2 is a schematic diagram of an application example. DETAILED DESCRIPTION
[0016] like Figure 1 As shown, this embodiment relates to a method for sound directional perception and enhancement, comprising the following steps:
[0017] Step 1: Use a microwave transceiver and microphone to synchronously sense the voice signal in the environment. Specifically, use one speaker to play the voice signal and two speakers to play the voice interference signal at the same time to simulate a test scenario with complex noise interference. Then, the microphone and microwave transceiver synchronously collect the sound signal in the measurement environment. The microphone is used to collect and capture the sound signal in the test environment; the antenna of the microwave transceiver is oriented toward the sound source to test the micro-vibration of the sound source target. The microwave transceiver obtains the baseband signal by transmitting and receiving the microwave signal, and inverts the vibration displacement information of the sound source target. Then the vibration signal is processed by high-pass filtering to invert the sound signal, where λ is the wavelength of the microwave signal, is the extracted interference phase evolution information.
[0018] Step 2: Fusing the directional voice signal extracted by the microwave transceiver with the voice signal captured by the microphone to achieve sound directional perception and enhancement, specifically including:
[0019] Step 2.1, perform time-frequency analysis on the sound signal sensed and inverted by the microwave transceiver to obtain its time-frequency characteristics, such as Figure 2 shown.
[0020] Step 2.2, such as Figure 3 As shown in the figure, after performing time-frequency analysis on the speech signal collected by the microphone to obtain its time-frequency distribution, the window length L and window width W of the time-frequency unit are selected, and the time-frequency distribution of the sound signal inverted by the microwave transceiver and the sound signal collected by the microphone are divided into the same time-frequency unit TF microwave (g, h)(g∈[1,G],h∈[1,H]) and TF microphone (g, h)(g∈[1, G], h∈[1, H]), where: G and H are the number of rows and columns of the video unit formed after time-frequency division, g and h are the row index and column index of the selected time-frequency unit, respectively. microphone represents the time-frequency distribution;
[0021] Step 2.3: Process the baseband signal of the microwave transceiver in each time-frequency unit, extract the vibration information of the target sound source, and perform matched filtering on the microphone voice signal: TF fusioo (g, h) = TF microwave (g,h)*TF microphone (g, h), where: * is the matched filter operation. Figure 4 As shown in FIG, this is the output result after matched filtering and denoising. It can be seen that this method can well separate the sound signal of the target sound source of interest.
[0022] Step 2.4: Convert the matched filtering output from the time-frequency domain to the time domain signal.
[0023] like Figure 5As shown, the sound directional perception and enhancement system for implementing the above-mentioned method in this embodiment includes: a microwave transceiver, a microphone, a sound directional perception and enhancement signal processor, and a speech result output unit, wherein: the microwave transceiver is used to transmit and receive continuous wave microwave signals, and then extract the vibration displacement information of the target surface from the echo signal reflected by the target; the microphone is used to collect speech signals in the environment, the sound directional perception and enhancement signal processor processes the vibration displacement information of the target surface measured by the microwave transceiver and the speech signal collected by the microphone, and obtains a directionally enhanced speech signal by fusing the two sensed sound signals; and the speech result output unit is used to output the directionally enhanced speech signal.
[0024] Compared with the existing technology, this method realizes the directional perception of sound and reduces the interference of other sound sources and environmental noise.
[0025] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.
Claims
1. A method for sound directional perception and enhancement, characterized in that: While the microwave transceiver collects the vibration information of the sound source to be measured in a directionally controlled manner, the microphone simultaneously collects the ambient sound signal; the sound signal extracted from the vibration information is integrated with the ambient sound signal to achieve directional sound perception and enhancement; The fusion specifically includes: Step 2.1, perform time-frequency analysis on the sound signal sensed and inverted by the microwave transceiver to obtain its time-frequency characteristics; Step 2.2: After performing time-frequency analysis on the voice signal collected by the microphone to obtain its time-frequency distribution, select the window length L and window width W of the time-frequency unit, and divide the time-frequency distribution of the sound signal inverted by the microwave transceiver and the sound signal collected by the microphone into the same time-frequency unit. and , where: G and H are the number of rows and columns of the time-frequency unit formed after time-frequency division, g and h are the row index and column index of the selected time-frequency unit, respectively. represents the time-frequency distribution; Step 2.3: Process the baseband signal of the microwave transceiver within each time-frequency unit, extract the vibration information of the target sound source, and perform matched filtering on the microphone voice signal: , where: * is the matched filtering operation; Step 2.4: Convert the matched filtering output from the time-frequency domain to the time domain signal.
2. The method for sound directional perception and enhancement according to claim 1, wherein: The extraction refers to: performing high-pass filtering on the vibration signal of the sound source target to invert the sound signal of the directional sound source.
3. The method for sound directional perception and enhancement according to claim 1, wherein: The fusion refers to: performing time-frequency analysis on the directional sound signal and the ambient sound signal respectively to obtain their time-frequency distribution characteristics and dividing them into time-frequency units with a length of L and a width of W respectively, processing the baseband signal of the microwave transceiver within each time-frequency unit, extracting the vibration information of the target sound source, performing matched filtering on the ambient sound signal based on it, and then converting the filtering result from the time-frequency domain to the time domain signal.
4. The method for sound directional perception and enhancement according to claim 1, wherein: The microwave transceiver includes: a continuous wave microwave signal source, a power divider, a power amplifier, a mixer, a low-pass filter, a conditioning circuit, a transceiver antenna and a processor, wherein: the continuous wave microwave signal source generates a continuous wave microwave signal, which is output to the power amplifier through the power divider to be emitted by the transmitting antenna, and output to the mixer to be mixed with the reflected signal received by the receiving antenna for processing, the mixer is connected to the low-pass filter and the mixed signal is input to the low-pass filter, the low-pass filter is connected to the conditioning circuit, the conditioning circuit is connected to the processor, and the processor collects and processes the baseband signal output by the conditioning circuit into directional vibration information of the sound source target.
5. The method for sound directional perception and enhancement according to any one of claims 1 to 4, wherein: include: Step 1: Use a microwave transceiver and a microphone to synchronously sense the voice signal in the environment. Specifically, the microphone is used to collect and capture the sound signal in the test environment; the antenna of the microwave transceiver is directed toward the sound source to be tested to directionally test the micro-vibration of the sound source target. The microwave transceiver obtains the baseband signal by transmitting and receiving microwave signals, and inverts the vibration displacement information of the sound source target. , and then perform high-pass filtering on the vibration signal to extract the sound signal, where: is the wavelength of the microwave signal, is the extracted interferometric phase evolution information; Step 2: Fuse the directional voice signal extracted by the microwave transceiver with the voice signal captured by the microphone to achieve sound directional perception and enhancement.
6. A sound directional perception and enhancement system for implementing the method according to any one of claims 1 to 5, characterized in that: include: A microwave transceiver, a microphone, a sound directional perception and enhancement signal processor, and a speech result output unit, wherein: the microwave transceiver is used to transmit and receive continuous wave microwave signals, and then extract the vibration displacement information of the target surface from the echo signal reflected by the target; the microphone is used to collect speech signals in the environment; the sound directional perception and enhancement signal processor processes the vibration displacement information of the target surface measured by the microwave transceiver and the speech signal collected by the microphone, and obtains a directionally enhanced speech signal by fusing the two sensed sound signals; and the speech result output unit is used to output the directionally enhanced speech signal.
Citation Information
Patent Citations
Vibration monitoring system and signal processing method based on LFMCW radar
CN107607923A
Device and method of performing automatic audio focusing on multiple objects
US10917721B1