Microphone array based directional noise reduction and speech enhancement system
By using a microphone array-based directional noise reduction and speech enhancement system, the noise aliasing problem of traditional omnidirectional acquisition methods is solved, achieving more efficient speech quality improvement, and is suitable for speech acquisition and processing in complex acoustic environments.
Patent Information
- Application Number
- CN202510546514.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Traditional speech enhancement techniques use a fixed omnidirectional acquisition method, which results in a high degree of overlap between the noise spectrum and human voice, leading to poor back-end processing and difficulty in effectively improving speech quality.
A microphone array-based directional noise reduction and speech enhancement system is adopted. By combining an umbrella-shaped acquisition head and a microphone array, audio acquisition and directional noise reduction in a specified direction are achieved, and speech enhancement is performed in conjunction with a processing device.
It improves the relevance and quality of speech signals, effectively suppresses noise, enhances speech enhancement effects, and ensures the clarity and reliability of target speech.
Smart Images

Figure CN120071952B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of voice acquisition and processing, and artificial intelligence, and in particular to a directional noise reduction and voice enhancement system based on a microphone array. Background Technology
[0002] Speech recognition accuracy is the cornerstone of reliability in smart home and in-vehicle interaction scenarios, while speech quality directly affects recognition performance. Traditional speech enhancement technology employs a separate architecture of "fixed omnidirectional acquisition + back-end processing." Fixed omnidirectional acquisition uses a fixed acquisition mode and performs omnidirectional acquisition. In complex acoustic environments, signals acquired in a fixed acquisition mode suffer from noise spectrum mixing with human voices. This omnidirectional acquisition method means that the acquisition device receives sound waves from all directions indiscriminately, including a large amount of noise signals. These noise signals are then mixed with the desired speech signal and acquired. When this audio, with its highly mixed noise spectrum and human voices, is transmitted to the back-end for speech quality enhancement, the back-end processing algorithm faces a significant challenge in improving speech quality, resulting in poor performance. Summary of the Invention
[0003] Based on this, it is necessary to address the technical problem that the existing technology's separate architecture of "fixed omnidirectional acquisition + back-end processing" does not effectively improve speech quality. Therefore, a directional noise reduction and speech enhancement system based on a microphone array is proposed.
[0004] In a first aspect, a directional noise reduction and speech enhancement system based on a microphone array is provided, the system comprising:
[0005] The acquisition device is equipped with a microphone array, which is used to control the microphone array to acquire sound waves to form a first audio frequency based on acquisition control data corresponding to the acquisition direction data.
[0006] The processing device is used to perform directional noise reduction and speech enhancement based on the first audio and the acquisition direction data to obtain the target speech;
[0007] The acquisition device has an umbrella-shaped acquisition head at one end, which includes at least four concave acquisition parts. The concave acquisition parts form a sound wave reflection and focusing structure, and the receiving end of at least one microphone in the microphone array is located within one of the sound wave reflection and focusing structures.
[0008] This application discloses a microphone array-based directional noise reduction and speech enhancement system. First, the acquisition device controls the microphone array to acquire audio signals based on acquisition control data corresponding to the acquisition direction data. This targeted acquisition method avoids the drawback of traditional omnidirectional acquisition, which captures excessive noise. It selectively acquires audio signals from the target direction (main acquisition direction) and reduces noise acquisition from non-target directions (auxiliary acquisition directions). Second, the processing device performs directional noise reduction and speech enhancement based on the first audio signal and the acquisition direction data. Backend operations based on the acquisition direction data significantly improve the noise reduction effect, making noise in the speech more effectively suppressed. Third, the acquisition head at one end of the acquisition device has a recessed acquisition section forming a sound wave reflection and focusing structure, and the receiving end of at least one microphone in the microphone array is located within this structure. This specific configuration of the acquisition section effectively complements the targeted acquisition, further realizing directional acquisition and making the acquired audio signals more directionally targeted. This application, through the acquisition direction data and the special configuration of the umbrella-shaped acquisition head, comprehensively improves the system's directional noise reduction and speech enhancement capabilities, thereby enhancing the quality of the target speech. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] in:
[0011] Figure 1 This is a schematic diagram of a microphone array-based directional noise reduction and speech enhancement system in one embodiment;
[0012] Figure 2 This is a block diagram of a microphone array-based directional noise reduction and speech enhancement system in one embodiment;
[0013] Figure 3 This is a schematic diagram of the acquisition device in one embodiment;
[0014] Figure 4 This is a top view of the data acquisition device in one embodiment;
[0015] Figure 5 This is a schematic diagram of the acquisition unit of the acquisition device in one embodiment.
[0016] The following is a description of the main structure of this application:
[0017] 1. Acquisition device; 11. Handheld part; 12. Umbrella-shaped acquisition head; 121. Acquisition part; 1211. Open end; 1221. Receiver end of the said microphone of type cardioid; 1222. Receiver end of the said microphone of type omnidirectional; 1223. Sound wave reflection and focusing structure; 13. First controller; 14. Second input component; 15. Indicator light; 16. Microphone array; 17. Direction detection component; 18. Camera component; 19. First communication component; 2. Processing device; 21. Second controller; 22. Playback component; 23. Second communication component; 24. First input component; 3. Sound wave. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] The microphone array-based directional noise reduction and speech enhancement system of this application is mainly used in complex scenarios such as conferences, interviews, and choral singing.
[0020] In large conference rooms, background noise is often generated by people moving around and equipment operating (such as air conditioners and projector fans). This system can target the speaker's voice based on their direction, focusing on capturing the speaker's voice, and through targeted noise reduction and voice enhancement, ensure that participants can hear the content clearly and avoid being disturbed by surrounding noise. In remote video conferencing, participants may be in various environments with significant differences in ambient noise. The system can target and optimize the voice of each participant, improving the quality of voice transmission and making remote communication smoother.
[0021] In interview scenarios, whether it is a noisy street, a bustling event, or a special place such as a factory with machine noise, the system can collect the interviewee's voice in a targeted manner, reduce the surrounding environmental noise, ensure that the interview content is clearly audible, and facilitate post-production and dissemination.
[0022] In venues such as choir rehearsal halls or concert halls, although the environment is relatively quiet, there may still be some interfering sounds, such as the audience's slight coughs or the sound of air conditioning running. A microphone array-based directional noise reduction and voice enhancement system can directionally collect the choir's voice, reduce background noise, and enhance the overall sound of the choir, making the harmonies clearer and fuller. This helps the conductor control the choir's performance and provides the audience with a better auditory experience.
[0023] Outdoor choral performances face more noise interference, such as wind and crowd noise. This system can focus on the singers' voices, effectively eliminating external noise, ensuring the purity and audibility of the choral sound, and guaranteeing the quality of the choral performance.
[0024] In this application, audio refers to all audible sound wave physical signals, including speech, music, and environmental noise. Speech specifically refers to semantic sound waves produced by humans through their vocal organs, which carry linguistic information (such as Chinese and English) and have a clear phoneme, intonation, and rhythmic structure.
[0025] Please see Figures 1 to 5 As shown, this application discloses a directional noise reduction and speech enhancement system based on a microphone array, the system comprising:
[0026] Acquisition device 1, wherein the acquisition device 1 is equipped with a microphone array 16, which is used to control the microphone array 16 to acquire sound waves to form a first audio based on acquisition control data corresponding to the acquisition direction data;
[0027] Processing device 2 is used to perform directional noise reduction and speech enhancement based on the first audio and the acquisition direction data to obtain the target speech;
[0028] The acquisition device 1 has an umbrella-shaped acquisition head 12 at one end, which includes at least four concave acquisition portions 121. The concave acquisition portions 121 form a sound wave reflection and focusing structure 1223. At least one microphone in the microphone array 16 ( Figure 3 The receiving end of 1221, 1222 is located within one of the acoustic wave reflection focusing structures 1223.
[0029] The microphone array 16 includes multiple microphones.
[0030] The acquisition device 1 includes a first controller 13, which is the core control component of the acquisition device 1. The first controller 13 is responsible for coordinating the operation of various electronic control components within the acquisition device 1. For example, the first controller 13 can control the opening and closing sequence of each microphone in the microphone array 16 and adjust its sensitivity. Through precise control of these electronic control components, the acquisition device 1 ensures that it acquires sound waves 3 in a predetermined manner to form audio (i.e., the first audio).
[0031] The first controller 13 is implemented through hardware circuitry and software programming. The hardware includes a microprocessor, control circuitry, etc. The microprocessor serves as the core, while the software programming sets the operating parameters of each electronically controlled component, such as defining the microphone's operating mode through code. The control circuitry is responsible for transmitting instructions, enabling the microprocessor to accurately control each component and achieve functions such as sequentially activating the microphones, ensuring that the microphone array 16 efficiently collects sound waves 3.
[0032] The processing device 2 is equipped with a second controller 21, which is used to control the operation of each electronic control component in the processing device 2.
[0033] The second controller 21 is implemented based on the collaboration of hardware and software. The hardware includes a processing chip and a storage unit. The processing chip executes instructions in the software, and the software writes algorithm logic according to processing requirements. The storage unit stores data and programs. It determines the processing flow of each electronic control component through software algorithms. For example, during noise reduction processing, it controls the operation of the components according to the algorithm to achieve effective processing of the input audio.
[0034] The first communication component 19 (i.e., the electronic control component) of the acquisition device 1 and the second communication component 23 (i.e., the electronic control component) of the processing device 2 are connected via wireless communication technology (e.g., Bluetooth) or wired communication technology (e.g., data cable). The first controller 13 is electrically connected to the first communication component 19, and the second controller 21 is electrically connected to the second communication component 23.
[0035] The data acquisition device 1 and the processing device 2 are communicatively connected. The data acquisition device 1 and the processing device 2 can be set up independently or integrated into the same device.
[0036] The acquisition direction data includes: primary acquisition direction and secondary acquisition direction. The number of primary acquisition directions can be one or more, and the number of secondary acquisition directions can be any one of zero, one, or more.
[0037] The acquisition direction data can be preset by the user, input by the user during the acquisition process, or automatically determined by the program implementing this application based on a preset direction determination algorithm.
[0038] The acquisition control data includes: main acquisition data and auxiliary acquisition data. The main acquisition data controls the gain of each microphone corresponding to the main acquisition direction. The auxiliary acquisition data controls whether the microphones corresponding to the auxiliary acquisition direction are turned off or have their noise reduction intensity increased for acquisition.
[0039] The acquisition device 1 determines the acquisition control data based on the acquisition direction data.
[0040] Optionally, the acquisition device 1 determines the acquisition control data using a lookup table method based on the acquisition direction data.
[0041] Optionally, the acquisition device 1 inputs the acquisition direction data into the pre-trained first model for classification and prediction, selects the vector element with the largest value from the predicted vector, and uses the control data corresponding to the selected vector element as the acquisition control data.
[0042] The first pre-trained model is a pre-trained multi-class classification model, and the model structure and training method of the first model can be selected from existing technologies.
[0043] The processing device 2 can determine a noise reduction scheme and a speech enhancement scheme based on the acquired direction data. Based on the determined noise reduction scheme, it performs directional noise reduction on the first audio. Based on the determined speech enhancement scheme, it performs human voice separation and speech enhancement on the directionally denoised audio (that is, the first audio), thereby obtaining high-quality speech, which is then used as the target speech.
[0044] Directional noise reduction involves reducing the noise reduction intensity for useful sound from the desired direction (i.e., the main acquisition direction) to ensure the clarity of the useful sound; while for audio from the unwanted direction (i.e., the auxiliary acquisition direction), the noise reduction intensity is increased according to the settings, and noise is suppressed through signal processing technology, thereby achieving different noise reduction effects in different directions.
[0045] Optionally, the processing device 2 uses a lookup table method to determine a noise reduction scheme based on the acquisition direction data, and uses a lookup table method to determine a speech enhancement scheme based on the acquisition direction data.
[0046] Optionally, the collected direction data is input into a pre-trained second model for classification prediction; from the predicted vectors, the vector element with the largest value is selected from each vector element corresponding to denoising, and the scheme corresponding to the classification category (i.e., scheme identifier) of the selected vector element is taken as the denoising scheme; from the predicted vectors, the vector element with the largest value is selected from each vector element corresponding to speech enhancement, and the scheme corresponding to the classification category (i.e., scheme identifier) of the selected vector element is taken as the speech enhancement scheme.
[0047] The pre-trained second model is a pre-trained multi-class classification model, and the model structure and training method of the second model can be selected from existing technologies.
[0048] A scheme identifier is data that uniquely identifies a scheme, such as the scheme name or scheme ID.
[0049] The acquisition unit 121 is recessed to form a sound wave reflection and focusing structure 1223, so that sound waves in a specific direction range can enter the sound wave 3 reflection and focusing structure 1223, while sound waves 3 outside the specific direction range will be blocked outside the sound wave reflection and focusing structure 1223. This enables sound wave screening in the structural setting of the acquisition unit 121, reducing the difficulty of directional noise reduction for the processing device 2.
[0050] The receiving end of at least one microphone in the microphone array 16 is located within one of the sound wave reflection and focusing structures 1223, so that the receiving end of the microphone within the sound wave reflection and focusing structure 1223 only receives the sound waves within the sound wave reflection and focusing structure 1223, thereby realizing directional acquisition.
[0051] Adjacent acquisition units 121 can be seamlessly connected. Alternatively, adjacent acquisition units 121 can be connected by flexible sound-insulating materials (such as silicone sealing strips) to avoid sound wave interference caused by structural resonance.
[0052] The opening diameter of the umbrella-shaped sampling head 12 can be set according to requirements and is not limited here. The depth of the umbrella-shaped sampling head 12 can be set according to requirements and is not limited here.
[0053] The entire umbrella-shaped acquisition head 12 can be printed using 3D printing technology, or the acquisition part 121 can be printed separately and then assembled into the umbrella-shaped acquisition head 12. Alternatively, the umbrella-shaped acquisition head 12 can be integrally molded using a mold, or the acquisition part 121 can be manufactured separately and then assembled into the umbrella-shaped acquisition head 12. The umbrella-shaped acquisition head 12 printed using 3D printing technology can avoid the manufacturing errors caused by mold manufacturing that affect the directional acquisition effect.
[0054] First, a 3D model of the umbrella-shaped acquisition head 12 is created using 3D modeling software, precisely setting its shape, size, and internal structure parameters to form a model file. Then, the model file is imported into a 3D printer. A suitable printing material, such as durable plastic, is selected. The 3D printer then deposits material layer by layer according to the model file, ultimately printing the physical umbrella-shaped acquisition head 12.
[0055] Optionally, the acquisition unit 121 has a concave structure that is parabolic or spherical (preferably parabolic), which is also the sound wave reflection and focusing structure 1223, ensuring that the sound waves are reflected and focused onto the microphone receiving end. That is, the sound wave reflection and focusing structure 1223 includes a cavity enclosed by an inner wall. The microphone receives the sound waves in this cavity.
[0056] The other end of the data acquisition device 1 can be a handheld part 11 or a support base. The data acquisition device 1 is placed on a flat surface (e.g., the ground or a table) via the support base.
[0057] Understandably, please refer to Figure 1 , Figure 3 , Figure 4 and Figure 5 The illustration shows five collecting units 121, with one collecting unit 121 at the top and four collecting units 121 arranged around the top collecting unit 121. It can be understood that the preferred number of collecting units 121 is four to eight. Besides... Figure 1 , Figure 3 , Figure 4 and Figure 5 The distribution pattern is shown in the diagram. Other distribution patterns for each acquisition unit 121 are not limited here.
[0058] The shape of the opening end 1211 of the collecting part 121 can be circular or other shapes, and is not limited here.
[0059] Sound wave 3 enters the sound wave reflection and focusing structure 1223 from the outside of the acquisition unit 121 through the opening end 1211, and is received by the receiving end of the microphone in the acquisition unit 121 to form audio.
[0060] In this embodiment, firstly, the acquisition device 1 controls the microphone array 16 to acquire audio signals based on the acquisition control data corresponding to the acquisition direction data. This method of selectively acquiring audio signals in a specific direction avoids the drawback of traditional omnidirectional acquisition, which acquires too much noise. It can selectively acquire audio signals from the target direction (main acquisition direction) and reduce noise acquisition from non-target directions (auxiliary acquisition directions). Secondly, the processing device 2 performs directional noise reduction and speech enhancement based on the first audio signal and the acquisition direction data. The backend operation based on the acquisition direction data greatly improves the noise reduction effect, making noise in the speech more effectively suppressed. Furthermore, the acquisition part 121 of the umbrella-shaped acquisition head 12 at one end of the acquisition device 1 is recessed to form a sound wave reflection and focusing structure 1223, and the receiving end of at least one microphone in the microphone array 16 is located within this structure. This specific setting of the acquisition part 121 works well with the selective acquisition in a specific direction, further realizing directional acquisition and making the acquired audio signals more targeted in direction. This application improves the overall directional noise reduction and speech enhancement capabilities of the system and improves the quality of the target speech through the acquisition direction data and the special setting of the umbrella-shaped acquisition head 12.
[0061] Please see Figure 1 , Figure 3 , Figure 4 and Figure 5 In one embodiment, the microphone type includes cardioid and omnidirectional, and a cardioid microphone receiver 1221 is provided at the center of the acoustic wave reflection focusing structure 1223, and at least one omnidirectional microphone receiver 1222 is provided at the edge of the acoustic wave reflection focusing structure 1223.
[0062] It is understood that in the acquisition unit 121, the receiving end of the microphone is located between the opening end 1211 and the center position of the sound wave reflection focusing structure 1223.
[0063] Specifically, the concave structure of the acquisition unit 121 reflects the incident sound waves to the receiving end of the cardioid microphone located at the center.
[0064] Cardioid microphone receivers have unique sound pickup characteristics. They are most sensitive to sound waves directly in front of them, and less sensitive or less sensitive to sound waves outside of that direction, forming a cardioid pickup pattern. When receiving sound waves, they can focus well on sound waves coming from the front, effectively suppressing interference from sound waves from the rear and sides.
[0065] A cardioid microphone typically has a diaphragm in the center of its receiver, surrounded by a specially designed sound-insulating structure. The diaphragm is unobstructed in front, which is beneficial for receiving sound waves directly in front, while the sides and rear have reflective or sound-absorbing designs, making it sensitive to sound waves in front and reducing the reception of sound waves from the sides and rear, thus forming a cardioid pickup pattern.
[0066] An omnidirectional microphone's receiver can receive sound waves from all directions. It has relatively uniform sensitivity to sound waves in all directions, and can receive sound waves from any horizontal angle or from above or below within a certain vertical range. In other words, the receiver of an omnidirectional microphone uses an open structure design.
[0067] The diaphragm of an omnidirectional microphone is usually designed in a relatively central position. The surrounding structure is relatively balanced, without any special obstruction or reinforcement design for any particular direction, so that sound waves can reach the diaphragm evenly from all directions, thus achieving the function of receiving sound waves from all directions.
[0068] Optionally, the inner wall of the sound wave reflection and focusing structure 1223 of the acquisition unit 121 is covered with a sound-absorbing material layer, which is a uniformly distributed sound-absorbing material. The sound-absorbing material layer can receive sound waves that enter the cavity of the sound wave reflection and focusing structure 1223 and propagate to the inner wall.
[0069] Optionally, the sound-absorbing layer can be made of nanofiber composite materials. The high porosity and inter-fiber frictional loss of nanofiber composite materials can absorb high-frequency sound waves. The sound-absorbing layer is used to suppress high-frequency standing waves while avoiding excessive attenuation of the low-frequency energy of speech.
[0070] Optionally, the sound-absorbing material layer can be made of graphene or graphene oxide film. The thickness of the graphene or graphene oxide film can be less than 0.1 mm, utilizing the material's flexibility and internal damping properties to absorb high-frequency vibrations.
[0071] In this embodiment, the sound wave reflection and focusing effect of the concave structure achieves sound wave direction filtering and physical noise reduction at the hardware level. The concave structure accurately reflects sound waves in a specific direction range to the central cardioid microphone to achieve focusing gain, while blocking sound waves outside the specific direction range from entering the sound wave reflection and focusing structure 1223. The signal fusion of the cardioid microphone (main receiver) and the edge omnidirectional microphone (auxiliary noise reference), combined with the spatial selectivity of the sound wave reflection and focusing structure 1223, ultimately achieves hardware-level directional noise reduction and sound source separation, further improving the overall system's directional noise reduction and speech enhancement capabilities, and improving the quality of the target speech.
[0072] Please see Figure 2 In one embodiment, each of the acquisition units 121 is provided with an indicator light 15, and the acquisition device 1 integrates an FPGA-based real-time signal quality assessment module;
[0073] The real-time signal quality assessment module is used for:
[0074] The audio formed by the sound waves collected by the microphone array 16 is acquired at a first time interval and used as the second audio. The real-time signal-to-noise ratio and real-time signal strength of each acquisition unit 121 are calculated based on the second audio.
[0075] Based on a preset dynamic signal-to-noise ratio threshold and a preset dynamic signal strength threshold, the N acquisition units 121 with the highest weighted scores of the real-time signal-to-noise ratio and the real-time signal strength are selected, where 0 < N < P, and P is the total number of acquisition units 121.
[0076] The indicator light 15 corresponding to each of the selected acquisition units 121 is turned on, and the indicator lights 15 of the remaining acquisition units 121 are turned off or dimmed.
[0077] The dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold are adaptively updated according to the ambient noise level of the environment where the acquisition device 1 is located.
[0078] Optionally, the FPGA (Field Programmable Gate Array)-based real-time signal quality assessment module is a circuit module specifically designed for real-time assessment of signal quality.
[0079] Specifically, the audio formed by the latest sound wave collected by the microphone array 16 is acquired at a first time interval and used as the second audio.
[0080] The first controller 13 controls the indicator light 15.
[0081] Indicator light 15 can be set on the inner wall of the sound wave reflection and focusing structure 1223 to illuminate the cavity of the sound wave reflection and focusing structure 1223.
[0082] To extract the average signal strength of the second audio signal as the real-time signal strength, the signal power and noise power in the second audio signal are first determined, and the logarithm of their ratio is used to obtain the signal-to-noise ratio (SNR). This SNR is then used as the real-time SNR. For each acquisition unit 121, its real-time SNR is compared with a dynamic SNR threshold (resulting in an SNR comparison score), and its real-time signal strength is compared with a dynamic signal strength threshold (resulting in an intensity comparison score). The SNR comparison score and the intensity comparison score are then weighted and summed to obtain a comprehensive score. Finally, the comprehensive scores of all acquisition units 121 are sorted, and N acquisition units 121 are selected from highest to lowest. Based on a preset control strategy, the indicator light 15 corresponding to each selected acquisition unit 121 is turned on, while the indicator lights 15 of the remaining unselected acquisition units 121 are turned off or dimmed.
[0083] The audio generated by the sound waves collected by the microphone array 16 is acquired at a second time interval and used as the audio to be analyzed. The short-time energy average of the noise in the audio to be analyzed is calculated and used as the environmental noise reference value. ;
[0084] Dynamic signal-to-noise ratio threshold , As the reference noise, It is a logarithmic function to base 10. , This is an empirical coefficient;
[0085] Dynamic signal strength threshold The value of γ ranges from 1.5 to 2.0, and γ is adaptively adjusted according to noise fluctuations.
[0086] In this embodiment, the signal-to-noise ratio and signal strength of the audio calculation acquisition unit 121 are acquired at a first time interval, and then the N acquisition units 121 with the highest weighted scores are selected. By controlling the on / off state or dimming of the indicator light 15, the status of each acquisition unit 121 is intuitively displayed, making it convenient for users to quickly identify and promptly address any issues. Moreover, the dynamic thresholds (i.e., the dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold) are adaptively updated according to the environmental noise, making the evaluation more consistent with the actual environment and avoiding misjudgments due to fixed thresholds in different environments. This helps to improve the reliability and stability of the audio generated by the entire system's sound waves.
[0087] In one embodiment, the real-time signal quality assessment module is used for:
[0088] The audio formed by the sound waves collected by the microphone array 16 is acquired at a first time interval and used as the second audio. The real-time signal-to-noise ratio and real-time signal strength of each acquisition unit 121 are calculated based on the second audio.
[0089] Based on a preset fixed signal-to-noise ratio threshold and a preset fixed signal strength threshold, the N acquisition units 121 with the highest weighted scores of the real-time signal-to-noise ratio and the real-time signal strength are selected, where 0 < N < P, and P is the total number of acquisition units 121.
[0090] The indicator light 15 corresponding to each of the selected acquisition units 121 is turned on, and the indicator lights 15 of the remaining acquisition units 121 are turned off or dimmed.
[0091] The preset fixed signal-to-noise ratio threshold is a pre-defined threshold. The preset fixed signal strength threshold is also a pre-defined threshold. Using pre-defined thresholds reduces computational resources.
[0092] Please see Figure 2 In one embodiment, the processing device 2 is provided with a first input component 24, and the acquisition device 1 is also provided with a second input component 14, a direction detection component 17 and a camera component 18. The system also includes a remote controller, which is communicatively connected to the acquisition device 1 or the processing device 2. The acquisition device 1 integrates an acquisition control module.
[0093] The acquisition and control module is used for:
[0094] The user sets a primary direction angle using any one of the first input component 24, the second input component 14, and the remote control as the first direction.
[0095] The current pointing direction of the acquisition device 1 detected by the direction detection component 17 is taken as the second direction;
[0096] Based on the image data captured by the camera component 18, the face orientation is determined by a face recognition algorithm based on the image data, and used as the third orientation;
[0097] The main acquisition direction and the auxiliary acquisition direction are determined based on the first direction, the second direction and the third direction, and used as initial direction data;
[0098] Determine whether the acquisition control data needs to be updated based on the deviation between the initial direction data and the acquisition direction data corresponding to the acquisition control data.
[0099] If necessary, control data is determined based on the initial direction data and used as the updated acquisition control data; the acquisition direction data is then updated based on the initial direction data.
[0100] Based on the acquisition control data, the microphone array 16 is controlled to acquire sound waves to form audio, which is used as the first audio.
[0101] The acquisition control data includes: main acquisition data and auxiliary acquisition data.
[0102] The second input component 14 can be provided on the handheld part 11. The orientation detection component 17 can be provided in the receiving cavity inside the handheld part 11.
[0103] The position of the camera component 18 can be set according to requirements. The camera component 18 can capture images of the surrounding environment of the acquisition device 1. The camera component 18 can be mounted on a pan-tilt head to capture images of the surrounding environment of the acquisition device 1. Alternatively, multiple camera components 18 can be mounted on the acquisition device 1 to capture images of the surrounding environment of the acquisition device 1. The second input component 14 and its supporting structure (such as a pan-tilt head) can be selected from existing technologies and will not be described in detail here.
[0104] The first input component 24 can be a button or a touch screen. The second input component 14 can be a button or a touch screen. The first controller 13 controls the operation of the second input component 14, the orientation detection component 17, and the camera component 18, while the second controller 21 controls the operation of the first input component 24.
[0105] The remote control communicates with the acquisition device 1 or the processing device 2 via Bluetooth technology.
[0106] The orientation detection component 17 uses a gyroscope, magnetometer, etc.
[0107] The orientation detection component 17 is used to acquire the spatial orientation data of the acquisition device 1 in real time. Based on the spatial orientation data of the acquisition device 1 and the correspondence between the spatial orientation data and the default acquisition direction, it determines the sound wave source direction of the acquisition unit 121 corresponding to the default acquisition direction and uses this sound wave source direction as the second direction. The orientation detection component 17 typically includes a MEMS gyroscope (measuring angular velocity), an accelerometer (detecting three-dimensional acceleration), and a magnetometer (sensing the direction of the geomagnetic field). It calculates the pitch angle, yaw angle, and roll angle of the acquisition device 1 (accuracy ≤ 1°) using a sensor fusion algorithm (such as Kalman filtering). The orientation detection component 17 outputs orientation data at a frequency of ≥ 100 Hz.
[0108] A MEMS gyroscope is a miniature sensor manufactured based on microelectromechanical systems (MEMS) technology, primarily used to measure rotational angular velocity.
[0109] The default acquisition direction corresponds to the acquisition unit 121, which is intended to acquire the sound waves generated by the speaker's voice. The default acquisition direction can be preset when the acquisition device 1 is set at the factory, or it can be preset by the user according to their own usage habits.
[0110] By analyzing the image data captured in real time by the camera component 18, the face recognition algorithm first locates facial feature points. Then, based on the relative positions of key feature points such as the eyes, nose, and mouth, it calculates the angle of the face relative to the camera, thereby determining the face orientation.
[0111] Optionally, the first direction can be an initially set direction or a real-time set direction. The second direction comes from the direction detection component 17, and the third direction is related to the face. The relationship between the three directions and preset rules is compared. For example, the direction pointing to the face is selected as the primary acquisition direction, and other directions are determined as auxiliary acquisition directions after weighting. If the first direction meets the primary acquisition requirements, it is directly determined, and then auxiliary acquisition directions are supplemented based on the second and third directions. The integrated data is used as the initial direction data.
[0112] Optionally, the vectors (normalized) of the first direction, the second direction, and the third direction are superimposed according to weights using a vector synthesis method, and the direction of the synthesized vector is taken as the main acquisition direction. A preset angle range is symmetrically extended on both sides of the main acquisition direction to form auxiliary acquisition directions.
[0113] Optionally, when the deviation between the first direction, the second direction, and the third direction is greater than 60°, the third direction is preferentially adopted as the main acquisition direction, and a preset angle range is symmetrically extended on both sides of the main acquisition direction to form an auxiliary acquisition direction.
[0114] Optionally, when the deviation between the first direction, the second direction, and the third direction is greater than 60°, the indicator lights 15 can be controlled to illuminate different colors to remind the user to adjust the first or second direction. That is, the indicator light 15 of the acquisition unit 121 corresponding to the first direction needs to display the first color, the indicator light 15 of the acquisition unit 121 corresponding to the second direction needs to display the second color, and the indicator light 15 of the acquisition unit 121 corresponding to the third direction needs to display the third color.
[0115] In this embodiment, firstly, the user can flexibly set the primary direction angle through various input components and a remote control, improving the convenience and versatility of the acquisition direction setting. Secondly, by combining the detection of the direction detection component 17 of the acquisition device 1 itself and the face recognition of the camera component 18 to determine multi-directional data, the determination of the primary and secondary acquisition directions becomes more accurate. Furthermore, the acquisition control data is updated based on the deviation between the initial direction data and existing acquisition direction data, achieving adaptive adjustment of the acquisition control data and avoiding frequent updates while ensuring that the acquisition direction always meets the requirements. Finally, the microphone array 16 is controlled to acquire audio based on the updated acquisition control data, enabling more accurate acquisition of audio from the target directions (primary and secondary acquisition directions), improving the accuracy and effectiveness of audio acquisition.
[0116] Please see Figure 2 In one embodiment, the acquisition device 1 has at least three directional acquisition positions, and the receiving end of an omnidirectional microphone in the microphone array is located at one of the directional acquisition positions;
[0117] The step of determining the main acquisition direction and auxiliary acquisition direction based on the first direction, the second direction, and the third direction in the acquisition control module as initial direction data includes:
[0118] The audio formed by the sound wave is collected by the microphone corresponding to the receiving end in each of the aforementioned direction acquisition positions, and is used as the positioning audio. The location of the sound source is determined based on the positioning audio, which is used as the fourth direction.
[0119] The first direction, the second direction, the third direction, and the fourth direction are used to determine the main acquisition direction and the auxiliary acquisition direction, which are then used as the initial direction data.
[0120] The first direction, the second direction, the third direction, and the fourth direction all use the same reference system (e.g., the device coordinate system).
[0121] The directional acquisition position can be set entirely on the umbrella-shaped acquisition head 12, or partially on the umbrella-shaped acquisition head 12 and partially on the handheld part 11.
[0122] Optionally, the acquisition device 1 has eight directional acquisition positions distributed in a ring around its periphery, with each acquisition position spaced 45° apart, covering a 360° omnidirectional range. Each directional acquisition position is equipped with an omnidirectional microphone, with its receiving end facing the center of the corresponding angle (e.g., 0°, 45°, 90°…315°).
[0123] Optionally, a preset time period is used to control the microphones corresponding to the receiving ends in each of the aforementioned directional acquisition positions to acquire sound waves and form audio, which is then used as the positioning audio.
[0124] The method of determining the location of a sound source by locating each sub-audio in the audio (audio formed by the sound waves collected by the microphone corresponding to the receiving end in a directional acquisition position) can be selected from existing technologies.
[0125] Optionally, the raw audio signal (i.e., the positioning audio) acquired by each omnidirectional microphone is bandpass filtered (85Hz-8kHz), then subjected to frame processing (frame length 20ms, frame shift 10ms), short-time energy and zero-crossing rate are extracted, and environmental noise is initially filtered out to obtain the preprocessed result; the time delay difference of the preprocessed results of adjacent microphone pairs (i.e., the time difference between the two microphones receiving the same sound wave) is calculated using the generalized cross-correlation (GCC-PHAT) algorithm; based on the geometric position of the microphone pair (spacing d=50mm) and the time delay difference, the sound source direction angle is solved by the least squares method; the sound source orientation (also known as the sound source direction) is determined based on the sound source direction angle, and each sound source orientation corresponding to the positioning audio is taken as the fourth direction.
[0126] Optionally, the fourth direction includes at least 0 sound source orientations.
[0127] Optionally, the first direction can be an initially set direction or a real-time set direction. The second direction may come from the direction detection component 17, the third direction is related to the face, and the fourth direction is related to lip movements. The relationship between these four directions and preset rules is compared. For example, the direction pointing to lip movements is prioritized as the primary acquisition direction, and other directions are determined as auxiliary acquisition directions after weighting. If the fourth direction meets the primary acquisition requirements, it is directly determined, and then auxiliary acquisition directions are supplemented based on the first, second, and third directions. The integrated data is then used as the initial direction data.
[0128] Optionally, the vectors (after normalization) of the first direction, the second direction, the third direction, and the fourth direction are superimposed according to weights using a vector synthesis method, and the direction of the synthesized vector is taken as the main acquisition direction. A preset angle range is symmetrically extended on both sides of the main acquisition direction to form auxiliary acquisition directions.
[0129] Optionally, when the deviation between two directions of each sound source direction in the first direction, the second direction, the third direction, and the fourth direction is greater than 60°, the sound source direction in the fourth direction is preferentially adopted as the main acquisition direction, and a preset angle range is symmetrically extended on both sides of the main acquisition direction to form an auxiliary acquisition direction.
[0130] Optionally, when the deviation between two of the first, second, third, and fourth directions is greater than 60°, the indicator lights 15 can be controlled to illuminate different colors to remind the user to adjust the first or second direction. That is, the indicator lights 15 of the acquisition unit 121 corresponding to the first direction need to display the first color, the indicator lights 15 of the acquisition unit 121 corresponding to the second direction need to display the second color, the indicator lights 15 of the acquisition unit 121 corresponding to the third direction need to display the third color, and the indicator lights 15 of the acquisition unit 121 corresponding to the fourth direction need to display the fourth color.
[0131] In this embodiment, the microphone corresponding to the receiving end in the directional acquisition position collects the audio generated by the sound waves to determine the location of the sound source, i.e., the fourth direction. The fourth direction, combined with the first, second, and third directions, is used to determine the acquisition direction data. This helps to more accurately locate the acquisition source and improve acquisition accuracy. It can adapt to the acquisition needs in complex environments, such as when multiple sound sources or interference sources are present. By comprehensively considering multiple directions, it avoids errors due to judgment of a single direction, ensuring that the main acquisition direction is aligned with the speaker, and the auxiliary acquisition direction is effectively supplemented, thus optimizing the overall performance of the acquisition device 1, better acquiring relevant data of the speaker, and improving acquisition efficiency and quality.
[0132] Please see Figure 2 In one embodiment, the processing device 2 integrates an aging compensation module, which is used for:
[0133] Obtain the compensation signal;
[0134] In response to the compensation signal, preset control data is acquired as debugging control data, and the acquisition control data of the acquisition device 1 is updated according to the debugging control data;
[0135] The processing device 2 is controlled to play a preset standard test audio, and the acquisition device 1 acquires the response audio corresponding to the standard test audio.
[0136] The frequency response curves of the response audio and the standard test audio are compared to obtain frequency response difference data.
[0137] Based on the frequency response difference data, aging compensation data is generated;
[0138] The step of determining control data based on the initial direction data in the acquisition and control module, as the updated acquisition and control data, includes:
[0139] Control data is determined based on the initial direction data and used as the initial control data.
[0140] Based on the aging compensation data, the initial control data is corrected to obtain the updated acquisition control data.
[0141] The compensation signal is the signal that initiates and determines the aging compensation data.
[0142] The user can trigger the compensation signal through any one of the first input component 24, the second input component 14, and the remote control, or the compensation signal can be automatically triggered by the system based on preset time data.
[0143] Specifically, firstly, the response audio and the standard test audio are converted into frequency domain signals, and the corresponding frequency response curves for each are obtained. For each frequency point, the difference between the amplitude value of the frequency response curve of the response audio and the amplitude value of the frequency response curve of the standard test audio is calculated. These differences constitute the frequency response difference data. The audio signals (i.e., the response audio and the standard test audio) can be converted into frequency domain signals using Discrete Fourier Transform (DFT) or Fast Fourier Transform (FFT). Then, at the same frequency scale, the amplitudes of the two are compared point by point to accurately obtain the frequency response difference data, thereby reflecting the changes in the audio acquisition characteristics of the acquisition device 1 due to factors such as aging.
[0144] The playback component 22 in the processing device 2 is controlled to play a preset standard test audio. The second controller 21 controls the operation of the playback component 22. The playback component 22 may be a speaker.
[0145] Optionally, aging compensation data can be determined using a lookup table method based on the frequency response difference data.
[0146] Optionally, the frequency response difference data is input into a pre-trained third model for classification and prediction. The vector element with the largest value is extracted from the predicted vector, and the compensation data corresponding to the extracted vector element is used as aging compensation data.
[0147] Aging compensation data can be determined based on limited experiments, and will not be elaborated here.
[0148] The pre-trained third model is a pre-trained multi-class classification model. The model structure and training method of the third model can be selected from existing technologies.
[0149] Based on a preset correction formula, the initial control data is corrected according to the aging compensation data to become the updated acquisition control data.
[0150] The preset correction formula can be obtained by fitting multiple experimental data, which will not be elaborated here.
[0151] In this embodiment, the aging compensation module updates the acquisition control data by acquiring compensation signals, ensuring that the acquisition device 1 adapts to aging conditions. Playing standard test audio and comparing the response audio yields frequency response difference data, which in turn generates aging compensation data. This helps to accurately compensate for performance deviations in the acquisition device 1 caused by aging. When determining the acquisition control data, the aging compensation data is integrated with the initial direction data, and the initial control data is also corrected, improving the accuracy and reliability of the acquisition. This ensures that the acquisition device 1 maintains stable performance during long-term use and reduces the impact of aging on the acquisition effect.
[0152] Please see Figure 2In one embodiment, the processing device 2 integrates a directional noise reduction and voice enhancement module;
[0153] The directional noise reduction and speech enhancement module is used for:
[0154] Based on the collected directional data, a directional noise reduction scheme is determined;
[0155] According to the directional noise reduction scheme, the first audio is subjected to directional noise reduction to obtain the third audio.
[0156] The third audio is subjected to voice separation and speech enhancement to obtain the target speech.
[0157] Specifically, based on the collected directional data, a lookup table method is used to determine the directional noise reduction scheme.
[0158] A preset voice separation method is used to separate the voices in the third audio file. A preset speech enhancement method is used to enhance the speech data obtained from the voice separation. The enhanced speech is then used as the target speech.
[0159] The directional noise reduction scheme includes the following types of data: acquisition direction data and noise reduction control parameters corresponding to each acquisition direction in the acquisition direction data.
[0160] In this embodiment, by collecting directional data to determine the directional noise reduction scheme, the first audio can be accurately denoised to obtain the third audio, effectively reducing noise interference outside the main acquisition direction, making the noise in the speech more effectively suppressed, and improving the quality of the target speech.
[0161] In one embodiment, the step of performing directional noise reduction on the first audio to obtain a third audio in the directional noise reduction and speech enhancement module according to the directional noise reduction scheme includes:
[0162] Using a multi-stage NLMS filter, the first audio is subjected to directional noise reduction according to the directional noise reduction scheme to obtain the third audio.
[0163] The multi-stage NLMS filter comprises three parallel target filters and a fusion unit. The target filters sequentially include a Butterworth bandpass filter and an NLMS filter. The fusion unit is used to perform weighted fusion of the outputs of each of the target filters.
[0164] The target filter in the first stage processes frequency band P1, where 0Hz < P1 ≤ the value of the first frequency band, and the step size factor is calculated according to the formula. Dynamic adjustment and It is an empirical constant, and its It is the signal-to-noise ratio of the data input to the target filter in the first frequency band from 0Hz;
[0165] The target filter in the second stage processes frequency band P2, where the value of the first frequency band < P2 ≤ the value of the second frequency band, and the step size factor is calculated according to the formula. Dynamic adjustment and It is an empirical constant, and its It is the signal-to-noise ratio of the input data to the target filter between the first frequency band value and the second frequency band value;
[0166] The target filter in the third stage processes frequency band P3, where the value of the second frequency band < P3 ≤ the value of the third frequency band, and the step size factor is calculated according to the formula. Dynamic adjustment and It is an empirical constant, and its It is the signal-to-noise ratio of the input data to the target filter in the second frequency band to the third frequency band.
[0167] The specific values for the first, second, and third frequency bands can be set according to requirements and are not limited here.
[0168] Optionally, the first frequency band value is set to 1kHz, the second frequency band value is set to 4kHz, and the third frequency band value is set to 8kHz.
[0169] Specifically, according to the directional noise reduction scheme, the noise reduction control parameters of each sub-audio in the first audio are determined. When the multi-level NLMS filter processes the sub-audio, the NLMS filter adopts the noise reduction control parameters corresponding to the sub-audio.
[0170] Butterworth bandpass filters are primarily used to select a specific frequency range, allowing signals within that band to pass while suppressing signals outside that band. In this multi-stage NLMS filter system, it serves as the front-end of the target filter, performing initial frequency filtering on the input audio signal within the corresponding target frequency band (e.g., 0Hz~1kHz, 1~4kHz, 4~8kHz), enabling subsequent NLMS filters to perform better noise reduction based on the filtered frequency band signal.
[0171] NLMS (Normalized Least Mean Square) filters are used to adaptively adjust the filter coefficients to minimize the error between the desired signal and the filter output. In this system, it receives the signal processed by a Butterworth bandpass filter and dynamically adjusts its coefficients according to the statistical characteristics of the input signal, thereby effectively removing noise components in the corresponding frequency band, improving the quality of the speech signal in that band, and ultimately achieving directional noise reduction of the first audio signal in different frequency bands.
[0172] The target filter of the first level mainly deals with environmental noise and uses a large step size factor to converge quickly.
[0173] The second-level target filter enhances speech clarity (the main energy region of human voice) with a moderate step size factor.
[0174] The third-level target filter suppresses sharp noise (such as metallic clanging) and uses a smaller step size to avoid distortion.
[0175] The fusion unit first receives the outputs of the NLMS filters from the three target filters. Then, based on pre-set weighting coefficients (which can be set according to noise reduction requirements and the importance of different frequency bands), the outputs of each target filter are weighted, and the weighted results are summed to obtain the fused output (i.e., the third audio signal), thus achieving effective fusion of noise reduction results from different frequency bands.
[0176] The step size factor of the target filter clarifies the negative correlation between step size and signal-to-noise ratio, which is beneficial for reducing the convergence speed in high-noise frequency bands and avoiding signal distortion.
[0177] and It can be obtained through limited experiments. and It can be obtained through limited experiments. and It can be obtained through limited experiments.
[0178] In this embodiment, a multi-stage NLMS filter is used for directional noise reduction, enabling precise processing of different frequency bands. Different target filters process different frequency bands, such as 0Hz to the first band value, the first band value to the second band value, and the second band value to the third band value. The step size factor is dynamically adjusted based on the signal-to-noise ratio of each band, effectively adapting to different input audio conditions. By weighted fusing the outputs of each filter stage through a fusion unit, the noise reduction advantages of different frequency bands can be combined, resulting in more precise noise suppression and enhanced target speech. This improves speech quality during directional noise reduction, yielding a purer and more indistinguishable third audio signal.
[0179] In one embodiment, the step of performing voice separation and speech enhancement on the third audio to obtain the target speech in the directional noise reduction and speech enhancement module includes:
[0180] Obtain hearing preference parameters, wherein the hearing preference parameters include: sensitive frequency band and volume tolerance threshold;
[0181] Based on the hearing preference parameters and the acquisition direction data, a speech enhancement scheme is determined;
[0182] According to the speech enhancement scheme, the third audio is subjected to voice separation and speech enhancement to obtain the target speech.
[0183] The listening preference parameters are data that users input and store in advance.
[0184] Optionally, a speech enhancement scheme can be determined using a lookup table method based on the hearing preference parameters and the acquisition direction data.
[0185] Optionally, the listening preference parameters and the acquisition direction data are concatenated and input into a pre-trained fourth model for classification and prediction. The vector element with the largest value is extracted from the predicted vector, and the scheme corresponding to the extracted vector element is used as the speech enhancement scheme.
[0186] Hearing preference parameters are obtained through the user input interface or preset configuration files.
[0187] The pre-trained fourth model is a pre-trained multi-class classification model. The model structure and training method of the fourth model can be selected from existing technologies.
[0188] The speech enhancement scheme includes a primary speech enhancement scheme in the main direction and a secondary speech enhancement scheme in the main direction. The primary speech enhancement scheme in the main direction is used to enhance the speech in the main acquisition direction after voice separation to obtain the first speech. The secondary speech enhancement scheme in the main direction is used to correct the first speech in the secondary acquisition direction after voice separation to obtain the target speech.
[0189] In this embodiment, firstly, by introducing hearing preference parameters, a customized speech enhancement scheme can be developed based on an individual's sensitive frequency bands and volume tolerance threshold, improving the personalization of the solution. Secondly, by combining the collected directional data, speech from specific directions can be processed more accurately. During voice separation and speech enhancement, the desired speech can be effectively enhanced, improving clarity and intelligibility. For users with special hearing needs, such as those sensitive to certain frequency bands or with limited volume tolerance, target speech that better matches their hearing characteristics can be obtained, improving the user's speech experience.
[0190] In one embodiment, the outer surface of the other end of the acquisition device 1 is coated with a conductive coating; the conductive coating is connected to the ground wire of the microphone array 16 to form an electromagnetic shielding loop.
[0191] Specifically, a conductive coating is applied to the outer surface of the other end of the acquisition device 1 using a spraying process.
[0192] Optionally, in an environment of 25°C and 50% humidity, the surface resistivity of the conductive coating is ≤5Ω / cm².
[0193] In this embodiment, the conductive coating is made of graphene composite material with low surface resistivity, which can effectively conduct current. After the conductive coating is connected to the ground wire of the microphone array 16 to form an electromagnetic shielding loop, it can play a good electromagnetic shielding role. This technology can reduce the influence of external electromagnetic interference on the acquisition device 1, ensure the purity of the acquired audio signal, improve the accuracy and stability of audio acquisition, and enhance the overall audio acquisition quality.
[0194] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0195] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A directional noise reduction and speech enhancement system based on a microphone array, characterized in that, The system includes: The acquisition device is equipped with a microphone array, which is used to control the microphone array to acquire sound waves to form a first audio frequency based on acquisition control data corresponding to the acquisition direction data. The processing device is used to perform directional noise reduction and speech enhancement based on the first audio and the acquisition direction data to obtain the target speech; The acquisition device has an umbrella-shaped acquisition head at one end, which includes at least four concave acquisition parts. The concave acquisition parts form a sound wave reflection and focusing structure. The receiving end of at least one microphone in the microphone array is located within one of the sound wave reflection and focusing structures. The processing device is provided with a first input component, and the acquisition device is also provided with a second input component, a direction detection component, and a camera component. The system also includes a remote controller, which is communicatively connected to the acquisition device or the processing device. The acquisition device integrates an acquisition control module. The acquisition and control module is used for: The first direction is obtained by the user setting a main direction angle through any one of the first input component, the second input component, and the remote control. The current pointing direction of the acquisition device detected by the direction detection component is taken as the second direction; Based on the image data captured by the camera component, the face orientation is determined by a face recognition algorithm based on the image data, and used as the third orientation; The main acquisition direction and the auxiliary acquisition direction are determined based on the first direction, the second direction and the third direction, and used as initial direction data; Determine whether the acquisition control data needs to be updated based on the deviation between the initial direction data and the acquisition direction data corresponding to the acquisition control data. If necessary, control data is determined based on the initial direction data and used as the updated acquisition control data; the acquisition direction data is then updated based on the initial direction data. Based on the acquisition control data, the microphone array is controlled to acquire sound waves to form audio, which is used as the first audio. The acquisition control data includes: main acquisition data and auxiliary acquisition data.
2. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, The microphone types include cardioid and omnidirectional. The center of the acoustic wave reflection and focusing structure is provided with a receiving end of the cardioid microphone, and the edge of the acoustic wave reflection and focusing structure is provided with at least one receiving end of the omnidirectional microphone.
3. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, Each of the acquisition units is equipped with an indicator light, and the acquisition device integrates an FPGA-based real-time signal quality assessment module; The real-time signal quality assessment module is used for: The audio generated by the sound waves collected by the microphone array is acquired at a first time interval and used as the second audio. The real-time signal-to-noise ratio and real-time signal strength of each acquisition unit are calculated based on the second audio. Based on preset dynamic signal-to-noise ratio thresholds and preset dynamic signal strength thresholds, the N acquisition units with the highest weighted scores of real-time signal-to-noise ratio and real-time signal strength are selected, where 0 < N < P, and P is the total number of acquisition units; the indicator lights corresponding to each of the selected acquisition units are turned on, and the indicator lights of the remaining acquisition units are turned off or dimmed. The dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold are adaptively updated according to the ambient noise level of the environment where the acquisition device is located.
4. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, The acquisition device has at least three directional acquisition positions, and the receiving end of one of the omnidirectional microphones in the microphone array is located at one of the directional acquisition positions; The step of determining the main acquisition direction and auxiliary acquisition direction based on the first direction, the second direction, and the third direction in the acquisition control module as initial direction data includes: The audio formed by the sound wave is collected by the microphone corresponding to the receiving end in each of the aforementioned direction acquisition positions, and is used as the positioning audio. The location of the sound source is determined based on the positioning audio, which is used as the fourth direction. The first direction, the second direction, the third direction, and the fourth direction are used to determine the main acquisition direction and the auxiliary acquisition direction, which are then used as the initial direction data.
5. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, The processing device integrates an aging compensation module, which is used for: Obtain the compensation signal; In response to the compensation signal, preset control data is acquired as debugging control data, and the acquisition control data of the acquisition device is updated according to the debugging control data; The processing device is controlled to play a preset standard test audio, and the acquisition device acquires the response audio corresponding to the standard test audio. The frequency response curves of the response audio and the standard test audio are compared to obtain frequency response difference data. Based on the frequency response difference data, aging compensation data is generated; The step of determining control data based on the initial direction data in the acquisition and control module, as the updated acquisition and control data, includes: Control data is determined based on the initial direction data and used as the initial control data. Based on the aging compensation data, the initial control data is corrected to obtain the updated acquisition control data.
6. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, The processing device integrates a directional noise reduction and voice enhancement module; The directional noise reduction and speech enhancement module is used for: Based on the collected directional data, a directional noise reduction scheme is determined; According to the directional noise reduction scheme, the first audio is subjected to directional noise reduction to obtain the third audio. The third audio is subjected to voice separation and speech enhancement to obtain the target speech.
7. The directional noise reduction and speech enhancement system based on a microphone array according to claim 6, characterized in that, The step of performing directional noise reduction on the first audio to obtain the third audio in the directional noise reduction and speech enhancement module according to the directional noise reduction scheme includes: Using a multi-stage NLMS filter, the first audio is subjected to directional noise reduction according to the directional noise reduction scheme to obtain the third audio. The multi-stage NLMS filter comprises three parallel target filters and a fusion unit. The target filters sequentially include a Butterworth bandpass filter and an NLMS filter. The fusion unit is used to perform weighted fusion of the outputs of each of the target filters. The target filter in the first stage processes the frequency band P1, where 0Hz < P1 ≤ the value of the first frequency band. The step size factor is dynamically adjusted according to the formula μ_1 = α_1 / (SNR_1 + β_1), where α_1 and β_1 are empirical constants, and SNR_1 is the signal-to-noise ratio of the data input to the target filter in the range of 0Hz to the value of the first frequency band. The target filter in the second stage processes frequency band P2, where the value of the first frequency band < P2 ≤ the value of the second frequency band. The step size factor is dynamically adjusted according to the formula μ_2=α_2 / (SNR_2+β_2), where α_2 and β_2 are empirical constants, and SNR_2 is the signal-to-noise ratio of the data input to the target filter from the value of the first frequency band to the value of the second frequency band. The target filter in the third stage processes frequency band P3, where the second frequency band value < P3 ≤ the third frequency band value. The step size factor is dynamically adjusted according to the formula μ_3=α_3 / (SNR_3+β_3), where α_3 and β_3 are empirical constants, and SNR_3 is the signal-to-noise ratio of the data input to the target filter between the second and third frequency band values.
8. The directional noise reduction and speech enhancement system based on a microphone array according to claim 6, characterized in that, The directional noise reduction and speech enhancement module's step of performing voice separation and speech enhancement on the third audio to obtain the target speech includes: Obtain hearing preference parameters, wherein the hearing preference parameters include: sensitive frequency band and volume tolerance threshold; Based on the hearing preference parameters and the acquisition direction data, a speech enhancement scheme is determined; According to the speech enhancement scheme, the third audio is subjected to voice separation and speech enhancement to obtain the target speech.
9. The directional noise reduction and speech enhancement system based on a microphone array according to claim 1, characterized in that, The outer surface of the other end of the acquisition device is coated with a conductive coating; the conductive coating is connected to the ground wire of the microphone array to form an electromagnetic shielding circuit.
Citation Information
Patent Citations
Multi-microphone array beamforming signal enhancement method and device
CN119811408A
Acoustic-thermal detection device for high-voltage equipment
CN214951816U