Directional noise reduction and speech enhancement system based on microphone array

Through the directional noise reduction and voice enhancement system based on the microphone array, the microphone array is controlled to collect sound waves using the acquisition direction data, and directional noise reduction and voice enhancement are performed in the processing device, which solves the problem of high aliasing of noise spectrum and vocals in traditional technology, and achieves more efficient voice quality improvement.

CN120071952AActive Publication Date: 2025-05-30GUANGZHOU MEILU ELECTRONICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510546514.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The noise spectrum and vocals are highly aliased by the signals collected by traditional speech enhancement technology in complex acoustic environments, resulting in poor results in improving speech quality by back-end processing.

Method used

The directional noise reduction and voice enhancement system based on the microphone array is adopted to control the microphone array to collect sound waves by collecting direction data to reduce noise acquisition in non-target directions, and directional noise reduction and voice enhancement are performed based on the acquisition direction data in the processing device.

Benefits of technology

It effectively avoids the noise problem of traditional omnidirectional acquisition, improves the voice quality, and enhances the clarity and reliability of the target voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071952A_ABST
    Figure CN120071952A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice acquisition and processing, and discloses a directional noise reduction and voice enhancement system based on a microphone array, and the system comprises an acquisition device which is provided with the microphone array and is used for acquiring control data corresponding to acquisition direction data based on the acquisition control data corresponding to the acquisition direction data; controlling the microphone array to collect sound waves to form a first audio frequency; the processing device is used for performing directional noise reduction and voice enhancement according to the first audio and the acquisition direction data to obtain target voice; wherein one end of the collection device is provided with an umbrella-shaped collection head, the umbrella-shaped collection head comprises at least four concave collection parts, the collection parts are concave to form sound wave reflection focusing structures, and the receiving end of at least one microphone in the microphone array is located in one sound wave reflection focusing structure. Through the collection of direction data and the special arrangement of the umbrella-shaped collection head, the directional noise reduction and speech enhancement capabilities of the system are integrally improved, and the quality of target speech is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of voice acquisition and processing, and artificial intelligence, and in particular, to a directional noise reduction and voice enhancement system based on a microphone array. Background Art

[0002] The accuracy of speech recognition is the reliability cornerstone in scenarios such as smart homes and in-vehicle interactions, and the voice quality directly affects the recognition performance. Traditional voice enhancement technologies adopt a separate architecture of "fixed omnidirectional acquisition + backend processing". Fixed omnidirectional acquisition uses a fixed acquisition mode and conducts omnidirectional acquisition. In a complex acoustic environment, the signals collected in the fixed acquisition mode have a high degree of overlap between the noise spectrum and the human voice. This omnidirectional acquisition method means that the acquisition device will receive sound waves from all directions without discrimination, including a large amount of noise signals, and the noise signals are mixed with the desired voice signals and collected. When the audio with a high degree of overlap between the noise spectrum and the human voice and containing a large amount of noise signals is transmitted to the backend for voice quality improvement processing, the backend processing algorithm faces great challenges in improving the voice quality, resulting in poor effects of the backend processing in improving the voice quality. Summary of the Invention

[0003] Based on this, in view of the technical problem that the existing technology adopts a separate architecture of "fixed omnidirectional acquisition + backend processing" and has poor effects in improving the voice quality, a directional noise reduction and voice enhancement system based on a microphone array is proposed.

[0004] In a first aspect, a directional noise reduction and voice enhancement system based on a microphone array is provided. The system includes: An acquisition device, which is provided with a microphone array and is used to control the microphone array to collect sound waves to form a first audio based on acquisition control data corresponding to acquisition direction data; A processing device, which is used to perform directional noise reduction and voice enhancement according to the first audio and the acquisition direction data to obtain a target voice; Wherein, one end of the acquisition device is provided with an umbrella-shaped acquisition head, the umbrella-shaped acquisition head includes at least four concave acquisition parts, the acquisition parts are concave to form a sound wave reflection focusing structure, and the receiving end of at least one microphone in the microphone array is located in one of the sound wave reflection focusing structures.

[0005] The directional noise reduction and speech enhancement system based on a microphone array of the present application first has the acquisition device control the microphone array to acquire based on the acquisition control data corresponding to the acquisition direction data. This method of focusing on acquisition in a specified direction avoids the drawback of excessive noise acquisition in traditional omnidirectional acquisition, can selectively obtain the audio signal in the target direction (main acquisition direction), and reduce the noise acquisition in the non-target direction (auxiliary acquisition direction). Secondly, the processing device performs directional noise reduction and speech enhancement based on the first audio and the acquisition direction data. The operation based on the acquisition direction data at the backend greatly improves the noise reduction effect, enabling more effective suppression of the noise in the speech. Furthermore, the acquisition part of the umbrella-shaped acquisition head at one end of the acquisition device is recessed to form a sound wave reflection focusing structure, and the receiving end of at least one microphone in the microphone array is located within this structure. This specific setting of the acquisition part well cooperates with the focused acquisition in the specified direction, further realizing directional acquisition and making the acquired audio signal more targeted in terms of direction. Through the acquisition direction data and the special setting of the umbrella-shaped acquisition head, the present application overall improves the ability of the system for directional noise reduction and speech enhancement, and improves the quality of the target speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0007] Among them: Figure 1 FIG. [X] is a schematic structural diagram of a directional noise reduction and speech enhancement system based on a microphone array in an embodiment; Figure 2 FIG. [X] is a structural block diagram of a directional noise reduction and speech enhancement system based on a microphone array in an embodiment; Figure 3 FIG. [X] is a schematic structural diagram of an acquisition device in an embodiment; Figure 4 FIG. [X] is a top view schematic diagram of an acquisition device in an embodiment; Figure 5 FIG. [X] is a schematic structural diagram of the acquisition part of an acquisition device in an embodiment.

[0008] The following is an explanation of the main structures of the present application: 1. Acquisition device; 11. Handheld part; 12. Umbrella-shaped acquisition head; 121. Acquisition part; 1211. Open end; 1221. Receiving end of the microphone of heart shape type; 1222. Receiving end of the omnidirectional microphone; 1223. Acoustic wave reflection focusing structure; 13. First controller; 14. Second input component; 15. Indicator light; 16. Microphone array; 17. Direction detection component; 18. Imaging component; 19. First communication component; 2. Processing device; 21. Second controller; 22. Playback component; 23. Second communication component; 24. First input component; 3. Acoustic wave. Detailed implementation manners

[0009] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0010] The directional noise reduction and speech enhancement system based on microphone array of the present application is mainly applied to complex scenarios such as conference scenarios, interview scenarios, and chorus scenarios.

[0011] In a large conference room, there are often background noises generated by people walking around and equipment operation (such as air conditioners, projector fans, etc.). This system can perform directional acquisition according to the direction of the speaker, focus on collecting the voice of the speaker, and through directional noise reduction and speech enhancement, enable the participants to clearly hear the speech content and avoid being interfered by the surrounding noises. In a remote video conference, the participants may be in various different environments with large differences in environmental noises. The system can perform directional acquisition and optimization processing on the speech of each participant, improve the quality of speech transmission, and make the communication in the remote conference smoother.

[0012] In an interview scenario, whether it is a news interview in a noisy street, a lively event site, or a special place such as a factory with machine noises, this system can directionally collect the voice of the interviewee, reduce the surrounding environmental noises, ensure that the interview content is clearly audible, and facilitate post-production and dissemination.

[0013] In places such as a chorus rehearsal hall or a concert hall, although the environment is relatively quiet, there may still be some interfering sounds, such as the slight coughing sounds of the audience and the operation sounds of air conditioners. The directional noise reduction and speech enhancement system based on microphone array can directionally collect the voice of the chorus team, reduce the background noise, and at the same time enhance the overall effect of the chorus sound, making the chorus harmony clearer and fuller, which helps the conductor to control the chorus effect and also enables the audience to obtain a better auditory experience.

[0014] During outdoor choral performances, there are more noise interferences, such as wind noise, crowd noise, etc. This system can be directionally focused on the voices of the choristers, effectively excluding external noises, ensuring the purity and audibility of the choral sound, and guaranteeing the quality of the choral performance.

[0015] In this application, audio refers to all physical signals of audible sound waves in general, including speech, music, environmental noise, etc. Speech refers to the semantic sound waves specifically generated by the human vocal organs, which corely carry language information (such as Chinese, English), and have clear phonemes, intonations, and rhythm structures.

[0016] Please refer to Figures 1 to 5 As shown, this application discloses a directional noise reduction and speech enhancement system based on a microphone array. The system includes: A collection device 1, which is provided with a microphone array 16, and is used to control the microphone array 16 to collect sound waves to form a first audio based on the collection control data corresponding to the collection direction data; A processing device 2, which is used to perform directional noise reduction and speech enhancement according to the first audio and the collection direction data to obtain the target speech; Among them, one end of the collection device 1 is provided with an umbrella-shaped collection head 12. The umbrella-shaped collection head 12 includes at least four concave collection parts 121. The collection part 121 is concave to form a sound wave reflection focusing structure 1223. At least one microphone ( Figure 3 1221, 1222 in it) in the microphone array 16 has its receiving end located within one of the sound wave reflection focusing structures 1223.

[0017] The microphone array 16 includes multiple microphones.

[0018] The collection device 1 is provided with a first controller 13. The first controller 13 is the core control component of the collection device 1. The first controller 13 is responsible for coordinating the operation of each electronic control component in the collection device 1. For example, the first controller 13 can control the on-off sequence of each microphone in the microphone array 16 and adjust its sensitivity, etc. Through precise control of these electronic control components, it is ensured that the collection device 1 collects sound waves 3 in a predetermined manner to form an audio (that is, the first audio).

[0019] The first controller 13 is realized through hardware circuits and software programming. The hardware includes a microprocessor, a control circuit, etc. The microprocessor is the core, and the software programming sets the working parameters of each electronic control component, such as defining the working mode of the microphone through code. The control circuit is responsible for transmitting instructions, enabling the microprocessor to accurately control each component, and realizing functions such as activating the microphones in sequence, ensuring that the microphone array 16 efficiently collects sound waves 3.

[0020] The processing device 2 is provided with a second controller 21, and the second controller 21 is used to control the operation of each electrically controlled component in the processing device 2.

[0021] The implementation of the second controller 21 is based on the cooperation of hardware and software. The hardware includes a processing chip, a storage unit, etc. The processing chip executes the instructions in the software, and the software writes algorithm logic according to the processing requirements. The storage unit stores data and programs. It determines the processing flow of each electrically controlled component through software algorithms. For example, during noise reduction processing, it controls the operation of components according to the algorithm to achieve effective processing of the input audio.

[0022] The first communication component 19 (which is an electrically controlled component) of the acquisition device 1 and the second communication component 23 (which is an electrically controlled component) of the processing device 2 are communicatively connected through wireless communication technology (such as Bluetooth) or wired communication technology (such as a data cable). The first controller 13 is electrically connected to the first communication component 19, and the second controller 21 is electrically connected to the second communication component 23.

[0023] The acquisition device 1 is communicatively connected to the processing device 2. The acquisition device 1 and the processing device 2 can be set independently or integrated in the same device.

[0024] The acquisition direction data includes: a main acquisition direction and an auxiliary acquisition direction. The number of main acquisition directions is one or more, and the number of auxiliary acquisition directions is any one of 0, 1, or multiple.

[0025] The acquisition direction data can be pre-set by the user, can also be input by the user during the acquisition process, or can be automatically determined by the program implementing the present application according to the preset direction determination algorithm.

[0026] The acquisition control data includes: main acquisition data and auxiliary acquisition data. The main acquisition data controls the gain of each microphone corresponding to the main acquisition direction. The auxiliary acquisition data controls the microphone corresponding to the auxiliary acquisition direction to be turned off or to increase the noise reduction strength for acquisition.

[0027] The acquisition device 1 determines the acquisition control data according to the acquisition direction data.

[0028] Optionally, the acquisition device 1 determines the acquisition control data by using a look-up table method according to the acquisition direction data.

[0029] Optionally, the acquisition device 1 inputs the acquisition direction data into a pre-trained first model for classification prediction, selects the vector element with the maximum value from the predicted vectors, and uses the control data corresponding to the selected vector element as the acquisition control data.

[0030] The pre-trained first model is a pre-trained multi-classification model, and the model structure and model training method of the first model can be selected from the prior art.

[0031] The processing device 2 can determine a noise reduction scheme and a voice enhancement scheme according to the collected direction data, perform directional noise reduction on the first audio based on the determined noise reduction scheme, and perform voice separation and voice enhancement on the audio after directional noise reduction (i.e., the first audio) based on the determined voice enhancement scheme, so as to obtain high-quality voice, and use this voice as the target voice.

[0032] Directional noise reduction means that for the useful sound in the desired direction (i.e., the main collection direction), the noise reduction intensity is reduced to ensure the clarity of the useful sound; while for the audio in the unwanted direction (i.e., the auxiliary collection direction), according to the set increased noise reduction intensity, the noise is suppressed through signal processing technology, so as to achieve different noise reduction effects in different directions.

[0033] Optionally, the processing device 2 determines the noise reduction scheme by using a look-up table method according to the collected direction data, and determines the voice enhancement scheme by using a look-up table method according to the collected direction data.

[0034] Optionally, the collected direction data is input into a pre-trained second model for classification prediction; from each vector element corresponding to noise reduction in the predicted vector, the vector element with the largest value is selected, and the scheme corresponding to the classification category (i.e., the scheme identifier) corresponding to the selected vector element is used as the noise reduction scheme; from each vector element corresponding to voice enhancement in the predicted vector, the vector element with the largest value is selected, and the scheme corresponding to the classification category (i.e., the scheme identifier) corresponding to the selected vector element is used as the voice enhancement scheme.

[0035] The pre-trained second model is a pre-trained multi-classification model, and the model structure and model training method of the second model can be selected from the prior art.

[0036] The scheme identifier is data such as a scheme name, a scheme ID, etc. that uniquely identifies a scheme.

[0037] The collection part 121 is recessed to form a sound wave reflection focusing structure 1223, so that sound waves within a specific direction range can enter the sound wave 3 reflection focusing structure 1223, and sound waves 3 outside the specific direction range will be blocked outside the sound wave reflection focusing structure 1223, thereby realizing sound wave screening in the structural setting of the collection part 121 and reducing the difficulty of directional noise reduction for the processing device 2.

[0038] The receiving end of at least one microphone in the microphone array 16 is located within one of the sound wave reflection focusing structures 1223, so that the receiving end of the microphone within the sound wave reflection focusing structure 1223 only receives the sound waves within the sound wave reflection focusing structure 1223, thereby realizing directional collection.

[0039] The adjacent acquisition parts 121 can be seamlessly connected. A flexible sound insulation material (such as a silicone sealing strip) can also be provided between the adjacent acquisition parts 121 for connection, so as to avoid acoustic interference caused by structural resonance.

[0040] The opening diameter of the umbrella-shaped acquisition head 12 can be set according to requirements and is not limited herein. The depth of the umbrella-shaped acquisition head 12 can be set according to requirements and is not limited herein.

[0041] The entire umbrella-shaped acquisition head 12 can be printed using 3D (three-dimensional) printing technology, or the acquisition parts 121 can be printed separately and then assembled into the umbrella-shaped acquisition head 12. A mold can also be opened for integrally forming the umbrella-shaped acquisition head 12, or the acquisition parts 121 can be manufactured separately and then assembled into the umbrella-shaped acquisition head 12. The umbrella-shaped acquisition head 12 printed by 3D printing technology can avoid the influence of manufacturing errors caused by mold manufacturing on the effect of directional acquisition.

[0042] First, use 3D modeling software to create a three-dimensional model of the umbrella-shaped acquisition head 12, and accurately set parameters such as its shape, size, and internal structure to form a model file. Then, import the model file into a 3D printer. Select a suitable printing material, such as a durable plastic. The 3D printer stacks the materials layer by layer and finally prints out the physical object of the umbrella-shaped acquisition head 12.

[0043] Optionally, the inner part of the acquisition part 121 is recessed into a concave structure in the shape of a paraboloid or a spherical surface (preferably a paraboloid). The concave structure is also the acoustic wave reflection and focusing structure 1223, which ensures that the acoustic waves are reflected and focused onto the microphone receiving end. That is to say, the acoustic wave reflection and focusing structure 1223 includes a cavity surrounded by an inner wall. The microphone receives the acoustic waves in this cavity.

[0044] The other end of the acquisition device 1 can be a handheld part 11 for holding, or it can also be a support base. The acquisition device 1 is placed on a plane (such as the ground, a tabletop) through the support base.

[0045] It can be understood that, please refer to Figure 1 、 Figure 3 、 Figure 4 and Figure 5 It is shown that there are 5 acquisition parts 121, one acquisition part 121 at the top, and 4 acquisition parts 121 are arranged around the acquisition part 121 at the top. It can be understood that the number of acquisition parts 121 is preferably 4 to 8. In addition to Figure 1 、 Figure 3 、 Figure 4 and Figure 5 shown in the distribution method, the distribution methods of the respective acquisition parts 121 can also be other, and are not limited herein.

[0046] The shape of the open end 1211 of the collection part 121 can be circular or other shapes, which is not limited here.

[0047] The sound wave 3 enters the sound wave reflection focusing structure 1223 from the outside of the collection part 121 through the open end 1211 and is received by the receiving end of the microphone in the collection part 121 to form an audio.

[0048] In this embodiment, first, the acquisition device 1 controls the microphone array 16 to acquire based on the acquisition control data corresponding to the acquisition direction data. This method of focusing on acquisition in a specified direction avoids the drawback of excessive noise acquisition in traditional omnidirectional acquisition, can selectively obtain the audio signal in the target direction (main acquisition direction), and reduces the noise acquisition in the non-target direction (auxiliary acquisition direction). Second, the processing device 2 performs directional noise reduction and speech enhancement according to the first audio and the acquisition direction data. The operation based on the acquisition direction data at the backend greatly improves the noise reduction effect and effectively suppresses the noise in the speech. Third, the collection part 121 of the umbrella-shaped collection head 12 at one end of the acquisition device 1 is concavely formed into a sound wave reflection focusing structure 1223, and the receiving end of at least one microphone in the microphone array 16 is located inside this structure. This specific setting of the collection part 121 well cooperates with the focused acquisition in the specified direction, further realizes directional acquisition, and makes the acquired audio signal more targeted in terms of direction. Through the acquisition direction data and the special setting of the umbrella-shaped collection head 12, the present application improves the overall ability of the system for directional noise reduction and speech enhancement and improves the quality of the target speech.

[0049] Please refer to Figure 1 、 Figure 3 、 Figure 4 and Figure 5 , in one embodiment, the types of the microphones include cardioid and omnidirectional. The receiving end 1221 of a cardioid-type microphone is provided at the central position of the sound wave reflection focusing structure 1223, and at least one receiving end 1222 of an omnidirectional-type microphone is provided at the edge position of the sound wave reflection focusing structure 1223.

[0050] It can be understood that in the collection part 121, the receiving end of the microphone is located between the open end 1211 and the central position of the sound wave reflection focusing structure 1223.

[0051] Specifically, the concave structure of the collection part 121 reflects the incident sound wave to the receiving end of the cardioid-type microphone located at the central position.

[0052] The receiving end of a cardioid microphone has unique sound pickup characteristics. It is most sensitive to sound waves directly in front, less sensitive or with reduced sensitivity to sound waves outside the front, forming a sound pickup pattern similar to a heart shape. When receiving sound waves, it can better focus on the sound waves coming from the front and effectively suppress the interference of sound waves from the rear and sides.

[0053] The receiving end of a cardioid microphone usually has a diaphragm located at the center position, with a specially designed sound insulation structure around it. There is no obstruction in front of the diaphragm, which is beneficial for receiving sound waves directly in front. There are reflection or sound absorption designs on both sides and at the rear, making it sensitive to sound waves from the front and reducing the reception of sound waves from the sides and the rear, forming a cardioid sound pickup pattern.

[0054] The receiving end of an omnidirectional microphone can receive sound waves omnidirectionally. Its sensitivity to sound waves is relatively uniform in all directions. Whether the sound waves come from any angle in the horizontal direction or from above and below within a certain vertical range, they can be received. That is to say, the receiving end of an omnidirectional microphone adopts an open structure design.

[0055] The diaphragm of the receiving end of an omnidirectional microphone is mostly designed at a relatively central position. The surrounding structure is relatively balanced, without special obstruction or enhancement design for a certain direction, enabling sound waves to reach the diaphragm evenly from all directions, thus realizing the function of receiving sound waves omnidirectionally.

[0056] Optionally, the inner wall of the sound wave reflection and focusing structure 1223 of the acquisition part 121 is covered with a sound absorption material layer, and the sound absorption material layer is a uniformly covered sound absorption material. The sound absorption material layer can receive the sound waves that enter the cavity of the sound wave reflection and focusing structure 1223 and propagate to the inner wall.

[0057] Optionally, the sound absorption material layer can be made of a nanofiber composite material. The high porosity of the nanofiber composite material and the frictional loss between fibers can absorb high-frequency sound waves. The sound absorption layer is used to suppress high-frequency standing waves and at the same time avoid excessive attenuation of the low-frequency energy of speech.

[0058] Optionally, the sound absorption material layer can be made of graphene or graphene oxide film. The thickness of the graphene or graphene oxide film can be less than 0.1 mm, and the flexibility and internal damping characteristics of the material are used to absorb high-frequency vibrations.

[0059] In this embodiment, through the sound wave reflection and focusing effect of the concave structure, sound wave direction screening and physical noise reduction are realized at the hardware level - the concave structure precisely reflects sound waves within a specific direction range to the heart-shaped microphone at the center to achieve focusing gain, while blocking sound waves outside the specific direction range from entering the sound wave reflection and focusing structure 1223; the signal fusion of the heart-shaped microphone (main receiver) and the omnidirectional microphone at the edge (auxiliary noise reference), combined with the spatial selectivity of the sound wave reflection and focusing structure 1223, finally realizes hardware-level directional noise reduction and sound source separation, further improving the system's directional noise reduction and speech enhancement capabilities as a whole, and improving the quality of the target speech.

[0060] Please refer to Figure 2 , in one embodiment, an indicator light 15 is provided for each of the acquisition units 121, and a real-time signal quality evaluation module based on FPGA is integrated in the acquisition device 1; The real-time signal quality evaluation module is used for: Obtaining, at a first time interval, the audio formed by the sound waves collected by the microphone array 16 as a second audio, and calculating the real-time signal-to-noise ratio and real-time signal strength of each of the acquisition units 121 according to the second audio; Based on a preset dynamic signal-to-noise ratio threshold and a preset dynamic signal strength threshold, screening out N of the acquisition units 121 with the highest weighted scores of the real-time signal-to-noise ratio and the real-time signal strength, where 0 < N < P, and P is the total number of the acquisition units 121; Controlling the indicator light 15 corresponding to each of the screened acquisition units 121 to be lit, and controlling the indicator lights 15 of the remaining acquisition units 121 to be turned off or dimmed; Wherein, the dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold are adaptively updated according to the environmental noise level of the environment where the acquisition device 1 is located.

[0061] Optionally, the real-time signal quality evaluation module based on FPGA (Field Programmable Gate Array) is a circuit module specifically designed for real-time evaluation of signal quality.

[0062] Specifically, obtaining, at a first time interval, the audio formed by the latest sound waves collected by the microphone array 16 as a second audio.

[0063] The first controller 13 controls the indicator light 15.

[0064] The indicator light 15 can be provided on the inner wall of the sound wave reflection and focusing structure 1223, so as to illuminate the cavity of the sound wave reflection and focusing structure 1223.

[0065] Extract the average signal strength of the second audio as the real-time signal strength. First, determine the signal power and noise power in the second audio, take the logarithm of the ratio of the two to obtain the signal-to-noise ratio, and use this signal-to-noise ratio as the real-time signal-to-noise ratio. For each acquisition unit 121, compare its real-time signal-to-noise ratio with the dynamic signal-to-noise ratio threshold (to obtain a signal-to-noise ratio comparison score), and compare its real-time signal strength with the dynamic signal strength threshold (to obtain a strength comparison score). Then, perform a weighted sum of the signal-to-noise ratio comparison score and the strength comparison score to obtain a comprehensive score. Finally, sort the comprehensive scores of all acquisition units 121, and select N acquisition units 121 in the direction from the highest to the lowest. Based on a preset control strategy, control the indicator lights 15 corresponding to each selected acquisition unit 121 to light up, and control the indicator lights 15 of the remaining unselected acquisition units 121 to turn off or dim.

[0066] Obtain the audio formed by the sound waves collected by the microphone array 16 at the second time interval as the audio to be analyzed, and calculate the short-time energy mean of the noise in the audio to be analyzed as the environmental noise reference value. ; Dynamic signal-to-noise ratio threshold , is the reference noise, is the logarithmic function with base 10, , is the empirical coefficient; Dynamic signal strength threshold , the value range of γ is 1.5 - 2.0, and γ is adaptively adjusted according to the noise fluctuation.

[0067] In this embodiment, the audio collected by the acquisition unit 121 is obtained at the first time interval to calculate the signal-to-noise ratio and signal strength of the acquisition unit 121, and then the N acquisition units 121 with the highest weighted scores are selected. By controlling the on / off or brightness of the indicator light 15, the status of each acquisition unit 121 is visually displayed, which is convenient for users to quickly identify and timely discover problems for adjustment. Moreover, the dynamic thresholds (that is, the dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold) are adaptively updated according to the environmental noise, making the evaluation more suitable for the actual environment, avoiding misjudgment of the fixed threshold in different environments, and helping to improve the reliability and stability of the audio formed by the sound waves collected by the entire system.

[0068] In one embodiment, the real-time signal quality evaluation module is used for: Obtain the audio formed by the sound waves collected by the microphone array 16 at the first time interval as the second audio, and calculate the real-time signal-to-noise ratio and real-time signal strength of each acquisition unit 121 according to the second audio; Based on a preset fixed signal-to-noise ratio threshold and a preset fixed signal strength threshold, N of the acquisition units 121 with the highest weighted scores of the real-time signal-to-noise ratio and the real-time signal strength are selected, where 0 < N < P, and P is the total number of the acquisition units 121; Control the indicator lights 15 corresponding to each of the selected acquisition units 121 to light up, and control the indicator lights 15 of the remaining acquisition units 121 to turn off or dim.

[0069] The preset fixed signal-to-noise ratio threshold is a preset threshold. The preset fixed signal strength threshold is a preset threshold. Using the preset thresholds reduces computing resources.

[0070] Please refer to Figure 2 , in one embodiment, the processing device 2 is provided with a first input component 24, the acquisition device 1 is further provided with a second input component 14, a direction detection component 17, and a camera component 18, the system further includes a remote control, the remote control is communicatively connected to the acquisition device 1 or the processing device 2, and an acquisition control module is integrated in the acquisition device 1; The acquisition control module is configured to: Obtain the main direction angle set by the user through any one of the first input component 24, the second input component 14, and the remote control as the first direction; Take the current pointing direction of the acquisition device 1 detected by the direction detection component 17 as the second direction; Based on the image data captured by the camera component 18, determine the face direction through a face recognition algorithm as the third direction; Determine the main acquisition direction and the auxiliary acquisition direction according to the first direction, the second direction, and the third direction as the initial direction data; Judge whether it is necessary to update the acquisition control data according to the deviation between the initial direction data and the acquisition direction data corresponding to the acquisition control data; If necessary, determine the control data according to the initial direction data as the updated acquisition control data, and update the acquisition direction data according to the initial direction data; According to the acquisition control data, control the microphone array 16 to collect sound waves to form an audio as the first audio; Wherein, the acquisition control data includes: main acquisition data and auxiliary acquisition data.

[0071] The second input component 14 can be arranged on the handheld part 11. The direction detection component 17 can be arranged in the accommodation cavity inside the handheld part 11.

[0072] The position of the imaging component 18 can be set according to requirements, and the imaging component 18 can capture the surrounding environment of the acquisition device 1. The imaging component 18 can be installed on the pan-tilt to capture the surrounding environment of the acquisition device 1. Multiple imaging components 18 can also be installed on the acquisition device 1 to capture the surrounding environment of the acquisition device 1. The second input component 14 and its supporting structure (such as a pan-tilt) can be selected from the prior art and will not be elaborated here.

[0073] The first input component 24 can be a button or a touch screen. The second input component 14 can be a button or a touch screen. The first controller 13 controls the second input component 14, the direction detection component 17, and the imaging component 18 to work, and the second controller 21 controls the first input component 24 to work.

[0074] The remote control is communicatively connected to the acquisition device 1 or the processing device 2 through Bluetooth technology.

[0075] The direction detection component 17 uses a gyroscope, a magnetometer, etc.

[0076] The direction detection component 17 is used to obtain the spatial orientation data of the acquisition device 1 in real time. According to the spatial orientation data of the acquisition device 1 and the corresponding relationship between the spatial orientation data and the default acquisition direction, determine the sound wave source direction of the acquisition unit 121 corresponding to the default acquisition direction, and use this sound wave source direction as the second direction. The direction detection component 17 usually includes a MEMS gyroscope (measuring angular velocity), an accelerometer (detecting three-dimensional acceleration), and a magnetometer (perceiving the direction of the geomagnetic field), and calculates the pitch angle, yaw angle, and roll angle (accuracy ≤ 1°) of the acquisition device 1 through a sensor fusion algorithm (such as Kalman filtering). The direction detection component 17 outputs direction data at a frequency of ≥ 100Hz.

[0077] The MEMS gyroscope (microelectromechanical gyroscope) is a micro sensor manufactured based on microelectromechanical system (MEMS) technology and is mainly used to measure the rotational angular velocity.

[0078] The acquisition unit 121 corresponding to the default acquisition direction is the acquisition unit 121 that wants to collect the sound waves formed by the speaker's voice. The default acquisition direction can be preset when the acquisition device 1 leaves the factory or can be pre-input into the acquisition device 1 by the user according to their own usage habits.

[0079] By analyzing the image data captured by the imaging component 18 in real time, the face recognition algorithm first locates the facial feature points. Then, according to the relative position relationship of key feature points such as the eyes, nose, and mouth, calculate the angle of the face relative to the camera, thereby determining the face direction.

[0080] Optionally, the first direction can be the initially set direction or the direction set in real time. The second direction comes from the direction detection component 17, and the third direction is related to the human face. Compare the relationships between the three with the preset rules. For example, preferentially select the direction pointing to the human face as the main acquisition direction, and determine the other directions as auxiliary acquisition directions after considering weights. If the first direction meets the main acquisition requirements, it is directly determined, and then the second and third directions are used to supplement the auxiliary acquisition directions, and the integrated result is used as the initial direction data.

[0081] Optionally, the vectors of the first direction, the second direction, and the third direction (after unitization) are superimposed according to weights by the vector synthesis method, and the direction of the synthesized vector is taken as the main acquisition direction. A preset angle range is symmetrically extended on both sides of the main acquisition direction to form the auxiliary acquisition directions.

[0082] Optionally, when the deviation between two of the first direction, the second direction, and the third direction > 60°, preferentially use the third direction as the main acquisition direction, and symmetrically extend a preset angle range on both sides of the main acquisition direction to form the auxiliary acquisition directions.

[0083] Optionally, when the deviation between two of the first direction, the second direction, and the third direction > 60°, the indicator light 15 can also be controlled to light up different colors to remind the user to adjust the first direction or the second direction. That is, it is required that the indicator light 15 of the acquisition unit 121 corresponding to the first direction displays the first color, the indicator light 15 of the acquisition unit 121 corresponding to the second direction displays the second color, and the indicator light 15 of the acquisition unit 121 corresponding to the third direction displays the third color.

[0084] In this embodiment, first, through various input components and the remote control, the user can flexibly set the main direction angle, improving the convenience and diversity of the acquisition direction setting. Second, by combining the detection of the direction detection component 17 of the acquisition device 1 itself and the face recognition of the imaging component 18, multi-direction data is determined, making the determination of the main acquisition direction and the auxiliary acquisition directions more accurate. Third, according to the deviation between the initial direction data and the existing acquisition direction data, it is judged whether to update the acquisition control data, realizing the adaptive adjustment of the acquisition control data, and avoiding frequent update of the acquisition control data while ensuring that the acquisition direction always meets the requirements. Finally, according to the updated acquisition control data, the microphone array 16 is controlled to collect audio, and the audio in the target direction (the main acquisition direction and the auxiliary acquisition directions) can be collected more accurately, improving the accuracy and effectiveness of audio acquisition.

[0085] Please refer to Figure 2 , in one embodiment, the acquisition device 1 is provided with at least three direction acquisition positions, and the receiving end of one type of omnidirectional microphone in the microphone array is located at the direction acquisition position; Determining a main acquisition direction and an auxiliary acquisition direction based on the first direction, the second direction, and the third direction in the acquisition control module as initial direction data includes: Collecting, as positioning audio, the audio formed by the microphone corresponding to the receiving end in each of the direction acquisition positions collecting sound waves, and determining the sound source orientation based on the positioning audio as the fourth direction; Determining the main acquisition direction and the auxiliary acquisition direction from the first direction, the second direction, the third direction, and the fourth direction as the initial direction data.

[0086] The first direction, the second direction, the third direction, and the fourth direction adopt the same reference system (for example, the device coordinate system).

[0087] All the direction acquisition positions can be set on the umbrella-shaped acquisition head 12, or some can be set on the umbrella-shaped acquisition head 12 and some on the handheld part 11.

[0088] Optionally, 8 direction acquisition positions are distributed annularly on the periphery of the acquisition device 1, with each acquisition position spaced 45°, covering a 360° omnidirectional range. Each direction acquisition position is equipped with an omnidirectional microphone, and its receiving end faces the corresponding angle center (such as 0°, 45°, 90°... 315°).

[0089] Optionally, a preset time period is used to control the microphone corresponding to the receiving end in each of the direction acquisition positions to collect sound waves to form an audio, and this audio is used as the positioning audio.

[0090] The method for determining the sound source orientation from each sub-audio in the positioning audio (the audio formed by the microphone corresponding to the receiving end in one direction acquisition position) can be selected from the prior art.

[0091] Optionally, the original audio signal collected by each omnidirectional microphone (i.e., the positioning audio) is band-pass filtered (85 Hz - 8 kHz), then frame processed (frame length 20 ms, frame shift 10 ms), the short-time energy and zero-crossing rate are extracted to preliminarily filter out environmental noise to obtain a preprocessing result; the generalized cross-correlation (GCC-PHAT) algorithm is used to calculate the time delay difference of the preprocessing results of adjacent microphone pairs (i.e., the time difference between the two microphones receiving the same sound wave); based on the geometric position of the microphone pair (spacing d = 50 mm) and the time delay difference, the sound source direction angle is solved by the least squares method; the sound source orientation (also called the sound source direction) is determined based on the sound source direction angle, and each sound source orientation corresponding to the positioning audio is used as the fourth direction.

[0092] Optionally, the fourth direction includes at least 0 sound source orientations.

[0093] Optionally, the first direction can be the initially set direction or the direction set in real time. The second direction is, for example, from the direction detection component 17. The third direction is related to the human face, and the fourth direction is related to lip movement. Compare the relationships between these four and the preset rules. For example, preferably select the direction pointing to lip movement as the main acquisition direction, and determine the other directions as auxiliary acquisition directions after considering weights. If the fourth direction meets the main acquisition requirements, it is directly determined, and then the first, second, and third directions are used to supplement the auxiliary acquisition directions, and after integration, it is used as the initial direction data.

[0094] Optionally, the vectors of the first direction, the second direction, the third direction, and the fourth direction (after unitization) are superimposed according to weights by the vector synthesis method, and the direction of the synthesized vector is taken as the main acquisition direction. Symmetrically expand a preset angle range on both sides of the main acquisition direction to form auxiliary acquisition directions.

[0095] Optionally, when the deviation between two directions of each sound source direction among the first direction, the second direction, the third direction, and the fourth direction > 60°, preferably use the sound source direction in the fourth direction as the main acquisition direction, and symmetrically expand a preset angle range on both sides of the main acquisition direction to form auxiliary acquisition directions.

[0096] Optionally, when the deviation between two directions of the first direction, the second direction, the third direction, and the fourth direction > 60°, it is also possible to control the indicator light 15 to light up different colors to remind the user to adjust the first direction or the second direction. That is to say, it is required that the indicator light 15 of the acquisition part 121 corresponding to the first direction shows the first color, the indicator light 15 of the acquisition part 121 corresponding to the second direction shows the second color, the indicator light 15 of the acquisition part 121 corresponding to the third direction shows the third color, and the indicator light 15 of the acquisition part 121 corresponding to the fourth direction shows the fourth color.

[0097] In this embodiment, the audio formed by the microphone corresponding to the receiving end in the direction acquisition position collects sound waves to determine the sound source orientation, that is, the fourth direction. The fourth direction is combined with the first, second, and third directions to determine the acquisition direction data. This helps to more accurately locate the acquisition source and improve the acquisition accuracy. It can meet the acquisition requirements in complex environments. For example, when there are multiple sound sources and interference sources, through comprehensive consideration of multiple directions, it can avoid misjudgment due to a single direction, align the main acquisition direction with the speaker, effectively supplement the auxiliary acquisition directions, optimize the overall performance of the acquisition device 1, better obtain the relevant data of the speaker, and improve the acquisition efficiency and quality.

[0098] Please refer to Figure 2 , in one embodiment, the processing device 2 is integrated with an aging compensation module, and the aging compensation module is used for: Obtain a compensation signal; In response to the compensation signal, obtain preset control data as debugging control data, and update the acquisition control data of the acquisition device 1 according to the debugging control data; Control the processing device 2 to play a preset standard test audio, and obtain a response audio corresponding to the standard test audio collected by the acquisition device 1; Compare the frequency response curves of the response audio and the standard test audio to obtain frequency response difference data; Generate aging compensation data according to the frequency response difference data; The determining control data according to the initial direction data in the acquisition control module as the updated acquisition control data includes: Determine control data according to the initial direction data as initial control data; Correct the initial control data according to the aging compensation data as the updated acquisition control data.

[0099] The compensation signal is a signal for starting to determine the aging compensation data.

[0100] The user triggers the compensation signal through any one of the first input component 24, the second input component 14, and the remote control, or it can also be a compensation signal automatically triggered by the system according to preset time data.

[0101] Specifically, first, convert the response audio and the standard test audio into frequency-domain signals, and obtain the frequency response curves corresponding to the response audio and the standard test audio respectively. For each frequency point, calculate the difference between the amplitude value of the frequency response curve of the response audio and the amplitude value of the frequency response curve of the standard test audio. These differences constitute the frequency response difference data. The discrete Fourier transform (DFT) or the fast Fourier transform (FFT) can be used to convert the audio signals (that is, the response audio, the standard test audio) into frequency-domain signals. Then, under the same frequency scale, compare the amplitudes of the two point by point to accurately obtain the frequency response difference data, so as to reflect the change of the audio acquisition characteristics of the acquisition device 1 due to factors such as aging.

[0102] Control the playback component 22 in the processing device 2 to play a preset standard test audio. The second controller 21 controls the operation of the playback component 22. The playback component 22 can use a speaker.

[0103] Optionally, determine the aging compensation data by using a look-up table method according to the frequency response difference data.

[0104] Optionally, input the frequency response difference data into a pre-trained third model for classification prediction, extract the vector element with the largest value from the predicted vectors, and use the compensation data corresponding to the extracted vector element as the aging compensation data.

[0105] The aging compensation data can be determined based on limited experiments, which will not be elaborated here.

[0106] The pre-trained third model is a pre-trained multi-classification model. The model structure and model training method of the third model can be selected from the prior art.

[0107] Based on a preset correction formula, the initial control data is corrected according to the aging compensation data to obtain the updated acquisition control data.

[0108] The preset correction formula can be obtained by fitting based on multiple trial data, which will not be elaborated here.

[0109] In this embodiment, the setting of the aging compensation module can update the acquisition control data by obtaining a compensation signal, ensuring that the acquisition device 1 adapts to the aging situation. By playing a standard test audio and comparing the response audio to obtain frequency response difference data, and then generating aging compensation data, this helps to accurately compensate for the performance deviation of the acquisition device 1 caused by aging. When determining the acquisition control data, by integrating the aging compensation data and the initial direction data, the initial control data is also corrected, improving the accuracy and reliability of the acquisition, enabling the acquisition device 1 to maintain stable performance during long-term use, and reducing the impact of aging on the acquisition effect.

[0110] Please refer to Figure 2 , in one embodiment, the processing device 2 is integrated with a directional noise reduction and speech enhancement module; The directional noise reduction and speech enhancement module is used for: Determining a directional noise reduction scheme according to the acquisition direction data; Performing directional noise reduction on the first audio according to the directional noise reduction scheme to obtain a third audio; Performing voice separation and speech enhancement on the third audio to obtain the target speech.

[0111] Specifically, according to the acquisition direction data, a lookup table method is used to determine the directional noise reduction scheme.

[0112] Using a preset voice separation method to perform voice separation on the third audio, using a preset speech enhancement method to perform speech enhancement on the data obtained by voice separation, and taking the enhanced speech as the target speech.

[0113] The directional noise reduction scheme includes the following types of data: acquisition direction data and noise reduction control parameters corresponding to each acquisition direction in the acquisition direction data.

[0114] In this embodiment, by collecting direction data to determine a directional noise reduction scheme, the first audio can be accurately directionally noise-reduced to obtain a third audio, effectively reducing noise interference outside the main collection direction, enabling more effective suppression of noise in the speech, and improving the quality of the target speech.

[0115] In one embodiment, the step of directionally noise-reducing the first audio according to the directional noise reduction scheme in the directional noise reduction and speech enhancement module to obtain a third audio includes: Using a multi-stage NLMS filter to directionally noise-reduce the first audio according to the directional noise reduction scheme to obtain the third audio; Among them, the multi-stage NLMS filter includes three parallel target filters and a combiner. The target filters sequentially include: a Butterworth band-pass filter and an NLMS filter; the combiner is used for weighted fusion of the outputs of the respective target filters; The first-stage target filter processes the frequency band P1, 0 Hz < P1 ≤ the first frequency band value, and the step factor is adjusted dynamically according to the formula where and are empirical constants, and is the signal-to-noise ratio of the data input to the target filter in the range of 0 Hz to the first frequency band value; The second-stage target filter processes the frequency band P2, the first frequency band value < P2 ≤ the second frequency band value, and the step factor is adjusted dynamically according to the formula where and are empirical constants, and is the signal-to-noise ratio of the data input to the target filter in the range of the first frequency band value to the second frequency band value; The third-stage target filter processes the frequency band P3, the second frequency band value < P3 ≤ the third frequency band value, and the step factor is adjusted dynamically according to the formula where and are empirical constants, and is the signal-to-noise ratio of the data input to the target filter in the range of the second frequency band value to the third frequency band value.

[0116] The specific values of the first frequency band value, the second frequency band value, and the third frequency band value can be set according to requirements and are not limited herein.

[0117] Optionally, the first frequency band value is set to 1 kHz, the second frequency band value is set to 4 kHz, and the third frequency band value is set to 8 kHz.

[0118] Specifically, according to the described directional noise reduction scheme, the noise reduction control parameters of each sub-audio in the first audio are determined, and when the multi-stage NLMS filter processes this sub-audio, the NLMS filter uses the noise reduction control parameter corresponding to this sub-audio.

[0119] The Butterworth band-pass filter is mainly used to select a specific frequency range, allowing signals within this frequency band to pass through while suppressing signals outside the frequency band. In this multi-stage NLMS filter system, it serves as the front-end part of the target filter and is used to perform preliminary frequency screening on the input first audio in the corresponding target frequency bands (such as 0 Hz to 1 kHz, 1 to 4 kHz, 4 to 8 kHz), enabling the subsequent NLMS filter to better perform noise reduction operations based on the screened frequency band signals.

[0120] The NLMS filter (Normalized Least Mean Square filter) is used to adaptively adjust the coefficients of the filter to minimize the error between the desired signal and the filter output. In this system, it receives the signal processed by the Butterworth band-pass filter and dynamically adjusts its own coefficients according to the statistical characteristics of the input signal, thereby effectively removing the noise components within the corresponding frequency band, improving the quality of the speech signal in this frequency band, and ultimately achieving the directional noise reduction effect on the first audio in different frequency bands.

[0121] The target filter at the first level mainly processes environmental noise and uses a larger step factor for rapid convergence.

[0122] The target filter at the second level enhances speech clarity (the main energy area of human voices) and has a moderate step factor.

[0123] The target filter at the third level suppresses sharp noises (such as metal collision sounds) and uses a smaller step to avoid distortion.

[0124] The combiner first receives the outputs of the NLMS filters from the three target filters. Then, based on the pre-set weight coefficients, which can be set according to the noise reduction requirements and the importance of different frequency bands, it performs a weighting operation on the outputs of each level of the target filter, adds the weighted results together, and finally obtains the combined output (i.e., the third audio), achieving an effective combination of the noise reduction results in different frequency bands.

[0125] The step factor of the target filter clarifies the negative correlation between the step and the signal-to-noise ratio, which is beneficial for reducing the convergence speed in high-noise frequency bands and avoiding signal distortion.

[0126] and can be obtained through limited experiments. and can be obtained through limited experiments. and It can be obtained through limited experiments.

[0127] In this embodiment, using such a multi-stage NLMS filter for directional noise reduction can precisely process different frequency bands. In frequency bands such as 0 Hz to the first frequency band value, the first frequency band value to the second frequency band value, and the second frequency band value to the third frequency band value, they are processed by different target filters respectively, and the step factor is dynamically adjusted according to the signal-to-noise ratio of the frequency band, which can effectively adapt to different input audio situations. By weighting and fusing the outputs of each stage of the filter through a fusion device, the noise reduction advantages of different frequency bands can be integrated, the noise can be more precisely suppressed, which is beneficial to enhancing the target speech, thereby improving the speech quality during the directional noise reduction process and obtaining a purer and more recognizable third audio.

[0128] In one embodiment, the process of separating the human voice and enhancing the speech from the third audio in the directional noise reduction and speech enhancement module to obtain the target speech includes: Obtain hearing preference parameters, where the hearing preference parameters include: sensitive frequency bands and volume tolerance thresholds; Determine a speech enhancement scheme according to the hearing preference parameters and the acquisition direction data; According to the speech enhancement scheme, perform human voice separation and speech enhancement on the third audio to obtain the target speech.

[0129] The hearing preference parameters are data pre-input and stored by the user.

[0130] Optionally, a speech enhancement scheme is determined by using a look-up table method according to the hearing preference parameters and the acquisition direction data.

[0131] Optionally, the hearing preference parameters and the acquisition direction data are concatenated and then input into a pre-trained fourth model for classification prediction, the vector element with the largest value is extracted from the predicted vectors, and the scheme corresponding to the extracted vector element is used as the speech enhancement scheme.

[0132] The hearing preference parameters are obtained through a user input interface or a preset configuration file.

[0133] The pre-trained fourth model is a pre-trained multi-classification model. The model structure and model training method of the fourth model can be selected from the prior art.

[0134] The speech enhancement scheme includes a main direction speech main enhancement scheme and a main direction speech auxiliary enhancement scheme. The main direction speech main enhancement scheme is used to enhance the speech in the main acquisition direction after human voice separation to obtain the first speech, and the main direction speech auxiliary enhancement scheme is used to correct the first speech for the speech in the auxiliary acquisition direction after human voice separation to obtain the target speech.

[0135] In this embodiment, first, by introducing hearing preference parameters, a voice enhancement scheme can be customized according to the sensitive frequency band of an individual's hearing and the volume tolerance threshold, improving the personalization of the scheme. Second, by combining the acquisition direction data, the voice in a specific direction can be processed more precisely. During the process of voice separation and voice enhancement, the desired voice can be effectively enhanced, improving the clarity and intelligibility of the voice. For users with special hearing needs, such as being sensitive to certain frequency bands or having limited volume tolerance, a target voice that better suits their own hearing characteristics can be obtained, improving the user's voice experience.

[0136] In one embodiment, the outer surface of the other end of the acquisition device 1 is coated with a conductive coating; the conductive coating is connected to the ground wire of the microphone array 16 to form an electromagnetic shielding circuit.

[0137] Specifically, the outer surface of the other end of the acquisition device 1 is coated with a conductive coating by using a spraying process.

[0138] Optionally, in an environment of 25°C and 50% humidity, the surface resistivity of the conductive coating ≤ 5 Ω / cm².

[0139] In this embodiment, the conductive coating is made of a graphene composite material and has a low surface resistivity, which can effectively conduct current. After the conductive coating is connected to the ground wire of the microphone array 16 to form an electromagnetic shielding circuit, it can play a good electromagnetic shielding role. This technology can reduce the influence of external electromagnetic interference on the acquisition device 1, ensure the purity of the acquired audio signal, improve the accuracy and stability of audio acquisition, and enhance the overall audio acquisition quality.

[0140] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0141] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A directional noise reduction and speech enhancement system based on a microphone array, characterized in that: The system comprises: A collection device, wherein the collection device is provided with a microphone array, and is used to control the microphone array to collect sound waves to form a first audio based on collection control data corresponding to the collection direction data; A processing device, configured to perform directional noise reduction and speech enhancement according to the first audio and the acquisition direction data to obtain a target speech; Among them, an umbrella-shaped collecting head is provided at one end of the collecting device, and the umbrella-shaped collecting head includes at least four concave collecting parts, and the collecting parts are concave to form a sound wave reflection focusing structure. The receiving end of at least one microphone in the microphone array is located in one of the sound wave reflection focusing structures.

2. The directional noise reduction and speech enhancement system based on microphone array according to claim 1, characterized in that: The types of microphones include cardioid and omnidirectional. A receiving end of a cardioid microphone is provided at the center of the sound wave reflection and focusing structure, and at least one receiving end of an omnidirectional microphone is provided at the edge of the sound wave reflection and focusing structure.

3. The directional noise reduction and speech enhancement system based on microphone array according to claim 1, characterized in that: Each of the acquisition units is provided with an indicator light, and the acquisition device is integrated with a real-time signal quality assessment module based on FPGA; The real-time signal quality assessment module is used for: Acquire the audio formed by the sound waves collected by the microphone array at a first time interval as the second audio, and calculate the real-time signal-to-noise ratio and real-time signal strength of each of the collection units according to the second audio; Based on a preset dynamic signal-to-noise ratio threshold and a preset dynamic signal strength threshold, screen out the N acquisition units with the highest weighted scores of the real-time signal-to-noise ratio and the real-time signal strength, 0<N<P, where P is the total number of the acquisition units; Control the indicator lights corresponding to each of the screened collection parts to light up, and control the indicator lights of the remaining collection parts to turn off or dim; The dynamic signal-to-noise ratio threshold and the dynamic signal strength threshold are adaptively updated according to the ambient noise level of the environment where the acquisition device is located.

4. The directional noise reduction and speech enhancement system based on microphone array according to claim 3, characterized in that: The processing device is provided with a first input component, the acquisition device is further provided with a second input component, a direction detection component and a camera component, the system further comprises a remote controller, the remote controller is communicatively connected with the acquisition device or the processing device, and the acquisition device is integrated with an acquisition control module; The acquisition control module is used for: Acquire a main direction angle set by a user through any one of the first input component, the second input component and the remote controller as a first direction; Using the current pointing direction of the acquisition device detected by the direction detection component as a second direction; Based on the image data captured by the camera component, a face direction is determined by a face recognition algorithm according to the image data as a third direction; Determine a main acquisition direction and an auxiliary acquisition direction according to the first direction, the second direction and the third direction as initial direction data; Determining whether the acquisition control data needs to be updated according to a deviation between the initial direction data and the acquisition direction data corresponding to the acquisition control data; If necessary, determining control data according to the initial direction data as the updated acquisition control data, and updating the acquisition direction data according to the initial direction data; According to the acquisition control data, control the microphone array to acquire sound waves to form audio as the first audio; The acquisition control data includes: main acquisition data and auxiliary acquisition data.

5. The directional noise reduction and speech enhancement system based on microphone array according to claim 4, characterized in that: The collection device is provided with at least three directional collection positions, and a receiving end of one omnidirectional microphone in the microphone array is located at the directional collection position; The determining of the main acquisition direction and the auxiliary acquisition direction according to the first direction, the second direction and the third direction in the acquisition control module as initial direction data includes: The microphone corresponding to the receiving end in each of the direction collection positions collects audio formed by sound waves as positioning audio, and determines the direction of the sound source according to the positioning audio as a fourth direction; The first direction, the second direction, the third direction and the fourth direction are used to determine a main acquisition direction and an auxiliary acquisition direction as the initial direction data.

6. The directional noise reduction and speech enhancement system based on microphone array according to claim 4, characterized in that: The processing device is integrated with an aging compensation module, and the aging compensation module is used to: Obtain compensation signal; In response to the compensation signal, obtaining preset control data as debugging control data, and updating the acquisition control data of the acquisition device according to the debugging control data; Controlling the processing device to play a preset standard test audio, and obtaining a response audio corresponding to the standard test audio collected by the collection device; Comparing the frequency response curves of the response audio and the standard test audio to obtain frequency response difference data; generating aging compensation data according to the frequency response difference data; The control data determined according to the initial direction data in the acquisition control module as the updated acquisition control data includes: Determine control data according to the initial direction data as initial control data; The initial control data is corrected according to the aging compensation data to obtain the updated acquisition control data.

7. The directional noise reduction and speech enhancement system based on microphone array according to claim 1, characterized in that: The processing device is integrated with a directional noise reduction and speech enhancement module; The directional noise reduction and speech enhancement module is used for: Determining a directional noise reduction scheme according to the collected direction data; According to the directional noise reduction scheme, perform directional noise reduction on the first audio to obtain a third audio; Performing human voice separation and speech enhancement on the third audio to obtain the target speech.

8. The directional noise reduction and speech enhancement system based on microphone array according to claim 7, characterized in that: The step of performing directional noise reduction on the first audio according to the directional noise reduction scheme in the directional noise reduction and speech enhancement module to obtain a third audio includes: Using a multi-stage NLMS filter, according to the directional noise reduction scheme, directional noise reduction is performed on the first audio to obtain the third audio; The multi-stage NLMS filter includes three-stage parallel target filters and a fusion device, wherein the target filters include: a Butterworth bandpass filter and an NLMS filter in sequence; the fusion device is used to perform weighted fusion on the outputs of each of the target filters; The target filter of the first stage processes the frequency band P1, 0Hz<P1≤the value of the first frequency band, and the step factor is calculated according to the formula Dynamic adjustment, and is an empirical constant, is the signal-to-noise ratio of the data input to the target filter from 0 Hz to the first frequency band value; The target filter of the second stage processes the frequency band P2, the first frequency band value < P2 ≤ the second frequency band value, and the step factor is calculated according to the formula Dynamic adjustment, and is an empirical constant, is the signal-to-noise ratio of the data input to the target filter between the first frequency band value and the second frequency band value; The target filter of the third level processes the frequency band P3, the second frequency band value < P3 ≤ the third frequency band value, and the step factor is calculated according to the formula Dynamic adjustment, and is an empirical constant, It is the signal-to-noise ratio of the data input to the target filter in the second frequency band value ~ the third frequency band value.

9. The directional noise reduction and speech enhancement system based on microphone array according to claim 7, characterized in that: The step of performing human voice separation and voice enhancement on the third audio in the directional noise reduction and voice enhancement module to obtain the target voice includes: Acquiring hearing preference parameters, wherein the hearing preference parameters include: sensitive frequency bands and volume tolerance thresholds; Determining a speech enhancement scheme according to the hearing preference parameter and the acquisition direction data; According to the speech enhancement scheme, the third audio is subjected to voice separation and speech enhancement to obtain the target speech.

10. The directional noise reduction and speech enhancement system based on microphone array according to claim 1, characterized in that: The outer surface of the other end of the collection device is coated with a conductive coating; the conductive coating is connected to the ground wire of the microphone array to form an electromagnetic shielding loop.

Citation Information

Patent Citations

  • Multi-microphone array beamforming signal enhancement method and device

    CN119811408A

  • Acoustic-thermal detection device for high-voltage equipment

    CN214951816U

  • Directional noise-cancelling and sound detection system and method for sound targeted hearing and imaging

    US10553196B1