Voice enhancement and directional area rendering sound reinforcement system and method based on microphone array

Through the combination of microphone array and speaker array, the problems of teacher activities restricted and sound wave interference in the classroom sound reinforcement system are solved, efficient interaction and clear voice transmission between teachers and students are achieved, and the auditory experience is improved.

CN120264201AInactive Publication Date: 2025-07-04SHENZHEN JUSHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327747.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing classroom sound reinforcement system has problems such as limited teacher activity range, serious sound wave interference, microphones are susceptible to interference and packet loss, and cannot achieve efficient voice pickup and clarity guarantee.

Method used

Using microphone arrays, voice signal processing devices and speaker arrays, sound signals are collected through microphone arrays. The voice signal processing device separates and enhances the voice signals of teachers and students, and is rendered in a directional manner through the speaker array to ensure clear sound transmission and regional isolation.

Benefits of technology

It achieves more convenient and efficient interaction between teachers and students, avoids sound wave interference, improves speech clarity and auditory experience, and makes the sound field more uniform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264201A_ABST
    Figure CN120264201A_ABST
Patent Text Reader

Abstract

The invention provides a voice enhancement and directional area rendering sound reinforcement system and method based on a microphone array, and the system comprises the microphone array, a voice signal processing device, a first loudspeaker array, and a second loudspeaker array. The microphone array is used for collecting sound signals in a target space; the voice signal processing device is used for determining a first voice signal of a first area and a second voice signal of a second area in a target space from the sound signals, and processing the first voice signal and the second voice signal to obtain a first audio output signal and a second audio output signal; the first loudspeaker array outputs according to the first audio output signal, the output bright area is arranged in the second area, and the output dark area is arranged in the first area; the second loudspeaker array is used for outputting according to the second audio output signal, a bright area output by the second loudspeaker array is arranged in the first area, and a dark area output by the second loudspeaker array is arranged in the second area. According to the invention, the isolation effect during inter-talking in the scene is improved, and the sound wave interference phenomenon is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal processing, and in particular to a voice enhancement and directional area rendering sound reinforcement system and method based on a microphone array. Background Art

[0002] The main configuration of the current classroom sound reinforcement system usually adopts a fixed wired microphone, which is placed at the podium position to realize the function of collecting and amplifying the teacher's voice. The working process of this typical sound reinforcement system is as follows: after the microphone converts the sound signal into an electrical signal, it is processed by an audio processor for gain adjustment and the like, and then the signal power is increased through a power amplifier module, and finally the same audio signal is copied and transmitted to multiple speakers in the classroom for playback, so as to cover the entire classroom area. Limited by this technical solution, the teacher's activity range is strictly restricted within the effective sound pickup area of the microphone, resulting in significant constraints on their teaching behavior. In addition, when the same audio signal is copied and distributed to multiple speakers, interference will be formed during the propagation in the physical space, resulting in a serious wave crest and wave trough phenomenon. This phenomenon will lead to obvious sound field non-uniformity in the classroom, making students in some areas face too strong or too weak sound environments. For example, students in the wave crest area will hear very large sound waves, while students in the wave trough area may not hear clearly.

[0003] Currently, some classroom sound reinforcement systems also use portable microphones to pick up the teacher's speech, which is wirelessly transmitted (such as infrared, Bluetooth, radio frequency) to the audio processor for amplification, and finally output to the speaker for playback. Although this solution has a certain degree of freedom of movement, there are problems that the microphone is easily interfered by others, resulting in packet loss during transmission and loss of sound quality. Secondly, it has high power consumption and requires frequent battery replacement, which is inconvenient to use. In addition, this microphone cannot pick up the speech of students from a long distance and can only achieve one-way output.

[0004] Another existing sound reinforcement system is to set a voice receiving module and a sound playing module. The voice receiving module includes a direction determination module and a microphone array. The microphone array is used to obtain the voice signal in the target space, and the direction determination module is used to determine the sound source direction. The sound playing module includes a beam adjustment module and a speaker array. The voice signal of the sound source in the target space is collected through the microphone array, the sound source direction is determined through the direction determination module, and then the sound pickup beam of the microphone array can be controlled to point to the sound source. Although this solution realizes the directional propagation of sound to some people, it only plays the sound to the speaker after identifying the position of the speaker, and is not suitable for the classroom sound reinforcement scenario. Summary of the Invention

[0005] The first object of the present invention is to provide a voice enhancement and directional area rendering sound reinforcement system based on a microphone array, which can improve the sound reinforcement effect, avoid the phenomenon of sound wave interference, and improve the clarity of speech.

[0006] The second object of the present invention is to provide a voice enhancement and directional area rendering sound reinforcement method based on a microphone array, which can improve the sound reinforcement effect and ensure a certain degree of isolation and clarity during crosstalk.

[0007] To achieve the above first object, the present invention provides a voice enhancement and directional area rendering sound reinforcement system based on a microphone array, which includes: a microphone array, a voice signal processing device, a first speaker array, and a second speaker array; a voice signal processing system runs on the voice signal processing device; the microphone array is used to collect sound signals in the target space; the voice signal processing system is used to determine a first voice signal in a first area and a second voice signal in a second area from the sound signals, and process the first voice signal to obtain a first audio output signal and process the second voice signal to obtain a second audio output signal; the first speaker array is used to output according to the first audio output signal, the bright area output by the first speaker array is set in the second area, and the dark area output by the first speaker array is set in the first area; the second speaker array is used to output according to the second audio output signal, the bright area output by the second speaker array is set in the first area, and the dark area output by the second speaker array is set in the second area.

[0008] As can be seen from the above solution, the present invention can cover a wider sound pickup area by setting a microphone array to collect sound signals in the target scene, significantly enhance the sound pickup effect, and accurately capture all sound signals in the target space. By setting a voice signal processing device to extract the first voice signal in the first area and the second voice signal in the second area from the sound signals, and then performing enhancement and directional rendering, the set first speaker array and second speaker array respectively render the sounds in the corresponding areas. The present invention can accurately distinguish the sound source position by processing the sound signals, ensuring clear transmission of sound. In addition, through the rendering method of two speaker arrays, the audio content is rendered to the specified areas respectively, realizing isolation when people in the first area and people in the second area talk to each other, and being able to pick up the speeches of people in the first area and people in the second area at the same time, making the interaction more convenient and efficient, avoiding the phenomenon of sound wave interference, and ensuring the clarity and intelligibility of speech. The present invention can also make the sound field more uniform, improving the overall auditory experience.

[0009] A further solution is that the microphone array includes a main microphone, and the voice signal processing system includes a voice activity detection module; the voice activity detection module is used to determine whether there is a normal voice signal in the sound signal according to the acquisition data of the main microphone, and after determining that there is a normal voice signal in the sound signal, determine a first voice signal and a second voice signal from the sound signal.

[0010] It can be seen that only using the main microphone as the input of the voice activity detection module can confirm whether the sound signal includes a normal voice signal, so there is no need to judge each microphone, avoiding waste of resources.

[0011] A further solution is that the voice signal processing system includes a sound source localization module and a beamforming module, and the sound source localization module is connected to the beamforming module; the sound source localization module is used to determine a first voice signal and a second voice signal from the sound signal, and the beamforming module is used to perform beamforming processing on the determined first voice signal and second voice signal according to the sound source position direction information obtained by the sound source localization module, respectively obtaining a first output audio signal and a second output audio signal.

[0012] A further solution is that the voice signal processing system includes a voice enhancement module, and the voice enhancement module is connected to the beamforming module; the voice enhancement module is used to perform voice enhancement processing on the first voice signal and the second voice signal after beam processing, respectively obtaining a first output audio signal and a second output audio signal.

[0013] A further solution is that the voice signal processing system includes a howling suppression module, and the howling suppression module is connected to the voice enhancement module. The howling suppression module is used to perform howling suppression processing on the first voice signal and the second voice signal after voice enhancement processing, respectively obtaining a first output audio signal and a second output audio signal.

[0014] It can be seen that by performing specific processing on each microphone signal, the pickup distance, the ability to distinguish the sound source position, and the signal strength can be greatly enhanced. Through specific processing of the signal, the position of the sound source can also be accurately distinguished, ensuring that the voice of the speaker can be clearly transmitted.

[0015] A further solution is that the target space is a classroom, the first area is the student area, the first voice signal is the student voice signal, the second area is the teacher area, and the second voice signal is the teacher voice signal.

[0016] It can be seen that the present invention can simultaneously pick up and process the speeches of students, making the interaction in the classroom more convenient and efficient, and improving the quality of classroom teaching.

[0017] A further solution is that the microphone array is arranged on the ceiling of the classroom.

[0018] Thus, it can make the wiring of the entire microphone array simple and tidy, and will not limit the activity range of the teacher.

[0019] To achieve the above second object, a voice enhancement and directional area rendering sound reinforcement method based on a microphone array provided by the present invention includes the following steps: obtaining a sound signal in a target space from the microphone array; determining a first voice signal in a first area and a second voice signal in a second area from the sound signal; processing the first voice signal to obtain a first output audio signal; processing the second voice signal to obtain a second output audio signal; outputting the first output audio signal to a first speaker array, where the bright area output by the first speaker array is set in the second area, and the dark area output by the first speaker array is set in the first area; outputting the second output audio signal to a second speaker array, where the bright area output by the second speaker array is set in the first area, and the dark area output by the second speaker array is set in the second area.

[0020] As can be seen from the above solution, through the extraction and processing of the sound signal, the present invention enables the speech in the first area to be better played in the second area through the first speaker array, enables the speech in the second area to be better played in the first area through the second speaker array, and the speech in the first area and the speech in the second area can be carried out simultaneously, improving the sound reinforcement effect and ensuring a certain degree of isolation and clarity during crosstalk.

[0021] A further solution is that after obtaining the sound signal in the target space from the microphone array, it is determined whether there is a normal voice signal in the acquisition data of the main microphone of the microphone array. When it is determined that there is a normal voice signal, then the first voice signal in the first area and the second voice signal in the second area in the sound signal are determined.

[0022] A further solution is that when processing the first voice signal and processing the second voice signal, both include beamforming processing, voice enhancement processing, howling suppression processing, and mixing processing performed in sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a system block diagram of an embodiment of the voice enhancement and directional area rendering sound reinforcement system of the microphone array of the present invention.

[0024] Figure 2 is a signal processing block diagram in an embodiment of the voice enhancement and directional area rendering sound reinforcement system of the microphone array of the present invention.

[0025] Figure 3 is a signal processing block diagram among the mixing module, the first speaker array, and the second speaker array in an embodiment of the voice enhancement and directional area rendering sound reinforcement system of the microphone array of the present invention.

[0026] Figure 4 It is a schematic diagram of the first voice signal rendering in the voice enhancement and directional area rendering sound reinforcement system embodiment of the microphone array of the present invention.

[0027] Figure 5 It is a schematic diagram of the second voice signal rendering in the voice enhancement and directional area rendering sound reinforcement system embodiment of the microphone array of the present invention.

[0028] Figure 6 It is a flowchart of the method embodiment for voice enhancement and directional area rendering sound reinforcement of the microphone array of the present invention.

[0029] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. Specific embodiments

[0030] The voice enhancement and directional area rendering sound reinforcement system and method based on a microphone array of the present invention can pick up voice signals in two areas simultaneously, and then render the audio content to the specified areas respectively through the rendering methods of two speaker arrays, improving the auditory experience.

[0031] Embodiment of the voice enhancement and directional area rendering sound reinforcement system based on a microphone array:

[0032] Refer to Figure 1 , the voice enhancement and directional area rendering sound reinforcement system based on a microphone array in this embodiment includes a microphone array 1, a voice signal processing device 2, a first speaker array 3, and a second speaker array 4. The microphone array 1 is connected to the voice signal processing device 2, and the voice signal processing system 2 is respectively connected to the first speaker array 3 and the second speaker array 4.

[0033] The connection manners between the microphone array 1, the voice signal processing device 2, the first speaker array 3, and the second speaker array 4 are preferably wired connection manners to improve signal stability. According to actual needs, all or part of them can also be connected by wireless connection manners (such as infrared, Bluetooth, radio frequency).

[0034] The microphone array 1 is used to collect sound signals in the target space.

[0035] The voice signal processing device 2 determines the first voice signal in the first area and the second voice signal in the second area from the sound signals, processes the first voice signal to obtain a first audio output signal, processes the second voice signal to obtain a second audio output signal, and outputs the processed second output audio signal to the second speaker array 4.

[0036] The first speaker array 3 is used to output according to the first audio output signal. The bright zone output by the first speaker array is set in the second area, and the dark zone output by the first speaker array is set in the first area.

[0037] The second speaker array 4 is used to output according to the second audio output signal. The bright zone output by the second speaker array is set in the first area, and the dark zone output by the second speaker array is set in the second area.

[0038] The bright zone (Bright Zone) means that when the sound wave played by the speaker array reaches this area, most of them form constructive interference, enhancing the amplitude of the sound wave, so the sound pressure is relatively large. The dark zone (Dark Zone) means that when the sound wave played by the speaker array reaches this area, most of them form destructive interference, weakening the amplitude of the sound wave, so the sound pressure is relatively small.

[0039] In this embodiment, the target space is a classroom, the first area is the student area, the first voice signal is the student voice signal, the second area is the teacher area, and the second voice signal is the teacher voice signal. Specifically, the student area is the range where students move in the classroom during class, and the teacher area is the range where the teacher moves in the classroom during class. For example, the student area can be the space range encompassing all student desks and chairs, and the teacher area can be the space range encompassing the podium and the blackboard; the student voice signal is the voice signal determined after the sound made by students in the student area is collected by the microphone array 1, and the teacher voice signal is the voice signal determined after the sound made by the teacher in the teacher area is collected by the microphone array 1.

[0040] Those skilled in the art can understand that in addition to the scenario of classroom communication between teachers and students applied in this embodiment, the present invention can also be applied in other similar scenarios in other embodiments, such as the scenario of communication between the keynote speaker and the participants in a conference room.

[0041] The microphone array 1 in this embodiment includes a plurality of microphones, which are fixedly arranged on the ceiling of the classroom, so as to collect the student voice signal in the student area and the teacher voice signal in the teacher area at the same time. Among them, the microphone array 1 includes a main microphone. The difference between the main microphone and other microphones in the microphone array 1 is that it is also used for voice activity detection, so that the voice signal processing device 2 can determine whether there is a normal voice signal in the current target space, that is, whether someone is currently speaking in the classroom.

[0042] In this embodiment, a scenario where both the teacher and students in the classroom are speaking, that is, the voice signal processing device 2 can determine the student voice signal and the teacher voice signal from the sound signals collected by the microphone array 1, will be described. Specifically, after the voice signal processing device 2 determines the student voice signal and the teacher voice signal from the sound signals, it processes the student voice signal according to a preset processing method to obtain a first output audio signal, and processes the teacher voice signal to obtain a second output audio signal. The first output audio signal is output to the first speaker array 3, and the second output audio signal is output to the second speaker array 4.

[0043] The bright area designed to be output by the first speaker array 3 is the teacher area, and the dark area output is the student area. The bright area designed to be output by the second speaker array 4 is the student area, and the dark area output is the teacher area. Thus, the isolation degree between the teacher and the students can be effectively improved, and the interactivity can be increased.

[0044] Specifically, the processing of the sound signals collected by the microphone array 1 is implemented through the voice signal processing system running on the voice signal processing device 2, and the first output audio signal and the second output audio signal are output based on the student voice signal and the teacher voice signal.

[0045] See Figure 2 , the voice signal processing system 21 includes a voice activity detection module 211, a sound source localization module 212, a beamforming module 213, a voice enhancement module 214, a howling suppression module 215, and a mixing module 216.

[0046] The voice activity detection module 211 is used to obtain the acquisition data collected by the main microphone 11, perform voice activity detection (VAD) based on the acquisition data of the main microphone 11, and determine whether there is a normal voice signal in the sound signals collected by the microphone array, that is, determine whether the teacher or the students in the classroom are speaking. Specifically, the acquisition data of each frame of the main microphone 11 will be sent into the voice activity detection module 211. If the acquisition data of the current frame of the main microphone 11 does not detect a normal voice signal, the mixing module 216 does not output the corresponding output audio signal, and the first speaker array 3 and the second speaker array 4 are muted; if the acquisition data of the current frame of the main microphone 11 detects a normal voice signal, it is determined that there is a normal voice signal in the speaker array, and then the sound source localization module 212 starts to work. In this embodiment, the main microphone 11 is an omnidirectional microphone with good sound collection performance.

[0047] Among them, voice activity detection can be achieved through conventional voice activity detection methods. For example, by extracting features from the time-domain audio signal (such as obtaining the frequency-domain signal through Fourier transform), and using spectral characteristics to analyze the fundamental frequency and harmonic components of the voice to determine whether there is a normal voice signal.

[0048] The sound source localization module 212 is used to perform sound source localization (Sound Localization) processing on the sound signals collected by the microphone array 1, so as to determine the direction of the sound source, and determine whether it is specifically from the teacher area or the student area according to the direction of the sound source, that is, to determine the classroom voice signal and the student voice signal. The sound source localization processing can adopt conventional sound source localization methods, such as the GCC-PHAT (Generalized Cross-Correlation with Phase Transform) method or the MUSIC (Multiple Signal Classification) method. The MUSIC method calculates the sound source direction through the incident function and covariance matrix of the microphone array, and the calculation amount is relatively low, which is suitable for relatively open environments. The GCC-PHAT method calculates the sound source position through the cross-correlation signal of the microphone array, and is especially good at dealing with the sound source localization problem in reverberant and noisy environments. Its calculation amount is relatively high, but it performs well in environments with large reverberation and noise.

[0049] The beamforming module 213 is used to adjust the delay and gain of the signals collected by each microphone to enhance the voice signals in a specific direction. The number of delay points can be calculated by combining the sound source position direction information calculated by the sound source localization module 212 with the transfer function of the incident angle.

[0050] Finally, the beamforming module 213 will output audio signals of two channels, corresponding to the teacher voice signal in the teacher area and the student voice signal in the student area respectively. The two-channel audio output at each frequency point can be expressed as (note: the representation symbol of each frequency point k is omitted in this formula):

[0051]

[0052] Among them, are the voice signals in the teacher direction and the student direction respectively, a(θ t ) and a(θ s ) are the transfer functions of the incident angles in the teacher direction and the student direction respectively, s M represents the audio signal of the Mth microphone, · H represents the conjugate transpose operation of the matrix.

[0053] The voice enhancement module 214 is used to perform speech enhancement (SE) processing on the student voice signal and the teacher voice signal after beamforming processing.

[0054] For speech enhancement processing, conventional speech enhancement methods can be adopted. In this embodiment, common sound reinforcement techniques are used. For example, with the dynamic range compressor (DRC) algorithm, the dynamic range of the voice is reduced, and then the overall volume is increased. This method can effectively prevent the digital audio signal from exceeding the maximum amplitude and avoid clipping. The output of the voice enhancement module 214 is expressed as:

[0055]

[0056] When the amplified sound is played through the speaker, it is likely to cause howling. Howling occurs because when the output signal is captured by the microphone and amplified again after being played and output by the speaker after amplification, this process repeats multiple times, resulting in continuous enhancement of the sound and ultimately generating high-frequency howling sounds.

[0057] The howling suppression module 215 is used to perform howling suppression processing on the student voice signal and the teacher voice signal after voice enhancement processing.

[0058] For howling suppression processing, conventional howling suppression methods can be adopted. For example, methods such as adaptive feedback path detection, frequency shift, and notch filters can effectively suppress howling.

[0059] In this embodiment, the howling suppression module 215 finds the frequency band prone to howling and designs a notch filter to reduce the sound in that frequency band. The howling frequency band depends on the impulse response of the listening environment, and the notch filter can be calculated and designed in advance. The output of the howling suppression module is expressed as:

[0060]

[0061] The mixing module 216 is used to perform mixing processing on the student voice signal and the teacher voice signal after howling suppression processing.

[0062] Through the mixing processing, the student voice signal and the teacher voice signal after howling suppression processing are respectively convolved with filters that can improve the isolation degree, and then the resulting signals are assigned to the corresponding speaker arrays for playback. See Figure 3, the mixing module 216 includes a first filter 2161 and a second filter 2162. The student voice signal after howling suppression processing enters the first filter 2161 and then outputs the corresponding first output audio signal to the first speaker array 3. The teacher voice signal after howling suppression processing enters the second filter 2162 and then outputs the corresponding second output audio signal to the second speaker array 4.

[0063] Specifically, the mixing module 216 is used to render the voice signal into two regions. One region has a larger sound pressure and is called the bright region, and the other region has a smaller sound pressure and is called the dark region. Different from the traditional beamforming rendering technology, this scheme is not restricted by the regional position, and the bright region and the dark region can be set at any position. The method of multi-zone reconstruction can refer to the paper "Spatial multizone soundfield reproduction: Theory and design".

[0064] See Figure 4 , through mixing processing, when rendering the student voice signal after howling suppression processing in the first speaker array 3, the bright region is designed in the student area and the dark region is designed in the teacher area.

[0065] See Figure 5 , through mixing processing, when rendering the teacher voice signal after howling suppression processing in the second speaker array 4, the bright region is designed in the teacher area and the dark region is designed in the student area.

[0066] Let the room impulse response of each speaker in the first speaker array 3 to the teacher area be H t , and the impulse response to the student area be H s . The first filter for rendering and enhancing the student voice in the teacher area is C s , and the second filter for rendering and enhancing the teacher voice in the student area is called C t .

[0067] Through the following formula for optimization, C can be deduced s (The representation of each frequency point k is omitted in the formula):

[0068]

[0069] s.t. H t C s x s > H s C s x s +σ;

[0070] Similarly, through the following formula for optimization, C can be deduced t (The representation of each frequency point k is omitted in the formula):

[0071]

[0072] s.t. H s C t x t > H t C t x t + σ;

[0073] where σ is the minimum threshold for comparing the sound pressure amplitudes of two regions.

[0074] Finally, the first output audio signal obtained by the first speaker array 3 is: y1 = C s x s and the output audio signal of the second speaker array 4 is: y2 = C t x t .

[0075] Embodiment of a voice enhancement and directional area rendering sound reinforcement system based on a microphone array:

[0076] This embodiment is implemented based on the above embodiment. First, the main microphone of the microphone array is used to detect whether there is a normal voice signal. If no normal voice signal is detected, the output channel will be set to zero and the speakers will be muted. If a voice signal is detected, first analyze the signals picked up by the microphones in the array to find the sound source positions of the teacher and the students. After obtaining the position information, beamforming filters can be designed for each direction to enhance the voice signal in that direction while suppressing interference signals from other directions and improving the signal-to-noise ratio. Then, use the voice enhancement module to amplify the voice frequency band signals output by beamforming, and then use the howling suppression module to reduce the signals in the howling frequency band. Finally, the signals are convolved with the designed filters through the mixing module to improve the isolation degree and output and feed them to the corresponding speaker arrays for playback.

[0077] See Figure 6 , the voice signal processing system is implemented through a computer program and includes the following steps:

[0078] S101: Determine whether the main microphone detects a normal voice signal.

[0079] Among them, when it is determined that the main microphone detects a normal voice signal, it means that someone is speaking in the classroom, and then step S102 is continued; when it is determined that the main microphone does not detect a normal voice signal, it means that no one is speaking in the classroom at this time, and then step S108 is executed to control the first speaker and the second speaker to remain muted.

[0080] S102: Determine the first voice signal and the second voice signal through sound source localization processing, and determine the sound source position and direction information of the first voice signal and the second voice signal.

[0081] Among them, through sound source localization processing of the sound signals collected by the microphone array, the student voice signal and the teacher voice signal are determined from the sound signals, and the sound source position and direction information of the first voice signal and the second voice signal are determined.

[0082] S103: Perform beamforming processing on the first voice signal and the second voice signal according to the sound source position and direction information of the first voice signal and the second voice signal.

[0083] Among them, according to the sound source position and direction information corresponding to the student voice signal, beamforming processing can be performed on the student voice signal; according to the sound source position and direction information corresponding to the teacher voice signal, beamforming processing can be performed on the teacher voice signal.

[0084] S104: Perform voice enhancement processing on the first voice signal and the second voice signal after beamforming processing.

[0085] Among them, perform voice enhancement processing on the student voice signal after beamforming processing, and perform voice enhancement processing on the teacher voice signal after beamforming processing.

[0086] S105: Perform howling suppression processing on the first voice signal and the second voice signal after voice enhancement processing.

[0087] Among them, perform howling suppression processing on the student voice signal after voice enhancement processing, and perform howling suppression processing on the teacher voice signal after voice enhancement processing.

[0088] S106: Perform mixing processing on the first voice signal and the second voice signal after howling suppression processing to obtain the first output audio signal and the second output audio signal.

[0089] Among them, perform mixing processing on the student voice signal after howling suppression processing to obtain the first output audio signal; perform mixing processing on the teacher voice signal after howling suppression processing to obtain the second output audio signal.

[0090] S107: Output the first output audio signal to the first speaker array, and output the second output audio signal to the second speaker array.

[0091] Thus, after the first output audio signal is output to the first speaker array, the words spoken by students in the student area can be clearly heard by teachers in the teacher area; after the second output audio signal is output to the second speaker array, the words spoken by teachers in the teacher area can be clearly heard by each student in the student area.

[0092] In summary, by processing the sound signal, the present invention can accurately identify the sound source position and ensure clear sound transmission. In addition, through the rendering method of two speaker arrays, that is, rendering the audio content to the specified area respectively, it realizes the isolation when people in the first area and the second area talk to each other, and can simultaneously pick up the speeches of people in the first area and the second area, making the interaction more convenient and efficient, avoiding the phenomenon of sound wave interference, and ensuring the clarity and intelligibility of speech. The present invention can also make the sound field more uniform and improve the overall auditory experience.

[0093] Finally, it should be emphasized that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A voice enhancement and directional area rendering sound reinforcement system based on a microphone array, characterized in that Including: A microphone array, a voice signal processing device, a first speaker array, and a second speaker array; A voice signal processing system runs on the voice signal processing device; The microphone array is used to collect sound signals in the target space; The voice signal processing system is used to determine a first voice signal in a first area and a second voice signal in a second area from the sound signals, and process the first voice signal to obtain a first audio output signal and process the second voice signal to obtain a second audio output signal; The first speaker array is used to output according to the first audio output signal. The bright area output by the first speaker array is set in the second area, and the dark area output by the first speaker array is set in the first area; The second speaker array is used to output according to the second audio output signal. The bright area output by the second speaker array is set in the first area, and the dark area output by the second speaker array is set in the second area.

2. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to claim 1, wherein: The microphone array includes a main microphone, and the voice signal processing system includes a voice activity detection module; The voice activity detection module is used to determine whether there is a normal voice signal in the sound signal according to the acquisition data of the main microphone, and after determining that there is the normal voice signal in the sound signal, determine the first voice signal and / or the second voice signal from the sound signal.

3. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to claim 1, wherein: The voice signal processing system includes a sound source localization module and a beamforming module, and the sound source localization module is connected to the beamforming module; The sound source localization module is used to determine the first voice signal and the second voice signal from the sound signal, and the beamforming module is used to perform beamforming processing on the determined first voice signal and second voice signal according to the sound source position and direction information obtained by the sound source localization module to obtain the first output audio signal and the second output audio signal respectively.

4. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to claim 3, wherein: The voice signal processing system includes a voice enhancement module, and the voice enhancement module is connected to the beamforming module; The voice enhancement module is used to perform voice enhancement processing on the first voice signal and the second voice signal after beam processing respectively to obtain the first output audio signal and the second output audio signal respectively.

5. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to claim 4, wherein: The voice signal processing system includes a howling suppression module, and the howling suppression module is connected to the voice enhancement module; The howling suppression module is configured to perform howling suppression processing on the first voice signal and the second voice signal after voice enhancement processing respectively, so as to obtain the first output audio signal and the second output audio signal respectively.

6. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to any one of claims 1 to 5, wherein: The target space is a classroom, the first area is the student area, the first voice signal is the student voice signal, the second area is the teacher area, and the second voice signal is the teacher voice signal.

7. The microphone array-based voice enhancement and directional area rendering sound reinforcement system according to claim 6, wherein: The microphone array is arranged on the ceiling of the classroom.

8. A voice enhancement and directional area rendering sound reinforcement method based on a microphone array, characterized in that, It includes the following steps: Obtain the sound signal in the target space from the microphone array; Determine the first voice signal in the first area and the second voice signal in the second area in the target space from the sound signal; Process the first voice signal to obtain a first output audio signal; Process the second voice signal to obtain a second output audio signal; Output the first output audio signal to the first speaker array, the bright area output by the first speaker array is arranged in the second area, and the dark area output by the first speaker array is arranged in the first area; Output the second output audio signal to the second speaker array, the bright area output by the second speaker array is arranged in the first area, and the dark area output by the second speaker array is arranged in the second area.

9. The microphone array-based voice enhancement and directional area rendering sound reinforcement method according to claim 8, wherein: After obtaining the sound signal in the target space from the microphone array, determine whether there is a normal voice signal in the acquisition data of the main microphone of the microphone array. When it is determined that there is a normal voice signal, then determine the first voice signal in the first area and the second voice signal in the second area in the target space from the sound signal.

10. The microphone array-based voice enhancement and directional area rendering sound reinforcement method according to claim 8, wherein: When processing the first voice signal and processing the second voice signal, both include beamforming processing, voice enhancement processing, howling suppression processing, and mixing processing performed in sequence.