Sound reinforcement in multiple sound zone environments
Adaptive filters in vehicle multimedia systems address reverberation and feedback issues by enhancing microphone signals, improving communication and karaoke experiences through echo and feedback cancellation, ensuring high-quality sound reinforcement in multiple sound zones.
Patent Information
- Application Number
- JP2024534696
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-30
- Filing Date
- 2022-02-17
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-02-17
AI Technical Summary
Modern vehicle multimedia systems face issues with reverberation and feedback in voice communication and karaoke systems due to sound being picked up by microphones from loudspeakers, leading to poor communication and entertainment quality.
Implementing acoustic echo cancellation (AEC) and acoustic feedback cancellation (AFC) using adaptive filters to process microphone signals, enhancing speech and audio signals for playback in vehicle environments, and applying sound reinforcement techniques to improve communication and karaoke experiences.
Reduces reverberation and feedback, enhancing voice communication and karaoke performance by improving signal quality and stability, allowing for effective sound reinforcement in multiple sound zones.
Smart Images

Figure 0007734850000004 
Figure 0007734850000005 
Figure 0007734850000006
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Application Serial No. 63 / 295,062, filed December 30, 2021, the disclosure of which is incorporated herein by reference in its entirety.
[0002] Aspects of the present disclosure generally relate to sound reinforcement in a multiple sound zone environment. [Background technology]
[0003] Modern vehicle multimedia systems often include an in-vehicle communication (voice processor) system to improve communication between occupants, especially when high background noise levels are present. It is particularly important to provide a means to improve communication between rear-seat and front-seat passengers. To improve communication, passenger speech is recorded by one or more microphones and played back through loudspeakers positioned in close proximity to the listening passengers. As a result, the sound emitted from the loudspeakers is picked up by the microphones, which can result in reverberation / echo and feedback. Loudspeakers may also be used to reproduce audio signals from audio sources such as radios, compact disc (CD) players, and navigation systems. Again, these audio signal components are picked up by the microphones and output through the loudspeakers.
[0004] Additionally, vehicle passengers may seek entertainment while traveling. For this purpose, karaoke systems may be installed in the vehicle. Such karaoke systems suffer from the same drawbacks as vehicle voice processor systems: playback of the voices from the singing passengers is prone to reverb and feedback. Summary of the Invention
[0005] In one or more illustrative examples, a microphone signal is received from at least one microphone. Acoustic echo cancellation (AEC) is performed on the microphone signal to generate an echo-canceled microphone signal. The AEC uses a first adaptive filter to estimate and cancel feedback resulting from the environment. Acoustic feedback cancellation (AFC) is performed on the echo-canceled microphone signal to generate an echo- and feedback-canceled microphone signal. The AFC uses a second adaptive filter to estimate and cancel feedback resulting from application of a reinforcement voice signal in the environment. Speech from the echo- and feedback-canceled microphone signal is applied to generate a reinforcement voice signal. The reinforcement voice signal and the audio signal are applied to a speaker for playback in the environment.
[0006] In one or more exemplary embodiments, a method for processing a sound signal in a vehicle multimedia system is provided. A microphone signal is received from at least one microphone. The microphone signal includes a first voice signal component corresponding to spoken speech, a second voice signal component corresponding to an augmented voice signal to be reproduced by a speaker in an environment, and an audio signal component corresponding to an audio signal to be reproduced by a loudspeaker. AEC of the microphone signal is performed to generate an echo-canceled microphone signal, the AEC using a first adaptive filter to estimate and cancel feedback resulting from the environment. AFC of the echo-canceled microphone signal is performed to generate a processed microphone signal, the AFC using a second adaptive filter to estimate and cancel feedback resulting from application of the augmented voice signal in the environment. The speech in the processed microphone signal is augmented to generate the augmented voice signal. The augmented voice signal and the audio signal are applied to the loudspeaker for reproduction in the environment.
[0007] In one or more example embodiments, a non-transitory computer-readable medium includes instructions for sound signal processing in a vehicle multimedia system that, when executed by a voice processor system, include receiving a microphone signal from at least one microphone, the microphone signal including a first voice signal component corresponding to spoken speech, a second voice signal component corresponding to an augmentation voice signal to be reproduced by a loudspeaker in an environment, and an audio signal component corresponding to the audio signal to be reproduced by the loudspeaker; performing AEC of the microphone signal to generate an echo-canceling microphone signal, the AEC using a first adaptive filter to estimate and cancel feedback resulting from the environment; performing AFC of the echo-canceling microphone signal to generate a processed microphone signal, the AFC using a second adaptive filter to estimate and cancel feedback caused by applying the augmentation voice signal in the environment; augmenting the spoken speech with the processed microphone signal to generate an augmentation voice signal, and applying the augmentation voice signal and the audio signal to the loudspeaker for reproduction in the environment. [Brief explanation of the drawings]
[0008] [Figure 1] Figure 1 shows an example of a multi-channel sound system providing sound reinforcement in an environment with multiple sound zones; [Figure 2] Figure 2 shows further aspects of the operation of the voice processor system; [Figure 3] FIG. 3 is a partial view of an example of a multi-channel sound system showing an example of electroacoustic feedback within the multi-channel sound system; [Figure 4] FIG. 4 is a partial view of an example multi-channel sound system illustrating the use of acoustic feedback cancellation to combat electro-acoustic feedback within the multi-channel sound system; [Figure 5]FIG. 5 is a diagram illustrating an example of a portion of a multi-channel sound system showing step size control for acoustic feedback cancellation with artificial reverberation; [Figure 6] Figure 6 shows an example graph of the local speech and loudspeaker signals showing artificially added reverberation; [Figure 7] FIG. 7 illustrates an example process for providing sound reinforcement in an environment with multiple sound zones. [Figure 8] FIG. 8 illustrates an example process for the operation of acoustic feedback cancellation in an audio processing system. DETAILED DESCRIPTION OF THE INVENTION
[0009] Where necessary, detailed embodiments of the present invention are disclosed herein; however, it should be understood that the disclosed embodiments are merely exemplary of the invention, which may be embodied in various alternative forms. The figures are not necessarily to scale, and some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be construed as limiting, but merely as a representative basis for teaching those skilled in the art how to employ the present invention in various ways.
[0010] 1 illustrates an example of a multi-channel sound system 100 for providing sound reinforcement in an environment 102 having multiple sound zones 104. The multi-channel sound system 100 may include an audio source 106, a loudspeaker 108, a microphone 110, a voice processor system 114, and a sound reinforcement application 120. The sound reinforcement application 120 may be programmed to control the voice processor system 114 to facilitate sound reinforcement in the environment 102. As described in detail herein, the sound reinforcement application 120 may activate and control functions for signal processing that cause the voice processor system 114 to utilize amplification and reverb or other sound effects to reinforce voice signals captured by the microphone 110 in the multiple sound zone environment 102. The reinforcement may include localizing the voice signal within the multiple sound zone environment 102, identifying the loudspeaker 108 closest to the speaker, and using that feedback to reinforce the sound output using the identified loudspeaker 108.
[0011] The environment 102 may be a room or other enclosed area, such as a concert hall, stadium, restaurant, auditorium, or vehicle cabin. In another example, the environment 102 may be an outdoor or at least partially unenclosed area or structure, such as an amphitheater or stage. In many examples, the environment 102 may include multiple sound zones 104. A sound zone 104 may refer to an acoustic section of the environment 102 that can reproduce different sounds. Using a vehicle as an example, the environment 102 may include a sound zone 104 for each seating position within the vehicle.
[0012] Audio source 106 may be any form of one or more devices capable of generating and outputting different media signals including one or more channels of audio. Examples of audio source 106 may include a media player (such as a compact disc, video disc, digital versatile disc (DVD), or Blu-ray disc player), a video system, a radio, a cassette tape player, a wireless or wired communication device, a navigation system, a personal computer, a portable music player device, a mobile phone, a musical instrument such as a keyboard or electric guitar, or any other form of media device capable of outputting a media signal.
[0013] The loudspeakers 108 can include a variety of devices configured to convert electrical signals into acoustic signals. The loudspeakers 108 can be positioned throughout the environment 102 to provide acoustic output across various sound zones 104 of the environment 102. Among some possibilities, the loudspeakers 108 can include dynamic drivers operating in a magnetic field and having a coil connected to a diaphragm, where application of an electrical signal to the coil causes the coil to move via induction, powering the diaphragm. Among other possibilities, the loudspeakers 108 can include other types of drivers, such as piezoelectric, electrostatic, ribbon, or planar elements. As an example, each of the sound zones 104 can be associated with one or more of the loudspeakers 108 to provide an audible output to the respective sound zone 104.
[0014] The microphones 110 may include various devices configured to convert acoustic signals into electrical signals. These electrical signals may be referred to as microphone signals 112. Microphones 110 may also be positioned throughout the sound zones 104 of the environment 102 to capture audio input from users throughout the multi-channel sound system 100. For example, microphones 110 may be available in the multi-channel sound system 100 to provide voice communication, such as hands-free phone calls and / or interaction with a voice assistant application. As an example, each of the sound zones 104 may include a microphone 110 or an array of microphones 110 to capture audio from the respective sound zones 104. In an embodiment, multiple microphones 110 may be positioned at each sound zone 104 to obtain beamformed signals for the respective sound zones 104. This allows the voice processor system 114 to receive directional detected sound signals for the respective sound zones 104 (e.g., when a speaker is detected within the sound zone 104). The beamformed signals may be used to derive information about whether a user is actively speaking in each sound zone 104. Additional voice activity detection techniques, such as changes in the energy, spectrum, or cepstral distance of the captured microphone signal 112, may also be used to determine whether a speaker is present.
[0015] The voice processor system 114 may be configured to use the loudspeaker 108 and the microphone 110 for sound reinforcement in the environment 102. The voice processor system 114 may be configured to receive a microphone signal 112 from the microphone 110 that may be used by the voice processor system 114 to identify audio content in the environment 102. The voice processor system 114 may also be configured to receive a reference signal 116 from the audio source 106 that is indicative of the audio to be reproduced by the loudspeaker 108.
[0016] As described in further detail below, the voice processor system 114 may use the reference signal 116 to perform AEC and / or AFC on the microphone signal 112 to generate a processed microphone signal 118. The processed microphone signal 118 may be provided to a sound reinforcement application 120.
[0017] In a vehicle use case, the sound reinforcement application 120 can support communication between sound zones 104. For example, a passenger in the vehicle can use the voice processor system 114 to communicate between the front and rear seats. In such an example, the sound reinforcement application 120 can instruct the voice processor system 114 to generate a sound processor output signal 122 including the passenger's voice for playback to other passengers in the vehicle via the loudspeakers 108.
[0018] In another example, the sound reinforcement application 120 may support use of the voice processor system 114 as a sound monitor. For example, a passenger in a vehicle may use the voice processor system 114 to sing karaoke. In such an example, the sound reinforcement application 120 may instruct the voice processor system 114 to provide a sound processor output signal 122 including the passenger's voice to the same passenger in the vehicle for playback over the loudspeakers 108. Further details of an embodiment of karaoke in a vehicle environment are described in detail in European Patent Application EP 2018034 B1, filed July 16, 2007, entitled "METHOD AND SYSTEM FOR PROCESSING SOUND SIGNALS IN A VEHICLE MULTIMEDIA SYSTEM," the disclosure of which is incorporated herein by reference in its entirety.
[0019] The audio processor output signal 122 may be applied to a summer 124 along with a reference signal 116 from the audio source 106, and the combined output to the summer 124 is provided to the speaker 108 for playback.
[0020] FIG. 2 illustrates further aspects of the operation of the voice processor system 114. As shown in FIG. 2, and with continued reference to FIG. 1, the voice processor system 114 can apply various types of speech enhancement (SE) 202 to the microphone signal 112. SE 202 can be performed at the beginning of voice processing to improve the quality of the received voice signal. These SE 202 can include techniques such as noise reduction, equalization, noise-dependent gain control, adaptive gain control, etc. These processed microphone signals 118 may be provided to a voice reinforcement application 120 for processing.
[0021] The speech reinforcement application 120 may be configured to control the mixer 204. The mixer 204 may be configured to receive the enhanced microphone signals 112 from the SE 202 module and apply gain to the received microphone signals 112 under the direction of the speech reinforcement application 120. For example, the speech reinforcement application 120 may instruct the mixer 204 to pass one or more of the microphone signals 112 for amplification and playback by the loudspeakers 108. The output of the mixer 204 may be referred to as speech reinforcement.
[0022] The sound reinforcement application 120 may be configured to control the application of one or more vocal effects 206 to the mixer 204 output. These effects may include, for example, reverb, chorus, etc. applied to the sound reinforcement output of the mixer 204. The results of the vocal effects 206 may be referred to as per-channel sound outputs 208. In some multi-channel sound systems 100, multi-channel effects 210 may be applied to the per-channel sound outputs 208 for playback within the environment 102. These multi-channel effects 210 may include panning, doubling, etc., as some examples. After mixing and applying the effects, the result may be provided as an audio processor output signal 122 for playback by the loudspeakers 108. Some sound effects (e.g., vocal effects 206) may be applied by single-channel processing to keep central processing unit (CPU) and memory costs low. Other effects may be applied as multi-channel effects 210 to enrich the listening experience.
[0023] The voice processor system 114 may also perform signal processing to improve system stability to compensate for acoustic feedback in the closed acoustic loop of the environment 102. In one example, the voice processor system 114 may utilize an AEC 212 to counteract feedback that is a result of the environment 102.
[0024] As discussed herein, the microphone signal 112 may include speech received from a user within the sound zone 104 of the environment 102. However, the microphone signal 112 may also capture sound output from the loudspeaker 108. This output of the loudspeaker 108, as sensed at least in part by the microphone 110, may be referred to as echo. In response, the AEC 212 may receive a reference signal 116 from the audio source 106 that is indicative of the sound being reproduced by the loudspeaker 108. Due to the slow propagation speed of sound compared to electrical signals, the AEC 212 may receive the reference signal 116 earlier in time than the echo captured in the microphone signal 112.
[0025] The AEC 212 can apply an adaptive filter to the reference signal 116 that estimates the linear acoustic impulse response of the loudspeakers 108 in the environment 102 to the microphone 110 system. Based on this echo estimate, the AEC 212 can generate an echo cancellation signal that is added to the microphone signal 112 to reduce echo. In one example, the AEC 212 can be performed on each channel of the reference signal 116 to generate a channel echo cancellation signal. These channel signals are applied to a summer 214 to generate an overall echo cancellation signal. This overall echo cancellation signal is applied to each of the microphone signals 112 via a summer 216.
[0026] The voice processor system 114 can utilize the AFC 218 to combat feedback resulting from the operation of the voice reinforcement application 120 to enhance the voice signal in the environment 102. For each microphone 110, the AFC 218 component can receive the echo cancellation microphone signal 112 corresponding to that microphone 110. The AFC 218 can also receive the per-channel audio output 208 of the vocal effects 206 as a reference. The AFC 218 can apply an adaptive filter to estimate the acoustic impulse response of the loudspeakers 108 in the environment 102 to the microphone 110 system of the per-channel audio output 208. Based on the estimate, the AFC 218 can generate a feedback cancellation signal that is summed by a summer 220 with the microphone signal 112 input to the SE 202 to combat feedback. Further aspects of the operation of the AFC 218 are described in detail below with reference to FIGS. 3-6.
[0027] In some examples, the voice reinforcement application 120 may be controllable using a voice interface using input from the microphone 110. However, the microphone signal 112 may additionally include acoustic echo of the playback of the audio source 106 and acoustic feedback of the voice playback (subject to reverberation or other effects) from the voice processor output signal 122. If a passenger stops singing and wishes to use a voice assistant (as one example), the voice processor system 114 and its vocal effects 206 and multi-channel effects 210 may continue to run. These effects may degrade voice recognition performance. Thus, the described voice processor system 114 may provide the processed microphone signal 118 to the voice reinforcement application 120 before the vocal effects 206 and / or multi-channel effects 210 are applied, but after echo suppression and voice feedback compensation using the AEC 212, after effects via the AFC 218, and after voice enhancements that may improve voice recognition performance due to their noise removal, signal conditioning, etc.
[0028] Using the processed microphone signal 118, the voice reinforcement application 120 can determine the sound zone 104 (shown in FIG. 1 ) of the speaking user and the user-specific speech dialog can engage that sound zone 104. In one example, automatic speech recognition (ASR) may be used to control the voice reinforcement application 120, such as to skip a song, repeat a song, repeat a section, adjust vocal effects 206 and / or multi-channel effects 210, add a user for voice reinforcement, turn a user off for voice reinforcement, turn off voice reinforcement for all users, request that a voice processor mode be turned on to send voice to other users, etc.
[0029] The voice processor system 114 can be configured to support any subset of the sound zones 104 that utilize sound reinforcement. For example, selected sound zones 104 can be added to or removed from the sound reinforcement. This can be accomplished by a user configuring the mixer 204 to pass a selected subset of its processed microphone signals 118 using the audio interface of the sound reinforcement application 120 or other user interface. Thus, a user can select from or ignore the processed microphone signals 118 from particular sound zones 104. In one example, the sound reinforcement application 120 can be used to control the mixer 204 to support two or more singers simultaneously, enabling duets and polyphonic performances.
[0030] In some examples, the voice reinforcement application 120 can provide performance quality assessment. For example, speaker separation may be applied to separate each user's vocal signal. This separated vocal signal (which may include singing) may be used for performance assessment (e.g., pitch estimation and assessment against a reference pitch). These assessments may be performed individually for each individual sound zone 104 or user. For example, performances from multiple users may be compared across participants across multiple sound zones 104. The best singer may be detected as the singer who, on average, is closest to the reference pitch in the audio content played from the audio source 106.
[0031] When a set of loudspeakers 108 in the environment 102 is used for playback of audio sources 106 as well as for reinforced audio playback, it may be possible to combine hardware implementing the AEC 212 and AFC 218 functions. However, in many applications, a different set of loudspeakers 108 may be used for echo cancellation compared to feedback cancellation, and therefore the channels for the AEC 212 and the AFC 218 may be different. For example, there may be many loudspeakers 108 in the environment 102 used for audio playback, but utilizing all of these loudspeakers 108 for audio reinforcement may be impractical due to their processing requirements. As a result, a common adaptive filter may not be a viable solution, and separate adaptive filters with separate adaptive controls may be used for the AEC 212 and AFC 218 functions.
[0032] As shown in Figure 2, the illustrated voice processor system 114 incorporates separate methods for AEC 212 and AFC 218. Thus, the voice processor system 114 can use a different subset of the loudspeakers 108 in the environment 102 for sound reinforcement compared to entertainment playback (as shown in the example of Figure 2, half of the loudspeaker outputs 222 to the loudspeakers 108 are used for sound reinforcement and the other half are unused). Music is processed by AEC 212, and voice is processed by AFC 218 (and / or other methods, such as feedback suppression).
[0033] As mentioned above, the sound reinforcement application 120 may be configured to perform sound processor functions. In such an example, the sound reinforcement application 120 may select a loudspeaker 108 that is far in the environment 102 from a user speaking into the microphone 110. This may be done to avoid acoustic feedback of the loudspeaker 108 returning to the microphone 110 in combination with the sound reinforcement.
[0034] However, in voice reinforcement use cases such as karaoke, it may be desirable to provide voice reinforcement using a loudspeaker 108 local to the speaking user. For example, a singer may want to use the loudspeaker 108 as a sound monitor to hear their own voice. In such an example, the voice reinforcement application 120 can determine the sound zone 104 corresponding to the user and direct sound reinforcement to the loudspeaker 108 in the corresponding sound zone 104. In voice reinforcement, the distance between the loudspeaker 108 and its associated open microphone 110 is small compared to the distance in voice processor use cases. This can result in high acoustic coupling and increased risk of instability. Therefore, additional aspects may be required to counter acoustic feedback in karaoke or other voice reinforcement applications 120 where the speaker is close to the loudspeaker 108. These additional aspects may include, for example, step size control for acoustic feedback cancellation through artificially added reverberation (or other vocal effects 206).
[0035] 3 shows an example portion 300 of the multi-channel sound system 100 illustrating an example of electro-acoustic feedback within the multi-channel sound system 100. With electro-acoustic feedback, the voice processor system 114 may operate in a closed electro-acoustic loop. If the gain of the voice processor system 114 exceeds the stability limit of the multi-channel sound system 100, instability may occur. Mathematically, the transfer function of a resonance is defined as follows: JPEG0007734850000001.jpg18147 here f is the continuous frequency of resonance. S(f) is the local vocalization signal from the user in the sound zone 104. x(f) is the signal from the speaker 108. H(f) is the transfer function of the path between the speaker 108 and the microphone 110 . H icc(f) is the transfer function of the voice processor system 114; In such an instance, the stability limit is mathematically defined as follows: JPEG0007734850000002.jpg12147Therefore, as long as the open-loop gain is less than unity, the system is likely to be stable.
[0036] 4 shows an example portion 400 of the multi-channel sound system 100 illustrating an example of the use of the AFC 218 to combat electro-acoustic feedback within the multi-channel sound system 100. The cancellation of the acoustic feedback may be performed, in one example, by estimating the impulse response of the environment 102 using an adaptive filter (e.g., a normalized least mean squares (NLMS) algorithm).
[0037] 4 more specifically, let n be a discrete time index. s(n) may refer to, for example, a local speech signal from a user in the sound zone 104. ŝ(n) may refer to an estimate of the local speech signal (with feedback removed). x(n) may refer to the loudspeaker output 222 signal for driving the loudspeaker 108. h(n) may refer to the actual impulse response from the loudspeaker 108 to the microphone 110. ĥ(n) refers to an estimate of the impulse response from the loudspeaker 108 to the microphone 110. icc The (n) may refer to the impulse response of the voice processor system 114. It should be noted that in other examples, the adaptive filter algorithm may be implemented in the frequency domain, for example, using frequency domain signal processing.
[0038] In general, the adaptive filter converges best when s(n) and x(n) are orthogonal. However, in speech reinforcement applications, the local speech may be intentionally equal to, or at least highly correlated with, the signal to the loudspeaker 108. In such conditions, the high correlation between the local signal and the excitation signal can cause the adaptive filter to converge toward the bias.
[0039] 5 is a diagram illustrating an example portion 500 of a multi-channel sound system 100, illustrating step size control for acoustic feedback cancellation with artificial reverberation. Reverberation is an important vocal effect 206 used in various styles of music. Therefore, the sound of a voice reinforcement application 120 can be improved by adding artificial reverberation to the speaker or singer's voice. This reverberation effect can be applied to the microphone signal 118 processed in the voice processor system 114 by the vocal effect 206, as described above.
[0040] Importantly, the artificially added reverberation is used to improve the convergence of the adaptive filters used for feedback cancellation. When the singer stops, only the reverberation is played back through the loudspeaker 108. Mathematically, JPEG0007734850000003.jpg20147
[0041] 6 shows an example graph 600 of local speech s(n) and speaker signal x(n). Importantly, speaker 108 continues to generate artificial reverberation for a period of time after the speaker goes silent. When local speech stops, the reverberant energy from vocal effects 206 decays exponentially. During this time, there is still a signal from loudspeaker 108, but no local speech. During this reverberant period when the user is not speaking or singing, s(n) and x(n) are uncorrelated.
[0042] By using this remaining reverberation energy when the speaker is silent, an adaptive algorithm such as NLMS can quickly converge to a desired solution during these periods. A step-size control mechanism can be utilized to increase the adaptation process during reverberant periods and slow down the adaptation process during localized speaking / singing. For example, if reverberation is detected in the microphone signal 112 and / or if no speech is detected in the microphone signal 112, the adaptation step size can be increased to allow the adaptive algorithm to converge. However, if speech is detected in the microphone signal 112, the adaptation step size can be slowed to reduce the chance of converging to a bias due to high correlation between the local signal and the excitation signal.
[0043] With this additional enhancement, the reverb applied to the processed microphone signal 118 can be used to both improve the target sound of the voice reinforcement and to improve the overall operation of the AFC 218. It should be noted that although this step size control technique is described with respect to reverb, it is possible to implement a similar technique based on the use of other effects such as delay or chorus.
[0044] 7 shows an example process 700 for providing sound reinforcement in an environment 102 having multiple sound zones 104. In one example, the process 700 may be performed by the voice processor system 114 in the context of a multi-channel sound system 100. For example, the process 700 may be performed by the voice processor system 114 to provide karaoke in a vehicle environment 102.
[0045] At operation 702, the voice processor system 114 receives audio from an audio source 106. The audio source 106 may be any form of one or more devices capable of generating and outputting different media signals including one or more channels of audio. The audio from the audio source 106 may be received as a reference signal 116 for processing by the voice processor system 114.
[0046] In operation 704, the voice processor system 114 receives the microphone signal 112 from the microphone 110. In one example, each of the sound zones 104 may include a microphone 110 or an array of microphones 110 for capturing the voice signal of the respective sound zone 104.
[0047] At operation 706, the voice processor system 114 executes the AEC 212 to generate an echo-canceled microphone signal. In one example, the AEC 212 can apply an adaptive filter to the microphone 110 system to estimate the linear acoustic impulse response of the speaker 108 in the environment 102 relative to the reference signal 116. Based on this echo estimate, the AEC 212 can generate an echo-canceling signal that is added to the microphone signal 112 to reduce echo.
[0048] At operation 708, the voice processor system 114 performs AFC 218 on the echo-canceled microphone signal. In one example, the voice processor system 114 may utilize the AFC 218 to combat feedback that is a result of the operation of a voice reinforcement application 120 to enhance the voice signal in the environment 102. The AFC 218 may generate a processed microphone signal 118 for further use. Further aspects of the operation of the AFC 218 are discussed with respect to FIG. 8 below.
[0049] At operation 710, the voice processor system 114 generates speech reinforcement. In one example, the voice reinforcement application 120 may receive commands from a user of the voice processor system 114 in the environment 102. These commands may enable the voice reinforcement application 120 to configure the mixer 204 to generate speech reinforcement for one or more users in one or more sound zones 104. For example, the voice reinforcement application 120 may instruct the mixer 204 to pass one or more microphone signals 112 for amplification and playback by the loudspeakers 108.
[0050] At operation 712, the voice processor system 114 applies vocal effects 206 to the audio enhancements to generate per-channel audio outputs 208. In many embodiments, these vocal effects 206 can include reverb. Additionally or alternatively, these vocal effects 206 can include chorus, pitch correction, introduction of sound effects, etc.
[0051] In operation 714, the voice processor system 114 provides the loudspeaker output 222 and the audio from the audio source 106 to the loudspeaker 108 for playback in the environment 102. Thus, users in the sound zones 104 of the environment 102 can enjoy playback of the audio enhancements with minimal feedback.
[0052] After operation 714, process 700 ends. Although process 700 is depicted as a linear process, it should be noted that process 700 may be performed sequentially. It should also be noted that one or more operations of process 700 may be performed simultaneously and / or out of order from the description of process 700.
[0053] 8 shows an exemplary process 800 for operation of the AFC 218 of the voice processor system 114. Similar to the process 700, the process 800 may be performed by the voice processor system 114 in the context of the multi-channel sound system 100.
[0054] At operation 802, similar to operation 704, the voice processor system 114 receives the microphone signal 112. At operation 804, the voice processor system 114 determines whether reverberation is present in the microphone signal 112 and / or whether an absence of speech is detected. Determining whether reverberation is present can be done using various techniques. As an example, determining whether reverberation is present can involve measuring the persistence of sound (echo), such as measuring how quickly the sound level drops when a loud sound is emitted (e.g., the time it takes for the sound energy to drop by 60 dB or another factor). Determining whether speech is present in the microphone signal 112 can be performed using various techniques described herein, such as capturing beamforming signals for each sound zone 104 location to determine the location of the speakers, or analyzing the microphone signal 112 to identify changes in the energy, spectrum, or septal distance of the captured microphone signal 112. If reverberation and / or speech are not detected, control proceeds to operation 806. However, if speech is detected, control passes to operation 808.
[0055] In operation 806, the voice processor system 114 increases the step size of the adaptive algorithm of the AFC 218. At this point, the microphone signal 112 may not contain any speech, but it may still contain reverberant energy applied by the vocal effects 206. Because this signal is no longer correlated with local speech, an adaptive algorithm such as NLMS can quickly converge to a desired solution during this time. In this way, the reverberant effect added to improve voice quality can be used to improve the tuning of the AFC filter via reverberation-based step size control. After operation 806, control returns to operation 802.
[0056] In operation 808, the voice processor system 114 decreases the step size of the adaptation algorithm of the AFC 218. Thus, if speech is detected in the microphone signal 112, the adaptation step size can be slowed down to reduce the likelihood of converging towards a bias due to high correlation between the local signal and the excitation signal. After operation 808, control returns to operation 802.
[0057] The signal processing means described herein may be implemented as software on a digital signal processor, or may be provided as a separate processing chip implemented, for example, on a card connectable to the multimedia bus system of a computing device, or in any other form known to those skilled in the art.
[0058] Computing devices described herein generally include computer-executable instructions, which may be executable by one or more computing devices such as those described above. The computer-executable instructions may be compiled or interpreted from computer programs written using a variety of programming languages and / or technologies, including, but not limited to, Java, C, C++, C#, Visual Basic, Java Script, Perl, etc., alone or in combination. Generally, a processor (e.g., a microprocessor) receives instructions from, for example, a memory, a computer-readable medium, etc., and executes these instructions to thereby perform one or more processes, including one or more processes described herein. Such instructions and other data may be stored and transmitted using a variety of computer-readable media.
[0059] While exemplary embodiments have been described above, it is not intended that these embodiments describe all possible forms of the invention. Rather, the words used herein are words of description rather than limitation, and it will be understood that various changes may be made without departing from the spirit and scope of the invention. Furthermore, features of various embodiments may be combined to form further embodiments of the invention.
Claims
1. A sound signal processing system in a vehicle multimedia system, comprising: a loudspeaker configured to reproduce within the environment the audio signal from the audio source and the reinforcement voice signal; at least one microphone for detecting a microphone signal including a first voice signal component corresponding to spoken speech, a second voice signal component corresponding to the reinforcement voice signal reproduced by the loudspeaker, and an audio signal component corresponding to the audio signal reproduced by the loudspeaker; receiving the microphone signal from the at least one microphone; performing acoustic echo cancellation (AEC) of the microphone signal to generate an echo-canceled microphone signal, wherein the AEC uses a first adaptive filter to estimate and cancel feedback resulting from the environment; performing acoustic feedback cancellation (AFC) of the echo-canceled microphone signal to generate a processed microphone signal, the AFC using a second adaptive filter to estimate and cancel feedback resulting from application of the reinforcement voice signal in the environment; augmenting the spoken voice with the processed microphone signal to generate the augmented voice signal; a voice processor system configured to apply the reinforcement voice signal and the audio signal to the loudspeaker for reproduction in the environment.
2. 2. The sound signal processing system of claim 1, wherein the AEC is performed using a first subset of the loudspeakers and the AFC is performed using a second, different subset of the loudspeakers.
3. 10. The sound signal processing system of claim 1, wherein the voice processor system is further configured to perform automatic speech recognition (ASR) on the processed microphone signal and to receive commands to control the voice processor system.
4. 4. The sound signal processing system of claim 3, wherein the commands include one or more of: skip a song, repeat a song, repeat a section, adjust vocal effects and / or multi-channel effects, add a user for voice reinforcement, turn a user off for voice reinforcement, turn off voice reinforcement for all users, or request to turn on a voice processor mode to transmit spoken audio from one user to another user.
5. the environment includes a plurality of sound zones; the at least one microphone includes a first microphone in a first sound zone of the plurality of sound zones and a second microphone in a second sound zone of the plurality of sound zones; The voice processor system comprises: augmenting a first sound received by the first microphone and generating a first aspect of the augmented voice signal in the first sound zone; 2. The sound signal processing system of claim 1, configured to reinforce a second sound received by the second microphone and generate a second component of the reinforcement voice signal in the second sound zone.
6. The first voice is received from a first singer and the second voice is received from a second singer, and the voice processor system is further configured as follows: evaluating the pitch of each of the first voice and the second voice relative to a reference pitch; The sound signal processing system of claim 5 , further comprising: identifying whether the first singer or the second singer provided a performance that was closest to the reference pitch.
7. 2. The sound signal processing system of claim 1, wherein the voice processor system is further configured to apply vocal effects to the reinforcement voice signal, the vocal effects including the addition of artificial reverberation.
8. The voice processor system comprises: increasing a step size of adjustment of the second adaptive filter in response to detecting reverberation in the processed microphone signal; 8. The sound signal processing system of claim 7, configured to reduce a step size of the adjustment of the second adaptive filter in response to a lack of reverberation in the processed microphone signal.
9. A sound signal processing method in a vehicle multimedia system, comprising: receiving a microphone signal from at least one microphone, the microphone signal including a first voice signal component corresponding to spoken speech, a second voice signal component corresponding to a reinforcement voice signal reproduced by a loudspeaker in the environment, and an audio signal component corresponding to an audio signal reproduced by said loudspeaker; performing acoustic echo cancellation (AEC) of the microphone signal to generate an echo-canceled microphone signal, the AEC using a first adaptive filter to estimate and cancel feedback resulting from the environment; performing acoustic feedback cancellation (AFC) of the echo-canceled microphone signal to generate a processed microphone signal, the AFC using a second adaptive filter to estimate and cancel feedback resulting from application of the reinforcement voice signal in the environment; augmenting the spoken voice with the processed microphone signal to generate the augmented voice signal; A sound signal processing method applying said reinforcement voice signal and said audio signal to said loudspeakers for reproduction in said environment.
10. 10. The method of claim 9, wherein the AEC is performed using a first subset of the loudspeakers and the AFC is performed using a second, different subset of the loudspeakers.
11. 10. The method of processing a sound signal of claim 9, further comprising performing automatic speech recognition (ASR) on the processed microphone signal and receiving commands to control the vehicle multimedia system.
12. 12. The method of claim 11, wherein the commands include one or more of: skipping a song, repeating a song, repeating a section, adjusting vocal effects and / or multi-channel effects, adding a user for voice reinforcement, turning a user off for voice reinforcement, turning off voice reinforcement for all users, or requesting to turn on a voice processor mode for transmitting spoken audio from one user to another.
13. the environment comprises a plurality of sound zones; the at least one microphone includes a first microphone in a first sound zone of the plurality of sound zones and a second microphone in a second sound zone of the plurality of sound zones; augmenting a first sound received by the first microphone and generating a first aspect of the augmented voice signal in the first sound zone; 10. The method of claim 9, further comprising reinforcing a second sound received by the second microphone to generate a second component of the reinforcing voice signal in the second sound zone.
14. the first audio is received from a first singer and the second audio is received from a second singer; evaluating the pitch of each of the first voice and the second voice relative to a reference pitch; 14. The method of processing a sound signal according to claim 13, further comprising identifying whether the first singer or the second singer provided a performance that was closest to the reference pitch.
15. 10. The method of claim 9, further comprising applying a vocal effect to the reinforcement voice signal, the vocal effect comprising adding artificial reverberation.
16. increasing a step size of the adjustment of the second adaptive filter in response to detecting reverberation in the microphone signal; 16. The method of claim 15, further comprising reducing a step size of the adjustment of the second adaptive filter in response to a lack of reverberation in the microphone signal.
17. 1. A non-transitory computer readable medium containing instructions for sound signal processing in a vehicle multimedia system, the instructions, when executed by a voice processor system, causing the voice processor system to: receiving a microphone signal from at least one microphone, the microphone signal including a first voice signal component corresponding to spoken speech, a second voice signal component corresponding to a reinforcement voice signal reproduced by a loudspeaker in the environment, and an audio signal component corresponding to an audio signal reproduced by said loudspeaker; performing acoustic echo cancellation (AEC) of the microphone signal to generate an echo-canceled microphone signal, the AEC using a first adaptive filter to estimate and cancel feedback resulting from the environment; performing acoustic feedback cancellation (AFC) of the echo-canceled microphone signal to generate a processed microphone signal, the AFC using a second adaptive filter to estimate and cancel feedback resulting from application of the reinforcement voice signal in the environment; augmenting the spoken voice with the processed microphone signal to generate the augmented voice signal; A non-transitory computer-readable medium for performing an operation of applying the reinforcement voice signal and the audio signal to the loudspeaker for reproduction within the environment.
18. 20. The non-transitory computer-readable medium of claim 17, further comprising instructions that, when executed by the voice processor system, cause the voice processor system to perform operations including performing AEC using a first subset of the loudspeakers and performing AFC using a second, different subset of the loudspeakers.
19. 20. The non-transitory computer-readable medium of claim 17, further comprising instructions that, when executed by the voice processor system, cause the voice processor system to perform operations including performing automatic speech recognition (ASR) on the processed microphone signal to receive commands to control the vehicle multimedia system.
20. 20. The non-transitory computer-readable medium of claim 19, wherein the commands include one or more of: skip a song, repeat a song, repeat a section, adjust vocal and / or multi-channel effects, add a user for voice reinforcement, turn a user off for voice reinforcement, turn off voice reinforcement for all users, or request to turn on a voice processor mode to transmit spoken voice from one user to another user.
21. the environment includes a plurality of sound zones; the at least one microphone includes a first microphone in a first sound zone of the plurality of sound zones and a second microphone in a second sound zone of the plurality of sound zones; When executed by the voice processor system, the voice processor system: A first sound received by the first microphone is reinforced to generate a first aspect of the reinforcement voice signal in the first sound zone.
20. The non-transitory computer-readable medium of claim 17, further comprising instructions to perform operations including reinforcing a second sound received at the second microphone and generating a second component of the reinforcing voice signal at the second sound zone.
22. the first audio is received from a first singer and the second audio is received from a second singer; When executed by the voice processor system, the voice processor system: evaluating the pitch of each of the first voice and the second voice relative to a reference pitch; 22. The non-transitory computer-readable medium of claim 21, further comprising instructions to perform operations including identifying whether the first singer or the second singer provided a performance that most closely matched the reference pitch.
23. 20. The non-transitory computer-readable medium of claim 17, further comprising instructions that, when executed by the voice processor system, cause the voice processor system to perform operations including adding vocal effects to the reinforcement voice signal, the vocal effects including adding artificial reverberation.
24. When executed by the voice processor system, the voice processor system: increasing a step size of adjustment of the second adaptive filter in response to detecting reverberation in the processed microphone signal; 24. The non-transitory computer-readable medium of claim 23, further comprising instructions to perform an action including decreasing a step size of adjustment of the second adaptive filter in response to a lack of reverberation in the processed microphone signal.
Citation Information
Patent Citations
Multi-zone interference cancellation for voice communication in a vehicle
EP3346466A1
In-car communication howling prevention
US20180068672A1
Voice interface and vocal entertainment system
US20180190306A1
Device for assisting two-way conversation and method for assisting two-way conversation
WO2017064839A1