Loudspeaker system with dynamic voice equalization
By using a dynamic speech equalization method, which combines adaptive and transparent filters to dynamically adjust the frequency response, the problem of poor frequency content of speech signals in long-distance communication is solved, thereby improving speech intelligibility and sound quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SENNHEISER COMM
- Filing Date
- 2021-07-21
- Publication Date
- 2026-05-15
AI Technical Summary
In remote communication, poor frequency content of voice signals leads to reduced voice intelligibility and affects the conversation experience, especially in noisy environments and damaged communication channels.
A dynamic speech equalization method is adopted. The speech analysis module identifies the frequency characteristics of the speech signal, and the combination of adaptive filter and transparent filter is used to dynamically adjust the frequency response to correct the speech signal and enhance speech intelligibility.
It improves the frequency correction effect of voice signals, enhances voice intelligibility and sound quality, and improves the conversation experience.
Smart Images

Figure CN113965852B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to loudspeaker systems, such as loudspeaker systems operating under power-limited conditions. More particularly, this invention relates to loudspeaker systems comprising loudspeaker units, equalization units, and a user interface. Preferably, the loudspeaker is part of or an amplifier.
[0002] The teachings of this invention can be used, for example, in the following applications: hands-free telephone systems, mobile phones, remote conferencing systems (such as loudspeakers), headphones, headsets, broadcasting systems, gaming devices, karaoke systems, classroom amplification systems, etc. Background Technology
[0003] In long-distance communication, incoming voice signals often have poor frequency content / spectrum.
[0004] Faulty microphones, dirty microphone inlets, poorly designed microphones, improperly used microphones (e.g., headset boom microphones placed under the chin, holding a mobile phone at an unfavorable angle, Bluetooth headsets pointed downwards instead of towards the mouth), noisy environments, wind noise and subsequent noise reduction, damaged communication channels, narrowband channels, bad codecs, etc., are all factors that negatively impact listening effort or / or speech intelligibility to varying degrees because they introduce non-flat filtering of the signal. Similarly, a speaker may have an unfavorable amplitude response, such as having too little mid- and high-frequency content when referring to achieving optimal speech intelligibility.
[0005] The damage described above can result in a poor conversation experience due to misunderstandings and the need for greater effort to understand what is being said.
[0006] Similarly, noise and interference in the receiver's environment can also affect speech intelligibility.
[0007] exist Figure 1 The diagram shows the communication channels and indicates where degradation might occur. Summary of the Invention
[0008] The specific solution to the problem mentioned above is to dynamically correct the frequency response based on the frequency content of the signal. A method is proposed to dynamically improve the above-mentioned defects during a call.
[0009] One aspect of the present invention is to provide a method, apparatus, and system for dynamically correcting frequency response based on the frequency content of an audio input signal, which, alone or in combination, alleviates, mitigates, or eliminates one or more deficiencies of the prior art mentioned above.
[0010] According to one aspect, a method for dynamic speech equalization includes: receiving an input signal; processing the input signal according to a frequency; providing an equalized electro-acoustic signal according to an equalization function, wherein the equalization function includes at least an actuator portion configured to dynamically apply a first compensation filter to the received input signal, dynamically apply a second compensation filter to the received input signal, and transmit an output signal representing the electro-acoustic input signal or a processed version thereof, which can be perceived by a user as sound. The second compensation filter may be a transparency filter.
[0011] Speech can be clearly characterized by, for example, modulation index, frequency response, and crest factor. In communications over bandwidth-limited channels such as softphones and / or voice over telephone or data, the spectral information in transmitted speech is often distorted. As will be described herein, by utilizing knowledge of “ideal” or approximate ideal speech, it is possible to at least partially correct the distorted information. One goal of such algorithms is to increase intelligibility and / or perceived sound quality.
[0012] The method described herein and / or one version of the loudspeaker amplifier include a speech analysis module and frequency-varying adjustments. The speech analysis module provides information such as the presence of speech. In other versions, the speech analysis module analyzes the incoming speech by comparing the incoming signal with known characteristics of speech, and possibly, analyzing deviations from those known characteristics. The speech can be compared using the spectrum of the incoming speech with one or more stored (e.g., reference) speech spectra, such as averaged or idealized speech spectra. The analysis can be based on analyzing the spectral shape of the incoming speech. The incoming speech can be analyzed using artificial intelligence. Based on the comparison, one or more corrections, such as a set of corrections, can be determined. The determined or estimated corrections can then be applied to the incoming speech to obtain an improved or processed incoming speech signal.
[0013] Speech analysis can be divided by frequency bands; for example, the first analysis path is a high-frequency path, and the second analysis path is a low-frequency path. The frequency band division into high-frequency and low-frequency paths can be determined by the interval between the cutoff frequency of the high-frequency filter and the low-frequency filter used for the corresponding high-frequency analysis path. In one example, the low-frequency path can analyze signals below 500 Hz.
[0014] In the analysis module described herein, an energy estimator may follow the low-pass filter. This will determine the energy of the incoming speech in the low-frequency range, such as below 500 Hz. In the analysis module described herein, an energy estimator may follow the high-pass filter. This part will determine the energy of the incoming speech in the high-frequency range, such as above 2000 Hz.
[0015] A mapper module can be provided that utilizes prior knowledge or speech characteristics from the "ideal speech" to map the relationship between the two high- and low-frequency energy estimates mentioned above to calculate the correction to be applied to improve the incoming speech signal. A simple example could be that the relationship between the high and low estimators indicates that the incoming signal lacks energy above a given frequency, such as above 2000 Hz; therefore, the correction could include adding the corresponding energy to the signal in the upper frequency range to better match the "ideal speech," thus providing a corrected speech signal.
[0016] Frequency-dependent adjustment can be implemented in a module with two inputs: 1) the incoming damaged speech, and 2) a correction to improve the quality estimate of the incoming speech. This frequency-dependent adjustment module can implement adaptive correction functions, such as an adaptive filter, which applies the correction received from the speech analysis module and outputs an improved speech signal for further processing or presentation to the user.
[0017] An adaptive filter can be constructed using two filters: filter 1 and filter 2. These filters can have different frequency responses. The output filters 1 and 2 can be weighted and summed together to form a corrected, for example, desired output. If, for example, filter 2 has an almost flat response, and filter 1 has a similar response, except for an increased boost above a given frequency such as 2000 Hz, then by making filter 1 weighted higher than filter 2, it is possible to boost the higher frequencies. Naturally, this boost can be adjusted by choosing different filters and by selecting appropriate weights. Thus, using two static filters in the described configuration, it is possible to control the output of the adaptive filter using a single scalar. The example above naturally scales to N filters; however, adding many filters in such a structure may not be the most efficient way to achieve the required degrees of freedom. For the applications described in this specification, two filters generally produce acceptable results.
[0018] In one aspect, the present invention provides a method of operating a loudspeaker (also known as a loudspeaker telephone), wherein the loudspeaker may include a microphone system, an output system, and a communication system, wherein the communication system is configured to communicate with a remote speaker. The loudspeaker may be configured to receive signals via the communication system. The loudspeaker may be configured to analyze the received signals to determine whether the signals include speech. If the signals include speech, the signals are processed by frequency-varying processing. A speech analysis module may be configured to perform the analysis of whether the signals include speech. The speech analysis module may include a low-frequency processing path and a high-frequency processing path. Each of the low-frequency processing path and the high-frequency processing path includes a corresponding low-pass or high-pass filter. Furthermore, each of the low-frequency and high-frequency filtering paths may include an energy estimator configured to provide an estimate or measure of the energy from the corresponding low-pass or high-pass filter. The estimated or measured energy level may be provided to a mapping module configured to characterize the corresponding signal relative to speech, such as ideal speech, a speech spectrogram. The characterization may include one or more of the following: direct comparison, correlation calculation, difference calculation, conversion. The low-frequency path may have a cutoff frequency of 500 Hz, for example, in the range of 400-600 Hz. The high-frequency path may have a cutoff frequency of 2000 Hz, for example, in the range of 1500-2500 Hz.
[0019] When a speech signal is identified in the input signal and it is determined that correction or improvement is needed, the signal can undergo frequency-varying adjustment, such as adaptive frequency-varying adjustment. Frequency-varying adjustment can be achieved via at least a first filter and a second filter. The first filter may have a different frequency distribution than the second filter. The first and second filters preferably operate within the same frequency range, for example, at least substantially within the same frequency range or at least partially overlapping frequency ranges. The outputs of the first and second filters can be combined, such as by summation. The combination or summation of the outputs of the first and second filters can be weighted. The weight of the combined filter components can be 1. The weighting can be frequency-varying, such that a frequency window from the first and second filters does not have the same weight distribution as different frequency windows.
[0020] This invention provides a loudspeaker amplifier configured to perform the methods described above. Such a loudspeaker amplifier includes at least a communication unit configured to receive communication signals, a processor for processing the communication signals, and an output converter configured to provide signals based on the processed communication signals. The loudspeaker amplifier can be configured to perform the steps mentioned in conjunction with the methods described above. The loudspeaker amplifier may include any or all of the additional features disclosed in this invention.
[0021] On the one hand, the compensation filter applies compensation to the damaged signal in the worst case, while the transparent filter applies compensation to the damaged signal in the best case.
[0022] In one aspect, the method of the present invention includes mixing a first output signal from a compensation filter and a second output signal from a transparency filter based on first and second dynamic weights, respectively.
[0023] In one aspect, the method of the present invention includes continuously updating the first and second dynamic weights, respectively.
[0024] On the one hand, a high-quality speech channel will have a high weighting on the output of the transparency filter, a low-quality speech channel will have a high weighting on the output of the compensation filter, and a moderately impaired speech channel may have a weighting of 0.5 on both.
[0025] The equalization function may include an analysis section configured to determine the mixing weights and apply a first filter to the received input signal and / or apply a second filter to the received input signal.
[0026] The analysis section may include mapping functions for mixing weights and updating the weights of each input signal.
[0027] According to another aspect, a loudspeaker system or hearing device is provided, comprising: an input unit for receiving an audio input signal having a first dynamic range level representing a sound signal with time and frequency variations and providing an electro-audio input signal, the electro-audio input signal including a target signal and / or a noise signal; and a processing unit for modifying the input audio signal according to a frequency and providing an equalized electro-audio signal according to an equalization function. Advantageously, the input unit receives the audio signal from a distant sound source, such as a distant speaker. The equalization function includes at least an actuator portion comprising a compensation filter dynamically applied to the received input signal and a transparency filter dynamically applied to the received input signal. The loudspeaker system further includes an output unit for providing an output signal that can be perceived by a user as sound, representing the electro-acoustic input signal or a processed version thereof.
[0028] In addition, manufactured articles such as communication devices are provided, including the speaker systems described above, described in detail below, and shown in the figures.
[0029] The manufactured article is capable of implementing a communication device, which includes:
[0030] - First microphone signal path, including
[0031] -- Microphone unit;
[0032] -- First signal processing unit; and
[0033] -- Transmitter unit;
[0034] The aforementioned units are interconnected and configured to transmit processed signals originating from the input sound picked up by the microphone unit; and
[0035] - Second speaker signal path, including
[0036] -- Receiver unit;
[0037] -- Second signal processing unit; and
[0038] -- Speaker unit;
[0039] These units are interconnected and configured to provide acoustic sound signals derived from signals received by the receiver unit.
[0040] This allows the implementation of a loudspeaker (also known as a speakerphone) or headset that includes a loudspeaker system according to the invention.
[0041] The communication device may include at least one audio interface to a switching network and at least one audio interface to an audio transmission device. The (unidirectional) audio transmission device may be a music player or any other entertainment device that provides audio signals. The communication device (its speaker system) may be configured to enter or be in a first full-bandwidth mode when connected to the (unidirectional) audio transmission device via the audio interface. The communication device (its speaker system) may be configured to enter or be in a second wired bandwidth mode when connected to the communication device via the audio interface to establish a (bidirectional) connection to another communication device via the switching network. Alternatively, such mode changes may be initiated via the user interface of the communication device.
[0042] Communication devices may include loudspeakers or mobile (such as cellular) phones (e.g., smartphones) or headsets or hearing aids.
[0043] The manufactured items may include or may be headphones or gaming devices.
[0044] Speaker systems can be used, for example, in teleconferences, where the audio gateway can be a laptop or another type of computer device or smartphone connected to the Internet or telephone network via a first communication link.
[0045] On the other hand, headsets or headphones are also provided, including the speaker systems described above and shown in the figures.
[0046] The loudspeaker system according to the invention is generally applicable to any device or system including a specific electroacoustic system (comprising a loudspeaker and mechanical parts communicating with the loudspeaker) having a transfer function, which (at a specific sound output) exhibits a sharp drop in low and / or high frequencies (e.g., Figure 3 (C is schematically shown for a loudspeaker unit). Applying a loudspeaker system according to the invention can advantageously contribute to compensation for the loss of low and / or high frequency components in the sound output to the environment (e.g., due to leakage).
[0047] In headsets or headphones, the dropout is primarily determined by the electroacoustic system (including A. speaker design and B. ear cup / ear pad / earplug design).
[0048] Headphones or headsets can be open-back (meaning they exchange a certain amount or a large amount of sound with the environment) or closed-back (meaning they aim to limit sound exchange with the environment).
[0049] In this specification, the term "open" refers to a considerably high acoustic leakage between the surrounding environment and the space confined by the ear / ear canal / eardrum and the ear cup / ear pad / ear plug that covers or blocks the ear / ear canal. In closed-back headphones or headsets, the leakage will be (substantially) lower than in open-back headphones (but some leakage will generally be present, which can be compensated for by the speaker system according to the invention).
[0050] Headphones or headsets may include active noise cancellation systems.
[0051] On the other hand, the speaker system described above and illustrated in the figures is provided for use in loudspeaker amplifiers or mobile (such as cellular) phones (such as smartphones) or gaming devices or headphones or headsets or hearing aids.
[0052] The communication device may be a portable device, such as a device that includes a local power source, such as a battery, for example a rechargeable battery. The hearing aid device may be a low-power device, and the term "low-power device" in this specification means a device whose energy budget is limited, for example because it is, for example, a portable device that includes a power source, and the duration of that power is limited (the limited duration is, for example, hours or days) without replacement or recharging.
[0053] The communication device may include an analog-to-digital converter (ADC) to digitize an analog input using a predetermined sampling rate such as 20 kHz. The communication device may also include a digital-to-analog converter (DAC) to convert a digital signal into an analog output signal, for example, for presentation to a user via an output converter.
[0054] The communication device considers the minimum frequency f min up to the maximum frequency f max The frequency range includes a portion of the typical human hearing range from 20 Hz to 20 kHz, such as a portion of the range from 20 Hz to 12 kHz.
[0055] Communication devices may include a voice detector (VD) for determining whether an input signal (at a given point in time) includes a voice signal. Such a detector can help determine the appropriate operating mode for the speaker system. The voice detector may output a probability of the presence of voice or a binary decision indicating whether voice is present.
[0056] The communication device may include an acoustic (and / or mechanical) feedback suppression system. The communication device may also include other relevant functions for the application in question, such as compression and noise reduction.
[0057] Communication devices may include mobile phones or loudspeakers. Communication devices may include or may be listening devices, such as entertainment devices, music players, hearing aids, hearing instruments, headsets, earphones, ear protection devices, or combinations thereof.
[0058] When the speaker system is a headphone, one of the advantages of this invention is that audio signals, including voice or music, or any sound, can be easily distributed between the two output converters of the communication unit located at the corresponding ears of the user / wearer.
[0059] Audio signals transmitted via the communication link may include voices from one or more users, i.e., speakers at a distance, from another audio gateway device connected to the other end of the first communication link. Alternatively, the audio signal may include music. In this example, music may be played simultaneously on the communication unit.
[0060] A hearing device may be or include a hearing aid suitable for improving or enhancing a user's hearing ability, which is achieved by receiving sound signals from the user's environment, generating corresponding audio signals, possibly modifying the audio signals, and providing the possibly modified audio signals as audible signals to at least one ear of the user. "Hearing device" may also, or alternatively, refer to a device suitable for electronically receiving audio signals, such as a wearable device, headset, or earphone, which may modify the audio signals and provide the possibly modified audio signals as audible signals to at least one ear of the user. The audible signals may be provided in the following forms: sound signals radiated into the user's outer ear; sound signals transmitted as mechanical vibrations through the bone structures of the user's head and / or through parts of the user's middle ear to the user's inner ear; or electrical signals transmitted directly or indirectly to the user's cochlear nerve and / or auditory cortex.
[0061] Hearing devices can be adapted to be worn in any known manner. This may include:
[0062] i) The hearing device unit is positioned behind the ear (having a tube to guide acoustic signals into the ear canal or having a receiver / speaker positioned close to or within the ear canal), such as a behind-the-ear hearing aid; and / or
[0063] ii) The hearing device is placed entirely or partially in the user's auricle and / or ear canal, such as an in-ear hearing aid or an in-the-canal / deep-canal hearing aid; or
[0064] iii) The hearing device is configured to connect to a fixation device implanted in the skull, such as a bone-anchored hearing aid or a cochlear implant; or
[0065] iv) The hearing device unit is configured as a fully or partially implanted unit, such as a bone-anchored hearing aid or a cochlear implant; or
[0066] v) Place the ear cups in one or both of the user's ears so that sound can be radiated into the corresponding ear canals.
[0067] In the case where the present invention is implemented in a loudspeaker, the aforementioned loudspeaker will include a loudspeaker housing, an output converter, an input converter, and a processor.
[0068] A “hearing system” refers to a system comprising one or two hearing devices. A “binaural hearing system” refers to a system comprising two hearing devices adapted to collaboratively provide audible signals to both of a user’s ears. A hearing system or a binaural hearing system may also include an auxiliary device that communicates with at least one hearing device, which affects the operation of the hearing device and / or benefits from its functionality. A wired or wireless communication link is established between at least one hearing device and the auxiliary device to allow the exchange of information (such as control and status signals, possibly audio signals). The auxiliary device may include at least one of the following: a remote control, a remote microphone, an audio gateway device, a mobile phone, a broadcasting system, a car audio system, a music player, or a combination thereof. The audio gateway device is adapted to receive multiple audio signals, such as from an entertainment device (e.g., a TV or music player), a telephone device (e.g., a mobile phone), or a computer (e.g., a PC). The audio gateway device is also adapted to select and / or combine appropriate signals from the received audio signals (or combinations of signals) to transmit to at least one hearing device. The remote control is adapted to control the function and operation of at least one hearing device. The remote control functionality can be implemented in a smartphone or another electronic device, which may run an application that controls the functionality of at least one hearing device.
[0069] Generally, hearing devices include:
[0070] i) An input unit, such as a microphone, for receiving sound signals from the user's surroundings and providing corresponding input audio signals; and / or
[0071] ii) A receiving unit for receiving input audio signals electronically.
[0072] The hearing device also includes a signal processing unit for processing the input audio signal and an output unit for providing the user with an audible signal based on the processed audio signal.
[0073] The input unit may include multiple input microphones, for example, for providing direction-dependent audio signal processing. The aforementioned directional microphone system is adapted to amplify a target sound source among multiple sound sources in a user's environment. In one aspect, the directional system is adapted to detect (e.g., adaptive detection) the direction from which a specific portion of the microphone signal originates. This can be achieved using methods conventionally known. The signal processing unit may include an amplifier adapted to apply a frequency-dependent gain to the input audio signal. The signal processing unit may also be adapted to provide other suitable functions such as compression, noise reduction, etc. The output unit may include an output converter, such as a speaker / receiver for providing airborne acoustic signals transdermally or through the skin to the skull, or a vibrator for providing acoustic signals propagating through structures or fluids. In some hearing devices, the output unit may include one or more output electrodes for providing electrical signals, such as in cochlear implants. Attached Figure Description
[0074] Various aspects of the invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the invention while omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Features of each aspect may be combined with any or all features of other aspects. These and other aspects, features, and / or technical effects will be apparent from and illustrated in the following figures, wherein:
[0075] Figure 1 A communication channel with amplitude response degradation is shown;
[0076] Figure 2 A hearing device / speaker system according to the present invention is shown;
[0077] Figure 3 A dynamic voice equalizer for a loudspeaker system according to the invention is shown;
[0078] Figure 4 A simplified block diagram of dynamic speech equalization according to the present invention is shown;
[0079] Figures 5a-5b Two applications of a loudspeaker amplifier or headset, including a loudspeaker system according to the present invention, are shown;
[0080] Figure 6 A loudspeaker or headset including a loudspeaker system according to the invention is shown;
[0081] Figure 7 A cross-sectional view is shown of a loudspeaker amplifier that can advantageously include a loudspeaker system according to the invention.
[0082] The further applicability of the invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and specific examples illustrate preferred embodiments of the invention, they are given for illustrative purposes only. Other embodiments of the invention will become apparent to those skilled in the art based on the following detailed description. Detailed Implementation
[0083] The detailed description below, taken in conjunction with the accompanying drawings, serves as a description of various different configurations. This detailed description includes specific details to provide a thorough understanding of several different concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively, “elements”). Depending on the specific application, design constraints, or other reasons, these elements may be implemented using electronic hardware, computer programs, or any combination thereof.
[0084] Electronic hardware may include microelectromechanical systems (MEMS), (e.g., application-specific integrated circuits), microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), gating logic, discrete hardware circuits, printed circuit boards (PCBs) (e.g., flexible PCBs), and other suitable hardware configured to perform the various functions described in this specification, such as sensors for sensing and / or recording the physical properties of the environment, devices, users, etc. Computer programs should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or other names.
[0085] Generally, a hearing device includes i) an input unit, such as a microphone, for receiving sound signals from the user's surroundings and / or the user's own speech and providing a corresponding input audio signal; and / or ii) a receiving unit for electronically receiving the input audio signal. The hearing device also includes a signal processing unit for processing the input audio signal and an output unit for providing an audible signal to the user based on the processed audio signal.
[0086] In long-distance communication, incoming voice signals often have undesirable frequency content / spectral density. A method is proposed to dynamically improve this deficiency during a call.
[0087] Faulty microphones, dirty microphone inlets, poorly designed microphones, improperly used microphones (e.g., headset boom microphones placed under the chin, holding mobile phones at an unfavorable angle, BT headsets pointed downwards instead of towards the mouth, etc.), noisy environments, wind noise and subsequent noise reduction, damaged communication channels, narrowband channels, bad codecs, etc., are all factors that have a greater or lesser negative impact on listening effort or / or speech intelligibility because they introduce non-flat filtering of the signal.
[0088] Similarly, a speaker may have an unfavorable amplitude response, such as having too little mid- and high-frequency content when referring to achieving optimal speech intelligibility.
[0089] The damage described above can result in a poor conversation experience due to misunderstandings and the need for greater effort to understand what is being said.
[0090] Similarly, noise and interference in the receiver's environment can also affect speech intelligibility.
[0091] Figure 1 The communication channel is shown, with areas where degradation may occur marked.
[0092] Figure 2 A loudspeaker system 10 is shown, including an input unit 11 (IU) that provides an electrical input audio signal 12 (eIN) based on an input signal IN. The input signal IN may be, for example, an ambient sound signal (in which case, the input unit IU includes a microphone) or an electrical signal received from a component of the loudspeaker system or from another device, such as via a data communication channel or telephone connection, or a combination thereof. The input unit IU includes an audio interface. The input signal IN (where it is an electrical signal) may be an analog signal (such as an audio signal from an audio jack interface) or a digital signal (such as an audio signal from a USB audio interface). The input unit 11 (IU) may include, for example, an analog-to-digital converter (ADC) to convert the analog electrical signal into a digital electrical signal (using an appropriate audio sampling frequency, such as 20 kHz). The loudspeaker system 10 also includes a processing unit 13, which includes a dynamic voice equalization unit 100 (EQ) for modifying the electrical input audio signal 12 (eIN) (or its processed version) according to frequency and providing a processed electrical audio signal 14 (eINeq) according to a predetermined dynamic voice equalization function. The dynamic voice equalization unit 100 (EQ) in... Figure 3Further description is provided below. The speaker system 10 also includes a speaker unit 15 (SPK) for converting the processed electrical audio signal eINeq into an acoustic output sound signal OUT. Alternatively, the speaker unit 14 may be a mechanical vibrator of a bone-anchored hearing aid. In a specific operating mode of the speaker system 10, the processing unit 13 is configured to apply a specific dynamic speech equalization function to the electrical input audio signal. The speaker system 10 includes a dynamic speech equalization unit 100, which specifies predetermined frequency-varying gains w1, w2 (e.g., attenuation) to be applied to the input audio signal 12.
[0093] Figure 3 A dynamic speech equalization function 100 or algorithm is shown, which dynamically corrects the frequency response based on the frequency content of the acoustic input audio signal 12. This algorithm enhances speech intelligibility by dynamically applying a compensation filter 111 to the incoming acoustic audio signal 12 to the necessary extent when necessary.
[0094] The algorithm is divided into an actuator section 110 and an analysis section 120. In the actuator section 110, the audio input signal 12 is passed through a compensation filter 111 and a transparency filter 112. The compensation filter 111 applies compensation to the worst-case damaged signal. The output 112a of the transparency filter 112 is designed for the best-case signal. The transparency filter 112 does not necessarily need to be transparent. The two outputs 111a and 112a are mixed together with dynamic weights w1 and w2, which are continuously updated by the analysis section 120 (for each input sample). A high-quality speech channel will have a high weight from the transparency filter output 112a, a low-quality speech channel will have a high weight from the compensation filter output 111a, and a moderately damaged speech channel may use a weight of 0.5 for both.
[0095] The audio input signal 12 is also fed into an analysis section 120 that determines the mixing weights w1, w2. In the analysis section 120, the audio signal 12 is filtered at least by a first filter 121 and a second filter 122. Filters 121 and 122 can be of any type. In one example, the first filter 121 is a low-pass filter (LP) and the second filter 122 is a high-pass filter (HP). The average energy of the second filter 122 is divided by the average energy of the first filter 121, and the result of the division is also an average value, which is fed into a mapping function 125. The mapping function outputs (126) the mixing weights w1, w2, which are updated for each input sample. The mapping function 125 can be selectively invoked only when speech is detected in the audio signal to avoid compensation based on non-speech signals.
[0096] The algorithm described herein aims to enhance speech intelligibility by dynamically applying compensation filters to the audio signal as needed and to the necessary degree. The algorithm can be divided into an actuator section and an analysis section. In the actuator section, the audio input can be passed through a compensation filter and a transparency filter. The compensation filter can be configured to apply compensation to the worst-case impaired signal. The output of the transparency filter can be configured to apply compensation to the best-case signal. The transparency filter does not necessarily have to be transparent; other settings can be applied. The two outputs can be mixed together using dynamic weights that are continuously updated by the analysis section (for each input sample). High-quality speech channels will have high weights from the transparency filter output, low-quality speech channels will have high weights from the compensation filter output, and moderately impaired speech channels may have a weight of 0.5 for both.
[0097] Audio input can be fed into an analysis section that determines the mixing weights, for example, continuously or intermittently and / or adjustablely. In the analysis section, the audio signal is filtered by two filters, referred to herein as LP and HP. The filters can be of any type in principle. In one example, they are a low-pass filter and a high-pass filter. The average energy of the HP filter is divided by the average energy of the LP filter, and the result is also averaged and fed into a mapping function. The mapping function outputs the mixing weights and can be updated for each input sample, or for a time interval of multiple samples, such as for every fifth sample, etc. The mapping function can be selected to be enabled only when speech is detected in the audio signal, which helps avoid compensation based on non-speech signals.
[0098] Figure 4 A method S100 for dynamic speech equalization is shown, comprising: step S10, receiving an acoustic input signal 12; and step S20, processing the acoustic input signal using a dynamic speech equalization function 100 for enhancing speech intelligibility, which includes at least an actuator portion 110. The method further comprises step S30, transmitting an output signal based on the processed input signal.
[0099] The process also includes step S21, dynamically applying the compensation filter 111 to the received acoustic input signal, and step S22, dynamically applying the transparency filter 112 to the received input signal.
[0100] Figure 5a The diagram shows a communication device CD (such as a speaker amplifier) including two wired or wireless audio interfaces to other devices, such as: a) a CellPh wireless telephone (such as a mobile phone, e.g., a smartphone). Figure 5a ) or one-way audio transmission device ( Figure 5bThe audio interface to the computer PC (e.g., a music player); and b) a computer PC (e.g., a personal computer). The audio interface to the computer PC can be a wireless or wired interface, such as a USB (audio) interface including a cable and a USB connector, for connecting the communication device to the computer and enabling two-way audio exchange between the communication device CD and the computer. The audio interface to the cordless phone CellPh may include a cable and a telephone connector PhCon for directly connecting the communication device to the cordless phone and enabling two-way audio exchange between the communication device and the cordless phone. The communication device CD includes multiple activation elements (B1, B2, B3) such as buttons (or, alternatively, a touch-sensitive display) to enable control of the communication device and / or the functions of devices connected to the communication device. One of the activation elements (e.g., B1) may be configured to connect the cordless phone CellPh to the communication device (off-hook, answering a call) and / or disconnect it (on-hook, ending a call). One of the activation elements (e.g., B2) may be configured to enable the user to control the volume of the speaker output. One of the activation elements (e.g., B3) may be configured to enable the user to control the operating mode of the speaker system of the communication device.
[0101] Figure 5a The scenario depicted is a remote conference between users U1 and U2 near the communication device CD and users RU1, RU2, and RU3 at two (different) remote locations. Remote user RU1 is connected to the communication device CD via a cellphone (CellPh) and a wireless connection (WL1) to the network NET. Remote users RU2 and RU3 are connected to the communication device CD via a computer (PC) and a wired connection (WI1) to the network NET.
[0102] Figure 5b It shows a difference Figure 5a In certain situations. Figure 5b The reception (and, optionally, mixing) of audio signals from multiple different audio transmission devices (music players and PCs) connected to the communication device CD is shown. The communication device CD includes two (bidirectional) audio interfaces embodied in I / O units IU1 / OU1 and IU2 / OU2, respectively.
[0103] Figure 5bThe communication device includes a speaker signal path SSP, a microphone signal path MSP, and a control unit CONT for dynamically controlling the signal processing of the two signal paths. The speaker signal path SSP includes receiver units IU1 and IU2 for receiving electrical signals from connected devices and providing them as electrically received input signals S-IN1 and S-IN2; an SSP signal processing unit 13a for processing (including dynamic voice equalization) the electrically received input signals S-IN1 and S-IN2 and providing a processed output signal S-OUT; and a speaker unit 15 that is connected to each other during operation and configured to convert the processed output signal S-OUT into an acoustic sound signal OS originating from the signals received by receiver units IU1 and IU2. The speaker signal path SSP also includes a selector-mixing unit SEL-MIX for selecting one of the two input audio signals (or mixing them) and providing the resulting signal S-IN to the SSP signal processing unit 13a. The microphone signal path (MSP) includes a microphone unit (MIC) for converting the acoustic input sound IS into an electro-microphone input signal M-IN, an MSP signal processing unit (13b) for processing the electro-microphone input signal M-IN and providing a processed output signal M-OUT, and corresponding transmitter units (OU1, OU2) connected to each other during operation and configured to transmit the processed signal M-OUT from the input sound IS picked up by the microphone unit (MIC) to the connected device. The control unit (CONT) is configured to dynamically control the processing of the SSP signal processing unit and the MSP signal processing unit (13a and 13b, respectively), including mode selection and equalization in the SSP path.
[0104] The speaker signal path SSP is divided into two IU1 and IU2 paths for receiving input signals from the corresponding audio devices (music player and PC). Similarly, the microphone signal path MSP is divided into two OU1 and OU2 paths for transmitting output signals to the corresponding audio devices (music player (not closely related) and PC). One-way and two-way audio connections between the communication devices (units IU1, IU2 and OU1, OU2) and the two audio devices (here, the music player and PC) can be established via wired or wireless connections, respectively.
[0105] Figure 6 The diagram illustrates a communication device CD (here, a loudspeaker or headset) including a loudspeaker system according to the invention, its units and functions, and their combination. Figures 5a-5b The same applies (therefore, the corresponding operating mode can be represented, wherein the equalization according to the invention, which varies with dynamic speech, is advantageously applied).
[0106] exist Figure 6In this configuration, the audio interface is included in the I / O control unit I / O-CNT, which receives input signals from devices connected to the corresponding audio interface and transmits output signals to the connected devices, if appropriate. In listening mode, music or other broadband audio signals are received from one or both of the audio transmission devices (music player and PC), assuming no audio signal is transmitted from the communication device CD to the connected audio transmission device. This listening mode is therefore equivalent to the previously described mode. The I / O control unit I / O-CNT is connected to the power control unit PWR-C and the battery BAT. The power control unit PWR-C receives signals from the I / O control unit I / O-CNT, enabling it to detect possible power signals from the audio interface and, if such power is present, to begin recharging the rechargeable battery BAT, if appropriate. It also provides a control signal CHc to the control unit, indicating whether the current power supply is based on a remote source (such as via the audio interface or via trunk power reception) or whether the local power supply BAT is currently being used. This information can be used in the control unit CONT to determine the appropriate operating mode, and also regarding equalization that varies with dynamic speech. For battery mode and external power mode, a specific set of equalization functions that vary with the voice can be defined, and the appropriate set of parameters for implementing the corresponding equalization functions can be stored in the communication device.
[0107] Figure 7 A speaker amplifier CD is shown that can advantageously include a speaker system according to the invention. The speaker amplifier includes a centrally located speaker unit (not shown). The speaker amplifier CD also includes a centrally located user interface (UI) in the form of an activation element (button) for, for example, changing the operating mode of the device (or for activating an on or off state, etc.).
[0108] When appropriately replaced by a corresponding process, the structural features of the apparatus described above, in detail in the "Detailed Description" section, and as defined in the claims can be combined with the steps of the method of the present invention.
[0109] Unless explicitly stated otherwise, the singular forms “a” and “the” used herein include the plural forms (i.e., meaning “at least one”). It should be further understood that the terms “having,” “comprising,” and / or “including” as used in the specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that, unless explicitly stated otherwise, when an element is referred to as “connected” or “coupled” to another element, it may be a direct connection or coupling to the other element, or there may be intermediate inserting elements. The term “and / or” as used herein includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed.
[0110] It should be understood that references to "an embodiment," "an embodiment," "an aspect," or "may" in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Furthermore, particular features, structures, or characteristics may be suitably combined in one or more embodiments of the invention. The foregoing description is provided to enable those skilled in the art to implement the various aspects described herein. Various modifications will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.
[0111] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the language of the claims, wherein, unless expressly stated, an element referred to in the singular does not mean "one and only one," but rather "one or more." Unless expressly stated, the term "some" means one or more.
[0112] The scope of this invention should be determined based on the claims.
Claims
1. A method for dynamic speech equalization, comprising: Receives electrical input audio signal (12); The electrical input audio signal (12) is processed according to frequency and an equalized electrical audio signal (14) is provided according to an equalization function (100), wherein the equalization function includes an actuator section (110) and an analysis section (120), the actuator section being configured to - Apply the first compensation filter (111) to the received electrical input audio signal; and - Apply a second compensation filter (112) to the received electrical input audio signal, wherein the second compensation filter is a transparent filter; The analysis section (120) is configured to determine dynamic weights (w1, w2) and includes - Apply a first filter (121) to the received electrical input audio signal (12), wherein the first filter is a low-pass filter; and - Apply a second filter (122) to the received electrical input audio signal (12), wherein the second filter is a high-pass filter; The first output signal (111a) from the first compensation filter (111) and the second output signal (112a) from the second compensation filter (112) are respectively mixed based on the first and second dynamic weights. The transmission may be perceived by the user as sound, representing the electrical input audio signal (12) or an output signal (OUT) of its processed version; The analysis section (120) includes a mapping function (125) for mixing the dynamic weights (w1, w2) and updating the dynamic weights (w1, w2) for each electrical input audio signal (12), wherein the average energy of the second filter (122) is divided by the average energy of the first filter (121), and the result of the division is averaged and then entered into the mapping function (125).
2. The method according to claim 1, wherein, The first compensation filter (111) applies compensation to the damaged signal in the worst case, and the second compensation filter (112) applies compensation to the damaged signal in the best case.
3. The method according to claim 1, wherein, The method includes continuously updating the first and second dynamic weights, respectively.
4. The method according to any one of claims 1-3, wherein, The high-quality voice channel will have a high weight for the second output signal (112a), the low-quality voice channel will have a high weight for the first output signal (111a), and the moderately impaired voice channel will have a weight of 0.5 for both.
5. The method according to claim 1, wherein, The mapping function (125) is enabled when speech is detected in the audio signal.
6. A loudspeaker amplifier (200), comprising: The input unit (11) is used to receive an audio input signal (IN) having a first dynamic range level representing a sound signal with time and frequency variations and to provide an electrical input audio signal (12), the audio input signal including a target signal and / or a noise signal; A processing unit (13) is configured to modify the electrical input audio signal (12) according to its frequency and provide an equalized electrical audio signal (14) according to an equalization function (100), wherein the equalization function includes an actuator section (110) and an analysis section (120), the actuator section including... - A first compensation filter (111) dynamically applied to the received electrical input audio signal (12), the first compensation filter providing a first output signal (111a); and - A second compensation filter (112) is dynamically applied to the received electrical input audio signal (12), the second compensation filter being a transparent filter, the second compensation filter providing a second output signal (112a). The analysis section (120) is configured to determine dynamic weights (w1, w2) and includes - A first filter (121) applied to the received electrical input audio signal (12), wherein the first filter is a low-pass filter; and - A second filter (122) applied to the received electrical input audio signal (12), the second filter being a high-pass filter; The actuator portion includes a mixing unit (113) for mixing a first output signal (111a) from the first compensation filter (111) and a second output signal (112a) from the second compensation filter (112) based on first and second dynamic weights, respectively. Output unit (15) is used to provide an output signal (OUT) that can be perceived by the user as sound, representing the electrical input audio signal (12) or a processed version thereof; The analysis section (120) includes a mapping function (125) for mixing the dynamic weights (w1, w2) and updating the dynamic weights (w1, w2) for each electrical input audio signal (12), wherein the average energy of the second filter (122) is divided by the average energy of the first filter (121), and the result of the division is averaged and then entered into the mapping function (125).
7. The loudspeaker amplifier according to claim 6, wherein, The first compensation filter (111) applies compensation to the damaged signal in the worst case, and the second compensation filter (112) applies compensation to the damaged signal in the best case.
8. The loudspeaker amplifier according to claim 6, wherein, The first and second dynamic weights are continuously updated.
9. The loudspeaker amplifier according to any one of claims 6-8, wherein, The high-quality voice channel will have a high weight for the second output signal (112a), the low-quality voice channel will have a high weight for the first output signal (111a), and the moderately impaired voice channel will have a weight of 0.5 for both.
10. The loudspeaker amplifier according to claim 6, wherein, The mapping function (125) is enabled when speech is detected in the audio signal.