Method and system for intelligent dynamic voice emphasis

By performing voice detection and intelligent modulation gain control in multi-channel audio sources, dynamically adjusting voice gain is solved, and dialogue clarity in the prior art is difficult to maintain in the surround sound experience, achieving better conversation quality and auditory experience.

JP2025072287APending Publication Date: 2025-05-09HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024145185
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-24
Filing Date
2024-08-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When the prior art pursues a multi-channel surround sound experience, it often sacrifices dialogue clarity, making it difficult to improve the quality and clarity of the conversation while maintaining the surround sound effect.

Method used

Voice gain dynamically adjusts voice gain to optimize voice gain settings based on the ratio of the central channel signal to other channel signals and system volume levels by performing voice detection and intelligent modulation gain control in multi-channel audio sources.

Benefits of technology

It achieves improving the quality and clarity of conversations while maintaining the surround sound effect, avoiding auditory imbalance and dynamic loss caused by speech gain at high volumes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072287000001_ABST
    Figure 2025072287000001_ABST
Patent Text Reader

Abstract

To provide a method and a system for intelligent dynamic voice emphasis of an audio source.SOLUTION: A method for intelligent dynamic voice emphasis for an audio source includes steps of: detecting voice and performing intelligent emphasis gain control on multichannel audio source input for determining a voice emphasis gain; and applying the voice emphasis gain in dynamic loudness equalization performed in multichannel audio source input. Intelligent emphasis gain control includes steps of: setting the voice emphasis gain based on a ratio of signal power strength of a center channel signal with respect to a total of other channel signals; and setting the voice emphasis gain based on a level of a volume of the system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present subject matter relates generally to the field of audio signal processing, and more particularly to methods and systems for intelligent dynamic speech enhancement of an audio source. [Background technology]

[0002] With the prevalence of high definition cable television, online streaming, and large screen display devices, home theater experience is becoming more and more popular in the market. These media sources are usually equipped with multi-channel audio to provide users with a more immersive surround sound experience. However, since the main purpose of movie formats is to provide particularly immersive sound effects, audio clarity is often sacrificed in pursuit of surround sound.

[0003] Speech enhancement technologies play an important role in the film production process and are designed to improve the quality, audibility, and intelligibility of dialogue in film material.

[0004] Developments in existing speech enhancement technologies can be derived from surround sound technologies. For example, surround sound providers such as Dolby, THX, and DTS bring multi-channel audio encoding technology for better spatial resolution and stereo sound experience. This technology immerses the audience in a surround sound environment, but mixes dialogue, sometimes with less clarity. Furthermore, film production includes a mixing process in which a sound mixer is responsible for adjusting the volume, balance, and spatial placement of various audio elements, including dialogue, music, and sound effects, to create a realistic listening effect. However, the pursuit of an immersive surround sound experience tends to compromise dialogue clarity.

[0005] Therefore, there is a need for a method and system for intelligent dynamic speech enhancement of an audio source to overcome the above-mentioned shortcomings of existing solutions. Summary of the Invention [Means for solving the problem]

[0006] On the one hand, the subject matter of the invention provides a method for intelligent dynamic speech enhancement, the method for intelligent dynamic speech enhancement includes performing speech detection and intelligent enhancement gain control on a multi-channel audio source input to determine a speech enhancement gain, the multi-channel audio source input including a center channel signal and other channel signals. The method for intelligent dynamic speech enhancement further includes applying the speech enhancement gain in a dynamic loudness equalization performed on the multi-channel audio source input, the intelligent enhancement gain control including setting the speech enhancement gain based on a ratio of signal power strength of the center channel signal to a sum of the other channel signals, and setting the speech enhancement gain based on a system volume level.

[0007] In one embodiment, the subject invention provides a system for intelligent dynamic speech enhancement, the system for intelligent dynamic speech enhancement including a memory configured to store computer-executable instructions and a processor configured to execute the computer-executable instructions to implement a method for intelligent dynamic speech enhancement. The present invention provides, for example, the following items. (Item 1) 1. A method for intelligent dynamic speech enhancement, comprising: performing speech detection and intelligent emphasis gain control on a multi-channel audio source input to determine a speech emphasis gain, the multi-channel audio source input including a center channel signal and other channel signals; applying the speech enhancement gains in a dynamic loudness equalization performed on the multi-channel audio source input; The intelligent emphasis gain control includes: setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels; and setting the speech enhancement gain based on a system volume level. (Item 2) Setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels comprises: setting the speech enhancement gain high if the ratio of the signal power strength of the center channel to the sum of the other channels is small; and 5. The method of claim 1, further comprising setting the speech enhancement gain low if the ratio of the signal power strength of the center channel to the sum of the other channels is large. (Item 3) Setting the speech enhancement gain based on a level of system volume may include: recognizing a level of the system volume and setting different speech enhancement gains when the recognized system volume level is within different volume ranges; When the system volume level is within a low range, the speech enhancement gain is set high; 13. The method of claim 12, wherein the speech enhancement gain is set low when the system volume level is in a high range. (Item 4) The voice detection includes: extracting the signal of the center channel from the multi-channel audio source input; performing a normalization on the signal of the center channel; and 2. A method as claimed in claim 1, further comprising: performing a fast autocorrelation on the normalized signal of the centre channel, the result of the fast autocorrelation representing a detection confidence that indicates the likelihood that speech is present in the signal of the centre channel. (Item 5) The intelligent emphasis gain control includes: converting the detection confidence into the speech enhancement gain; and 13. The method of any of the preceding claims, further comprising performing a smoothing operation on the speech enhancement gains. (Item 6) 5. The method of claim 1, wherein the intelligent emphasis gain control further comprises performing a soft limiting process on the set speech emphasis gain. (Item 7) The dynamic loudness equalization performed on the multi-channel audio source input comprises: enhancing the loudness of the signal in the center channel and attenuating the loudness of the signals in the other channels based on the set speech enhancement gain; and 2. The method of claim 1, further comprising: performing a combining and mixing process on the emphasized signal of the center channel and the attenuated signals of the other channels to generate an output signal. (Item 8) 5. The method of claim 1, further comprising performing a crossover filtering process on the multi-channel audio source input prior to the dynamic loudness equalization performed on the multi-channel audio source input. (Item 9) performing the dynamic loudness equalization only on the multi-channel audio source input within a mid-frequency range; and 5. The method of claim 1, further comprising: combining and mixing the multi-channel audio source input in the mid-frequency range that has undergone the dynamic loudness balancing with the multi-channel audio source input in the low and high frequency ranges to generate an output signal. (Item 10) 1. A system for intelligent dynamic speech enhancement, comprising: a memory configured to store computer-executable instructions; and 2. The system comprising: one or more processors configured to execute the computer-executable instructions to implement the method for intelligent dynamic speech enhancement described in any of the preceding items. (Summary) The inventive subject matter provides a method and system for intelligent dynamic speech enhancement for an audio source, comprising performing speech detection and intelligent enhancement gain control on a multi-channel audio source input to determine a speech enhancement gain, and further comprising applying the speech enhancement gain in a dynamic loudness equalization performed on the multi-channel audio source input, the intelligent enhancement gain control including setting the speech enhancement gain based on a ratio of signal power strength of a center channel signal to a sum of the other channel signals, and setting the speech enhancement gain based on a system volume level.

[0008] The subject matter of the invention can be better understood by reading the following description of non-limiting embodiments with reference to the attached drawings, in which: [Brief description of the drawings]

[0009] [Figure 1] 1 illustrates a block diagram of a principle of intelligent speech enhancement according to one or more embodiments of the inventive subject matter; [Diagram 2] 1 exemplarily illustrates a schematic block diagram of a principle of audio detection according to one or more embodiments of the inventive subject matter. [Diagram 3]1 illustrates, by way of example, a schematic block diagram of the principle of intelligently controlling speech enhancement gain based on speech detection, in accordance with one or more embodiments of the inventive subject matter. [Figure 4] 10A-10C are exemplary illustrations of plots of non-linear volume correlation functions in accordance with one or more of the inventive subject matter. [Diagram 5] 1 illustrates an exemplary schematic diagram of a dynamic loudness balancing process in accordance with one or more embodiments of the inventive subject matter. [Figure 6] 1 illustrates generally a flowchart of a method for intelligent dynamic speech enhancement, in accordance with one or more embodiments of the inventive subject matter. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] It should be understood that the following description of the embodiments is given for illustrative purposes only and is not limiting. The division of the example functional blocks, modules or units shown in the accompanying drawings should not be construed as indicating that these functional blocks, modules or units must be implemented as physically separated units. The illustrated or described functional blocks, modules or units may be implemented as separate units, circuits, chips, functional blocks, modules or circuit elements. One or more functional blocks or units may also be implemented in a common circuit, chip, circuit element or unit.

[0011] Use of singular terms (such as, but not limited to, "one") is not intended to limit the number of items. Use of relative terms, such as "top," "bottom," "left," "right," "upper," "lower," "upper," "side," "first," "second" (e.g., "third"), "inlet," "outlet," and the like, are used in the written description to clarify specific reference to the accompanying drawings, and are not intended to limit the scope of the inventive subject matter or the appended claims, unless otherwise indicated. The terms "couple," "connecting," "coupled," "coupled," "coupler," and similar terms are used broadly herein and may include, without limitation, any method or apparatus that fastens, couples, adheres, fastens, attaches, combines, inserts, forms on or within, and communicates with, or is otherwise mechanically, magnetically, electrically, chemically, or operatively associated directly or indirectly with an intermediate element or one or more members, or may include one member that is integrally formed in a uniform manner with another member. Coupling can occur in any orientation, including rotational manner. The terms "comprise" and "such as" are exemplary rather than limiting, and unless otherwise specified, the term "may" means "may, but need not be." Regardless of the use of any other language in the subject matter, the embodiments illustrated in the accompanying drawings are examples given for purposes of illustration and description and are not the only embodiments of the subject matter herein.

[0012] In order to improve the quality of the audio output and thereby provide the user with a better listening experience, the subject matter of the present invention proposes a solution for active detection of human voice based on a detection confidence score and intelligent dynamic enhancement of the voice loudness of an audio source (e.g., a theater audio source).

[0013] For example, one method for voice enhancement to improve the quality, audibility, and clarity of dialogue in movie productions utilizes static equalization techniques. This method typically uses a static equalizer that adjusts the frequency response within the range of 200Hz to 4kHz to increase the volume and clarity of the dialogue range, thereby enhancing the voice dialogue. However, a drawback of this approach is that audio processing is active throughout the audio clipping, resulting in unbalanced sound even when there is no dialogue, while amplifying background noise. Another method for dynamic voice enhancement is to detect the voice signal at each time frame and apply adaptive audio processing based on the detection result. This method allows for sound enhancement when voice is detected to improve the clarity and intelligibility of the dialogue. However, this method requires a very accurate and fast voice detection algorithm to process the sound quickly.

[0014] However, there are some side effects of continuous speech enhancement when speech enhancement is applied continuously at high system volumes. For example, it may lead to perceptual imbalance and loss of dynamics of the listening experience. Moreover, in audio clips where only speech signals are present, the speech may already be clear and not require further enhancement. Therefore, to obtain a rational audio processing and enhancement method that improves the quality of dialogue in movie productions and allows listeners to better understand and recognize the movie content, the subject matter disclosed provides a method and system for further optimized intelligent dynamic speech enhancement.

[0015] The subject matter of the present invention focuses on multi-channel audio source inputs where the voice and environmental channels of the audio source can be easily extracted and directly analyzed.

[0016] 1 shows a block diagram 100 of the principles of intelligent speech enhancement according to one or more embodiments of the inventive subject matter. For ease of understanding, the inventive subject matter is illustrated mainly with reference to certain modules. It should be understood that the description referring to modules is intended to more clearly describe the solution and is not intended to be limiting.

[0017] The audio source inputs that may be processed by the system may include single channel source inputs, dual channel source inputs, and multi-channel audio source inputs. In a home theater example, the audio source inputs may be considered to be configured multi-channel audio source inputs. For example, in a 5.1 Dolby Surround Sound theater, the audio source inputs may typically include one bass channel and five surround channels, including a center channel, a left front channel, a right front channel, a left rear channel, and a right rear channel. In the case of a multi-channel audio source input, most of the audio signal is typically present in the center channel, and the other channels (i.e., the left front channel, the right front channel, the left rear channel, the right rear channel, etc.) may be considered to include surrounding environment channels. Thus, in the subject matter example, the audio detection processing of the audio detection module uses a method for detecting audio in the center channel.

[0018] 1, a multi-channel audio source signal first enters the system for intelligent dynamic speech enhancement from a multi-channel audio source input module 102. This multi-channel audio source input may enter a speech detection module 104 for speech detection.

[0019] FIG. 2 exemplarily illustrates a schematic block diagram of the principle of voice detection according to one or more embodiments of the inventive subject matter. The voice detection process of FIG. 2 may be performed, for example, by the voice detection module 104 of FIG. 1. In this voice detection module 104, a center channel extraction is first performed on the multi-channel audio source input, where most of the voice signal is present in the center channel. Then, normalization is performed on the extracted signal of the center channel, such that the input signals are adjusted to similar levels on a proportional basis. The normalized signal may be expressed, for example, by the following equation:

number

[0020] Then, a fast autocorrelation process is performed on the normalized signal, and an autocorrelation result is output. For example, in the fast autocorrelation process, a Fourier transform may be first performed on the normalized input signal using a short-time Fourier transform (STFT) method, and fast autocorrelation may be performed on the Fourier transformed signal. For example, in the case of the fast autocorrelation process, the following formulas (2) to (4) are referred to.

number

number

number

number

[0021] Next, Fig. 3 exemplarily illustrates a schematic block diagram of the principle of intelligently controlling speech enhancement gain based on speech detection according to one or more embodiments of the inventive subject matter. The processing of this step may correspond to the processing of the intelligent enhancement gain control module 106 of Fig. 1. As can be seen from the above description, the input to this intelligent enhancement gain control module is the detection confidence, which is itself the signal correlation result that needs to be converted into a speech enhancement gain.

[0022] The dynamic range of the enhancement gain may be defined as follows:

number

[0023] In some cases, the emphasis gain G is used to reduce audio distortion. i Further processing of G is required. i A smoothing process can be performed on the

[0024] In one or more embodiments of the inventive subject matter, the gain G iRegarding the above, when speech enhancement is applied continuously at high system volumes, the continuous enhancement of speech may lead to side effects in the listening experience, such as perceptual imbalance and loss of dynamics. Furthermore, in scenarios where only speech signals are present, for those audio clips where the speech is already clear, there is no longer any need to apply further gain processing to enhance the speech signal. Therefore, with regard to these raised side effects, the subject matter of the invention further provides intelligent logic decisions and intelligent enhancement gain processing to intelligently optimize the speech enhancement gain processing.

[0025] In an example, in an intelligent logical decision, the setting of the speech enhancement gain is determined by comparing the signal power strength of the center channel with the other channels. Specifically, in one or more embodiments of the present subject matter, the energy of the signals of the other channels is first extracted, summed, and then compared with the energy of the center channel. For example, if there is a voice signal in the center channel and the ambient sound of the other channels is loud, the speech enhancement gain needs to be set high, while, for example, if there is a voice signal in the center channel but the ambient sound of the other channels is weak, the speech enhancement gain needs to be set low. As represented by the following formula:

number

number

[0026] In other words, if the ratio of the signal power strength of the center channel to the sum of the other channels is small, the speech enhancement gain should be set high, and if the ratio of the signal power strength of the center channel to the sum of the other channels is large, the speech enhancement gain can be set low.

[0027] By performing a logical judgment of the ratio of the signal power strength of the center channel to the sum of the other channels and comparing the relative intensity of the environmental sound based on the voice dialogue of the audio clip, if the problem of perceptual imbalance in emphasizing audio sound is not resolved and the audibility and clarity is poor due to the environmental sound being strong and the human voice not being prominent enough, the user can obtain a better listening perception by further intelligently optimizing the gain of the human voice dialogue.

[0028] In an example, in the intelligent logic decision, the setting of the speech enhancement gain is determined by the system volume of the audio source, e.g., a movie production. The current system volume level can be recognized, and then different speech enhancement gains can be set within the recognized different volume ranges. For example, the current system volume level can be recognized, and if the system volume level is in a low range, the speech enhancement gain can be set high, while if the system volume level is in a high range, the speech enhancement gain should be set low. For example, if the current system volume level exceeds a threshold, the speech enhancement gain can be set to 1, i.e., the loudness level of the audio signal is not changed.

[0029] 4 shows an exemplary plot of a non-linear volume correlation function in accordance with one or more embodiments of the inventive subject matter. The volume ranges for the speech enhancement gains may be defined as follows:

number

[0030] Finally, a soft limiter can also be used for the limiting process. For example, a tanh function can be used as a soft limiter to reduce the enhancement gain after limiting.

number

number

[0031] α, β, and γ are limiter parameters that depend on the system configuration, where α can be a real number greater than zero, and β and γ can be non-zero real numbers. At this point,

number

[0032] By recognizing the system volume level of the film production and according to the system volume intensity of the audio clip, logical decision making as described above can be performed to avoid audio distortion caused by overemphasizing voice dialogue when the volume is emphasized to a greater intensity, thus enabling the dynamic audio gain to be more intelligently optimized, thereby improving the user's hearing.

[0033] Through the above logical judgment of two aspects, namely the ratio of the signal power strength of the center channel to the sum of the other channels, and the level of the system volume, the dynamic audio enhancement process can be intelligently optimized to further improve the user experience, thereby intelligently adjusting the audio enhancement for different scenarios in various film production sources. Therefore, the method can definitely enable users to obtain a better movie listening experience.

[0034] The intelligent dynamic processing defined in the subject matter of the invention is all aimed at voice enhancement. Since the frequency range of the human voice is also substantially within the mid-frequency range, e.g., 250-4000 Hz, the dynamic loudness equalization module 108 of the subject matter of the invention may mainly focus on the mid-frequency range of the input audio source signal for processing. Thus, a crossover filter may be used, for example, in FIG. 1, to first crossover the input audio signal, so that the input audio signal enters the crossover filter module 110 for crossover filtering to distinguish signals of different frequency ranges, such as low frequency signals, mid frequency signals, and high frequency signals, and then the separated mid-frequency signals may be made to enter the dynamic loudness equalization module 108 for dynamic loudness equalization processing. Furthermore, the intelligent emphasis gain control module may also be applied only to the voice signals within the mid-frequency range of the multi-channel audio source input and output an emphasis gain, while signals of other frequency ranges of the input signal remain unchanged. Separate processing of mid-frequency range signals and other high and low frequency range signals in the multi-channel audio source input can be achieved by crossover filtering processes, thereby reducing audio distortion that is not within the audio frequency range as much as possible.

[0035] FIG. 5 exemplarily illustrates a schematic diagram of a dynamic loudness equalization process according to one or more embodiments of the inventive subject matter. Optionally, the signal of the input multi-channel audio source after crossover filtering may be divided into signals in low, mid, and high frequency ranges by crossover. The dynamic loudness equalization module may perform dynamic loudness equalization only on signals in the mid frequency range. Here, for the center channel signal extracted by channel extraction, the loudness of the center channel signal may be enhanced, for example, based on an intelligent enhancement gain, while, optionally, the loudness of the signals of the other channels may be attenuated. For example, based on the intelligent enhancement gain, the center channel signal and the signals of the other channels may be separately enhanced and / or attenuated in different ratios, and then concatenated and mixed to generate a signal output. In addition, the low-frequency range and high-frequency range signals of the input multi-channel audio source after the crossover filtering process are not subjected to the dynamic loudness equalization process here, but are directly combined and mixed with the mid-frequency range signal after the dynamic loudness equalization to generate a signal output, for example, the signal is output by the signal output module 112 shown in Fig. 1. As a result, the loss of audibility or clarity caused by non-voice signals can be better reduced.

[0036] Returning to FIG. 1, as can be seen, the method for intelligent dynamic speech enhancement in the subject matter of the invention is mainly divided into two paths for processing, the upper path includes audio source input, speech detection, and intelligent enhancement gain, and this path is considered as a side chain processing flow for performing detection. The lower path includes a main chain processing flow consisting of audio source input, crossover filtering, and dynamic loudness equalization. The upper path and the lower path can be performed synchronously or asynchronously, which depends on the actual system capabilities and latency requirements. This approach of performing intelligent dynamic speech enhancement in a two-path manner can minimize latency and prevent audio distortion. When these two-path layers operate asynchronously, the signal can travel very fast throughout the system with little or no latency. On the other hand, estimating the enhancement gain at a relatively low rate can have higher accuracy and smoothness, which is quite helpful in preventing audio distortion.

[0037] FIG. 6 generally illustrates a flowchart 600 of a method for intelligent dynamic speech enhancement, in accordance with one or more embodiments of the inventive subject matter. As shown in FIG. 6, the method for intelligent dynamic speech enhancement may include receiving a multi-channel audio source input in step S610. Then, in step S620, speech detection is performed on the multi-channel audio source input in the side chain processing flow path. Then, in step S630, intelligent enhancement gain control is performed on the detected speech to determine a speech enhancement gain. In step S640, the determined speech enhancement gain is used to perform dynamic loudness balancing of the multi-channel audio source input in the main chain processing flow, and finally, in step S650, the method for intelligent dynamic speech enhancement may provide an intelligent speech enhancement audio signal output.

[0038] Additionally or alternatively, the steps of the method for intelligent dynamic voice enhancement shown in FIG. 6 may be implemented by one or more processors. The processor may be implemented as a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), a digital signal processor (DSP), discrete logic, or a combination thereof, and / or other types of circuits or logic. Similarly, instructions for implementing the method for intelligent dynamic voice enhancement shown in FIG. 6 may be stored in a memory, which may be a DRAM, SRAM, flash memory, or other type of memory. The parameters (e.g., conditions and thresholds) and other data structures may be stored and managed separately, may be integrated into a single memory or database, or may be logically and physically organized in many different ways. The programs and instruction sets may be part of a single program, or may be separate programs, or may be distributed across multiple memories and processors.

[0039] The method and system for intelligent dynamic speech enhancement defined in the subject matter of the invention can be used not only in consumer products such as sound bars and stereo speakers, but also in applications in occasions such as theaters and music halls. Compared with static audio equalization and static speech enhancement technology, the method and system for intelligent dynamic speech enhancement defined in the subject matter of the invention can improve speech clarity through intelligent gain control. In the method for intelligent dynamic speech enhancement defined in the subject matter of the invention, optimization is performed on existing speech enhancement technology to intelligently incorporate dynamic speech enhancement gain control in applications such as theaters, thereby enabling improved user experience of various movie clips, including but not limited to quick implementation of voice activity detection based on center channel, intelligent enhancement gain control that intelligently adjusts speech enhancement gain level according to movie content, intelligently adjusts speech enhancement gain level according to the volume of the playback system, and implementation of multi-channel intelligent dynamic loudness balancing using dual processing paths.

[0040] One or more example embodiments of the inventive subject matter are set forth in the following clauses.

[0041] Clause 1. performing speech detection and intelligent emphasis gain control on a multi-channel audio source input to determine a speech emphasis gain, the multi-channel audio source input including a center channel signal and other channel signals; applying the speech enhancement gain in a dynamic loudness equalization performed on the multi-channel audio source input, the intelligent enhancement gain control comprising: setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels; and and setting the speech enhancement gain based on a system volume level.

[0042] Clause 2. Setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels, setting the speech enhancement gain high if the ratio of the signal power strength of the center channel to the sum of the other channels is small; and 2. The method for intelligent dynamic speech enhancement described in clause 1, comprising setting the speech enhancement gain low if a ratio of the signal power strength of the center channel to the sum of the other channels is large.

[0043] Clause 3. Setting the speech enhancement gain based on a level of a system volume, recognizing a level of the system volume and setting different speech enhancement gains when the recognized system volume level is within different volume ranges; When the system volume level is within a low range, the speech enhancement gain is set high; 3. The method for intelligent dynamic speech enhancement described in clause 1 or 2, wherein the speech enhancement gain is set low when the system volume level is in a high range.

[0044] Clause 4. The voice detection extracting the signal of the center channel from the multi-channel audio source input; performing a normalization on the signal of the center channel; and 4. A method for intelligent dynamic speech enhancement as described in any one of clauses 1 to 3, comprising performing a fast autocorrelation on the normalized signal of the center channel, the result of the fast autocorrelation representing a detection confidence indicative of the likelihood that speech is present in the signal of the center channel.

[0045] Clause 5. The intelligent emphasis gain control comprises: converting the detection confidence into the speech enhancement gain; and 5. The method for intelligent dynamic speech enhancement of any one of claims 1 to 4, further comprising performing a smoothing process on the speech enhancement gains.

[0046] Clause 6. A method for intelligent dynamic speech enhancement as described in any one of clauses 1 to 5, wherein the intelligent enhancement gain control further comprises performing a soft limiting process on the set speech enhancement gain.

[0047] Clause 7. The dynamic loudness equalization performed on the multi-channel audio source input comprises: enhancing the loudness of the signal in the center channel and attenuating the loudness of the signals in the other channels based on the set speech enhancement gain; and 7. The method for intelligent dynamic speech enhancement described in any one of clauses 1 to 6, comprising performing a concatenation and mixing process on the enhanced signal of the center channel and the attenuated signals of the other channels to generate an output signal.

[0048] Clause 8. A method for intelligent dynamic speech enhancement as described in any one of clauses 1 to 7, further comprising performing a crossover filtering process on the multi-channel audio source input prior to the dynamic loudness equalization performed on the multi-channel audio source input.

[0049] Clause 9. Performing said dynamic loudness equalization only on said multi-channel audio source input within a mid-frequency range; and 9. The method for intelligent dynamic speech enhancement of any one of claims 1 to 8, further comprising combining and mixing the multi-channel audio source input in the mid-frequency range that has undergone the dynamic loudness balancing with the multi-channel audio source input in the low frequency range and the high frequency range to generate an output signal.

[0050] Clause 10. A system for intelligent dynamic speech enhancement, comprising: a memory configured to store computer-executable instructions; and The system comprises one or more processors configured to execute the stored computer-executable instructions to implement the method for intelligent dynamic speech enhancement described in any one of clauses 1 to 9.

[0051] The description of the embodiments is presented for purposes of illustration and description. Suitable modifications and alterations to the embodiments may be made in accordance with the above description or may be acquired by practicing the methods. For example, unless otherwise specified, one or more of the described methods may be performed by suitable devices and / or combinations of devices. The methods may be performed by executing instructions stored on one or more logic devices (e.g., processors) in conjunction with one or more additional hardware elements (storage devices, memories, hardware network interfaces / antennas, switches, actuators, clock circuits, etc.). The described methods and related acts may also be performed in various orders, in parallel, and / or simultaneously in addition to the order described in this application. The described systems are exemplary in nature and may include additional elements and / or omit elements. The inventive subject matter of the present disclosure includes all novel and non-obvious combinations of the various systems and configurations and components of the combinations, as well as other features, functions, and / or properties disclosed.

[0052] The various real-time scheme elements of the modules, elements, and components for implementing the methods defined in the subject invention may be fabricated as one or more electronic devices, including, but not limited to, arrays of fixed or programmable logic elements (e.g., transistors, gates, etc.) residing on the same chip or chipset. One or more elements of the various embodiments of the devices described herein may also be implemented in whole or in part as one or more instruction sets, which may be arranged to execute on one or more arrays of fixed or programmable logic elements (e.g., microprocessors, embedded processors, IP cores, digital signal processors, FPGAs, ASSPs, ASICs, etc.).

[0053] The terminology used herein has been selected to best explain the principles of the embodiments, practical applications, or improvements to the art found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0054] In the foregoing, embodiments presented in the inventive subject matter are identified by reference. However, the scope of the inventive subject matter is not limited to the specifically described embodiments. Rather, any combination of the above-described features and elements, whether with or without different embodiments, is contemplated to implement and practice the contemplated embodiments.

[0055] Moreover, while the embodiments disclosed herein may achieve advantages over other possible solutions and over the prior art, whether or not a given embodiment achieves a particular advantage does not limit the scope of the inventive subject matter. Thus, the foregoing aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly recited in the claims.

[0056] While the foregoing is directed to embodiments of the inventive subject matter, other and further embodiments of the inventive subject matter may be devised without departing from the essential scope of the inventive subject matter, the scope of the inventive subject matter being determined by the appended claims.

Claims

1. 1. A method for intelligent dynamic speech enhancement, comprising: performing speech detection and intelligent emphasis gain control on a multi-channel audio source input to determine a speech emphasis gain, the multi-channel audio source input including a center channel signal and other channel signals; applying the speech enhancement gains in a dynamic loudness equalization performed on the multi-channel audio source input; The intelligent emphasis gain control includes: setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels; and setting the speech enhancement gain based on a system volume level.

2. Setting the speech enhancement gain based on a ratio of signal power strength of the center channel to the sum of the other channels comprises: setting the speech enhancement gain high if the ratio of the signal power strength of the center channel to the sum of the other channels is small; and 2. The method of claim 1, comprising setting the speech enhancement gain low if a ratio of the signal power strength of the center channel to the sum of the other channels is large.

3. Setting the speech enhancement gain based on a level of system volume may include: recognizing a level of the system volume and setting different speech enhancement gains when the recognized system volume level is within different volume ranges; When the system volume level is within a low range, the speech enhancement gain is set high; The method of claim 1 , wherein the speech enhancement gain is set low when the system volume level is in a high range.

4. The voice detection includes: extracting the signal of the center channel from the multi-channel audio source input; performing a normalization on the signal of the center channel; and 2. The method of claim 1, comprising performing a fast autocorrelation on the normalized signal of the center channel, the result of the fast autocorrelation representing a detection confidence that indicates the likelihood that speech is present in the signal of the center channel.

5. The intelligent emphasis gain control includes: converting the detection confidence into the speech enhancement gain; and The method of claim 4 , further comprising performing a smoothing process on the speech enhancement gains.

6. The method of claim 1 , wherein the intelligent emphasis gain control further comprises performing a soft limiting operation on the set speech emphasis gain.

7. The dynamic loudness equalization performed on the multi-channel audio source input comprises: enhancing the loudness of the signal in the center channel and attenuating the loudness of the signals in the other channels based on the set speech enhancement gain; and The method of claim 2 or 3, comprising performing a combining and mixing process on the emphasized signal of the center channel and the attenuated signals of the other channels to generate an output signal.

8. The method of claim 7 , further comprising performing a crossover filtering process on the multi-channel audio source input prior to the dynamic loudness equalization performed on the multi-channel audio source input.

9. performing the dynamic loudness equalization only on the multi-channel audio source input within a mid-frequency range; and 9. The method of claim 8, further comprising combining and mixing the multi-channel audio source input in the mid-frequency range that has undergone the dynamic loudness balancing with the multi-channel audio source input in a low frequency range and a high frequency range to generate an output signal.

10. 1. A system for intelligent dynamic speech enhancement, comprising: a memory configured to store computer-executable instructions; and 10. The system comprising one or more processors configured to execute the computer-executable instructions to implement the method for intelligent dynamic speech enhancement according to any one of claims 1 to 9.