Method and system for dynamic speech enhancement

The method and system for dynamic speech enhancement address the issue of speech distortion in surround sound systems by using dual-path processing for accurate speech detection and volume adjustment, enhancing speech clarity and user experience.

JP7867393B2Active Publication Date: 2026-05-29HARMAN INT IND INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HARMAN INT IND INC
Filing Date
2022-07-08
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing audio enhancement methods, particularly for surround sound systems, often distort speech due to inaccurate speech detection and require significant system resources, leading to tonal imbalances and background amplification, making it difficult to understand dialogue in video content.

Method used

A method and system for dynamic speech enhancement that performs simultaneous signal processing on two paths, including dynamic volume balancing based on gain control parameters and speech detection confidence levels, allowing for synchronous or asynchronous processing to enhance speech intelligibility.

Benefits of technology

The method and system effectively improve speech clarity and user experience by dynamically adjusting volume based on speech detection, minimizing distortion and maintaining overall audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007867393000003
    Figure 0007867393000003
  • Figure 0007867393000004
    Figure 0007867393000004
  • Figure 0007867393000005
    Figure 0007867393000005
Patent Text Reader

Abstract

To provide a method and system for dynamic voice enhancement of an audio source.SOLUTION: A method of processing audio source input signals includes: first path signal processing including receiving an audio source input and performing dynamic loudness balancing on the audio source input based on a first gain control parameter; and second path signal processing including performing voice detection on the audio source input and calculating a detection confidence, and calculating a second gain control parameter based on the detection confidence. The first gain control parameter is updated with the second gain control parameter calculated by a second processing path. The first path signal processing and the second path signal processing may be performed synchronously or asynchronously.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates in general to the field of audio signal processing, and more particularly to methods and systems for dynamic speech enhancement of audio sources. [Background technology]

[0002] With the emergence of new ways of consuming media such as high-definition cable television and online streaming, as well as large-screen televisions and displays, the movie experience is gaining popularity in the consumer market. These media sources often involve what is called multi-channel audio technology, or more commonly, surround sound technology. Surround sound providers such as Dolby, THX, and DTS have their own unique multi-channel audio encoding technologies that provide better spatial audio solutions for source content. Since the primary purpose of movie-format content is to provide an immersive surround sound experience, it is often preferred to opt for the surround experience at the expense of voice intelligibility. While this offers significant advantages in terms of immersion and spatial solutions, it often results in a degradation of sound quality, sometimes making it difficult to even understand the video content. To improve the quality of the audio source of the video content and enhance intelligibility and audibility, methods of audio enhancement are often applied to video content.

[0003] A common existing method for speech enhancement utilizes static equalization. This method applies static equalization only to audio channels between approximately 200Hz and 4kHz to increase the volume of the speech band. While this implementation requires minimal system resources, the distortion it produces is immediately apparent. This implementation always works even when there is no speech or dialogue in the clip, leading to tonal imbalances and background amplification. More advanced methods first detect speech within each time frame and then automatically process the audio signal based on the detection results. This one-way approach requires accurate speech detection and fast system processing response. However, some existing methods fail to detect speech quickly and accurately, often distorting the signal and resulting in an unpleasant listening experience.

[0004] Therefore, an improved technical solution is needed that overcomes the aforementioned shortcomings of existing solutions. [Overview of the project] [Means for solving the problem]

[0005] A method for dynamic speech enhancement is provided according to aspects of the present disclosure. The method may include performing first path signal processing, which includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter. The method may also include performing second path signal processing, which includes performing speech detection on the audio source input and calculating a detection confidence level indicating the likelihood of speech in the audio source input, and calculating a second gain control parameter based on the detection confidence level. The method may further include updating the first gain control parameter with the second gain control parameter and performing first path signal processing based on the updated first gain control parameter.

[0006] According to one or more embodiments, the audio source input may include a multi-channel source input, and performing speech detection on the audio source input and calculating the detection confidence may include extracting a center channel signal from the multi-channel source input, performing normalization on the center channel signal, and performing fast autocorrelation on the normalized center channel signal, the result of which represents the detection confidence.

[0007] According to one or more embodiments, calculating a second gain control parameter based on detection confidence may include calculating a second gain control parameter based on a logarithmic function of detection confidence, smoothing the calculated second gain control parameter, and limiting the smoothed second gain control parameter.

[0008] According to one or more embodiments, the audio source input may include a multi-channel source input, and performing dynamic volume balancing on the audio source input includes: extracting a center channel signal from the multi-channel source input; increasing the volume of the center channel signal; decreasing the volume of the other channel signals based on a first gain control parameter or an updated first gain control parameter; and combining and mixing the increased center channel signal and the decreased other channel signals to generate an output signal.

[0009] According to one or more embodiments, the method may also include performing crossover filtering on the audio source input before performing dynamic volume balancing.

[0010] According to one or more embodiments, the method may also include performing dynamic volume balancing only on the signal in the intermediate frequency range of the audio source input, and after dynamic volume balancing, combining and mixing the low-frequency and high-frequency signals of the audio source input with the signal in the intermediate frequency range of the audio source input to generate an output signal.

[0011] According to one or more embodiments, the audio source input also includes a dual-channel source input, and the method also includes generating a multi-channel source input based on the dual-channel source input.

[0012] According to one or more embodiments, generating a multi-channel source input based on a dual-channel source input may include performing cross-correlation between the left channel signal and the right channel signal from the dual-channel source input, and generating the multi-channel source input according to a combination ratio, the combination ratio of which depends on the result of the cross-correlation.

[0013] According to one or more embodiments, the first path signal processing and the second path signal processing are synchronous or asynchronous.

[0014] According to another aspect of this disclosure, a system for speech enhancement is provided, including memory and a processor. The memory is configured to store computer-executable instructions. The processor is configured to execute instructions and implement the method described above.

[0015] This disclosure can be better understood by referring to the attached drawings for a description of the following non-limiting implementations. For example, this application provides the following items. (Item 1) A method of dynamic speech enhancement, Performing a first path signal processing, wherein the first path signal processing includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter, The second path signal processing is performed, and the second path signal processing is performed Performing voice detection on the audio source input and calculating a detection reliability, the second path signal processing being performed, wherein the detection reliability indicates the likelihood of the voice in the audio source input, Calculating a second gain control parameter based on the detection reliability, Updating the first gain control parameter with the second gain control parameter and performing the first path signal processing based on the updated first gain control parameter, The method comprising the above. (Item 2) The audio source input includes a multi-channel source input, and performing voice detection on the audio source input and calculating a detection reliability includes Extracting a center channel signal from the multi-channel source input, Performing normalization on the center channel signal, Performing fast autocorrelation on the normalized center channel signal, the result of the fast autocorrelation representing the detection reliability, the method according to the above item including performing the fast autocorrelation. (Item 3) Calculating a second gain control parameter based on the detection reliability includes Calculating the second gain control parameter based on the logarithmic function of the detection reliability, Smoothing the calculated second gain control parameter, Limiting the smoothed second gain control parameter, the method according to any one of the above items including the above. (Item 4) The audio source input includes a multi-channel source input, and performing the dynamic volume balancing on the audio source input includes Extracting a center channel signal from the multi-channel source input, To increase the volume of the center channel signal and to reduce the volume of other channel signals based on the first gain control parameter or the updated first gain control parameter, The method according to any one of the above items, comprising combining and mixing the augmented center channel signal and the reduced other channel signals to generate an output signal. (Item 5) The method according to any one of the above items, further comprising performing crossover filtering on the audio source input before performing the dynamic volume balancing described above. (Item 6) The dynamic volume balancing described above is applied only to the signal in the intermediate frequency range of the audio source input. The method according to any one of the above items, further comprising: after the above dynamic volume balancing, combining and mixing the low-frequency and high-frequency signals of the audio source input with the intermediate-frequency signal of the audio source input to generate the above output signal. (Item 7) The method according to any one of the above items, wherein the above audio source input further includes a dual channel source input, and the method further includes generating a multi-channel source input based on the dual channel source input. (Item 8) The multi-channel source input is generated based on the dual-channel source input described above. This involves performing cross-correlation between the left channel signal and the right channel signal from the dual-channel source input described above, This includes generating the above multi-channel source input according to the combination ratio, The above combination ratio is determined by the method described in any one of the above items, based on the results of the above cross-correlation. (Item 9) The method according to any one of the above items, wherein the first path signal processing and the second path signal processing are synchronous or asynchronous. (Item 10) A dynamic speech enhancement system, Memory configured to store computer-executable instructions, The system comprises a processor configured to execute the above-mentioned computer-executable instructions in order to implement the method described in any one of the above items. (Summary) This disclosure provides a method and system for speech enhancement. The method and system of this disclosure can simultaneously perform signal processing on two paths with respect to an input signal. The first path signal processing includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter. The second path signal processing includes performing speech detection on the audio source input and calculating detection confidence and calculating a second gain control parameter based on detection confidence. The first and second path signal processing may be synchronous or asynchronous. The method of this disclosure includes updating the first gain control parameter with a second gain control parameter calculated by the second processing path and performing the first path signal processing based on the updated first gain control parameter. [Brief explanation of the drawing]

[0016] [Figure 1] A schematic block diagram of speech enhancement according to one or more embodiments of the implementation of the present disclosure is schematically shown. [Figure 2] A schematic block diagram illustrating one or more embodiments of the present disclosure of voice detection is shown as an example. [Figure 3] A schematic block diagram illustrating gain estimation based on speech detection according to one or more embodiments of the present disclosure is shown. [Figure 4] A schematic diagram illustrating a dynamic volume balancing process according to one or more embodiments of the present disclosure is shown. [Figure 5] A schematic diagram of speech enhancement according to one or more embodiments of another implementation of the present disclosure is shown. [Figure 6]Figure 5 shows a schematic diagram of a dynamic volume balancing process according to one or more embodiments of the implementation configuration. [Figure 7] The following schematic diagram illustrates a process by which one or more embodiments of the present disclosure generate a multichannel source input based on a dual-channel source input when the source input is a dual-channel source input. [Figure 8] A schematic representation of one or more embodiments of the present disclosure of a method for dynamic speech enhancement is provided. [Modes for carrying out the invention]

[0017] It should be understood that the following description of this embodiment is for illustrative purposes only and not limiting. The division of examples of functional blocks, modules, or units shown in the figures should not be construed as meaning that these functional blocks, modules, or units must be implemented as physically separate units. The illustrated or described functional blocks, modules, or units may be implemented as separate units, circuits, chips, functional blocks, modules, or circuit elements. One or more functional blocks or units may also be implemented in a common circuit, chip, circuit element, or unit.

[0018] The use of singular nouns (e.g., "a") is not intended to limit the number of items. Related terms, such as "top," "bottom," "left," "right," "upper," "lower," "down," "up," "side," "first," "second" ("third," etc.), "entrance," "exit," etc., but not limited to these, are used in descriptions provided to clarify the description in specific references to drawings, unless otherwise noted, and are not intended to limit the scope of the claims of this disclosure or the appended claims. The terms “couple,” “coupling,” “being coupled,” “coupled,” and “coupler,” and similar terms are used herein and may include, but are not limited to, any method or device for fixing, joining, bonding, fastening, attaching, relating, inserting, forming on or in, communicating, or otherwise directly or indirectly, mechanically, magnetically, electrically, chemically, and operably relating to an intermediate element and one or more members, or a single member integrally molded with another member in a unified manner. Coupling may occur in any direction, including rotation. The terms “including” and “such as” are illustrative, not restrictive, and the word “may” means “may, but not necessarily” unless otherwise stated. Any other terms used herein are given, but the illustrated embodiments are illustrative and explanatory examples and are not the only embodiments of the subject matter herein.

[0019] To overcome the shortcomings of existing technical solutions and improve the quality of audio output to provide a better user experience, this disclosure proposes a solution for actively detecting human voices in an audio source (e.g., a theater audio source) based on a detection confidence level indicating the possibility of a voice in the audio source input, and dynamically enhancing the volume of the voice. The method and system of this disclosure can simultaneously perform signal processing in two paths with respect to the input signal. The first path signal processing includes receiving the audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter. The second path signal processing includes performing voice detection on the audio source input and calculating a detection confidence level, and calculating a second gain control parameter based on the detection confidence level. The first and second path signal processing may be synchronous or asynchronous. The method of this disclosure also includes updating the first gain control parameter with a second gain control parameter calculated by the second processing path and performing the first path signal processing based on the updated first gain control parameter. The methods and systems disclosed herein can enhance speech intelligibility and improve the user experience when using audio products.

[0020] Methods and systems for dynamic speech enhancement according to various embodiments of various implementations of the present disclosure are described in detail below with reference to the accompanying drawings. Figure 1 shows a schematic block diagram of a speech method and system according to one or more embodiments of the implementations of the present disclosure. For ease of understanding, the present disclosure will be described with reference to several modules by which the main processing steps of the method and system are located. Those skilled in the art will understand that the reference to the module descriptions is for the purpose of making the solutions clearer and not for the purpose of limitation.

[0021] Figure 1 shows a schematic diagram of one or more embodiments of the implementation of the present disclosure. In one or more embodiments shown in Figure 1, the method and system for processing an audio source input signal according to the present disclosure includes a source input module 102, a dynamic volume balancing module 104, a signal output module 106, a voice detection module 108, and a gain control module 110. As can be seen from Figure 1, the method and system of the present disclosure can simultaneously perform signal processing on two paths to the input signal. The first path signal processing is primarily used to perform dynamic volume balancing on the received source input signal. The second path signal processing is used to perform voice detection on the received source input signal and estimate the gain. The first and second path signal processing may be performed synchronously or asynchronously. This depends on the processing capacity and the latency requirements of the actual system. This dual path processing design for the source input signal minimizes delay and prevents audio distortion. For example, if the first and second path signal processing are performed asynchronously, the signal can pass through the entire system quickly and with low latency. On the other hand, the gain can be estimated to be relatively low. As a result, the estimated gain is more accurate and smoother, which greatly helps in preventing audio distortion.

[0022] For example, referring to Figure 1, the first routing signal processing may include receiving an audio source input signal via the source input module 102 and performing dynamic balancing on the received audio source input signal based on the current gain control parameters using the dynamic volume balancing module 104. The second routing processing may include detecting the audio source input signal received from the input module 102 using the voice detection module 108 and calculating the detection confidence. For example, the second routing processing may also include enabling the gain control module 110 to estimate new gain control parameters based on the calculated detection confidence.

[0023] The dynamic volume balancing module 104 can update the currently used gain control parameters using the new gain control parameters estimated by the gain control module 110. Therefore, the dynamic volume balancing module 104 can perform first path signal processing based on the updated gain control parameters. That is, the dynamic volume balancing module 104 can perform dynamic volume balancing on the received audio source input signal based on the updated gain control parameters. The audio signal after dynamic volume balancing can be output via the signal output module 106.

[0024] Audio source inputs may include multi-channel source inputs, dual-channel source inputs, and single-channel source inputs. Processing procedures for different source inputs are described below with reference to the accompanying drawings. Figure 2 illustrates a schematic block diagram of speech detection according to one or more embodiments of the present disclosure, where the audio input source includes a multi-channel source input. The speech detection process shown in Figure 2 may be performed, for example, by the speech detection module 108 in Figure 1. As shown in Figure 2, center channel extraction is performed first; that is, the center channel signal is extracted from the multi-channel source input. Typically, most of the audio signal resides in the center channel. Next, the extracted center channel signal is normalized so that the input signals are adjusted to similar levels. For example, the normalized signal is represented by the following equation (1). x i_norm (n) = (x i (n)-μ i ) / σ i (1) In the formula, x i (n) represents the input signal at the nth sampling point of the i-th timeframe, and x i_norm (n) represents the output signal at the nth sampling point of the i-th time frame, i.e., the normalized signal μ. i and σ i These are the mean and variance of the input signal corresponding to the i-th time frame.

[0025] Next, perform high-speed autocorrelation processing on the normalized signal and output the result of the autocorrelation. For example, the high-speed autocorrelation processing may first perform a Fourier transform on the normalized input signal by using the short-time Fourier transform (STFT) method, and then perform high-speed autocorrelation on the Fourier transform signal. For example, the high-speed autocorrelation processing procedure is shown in the following equations (2) to (4).

Equation

Equation

[0026] FIG. 3 exemplarily shows a schematic block diagram of a method and system for estimating dynamic gain based on voice detection according to one or more embodiments of the present disclosure. The process for estimating dynamic gain based on voice detection shown in FIG. 3 may be performed, for example, by the gain control module 110 of FIG. 1. For example, referring to the process shown in FIG. 2, the detection reliability C i generated via the voice detection module 108 acts as an input to the gain control module 110. Based on the input, the gain for voice (hereinafter may also be referred to as a gain control parameter) is processed within the gain control module 110 and then output as an input to the dynamic volume balancing module 104. In some examples, the dynamic range of the gain is calculated by the following equation (5). G i =D0*ln(C i+D1) (5) G in the formula i represents the output of the dynamic control module, D0 and D1 are control parameters for the dynamic gain variation range, which can be real numbers greater than 0, and ln() is the natural logarithm function. In some examples, G i This can be provided to the dynamic volume balancing module 104 as the output from the gain control module 110.

[0027] In some other examples, G i This may be further processed and then act as an output from the gain control module 104. For example, G i This is smoothed to reduce audio distortion. In addition, a soft limiter is used to reduce the gain G. i_lim However, it is also possible to ensure that it remains within a reasonable range. For example, the tangent function in equation (6) below can be used as a soft limiter. G i_lim =tanh(αG i +β)+γ (6) In the formula, α, β, and γ are limiter parameters, which depend on the system configuration. α may be a real number greater than 0, and β and γ may be non-zero real numbers. At present, G i_lim This can act as an output from the gain control module 110.

[0028] Figure 4 illustrates a schematic diagram of a dynamic volume balancing process for each channel according to one or more embodiments of the present disclosure. The dynamic volume balancing process in Figure 4 may be performed by a dynamic volume balancing module 104. For example, after receiving a multi-channel source input, the dynamic volume balancing module 104 first performs channel extraction to extract the center channel signal. Next, it increases the volume of the center channel signal and decreases the volume of the other channel signals based on a gain control parameter. Then, it combines and mixes the increased center channel signal and the decreased other channel signals to generate an output signal. The gain control parameter may be a current gain control parameter or an updated gain parameter. For example, if the first path signal processing and the second path signal processing are synchronous, the gain control parameter used for dynamic volume balancing of the signal in the current time frame (e.g., the i-th time frame) may be a real-time updated calculated gain control parameter, e.g., a real-time updated G i or G i_lim Therefore, if the first path signal processing and the second path signal processing are asynchronous, the speed of the second path signal processing, which includes speech detection and gain estimation, is relatively slow, so the gain control parameter used for dynamic volume balancing of the signal in the current time frame (e.g., the i-th time frame) is G i-n or G i-n_limThese may be gain control parameters used for dynamic volume balancing of signals in previous timeframes, where n is an integer greater than 0, and its value may vary depending on the actual processing power of the system or the practical experience of the engineer. Furthermore, based on the current / updated gain control parameters, the signals of the center channel and the other channels may be boosted and reduced by different ratios. That is, boost control parameters for boosting the volume of the center channel signal and attenuation control parameters for reducing the volume of the center channel signal may be further determined based on the current / updated gain control parameters, respectively. For example, the boost and attenuation control parameters may be determined by proportional calculations, function calculations, or other calculation methods set by the engineer according to system requirements or experience. As a result, the overall system volume remains unchanged, while the volume of each channel is dynamically balanced.

[0029] Figure 5 shows a schematic diagram of a method and system according to one or more embodiments of another implementation of the present disclosure. In one or more embodiments shown in Figure 5, the method and system for processing an audio source input signal according to the present disclosure includes a source input module 502, a dynamic volume balancing module 504, a signal output module 506, a voice detection module 508, and a gain control module 510. These modules operate in substantially the same manner as the corresponding modules 102-110 in Figure 1. Furthermore, the method and system shown in Figure 5 may further include a crossover filtering module 512. It will be understood that the difference between the processing procedure shown in Figure 5 and the processing procedure described above with reference to Figures 1-4 is that crossover filtering is added to the first signal path. Thus, the source input signal received from the input module 502 is first processed by the crossover filtering module 512 and then processed by the dynamic volume balancing module 504 for dynamic volume balancing. Since the frequency range of the human voice is basically in the intermediate frequency range, a crossover filter is selected to process the input signal and distinguish signals in different frequency ranges. Therefore, gain control is applied only to signals in the intermediate frequency range of the input signal, while signals in other frequency ranges of the input signal remain unchanged. Through the added crossover filtering, dynamic volume balancing can be performed only on signals in the intermediate frequency range of the source input signal to avoid distortion as much as possible in the non-voice frequency range. To save space, only the different components of the embodiments shown in Figures 5 and 1 will be described below. For other identical components, please refer to Figures 1-4 and their related descriptions.

[0030] Figure 6 shows a schematic diagram of the dynamic volume balancing process according to one or more embodiments of the implementation configuration of Figure 5. As shown in Figure 6, the source input signal after crossover filtering may contain signals in the intermediate frequency, high frequency, and low frequency ranges. Next, dynamic volume balancing is performed only on the signals in the intermediate frequency range. Dynamic volume balancing includes channel extraction to extract the center channel signal. The volume of the center channel signal is then increased, and the volumes of the other channel signals are reduced based on the current / updated gain control parameters. The low-frequency and high-frequency signals of the multi-channel source input signal are not affected by dynamic volume balancing, but after dynamic volume balancing, they are directly coupled and mixed with the signal in the intermediate frequency range to generate the output signal. Therefore, distortion caused by non-voice signals may be preferable to avoid.

[0031] Several processing steps performed when the source input is a multi-channel source input with a center channel will be described with reference to Figures 1-6. Those skilled in the art will understand from this disclosure that when the source input is a single-channel input, the processing steps shown in Figures 1-6 can also be performed, and the center channel extraction process can be omitted. That is, the signal processing for the two paths described above is performed directly on the single-channel source input.

[0032] In cases where the source input is a dual-channel source input, it is necessary to add a center extraction process before implementing the methods and systems described above so that a multi-channel source input is generated based on the dual-channel source input. Figure 7 schematically shows a process for generating a multi-channel source input based on a dual-channel source input, according to one or more embodiments of the present disclosure.

[0033] The upmixing process shown in Figure 7 may incorporate a center extraction algorithm because it outputs a multi-channel source input based on a dual-channel source input. The center extraction algorithm may include, for example, calculating the cross-correlation between the left and right channel input signals and combining the left and right channel input signals into a center channel signal, the combination ratio depending on the cross-correlation, referencing equation (7) below. Center(n) = θ * corr(left(n), right(n)) * (left(n) + right(n)) (7) In the formula, left(n) is the left channel input signal, right(n) is the right channel input signal, center(n) is the center channel signal, corr() represents the cross-correlation function, and θ is actually an adjustment parameter, which is greater than 0 and less than or equal to 1.

[0034] Figure 8 schematically illustrates a method for dynamic speech enhancement according to one or more embodiments of the present disclosure. As shown in Figure 8, the method includes performing a first path signal processing. The first path signal processing includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter S802. The method also includes performing a second path signal processing. The second path signal processing includes performing speech detection on the audio source input and calculating a detection confidence S804, and calculating a second gain control parameter based on the detection confidence S806. The method may also include updating the first gain control parameter with a second gain control parameter S808, and performing the first path signal processing based on the updated first gain control parameter S802. The method shown in Figure 8 may be performed by a processor.

[0035] The methods and systems provided in this disclosure may be applied not only to consumer products such as soundbars and stereo speakers, but also to products for cinematic applications such as theaters and concert halls. The methods and systems provided in this disclosure can enhance speech intelligibility and improve the user experience of audio products and applications. The methods and systems described above in this disclosure with respect to the accompanying drawings may be implemented by a processor.

[0036] Article 1 A method for dynamic speech enhancement, comprising: performing a first path signal processing, the first path signal processing comprising receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter; performing a second path signal processing, the second path signal processing comprising performing speech detection on the audio source input and calculating a detection confidence, the detection confidence indicating the possibility of speech in the audio source input; calculating a second gain control parameter based on the detection confidence; updating the first gain control parameter with the second gain control parameter; and performing the first path signal processing based on the updated first gain control parameter.

[0037] Article 2 The method according to Clause 1, wherein the audio source input includes a multichannel source input, and the process of performing voice detection on the audio source input and calculating the detection confidence includes extracting a center channel signal from the multichannel source input, normalizing the center channel signal, and performing fast autocorrelation on the normalized center channel signal, the result of which represents the detection confidence.

[0038] Article 3 The method according to any one of clauses 1 to 2, wherein calculating a second gain control parameter based on detection confidence comprises calculating a second gain control parameter based on a logarithmic function of detection confidence, smoothing the calculated second gain control parameter, and limiting the smoothed second gain control parameter.

[0039] Article 4 The method according to any one of claims 1 to 3, wherein the audio source input includes a multi-channel source input, and performing dynamic volume balancing on the audio source input includes extracting a center channel signal from the multi-channel source input, increasing the volume of the center channel signal and decreasing the volume of the other channel signals based on a first gain control parameter or an updated first gain control parameter, and combining and mixing the increased center channel signal and the decreased other channel signals to generate an output signal.

[0040] Article 5 The method according to any one of clauses 1 to 4, further comprising performing crossover filtering on the audio source input before performing dynamic volume balancing.

[0041] Article 6 The method according to any one of clauses 1 to 5, further comprising: performing dynamic volume balancing only on the signal in the mid-frequency range of the audio source input; and after dynamic volume balancing, combining and mixing the low-frequency and high-frequency signals of the audio source input with the signal in the mid-frequency range of the audio source input to generate an output signal.

[0042] Article 7 The method according to any one of the clauses 1 to 6, wherein the audio source input further includes a dual-channel source input, and the method further includes generating a multi-channel source input based on the dual-channel source input.

[0043] Article 8 The method according to any one of the clauses 1 to 7, wherein generating a multichannel source input based on a dual-channel source input includes performing cross-correlation between the left channel signal and the right channel signal from the dual-channel source input, and generating a multichannel source input according to a combination ratio, wherein the combination ratio depends on the result of the cross-correlation.

[0044] Article 9 The method according to any one of the clauses 1 to 8, wherein the first path signaling and the second path signaling are synchronous or asynchronous.

[0045] Clause 10 A system for dynamic speech enhancement comprising: memory configured to store computer-executable instructions; and a processor configured to execute computer-executable instructions in order to implement the method described in any one of clauses 1 to 9.

[0046] The descriptions of implementations have been presented for illustrative and illustrative purposes. The implementations may be appropriately modified and altered in accordance with the above descriptions, or such modifications and alterations may be obtained by practicing the methods. For example, unless otherwise indicated, one or more of the described methods may be carried out by appropriate devices and / or combinations of devices. The methods may be carried out by using one or more logic devices (e.g., processors) in combination with one or more additional hardware elements (e.g., storage devices, memory, hardware network interfaces / antennas, switches, actuators, clock circuits) to execute stored instructions. The described methods and associated actions may also be performed in parallel and / or simultaneously in various orders other than those specified in this application. The described systems are essentially illustrative and may include and / or omit additional elements. The subject matter of this disclosure includes all disclosed various systems and configurations, as well as novel and non-obvious combinations of other features, functions, and / or characteristics.

[0047] The system may include additional or different logic and may be implemented in many different ways. The processor may be implemented as a microprocessor, microcontroller, application-specific integrated circuit (ASIC), digital signal processing processor (DSP), discrete logic, or a combination of these and / or other types of circuits or logic. Similarly, memory may be DRAM, SRAM, flash memory, or other types of memory. Parameters (e.g., conditions and thresholds) and other data structures may be stored and managed separately, combined in a single memory or database, or organized logically and physically in many different ways. The set of programs and instructions may be a single program or part of separate programs, or distributed across multiple memories and processors.

[0048] In this application, elements or steps listed in the singular and preceded by the word "a" or "one" should be understood not to exclude any multiple aforementioned elements or steps unless an exclusion is indicated. Furthermore, references in this disclosure to "one implementation" or "one example" are not intended to be construed as excluding the existence of additional implementations that also incorporate the enumerated features. The present invention has been described above with reference to specific implementations. However, those skilled in the art will understand that various modifications and changes can be made herein without departing from the broader spirit and scope of the invention as described in the appended claims.

Claims

1. A method of dynamic speech enhancement, Performing a first path signal processing, wherein the first path signal processing includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter, The second path signal processing is performed, wherein the second path signal processing is performed This includes performing voice detection on the audio source input and calculating the detection confidence level, wherein the detection confidence level indicates the possibility of voice in the audio source input. The second gain control parameter is calculated based on the aforementioned detection confidence level, Updating the first gain control parameter with the second gain control parameter, and performing the first path signal processing based on the updated first gain control parameter, Includes, The aforementioned audio source input includes a multi-channel source input, and the audio source input is used to perform the aforementioned voice detection and calculate the detection confidence level. Extracting the center channel signal from the aforementioned multi-channel source input, The center channel signal is normalized, A method comprising performing fast autocorrelation on the normalized center channel signal, wherein the result of the fast autocorrelation represents the detection confidence.

2. A method for dynamic speech enhancement, Performing a first path signal processing, wherein the first path signal processing includes receiving an audio source input and performing dynamic volume balancing on the audio source input based on a first gain control parameter, The second path signal processing is performed, wherein the second path signal processing is performed This includes performing voice detection on the audio source input and calculating the detection confidence level, wherein the detection confidence level indicates the possibility of voice in the audio source input. The second gain control parameter is calculated based on the aforementioned detection confidence level, Updating the first gain control parameter with the second gain control parameter, and performing the first path signal processing based on the updated first gain control parameter, Includes, The second gain control parameter is calculated based on the aforementioned detection confidence level. The second gain control parameter is calculated based on the logarithmic function of the detection confidence, The calculated second gain control parameter is smoothed, A method comprising limiting the smoothed second gain control parameter.

3. The audio source input includes a multi-channel source input, and the dynamic volume balancing is performed on the audio source input. Extracting the center channel signal from the aforementioned multi-channel source input, To increase the volume of the center channel signal and to reduce the volume of other channel signals based on the first gain control parameter or the updated first gain control parameter, The method according to claim 1, comprising combining and mixing the augmented center channel signal and the reduced other channel signals to generate an output signal.

4. The method according to claim 1, further comprising performing crossover filtering on the audio source input before performing the dynamic volume balancing.

5. The dynamic volume balancing is performed only on the signal in the intermediate frequency range of the audio source input. The method according to claim 4, further comprising: after the dynamic volume balancing, combining and mixing the low-frequency and high-frequency signals of the audio source input with the intermediate-frequency signal of the audio source input to generate an output signal.

6. The method according to claim 1, wherein the audio source input further includes a dual-channel source input, and the method further includes generating a multi-channel source input based on the dual-channel source input.

7. To generate the multi-channel source input based on the dual-channel source input, The cross-correlation between the left channel signal and the right channel signal from the dual-channel source input is performed, This includes generating the multi-channel source input according to the combination ratio, The method according to claim 6, wherein the combination ratio is determined by the results of the cross-correlation.

8. The method according to claim 1, wherein the first path signal processing and the second path signal processing are synchronous or asynchronous.

9. A dynamic speech enhancement system, Memory configured to store computer-executable instructions, A system comprising a processor configured to execute the computer-executable instructions in order to implement the method according to any one of claims 1 to 8.