Audio processing method and audio processing device

Through a combination of dynamic preset gain and short-time energy gain control, the problem of existing AGC technologies being difficult to balance volume control and sound quality is solved, achieving more efficient volume control and higher quality audio output.

CN114157254BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111465600.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-10
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Existing automatic gain control (AGC) technologies are difficult to balance the ability of audio volume control and processed audio quality, especially when dealing with different voice volume differences and distance from the device.

Method used

A combination of dynamic preset gain and short-time energy gain control is adopted to achieve more accurate volume control and higher sound quality by obtaining the energy and type of the current audio frame, calculating speech energy distribution data, and dynamically adjusting the gain.

Benefits of technology

It realizes better voice volume control capabilities, shortens gain convergence time, expands gain control range, and ensures stability and high quality of sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114157254B_ABST
    Figure CN114157254B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio processing method and an audio processing apparatus. The audio processing method includes the following steps: obtaining a current audio frame to be processed; determining the energy and type of the current audio frame, where the type includes one of a speech frame and a non-speech frame; obtaining speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, where the speech energy distribution data is used to statistically calculate the proportion of speech frames in different energy intervals; determining a first gain for the current audio frame according to the speech energy distribution data for the current audio frame; and applying the first gain to the current audio frame to obtain a first audio frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio technologies, and in particular, to an audio processing method and an audio processing apparatus for automatic gain control. Background Art

[0002] Automatic Gain Control (AGC) is a key technology in the field of audio processing and is widely used in fields such as real-time communication. Its basic purpose is to apply different magnitudes of gain to an audio signal according to the volume of the input audio signal, so that the volume of the output audio signal is stabilized within a certain range, and to avoid problems such as too loud or too soft sound caused by differences in the voice volumes of different speakers or the distance from the device. The AGC technology has relatively high requirements for both the volume control ability of the output audio and the audio quality after processing. Among them, the volume control ability is mainly reflected in the gain convergence time (i.e., the time required to calculate a reasonable volume audio for a section of audio with a stable volume) and the gain control range (i.e., the range of gain variation), and the audio quality is mainly reflected in the scores of objective evaluation indicators such as Perceptual evaluation of speech quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA). However, the existing AGC technologies are difficult to balance the ability to control the audio volume and the audio quality of the processed audio. Summary of the Invention

[0003] The present disclosure provides an audio processing method and an audio processing apparatus to at least solve the above-mentioned problems.

[0004] According to a first aspect of an embodiment of the present disclosure, an audio processing method is provided, which may include: obtaining a current audio frame to be processed; determining the energy and type of the current audio frame, where the type includes one of a speech frame and a non-speech frame; obtaining speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, where the speech energy distribution data is used to count the proportion of speech frames in different energy intervals; determining a first gain for the current audio frame according to the speech energy distribution data for the current audio frame; and applying the first gain to the current audio frame to obtain a first audio frame.

[0005] Optionally, obtaining voice energy distribution data for the current audio frame based on the energy and type of the current audio frame may include: when the energy of the current audio frame is less than a preset noise threshold or the current audio frame is a non-voice frame, using the voice energy distribution data of the previous audio frame of the current audio frame as the voice energy distribution data of the current audio frame; when the energy of the current audio frame is greater than or equal to the preset noise threshold and the current audio frame is a voice frame, updating the voice energy distribution data of the previous audio frame based on the energy of the current audio frame, where when the current audio frame is the first frame, updating the initial voice energy distribution data based on the energy of the current audio frame, and the proportion of voice frames occupied by each energy interval in the initial voice energy distribution data is evenly distributed.

[0006] Optionally, updating the voice energy distribution data of the previous audio frame based on the energy of the current audio frame may include: determining the energy interval to which the energy of the current audio frame belongs in the voice energy distribution data; increasing the proportion of voice frames in the energy interval corresponding to the determined energy interval in the voice energy distribution data of the previous audio frame; and decreasing the proportion of voice frames in the energy intervals in the voice energy distribution data of the previous audio frame that do not correspond to the determined energy interval.

[0007] Optionally, updating the voice energy distribution data of the previous audio frame based on the energy of the current audio frame may include: calculating the sum of the proportions of voice frames in each energy interval in the updated voice energy distribution data; determining a residual probability by comparing the sum of the proportions of voice frames with a preset value; and allocating the residual probability to each energy interval in the updated voice energy distribution data until the sum of the proportions of voice frames in each energy interval in the updated voice energy distribution data is the preset value.

[0008] Optionally, determining a first gain for the current audio frame according to the voice energy distribution data for the current audio frame may include: sequentially accumulating the proportions of voice frames in each energy interval starting from the first energy interval in the voice energy distribution data for the current audio frame until the accumulated sum is equal to or greater than a preset threshold; when the accumulated sum is equal to the preset threshold, using the upper limit of the energy interval where the accumulated sum reaches the preset threshold as the first energy limit; when the accumulated sum is greater than the preset threshold, using the lower limit of the energy interval where the accumulated sum exceeds the preset threshold as the first energy limit; and determining the first gain according to the target energy of the current audio frame and the first energy limit.

[0009] Optionally, determining the first gain according to the target energy of the current audio frame and the first energy limit may include: determining an initial first gain according to the target energy of the current audio frame and the first energy limit; determining the number of frames corresponding to the current audio frame according to the type of the current audio frame; adjusting the initial first gain by comparing the number of frames corresponding to the current audio frame with a preset number of frames, and using the adjusted initial first gain as the first gain.

[0010] Optionally, the audio processing method may further include: determining a second energy limit based on the first gain and the energy of the current audio frame; determining an initial second gain according to the target energy of the current audio frame and the second energy limit; obtaining a second gain vector based on the audio sampling points in the current audio frame and the initial second gain; applying the second gain to the first audio frame to obtain a second audio frame.

[0011] Optionally, obtaining a second gain vector based on the audio sampling points in the current audio frame and the initial second gain may include: calculating the gain for each audio sampling point in the current audio frame respectively based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain, to generate the second gain vector.

[0012] Optionally, applying the second gain to the first audio frame to obtain a second audio frame may include: applying each gain in the second gain vector to the corresponding audio sampling point of the first audio frame respectively to obtain a second audio frame; and performing clipping processing on the amplitude of the second audio frame.

[0013] According to a second aspect of the embodiments of the present disclosure, there is provided an audio processing apparatus, which may include: an acquisition module configured to acquire a current audio frame to be processed; a determination module configured to determine the energy and type of the current audio frame, where the type includes one of a voice frame and a non-voice frame; and obtain voice energy distribution data for the current audio frame based on the energy and type of the current audio frame, where the voice energy distribution data is used to count the proportion of voice frames in different energy intervals; a first gain module configured to determine a first gain for the current audio frame according to the voice energy distribution data for the current audio frame; and apply the first gain to the current audio frame to obtain a first audio frame.

[0014] Optionally, the determination module may be configured to: when the energy of the current audio frame is less than a preset noise threshold or the current audio frame is a non-speech frame, use the speech energy distribution data of the previous audio frame of the current audio frame as the speech energy distribution data of the current audio frame; when the energy of the current audio frame is greater than or equal to the preset noise threshold and the current audio frame is a speech frame, update the speech energy distribution data of the previous audio frame based on the energy of the current audio frame, wherein when the current audio frame is the first frame, update the initial speech energy distribution data based on the energy of the current audio frame, and the proportion of speech frames evenly distributed in each energy interval of the initial speech energy distribution data.

[0015] Optionally, the determination module may be configured to: determine the energy interval to which the energy of the current audio frame belongs in the speech energy distribution data; increase the proportion of speech frames in the energy interval corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame; decrease the proportion of speech frames in the energy intervals in the speech energy distribution data of the previous audio frame that do not correspond to the determined energy interval.

[0016] Optionally, the determination module may be configured to: calculate the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data; determine the residual probability by comparing the sum of the proportions of speech frames with a preset value; distribute the residual probability to each energy interval of the updated speech energy distribution data until the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data is the preset value.

[0017] Optionally, the first gain module may be configured to: sequentially accumulate the proportions of speech frames in each energy interval starting from the first energy interval of the speech energy distribution data of the current audio frame until the accumulated sum is equal to or greater than a preset threshold; when the accumulated sum is equal to the preset threshold, use the upper limit of the energy interval where the accumulated sum reaches the preset threshold as the first energy limit; when the accumulated sum is greater than the preset threshold, use the lower limit of the energy interval where the accumulated sum exceeds the preset threshold as the first energy limit; determine the first gain based on the target energy of the current audio frame and the first energy limit.

[0018] Optionally, the first gain module may be configured to: determine an initial first gain based on the target energy of the current audio frame and the first energy limit; determine the number of frames corresponding to the current audio frame according to the type of the current audio frame; adjust the initial first gain by comparing the number of frames corresponding to the current audio frame with a preset number of frames and use the adjusted initial first gain as the first gain.

[0019] Optionally, the audio processing device may further include a second gain module, configured to: determine a second energy threshold based on the first gain and the energy of the current audio frame; determine an initial second gain according to the target energy of the current audio frame and the second energy threshold; obtain a second gain vector based on the audio sampling points in the current audio frame and the initial second gain; and apply the second gain to the first audio frame to obtain a second audio frame.

[0020] Optionally, the second gain module may be configured to: calculate the gain for each audio sampling point in the current audio frame respectively based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain, so as to generate the second gain vector.

[0021] Optionally, the second gain module may be configured to: apply each gain in the second gain vector to the corresponding audio sampling point of the first audio frame respectively to obtain a second audio frame; and perform clipping processing on the amplitude of the second audio frame.

[0022] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, which may include: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the audio processing method as described above.

[0023] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing instructions, which when run by at least one processor, causes the at least one processor to execute the audio processing method as described above.

[0024] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, and the instructions in the computer program product are run by at least one processor in an electronic device to execute the audio processing method as described above.

[0025] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0026] It can better control the voice volume, and achieve a shorter gain convergence time and a larger gain control range. While ensuring better volume control, it can achieve relatively stable gain, and at the same time ensure higher-quality audio. In addition, by using the energy distribution data of the input voice, the dynamic preset gain is determined more accurately, so as to obtain higher-quality voice.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0028] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.

[0029] Figure 1 is a flowchart of an audio processing method according to an embodiment of the present disclosure;

[0030] Figure 2 is a schematic flow diagram of an audio processing method according to an embodiment of the present disclosure;

[0031] Figure 3 is a block diagram of an audio processing apparatus according to an embodiment of the present disclosure;

[0032] Figure 4 is a schematic structural diagram of an audio processing device according to an embodiment of the present disclosure;

[0033] Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure.

[0034] Throughout the drawings, it should be noted that the same reference numerals are used to denote the same or similar elements, features, and structures. Detailed Embodiments

[0035] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0036] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the embodiments of the present disclosure defined by the claims and their equivalents. Various specific details are included to assist in the understanding, but these details are only regarded as exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and structures are omitted for clarity and conciseness.

[0037] The terms and words used in the following description and claims are not limited to the written meanings, but are used by the inventors only to achieve a clear and consistent understanding of the present disclosure. Thus, those skilled in the art should clearly understand that the following descriptions of the various embodiments of the present disclosure are provided only for illustrative purposes and not for the purpose of limiting the present disclosure defined by the claims and their equivalents.

[0038] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0039] The existing AGC algorithm can preset a fixed high gain for the input audio and perform amplitude limiting protection on the audio. However, for the case of applying a high gain to a large-volume audio, this scheme will cause the amplitude limiting module to generate great distortion of the audio waveform, making it difficult to ensure high sound quality; or, it can refer to the audio energy size within a certain period of time and calculate the reasonable gain that needs to be applied to the audio currently. However, since the volume of the input audio changes both in the short term and the long term, this scheme usually has problems such as too large a gain change amplitude or being insensitive to the volume, resulting in a long time required for an audio segment to obtain a reasonable gain, and it is difficult to balance the audio volume control ability and the sound quality of the processed audio.

[0040] Aiming at the problems existing in the common schemes of the AGC algorithm, the present disclosure aims to propose an AGC method combining dynamic preset gain and short-time energy gain control, which can have a strong audio gain control ability while ensuring sound quality, and avoid problems such as slow audio gain convergence speed or large and small fluctuations in gain.

[0041] Hereinafter, according to various embodiments of the present disclosure, the methods, devices and systems of the present disclosure will be described in detail with reference to the drawings.

[0042] Figure 1 It is a flowchart of an audio processing method according to an embodiment of the present disclosure. The audio processing method according to the present disclosure can achieve automatic gain control with high sound quality.

[0043] The audio processing method according to the present disclosure can be executed by any electronic device with audio processing functions. The electronic device can be at least one of a smart phone, a tablet computer, a portable computer, a desktop computer, etc. The electronic device can be installed with a target application for performing automatic gain control on the input audio.

[0044] Refer to Figure 1, in step S101, obtain the current audio frame to be processed. For the input audio to be processed, frame division processing can be performed on the input audio, and then the operations described later are performed for each audio frame. Here, each audio frame may include a number of audio sampling points. For example, an audio frame may contain signal sampling points within a time period of 10 - 25 ms.

[0045] In step S102, determine the energy and type of the current audio frame. Here, the type may include one of a speech frame and a non - speech frame. For example, a voice activity detection algorithm can be used to detect whether the current audio frame is a speech frame or a non - speech frame (non - speech frames include cases such as noise or silence).

[0046] In step S103, obtain the speech energy distribution data for the current audio frame based on the energy and type of the current audio frame. Among them, the speech energy distribution data can be used to statistically analyze the proportion of speech frames in different energy intervals. The speech energy distribution data can be represented in the form of a histogram. For example, a speech energy histogram can represent the proportion of speech frames in different energy intervals.

[0047] Specifically, when the energy of the current audio frame is less than a preset noise threshold or the current audio frame is a non - speech frame, the speech energy distribution data of the previous audio frame of the current audio frame can be used as the speech energy distribution data of the current audio frame. When the energy of the current audio frame is greater than or equal to the preset noise threshold and the current audio frame is a speech frame, the speech energy distribution data of the previous audio frame of the current audio frame can be updated based on the energy of the current audio frame. Here, the preset noise threshold can be set differently according to the actual situation.

[0048] Here, when the current audio frame is the first frame of the input audio, the initial speech energy distribution data can be updated based on the energy of the current audio frame. The proportion of speech frames evenly distributed in each energy interval of the initial speech energy distribution data. For example, the initial speech energy distribution data can be divided into several energy intervals, the initial probabilities of each energy interval are set to be evenly distributed, and the sum of the initial probabilities of each energy interval is 1.

[0049] When updating the speech energy distribution data for the current audio frame, first determine the energy interval to which the energy of the current audio frame belongs in the speech energy distribution data, and then increase the proportion of speech frames in the energy interval corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame, and decrease the proportion of speech frames in the energy intervals that do not correspond to the determined energy interval in the speech energy distribution data of the previous audio frame.

[0050] After increasing or decreasing the proportion of speech frames in the corresponding energy range in the above manner, since it is necessary to ensure that the sum of all probabilities in the speech energy distribution data is 1, it is necessary to calculate the sum of the proportions of speech frames in each energy range in the updated speech energy distribution data, and determine the residual probability by comparing the sum of the proportions of speech frames with a preset value. Then, the residual probability is distributed to each energy range of the updated speech energy distribution data until the sum of the proportions of the speech frames in each energy range of the updated speech energy distribution data is equal to the preset value. For example, the preset value can be 1, but the present disclosure is not limited thereto.

[0051] In step S104, a first gain for the current audio frame is determined according to the speech energy distribution data for the current audio frame. In the present disclosure, the first gain may also be referred to as a primary gain. As an example, starting from the first energy range of the speech energy distribution data for the current audio frame, the proportions of speech frames in each energy range are sequentially accumulated until the sum is equal to or greater than a preset threshold. When the sum is equal to the preset threshold, the upper limit of the energy range where the sum reaches the preset threshold is used as the first energy boundary; when the sum is greater than the preset threshold, the lower limit of the energy range where the sum exceeds the preset threshold is used as the first energy boundary.

[0052] Next, an initial first gain is determined according to the target energy of the current audio frame and the first energy boundary. The number of frames corresponding to the current audio frame is determined according to the type of the current audio frame, and the initial first gain is adjusted by comparing the number of frames corresponding to the current audio frame with a preset number of frames, and the adjusted initial first gain is used as the first gain. Here, the preset threshold and the target energy can be set differently according to actual situations.

[0053] In addition, before adjusting the initial first gain to obtain the first gain, the initial first gain can be first controlled within a gain range to make the initial first gain meet the actual requirements.

[0054] In step S105, the first gain is applied to the current audio frame to obtain a first audio frame. After obtaining the first gain, the first gain can be applied to the original current audio frame, and thus a first audio signal can be obtained.

[0055] According to an embodiment of the present disclosure, after applying the first gain to the original current audio frame, a second gain can be applied to the audio frame to which the first gain has been applied, that is, in a manner of two-stage gain fusion, to obtain a final audio signal. In the present disclosure, the second gain may also be referred to as a secondary gain.

[0056] As an example, the second gain for the current audio frame can be determined based on the first gain and the energy of the current audio frame, and then the second gain is applied to the first audio frame to obtain a second audio frame.

[0057] For example, the second energy threshold may be determined based on the first gain and the energy of the current audio frame, and the initial second gain may be determined according to the target energy of the current audio frame and the second energy threshold. In addition, the second energy threshold may be smoothed first, and then the smoothed second energy threshold and the target energy of the current audio frame may be used to determine the initial second gain. The initial second gain is changed to a second gain vector according to the audio sampling points in the current audio frame. Here, the gain for each audio sampling point in the current audio frame may be calculated respectively based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain to generate the second gain vector. In addition, before generating the second gain vector, the initial second gain may be first controlled within a gain range so that the initial second gain meets the actual requirements.

[0058] Next, each gain in the second gain vector may be applied to the corresponding audio sampling of the first audio frame to obtain a second audio frame, and then the amplitude of the second audio frame may be clipped to output the final audio signal. The following will refer to Figure 2 describe in more detail the audio processing method according to an embodiment of the present disclosure.

[0059] Figure 2 is a schematic flowchart of an audio processing method according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the input audio may be frame-divided, and then the audio processing method shown in Figure 2 is applied to each audio frame of the input audio.

[0060] In the audio processing flow of the present disclosure, a voice energy calculation module, a voice activity detection module, a voice energy histogram statistics module, a dynamic prefabricated gain (primary gain) calculation module, a secondary gain calculation module, and a limiter module may be used to implement the audio processing method of the present disclosure.

[0061] For example, the voice energy calculation module is used to calculate the energy of the current input audio, the voice activity detection module is used to determine whether the audio at the current time is in the voice stage or the non-voice stage (such as noise or silence, etc.), the voice energy histogram statistics module is used to statistically analyze the voice energy distribution in the past period of time, the dynamic prefabricated gain (primary gain) calculation module calculates the dynamic prefabricated gain (i.e., the primary gain) that needs to be applied to the input audio currently according to the voice energy distribution data and the voice activity detection result, and applies the gain to the currently input audio, the secondary gain calculation module further adjusts the audio gain based on the audio with the primary gain applied, and the limiter module protects the audio from clipping distortion in some extreme cases.

[0062] Refer to Figure 2, the input audio is framed. The current audio frame (assumed to be the nth audio frame) is represented by x(n), where n ∈ N. The data contained in each audio frame can be the number of signal sampling points within 10 - 25 ms. That is, x(n) is a vector composed of a certain number of audio sampling points.

[0063] Input x(n) into the speech energy calculation module, and the energy of the nth audio frame can be calculated according to equation (1):

[0064]

[0065] where M is the number of audio sampling points contained in x(n). The unit of this energy is dBFS, and the value range of the calculation result can be (-∞, 0].

[0066] Input x(n) into the voice activity detection module to determine whether the current nth audio frame is in the speech stage or the non-speech stage (such as noise and silence). The two states can be represented by the following equation (2) respectively:

[0067]

[0068] where when vad(n) is 1, it means the current frame is a speech frame (speech active), and when vad(n) is 0, it means the current frame is a non-speech frame (speech inactive). Here, no restrictions are imposed on the VAD algorithm.

[0069] Next, input energyraw(n) and vad(n) into the speech energy histogram statistics module to statistically analyze the speech energy over a period of time (which can be set differently according to actual situations). The abscissa of the speech energy histogram is different energy intervals, and the width of each energy interval can be 1 dB. Its ordinate is the proportion of speech frames in each energy interval over a period of time. The speech energy distribution data of the current nth audio frame can be represented by HistogramEnergy(n). The specific statistical method is as follows:

[0070] First, HistogramEnergy(n) can be divided into several energy intervals. Here, taking the number of energy intervals as 100 and the width of the energy interval as 1 dB as an example, however, the present disclosure can be adjusted according to actual needs and is not limited thereto. HistogramEnergy(n) divided in this way can be represented by equation (3):

[0071] HistogramEnergy(n) = [e t (n), e 2 (n),......, e 100 (n)] (3) The energies corresponding to the subscripts of each energy interval increase in sequence, and the corresponding relationship is as follows:

[0072]

[0073] The initial probabilities (i.e., initial ratios) of each energy interval of the voice energy distribution data can be set to a uniform distribution. Taking 100 energy intervals as an example:

[0074]

[0075] When vad(n) = 0 or energyraw(n) < noisefloor, it means that the current frame is a non-voice frame or the audio energy is less than the noise threshold noisefloor (this value can be set to -50dBFS, but is not limited to this). The energy of the current nth frame of audio may not participate in the statistics of the voice energy distribution data. For example, the voice energy distribution data of the previous audio frame of the current audio frame can be used as the voice energy distribution data of the current audio frame, that is, HistogramEnergy(n) = HistogramEnergy(n - 1).

[0076] When vad(n) = 1 and energyraw(n) ≥ noisefloor, it means that the current frame is a voice frame, and HistogramEnergy(n) can be updated. For example, it can be determined which energy interval the energy of the current audio frame belongs to in the voice energy distribution data, increase the voice frame ratio of the energy interval corresponding to the determined energy interval in the voice energy distribution data of the previous audio frame of the current audio frame, and decrease the voice frame ratio of the energy intervals in the voice energy distribution data of the previous audio frame that do not correspond to the determined energy interval.

[0077] For example, first confirm the subscript of the energy interval of the energy energyraw(n) of the current audio frame in Equation (3), denoted as e x (n). The way to update the voice energy distribution data can be expressed as Equation (4):

[0078]

[0079] Among them, histSmooth is a smoothing factor used for the statistics of the voice energy distribution data. This smoothing factor can be set to 0.95, but can be adjusted according to requirements and specific situations. Methods such as selecting different parameters according to energyraw(n) in different energy intervals can also be adopted. The above examples are only exemplary, and the present disclosure is not limited thereto.

[0080] In addition, since it is necessary to ensure that the sum of all probabilities in HistogramEnergy(n) is 1, it is necessary to calculate the sum of the probabilities of the voice energy distribution data obtained in the above steps, and calculate the difference between this sum and 1 (i.e., the residual probability), and distribute this difference to the entire voice energy distribution data. Specifically, the residual probability residualPro(n) can be calculated according to the following equation (5):

[0081]

[0082] After obtaining the residual probability, the residual probability can be distributed according to the following equation (6):

[0083]

[0084] Repeat the above steps of distributing the residual probability until residualPro(n) = 0, at which time the update of HistogramEnergy(n) ends.

[0085] According to the embodiments of the present disclosure, the voice energy distribution data will be updated consistently starting from the first frame, and each audio frame of the input audio can update the corresponding voice energy distribution data according to the above equation (6).

[0086] Input the updated HistogramEnergy(n), vad(n), and x(n) into the dynamic preset gain (primary gain) module, and the gain gainPre(n) to be applied to the current audio frame can be calculated by integrating the energy distribution information and silence detection information over a period of time, so as to obtain the audio xGainPre(n) applied with the dynamic preset gain. The specific calculation method is as follows:

[0087] Judge the state of the current audio frame according to the vad(n) information, as shown in the following equation (7):

[0088]

[0089] Here, the state of the current audio frame may refer to the frame number corresponding to the current audio frame currently.

[0090] Next, starting from the first energy interval of the voice energy distribution data for the current audio frame, sequentially accumulate the voice frame ratios of each energy interval until the sum is equal to or greater than the preset threshold. When the sum is equal to the preset threshold, use the upper limit of the energy interval where the sum reaches the preset threshold as the first energy boundary; when the sum is greater than the preset threshold, use the lower limit of the energy interval where the sum exceeds the preset threshold as the first energy boundary.

[0091] As an example, according to the statistical distribution of HistogramEnergy(n), the boundary cnergyLevel (i.e., the first energy boundary) where the energy below the percentage threshold probThre (i.e., the preset threshold) is located is statistically determined. That is, all the energies in the speech energy distribution data statistically calculated over a period of time are below energyLevel. Here, probThre can be set to 95%, but it is not limited to this value.

[0092] The specific calculation process can refer to the following method: (i) First, set probSum = 0; (ii) Cumulatively add the energy probabilities in HistogramEnergy(n) in sequence, that is, probSum < probSum + e i (n), (i = 1, 2,......, 100); (iii) For each cumulative energy probability, judge the relationship between probSum and probThre. If probSum < probThre, continue to accumulate the next energy probability. If probSum = probThre, then energyLevel is the upper limit of the energy interval where e i (n) is located, and the calculation stops. If probSum > probThre, then energyLevel is the lower limit of the energy interval where e i (n) is located, and the calculation stops.

[0093] Calculate the initial dynamic preset gain (i.e., the initial first gain) gainPreRaw based on the energyLevel calculated above. It can be expressed as Equation (8):

[0094] gainPreRaw = EnergyTarget - energyLevel (8)

[0095] Where EnergyTarget is the energy that the current audio frame is expected to reach. Here, it can be set to -18dB, or it can be adjusted according to requirements. The gainPreRaw obtained at this time needs to be controlled within a certain gain range. According to the actual situation, the gain range is generally [-6dB, 12dB], or it can be adjusted according to requirements. gainPreRaw can be adjusted according to the following Equation (9):

[0096]

[0097] Calculate the dynamic preset gain gainPre(n) (i.e., the first gain) that needs to be applied to the current audio based on the initial dynamic preset gain gainPreRaw calculated above and silenceState(n).

[0098] The initial first gain can be adjusted by comparing the frame number corresponding to the current audio frame with a preset frame number, and the adjusted initial first gain is used as the first gain. The specific method is as follows:

[0099] If silenceState(n) ≥ silThre (where silThre is the number of audio frames corresponding to a period of time, and this period of time can be from 1 second to 2 seconds), then gainPre(n) = gainPreRaw;

[0100] If silenceState(n) < silThre and gainPreRaw ≥ gainPre(n - 1), then gainPre(n) = gainPreRaw×(1 - sAtt) + gainPre(n - 1)×sAtt, where sAtt is the follow-up smoothing factor, generally set to 0.9999, and can also be set according to the actual situation;

[0101] If silenceState(n) < silThre and gainPreRaw < gainPre(n - 1), then gainPre(n) = gainPreRaw×(1 - sRel) + gainPre(n - 1)×sRel, where sRel is the release smoothing factor, generally set to 0.99, and can also be set according to the actual situation.

[0102] After obtaining the first gain gainPre(n), apply it to the input original audio x(n), as shown in the following equation (10):

[0103]

[0104] The audio signal with the dynamic preset gain (primary gain) applied can be obtained.

[0105] Input gainPre(n), xGainPre(n), energyraw(n) and vad(n) into the secondary gain calculation module, and the second gain gainPost(n) that needs to be applied to the current audio frame can be further calculated. Considering that the gain corresponding to each sampling point in the current audio frame is different, the above gain can be regarded as a gain vector.

[0106] Specifically, the audio energy after applying the dynamic preset gain (i.e., the second energy limit) can be obtained according to gainPre(n) and energyraw(n), as shown in the following equation (11):

[0107] energyGainPreRaw = gainPre(n) + energyraw(n) (11)

[0108] Smooth the audio energy as shown in the following equation (12):

[0109] energyGainPreSmooth(n) = energyGainPreSmooth(n - 1) × smoothEnergy + energyGainPreRaw × (1 - smoothEnergy) (12)

[0110] Where smoothEnergy represents the smoothing factor, energyGainPreSmooth(n) represents the smoothed audio energy of the current audio frame, and energyGainPreSmooth(n - 1) is the smoothed audio energy of the previous audio frame. In the case where the current audio frame is the first frame, energyGainPreSmooth(n - 1) can be set to zero.

[0111] According to energyGainPreSmooth(n) and EnergyTarget, the expected value of the secondary gain of the current audio frame (i.e., the initial second gain) can be calculated as shown in the following equation (13):

[0112] gainPostRaw = EnergyTarget - energtGainPreSmooth(n) (13)

[0113] Similar to the calculation of gainPreRaw, gainPostRaw also needs to limit the gain range:

[0114]

[0115] According to the gain after the above gain control processing, the gain vector gainPost(n) of the current audio frame can be obtained. This vector has the same dimension as xGainPre(n), that is, if they both contain M elements, the above gain vector and xGainPre(n) can be specifically expressed in the following form:

[0116] gainPost(n) = [gainPost 1 (n), gainPost 2 (n),......, gainPost M (n)] T

[0117] xGainPre(n) = [xGainPre 1 (n), xGainPre 2 (n),......, xGainPreM (n)] T

[0118] Each element in the current secondary gain vector can be calculated according to the following equation (14):

[0119]

[0120] where i represents the i-th element in the current secondary gain vector, and gainPost M (n - 1) represents the gain at the M-th sample point of the previous audio frame (i.e., the last sample point of the previous audio frame).

[0121] The corresponding elements of the secondary gain vector gainPost(n) obtained above are multiplied by the corresponding elements of xGainPre(n) (note that the gain unit needs to be converted from dBFS to linear gain), and the audio signal xGainPre(n) with secondary gain applied can be obtained, as shown in the following equation (15):

[0122]

[0123] The output audio xGainPost(n) obtained above is input into the limiter module to ensure that the audio does not suffer from clipping distortion, as shown in the following equation (16):

[0124] y(n) = Limiter[xGainPost(n)] (16)

[0125] where Limiter[*] represents amplitude clipping protection for the input signal, and y(n) is the final output audio signal of one frame after AGC processing.

[0126] Figure 3 is a block diagram of an audio processing apparatus according to an embodiment of the present disclosure.

[0127] Referring to Figure 3 , the audio processing apparatus 300 may include an acquisition module 301, a determination module 302, a first gain module 303, and a second gain module 304. Each module in the audio processing apparatus 300 may be implemented by one or more modules, and the names of the corresponding modules may vary according to the type of the module. In various embodiments, some modules in the audio processing apparatus 300 may be omitted, or additional modules may also be included. In addition, the modules / components according to various embodiments of the present disclosure may be combined to form a single entity, and thus may equivalently perform the functions of the corresponding modules / components before combination.

[0128] The acquisition module 301 may acquire the current audio frame to be processed.

[0129] The determination module 302 can determine the energy and type of the current audio frame, and the type includes one of a speech frame and a non-speech frame.

[0130] The determination module 302 can obtain speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, where the speech energy distribution data can be used to count the proportion of speech frames in different energy intervals.

[0131] The first gain module 303 can determine a first gain for the current audio frame according to the speech energy distribution data for the current audio frame, and apply the first gain to the current audio frame to obtain a first audio frame.

[0132] The second gain module 304 can determine a second gain for the current audio frame based on the first gain and the energy of the current audio frame, and apply the second gain to the first audio frame to obtain a second audio frame.

[0133] When the energy of the current audio frame is less than a preset noise threshold or the current audio frame is a non-speech frame, the determination module 302 can use the speech energy distribution data of the previous audio frame of the current audio frame as the speech energy distribution data of the current audio frame. When the energy of the current audio frame is greater than or equal to the preset noise threshold and the current audio frame is a speech frame, the determination module 302 can update the speech energy distribution data of the previous audio frame based on the energy of the current audio frame. Among them, when the current audio frame is the first frame, the determination module 302 can update the initial speech energy distribution data based on the energy of the current audio frame, and the proportion of speech frames in each energy interval of the initial speech energy distribution data is evenly distributed.

[0134] The determination module 302 can determine the energy interval to which the energy of the current audio frame belongs in the speech energy distribution data, increase the proportion of speech frames in the energy interval corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame, and decrease the proportion of speech frames in the energy interval not corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame.

[0135] The determination module 302 can calculate the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data, determine the residual probability by comparing the sum of the speech frame proportions with a preset value, and distribute the residual probability to each energy interval of the updated speech energy distribution data until the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data is the preset value. For example, the preset value can be set to 1.

[0136] The first gain module 303 may sequentially accumulate the speech frame ratios of each energy interval starting from the first energy interval of the speech energy distribution data for the current audio frame until the sum of the accumulations is equal to or greater than a preset threshold. When the sum of the accumulations is equal to the preset threshold, the first gain module 303 may use the upper limit of the energy interval where the sum of the accumulations reaches the preset threshold as the first energy boundary. When the sum of the accumulations is greater than the preset threshold, the first gain module 303 may use the lower limit of the energy interval where the sum of the accumulations exceeds the preset threshold as the first energy boundary. Then, the first gain module 303 may determine the first gain based on the target energy of the current audio frame and the first energy boundary.

[0137] The first gain module 303 may determine an initial first gain based on the target energy of the current audio frame and the first energy boundary, determine the number of frames corresponding to the current audio frame according to the type of the current audio frame, adjust the initial first gain by comparing the number of frames corresponding to the current audio frame with a preset number of frames, and use the adjusted initial first gain as the first gain.

[0138] The second gain module 304 may determine a second energy boundary based on the first gain and the energy of the current audio frame, determine an initial second gain based on the target energy of the current audio frame and the second energy boundary, and change the initial second gain into a second gain vector according to the audio sampling points in the current audio frame.

[0139] The second gain module 304 may calculate the gain for each audio sampling point in the current audio frame based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain respectively to generate a second gain vector.

[0140] The second gain module 304 may apply each gain in the second gain vector to the corresponding audio sampling point of the first audio frame to obtain a second audio frame, and perform amplitude clipping processing on the second audio frame.

[0141] The above has Figure 1 and Figure 2 detailed the automatic gain control process according to the embodiments of the present disclosure, and will not be described here again.

[0142] Figure 4 is a schematic structural diagram of an audio processing device in the hardware operating environment of the embodiments of the present disclosure.

[0143] As Figure 4As shown, the audio processing device 400 may include: a processing component 401, a communication bus 402, a network interface 403, an input / output interface 404, a memory 405, and a power supply component 404. Among them, the communication bus 402 is used to enable connection and communication between these components. The input / output interface 404 may include a video display (such as a liquid crystal display), a microphone, a speaker, and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). Optionally, the input / output interface 404 may further include a standard wired interface and a wireless interface. The network interface 403 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 405 may be a high-speed random access memory or a stable non-volatile memory. The memory 405 may optionally also be a storage device independent of the aforementioned processing component 401.

[0144] Those skilled in the art can understand that Figure 4 the structure shown in does not constitute a limitation on the audio processing device 400, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0145] As Figure 4 shown, the memory 405, as a storage medium, may include an operating system (such as a MAC operating system), a data storage module, a network communication module, a user interface module, an audio processing program, and a database.

[0146] In Figure 4 the audio processing device 400 shown, the network interface 403 is mainly used for data communication with external electronic devices / terminals; the input / output interface 404 is mainly used for data interaction with users; the processing component 401 and the memory 405 in the audio processing device 400 may be disposed in the audio processing device 400. The audio processing device 400 calls the audio processing program, materials, and various APIs provided by the operating system stored in the memory 405 through the processing component 401 to execute the audio processing method provided by the embodiments of the present disclosure.

[0147] The processing component 401 may include at least one processor. A set of computer-executable instructions is stored in the memory 405. When the set of computer-executable instructions is executed by at least one processor, the audio processing method according to the embodiments of the present disclosure is executed. However, the above examples are only exemplary, and the present disclosure is not limited thereto.

[0148] The processing component 401 can obtain the current audio frame to be processed, determine the energy and type of the current audio frame, obtain the speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, determine the first gain for the current audio frame according to the speech energy distribution data for the current audio frame, apply the first gain to the current audio frame to obtain the first audio frame, and then can determine the second gain for the current audio frame based on the first gain and the energy of the current audio frame, and apply the second gain to the first audio frame to obtain the second audio frame.

[0149] The processing component 401 can control the components included in the audio processing device 400 by executing a program.

[0150] The audio processing device 400 can receive or output video and / or audio via the input / output interface 404. For example, the audio processing device 400 can output the audio signal after applying the gain via the input / output interface 404.

[0151] As an example, the audio processing device 400 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the audio processing device 400 does not have to be a single electronic device, and can also be an aggregate of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The audio processing device 400 can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0152] In the audio processing device 400, the processing component 401 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processing component 401 can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0153] The processing component 401 can run the instructions or code stored in the memory. Among them, the memory 405 can also store data. The instructions and data can also be sent and received via the network interface 403 through the network, where the network interface 403 can adopt any known transmission protocol.

[0154] The memory 405 may be integrated with the processing component 401. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 405 may include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory and the processing component 401 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processing component 401 can read the data stored in the memory 405.

[0155] According to an embodiment of the present disclosure, an electronic device may be provided. Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device 500 may include at least one memory 502 and at least one processor 501. The at least one memory 502 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor 501, an audio processing method according to an embodiment of the present disclosure is executed.

[0156] The processor 501 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 501 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0157] The memory 502, as a storage medium, may include an operating system (e.g., MAC operating system), a data storage module, a network communication module, a user interface module, an audio processing program, and a database.

[0158] The memory 502 may be integrated with the processor 501. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 502 may include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 502 and the processor 501 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 501 can read the files stored in the memory 502.

[0159] In addition, the electronic device 500 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 500 may be connected to each other via a bus and / or a network.

[0160] Those skilled in the art will understand that Figure 5 the structure shown in does not constitute a limitation on, and may include more or fewer components than shown, or combine certain components, or have a different component arrangement.

[0161] According to an embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are run by at least one processor, the at least one processor is caused to execute the audio processing method according to the present disclosure. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid state drive (SSD), card memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0162] In an embodiment according to the present disclosure, a computer program product may also be provided, and the instructions in the computer program product may be executed by a processor of a computer device to complete the above audio processing method.

[0163] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0164] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An audio processing method, comprising: Obtaining a current audio frame to be processed; Determining the energy and type of the current audio frame, where the type includes one of a speech frame and a non-speech frame; Obtaining speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, where the speech energy distribution data is used to statistically analyze the proportion of speech frames in different energy intervals; Determining a first gain for the current audio frame according to the speech energy distribution data for the current audio frame; Applying the first gain to the current audio frame to obtain a first audio frame, where obtaining the speech energy distribution data for the current audio frame based on the energy and type of the current audio frame includes: When the energy of the current audio frame is greater than or equal to a preset noise threshold and the current audio frame is a speech frame, updating the speech energy distribution data of the previous audio frame based on the energy of the current audio frame, where updating the speech energy distribution data of the previous audio frame based on the energy of the current audio frame includes: Calculating the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data; Determining a residual probability by comparing the sum of the speech frame proportions with a preset value; Allocating the residual probability to each energy interval of the updated speech energy distribution data until the sum of the proportions of speech frames in each energy interval of the updated speech energy distribution data is the preset value.

2. The audio processing method according to claim 1, characterized in that obtaining the speech energy distribution data for the current audio frame based on the energy and type of the current audio frame further includes: When the energy of the current audio frame is less than the preset noise threshold or the current audio frame is a non-speech frame, using the speech energy distribution data of the previous audio frame of the current audio frame as the speech energy distribution data of the current audio frame; where when the current audio frame is the first frame, updating the initial speech energy distribution data based on the energy of the current audio frame, and the proportions of speech frames in each energy interval of the initial speech energy distribution data are evenly distributed.

3. The audio processing method according to claim 1, characterized in that updating the speech energy distribution data of the previous audio frame based on the energy of the current audio frame further includes: Determining the energy interval in the speech energy distribution data to which the energy of the current audio frame belongs; Increasing the proportion of speech frames in the energy interval corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame; Decreasing the proportion of speech frames in the energy intervals in the speech energy distribution data of the previous audio frame that do not correspond to the determined energy interval.

4. The audio processing method according to claim 1, characterized in that determining the first gain for the current audio frame according to the speech energy distribution data for the current audio frame includes: Sequentially accumulating the proportions of speech frames in each energy interval starting from the first energy interval of the speech energy distribution data for the current audio frame until the sum is equal to or greater than a preset threshold; When the sum is equal to the preset threshold, the upper limit of the energy range accumulated to satisfy that the sum is equal to the preset threshold is used as the first energy boundary; When the sum is greater than the preset threshold, the lower limit of the energy range accumulated to satisfy that the sum is greater than the preset threshold is used as the first energy boundary; Determine the first gain according to the target energy of the current audio frame and the first energy boundary.

5. The audio processing method according to claim 4, wherein, Determining the first gain according to the target energy of the current audio frame and the first energy boundary includes: Determine an initial first gain according to the target energy of the current audio frame and the first energy boundary; Determine the number of frames corresponding to the current audio frame according to the type of the current audio frame; Adjust the initial first gain by comparing the number of frames corresponding to the current audio frame with a preset number of frames, and use the adjusted initial first gain as the first gain.

6. The audio processing method according to claim 1, wherein, further comprising: Determine a second energy boundary based on the first gain and the energy of the current audio frame; Determine an initial second gain according to the target energy of the current audio frame and the second energy boundary; Obtain a second gain vector based on the audio sampling points in the current audio frame and the initial second gain; Apply the second gain to the first audio frame to obtain a second audio frame.

7. The audio processing method according to claim 6, wherein, Obtaining a second gain vector based on the audio sampling points in the current audio frame and the initial second gain includes: Calculate the gain for each audio sampling point in the current audio frame respectively based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain to generate the second gain vector.

8. The audio processing method according to claim 6, wherein, Applying the second gain to the first audio frame to obtain a second audio frame includes: Apply each gain in the second gain vector to the corresponding audio sampling point of the first audio frame respectively to obtain a second audio frame; and Perform clipping processing on the amplitude of the second audio frame.

9. An audio processing device, comprising: An acquisition module configured to acquire a current audio frame to be processed; A determination module configured to determine the energy and type of the current audio frame, the type including one of a speech frame and a non-speech frame; and obtain speech energy distribution data for the current audio frame based on the energy and type of the current audio frame, wherein the speech energy distribution data is used to statistically calculate the proportion of speech frames in different energy ranges; A first gain module configured to determine a first gain for the current audio frame according to the speech energy distribution data for the current audio frame; and apply the first gain to the current audio frame to obtain a first audio frame, The determination module is configured to: when the energy of the current audio frame is greater than or equal to a preset noise threshold and the current audio frame is a speech frame, update the speech energy distribution data of the previous audio frame based on the energy of the current audio frame, The determination module is configured to: calculate the sum of the speech frame ratios of each energy interval in the updated speech energy distribution data; determine the residual probability by comparing the sum of the speech frame ratios with a preset value; allocate the residual probability to each energy interval of the updated speech energy distribution data until the sum of the proportions of the speech frames in each energy interval of the updated speech energy distribution data is the preset value.

10. The audio processing device according to claim 9, wherein, The determination module is configured to: when the energy of the current audio frame is less than the preset noise threshold or the current audio frame is a non-speech frame, use the speech energy distribution data of the previous audio frame of the current audio frame as the speech energy distribution data of the current audio frame; wherein, when the current audio frame is the first frame, update the initial speech energy distribution data based on the energy of the current audio frame, and the proportions of the speech frames evenly distributed in each energy interval of the initial speech energy distribution data.

11. The audio processing device according to claim 9, wherein, The determination module is configured to: determine the energy interval to which the energy of the current audio frame belongs in the speech energy distribution data; increase the speech frame ratio of the energy interval corresponding to the determined energy interval in the speech energy distribution data of the previous audio frame; decrease the speech frame ratio of the energy interval in the speech energy distribution data of the previous audio frame that does not correspond to the determined energy interval.

12. The audio processing device according to claim 12, wherein, The first gain module is configured to: accumulate the speech frame ratios of each energy interval in turn starting from the first energy interval of the speech energy distribution data for the current audio frame until the sum is equal to or greater than a preset threshold; when the sum is equal to the preset threshold, use the upper limit of the energy interval where the sum reaches the preset threshold as the first energy limit; when the sum is greater than the preset threshold, use the lower limit of the energy interval where the sum exceeds the preset threshold as the first energy limit; determine the first gain according to the target energy of the current audio frame and the first energy limit.

13. The audio processing device according to claim 12, wherein, The first gain module is configured to: determine an initial first gain according to the target energy of the current audio frame and the first energy limit; determine the number of frames corresponding to the current audio frame according to the type of the current audio frame; adjust the initial first gain by comparing the number of frames corresponding to the current audio frame with a preset number of frames and use the adjusted initial first gain as the first gain.

14. The audio processing device according to claim 9, wherein, further includes a second gain module, configured to: Determine a second energy threshold based on the first gain and the energy of the current audio frame; Determine an initial second gain according to the target energy of the current audio frame and the second energy threshold; Obtain a second gain vector based on the audio sampling points in the current audio frame and the initial second gain; Apply the second gain to the first audio frame to obtain a second audio frame.

15. The audio processing apparatus according to claim 14, wherein, the second gain module is configured to: Calculate the gain for each audio sampling point in the current audio frame respectively based on the gain of the last audio sampling point in the previous audio frame of the current audio frame and the initial second gain, so as to generate the second gain vector.

16. The audio processing apparatus according to claim 14, wherein, the second gain module is configured to: Apply each gain in the second gain vector to the corresponding audio sampling point of the first audio frame respectively to obtain a second audio frame; and Perform clipping processing on the amplitude of the second audio frame.

17. An electronic device, wherein, comprising: At least one processor; At least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the audio processing method according to any one of claims 1 to 8.

18. A computer-readable storage medium storing instructions, wherein, when the instructions are run by at least one processor, the at least one processor is caused to execute the audio processing method according to any one of claims 1 to 8.

19. A computer program product, wherein the instructions in the computer program product are run by at least one processor in an electronic device to execute the audio processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice signal processing method and device

    CN103544961A

  • Automatic gain control device and method

    CN104200810A