Audio frame processing method and device, electronic equipment and storage medium

By smoothing the amplitude gain of audio frames, the problem of abrupt gain changes between adjacent frames is solved, improving the coherence and auditory quality of the audio.

CN115910094BActive Publication Date: 2025-10-21BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211168036.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-10-21
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

In existing technologies, gain adjustment of audio frames causes abrupt changes in gain values ​​at the junction of adjacent frames, resulting in audio discontinuity and affecting subjective listening quality.

Method used

By acquiring the amplitude gain of the first audio frame and smoothing it based on the amplitude gains of the adjacent second and third audio frames, the amplitude gain of the first audio frame is adjusted to change slowly, thus solving the problem of inconsistent gain.

Benefits of technology

It achieves a smooth transition between adjacent audio frames, improves the subjective listening quality of audio, and eliminates the phenomenon of gain discontinuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910094B_ABST
    Figure CN115910094B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an audio frame processing method and device, electronic equipment and a storage medium. The method comprises: obtaining a first audio frame, determining a first amplitude gain corresponding to the first audio frame; performing smoothing processing on the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and / or performing smoothing processing on the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to the first audio frame and located before the first audio frame, and the third audio frame is an audio frame adjacent to the first audio frame and located after the first audio frame; and adjusting the amplitude of the first audio frame based on the first amplitude gain after the smoothing processing to obtain a target audio frame. Through the technical scheme of the embodiments of the present application, the problem of audio discontinuity caused by the gain adjustment of the audio frame is solved, and the subjective auditory quality of the audio is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to audio processing technology, and in particular to an audio frame processing method, device, electronic device and storage medium. Background Art

[0002] The Automatic Gain Control (AGC) algorithm expands or compresses an audio signal by a certain gain value to increase or decrease the audio volume. In the prior art, a gain value is determined for each audio frame and used to adjust the volume of the corresponding audio frame.

[0003] In the process of implementing the present invention, the inventors discovered that the prior art has at least the following problems:

[0004] The above-mentioned method of the prior art easily leads to a sudden change in gain value at the junction of adjacent audio frames, thereby causing a large instantaneous change in the audio, resulting in audio discontinuity and affecting the subjective auditory quality of the audio. Summary of the Invention

[0005] Embodiments of the present invention provide an audio frame processing method, device, electronic device, and storage medium to solve the problem of audio incoherence caused by audio frame gain adjustment and improve the subjective auditory quality of the audio.

[0006] In a first aspect, an embodiment of the present invention provides an audio frame processing method, the method comprising:

[0007] Acquire a first audio frame, and determine a first amplitude gain corresponding to the first audio frame;

[0008] performing smoothing on the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and / or performing smoothing on the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to and preceding the first audio frame, and the third audio frame is an audio frame adjacent to and following the first audio frame;

[0009] The amplitude of the first audio frame is adjusted based on the first amplitude gain after smoothing to obtain a target audio frame.

[0010] In a second aspect, an embodiment of the present invention further provides an audio frame processing device, the device comprising:

[0011] a first amplitude gain determining module, configured to obtain a first audio frame and determine a first amplitude gain corresponding to the first audio frame;

[0012] a first amplitude gain smoothing module, configured to smooth the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and / or smooth the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to and preceding the first audio frame, and the third audio frame is an audio frame adjacent to and following the first audio frame;

[0013] The first audio frame gain module is configured to adjust the amplitude of the first audio frame based on the first amplitude gain after smoothing to obtain a target audio frame.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:

[0015] one or more processors;

[0016] a storage device for storing one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio frame processing method as described in any one of the embodiments of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio frame processing method as described in any one of the embodiments of the present invention.

[0019] The technical solution of the embodiment of the present invention obtains a first audio frame, determines a first amplitude gain corresponding to the first audio frame, and determines amplitude gain values ​​corresponding to each sampling point in the first audio frame. The first amplitude gain is smoothed according to a second amplitude gain corresponding to the second audio frame, and / or the first amplitude gain is smoothed according to a third amplitude gain corresponding to the third audio frame to smooth the first amplitude gain so that the amplitude gain value changes slowly. The amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain a target audio frame. This solves the problem of sudden changes in amplitude gain values ​​at the junction of adjacent audio frames, which leads to audio discontinuity and poor audio quality. The gain discontinuity is eliminated, thereby improving the subjective auditory quality of the audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings introduced here only illustrate some of the embodiments to be described by the present invention, and are not exhaustive. A person skilled in the art can derive other drawings based on these drawings without inventive effort.

[0021] Figure 1 A flowchart of an audio frame processing method provided by an embodiment of the present invention;

[0022] Figure 2 A schematic diagram of gain smoothing provided by an embodiment of the present invention when the first amplitude gain is greater than the second amplitude gain;

[0023] Figure 3 A schematic diagram of gain smoothing provided by an embodiment of the present invention when the first amplitude gain is greater than the third amplitude gain;

[0024] Figure 4 A schematic flow chart of another audio frame processing method provided by an embodiment of the present invention;

[0025] Figure 5 A schematic flow chart of another audio frame processing method provided by an embodiment of the present invention;

[0026] Figure 6 A schematic structural diagram of an audio frame processing device provided by an embodiment of the present invention;

[0027] Figure 7 The present invention provides a schematic structural diagram of an electronic device. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0029] Figure 1 A flow chart of an audio frame processing method provided in an embodiment of the present invention is provided. This embodiment is applicable to the case of performing gain adjustment on audio frames. The method can be executed by an audio frame processing device. The device can be implemented in the form of software and / or hardware. The hardware can be an electronic device. Optionally, the electronic device can be a mobile terminal, a PC, a server, etc.

[0030] like Figure 1 As shown, the method of this embodiment may specifically include:

[0031] S110: Acquire a first audio frame, and determine a first amplitude gain corresponding to the first audio frame.

[0032] The first audio frame may be an audio frame currently to be processed, may be an original audio frame, or may be an audio frame that has undergone other audio processing other than gain processing. The first amplitude gain may be an amplitude gain value corresponding to the first audio frame obtained through data processing or other means. It is understood that the first amplitude gain is an amplitude gain value that corresponds to the amplitude of each sampling point in the first audio frame.

[0033] Specifically, during the process of processing the audio frame, the first audio frame can be obtained so as to subsequently process the first audio frame. Furthermore, the amplitude gain value corresponding to the first audio frame can be obtained by processing the first audio frame, and the amplitude gain value can be used as the first amplitude gain.

[0034] In a specific implementation, in order to improve the accuracy of the first amplitude gain and process the audio frame more accurately, the first amplitude gain corresponding to the first audio frame may be determined by the following steps:

[0035] Step 1: Determine a first average amplitude and a first maximum amplitude of the first audio frame according to the amplitude of each sampling point in the first audio frame.

[0036] The first average amplitude may be an average of the amplitudes of the sampling points in the first audio frame, and the first maximum amplitude may be a maximum value among the amplitudes of the sampling points in the first audio frame.

[0037] Specifically, after obtaining the first audio frame, the amplitudes of the sampling points in the first audio frame may be determined, and the average of the amplitudes of the sampling points may be determined as the first average amplitude, and the maximum of the amplitudes of the sampling points may be determined as the first maximum amplitude.

[0038] It should be noted that the first average amplitude and the first maximum amplitude can be determined by the following formula:

[0039]

[0040]

[0041] Among them, x k (n) represents the amplitude of the nth sampling point in the first audio frame, N represents the number of sampling points in the first audio frame, represents the first average amplitude of the first audio frame, Indicates the first maximum amplitude of the first audio frame.

[0042] Step 2: Determine a first average amplitude gain according to the first average amplitude and a preset average amplitude.

[0043] The preset average amplitude may be a received, manually set average amplitude. It should be noted that the preset average amplitude may correspond to the first audio frame, i.e., the preset average amplitudes corresponding to different first audio frames may be the same or different. The preset average amplitude may also correspond to the entire audio, i.e., the preset average amplitudes corresponding to each first audio frame are the same. The first average amplitude gain may be a gain obtained by adjusting the first average amplitude.

[0044] Specifically, data processing is performed on the first average amplitude and the preset average amplitude to obtain a first average amplitude gain, wherein the data processing may be performed in a manner such as calculating a ratio.

[0045] Exemplarily, the first average amplitude gain may be determined according to the first average amplitude and the preset average amplitude using the following formula:

[0046]

[0047] in, represents the first average amplitude of the first audio frame, Indicates the preset average amplitude, Represents the first average amplitude gain.

[0048] Step 3: Determine a first maximum amplitude gain according to the first maximum amplitude and the preset maximum amplitude.

[0049] The preset maximum amplitude may be a received manually set maximum amplitude. It should be noted that the preset maximum amplitude may correspond to the first audio frame, i.e., the preset maximum amplitudes corresponding to different first audio frames may be the same or different. The preset maximum amplitude may also correspond to the entire audio, i.e., the preset maximum amplitudes corresponding to all first audio frames are the same. The first maximum amplitude gain may be a gain obtained by adjusting the first maximum amplitude.

[0050] Specifically, data processing is performed on the first maximum amplitude and the preset maximum amplitude to obtain the first maximum amplitude gain, wherein the data processing may be performed in a manner such as calculating a ratio.

[0051] Exemplarily, the first maximum amplitude gain may be determined according to the first maximum amplitude and the preset maximum amplitude using the following formula:

[0052]

[0053] in, represents the first maximum amplitude of the first audio frame, Indicates the preset maximum amplitude. Indicates the first maximum amplitude gain.

[0054] Step 4: Determine a first amplitude gain corresponding to the first audio frame according to the first average amplitude gain and the first maximum amplitude gain.

[0055] Specifically, data processing can be performed based on the first average amplitude gain and the first maximum amplitude gain to analyze the first average amplitude gain and the first maximum amplitude gain, for example, by using a preset algorithm or a preset model, and the analysis result is used as the first amplitude gain corresponding to the first audio frame.

[0056] For example, the minimum value between the first average amplitude gain and the first maximum amplitude gain may be taken as the first amplitude gain corresponding to the first audio frame, that is, in, represents the first average amplitude gain, Indicates the first maximum amplitude gain, G k Indicates the first amplitude gain corresponding to the first audio frame.

[0057] It should be noted that the first amplitude gain can also be determined according to other methods, for example: by obtaining a preset audio gain, or by determining the first amplitude gain according to the current volume through a pre-established correspondence between volume and audio gain, or by determining the first amplitude gain through a pre-established deep learning gain model, etc.

[0058] S120: Smoothing the first amplitude gain according to the second amplitude gain corresponding to the second audio frame; and / or, smoothing the first amplitude gain according to the third amplitude gain corresponding to the third audio frame.

[0059] The second audio frame is an audio frame adjacent to and preceding the first audio frame, and the third audio frame is an audio frame adjacent to and following the first audio frame. The second amplitude gain may be an amplitude gain value corresponding to the second audio frame obtained by processing the second audio frame, which may be understood as the amplitude gain value corresponding to the second audio frame cached after processing the second audio frame. The third amplitude gain may be an amplitude gain value corresponding to the third audio frame obtained by processing the third audio frame, which may be understood as the amplitude gain value corresponding to the third audio frame cached after processing the third audio frame. Smoothing may be data processing performed according to a preset smoothing algorithm or a preset curve form.

[0060] Specifically, the first amplitude gain can be smoothed based on the second amplitude gain corresponding to the second audio frame to obtain amplitude gain values ​​corresponding to each sampling point in the first audio frame. In this case, the amplitude gain values ​​corresponding to each sampling point are smooth, thereby ensuring a smooth transition between the amplitude gain values ​​of the first audio frame and the amplitude gain values ​​of the second audio frame. Alternatively, the first amplitude gain can be smoothed based on the third amplitude gain corresponding to the third audio frame to obtain amplitude gain values ​​corresponding to each sampling point in the first audio frame. In this case, the amplitude gain values ​​corresponding to each sampling point are smooth, thereby ensuring a smooth transition between the amplitude gain values ​​of the first audio frame and the amplitude gain values ​​of the third audio frame. Amplitude gain values ​​corresponding to some sampling points in the first audio may also be smoothed according to the second amplitude gain, and amplitude gain values ​​corresponding to some sampling points in the first audio may also be smoothed according to the third amplitude gain to obtain amplitude gain values ​​corresponding to each sampling point in the first audio frame. In this case, the amplitude gain values ​​corresponding to each sampling point are smooth, and therefore, the amplitude gain value of the first audio frame can be smoothly transitioned with the amplitude gain values ​​of the second audio frame and the amplitude gain values ​​of the third audio frame, respectively.

[0061] Exemplarily, the amplitude gain values ​​of the sampling points in the first half of the first audio frame are processed according to the second amplitude gain, and the amplitude gain values ​​of the sampling points in the second half of the first audio frame are processed according to the third amplitude gain. In this case, the sum of the number of sampling points in the first half and the number of sampling points in the second half is less than or equal to the total number of sampling points in the first audio frame.

[0062] Based on the above example, in order to effectively limit the range of sampling points in the first audio frame for smoothing the gain with the second audio frame, and accurately determine the amplitude gain value of the sampling points to be smoothed in the first audio frame, it can be:

[0063] A first gain value in the first amplitude gain is smoothed according to a second amplitude gain corresponding to the second audio frame.

[0064] The first gain value is the amplitude gain value of a first preset number of sampling points adjacent to the second audio frame in the first audio frame. It is understood that the first gain value is the amplitude gain value of a first preset number of sampling points located at the front of the first audio frame after the sampling points in the first audio frame are sorted by timestamp. The first preset number may be a preset number of amplitude gain values ​​of the sampling points located at the front of the first audio frame to be smoothed.

[0065] Specifically, from the amplitude gain values ​​corresponding to the sampling points in the first audio frame, the amplitude gain values ​​of a first preset number of sampling points at the front are determined as the first gain value. Furthermore, the first gain value is smoothed according to the second amplitude gain. If the first preset number is less than the total number of sampling points in the first audio frame, the amplitude gain values ​​of the first portion of the sampling points in the first audio frame are smoothed. If the first preset number is equal to the total number of sampling points in the first audio frame, the amplitude gain values ​​of all the sampling points in the first audio frame are smoothed.

[0066] Based on the above example, when the first amplitude gain is greater than the second amplitude gain corresponding to the second audio frame, the first gain value in the first amplitude gain may be smoothed according to the second amplitude gain.

[0067] Specifically, when the first amplitude gain is greater than the second amplitude gain, the first gain value in the first amplitude gain is smoothed; when the first amplitude gain is less than or equal to the second amplitude gain, the first gain value in the first amplitude gain is not smoothed.

[0068] It should be noted that the reason for setting the above-mentioned conditional restrictions is that, when the first amplitude gain is greater than the second amplitude gain, since the second amplitude gain is smaller than the first amplitude gain, when smoothing the second audio frame, the amplitude gain value in the second audio frame may also be smoothed according to the first amplitude gain. By setting the above-mentioned conditions, the problem of poor audio quality caused by repeated smoothing can be avoided.

[0069] For example, the gain smoothing diagram when the first amplitude gain is greater than the second amplitude gain is as follows: Figure 2 As shown. Among them, G k represents the first amplitude gain corresponding to the first audio frame, G k-1 Indicates the second amplitude gain corresponding to the second audio frame.

[0070] Based on the above example, the smoothing method can be based on the sigmoid smoothing algorithm, that is, G k =sigmoid(G k , G k-1 ). Exemplarily, the first gain value in the first amplitude gain may be smoothed according to the second amplitude gain based on the following formula:

[0071] ΔG k =G k -G k-1

[0072]

[0073]

[0074] G k (m) = max{G k (m),1.0}

[0075] Among them, G k represents the first amplitude gain corresponding to the first audio frame, G k-1 represents the second amplitude gain corresponding to the second audio frame, ΔG k is the difference between the first amplitude gain and the second amplitude gain, M is the first preset number, x(m) is the smoothing coefficient corresponding to the mth sampling point, α and β are the preset smoothing parameter values, G k (m) is the amplitude gain value after smoothing corresponding to the m-th sampling point.

[0076] It should be noted that the smoothing parameter values ​​in the above formula can be set according to actual needs, for example: α=-5.0, β=10.0, etc. The specific values ​​are not specifically limited in this embodiment.

[0077] It should also be noted that when smoothing is performed in the above manner, the second amplitude gain corresponding to the second audio frame may be cached in advance.

[0078] Based on the above example, in order to effectively limit the range of sampling points in the first audio frame for smoothing the gain with the third audio frame, and accurately determine the amplitude gain value of the sampling points to be smoothed in the first audio frame, it can be:

[0079] The second gain value in the first amplitude gain is smoothed according to the third amplitude gain corresponding to the third audio frame.

[0080] The second gain value is the amplitude gain value of a second preset number of sampling points adjacent to the third audio frame in the first audio frame. It is understood that the second gain value is the amplitude gain value of a second preset number of sampling points located at the end of the first audio frame after the sampling points are sorted according to their timestamps. The second preset number may be a preset number of amplitude gain values ​​of the sampling points located at the end of the first audio frame to be smoothed.

[0081] Specifically, from the amplitude gain values ​​corresponding to the sampling points in the first audio frame, the amplitude gain values ​​of a second preset number of sampling points at the rear end are determined as the second gain value. Furthermore, the second gain value is smoothed according to the third amplitude gain. If the second preset number is less than the total number of sampling points in the first audio frame, the amplitude gain values ​​of the sampling points at the rear end of the first audio frame are smoothed. If the second preset number is equal to the total number of sampling points in the first audio frame, the amplitude gain values ​​of all sampling points in the first audio frame are smoothed.

[0082] Based on the above example, when the first amplitude gain is greater than the corresponding third amplitude gain of the third audio frame, the second gain value in the first amplitude gain may be smoothed according to the third amplitude gain.

[0083] Specifically, when the first amplitude gain is greater than the third amplitude gain, the second gain value in the first amplitude gain is smoothed; when the first amplitude gain is less than or equal to the third amplitude gain, the second gain value in the first amplitude gain is not smoothed.

[0084] It should be noted that the reason for setting the above-mentioned conditional restrictions is that, when the first amplitude gain is greater than the third amplitude gain, since the third amplitude gain is less than the first amplitude gain, when smoothing the third audio frame, the amplitude gain value in the third audio frame may also be smoothed according to the first amplitude gain. By setting the above-mentioned conditions, the problem of poor audio quality caused by repeated smoothing can be avoided.

[0085] For example, the gain smoothing diagram when the first amplitude gain is greater than the third amplitude gain is as follows: Figure 3 As shown. Among them, G k represents the first amplitude gain corresponding to the first audio frame, G k+1 Indicates the third amplitude gain corresponding to the third audio frame.

[0086] Based on the above example, the smoothing method can be based on the sigmoid smoothing algorithm, that is, G k =sigmoid(G k , G k+1 ). Exemplarily, the second gain value in the first amplitude gain may be smoothed according to the third amplitude gain based on the following formula:

[0087] ΔG k =G k+1 -G k

[0088]

[0089]

[0090] G k (m) = max{G k (m),1.0}

[0091] Among them, G k represents the first amplitude gain corresponding to the first audio frame, G k+1 Indicates the third amplitude gain corresponding to the third frequency frame, ΔG kis the difference between the third amplitude gain and the first amplitude gain, M is the second preset number, x(m) is the smoothing coefficient corresponding to the mth sampling point in the second preset number, α and β are preset smoothing parameter values, G k (m) is the amplitude gain value after smoothing corresponding to the m-th sampling point in the second preset number.

[0092] It should be noted that the smoothing parameter values ​​in the above formula can be set according to actual needs, for example: α=-5.0, β=10.0, etc. The specific values ​​are not specifically limited in this embodiment.

[0093] It should also be noted that when smoothing is performed in the above manner, the third amplitude gain corresponding to the third audio frame may be cached in advance.

[0094] S130 : Adjust the amplitude of the first audio frame based on the smoothed first amplitude gain to obtain a target audio frame.

[0095] The target audio frame may be an audio frame obtained by performing gain processing on the amplitude of the first audio frame, that is, the first audio frame after gain processing.

[0096] Specifically, an amplitude gain value corresponding to each sampling point in the first audio frame is determined based on the smoothed first amplitude gain. Furthermore, the amplitude of each sampling point in the first audio frame is multiplied by the amplitude gain value corresponding to each sampling point in the first audio frame to obtain an amplitude value after gain processing is performed on each sampling point. The first audio frame after gain processing is performed on each sampling point is used as the target audio frame.

[0097] The technical solution of the embodiment of the present invention obtains a first audio frame, determines a first amplitude gain corresponding to the first audio frame, and determines amplitude gain values ​​corresponding to each sampling point in the first audio frame. The first amplitude gain is smoothed according to a second amplitude gain corresponding to the second audio frame, and / or the first amplitude gain is smoothed according to a third amplitude gain corresponding to the third audio frame to smooth the first amplitude gain so that the amplitude gain value changes slowly. The amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain a target audio frame. This solves the problem of sudden changes in amplitude gain values ​​at the junction of adjacent audio frames, which leads to audio discontinuity and poor audio quality. The gain discontinuity is eliminated, thereby improving the subjective auditory quality of the audio.

[0098] Figure 4This is a flow chart of another audio frame processing method provided by an embodiment of the present invention. This embodiment is optimized based on the above-mentioned technical solutions. Optionally, after determining the first amplitude gain, before smoothing the first amplitude gain, the first amplitude gain may be updated; before determining the first amplitude gain corresponding to the first audio frame, a processing method for the first audio frame may be determined using a speech presence algorithm. The explanations of terms that are identical or corresponding to those in the above-mentioned embodiments are not repeated here.

[0099] like Figure 4 As shown, the method of this embodiment may specifically include:

[0100] S210: Obtain a first audio frame, and determine a speech presence probability of the first audio frame according to a preset speech presence algorithm.

[0101] The preset speech presence algorithm may be a predetermined method for determining the probability of speech presence in the first audio frame, and may be a traditional signal processing method or a neural network model method. The speech presence probability may be a probability value of speech presence in the first audio frame, i.e., a result of processing the first audio frame by the preset speech presence algorithm.

[0102] Specifically, a first audio frame is obtained, and the first audio frame is input into a preset voice presence algorithm to obtain an output result, which is used as the voice presence probability of the first audio frame.

[0103] Exemplarily, the probability of speech presence in the first audio frame is determined by the following formula:

[0104] p k =f(x k (n)),n=0,…,N-1

[0105] Among them, x k (n) represents the amplitude sequence of each sampling point in the first audio frame, N represents the number of sampling points in the first audio frame, f(·) represents the preset voice presence algorithm, p k Indicates the probability of speech presence in the first audio frame.

[0106] S220: Determine whether the voice existence probability reaches a preset voice threshold. If so, execute S230; if not, execute S270.

[0107] The preset voice threshold may be a pre-set threshold for determining whether voice exists in the first audio frame.

[0108] Specifically, whether the probability of speech presence reaches a preset speech threshold can be determined by:

[0109] b k =(pk >p th )? 1:0

[0110] Among them, b k A flag indicating whether there is speech in the first audio frame. "1" indicates that there is speech, and "0" indicates that there is no speech. k represents the probability of speech presence in the first audio frame, p th Indicates the preset voice threshold.

[0111] In practical applications, we can also k Performs a certain amount of smoothing control to eliminate the impact of short pauses or gaps in speech.

[0112] It should be noted that before performing gain smoothing processing, a preset speech existence algorithm can be used to first determine whether there is speech to be gain processed in the first audio frame. If so, the first amplitude gain is determined and the first amplitude gain is smoothed. If not, it indicates that there is no meaningful speech in the first audio frame, and gain processing is not required to reduce the amount of data processing in the audio frame processing process and improve processing efficiency.

[0113] S230: Determine a first amplitude gain corresponding to the first audio frame, and execute S240.

[0114] S240 . Smoothing the amplitude gain value of each sampling point in the first amplitude gain according to the second amplitude gain corresponding to the second audio frame to update the first amplitude gain, and executing S250 .

[0115] Specifically, the smoothing process in this step may be a weighted average process, in which the second amplitude gain and the first amplitude gain are weighted averaged, and the amplitude gain value obtained by the weighted average is used as the amplitude gain value for smoothing the amplitude gain value of each sampling point in the first amplitude gain. Then, the first amplitude gain is updated based on the amplitude gain value of the smoothed amplitude gain of each sampling point. The first amplitude gain may be updated based on the following formula:

[0116] G k =α·G k-1 +(1-α)·G k

[0117] Among them, G k Represents the first amplitude gain, G k-1 represents the second amplitude gain, and α represents a preset weighting coefficient.

[0118] It should be noted that the preset weighting coefficient can be set according to actual needs, for example, it can be 0.5, 0.6, etc. The specific value is not specifically limited in this embodiment.

[0119] Based on the above example, in order to avoid the problem of voice distortion caused by the first amplitude gain being too large and the problem of inaccurate determined amplitude gain, the first amplitude gain can be clipped using the first maximum amplitude gain of the first audio frame. Specifically, it can be:

[0120] The first amplitude gain is updated according to the first amplitude gain and a first maximum amplitude gain of the first audio frame.

[0121] The first maximum amplitude gain is a gain value determined according to a ratio of a preset maximum amplitude to a first maximum amplitude of the first audio frame.

[0122] Specifically, after determining the first amplitude gain and the first maximum amplitude gain of the first audio frame, the first amplitude gain may be processed so that the first amplitude gain is less than or equal to the first maximum amplitude gain. Furthermore, the first amplitude gain is updated according to the processed first amplitude gain.

[0123] Optionally, in order to avoid resource occupation and time consumption caused by large-scale data processing, the minimum value between the first amplitude gain and the first maximum amplitude gain may be selected as the updated first amplitude gain.

[0124] Specifically, it can be:

[0125] The first amplitude gain is updated according to a minimum value between the first amplitude gain and a first maximum amplitude gain of the first audio frame.

[0126] Specifically, if the first amplitude gain is less than or equal to the first maximum amplitude gain, the first amplitude gain is kept unchanged; if the first amplitude gain is greater than the first maximum amplitude gain, the first maximum amplitude gain is used as the updated first amplitude gain. The first amplitude gain may be updated based on the following formula:

[0127]

[0128] Among them, G k represents the first amplitude gain, Indicates the first maximum amplitude gain.

[0129] It should be noted that, based on the first amplitude gain and the first maximum amplitude gain of the first audio frame, updating the first amplitude gain may be performed between S230 and S240 and / or between S240 and S250 .

[0130] It should also be noted that the clipping process may cause discontinuity in the amplitude gain values ​​of two adjacent frames of speech, thereby causing discontinuity in the audio. Therefore, the amplitude gain value process may continue.

[0131] S250 . Smooth the first amplitude gain according to the second amplitude gain corresponding to the second audio frame; and / or smooth the first amplitude gain according to the third amplitude gain corresponding to the third audio frame, and execute S260 .

[0132] S260: Adjust the amplitude of the first audio frame based on the first amplitude gain after smoothing to obtain a target audio frame.

[0133] S270: Adjust the amplitude of the first audio frame based on a preset gain to obtain a target audio frame.

[0134] The preset gain may be an amplitude gain value set for the first audio frame in which no speech exists.

[0135] Specifically, when the speech existence probability does not reach the preset speech threshold, it can be determined that there is no speech in the first audio frame, and the target audio frame can be obtained by multiplying the amplitude of the first audio frame based on the preset gain.

[0136] It should be noted that, if the preset gain is 1, it indicates that no gain processing is performed on the first audio frame, and the first audio frame can be directly used as the target audio frame without adjusting the amplitude of the first audio frame.

[0137] The technical solution of the embodiment of the present invention is to obtain a first audio frame, determine the voice existence probability of the first audio frame according to a preset voice existence algorithm, and judge whether the voice existence probability reaches a preset voice threshold value. If not, the amplitude of the first audio frame is adjusted based on a preset gain to obtain a target audio frame, so as to quickly process the first audio frame and reduce resource occupation and time consumption. If it exists, a first amplitude gain corresponding to the first audio frame is determined, and the amplitude gain value of each sampling point in the first amplitude gain is smoothed according to the second amplitude gain corresponding to the second audio frame to update the first amplitude gain so that the first amplitude gain is greater than the first amplitude gain. The first amplitude gain is initially smoothed based on the second amplitude gain corresponding to the second audio frame; and / or the first amplitude gain is smoothed based on the third amplitude gain corresponding to the third audio frame to further smooth the first amplitude gain so that the amplitude gain value changes slowly. The amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain a target audio frame. This solves the problem of a sudden change in amplitude gain value at the junction of adjacent audio frames, which leads to audio discontinuity and poor audio quality. The gain discontinuity is eliminated, thereby improving the subjective auditory quality of the audio.

[0138] Figure 5 FIG. 1 is a flow chart of another audio frame processing method provided by an embodiment of the present invention. Figure 5 As shown, the method of this embodiment may specifically be:

[0139] A first audio frame is acquired and processed according to a preset speech presence algorithm to calculate a speech presence probability. The speech presence probability is compared with a preset speech threshold to determine whether speech is present, which is then used to control gain adjustment of the valid speech portion.

[0140] If the voice does not exist, a preset gain is obtained, and the first audio frame is adjusted according to the preset gain to obtain a target audio frame.

[0141] If speech exists, the first average amplitude gain and the first maximum amplitude gain of the first audio frame are calculated based on the first audio frame, and the first amplitude gain of the first audio frame is calculated based on the first average amplitude gain and the first maximum amplitude gain. At this time, the first amplitude gain can be cached so that the amplitude gains of adjacent audio frames can be smoothed later. The first amplitude gain is updated by the second amplitude gain corresponding to the second audio frame, and / or the first amplitude gain is clipped by the first maximum amplitude gain and updated to make the gains of adjacent frames change slowly. The first amplitude gain is smoothed according to the second amplitude gain corresponding to the second audio frame; and / or the first amplitude gain is smoothed according to the third amplitude gain corresponding to the third audio frame, and the amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain the target audio frame.

[0142] The technical solution of the embodiment of the present invention obtains a first audio frame and processes the first audio frame according to a preset speech presence algorithm. If speech is not present, the first audio frame is adjusted according to a preset gain to obtain a target audio frame. If speech is present, a first average amplitude gain and a first maximum amplitude gain of the first audio frame are calculated to calculate a first amplitude gain of the first audio frame. The first amplitude gain is updated using a second amplitude gain corresponding to a second audio frame, and / or the first amplitude gain is clipped using the first maximum amplitude gain to update the first amplitude gain. The first amplitude gain is then smoothed according to the second amplitude gain corresponding to the second audio frame, and / or the first amplitude gain is smoothed according to a third amplitude gain corresponding to a third audio frame. The amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain the target audio frame. This solves the problem of a sudden change in amplitude gain value at the junction of adjacent audio frames, which causes audio discontinuity and poor audio quality. The gain discontinuity is eliminated, thereby improving the subjective auditory quality of the audio.

[0143] Figure 6 This is a structural diagram of an audio frame processing device provided by an embodiment of the present invention. The device includes: a first amplitude gain determination module 310 , a first amplitude gain smoothing module 320 and a first audio frame gain module 330 .

[0144] Among them, the first amplitude gain determination module 310 is used to obtain a first audio frame and determine a first amplitude gain corresponding to the first audio frame; the first amplitude gain smoothing module 320 is used to smooth the first amplitude gain according to the second amplitude gain corresponding to the second audio frame; and / or, to smooth the first amplitude gain according to the third amplitude gain corresponding to the third audio frame; wherein, the second audio frame is an audio frame adjacent to the first audio frame and located before the first audio frame, and the third audio frame is an audio frame adjacent to the first audio frame and located after the first audio frame; the first audio frame gain module 330 is used to adjust the amplitude of the first audio frame based on the first amplitude gain after smoothing to obtain a target audio frame.

[0145] Optionally, the first amplitude gain smoothing module 320 is further used to smooth the first gain value in the first amplitude gain according to the second amplitude gain corresponding to the second audio frame; wherein the first gain value is the amplitude gain value of a first preset number of sampling points adjacent to the second audio frame in the first audio frame.

[0146] Optionally, the first amplitude gain smoothing module 320 is further configured to smooth the first gain value in the first amplitude gain according to the second amplitude gain when the first amplitude gain is greater than the second amplitude gain corresponding to the second audio frame.

[0147] Optionally, the first amplitude gain smoothing module 320 is further used to smooth the second gain value in the first amplitude gain according to the third amplitude gain corresponding to the third audio frame; wherein the second gain value is the amplitude gain value of a second preset number of sampling points in the first audio frame that are adjacent to the third audio frame.

[0148] Optionally, the first amplitude gain smoothing module 320 is further configured to smooth the second gain value in the first amplitude gain according to the third amplitude gain when the first amplitude gain is greater than a third amplitude gain corresponding to the third audio frame.

[0149] Optionally, the device further includes: a first updating module, configured to smooth the amplitude gain value of each sampling point in the first amplitude gain according to the second amplitude gain corresponding to the second audio frame, so as to update the first amplitude gain.

[0150] Optionally, the device also includes: a second updating module, used to update the first amplitude gain based on the first amplitude gain and the first maximum amplitude gain of the first audio frame, wherein the first maximum amplitude gain is a gain value determined based on the ratio of a preset maximum amplitude to the first maximum amplitude of the first audio frame.

[0151] Optionally, the second updating module is further configured to update the first amplitude gain according to a minimum value between the first amplitude gain and a first maximum amplitude gain of the first audio frame.

[0152] Optionally, the first amplitude gain determination module 310 is further used to determine a first average amplitude and a first maximum amplitude of the first audio frame based on the amplitudes of each sampling point in the first audio frame; determine a first average amplitude gain based on the first average amplitude and the preset average amplitude; determine a first maximum amplitude gain based on the first maximum amplitude and the preset maximum amplitude; and determine a first amplitude gain corresponding to the first audio frame based on the first average amplitude gain and the first maximum amplitude gain.

[0153] Optionally, the device also includes: a speech presence judgment module, which is used to determine the speech presence probability of the first audio frame according to a preset speech presence algorithm; when the speech presence probability reaches a preset speech threshold value, triggering the execution of an operation to determine the first amplitude gain corresponding to the first audio frame.

[0154] Optionally, the device further includes: a preset processing module, configured to adjust the amplitude of the first audio frame based on a preset gain to obtain a target audio frame when the probability of the speech existence does not reach a preset speech threshold value.

[0155] The technical solution of the embodiment of the present invention obtains a first audio frame, determines a first amplitude gain corresponding to the first audio frame, and determines amplitude gain values ​​corresponding to each sampling point in the first audio frame. The first amplitude gain is smoothed according to a second amplitude gain corresponding to the second audio frame, and / or the first amplitude gain is smoothed according to a third amplitude gain corresponding to the third audio frame to smooth the first amplitude gain so that the amplitude gain value changes slowly. The amplitude of the first audio frame is adjusted based on the smoothed first amplitude gain to obtain a target audio frame. This solves the problem of sudden changes in amplitude gain values ​​at the junction of adjacent audio frames, which leads to audio discontinuity and poor audio quality. The gain discontinuity is eliminated, thereby improving the subjective auditory quality of the audio.

[0156] The audio frame processing device provided in the embodiment of the present invention can execute the audio frame processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0157] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of the present invention.

[0158] Figure 7 The present invention provides a schematic structural diagram of an electronic device. Figure 7 A block diagram of an exemplary electronic device 40 suitable for implementing exemplary embodiments of the present invention is shown. Figure 7 The electronic device 40 shown is only an example and should not limit the functionality and scope of use of the embodiments of the present invention.

[0159] like Figure 7 As shown, electronic device 40 is a general-purpose computing device. Components of electronic device 40 may include, but are not limited to, one or more processors or processing units 401, system memory 402, and a bus 403 connecting various system components (including system memory 402 and processing unit 401).

[0160] Bus 403 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0161] The electronic device 40 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 40, including volatile and non-volatile media, removable and non-removable media.

[0162] System memory 402 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 404 and / or cache 405. Electronic device 40 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 406 may be used to read and write non-removable, non-volatile magnetic media ( Figure 7 Not shown, often called a "hard drive"). Although Figure 7Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 403 via one or more data media interfaces. System memory 402 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0163] A program / utility 408 having a set (at least one) of program modules 407 may be stored, for example, in system memory 402. Such program modules 407 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 407 generally perform the functions and / or methods of the embodiments described herein.

[0164] The electronic device 40 may also communicate with one or more external devices 409 (e.g., keyboard, pointing device, display 410, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 40, and / or communicate with any device that enables the electronic device 40 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through an I / O interface (input / output interface) 411. Furthermore, the electronic device 40 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 412. As shown, the network adapter 412 communicates with other modules of the electronic device 40 via the bus 403. It should be understood that although Figure 7 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 40, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0165] The processing unit 401 executes various functional applications and data processing by running programs stored in the system memory 402, such as implementing the audio frame processing method provided in the embodiment of the present invention.

[0166] An embodiment of the present invention further provides a storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform an audio frame processing method, the method comprising:

[0167] Acquire a first audio frame, and determine a first amplitude gain corresponding to the first audio frame;

[0168] performing smoothing on the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and / or performing smoothing on the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to and preceding the first audio frame, and the third audio frame is an audio frame adjacent to and following the first audio frame;

[0169] The amplitude of the first audio frame is adjusted based on the first amplitude gain after smoothing to obtain a target audio frame.

[0170] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0171] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0172] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0173] The computer program code for performing the operations of the embodiments of the present invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0174] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. An audio frame processing method, characterized in that: include: Acquire a first audio frame, and determine a first amplitude gain corresponding to the first audio frame; performing smoothing on a first gain value in the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and performing smoothing on a second gain value in the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to and before the first audio frame, and the third audio frame is an audio frame adjacent to and after the first audio frame; the first gain value is an amplitude gain value of a first preset number of sampling points in the first audio frame adjacent to the second audio frame; and the second gain value is an amplitude gain value of a second preset number of sampling points in the first audio frame adjacent to the third audio frame; The amplitude of the first audio frame is adjusted based on the first amplitude gain after smoothing to obtain a target audio frame.

2. The method according to claim 1, characterized in that The smoothing of the first gain value in the first amplitude gain according to the second amplitude gain corresponding to the second audio frame includes: In a case where the first amplitude gain is greater than a second amplitude gain corresponding to the second audio frame, a first gain value in the first amplitude gain is smoothed according to the second amplitude gain.

3. The method according to claim 1, characterized in that The smoothing of the second gain value in the first amplitude gain according to the third amplitude gain corresponding to the third audio frame includes: In a case where the first amplitude gain is greater than a corresponding third amplitude gain of the third audio frame, a second gain value in the first amplitude gain is smoothed according to the third amplitude gain.

4. The method according to claim 1, wherein After acquiring the first audio frame and determining the first amplitude gain corresponding to the first audio frame, performing smoothing processing on the first amplitude gain according to the second amplitude gain corresponding to the second audio frame; And / or, before smoothing the first amplitude gain according to the third amplitude gain corresponding to the third audio frame, the method further includes: According to the second amplitude gain corresponding to the second audio frame, the amplitude gain value of each sampling point in the first amplitude gain is smoothed to update the first amplitude gain.

5. The method according to claim 1, wherein After obtaining the first audio frame and determining the first amplitude gain corresponding to the first audio frame, the method further includes: The first amplitude gain is updated according to the first amplitude gain and a first maximum amplitude gain of the first audio frame, wherein the first maximum amplitude gain is a gain value determined according to a ratio of a preset maximum amplitude to the first maximum amplitude of the first audio frame.

6. The method according to claim 5, characterized in that Updating the first amplitude gain according to the first amplitude gain and the first maximum amplitude gain of the first audio frame includes: The first amplitude gain is updated according to a minimum value between the first amplitude gain and a first maximum amplitude gain of the first audio frame.

7. The method according to claim 1, characterized in that The determining a first amplitude gain corresponding to the first audio frame includes: determining a first average amplitude and a first maximum amplitude of the first audio frame according to the amplitude of each sampling point in the first audio frame; determining a first average amplitude gain according to the first average amplitude and a preset average amplitude; determining a first maximum amplitude gain according to the first maximum amplitude and a preset maximum amplitude; A first amplitude gain corresponding to the first audio frame is determined according to the first average amplitude gain and the first maximum amplitude gain.

8. The method according to claim 1, characterized in that Before determining the first amplitude gain corresponding to the first audio frame, the method further includes: determining a speech presence probability of the first audio frame according to a preset speech presence algorithm; When the speech existence probability reaches a preset speech threshold, an operation of determining a first amplitude gain corresponding to the first audio frame is triggered.

9. The method according to claim 8, characterized in that The method further comprises: When the speech existence probability does not reach a preset speech threshold, the amplitude of the first audio frame is adjusted based on a preset gain to obtain a target audio frame.

10. An audio frame processing device, characterized in that: include: a first amplitude gain determining module, configured to obtain a first audio frame and determine a first amplitude gain corresponding to the first audio frame; a first amplitude gain smoothing module, configured to smooth a first gain value in the first amplitude gain according to a second amplitude gain corresponding to a second audio frame; and smooth a second gain value in the first amplitude gain according to a third amplitude gain corresponding to a third audio frame; wherein the second audio frame is an audio frame adjacent to and before the first audio frame, and the third audio frame is an audio frame adjacent to and after the first audio frame; the first gain value is an amplitude gain value of a first preset number of sampling points in the first audio frame adjacent to the second audio frame; and the second gain value is an amplitude gain value of a second preset number of sampling points in the first audio frame adjacent to the third audio frame; The first audio frame gain module is configured to adjust the amplitude of the first audio frame based on the first amplitude gain after smoothing to obtain a target audio frame.

11. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the audio frame processing method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the audio frame processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Automatic gain control method for audio signal and apparatus thereof

    CN101110217A

  • Sound signal processing method, device and equipment

    CN112133299A