An audio signal processing method, apparatus, device and medium
By determining the predicted noise segments and noise thresholds through Fourier transform and frequency cutting, and combining smoothing filters and inverse Fourier transform for noise reduction, and performing loudness gain operation on each frame of audio data, the problems of inaccurate noise estimation and reduced sound quality and volume are solved, achieving high signal-to-noise ratio and high audio loudness fidelity audio signal processing.
Patent Information
- Application Number
- CN202310311891.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-03-21
AI Technical Summary
Existing audio signal processing methods suffer from inaccurate noise estimation and reduced sound quality and volume after noise reduction, resulting in cumbersome processing and distortion.
The predicted noise segments and noise thresholds are determined by Fourier transform and frequency cutting. Noise reduction is performed by combining smoothing filters and inverse Fourier transform. Loudness gain is applied to each frame of audio data to achieve adaptive noise selection and automatic gain for each frame.
It improves the general noise reduction quality of audio signal processing, ensures improved signal-to-noise ratio and preserved audio loudness, and avoids the distortion and volume reduction problems of traditional methods.
Smart Images

Figure CN116524950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of audio signal processing technology of audio or voice, video software, and particularly relate to an audio signal processing method, device, equipment and medium. BACKGROUND
[0002] In the field of audio signal processing technology, general noise reduction is a relatively important part. The goal of general noise reduction is to remove some low-frequency noise in recording or speaking sound. In the editing of voice and video software, general noise reduction is often used to help editors to preliminarily process audio and improve the signal-to-noise ratio of audio.
[0003] In a common audio workstation, a noise estimation module is provided to allow a user to select a piece of audio as a pre-noise estimation piece of audio for the entire audio. However, the user cannot accurately grasp the noise piece, and needs to repeatedly listen to confirm. Meanwhile, the audio after general noise reduction often has the problem of decreased sound quality and volume. The traditional method is to use a specific gain to increase the volume of the entire audio signal, which may cause distortion. Therefore, the selection of noise estimation and the gain processing after noise reduction in the current audio signal processing method are not perfect. SUMMARY
[0004] To solve the problems in the related art, embodiments of the present application provide an audio signal processing method, which can consider the adaptive selection of noise pieces and automatic gain, and greatly improve the general noise reduction quality of audio signal processing.
[0005] In a first aspect, embodiments of the present application provide an audio signal processing method, which can include: obtaining first audio data, the first audio data including a plurality of first audio data pieces, each first audio piece including a plurality of first audio data frames; preprocessing the plurality of first audio data pieces to obtain a plurality of second audio data pieces, each second audio piece including a plurality of second audio data frames; determining a pre-estimated noise piece and a noise threshold in the plurality of second audio data pieces based on a plurality of power spectrums corresponding to the plurality of second audio data pieces and a silent condition; performing noise reduction processing on each second audio data piece based on the plurality of second audio data frames of each second audio data piece and the noise threshold to obtain a plurality of third audio data pieces, each third audio piece including a plurality of third audio data frames; and performing a loudness gain operation on each third audio data frame in each third audio data piece to obtain target audio data.
[0006] Further, the preprocessing the plurality of first audio data pieces to obtain a plurality of second audio data pieces, each second audio piece including a plurality of second audio data frames, includes: performing Fourier transform on the plurality of first audio data pieces and performing frequency cutting to obtain the plurality of second audio data pieces.
[0007] Further, the determining the estimated noise segment and the noise threshold in the second audio segments based on the power spectrum corresponding to each of the second audio data segments and the noise condition comprises: sorting the power spectrum corresponding to each of the second audio data segments based on the size; determining the estimated noise segment based on the sorted power spectrum and the noise condition; and calculating the noise threshold based on the estimated noise segment.
[0008] Further, the noise reduction processing the second audio data segments based on the second audio data frames of each of the second audio data segments and the noise threshold to obtain third audio segments, wherein each of the third audio segments comprises third audio data frames, comprises: calculating a masking matrix based on the comparison between each of the second audio data frames and the noise threshold; calculating a smoothing filter based on the second audio data segments; and processing the second audio data segments based on the masking matrix and the smoothing filter and performing inverse Fourier transform to obtain the third audio segments.
[0009] Further, the loudness gain operation on each of the third audio data frames in each of the third audio segments to obtain target audio data comprises: calculating a difference between a loudness value of each of the third audio data frames and a loudness value of each of the first audio data frames corresponding to the third audio data frames; calculating a loudness gain of each of the third audio data frames based on the comparison between each of the differences; and calculating the target audio data based on each of the loudness gains and each of the third audio data frames.
[0010] Further, the calculating the noise threshold based on the estimated noise segment comprises: calculating a mean value and a variance of the estimated noise segment; and determining the noise threshold as a sum of the mean value and a product of the variance and a balance constant.
[0011] Further, the calculating the loudness gain of each of the third audio data frames based on the comparison between each of the differences comprises: calculating a minimum value and a mean value of the differences; if any of the differences is greater than the mean value, assigning the mean value as the loudness gain of the third audio data frame corresponding to the difference; and if any of the differences is less than the mean value, assigning the minimum value as the loudness gain of the third audio data frame corresponding to the difference.
[0012] In a second aspect, an embodiment of the present application further provides an audio signal processing device, which can comprise:
[0013] The acquisition module is configured to acquire first audio data, the first audio data comprising a plurality of first audio data segments, each first audio data segment comprising a plurality of first audio data frames; the preprocessing module is configured to preprocess the plurality of first audio data segments to obtain a plurality of second audio data segments, each second audio data segment comprising a plurality of second audio data frames; the noise estimation module is configured to determine an estimated noise segment and a noise threshold based on a plurality of power spectrums corresponding to the plurality of first audio data segments and a silent condition; the noise reduction processing module is configured to calculate a plurality of third audio data segments based on the plurality of second audio data frames of each second audio data segment and the noise threshold, each third audio data segment comprising a plurality of third audio data frames; and the loudness gain operation module is configured to perform a loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data.
[0014] In a third aspect, an embodiment of the present application further provides a computer device, comprising: a memory and a processor, the memory being configured to store and support the processor to execute a program of any of the methods in the first aspect, and the processor being configured to execute the program stored in the memory.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer readable medium having non-volatile program codes executable by a processor, wherein the program codes cause the processor to execute any of the methods in the first aspect.
[0016] In the embodiments of the present application, the first audio data is acquired, the first audio data comprising a plurality of first audio data segments, each first audio data segment comprising a plurality of first audio data frames, the plurality of first audio data segments are preprocessed to obtain a plurality of second audio data segments, each second audio data segment comprising a plurality of second audio data frames, an estimated noise segment and a noise threshold in the plurality of second audio data segments are determined based on a plurality of power spectrums corresponding to the plurality of first audio data segments and a silent condition, the adaptive selection of the estimated noise segment of the first audio data is realized, the process complexity and estimation errors caused by repeated listening and selection of the user are avoided, a plurality of third audio data segments are obtained by noise reduction processing based on the plurality of second audio data frames of each second audio data segment and the noise threshold, each third audio data segment comprising a plurality of third audio data frames, the target audio data is obtained by performing a loudness gain operation on each third audio data frame in each third audio data segment, that is, the loudness gain operation is performed on each third audio data frame in each third audio data segment, the distortion caused by assigning the same gain coefficient to the entire audio data is avoided, and the frame-based automatic gain of the audio data is realized. Thus, in the general noise reduction of audio signal processing, a high signal-to-noise ratio is realized, the loudness fidelity before and after the audio signal noise reduction is ensured, and the quality of the processed audio signal is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from the structures shown in the drawings without creative labor.
[0018] Figure 1 A flowchart of an audio signal processing method provided by an embodiment of the present application;
[0019] FIG. 2(a) is a spectrum diagram of an original audio sample provided by an embodiment of the present application;
[0020] FIG. 2(b) is a spectrum diagram of a preprocessed audio sample provided by an embodiment of the present application;
[0021] Figure 3 A flowchart of an audio signal processing method provided by another embodiment of the present application;
[0022] Figure 4 A flowchart of an audio signal processing method provided by another embodiment of the present application;
[0023] Figure 5 A flowchart of an audio signal processing method provided by another embodiment of the present application;
[0024] Figure 6 A flowchart of an audio signal processing method provided by another embodiment of the present application;
[0025] Figure 7 A flowchart of an audio signal processing method provided by another embodiment of the present application;
[0026] Figure 8 A schematic block diagram of an audio signal processing device provided by an embodiment of the present application;
[0027] Figure 9 A schematic block diagram of a computer device provided by an embodiment of the present application.
[0028] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION
[0029] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of the present application.
[0030] It should be understood that the terms “comprise” and “include” as used in the specification and the appended claims indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0031] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms “a”, “an” and “the” are intended to include the plural forms as well.
[0032] It should be further understood that the term “and / or” as used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0033] It should be still further understood that the terms “first”, “second” in the present application specification and claims and the above-described drawings are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with “first”, “second” can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of “multiple” is two or more. The specific meaning of the above terms in the present application can be understood by those of ordinary skill in the art according to the specific circumstances.
[0034] In the technical field of audio signal processing method, general noise reduction is an important part. The goal of general noise reduction is to remove some low-frequency noise in recording or speaking sound. In the editing of voice and video software, general noise reduction is often used to help editors to preliminarily process audio and improve the signal-to-noise ratio of audio.
[0035] In a common audio workstation, such as a Digital Audio Workstation (DAW), a noise estimation module, such as an adobe audiotion module or an Audacity module, allows a user to select a piece of audio as a noise estimation piece of the entire audio. In actual applications, the user usually cannot accurately select the noise estimation piece, and the human selection will affect the effect of audio signal processing. Usually, it needs to be repeatedly listened to confirm, which leads to the complexity and uncertainty of the processing process. In addition, the audio signal after general noise reduction often has the problem of sound quality and volume decrease. The traditional method is to use a selected gain coefficient after processing to improve the loudness of the entire audio signal, which has a negative effect on the audio signal piece with low noise reduction amplitude, resulting in distortion problem.
[0036] Therefore, the selection of noise estimation and the gain processing after noise reduction in the current audio signal processing method are not perfect.
[0037] Therefore, the selection of noise estimation and the gain processing after noise reduction in the current audio signal processing method are not perfect.
[0038] Firstly, the related concepts that may be involved in the present application are briefly introduced.
[0039] Audio data: There are two data forms for describing audio signals, time domain signal and frequency domain signal. Time domain is a description of the relationship between mathematical functions or physical signals and time. For example, the time domain waveform of an audio signal can express the change of the audio signal with time. Frequency domain is a coordinate system used to describe the characteristics of signals in frequency. The frequency domain graph shows the amount of audio signal in each given frequency band within a frequency range. The transformation of audio signal from time domain to frequency domain is mainly realized through Fourier series and Fourier transform.
[0040] Audio data piece and audio data frame: Audio data can be divided into pieces according to specific nodes, and the divided pieces can be further divided into frames. Usually, the node is a time node.
[0041] Power spectrum: the power spectrum is a brief for power spectral density function, which is defined as the signal power in a unit frequency band. It shows the change of signal power with frequency, i.e. the distribution of signal power in frequency domain. The power spectrum is calculated as the sum of the square of the amplitude value and the square of the phase value of the frequency domain signal, and then the sum is opened.
[0042] VAD (Voice Activity Detection) algorithm: the main principle of the algorithm is to divide the signal into multiple frequency bands in the frequency spectrum, calculate the energy of each frequency band as a feature, and construct a mixed Gaussian distribution model of non-sound and sound for each sub-band by hypothesis testing. Through maximum likelihood estimation, the model is adaptively learned and optimized, and the probability ratio decision is used to infer the non-sound and sound two segments.
[0043] Non-sound condition: the non-sound condition in the present application is defined as finding a sound segment that meets the sound condition in a piece of audio signal to be processed, and judging the remaining segment as a non-sound segment after removing the sound segment.
[0044] Loudness: sound pressure level (SPL) is commonly used to describe the loudness of sound, with unit of dB. The measured sound pressure value is divided by the standard sound pressure value (1 kHz), and the logarithm is taken and multiplied by 10 to get SPL. The measured sound pressure value is generally recorded in the time domain signal of the audio signal.
[0045] The above is only a description of the technical principles and exemplary application framework of the embodiments of the present application. The specific technical solutions of the embodiments of the present application will be described in detail through multiple embodiments. Please refer to Figure 1 An embodiment of an audio signal processing method in the embodiments of the present application can include:
[0046] S100: acquiring first audio data, the first audio data including a plurality of first audio data segments, the first audio segment including a plurality of first audio data frames.
[0047] In some embodiments, the first audio data is acquired, the first audio data is a time domain signal, time is taken as a decomposition node, the first audio data can be decomposed into first audio data segments, and time is taken as a decomposition node, which can be further decomposed into first audio data frames.
[0048] In some embodiments, the first audio data is divided into a plurality of first audio data frames with a time node of 25 ms. It can be understood that 25 ms is a classical windowing method based on a sampling rate of 16000 Hz. The first audio data can include a plurality of first audio data frames, and 50 first audio data frames are taken as a first audio data segment of the first audio data. The audio duration of 50 is taken from a general audio event detection model, such as the Google event detection model.
[0049] In some embodiments, the first audio data is divided into a plurality of first audio data segments with a time node of 50 multiplied by 25 ms. Each first audio data segment of 50 multiplied by 25 ms is sequentially decomposed into the first audio data. If the last segment does not meet the duration, a first audio segment is supplemented by a 0 operation. After all steps are completed, the 0 part is deleted to obtain target audio data aligned with the time length of the first audio data.
[0050] S200: Preprocessing the plurality of first audio data segments to obtain a plurality of second audio data segments, the second audio data segments including a plurality of second audio data frames.
[0051] Specifically, the plurality of first audio data segments are preprocessed to obtain a plurality of second audio data segments. The preprocessing between the plurality of first audio data segments and the plurality of second audio data segments does not involve a change in length, but each first audio data segment in the plurality of first audio data segments is preprocessed to obtain a one-to-one corresponding plurality of second audio data segments, and a one-to-one corresponding plurality of second audio data frames are obtained.
[0052] More specifically, the preprocessing operation includes Fourier transform and frequency cutting of the plurality of first audio data segments to obtain a plurality of second audio data segments. Wherein, the Fourier transform of the plurality of first audio data segments is to convert the plurality of first audio data segments from time domain signals to frequency domain signals, i.e. frequency domain signals. After conversion to frequency domain signals, frequency cutting is performed according to frequency characteristics to obtain a plurality of second audio data segments. The second audio data segments correspondingly include a plurality of second audio data frames.
[0053] In some embodiments, after the Fourier transform of the plurality of first audio data segments, i.e. each first audio data segment is transformed from a time domain signal to a frequency domain signal, since the general noise reduction of audio signal processing mainly acts on the low frequency domain, further, the frequency cutting is performed on the converted frequency domain signal, in the embodiment, the method of frequency cutting uses the method of low-pass filter to filter the high frequency signal. Please refer to FIG. 2(a), which is a frequency spectrum diagram of the original audio sample provided by the embodiment of the present application, for example, the frequency spectrum diagram obtained after the Fourier transform of the plurality of first audio data segments is shown in FIG. 2(a), the frequency domain is from 0 to K1 (Hz), and after using the low-pass filter to filter the high frequency part, the frequency spectrum diagram shown in FIG. 2(b) is obtained, at this time, the frequency domain is cut from 0 to K1 (Hz) to 0 to K2 (Hz), and FIG. 2(b) is a frequency spectrum diagram of the preprocessed audio sample provided by the embodiment of the present application. The subsequent audio signal processing steps operate on the plurality of second audio data segments obtained after the frequency cutting, and it can be understood that each second audio segment includes a plurality of second audio data frames, and each second audio data frame is also a frequency domain signal after frequency cutting.
[0054] S300: determining the estimated noise segment and the noise threshold in the plurality of second audio segments based on the plurality of power spectrums corresponding to the plurality of first audio data segments and the silent condition.
[0055] In a further embodiment, in order to realize adaptive selection of the estimated noise segment, in the embodiment, the estimated noise segment and the noise threshold in the second audio segment are determined based on the plurality of power spectrums corresponding to the plurality of first audio data segments and the silent condition. It can be understood that the embodiment determines the estimated noise segment and the noise threshold in the plurality of second audio segments by using the plurality of power spectrums corresponding to the plurality of first audio data segments and the silent condition, avoids the process redundancy caused by artificial selection of the noise estimated segment and the estimation error caused thereby, and improves the efficiency and accuracy of the audio signal processing.
[0056] In some embodiments, as shown in Figure 3 Figure 3 is a flowchart of an audio signal processing method provided by another embodiment of the present application. It includes: performing size sorting based on the plurality of power spectrums corresponding to the plurality of second audio data segments; determining the estimated noise segment based on the sorted plurality of power spectrums and the silent condition; and calculating the noise threshold based on the estimated noise segment.
[0057] More specifically, after power spectrum calculation is performed on the frequency domain signal of each of the second audio data segments, the corresponding plurality of power spectra are sorted from large to small, and the voiceless segments in the plurality of second audio data segments are determined using a voiceless condition. For example, the time indexes of the voiced segments in the plurality of first audio data segments are calculated using the VAD algorithm provided in the MATLAB software, and the voiceless segments in the plurality of first audio data segments are obtained by removing the voiced segments. The time indexes of the voiceless segments in the plurality of second audio data segments are corresponded one by one, and the voiceless segments in the plurality of second audio data segments are obtained. The second audio data segments with high power spectrum are determined as the estimated noise segments, and the noise threshold is calculated according to the frequency domain signal of the determined estimated noise segments. It can be understood that the corresponding plurality of power spectra are sorted by size, and the part with a power spectrum greater than a certain threshold is determined as high, or the first third part after sorting is determined as high. The high power spectrum indicates that there is a frequency vibration, that is, there is a sound vibration other than human voice, which can be estimated as noise. The definition of the high power spectrum is not limited here, and is subject to the actual processing requirements. It can be understood that in the example of using other algorithms, the voiceless segments in the plurality of second audio data segments can be directly confirmed without corresponding to the voiceless segments in the plurality of second audio data segments again through the time indexes of the voiceless segments in the plurality of first audio data segments.
[0058] In some embodiments, as shown in Figure 4 , Figure 4 is a flowchart of an audio signal processing method provided by another embodiment of the present application. The noise threshold is calculated based on the estimated noise segment, including: calculating the mean and variance of the noise estimated segment; and determining the sum of the mean and the product of the variance and a balance constant as the noise threshold. It can be understood that the noise threshold method is a preferred scheme of the present embodiment, and in other embodiments, the calculation method can be determined according to actual requirements.
[0059] Specifically, the estimated noise segment can be one or more of the plurality of second audio data segments. However, the noise threshold can be calculated through the estimated noise segment regardless of the number of the estimated noise segment. Since the estimated noise segment is a frequency domain signal, the mean and variance of the frequency domain signal of the noise estimated segment are calculated, and the noise threshold noise_thresh is:
[0060] noise_thresh = mean_noise + noise_std * std_thresh_stationary (1)
[0061] wherein the balance constant std_thresh_stationary is used to ensure that the standard deviation of the noise threshold is higher than the average value to place the noise threshold between the valid signal and the noise of the audio signal, and the balance constant in this embodiment is 1.5.
[0062] S400: obtaining a plurality of third audio data segments based on the plurality of second audio data frames of each second audio data segment and the noise threshold, wherein each third audio segment comprises a plurality of third audio data frames. In this step, each second audio data segment is processed to obtain a corresponding third audio data segment after noise reduction using the plurality of second audio data frames of each second audio data segment and the noise threshold obtained in the above step, so that noise reduction is realized for each second audio data segment respectively, avoiding distortion caused by using a specific noise threshold to reduce noise of all audio data segments, and effectively ensuring the quality of the audio signal after noise reduction.
[0063] In some embodiments, as shown in Figure 5 , the method comprises the following steps. Figure 5 Fig. 2 shows a flowchart of an audio signal processing method according to another embodiment of the present application. The method comprises the following steps.
[0064] Specifically, after determining the noise threshold in the above step, a valve is obtained. In this embodiment of the present application, for each second audio data segment, a mask matrix is calculated by comparing the frequency domain signal of each second audio data frame of the second audio data segment with the noise threshold. Specifically, each second audio data frame is compared with the noise threshold one by one. If the former is smaller than the latter, it means that the frequency domain signal of the second audio data frame needs to be weakened, and if the former is larger than the latter, it remains unchanged. In this way, the second audio data frame smaller than the noise threshold in the second audio data segment is recorded as 0, and the second audio data frame larger than the noise threshold in the second audio data segment is recorded as 1, so as to obtain a mask matrix sig_mask, wherein the size of the mask matrix sig_mask is consistent with the size of each second audio data segment.
[0065] Further, the smoothing filter is calculated based on the plurality of second audio data segments. Since the plurality of second audio data segments are frequency domain signals, various smoothing filters can be used for smoothing, such as neighborhood smoothing filter, median filter, etc. In this embodiment, a smoothing filter smoothing_filter is set according to the plurality of second audio data segments, which is used to smooth the noise value points in the above-mentioned masking matrix sig_mask, i.e. the second audio data frames that need to be weakened. The smoothing filter smoothing_filter is determined by the Fourier point number and the window length in the Fourier transformation of the plurality of second audio data segments, and is realized by using a combination matrix of four arithmetic sequences, which are the arithmetic sequences of 0 to 1, the arithmetic sequences of 1 to 0, the point number of which is the sampling rate sr divided by the Fourier point number divided by 2, i.e. sr / n_fft / 2, and the arithmetic sequences of 0 to 1, the arithmetic sequences of 1 to 0, the point number of which is the Fourier window length divided by the sampling rate, i.e. hop_length / sr. Thus, by setting the smoothing filter, the fluctuation curve of noise reduction can be smoothed, ensuring that the subsequent inverse Fourier transformation will not produce distortion problems, and at the same time ensuring that the sound quality change of the audio signal will not be high and low, maintaining the good change of the linearity of the audio signal curve. In combination with the smoothing filter smoothing_filter and the masking matrix sig_mask, a new masking matrix new_sig_mask with a numerical range of 0-1 can be obtained, sig_stft_denoised is the second audio data segment after noise reduction processing, and sig_stft is the second audio data segment before noise reduction:
[0066] sig_stft_denoised = sig_stft * new_sig_mask (2)
[0067] Thus, after obtaining the plurality of second audio data segments after noise reduction, inverse Fourier transformation is performed to obtain a plurality of third audio segments corresponding one by one. At this time, the third audio segment is a time domain signal after noise reduction processing, and each third audio data segment includes a plurality of third audio data frames.
[0068] S500: performing a loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data. Thus, by performing a loudness gain operation on each third audio data frame of each third audio segment respectively, distortion caused by assigning the same gain coefficient to the entire audio data is avoided, realizing automatic gain of audio data in frames, making the loudness fidelity of audio signal before and after noise reduction, and ensuring the quality of the processed audio signal.
[0069] Please refer to Figure 6 , Figure 6is a flowchart of an audio signal processing method provided by another embodiment of the present application. In this embodiment, the loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data includes: calculating a difference value between the loudness value of each third audio data frame and the loudness value of a corresponding first audio data frame; calculating the loudness gain of each third audio data frame based on each difference value; and calculating target audio data based on each loudness gain and each third audio data frame.
[0070] Specifically, for each first audio data segment and a corresponding third audio data segment, i.e., the time domain signals before and after noise reduction, the loudness value origin_db of the first audio data segment and the loudness value denoised_db of the third audio data segment are calculated respectively, where the loudness calculation formula is:
[0071] X db = 20*log10(data) (3)
[0072] where data represents the audio data segment to be calculated, X db is the loudness value, and the number of output loudness values is consistent with the number of data frames of the input audio data segment. Further, a difference value off_db between the loudness value of each first audio data frame and the loudness value of a corresponding third audio data frame is obtained by subtraction:
[0073] origin_db-denoised_db = off_db (4)
[0074] Further, the loudness gain of each third audio data frame is calculated based on each difference value. The calculation method of the loudness gain of each third audio data frame is designed according to the actual required error range.
[0075] In some embodiments, please refer to Figure 7 , Figure 7 is a flowchart of an audio signal processing method provided by another embodiment of the present application. The calculation of the loudness gain of each third audio data frame based on each difference value further includes: calculating the minimum value and the average value of the plurality of difference values, if any difference value is greater than the average value, assigning the average value as the loudness gain of the third audio data frame corresponding to the difference value; and if any difference value is less than the average value, assigning the minimum value as the loudness gain of the third audio data frame corresponding to the difference value.
[0076] Specifically, based on each third data segment, the minimum value off_db_min and the average value off_db_mean of the plurality of difference values obtained in the above step are calculated, and the following steps are performed:
[0077] The size of any difference value off_db and off_db_min and off_db_mean is compared, and if the difference value off_db>off_db_mean, it indicates that the loudness value of the third audio data frame corresponding to the difference value decreases more after noise reduction, and then off_db_mean is added to it. If the difference value off_db<off_db_mean, in order to prevent the loudness from exceeding the boundary, off_db_min is added to it.
[0078] In this way, for each third audio data segment, the modified loudness gain enhance_db is obtained, and further, the loudness gain enhance_db is converted to the linear domain:
[0079]
[0080] Further, each third audio data segment is multiplied by each loudness gain corresponding to the above formula to finally obtain the target audio data.
[0081] In some embodiments, in order to further check the accuracy of the target audio data, it can also include: confirming the time sequence number of each target audio data segment in the target audio data, and splicing according to the time sequence number.
[0082] In some embodiments, in the last target audio data segment of the target audio data, since the 0 filling operation is performed in the first step, the filled part needs to be further cut off to obtain the target audio data aligned with the first audio data.
[0083] In summary, in the embodiments of the present application, the first audio data is obtained, the first audio data includes a plurality of first audio data segments, the first audio data segment includes a plurality of first audio data frames, the plurality of first audio data segments are preprocessed to obtain a plurality of second audio data segments, the second audio data segment includes a plurality of second audio data frames, the estimated noise segment and the noise threshold in the plurality of second audio data segments are determined based on the plurality of power spectrums corresponding to the plurality of first audio data segments and the silent condition, the adaptive selection of the estimated noise segment of the first audio data is realized, the process complexity and the estimation error caused by repeated listening and selection of the user are avoided, the plurality of third audio data segments are obtained by noise reduction processing based on the plurality of second audio data frames of each second audio data segment and the noise threshold, the third audio data segment includes a plurality of third audio data frames, the target audio data is obtained by performing the loudness gain operation on each third audio data frame in each third audio data segment, that is, the loudness gain operation is performed on each third audio data frame of each third audio data segment, the distortion caused by assigning the same gain coefficient to the entire audio data is avoided, and the frame-based automatic gain of the audio data is realized. Thus, in the general noise reduction of the audio signal processing, a high signal-to-noise ratio is realized, the loudness fidelity before and after the audio signal noise reduction is realized, and the quality of the processed audio signal is ensured.
[0084] Figure 8 is a schematic block diagram of an audio signal processing device provided by the embodiments of the present application. As shown in Figure 8 corresponding to the above audio signal processing method, the present application also provides an audio signal processing device 100. The audio signal processing device 100 includes units for executing the above-mentioned audio signal processing method, and the device can be configured in a desktop computer, a tablet computer, a laptop computer, and the like. Specifically, please refer to Figure 8 , the audio signal processing device 100 includes an acquisition module 101, a preprocessing module 102, a noise estimation module 103, a noise reduction processing module 104, and a loudness gain operation module 105, wherein:
[0085] The acquisition module 101 is configured to acquire first audio data, the first audio data includes a plurality of first audio data segments, and the first audio data segment includes a plurality of first audio data frames.
[0086] The preprocessing module 102 is configured to preprocess the plurality of first audio data segments to obtain a plurality of second audio data segments, and the second audio data segment includes a plurality of second audio data frames.
[0087] The noise estimation module 103 is configured to determine an estimated noise segment and a noise threshold based on a plurality of power spectrums corresponding to the plurality of first audio data segments and a silent condition.
[0088] The noise reduction processing module 104 is configured to calculate a plurality of third audio data segments based on the plurality of second audio data frames of each second audio data segment and the noise threshold, wherein each third audio segment comprises a plurality of third audio data frames.
[0089] The loudness gain operation module 105 is configured to perform a loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data.
[0090] In some embodiments, the preprocessing module 102 is configured to perform the preprocessing on the plurality of first audio data segments to obtain the plurality of second audio data segments, wherein each second audio segment comprises a plurality of second audio data frames, by performing Fourier transform on the plurality of first audio data segments and then performing frequency cutting.
[0091] In some embodiments, the noise estimation module 103 is configured to determine the estimated noise segment and the noise threshold based on the one-to-one corresponding plurality of power spectrums of the plurality of second audio data segments and the noiseless condition, by performing size sorting based on the one-to-one corresponding plurality of power spectrums of the plurality of first audio data segments, determining the estimated noise segment based on the sorted plurality of power spectrums and the noiseless condition, and calculating the noise threshold based on the estimated noise segment.
[0092] In some embodiments, the noise estimation module 103 is configured to calculate the noise threshold based on the estimated noise segment, by calculating a mean value and a variance of the estimated noise segment, and determining the sum of the mean value and the product of the variance and a balance constant as the noise threshold.
[0093] In some embodiments, the noise reduction processing module 104 is configured to perform the noise reduction processing on the plurality of second audio data segments based on the plurality of second audio data frames of each second audio data segment and the noise threshold to obtain the plurality of third audio data segments, wherein each third audio segment comprises a plurality of third audio data frames, by comparing each second audio data frame with the noise threshold to calculate a masking matrix, calculating a smoothing filter based on the plurality of second audio data segments, and processing the plurality of second audio data segments in combination with the masking matrix and the smoothing filter and then performing inverse Fourier transform to obtain the plurality of third audio segments.
[0094] In some embodiments, the loudness gain operation module 105 performs the loudness gain operation on each third audio data frame in each of the third audio data segments to obtain target audio data, including: calculating a difference value between a loudness value of each third audio data frame and a corresponding loudness value of each first audio data frame; calculating a loudness gain of each third audio data frame based on each difference value comparison; and calculating target audio data based on each loudness gain and each third audio data frame.
[0095] In some embodiments, the loudness gain operation module 105 performs the loudness gain operation on each third audio data frame in each of the third audio data segments to obtain target audio data, including: calculating a difference value between a loudness value of each third audio data frame and a corresponding loudness value of each first audio data frame; calculating a loudness gain of each third audio data frame based on each difference value comparison; and calculating target audio data based on each loudness gain and each third audio data frame.
[0096] It should be noted that the specific implementation process of the audio signal processing apparatus 100 and each unit can be clearly understood by those skilled in the art, which can be referred to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, it will not be repeated here.
[0097] The audio signal processing apparatus 100 described above can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 9 .
[0098] Please refer to Figure 9 , Figure 9 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 200 can be a terminal or a server, wherein the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, and other electronic devices with communication functions. The server can be a stand-alone server or a server cluster composed of multiple servers.
[0099] Please refer to Figure 9 , the computer device 200 includes a processor 202, a memory, and a network interface 205 connected through a system bus 201, wherein the memory can include a computer readable medium 203 storing non-volatile program code and an internal memory 204.
[0100] The computer readable medium 203 of non-volatile program code can store an operating system 2031 and a computer program 2032. The computer program 2032 includes program instructions which, when executed, can cause the processor 202 to perform an audio signal processing method.
[0101] The processor 202 is configured to provide computing and control capabilities to support the operation of the entire computer device 200.
[0102] The memory 204 provides an environment for the computer program 2032 in the computer readable medium 203 of non-volatile program code, which, when executed by the processor 202, can cause the processor 202 to perform an audio signal processing method.
[0103] The network interface 205 is configured to communicate with other devices via a network. Those skilled in the art can understand that the network interface 205 can be configured to communicate with other devices via a wired or wireless network. Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 200 to which the scheme of the present application is applied. The specific computer device 200 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0104] The processor 202 is configured to run the computer program 2032 stored in the memory to implement the following steps:
[0105] Obtain first audio data, the first audio data including a plurality of first audio data segments, the first audio data segment including a plurality of first audio data frames; preprocess the plurality of first audio data segments to obtain a plurality of second audio data segments, the second audio data segment including a plurality of second audio data frames; determine estimated noise segments and a noise threshold in the plurality of second audio data segments based on a plurality of power spectra and a noise condition corresponding to the plurality of second audio data segments; perform noise reduction processing based on the plurality of second audio data frames of each second audio data segment and the noise threshold to obtain a plurality of third audio data segments, the third audio data segment including a plurality of third audio data frames; and perform a loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data.
[0106] In some embodiments, the preprocessing of the plurality of first audio data segments to obtain a plurality of second audio data segments, the second audio data segment including a plurality of second audio data frames, includes performing Fourier transform on the plurality of first audio data segments and then performing frequency cutting to obtain the plurality of second audio data segments.
[0107] In some embodiments, the determining the estimated noise segment and the noise threshold based on the one-to-one corresponding power spectrums of the plurality of second audio data segments and the noiseless condition comprises: sorting the one-to-one corresponding power spectrums of the plurality of second audio data segments based on the size; determining the estimated noise segment based on the sorted power spectrums and the noiseless condition; and calculating the noise threshold based on the estimated noise segment.
[0108] In some embodiments, the noise reduction processing the plurality of second audio data segments based on the plurality of second audio data frames of each second audio data segment and the noise threshold to obtain a plurality of third audio data segments, each third audio data segment comprising a plurality of third audio data frames, comprises: comparing each second audio data frame with the noise threshold to obtain a masking matrix; calculating a smoothing filter based on the plurality of second audio data segments; and processing the plurality of second audio data segments based on the masking matrix and the smoothing filter and performing inverse Fourier transform to obtain the plurality of third audio data segments.
[0109] In some embodiments, the performing loudness gain operation on each third audio data frame in each third audio data segment to obtain target audio data comprises: calculating a difference between a loudness value of each third audio data frame and a loudness value of the corresponding first audio data frame; comparing each difference to calculate a corresponding loudness gain of each third audio data frame; and calculating target audio data based on each loudness gain and each third audio data frame.
[0110] In some embodiments, the calculating the noise threshold based on the estimated noise segment comprises: calculating a mean value and a variance of the estimated noise segment; and determining the noise threshold as a sum of the mean value and a product of the variance and a balance constant.
[0111] Further, the comparing each difference to calculate a corresponding loudness gain of each third audio data frame further comprises: calculating a minimum value and a mean value of the plurality of differences; if any difference is greater than the mean value, assigning the mean value as the loudness gain of the third audio data frame corresponding to the difference; and if any difference is less than the mean value, assigning the minimum value as the loudness gain of the third audio data frame corresponding to the difference.
[0112] It should be understood that, in the embodiments of the present application, the processor 202 can be a central processing unit (CPU), and the processor 202 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0113] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various computer readable storage media that can store program codes.
[0114] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0115] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0116] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the apparatus embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0117] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a storage medium. Based on such an understanding, the technical solutions of the present application essentially or say the part that contributes to the related art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application.
[0118] The above merely provides a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements shall be encompassed within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An audio signal processing method, characterized in that, include: Acquire first audio data, the first audio data including multiple first audio data segments, the first audio data segments including multiple first audio data frames; Preprocessing the plurality of first audio data segments yields a plurality of second audio data segments, each of which includes a plurality of second audio data frames; Based on the multiple power spectra and silent conditions corresponding to the multiple second audio data segments, the estimated noise segments and noise thresholds in the multiple second audio data segments are determined; Multiple third audio data segments are obtained based on multiple second audio data frames of each second audio data segment and the noise threshold denoising process, wherein the third audio data segments include multiple third audio data frames; The target audio data is obtained by performing a loudness gain operation on each third audio data frame in each of the third audio data segments.
2. The method according to claim 1, characterized in that, The preprocessing of the plurality of first audio data segments yields a plurality of second audio data segments, each second audio data segment comprising a plurality of second audio data frames, including: After performing Fourier transform on the plurality of first audio data segments, frequency segmentation is performed to obtain a plurality of second audio data segments.
3. The method according to claim 1, characterized in that, The step of determining the estimated noise segments and noise thresholds in the plurality of second audio data segments based on the power spectra and silence conditions corresponding to the plurality of second audio data segments includes: The power spectra corresponding to the multiple first audio data segments are sorted by size. The predicted noise segment is determined based on multiple power spectra and silent conditions in a sorted manner; The noise threshold is calculated based on the estimated noise segment.
4. The method according to claim 1, characterized in that, The process of obtaining multiple third audio data segments based on multiple second audio data frames of each second audio data segment and the noise threshold denoising process, wherein each third audio data segment includes multiple third audio data frames, including: The masking matrix is calculated by comparing each second audio data frame with the noise threshold. A smoothing filter is calculated based on the multiple second audio data segments; The multiple second audio data segments are processed by combining a masking matrix and a smoothing filter, and then subjected to an inverse Fourier transform to obtain the multiple third audio data segments.
5. The method according to claim 1, characterized in that, The step of performing a loudness gain operation on each third audio data frame in each of the third audio data segments to obtain the target audio data includes: Calculate the difference between the loudness value of each third audio data frame and the loudness value of each corresponding first audio data frame; The loudness gain of each third audio data frame is calculated based on each of the aforementioned difference comparisons. The target audio data is calculated based on each loudness gain and each third audio data frame.
6. The method according to claim 3, characterized in that, The step of calculating the noise threshold based on the estimated noise segment includes: Calculate the mean and variance of the noise prediction segment; The noise threshold is determined by the sum of the product of the average value, the variance, and the balance constant.
7. The method according to claim 5, characterized in that, The calculation of the loudness gain of each corresponding third audio data frame based on each difference comparison further includes: Calculate the minimum and average values of the multiple differences; If any difference is greater than the average value, the average value is assigned as the loudness gain of the third audio data frame corresponding to that difference; if any difference is less than the average value, the minimum value is assigned as the loudness gain of the third audio data frame corresponding to that difference.
8. An audio signal processing device, characterized in that, include: The acquisition module is used to acquire first audio data, the first audio data including multiple first audio data segments, and the first audio data segments including multiple first audio data frames; The preprocessing module is used to preprocess the plurality of first audio data segments to obtain a plurality of second audio data segments, wherein the second audio data segments include a plurality of second audio data frames; The noise prediction module is used to determine the predicted noise segment and noise threshold based on the multiple power spectra and silent conditions corresponding to the multiple first audio data segments; A noise reduction processing module is used to calculate multiple third audio data segments based on multiple second audio data frames of each second audio data segment and the noise threshold, wherein the third audio data segment includes multiple third audio data frames; as well as The loudness gain operation module is used to perform loudness gain operation on each third audio data frame in each of the third audio data segments to obtain the target audio data.
9. A computer device, characterized in that, include: A memory and a processor, the memory being used to store and support the processor in executing a program of any one of claims 1 to 7, the processor being configured to execute the program stored in the memory.
10. A computer-readable medium having processor-executable non-volatile program code, characterized in that, The program code causes the processor to execute the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Audio signal processing method and device, equipment and storage medium
CN110782911A
Apparatus, methods and computer programs for controlling noise reduction
CN113454716A