An audio noise reduction method, device, system and computer-readable storage medium
By performing power spectrum analysis and initialization of noise estimation parameters on the audio data, combined with minimum value tracking and signal-to-noise ratio calculation, the problem of poor noise reduction effect in the prior art is solved, and more efficient audio noise reduction processing is achieved.
Patent Information
- Application Number
- CN202210034896.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-12
AI Technical Summary
The prior art cannot effectively reduce noise during audio data processing, affecting audio quality and naturalness.
By obtaining the power spectrum of the data to be reduced, initializing the noise estimation parameters, performing minimum value tracking, calculating the posterior signal-to-noise ratio and prior signal-to-noise ratio, estimating the no-sound probability and the sound probability, updating the noise spectrum, calculating the gain estimate value, and performing noise reduction processing.
It improves the computing speed and efficiency of audio noise reduction, and enhances the noise reduction effect of noise reduction data.
Smart Images

Figure CN114495962B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and particularly relates to an audio noise reduction method, device, system, and computer-readable storage medium. Background Art
[0002] During the process of voice communication, audio data is affected by various interferences to varying degrees, which affects the quality and naturalness of the audio. Therefore, it is necessary to extract as pure original audio data as possible from the noisy audio data, that is, to perform audio noise reduction processing on the noisy audio data to achieve an anti-noise effect. However, in the existing process of processing audio data, a good audio noise reduction effect cannot be achieved. Summary of the Invention
[0003] The present application provides an audio noise reduction method, device, system, and computer-readable storage medium, which can improve the noise reduction effect on audio data.
[0004] To solve the above technical problems, the technical solution adopted by the present application is: to provide an audio noise reduction method, which includes: obtaining data to be noise-reduced, and calculating the power spectrum of the data to be noise-reduced, where the data to be noise-reduced includes noise data and noise-free data; initializing noise estimation parameters based on the power spectrum to obtain the noise spectrum of the noise data; performing minimum value tracking on the initialized noise estimation parameters in the first time period to obtain a first array; calculating the posterior signal-to-noise ratio and the prior signal-to-noise ratio of the data to be noise-reduced based on the first array; calculating an estimated value of the probability of no sound, an estimated value of the probability of sound, and an estimated value of the noise power spectrum based on the confidence level of the posterior signal-to-noise ratio; performing minimum value tracking on the initialized noise estimation parameters in the second time period to obtain a second array; calculating an estimated value of the gain of the noise-free data based on the second array, the estimated value of the probability of no sound, the estimated value of the probability of sound, and the estimated value of the noise power spectrum; and performing noise reduction processing on the data to be noise-reduced based on the estimated value of the gain, the noise spectrum, and the estimated value of the noise power spectrum to obtain noise-free data.
[0005] To solve the above technical problems, another technical solution adopted by the present application is: to provide an audio noise reduction device, which includes a memory and a processor connected to each other. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the audio noise reduction method in the above technical solution.
[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide an audio noise reduction device, which is used to perform a receiving task and a noise reduction task simultaneously. The audio noise reduction device includes a scheduling circuit and a noise reduction circuit. The scheduling circuit is used to receive the data to be noise-reduced corresponding to the receiving task, and perform a splitting process on the data to be noise-reduced to obtain multiple sub-audio data; the noise reduction circuit is connected to the scheduling circuit and is used to perform parallel noise reduction processing on all sub-audio data corresponding to the noise reduction task to obtain noise-reduced audio data; wherein, the noise reduction circuit is used to implement the audio noise reduction method in the above technical solution.
[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide an audio noise reduction system, which includes an audio acquisition device and an audio noise reduction device. The audio acquisition device is used to acquire the sound in the target scene to obtain the data to be noise-reduced; the audio noise reduction device is connected to the audio acquisition device and is used to perform noise reduction processing on the data to be noise-reduced to obtain noise-reduced audio data; wherein, the audio noise reduction device is the audio noise reduction device in the above technical solution.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, it is used to implement the audio noise reduction method in the above technical solution.
[0009] Through the above solution, the beneficial effect of this application is: the audio noise reduction method provided by this application can improve the operation speed of audio noise reduction by pipelining the steps of the audio noise reduction method, thereby improving the efficiency of audio noise reduction; at the same time, by performing multiple minimum value tracking operations on the initialized noise estimation parameters, the accuracy of the subsequent calculated gain estimation value can be improved, and further the noise reduction effect on the data to be noise-reduced can be improved. Description of the Drawings
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0011] Figure 1 is a schematic flowchart of the audio noise reduction method provided by this application;
[0012] Figure 2 is a schematic diagram of voice noise reduction provided by this application;
[0013] Figure 3 is a schematic diagram of the implementation process of FFT and IFFT provided by this application;
[0014] Figure 4 It is a schematic diagram of the butterfly operation in FFT and IFFT provided by this application;
[0015] Figure 5 It is a schematic diagram of processing the operation result of IFFT provided by this application;
[0016] Figure 6 It is a schematic diagram of the noise spectrum estimation principle provided by this application;
[0017] Figure 7 It is a schematic diagram of the gain calculation principle provided by this application;
[0018] Figure 8 It is a schematic flowchart of an embodiment of the audio noise reduction device provided by this application;
[0019] Figure 9 It is a schematic structural diagram of an embodiment of the audio noise reduction device provided by this application;
[0020] Figure 10 It is a schematic structural diagram of another embodiment of the audio noise reduction device provided by this application;
[0021] Figure 11 is Figure 10 A schematic connection diagram of the shunt module and the scheduling module in the illustrated embodiment;
[0022] Figure 12 It is a schematic structural diagram of an embodiment of the audio noise reduction system provided by this application;
[0023] Figure 13 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners
[0024] Next, with reference to the accompanying drawings and embodiments, the present application will be further described in detail. It should be specifically noted that the following embodiments are only used to illustrate the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only partial embodiments of the present application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0025] Referring to "embodiment" in the present application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0026] It should be noted that the terms "first", "second", and "third" in this application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0027] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the audio noise reduction method provided by this application. The audio noise reduction method includes:
[0028] Step 11: Obtain the data to be noise-reduced and calculate the power spectrum of the data to be noise-reduced.
[0029] The data to be noise-reduced includes noise data and noise-free data. Specifically, the power spectrum of the first half of the data to be noise-reduced (such as the first 128 data points) can be calculated first to obtain the first power spectrum, and then based on symmetry, the power spectrum of the second half of the data to be noise-reduced (such as the last 128 data points) can be calculated to obtain the second power spectrum; then the first power spectrum and the second power spectrum are combined to obtain the power spectrum. It can be understood that before calculating the power spectrum of the data to be noise-reduced, the data to be noise-reduced can be first subjected to a fast Fourier transform (Fast Fourier Transform, FFT), and then the power spectrum of the first half of the data obtained by the FFT transform is calculated first, and then the power spectrum of the second half of the data is calculated based on symmetry to obtain the power spectrum of the entire data, so that it is not necessary to calculate the power spectrum for all data, which can improve the operation efficiency.
[0030] Step 12: Initialize the noise estimation parameters based on the power spectrum to obtain the noise spectrum of the noise data.
[0031] After calculating the power spectrum of the data to be noise-reduced, the noise estimation parameters can be initialized based on the power spectrum to obtain the noise spectrum of the noise data, and the power spectrum is assigned to the noise spectrum. Specifically, the power spectrum can be smoothed first using the three-point mean filtering method to obtain the smoothed power spectrum, and then the noise estimation parameters are initialized based on the smoothed power spectrum to obtain the noise spectrum. It can be understood that the three-point mean filtering method is a conventional operation method in the technical field and will not be elaborated here.
[0032] Step 13: Track the minimum value of the initialized noise estimation parameters in the first time period to obtain the first array.
[0033] In a specific embodiment, the minimum value of the initialized noise estimation parameters may be calculated first to obtain the fourth array; then the minimum value of the fourth array is tracked to obtain the first array; by tracking the minimum value of the fourth array, the accuracy of the fourth array can be verified to determine whether the fourth array is the minimum value of the initialized noise estimation parameters. If the fourth array is not the minimum value of the initialized noise estimation parameters, the minimum value tracking can be performed to update the actual minimum value to the fourth array.
[0034] Step 14: Based on the first array, calculate the posterior signal-to-noise ratio and the prior signal-to-noise ratio of the data to be denoised.
[0035] The posterior signal-to-noise ratio can be calculated based on the first array, and the minimum value of the posterior signal-to-noise ratio is obtained to get the third array, and then the prior signal-to-noise ratio is calculated based on the third array.
[0036] Step 15: Based on the confidence level of the posterior signal-to-noise ratio, calculate the unvoiced probability estimate, the voiced probability estimate, and the noise power spectrum estimate.
[0037] Step 16: Track the minimum value of the initialized noise estimation parameters in the second time period to obtain the second array.
[0038] The second time period is after the first time period, that is, after the end of the first time period, the second time period starts. The minimum value tracking of the initialized noise estimation parameters is performed once in the first time period, and after the end of the first time period, the minimum value tracking of the initialized noise estimation parameters is performed again in the second time period; it can be understood that the minimum value tracking operation performed in step 16 is the same as the minimum value tracking operation in step 13 above and will not be elaborated here.
[0039] In a specific embodiment, the step of tracking the minimum value of the initialized noise estimation parameters can be performed two or more times. The data to be denoised may include multiple frames of audio data, and the minimum value tracking operation can be performed once at intervals of a preset number of frames of audio data. By performing the minimum value tracking multiple times, the accuracy of the subsequent calculated gain estimate can be improved, thereby improving the denoising effect on the data to be denoised.
[0040] Step 17: Based on the second array, the unvoiced probability estimate, the voiced probability estimate, and the noise power spectrum estimate, calculate the gain estimate of the noise-free data.
[0041] The unvoiced probability estimate is the probability estimate of the non-existence of the prior speech, and the voiced probability estimate is the probability estimate of the existence of the conditional speech.
[0042] Step 18: Denoise the data to be denoised based on the gain estimate value, the noise spectrum, and the noise power spectrum estimate value to obtain noise-free data.
[0043] In a specific embodiment, after obtaining the second array, the voiceless probability estimate value can be updated based on the second array to obtain an updated voiceless probability estimate value; then the noise spectrum can be updated based on the noise power spectrum estimate value to obtain an updated noise spectrum; thereby, based on the prior signal-to-noise ratio, the posterior signal-to-noise ratio, the updated voiceless probability estimate, and the voiced probability estimate, the gain estimate value is calculated; finally, based on the gain estimate value and the updated noise spectrum, the data to be denoised is denoised to obtain noise-free data.
[0044] Specifically, the step of updating the voiceless probability estimate value based on the second array to obtain an updated voiceless probability estimate value may include: calculating a first voiceless probability estimate value based on the second array; smoothing the first voiceless probability estimate value using a three-point mean filtering method to obtain a second voiceless probability estimate value; windowing the second voiceless probability estimate value to obtain a third voiceless probability estimate value; and performing a numerical transformation on the third voiceless probability estimate value to obtain an updated voiceless probability estimate value.
[0045] In this embodiment, by pipelining the steps of the audio denoising method, the operation speed of audio denoising can be improved, thereby improving the efficiency of audio denoising; at the same time, by performing multiple minimum value tracking operations on the initialized noise estimation parameters, the accuracy of calculating the gain estimate value in the subsequent stage can be improved, and further the denoising effect on the data to be denoised can be improved.
[0046] In a specific embodiment, the audio denoising method in the above embodiment can also be used to implement parallel processing of the data to be denoised to improve the denoising rate; specifically, the data to be denoised can be first split to obtain multiple sub-audio data; then the noise reduction processing method is used to perform parallel denoising on all sub-audio data to obtain denoised sub-audio data, and then the denoised sub-audio data is merged to obtain noise-free data; where the noise reduction processing method is the audio denoising method in the above embodiment.
[0047] Furthermore, the audio denoising method in the above embodiment can be applied to a Field Programmable Gate Array (FPGA) to implement parallel processing of the data to be denoised. The steps of the audio denoising method based on the FPGA platform are specifically introduced below:
[0048] First, as Figure 2As shown, the implementation process of voice noise reduction mainly includes the following parts: windowing in the time domain, FFT transformation, noise reduction operation (Log-Spectral Amplitude estimator, LSA), inverse Fourier transform (Inverse Fast Fourier Transform, IFFT), weight superposition, and storage, etc.; specifically, before performing the IFFT operation, the audio data (i.e., the data to be noise-reduced in the above embodiments) can be preprocessed, such as: removing low frequencies, removing high frequencies, or multiplying by a gain coefficient, etc.; after the IFFT operation, the audio time-domain data is restored by using weighted overlap-add, windowing, and removing calibration; among them, since the input audio data is frame-divided (1 frame is divided into 4 frames) before the FFT operation, it is necessary to perform weight superposition and other processing after the IFFT operation to convert these 4 frames of data into one frame to ensure that the number of input audio data and output audio data is the same. The storage of the audio data after weight superposition processing is similar to the storage of the audio data input before the FFT operation, and the front and back data are correlated.
[0049] Specifically, the audio data to be noise-reduced can be a noisy signal, which includes a noise signal and a clean signal. Audio noise reduction is achieved by extracting the clean signal. The steps of voice noise reduction can include: 1) performing frame division and windowing processing on the input audio data (i.e., the noisy signal); 2) performing FFT operation on each frame of the noisy signal; 3) first estimating the posterior signal-to-noise ratio, and then using the decision-directed method to estimate the prior signal-to-noise ratio, where the energy spectrum of the noise is estimated in non-speech segments (such as a few frames before the start of speech or speech gaps); 4) using the optimal MMSE-LSA estimator to estimate the intensity of the enhanced signal (equivalent to the step of calculating the gain below); 5) reconstructing the enhanced signal spectrum, and then performing IFFT operation on the enhanced signal spectrum to obtain the corresponding time-domain signal of the enhanced speech (i.e., the noise-reduced audio data).
[0050] (1) Since the human ear's perception of sound intensity is proportional to the logarithm of the spectral amplitude, assuming that the noise signal and the speech signal are uncorrelated, the noisy signal can be expressed as y = x + d, where y is the noisy speech, x is the clean speech, and d is the additive stationary noise.
[0051] First, Y k , X k and D k can be used to represent the k-th spectral component of the above y, x, and d after FFT operation respectively. Y k , B k can be calculated using the following formulas (1) and (2):
[0052]
[0053]
[0054] Among them, R in the above formula (1) and formula (2) k and B k are respectively the amplitudes of the noisy speech and the clean speech at frequency point k, and θ k and α k are respectively the phases of the noisy speech and the clean speech at frequency point k.
[0055] Then, use the following formulas (3) to (6) to estimate B from Y k : k
[0056]
[0057]
[0058]
[0059]
[0060] Among them, in the above formula (3), is the estimate of B k , and the above formula (4) can be derived from the above formula (3). In the above formula (6), G(ξ k , γ k ) is the gain function, ξ k and γ k are respectively the prior signal-to-noise ratio (Signal-Noise Ratio, SNR) and the posterior signal-to-noise ratio, v k =(ξ k / 1 + ξ k ) * γ k , ξ k =λ s (k) / λ n (k), γ k =R k 2 / λ n (k), λ s is the clean speech variance, and λ n is the noise variance; it can be seen from the above formula (5) that multiplying R k 2 by the gain function shown in formula (6) can obtain the clean speech estimate
[0061] (2) For the operations of FFT and IFFT, they can be implemented by transplanting the C source code and adopting the radix-2 frequency-domain / time-domain decimation method, such as Figure 3 As shown in the figure, the implementation process of FFT and IFFT operations under the FPGA platform is as follows:
[0062] It can be achieved through Figure 3 the Norml_ram (random access memory) shown in the figure stores the original audio data before the first butterfly operation. The Norml_ram can include a real part memory (not shown in the figure) and an imaginary part memory (not shown in the figure). The size of each real part / imaginary part memory can be 256 * 16 bits. Figure 3 The DATA RAM0 and DATA RAM1 shown in the figure can store the real part data and imaginary part data in the FFT or IFFT operations respectively. The size of DATA RAM0 and DATA RAM1 can be 256 * 32 bits.
[0063] When performing the first butterfly operation, the audio data can be directly input into the butterfly operation module without using a memory cache, which can save the operation time. Among them, the sizes of the real part data and imaginary part data of the input audio data are both 32 bits, and the sizes of the real part data and imaginary part data of the output audio data after FFT or IFFT processing are both 16 bits. When performing FFT or IFFT operations, the sine and cosine numbers in the FFT or IFFT operations can be obtained by looking up the FFT cosine / sine table or IFFT cosine / sine table, so as to adjust the data order of the butterfly operation. For example, the FFT cosine / sine table is FFTg_FFTCos or g_FFTReverse, and the sizes of FFTg_FFTCos and g_FFTReverse can be 512 * 16 bits and 256 * 16 bits respectively.
[0064] Furthermore, as Figure 4 shown in the figure, during the butterfly operation process of FFT, the numerical values of the audio data of the first 128 points and the last 128 points are calculated separately. The FFT values of the audio data of the first 128 points can be calculated first, and the audio data of the first 64 points can be calibrated and complex operation processed; then according to symmetry, the FFT values of the audio data of the last 128 points can be calculated; while during the butterfly operation process of IFFT, it is divided into two butterfly operations to directly calculate the audio data of 256 points, so as to obtain the IFFT values of the audio data of 256 points. Among them, the multiplication and addition operations during the FFT / IFFT operation process are both implemented by using a multiplier intellectual property (IP) core and an adder IP core. Since it involves signed number operations, the delays of the multiplier and adder used can be set to 2 clock cycles.
[0065] It can be understood that as Figure 5As shown, the data after IFFT operation has 256 points. The current operation result needs to be added to the previous operation result, and then the lower 64 bits of the data are output as the result of the current entire noise reduction algorithm. After outputting the operation result, the data from bits 64 - 255 are shifted to bits 0 - 191, and the upper 64 bits of the data are filled with 0 for the next operation.
[0066] (3) The implementation process of the noise reduction operation is actually to calculate the noise spectrum estimation and gain. Specifically, it can be based on, for example, Figure 6 and Figure 7 shown noise spectrum estimation principle and gain calculation principle to achieve noise reduction. Among them, |Y| 2 represents the energy of the speech signal (i.e., the data after FFT operation), λ d represents the noise spectrum, represents the noise spectrum estimation value, G represents the gain, and Y a 2 represents the power spectrum. Specifically, the steps for calculating the noise spectrum estimation value and the gain G are introduced below (i.e., the audio noise reduction method in the above embodiment):
[0067] 1) Calculate the power spectrum of the data of the first 128 points, that is, Y a 2 = Real 2 + Image 2 , and restore the data according to the scaling value in the FFT operation process to obtain the value of the power spectrum, and store it in the corresponding memory. The size of the memory can be 129 * 32bit; at the same time, perform three - point mean filtering on the power spectrum, and set the weights of the data of the previous point, the middle point, and the next point to 1 / 4, 1 / 2, and 1 / 4 respectively to achieve S f [i]=(Ya 2 [i - 1]>>2)+(Ya 2 [i]>>1)+(Ya 2 [i + 1]>>2), and then set the initial values of some intermediate array - type variables to the values after three - point mean filtering.
[0068] 2) Initialize the noise spectrum based on the power spectrum, that is, λ d = Y a 2 , and then initialize some intermediate array - type variables (i.e., noise estimation parameters) in the memory (Blockram) according to the initialized noise spectrum. For example: nShiftYa2(m_nShiftYa2) and eta(m_eta).
[0069] 3) Limited by the number of adders and multipliers, the minimum value of some array-type variables can be calculated first, and then the first minimum value search in the array-type variables is performed. Among them, the operation / assignment methods of the array-type variables in the first operation, the first 14 frames, and the operations / assignments after 14 frames can be different.
[0070] 4) Calculate the minimum value of the posterior SNR and calculate the prior SNR.
[0071] 5) According to the confidence level (i.e., value range) of the posterior SNR, calculate the probability estimate of the absence of prior speech, the probability estimate of the presence of conditional speech, and calculate the noise power spectrum estimate by recursive averaging.
[0072] 6) Perform the second minimum value search in the array-type variables. Except for the special processing of the data in the 10th frame, calculate the minimum value every 10 frames for the other data. Keep 5 values for each frequency point and then store them in 5 memories (Blockram) of 129 * 32 bits to obtain the minimum value of each frequency point.
[0073] 7) Calculate the probability estimate of the absence of prior speech. Update the noise spectrum according to the noise power spectrum estimate. Apply a local window (three-point mean filtering) and a global window (windowing and restoring the data according to the calibration) to the probability estimate respectively. Finally, through column transformation and other operations, obtain the probability estimate of the absence of prior speech.
[0074] 8) Update some intermediate array-type variables for calculating the noise spectrum, such as m_gamma & m_eta & m_v.
[0075] 9) Update the intermediate variable values for calculating the gain estimate, calculate the minimum gain estimate value, and thus calculate the gain G. Among them, during the calculation process, 5 temporary array variables are involved: ivUInt32m_min_temp
[129] , m_lambda_d_global
[129] , m_GH0
[129] , m_GH1
[129] , and ivUInt16m_PH1
[129] . The gain G and the temporary array variable m_PH1 can share a memory (Blockram), and the depth of the memory is 129 * 16 bits. The temporary array variable m_GH1 and the power spectrum Ya 2 Share a memory (Blockram).
[0076] 11) Finally, update the calculation of the noise spectrum estimate And the intermediate variable eta_2term of the gain G to complete the noise reduction process.
[0077] The audio noise reduction method based on the FPGA platform in this embodiment can achieve source code segmentation according to the coupling degree of functions and contexts by means of porting the source C, so as to reduce the coupling degree between each noise reduction and / or operation module, thereby realizing parallel operation of the modules, greatly improving the effect and efficiency of noise reduction processing; in addition, all operations such as multiplication-addition or division are executed in a pipeline manner, which can improve the operation efficiency; in addition, some memory and operation modules can be reused, which can save the resource occupancy of the platform, and a clock frequency above 100M can also be adopted in the platform to further improve the operation speed.
[0078] Please refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of an embodiment of the audio noise reduction device provided by the present application. The audio noise reduction device 80 includes a memory 81 and a processor 82 connected to each other. Among them, the memory 81 is used to store a computer program, and when the computer program is executed by the processor 82, it is used to implement the audio noise reduction method in the above embodiment.
[0079] Please refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an embodiment of the audio noise reduction device provided by the present application. The audio noise reduction device 10 includes: a scheduling circuit 11 and a noise reduction circuit 12.
[0080] The scheduling circuit 11 is used to receive the data to be noise-reduced corresponding to the received task, and perform splitting processing on the data to be noise-reduced to obtain multiple sub-audio data; the noise reduction circuit 12 is connected to the scheduling circuit 11, and is used to perform parallel noise reduction processing on all sub-audio data corresponding to the noise reduction task to obtain the noise-reduced data to be noise-reduced.
[0081] In a specific implementation manner, multiple sub-noise reduction circuits (not shown in the figure) can be set, and then the multiple sub-noise reduction circuits are used to process the multiple sub-audio data respectively to realize parallel noise reduction processing of the sub-audio data. Among them, the number of sub-noise reduction circuits can be set according to actual needs and is not limited here. It can be understood that the "noise reduction" in this embodiment is not only limited to the noise reduction processing of the data to be noise-reduced, but may also include operations such as separating or reverberation removal of the data to be noise-reduced.
[0082] Specifically, the sub-audio data can be the data to be noise-reduced in the data to be noise-reduced that needs to undergo noise reduction processing. The scheduling circuit 11 can split the received data to be noise-reduced, and at the same time, can identify the sub-audio data, and then divide the data to be noise-reduced into multiple sub-audio data; among them, "multiple paths" can be understood as "multiple channels", that is, each sub-audio data is transmitted in parallel to the noise reduction circuit 12 through multiple different transmission channels to improve the data transmission efficiency. It can be understood that the data to be noise-reduced can be divided into four sub-audio data or eight sub-audio data, etc. The number of channels of the sub-audio data can be increased or decreased according to actual needs, and no limitation is made here.
[0083] Furthermore, the receiving task and the noise reduction task can be executed simultaneously, that is, the audio noise reduction device 10 can perform real-time noise reduction processing on the received data to be noise-reduced. Among them, the data to be noise-reduced can be data containing speech that is collected in real time by an audio collection device (such as a pickup, etc.). In this case, the scheduling circuit 11 can directly transmit the collected sub-audio data to the noise reduction circuit 12 at the first time, so that the noise reduction circuit 12 can perform parallel noise reduction processing on all sub-audio data to obtain the noise-reduced data to be noise-reduced, without passing through the computer terminal to transfer the data to be noise-reduced, saving data transmission time, and greatly improving the real-time performance of the processing of the data to be noise-reduced. Through actual application verification, compared with the existing solution, the solution of this embodiment can save 4 - 8 ms of time; or, in other embodiments, the data to be noise-reduced can also be the previously collected data to be noise-reduced pre-stored in the audio collection device, and no limitation is made here.
[0084] In a specific embodiment, for each transmission channel, the scheduling circuit 11 can serially transmit the corresponding sub-audio data, and each time a preset number (such as 64 data points) of partial sub-audio data is transmitted to the noise reduction circuit 12 for noise reduction processing by the noise reduction circuit 12; then the next preset number of partial sub-audio data is transmitted to the noise reduction circuit 12, and so on, until the noise reduction processing of all sub-audio data is completed; through this method of transmitting and processing simultaneously, the processing pressure on the noise reduction circuit 12 can be reduced, and the audio noise reduction efficiency can be improved. Through actual application verification, the noise reduction time of the noise reduction circuit 12 for each sub-audio data is about 250 us, which is far lower than the conventional 2 ms - 6 ms noise reduction processing time in the related art.
[0085] In this embodiment, a scheduling circuit is adopted to execute the receiving task, that is, to receive the data to be denoised, and perform a splitting process on the data to be denoised to obtain multiple sub-audio data, so as to transmit the multiple sub-audio data to the noise reduction circuit in parallel; the noise reduction circuit is used to execute the noise reduction task, that is, to perform parallel noise reduction processing on all sub-audio data to obtain the data to be denoised after noise reduction, which can greatly save the time of data transmission and noise reduction processing, thereby improving the efficiency of noise reduction processing; moreover, the receiving task and the noise reduction task are executed simultaneously, and the data to be denoised collected can be received in the first time, and then the data to be denoised is transmitted to the noise reduction circuit without relaying the data to be denoised, which can greatly improve the real-time performance of the data to be denoised processing.
[0086] Please refer to Figure 10 , Figure 10 FIG. is a schematic structural diagram of another embodiment of the audio noise reduction device provided by the present application. The audio noise reduction device 20 includes: a scheduling circuit 21 and a noise reduction circuit 22.
[0087] The scheduling circuit 21 includes a splitting module 211 and a scheduling module 212. The splitting module 211 is used to perform a splitting process on the data to be denoised to obtain multiple pieces of data to be denoised; the scheduling module 212 is connected to the splitting module 211 and the noise reduction circuit 22, and is used to perform a scheduling process on the multiple pieces of data to be denoised to obtain multiple sub-audio data to the noise reduction circuit 22; specifically, the splitting module 211 can be a data selector (multiplexer, MUX).
[0088] In a specific implementation manner, the data to be denoised may include identification information and original audio information. The identification information includes a channel number identification and a noise reduction information identification. The channel number identification is used to identify the channel number of the data to be denoised, and each bit of the data to be denoised may correspond to a channel number. The noise reduction information identification is used to identify the noise reduction information of the data to be denoised, where the noise reduction information is used to indicate whether the data to be denoised needs to be subjected to noise reduction processing.
[0089] Specifically, the splitting module 211 can identify the channel number in the channel number identification and transmit the data to be denoised using the transmission channel corresponding to the channel number; the scheduling module 212 can also identify the noise reduction information in the noise reduction information identification and determine whether the data to be denoised is sub-audio data based on the noise reduction information, that is, determine whether the data to be denoised needs to be subjected to noise reduction processing. If the data to be denoised is sub-audio data, the sub-audio data is input to the noise reduction circuit 22; if the data to be denoised is not sub-audio data, the data to be denoised can be input to other circuits to implement other processing operations or directly output.
[0090] Such as Figure 11As shown, the audio noise reduction device 20 further includes a storage module 23, which is connected to the demultiplexing module 211 and is used to store multiple paths of data to be noise-reduced. Specifically, the storage module 23 may include multiple sub-storage modules 231, and each sub-storage module 231 corresponds to the channel number in the channel number identifier. The demultiplexing module 211 may also output the data to be noise-reduced to the corresponding sub-storage module 231 based on the channel number.
[0091] The noise reduction circuit 22 may also include multiple sub-noise reduction circuits 221( Figure 10 Taking three sub-noise reduction circuits 221 as an example), the data to be noise-reduced after noise reduction may include multiple sub-noise reduction audio data. Each sub-noise reduction circuit 221 is connected to the scheduling circuit 21. By performing noise reduction processing on the corresponding data to be processed and noise-reduced through each sub-noise reduction circuit 221, the corresponding sub-noise reduction audio data can be obtained. In other embodiments, the audio noise reduction device 20 may further include a demultiplexer (DEMUX) to output multiple sub-noise reduction audio data through the DEMUX.
[0092] In a specific embodiment, the audio noise reduction device 20 may be implemented based on an FPGA. After actual tests, the size of the audio noise reduction device 20 based on the FPGA can be controlled within 50mm×40mm, and the audio noise reduction device 20 can be integrated into an IP core. In different application scenarios, it can be flexibly transplanted, enabling the audio noise reduction device 20 to be widely applicable to any FPGA platform. Moreover, due to the advantages of low cost, large scale, high integration, low power consumption, high flexibility, and short development cycle of the FPGA, and its processing power consumption not exceeding 3W, the audio noise reduction device 20 based on the FPGA platform can save the noise reduction cost, improve the real-time performance and efficiency of data noise reduction.
[0093] Furthermore, in the audio noise reduction scheme based on other platforms (such as IMAX6Q or STM32, etc.), after the data to be noise-reduced is collected, it is necessary to wait until a frame of data to be noise-reduced is received and the data is updated before starting the noise reduction processing. Taking a frame of data to be noise-reduced containing 256 data points as an example, only 64 data points can be updated each time, so it is necessary to update four times before starting the noise reduction processing. However, in the audio noise reduction device 20 based on the FPGA in this embodiment, the collection of the data to be noise-reduced is implemented in the FPGA, and there is no need to wait until a frame of data to be noise-reduced is updated before starting the dimensionality reduction processing. The audio noise reduction processing can start when 1 / 4 frame of the data to be noise-reduced is received, thus greatly improving the efficiency of data transmission and noise reduction processing of the data to be noise-reduced.
[0094] Specifically, the noise reduction circuit 22 can be used to implement the audio noise reduction method in the above embodiments. Compared with other speech noise reduction algorithms (such as spectral subtraction, adaptive filtering, Wiener filtering, or minimum mean square error estimation (MMSE), etc.), using the audio noise reduction algorithm in the above embodiments for noise reduction processing can optimize the effect of speech noise reduction processing, achieve higher suppression of background noise, and make the distortion degree of the data to be noise-reduced lower.
[0095] In this embodiment, through the shunt module and the scheduling module, the shunting and scheduling of the data to be noise-reduced are realized. In different transmission channels, multiple paths of data to be noise-reduced that need to be noise-reduced are transmitted in parallel to multiple subsequent sub-noise reduction circuits. By setting multiple sub-noise reduction circuits, parallel noise reduction processing of multiple paths of sub-audio data is realized, thereby greatly improving the efficiency of audio noise reduction. Moreover, the audio noise reduction device in this embodiment can be implemented on the FPGA platform, and at the same time, the parallelism and rapidity of data processing on the FPGA platform are used to support the parallel processing of the multi-channel audio noise reduction method, which can ensure the processing speed while improving the noise reduction effect, and can solve the problems of long time consumption and slow response of audio noise reduction, and further improve the efficiency of audio noise reduction.
[0096] Please refer to Figure 12 , Figure 12 FIG. is a schematic structural diagram of an embodiment of an audio noise reduction system provided by the present application. The audio noise reduction system 120 includes an audio acquisition device 121 and an audio noise reduction device 122. The audio acquisition device 121 is used to acquire the sound in the target scene to obtain the data to be noise-reduced. The audio noise reduction device 122 is connected to the audio acquisition device 121 and is used to perform noise reduction processing on the data to be noise-reduced to obtain the noise-reduced audio data. Among them, the audio noise reduction device 122 is the audio noise reduction device in the above embodiment.
[0097] Please refer to Figure 13 , Figure 13 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 130 is used to store a computer program 131. When the computer program 131 is executed by a processor, it is used to implement the audio noise reduction method in the above embodiment.
[0098] The computer-readable storage medium 130 can be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.
[0099] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0100] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0101] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0102] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. An audio noise reduction method, characterized in that, the method includes: obtaining the data to be noise-reduced, and calculating the power spectrum of the data to be noise-reduced, where the data to be noise-reduced includes noise data and noise-free data; wherein, the data to be noise-reduced contains multiple frames of audio data; initializing the noise estimation parameters based on the power spectrum to assign a noise spectrum based on the power spectrum, and obtaining the noise spectrum of the noise data; wherein, the noise estimation parameters are intermediate array-type variables; tracking the minimum values of several data regarding a preset number of frames of audio data in the initialized noise estimation parameters respectively in a first time period, and obtaining a first array; calculating the posterior signal-to-noise ratio of the data to be noise-reduced based on the first array; calculating an estimated value of the probability of no sound, an estimated value of the probability of sound, and an estimated value of the noise power spectrum based on the confidence level of the posterior signal-to-noise ratio; tracking the minimum values of several data regarding a preset number of frames of audio data in the initialized noise estimation parameters respectively in a second time period, and obtaining a second array; calculating an estimated value of the gain of the noise-free data based on the second array, the estimated value of the probability of no sound, the estimated value of the probability of sound, and the estimated value of the noise power spectrum; performing noise reduction processing on the data to be noise-reduced based on the estimated value of the gain, the noise spectrum, and the estimated value of the noise power spectrum, and obtaining the noise-free data.
2. The audio noise reduction method according to claim 1, characterized in that, after the step of calculating the posterior signal-to-noise ratio of the data to be noise-reduced based on the first array, the method further includes: obtaining the minimum value of the posterior signal-to-noise ratio, and obtaining a third array; calculating the prior signal-to-noise ratio based on the third array.
3. The audio noise reduction method according to claim 2, characterized in that, the method further includes: updating the estimated value of the probability of no sound based on the second array, and obtaining an updated estimated value of the probability of no sound; updating the noise spectrum based on the estimated value of the noise power spectrum, and obtaining an updated noise spectrum; calculating the estimated value of the gain based on the prior signal-to-noise ratio, the posterior signal-to-noise ratio, the updated estimated value of the probability of no sound, and the estimated value of the probability of sound; performing noise reduction processing on the data to be noise-reduced based on the estimated value of the gain and the updated noise spectrum, and obtaining the noise-free data.
4. The audio noise reduction method according to claim 1, characterized in that, the step of initializing the noise estimation parameters based on the power spectrum to obtain the noise spectrum of the noise data includes: performing smoothing processing on the power spectrum by using a three-point mean filtering method, and obtaining a smoothed power spectrum; initializing the noise estimation parameters based on the smoothed power spectrum, and obtaining the noise spectrum.
5. The audio noise reduction method according to claim 1, characterized in that, the step of calculating the power spectrum of the data to be noise-reduced includes: calculating the power spectrum of the first half of the data to be noise-reduced, and obtaining a first power spectrum; calculating the power spectrum of the second half of the data to be noise-reduced based on symmetry, and obtaining a second power spectrum; Merge the first power spectrum and the second power spectrum to obtain the power spectrum.
6. The audio noise reduction method according to claim 1, wherein, the step of respectively tracking the minimum values of several numerical values of the initialized noise estimation parameters for a preset number of frames of audio data in the first time period includes: Calculating the minimum value of the initialized noise estimation parameters to obtain a fourth array; Tracking the minimum value of the fourth array to obtain the first array.
7. The audio noise reduction method according to claim 3, wherein, the method includes: The step of updating the voice absence probability estimate value based on the second array to obtain an updated voice absence probability estimate value includes: Calculating a first voice absence probability estimate value based on the second array; Smoothing the first voice absence probability estimate value by using a three-point mean filtering method to obtain a second voice absence probability estimate value; Performing windowing processing on the second voice absence probability estimate value to obtain a third voice absence probability estimate value; Performing numerical transformation processing on the third voice absence probability estimate value to obtain an updated voice absence probability estimate value.
8. The audio noise reduction method according to claim 1, wherein, the method further includes: Performing branch processing on the data to be noise-reduced to obtain multiple sub-audio data; Performing parallel noise reduction processing on all the sub-audio data by using a noise reduction processing method to obtain noise-reduced sub-audio data, and the noise reduction processing method is the audio noise reduction method according to any one of claims 1-7; Merging the noise-reduced sub-audio data to obtain the noise-free data.
9. An audio noise reduction device, wherein, it includes a memory and a processor connected to each other. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the audio noise reduction method according to any one of claims 1-8.
10. An audio noise reduction device, wherein, used to simultaneously execute a reception task and a noise reduction task, and the audio noise reduction device includes: A scheduling circuit, configured to receive the data to be noise-reduced corresponding to the reception task, and perform branch processing on the data to be noise-reduced to obtain multiple sub-audio data; A noise reduction circuit, connected to the scheduling circuit, configured to perform parallel noise reduction processing on all the sub-audio data corresponding to the noise reduction task to obtain noise-reduced audio data; wherein, the noise reduction circuit is used to implement the audio noise reduction method according to any one of claims 1-8.
11. The audio noise reduction device according to claim 10, wherein, the noise-reduced audio data includes multiple sub-noise-reduced audio data; the noise reduction circuit includes multiple sub-noise reduction circuits, and each sub-noise reduction circuit is connected to the scheduling circuit and is configured to perform noise reduction processing on the corresponding sub-audio data to obtain the sub-noise-reduced audio data.
12. The audio noise reduction device according to claim 10, wherein, the scheduling circuit includes: A branch module, configured to perform branch processing on the data to be noise-reduced to obtain multiple paths of data to be noise-reduced; A scheduling module, connected to the branching module and the noise reduction circuit, for scheduling and processing the multi-channel data to be noise-reduced to obtain the multi-channel sub-audio data to the noise reduction circuit.
13. The audio noise reduction device according to claim 12, wherein, the data to be noise-reduced includes identification information and original audio information, and the identification information includes a channel number identifier and a noise reduction information identifier; the branching module is further configured to identify the channel number in the channel number identifier and transmit the data to be noise-reduced using a transmission channel corresponding to the channel number.
14. The audio noise reduction device according to claim 13, wherein, the scheduling module is further configured to identify the noise reduction information in the noise reduction information identifier and determine whether the data to be noise-reduced is the sub-audio data based on the noise reduction information; if so, input the sub-audio data into the noise reduction circuit.
15. The audio noise reduction device according to claim 13, wherein, the audio noise reduction device further includes a storage module, connected to the branching module, for storing the multi-channel data to be noise-reduced.
16. The audio noise reduction device according to claim 15, wherein, the storage module includes a plurality of sub-storage modules, each sub-storage module corresponding to the channel number in the channel number identifier, and the branching module is further configured to output the multi-channel data to be noise-reduced to the corresponding sub-storage module based on the channel number.
17. An audio noise reduction system, wherein, the audio noise reduction system includes: an audio collection device, for collecting sounds in a target scene to obtain data to be noise-reduced; an audio noise reduction device, connected to the audio collection device, for performing noise reduction processing on the data to be noise-reduced to obtain noise-reduced audio data; wherein, the audio noise reduction device is the audio noise reduction device according to any one of claims 10-16 above.
18. A computer-readable storage medium, for storing a computer program, wherein, when the computer program is executed by a processor, it is used to implement the audio noise reduction method according to any one of claims 1-8.
Citation Information
Patent Citations
Noise suppression method and device for quickly calculating voice existence probability, storage medium and terminal
CN111899752A
Audio signal processing method and device and storage medium
CN111968662A