Audio signal processing method, device, equipment and computer-readable storage medium

Through the methods of band decomposition and gain calculation, the problem of high complexity of speech processing algorithms is solved, efficient speech signal processing is achieved, and the computing burden and energy consumption of mobile devices are reduced.

CN114822569BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110081032.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-21
Publication Date
2025-07-25
Estimated Expiration
2041-01-21

AI Technical Summary

Technical Problem

When existing voice processing technologies increase the voice bandwidth, the algorithm complexity increases, resulting in excessive CPU consumption of mobile devices and increased power consumption, affecting system stability, especially when applying deep neural networks.

Method used

The band decomposition method is used to decompose the audio signal into low-band signals and high-band signals, and the gain of the low-band signals is used to calculate the gain of the high-band signals, and the corresponding processing is carried out to reduce the calculation amount and improve the signal processing efficiency.

Benefits of technology

By reducing the complexity of signal processing algorithms, voice processing efficiency is improved, computing burden and energy consumption of mobile devices are reduced, while maintaining the denoising effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114822569B_ABST
    Figure CN114822569B_ABST
Patent Text Reader

Abstract

The present application provides an audio signal processing method, apparatus, device, and computer-readable storage medium. The method includes: obtaining an audio signal to be processed; performing frequency band decomposition on the audio signal to obtain a first frequency band signal and a second frequency band signal, where the frequency of the first frequency band signal is lower than that of the second frequency band signal; determining a first signal gain corresponding to the first frequency band signal, and determining a second signal gain corresponding to the second frequency band signal based on the first signal gain; determining the processed first frequency band signal based on the first signal gain and the first frequency band signal, and determining the processed second frequency band signal based on the second signal gain and the second frequency band signal; performing frequency band synthesis on the processed first frequency band signal and the processed second frequency band signal to obtain a processed audio signal. Through the present application, the speech processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to signal processing technologies, and in particular to an audio signal processing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] In a voice communication system, such as cellular communication or Voice over Internet Protocol (VoIP) communication, voice signals have been upgraded from narrowband signals with a bandwidth of about 4 kilohertz (kHz) to broadband (high-definition) signals with a bandwidth of about 8 kHz, and are currently gradually upgraded to ultra-wideband (ultra-high-definition) signals with a bandwidth above 10 kHz, thereby improving the voice fidelity of calls. The bandwidth of typical ultra-high-definition voice signals is 12 kHz, 16 kHz, 24 kHz, etc. However, on the one hand, while the voice bandwidth is increased, the complexity of various voice processing algorithms also increases. Coupled with the application of deep neural network models in voice denoising and other processing, the complexity further increases. Excessive complexity may lead to excessive consumption of the central processing unit (CPU) of a mobile device, increased power consumption, and even affect system stability, such as increasing the jitter phenomenon during a voice call. Summary of the Invention

[0003] Embodiments of the present application provide an audio signal processing method, apparatus, and computer-readable storage medium, which can improve the voice processing efficiency.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] Embodiments of the present application provide a method for processing an audio signal, including:

[0006] Obtaining an audio signal to be processed;

[0007] Performing band decomposition on the audio signal to obtain a first band signal and a second band signal, where the frequency of the first band signal is lower than that of the second band signal;

[0008] Determining a first signal gain corresponding to the first band signal, and determining a second signal gain corresponding to the second band signal based on the first signal gain;

[0009] Determining a processed first band signal based on the first signal gain and the first band signal, and determining a processed second band signal based on the second signal gain and the second band signal;

[0010] Performing band synthesis on the processed first band signal and the processed second band signal to obtain a processed audio signal.

[0011] An embodiment of the present application provides an audio signal processing device, including:

[0012] A first acquisition module, configured to acquire an audio signal to be processed;

[0013] A frequency band decomposition module, configured to perform frequency band decomposition on the audio signal to obtain a first frequency band signal and a second frequency band signal, where the frequency of the first frequency band signal is lower than that of the second frequency band signal;

[0014] A first determination module, configured to determine a first signal gain corresponding to the first frequency band signal, and determine a second signal gain corresponding to the second frequency band signal based on the first signal gain;

[0015] A second determination module, configured to determine a processed first frequency band signal based on the first signal gain and the first frequency band signal, and determine a processed second frequency band signal based on the second signal gain and the second frequency band signal;

[0016] A frequency band synthesis module, configured to perform frequency band synthesis on the processed first frequency band signal and the processed second frequency band signal to obtain a processed audio signal.

[0017] In some embodiments, the first determination module is further configured to:

[0018] Determine a first sub-signal gain corresponding to a first echo cancellation module in a first signal processing link, where the first signal processing link at least includes the first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module for processing the first frequency band signal;

[0019] Determine a second sub-signal gain corresponding to the first noise suppression module, determine a third sub-signal gain corresponding to the first howling control module, and determine a fourth sub-signal gain corresponding to the first gain control module.

[0020] In some embodiments, the first determination module is further configured to:

[0021] Acquire the first frequency band signal input to the first noise suppression module;

[0022] Perform time-frequency conversion on the first frequency band signal to obtain spectral data of the first frequency band signal;

[0023] Input the spectral data into a statistical model to obtain a statistical model gain; input the spectral data into a trained neural network model to obtain a network model gain;

[0024] Determine the second sub-signal gain corresponding to the first noise suppression module based on the statistical model gain and the network model gain.

[0025] In some embodiments, the first determination module is further configured to:

[0026] Determine the smaller value of the statistical model gain and the network model gain as the second sub-signal gain corresponding to the first noise suppression module; or,

[0027] Obtain a first weight corresponding to the statistical model gain and a second weight corresponding to the network model gain;

[0028] Perform a weighted sum of the statistical model gain and the network model gain using the first weight and the second weight to obtain the second sub-signal gain corresponding to the first noise suppression module.

[0029] In some embodiments, the first determination module is further configured to:

[0030] Obtain a first prediction probability that there is speech in the first band signal by the statistical model;

[0031] Obtain a second prediction probability that there is speech in the first band signal by the trained neural network model;

[0032] Determine the first prediction probability as the first weight and the second prediction probability as the second weight; or;

[0033] Obtain a preset first weight and a preset second weight.

[0034] In some embodiments, the first determination module is further configured to:

[0035] Determine a fifth sub-signal gain corresponding to the second band signal based on the first sub-signal gain;

[0036] Determine a sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain;

[0037] Determine a seventh sub-signal gain corresponding to the second band signal based on the third sub-signal gain;

[0038] Determine an eighth sub-signal gain corresponding to the second band signal based on the fourth sub-signal gain.

[0039] Determine the second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain.

[0040] In some embodiments, the first determination module is further configured to:

[0041] Determine the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain as the second signal gain corresponding to the second band signal;

[0042] Correspondingly, this second determination module is further configured to:

[0043] Determine the product of the second band signal and the second signal gain as the processed second band signal.

[0044] In some embodiments, this first determination module is further configured to:

[0045] Determine the fifth sub-signal gain as the signal gain of the second echo cancellation module in the second signal processing link; the second signal processing link at least includes a second echo cancellation module, a second noise suppression module, a second howling control module, and a second gain control module for processing the second band signal;

[0046] Determine the sixth sub-signal gain as the signal gain of the second noise suppression module;

[0047] Determine the seventh sub-signal gain as the signal gain of the second howling control module;

[0048] Determine the eighth sub-signal gain as the signal gain of the second gain control module.

[0049] In some embodiments, this second determination module is further configured to:

[0050] Obtain a first output signal obtained by the second echo cancellation module based on the second band signal and the fifth sub-signal gain;

[0051] Obtain a second output signal obtained by the second noise suppression module based on the first output signal and the sixth sub-signal gain;

[0052] Obtain a third output signal obtained by the second howling control module based on the second output signal and the seventh sub-signal gain;

[0053] Obtain the processed second band signal obtained by the second gain control module based on the third output signal and the eighth sub-signal gain.

[0054] In some embodiments, the second sub-signal gain is a gain vector including K gain values, and this determination module is further configured to:

[0055] Determine the top P target frequency points with the highest frequencies from the K frequency points corresponding to the K gain values;

[0056] Determine the gain values corresponding to the P target frequency points as P target gain values;

[0057] Determine the minimum value among the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0058] An embodiment of the present application provides an audio signal processing device, including:

[0059] A memory for storing executable instructions;

[0060] A processor, when executing the executable instructions stored in the memory, implements the method provided by the embodiment of the present application.

[0061] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to implement the method provided by the embodiment of the present application when executed.

[0062] The embodiments of the present application have the following beneficial effects:

[0063] After obtaining the audio signal to be processed, first decompose the audio signal into a first band signal and a second band signal. The frequency of the first band signal is lower than that of the second band signal, that is, the first band signal is a low-band signal and the second band signal is a high-band signal. Then determine the first signal gain corresponding to the first band signal, and determine the second signal gain corresponding to the second band signal based on the first signal gain. Next, determine the processed first band signal based on the first signal gain and the first band signal, and determine the processed second band signal based on the second signal gain and the second band signal; finally, perform band synthesis on the processed first band signal and the processed second band signal to obtain the processed audio signal; in this way, the gain of the high-frequency second band signal is deduced from the first signal gain of the low-frequency first band signal, so the algorithm complexity of signal processing can be reduced, thereby improving the signal processing efficiency. Description of the Drawings

[0064] Figure 1A It is a schematic diagram of the framework structure of a speech denoising system in the related art;

[0065] Figure 1B It is a schematic diagram of the framework structure of a neural network denoising in the related art;

[0066] Figure 2 It is a schematic diagram of the network architecture of the voice call system provided by the embodiment of the present application;

[0067] Figure 3 It is a schematic diagram of the structure of the second terminal provided by the embodiment of the present application;

[0068] Figure 4 Schematic diagram of an implementation process of the audio signal processing method provided by an embodiment of the present application;

[0069] Figure 5 Schematic diagram of an implementation process of determining a first signal gain and a second signal gain provided by an embodiment of the present application;

[0070] Figure 6 Another schematic diagram of an implementation process of the audio signal processing method provided by an embodiment of the present application;

[0071] Figure 7 Schematic diagram of the implementation framework structure of the combined denoising algorithm provided by an embodiment of the present application;

[0072] Figure 8 Schematic diagram of the implementation framework structure of the audio signal processing method proposed by an embodiment of the present application;

[0073] Figure 9A Schematic diagram of band decomposition based on quadrature mirror filtering provided by an embodiment of the present application;

[0074] Figure 9B Spectrum response schematic diagram provided by an embodiment of the present application;

[0075] Figure 10 Schematic diagram of band synthesis based on quadrature mirror filtering provided by the present application;

[0076] Figure 11 Schematic diagram of the structure of a signal processing system in the related art;

[0077] Figure 12 Schematic diagram of the structure of a signal processing system provided by an embodiment of the present application;

[0078] Figure 13 Another schematic diagram of the structure of a signal processing system provided by an embodiment of the present application. Detailed implementation manners

[0079] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0080] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0081] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0083] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0084] 1) Fourier transform is a method for analyzing signals, which can analyze the components of signals and also synthesize signals using these components. Many waveforms can be used as the components of signals, such as sine waves, square waves, sawtooth waves, etc. The Fourier transform uses sine waves as the components of signals.

[0085] 2) Short-time Fourier transform is a mathematical transform related to the Fourier transform, used to determine the frequency and phase of the sine waves in the local region of a time-varying signal. The idea is to select a window function for time-frequency localization. It is assumed that the analysis window function g(t) is stationary (pseudo-stationary) within a short time interval. The window function is moved so that f(t)g(t) is a stationary signal within different finite time widths, and thus the power spectrum at each different moment is calculated.

[0086] 3) Echo cancellation technology uses the echo cancellation method, that is, it estimates the magnitude of the echo signal through an adaptive method and then subtracts this estimated value from the received signal to cancel the echo.

[0087] 4) Noise suppression, also known as speech enhancement, refers to the technology of extracting useful speech signals from the noise background and suppressing and reducing noise interference when the speech signal is interfered with or even submerged by various noises. That is, to extract as pure an original speech as possible from the noisy speech.

[0088] 5) Howling control, also known as howling suppression. The generation of howling belongs to positive feedback. The sound of the audio is picked up by the microphone again, resulting in self-oscillation and causing howling. Howling not only affects hearing but also burns out audio equipment. Howling suppression is also the technology of suppressing and reducing howling from audio data.

[0089] To better understand the speech signal processing method provided by the embodiments of the present application, first, the speech signal processing methods and their existing drawbacks in the related art are described.

[0090] Currently, broadband (HD) voice calls over mobile networks and VoIP have become popular and are transitioning to ultra-broadband (UHD) voice. The frequency range of broadband voice is approximately 0 - 8 kHz, which can cover most of the energy of human speech, and the voice quality has been greatly improved compared to the narrowband voice quality of early fixed telephones. With the increase in network bandwidth, the voice bandwidth is gradually being increased to ultra-broadband, such as typically including 12 kHz, 16 kHz, etc., and even full-band (the highest frequency that the human ear can hear is approximately 20 kHz, and an audio signal with a bandwidth close to or exceeding this frequency is a full-band signal), such as the commonly used 24 kHz.

[0091] Figure 1A It is a schematic diagram of the framework structure of a voice denoising system in related technologies, as Figure 1A shown. This framework structure includes: a time-domain to frequency-domain conversion module 001A, a statistical model 002A, and a frequency-domain to time-domain conversion module 003A, where:

[0092] The time-domain - frequency-domain conversion module 001A is used to transform each frame of the broadband voice signal (usually 5 - 20 milliseconds of voice as one frame) from the time domain to the frequency domain using the short-time Fourier transform (STFT, Short-Time Fourier Transform), and then obtain the spectrum of the voice frame. For example, for a voice frame signal x(n) of N voice samples, n = 1, 2,..., N, perform time-frequency conversion to obtain the spectrum X(k) of N / 2 + 1 frequency points.

[0093] The statistical model gain calculation module 002A is used to calculate the spectrum gain G1(k) from the spectrum X(k) through a traditional statistical model, where k is each frequency point on the spectrum, k = 1, 2,..., N / 2 + 1. When performing frequency-domain conversion on a frame of signal, the larger N is, the higher the frequency resolution. Typically, the values of N are 256, 512, etc., and the corresponding number of frequency points are 129, 257, etc.

[0094] The frequency-domain - time-domain conversion module 003A is used to multiply the gain G1(k) obtained by the statistical model 002A with the spectrum X(k) to obtain the denoised spectrum Xout1(k), that is, Xout1(k) = X(k) * G1(k). Then, perform the inverse short-time Fourier transform (ISTFT, Inverse Short Time Fourier Transform) to reconstruct the denoised voice frame signal.

[0095] In the above voice denoising process, for each voice frame of the wideband voice signal, the voice and noise are estimated spectrally, and according to the relative intensities of the voice spectrum and the noise spectrum in the most recent time period, for example, according to the a priori signal-to-noise ratio or the a posteriori signal-to-noise ratio or some combination of the two, the noise component is suppressed as much as possible from the noisy voice spectrum and the voice component is retained. For example, for the part with a higher signal-to-noise ratio at a certain frequency point, a larger gain is applied at that frequency point, and the part with a lower signal-to-noise ratio indicates that it is more likely to contain only noise, so a smaller gain is applied for suppression. The methods for noise estimation include but are not limited to the minimum tracking method, the minima controlled recursive averaging (MCRA) method, etc., and the methods for voice estimation include but are not limited to the likelihood ratio factor (LRF) method, the optimized modified log spectral amplitude (OMSLA) method, etc.

[0096] Traditional voice denoising algorithms usually utilize the stationarity of noise when estimating noise, that is, only the relatively stationary signal segments are considered as noise and then noise estimation is carried out. This also means that for non-stationary signals with rapid changes, such as keyboard sounds, knocking sounds, etc., traditional statistics-based algorithms tend to regard non-stationary noise as voice signals and thus cannot suppress them well.

[0097] Due to the development of deep learning technology, deep neural networks can better learn the characteristics of voice and noise, and thus can better distinguish voice and noise, including non-stationary noise, so as to better suppress noise.

[0098] Figure 1B It is a schematic diagram of a neural network denoising framework structure in the related technology, as Figure 1B shown. This framework includes: a time-domain to frequency-domain conversion module 001B, a voice feature extraction module 002B, a neural network model 003B, and a frequency-domain to time-domain conversion module 004B, where:

[0099] The time-domain - frequency-domain conversion module 001B is used to transform each frame of the wideband voice signal (usually 5 - 20 milliseconds of voice as one frame) from the time domain to the frequency domain by using the short-time Fourier transform (STFT), and then obtain the spectrum of the voice frame. For example, for the voice frame signal x(n) of N voice samples, n = 1, 2,..., N, time-frequency conversion is performed to obtain the spectrum X(k) of N / 2 + 1 frequency points.

[0100] The voice feature extraction module 002B is used to calculate the required voice feature vector from the spectrum, and this feature vector serves as the input of the neural network.

[0101] Common voice features include spectral amplitude value vectors, spectral logarithmic energy value vectors, Mel Frequency Cepstrum Coefficient (MFCC) vectors, Fbanks vectors, Bark Frequency Cepstrum Coefficient (BFCC) vectors, pitch periods, etc., as well as the first-order or second-order differences of certain of these feature vectors in the time domain to reflect the dynamic change characteristics of the features over time. Finally, the feature vectors input into the neural network model may be one or a combination of multiple of the above-mentioned feature vectors.

[0102] Neural network model 003B is used to calculate the spectral gain G2(k) using voice feature vectors, where k is each frequency point on the spectrum, k = 1, 2, …, N / 2 + 1. Therefore, G2 represents a total of N / 2 + 1 gains. When performing frequency-domain conversion on a frame of signal, the larger N is, the higher the frequency resolution. Typical values of N are 256, 512, etc., and the corresponding number of frequency points are 129, 257, etc.

[0103] Frequency-domain to time-domain conversion module 004B is used to multiply the gain G2(k) obtained by the neural network model with the spectrum X(k) to obtain the denoised spectrum Xout2(k) = X(k) * G2(k), and then inverse Fourier transform the denoised spectrum back to the time-domain voice frame.

[0104] The selected neural network model 003B can be a forward fully connected deep neural network (DNN, Deep Neural Networks), or a certain recurrent neural network (RNN, Recurrent Neural Network), such as LSTM, GRU, etc., or a convolutional neural network (CNN, Convolutional Neural Networks), or a combined form of these networks. For example, some network layers are fully connected layers, some layers are RNN network layers, and some layers are CNN layers. The deep neural network includes an input layer, intermediate hidden layers, and an output layer. The number of neurons in the input layer is generally the same as the length of the input feature vector. For example, if the input feature vector includes 129 spectral logarithmic energy values and a pitch period value, that is, a total of 130 values, then the neural network input layer has 130 neurons. The number of layers of the intermediate hidden layer and the number of neurons in each layer are determined according to the scale of the training data and the computing resources. If less computing resources are required, fewer layers and fewer neurons are used. If the scale of the training data is large, a larger network scale may achieve better results, which needs to be considered comprehensively. The number of neurons in the output layer is generally related to the number of gains to be calculated. For example, if the gain of each frequency point needs to be calculated here, then the output gain G2(k), k = 1, 2,..., N / 2 + 1, and the number of neurons in the output layer is N / 2 + 1. In other implementation schemes, the number of neurons in the output layer can also be less than N / 2 + 1. For example, if the N / 2 + 1 frequency points are divided into different frequency subbands, each neuron in the output layer only needs to predict the gain of each subband.

[0105] Since the current speech signal has been upgraded from a narrowband signal with a bandwidth of about 4 kHz to a wideband (high-definition) signal with a bandwidth of about 8 kHz, and is currently gradually upgraded to an ultra-wideband (ultra-high definition) signal with a bandwidth above 10 kHz, while the bandwidth of the speech is increased, the complexity of various speech processing algorithms also increases accordingly. Coupled with the application of deep neural network models in speech denoising and other processing, the complexity further increases.

[0106] In another speech signal processing method in the related art, the ultra-wideband signal or full-band signal collected by the microphone is subjected to sub-band processing, and then speech enhancement, echo cancellation, and encoding are performed according to the sub-bands. However, in this implementation method, the two sub-bands are processed independently, and the computational complexity is still relatively high.

[0107] Based on the above problems, the embodiments of the present application provide an audio signal processing method, which uses a neural network for speech denoising in the low frequency band and performs denoising by referring to the results of the low frequency band in the high frequency band to obtain the final ultra-high definition signal, which can not only ensure the denoising effect, but also reduce the computational amount, thereby improving the signal processing efficiency.

[0108] The following describes an exemplary application of the audio signal processing device provided in the embodiments of the present application. The device provided in the embodiments of the present application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), etc. Below, the exemplary application when the device is implemented as a terminal will be described.

[0109] Refer to Figure 2 , Figure 2 which is a schematic diagram of the network architecture of the voice call system 100 provided in the embodiments of the present application. As Figure 2 shown, the voice call system 100 includes a first terminal 200, a network 300, and a second terminal 400, where the first terminal 200 and the second terminal 400 are connected through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. The network 300 can also be a cellular communication network.

[0110] An application program capable of making voice calls can be installed on the first terminal 200 and the second terminal 400. For example, it can be an instant messaging application program to make voice calls through this instant messaging program. Of course, the first terminal 200 and the second terminal 400 can only have cellular network communication functions and make voice calls with each other by dialing. In the Figure 2 shown network architecture, taking the voice call between the first terminal 200 and the second terminal 400 through the network 300 as an example, assuming that the second terminal 400 can implement the audio signal processing method provided in the embodiments of the present application. First, the second terminal 400 collects an audio signal through a voice input device (e.g., a microphone), then performs frequency division processing on the audio signal to obtain a first band signal (low-frequency band signal) and a second band signal (high-frequency band signal), processes the first band signal through each module of the first band signal processing link, and determines a first gain corresponding to the first band signal. Then, according to the first gain, a second gain corresponding to the second band signal is determined. The first band signal is processed using the first gain to obtain a processed first band signal, and the second band signal is processed using the second gain to obtain a processed second band signal. Then, the first band signal and the second band signal are subjected to band synthesis to obtain a processed audio signal. Finally, the processed audio signal is encoded and sent. In this way, only the gain of the low-frequency band signal is calculated, and the gain of the high-frequency band is calculated using the gain of the low-frequency band signal, which can reduce the calculation amount and improve the signal processing efficiency while ensuring the signal processing effect.

[0111] Refer to Figure 3 , Figure 3 which is a schematic diagram of the structure of the second terminal 400 provided in the embodiments of the present application.Figure 3 The second terminal 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the second terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to implement connection communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 the various buses are all labeled as the bus system 440.

[0112] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0113] The user interface 430 includes one or more output devices 431 capable of presenting media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as keyboards, mice, microphones, touch screen displays, cameras, other input buttons, and controls.

[0114] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.

[0115] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0116] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of these data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0117] An operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0118] A network communication module 452 for reaching other computing devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;

[0119] A presentation module 453 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.);

[0120] An input processing module 454 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432.

[0121] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 3 An audio signal processing device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a first acquisition module 4551, a frequency band decomposition module 4552, a first determination module 4553, a second determination module 4554, and a frequency band synthesis module 4555. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented.

[0122] The functions of each module will be described below.

[0123] In other embodiments, the device provided by the embodiments of the present application can be implemented in hardware. As an example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio signal processing method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more Application Specific Integrated Circuits (ASICs), DSPs, Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), or other electronic components.

[0124] The audio signal processing method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the terminals provided by the embodiments of the present application.

[0125] See Figure 4 , Figure 4 which is a schematic diagram of an implementation process of the audio signal processing method provided by an embodiment of the present application, and will be described in combination with the steps shown in Figure 4 .

[0126] Step S101: Obtain the audio signal to be processed.

[0127] Here, in implementation, it may be to obtain the audio signal to be processed collected by the audio acquisition device of the terminal. In some embodiments, the audio acquisition device may be a microphone. The audio signal to be processed may be an audio signal with a bandwidth exceeding a preset value. For example, it may be an ultra-high-definition audio signal with a bandwidth exceeding 10 kHz.

[0128] In some embodiments, after obtaining the audio signal to be processed, it is first necessary to divide the audio signal to be processed into multiple audio frames. For example, it can be divided into multiple 25-ms audio frames. The following processing of the audio signal to be processed is for each audio frame.

[0129] Step S102: Perform band decomposition on the audio signal to obtain a first band signal and a second band signal.

[0130] Here, the frequency of the first band signal is lower than the frequency of the second band signal. That is to say, the first band signal is a low-band signal, and the second band signal is a high-band signal.

[0131] In implementation, step S102 can use methods such as discrete Fourier transform, wavelet decomposition, or filter bank for band decomposition. Since the audio signal to be processed is decomposed into a first band signal and a second band signal. For example, an ultra-clear signal with a bandwidth of 16 kHz is decomposed into two 8-kHz bandwidth signals, a low frequency and a high frequency. This frequency division method is also called two-way frequency division. At this time, an efficient band decomposition method commonly used is quadrature mirror filter (QMF) decomposition, which uses an anti-aliasing low-pass filter H0(z) and a high-pass filter H1(z) to achieve band decomposition.

[0132] The cut-off frequency of H0(z) is around π / 2 (that is, half of the bandwidth frequency). H0(z) and H1(z) are symmetric with each other at the orthogonal frequency π / 2 in the frequency spectrum. The signal after passing through H0(z) and H1(z) is V b (z)=H b (z)*X(z), b = 0, 1, and X(z) is the frequency spectrum of the audio signal x(n) to be processed. The signal after downsampling is Therefore, when b = 0, U0(z) is the spectrum of the first band signal (i.e., the high-definition voice signal) output. When k = 1, U1(z) is the spectrum of the second band signal output.

[0133] Step S103: Determine the first signal gain corresponding to the first band signal, and determine the second signal gain corresponding to the second band signal based on the first signal gain.

[0134] The first signal processing link in the terminal processes the first band signal. The first signal processing link at least includes a first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module to perform processing such as echo cancellation, noise suppression, howling control, and gain control on the first band signal. Each of the above modules outputs a sub-signal gain. In the embodiments of the present application, the sub-signal gains output by each module constitute the first signal gain corresponding to the first band signal.

[0135] Since the second band signal is a high-frequency band signal and the high-frequency band part has a small impact on the auditory quality, in the embodiments of the present application, in order to reduce the calculation amount and improve the signal processing efficiency, the second signal gain corresponding to the first band signal is determined by inference based on the first signal gain.

[0136] When determining the second signal gain corresponding to the second band signal based on the first signal gain, the sub-signal gains corresponding to the second band signal can be determined based on the sub-signal gains in the first signal gain, so as to determine the second signal gain corresponding to the second band signal based on the sub-signal gains corresponding to the second band signal.

[0137] Step S104: Determine the processed first band signal based on the first signal gain and the first band signal, and determine the processed second band signal based on the second signal gain and the second band signal.

[0138] Here, when step S104 is implemented, the first echo cancellation module, the first noise suppression module, the first howling control module, and the first gain control module in the first signal link process the first band signal based on the original input first band signal and the sub-signal gains corresponding to each module. For example, when the first band signal is input into the first echo cancellation module, the first echo cancellation module performs time-frequency conversion on the first band signal to obtain the spectrum of the first band signal, then obtains the first sub-signal gain corresponding to the first echo cancellation module, multiplies the first sub-signal gain by the spectrum, and performs frequency-time conversion on the product result to obtain the fourth output signal of the first echo cancellation module. The fourth output signal is still a time-domain signal. The fourth output signal is input into the first noise suppression module. The first noise suppression module performs time-frequency conversion on the fourth output signal to obtain the spectrum of the fourth output signal, then inputs the spectrum of the fourth output signal into the statistical model and the neural network model, comprehensively calculates to obtain the second sub-signal gain, multiplies the second sub-signal gain by the spectrum of the first output signal, and performs frequency-time conversion on the obtained product result to obtain the fifth output signal of the first noise suppression module.

[0139] When determining the processed second band signal based on the second signal gain and the second band signal, it can be to directly multiply the second band signal and the second signal gain in the time domain to obtain the processed second band signal, thereby improving the signal processing efficiency.

[0140] Step S105, perform band synthesis on the processed first band signal and the processed second band signal to obtain the processed audio signal.

[0141] Here, band synthesis is the inverse process of band decomposition in step S102. When implemented, first perform time-frequency conversion on the processed first band signal and the processed second band signal to obtain the spectrum of the processed first band signal and the spectrum of the processed second band signal. Then, perform upsampling on the spectrum of the processed first band signal and the spectrum of the processed second band signal respectively. The spectrum after sampling is V' b (z) = U' b (z 2 )), b = 1, 2. Here, U'0(z) is the processed first band signal, and U'1(z) is the processed second band signal. Then input U'0(z) into the filter F0(z), and input U'1(z) into the filter F1(z). Among them, F0(z) = H1(-z) = H0(z), F1(z) = -H1(z) = -H0(-z). The filter coefficients can be set in advance. The finally output signal spectrum is X'(z) = F0(z)V'0(z) + F1(z)V'1(z). The time-domain representation of X'(z) is the processed audio signal.

[0142] In the audio signal processing method provided by the embodiment of the present application, after obtaining the audio signal to be processed, the audio signal is first decomposed into frequency bands to obtain a first frequency band signal and a second frequency band signal. The frequency of the first frequency band signal is lower than that of the second frequency band signal, that is, the first frequency band signal is a low-frequency band signal and the second frequency band signal is a high-frequency band signal. Then, the first signal gain corresponding to the first frequency band signal is determined, and the second signal gain corresponding to the second frequency band signal is determined based on the first signal gain. Next, the processed first frequency band signal is determined based on the first signal gain and the first frequency band signal, and the processed second frequency band signal is determined based on the second signal gain and the second frequency band signal. Finally, the processed first frequency band signal and the processed second frequency band signal are synthesized in frequency bands to obtain the processed audio signal. In this way, the gain of the high-frequency second frequency band signal is deduced from the first signal gain of the low-frequency first frequency band signal, so that the algorithm complexity of signal processing can be reduced, thereby improving the signal processing efficiency.

[0143] In some embodiments, Figure 4 "Determining the first signal gain corresponding to the first frequency band signal" in step S103 shown can be implemented through steps S1031 to S1032 shown as follows: Figure 5 shown as follows:

[0144] Step S1031, determining the first sub-signal gain corresponding to the first echo cancellation module in the first signal processing link.

[0145] Here, the first signal processing link at least includes a first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module for processing the first frequency band signal. In the embodiment of the present application, the output signal of the first echo cancellation module may be input to the first noise suppression module, the output signal of the first noise suppression module may be input to the first howling control module, and the output signal of the first howling control module may be input to the first gain control module. Of course, in actual implementation, the first echo cancellation module, the first noise suppression module, the first howling control module, and the first gain module may also have other input-output sequences.

[0146] Step S1032, determining the second sub-signal gain corresponding to the first noise suppression module, determining the third sub-signal gain corresponding to the first howling control module, and determining the fourth sub-signal gain corresponding to the first gain control module.

[0147] Here, the second sub-signal gain, the third sub-signal gain, and the fourth sub-signal gain are the signal gains output by the first noise suppression module, the first howling control module, and the first gain control module.

[0148] In some embodiments, "determining the second sub-signal gain corresponding to the first noise suppression module" in step S1032 may be implemented through the following steps:

[0149] Step S321, obtain a first band signal input to the first noise suppression module.

[0150] Here, if the first echo cancellation module has processed the first band signal obtained by frequency division before the first noise suppression module, then the first band signal input to the first noise suppression module is also the signal output by the first echo cancellation module. If there is no other processing module before the first noise suppression module, then what is input to the first noise suppression module is the first band signal after frequency division.

[0151] Step S322, perform time-frequency conversion on the first band signal to obtain spectral data of the first band signal.

[0152] Here, it may be to perform a Fourier transform on the first band signal, for example, it may be STFT, so as to convert the continuous first band signal into a discrete frequency-domain signal, realize the time-domain to frequency-domain conversion of the first band signal, and obtain the spectral data of the first band.

[0153] For example, perform time-frequency conversion on the first band signal x(n) of N speech samples, where n = 1, 2,..., N, to obtain the spectrum X(k) of N / 2 + 1 frequency points, where k = 1, 2,..., N / 2 + 1. When performing frequency-domain conversion on a frame of signal, the larger N is, the higher the frequency resolution. N can take values such as 256 and 512, and the corresponding number of frequency points k is 129 and 257.

[0154] Step S323, input the spectral data into a statistical model to obtain a statistical model gain, and input the spectral data into a trained neural network model to obtain a network model gain.

[0155] Here, the spectral data is used to calculate the statistical model gain G1(k) through a statistical model, where k is each frequency point on the spectrum, that is, the statistical model gain is a gain vector. When calculating the gain, the statistical model can suppress the noise component as much as possible and retain the speech component from the noisy speech spectrum according to the prior signal-to-noise ratio or the posterior signal-to-noise ratio or a certain combination of the two. For example, at a certain frequency point where the signal-to-noise ratio is higher, a larger gain is applied at that frequency point, and at a lower signal-to-noise ratio, it indicates that there is likely only noise, so a smaller gain is applied for suppression.

[0156] The trained neural network model can be a deep learning neural network model or a convolutional neural network model. When determining the network model gain, the trained neural network model first extracts the feature vectors of the spectral data of the first band signal, and then calculates the gain based on the feature vectors to obtain the network model gain, which is also a gain vector.

[0157] Step S324: Determine the second sub-signal gain corresponding to the first noise suppression module based on the statistical model gain and the network model gain.

[0158] In step S324, gain fusion is performed based on the statistical model gain and the network model gain to determine the second sub-signal gain corresponding to the first noise suppression module. In actual implementation, step S324 can have at least the following two implementation methods:

[0159] The first implementation method: Determine the smaller value of the statistical model gain and the network model gain as the second sub-signal gain corresponding to the first noise suppression module.

[0160] The second implementation method: Perform weighted summation on the statistical model gain and the network model gain to determine the second sub-signal gain. This implementation method can be achieved through the following steps:

[0161] Step S3241: Obtain the first weight corresponding to the statistical model gain and the second weight corresponding to the network model gain.

[0162] When implementing step S3241, there are at least the following two implementation methods:

[0163] Method A: Obtain a preset first weight and a preset second weight. In this method, the first weight and the second weight can be weights preset according to the empirical probability of the presence of speech.

[0164] Method B: First, obtain the first prediction probability of the statistical model for the presence of speech in the first band signal; and obtain the second prediction probability of the trained neural network model for the presence of speech in the first band signal; determine the first prediction probability as the first weight and the second prediction probability as the second weight.

[0165] Step S3242: Use the first weight and the second weight to perform weighted summation on the statistical model gain and the network model gain to obtain the second sub-signal gain corresponding to the first noise suppression module.

[0166] The second sub-signal gain corresponding to the first noise suppression module is calculated by the first method. The calculation method is simple and efficient. The second sub-signal gain corresponding to the first noise suppression module is calculated by the second method, which combines the statistical model gain and the network model gain, and has a higher accuracy.

[0167] When calculating the statistical model gain using the statistical model and estimating the noise, the stationarity of the noise is usually utilized, that is, only relatively stationary signal segments are considered as noise, and then noise estimation is performed. This also means that for non-stationary signals with rapid changes, such as keyboard sounds and tapping sounds, this statistical model tends to regard non-stationary noise as speech signals, and thus cannot suppress them well. Due to the development of deep learning technology, deep neural networks can better learn the characteristics of speech and noise, and thus can better distinguish speech and noise, including non-stationary noise, so as to better suppress the noise. Therefore, in the embodiments of the present application, when determining the second sub-signal gain corresponding to the first noise suppression module, the advantages of the statistical model and the deep learning algorithm are comprehensively utilized, that is, the neural network is more effective in suppressing non-stationary noise, and the statistical model has low risk, small computational complexity, and better prediction of stationary noise. Combining the statistical model and the neural network model can better realize low-risk and high-efficiency product applications.

[0168] In some embodiments, Figure 4 "Determining the second signal gain corresponding to the second band signal based on the first signal gain" in step S103 shown can be achieved by Figure 5 the following steps from step S1033 to step S1037 shown:

[0169] Step S1033, determining the fifth sub-signal gain corresponding to the second band signal based on the first sub-signal gain.

[0170] Here, if the first sub-signal gain is a gain value rather than a gain vector, when implementing step S1033, the first sub-signal gain can be directly determined as the fifth sub-signal gain; if the first sub-signal gain is a gain vector, the fifth sub-signal gain can be determined based on the gain values of the top P frequency points with the highest gain in the gain vector. For example, the smallest gain value among the top P frequency points can be determined as the fifth sub-signal gain, or the gain values of the top P frequency points can be averaged to obtain the fifth sub-signal gain.

[0171] Step S1034, determining the sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain.

[0172] Here, the second sub-signal gain is a gain vector including K gain values. When implementing step S1034, it may be: First, determine the top P target frequency points with the highest values from the K frequency points corresponding to the K gain values; then determine the gain values corresponding to the P target frequency points as P target gain values; and determine the minimum value among the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0173] In some embodiments, it may also be to determine the average value of the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0174] Step S1035, determine the seventh sub-signal gain corresponding to the second band signal based on the third sub-signal gain.

[0175] Step S1036, determine the eighth sub-signal gain corresponding to the second band signal based on the fourth sub-signal gain.

[0176] The implementation manners of step S1035 and step S1036 are similar to the implementation manner of step S1031, and the implementation process can refer to step S1031.

[0177] Step S1037, determine the second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain.

[0178] Here, when implementing step S1037, there are two implementation manners based on the module structure of the signal processing link in the terminal:

[0179] The first manner: When the signal processing link module includes a first signal processing link for processing the first band signal and a second signal processing link for processing the second band signal, corresponding to the first signal processing link, the second signal processing link includes a second echo cancellation module, a second noise suppression module, a second howling control module, and a second gain control module. At this time, step S1037 can be implemented through the following steps:

[0180] Step S371A, determine the fifth sub-signal gain as the signal gain of the second echo cancellation module in the second signal processing link.

[0181] Step S372A, determine the sixth sub-signal gain as the signal gain of the second noise suppression module.

[0182] Step S373A, determine the seventh sub-signal gain as the signal gain of the second howling control module.

[0183] Step S374A, determine the eighth sub-signal gain as the signal gain of the second gain control module.

[0184] Correspondingly, Figure 4 "Determining the processed second band signal based on the second signal gain and the second band signal" in step S104 shown can be implemented through the following steps:

[0185] Step S1041, obtaining a first output signal obtained by the second echo cancellation module based on the second band signal and the fifth sub-signal gain.

[0186] Step S1042, obtaining a second output signal obtained by the second noise suppression module based on the first output signal and the sixth sub-signal gain.

[0187] Step S1043, obtaining a third output signal obtained by the second howling control module based on the second output signal and the seventh sub-signal gain;

[0188] Step S1044, obtaining the processed second band signal obtained by the second gain control module based on the third output signal and the eighth sub-signal gain.

[0189] In the above steps S371A to S374A and in S1041 to S1044, since the second signal processing link includes a second echo cancellation module, a second noise suppression module, a second howling control module, and a second gain control module, the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain are determined as the signal gains of each processing model, and then each module adjusts the gain of the second band signal based on the corresponding sub-signal gain to obtain the processed second band signal, without each module calculating the corresponding sub-signal gain by itself, which can reduce the calculation amount and improve the signal processing efficiency.

[0190] The second method: When the signal processing link module only includes a first signal processing link for processing the first band signal, step S1037 can be implemented through the following steps:

[0191] Step S371B, determining the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain as the second signal gain corresponding to the second band signal.

[0192] In the embodiments of the present application, the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain refer to the linear gain for the time-domain signal. Therefore, in step S371B, the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain is determined as the second signal gain.

[0193] Correspondingly, Figure 4"Determining the processed second band signal based on the second signal gain and the second band signal" in step S104 shown can be implemented as: determining the product of the second band signal and the second signal gain as the processed second band signal.

[0194] Compared with the first implementation, it is not necessary for each module in the second signal processing link to process the second band signal. In this implementation, the second band signal (time domain signal) obtained by frequency division is directly multiplied by the second signal gain to obtain the processed second band signal, which can further reduce the amount of calculation.

[0195] Based on the foregoing embodiments, the embodiments of the present application further provide an audio signal processing method, which is applied to Figure 2 the network architecture shown. Figure 6 Another schematic diagram of the implementation process of the audio signal processing method provided by the embodiments of the present application is shown in Figure 6 As shown, this process includes:

[0196] Step S601, the second terminal collects the audio signal to be processed through the voice input device.

[0197] Here, a call connection is established between the second terminal and the first terminal. This call connection can be established through an instant messaging application or through a dialing program. Through this call connection, users of the first terminal and the second terminal can make voice or video calls. Assume that the second terminal supports the audio data processing method provided by the embodiments of the present application. The second terminal collects the audio signal through the voice input device (microphone). This audio signal can include the voice signal emitted by the user and can also include some other noise signals.

[0198] Step S602, the second terminal performs band decomposition on the audio signal to obtain a first band signal and a second band signal.

[0199] Here, the highest frequency of the first band signal is lower than the lowest frequency of the second band signal, that is, the first band signal is a low-frequency band signal, for example, an audio signal of 0-8 kHz; the second band signal is a high-frequency band signal, for example, an audio signal of 8-16 kHz. Generally, people are more sensitive to low-frequency band signals and less sensitive to high-frequency band signals.

[0200] Step S603, the second terminal determines the first sub-signal gain corresponding to the first echo cancellation module in the first signal processing link.

[0201] Here, the first signal processing link at least includes a first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module for processing the first band signal.

[0202] Step S604, the second terminal determines the second sub-signal gain corresponding to the first noise suppression module, determines the third sub-signal gain corresponding to the first howling control module, and determines the fourth sub-signal gain corresponding to the first gain control module.

[0203] Step S605, the second terminal determines the fifth sub-signal gain corresponding to the second band signal based on the first sub-signal gain.

[0204] When the first sub-signal gain is a single gain value rather than a gain vector, in implementing step S1033, it may be directly determining the first sub-signal gain as the fifth sub-signal gain; when the first sub-signal gain is a gain vector, it may be determining the fifth sub-signal gain based on the gain values of the top P frequency points with the highest gain in the gain vector. For example, determining the minimum gain value among the top P frequency points as the fifth sub-signal gain, or it may be averaging the gain values of the top P frequency points to obtain the fifth sub-signal gain.

[0205] Step S606, the second terminal determines the sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain.

[0206] The second sub-signal gain is a gain vector including K gain values. In implementing step S1034, it may be: first, determining the top P target frequency points with the highest gain from the K frequency points corresponding to the K gain values; then determining the gain values corresponding to the P target frequency points as P target gain values; and determining the minimum value among the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0207] In some embodiments, it may also be determining the average value of the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0208] Step S607, the second terminal determines the seventh sub-signal gain corresponding to the second band signal based on the third sub-signal gain.

[0209] Step S608, the second terminal determines the eighth sub-signal gain corresponding to the second band signal based on the fourth sub-signal gain.

[0210] Step S609, the second terminal determines the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain as the second signal gain corresponding to the second band signal.

[0211] Step S610, the second terminal determines the processed first band signal based on the first signal gain and the first band signal, and determines the processed second band signal based on the second signal gain and the second band signal.

[0212] Step S611: The second terminal performs band synthesis on the processed first band signal and the processed second band signal to obtain a processed audio signal.

[0213] Step S612: The second terminal encodes the processed audio signal to obtain an encoded audio signal.

[0214] Since there is data redundancy in the audio signal, an encoder is used for encoding when sending the audio - video signal. When implementing step S612, an encoder for ultra - high - definition voice can be used to encode the processed audio signal, and the encoding method can be Advanced Audio Coding (AAC) encoding, etc.

[0215] Step S613: The second terminal sends the encoded audio signal to the first terminal.

[0216] Step S614: The first terminal decodes the encoded audio signal to obtain a decoded audio signal.

[0217] Here, the first terminal decodes the encoded audio signal using a decoding method corresponding to the encoding method to restore the audio signal.

[0218] Step S615: The first terminal outputs the decoded audio signal using its own audio output device.

[0219] In the audio signal processing method provided in the embodiments of the present application, after obtaining the audio signal to be processed, first, the audio signal is decomposed into a first band signal and a second band signal. The frequency of the first band signal is lower than that of the second band signal, that is, the first band signal is a low - frequency band signal and the second band signal is a high - frequency band signal. Then, the first sub - gain corresponding to the first echo cancellation module, the second sub - gain corresponding to the first noise suppression module, the third sub - gain corresponding to the first howling control module, and the fourth sub - gain corresponding to the first gain control module in the first signal processing link for processing the first band signal are determined. Then, the second signal gain corresponding to the second band signal is calculated based on the first sub - gain, the second sub - gain, the third sub - gain, and the fourth sub - gain, which can reduce the calculation amount and thus improve the signal processing efficiency. After that, the first band signal and the second band signal are respectively processed using the first signal gain and the second signal gain to obtain processed sub - band signals, and then the processed sub - band signals are subjected to band synthesis to obtain a processed ultra - high - definition voice signal. Finally, the processed ultra - high - definition voice signal is encoded and sent.

[0220] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0221] In order to comprehensively utilize the advantages of statistical algorithms and deep learning algorithms, that is, neural networks are more effective in suppressing non-stationary noise, and statistical algorithms have low risk, small computational complexity, and better prediction of stationary noise. Therefore, in the embodiments of the present application, the statistical model and the neural network model are combined to better achieve low-risk and high-efficiency product applications.

[0222] Figure 7 It is a schematic structural diagram of the implementation framework of the combined denoising algorithm provided by the embodiments of the present application. As Figure 7 shown, this framework includes: a time-domain to frequency-domain conversion module 701, a statistical model gain calculation module 702, a speech feature extraction module 703, a neural network gain calculation module 704, a gain fusion module 705, and a frequency-domain to time-domain conversion module 706. The combined denoising method will be described below in conjunction with each module.

[0223] After the audio frame is converted into the spectral signal X(k) through the time-domain to frequency-domain conversion module, X(k) is respectively input into the statistical model gain calculation module 702 and the speech feature extraction module 703. The speech feature vector extracted by the speech feature extraction module is input into the neural network gain calculation module 704, and finally the spectral gains G1(k) and G2(k) are obtained respectively. These two gain vectors are sent to the gain fusion module 705 for gain fusion to obtain the final spectral gain G3(k). G3(k) is multiplied by the spectrum X(k) to obtain the final spectral output Xout3(k) after denoising.

[0224] Among them, when the gain fusion module 705 calculates G3(k) based on G1(k) and G2(k), it can use G3(k)=min(G1(k), G2(k)), that is, take the smaller value of the two to obtain G3(k); it can also perform weighted summation of the two gains according to the detection results of the probability of the presence of speech in the signal by the statistical model or the neural network model: G3(k)=a*G1(k)+b*G2(k), where a and b are factor parameters preset according to the speech presence probability.

[0225] In order to enable the above combined denoising method to not only support the denoising of high-definition speech, but also support the denoising of ultra-high-definition speech, one way is to directly use the ultra-high-definition signal as the input of both the statistical model and the neural network model, which also means that the computational complexity may increase significantly.

[0226] In order to reduce the computational complexity and be able to continue using the previous high-definition speech denoising algorithm, the implementation framework structure of the audio signal processing method proposed by the embodiments of the present application is as Figure 8As shown in the figure, it includes: a frequency band decomposition module 801, a low-frequency band noise suppression module 802, a high-frequency band noise suppression module 803, and a frequency band synthesis module 804. The audio signal processing method provided by the embodiments of the present application will be described below in combination with each module.

[0227] The frequency band decomposition module 801 decomposes the input ultra-high-definition voice signal through frequency band decomposition. Further, it is decomposed into a low-frequency band signal and a high-frequency band signal. Among them, the low-frequency band is the frequency band where the previous high-definition voice is located. For example, if the frequency range (bandwidth) of the ultra-high-definition voice includes 0 to 12 kHz, then the low-frequency band is 0 to 8 kHz, and the high-frequency band is 8 to 12 kHz; if the frequency range of the ultra-high-definition voice includes 0 to 16 kHz, then the low-frequency band is 0 to 8 kHz, and the high-frequency band is 8 to 16 kHz.

[0228] The low-frequency signal enters the low-frequency band noise suppression module 802, and the low-frequency voice signal after noise suppression is output; the high-frequency signal enters the high-frequency band noise suppression module 803, and the high-frequency voice signal after noise suppression is output; then the high-frequency and low-frequency signals after noise suppression enter the frequency band synthesis module 804 at the same time, and are re-synthesized into the ultra-high-definition voice frame after denoising.

[0229] Figure 8 The frequency band decomposition module 801 shown in the figure can adopt methods such as discrete Fourier transform, wavelet decomposition, or sub-band decomposition based on filter banks. For two-way frequency division, for example, decomposing a 16 kHz bandwidth ultra-high-definition signal into two 8 kHz bandwidth signals of low frequency and high frequency, an efficient frequency band decomposition method commonly used is quadrature mirror filter (QMF) decomposition, as Figure 9A shown in the figure, where H0(z) is an anti-aliasing low-pass filter with a cut-off frequency around π / 2 (that is, half of the bandwidth frequency), H1(z) is a high-pass filter, H0(z) and H1(z) are symmetric to each other at the orthogonal frequency π / 2 in the frequency spectrum, and the schematic diagram of the frequency spectrum response of the two is as Figure 9B shown in the figure. M = 2 represents 2-fold downsampling. After passing through the low-pass and high-pass filters H b (z), b = 0, 1, the signal is V b (z) = H b (z)*X(z), where X(z) is the spectrum of the input speech frame signal x(n). The signal after downsampling is Therefore, when b = 0, U0(z) is the spectrum of the output low-frequency band signal (that is, the high-definition voice signal), and when k = 1, U1(z) is the spectrum of the output high-frequency band signal.

[0230] Figure 8 The frequency band synthesis module 804 in the figure is the inverse transformation of the frequency band decomposition module 801, asFigure 10 As shown. Where L = 2, indicating 2x upsampling. The spectrum after upsampling is V' b (z) = U' b (z 2 ), b = 1, 2, where U'0(z) is the signal after denoising in the low-frequency band, and U'1(z) is the signal after denoising in the high-frequency band. Figure 10 The relationship between the filter F in [Figure] and the filter H in Figure 9 is: F0(z) = H1(-z) = H0(z), F1(z) = -H1(z) = -H0(-z). The filter coefficients can all be set in advance. The finally output signal spectrum is X'(z) = F0(z)V'0(z) + F1(z)V'1(z). The time-domain representation of X'(z) is Figure 10 in

[0231] Figure 8 The low-frequency band noise suppression module 802 in [Figure] is also Figure 7 the implementation framework of the combined denoising algorithm shown.

[0232] Figure 8 The high-frequency band noise suppression module 803 in [Figure] suppresses the noise in the high-frequency band. Since the importance of the high-frequency band to the perceived quality of the human ear is significantly lower than that of the low-frequency band, a low computational complexity algorithm can be used for denoising. The low-frequency band noise suppression module 802 will calculate the spectral gain G3(k) during the noise suppression process, so the information in G3(k) can be directly borrowed to adjust the gain of the high-frequency signal to achieve the purpose of suppressing most of the noise.

[0233] For example, if the spectrum in the low-frequency band contains N / 2 + 1 frequency points k = 1, 2,..., N / 2 + 1, the larger k is, the higher the frequency. The frequencies of the highest P frequency points are obviously closest to the frequency range of the high-frequency band, and the noise is also closest to the noise in the high-frequency band. Therefore, we select the smallest gain among G3(N / 2 + 1 - P), G3(N / 2 + 1 - P + 1),..., G3(N / 2 + 1) as the gain of the high-frequency band, thus saving the cost of using a complex algorithm to denoise the high-frequency band. Therefore Figure 8 the dashed line in [Figure] represents a gain G4 = Min(G3(N / 2 + 1 - P), G3(N / 2 + 1 - P + 1),..., G3(N / 2 + 1)), where P is an integer less than (N / 2 + 1) and greater than or equal to 0. Therefore, the output of the high-frequency band u'1(n) = u1(n) * G4, where u1(n) is the speech frame signal in the high-frequency band obtained after the input speech frame x(n) is decomposed by frequency band.

[0234] Figure 11 is the structural schematic diagram of a signal processing system in the related art, such as Figure 11As shown, it includes but is not limited to audio acquisition 1101, voice processing link 1102 (including modules such as echo cancellation, noise suppression, howling suppression, gain control, etc.), encoding and transmission 1103, etc. The order of some voice processing modules may be different.

[0235] When the voice communication system upgrades from high-definition voice to ultra-high-definition voice, all voice signal processing modules in the system need to upgrade from supporting high-definition voice to supporting ultra-high-definition voice input, and the computational load of the entire system may increase significantly. Considering that the high-frequency band part has a relatively small impact on the auditory quality, it can be similar to Figure 8 Decompose the input voice frame into frequency bands to obtain low-frequency band and high-frequency band signals and then process them separately, as Figure 12 shown.

[0236] In Figure 12 after the audio acquisition hardware interface 1201 acquires the voice signal, it uses the frequency band decomposition module 1202 to decompose it into a low-frequency band signal and a high-frequency band signal. The signal processing link 1203 includes a low-frequency band signal processing link 12031 and a high-frequency band signal processing link 12032. The low-frequency band signal passes through the low-frequency band signal processing link 12031, and the high-frequency band information passes through the high-frequency band signal processing link 12032. Among them, the low-frequency band signal processing link 12031 follows the Figure 11 shown signal processing link 1102, and at the same time each module outputs a gain value. For example, the low-frequency band noise suppression in Figure 8 before can calculate G4, and the high-frequency band noise suppression part uses this gain to adjust the amplitude of the high-frequency band signal. Using the same method, gain values such as G5, G6, G7, etc. are obtained corresponding to the traditional low-frequency band echo cancellation, low-frequency band howling suppression, low-frequency band gain control, etc. modules, and the high-frequency band signals of each module are adjusted for gain. Finally, the high-frequency band and low-frequency band signal frames after passing through the voice processing link 1203 are sent to the frequency band synthesis module 1204 to re-synthesize the ultra-high-definition voice signal, and then the voice is encoded and sent through the ultra-high-definition voice encoding and transmission module 1205.

[0237] To further reduce the computational load, another implementation method of the microphone ultra-high-definition voice signal processing flow can utilize, such as Figure 13The signal processing system shown implements as follows: The audio acquisition hardware interface 1301 acquires a voice signal, and then performs band decomposition through the band decomposition module 1302 to obtain a low-band signal and a high-band signal. The low-band signal is input to the low-band signal link 1303. After each module in the low-band signal link 1303 obtains each gain value G4 to G7, all the gain values are sent to the high-band gain calculation module 1304, and the high-band gain value G8 = G4 * G5 * G6 * G7 is calculated. Then, the processed high-band signal is obtained by applying this gain at one time. Finally, the high- and low-band signals are sent to the band synthesis module 1305 to re-synthesize a super-clear voice signal, and then the voice is encoded and sent through the encoding and sending module 1306 of the super-clear voice.

[0238] In the embodiment of the present application, the super-clear voice signal is divided into a high-band signal and a low-band signal. For example, if the bandwidth of the super-clear voice signal is 16 kHz, the low-band can be the voice signal within the bandwidth of 0 to 8 kHz, and the high-band can be the voice signal within the bandwidth of 8 kHz to 16 kHz. The human ear is significantly less sensitive to the high-band signal than to the low-band signal. Therefore, in the embodiment of the present application, a deep neural network and a voice enhancement module based on a statistical algorithm are used to eliminate noise (usually also called voice enhancement) in the low frequency band, and the degree of noise elimination in the high frequency band depends on the degree of noise elimination in the low frequency band; other voice processing modules, such as echo cancellation, howling suppression, automatic gain control, etc., also process the high-band and low-band signals separately; all voice processing modules are connected in series to form a voice processing system. After each module in the system processes the voice in the low frequency band separately, it provides reference information for the processing of the high frequency band. Summarizing these reference information can deduce what kind of processing needs to be performed on the high frequency band. In this way, while improving the voice denoising effect, the calculation amount can be greatly reduced, so that the system can also be widely applied on resource-constrained ARM chip platforms such as Android mobile phones.

[0239] Next, the implementation of the audio signal processing device 455 provided in the embodiment of the present application as a software module will be further described. In some embodiments, as Figure 3 shown, the software module stored in the audio signal processing device 455 in the memory 440 may include:

[0240] The first acquisition module 4551 is used to acquire the audio signal to be processed;

[0241] The band decomposition module 4552 is used to perform band decomposition on the audio signal to obtain a first band signal and a second band signal, and the frequency of the first band signal is lower than the frequency of the second band signal;

[0242] The first determination module 4553 is configured to determine a first signal gain corresponding to a first band signal, and determine a second signal gain corresponding to a second band signal based on the first signal gain;

[0243] The second determination module 4554 is configured to determine a processed first band signal based on the first signal gain and the first band signal, and determine a processed second band signal based on the second signal gain and the second band signal;

[0244] The band synthesis module 4555 is configured to perform band synthesis on the processed first band signal and the processed second band signal to obtain a processed audio signal.

[0245] In some embodiments, the first determination module is further configured to:

[0246] Determine a first sub-signal gain corresponding to a first echo cancellation module in a first signal processing link, where the first signal processing link at least includes a first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module for processing the first band signal;

[0247] Determine a second sub-signal gain corresponding to the first noise suppression module, determine a third sub-signal gain corresponding to the first howling control module, and determine a fourth sub-signal gain corresponding to the first gain control module.

[0248] In some embodiments, the first determination module is further configured to:

[0249] Obtain a first band signal input to the first noise suppression module;

[0250] Perform time-frequency conversion on the first band signal to obtain spectral data of the first band signal;

[0251] Input the spectral data into a statistical model to obtain a statistical model gain; input the spectral data into a trained neural network model to obtain a network model gain;

[0252] Determine a second sub-signal gain corresponding to the first noise suppression module based on the statistical model gain and the network model gain.

[0253] In some embodiments, the first determination module is further configured to:

[0254] Determine the smaller value of the statistical model gain and the network model gain as the second sub-signal gain corresponding to the first noise suppression module; or,

[0255] Obtain a first weight value corresponding to the statistical model gain and a second weight value corresponding to the network model gain;

[0256] Use the first weight value and the second weight value to perform a weighted sum of the statistical model gain and the network model gain to obtain the second sub-signal gain corresponding to the first noise suppression module.

[0257] In some embodiments, the first determination module is further configured to:

[0258] Obtain a first prediction probability that the statistical model has speech in the first band signal;

[0259] Obtain a second prediction probability that the trained neural network model has speech in the first band signal;

[0260] Determine the first prediction probability as the first weight value and the second prediction probability as the second weight value; or;

[0261] Obtain a preset first weight value and a preset second weight value.

[0262] In some embodiments, the first determination module is further configured to:

[0263] Determine a fifth sub-signal gain corresponding to the second band signal based on the first sub-signal gain;

[0264] Determine a sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain;

[0265] Determine a seventh sub-signal gain corresponding to the second band signal based on the third sub-signal gain;

[0266] Determine an eighth sub-signal gain corresponding to the second band signal based on the fourth sub-signal gain.

[0267] Determine a second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain.

[0268] In some embodiments, the first determination module is further configured to:

[0269] Determine the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain as the second signal gain corresponding to the second band signal;

[0270] Correspondingly, the second determination module is further configured to:

[0271] Determine the product of the second band signal and the second signal gain as the processed second band signal.

[0272] In some embodiments, the first determination module is further configured to:

[0273] Determine the fifth sub-signal gain as the signal gain of the second echo cancellation module in the second signal processing link; the second signal processing link at least includes a second echo cancellation module, a second noise suppression module, a second howling control module, and a second gain control module for processing the second band signal;

[0274] Determine the sixth sub-signal gain as the signal gain of the second noise suppression module;

[0275] Determine the seventh sub-signal gain as the signal gain of the second howling control module;

[0276] Determine the eighth sub-signal gain as the signal gain of the second gain control module.

[0277] In some embodiments, the second determination module is further configured to:

[0278] Obtain a first output signal obtained by the second echo cancellation module based on the second band signal and the fifth sub-signal gain;

[0279] Obtain a second output signal obtained by the second noise suppression module based on the first output signal and the sixth sub-signal gain;

[0280] Obtain a third output signal obtained by the second howling control module based on the second output signal and the seventh sub-signal gain;

[0281] Obtain a processed second band signal obtained by the second gain control module based on the third output signal and the eighth sub-signal gain.

[0282] In some embodiments, the second sub-signal gain is a gain vector including K gain values, and the determination module is further configured to:

[0283] Determine the top P target frequency points with the highest frequencies from the K frequency points corresponding to the K gain values;

[0284] Determine the gain values corresponding to the P target frequency points as P target gain values;

[0285] Determine the minimum value of the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

[0286] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the audio signal processing method described above in the embodiments of the present application.

[0287] An embodiment of the present application provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the embodiment of the present application. For example, as Figure 4 , Figure 5 and Figure 6 the methods shown.

[0288] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0289] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0290] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file storing other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (e.g., files storing one or more modules, subroutines, or code portions).

[0291] As an example, the executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0292] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. An audio signal processing method, characterized in that Including: Obtain an audio signal to be processed; Perform frequency band decomposition on the audio signal to obtain a first frequency band signal and a second frequency band signal. The frequency of the first frequency band signal is lower than that of the second frequency band signal. The first frequency band signal is a low-frequency band signal, and the second frequency band signal is a high-frequency band signal; Determine a first signal gain corresponding to the first frequency band signal. Based on each sub-signal gain in the first signal gain, confirm each sub-signal gain corresponding to the second frequency band signal. Based on each sub-signal gain corresponding to the second frequency band signal, confirm a second signal gain corresponding to the second frequency band signal; Determine a processed first frequency band signal based on the first signal gain and the first frequency band signal, and determine a processed second frequency band signal based on the second signal gain and the second frequency band signal; Perform frequency band synthesis on the processed first frequency band signal and the processed second frequency band signal to obtain a processed audio signal.

2. The method according to claim 1, characterized in that, The determining the first signal gain corresponding to the first frequency band signal includes: Determine a first sub-signal gain corresponding to a first echo cancellation module in a first signal processing link. The first signal processing link at least includes a first echo cancellation module, a first noise suppression module, a first howling control module, and a first gain control module for processing the first frequency band signal; Determine a second sub-signal gain corresponding to the first noise suppression module, determine a third sub-signal gain corresponding to the first howling control module, and determine a fourth sub-signal gain corresponding to the first gain control module.

3. The method according to claim 2, wherein The determining the second sub-signal gain corresponding to the first noise suppression module includes: Obtain the first frequency band signal input to the first noise suppression module; Perform time-frequency conversion on the first frequency band signal to obtain spectral data of the first frequency band signal; Input the spectral data into a statistical model to obtain a statistical model gain; input the spectral data into a trained neural network model to obtain a network model gain; Determine the second sub-signal gain corresponding to the first noise suppression module based on the statistical model gain and the network model gain.

4. The method according to claim 3, wherein The determining the second sub-signal gain corresponding to the first noise suppression module based on the statistical model gain and the network model gain includes: Determine the smaller value of the statistical model gain and the network model gain as the second sub-signal gain corresponding to the first noise suppression module; or, Obtain a first weight value corresponding to the statistical model gain and a second weight value corresponding to the network model gain; Perform weighted summation on the statistical model gain and the network model gain using the first weight value and the second weight value to obtain the second sub-signal gain corresponding to the first noise suppression module.

5. The method according to claim 4, wherein The obtaining the first weight value corresponding to the statistical model gain and the second weight value corresponding to the network model gain includes: Obtain a first prediction probability that the statistical model has speech in the first frequency band signal; Obtain a second prediction probability that the trained neural network model has speech in the first frequency band signal; Determine the first prediction probability as the first weight value and the second prediction probability as the second weight value; or, Obtain a preset first weight value and a preset second weight value.

6. The method according to claim 2, wherein The confirmation of the respective sub-signal gains corresponding to the second band signal based on the respective sub-signal gains in the first signal gain, and the confirmation of the second signal gain corresponding to the second band signal based on the respective sub-signal gains corresponding to the second band signal, includes: Determine a fifth sub-signal gain corresponding to the second band signal based on the first sub-signal gain; Determine a sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain; Determine a seventh sub-signal gain corresponding to the second band signal based on the third sub-signal gain; Determine an eighth sub-signal gain corresponding to the second band signal based on the fourth sub-signal gain; Determine the second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain.

7. The method according to claim 6, characterized in that The determination of the second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain, includes: Determine the product of the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain as the second signal gain corresponding to the second band signal; Correspondingly, the determination of the processed second band signal based on the second signal gain and the second band signal includes: Determine the product of the second band signal and the second signal gain as the processed second band signal.

8. The method according to claim 6, wherein The determination of the second signal gain corresponding to the second band signal according to the fifth sub-signal gain, the sixth sub-signal gain, the seventh sub-signal gain, and the eighth sub-signal gain, includes: Determine the fifth sub-signal gain as the signal gain of the second echo cancellation module in the second signal processing link; the second signal processing link at least includes a second echo cancellation module, a second noise suppression module, a second howling control module, and a second gain control module for processing the second band signal; Determine the sixth sub-signal gain as the signal gain of the second noise suppression module; Determine the seventh sub-signal gain as the signal gain of the second howling control module; Determine the eighth sub-signal gain as the signal gain of the second gain control module.

9. The method according to claim 8, wherein The determination of the processed second band signal based on the second signal gain and the second band signal includes: Obtain a first output signal obtained by the second echo cancellation module based on the second band signal and the fifth sub-signal gain; Obtain a second output signal obtained by the second noise suppression module based on the first output signal and the sixth sub-signal gain; Obtain a third output signal obtained by the second howling control module based on the second output signal and the seventh sub-signal gain; Obtain the processed second band signal obtained by the second gain control module based on the third output signal and the eighth sub-signal gain.

10. The method according to claim 6, wherein The second sub-signal gain is a gain vector including K gain values. Determining the sixth sub-signal gain corresponding to the second band signal based on the second sub-signal gain includes: Determining the top P target frequency points with the highest frequencies from the K frequency points corresponding to the K gain values, where K is a positive integer greater than 2 and P is a positive integer less than K; Determining the gain values corresponding to the P target frequency points as P target gain values; Determining the minimum value among the P target gain values as the sixth sub-signal gain corresponding to the second band signal.

11. An audio signal processing device, characterized in that, Including: A first acquisition module, configured to acquire an audio signal to be processed; A frequency band decomposition module, configured to perform frequency band decomposition on the audio signal to obtain a first band signal and a second band signal, where the frequency of the first band signal is lower than that of the second band signal, the first band signal is a low-frequency band signal, and the second band signal is a high-frequency band signal; A first determination module, configured to determine a first signal gain corresponding to the first band signal, confirm each sub-signal gain corresponding to the second band signal based on each sub-signal gain in the first signal gain, and confirm a second signal gain corresponding to the second band signal based on each sub-signal gain corresponding to the second band signal; A second determination module, configured to determine a processed first band signal based on the first signal gain and the first band signal, and determine a processed second band signal based on the second signal gain and the second band signal; A frequency band synthesis module, configured to perform frequency band synthesis on the processed first band signal and the processed second band signal to obtain a processed audio signal.

12. An audio signal processing device, characterized in that, Including: A memory, configured to store executable instructions; A processor, configured to implement the method according to any one of claims 1 to 10 when executing the executable instructions stored in the memory.

13. A computer-readable storage medium, characterized in that, Stored with executable instructions, configured to implement the method according to any one of claims 1 to 10 when being executed by a processor.

14. A computer program product, characterized in that, Stored with computer instructions, configured to implement the method according to any one of claims 1 to 10 when being executed by a processor.

Citation Information

Patent Citations

  • Speech enhancement method and device, electronic equipment and storage medium

    CN113299308A