Audio processing method, de-howling model training method and electronic equipment
Through the audio processing methods of deep neural network model and diffusion model, the whistling problem in the amplification device is solved, efficient and accurate whistling removal is achieved, and the audio quality is improved.
Patent Information
- Application Number
- CN202410114593.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to effectively remove the howling phenomenon caused by feedback loops in the amplifying device, affecting the audio quality.
The deep neural network model is used for audio processing, and Gaussian noise is removed through multiple denoising processing to eliminate howling. The diffusion model is used for audio de-howling training, and the model parameters are optimized in combination with forward and reverse diffusion processes.
It improves the accuracy of howling removal and audio processing efficiency, reduces irreversible damage to the audio, and improves audio quality.
Smart Images

Figure CN120416733A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to an audio processing method, a howling elimination model training method, and an electronic device. Background Art
[0002] Artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0003] Machine learning is an important branch of artificial intelligence, and deep learning is an important branch of machine learning. Deep learning refers to using a multi-layer neural network structure to learn the representation forms that can be directly used for computer calculation of various things in the real world from big data (for example, things in images, sounds in audio, etc.).
[0004] A sound amplification device may include a microphone, an amplifier, and a speaker. The microphone is used to convert a sound signal into an electrical signal. The amplifier is used to amplify the power of the electrical signal. The speaker is used to convert the amplified electrical signal into a sound signal. The power of the sound signal output by the speaker is greater than the power of the sound signal collected by the microphone.
[0005] When the microphone and the speaker are used simultaneously, the sound emitted by the speaker may be fed back to the microphone through space transmission. That is to say, the sound signal collected by the microphone is played through the speaker and transmitted back to the microphone through space feedback. When the gain of the sound signal fed back to the microphone relative to the sound signal collected by the microphone (i.e., the feedback gain) is greater than 1, and the phase difference between the sound signal fed back to the microphone and the sound signal collected by the microphone is an integer multiple of 360 degrees, the sound amplification device will generate howling. How to accurately remove howling in audio is an urgent problem to be solved. Summary of the Invention
[0006] The present application provides an audio processing method, an audio processing model training method, and an electronic device, which can improve the accuracy of howling elimination.
[0007] In a first aspect, an audio processing method is provided. An audio to be processed is obtained, and the audio to be processed is obtained from the audio collected by a microphone in a sound reinforcement system. A denoising model is used to perform denoising processing for removing Gaussian noise on multiple first audios in sequence to obtain multiple second audios. The first audio corresponding to the first denoising processing is the audio to be processed. The first audios corresponding to the other denoising processes except the first one are the second audios obtained from the previous denoising processing. The second audio obtained from the last denoising processing is the audio with howling removed. The denoising model is a neural network model obtained through training. The loudspeaker in the sound reinforcement system is used to convert the audio with howling removed into a sound signal. [[ID=@2]]
[0008] The audio processing method provided in this application uses a denoising model to perform denoising processing for removing Gaussian noise on the audio to be processed to obtain a first audio, and uses the denoising model to perform denoising processing for removing Gaussian noise on the first audio to obtain a second audio. The obtained second audio can be used as the first audio again to perform one or more times of denoising processing for removing Gaussian noise. The audio obtained from each denoising processing has a certain similarity with the audio before the denoising processing, so that the obtained audio with howling removed is more accurate.
[0009] Moreover, the denoising model is used to perform denoising processing on the audio to be processed, rather than simply using the audio to be processed as conditional information of the denoising model to perform denoising processing on the noisy audio. Compared with the noisy audio, the audio to be processed is closer to the original audio without howling. Therefore, an accurate audio with howling removed can be obtained with fewer denoising processing times, improving the efficiency of audio processing.
[0010] In some possible implementation manners, the audio to be processed is a preprocessed audio obtained by performing howling removal processing on the audio collected by the microphone.
[0011] The audio after howling removal processing is closer to the original audio before howling occurred. In the process of using the denoising model to perform denoising processing on the audio to be processed, and using the denoising model to perform denoising processing on the second audio obtained from the previous denoising processing, and obtaining the audio with howling removed through multiple denoising processes of the denoising model, the preprocessed audio obtained by howling removal processing is used as the audio to be processed. Therefore, compared with using the audio without howling removal processing as the audio to be processed, using the preprocessed audio obtained by howling removal processing as the audio to be processed can obtain an accurate audio with howling removed with fewer denoising processing times of the denoising model. That is to say, using the preprocessed audio obtained by howling removal processing as the audio to be processed can reduce the number of denoising processing times using the denoising model without affecting the accuracy of audio processing, improving the efficiency of audio processing.
[0012] In some possible implementation manners, each denoising process in the multiple denoising processes includes: using the anti-whistling model to perform the denoising process on the first audio corresponding to the denoising process according to the conditional information, where the conditional information includes conditional audio obtained from the audio collected by the microphone.
[0013] In the process of performing the denoising process using the anti-whistling model each time, using the conditional audio obtained from the audio collected by the microphone as a part of the input of the anti-whistling model can improve the accuracy of the anti-whistling audio obtained by the process.
[0014] In some possible implementation manners, the audio to be processed is preprocessed audio obtained by performing an anti-whistling process on the audio collected by the microphone, and the conditional audio is audio that has not been subjected to an anti-whistling process.
[0015] Performing an anti-whistling process on the audio may cause irreversible damage to the audio. Using the audio that has not been subjected to an anti-whistling process as the conditional audio enables the process of performing the denoising process using the anti-whistling model to comprehensively consider the influence of the audio that has not been subjected to an anti-whistling process, making the obtained anti-whistling audio more accurate.
[0016] In some possible implementation manners, the anti-whistling model is obtained by adjusting the parameters of the initial anti-whistling model according to the difference between the second training audio corresponding to each first training audio in a plurality of first training audio and the second training audio corresponding to the first training audio; the third training audio corresponding to each first training audio is obtained by using the initial anti-whistling model to perform a denoising process for removing Gaussian noise on the plurality of first training audio respectively; the plurality of first training audio includes training initial audio, the training initial audio is obtained from the training whistling audio with whistling corresponding to the labeled anti-whistling audio, the labeled anti-whistling audio does not include whistling, each first training audio is expressed as being obtained by performing a noise addition process for adding Gaussian noise on the second training audio corresponding to the first training audio, the second training audio corresponding to the first noise addition process is the labeled anti-whistling audio, the second training audio corresponding to the other noise addition processes except the first time is the first training audio obtained by the previous noise addition process, different training audio other than the training initial audio in the plurality of training audio are obtained by different times of noise addition processes, and the first training audio obtained by the last noise addition process is the training initial audio.
[0017] In some possible implementation manners, the training initial audio is training preprocessed audio obtained by performing an anti-whistling process on the training whistling audio.
[0018] In some possible implementation manners, the third training audio corresponding to each first training audio is obtained by performing the denoising process on the first training audio according to the training conditions by using the initial dewhistling model, and the training conditions include a training condition audio obtained according to the training whistling audio.
[0019] In some possible implementation manners, the training initial audio is a training preprocessing audio obtained by performing a dewhistling process on the training whistling audio, and the training condition audio is an audio without a dewhistling process.
[0020] In a second aspect, a method for training a dewhistling model is provided, including: obtaining a plurality of first training audios and a labeled dewhistling audio, where the plurality of training audios include a training initial audio, the labeled dewhistling audio does not include whistling, the training initial audio is obtained according to a training whistling audio with whistling corresponding to the labeled dewhistling audio, each first training audio is expressed as being obtained by performing a noise addition process for adding Gaussian noise on a second training audio corresponding to the first training audio, the second training audio corresponding to the first noise addition process is the labeled dewhistling audio, the second training audio corresponding to other noise addition processes except the first time is the first training audio obtained by the previous noise addition process, different training audios among the plurality of training audios except the training initial audio are obtained by different times of noise addition processes, and the first training audio obtained by the last noise addition process is the training initial audio; using the initial dewhistling model to perform a denoising process for removing Gaussian noise on the plurality of first training audios respectively to obtain a third training audio corresponding to each first training audio; and adjusting parameters of the initial dewhistling model according to a difference between the second training audio corresponding to each first training audio among the plurality of first training audios and the second training audio corresponding to the first training audio to obtain a dewhistling model.
[0021] In some possible implementation manners, the training initial audio is a training preprocessing audio obtained by performing a dewhistling process on the training whistling audio.
[0022] In some possible implementation manners, the using the initial dewhistling model to perform a denoising process on the plurality of first training audios respectively to obtain a third training audio corresponding to each first training audio includes: for each first training audio, using the initial dewhistling model to perform a denoising process on the first training audio according to the training conditions to obtain the third training audio corresponding to the first training audio, where the training conditions include a training condition audio obtained according to the training whistling audio.
[0023] In some possible implementation manners, the initial training audio is the preprocessed training audio obtained by performing de-whistling processing on the training whistling audio, and the conditional training audio is the audio without de-whistling processing.
[0024] In a third aspect, an audio processing apparatus is provided, including units for performing the method according to the first aspect or the second aspect. The apparatus may be a terminal device or a chip within the terminal device.
[0025] In a fourth aspect, an electronic device is provided, including one or more processors and a memory. The memory is coupled to the one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to perform the method according to the first aspect and / or the second aspect.
[0026] In a fifth aspect, a chip system is provided. The chip system is applied to an electronic device and includes one or more processors. The one or more processors are used to call computer instructions to cause the electronic device to perform the method according to the first aspect and / or the second aspect.
[0027] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes instructions. When the instructions run on an electronic device, the electronic device is caused to perform the method according to the first aspect and / or the second aspect.
[0028] In a seventh aspect, a computer program product is provided. The computer program product includes computer program code. When the computer program code runs on an electronic device, the electronic device performs the method according to the first aspect and / or the second aspect. Description of the Drawings
[0029] Figures 1 to 4 is a schematic diagram of a sound reinforcement system;
[0030] Figure 5 is a schematic flowchart of an audio processing method provided by an embodiment of the present application;
[0031] Figure 6 is a schematic diagram of a forward diffusion process and a reverse diffusion process;
[0032] Figure 7 is a schematic flowchart of an audio processing system provided by an embodiment of the present application;
[0033] Figure 8 is a schematic diagram of a graphical user interface provided by an embodiment of the present application;
[0034] Figure 9It is a schematic diagram of another graphical user interface provided by an embodiment of the present application;
[0035] Figure 10 It is a schematic structural diagram of a howling elimination model training method provided by an embodiment of the application;
[0036] Figure 11 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;
[0037] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0038] Figure 13 It is a schematic diagram of the software system of an electronic device provided by an embodiment of the present application;
[0039] Figure 14 It is a schematic diagram of a system architecture provided by an embodiment of the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings.
[0041] In daily life, people often encounter scenarios where sound amplification is required, such as: relatively noisy scenarios, meeting scenarios, or classroom scenarios, etc.
[0042] Figure 1 It is a schematic diagram of a sound reinforcement system.
[0043] The sound reinforcement system 200 may include a microphone 210, a speaker 220, and an amplifier 230.
[0044] The microphone 210, also known as a microphone or a transmitter, is used to convert a sound signal into an electrical signal to achieve the acquisition of an audio electrical signal. The sound signal can also be referred to as a voice signal or a sound wave signal.
[0045] The amplifier 230 is used to amplify the power of the audio electrical signal collected by the microphone 210.
[0046] The speaker 220, also known as a loudspeaker, is used to convert the audio electrical signal amplified by the amplifier 230 into a sound signal.
[0047] The microphone 210, the speaker 220, and the amplifier 230 may be located in the same or different electronic devices.
[0048] Such as Figure 2As shown, the public address system 200 can be a sound amplification device. The speaker 220 and the amplifier 230 are arranged in the housing 240 of the sound amplification device. The microphone 210 is connected to the amplifier 230 in the housing 240 in a wired manner. This scenario can be a class or a meeting. The user approaches the microphone 210 and speaks. The microphone 210 collects the user's voice signal and transmits it to the amplifier 230 in the housing 240 in a wired manner.
[0049] An electronic device can serve as the public address system 200. Thus, the user can speak to the mobile phone, and the mobile phone can collect the user's voice signal through the microphone and amplify the voice signal. Exemplarily, this scenario can be a live broadcast scenario or a speech scenario, that is, when giving a speech, the user can amplify the voice through the mobile phone without carrying a special sound amplification device.
[0050] As Figure 3 As shown, the public address system 200 can include an electronic device 310 and a sound device 320. The electronic device 310 can include a microphone 210, and the sound device 320 can include a speaker 220 and an amplifier 230. The electronic device 310 and the sound device 320 can communicate in a wired or wireless manner. Exemplarily, this scenario can also be a live broadcast scenario or a speech scenario, that is, when giving a speech, the user approaches the microphone in the electronic device 310 and speaks. The microphone in the electronic device 310 collects the user's voice signal, and the sound device 320 amplifies the user's voice signal.
[0051] As Figure 4 As shown, the public address system 200 can include a headphone 330, an electronic device 310, and a sound device 320. The headphone 330 can include a microphone 210, and the sound device 320 can include a speaker 220 and an amplifier 230. The electronic device 310 can communicate with the headphone 330 in a wired or wireless manner. If the headphone 330 is a wired headphone, the headphone 330 is connected to the electronic device 310 through a headphone cable. If the headphone 330 is a wireless headphone, such as a Bluetooth headphone, the electronic device 310 establishes a wireless connection with the headphone 330, such as a Bluetooth connection. The electronic device 310 and the sound device 320 can communicate in a wired or wireless manner.
[0052] This scenario can be a live broadcast, a speech, a class, a meeting, or other scenarios. The user speaks to the microphone of the headphone, and the headphone 330 can send the user's voice signal collected by the microphone 210 in the headphone 330 to the electronic device 310. Thus, the electronic device 310 can forward the user's voice signal collected by the microphone 210 to the sound device 320. The sound device 320 amplifies the user's voice signal.
[0053] When having an online meeting or class, or at a large conference, after the speaker turns on the sound reinforcement system 200, the speaker 220 may emit a sharp sound, similar to the sound of a whistle but even sharper and harsher. After moving the microphone 210 in the sound reinforcement system 200, this harsh sound may disappear, but it may reappear when the microphone 210 is in some other positions. This sharp and harsh sound is called howling.
[0054] When the microphone 210 and the speaker 220 are used simultaneously, the sound emitted by the speaker 220 may be transmitted through space and fed back to the microphone 210, forming a closed-loop circuit. When the Nyquist stability criterion is met, that is, at a certain frequency point, the closed-loop circuit gain is greater than 1 and the phase difference between the sound wave signal fed back to the microphone 210 and the original sound wave signal input by the microphone 210 is an integer multiple of 360 degrees, the system is unstable at this frequency point, generating self-excited oscillation, and the sound output by the speaker 220 is a harsh howl.
[0055] The closed-loop circuit gain can be understood as the amplification factor of the sound reinforcement system 200 for the audio collected by the microphone 210.
[0056] To eliminate howling, an embodiment of the present application provides an audio processing method.
[0057] The following combines Figure 5 A detailed description is given to the audio processing method provided by the embodiment of the present application. The execution subject of the method provided by the present application can be an electronic device, or a software / hardware module in the electronic device capable of performing audio processing. For the convenience of description, the following embodiments are described by taking the electronic device as an example.
[0058] Figure 5 It is a schematic flowchart of the audio processing method provided by the embodiment of the present application. The method may include steps S510 to S520, and the following gives a detailed description of these steps respectively.
[0059] Step S510, obtain the audio to be processed, and the audio to be processed is obtained according to the audio collected by the microphone in the sound reinforcement system.
[0060] The audio to be processed may include one frame or multiple frames.
[0061] The audio to be processed may be the audio collected by the microphone in the sound reinforcement system.
[0062] The audio to be processed may also be the audio obtained by performing howling elimination processing on the audio collected by the microphone, or may be the audio obtained by performing howling elimination processing on the amplified audio obtained by amplifying the audio collected by the microphone through the amplifier in the sound reinforcement system.
[0063] The audio to be processed may also be the amplified audio obtained by amplifying the audio collected by the microphone by an amplifier in the sound reinforcement system, or the audio obtained by performing a howling removal process on the audio collected by the microphone and amplifying it by the amplifier in the sound reinforcement system.
[0064] The anti-howling process can be implemented by one or more of a phase shift method, a notch filter method, an adaptive filter method, a neural network method, and the like.
[0065] Phase shifting allows for audio framing and spectrum analysis, identifying the frequency points where howling occurs. Phase shifting is then applied to these frequencies, allowing the speaker to convert the phase-shifted audio into a sound wave signal. This sound wave signal, fed back to the microphone, is no longer phase-shifted by an integer multiple of 360 degrees from the original sound wave signal input from the microphone, eliminating the howling. The sound wave signal input from the original sound source can be understood as the sound wave signal corresponding to the audio being framed and spectrum analyzed.
[0066] By using the notch filter method, the audio frame can be divided into spectrum analysis to determine the frequency point where howling occurs and design a corresponding notch filter to reduce the gain of the howling frequency point so that the closed-loop gain is less than or equal to 1, thereby eliminating the howling.
[0067] The phase shift method and notch filter method adjust the phase or gain of the frequency point where the howling occurs so that the frequency point where the howling occurs in the audio no longer meets the Nyquist stability criterion, destroying the formation of self-excited oscillation, thereby eliminating the howling in the audio.
[0068] Through the adaptive filtering method, an adaptive filter can be used to simulate the feedback path in the sound field, and the signal emitted by the speaker through the simulated feedback path can be subtracted from the audio to remove the howling in the audio.
[0069] Using a neural network method, a neural network can be used to process audio. Audio resulting from the howling removal process can be determined based on the output of the neural network processing of the audio. For example, the output of the neural network processing of the audio can be audio resulting from the howling removal process, or difference information, which can represent the difference between the audio input to the neural network and the audio resulting from the howling removal process.
[0070] Since the detection of the howling point is not accurate enough, the audio frequency howling removal processing by the phase shift method may result in the phase shifting of the non-howling point, so that the non-howling point meets the Nyquist stability criterion and generates howling.
[0071] Due to inaccurate detection of the howling point, when using the notch filtering method to perform howling removal on the audio, it may notch-filter non-howling points. Moreover, due to the bandwidth design limitation of the notch filtering method, the signals near the howling frequency points will also be affected by the reduced gain. Therefore, using the notch filtering method will cause significant damage to the sound quality of the processed audio.
[0072] The adaptive filtering method has problems such as slow convergence speed, untimely tracking, and steady-state offset error, which will affect both the subjective perception of people and the sound quality. Therefore, the audio after howling removal processing can be used as the audio to be processed.
[0073] Generally, the audio after howling removal processing is closer to the original audio before howling occurs.
[0074] Performing howling removal processing on the audio may cause irreversible damage to the audio.
[0075] Step S520: Use the howling removal model to sequentially perform denoising processing for removing Gaussian noise on multiple first audios to obtain multiple second audios. The first audio corresponding to the first denoising processing is the audio to be processed. The first audios corresponding to the other denoising processes except the first one are the second audios obtained from the previous denoising processing. The second audio obtained from the last denoising processing is the howling removal audio. The howling removal model is a neural network model obtained through training. The loudspeaker in the sound reinforcement system is used to convert the howling removal audio into a sound signal.
[0076] The denoising processing is used to remove Gaussian noise. Thus, each time the howling removal model performs the denoising processing on the first audio corresponding to this denoising processing, the Gaussian noise in this first audio can be removed.
[0077] The howling removal model can be understood as a diffusion model. The diffusion model is a type of deep neural network (DNN).
[0078] The neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and the intercept 1 as inputs. The output of this operation unit can be:
[0079]
[0080] where s = 1, 2, …… n, n is a natural number greater than 1, and W s is x sThe weight is \(w\), \(b\) is the bias of the neuron. \(f\) is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0081] DNN can also be called a multi - layer neural network and can be understood as a neural network with multiple hidden layers. Classifying DNN according to the positions of different layers, the neural networks inside DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is to say, any neuron in the \(i\) - th layer must be connected to any neuron in the \((i + 1)\) - th layer.
[0082] Although DNN looks very complex, in terms of the work of each layer, it is actually not. The work of each layer in a deep neural network can be described by the mathematical expression as follows: From a physical level, the work of each layer in a deep neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / reduction; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by \(W\), the operation of 4 is completed by \(+b\), and the operation of 5 is realized by \(a()\). The reason for using the word "space" here is that the objects to be classified are not single things, but a class of things, and the space refers to the set of all individuals of this class of things. Among them, \(W\) is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector \(W\) determines the space transformation from the input space to the output space described above, that is, the weight \(W\) of each layer controls how to transform the space. The purpose of training a deep neural network is to finally obtain the weight matrices of all layers of the trained neural network (the weight matrix formed by vectors \(W\) of many layers). Therefore, the training process of a neural network is essentially the process of learning the way to control space transformation, and more specifically, it is learning the weight matrix.
[0083] Therefore, DNN can be simply expressed by the following linear relationship expression: Among them, \(\mathbf{x}\) is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer is simply an operation on the input vector to obtain the output vector through such a simple operation Due to the large number of layers in the DNN, the number of coefficients W and offset vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output in the third layer and the index 4 of the input in the second layer.
[0084] In summary, the coefficient from the kth neuron in the (L - 1)th layer to the jth neuron in the Lth layer is defined as
[0085] It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0086] The framework of the diffusion model can adopt a denoising diffusion probabilistic model (DDPM), a score-based generative model (SGM), or a stochastic differential equation (SDE), etc.
[0087] The model structure of the diffusion model can be a convolutional neural network, such as a U-Net, a noise conditional score network (NCSN), an upgraded version of NCSN (NCSN++), etc.
[0088] A convolutional neural network is a deep neural network with a convolutional structure and is a deep learning architecture. A deep learning architecture refers to multiple levels of learning at different levels of abstraction through machine learning algorithms. As a deep learning architecture, a CNN is a feed-forward artificial neural network in which each neuron can respond to the data input into it.
[0089] A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve with an input data or a convolutional feature plane. A convolutional layer refers to the layer of neurons in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected only to some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weights here are the convolutional kernels. Taking the input data as an image as an example, sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of the image is the same as that of other parts. That is to say, the image information learned in a certain part can also be used in another part. Therefore, for all positions on the image, we can use the same learned image information. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels, the richer the image information reflected by the convolution operation.
[0090] The convolutional kernel can be initialized in the form of a matrix of random size, and the convolutional kernel can learn reasonable weights during the training process of the convolutional neural network. In addition, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0091] A convolutional neural network can include an input layer, convolutional layers, and neural network layers. A convolutional neural network can also include pooling layers.
[0092] A convolutional layer can include many convolutional operators, which are also called kernels. Their role in natural language processing is equivalent to a filter that extracts specific information from the input speech or semantic information. A convolutional operator can essentially be a weight matrix, and this weight matrix is usually predefined.
[0093] The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can extract information from the input data, thereby helping the convolutional neural network to make correct predictions.
[0094] When a convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network increases, the features extracted by the subsequent convolutional layers become more and more complex, such as high-level semantic features, and the higher the semantic features, the more suitable they are for the problem to be solved.
[0095] Since it is often necessary to reduce the number of training parameters, a pooling layer is often introduced periodically after the convolutional layer. It can be one convolutional layer followed by one pooling layer, or multiple convolutional layers followed by one or more pooling layers.
[0096] During the audio processing, the only purpose of the pooling layer is to reduce the spatial size of the audio. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input audio features to obtain smaller-sized audio features.
[0097] After being processed by the convolutional layer / pooling layer, the convolutional neural network is not sufficient to output the required output information. As mentioned above, the convolutional layer / pooling layer only extracts features and reduces the parameters brought by the input data. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network needs to use neural network layers to generate one or a group of outputs with the number of classes required. Therefore, the neural network layer can include multiple hidden layers and an output layer. The parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include speech or semantic recognition, classification or generation, etc.
[0098] After the multiple hidden layers in the neural network layer, that is, the last layer of the entire convolutional neural network is the output layer, which has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the backpropagation will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network, that is, the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0099] It should be understood that the above introduction to the convolutional neural network is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models. For example, multiple convolutional layers / pooling layers are parallel, and the features extracted separately are all input to the fully neural network layer for processing, etc.
[0100] The use of diffusion models involves a forward diffusion process and a reverse diffusion process.
[0101] The forward diffusion process can also be called the forward diffusion process or the diffusion process. The reverse diffusion process can also be called the inverse diffusion process, the abstraction process, or the denoising process. In the forward diffusion process, a Markov chain for generating samples is created by slowly adding noise. In the reverse diffusion process, a Markov chain for generating samples is created by slowly removing noise.
[0102] A Markov chain consists of a series of states and a series of transition probabilities. Here, the states refer to data with different noise levels, and the transition probabilities refer to the probabilities of changing from the current state to the next state, which are implemented using a transition matrix.
[0103] For example, as Figure 6 shown, for data x0, in the forward diffusion process, at the t-th step of the T steps of the forward diffusion process, a small amount of Gaussian noise is added to data x t-1 to obtain data x t , where data x t represents the data obtained after t steps. Here, t = 1, 2 …, T, and T is a positive integer. That is, for data x0, after T steps, the data obtained at each step are x1, x2, …, x T . In the forward diffusion process, the parameter t can represent the number of iterations.
[0104] The forward diffusion process can be expressed as:
[0105] p 0t (x <4000009>|x0,y) = N c (x t , μ(x0,y,t), σ(t) 2 I).
[0106] Among them, p 0t (x t |x0,y) represents the expression of data x t at time t, y represents the condition, N c represents the Gaussian distribution, μ(x0,y,t) represents the mean of the Gaussian distribution, σ(t) represents the standard deviation of the Gaussian distribution, and σ(t) 2 represents the variance of the Gaussian distribution, and I represents the identity matrix.
[0107] The mean μ(x0, y, t) of the Gaussian distribution can be determined according to the condition y, the parameter t representing the number of iterations, and the normalization constant γ. The mean μ(x0, y, t) of the Gaussian distribution can be expressed as:
[0108] μ(x0, y, t) = e -γt x0 + (1 - e -γ t)y.
[0109] The standard deviation σ(t) of the Gaussian distribution can be determined according to the parameter t representing the number of iterations, the normalization constant γ, the preset maximum variance σ max and the preset minimum variance σ min . The square of the standard deviation σ(t) of the Gaussian distribution is the variance of the Gaussian distribution, and the variance σ(t) 2 of the Gaussian distribution can be expressed as:
[0110]
[0111] According to the expression of the data x t at the time t, the data x t can be determined.
[0112] The forward diffusion process can be understood as the training process of the diffusion model. Using the initial diffusion model to process the training samples when the parameter t is t = 1, 2,..., T to obtain the training denoised data, and adjusting the parameters of the initial diffusion model according to the difference between the training denoised data and the data x t-1 . The training samples of the initial diffusion model include the data x t , and the diffusion model is the initial diffusion model after parameter adjustment.
[0113] The training denoised data can be the output of the initial diffusion model, or can be obtained by removing the noise represented by the output of the initial diffusion model from the data x t .
[0114] When the training denoised data is obtained by removing the noise represented by the output of the initial diffusion model from the data x t , adjusting the parameters of the initial diffusion model according to the difference between the training denoised data and the data x t-1 can be understood as adjusting the parameters of the initial diffusion model according to the difference between the output of the initial diffusion model and the noise added to the data x t relative to the data x t-1 .
[0115] According to the difference between the output of the initial diffusion model and the data x t-1 when t is t = 1, 2,..., T, adjusting the parameters of the initial diffusion model, we can obtain which can be expressed as:
[0116]
[0117] Among them, s θ represents the output of the initial diffusion model, represents the noise added to the data x t-1 compared with the data x t in the data x, arg θ is a preset constant, |||| represents the norm, represents the expectation for multiple t in the case of t = 1, 2..., T. The norm is usually used to measure the length or size of each vector in a certain vector space (or matrix).
[0118] The reverse diffusion process can be understood as the inference process of the diffusion model. In the reverse diffusion process, t can be understood as the remaining number of iterations.
[0119] The diffusion model has good robustness. Repeating the forward diffusion process a large number of times during training can enhance the robustness of the model.
[0120] For the de-whistling model, during training, audio without whistling can be used as the label to de-whistle the audio. According to the training whistling audio with whistling corresponding to the label de-whistled audio, the training initial audio can be determined. That is to say, the training initial audio can be obtained based on the training whistling audio with whistling corresponding to the label de-whistled audio.
[0121] Exemplarily, the training diffusion system may include a training microphone and a training speaker. The training microphone is used for sound collection and converts the sound signal into an electrical signal. The training speaker is used to convert the electrical signal collected by the training microphone into a sound signal.
[0122] The training diffusion system may further include other speakers besides the training speaker, and the other speakers can be used to convert the label de-whistled audio into a sound signal, so that the signal collected by the training microphone can record the sound recorded in the label de-whistled audio.
[0123] The sound signal output by the training speaker is obtained by converting the electrical signal collected by the training microphone. When there is whistling in the sound signal output by the training speaker, the electrical signal collected by the training microphone is the training whistling audio with whistling corresponding to the label de-whistled audio. Thus, according to the electrical signal output by the training microphone, the training initial audio can be determined.
[0124] Alternatively, the training whistling audio with whistling corresponding to the label de-whistled audio can be obtained by superimposing the signals at one or more frequency points of the label de-whistled audio on the label de-whistled audio.
[0125] The initial training audio can be a training whistling audio. Alternatively, the initial training audio can be a preprocessed training audio obtained by removing the whistling from the training whistling audio.
[0126] When the initial training audio is a training whistling audio, the audio to be processed can be the audio collected by the microphone and not processed for removing the whistling.
[0127] When the initial training audio is a preprocessed training audio obtained by removing the whistling from the training whistling audio, the audio to be processed can be the preprocessed audio. The preprocessed audio can be understood as the audio obtained by removing the whistling from the audio collected by the microphone.
[0128] Although the process of removing the whistling will affect the sound quality of the audio, the audio after removing the whistling is closer to the original audio before the whistling occurred.
[0129] In the process of using the de-whistling model to perform denoising processing for removing Gaussian noise on the audio to be processed, and using the de-whistling model to perform denoising processing for removing Gaussian noise on the second audio obtained from the previous denoising processing, and obtaining the de-whistled audio through multiple denoising processes of the de-whistling model, the preprocessed audio obtained by the de-whistling process is used as the audio to be processed. Therefore, compared with using the audio not processed for removing the whistling as the audio to be processed, using the preprocessed audio processed for removing the whistling as the audio to be processed can obtain an accurate de-whistled audio after fewer denoising processes of the de-whistling model. That is to say, using the preprocessed audio processed for removing the whistling as the audio to be processed can reduce the number of denoising processes using the de-whistling model without affecting the accuracy of audio processing, that is, a smaller parameter T can be set, and the total number of steps in the forward diffusion process and the reverse diffusion process can be reduced.
[0130] During the training process of the de-whistling model, compared with using the training whistling audio with whistling corresponding to the labeled de-whistled audio as the initial training audio, using the preprocessed training audio as the initial training audio makes the difference between the initial training audio and the labeled de-whistled audio smaller, and a smaller parameter T can also be set.
[0131] Reducing the total number of steps represented by the parameter T in the forward diffusion process can reduce the training efficiency of the de-whistling model. Reducing the total number of steps in the reverse diffusion process, that is, reducing the number of denoising processes using the de-whistling model, can improve the audio processing efficiency.
[0132] When using unprocessed audio as the pre-processed audio, the number of denoising operations performed using the whistling removal model is reduced, meaning the total number of steps in the reverse diffusion process can be set to 30 to 40. When using pre-processed audio as the pre-processed audio, setting the total number of steps in the reverse diffusion process to 5 achieves similar audio processing results to the 30 to 40 steps used when using unprocessed audio as the pre-processed audio, effectively removing howling from the pre-processed audio. Using pre-processed audio that has undergone whistling removal as the pre-processed audio significantly reduces the number of denoising operations performed using the whistling removal model.
[0133] Based on the initial training audio and the labeled whistling-free audio, multiple first training audios can be obtained. Each first training audio can be represented as a second training audio corresponding to the first training audio subjected to a noise addition process for adding Gaussian noise. The second training audio corresponding to the first noise addition process is the labeled whistling-free audio, and the second training audio corresponding to the noise addition process other than the first process is the first training audio obtained by the previous noise addition process. The first training audio obtained by the last noise addition process is the initial training audio. In other words, the multiple first training audios include the initial training audio.
[0134] By performing denoising processing on the plurality of first training audios respectively using the initial whistle removal model, a third training audio corresponding to each first training audio can be obtained.
[0135] The anti-howling model can be obtained by adjusting the parameters of the initial anti-howling model according to the difference between the second training audio corresponding to each first training audio and the second training audio corresponding to the first training audio.
[0136] The label-free howling audio can be used as data x0, and the training initial audio can be used as data x T .
[0137] According to data x0 and data x T , and the preset number of steps T, the data x1, x2, ..., x T-1 . Among them, the data x1, x2, ..., x T-1 The data x in t is in the data x t-1 Gaussian noise is added, t=1,2,…,T-1. Data x T It can also be understood as the data x T-1 Add Gaussian noise. That is, the data x1, x2, ..., x T-1 The data x in t It can be understood as the data x t-1 Obtained by adding Gaussian noise, t=1,2,…,T.
[0138] In data x t-1 The mean of the added Gaussian noise can be the average μ(x0, y, t) of the Gaussian distribution. In data x t-1 The standard deviation of the added Gaussian noise can be expressed as the standard deviation σ(t) of the Gaussian distribution.
[0139] Thus, in step S520, using the trained de-whistling model, multiple denoising processes can be performed to obtain the de-whistling audio.
[0140] Step S520 can be understood as the reverse diffusion process of the diffusion model. In the reverse diffusion process, the audio to be processed can be understood as data x T , and the de-whistling audio can be understood as data x0.
[0141] The number of the first training audios used in the training process can be equal to the number of the second audios in step S520. That is to say, the number of times of adding noise in the training process can be equal to the number of times of denoising using the de-whistling model in step S520.
[0142] In the training process of the de-whistling model, the data used can include or not include the training conditions. The training conditions can be understood as the conditions y in the forward diffusion process.
[0143] The training conditions can include the training condition audios determined according to the training whistling audios with whistling corresponding to the labeled de-whistling audios.
[0144] When the data used in the training process of the de-whistling model includes the training conditions, in the training process of the de-whistling model, the initial de-whistling model can be used to perform denoising processing on multiple first training audios respectively according to the training conditions to obtain the third training audios corresponding to each first training audio.
[0145] When the data used in the training process of the de-whistling model includes the training conditions, in step S520, using the de-whistling model, according to the condition information, performing the denoising processing on the first audio corresponding to each denoising processing can obtain the second audio corresponding to the denoising processing.
[0146] The condition information can include the condition audios obtained from the audio collected by the microphone.
[0147] In the process of each denoising using the de-whistling model, taking the condition audio obtained from the audio collected by the microphone as a part of the input of the de-whistling model can improve the accuracy of the obtained de-whistling audio.
[0148] In some embodiments, the training audio can be the pre - processed training audio obtained by performing de - whistling processing on the training whistling audio. The conditional audio can be the pre - processed audio obtained by performing de - whistling processing on the audio collected by the microphone.
[0149] In other embodiments, the training conditional audio can be the training whistling audio. The conditional audio can be the audio collected by the microphone without de - whistling processing.
[0150] Performing de - whistling processing on the audio may cause irreversible damage to the audio. Using the audio without de - whistling processing as the conditional audio enables the denoising process using the de - whistling model to comprehensively consider the influence of the audio without de - whistling processing, making the obtained de - whistled audio more accurate.
[0151] The training conditional audio can also be the audio without environmental noise reduction and enhancement processing. The conditional audio can also be the audio without environmental noise reduction and enhancement processing.
[0152] Environmental noise reduction and enhancement processing can be used to reduce the noise signals entering the microphone, such as wind noise, environmental noise, and other interference signals, and can moderately enhance the user's voice signal to improve the clarity and signal - to - noise ratio of the human voice.
[0153] Environmental noise reduction and enhancement processing may cause irreversible damage to the audio noise.
[0154] During each denoising process using the de - whistling model in step S520, the conditional information may or may not include the information about the remaining number of denoising processes.
[0155] During the training process of the de - whistling model, the training conditions can also include the number of times of adding noise to the labeled de - whistling audio. The number of times of adding noise to the labeled de - whistling audio can be understood as the parameter t in the forward diffusion process.
[0156] In the case where the training conditions include the number of times of adding noise to the labeled de - whistling audio, during each denoising process using the de - whistling model in step S520, the conditional information can include the information about the remaining number of denoising processes. The information about the remaining number of denoising processes can be understood as the parameter t in the reverse diffusion process.
[0157] The sound reinforcement system can also include an amplifier. Training the sound reinforcement system can include training the amplifier.
[0158] The training initial audio can be obtained by amplifying the electrical signal output by the training microphone in the training sound reinforcement system through the training amplifier.
[0159] The audio to be processed can be obtained by an amplifier in a sound reinforcement system amplifying the audio collected by a microphone in the sound reinforcement system. A speaker in the sound reinforcement system can convert the second audio, i.e., the anti-whistling audio, obtained by performing the last denoising process using the anti-whistling model, into a sound signal.
[0160] That is to say, the audio to be processed can be the amplified audio obtained by an amplifier in the sound reinforcement system amplifying the audio collected by the microphone, or the audio obtained by performing anti-whistling processing on the audio collected by the microphone, and then amplified by the amplifier in the sound reinforcement system.
[0161] Alternatively, the initial audio for training can be the audio that has not been amplified by the training amplifier.
[0162] The audio to be processed can be the audio that has not been amplified by the amplifier in the sound reinforcement system. That is to say, the audio to be processed can be the audio collected by the microphone in the sound reinforcement system, or the audio obtained by performing anti-whistling processing on the audio collected by the microphone.
[0163] Exemplarily, as Figure 7 shown, the audio processing system 700 can also be referred to as a sound amplification system. The audio processing system can include a microphone 210, a speaker 220, an amplifier 230, an audio preprocessing module 710, and an anti-whistling model 720.
[0164] The audio collected by the microphone 210 in the sound reinforcement system 200 can be used as the conditional audio.
[0165] Before performing step S520, the audio preprocessing module 710 can perform anti-whistling processing on the audio collected by the microphone 210, i.e., the conditional audio, to obtain the preprocessed audio. The preprocessed audio can be used as the audio to be processed.
[0166] The audio preprocessing module 710 can also perform other preprocessing on the audio collected by the microphone 210, such as environmental noise reduction and enhancement processing, sampling rate and bit width conversion processing, and other preprocessing.
[0167] The sampling rate and bit width conversion processing can be used to convert the speech signal into a sampling rate and bit width compatible with the system.
[0168] Step S520 may include step S521 and step S522. In step S521, the de-whistling model 720 may be utilized to perform noise reduction processing on the audio to be processed according to the condition information to obtain a second audio. The condition information includes a condition audio. In step S522, the second audio may be used as the first audio, and the de-whistling model 720 may be utilized to perform noise reduction processing on the first audio according to the condition information to obtain a new second audio. Step S522 may be performed once or multiple times. The second audio obtained by performing step S522 for the last time may be used as the de-whistled audio.
[0169] The amplifier 230 in the sound reinforcement system 200 may amplify the de-whistled audio. The speaker 220 in the sound reinforcement system may convert the amplified de-whistled audio into a sound signal.
[0170] In some embodiments, before performing step S510, it may also be determined whether there is whistling in the audio to be processed. In the case where there is whistling in the audio to be processed, steps S510 to S520 may be performed. Conversely, in the case where there is no whistling in the audio to be processed, steps S510 to S520 may no longer be performed. The speaker is used to convert the audio to be processed without whistling into a sound signal.
[0171] By performing whistling detection on the audio to be processed, it may be determined whether there is whistling in the audio to be processed.
[0172] During the whistling detection process, the audio to be processed may be transformed from the time domain to the frequency domain to obtain the frequency domain signal corresponding to the audio to be processed, and the first ratio between the first signal energy at the first frequency point and the total energy of the frequency domain signal may be calculated. The first frequency point may be any frequency point within the frequency domain where the frequency domain signal is located. In the case where the first ratio is greater than the preset whistling threshold value, it may be determined that there is whistling in the audio signal to be processed.
[0173] Alternatively, the whistling detection may also be implemented by other means, and the embodiments of the present application do not limit this.
[0174] During the whistling detection process, the frequency point with the largest signal energy may be used as the first frequency point, or multiple frequency points within the frequency domain where the frequency domain signal is located may be used as the first frequency point for detection respectively.
[0175] The method provided by the embodiments of the present application utilizes the de-whistling model to perform multiple noise reduction processes for removing Gaussian noise on the audio to be processed to obtain the de-whistled audio. The audio obtained by each noise reduction process has a certain similarity with the audio before the noise reduction process, so that the obtained de-whistled audio is more accurate, that is, the method provided by the embodiments of the present application can obtain an audio with higher quality.
[0176] Moreover, in the audio processing method provided in this application, the anti-whistling model is used to denoise the audio to be processed, rather than simply using the audio to be processed as the conditional information of the anti-whistling model to denoise the noisy audio. Compared with the noisy audio, the audio to be processed is closer to the original audio without whistling. Therefore, an accurate anti-whistling audio can be obtained by using the anti-whistling model for denoising a smaller number of times, improving the efficiency of audio processing.
[0177] It should be understood that Figure 5 the method shown can be processed by a central processing unit (CPU) in an electronic device, or can be jointly processed by a CPU and a neural-network processing unit (NPU), or other processors suitable for neural network computing can be used instead of the NPU, and this application does not make any restrictions.
[0178] Exemplarily, the electronic device executing Figure 5 the method shown may include Figure 2 the housing 240 shown and an amplifier 230 and a speaker 220 disposed in the housing. The electronic device executing Figure 5 the method shown can be connected to a microphone 210 through an interface to receive the audio collected by the microphone 210.
[0179] Or, the electronic device executing Figure 5 the method shown can be Figure 3 the electronic device 310 shown. A microphone may be provided in the electronic device 310.
[0180] Alternatively, the electronic device executing Figure 5 the method shown can be Figure 4 the electronic device 310 shown. The electronic device 310 can communicate with a headset 330 by wired or wireless means to receive the audio collected by the microphone in the headset 330.
[0181] Or, the electronic device executing Figure 5 the method shown can be Figure 3 or Figure 4 the audio device 320 shown. The audio device 320 can communicate with the electronic device 310 by wired or wireless means to receive the audio collected by the microphone in the electronic device 310, or receive the audio collected by the microphone in the headset 330 forwarded by the electronic device 310.
[0182] The electronic device executing Figure 5 the method shown can determine the audio to be processed according to the audio collected by the microphone, and perform audio processing on the audio to be processed to obtain an anti-whistling audio.
[0183] Execute Figure 5 The speaker provided in the electronic device that executes the method shown can convert the anti-whistling audio into a sound signal. Or, execute Figure 5 The electronic device that executes the method shown can communicate with the audio device 320 provided with a speaker in a wired or wireless manner, and send the anti-whistling audio to the audio device 320, so that the speaker in the audio device 320 can convert the anti-whistling audio into a sound signal.
[0184] For executing Figure 5 The electronic device that executes the method shown can be Figure 10 The electronic device 100 shown.
[0185] Exemplarily, when the user needs to use the electronic device to achieve sound amplification, the user can turn on the sound amplification function of the terminal in the setting interface of the electronic device. Figure 8 (a) in shows a graphical user interface (GUI) of the electronic device, and this GUI is the desktop 1210 of the electronic device. When detecting the setting start operation of the user clicking the setting icon 1211 of the setting application (APP) on the desktop 1210, the electronic device can start the setting application and display as Figure 8 (b) in shows another GUI, and this GUI can be called the setting interface 1220. The setting interface 1220 can include a sound amplification switch icon 1221.
[0186] Or, when the user needs to use the electronic device to achieve sound amplification, the user can turn on the sound amplification function of the terminal in the down-dragging notification bar interface of the electronic device. Figure 9 shows a graphical user interface (GUI) of the electronic device, and this GUI is the down-dragging notification bar interface 1310 of the electronic device. The down-dragging notification bar interface 1310 can include a sound amplification switch icon 1221.
[0187] When the sound amplification switch icon 1221 in the setting interface 1220 or the down-dragging notification bar interface 1310 indicates that the sound amplification function is in the on state, if it is detected that the user clicks the sound amplification switch icon 1221, the electronic device can set the state of the sound amplification function indicated by the sound amplification switch icon 1221 to the off state; on the contrary, when the sound amplification switch icon 1221 indicates that the sound amplification function is in the off state, if it is detected that the user clicks the sound amplification switch icon 1221, the electronic device can set the state of the sound amplification function indicated by the sound amplification switch icon 1221 to the on state.
[0188] When the sound amplification function of the electronic device is in the on state, as a part of the sound amplification system, the electronic device can execute Figure 5The method shown processes the audio collected by the microphone and transmits the processed audio to the speaker.
[0189] After the wireless sound amplification function of the electronic device is turned on, the electronic device can obtain the user's voice signal, and pass the obtained user's voice signal through the processor or audio module of the electronic device, such as Figure 12 the audio module 170 in it, and after processing the user's voice signal, output the processed voice signal.
[0190] Exemplarily, the electronic device can determine the audio to be processed according to the voice signal collected by the microphone, and perform multiple denoising processes on the audio to be processed using a de-whistling model, so as to obtain a de-whistling audio. The de-whistling audio can be understood as the processed voice signal output by the electronic device. The speaker can convert the processed voice signal into a sound signal.
[0191] The microphone and the speaker can be located in the electronic device that displays the setting interface 1220, or can be located in other electronic devices.
[0192] Exemplarily, the device for collecting the user's voice signal can be the microphone built in the electronic device. When the electronic device is connected to the earphone, the device for collecting the user's voice signal can also be the microphone in the earphone.
[0193] The electronic device can amplify the processed voice signal through its own speaker, or, when the electronic device is connected to an external device, it can also output the processed voice signal to the external device and amplify the voice signal through the external device. The external device can be an audio device, for example, a Bluetooth audio device, or the audio device of a TV, etc.
[0194] Figure 5 The method described above can be specifically executed by an execution device 1110 as shown in Figure 14 The audio to be processed in the method described above can be the input data given by a client device 1140 as shown in Figure 5 The preprocessing module 1113 in the execution device 1110 can be used to process the audio to be processed to obtain preprocessed audio. The calculation module 1111 in the execution device 1110 can be used to execute steps S510 to S520. Figure 14 shown.
[0195] Next, in combination with Figure 10 , for Figure 5 the training method of the de-whistling model used in the audio processing method shown is described.
[0196] Figure 10 is a schematic flowchart of a de-whistling model training method provided by an embodiment of the present application. Figure 10The method shown includes steps S810 to S830.
[0197] Step S810: Obtain a plurality of first training audio and labeled denoised audio. The plurality of training audio includes training initial audio. The labeled denoised audio does not include whistling. The training initial audio is obtained from the training whistling audio with whistling corresponding to the labeled denoised audio. Each first training audio is obtained by performing noise addition processing for adding Gaussian noise to the second training audio corresponding to the first training audio. The second training audio corresponding to the first noise addition processing is the labeled denoised audio. The second training audio corresponding to the other noise addition processes except the first time is the first training audio obtained from the previous noise addition processing. Different training audio among the plurality of training audio except the training initial audio are obtained from different times of noise addition processing. The first training audio obtained from the last noise addition processing is the training initial audio.
[0198] Step S820: Use the initial denoising model to perform noise removal processing for removing Gaussian noise on the plurality of first training audio respectively to obtain the third training audio corresponding to each first training audio.
[0199] Step S830: Adjust the parameters of the initial denoising model according to the difference between the second training audio corresponding to each first training audio among the plurality of first training audio and the second training audio corresponding to the first training audio to obtain a denoising model.
[0200] Optionally, the training initial audio is training preprocessed audio obtained by performing denoising processing on the training whistling audio.
[0201] Optionally, the use of the initial denoising model to perform noise removal processing for removing Gaussian noise on the plurality of first training audio respectively to obtain the third training audio corresponding to each first training audio includes: For each first training audio, use the initial denoising model to perform noise removal processing on the first training audio according to training conditions to obtain the third training audio corresponding to the first training audio. The training conditions include training condition audio obtained from the training whistling audio.
[0202] Optionally, the training initial audio is training preprocessed audio obtained by performing denoising processing on the training whistling audio, and the training condition audio is audio that has not undergone denoising processing.
[0203] Through Figure 10 The denoising model trained by the method shown can be applied in Figure 5 The audio processing method shown. Figure 5 The audio processing method shown andFigure 10 The de-whistling model training method shown can be performed by the same or different electronic devices. For example, a terminal device can be used to perform Figure 5 the audio processing method shown, and a server can be used to Figure 10 perform the de-whistling model training method shown. The terminal device can be Figure 12 the electronic device 100 shown.
[0204] Perform Figure 10 The device that performs the Figure 14 method shown can also be Figure 14 the training device 1120 shown. The multiple first training audios and labeled de-whistling audios used in step S810 can be the training data maintained in
[0205] Optionally, Figure 10 steps S810 to S830 of the method shown can be performed in the training device 1120, or can be pre-executed by other functional modules before the training device 1120, that is, first preprocess the training data received or obtained from the database 1130, such as determining other first training audios except the training initial audio in the multiple first training audios according to the labeled de-whistling audio and the training initial audio, and using the multiple first training audios, the labeled de-whistling audio as the input of the training device 1120, and the training device 1120 performs S810 to S830.
[0206] Optionally, Figure 10 the method shown can be processed by a CPU, or can be jointly processed by a CPU and a GPU, or can use other processors suitable for neural network computing instead of a GPU, and the present application does not make any restrictions.
[0207] The server can send the trained de-whistling model to the electronic device. The electronic device can, when obtaining the audio to be processed, use the de-whistling model to process the obtained data to obtain the de-whistling audio.
[0208] It should be understood that the above examples are for helping those skilled in the art to understand the embodiments of the present application, rather than limiting the embodiments of the present application to the specific values or specific scenarios illustrated. Those skilled in the art can clearly make various equivalent modifications or changes according to the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0209] As described above in conjunction with Figures 5 to 10 , the audio processing method and the de-whistling model training method of the embodiments of the present application have been described in detail. Next, in conjunction with Figures 11 to 14, an apparatus embodiment of the present application is described in detail. It should be understood that the audio processing apparatus in the embodiments of the present application can execute the audio processing method and the anti-whistling model training method in the foregoing embodiments of the present application. That is, the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.
[0210] Figure 11 is a schematic diagram of an audio processing apparatus provided by an embodiment of the present application.
[0211] The audio processing apparatus 900 includes an acquisition unit 910 and a processing unit 920.
[0212] In some embodiments, the audio processing apparatus 900 can execute Figure 5 the audio processing method shown.
[0213] The acquisition unit 910 is configured to acquire an audio to be processed, where the audio to be processed is obtained based on the audio collected by a microphone in a sound reinforcement system.
[0214] The processing unit 920 is configured to use the anti-whistling model to perform denoising processing for removing Gaussian noise on a plurality of first audios in sequence to obtain a plurality of second audios. The first audio corresponding to the first denoising processing is the audio to be processed. The first audio corresponding to other denoising processes except the first time is the second audio obtained by the previous denoising process. The second audio obtained by the last denoising process is the anti-whistling audio. The anti-whistling model is a neural network model obtained by training. The loudspeaker in the sound reinforcement system is configured to convert the anti-whistling audio into a sound signal.
[0215] Optionally, the audio to be processed is a preprocessed audio obtained by performing anti-whistling processing on the audio collected by the microphone.
[0216] Optionally, the processing unit 920 is configured to use the anti-whistling model to perform denoising processing on the first audio corresponding to each denoising process according to condition information, where the condition information includes a condition audio obtained based on the audio collected by the microphone.
[0217] Optionally, the audio to be processed is a preprocessed audio obtained by performing anti-whistling processing on the audio collected by the microphone, and the condition audio is an audio that has not been subjected to anti-whistling processing.
[0218] In other embodiments, the audio processing apparatus 900 can execute Figure 10 the anti-whistling model training method shown.
[0219] The acquisition unit 910 is configured to acquire a plurality of first training audio and label denoised audio. The plurality of first training audio includes training initial audio. The label denoised audio does not include whistling. The training initial audio is obtained from the training whistling audio with whistling corresponding to the label denoised audio. Each first training audio is obtained by performing noise addition processing for adding Gaussian noise on the second training audio corresponding to the first training audio. The second training audio corresponding to the first noise addition processing is the label denoised audio, and the second training audio corresponding to the noise addition processing other than the first time is the first training audio obtained from the previous noise addition processing. Different training audio among the plurality of training audio except the training initial audio are obtained from different times of noise addition processing, and the first training audio obtained from the last noise addition processing is the training initial audio.
[0220] The processing unit 920 is configured to use the initial denoising model to perform denoising processing for removing Gaussian noise on the plurality of first training audio respectively, so as to obtain a third training audio corresponding to each first training audio.
[0221] The processing unit 920 is further configured to adjust the parameters of the initial denoising model according to the difference between the second training audio corresponding to each first training audio in the plurality of first training audio and the second training audio corresponding to the first training audio, so as to obtain a denoising model.
[0222] Optionally, the training initial audio is a training preprocessing audio obtained by performing denoising processing on the training whistling audio.
[0223] Optionally, the processing unit 920 is specifically configured to, for each first training audio, use the initial denoising model to perform denoising processing on the first training audio according to training conditions to obtain the third training audio corresponding to the first training audio. The training conditions include a training condition audio obtained from the training whistling audio.
[0224] Optionally, the training initial audio is a training preprocessing audio obtained by performing denoising processing on the training whistling audio, and the training condition audio is an audio without denoising processing.
[0225] It should be noted that the above audio processing device 900 is embodied in the form of functional units. The term "unit" here can be implemented in software and / or hardware forms, and no specific limitation is made thereto.
[0226] For example, a "unit" may be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor, or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a combined logic circuit, and / or other suitable components that support the described functions.
[0227] Therefore, the units of each example described in the embodiments of the present application can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0228] Figure 12 A hardware system of an electronic device applicable to the present application is shown.
[0229] The electronic device 100 may be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, etc. The embodiments of the present application do not impose any restrictions on the specific type of the electronic device 100.
[0230] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0231] Figure 12 The connection relationships shown among the various modules are only illustrative descriptions and do not constitute limitations on the connection relationships among the various modules of the electronic device 100. Optionally, the various modules of the electronic device 100 may also adopt combinations of various connection methods in the above embodiments.
[0232] It should be noted that Figure 1 the structure shown does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than Figure 1 the components shown, or the electronic device 100 may include Figure 12 combinations of some of the components shown, or the electronic device 100 may include Figure 12 sub-components of some of the components shown. Figure 1 The components shown may be implemented in hardware, software, or a combination of software and hardware.
[0233] The processor 110 may include one or more processing units. For example, the processor 110 may include at least one of the following processing units: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and a neural-network processing unit (NPU). Among them, different processing units may be independent devices or integrated devices.
[0234] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0235] The NPU is a processor that draws on the structure of a biological neural network. For example, it quickly processes input information by drawing on the transmission pattern between human brain neurons and can also continuously self-learn. Through the NPU, functions such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, voice recognition, and text understanding.
[0236] The external memory interface 120 can be used to connect an external memory card, such as a secure digital (SD) card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0237] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as a sound playback function and an image playback function). The data storage area can store data created during the use of the electronic device 100 (such as audio data and a phone book). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as: at least one disk storage device, a flash memory device, and a universal flash storage (UFS), etc. The processor 110 executes various processing methods of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0238] The electronic device 100 can implement audio functions, such as music playback and recording, through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.
[0239] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and can also be used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 or some functional modules of the audio module 170 can be provided in the processor 110.
[0240] The headphone jack 170D is used to connect a wired headphone.
[0241] The touch sensor 180K, also known as a touch control device. The touch sensor 180K can be provided on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, and the touch screen is also known as a touch control screen. The touch sensor 180K is used to detect a touch operation acting on it or near it. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be provided on the surface of the electronic device 100 and be provided at a different position from the display screen 194.
[0242] Figure 12 The hardware system of the electronic device 100 is described in detail. Next, the software system of the electronic device 100 is introduced. The software system can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the layered architecture is taken as an example to exemplarily describe the software system of the electronic device 100.
[0243] As Figure 13 shown, a software system adopting a layered architecture is divided into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into four layers, from top to bottom, namely the application layer, the application framework layer, Android Runtime and system libraries, and the kernel layer.
[0244] The application layer may include applications such as a camera, a gallery, a calendar, a call, a map, a navigation, a WLAN, a Bluetooth, music, a video, a short message, etc.
[0245] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer may include some predefined functions.
[0246] For example, the application framework layer includes a window manager, a content provider, a view system, a phone manager, a resource manager, and a notification manager.
[0247] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, a locked screen, and capture the screen.
[0248] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audios, dialed and answered calls, browsing histories and bookmarks, and phone books.
[0249] The view system includes visible controls, such as controls for displaying text and controls for displaying pictures. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.
[0250] The phone manager is used to provide the communication function of the electronic device 100, such as the management of call status (connected or hung up).
[0251] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, and video files.
[0252] The notification manager enables applications to display notification information in the status bar, can be used to convey notification-type messages, and can disappear automatically after a short stay without user interaction.
[0253] The Android Runtime includes core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0254] The core libraries consist of two parts: one is the functional functions that the Java language needs to call, and the other is the core libraries of Android.
[0255] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0256] The system libraries can include multiple functional modules, such as: surface manager, Media Libraries, 3D graphics processing library, and 2D graphics engine.
[0257] The surface manager is used to manage the display subsystem and provides the fusion of 2D layers and 3D layers for multiple applications.
[0258] The media libraries support the playback and recording of multiple audio formats, the playback and recording of multiple video formats, and static image files. The media libraries can support multiple audio and video coding formats.
[0259] The 3D graphics processing library can be used to implement 3D graphics drawing, image rendering, synthesis, and layer processing.
[0260] The 2D graphics engine is the drawing engine for 2D drawing.
[0261] The kernel layer is the layer between the hardware and the software. The kernel layer can include driver modules such as display drivers, camera drivers, audio drivers, and sensor drivers.
[0262] It should be noted that the above Figure 12 illustrates the structure diagram of the electronic device by way of Figure 13 illustrates the software architecture diagram of the electronic device by way of; this application makes no limitations in this regard.
[0263] The settings APP or other application programs in the application layer of the electronic device 100 can be used to execute the Figure 5 audio processing method shown.
[0264] Figure 14 is a system architecture provided by an embodiment of this application.
[0265] As shown in the system architecture 1100, the data acquisition device 1160 is used to acquire training data. In the embodiments of the present application, the training data includes: a plurality of first training audios and label denoised audios. The data acquisition device 1160 is further used to store the training data in the database 1130.
[0266] The training device 1120 trains to obtain the target model / rule 1101 based on the training data maintained in the database 1130. The target model / rule 1101 may be a trained denoising model. Figure 10 The method for training the denoising model describes in detail how the training device 1120 obtains the target model / rule 1101 based on the training data. The target model / rule 1101 can be used to implement the audio processing method provided in the embodiments of the present application, that is, using the target model / rule 1101 to perform denoising processing on a plurality of first audios in sequence to obtain a plurality of second audios. The first audio corresponding to the first denoising processing is the audio to be processed, and the first audios corresponding to the other denoising processes except the first time are the second audios obtained from the previous denoising process. Therefore, the second audio obtained from the last denoising process is the denoised audio.
[0267] It should be noted that in actual applications, the training data maintained in the database 1130 may not necessarily come from the acquisition of the data acquisition device 1160, and it may also be received from other devices. Additionally, it should be noted that the training device 1120 may not necessarily train the target model / rule 1101 completely based on the training data maintained in the database 1130, and it may also obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiments of the present application.
[0268] The target model / rule 1101 trained according to the training device 1120 can be applied to different systems or devices, such as applied to Figure 14 the execution device 1110 shown. The execution device 1110 may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, AR / VR, a vehicle-mounted terminal, etc., or may also be a server or the cloud, etc. In the appendix Figure 14 the execution device 1110 is configured with an I / O interface 1112 for data interaction with external devices. The user can input data to the I / O interface 1112 through the client device 1140. The input data in the embodiments of the present application may include: the audio collected by the microphone in the sound reinforcement system.
[0269] The preprocessing modules 1113 and 1114 are used to perform preprocessing on the input data received by the I / O interface 1112 (such as the audio collected by the microphone in the public address system). In the embodiments of the present application, the preprocessing module 1113 can be used to perform howling suppression processing on the audio collected by the microphone in the public address system to obtain the preprocessed audio.
[0270] In the embodiments of the present application, it is also possible that there are no preprocessing modules 1113 and 1114 (or only one of the preprocessing modules), and the computing module 1111 is directly used to process the input data.
[0271] When the execution device 1110 performs preprocessing on the input data, or when the computing module 1111 of the execution device 1110 performs calculations and other related processing, the execution device 1110 can call the data, code, etc. in the data storage system 1,150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 1150.
[0272] Finally, the I / O interface 1112 returns the processing result, such as the above-mentioned howling suppression audio, to the client device 1140, and thus provides it to the user.
[0273] It should be noted that the training device 1120 can generate corresponding target models / rules 1101 based on different training data for different targets or tasks, and the corresponding target models / rules 1101 can be used to achieve the above-mentioned targets or complete the above-mentioned tasks, so as to provide the required results for the user.
[0274] In the attached Figure 14 In the case shown in the figure, the user can manually give the input data, and this manual giving can be operated through the interface provided by the I / O interface 1112. In another case, the client device 1140 can automatically send the input data to the I / O interface 1112. If the client device 1140 is required to automatically send the input data and user authorization is required, the user can set the corresponding permissions in the client device 1140. The user can view the results output by the execution device 1110 on the client device 1140, and the specific presentation form can be specific ways such as display, sound, and action. The client device 1140 can also be used as a data acquisition end to collect the input data input to the I / O interface 1112 and the output result of the output I / O interface 1112 as new sample data and store them in the database 1130. Of course, it is also possible not to collect through the client device 1140, but directly store the input data input to the I / O interface 1112 and the output result of the output I / O interface 1112 shown in the figure as new sample data into the database 1130.
[0275] It should be noted that the appended Figure 14 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitations. For example, in the appended Figure 14 , the data storage system 1150 is an external memory relative to the execution device 1110. In other cases, the data storage system 1150 can also be placed in the execution device 1110.
[0276] As Figure 14 shown, the target model / rule 1101 is trained according to the training device 1120. The target model / rule 1101 can be a de-noising model in an embodiment of the present application. Specifically, the de-noising model provided by the embodiment of the present application is a U-shaped network (U-NET), a noise conditional score network (NCSN) based on gradient denoising, an upgraded version of NCSN (NCSN++), etc. The de-noising model can be a convolutional neural network.
[0277] The present application also provides a chip, which includes a data interface and one or more processors. When the one or more processors execute instructions, the one or more processors read the instructions stored on the memory through the data interface to implement the audio processing method and / or the de-noising model training method described in the above method embodiments.
[0278] The one or more processors can be general-purpose processors or dedicated processors. For example, the one or more processors can be a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gate, transistor logic devices, or discrete hardware components.
[0279] The chip can be a component of a terminal device or other electronic devices. For example, the chip can be located in the electronic device 100.
[0280] The processor and the memory can be set separately or integrated together. For example, the processor and the memory can be integrated on a system on chip (SOC) of the terminal device. That is to say, the chip can also include a memory.
[0281] Programs can be stored in the memory, and the programs can be run by the processor to generate instructions, enabling the processor to execute the audio processing method and / or the de-whistling model training method described in the above method embodiments according to the instructions.
[0282] Optionally, data can also be stored in the memory. Optionally, the processor can also read the data stored in the memory. The data can be stored at the same storage address as the program, or it can be stored at a different storage address from the program.
[0283] Exemplarily, the memory can be used to store the relevant programs of the audio processing method provided in the embodiments of the present application, and the processor can be used to call the relevant programs of the audio processing method stored in the memory to implement the audio processing method of the embodiments of the present application.
[0284] For example, obtain the audio to be processed, where the audio to be processed is obtained from the audio collected by the microphone in the sound reinforcement system; use the de-whistling model to perform denoising processing on multiple first audios in sequence to obtain multiple second audios. The first audio corresponding to the first denoising processing is the audio to be processed, the first audios corresponding to the other denoising processes except the first time are the second audios obtained from the previous denoising processing, and the second audio obtained from the last denoising processing is the de-whistled audio. The de-whistling model is a neural network model obtained by training, and the speaker in the sound reinforcement system is used to convert the de-whistled audio into a sound signal.
[0285] Exemplarily, the memory can be used to store the relevant programs of the de-whistling model training method provided in the embodiments of the present application, and the processor can be used to call the relevant programs of the de-whistling model training method stored in the memory to implement the de-whistling model training method of the embodiments of the present application.
[0286] For example, multiple first training audio and label denoised audio are obtained. The multiple first training audio include training initial audio. The label denoised audio does not include whistling. The training initial audio is obtained from the training whistling audio with whistling corresponding to the label denoised audio. Each first training audio is obtained by performing noise addition processing for adding Gaussian noise to the second training audio corresponding to the first training audio. The second training audio corresponding to the first noise addition processing is the label denoised audio. The second training audio corresponding to other noise addition processing except the first time is the first training audio obtained from the previous noise addition processing. Different training audio except the training initial audio among the multiple training audio are obtained from different times of noise addition processing. The first training audio obtained from the last noise addition processing is the training initial audio. The initial denoising model is used to perform denoising processing on the multiple first training audio respectively to obtain the third training audio corresponding to each first training audio. According to the difference between the second training audio corresponding to each first training audio among the multiple first training audio and the second training audio corresponding to the first training audio, the parameters of the initial denoising model are adjusted to obtain the denoising model.
[0287] The chip can be arranged in an electronic device.
[0288] The present application also provides a computer program product, which when executed by a processor implements the touch recognition method described in any method embodiment of the present application.
[0289] The computer program product can be stored in a memory. For example, it is a program, and after processes such as preprocessing, compilation, assembly, and linking, it is finally converted into an executable target file that can be executed by the processor.
[0290] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, it implements the touch recognition method described in any method embodiment of the present application. The computer program can be a high-level language program or an executable target program.
[0291] The computer-readable storage medium is, for example, a memory. The memory can be a volatile memory or a non-volatile memory, or the memory can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0292] In the embodiments of the present application, the use of user data may be involved. In practical applications, user-specific personal data can be used in the solutions described herein within the scope permitted by applicable laws and regulations, provided that the requirements of the applicable laws and regulations of the country where the user is located are met (for example, the user gives explicit consent, is effectively notified to the user, etc.).
[0293] In the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance, as well as a specific order or sequence. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0294] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items" or a similar expression means any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0295] It should be understood that in various embodiments of the present application, the magnitude of the sequence numbers of the above processes does not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0296] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0297] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0298] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.
[0299] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0300] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0301] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An audio processing method, characterized in that Including: Obtain an audio to be processed, where the audio to be processed is obtained based on the audio collected by a microphone in a sound reinforcement system; Use a de-whistling model to sequentially perform denoising processing for removing Gaussian noise on multiple first audios to obtain multiple second audios. The first audio corresponding to the first denoising processing is the audio to be processed. The first audios corresponding to the denoising processing other than the first time are the second audios obtained from the previous denoising processing. The second audio obtained from the last denoising processing is the de-whistling audio. The de-whistling model is a neural network model obtained by training. The speaker in the sound reinforcement system is used to convert the de-whistling audio into a sound signal.
2. The method according to claim 1, characterized in that, The audio to be processed is a preprocessed audio obtained by performing de-whistling processing on the audio collected by the microphone.
3. The method according to claim 1 or 2, characterized in that, Each denoising processing includes: using the de-whistling model, according to condition information, performing the denoising processing on the first audio corresponding to the denoising processing, where the condition information includes a conditional audio obtained based on the audio collected by the microphone.
4. The method according to claim 3, wherein The audio to be processed is a preprocessed audio obtained by performing de-whistling processing on the audio collected by the microphone, and the conditional audio is an audio that has not undergone de-whistling processing.
5. A method for training a de-whistling model, characterized in that, The method includes: Obtain multiple first training audios and a labeled de-whistling audio. The multiple first training audios include training initial audios. The labeled de-whistling audio does not include whistling. The training initial audios are obtained based on the training whistling audios with whistling corresponding to the labeled de-whistling audio. Each first training audio is represented as being obtained by performing noise addition processing for adding Gaussian noise to the second training audio corresponding to the first training audio. The second training audio corresponding to the first noise addition processing is the labeled de-whistling audio. The second training audios corresponding to the noise addition processing other than the first time are the first training audios obtained from the previous noise addition processing. Different training audios among the multiple training audios other than the training initial audios are obtained from different times of noise addition processing. The first training audio obtained from the last noise addition processing is the training initial audio; Use an initial de-whistling model to respectively perform denoising processing for removing Gaussian noise on the multiple first training audios to obtain a third training audio corresponding to each first training audio; According to the difference between the second training audio corresponding to each first training audio in the multiple first training audios and the second training audio corresponding to the first training audio, adjust the parameters of the initial de-whistling model to obtain a de-whistling model.
6. The method according to claim 5, wherein The training initial audio is a training preprocessed audio obtained by performing de-whistling processing on the training whistling audio.
7. The method according to claim 5 or 6, characterized in that, The step of using the initial de-whistling model to perform denoising processing for removing Gaussian noise on the multiple first training audios respectively to obtain a third training audio corresponding to each first training audio includes: for each first training audio, using the initial de-whistling model, according to the training conditions, perform the denoising processing on the first training audio to obtain the third training audio corresponding to the first training audio, where the training conditions include a training condition audio obtained according to the training whistling audio.
8. The method according to claim 7, wherein The training initial audio is a training preprocessing audio obtained by performing de-whistling processing on the training whistling audio, and the training condition audio is an audio without de-whistling processing.
9. An electronic device, characterized in that, The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
10. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions, and when the instructions run on an electronic device, the instructions cause the electronic device to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Audio processing method, sound separation model training method and electronic equipment
CN120954432A
Audio processing methods, sound separation model training methods, and electronic devices
CN120954432B