Audio processing method, model training method, and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-08-07
AI Technical Summary
[0006]第一方面提供的音频处理方法,用户可以输入第一指令,即用户可以根据需求,自主选择降噪强度和/或去混响强度,实现音频效果的自定义,解决降噪和/或去混响的效果较为单一的问题,提高用户体验;而且,该方法基于扩散模型处理音频,得到的目标音频质量更高、听感更好,能够有效提升用户体验
Smart Images

Figure CN120431958B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, specifically to an audio processing method, a model training method, and an electronic device. Background Technology
[0002] After electronic devices collect or store audio in real time, noise reduction and / or de-reverberation can be applied to improve the listening experience.
[0003] The audio processing methods in related technologies offer limited noise reduction and / or dereverberation effects, resulting in a poor user experience. Summary of the Invention
[0004] This application provides an audio processing method, a model training method, and an electronic device, allowing users to independently select the noise reduction intensity and / or dereverberation intensity of audio, thereby improving the user experience.
[0005] In a first aspect, this application provides an audio processing method executed by an electronic device. The method includes: acquiring audio to be processed; receiving a first instruction input by a user, the first instruction indicating adjustment of noise reduction intensity and / or dereverberation intensity; encoding the first instruction by a first encoder to obtain a first instruction feature vector; inputting the first audio as input and the first instruction feature vector as conditional input into an audio processing model; the first audio being a segment of the audio to be processed; the audio processing model being a model trained based on an initial diffusion model; and the audio processing model performing a back-diffusion process to perform noise reduction processing and / or dereverberation processing on the first audio according to the first instruction feature vector, and outputting a first target audio.
[0006] The first aspect provides an audio processing method where users can input a first command, allowing them to customize the audio effect by selecting the noise reduction and / or dereverberation intensity according to their needs. This addresses the issue of relatively limited noise reduction and / or dereverberation effects and improves the user experience. Furthermore, this method processes audio based on a diffusion model, resulting in higher quality and better listening experience, effectively enhancing the user experience.
[0007] In one possible implementation, the first instruction includes one or two sets of instruction information. Each set of instruction information includes a function type and a corresponding adjustment method. The function type is noise reduction or dereverberation, and the adjustment method is addition, subtraction, or n%. The function type of noise reduction indicates that the audio processing method is noise reduction, and the function type of dereverberation indicates that the audio processing method is dereverberation. The adjustment method of addition indicates that the processing intensity is increased, the adjustment method of subtraction indicates that the processing intensity is decreased, and the adjustment method of n% is used to indicate that the processing intensity is adjusted to n% of the preset adjustable range of processing intensity, where n is a value greater than 0 and less than 100.
[0008] In this implementation, the first instruction can indicate that the noise reduction intensity is increased, decreased, or set to n% of the preset adjustable range of noise reduction intensity, and can also indicate that the dereverberation intensity is increased, decreased, or set to n% of the preset adjustable range of dereverberation intensity, providing users with multiple ways to adjust the noise reduction intensity and dereverberation intensity, meeting different user needs, and further improving the user experience.
[0009] In one possible implementation, after acquiring the audio to be processed and before receiving the first instruction input by the user, the method further includes: acquiring a second instruction, the second instruction being used to indicate adjustment of noise reduction intensity and / or dereverberation intensity; encoding the second instruction through a first encoder to obtain a second instruction feature vector; inputting the second audio as input and the second instruction feature vector as conditional input into an audio processing model; the second audio being a segment preceding the first audio in the audio to be processed; the audio processing model performing a backdiffusion process, performing noise reduction processing and / or dereverberation processing on the second audio according to the second instruction feature vector, and outputting the second target audio.
[0010] Optionally, the first command can be the default intensity adjustment command or a historical intensity adjustment command. That is, after acquiring the audio to be processed, the audio can be processed based on the default or historical intensity adjustment command. If needed, the user can then input a second command to adjust the noise reduction intensity and / or de-reverberation intensity.
[0011] In one possible implementation, the first instruction includes instruction information with the function type being noise reduction and the adjustment method being addition, and the first signal-to-noise ratio difference is greater than the second signal-to-noise ratio difference; the first signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the first target audio and the signal-to-noise ratio of the first audio, and the second signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the second target audio and the signal-to-noise ratio of the second audio.
[0012] Since the signal-to-noise ratio (SNR) is negatively correlated with noise power, increasing the noise reduction intensity reduces noise in the audio, decreases noise power, and increases the SNR. Therefore, the change in SNR is greater, meaning the first SNR difference is greater than the second SNR difference. Conversely, a first SNR difference greater than the second SNR difference indicates an increase in noise reduction intensity. Thus, in this implementation, when the first instruction includes instructions with the function type "noise reduction" and the adjustment method "addition," the noise reduction intensity is increased.
[0013] In one possible implementation, the first instruction includes instruction information with the function type being noise reduction and the adjustment method being reduction, and the first signal-to-noise ratio difference is less than the second signal-to-noise ratio difference; the first signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the first target audio and the signal-to-noise ratio of the first audio, and the second signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the second target audio and the signal-to-noise ratio of the second audio.
[0014] Since the signal-to-noise ratio (SNR) is negatively correlated with noise power, a decrease in noise reduction intensity leads to an increase in noise in the audio, resulting in increased noise power and a decrease in the SNR. Consequently, the change in SNR decreases, meaning the first SNR difference is less than the second SNR difference. Conversely, a first SNR difference less than the second SNR difference indicates a decrease in noise reduction intensity. Therefore, in this implementation, when the first instruction includes instructions with the function type of noise reduction and the adjustment method of reduction, the noise reduction intensity is reduced.
[0015] In one possible implementation, the absolute value of the first signal-to-noise ratio difference and the second signal-to-noise ratio difference is a first preset value.
[0016] In other words, the amount by which the noise reduction intensity is increased or decreased is preset to a fixed value. By quantitatively increasing or decreasing the noise reduction intensity, the noise reduction intensity becomes highly controllable, making it easy for users to operate and further improving the user experience.
[0017] In one possible implementation, the first instruction includes instruction information with the function type being noise reduction and the adjustment method being n%. The signal-to-noise ratio of the first target audio is the sum of the product of the range difference of the preset signal-to-noise ratio range and n%, and the range lower limit of the preset signal-to-noise ratio range.
[0018] In this implementation, when the first instruction includes instruction information with the function type being noise reduction and the adjustment method being n%, the noise reduction intensity is changed to n% of the preset noise reduction intensity range.
[0019] In one possible implementation, the first instruction includes instruction information with function type "de-reverb" and adjustment method "addition", and the first reverb time difference is less than the second reverb time difference; the first reverb time difference is the difference between the reverb time of the first target audio and the reverb time of the first audio, and the second reverb time difference is the difference between the reverb time of the second target audio and the reverb time of the second audio.
[0020] Since reverberation time is positively correlated with reverberation effect, increasing the demeveraging intensity decreases the audio reverberation time. Therefore, the change in reverberation time decreases, meaning the first reverberation time difference is less than the second reverberation time difference. Conversely, a first reverberation time difference less than a second reverberation time difference indicates an increase in demeveraging intensity. Thus, in this implementation, when the first instruction includes instructions with the function type "demeveraging" and the adjustment method "addition," the demeveraging intensity is increased.
[0021] In one possible implementation, the first instruction includes instruction information with function type "de-reverb" and adjustment method "reduction", and the first reverb time difference is greater than the second reverb time difference; the reverb time difference is the difference between the reverb time of the first target audio and the reverb time of the first audio, and the second reverb time difference is the difference between the reverb time of the second target audio and the reverb time of the second audio.
[0022] Since reverberation time is positively correlated with reverberation effect, a decrease in demeveraging intensity leads to an increase in audio reverberation time. Consequently, the change in reverberation time increases, meaning the first reverberation time difference is greater than the second reverberation time difference. Conversely, a first reverberation time difference greater than a second reverberation time difference indicates a decrease in demeveraging intensity. Therefore, in this implementation, when the first instruction includes instructions with the function type "demeveraging" and the adjustment method "addition," the demeveraging intensity is reduced.
[0023] In one possible implementation, the absolute value of the first reverberation time difference and the second reverberation time difference is a second preset value.
[0024] In other words, the amount by which the dereverberation intensity is increased or decreased is preset to a fixed value. By quantitatively increasing or decreasing the dereverberation intensity, the dereverberation intensity becomes highly controllable, facilitating user operation and further improving the user experience.
[0025] In one possible implementation, the first instruction includes instruction information with function type "de-reverb" and adjustment method "n%", and the reverb time of the first target audio is the sum of the product of the range difference of the preset reverb time range and "n%" and the lower limit of the preset reverb time range.
[0026] In this implementation, when the first instruction includes instruction information with function type "de-reverb" and adjustment method "n%", the de-reverb intensity change is achieved to n% of the preset de-reverb intensity range.
[0027] In one possible implementation, the first encoder is a content encoder or a text encoder, and the first instruction feature vector is an instruction embedding.
[0028] Instruction embedding can essentially be text embedding, or content embedding, which contains semantic and data information about the instructions input to the content encoder or text encoder. Instruction embedding allows for a simple and accurate representation of instruction feature vectors.
[0029] In one possible implementation, receiving the first instruction input by the user includes: receiving instruction audio input by the user's voice; recognizing the speech in the instruction audio and converting the recognized speech into text; recognizing the instruction in the text to obtain instruction text; and converting the instruction text into a preset format to obtain the first instruction.
[0030] In this implementation, users can input their first command via voice, which improves the intelligence of the operation and further enhances the user experience.
[0031] In one possible implementation, receiving a first instruction input by the user includes: displaying the interface of a target application; the interface of the target application including the audio to be processed; displaying a first control in response to a first operation by the user on the audio to be processed; and receiving the first instruction input by the user through the first control.
[0032] The first control can be a specific implementation method. Figure 8 or Figure 9 The controls in the noise reduction and de-reverb settings areas.
[0033] In one possible implementation, the target application is a voice recorder application, a note-taking application, a call application, an instant messaging application, or a settings application.
[0034] Secondly, this application provides a model training method, which is executed by an electronic device. The method includes: acquiring multiple audio samples; performing a training process; the training process includes: setting a first instruction sample, encoding the instruction sample through a second encoder to obtain a sample instruction feature vector; inputting the first audio sample as input and the sample instruction feature vector as conditional input into an initial diffusion model; performing noise addition and / or reverberation addition on the first audio sample according to the first instruction sample to obtain a first target audio sample; using the first target audio sample as a target label, performing a forward diffusion process through the initial diffusion model to train the initial diffusion model; the first audio sample is any one of multiple audio samples; using other audio samples from the multiple audio samples as the first audio sample respectively, repeatedly performing the training process until the initial diffusion model converges or the number of training iterations reaches a preset threshold, thereby obtaining an audio processing model.
[0035] The second aspect provides a model training method where clean audio samples serve as the target domain signal, or target distribution, for the diffusion model. The target label after the diffusion model performs a forward diffusion process serves as the target audio sample, or noise distribution. In other words, the initial diffusion model performs a forward diffusion process to transform the target distribution into a noise distribution. During this transformation, the model learns the ability to transform the noise distribution into the target distribution based on conditional inputs; that is, the model learns the ability to transform noisy audio into clean audio. However, the transformation process and direction are constrained by the conditional inputs. Therefore, the resulting audio is not completely clean, but rather audio that matches the effect defined in the conditional inputs. This yields an audio processing model capable of processing audio according to intensity adjustment instructions. When an audio signal and intensity adjustment instructions are input into the audio processing model, the model performs a reverse diffusion process, processing the audio according to the intensity specified by the intensity adjustment instructions to obtain the target audio. This allows for selectable audio effects and improves the quality and listening experience of the target audio, enhancing the user experience.
[0036] In one possible implementation, the first instruction sample includes one or two sets of instruction information. Each set of instruction information includes a function type and a corresponding adjustment method. The function type is noise reduction or dereverberation, and the adjustment method is addition, subtraction, or n%. The function type of noise reduction is used to characterize the audio processing method as noise reduction, and the function type of dereverberation is used to characterize the audio processing method as dereverberation. The adjustment method of addition is used to characterize the increase of processing intensity, the adjustment method of subtraction is used to characterize the decrease of processing intensity, and the adjustment method of n% is used to characterize the adjustment of processing intensity to n% of the preset adjustable range of processing intensity, where n is a value greater than 0 and less than 100.
[0037] In this implementation, the first instruction sample can indicate the addition, subtraction, and setting of the noise reduction intensity to n% of the preset adjustable range of noise reduction intensity, and can also indicate the addition, subtraction, and setting of the dereverberation intensity to n% of the preset adjustable range of dereverberation intensity, providing users with multiple ways to adjust the noise reduction intensity and dereverberation intensity, so that the trained audio processing model can meet the different needs of users and further improve the user experience.
[0038] In one possible implementation, the first audio sample is subjected to noise addition and / or reverberation processing based on the first instruction sample, including: if the first instruction sample includes instruction information with the function type of noise reduction and the adjustment method of addition, then the first audio sample is subjected to noise addition processing; the signal-to-noise ratio (SNR) of the first target audio sample is smaller than the SNR of the first audio sample by a first preset value; if the first instruction sample includes instruction information with the function type of noise reduction and the adjustment method of subtraction, then the first audio sample is subjected to noise addition processing; the SNR of the first target audio sample is larger than the SNR of the first audio sample by a first preset value; if the first instruction sample includes instruction information with the function type of noise reduction and the adjustment method of n%, then the first audio sample is subjected to noise addition processing; the SNR of the first target audio sample is: the product of the range difference of the preset SNR range and n%, and the sum of the lower limit of the preset SNR range.
[0039] In this implementation, when the function type included in the first instruction sample is noise reduction, the first audio sample is subjected to different degrees of noise addition processing according to the adjustment method, so that the signal-to-noise ratio of the first target audio sample matches that of the first instruction sample. In this way, the model can learn to perform noise reduction according to the noise reduction intensity indicated by the instruction, enabling the obtained audio processing model to achieve selectability of noise reduction intensity.
[0040] In one possible implementation, the first audio sample is subjected to noise addition and / or reverberation processing according to the instruction sample, including: if the first instruction sample includes instruction information with function type "de-reverberation" and adjustment method "addition", then the first audio sample is subjected to reverberation processing; the reverberation time of the first target audio sample is greater than the reverberation time of the first audio sample by a second preset value; if the first instruction sample includes instruction information with function type "de-reverberation" and adjustment method "subtraction", then the first audio sample is subjected to reverberation processing; the reverberation time of the first target audio sample is less than the reverberation time of the first audio sample by a second preset value; if the first instruction sample includes instruction information with function type "de-reverberation" and adjustment method "n%", then the first audio sample is subjected to reverberation processing; the reverberation time of the first target audio sample is: the product of the range difference of the preset reverberation time range and "n%", and the sum of the lower limit of the preset reverberation time range.
[0041] In this implementation, when the function type included in the first instruction sample is dreverb, the first audio sample is subjected to different degrees of reverb processing according to the adjustment method, so that the reverb time of the resulting first target audio sample matches that of the first instruction sample. In this way, the model can learn to perform dreverb according to the dreverb intensity indicated by the instruction, enabling the obtained audio processing model to select the dreverb intensity.
[0042] In one possible implementation, the second encoder is a content encoder or a text encoder, and the sample instruction feature vector is the instruction embedding.
[0043] Thirdly, this application provides an apparatus included in an electronic device, which has the function of implementing the behaviors of the electronic device in the first aspect and its possible implementations, and the second aspect and its possible implementations. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a receiving module or unit, a processing module or unit, etc.
[0044] Fourthly, this application provides an electronic device, which includes a processor, a memory, and an interface; the processor, memory, and interface cooperate with each other to enable the electronic device to execute any one of the technical solutions of the first and second aspects.
[0045] Fifthly, this application provides a chip system including a processor. The processor is configured to read and execute a computer program stored in a memory to perform the methods of the first aspect and any possible implementation thereof, and the second aspect and any possible implementation thereof.
[0046] Optionally, the chip system may also include memory, which is connected to the processor via circuitry or wires.
[0047] Alternatively, the chip system may also include a communication interface.
[0048] Sixthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform any one of the methods in the first and second aspects of the technical solutions.
[0049] In a seventh aspect, this application provides a computer program product comprising: computer program code, which, when executed on an electronic device, causes the electronic device to perform any one of the technical solutions of the first and second aspects. Attached Figure Description
[0050] Figure 1 This is an application scenario diagram of an audio processing method provided in an embodiment of this application;
[0051] Figure 2 This is an application scenario diagram of another audio processing method provided in the embodiments of this application;
[0052] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0053] Figure 4 This is a software structure block diagram of an electronic device provided in an embodiment of this application;
[0054] Figure 5 This is a schematic diagram illustrating the working principle of an example diffusion model provided in an embodiment of this application;
[0055] Figure 6 This is a schematic diagram illustrating the principle of an example of training an audio processing model provided in an embodiment of this application;
[0056] Figure 7 This is a flowchart illustrating an example of a model training method provided in an embodiment of this application;
[0057] Figure 8 This is another application scenario diagram of the audio processing method provided in the embodiments of this application;
[0058] Figure 9 This is another application scenario diagram of the audio processing method provided in the embodiments of this application;
[0059] Figure 10 This is a schematic diagram illustrating the reverse diffusion principle of an audio processing model provided in an embodiment of this application;
[0060] Figure 11 This is a flowchart illustrating an example of an audio processing method provided in an embodiment of this application;
[0061] Figure 12 This is another application scenario diagram of the audio processing method provided in the embodiments of this application;
[0062] Figure 13 This is a schematic diagram illustrating the back diffusion principle of an audio processing model provided in an embodiment of this application;
[0063] Figure 14 This is a flowchart illustrating an example of an audio processing method provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0065] Hereinafter, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more of that feature.
[0066] References to "one embodiment" or "some embodiments" as described in this application specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this application specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0067] To better understand the embodiments of this application, the terms or concepts that may be involved in the embodiments are explained below.
[0068] 1. Signal-to-noise ratio (SNR)
[0069] Signal-to-noise ratio (SNR) is the ratio of signal power to noise power, usually expressed in decibels (dB). It is a key indicator of signal quality, used to assess the relative strength of useful information (signal) and interference information (noise) in a signal.
[0070] 2. Reverb and Reverb Time
[0071] In the field of audio, reverberation is an important acoustic phenomenon and audio effect, referring to a series of gradually decaying reflected sounds created by sound reflections within an enclosed space. When a sound source emits sound, the sound is continuously reflected by surfaces such as walls, ceilings, and floors of a room. These reflected sounds mix with the original sound to form the reverberant sound we hear.
[0072] Reverberation time refers to the time required for the sound energy density in a closed space to decay to one millionth of its original value after sound stops emitting. The unit of reverberation time can be seconds (s).
[0073] Reverberation time is inversely proportional to audio clarity and intelligibility. A shorter reverberation time reduces interference from reflected sound, allowing listeners to hear the audio content more clearly; therefore, a shorter reverberation time contributes to improved clarity and intelligibility. Additionally, a certain level of reverberation time can enhance the ambiance of audio. However, excessively long reverberation times may mask details in the original audio, negatively impacting clarity and intelligibility.
[0074] In audio processing software or devices, adjusting reverberation time is a crucial method for altering reverberation effects. Increasing reverberation time can create sound effects reminiscent of a cathedral or large concert hall; conversely, decreasing reverberation time can simulate sound in a small room or near-field environment. If the initial reverberation time is set to 1 second, the sound might have the feel of a medium-sized room. Increasing the reverberation time to 3 seconds will give the sound the characteristics of a large space, as if played in a spacious auditorium. In this embodiment, a longer reverberation time is referred to as a more pronounced reverberation effect, and a shorter reverberation time is referred to as a less pronounced reverberation effect.
[0075] 3. Input and conditional input
[0076] In the field of modeling, input refers to the data provided to the model; this data is the raw material the model processes. The model's input is the direct object on which the model performs calculations, inferences, and other operations; it is also called the primary input. For example, in an image recognition model, the input can be a series of images; in a natural language processing model, the input might be sentences, paragraphs, or other text content. The data types of input are diverse, including numerical (such as temperature and pressure data collected by sensors), text, images, and audio. The model output is mainly generated based on the features of the input data and the model's internal structure (such as the weights of a neural network and the branching rules of a decision tree). Different inputs will lead to different model outputs.
[0077] Conditional inputs, also known as constraints, conditions, or input conditions, are not the primary objects the model processes (i.e., the main inputs) themselves, but rather factors that influence how the model processes the main inputs. Conditional inputs are primarily used to adjust the model's behavior. Compared to regular inputs, conditional inputs emphasize the conditional nature of these input data; they are data used to conditionally limit the model's behavior, state, or output range. Conditional inputs alter the pattern or range of the model's output.
[0078] Before describing the audio processing method and model training method provided in the embodiments of this application, the application scenarios and technical problems faced by this application will be described first.
[0079] During real-time audio acquisition or after audio is saved, electronic devices can perform noise reduction (or noise cancellation) and / or de-reverberation processing on the audio to improve its listening quality and enhance the user experience. For ease of explanation, the scenario of real-time audio acquisition is referred to as the real-time sound pickup scenario, and the scenario of pre-saving audio is referred to as the pre-recording scenario. The real-time sound pickup scenario and the pre-recording scenario are explained separately below.
[0080] 1. Real-time sound pickup scenario
[0081] Currently, many applications (apps) involve noise reduction and / or dereverberation processing of audio in real-time audio pickup scenarios. For example, when recording audio using a recorder app, note-taking app, or call app, noise reduction and / or dereverberation processing is required for the recorded audio; similarly, when sending voice messages through an instant messaging app or recording video using a camera app, noise reduction and / or dereverberation processing is also necessary for the recorded audio.
[0082] For example, Figure 1 This diagram illustrates an application scenario of an audio processing method provided for the implementation of this application. The example given is sending voice messages via an instant messaging app. Figure 1 As shown in Figure (a), the electronic device displays the main interface 101 of an instant messaging app. The main interface 101 of the instant messaging app can display chat options with multiple users. Assuming a user clicks the chat option 102 with user B, the electronic device responds to the user's click and displays the chat interface 103 with user B, as shown below. Figure 1 As shown in Figure (b), the dialogue interface 103 includes a voice input control 104. The user presses the voice input control 104 and inputs a voice message (1). During the user's voice input, the electronic device captures the user's voice input in real time and processes the voice, including but not limited to noise reduction and dereverberation. During the voice input process, the electronic device displays the following: Figure 1 Interface 105 is shown in Figure (c). Afterwards, the user raises their hand, ending the voice input 1. During the voice input process, the electronic device can perform noise reduction and dereverberation processing on the input voice, ultimately completing the processing of all voice inputs to obtain voice 2. Then, the electronic device sends voice 2 to user B, and the electronic device displays as shown... Figure 1 The interface 106 is shown in Figure (d). Speech 2 is obtained after noise reduction and dereverberation processing, resulting in high clarity and intelligibility, with moderate reverberation for a good listening experience. It is evident that noise reduction and dereverberation processing of real-time audio acquisition can improve audio quality and enhance user experience.
[0083] The audio processing during real-time audio pickup using apps such as recorder apps, note-taking apps, and call apps is similar to the process described above and will not be repeated here.
[0084] 2. Pre-record the scene
[0085] In this application embodiment, noise reduction and / or de-reverberation processing can also be performed on pre-recorded audio stored in apps such as recorder apps and note-taking apps, or audio from pre-shot videos stored in photo album apps.
[0086] For example, Figure 2 This is a schematic diagram illustrating another application scenario of the audio processing method provided in this application embodiment. The example used is audio stored in a recorder app. Figure 2 As shown in Figure (a), the main interface of the electronic device may include an icon 201 for a recorder app. When the user clicks the recorder app icon 201, the electronic device displays the main interface 202 of the recorder app, as shown in Figure (a). Figure 2 As shown in Figure (b) above, the main interface 202 of the recorder app includes multiple recordings. Taking recording 3 as an example, when the user long-presses recording 3, the electronic device displays the following in response to the user's operation: Figure 2 The interface 203 is shown in Figure (c). Interface 203 includes multiple processing options for recording 3, including a noise reduction option 204 and a de-reverb option 205. In response to the user clicking the noise reduction option 204, the electronic device can perform noise reduction processing on recording 3. During the processing, the electronic device can display... Figure 2 The interface is shown in Figure (d). Similarly, in response to the user clicking the de-reverb option 205, the electronic device can perform de-reverb processing on the recording 3, as shown in the specific interface. Figure 2 (Not shown in the image). After processing, the electronic device can save the processed audio again. Noise reduction and / or dérarization of the stored audio can further improve its clarity and intelligibility, alter reverberation effects, and make the audio sound more satisfying to the user, thus enhancing the user experience.
[0087] For audio noise reduction and dereverberation, related technologies generally employ signal-based processing methods or neural network (NN)-based methods. Signal-based methods use fixed parameters, resulting in relatively uniform audio effects. NN-based methods typically use a fixed NN model, also leading to relatively uniform audio effects. However, different users have different needs for noise reduction and dereverberation. Some users prefer very clean noise reduction and thorough dereverberation, even to the point of over-elimination, which may result in incomplete speech. Others prioritize speech integrity, accepting some residual noise or noticeable reverberation rather than incomplete speech. Therefore, audio processing methods based on fixed parameters or a fixed NN model often fail to meet user needs, resulting in a poor user experience.
[0088] To address the aforementioned issues, a weighted summation-based audio processing method exists in related technologies. This method includes: saving the audio signal to be processed; processing the audio signal separately based on different noise parameters and different reverberation parameters to obtain multiple audio samples with different noise residues and different reverberation effects; selecting the audio sample with the corresponding noise residue and the audio sample with the corresponding reverberation effect according to user needs, and weighting and summing them proportionally to fuse the two audio samples to obtain the processed audio. This method can relatively easily achieve selectability of audio effects; however, since this method involves signal fusion, the quality of the processed audio is poor, affecting the user's listening experience and failing to significantly improve the user experience.
[0089] Based on this, this application provides an audio processing method. A user can input an intensity adjustment command into an electronic device, which instructs the user to adjust the noise reduction intensity and / or dereverberation intensity. The electronic device encodes the user-input intensity adjustment command to obtain a command feature vector. Then, the audio to be processed is input into an audio processing model, and the command feature vector is used as a conditional input to a diffusion model. The audio processing model is a model trained based on an initial diffusion model. The audio processing model performs a back-diffusion process, performing noise reduction and / or dereverberation processing on the audio according to the processing intensity indicated by the intensity adjustment command, to obtain the target audio. Compared to the audio to be processed, the target audio has the same noise reduction intensity and dereverberation effect as indicated by the intensity adjustment command. This method allows users to customize the audio effect by selecting the noise reduction and dereverberation intensity according to their needs; moreover, this method, based on a diffusion model, produces higher quality target audio with a better listening experience, effectively improving the user experience.
[0090] Furthermore, this application also provides a model training method. This method, based on multiple audio samples and multiple intensity adjustment command samples, trains an initial diffusion model by performing a forward diffusion process, thereby obtaining an audio processing model. Through training, the model learns the ability to perform audio noise reduction and / or dereverberation based on intensity adjustment command samples. Thus, an audio processing model capable of improving the quality and listening experience of the target audio is obtained, enhancing the user experience.
[0091] The structure of the electronic device to which the audio processing method provided in the embodiments of this application is applicable will be described below.
[0092] The audio processing method provided in this application can be applied to electronic devices capable of receiving voice signals, such as mobile phones, tablets, wearable devices, recording devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.
[0093] For example, Figure 3 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application. The electronic device 100 may include a processor 110, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a sensor module 180, a display screen 194, etc. The sensor module 180 may include a touch sensor 180K.
[0094] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0095] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0096] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0097] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0098] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0099] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may include submodules for noise reduction and / or dereverberation processing of the audio. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0100] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0101] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0102] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0103] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0104] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0105] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.
[0106] Figure 4 This is a software structure block diagram of an electronic device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.
[0107] like Figure 4 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS (not shown in the figure).
[0108] In this embodiment, the application layer may further include an audio processing module. Optionally, the audio processing module may be a separate app or a module within an app, such as a module within a recorder app or an instant messaging app.
[0109] like Figure 4 As shown, the audio processing module may include a first encoder and an audio processing model. The first encoder encodes the intensity adjustment command input by the user to obtain a command feature vector. The first encoder uses the command feature vector as conditional input and inputs it into the audio processing model. Other modules in the application layer input the audio to be processed into the audio processing model. The audio processing model performs a backdiffusion process, performs noise reduction and / or reverberation processing on the audio to be processed according to the intensity adjustment command, and outputs the target audio.
[0110] As one possible implementation, the audio processing module may further include an automatic speech recognition (ASR) unit, an instruction recognition unit, and an instruction conversion unit. The ASR unit is used to recognize speech in the audio and convert it into text. The instruction recognition unit is used to extract keywords from the text output by the ASR unit and match the extracted keywords with preset instruction words to recognize intensity adjustment instructions in the text. The instruction conversion unit is used to convert the intensity adjustment instructions into a preset format to obtain intensity adjustment instructions that can be recognized by the first encoder. The instruction conversion unit then inputs the intensity adjustment instructions into the first encoder.
[0111] As an optional implementation, the audio processing module may also include a model training module. The model training module is used to train the initial diffusion model to obtain the aforementioned audio processing model. Optionally, the model training module may include a data loading unit, an instruction setting unit, an audio enhancement unit, a second encoder, the initial diffusion model, and a training algorithm unit.
[0112] The data loading unit loads a sample library, retrieves clean audio samples from it, and inputs these clean audio samples into the initial diffusion model and the audio enhancement unit. The instruction setting unit sets intensity adjustment instruction samples and inputs them into the second encoder and the audio enhancement unit. The audio enhancement unit enhances the clean audio samples based on the intensity adjustment instruction samples to obtain the target audio sample. The audio enhancement unit inputs the target audio sample into the training algorithm unit. The second encoding unit encodes the intensity adjustment instruction samples to obtain instruction feature vector samples. The second encoder inputs the instruction feature vector samples as conditional input into the initial diffusion model. The initial diffusion model performs a forward diffusion process and outputs predicted audio. The initial diffusion model inputs the predicted audio into the training algorithm unit. The training algorithm unit uses the target audio sample as the target label, calculates the loss between the predicted audio and the target label, and adjusts the parameters of the initial diffusion model based on the loss.
[0113] Optionally, in other embodiments, the electronic device may not include a model training module. Instead, the audio processing model is trained on another device and then imported into the electronic device. This application does not limit this aspect.
[0114] It is understandable that applications and modules at the application layer can implement their functions by calling the relevant modules at the underlying level.
[0115] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0116] like Figure 4 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0117] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0118] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0119] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.
[0120] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0121] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0122] A system library can include multiple functional modules. For example, media libraries.
[0123] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0124] The kernel layer is the layer between hardware and software. The kernel layer contains at least an audio driver. This audio driver can include audio input drivers and audio output drivers. The audio input driver is also called the microphone driver. The microphone driver is used to identify and drive the microphone, transmitting the microphone input signal to the upper layer. The microphone resides in the hardware layer of the electronic device. Figure 4 Not shown in the image.
[0125] The following embodiments of this application will be used to illustrate having Figure 3 and Figure 4 Taking the electronic device with the structure shown as an example, and in conjunction with the accompanying drawings and application scenarios, the audio processing method provided in this application embodiment will be specifically described.
[0126] To facilitate understanding, we will first introduce the diffusion model.
[0127] The diffusion model is a type of deep neural network (DNN).
[0128] Alternatively, the diffusion model framework can adopt a denoising diffusion probabilistic model (DDPM), a score-based generative model (SGM), or a stochastic differential equation (SDE), etc.
[0129] Optionally, the model structure of the diffusion model can be a convolutional neural network, such as a U-NET, a gradient-based noise conditional score network (NCSN), or an upgraded version of NCSN (NCSN++).
[0130] A diffusion model is a generative model that, given independent and identically distributed sample data, learns to approximate an unknown data distribution. The sample data originates from location data distributions. The use of diffusion models involves forward diffusion and reverse diffusion processes.
[0131] Forward diffusion can also be called forward diffusion or diffusion process. Backward diffusion can also be called inference process, reverse diffusion process, abstraction process, or denoising process. In forward diffusion, Gaussian noise is added to the data step by step in an orderly manner to generate a Markov chain of the data. In backward diffusion, noise is removed from a noisy data point step by step to generate a Markov chain of the noisy data.
[0132] A Markov chain consists of a series of states and a series of change probabilities. Here, a state refers to data with different noise levels, and a change probability refers to the probability of changing from the current state to the next state, which is implemented using a change matrix.
[0133] For example, such as Figure 5 As shown in Figure (a), for data x0, Gaussian noise is gradually added during the forward diffusion process. The data obtained at step t-1 is x. t-1 In the t-th step of the T steps (total steps) of the forward diffusion process, data x is... t-1 Add a small amount of Gaussian noise to obtain data x. t Data x t Let represent the data obtained after t steps. Here, t = 1, 2, ..., T, where T is a positive integer. That is, for data x0, after T steps, the data obtained in each step are x1, x2, ..., xt, respectively. TIn the forward diffusion process, the parameter t can represent the number of steps to add Gaussian noise, or the number of iterations.
[0134] like Figure 5 As shown in Figure (a), the reverse diffusion process is the opposite of the forward diffusion process, gradually removing noise. In the reverse diffusion process, the parameter t can be understood as the remaining number of iterations. For data x... T After T steps, the data obtained in each step are x. T x T-1 , ..., x1.
[0135] like Figure 5 As shown in Figure (b), taking image data as an example, during the forward diffusion process, Gaussian noise is gradually added to the image data, causing the image to gradually blur. During the reverse diffusion process, the noise gradually decreases, and the image gradually becomes clearer until the original image data is restored.
[0136] like Figure 5 As shown in Figure (c), taking audio data as an example, during the forward diffusion process, Gaussian noise is gradually added to the clean audio data, and the noise in the audio data gradually increases. During the reverse diffusion process, the noise in the audio gradually decreases until the original clean audio data is restored.
[0137] The forward diffusion process can be represented by formula (1):
[0138] q(x t |x t-1 ,y)(1)
[0139] Where q(x) t |x t-1 ,y) represents the expression of data x0 after the t-th iteration, and y represents the conditional input.
[0140] It can be understood that the data x0 after the tth iteration can be directly calculated from the data x0, as shown in the following formula (2):
[0141] q(x t |x t-1 ,y)=(x t |x0,y)=N c (x t ,μ(x0,y,t),σ(t) 2 I)(2)
[0142] Where, N c Let μ(x0,y,t) represent the Gaussian distribution, μ(x0,y,t) represent the mean of the Gaussian distribution, and σ(t) represent the standard deviation of the Gaussian distribution. 2 Let I represent the variance of the Gaussian distribution, and let I represent the identity matrix.
[0143] The mean value μ(x0,y,t) of the Gaussian distribution can be determined based on the condition y, the parameter t used to represent the number of iterations, and the normalization constant γ. The mean value μ(x0,y,t) of the Gaussian distribution can be expressed as formula (3):
[0144] μ(x0,y,t)=e -γt x0+(1-e -γt )y(3)
[0145] The standard deviation σ(t) of the Gaussian distribution can be determined based on the parameter t used to represent the number of iterations, the normalization constant γ, and the preset maximum variance σ. max and the preset minimum variance σ min Confirmed. The square of the standard deviation σ(t) of a Gaussian distribution is the variance of the Gaussian distribution. 2 This can be expressed as formula (4):
[0146]
[0147] According to data x t The expression at time t can determine the data x. t .
[0148] The forward diffusion process can be understood as the training process of the diffusion model. The initial diffusion model is used to process training samples with parameters t = 1, 2, ..., T to obtain denoised training data. Then, the denoised training data is compared with the data x. t-1 The differences between the data are used to adjust the parameters of the initial diffusion model. The training samples for the initial diffusion model include data x. t The diffusion model is the initial diffusion model after parameter adjustment.
[0149] The training data for denoising can be the output of the initial diffusion model, or it can be data x. t It is obtained by removing noise from the output representation of the initial diffusion model.
[0150] The training data for denoising is data x. t Based on the noise removal from the output representation of the initial diffusion model, and then based on the training denoised data and data x... t-1 The difference between the parameters of the initial diffusion model and the data x can be understood as the difference between the output of the initial diffusion model and the data x. t relative data x t-1 The parameters of the initial diffusion model are adjusted based on the differences between the added noise.
[0151] Based on the output of the initial diffusion model and the data x when t = 1, 2, ..., T respectively... t-1The difference between them is used to adjust the parameters of the initial diffusion model, which can be expressed as formula (5):
[0152]
[0153] Among them, s θ This represents the output of the initial diffusion model. Represents data x t-1 Compared to data x t The noise added, arg θ Let || be a preset constant, and || denote the norm. It represents the expectation of a plurality of t for the case t = 1, 2, ..., T. Norms are commonly used to measure the length or size of each vector in a vector space (or matrix).
[0154] The reverse diffusion process can be understood as the reasoning process of the diffusion model. The reverse diffusion process can be expressed as formula (6):
[0155] p θ (x t-1 |x t ,y)(6)
[0156] Based on the above formulas for representing the forward and reverse diffusion processes, the data x0 can be represented as q(x0), and the data x obtained after T iterations can be... T Represented as p θ (x T ), where p θ (x T ) = N c (0, I).
[0157] As a high-quality generative model, the diffusion model generates high-quality speech signals with a good listening experience. Moreover, the diffusion model has good robustness.
[0158] The audio processing method and model training method provided in the embodiments of this application will be described next. The model training method is the training process of the audio processing model, and the audio processing method is the process of using the audio processing model. For ease of distinction, the word "sample" can be added before or after some names involved in the model training process, such as target audio and instruction feature vectors, to indicate that they are data from the model training process. It should be understood that these sample data and the corresponding non-sample data during model usage can be essentially the same.
[0159] First, we will introduce the training process of the audio processing model.
[0160] For example, Figure 6This is a schematic diagram illustrating the principle of training an audio processing model according to an embodiment of this application. Figure 7 Please refer to the flowchart of an example model training method provided in this application embodiment. Figure 6 and Figure 7 The method includes the following steps S101 to S111. The entity performing the following steps can be... Figure 4 The model training module in the application layer.
[0161] S101, The data loading unit loads the sample library, which includes multiple clean audio samples.
[0162] Clean audio samples are also called clean signals. It is understood that the term "clean audio" in this application embodiment is a relative concept, referring to an audio signal that is almost free of noise and has almost no reverberation. Clean audio samples have a higher SNR and a shorter reverberation time, closer to 0.
[0163] It should be noted that in some other embodiments, the audio samples in the sample library may not be clean audio samples. Using clean audio samples to train the model makes it easier for the model to converge and the accuracy of the trained model is also higher.
[0164] S102, the data loading unit takes any clean audio sample A from the sample library as input and inputs it into the initial diffusion model and the audio enhancement unit respectively.
[0165] S103, Instruction setting unit sets intensity adjustment instruction sample.
[0166] The intensity adjustment command sample is sample data of the intensity adjustment command. The following is an explanation of the intensity adjustment command; the intensity adjustment command sample and the intensity adjustment command are essentially the same.
[0167] Intensity adjustment commands are used to instruct adjustments to the noise reduction intensity during audio noise reduction processing, and / or, to instruct the dereverberation intensity during audio demeverberation processing, thereby altering the audio effect. Noise reduction intensity, also known as noise reduction depth or noise reduction strength, can be characterized by changes in noise levels in the audio (e.g., changes in noise power) or changes in the noise reduction effect (e.g., changes in SNR). Dereverberation intensity, also known as dereverberation depth or dereverberation strength, can be characterized by changes in the room spectrum of the audio or changes in the dereverberation effect (e.g., changes in reverberation time). Noise reduction intensity and dereverberation intensity can be collectively referred to as the audio processing intensity.
[0168] Optionally, each intensity adjustment command may include one or two sets of command information. Each set of command information includes a function type and the corresponding adjustment method. The function type indicates the method of audio processing, such as noise reduction or dereverberation. The adjustment method indicates how the processing intensity is changed. The adjustment method determines the processing intensity.
[0169] Optionally, the adjustment method can be plus, minus, or percentage. The plus adjustment method indicates a further increase in processing intensity based on the current processing intensity. Specifically, the intensity adjustment command includes a function type of noise reduction and an adjustment method of plus, indicating an increase in noise reduction intensity based on the current noise reduction intensity; the intensity adjustment command also includes a function type of demeverage and an adjustment method of plus, indicating an increase in demeverage intensity based on the current demeverage intensity. Optionally, the increase in noise reduction intensity can be a preset fixed value.
[0170] The adjustment method is "decrease," indicating a reduction in processing intensity based on the current processing strength. Specifically, the intensity adjustment command includes instructions with the function type "noise reduction" and the adjustment method "decrease," indicating a reduction in noise reduction intensity based on the current noise reduction strength; the intensity adjustment command also includes instructions with the function type "de-reverberation" and the adjustment method "decrease," indicating a reduction in reverberation intensity based on the current de-reverberation strength. Optionally, the reduction value of the noise reduction intensity can be a preset fixed value.
[0171] The adjustment method is a percentage, used to indicate that the processing intensity is set to a certain percentage of the preset adjustable range. Specifically, taking an adjustment method of n% as an example (n is a value greater than 0 and less than 100), the intensity adjustment command includes instruction information with function type "noise reduction" and adjustment method "n%", indicating that the noise reduction intensity is adjusted to the noise reduction intensity corresponding to n% of the preset adjustable range. Specifically, based on the lower limit of the preset adjustable range, the range difference of the adjustable range is increased by n%. The intensity adjustment command also includes instruction information with function type "de-reverberation" and adjustment method "n%", indicating that the de-reverberation intensity is adjusted to the reverberation intensity corresponding to n% of the preset adjustable range. Specifically, based on the lower limit of the preset adjustable range, the range difference of the adjustable range is increased by n%. In this embodiment, both the adjustable range of noise reduction intensity and the adjustable range of de-reverberation intensity can be set according to actual usage requirements. The range difference refers to the difference between the upper and lower limits of the adjustable range.
[0172] For example, Table 1 lists some samples of intensity adjustment instructions and the instruction information they contain.
[0173] Table 1
[0174]
[0175] It should be noted that Table 1 is only an example and is not intended to limit the scope of the information.
[0176] Optionally, different information in the intensity adjustment command can be represented by the values of different fields. For example, the intensity adjustment command may include a function type field, an add field, a subtract field, and a percentage field. A function type field value of preset value 1 (e.g., 0) indicates that the function type is noise reduction; a function type field value of preset value 2 (e.g., 1) indicates that the function type is de-reverb. An add field value of preset value 3 (e.g., true) indicates that the adjustment method is add; an add field value of preset value 4 (e.g., false) indicates that the adjustment method is not add. A subtract field value of preset value 3 (e.g., true) indicates that the adjustment method is subtract; a subtract field value of preset value 4 (e.g., false) indicates that the adjustment method is not subtract. A percentage field value of preset value 5 (e.g., false) indicates that the adjustment method is not percentage; a percentage field value other than preset value 5 indicates that the adjustment method is the corresponding value of that field. For example, a percentage field value of 50% indicates that the adjustment method is 50%.
[0177] Optionally, the instruction setting unit can randomly set intensity adjustment instruction samples, or it can set intensity adjustment instruction samples according to a certain pattern. For different clean audio samples, the intensity adjustment instruction samples set by the instruction setting unit can be the same or different.
[0178] S104, the instruction setting unit inputs intensity adjustment instruction samples to the second encoder and the audio enhancement unit respectively.
[0179] S105. The second encoder encodes the intensity adjustment instruction sample, extracts the feature vector from the intensity adjustment instruction sample, and obtains the sample instruction embedding.
[0180] Optionally, the instruction setting unit can concatenate the function type and adjustment method from the intensity adjustment instruction sample and input them into the second encoder. For example, if the intensity adjustment instruction sample is Intensity Adjustment Instruction Sample 1 in Table 1, the function type "Noise Reduction" and the adjustment method "60%" from Intensity Adjustment Instruction Sample 4 can be concatenated, and "De-reverberation" and "Add" can be concatenated and input into the second encoder. After the second encoder encodes the concatenated information, the output sample instruction embedding also includes two parts, one part representing the function type and the other part representing the adjustment method.
[0181] An encoder is a tool or component that transforms or encodes input information according to specific rules. The input to an encoder can include signals, images, positions, text, etc. The encoder encodes the input information, converting it into a form that a computer can understand and process, and then outputs it. Optionally, the encoder's output can be a vector.
[0182] Specifically, in the embodiments of this application, the first encoder and the second encoder can be content encoders or text encoders. The first encoder and the second encoder can encode the intensity adjustment command, extract the semantics and data from the intensity adjustment command, and obtain a command form that the diffusion model can understand and process. Optionally, after encoding the intensity adjustment command, the first encoder and the second encoder can output a command embedding. The command embedding is a vector representation of the command. The command embedding contains information such as the semantics (e.g., function type, adjustment method) and data (e.g., percentage value) of the command input to the encoder. Optionally, the command embedding can be represented as a fixed-length vector or as a distributed representation.
[0183] It should be noted that in other embodiments, the data output by the first encoder and the second encoder can also be represented in other forms, as long as the intensity adjustment command can be quantized and described in vector form. The specific representation can be determined based on the principle and structure of the encoder used. For ease of understanding and description, in some embodiments of this application, the results obtained by the first encoder and the second encoder after encoding the command can also be collectively referred to as the command feature vector.
[0184] Optionally, the encoding methods used by the first encoder and the second encoder may include, but are not limited to: one-hot encoding, word2vec, global vectors for word representation (GloVe), fast text classifier, etc.
[0185] In this embodiment, the instruction embedding output by the second encoder is referred to as the sample instruction embedding.
[0186] S106. The second encoder takes the sample instruction embedding as a conditional input and inputs it into the initial diffusion model.
[0187] S107. The initial diffusion model performs a forward diffusion process and outputs the predicted audio A.
[0188] Specifically, a clean audio sample A is used as the input to the initial diffusion model, and the sample instruction embedding is used as the conditional input to the initial diffusion model. Based on this input and the conditional input, the initial diffusion model performs the following forward diffusion process:
[0189] Using a clean audio sample A as the starting point of the forward diffusion process (i.e., clean audio sample A as data x0), and referring to the above formulas (1) to (4), Gaussian noise is added to the current speech signal at each step t to simulate noise addition and reverberation. Here, the conditional input y is the sample instruction embedding. As the number of iterations increases, the noise features in the clean audio sample A gradually increase, the SNR gradually decreases, and the reverberation duration gradually increases, making the reverberation effect increasingly obvious. Moreover, during the iteration process, the intensity of noise addition and reverberation is limited by the sample instruction embedding. After T iterations, the predicted audio A is finally obtained.
[0190] The speech output by the initial diffusion model is called the predicted audio. Specifically, in this embodiment, when the input is audio sample A, the audio output by the initial diffusion model is called the predicted audio A.
[0191] S108, The initial diffusion model will predict audio A input to train the algorithm unit.
[0192] S109. The audio enhancement unit enhances the clean audio sample A based on the intensity adjustment instruction sample to obtain the target audio sample A.
[0193] The target audio sample is noisy audio. Noisy audio, also known as a noisy signal or a frequency with noise, refers to audio that contains noise and / or has a reverberation effect (i.e., the reverberation time is not zero).
[0194] The audio enhancement unit adds noise and / or reverberation to a clean audio sample A, resulting in noisy audio. The noise addition intensity matches the noise reduction intensity indicated in the intensity adjustment sample instruction, and the reverberation intensity matches the de-reverberation intensity indicated in the intensity adjustment sample instruction. This allows the model to learn the ability to process noisy audio according to the noise reduction and / or de-reverberation intensity, thereby obtaining a cleaner audio sample with a corresponding effect.
[0195] Optionally, the noise reduction effect can be characterized by SNR, and the noise reduction intensity can be characterized by the change in SNR (denoted as ΔSNR). ΔSNR can be understood as the difference between the SNR of the audio after noise reduction and the SNR of the audio before noise reduction. Optionally, the noise reduction intensity can be preset to an adjustable range [a0, a1], that is, the preset adjustable range of ΔSNR is [a0, a1]. In other words, during the noise reduction process, after the minimum noise reduction intensity, the change in the audio's SNR is a0, and after the maximum noise reduction intensity, the change in the audio's SNR is a1. Additionally, the preset step value 'a' for increasing or decreasing the noise reduction intensity corresponding to ΔSNR can be set, where 0 < a < a1 - a0. The specific values of a0, a1, and a can be set according to actual needs; for example, a0 can be -10dB, a1 can be 30dB, and a can be 5dB. The audio enhancement unit can perform noise addition processing on the audio based on the ΔSNR step value and adjustable range.
[0196] It can be understood that SNR is negatively correlated with noise power. Increasing noise reduction intensity reduces noise level and noise power, thus increasing SNR; conversely, decreasing noise reduction intensity increases noise level and noise power, thus decreasing SNR. The instruction sample for intensity adjustment, with function type "noise reduction" and adjustment method "addition," indicates an increase in noise reduction intensity. Therefore, in actual use, each time the noise reduction intensity is increased, ΔSNR increases by 'a'. For example, for a certain audio segment, if the signal-to-noise ratio (SNR) before noise reduction is SNR0, and the SNR of the audio obtained after noise reduction based on the current noise reduction intensity (i.e., before the noise reduction intensity was increased) is SNR1, then the difference in SNR between the audio before and after processing is ΔSNR1 = SNR1 - SNR0. If the noise reduction intensity is increased once, the signal-to-noise ratio (SNR) of the audio obtained after noise reduction based on the increased intensity is SNR2. The difference in SNR between the audio before and after processing is ΔSNR2 = SNR2 - SNR0. Therefore, ΔSNR2 is 'a' greater than ΔSNR1, i.e., ΔSNR2 = ΔSNR1 + a. Thus, SNR2 - SNR0 = SNR1 - SNR0 + a, and therefore, SNR2 = SNR1 + a. During training, the audio enhancement module's processing is essentially the reverse of the noise reduction process. Based on this, during training, when enhancing a clean audio sample A, if the intensity adjustment instruction contains a function type of noise reduction and an adjustment method of addition, the audio enhancement unit adds noise to the clean audio sample A, reducing the SNR of the target audio sample A by 'a', i.e., SNR(target audio sample A) = SNR(clean audio sample A) - a.
[0197] Similarly, in the intensity adjustment instruction sample, the function type is noise reduction and the adjustment method is decrease, indicating that the noise reduction intensity is reduced. Therefore, in actual use, each time the noise reduction intensity is reduced, the audio's ΔSNR decreases by 'a'. For example, for a certain audio segment, the signal-to-noise ratio (SNR) before noise reduction is SNR0. If the SNR of the audio obtained after noise reduction based on the current noise reduction intensity (i.e., before the noise reduction intensity increases) is SNR1, the difference in SNR between the audio before and after processing is ΔSNR1 = SNR1 - SNR0. If the noise reduction intensity is reduced once, the SNR of the audio obtained after noise reduction based on the reduced intensity is SNR2, and the difference in SNR between the audio before and after processing is ΔSNR2 = SNR2 - SNR0 (ΔSNR2 is also called the second SNR difference). Therefore, ΔSNR2 is 'a' less than ΔSNR1, i.e., ΔSNR2 = ΔSNR1 - a. Then, SNR2 - SNR0 = SNR1 - SNR0 - a, therefore, SNR2 = SNR1 - a. During training, the audio enhancement module's processing is essentially the reverse of the noise reduction process. Therefore, during training, when enhancing a clean audio sample A, if the intensity adjustment command contains a function type of noise reduction and an adjustment method of reduction, the audio enhancement unit will add noise to the clean audio sample A, increasing the signal-to-noise ratio (SNR) of the target audio sample A by 'a', i.e., SNR(target audio sample A) = SNR(clean audio sample A) + a.
[0198] If the intensity adjustment instruction sample contains instruction information with function type "noise reduction" and adjustment method "n%", in actual use, the audio's ΔSNR will be increased by n% of the range difference based on the lower limit of the preset adjustable range of ΔSNR, that is, ΔSNR = (a1 - a0) * n% + a0. Based on this, during the training process, the audio enhancement unit adds noise to the clean audio sample A, so that the signal-to-noise ratio (SNR) of the target audio sample A is (target audio sample A) = (a1 - a0) * n% + a0.
[0199] Optionally, the reverberation effect can be characterized by reverberation time, and the de-reverberation intensity can be represented by the change in reverberation time (Δrev_t). Δrev_t can be understood as the difference between the reverberation time of the audio after de-reverberation and the reverberation time of the audio before de-reverberation. Specifically, the de-reverberation intensity can be preset to an adjustable range [t0, t1], that is, the adjustable range of Δrev_t can be preset to [t0, t1]. In other words, during the de-reverberation process, after the minimum de-reverberation intensity, the change in the audio's reverberation time is t0, and after the maximum de-reverberation intensity, the change in the audio's reverberation time is t1. In addition, the step value t for increasing or decreasing the de-reverberation intensity corresponding to Δrev_t can be preset, where 0 < t < t1 - t0. The specific values of t0, t1, and t can be set according to actual needs; for example, t0 can be 0.4s, t1 can be 10s, and t can be 0.96s. The audio enhancement unit can add reverberation to audio based on the reverberation time step value and adjustable range.
[0200] It's understandable that reverberation time and reverberation effect are positively correlated: increasing de-reverberation intensity decreases reverberation time, and vice versa. The instruction information in the intensity adjustment command, with the function type being de-reverberation and the adjustment method being "add," indicates an increase in de-reverberation intensity. Therefore, in actual use, each time the de-reverberation intensity is increased, the audio's Δrev_t decreases by t. For example, for a certain audio segment, the reverberation time of the signal before de-reverberation is rev_t0. If the reverberation time of the audio obtained by de-reverberation based on the current de-reverberation intensity (i.e., before the increase in de-reverberation intensity) is rev_t1, the difference in reverberation time between the two audio segments is Δrev_t1 = rev_t1 - rev_t0. The dreverb intensity is increased once. The reverb time of the audio obtained by dreverberation based on the increased dreverb intensity is rev_t2. The difference in reverb time between the audio before and after processing is Δrev_t2 = rev_t2 - rev_t0. Therefore, Δrev_t2 is t less than Δrev_t1, that is, Δrev_t2 = Δrev_t - t. Then, rev_t2 - rev_t0 = rev_t1 - rev_t0 - t, therefore, rev_t2 = rev_t1 - t. During training, the processing of the audio enhancement module is essentially the reverse of the dreverb processing. Based on this, during training, when enhancing a clean audio sample A, if the intensity adjustment instruction contains the instruction information of function type dereverb and adjustment method add, then the audio enhancement unit adds reverb to the clean audio sample A, so that the reverb time of the target audio sample A is increased by t compared to the reverb time of the clean audio sample A, that is: rev_t(target audio sample A) = rev_t(clean audio sample A) + t.
[0201] Similarly, in the intensity adjustment instruction sample, the function type is de-reverb, and the adjustment method is decrease, indicating that the de-reverb intensity is reduced. Therefore, in actual use, each time the de-reverb intensity is reduced, the audio's Δrev_t decreases by t. For example, for a certain audio segment, the reverb time of the signal before de-reverb is rev_t 0. If the reverb time of the audio obtained by de-reverb based on the current de-reverb intensity (i.e., before the de-reverb intensity increases) is rev_t1, the difference in reverb time between the audio before and after processing is Δrev_t1 = rev_t1 - rev_t0. The dereverberation intensity is reduced once. The reverberation time of the audio obtained by dereverberation based on the reduced dereverberation intensity is rev_t2. The difference in reverberation time between the audio before and after processing is Δrev_t2 = rev_t2 - rev_t0. Therefore, Δrev_t2 is t greater than Δrev_t1, that is, Δrev_t2 = Δrev_t + t. Thus, rev_t2 - rev_t0 = rev_t1 - rev_t0 + t, and therefore rev_t2 = rev_t1 + t. Based on this, during training, when enhancing a clean audio sample A, if the intensity adjustment instruction contains the instruction information of function type dereverberation and adjustment method reduction, the audio enhancement unit will add reverberation to the clean audio sample A, so that the reverberation time of the target audio sample A is reduced by 'a' compared to the reverberation time of the clean audio sample A, that is: rev_t(target audio sample A) = rev_t(clean audio sample A) - t.
[0202] If the intensity adjustment instruction sample contains instruction information with function type "de-reverb" and adjustment method "n%", in actual use, the audio Δrev_t will be increased by n% of the range difference based on the lower limit of the preset adjustable range of Δrev_t, that is, Δrev_t = (t1-t0)*n% + t0. Based on this, during the training process, the audio enhancement unit adds reverb to the clean audio sample A, so that the reverb time of the target audio sample A is rev_t(target audio sample A) = (t1-t0)*n% + t0.
[0203] Optionally, the noise addition process can be specifically performed by superimposing the spectrum of the clean audio sample A with the noise signal corresponding to the SNR. The reverberation addition process can be specifically performed by multiplying the spectrum of the clean audio sample A with the room impulse response corresponding to the reverberation time.
[0204] S110, the audio enhancement unit inputs the target audio sample A into the training algorithm unit.
[0205] S111 The training algorithm unit adjusts the parameters of the initial diffusion model based on the loss of the predicted audio A and the target audio sample A.
[0206] Optionally, a loss function for the initial diffusion model can be preset. The loss function measures the loss (i.e., difference) between the predicted audio output by the initial diffusion model and the real audio (i.e., the target label) during the forward diffusion process. Specifically, in this step, the target audio sample A can be understood as a simulation of real noisy audio. The predicted audio A and the target audio sample A can be substituted into the preset loss function to calculate the loss value. Based on the magnitude of the loss value, the parameters of the initial diffusion model are adjusted according to formula (5), such as the weights of the neural network, to minimize the loss. The process of adjusting the parameters of the initial diffusion model to minimize the loss is also the process of the model's parameters gradually approaching the target value. This process can be understood as the model learning its ability to denoise and / or dereverberate audio.
[0207] In general, the above process starts with clean audio samples, uses sample instruction embedding as conditional input, takes target audio sample A as the endpoint (target label), performs a forward diffusion process through the initial diffusion model, trains the diffusion model, and constrains the direction of forward diffusion by sample instruction embedding, coupling sample instruction embedding with audio, so that the model learns how to process audio according to instruction embedding.
[0208] Optionally, the initial diffusion model can fuse the conditional inputs in various ways. For example, the conditional inputs can be concatenated with the inputs, or they can be concatenated and embedded into the bottleneck layer of the diffusion model. This application does not limit this approach and the choice can be made according to actual needs.
[0209] It is understandable that, based on each clean audio sample in the audio sample library, steps S102 to S111 are repeated until the loss is less than a preset threshold, i.e., the model converges, or the number of training iterations reaches a preset number, at which point training stops. The resulting diffusion model is called the audio processing model.
[0210] In summary, during model training, clean audio samples serve as the target domain signal, or target distribution, for the diffusion model. The target labels after the diffusion model performs a forward diffusion process are the target audio samples, or noise distribution. That is, the initial diffusion model performs a forward diffusion process to transform the target distribution into a noise distribution. During this transformation, the model learns the ability to transform the noise distribution into the target distribution based on conditional inputs; in other words, the model learns the ability to transform noisy audio into clean audio. However, the transformation process and direction are constrained by the conditional inputs. Therefore, the resulting audio is not completely clean, but rather audio that matches the effect specified in the conditional inputs. This results in an audio processing model capable of processing audio according to intensity adjustment instructions. When an audio signal and intensity adjustment instructions are input into the audio processing model, the model performs a reverse diffusion process, processing the audio according to the intensity specified by the intensity adjustment instructions to obtain the target audio. This allows for selectable audio effects, improves the target audio quality and listening experience, and enhances the user experience.
[0211] Next, the process of using the audio processing model for audio processing (i.e., the process of using the audio processing model) will be explained.
[0212] During audio processing based on the audio processing model, users can set the audio processing effects as needed, including turning noise reduction and dereverberation functions on or off, and adjusting the noise reduction and dereverberation intensity. In other words, users can input intensity adjustment commands as needed. To facilitate understanding, the method for inputting intensity adjustment commands will be explained first.
[0213] For real-time audio pickup scenarios, users can pre-input intensity adjustment commands. In one embodiment, the electronic device can provide a unified settings entry for inputting noise reduction intensity and dereverberation intensity, etc. The input information is valid for some or all of the apps involved in audio processing in real-time audio pickup scenarios (referred to as the target app). Optionally, the settings entry can be located in the settings app or in the negative one screen. As described in the foregoing embodiments, the target app may include a recorder app, a note-taking app, a call app, an instant messaging app, etc.
[0214] For example, Figure 8 This is a schematic diagram illustrating another application scenario of the audio processing method provided in this application embodiment. Taking the audio processing function settings entry located in the settings app as an example, as... Figure 8 As shown in Figure (a), the electronic device displays the desktop. The user can click the Settings app icon 801 on the desktop to enter the Settings app interface 802, as shown... Figure 8As shown in Figure (b). The settings interface 802 of the app can include multiple settings options, including an audio processing option 803. The user clicks on the audio processing option 803 to enter the audio processing settings interface 804, as shown... Figure 8 As shown in Figure (c).
[0215] The audio processing settings interface 804 may include a noise reduction function setting area 805 and a de-reverb function setting area 806. The noise reduction function setting area 805 includes a noise reduction function switch 8051, which the user can click to control the noise reduction function to be turned on or off. When the noise reduction function switch 8051 is on, the noise reduction function setting area 805 also includes a noise reduction intensity increase control 8052, a noise reduction intensity decrease control 8053, and a noise reduction intensity percentage control 8054. The noise reduction intensity increase control 8052 is used to increase the noise reduction intensity. When the noise reduction function switch 8051 is on and the user clicks the noise reduction intensity increase control 8052, it indicates that the input function type is noise reduction and the adjustment method is increase. The noise reduction intensity decrease control 8053 is used to decrease the noise reduction intensity. When the noise reduction function switch 8051 is on and the user clicks the noise reduction intensity decrease control 8053, it indicates that the input function type is noise reduction and the adjustment method is decrease. The noise reduction intensity percentage control 8054 is used to set the percentage of noise reduction intensity. With the noise reduction function switch 8051 in the "on" state, the user inputs the desired percentage (n%) through the noise reduction intensity percentage control 8054, indicating that the input function type is noise reduction and the adjustment method is n%. The electronic device responds to the user's input and saves each set of instruction information as part of the intensity adjustment instruction. Subsequently, after the electronic device acquires audio in real time, it obtains the intensity adjustment instruction and performs noise reduction processing on the audio based on the noise reduction-related instruction information in the intensity adjustment instruction.
[0216] Similarly, the dereverb setting area 806 includes a dereverb switch 8061. Users can control the dereverb function to be turned on and off by clicking the dereverb switch 8061. When the dereverb switch 8061 is on, the dereverb setting area 806 also includes a dereverb intensity increase control 8062, a dereverb intensity decrease control 8063, and a dereverb intensity percentage control 8064. The dereverb intensity increase control 8062 controls the dereverb intensity. When the dereverb switch 8061 is on and the user clicks the dereverb intensity increase control 8062, it indicates that the input function type is dereverb and the adjustment method is increase. The dereverb intensity decrease control 8063 controls the dereverb intensity to decrease. When the dereverb switch 8061 is on and the user clicks the dereverb intensity decrease control 8063, it indicates that the input function type is dereverb and the adjustment method is decrease. The dereverb intensity percentage control 8064 is used to set the percentage of dereverb intensity. With the dereverb function switch 8061 in the "on" state, the user inputs the desired percentage (n%) through the dereverb intensity percentage control 8064, indicating that the input function type is dereverb and the adjustment method is n%. The electronic device responds to the user's input and saves this instruction information as part of the intensity adjustment instruction. Subsequently, after the electronic device acquires audio in real time, it obtains the intensity adjustment instruction and performs dereverb processing on the audio based on the dereverb-related instruction information in the intensity adjustment instruction.
[0217] In another embodiment, each target app can provide its own audio processing settings entry. In this case, the user can input intensity adjustment commands in the target app, and the target app saves these commands. Then, when the target app captures audio, the audio processing module processes the audio based on the saved intensity adjustment commands and performs subsequent operations.
[0218] In a pre-recorded scenario, in one embodiment, the user can input an intensity adjustment command at any time before or during audio playback. The input intensity adjustment command applies to the unplayed portion of the audio. In this scenario, the input port for the intensity adjustment command can be located within the app involved in audio processing for pre-recorded scenarios. As described in the foregoing embodiments, the app involved in audio processing for pre-recorded scenarios can include a recorder app, a photo album app, etc.
[0219] For example, Figure 9 This is a schematic diagram illustrating another application scenario of the audio processing method provided in this application embodiment. Taking a recorder APP as an example, such as... Figure 9 As shown in Figure (a), Figure 9As shown in Figure (a), the main interface of the electronic device may include an icon 201 for a recorder app. When the user clicks the recorder app icon 201, the electronic device displays the main interface 202 of the recorder app, as shown in Figure (a). Figure 9 As shown in Figure (b) above, the main interface 202 of the recorder app includes multiple recordings. Taking the user clicking on recording 3 as an example, in response to the user's click, the electronic device performs noise reduction and dereverberation processing on recording 3 according to the initial intensity adjustment command. Afterwards, the processed audio is played and displayed as shown below. Figure 9 The recording playback interface 901 is shown in Figure (c). The initial intensity adjustment command can be either the default intensity adjustment command of the recorder app or the intensity adjustment command previously entered by the user; there are no restrictions on which.
[0220] The recording playback interface 901 includes a noise reduction function setting area 902 and a de-reverb function setting area 903. The noise reduction function setting area 902 includes a noise reduction function switch 9021, which the user can click to control the noise reduction function to be turned on or off. When the noise reduction function switch 9021 is on, the noise reduction function setting area 902 displays a noise reduction intensity increase control 9022, a noise reduction intensity decrease control 9023, and a noise reduction intensity adjustment bar 9024. The noise reduction intensity increase control 9022 is used to increase the noise reduction intensity. When the noise reduction function switch 9021 is on and the user clicks the noise reduction intensity increase control 9022, it indicates that the function type is noise reduction and the adjustment method is increase. When the noise reduction function switch 9021 is on and the user clicks the noise reduction intensity decrease control 9023, it indicates that the function type is noise reduction and the adjustment method is decrease. The length of the noise reduction intensity adjustment bar 9024 corresponds to the adjustable range of the noise reduction intensity. A slider 9025 is displayed on the noise reduction intensity adjustment bar 9024; the position of slider 9025 indicates the current percentage of noise reduction intensity. When the noise reduction function switch 9021 is in the on state, the user can input the desired percentage (n%) by dragging slider 9025 on the noise reduction intensity adjustment bar 9024, indicating that the function type is noise reduction and the adjustment method is (n%). The electronic device responds to the user's operation, performs noise reduction processing on the unplayed portion of the currently playing recording 3 based on the input command information, and then plays the noise-reduced audio.
[0221] Similarly, the dereverb setting area 903 includes a dereverb switch 9031. Users can control the dereverb function to be turned on and off by clicking the dereverb switch 9031. When the dereverb switch 9031 is on, the dereverb intensity increase control 9032, the dereverb intensity decrease control 9033, and the dereverb intensity adjustment bar 9034 are displayed in the dereverb intensity setting area 903. The dereverb intensity increase control 9032 controls the increase of the dereverb intensity. When the dereverb switch 9031 is on and the user clicks the dereverb intensity increase control 9032, it indicates that the function type is dereverb and the adjustment method is increase. When the dereverb switch 9031 is on and the user clicks the dereverb intensity decrease control 9033, it indicates that the function type is dereverb and the adjustment method is decrease. The length of the dereverb intensity adjustment bar 9034 corresponds to the adjustable range of the dereverb intensity. The dereverb intensity adjustment bar 9034 displays a slider 9035, the position of which indicates the current percentage of dereverb intensity. With the dereverb function switch 9031 in the "on" state, the user can input the desired percentage (n%) by dragging the slider 9035 on the dereverb intensity adjustment bar 9034, indicating a command message with the function type being dereverb and the adjustment method being percentage (n%). Responding to the user's operation, the electronic device performs dereverb processing on the unplayed portion of the currently playing recording 3 based on the input command message, and then plays the dereverb-processed audio.
[0222] Below, with Figure 9 The audio processing method will be explained using the scenario shown as an example.
[0223] For example, Figure 10 This is a schematic diagram illustrating the back diffusion principle of an audio processing model provided in an embodiment of this application. Figure 11 A flowchart illustrating an example audio processing method provided in this application embodiment is also available. Figure 10 and Figure 11 The method includes:
[0224] S201. The recorder APP responds to the user's playback command and obtains the initial intensity adjustment command (also known as the second command).
[0225] In one specific embodiment, each time the recorder app opens a recording for playback, it can process the audio using a default intensity adjustment command. In this case, after receiving the user's playback command, the recorder app directly obtains the default intensity adjustment command and uses it as the initial intensity adjustment command.
[0226] In another specific embodiment, the recorder app has a memory function for the intensity adjustment commands input by the user. In this case, after receiving the user's playback command, the recorder app can retrieve the previous intensity adjustment command input by the user and use that command as the initial intensity adjustment command.
[0227] S202, The recorder APP inputs the initial intensity adjustment command into the first encoder in the audio processing module.
[0228] S203. The first encoder encodes the initial intensity adjustment command and extracts the feature vector from the initial intensity adjustment command to obtain command embedding1 (also known as the second command feature vector).
[0229] S204. The first encoder takes the instruction embedding1 as a conditional input and inputs it into the audio processing model.
[0230] The specific execution process of steps S202 to S204 is similar to steps S104 to S106 in the model training process, and will not be described again here.
[0231] S205, the recorder APP takes each audio segment from recording 3 as input and inputs it into the audio processing model in the audio processing module segment by segment.
[0232] Recording 3 is input into the diffusion model segment by segment, that is, recording 3 is split into multiple segments to obtain multiple audio segments. Each audio segment is used as input to the audio processing model, and then input into the audio processing model sequentially. The length of each audio segment can be set according to requirements. The audio processing model also processes the audio segment by segment. The segment-by-segment audio input method can also be called streaming input.
[0233] It can be understood that each audio segment is an audio signal containing noise and / or having a reverberation effect, that is, noisy audio. In some embodiments, the audio input to the audio processing model is also referred to as the audio to be processed.
[0234] S206. The audio processing model performs a reverse diffusion process, processes the audio 1 (also known as the second audio) input by the recorder APP according to the instruction embedding1, and outputs the target audio 1 (also known as the second target audio).
[0235] As described in the above embodiments, through model training, the audio processing model learns the ability to process audio according to intensity adjustment instructions. Therefore, during backdiffusion, the audio processing model can perform noise reduction and / or dereverberation processing on the audio based on the instruction embedding, outputting an audio signal that is cleaner than the input audio and matches the processing intensity of the conditional input constraints. For ease of description, the audio output by the audio processing model is referred to as the target audio.
[0236] Specifically, the audio processing model performs the following back-diffusion process: The audio is used as the starting point of the back-diffusion process (i.e., the audio is used as data x). T Referring to formula (6) above, the audio is processed at each step t. Here, the conditional input y is the embedding instruction. As the number of iterations increases, the noise in the audio gradually decreases, and / or the reverberation effect gradually becomes less noticeable. After T iterations, the final noise reduction intensity is consistent with the noise reduction intensity indicated by the conditional input, and / or the final de-reverberation intensity is consistent with the de-reverberation intensity indicated by the conditional input. Thus, this matches the user's requirements for audio processing intensity, satisfying the user's needs.
[0237] S207, The audio processing model returns the target audio 1 to the recorder APP.
[0238] S208, Recorder APP plays target audio 1.
[0239] In theory, the noise reduction and déreverberation effects of target audio 1 match the initial intensity adjustment command. Specifically, the change in SNR of target audio 1 compared to audio 1 input by the recording app is consistent with the noise reduction-related command information in the initial intensity adjustment command; the change in reverberation time of target audio 1 compared to audio 1 input by the recording app is consistent with the déreverberation-related command information in the initial intensity adjustment command.
[0240] S209. In response to the user inputting a new intensity adjustment command (also known as a first command), the recorder APP inputs the new intensity adjustment command to the first encoder.
[0241] For inputting the new intensity adjustment command, see [link / reference]. Figure 9 I will not go into details.
[0242] S210. The first encoder encodes the new intensity adjustment instruction and extracts the feature vector from the new intensity adjustment instruction to obtain instruction embedding2 (also known as the first instruction feature vector).
[0243] S211. The first encoder takes the instruction embedding2 as a conditional input and inputs it into the audio processing model.
[0244] S212. The audio processing model performs a reverse diffusion process, processes the audio 2 (also known as the first audio) input by the recorder APP according to the instruction embedding2, and outputs the target audio 2 (also known as the first target audio).
[0245] For example, a recording app might split a recording into 10 audio segments. When the third segment finishes playing, the user inputs a new intensity adjustment command. The audio processing model can then begin processing the subsequent segments from the fourth segment onwards using the new intensity adjustment command.
[0246] The noise reduction and dereverberation effects of target audio 2 are matched with the new intensity adjustment command. Specifically, the change in SNR (ΔSNR) of target audio 2 compared to audio 2 input by the recording app is consistent with the noise reduction-related command information in the new intensity adjustment command: if the new intensity adjustment command includes a command with the function type of noise reduction and the adjustment method of addition, then ΔSNR2 (also known as the first SNR difference) is greater than ΔSNR1 (also known as the second SNR difference). Optionally, ΔSNR2 is a larger than ΔSNR1 (also known as the first preset value). ΔSNR2 is the difference between the SNR of target audio 2 and the SNR of audio 2, and ΔSNR1 is the difference between the SNR of target audio 1 and the SNR of audio 1. If the new intensity adjustment command includes a command with the function type of noise reduction and the adjustment method of subtraction, then ΔSNR2 is less than ΔSNR1. Optionally, ΔSNR2 is a smaller than ΔSNR1. If the new intensity adjustment instruction includes instruction information with function type as noise reduction and adjustment method as n%, then the signal-to-noise ratio of target audio 2 is: the product of the range difference of the preset adjustable range of signal-to-noise ratio (also referred to as the preset signal-to-noise ratio range) and n%, and the sum of the lower limit of the preset adjustable range of signal-to-noise ratio, that is, SNR(target audio 2) = (a1-a0)*n%+a0.
[0247] Similarly, the change in reverberation time of target audio 1 compared to audio 2 input by the recording app is consistent with the de-reverberation related instruction information in the new intensity adjustment instruction: if the new intensity adjustment instruction includes instruction information with function type "de-reverberation" and adjustment method "addition", then △rev_t2 (also known as the first reverberation time difference) is less than △rev_t1 (also known as the second reverberation time difference). Optionally, △rev_t2 is t smaller than △rev_t1 (also known as the second preset value). △rev_t2 is the difference between the reverberation time of target audio 2 and the reverberation time of audio 2, and △rev_t1 is the difference between the reverberation time of target audio 1 and the reverberation time of audio 1. If the new intensity adjustment instruction includes instruction information with function type "de-reverberation" and adjustment method "subtraction", then △rev_t2 is greater than △rev_t1. Optionally, △rev_t2 is t larger than △rev_t1. If the new intensity adjustment instruction includes instruction information with function type "de-reverb" and adjustment method "n%", then the reverb time of target audio 2 is: the product of the difference between the preset adjustable reverb time range (also referred to as the preset reverb time range) and n%, and the sum of the lower limit of the preset adjustable reverb time range, i.e., rev_t(target audio 2) = (t1-t0)*n%+t0.
[0248] This allows users to adjust the audio effects in real time to suit their preferences and usage habits, thus improving the user experience.
[0249] S213, The audio processing model returns the target audio 2 to the recorder APP.
[0250] S214, Recorder APP plays target audio 2.
[0251] The audio processing method provided in this application involves inputting the audio to be processed into an audio processing model and encoding the intensity adjustment command input by the user to obtain a command feature vector. This command feature vector is then used as a conditional input to a diffusion model. The audio processing model performs a back-diffusion process, performing noise reduction and / or dereverberation processing on the audio to be processed according to the intensity adjustment command to obtain the target audio. This method allows users to customize the degree of noise reduction and dereverberation according to their needs, achieving a customized audio effect. Furthermore, the audio processing model in this method is based on a diffusion model trained on it, eliminating signal fusion and resulting in higher quality and better listening experience of the target audio, effectively improving the user experience.
[0252] The above embodiment illustrates the use of a user inputting intensity adjustment commands through controls on the interface. In practical applications, users can also input intensity adjustment commands via voice or text. The following section further explains the processing procedures of the electronic device when a user inputs an intensity adjustment command via voice.
[0253] For example, Figure 12 This is another example scenario diagram provided in the embodiments of this application. For example... Figure 12 As shown, in this scenario, the recording playback interface can be as follows: Figure 12 As shown in 1201, the recording and playback interface 1201 may include an input control 1202 for command audio. Users can trigger the recording of command audio through the command audio input control 1202. The audio processing module can process the audio based on the command audio.
[0254] For example, Figure 13 This is a schematic diagram illustrating the back diffusion principle of another audio processing model provided in this application embodiment. Figure 14 Please refer to the flowchart of another audio processing method provided in this application embodiment. Figure 13 and Figure 14 In step S209 above, the following process can be replaced:
[0255] S301, The recorder APP receives audio commands input by the user via voice.
[0256] For example, a user inputs an audio command: "More noise reduction, 50% reverb."
[0257] S302, the recorder APP sends the instruction audio to the ASR unit in the audio processing module.
[0258] The S303 and ASR units recognize speech in the instruction audio and convert the speech into text.
[0259] S304, the ASR unit sends the converted text to the instruction recognition unit.
[0260] S305, The instruction recognition unit extracts keywords from the text and matches the extracted keywords with preset instruction words to determine the intensity adjustment instruction text from the keywords.
[0261] Preset command words refer to words that are pre-defined to describe the content of intensity adjustment commands. For example, noise reduction, noise cancellation, more, less, percentage, de-reverberation, reverberation, echo, etc.
[0262] Intensity adjustment command text refers to the words in the keywords that describe the content of the intensity adjustment command, obtained by comparing and matching keywords with preset command words. Continuing with the example in step S301 above, if the content of the command audio is "more noise reduction, 50% reverb", the intensity adjustment command text determined from the keywords can include: "noise reduction", "more", "reverb" and "50%".
[0263] S306, The instruction recognition unit sends the intensity adjustment instruction text to the instruction conversion unit.
[0264] S307, The instruction conversion unit converts the intensity adjustment instruction text into a preset format to obtain a new intensity adjustment instruction.
[0265] As described in step S103 above, the intensity adjustment instruction can have a preset format; for example, different information can be represented by the values of different fields. Based on this, the intensity adjustment instruction text can be converted, transforming its content into a preset format to obtain a new intensity adjustment instruction.
[0266] Continuing with the example above, the intensity adjustment command text includes "noise reduction," "multi-point," "reverb," and "50%." Following the method described in step S103, this intensity adjustment command text is converted to obtain two new intensity adjustment commands. One is an intensity adjustment command with the function type being noise reduction and the adjustment direction being increment. In this command, the function type field has a preset value of 1 (e.g., 0), indicating that the function type is noise reduction; the increment field has a preset value of 3 (e.g., true), indicating that the adjustment method is increment; the decrement field has a preset value of 4 (e.g., false), indicating that the adjustment method is not decrement; and the percentage field has a preset value of 5 (e.g., false), indicating that the adjustment method is not percentage. The other is an intensity adjustment command with the function type being dereverb and the adjustment direction being 50%. In this command, the function type field has a preset value of 2 (e.g., 1), indicating that the function type is dereverb; the increment field has a preset value of 4 (e.g., false), indicating that the adjustment method is not increment; the decrement field has a preset value of 4 (e.g., false), indicating that the adjustment method is not decrement; and the percentage field has a value of 50%, indicating that the adjustment method is 50%.
[0267] S308, the instruction conversion unit inputs the new intensity adjustment instruction into the first encoder.
[0268] After step S309, refer to the above. Figure 11 In the embodiment shown, steps S210 to S214 are performed, and will not be described again.
[0269] In this embodiment, users can adjust the intensity via voice input, which improves the intelligence of the operation and further enhances the user experience.
[0270] The foregoing has detailed examples of the audio processing methods and model training methods provided in the embodiments of this application. It is understood that, in order to achieve the above functions, the electronic device includes hardware and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.
[0271] This application embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into a separate functional module, such as a detection unit, a processing unit, a display unit, etc., or two or more functions can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0272] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0273] The electronic device provided in this embodiment is used to execute the above-described model training method and audio processing method, and therefore can achieve the same effect as the above-described implementation method.
[0274] When using integrated units, the electronic device may further include a processing module, a storage module, and a communication module. The processing module is used to control and manage the operation of the electronic device. The storage module supports the execution of stored program code and data. The communication module supports communication between the electronic device and other devices.
[0275] The processing module can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a digital signal processor (DSP), and a microprocessor, etc. The storage module can be a memory. The communication module can specifically be a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, or other devices that interact with other electronic devices.
[0276] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device having... Figure 3 The device with the structure shown.
[0277] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the model training method and audio processing method of any of the above embodiments.
[0278] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the model training method and audio processing method in the above embodiments.
[0279] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the model training method and audio processing method in the above-described method embodiments.
[0280] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0281] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0282] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0283] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0284] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0285] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0286] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An audio processing method, said method being executed by an electronic device, characterized in that, The method includes: Obtain the audio to be processed; The system receives a first instruction input by the user, which instructs the user to adjust the noise reduction intensity and / or dereverberation intensity. The first instruction includes one or two sets of instruction information, each set including a function type and a corresponding adjustment method. The function type is noise reduction or dereverberation, and the adjustment method is increment, decrement, or n%. The function type "noise reduction" indicates that the audio processing method is noise reduction, and the function type "dereverberation" indicates that the audio processing method is dereverberation. The adjustment method "increase" indicates an increase in processing intensity, and "decrease" indicates a decrease in processing intensity. The adjustment method "n%" indicates that the processing intensity is adjusted to n% of a preset adjustable range, where n is a value greater than 0 and less than 100. The first splicing result, which is the concatenation of the function type and the corresponding adjustment method in the first instruction, is input into the first encoder; the first encoder encodes the first splicing result to obtain a first instruction feature vector, which includes a function type part and an adjustment method part; The first audio file is used as input, and the first instruction feature vector is used as conditional input to input the audio processing model; the audio processing model is a model trained based on the initial diffusion model; the first audio file is a segment of the audio file to be processed. The audio processing model performs a back diffusion process, performs noise reduction and / or de-reverberation processing on the first audio based on the first instruction feature vector, and outputs the first target audio.
2. The method according to claim 1, characterized in that, After acquiring the audio to be processed and before receiving the first instruction input by the user, the method further includes: Obtain a second instruction, which is used to instruct the adjustment of noise reduction intensity and / or dereverberation intensity; The second instruction is encoded by the first encoder to obtain the feature vector of the second instruction; The second audio is used as input, and the second instruction feature vector is used as conditional input to the audio processing model; the second audio is a segment of the audio to be processed that precedes the first audio. The audio processing model performs a back diffusion process, performs noise reduction and / or de-reverberation processing on the second audio based on the second instruction feature vector, and outputs the second target audio.
3. The method according to claim 2, characterized in that, The first instruction includes instruction information with the function type being noise reduction and the adjustment method being addition, and the first signal-to-noise ratio difference being greater than the second signal-to-noise ratio difference; the first signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the first target audio and the signal-to-noise ratio of the first audio, and the second signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the second target audio and the signal-to-noise ratio of the second audio.
4. The method according to claim 2, characterized in that, The first instruction includes instruction information with the function type being noise reduction and the adjustment method being reduction. The first signal-to-noise ratio difference is less than the second signal-to-noise ratio difference. The first signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the first target audio and the signal-to-noise ratio of the first audio. The second signal-to-noise ratio difference is the difference between the signal-to-noise ratio of the second target audio and the signal-to-noise ratio of the second audio.
5. The method according to claim 3, characterized in that, The absolute value of the difference between the first signal-to-noise ratio difference and the second signal-to-noise ratio difference is a first preset value.
6. The method according to claim 2, characterized in that, The first instruction includes instruction information with function type "noise reduction" and adjustment method "n%". The signal-to-noise ratio of the first target audio is the sum of the product of the range difference of the preset signal-to-noise ratio range and "n%" and the lower limit of the preset signal-to-noise ratio range.
7. The method according to claim 2, characterized in that, The first instruction includes instruction information with function type "de-reverb" and adjustment method "addition", and the first reverb time difference is less than the second reverb time difference; the first reverb time difference is the difference between the reverb time of the first target audio and the reverb time of the first audio, and the second reverb time difference is the difference between the reverb time of the second target audio and the reverb time of the second audio.
8. The method according to claim 2, characterized in that, The first instruction includes instruction information with function type "de-reverb" and adjustment method "reduction", and the first reverb time difference is greater than the second reverb time difference; the reverb time difference is the difference between the reverb time of the first target audio and the reverb time of the first audio, and the second reverb time difference is the difference between the reverb time of the second target audio and the reverb time of the second audio.
9. The method according to claim 7, characterized in that, The absolute value of the difference between the first reverberation time difference and the second reverberation time difference is a second preset value.
10. The method according to claim 2, characterized in that, The first instruction includes instruction information with function type "de-reverb" and adjustment method "n%". The reverb time of the first target audio is the sum of the product of the range difference of the preset reverb time range and "n%" and the lower limit of the preset reverb time range.
11. The method according to claim 1, characterized in that, The first encoder is a content encoder or a text encoder, and the first instruction feature vector is an instruction embedding.
12. The method according to claim 1, characterized in that, The first instruction for receiving user input includes: Receives audio commands from user voice input; Identify the speech in the instruction audio and convert the identified speech into text; Identify the instructions in the text to obtain the instruction text; The instruction text is converted into a preset format to obtain the first instruction.
13. The method according to any one of claims 1 to 12, characterized in that, The first instruction for receiving user input includes: The target application's interface is displayed; the target application's interface includes the audio to be processed. In response to a user's first operation on the audio to be processed, a first control is displayed; Receive the first instruction input by the user through the first control.
14. The method according to claim 13, characterized in that, The target application is a voice recorder application, a note-taking application, a call application, an instant messaging application, or a settings application.
15. A model training method, said method being performed by an electronic device, characterized in that, The method includes: Acquire multiple audio samples; Perform the training process; The training process includes: setting a first instruction sample, which includes one or two sets of instruction information. Each set of instruction information includes a function type and a corresponding adjustment method. The function type is noise reduction or de-reverb, and the adjustment method is increment, decrement, or n%. The function type "noise reduction" represents the audio processing method as noise reduction, and the function type "de-reverb" represents the audio processing method as de-reverb. The adjustment method "increase" represents an increase in processing intensity, "decrease" represents a decrease in processing intensity, and "n%" represents adjusting the processing intensity to n% of a preset adjustable range of processing intensity, where n is a value greater than 0 and less than 100. The first instruction sample contains the function... The second splicing result obtained by concatenating the type and the corresponding adjustment method is input into the second encoder; the second encoder encodes the second splicing result to obtain a sample instruction feature vector, which includes a function type part and an adjustment method part; the first audio sample is used as input, and the sample instruction feature vector is used as a conditional input to input the initial diffusion model; according to the first instruction sample, the first audio sample is subjected to noise addition and / or reverberation addition to obtain a first target audio sample; the first target audio sample is used as a target label, and the initial diffusion model is trained by performing a forward diffusion process; the first audio sample is any one of the plurality of audio samples; The other audio samples among the plurality of audio samples are used as the first audio sample, and the training process is repeated until the initial diffusion model converges or the number of training iterations reaches a preset threshold, thereby obtaining the audio processing model.
16. The method according to claim 15, characterized in that, The step of adding noise and / or adding reverb to the first audio sample according to the first instruction sample includes: If the first instruction sample includes instruction information with the function type of noise reduction and the adjustment method of addition, then the first audio sample is subjected to noise addition processing; the signal-to-noise ratio of the first target audio sample is smaller than the signal-to-noise ratio of the first audio sample by a first preset value; If the first instruction sample includes instruction information with the function type of noise reduction and the adjustment method of reduction, then the first audio sample is subjected to noise addition processing; the signal-to-noise ratio of the first target audio sample is greater than the signal-to-noise ratio of the first audio sample by the first preset value. If the first instruction sample includes instruction information with function type of noise reduction and adjustment method of n%, then the first audio sample is subjected to noise addition processing; the signal-to-noise ratio of the first target audio sample is: the product of the range difference of the preset signal-to-noise ratio range and n%, and the sum of the lower limit of the preset signal-to-noise ratio range.
17. The method according to claim 15, characterized in that, The step of adding noise and / or adding reverb to the first audio sample according to the instruction sample includes: If the first instruction sample includes instruction information with function type "de-reverb" and adjustment method "add", then the first audio sample is subjected to reverb processing; the reverb time of the first target audio sample is greater than the reverb time of the first audio sample by a second preset value. If the first instruction sample includes instruction information with function type "de-reverb" and adjustment method "reduction", then the first audio sample is subjected to reverb processing; the reverb time of the first target audio sample is less than the second preset value than the reverb time of the first audio sample. If the first instruction sample includes instruction information with function type of de-reverb and adjustment method of n%, then the first audio sample is subjected to reverb processing; the reverb time of the first target audio sample is: the product of the range difference of the preset reverb time range and n%, and the sum of the lower limit of the preset reverb time range.
18. The method according to any one of claims 15 to 17, characterized in that, The second encoder is a content encoder or a text encoder, and the sample instruction feature vector is an instruction embedding.
19. An electronic device, characterized in that, The electronic device includes: one or more processors, and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 18.
20. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 18.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 18.
Citation Information
Patent Citations
Method of earphone de-noising and earphone
CN106686481A