Audio debugging method and device, electronic equipment and storage medium

By using a pre-trained audio debugging model, the audio debugging parameters are automatically adjusted, which solves the problem of time-consuming and labor-consuming manual parameters in the prior art, and improves audio quality and efficiency.

CN119943075APending Publication Date: 2025-05-06HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097123.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, audio debugging requires manual adjustment of multiple parameters, which makes it difficult for non-professional personnel to understand and adjust. Professional personnel need to spend a lot of time and effort, and it is difficult to adjust to the optimal parameter value.

Method used

The pre-trained audio debugging model is adopted to input the audio to be debugged, and the audio debugging parameters to be debugged according to the preset debugging method based on the training. The model parameters are adjusted to reduce the difference in audio quality.

Benefits of technology

It saves labor costs required for audio debugging, improves the audio quality after debugging, and reduces the time and effort of manually adjusting parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943075A_ABST
    Figure CN119943075A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio debugging method and device, electronic equipment and a storage medium, and belongs to the technical field of audio debugging. And inputting the to-be-debugged audio into a pre-trained audio debugging model, so that the audio debugging model carries out audio debugging on the to-be-debugged audio according to a preset debugging mode based on the parameter value of the audio debugging parameter obtained in the training process to obtain a debugged audio, and outputs the debugged audio. Due to the fact that the preset debugging mode can conduct audio debugging on the audio to be debugged based on the included parameters, the audio debugging parameters included in the audio debugging model can be made to be the same as the parameters included in the preset debugging mode. Thus, in the training process of the audio debugging model, the parameter value of the audio debugging parameter can be obtained, the parameter value of the audio debugging parameter does not need to be adjusted manually, the labor cost required by audio debugging can be saved, and the audio quality of the debugged audio is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio debugging, and in particular to an audio debugging method, device, electronic device and storage medium. Background Art

[0002] Audio debugging is an important technical task, which can perform various processing on audio, such as suppressing noise in audio, eliminating echo in audio, enhancing voice signal, etc., so as to improve audio quality. In the current related technology, it is necessary to manually adjust each parameter in the audio quality enhancement algorithm to the appropriate parameter value one by one, and then debug the audio based on the manually adjusted audio quality enhancement algorithm.

[0003] Since the audio debugging process involves many parameters and is highly professional, it is difficult for non-professionals to understand and grasp the meaning of various parameters and cannot accurately adjust the parameters. Professionals need to spend a lot of time and energy on debugging, and it is difficult to adjust to the optimal parameter value. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide an audio debugging method, device, electronic device and storage medium to save the manpower cost required for audio debugging and improve the audio quality of the debugged audio. The specific technical solution is as follows:

[0005] In a first aspect, an embodiment of the present application provides an audio debugging method, the method comprising:

[0006] Get the audio to be debugged;

[0007] Inputting the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, obtains the debugged audio, and outputs the debugged audio;

[0008] Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

[0009] Optionally, the training method of the audio debugging model includes:

[0010] Acquire multiple sample audios and a label audio corresponding to each sample audio, wherein noise included in the label audio is lower than a preset noise threshold;

[0011] Inputting the sample audio into an initial audio debugging model, so that the initial audio debugging model performs audio debugging on the sample audio based on parameter values ​​of current audio debugging parameters of the initial audio debugging model to obtain predicted audio;

[0012] Based on the difference between the predicted audio and its corresponding label audio, the parameter values ​​of the audio tuning parameters of the initial audio tuning model are adjusted by back propagation until the initial audio tuning model meets the convergence condition, thereby obtaining a trained audio tuning model.

[0013] Optionally, the step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation based on the difference between the predicted audio and the corresponding label audio includes:

[0014] Based on the difference between the predicted audio and the corresponding labeled audio, according to a preset audio quality assessment algorithm, calculate the audio quality assessment result corresponding to the predicted audio;

[0015] Determining a function value of a preset loss function according to the audio quality evaluation result, wherein the audio quality evaluation result is negatively correlated with the function value of the preset loss function;

[0016] The parameter values ​​of the audio tuning parameters of the initial audio tuning model are adjusted by back propagation in a direction that reduces the function value of the preset loss function.

[0017] Optionally, the step of calculating the audio quality assessment result corresponding to the predicted audio according to a preset audio quality assessment algorithm based on the difference between the predicted audio and the corresponding label audio includes:

[0018] Convert the predicted audio and its corresponding labeled audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the labeled audio in the frequency domain;

[0019] In each preset frequency band, determining a first energy of the predicted audio in the frequency domain and a second energy of the label audio in the frequency domain;

[0020] Calculating a distortion metric in the frequency domain based on a first preset weight and an energy difference corresponding to each preset frequency band, wherein the energy difference is a difference between a first energy and a second energy corresponding to the same preset frequency band;

[0021] According to the distortion metric, a function is determined according to a preset evaluation result to calculate an audio quality evaluation result corresponding to the predicted audio, wherein the distortion metric is negatively correlated with the audio quality evaluation result.

[0022] Optionally, the step of determining, in each preset frequency band, a first energy of the predicted audio in the frequency domain and a second energy of the label audio in the frequency domain includes:

[0023] In each preset frequency band, based on the predicted audio in the frequency domain and its corresponding second preset weight, the first energy of the predicted audio in the frequency domain is calculated according to the following formula: :

[0024] ;

[0025] In each preset frequency band, based on the label audio in the frequency domain and its corresponding third preset weight, the second energy of the label audio in the frequency domain is calculated according to the following formula: :

[0026] ;

[0027] in, For each preset frequency in the preset frequency band, Take the frequency Frequency value And the time is The predicted audio at Take the frequency Frequency value And the time is When the label audio, is the second preset weight, is the third preset weight.

[0028] Optionally, the step of calculating the distortion metric in the frequency domain based on the first preset weight and the energy difference corresponding to each preset frequency band includes:

[0029] Based on the difference between the modified first energy and the modified second energy corresponding to each preset frequency band and the first preset weight corresponding to each preset frequency band, the distortion metric in the frequency domain is calculated according to the following formula: :

[0030] ;

[0031] in, For the A first preset weight corresponding to a preset frequency band, For the The first energy corresponding to the preset frequency band, For the The second energy corresponding to a preset frequency band.

[0032] Optionally, the step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation in a direction that reduces the function value of the preset loss function includes:

[0033] For each audio tuning parameter, calculate the current partial derivative value of the loss function with respect to the audio tuning parameter as the gradient of the audio tuning parameter, wherein the current partial derivative value is the partial derivative value corresponding to when the parameter value of the audio tuning parameter is the current parameter value;

[0034] The parameter value of the audio debugging parameter is updated to the difference between the current parameter value and the updated value, wherein the updated value is the product of a preset learning rate and the gradient.

[0035] Optionally, the preset debugging method is an audio quality enhancement algorithm, and the audio debugging parameters include high-pass filter parameters, acoustic echo cancellation parameters, automatic noise suppression parameters, automatic gain control parameters, equalizer parameters and dynamic range control parameters.

[0036] In a second aspect, an embodiment of the present application provides an audio debugging device, the device comprising:

[0037] The audio acquisition module to be debugged is used to acquire the audio to be debugged;

[0038] An audio debugging module, used for inputting the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, obtains the debugged audio, and outputs the debugged audio;

[0039] Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

[0040] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0041] Memory, used to store computer programs;

[0042] The processor is used to implement any method described in the first aspect when executing a program stored in the memory.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any method described in the first aspect is implemented.

[0044] Beneficial effects of the embodiments of the present application:

[0045] In the scheme provided by the embodiment of the present application, the electronic device can obtain the audio to be debugged; input the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged based on the parameter values ​​of the audio debugging parameters obtained during the training process, according to the preset debugging method, obtains the debugged audio, and outputs the debugged audio; wherein the audio debugging parameters are the same as the parameters included in the preset debugging method, the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters, and the audio debugging model is: based on multiple sample audios and the label audio corresponding to each sample audio, in the direction of reducing the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio, the parameter value of the current audio debugging model is adjusted, and the debugged sample audio is debugged by the parameter value of the current audio debugging model. Since the preset debugging method can perform audio debugging on the audio to be debugged based on the included parameters, the audio debugging parameters included in the audio debugging model can be made the same as the parameters included in the preset debugging method. In this way, the parameters included in the preset debugging method can be used as trainable weights in the audio debugging model. During the training process of the audio debugging model, the parameter values ​​of the audio debugging parameters can be obtained. There is no need to manually adjust the parameter values ​​of the audio debugging parameters, which can save the manpower cost required for audio debugging and improve the audio quality of the debugged audio.

[0046] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0048] Figure 1 A flowchart of an audio debugging method provided in an embodiment of the present application;

[0049] Figure 2 Based on Figure 1 A schematic diagram of the voice quality enhancement algorithm of the illustrated embodiment debugging the audio;

[0050] Figure 3 Based on Figure 1 A flow chart of a training method of an audio debugging model in the illustrated embodiment;

[0051] Figure 4 for Figure 3 A specific flow chart of step S303 in the illustrated embodiment;

[0052] Figure 5 for Figure 4 A specific flow chart of step S401 in the illustrated embodiment;

[0053] Figure 6 for Figure 4 A specific flow chart of step S403 in the illustrated embodiment;

[0054] Figure 7 Based on Figure 1 A schematic flow chart of a method for determining parameter values ​​of audio debugging parameters in the illustrated embodiment;

[0055] Figure 8 Based on Figure 1 A schematic diagram of calculating the function value of the loss function in the illustrated embodiment;

[0056] Fig. 9 A schematic diagram of the structure of an audio debugging device provided in an embodiment of the present application;

[0057] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field based on the present application belong to the scope of protection of the present application.

[0059] In order to save the labor cost required for audio debugging and improve the audio quality of the debugged audio, the embodiments of the present application provide an audio debugging method, device, electronic device, computer-readable storage medium and computer program product. The following first introduces an audio debugging method provided by the embodiments of the present application.

[0060] The audio debugging method provided in the embodiment of the present application can be applied to any electronic device that needs to debug the audio, for example, an audio processor, a processing device, etc., which is not specifically limited here. For the sake of clarity, it is subsequently referred to as an electronic device in this article.

[0061] like Figure 1 As shown, an audio debugging method, the method comprising:

[0062] S101, obtaining the audio to be debugged;

[0063] S102, inputting the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on parameter values ​​of audio debugging parameters obtained during the training process, obtains debugged audio, and outputs the debugged audio.

[0064] Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

[0065] It can be seen that in the embodiment of the present application, the electronic device can obtain the audio to be debugged; input the audio to be debugged into the pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged based on the parameter values ​​of the audio debugging parameters obtained during the training process, according to the preset debugging method, obtains the debugged audio, and outputs the debugged audio; wherein the audio debugging parameters are the same as the parameters included in the preset debugging method, the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters, and the audio debugging model is: based on multiple sample audios and the label audio corresponding to each sample audio, in the direction of reducing the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio, the parameter value of the current audio debugging model is adjusted, and the debugged sample audio is debugged by the parameter value of the current audio debugging model. Since the preset debugging method can perform audio debugging on the audio to be debugged based on the included parameters, the audio debugging parameters included in the audio debugging model can be made the same as the parameters included in the preset debugging method. In this way, the parameters included in the preset debugging method can be used as trainable weights in the audio debugging model. During the training process of the audio debugging model, the parameter values ​​of the audio debugging parameters can be obtained. There is no need to manually adjust the parameter values ​​of the audio debugging parameters, which can save the manpower cost required for audio debugging and improve the audio quality of the debugged audio.

[0066] In step S101, the electronic device may obtain audio to be debugged, wherein the audio to be debugged is audio that needs to be debugged such as noise suppression, echo cancellation, and voice enhancement.

[0067] Next, the electronic device can input the audio to be debugged into the pre-trained audio debugging model, that is, execute step S102. The audio debugging model can then perform audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, thereby obtaining the debugged audio, and output the debugged audio.

[0068] The audio debugging model may be a deep learning model, specifically Tensorflow (an open source deep learning framework) and pytorch (an open source deep learning framework). The preset debugging method may perform audio debugging on the audio to be debugged based on the parameters included. In order to enable the audio debugging model to achieve the debugging effect of the preset debugging method, the audio debugging parameters may be the same as the parameters included in the preset debugging method.

[0069] In one implementation, the preset debugging mode is a voice quality enhancement algorithm (Voice Quality Enhancement, VQE), and the audio debugging parameters included in the audio debugging model may be parameters included in the voice quality enhancement algorithm.

[0070] like Figure 2 As shown, Figure 2 The figure is a schematic diagram of the voice quality enhancement algorithm debugging the audio. The voice quality enhancement algorithm may include the following virtual modules: High Pass Filter (HPF), Acoustic Echo Cancel (AEC), Automatic Noise Suppression (ANS), Automatic Gain Control (AGC), Equalizer (EQ) and Dynamic Range Control (DRC). The voice quality enhancement algorithm may debug the input audio based on the parameters included in the above virtual modules to obtain the output audio.

[0071] In this case, the audio debugging model includes audio debugging parameters such as high-pass filter parameters, acoustic echo cancellation parameters, automatic noise suppression parameters, automatic gain control parameters, equalizer parameters, and dynamic range control parameters. In this way, the voice quality enhancement algorithm can be used as a deep learning model as a whole, without distinguishing different virtual modules, and the parameters included in each virtual module can be used as trainable weights in the audio debugging model.

[0072] The above-mentioned audio debugging model can be obtained by adjusting the parameter value of the current audio debugging model based on multiple sample audios and the label audio corresponding to each sample audio, in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio.

[0073] The sample audio may be audio that needs to be debugged by the audio debugging model, and the label audio may be audio that represents ideal audio quality. The debugged sample audio may be obtained by debugging the sample audio using the parameter values ​​of the current audio debugging model.

[0074] Since the audio quality of the label audio is usually ideal, if the difference between the audio quality of the sample audio after debugging and the audio quality of the label audio is small, it means that the audio debugging model can debug the sample audio well. Therefore, in order to make the audio quality of the sample audio after debugging approach the label audio corresponding to the sample audio, the parameter value of the current audio debugging model can be adjusted in the direction of reducing the difference between the audio quality of the sample audio after debugging and the audio quality of the label audio corresponding to the sample audio.

[0075] It can be seen that in the embodiment of the present application, since the preset debugging method can perform audio debugging on the audio to be debugged based on the parameters included, the audio debugging parameters included in the audio debugging model can be made the same as the parameters included in the preset debugging method. In this way, the parameters included in the preset debugging method can be used as trainable weights in the audio debugging model, and in the training process of the audio debugging model, the parameter values ​​of the audio debugging parameters can be obtained without manually adjusting the parameter values ​​of the audio debugging parameters. Since the number of adjustable parameters in the preset debugging method is usually large, if it is done manually, it takes a lot of time to debug to find a group of parameter value combinations with better debugging effects, and it is difficult to find the parameter value combination with the best debugging effect. The embodiment of the present application determines the parameter values ​​of each parameter in the preset debugging method by machine learning, which can save the labor cost required for audio debugging and improve the audio quality of the debugged audio.

[0076] In addition, since different audio playback devices differ in performance and characteristics, if parameter values ​​are debugged for a certain type of audio playback device and the debugged parameter values ​​are applied to another type of audio playback device, the other audio playback device will not be able to achieve the same audio debugging effect as that of the audio playback device of that type. For example, if the audio debugged in a professional audio studio is played through an ordinary audio playback device, problems such as poor audio quality and unbalanced volume may occur. Therefore, in the current related technology, it is necessary to manually debug parameter values ​​for each type of audio playback device, which consumes a lot of manpower costs. In this scenario, the manpower cost required for audio debugging can be further reduced.

[0077] As an implementation method of the present application, Figure 3 As shown, the training method of the above audio debugging model may include:

[0078] S301, obtaining a plurality of sample audios and a label audio corresponding to each sample audio;

[0079] The electronic device may obtain multiple sample audios and label audios corresponding to each sample audio. The label audio includes noise lower than a preset noise threshold. The preset noise threshold may be set according to actual needs, and may be 15 dB, 10 dB, etc., which is not specifically limited here.

[0080] S302, inputting the sample audio into an initial audio debugging model, so that the initial audio debugging model performs audio debugging on the sample audio based on parameter values ​​of current audio debugging parameters of the initial audio debugging model to obtain predicted audio;

[0081] In order to detect the audio processing capability of the initial audio tuning model, the electronic device may input the sample audio into the initial audio tuning model, and the initial audio tuning model may perform audio tuning on the sample audio based on the parameter values ​​of the current audio tuning parameters of the initial audio tuning model to obtain the predicted audio.

[0082] S303, based on the difference between the predicted audio and its corresponding label audio, adjust the parameter values ​​of the audio tuning parameters of the initial audio tuning model by back propagation until the initial audio tuning model meets the convergence condition, thereby obtaining a trained audio tuning model.

[0083] Since the predicted audio is obtained by the initial audio debugging model through audio debugging of the sample audio, and the labeled audio represents audio with ideal audio quality, in order to enable the initial audio debugging model to output audio with ideal audio quality, the electronic device can adjust the parameter values ​​of the audio debugging parameters of the initial audio debugging model through back propagation based on the difference between the predicted audio and its corresponding labeled audio until the initial audio debugging model meets the convergence conditions, thereby obtaining a trained audio debugging model.

[0084] For example, the electronic device can obtain sample audio 1-sample audio 5 and label audio 1-label audio 5. First, the sample audio 1 is input into the initial audio debugging model. The initial audio debugging model can perform audio debugging on the sample audio 1 based on the parameter values ​​of the current audio debugging parameters of the initial audio debugging model to obtain predicted audio 1. Next, the parameter values ​​of the audio debugging parameters of the initial audio debugging model can be adjusted based on the difference between the predicted audio 1 and the label audio 1. For the sample audio that has not been input into the initial audio debugging model, the above steps are repeated until the initial audio debugging model meets the convergence conditions, and a trained audio debugging model can be obtained.

[0085] It can be seen that in the embodiment of the present application, the electronic device can obtain multiple sample audios and label audios corresponding to each sample audio, wherein the noise included in the label audio is lower than the preset noise threshold; the sample audio is input into the initial audio debugging model, so that the initial audio debugging model performs audio debugging on the sample audio based on the parameter values ​​of the current audio debugging parameters of the initial audio debugging model to obtain the predicted audio; based on the difference between the predicted audio and its corresponding label audio, the parameter values ​​of the audio debugging parameters of the initial audio debugging model are adjusted by back propagation until the initial audio debugging model meets the convergence conditions, thereby obtaining a trained audio debugging model. In this way, the parameter values ​​of the audio debugging parameters in the audio debugging model can be determined by machine learning, so that the debugging effect of the audio debugging based on the parameter values ​​of the audio debugging parameters can be better, thereby improving the audio quality of the debugged audio.

[0086] As an implementation method of the present application, Figure 4 As shown, the step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation based on the difference between the predicted audio and the corresponding label audio may include:

[0087] S401, calculating an audio quality assessment result corresponding to the predicted audio according to a preset audio quality assessment algorithm based on a difference between the predicted audio and the corresponding label audio;

[0088] In order to evaluate the audio debugging effect of the initial audio debugging model, the electronic device can calculate the audio quality evaluation result corresponding to the predicted audio according to a preset audio quality evaluation algorithm based on the difference between the predicted audio and its corresponding label audio.

[0089] Among them, the preset audio quality assessment algorithm can be divided into an algorithm that scores from objective indicators and an algorithm that scores from subjective indicators. For example, the algorithm that scores from objective indicators can be a perceptual evaluation of speech quality (PESQ) algorithm, a perceptual objective listening quality assessment algorithm (POLQA) algorithm, and the algorithm that scores from subjective indicators can be a deep noise suppression subjective opinion scoring algorithm (DNSmos).

[0090] S402, determining a function value of a preset loss function according to the audio quality evaluation result;

[0091] Since the audio quality assessment result is positively correlated with the audio quality of the predicted audio, and the function value of the loss function is negatively correlated with the audio quality of the predicted audio, the function value of the preset loss function can be determined based on the audio quality assessment result so that the audio quality assessment result is negatively correlated with the function value of the preset loss function.

[0092] The preset loss function may be pre-constructed, and the specific form is not limited as long as the function value of the preset loss function is negatively correlated with the audio quality evaluation result. For example, the preset loss function may be to change the audio quality evaluation result from a positive value to a negative value.

[0093] S403: adjusting the parameter values ​​of the audio tuning parameters of the initial audio tuning model by back propagation in a direction that reduces the function value of the preset loss function.

[0094] Next, the electronic device can adjust the parameter values ​​of the audio tuning parameters of the initial audio tuning model by back propagation in a direction that reduces the function value of the preset loss function.

[0095] It can be seen that in the embodiment of the present application, the electronic device can calculate the audio quality evaluation result corresponding to the predicted audio according to the preset audio quality evaluation algorithm based on the difference between the predicted audio and the corresponding label audio; determine the function value of the preset loss function according to the audio quality evaluation result, wherein the audio quality evaluation result is negatively correlated with the function value of the preset loss function; and adjust the parameter value of the audio debugging parameter of the initial audio debugging model by back propagation in the direction that reduces the function value of the preset loss function. Since the audio quality evaluation result is positively correlated with the audio quality of the predicted audio, and the function value of the loss function is negatively correlated with the audio quality of the predicted audio, the audio quality evaluation result can be converted by the preset loss function to obtain the function value of the preset loss function. In this way, the function value of the preset loss function can accurately characterize the difference between the predicted audio and the corresponding label audio.

[0096] Exemplarily, the preset audio quality assessment algorithm used in the subsequent embodiments is the above-mentioned perceptual speech quality assessment algorithm. Since the specific calculation method of the perceptual speech quality assessment algorithm is relatively complex and involves multiple steps such as signal processing and perceptual models, in order to further improve the debugging efficiency, the embodiment of the present application adopts a relatively simplified overview of the perceptual speech quality assessment algorithm. The complete calculation method can refer to standard documents, for example, the P.862 evaluation standard proposed by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T).

[0097] As an implementation method of the present application, Figure 5 As shown, the step of calculating the audio quality assessment result corresponding to the predicted audio according to a preset audio quality assessment algorithm based on the difference between the predicted audio and the corresponding label audio may include:

[0098] S501, converting the predicted audio and its corresponding label audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the label audio in the frequency domain;

[0099] In order to accurately capture the spectrum information of the predicted audio and the labeled audio, the electronic device can convert the predicted audio and its corresponding labeled audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the labeled audio in the frequency domain. The specific conversion method can be Short-Time Fourier Transform (STFT).

[0100] The above process can be expressed in the form of a formula: , among which, in To predict audio hour, For the predicted audio in the frequency domain , is the frequency of the predicted audio in the frequency domain; Tag Audio hour, is the labeled audio in the frequency domain , is the frequency of the label audio in the frequency domain.

[0101] In one implementation, before executing step S501, the electronic device may also pre-process the predicted audio and the labeled audio, and the specific processing method may include filtering and frame synchronization.

[0102] S502, in each preset frequency band, determining a first energy of the predicted audio in the frequency domain and a second energy of the label audio in the frequency domain;

[0103] In order to simulate the perceptual characteristics of the human auditory system, the predicted audio in the frequency domain and the labeled audio in the frequency domain can be mapped to the Bark Scale to determine the energy distribution corresponding to each preset frequency band in the Bark Scale. Specifically, the electronic device can determine the first energy of the predicted audio in the frequency domain and the second energy of the labeled audio in the frequency domain in each preset frequency band. The above-mentioned preset frequency band can be obtained by dividing 20Hz to 20000Hz into 24 frequency bands with different bandwidths.

[0104] S503, calculating a distortion metric in the frequency domain based on a first preset weight and an energy difference corresponding to each preset frequency band;

[0105] The electronic device can calculate the distortion metric in the frequency domain based on the first preset weight corresponding to each preset frequency band and the energy difference, thereby quantifying the difference between the predicted audio and the labeled audio, and generating an evaluation index consistent with the perceptual characteristics of the human auditory system.

[0106] The energy difference is the difference between the first energy and the second energy corresponding to the same preset frequency band. For example, assuming that the preset frequency band includes frequency band 1 to frequency band 20, the energy difference can be the difference between the first energy and the second energy corresponding to frequency band 1, or the difference between the first energy and the second energy corresponding to frequency band 2, and so on.

[0107] As an implementation method, before executing step S503, the electronic device may also align and calculate the mismatch between the predicted audio in the frequency domain and the labeled audio in the frequency domain through a compensation function, thereby calculating and correcting the alignment difference in time and frequency between the predicted audio and the labeled audio. Specifically, the alignment difference can be calculated by the following formula: :

[0108] .

[0109] The above formula makes the error function Get the minimum alignment difference , represents the compensation function, error function It can be as follows:

[0110] .

[0111] Next, you can use The time of the predicted audio is corrected so that the predicted audio is aligned with the labeled audio in the time domain.

[0112] S504: According to the distortion metric and a preset evaluation result determining function, an audio quality evaluation result corresponding to the predicted audio is calculated.

[0113] Next, the electronic device can determine the function according to the preset evaluation result based on the distortion metric and calculate the audio quality evaluation result corresponding to the predicted audio. Since the distortion metric represents the difference between the predicted audio and the labeled audio, the smaller the difference, the better the audio quality of the predicted audio, and the larger the audio quality evaluation result, the better the audio quality of the predicted audio, so the distortion metric is negatively correlated with the audio quality evaluation result. The audio quality evaluation result can be represented by a numerical value, for example, the value range can be 1 to 4.5.

[0114] The specific form of the preset evaluation result determination function can be various, as long as the distortion metric is negatively correlated with the audio quality evaluation result. For example, the preset evaluation result determination function can be MOS (Mean Opinion Score) represents the audio quality evaluation result, D is the distortion metric, and a, b, and c are preset constants.

[0115] It can be seen that in the embodiment of the present application, the electronic device can convert the predicted audio and its corresponding label audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the label audio in the frequency domain; in each preset frequency band, determine the first energy of the predicted audio in the frequency domain and the second energy of the label audio in the frequency domain; based on the first preset weight and energy difference corresponding to each preset frequency band, calculate the distortion metric in the frequency domain, wherein the energy difference is the difference between the first energy and the second energy corresponding to the same preset frequency band; according to the distortion metric, determine the function according to the preset evaluation result, and calculate the audio quality evaluation result corresponding to the predicted audio, wherein the distortion metric is negatively correlated with the audio quality evaluation result. Through the above implementation, the audio quality evaluation result can be calculated quickly and accurately.

[0116] As an implementation manner of an embodiment of the present application, the above-mentioned step of determining the first energy of the predicted audio in the frequency domain and the second energy of the label audio in the frequency domain in each preset frequency band may include the following two implementation manners:

[0117] In a first embodiment, the electronic device can calculate the first energy of the predicted audio in the frequency domain in each preset frequency band based on the predicted audio in the frequency domain and its corresponding second preset weight according to the following formula: :

[0118] .

[0119] in, The preset frequency in each preset frequency band, specifically, the preset frequency may be a middle value, a 1 / 4 value, or a 3 / 4 value of the preset frequency band, etc. For example, assuming that a preset frequency band is 20 Hz to 820 Hz, the preset frequency corresponding to the preset frequency band may be 410 Hz. Take the frequency Frequency value And the time is The predicted audio at is the second preset weight, indicating the frequency value Relative to preset frequency The weight of If the frequency range belongs to the sensitive frequency range of human ear, the value of the second preset weight is larger.

[0120] In a second embodiment, the electronic device can calculate the second energy of the tagged audio in the frequency domain in each preset frequency band based on the tagged audio in the frequency domain and its corresponding third preset weight according to the following formula: :

[0121] .

[0122] in, Take the frequency Frequency value And the time is When the label audio, is the third preset weight, indicating the frequency value Relative to preset frequency The weight of If the frequency range belongs to the sensitive frequency range of human ear, the value of the second preset weight is larger.

[0123] It can be seen that in the embodiment of the present application, the electronic device can calculate the first energy and the second energy respectively through the above two implementations. In this way, the label audio in the frequency domain and the predicted audio in the frequency domain can be mapped to the Bark frequency scale, thereby reflecting the human ear auditory perception characteristics of the label audio in the frequency domain and the predicted audio in the frequency domain.

[0124] As an implementation manner of an embodiment of the present application, the step of calculating the distortion metric in the frequency domain based on the first preset weight and energy difference corresponding to each preset frequency band may include:

[0125] Based on the difference between the modified first energy and the modified second energy corresponding to each preset frequency band and the first preset weight corresponding to each preset frequency band, the distortion metric in the frequency domain is calculated according to the following formula: :

[0126] .

[0127] in, For the A first preset weight corresponding to a preset frequency band is set. Assuming that the human ear is more sensitive to the audio of a preset frequency band, the value of the first preset weight corresponding to the preset frequency band is larger. For the The first energy corresponding to the preset frequency band, For the The second energy corresponding to a preset frequency band.

[0128] For example, assuming there are preset frequency bands 1 to 20, then The value of is 1 to 20. When the value of is 1, represents the first preset weight corresponding to the preset frequency band 1, is the first energy corresponding to the preset frequency band 1, is the second energy corresponding to the preset frequency band 1. When taking other values, The same applies when the value of is 1, which will not be repeated here.

[0129] It can be seen that in the embodiment of the present application, through the above implementation, the electronic device can calculate the distortion metric in the frequency domain. In this way, the difference between the predicted audio and the labeled audio can be quantified, thereby generating a quality assessment result consistent with the subjective perception of the human ear.

[0130] As an implementation method of the present application, Figure 6 As shown, the step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation in a direction that reduces the function value of the preset loss function may include:

[0131] S601, for each audio tuning parameter, calculating a current partial derivative value of the loss function with respect to the audio tuning parameter as a gradient of the audio tuning parameter;

[0132] The electronic device can calculate the current partial derivative value of the loss function for each audio tuning parameter as the gradient of the audio tuning parameter, wherein the current partial derivative value is the partial derivative value corresponding to the current parameter value of the audio tuning parameter.

[0133] Since there is a nested functional relationship between audio tuning parameters and output in the neural network, the derivative of the composite function can be calculated using the chain rule.

[0134] S602: Update the parameter value of the audio debugging parameter to the difference between the current parameter value and the updated value.

[0135] After calculating the gradient of the audio tuning parameter, the electronic device can update the parameter value of the audio tuning parameter by a gradient descent algorithm. Specifically, the electronic device can update the parameter value of the audio tuning parameter to the difference between the current parameter value and the updated value. The updated value is the product of the preset learning rate and the gradient. The preset learning rate can be set according to actual needs, for example, it can be 60%, 70%, 80%, etc., which is not specifically limited here.

[0136] If the above process is expressed in the form of a formula, we can get: .in, is the parameter value of the audio debugging parameter. is the preset learning rate, is the loss function, The gradient of the tuning parameter for this audio.

[0137] When the parameter value of the audio tuning parameter satisfies the convergence condition, the parameter value of the audio tuning parameter can be stopped from being updated and the parameter value of the audio tuning parameter can be saved. The convergence condition may include that the number of iterations reaches a preset number, the function value of the loss function no longer decreases, and so on.

[0138] It can be seen that in the embodiment of the present application, the electronic device can calculate the current partial derivative value of the loss function for each audio debugging parameter as the gradient of the audio debugging parameter, wherein the current partial derivative value is the partial derivative value corresponding to the current parameter value of the audio debugging parameter; the parameter value of the audio debugging parameter is updated to the difference between the current parameter value and the updated value, wherein the updated value is the product of the preset learning rate and the gradient. In this way, the parameter value with the best debugging effect corresponding to each audio debugging parameter can be quickly and accurately determined.

[0139] As an implementation method of the embodiment of the present application, a flow chart of a method for determining the parameter value of the audio debugging parameter can be as follows: Figure 7 As shown, the following steps may be specifically included:

[0140] S701, calculating the function value of the loss function;

[0141] The electronic device can process the sample audio in a preset debugging mode to obtain the predicted audio. Next, based on the difference between the label audio corresponding to the sample audio and the predicted audio, the function value of the loss function can be calculated by forward reasoning. The loss function can specifically be a perceptual speech quality assessment algorithm.

[0142] For example, a schematic diagram for calculating the function value of the loss function can be as follows Figure 8 As shown, for Figure 8 In the upper part, the sample audio can be processed by the preset debugging method to obtain the predicted audio. Figure 8 In the lower part, the difference between the labeled audio and the predicted audio can be calculated through the perceptual speech quality assessment algorithm to obtain the function value of the loss function.

[0143] S702, updating parameter values ​​by back propagation;

[0144] For each audio tuning parameter, the electronic device may update the parameter value of the audio tuning parameter in a direction that reduces the function value of the loss function by means of back propagation.

[0145] S703: Save the parameter value of the audio debugging parameter.

[0146] When the function value of the loss function is trained to converge, the electronic device can save the parameter value of each audio debugging parameter.

[0147] In the technical solution of this application, the operations involved in obtaining, storing, using, processing, transmitting, providing and disclosing user personal information are all carried out with the user's authorization.

[0148] Corresponding to the above-mentioned audio debugging method, the embodiment of the present application further provides an audio debugging device. The following is an introduction to the audio debugging device provided by the embodiment of the present application.

[0149] like Fig. 9 As shown, an audio debugging device, the device comprising:

[0150] The audio to be debugged acquisition module 901 is used to acquire the audio to be debugged;

[0151] The audio debugging module 902 is used to input the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, obtains the debugged audio, and outputs the debugged audio;

[0152] Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

[0153] It can be seen that in the scheme provided by the embodiment of the present application, the electronic device can obtain the audio to be debugged; input the audio to be debugged into the pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged based on the parameter values ​​of the audio debugging parameters obtained during the training process, according to the preset debugging method, obtains the debugged audio, and outputs the debugged audio; wherein the audio debugging parameters are the same as the parameters included in the preset debugging method, the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters, and the audio debugging model is: based on multiple sample audios and the label audio corresponding to each sample audio, in the direction of reducing the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio, the parameter value of the current audio debugging model is adjusted, and the debugged sample audio is debugged by the parameter value of the current audio debugging model. Since the preset debugging method can perform audio debugging on the audio to be debugged based on the included parameters, the audio debugging parameters included in the audio debugging model can be made the same as the parameters included in the preset debugging method. In this way, the parameters included in the preset debugging method can be used as trainable weights in the audio debugging model. During the training process of the audio debugging model, the parameter values ​​of the audio debugging parameters can be obtained. There is no need to manually adjust the parameter values ​​of the audio debugging parameters, which can save the manpower cost required for audio debugging and improve the audio quality of the debugged audio.

[0154] As an implementation of the embodiment of the present application, the above device may further include:

[0155] A training audio acquisition module, used to acquire a plurality of sample audios and a label audio corresponding to each sample audio, wherein the noise included in the label audio is lower than a preset noise threshold;

[0156] A predicted audio output module, used for inputting the sample audio into an initial audio debugging model, so that the initial audio debugging model performs audio debugging on the sample audio based on the parameter value of the current audio debugging parameter of the initial audio debugging model to obtain predicted audio;

[0157] A parameter value adjustment module is used to adjust the parameter values ​​of the audio tuning parameters of the initial audio tuning model by back propagation based on the difference between the predicted audio and the corresponding label audio, until the initial audio tuning model meets the convergence condition, thereby obtaining a trained audio tuning model.

[0158] As an implementation of the embodiment of the present application, the parameter value adjustment module may include:

[0159] An evaluation result acquisition submodule, used to calculate an audio quality evaluation result corresponding to the predicted audio according to a preset audio quality evaluation algorithm based on the difference between the predicted audio and the corresponding labeled audio;

[0160] A function value determination submodule, used to determine a function value of a preset loss function according to the audio quality evaluation result, wherein the audio quality evaluation result is negatively correlated with the function value of the preset loss function;

[0161] The parameter value adjustment submodule is used to adjust the parameter value of the audio debugging parameter of the initial audio debugging model by back propagation in a direction that reduces the function value of the preset loss function.

[0162] As an implementation of an embodiment of the present application, the above-mentioned evaluation result acquisition submodule may include:

[0163] A frequency domain conversion unit, used to convert the predicted audio and its corresponding label audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the label audio in the frequency domain;

[0164] An energy calculation unit, configured to determine, in each preset frequency band, a first energy of the predicted audio in the frequency domain and a second energy of the label audio in the frequency domain;

[0165] A distortion metric calculation unit, configured to calculate a distortion metric in the frequency domain based on a first preset weight corresponding to each preset frequency band and an energy difference, wherein the energy difference is a difference between a first energy and a second energy corresponding to the same preset frequency band;

[0166] An evaluation result calculation unit is used to calculate the audio quality evaluation result corresponding to the predicted audio according to the distortion metric and a preset evaluation result determination function, wherein the distortion metric is negatively correlated with the audio quality evaluation result.

[0167] As an implementation of an embodiment of the present application, the energy calculation unit may include:

[0168] The first energy calculation subunit is used to calculate the first energy of the predicted audio in the frequency domain in each preset frequency band based on the predicted audio in the frequency domain and its corresponding second preset weight according to the following formula: :

[0169] ;

[0170] The second energy calculation subunit is used to calculate the second energy of the label audio in the frequency domain in each preset frequency band based on the label audio in the frequency domain and its corresponding third preset weight according to the following formula: :

[0171] ;

[0172] in, For each preset frequency in the preset frequency band, Take the frequency Frequency value And the time is The predicted audio at Take the frequency Frequency value And the time is When the label audio, is the second preset weight, is the third preset weight.

[0173] As an implementation manner of the embodiment of the present application, the above-mentioned distortion metric calculation unit may include:

[0174] The distortion metric calculation subunit is used to calculate the distortion metric in the frequency domain according to the following formula based on the difference between the modified first energy and the modified second energy corresponding to each preset frequency band and the first preset weight corresponding to each preset frequency band: :

[0175] ;

[0176] in, For the A first preset weight corresponding to a preset frequency band, For the The first energy corresponding to the preset frequency band, For the The second energy corresponding to a preset frequency band.

[0177] As an implementation of the embodiment of the present application, the above parameter value adjustment submodule may include:

[0178] A gradient calculation unit, used to calculate, for each audio tuning parameter, a current partial derivative value of the loss function with respect to the audio tuning parameter as the gradient of the audio tuning parameter, wherein the current partial derivative value is a partial derivative value corresponding to when the parameter value of the audio tuning parameter is the current parameter value;

[0179] The parameter value updating unit is used to update the parameter value of the audio debugging parameter to the difference between the current parameter value and the updated value, wherein the updated value is the product of a preset learning rate and the gradient.

[0180] As an implementation method of an embodiment of the present application, the above-mentioned preset debugging method can be an audio quality enhancement algorithm, and the above-mentioned audio debugging parameters can include high-pass filter parameters, acoustic echo cancellation parameters, automatic noise suppression parameters, automatic gain control parameters, equalizer parameters and dynamic range control parameters.

[0181] The present application also provides an electronic device, such as Fig.10 As shown, including:

[0182] Memory 1001, used for storing computer programs;

[0183] The processor 1002 is used to implement the audio debugging method described in any of the above embodiments when executing the program stored in the memory 1001.

[0184] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 1002, the communication interface, and the memory 1001 communicate with each other via the communication bus.

[0185] It can be seen that in the scheme provided by the embodiment of the present application, the electronic device can obtain the audio to be debugged; input the audio to be debugged into the pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged based on the parameter values ​​of the audio debugging parameters obtained during the training process, according to the preset debugging method, obtains the debugged audio, and outputs the debugged audio; wherein the audio debugging parameters are the same as the parameters included in the preset debugging method, the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters, and the audio debugging model is: based on multiple sample audios and the label audio corresponding to each sample audio, in the direction of reducing the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio, the parameter value of the current audio debugging model is adjusted, and the debugged sample audio is debugged by the parameter value of the current audio debugging model. Since the preset debugging method can perform audio debugging on the audio to be debugged based on the included parameters, the audio debugging parameters included in the audio debugging model can be made the same as the parameters included in the preset debugging method. In this way, the parameters included in the preset debugging method can be used as trainable weights in the audio debugging model. During the training process of the audio debugging model, the parameter values ​​of the audio debugging parameters can be obtained. There is no need to manually adjust the parameter values ​​of the audio debugging parameters, which can save the manpower cost required for audio debugging and improve the audio quality of the debugged audio.

[0186] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0187] The communication interface is used for communication between the above electronic device and other devices.

[0188] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0189] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0190] In another embodiment provided in the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above audio debugging methods are implemented.

[0191] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any of the audio debugging methods in the above embodiments.

[0192] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state hard disk (SSD), etc.

[0193] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0194] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device, computer-readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0195] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. An audio debugging method, characterized in that: The method comprises: Get the audio to be debugged; Inputting the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, obtains the debugged audio, and outputs the debugged audio; Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

2. The method according to claim 1, characterized in that The training method of the audio debugging model includes: Acquire multiple sample audios and a label audio corresponding to each sample audio, wherein noise included in the label audio is lower than a preset noise threshold; Inputting the sample audio into an initial audio debugging model, so that the initial audio debugging model performs audio debugging on the sample audio based on parameter values ​​of current audio debugging parameters of the initial audio debugging model to obtain predicted audio; Based on the difference between the predicted audio and its corresponding label audio, the parameter values ​​of the audio tuning parameters of the initial audio tuning model are adjusted by back propagation until the initial audio tuning model meets the convergence condition, thereby obtaining a trained audio tuning model.

3. The method according to claim 2, characterized in that The step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation based on the difference between the predicted audio and the corresponding label audio includes: Based on the difference between the predicted audio and the corresponding labeled audio, according to a preset audio quality assessment algorithm, calculate the audio quality assessment result corresponding to the predicted audio; Determining a function value of a preset loss function according to the audio quality evaluation result, wherein the audio quality evaluation result is negatively correlated with the function value of the preset loss function; The parameter values ​​of the audio tuning parameters of the initial audio tuning model are adjusted by back propagation in a direction that reduces the function value of the preset loss function.

4. The method according to claim 3, characterized in that The step of calculating the audio quality evaluation result corresponding to the predicted audio according to a preset audio quality evaluation algorithm based on the difference between the predicted audio and the corresponding label audio includes: Convert the predicted audio and its corresponding labeled audio from the time domain to the frequency domain to obtain the predicted audio in the frequency domain and the labeled audio in the frequency domain; In each preset frequency band, determining a first energy of the predicted audio in the frequency domain and a second energy of the label audio in the frequency domain; Calculating a distortion metric in the frequency domain based on a first preset weight and an energy difference corresponding to each preset frequency band, wherein the energy difference is a difference between a first energy and a second energy corresponding to the same preset frequency band; According to the distortion metric, a function is determined according to a preset evaluation result to calculate an audio quality evaluation result corresponding to the predicted audio, wherein the distortion metric is negatively correlated with the audio quality evaluation result.

5. The method according to claim 4, characterized in that The step of determining the first energy of the predicted audio in the frequency domain and the second energy of the label audio in the frequency domain in each preset frequency band includes: In each preset frequency band, based on the predicted audio in the frequency domain and its corresponding second preset weight, the first energy of the predicted audio in the frequency domain is calculated according to the following formula: : ; In each preset frequency band, based on the label audio in the frequency domain and its corresponding third preset weight, the second energy of the label audio in the frequency domain is calculated according to the following formula: : ; in, For each preset frequency in the preset frequency band, Take the frequency Frequency value And the time is The predicted audio at Take the frequency Frequency value And the time is When the label audio, is the second preset weight, is the third preset weight.

6. The method according to claim 4, characterized in that The step of calculating the distortion metric in the frequency domain based on the first preset weight and energy difference corresponding to each preset frequency band includes: Based on the difference between the modified first energy and the modified second energy corresponding to each preset frequency band and the first preset weight corresponding to each preset frequency band, the distortion metric in the frequency domain is calculated according to the following formula: : ; in, For the A first preset weight corresponding to a preset frequency band, For the The first energy corresponding to the preset frequency band, For the The second energy corresponding to a preset frequency band.

7. The method according to any one of claims 3 to 6, characterized in that: The step of adjusting the parameter value of the audio tuning parameter of the initial audio tuning model by back propagation in a direction that reduces the function value of the preset loss function comprises: For each audio tuning parameter, calculate the current partial derivative value of the loss function with respect to the audio tuning parameter as the gradient of the audio tuning parameter, wherein the current partial derivative value is the partial derivative value corresponding to when the parameter value of the audio tuning parameter is the current parameter value; The parameter value of the audio debugging parameter is updated to the difference between the current parameter value and the updated value, wherein the updated value is the product of a preset learning rate and the gradient.

8. The method according to any one of claims 1 to 6, characterized in that: The preset debugging mode is an audio quality enhancement algorithm, and the audio debugging parameters include high-pass filter parameters, acoustic echo cancellation parameters, automatic noise suppression parameters, automatic gain control parameters, equalizer parameters, and dynamic range control parameters.

9. An audio debugging device, characterized in that: The device comprises: The audio acquisition module to be debugged is used to acquire the audio to be debugged; An audio debugging module, used for inputting the audio to be debugged into a pre-trained audio debugging model, so that the audio debugging model performs audio debugging on the audio to be debugged according to a preset debugging method based on the parameter values ​​of the audio debugging parameters obtained during the training process, obtains the debugged audio, and outputs the debugged audio; Among them, the audio debugging parameters are the same as the parameters included in the preset debugging method, and the preset debugging method is used to perform audio debugging on the audio to be debugged based on the included parameters. The audio debugging model is: based on multiple sample audios and a label audio corresponding to each sample audio, the parameter value of the current audio debugging model is adjusted in a direction that reduces the difference between the audio quality of the debugged sample audio and the audio quality of the label audio corresponding to the sample audio. The debugged sample audio is obtained by debugging the sample audio using the parameter value of the current audio debugging model.

10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-8 when executing a program stored in a memory.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.