Effector parameter evaluation model generation method and audio processing method with sound effects
By generating the effector parameter evaluation model, the deep learning neural network is used to extract parameter estimates from the two-dimensional spectral diagram of the audio effector, solving the problem of complex use of audio effectors and inaccurate parameter settings, and improving the sound effect and improving the user experience.
Patent Information
- Application Number
- CN202111348150.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-15
AI Technical Summary
The use of existing audio effects is complicated, making it difficult for users to understand and set the correct parameters, resulting in poor sound effects processing, and excessive parameter settings may lead to unstable working state of the effects.
By generating the effector parameter evaluation model, a deep learning neural network extracts parameter estimates from a two-dimensional spectrogram with sound-effect audio, and iterates over the network until training is completed, generating a model for evaluating the audio effector parameters.
This method can effectively evaluate key mixing parameters, accurately estimate the parameter value size, and identify important parameters, thereby reducing user threshold, improving the user experience of audio processing, and improving the working stability of the audio effector.
Smart Images

Figure CN114067838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a method for generating an effector parameter evaluation model and a method for processing audio with sound effects. In addition, the present invention also relates to related electronic equipment and storage media. Background Art
[0002] Audio effects are important tools for audio processing. They can mix various sound effects into the original audio to obtain new audio playback effects. Audio effects may be used in the entire process from music production to playback. For example, in the music production and mixing stages, audio effects are important tools for professionals in mixing. In the music playback end, such as in music playback applications, audio effects can effectively improve the user experience of the majority of users listening to music.
[0003] With the development of sound processing technology, the types of audio effects have gradually increased. Different types of audio effects require a variety of parameters to be set when used, which greatly increases the difficulty of use for users and makes it difficult for users to understand the physical meaning of various setting parameters. Most users can only set parameters based on subjective hearing and experience. In addition, faced with a variety of setting parameters, users may not be able to distinguish the importance of effect parameters, and inaccurate settings may result in poor sound processing effects. Furthermore, if too many effect parameters are set, the working state of the audio effect may become unstable, which in turn affects the mixing effect.
[0004] The content of this background technology description is only for facilitating understanding of the relevant technology in this field and is not regarded as an admission of the prior art. Summary of the invention
[0005] Therefore, the embodiments of the present invention intend to provide a method and device for generating an effector parameter evaluation model, a method and device for processing audio with sound effects, and related electronic devices and storage media. The scheme provided can be used to effectively evaluate key mixing parameters, and also provides the possibility for accurately estimating the numerical values of related parameters.
[0006] In a first aspect, a method for generating an effector parameter evaluation model is provided, the method comprising:
[0007] Obtain multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtain parameter value labels of the audios with sound effects, wherein each type of audio effector has multiple effector parameters, and the parameter value labels of the effect sounds correspond to the effector parameters of the audio effector that generates the audios with sound effects;
[0008] Extract multiple audio segments from each of the audios with sound effects, and generate a two-dimensional spectrogram corresponding to each of the audio segments;
[0009] Inputting the two-dimensional spectrogram of the audio clip into a deep learning neural network, obtaining a plurality of parameter estimation values output by the deep learning neural network, and iteratively updating the deep learning neural network based on the difference between the parameter estimation values and the parameter value labels corresponding to the audio clip until the training is completed;
[0010] The trained deep learning neural network is used as an effector parameter evaluation model; wherein the effector parameter evaluation model has a weight factor corresponding to each effector parameter of each type of audio effector, and the weight factor is used to represent the degree of influence of the effector parameter on the generated audio with sound effects.
[0011] In an embodiment of the present invention, the step of obtaining a plurality of audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value labels of the audios with sound effects includes:
[0012] For each type of audio effector, randomly generate N groups of parameter values of the audio effector, where different parameter values in the same group correspond to different effector parameters of the audio effector of this type, and N is greater than 1;
[0013] Using the N groups of parameter values to set the audio effectors, obtaining N audio effectors;
[0014] N dry audios are respectively input into the N audio effectors to generate multiple audios with sound effects of this type of audio effectors, and the parameter values set by the audio effectors that generate the audios with sound effects are determined as parameter value labels of the audios with sound effects.
[0015] In the embodiment of the present invention, after randomly generating N groups of parameter values of the audio effector, before using the N groups of parameter values to set the audio effector, the method further includes:
[0016] Delete the parameter group that does not meet the screening condition, wherein the screening condition is that the parameter value of the predetermined parameter in the parameter group is non-zero or the number of the parameter values that is zero is less than a predetermined threshold.
[0017] In the embodiment of the present invention, for each type of audio effector, randomly generating N groups of parameter values of the audio effector includes:
[0018] Normalizing parameter ranges of multiple effector parameters of the audio effector;
[0019] N groups of parameter values are randomly generated within the normalized parameter range.
[0020] In an embodiment of the present invention, the step of obtaining a plurality of audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value labels of the audios with sound effects includes:
[0021] Obtain multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector;
[0022] Analyze each of the audios with sound effects to obtain multiple parameter values of an audio effector used to generate the audios with sound effects;
[0023] The parameter value obtained from analyzing each audio with sound effect is determined as the parameter value label of the audio with sound effect.
[0024] In an embodiment of the present invention, the deep learning neural network includes a shared network and a plurality of subtask modules connected to the shared network, and different subtask modules correspond to different effector parameters and different weight factors of the same audio effector;
[0025] The step of inputting the two-dimensional spectrogram of the audio segment into a deep learning neural network, obtaining a plurality of parameter estimation values output by the deep learning neural network, and iteratively updating the deep learning neural network based on the difference between the parameter estimation values and the parameter value labels corresponding to the audio segment until the training is completed comprises:
[0026] Inputting the two-dimensional spectrogram into a deep learning neural network to obtain parameter estimation values outputted by each of the subtask modules;
[0027] Calculate the subtask loss value of the subtask module based on the parameter estimation value and parameter value label corresponding to the same subtask module;
[0028] Based on the subtask loss values and weight factors corresponding to each subtask module, the target loss value is calculated by weighted summation;
[0029] The target loss value is used to iteratively update the deep learning neural network and the weight factors corresponding to each subtask module until a predetermined convergence condition is met.
[0030] Optionally, the subtask loss value is determined based on a subtask loss function, and the subtask loss function is selected from at least one of a mean square error loss function, a root mean square error loss function, a mean absolute error loss function, and a cross entropy loss function.
[0031] In the embodiment of the present invention, based on the subtask loss values and weight factors corresponding to each subtask module, the target loss value is calculated by weighted summation, including:
[0032] If the parameter value label of the effector parameter corresponding to the subtask module is 0, the subtask loss value of the subtask module does not participate in the calculation of the target loss value.
[0033] In the embodiment of the present invention, the weighted summation based on the subtask loss values and weight factors corresponding to each subtask module respectively calculates the target loss value, including:
[0034] The target loss value is calculated based on the target loss function described below:
[0035]
[0036] Among them, Loss is the target loss value, loss is the subtask loss value, α is the weight factor, n is the sequence number of the subtask module, N is the number of subtask modules; y is the parameter value label.
[0037] In an embodiment of the present invention, the shared network includes multiple convolutional layers, a flat layer connected to the convolutional layers, one or more time-related recurrent neural network layers connected to the flat layers, and a fully connected layer connected to the one or more time-related recurrent neural network layers.
[0038] In an embodiment of the present invention, the one or more time-related recurrent neural network layers include a first bidirectional gated recurrent unit (GRU) layer, a second GRU layer connected to the forward output of the first bidirectional GRU layer, and a third GRU layer connected to the backward output of the first bidirectional GRU layer.
[0039] In a second aspect, a method for processing audio with sound effects is provided, comprising:
[0040] Obtaining audio with sound effects and the type of an audio effector that generates the audio with sound effects, where the type of audio effector has multiple effector parameters;
[0041] Extract at least one audio segment from the audio data with sound effects, and generate a two-dimensional spectrogram of the audio segment;
[0042] The two-dimensional spectrogram is input into an effector parameter evaluation model, and parameter estimation values of each effector parameter of the audio effector and weight factors of the effector parameters are output; wherein the effector parameter evaluation model has a weight factor corresponding to each effector parameter of each type of audio effector in at least one type of audio effector, and the weight factor is used to represent the degree of influence of the effector parameter on the generated audio with sound effects.
[0043] In an embodiment of the present invention, the method further includes:
[0044] Based on the values of the weight factors, the plurality of effector parameters of the audio effector are graded to be divided into basic parameters and high-order parameters;
[0045] Based on the classification results, the display of the effector parameters is set.
[0046] In the embodiment of the present invention, the step of setting the display of the effector parameters based on the classification result includes:
[0047] In response to a user triggering operation on the audio effector icon, a basic setting page is displayed, wherein the basic parameters and the corresponding estimated values of the parameters and an icon for entering an advanced setting page are presented on the basic setting page;
[0048] In response to a user triggering operation on the icon for entering the advanced setting page, the advanced setting page is displayed, and the advanced parameters and corresponding parameter estimation values are presented in the advanced setting page.
[0049] In a third aspect, an electronic device is provided, comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute the method described in any embodiment of the present invention when running the computer program.
[0050] In a fourth aspect, a storage medium is provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the method described in any embodiment of the present invention when executed.
[0051] The effector parameter evaluation model trained by the effector parameter evaluation model training method proposed in the embodiment of the present application can not only accurately estimate the numerical value of each effector parameter, but also identify which parameters are crucial and which parameters determine the overall direction of the audio processing effect. By analyzing and evaluating the effector parameters, the user's understanding of the effector parameters is enhanced, thereby lowering the threshold for using the audio effector, and the user experience is better. In addition, the effector parameter evaluation model training method proposed in the embodiment of the present application has the ability to train an effector parameter evaluation model with excellent generalization, and the effector parameter evaluation model can be trained using various types of audio effector training data, and can therefore be used to evaluate the parameters of various types of audio effectors.
[0052] Optional features and other effects of the embodiments of the present invention are partially described below and partially understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The elements shown are not limited to the proportions shown in the drawings and the same or similar reference numerals in the drawings represent the same or similar elements, wherein:
[0054] Figure 1 A first exemplary flow chart of a method according to an embodiment of the present invention is shown;
[0055] Figure 2 A second exemplary flow chart of a method according to an embodiment of the present invention is shown;
[0056] Figure 3 A third exemplary flow chart of a method according to an embodiment of the present invention is shown;
[0057] Figure 4 A fourth exemplary flow chart of a method according to an embodiment of the present invention is shown;
[0058] Figure 5 A fifth exemplary flow chart of a method according to an embodiment of the present invention is shown;
[0059] Figure 6 A schematic diagram of a model implementing an embodiment of the present invention is shown;
[0060] Figure 7 A schematic diagram showing a neural network according to an embodiment of the present invention is shown.
[0061] Figure 8 A schematic diagram of an application program interface provided by an embodiment of the present invention is shown;
[0062] Fig. 9 A first structural schematic diagram of a device according to an embodiment of the present invention is shown;
[0063] Fig.10 A second structural schematic diagram of a device according to an embodiment of the present invention is shown;
[0064] Fig.11 A schematic diagram of an exemplary hardware structure of an electronic device capable of implementing the method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific implementation methods and drawings. Here, the exemplary implementation methods of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0066] The embodiments of the present invention provide a method and device for generating an effector parameter evaluation model and a method and device for processing audio with sound effects. The method can be implemented with the aid of one or more computers. In some embodiments, the device can be implemented by software, hardware, or a combination of software and hardware. In some embodiments, the electronic device or computer can be implemented by a computer described herein or other electronic devices that can implement corresponding functions.
[0067] The scheme (method, device, electronic device or storage medium, etc.) of the embodiment of the present invention can be applied to audio processing occasions, especially to audio effectors. The scheme of the embodiment of the present invention can be applied to the music production stage (production end), and can also be used in the music playback stage (playing end). For example, in music production such as the mixing stage, professionals can use audio effectors for mixing, and the audio effectors can utilize the scheme described in the embodiment of the present invention or be implemented by it, or have parameters (parameter values) obtained by the scheme described in the embodiment of the present invention. For example, when playing music, users can choose different audio effectors and can adjust the parameters of the effectors as needed. The audio effectors can also utilize the scheme described in the embodiment of the present invention or be implemented by it, or have parameters (parameter values) obtained by the scheme of the embodiment of the present invention.
[0068] like Figure 1 As shown, an embodiment of the present invention provides a method for generating an effector parameter evaluation model. The method specifically includes steps S101-S104:
[0069] S101: Obtain multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtain parameter value labels of the audios with sound effects, wherein each type of audio effector has multiple effector parameters, and the parameter value labels of the audios with sound effects correspond to the effector parameters of the audio effector that generates the audios with sound effects.
[0070] In an embodiment of the present invention, audio effectors cover various types of processing units / devices for adding sound effects to audio, which may include but are not limited to reverberators, compressors, pitch shifters, etc. In some embodiments of the present invention, the method can be used to generate an evaluation model that is applicable to at least one, preferably multiple types of audio effectors, which will be further described below. In an example below, one type of audio effector is a reverberator, such as an mverb reverberator (artificial reverberation).
[0071] In the embodiment of the present invention, the audio with sound effects is audio that is processed by an audio effector to add sound effects, whereby the processing by the audio effector adds audio effects to the processed audio.
[0072] In some embodiments of the present invention, the audio with sound effects includes reverberant audio.
[0073] In the embodiment of the present invention, each type of audio effector has a plurality of effector parameters. For example, the effector parameters of a reverberator such as an mverb reverberator (artificial reverberation) include high frequency attenuation degree, density, reverberation bandwidth range, attenuation size, late reverberation delay, space size, dry-wet mixing ratio, and early reflection to late reverberation ratio. Those skilled in the art will appreciate that different types of more or fewer effector parameters may be selected for different audio effectors.
[0074] In an embodiment of the present invention, different audio effectors have different effector parameters. As described below, the effector parameters of different types of audio effectors may not overlap or partially overlap, and the parameter range of the overlapping effectors of different types of audio effectors may be identical or unequal, which all falls within the scope of the present invention.
[0075] In an embodiment of the present invention, the parameter value label relates to the true value of the effector parameter to be predicted, and can be used to train the effector parameter evaluation model according to an embodiment of the present invention.
[0076] In some embodiments of the present invention, audio with sound effects, such as reverberation audio, can be obtained by combining randomly generated parameter values with dry music. In this case, the randomly generated parameter values can be used directly or after processing as multiple parameter value labels of the audio with sound effects, such as reverberation audio. This construction method of reverberation audio and its parameter value labels enables the trained model to have a more generalized ability.
[0077] In a specific embodiment, Figure 2 As shown, obtaining multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value labels of the audios with sound effects can be achieved through steps S201-S203.
[0078] S201: For each type of audio effector, randomly generate N groups of parameter values of the audio effector, where different parameter values in the same group correspond to different effector parameters of the type of audio effector, and N is greater than 1.
[0079] S202: Use the N groups of parameter values to set the audio effectors to obtain N audio effectors.
[0080] S203: Input N dry audios into the N audio effectors respectively to generate multiple audios with sound effects of the same type of audio effectors, and determine the parameter values set by the audio effectors for generating the audios with sound effects as parameter value labels of the audios with sound effects.
[0081] In some embodiments, the parameter value range within the group can be normalized, and such normalization has special benefits, such as not only saving training resources, but also making the generated evaluation model suitable for processing multiple different types of audio effectors. For example, as previously described, when different types of audio effectors have overlapping effector parameters, but the parameter ranges of the overlapping effector parameters are different, performing such normalization allows the effector parameter evaluation model to be trained with less training data.
[0082] Wherein, step S201 may specifically include two steps, namely: step A1 normalizes the parameter ranges of multiple effector parameters of the audio effector; and step A2 randomly generates N groups of parameter values within the normalized parameter range. Here, step S203 may specifically include directly using the randomly generated parameter value of the normalized parameter range as a parameter value label.
[0083] In a further embodiment, especially before determining the parameter value label and setting the audio effector, after obtaining the audio with sound effects and the corresponding multiple parameter value labels, the following steps may also be included: deleting parameter groups that do not meet the filtering conditions, wherein the filtering conditions are that the parameter values of the predetermined parameters in the parameter group are non-zero values or the number of parameter values that are zero is less than a predetermined threshold.
[0084] Combined with reference Figure 1 , Figure 2 and Figure 6 , describing an exemplary production process of audio with sound effects (training data) with parameter value labels.
[0085] As mentioned above, taking the mverb reverberator (artificial reverberation) as an example, its setting parameters include 8 parameters such as high-frequency attenuation degree, density, reverberation bandwidth range, attenuation size, late reverberation delay, space size, dry-wet mixing ratio, and early reflection and late reverberation ratio. Optionally, the values of these parameters can be normalized so that their value range is concentrated between 0 and 1. Then, the 8 parameters are randomly selected within this value range to obtain multiple groups of randomly generated parameter values. In this example, each group has 8 parameter values, and each parameter value corresponds to the aforementioned 8 parameters.
[0086] Optionally, it is determined whether each group of randomly generated parameter values meets the preset screening conditions. In some embodiments, it is determined whether the parameter value of the predetermined parameter in the parameter group is zero. For example, since the remaining parameters do not work when the parameter dry-wet mixing ratio is 0, and some parameters will not work when the parameter early reflection to late reverberation ratio is 0, the parameter values of 0 in the above-mentioned dry-wet mixing ratio and early reflection to late reverberation ratio can be ignored, that is, the corresponding parameter value group is deleted. In some embodiments, the parameter value group with too many zero values can be directly deleted. For example, the parameter group with the number of zero parameter values greater than or equal to 3, 4, 5 or 6 can be deleted.
[0087] The mverb reverberator of this type can be set with the aforementioned randomly generated and optionally screened multiple sets of parameter values, and multiple dry audios without reverberation effects (also referred to as dry signals) can be selected as inputs of the mverb reverberator set by the multiple sets of parameter values to obtain multiple reverberation audios. Repeating this process, combining different dry audio inputs with different reverberation effects (setting different sets of parameter values), sufficient reverberation data with multiple parameter value labels can be obtained (at this time, the parameter values used to set the reverberator are multiple parameter value labels, and the multiple parameter value labels correspond to multiple parameters of the reverberator) for neural network model training.
[0088] The training data of other audio effectors or multiple categories of audio effectors, that is, the generation of audio with sound effects and parameter value labels of corresponding audio effectors can be obtained similarly.
[0089] In other embodiments of the invention, the audio with sound effects can be directly obtained, and then the parameter values of the audio effector used in the audio with sound effects are extracted from the audio with sound effects, and then multiple parameter value tags of the audio with sound effects are determined. Figure 3 As shown, obtaining audio data with sound effects and multiple parameter value tags includes steps S301-S303.
[0090] S301: Obtain multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector.
[0091] S302: Analyze each of the audios with sound effects to obtain multiple parameter values of an audio effector used to generate the audios with sound effects.
[0092] S303: Determine multiple parameter value labels for each audio with sound effects based on the extracted multiple effector parameter values.
[0093] Specifically, audio with sound effects can be obtained, for example, by obtaining the usage data of users of the music APP or other music production data. As an example, a large amount of user-side usage data is deposited in the music APP, including data on how users call different effectors to process the original sound when listening to songs. Therefore, in some embodiments, step S302 also includes: extracting multiple effector parameter values used for each audio with sound effects respectively. Specifically, it can be obtained by obtaining the user's audio effector setting history. At this point, parameter value labels for audio with sound effects can be obtained based on these parameter values. Those skilled in the art will understand that the above-mentioned parameter value labels can be directly determined by the parameter values or by processing the parameter values (such as the above-mentioned normalization, etc.).
[0094] S102: extracting a plurality of audio segments from each of the audios with sound effects, and generating a two-dimensional spectrogram corresponding to each of the audio segments.
[0095] In some embodiments, the two-dimensional spectrogram is a speech frequency spectrum graph, which can be obtained by processing a time domain signal.
[0096] In some embodiments, a two-dimensional spectrogram can be obtained by a short-time Fourier transform (STFT). Specifically, a short-time Fourier transform can be performed on the audio with sound effects by continuously adding a window function to obtain a spectrogram of a frame (or a small segment) based on the window function. In some embodiments of the present invention, the segment duration can be consistent with the time length intercepted by the window function. In other embodiments, the segment duration can be different from the time length intercepted by the window function. For example, in some embodiments, the spectrogram intercepted by the window function can be synthesized into a spectrogram of a longer duration segment for input into the neural network below.
[0097] Exemplarily, the window function is a Hamming window, preferably the window length is 1024, the step length is 512, the obtained energy spectrum resolution is 512×1, and the time resolution can be 32ms. It should be noted that the values of the above parameters are only exemplary, and other window functions can be used, and different window lengths, step sizes and time resolutions can be selected.
[0098] In some embodiments, the method may further include step B1: before intercepting, downsampling the audio with sound effects according to a predetermined sampling rate.
[0099] S103: Inputting the two-dimensional spectrogram of the audio segment into a deep learning neural network, obtaining a plurality of parameter estimation values output by the deep learning neural network, and iteratively updating the deep learning neural network based on the difference between the parameter estimation values and the parameter value labels corresponding to the audio segment until the training is completed.
[0100] S104: Using the trained deep learning neural network as an effector parameter evaluation model; wherein the effector parameter evaluation model has a weight factor corresponding to each effector parameter of each type of audio effector, and the weight factor is used to represent the degree of influence of the effector parameter on the generated audio with sound effects.
[0101] In an embodiment of the present invention, the deep learning neural network includes a shared network and multiple subtask modules connected to the shared network, and different subtask modules correspond to different effector parameters and different weight factors of the same audio effector.
[0102] In some embodiments, the shared network includes a shared network including multiple convolutional layers, a flat layer connecting the convolutional layers, one or more time-dependent recurrent neural network layers connecting the flat layers, and a fully connected layer connecting the one or more time-dependent recurrent neural network layers.
[0103] like Figure 7 As shown, one or more time-series related recurrent neural network layers include a first bidirectional gate recurrent unit (GRU) layer, a second GRU layer connected to the forward output of the first bidirectional GRU layer, and a third GRU layer connected to the backward output of the first bidirectional GRU layer. Among them, GRU (Gate Recurrent Unit) is a type of recurrent neural network (RNN), which can be used to solve problems such as long-term memory and gradient disappearance in back propagation.
[0104] In a further embodiment, Figure 4 As shown, the two-dimensional spectrogram of the audio segment is input into a deep learning neural network, a plurality of parameter estimation values output by the deep learning neural network are obtained, and the deep learning neural network is iteratively updated based on the difference between the parameter estimation values and the parameter value labels corresponding to the audio segment until the training is completed, which may specifically include steps S401-S404. Among them, S401: inputting the two-dimensional spectrogram into a deep learning neural network to obtain parameter estimation values respectively output by each of the subtask modules; S402: calculating the subtask loss value of the subtask module based on the parameter estimation values and parameter value labels corresponding to the same subtask module; S403: calculating the target loss value by weighted summation based on the subtask loss values and weight factors respectively corresponding to each subtask module; S404: using the target loss value to iteratively update the deep learning neural network and the weight factors corresponding to each subtask module until the predetermined convergence condition is met.
[0105] In the embodiment of the present invention, the iterative update is implemented based on the gradient descent method.
[0106] In an embodiment of the present invention, the subtask module can be initially set to a non-zero 1*M matrix, where M can be determined according to the matrix dimension of the upstream neural network layer to which the subtask module is connected, for example, according to the matrix dimension of the aforementioned fully connected layer.
[0107] In some embodiments, the subtask loss value may be determined by a subtask loss function. In some embodiments, the subtask loss function may include at least one of a mean square error loss function, a root mean square error loss function, a mean absolute error loss function, and a cross entropy loss function.
[0108] In one embodiment, step S403 may specifically include step C1: if the parameter value label of the effector parameter corresponding to the subtask module is 0, the subtask loss value of the subtask module does not participate in the calculation of the target loss value.
[0109] In a further embodiment, step C1 may specifically include step C11: calculating the target loss value based on the following target loss function formula (1):
[0110]
[0111] Among them, Loss is the target loss value, loss is the subtask loss value, α is the weight factor, n is the sequence number of the subtask module, N is the number of subtask modules; y is the parameter value label.
[0112] Combination Figure 1-7 , the specific process of iterative training of neural networks is explained.
[0113] First, different training data is input according to the different types of audio effectors selected for training. For example, from audio with sound effects with parameter value labels (for example, for reverberators), n seconds of audio (signal) are taken as input each time. It should be noted that it is converted into a two-dimensional spectrogram before being input into the shared network layer of the neural network.
[0114] After passing through the shared network layer, it is split into subtasks with their respective effector parameters as the object, and each subtask estimates the value of its own effector parameter. The related multiple effector parameters are summarized into a group of effector setting parameters (i.e. parameter value group), and finally loaded into the audio effector.
[0115] Among them, according to the labels of the effector parameters and the estimated values of the effector parameters, the minimum mean square error of the output results of each subtask is calculated and recorded as the sub-loss (function) value loss. Different subtasks are distinguished by numbers 1, 2, 3, etc. The target loss (function) value is obtained by weighted summation between the sub-loss (function) values, that is, the following formula (2):
[0116] Loss=(α 1 *loss 1 -log(α 1 )+(α 2 *loss 2 -log(α 2 )+(α 3 *loss 3 -log(α 3 )+… (2)
[0117] Among them, α is the weight coefficient (or weight factor) of each subtask. It can be understood that the weight coefficient is updated through iteration.
[0118] In one example, when the parameter value label is 0, the corresponding subtask loss function value loss is not included in the calculation of the overall target loss function value. The target loss function can be formula (1). For example, when the third parameter value label (corresponding to subtask 3) is 0, the target loss function is in the form of formula (3):
[0119] Loss=(α 1 *loss 1 -log(α 1 )+(α 2 *loss 2 -log(α 2 )+.... (3)
[0120] In summary, the training of multi-task learning models can be achieved by using labeled training data, neural network models and loss functions.
[0121] In an embodiment of the present invention, by not including the corresponding subtask loss function value loss into the overall target loss function value Loss when the parameter value label is 0, the trained effector parameter evaluation model can be used for multiple different types of audio effectors, making the evaluation model more generalizable. In these embodiments, when training the effector parameter evaluation model, different training data can be fed to different audio effectors. In addition, considering that different effectors may have the same parameter types, the target loss function taught in an embodiment of the present invention allows the evaluation model to be trained with less data for multiple audio effectors with the same shared parameters. Here, the normalization of the aforementioned parameter range can further improve the processing efficiency for multiple audio effectors with the same shared parameters.
[0122] Thus, the method in some embodiments of the present invention can be used to evaluate multiple types of audio effectors, wherein the effector parameters do not overlap. At this time, multiple first subtask modules (and the first weight factor) correspond to multiple first effector parameters of the first (type) audio effector, and multiple second subtask modules (and the second weight factor) correspond to multiple second effector parameters of the second (type) audio effector (wherein the first and second subtask modules do not overlap).
[0123] It should be noted that the acquisition of audio with sound effects and its parameter value labels and the training of the neural network can be configured accordingly as follows to generate a method for evaluating multiple audio effectors.
[0124] In a specific implementation, obtaining multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value tags of the audios with sound effects may specifically include steps D1 and D2. Step D1: obtaining multiple first audios with sound effects and multiple first parameter value tags of each first audio with sound effects; step D2: obtaining multiple second audios with sound effects and multiple second parameter value tags of each second audio with sound effects.
[0125] Furthermore, multiple audio segments are intercepted from multiple audios with sound effects, and multiple two-dimensional spectrograms corresponding to the multiple audio segments are generated, including steps E1 and E2. Among them, step E1: intercept one or more first audio segments from the first audio with sound effects, and convert the first audio segments into a first two-dimensional spectrogram; step E2: intercept one or more second audio segments from the second audio with sound effects, and convert the second audio segments into a second two-dimensional spectrogram.
[0126] Furthermore, multiple two-dimensional spectrograms are input into a deep learning neural network, and the deep learning neural network is iteratively updated based on the difference between the output multiple parameter estimates and multiple parameter audio effect labels until the training is completed, which may include steps F1 and F2. Among them, step F1: input the first two-dimensional spectrogram into a multi-task deep learning neural network, and iteratively update the deep learning neural network based on the difference between the output multiple first parameter estimates and multiple first parameter value labels until convergence; step F2: input the second two-dimensional spectrogram into a multi-task deep learning neural network, and iteratively update the deep learning neural network based on the difference between the output second parameter estimates and multiple second audio effect parameter value labels until convergence.
[0127] Furthermore, the first two-dimensional spectrogram is input into a multi-task deep learning neural network, and the deep learning neural network is iteratively updated based on the difference between the output multiple first parameter estimates and multiple first parameter value labels until convergence, which may include: steps F11-F14. Among them, step F11: input the two-dimensional spectrogram obtained from the first audio with sound effects into the deep learning neural network, and obtain the first parameter estimates output by multiple parallel first subtask modules corresponding to the multiple first parameter value labels; step F12: calculate the subtask loss value of the first subtask module based on the parameter estimates and parameter value labels corresponding to the first subtask module; F13: calculate the target loss value by weighted summation based on the subtask loss values and weight factors corresponding to each subtask module of the first subtask module; step F14: use the target loss value to iteratively update the deep learning neural network and the weight factor corresponding to each first subtask module until the predetermined convergence condition is met.
[0128] In some embodiments, when using the first audio with sound effects for training, if the remaining subtasks are not taken into account when calculating the target loss function value, the parameter value labels other than the first parameter value label can be set to 0 in the aforementioned step D1; and / or, in step F12, only the subtask loss value corresponding to the first parameter value label is directly calculated.
[0129] The second audio band can be processed accordingly.
[0130] Additionally, the evaluation model generated according to the method of an embodiment of the present invention can be used to evaluate multiple types of audio effects, wherein effector parameters may overlap. At this point, multiple first subtask modules (and the first weight factor) correspond to multiple first effector parameters of the first audio effector, and multiple second subtask modules (and the second weight factor) correspond to multiple second effector parameters of the second audio effector (wherein the first and second subtask modules overlap). The method for generating a multi-audio effector evaluation model for effector parameters that may overlap can refer to the relevant embodiments of the method for generating a multi-audio effector evaluation model for effector parameters that do not overlap. For example, different training data (with sound effect audio and audio effector labels) are provided for different audio effects for iterative training, and when the training data for one of the audio effects is used to be trained, the subtasks corresponding to the parameters of other audio effects and their loss functions are not counted in the iteration (in this example, except for the overlapping parameters).
[0131] As mentioned above, in another embodiment of the present invention, a method for processing audio with sound effects may also be provided.
[0132] like Figure 5 As shown, the audio processing method with sound effects may include steps S501-S503. Among them:
[0133] S501: Obtain audio with sound effects and the type of an audio effector that generates the audio with sound effects, wherein the audio effector has multiple effector parameters.
[0134] S502: extract at least one audio segment from the audio data with sound effects, and generate a two-dimensional spectrogram of the audio segment.
[0135] S503: Input the two-dimensional spectrogram into an effector parameter evaluation model, and output parameter estimation values of each effector parameter of the audio effector and weight factors of the effector parameters.
[0136] The effector parameter evaluation model stores a weight factor corresponding to each effector parameter of each type of audio effector, and the weight factor is used to indicate the degree of influence of the effector parameter on the generated audio with sound effects.
[0137] In an embodiment of the present invention, the deep learning neural network is obtained by processing using the effector parameter evaluation model generation method of an embodiment of the present invention.
[0138] In a further embodiment, the audio processing method with sound effects may further include steps G1-G2. Step G1: grading the multiple effector parameters of the audio effector based on the value of the weight factor to divide them into basic parameters and high-order parameters; Step G2: setting the display of the effector parameters based on the grading results.
[0139] In a further embodiment, H2 may include step H21: in response to a user triggering an audio effector icon, displaying a basic settings page, presenting basic parameters and corresponding parameter estimates and an icon for entering an advanced settings page; and step H22: in response to a user triggering an icon for entering an advanced settings page, displaying an advanced settings page, presenting advanced parameters and corresponding parameter estimates.
[0140] As mentioned above, after completing the model training, the optimal value of each weight factor α can be obtained. The larger the α value, the more obvious the influence of the coefficient change corresponding to the subtask on the audio processing effect, and the smaller the α value, the opposite.
[0141] In one example, still taking the mverb reverberator as an example, the parameters are divided into gears according to the alpha values of each subtask. It is observed that among the 8 key parameters of the mverb reverberator, the alpha weights of four parameters, such as attenuation size, space size, dry-wet mixing ratio, and early reflection and late reverberation ratio, are in the same order of magnitude, belong to the same gear, and have relatively large values. Therefore, it can be considered that the settings of these four parameters will mainly affect the overall effect of the reverberator, and the remaining parameters are further fine-tuning of the reverberation effect based on the preliminary effect. Through the grading results, users can give priority to and focus on the settings of the above four parameters when using the mverb reverberator to produce sound effects. These four parameters determine the design direction of the reverberation effect.
[0142] Combined with reference Figure 8 For explanation. In one embodiment, when the user inputs an audio signal with a processing effect and an audio effector used to process the audio signal into a multi-task network model, the parameter estimation value of the effector and the grading result of the importance of each parameter can be obtained. The grading result is an evaluation of the parameter setting priority by the multi-task network model. Based on the evaluation result, the method of using the audio effector can be further optimized. Taking two levels as an example, the front-end interface can divide the relevant parameters into basic parameters and high-order parameters. Among them, the basic (general) parameters have a more obvious influence on the effector, which determines the design direction of the effector and can be used as the main adjustment target; while the high-order parameters describe more details of the effector, and ordinary users do not need to pay too much attention to them. Accordingly, a hierarchical display can be provided in the user interface. For example Figure 8 When the user clicks the audio effect icon, the basic setting page will be displayed, which displays a plurality of basic parameters and their estimated values, and an icon for instructing the user to enter the advanced setting page; when the user clicks the icon for instructing the user to enter the advanced user interface, the advanced page (not shown) will be displayed, which displays the advanced parameters and their estimated values. From the perspective of the user interface, the entire display interface will be more concise, clear and easy to operate.
[0143] like Fig. 9As shown, in some embodiments, an effector parameter evaluation model generation device 900 is also provided, which includes an acquisition unit 910, a conversion unit 920, a training unit 930 and a construction unit 940. The acquisition unit 910 is configured to acquire multiple audio with sound effects generated by each type of audio effector in at least one type of audio effector and acquire parameter value labels of the audio with sound effects, wherein each type of audio effector has multiple effector parameters, and the parameter value labels of the audio with sound effects correspond to the effector parameters of the audio effector that generates the audio with sound effects. The conversion unit 920 is configured to intercept multiple audio segments from each of the audio with sound effects and generate a two-dimensional spectrogram corresponding to each of the audio segments. The training unit 930 is configured to input the two-dimensional spectrogram of the audio segment into a deep learning neural network, obtain multiple parameter estimation values output by the deep learning neural network, and iteratively update the deep learning neural network based on the difference between the parameter estimation value and the parameter value label corresponding to the audio segment until the training is completed. The construction unit 940 is configured to use the trained deep learning neural network as an effector parameter evaluation model; wherein the effector parameter evaluation model has a weight factor corresponding to each effector parameter of each type of audio effector, and the weight factor is used to represent the degree of influence of the effector parameter on the generated audio with sound effects.
[0144] like Fig.10 As shown, in some embodiments, there is also provided an audio processing device 1000 with sound effects, which includes an acquisition unit 1010, a conversion unit 1020 and an input unit 1030. The acquisition unit 1010 is configured to obtain the audio with sound effects and the type of the audio effector that generates the audio with sound effects, and the audio effector has multiple effector parameters. The conversion unit 1020 is configured to intercept at least one audio segment from the audio with sound effects and generate a two-dimensional spectrogram of the audio segment. The input unit 1030 is configured to input the two-dimensional spectrogram into an effector parameter evaluation model, and output the parameter estimation values of each effector parameter of the audio effector and the weight factor of the effector parameter. In an embodiment of the present invention, the effector parameter evaluation model stores the weight factor corresponding to each effector parameter of each type of the audio effector, and the weight factor is used to indicate the degree of influence of the effector parameter on the generated audio with sound effects. In an embodiment of the present invention, the deep learning neural network can be obtained by pre-training using the method of any embodiment of the present invention.
[0145] In embodiments of the present invention, features directed to the method may be incorporated into the apparatus in a non-contradictory manner, and vice versa.
[0146] In an embodiment of the present invention, an electronic device is provided, which includes: a processor and a memory storing a computer program, wherein the processor is configured to implement any method according to an embodiment of the present invention when running the computer program. In addition, the electronic device can also be used to implement an apparatus according to any embodiment of the present invention.
[0147] Fig.11 A schematic diagram of an electronic device 1100 that can implement a method or implement an embodiment of the present invention is shown, and in some embodiments, more or fewer electronic devices than shown may be included. In some embodiments, it can be implemented using a single or multiple electronic devices. In some embodiments, it can be implemented using cloud or distributed electronic devices.
[0148] like Fig.11 As shown, the electronic device 1100 includes a central processing unit (CPU) 1101, which can perform various appropriate operations and processes according to the programs and / or data stored in the read-only memory (ROM) 1102 or the programs and / or data loaded from the storage part 1108 to the random access memory (RAM) 1103. CPU 1101 can be a multi-core processor, or it can include multiple processors. In some embodiments, CPU 1101 can include a general main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. CPU 1101, ROM 1102 and RAM 1103 are connected to each other via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0149] The processor and the memory are used together to execute the program stored in the memory. When the program is executed by the computer, the steps or functions of the sound effect data processing method or device described in the above embodiments can be realized.
[0150] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read therefrom is installed into the storage section 1108 as needed. Fig.11 Only some components are schematically shown in the figure, which does not mean that the computer system 1100 only includes Fig.11 Components shown.
[0151] The systems, devices, modules or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server or a combination thereof.
[0152] Although not shown, in some embodiments, a storage medium is further provided, storing a computer program, which is configured to execute any song identification method of the embodiments of the present invention when executed.
[0153] Storage media in embodiments of the present invention include permanent and non-permanent, removable and non-removable items that can be used to store information by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0154] The methods, programs, systems, devices, etc. of the embodiments of the present invention may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.
[0155] Those skilled in the art should understand that the embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, those skilled in the art can imagine that the implementation of the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware or a combination of software / hardware.
[0156] Unless explicitly stated, the actions or steps of the methods, programs, and embodiments of the present invention do not have to be performed in a specific order and can still achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0157] In this article, multiple embodiments of the present invention are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this article, "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" are meant to be applicable to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. In the absence of contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0158] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are merely examples of the best modes for implementing the present systems and methods. It will be appreciated by those skilled in the art that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A method for generating an effector parameter evaluation model, It is characterized in that The method comprises: Obtain multiple audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtain parameter value tags of the audios with sound effects, wherein each type of audio effector has multiple effector parameters, and the parameter value tags of the audios with sound effects correspond to the effector parameters of the audio effector that generates the audios with sound effects; Extract multiple audio segments from each of the audios with sound effects, and generate a two-dimensional spectrogram corresponding to each of the audio segments; Inputting the two-dimensional spectrogram of the audio clip into a deep learning neural network, obtaining a plurality of parameter estimation values output by the deep learning neural network, and iteratively updating the deep learning neural network based on the difference between the parameter estimation values and the parameter value labels corresponding to the audio clip until the training is completed, wherein the deep learning neural network includes a shared network and a plurality of subtask modules connected to the shared network; The trained deep learning neural network is used as an effector parameter evaluation model; wherein the effector parameter evaluation model has a weight factor corresponding to each effector parameter of each type of audio effector, and the weight factor is used to indicate the influence of the effector parameter on the generated audio with sound effects, and different subtask modules correspond to different effector parameters and different weight factors of the same audio effector.
2. The method according to claim 1, It is characterized in that The step of obtaining a plurality of audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value labels of the audios with sound effects comprises: For each type of audio effector, randomly generate N groups of parameter values of the audio effector, where different parameter values in the same group correspond to different effector parameters of the audio effector of this type, and N is greater than 1; Using the N groups of parameter values to set the audio effectors, obtaining N audio effectors; N dry audios are respectively input into the N audio effectors to generate multiple audios with sound effects of this type of audio effectors, and the parameter values set by the audio effectors that generate the audios with sound effects are determined as parameter value labels of the audios with sound effects.
3. The method according to claim 2, It is characterized in that After randomly generating N groups of parameter values of the audio effector, and before using the N groups of parameter values to set the audio effector, the method further includes: Delete the parameter group that does not meet the screening condition, wherein the screening condition is that the parameter value of the predetermined parameter in the parameter group is non-zero or the number of the parameter values that is zero is less than a predetermined threshold.
4. The method according to claim 2, It is characterized in that For each type of audio effector, N groups of parameter values of the audio effector are randomly generated, including: Normalizing parameter ranges of multiple effector parameters of the audio effector; N groups of parameter values are randomly generated within the normalized parameter range.
5. The method according to claim 1, It is characterized in that The step of obtaining a plurality of audios with sound effects generated by each type of audio effector in at least one type of audio effector and obtaining parameter value labels of the audios with sound effects comprises: Obtain multiple audio effects with sound for each type of audio effect in at least one type of audio effectors; Analyze each of the audio effects with sound to obtain multiple parameter values of the audio effector used to generate the audio effect with sound; Determine the parameter value tags of the audio effects with sound from the parameter values analyzed from each of the audio effects with sound.
6. The method according to claim 1, wherein, inputting the two-dimensional spectrogram of the audio segment into a deep learning neural network, obtaining multiple parameter estimates output by the deep learning neural network, and iteratively updating the deep learning neural network based on the difference between the parameter estimates and the parameter value tags corresponding to the audio segment until the training is completed, includes: Inputting the two-dimensional spectrogram into a deep learning neural network, and obtaining parameter estimates respectively output by each of the subtask modules; Calculating the subtask loss value of the subtask module based on the parameter estimates and parameter value tags corresponding to the same subtask module; Calculating the target loss value by weighted summation based on the subtask loss values and weight factors respectively corresponding to each subtask module; Iteratively updating the deep learning neural network and the weight factors corresponding to each subtask module by using the target loss value until a predetermined convergence condition is met.
7. The method according to claim 6, wherein, calculating the target loss value by weighted summation based on the subtask loss values and weight factors respectively corresponding to each subtask module, includes: If the parameter value tag of the effector parameter corresponding to the subtask module is 0, the subtask loss value of the subtask module does not participate in the calculation of the target loss value.
8. The method according to claim 7, wherein, calculating the target loss value by weighted summation based on the subtask loss values and weight factors respectively corresponding to each subtask module, includes: Calculating the target loss value based on the following target loss function: where Loss is the target loss value, loss is the subtask loss value, α is the weight factor, n is the serial number of the subtask module, N is the number of subtask modules; y is the parameter value tag.
9. The method according to claim 6, wherein, the shared network includes multiple convolutional layers, a flattening layer connecting the convolutional layers, one or more time-series related recurrent neural network layers connecting the flattening layer, and a fully connected layer connecting the one or more time-series related recurrent neural network layers.
10. The method according to claim 9, wherein, the one or more time-series related recurrent neural network layers include a first bidirectional gated recurrent unit layer, a second bidirectional gated recurrent unit layer connected to the forward output of the first bidirectional gated recurrent unit layer, and a third bidirectional gated recurrent unit layer connected to the backward output of the first bidirectional gated recurrent unit layer.
11. A method for processing audio effects with sound, wherein, includes: obtaining an audio effect with sound and the type of the audio effector that generates the audio effect with sound, the audio effector having multiple effector parameters; intercepting at least one audio segment from the audio effect with sound to generate a two-dimensional spectrogram of the audio segment; The two-dimensional spectrogram is input into an effector parameter evaluation model, and the parameter estimation values of each effector parameter of the audio effector and the weight factors of the effector parameters are output; wherein the effector parameter evaluation model stores the weight factors corresponding to each effector parameter of each type of the audio effector, and the weight factors are used to represent the degree of influence of the effector parameters on the generated audio with sound effects, and the effector parameter evaluation model includes the effector parameter evaluation model described in any one of claims 1 to 10.
12. The method according to claim 11, It is characterized in that Also includes: Based on the values of the weight factors, the plurality of effector parameters of the audio effector are graded to be divided into basic parameters and high-order parameters; Based on the classification results, the display of the effector parameters is set.
13. The method according to claim 12, It is characterized in that The step of setting the display of the effector parameters based on the classification result includes: In response to a user triggering operation on the audio effector icon, a basic setting page is displayed, wherein the basic parameters and the corresponding estimated values of the parameters and an icon for entering an advanced setting page are presented on the basic setting page; In response to a user triggering operation on the icon for entering the advanced setting page, the advanced setting page is displayed, and the advanced parameters and corresponding parameter estimation values are presented in the advanced setting page.
14. An electronic device, It is characterized in that include: A processor and a memory storing a computer program, wherein the processor is configured to perform the method according to any one of claims 1 to 13 when running the computer program.
15. A storage medium, It is characterized in that The storage medium stores a computer program, and the computer program is configured to perform the method according to any one of claims 1 to 13 when executed.
Citation Information
Patent Citations
Sound evaluation display method and device
CN107391076A
Method and system for intelligently adjusting sound effects
CN109905806A