Model training method and device, electronic equipment and storage medium

By entering audio frame parameters in the audio equalization model, outputting gain and updating model parameters, the problem of poor audio playback of micro speakers is solved, achieving more accurate audio equalization adjustment and better playback effects.

CN120343468APending Publication Date: 2025-07-18VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510456000.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the audio playback effect of the micro speaker is poor, which is manifested as low loudness, unbalanced frequency response and large noise, and the fixed equalization parameters lead to poor universality.

Method used

By inputting the audio frame parameters of the audio signal into the audio equalization model for processing, outputting gain and adjusting the spectrum amplitude, and updating the model parameters with the reward value, dynamic equalization adjustment is achieved.

Benefits of technology

It improves the effect of audio balance adjustment, meets users' balanced needs for audio signals, and improves the audio playback quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343468A_ABST
    Figure CN120343468A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, electronic equipment and a storage medium, and belongs to the technical field of audio processing, and the method comprises the steps: inputting a first audio frame parameter related to a first audio frame in a first audio signal into an audio equalization model for audio equalization processing, and outputting the gain of the first audio frame; adjusting the spectrum amplitude of the first audio frame by using the gain of the first audio frame to obtain an adjusted first audio frame; determining a first reward value based on the adjusted first audio frame; and on the basis of the first audio frame parameter, the gain of the first audio frame, the first reward value and a second audio frame parameter related to the second audio frame, updating the model parameter of the audio equalization model to obtain an updated audio equalization model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of model training, and particularly relates to a model training method, apparatus, electronic device, and storage medium. Background Art

[0002] With the popularization of electronic devices, the use of micro speakers is becoming more and more widespread. Due to the limited physical dimensions such as the cavity structure and the diaphragm vibration space of the micro speaker, there will be problems such as low loudness, unbalanced frequency response, and large noise of the audio played through the micro speaker, resulting in poor audio playback effect of the electronic device.

[0003] Currently, in order to improve the audio playback effect, professional personnel such as sound engineers can adjust the parameters for audio equalization processing, such as adjusting the gain parameter of the audio signal, or adding a filter in the above micro speaker and adjusting the parameters of the filter; thus obtaining a set of fixed equalization parameters to perform equalization adjustment on the audio signal through this set of fixed equalization parameters.

[0004] However, in the above method, since the equalization parameters used for equalization adjustment of the audio signal are a fixed set of parameters, the universality of this set of equalization parameters is poor, resulting in poor equalization adjustment effect of the electronic device on the audio signal. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a model training method, apparatus, electronic device, and storage medium, which can improve the effect of the electronic device performing audio equalization adjustment.

[0006] In a first aspect, the embodiments of this application provide a model training method, which includes: inputting the first audio frame parameters related to the first audio frame in the first audio signal into an audio equalization model for audio equalization processing, and outputting the gain of the first audio frame. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, and N is a positive integer; adjusting the spectral amplitude of the first audio frame with the gain of the first audio frame to obtain the adjusted first audio frame; determining the first reward value based on the adjusted first audio frame; updating the model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and the second audio frame parameters related to the second audio frame to obtain the updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal.

[0007] Second aspect, an embodiment of the present application provides a model inference method, which includes: when the third audio frame of the second audio signal is collected, obtaining third audio frame parameters related to the third audio frame, where the third audio frame is an audio frame that needs to be processed in real time, and the third audio frame parameters include: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer; inputting the third audio frame parameters into an audio equalization model for audio equalization processing, and outputting the gain of the third audio frame; adjusting the spectral amplitude of the third audio frame by using the gain of the third audio frame to obtain the adjusted third audio frame; and playing the adjusted third audio frame through a speaker.

[0008] Third aspect, an embodiment of the present application provides a model training device, which includes: a processing module, an adjustment module, a determination module, and an update module; the processing module is configured to input the first audio frame parameters related to the first audio frame in the first audio signal into an audio equalization model for audio equalization processing, and output the gain of the first audio frame, where the first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, and N is a positive integer; the adjustment module is configured to adjust the spectral amplitude of the first audio frame by using the gain of the first audio frame output by the processing module to obtain the adjusted first audio frame; the determination module is configured to determine a first reward value based on the first audio frame adjusted by the adjustment module; the update module is configured to update the model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value determined by the determination module, and the second audio frame parameters related to the second audio frame to obtain the updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal.

[0009] Fourth aspect, an embodiment of the present application provides a model inference device, which includes: an acquisition module, a processing module, an adjustment module, and a playback module; the acquisition module is configured to, when the third audio frame of the second audio signal is collected, obtain third audio frame parameters related to the third audio frame, where the third audio frame is an audio frame that needs to be processed in real time, and the third audio frame parameters include: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer; the processing module is configured to input the third audio frame parameters obtained by the acquisition module into an audio equalization model for audio equalization processing, and output the gain of the third audio frame; the adjustment module is configured to adjust the spectral amplitude of the third audio frame by using the gain of the third audio frame output by the processing module to obtain the adjusted third audio frame; the playback module is configured to play the adjusted third audio frame through a speaker.

[0010] Fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the method described in the first aspect.

[0011] Sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, it implements the steps of the method described in the first aspect.

[0012] Seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method described in the first aspect.

[0013] In the embodiment of the present application, the first audio frame parameters related to the first audio frame in the first audio signal are input into an audio equalization model for audio equalization processing, and the gain of the first audio frame is output. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, and N is a positive integer; the spectral amplitude of the first audio frame is adjusted by using the gain of the first audio frame to obtain the adjusted first audio frame; based on the adjusted first audio frame, a first reward value is determined; based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and the second audio frame parameters related to the second audio frame, the model parameters of the audio equalization model are updated to obtain an updated audio equalization model, and the second audio frame is the next audio frame of the first audio frame in the first audio signal. In this solution, since the electronic device inputs the parameters of the first audio frame and the historical frames of the first audio frame into the audio equalization model to obtain the gain of the first audio frame, the gain of the first audio frame is obtained based on the characteristics of the real-time input audio frame, that is, the gain can more accurately adjust the first audio frame, so as to improve the playback effect of the adjusted first audio frame. Then, since the electronic device can determine a reward value based on the adjusted first audio frame, and combine the parameters corresponding to the first audio frame and the next audio frame of the first audio frame to update the model of the audio equalization model, the adjusted audio equalization model can further meet the user's demand for equalization adjustment of the audio signal. In this way, the effect of the electronic device performing audio equalization adjustment can be improved. Description of the Drawings

[0014] Figure 1 is one of the flowcharts of a model training method provided by an embodiment of the present application;

[0015] Figure 2 is the second flowchart of a model training method provided by an embodiment of the present application;

[0016] Figure 3 It is the third flowchart of a model training method provided by an embodiment of the present application;

[0017] Figure 4 It is one of the schematic structural diagrams of an electronic device provided by an embodiment of the present application;

[0018] Figure 5 It is the fourth flowchart of a model training method provided by an embodiment of the present application;

[0019] Figure 6 It is the second schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0020] Figure 7 It is the first flowchart of a model inference method provided by an embodiment of the present application;

[0021] Figure 8 It is the third schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0022] Figure 9 It is the fifth flowchart of a model training method provided by an embodiment of the present application;

[0023] Figure 10 It is the second flowchart of a model inference method provided by an embodiment of the present application;

[0024] Figure 11 It is the schematic structural diagram of a model training device provided by an embodiment of the present application;

[0025] Figure 12 It is the schematic structural diagram of a model inference device provided by an embodiment of the present application;

[0026] Figure 13 It is the first schematic hardware structural diagram of an electronic device provided by an embodiment of the present application;

[0027] Figure 14 It is the second schematic hardware structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope protected by the present application.

[0029] The terms "first", "second", etc. in the specification of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0030] The terms "at least one (item)", "at least one of", etc. in the specification of this application refer to any one, any two or more combinations of the objects it contains. For example, at least one (item) of a, b, c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b and c", where a, b, c can be single or multiple. Similarly, "at least two (items)" means two or more, and its meaning is similar to that of "at least one (item)". The identifiers in this application are words, symbols, images, etc. used to indicate information, and can use controls or other containers as the carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.

[0031] The model training method provided by the embodiments of this application will be described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0032] The model training method provided by the embodiments of this application can be applied to the scenario of audio equalization processing of audio signals. For example, when the user needs to play a video, music, or received voice message through an electronic device, the electronic device can perform audio equalization processing on the audio frame 1 of the audio signal to be played based on the demand, so as to output the gain of the audio frame 1, and further adjust the spectral amplitude of the audio frame 1 through the gain of the audio frame 1. Then, the electronic device can update the model parameters in the audio equalization model based on the audio frame parameters, gain, and reward value corresponding to the audio 1 to obtain an updated audio equalization model.

[0033] It should be noted that the above examples only list some scenarios in which the embodiments of this application may be applied. In actual implementation, the embodiments of this application can also be applied to more scenarios that require playing audio signals through an electronic device and any other possible scenarios. The embodiments of this application are not limited here.

[0034] Based on the above scenarios to which the embodiments of the present application are applied, in the model training method provided by the embodiments of the present application, the first audio frame parameters related to the first audio frame in the first audio signal are input into an audio equalization model for audio equalization processing to output the gain of the first audio frame. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, where N is a positive integer; the spectral amplitude of the first audio frame is adjusted by using the gain of the first audio frame to obtain the adjusted first audio frame; a first reward value is determined based on the adjusted first audio frame; the model parameters of the audio equalization model are updated based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and the second audio frame parameters related to the second audio frame to obtain an updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal. In this solution, since the electronic device inputs the parameters of the first audio frame and the historical frames of the first audio frame into the audio equalization model to obtain the gain of the first audio frame, the gain of the first audio frame is obtained based on the characteristics of the audio frames input in real time, that is, the gain can more accurately adjust the first audio frame, so as to improve the playback effect of the adjusted first audio frame. Then, since the electronic device can determine a reward value based on the adjusted first audio frame and update the model of the audio equalization model by combining the parameters corresponding to the first audio frame and the next audio frame of the first audio frame, the adjusted audio equalization model can further meet the user's demand for equalization adjustment of the audio signal. In this way, the effect of the electronic device performing audio equalization adjustment can be improved.

[0035] The execution subject of the model training method provided by the embodiments of the present application is a model training device, which can be an electronic device, or a functional module or entity in the electronic device. The embodiments of the present application do not make any limitations in this regard. Hereinafter, the model training method provided by the embodiments of the present application will be exemplarily described by taking an electronic device as an example.

[0036] The embodiments of the present application provide a model training method, Figure 1 which shows a flowchart of a model training method provided by the embodiments of the present application. As Figure 1 shown, the model training method provided by the embodiments of the present application may include the following steps 201 to 204.

[0037] Step 201: The electronic device inputs the first audio frame parameters related to the first audio frame in the first audio signal into an audio equalization model for audio equalization processing, and outputs the gain of the first audio frame.

[0038] In some embodiments of the present application, the above-mentioned first audio signal may be an audio signal stored in the electronic device. For example, it may be an audio signal locally stored in the electronic device or an audio signal stored in the cloud of the electronic device. It may also be an audio signal collected by the electronic device in real time. For example, the electronic device can collect the audio signal through a microphone.

[0039] In some embodiments of the present application, after the electronic device collects an audio signal through a microphone or the like, the electronic device can process the audio signal to obtain the spectral amplitude of the audio frame involved later.

[0040] Exemplarily, the above-mentioned first audio signal may be any one of the audio signals included in the audio data training set locally stored in the electronic device.

[0041] In some embodiments of the present application, the type of the above-mentioned first audio signal may include but is not limited to any one of the following: human voice, music sound, animal sound, natural sound, etc.

[0042] Exemplarily, the above-mentioned human voice may be the sound emitted by a person. For example: the sound of a person speaking, the sound of a person singing;

[0043] The above-mentioned music sound may be the sound emitted by an instrument, and the instrument may include but is not limited to: piano, violin, zither, pipa, guitar;

[0044] The above-mentioned animal sound may be the sound emitted by an animal, and the animal may include but is not limited to: cat, dog, bird, frog;

[0045] The above-mentioned natural sound may be the sound generated in nature. For example: the sound of wind, the sound of rain, the sound of thunder.

[0046] In some embodiments of the present application, the above-mentioned first audio signal may be a time-domain signal.

[0047] In some embodiments of the present application, after the electronic device obtains the first audio signal, the electronic device can first perform frame division processing on the first audio signal to obtain at least one audio frame corresponding to the first audio signal; then, the electronic device can perform frequency-domain conversion on each audio frame in the at least one audio frame respectively and obtain the audio frame parameters corresponding to each audio frame.

[0048] In some embodiments of the present application, the electronic device can determine the frame length of each audio frame based on the sampling rate and the number of sampling points, so as to perform frame division processing. For example, the frame length of the audio frame = the number of sampling points per frame / the sampling rate.

[0049] It should be noted that the above-mentioned frame length represents the duration of each frame of the signal, usually expressed in milliseconds (ms).

[0050] Exemplarily, assume that the sampling rate of the above first audio signal is 48 kHz and the number of sampling points is 240. Then, the frame length of the above first audio frame can be 240÷48000 = 5 ms.

[0051] In some embodiments of the present application, the electronic device can also set a frame shift to achieve framing of the audio signal based on the determined frame length and frame shift.

[0052] It should be noted that the above frame shift is the overlapping part between two adjacent audio frames and is usually expressed in the number of sampling points.

[0053] It should be noted that the setting of the above frame length and frame shift can be determined according to actual needs, and the embodiments of the present application do not limit this here.

[0054] In some embodiments of the present application, the electronic device can perform a Fast Fourier Transformation (FFT) to convert the audio frame from a time-domain signal to a frequency-domain signal.

[0055] Exemplarily, the electronic device can perform a 512-point FFT on each of the at least one audio frame described above to obtain the frequency-domain representation corresponding to each audio frame, such as the amplitude of 512 frequency points.

[0056] It should be noted that when the electronic device performs an FFT, the specified number of FFT points can determine the spectral resolution and calculation efficiency. Therefore, the specific selection of the number of FFT points can be determined according to actual needs, and the embodiments of the present application do not limit this here.

[0057] In the embodiments of the present application, the above first audio frame parameters may include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N previous audio frames of the first audio frame, and the gain of the previous audio frame of the first audio frame.

[0058] Wherein, N is a positive integer.

[0059] In some embodiments of the present application, the above spectral amplitude can be understood as the amplitude of the frequency points corresponding to the audio frame, and the number of such frequency points is determined by the above number of FFT points. For example, the spectral amplitude of the above first audio frame can be represented by the amplitudes of 512 frequency points, that is, it can be understood as a 512-dimensional row vector.

[0060] In some embodiments of the present application, the above gain is the Equalizer (EQ) gain, which can represent the degree of amplification or attenuation of the signal at a certain center frequency of the equalizer and is usually expressed in decibels (dB).

[0061] In some embodiments of the present application, the gain of the above audio frame can be represented in the form of a vector, and the number of elements included in the vector can be determined by the number of frequency points of the audio frame, so that the electronic device can adjust the gain of the frequency points of the audio frame.

[0062] Exemplarily, when the number of frequency points of an audio frame is 512, the dimension of the gain vector of this audio frame can be 512; or, the dimension of the gain vector of this audio frame can be a value determined by appropriately compressing the dimension after dividing based on 512 frequency points.

[0063] In some embodiments of the present application, the electronic device can limit the range of the above gain to avoid the degradation of the playback effect of the audio signal due to excessive gain of the frequency points.

[0064] Exemplarily, the value range of the above gain can be [-12dB, 12dB].

[0065] In some embodiments of the present application, the above first audio frame parameter can be represented by s t =[F t ; G t-1 . Wherein, F t can represent the spectral amplitude of the above first audio frame, the information after splicing the spectral amplitudes of the first N audio frames of the first audio frame, and G t-1 can represent the gain of the previous audio frame of the above first audio frame.

[0066] It should be noted that when N is 3 and the number of FFT points is 512, this F t ∈R 4×512 , at this time F t can represent the spectral amplitudes of 4 audio frames, and each audio frame corresponds to 512 data; this G t-1 ∈R 1×512 , at this time G t-1 can represent that the gain of the audio frame contains the gains corresponding to 512 frequency points respectively; thus, this first audio frame parameter s t ∈R 5×512 , that is, it contains 5×512 data.

[0067] Exemplarily, assuming that the above first audio frame is the 10th audio frame in an audio signal and N is 3, the electronic device can obtain the spectral amplitude of the 10th audio frame, the spectral amplitudes of the 7th audio frame to the 9th audio frame, that is, F 10 ∈R 4×512 ; and the gain of the 9th audio frame, that is, G9, so as to obtain the audio frame parameter s 10 =[F 10 ; G9]∈R5×512 。

[0068] It should be noted that the value of N above can be determined according to actual needs, and the embodiments of the present application do not limit this here.

[0069] In some embodiments of the present application, when the first audio frame is any frame among the first N frames of the first audio signal, the electronic device can correspondingly fill in frames before the first frame of the first audio signal, so that the first audio frame becomes the (N + 1)-th audio frame.

[0070] Exemplarily, assuming that N above is 3, when the first audio frame is the first audio frame in the first audio signal, the electronic device can fill in three audio frames with zero data before the first audio frame, so that when the electronic device processes the first audio frame, it can successfully obtain the three audio frames before this audio frame for subsequent other operations.

[0071] In some embodiments of the present application, the above audio equalization model can be deployed in the online inference stage to meet the real-time requirement for obtaining audio frames. For example, the audio equalization model can be deployed on a Central Processing Unit (CPU).

[0072] In some embodiments of the present application, the above audio equalization model can be deployed in both the online inference stage and the offline training stage to meet the real-time requirement for obtaining audio frames and the offline training of the model. For example, the audio equalization model can be deployed in offline environments such as a CPU and a Graphics Processing Unit (GPU) cluster at the same time.

[0073] In some embodiments of the present application, the above audio equalization model can include at least one of the following: Multilayer Perceptron (MLP), Convolutional Neural Network (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM).

[0074] It should be noted that the above MLP is a supervised learning model composed of neurons at multiple levels, and the structure of the MLP can include an Input Layer, one or more artificially constructed Hidden Layers in the middle, and an Output Layer;

[0075] The structure of the above-mentioned CNN may include Convolutional Layers, Activation Functions, Pooling Layers, and Fully Connected Layers;

[0076] The above-mentioned RNN is a neural network architecture for processing sequential data. The structure of the RNN may include an input layer, a hidden layer, and an output layer. Its feature is that the connections between the hidden layers form a loop, enabling the network to retain information from previous time steps;

[0077] The above-mentioned LSTM is a type of recurrent neural network that addresses the vanishing gradient and exploding gradient problems in RNNs by introducing a gating mechanism. The structure of the LSTM may include an input gate, a forget gate, an output gate, and a memory cell (also known as the cell state).

[0078] In some embodiments of the present application, the above-mentioned audio equalization model includes an agent network. Combining Figure 1 , as Figure 2 shown, the above-mentioned step 201 can be specifically implemented through the following steps 201a to 201d.

[0079] Step 201a: The electronic device inputs the first audio frame parameters into the audio equalization model.

[0080] In some embodiments of the present application, the above-mentioned agent network can be a lightweight MLP network structure.

[0081] In some embodiments of the present application, the electronic device can input the above-mentioned first audio frame parameters into the input layer of the above-mentioned lightweight MLP network structure, and then input the first audio frame parameters into the fully connected hidden layer of the MLP network structure through the input layer to perform subsequent operations.

[0082] In some embodiments of the present application, the electronic device can first preprocess the first audio frame parameter s t and then input the preprocessed parameter into the above-mentioned equalization parameter.

[0083] Exemplarily, in the case of s t ∈R 5×512 , the electronic device can convert s t into a one-dimensional vector in a tiled manner, such as a vector containing 5 × 512 elements. Then, the electronic device can input the row vector or column vector containing 2560 elements into the MLP network structure.

[0084] Step 201b: The electronic device encodes the first audio frame parameters into a first feature vector through the fully-connected hidden layer of the agent network.

[0085] In some embodiments of the present application, the above-mentioned agent network may include one or more fully-connected hidden layers. Among them, each fully-connected hidden layer may use an activation function to encode the received parameters to obtain corresponding feature vectors. That is to say, by adjusting the number of neuron nodes and the number of layers of the fully-connected hidden layer of the agent network, the learning ability and expression ability of the agent network can be adjusted.

[0086] In some embodiments of the present application, the number of neuron nodes and the number of layers of the fully-connected hidden layer included in the above-mentioned agent network may be determined according to actual needs, and the embodiments of the present application do not limit this here.

[0087] In some embodiments of the present application, the above-mentioned activation function may be a Rectified Linear Unit (ReLU), LeakyReLU, or other feasible non-linear activation functions.

[0088] Exemplarily, in the case where the above-mentioned agent network includes multiple fully-connected hidden layers, the electronic device may first perform a non-linear transformation on the above-mentioned vector containing 2560 elements through the first fully-connected hidden layer to encode and obtain a higher-level feature representation; then, the audio equalization model may input the data output by the first fully-connected hidden layer of the agent network into the second fully-connected hidden layer of the agent network to perform similar steps. In this way, until the last fully-connected hidden layer in the audio equalization model outputs the above-mentioned first characteristic vector, such as a feature vector containing 512 elements.

[0089] Step 201c: The electronic device outputs the action probability parameters corresponding to the first feature vector through the fully-connected output layer of the agent network.

[0090] In some embodiments of the present application, the above-mentioned fully-connected output layer may include at least one output layer.

[0091] In the embodiments of the present application, the above-mentioned action probability parameters may include, but are not limited to, at least one of the following: the mean μ of the gain distribution probability of the first audio frame, the logarithmic standard deviation logσ of the gain distribution probability of the first audio frame.

[0092] In some embodiments of the present application, the above-mentioned agent network may be the Actor network in the SAC (fully called Soft Actor-Critic) algorithm.

[0093] In some embodiments of the present application, the above-mentioned Actor network may be responsible for generating an action policy according to the current state. That is, according to the above-mentioned first audio frame parameters, the gain of the subsequent corresponding first audio frame is output.

[0094] It should be noted that for a continuous action space, the Actor network usually outputs the mean μ and logarithmic standard deviation logσ of the action, so as to obtain the distribution of the action and its logarithmic probability.

[0095] Step 201d: The electronic device samples the action probability parameters corresponding to the first feature vector to obtain the gain of the first audio frame.

[0096] In some embodiments of the present application, the sampling method used in the above sampling process may be a Gaussian distribution sampling method. For example, the inverse transform method, rejection sampling, importance sampling, etc.

[0097] Exemplarily, the electronic device may sample any of the above Gaussian distribution sampling methods according to the mean μ and logarithmic standard deviation logσ, and obtain the gain a of the above first audio through the reparameterization technique. t . That is, a t =π φ (s t ; φ), where s t may represent the current state vector, that is, the above-mentioned first audio frame parameters; π φ may represent a neural network, that is, the above-mentioned lightweight MLP network structure; φ may represent the parameters of the neural network.

[0098] In the embodiments of the present application, the electronic device may extract the corresponding feature vector based on the received first audio frame and the parameters of the historical frame of the first audio frame. That is to say, the gain corresponding to the first audio frame obtained by the electronic device through the audio equalization model is essentially obtained based on the characteristics of the audio signal. In this way, the accuracy of the electronic device in obtaining the gain corresponding to the audio frame can be improved. Further, the electronic device may update the parameters in the audio equalization model based on the gain in the subsequent process, so as to improve the effect of the electronic device in performing audio equalization adjustment.

[0099] Step 202: The electronic device adjusts the spectral amplitude of the first audio frame by using the gain of the first audio frame to obtain the adjusted first audio frame.

[0100] In some embodiments of the present application, the electronic device may calculate the product of the spectral amplitude and the gain of the first audio frame to adjust the amplitude of at least one frequency point of the first audio frame, so as to achieve the equalization processing of the first audio frame.

[0101] Exemplarily, when the spectral amplitude of the first audio frame can be represented by the amplitudes of 512 frequency points, the gain of the first audio frame described above will also correspond to 512 elements. Thus, the electronic device can multiply the amplitude of each frequency point by the gain element corresponding to that frequency point to obtain the adjusted first audio frame described above.

[0102] Step 203: The electronic device determines a first reward value based on the adjusted first audio frame.

[0103] In some embodiments of the present application, the adjusted first audio frame described above can be a time-domain signal or a frequency-domain signal.

[0104] In some embodiments of the present application, the first reward value described above can be a value determined based on information in multiple dimensions. For example, at least one of the spectral amplitude of the first audio frame, the gain of the first audio frame, the spectral amplitudes of the N audio frames before the first audio frame, and the gain of the previous audio frame of the first audio frame.

[0105] In some embodiments of the present application, the first reward value described above can be used to update the model parameters of the audio equalization model.

[0106] In some embodiments of the present application, in combination Figure 1 , such as Figure 3 shown, step 203 can be specifically implemented through the following step 203a and step 203b.

[0107] Step 203a: The electronic device plays the adjusted first audio frame through the speaker of the electronic device and measures the output of the speaker to obtain a sound pressure signal and a distortion energy value.

[0108] In some embodiments of the present application, after the above step 202, the electronic device can first convert the adjusted first audio frame from a frequency-domain signal to a time-domain signal through an Inverse Fast Fourier Transform (IFFT) to achieve the reconstruction of the adjusted first audio frame. Thus, the electronic device can play the adjusted first audio frame through the speaker.

[0109] In some embodiments of the present application, the electronic device can play the adjusted first audio frame in a physical environment or a simulation environment.

[0110] It should be noted that the above physical environment can be an actual, touchable, and perceivable real-world environment; the above simulation environment can be a virtual environment created through electronic device simulation technology, aiming to imitate the physical characteristics and behaviors of the real world.

[0111] In some embodiments of the present application, the above sound pressure signal may be the pressure change generated when sound propagates in a medium such as air, and it is usually measured in Pascal (Pa).

[0112] It should be noted that the embodiments of the present application are only illustrated by taking a loudspeaker as an example. In actual applications, the above loudspeaker may be different types of loudspeakers, a multi-channel surround system, or any other device that can play audio.

[0113] In some embodiments of the present application, the electronic device may determine the frequency point sequence f2 increased due to playback distortion based on the masked frequency point sequence f0 in the adjusted first audio frame and the masked frequency point sequence f1 in the sound pressure signal measured by the above electronic device, so as to obtain the signal energy corresponding to the frequency point sequence f2, that is, the above distortion energy value.

[0114] It should be noted that the above f2 only exists in f0 and does not exist in f1, that is That is to say, the sound source corresponding to this f2 is a frequency signal that cannot be heard in the original sound source, but due to the frequency response and harmonic distortion of the loudspeaker, it is finally emitted and heard by the loudspeaker. It can be understood that the frequency point sequence f2 belongs to noise and has a negative impact on the listening experience.

[0115] In some embodiments of the present application, the electronic device obtaining the masked frequency point sequence f0 may include but is not limited to the following steps:

[0116] Time-frequency domain conversion, the electronic device can perform time-frequency conversion through FFT to obtain the amplitude A(f) of each frequency point;

[0117] Dividing critical bands;

[0118] Calculating the energy of each critical band;

[0119] Calculating the masking threshold T0(f);

[0120] Through the masking threshold, the masked frequency point sequence f0 in the adjusted first audio frame is obtained, which satisfies A(f0) < T0(f0).

[0121] It should be noted that after the above critical band division, the parameters shown in Table 1 can be obtained.

[0122] Table 1 Division method of critical bands based on the bark domain

[0123]

[0124] It should be noted that the steps of calculating the energy and masking threshold of each critical band based on the divided critical bands, and determining the sequence of masked frequency points f0 can all be executed through existing solutions, and are not elaborated in the embodiments of the present application.

[0125] In some embodiments of the present application, the method for the electronic device to obtain the sequence of masked frequency points f1 in the measured sound pressure signal is similar to the method for the electronic device to obtain the sequence of masked frequency points f0, and is not elaborated in the embodiments of the present application.

[0126] Step 203b: The electronic device determines a first reward value based on the sound pressure signal, the distortion energy value, the spectral difference value, and the gain change value.

[0127] In some embodiments of the present application, the electronic device can calculate the loudness gain r1 through the Zwicker calculation method of the Critical Band model.

[0128] It should be noted that the above critical band model can divide the auditory frequency into several intervals with equal perceptual bandwidths based on the auditory characteristics of the human ear, and each interval is called a critical band. This critical band model can reflect the perception of different frequencies by the human auditory system;

[0129] The concept of critical band comes from the masking of pure tones by noise; a pure tone can be masked by continuous noise centered on its central frequency and having a certain bandwidth. If the noise power within this bandwidth is equal to the power of the pure tone, the pure tone is then in the boundary state of just being audible, where this bandwidth is the critical bandwidth;

[0130] The above Zwicker calculation method is a loudness calculation method based on the critical band model, which is used to evaluate the human perception of sound loudness.

[0131] In some embodiments of the present application, the main steps of the above Zwicker calculation method may include but are not limited to at least one of the following:

[0132] Perform 1 / 3 octave spectrum analysis on the sound pressure signal to decompose the sound pressure signal into the energy distribution of different frequency components;

[0133] Calculate the Specific Loudness within each critical band;

[0134] Integrate the Specific Loudness over the 0 - 24 Bark domain to obtain the Total Loudness.

[0135] In some embodiments of the present application, the unit of the above loudness is sone, and the above characteristic loudness can reflect the distribution of loudness in the frequency domain, and its unit is sone / Bark. The characteristic loudness can be calculated from the principal loudness and the ramp loudness.

[0136] In some embodiments of the present application, the electronic device can determine an increased frequency point sequence f2 due to playback distortion, and calculate a distortion penalty r2 = -Σω2E by using the signal energy corresponding to the frequency point sequence f2, that is, the above distortion energy value. t =-∑ω2E f2 . Where ω2 can be a weight system corresponding to each frequency point; E t can represent the energy of the frequency points where obvious distortion is heard after playback through the speaker.

[0137] In the embodiments of the present application, the above spectral difference value is used to characterize the difference between the spectrum of the adjusted first audio frame and the spectrum of the preset audio frame.

[0138] In some embodiments of the present application, the electronic device can subtract the spectrum of the above adjusted first audio frame from the spectrum of the above preset audio frame and take the absolute value, and then perform weighted summation on each frequency point to obtain the above spectral difference value r3.

[0139] Exemplarily, the electronic device can calculate the above spectral difference value r3 based on the formula r3 = -Σω3|F t '-F target |. Where ω3 can be a weight system corresponding to each frequency point; F t ' can represent the spectral amplitude of the adjusted first audio frame; F target can represent the spectral amplitude of the preset audio frame.

[0140] In some embodiments of the present application, the above preset audio frame can be an audio frame whose preset spectral amplitude can meet the actual needs of the user.

[0141] In the embodiments of the present application, the above gain change value is used to characterize the change amplitude between the gain of the first audio frame and the gain of the previous audio frame of the first audio frame.

[0142] In some embodiments of the present application, the electronic device can subtract the gain of the above first audio frame from the gain of the previous audio frame of the above first audio frame and take the absolute value, and then perform weighted summation on each frequency point to obtain the above change amplitude r4.

[0143] Exemplarily, the electronic device can calculate based on the formula r4 = -Σω4|G t -G t-1|, calculate the above spectral difference value r4. Where ω4 can be the weight system corresponding to each frequency point; G t can represent the gain of the first audio frame; G t-1 can represent the gain of the previous audio frame of the first audio frame.

[0144] In some embodiments of the present application, the electronic device can perform weighted summation based on the above r1, r2, r3, and r4 to calculate the above first reward value r t . For example, r t =λ1r1 + λ2r2 + λ3r3 + λ4r4. Where r t ∈R, r i (i = 1, 2, 3, 4) can represent the reward value of a certain dimension that needs to be concerned, and λ i (i = 1, 2, 3, 4) can represent the weight coefficient corresponding to the reward value of each dimension.

[0145] It should be noted that the higher the A-weighting (A Weighting) loudness corresponding to the output loudness gain r1, the higher the reward value corresponding to this r1. Among them, the A-weighting is a standard weight curve for audio measurement, aiming to reflect the response characteristics of the human ear;

[0146] The reward value corresponding to the above r2 can be based on the distortion degree or noise amplitude of the speaker output measured by the microphone to punish overexcitation;

[0147] The smaller the weighted deviation of the frequency points between the spectrum of the adjusted first audio frame and the spectrum of the preset audio frame, the higher the reward value corresponding to this r3.

[0148] It should be noted that the smaller the change between the gain of the first audio frame and the gain of the previous audio frame, the higher the reward corresponding to this r4, so as to reduce the abnormal human ear listening experience caused by overly frequent or large-amplitude coefficient changes.

[0149] In some embodiments of the present application, the electronic device can introduce more dimensions of reward values, so that the subsequent trained audio equalization model can combine more dimensions of information and provide more accurate gain for the audio frame. For example, avoiding sibilance, reducing nonlinear distortion, reducing power consumption, etc.

[0150] In the embodiments of the present application, the electronic device can calculate the first reward value based on the set reward rules of multiple dimensions, so as to facilitate the subsequent update of the model parameters of the audio equalization model, so that the subsequent audio equalization model can combine the characteristics of real-time audio frames, ensure a high loudness, while reducing adverse effects such as spectral coloring and distortion, and can obtain the best sound quality experience under the condition of adjusting the equalization coefficient as smoothly as possible.

[0151] Further, since the sound pressure signal of the audio signal played by the speaker detected by the microphone is used in the process of calculating the above first reward value, that is to say, the calculation of the above first reward value takes into account the playback characteristics of the speaker of the electronic device. Therefore, even for the same audio frame, different first reward values can be obtained when the types of speakers of the electronic device are different, and different audio equalization models can be trained, thereby improving the flexibility of audio equalization processing.

[0152] In some embodiments of the present application, after the electronic device outputs the gain a of the first audio frame for the first audio frame based on the audio equalization model t the environment can obtain the corresponding reward value r based on the returned gain a t and start to obtain the audio frame parameters s corresponding to the audio frame after the first audio frame t . It can be understood that the audio equalization model and the environment have an interaction, which can be recorded as sample data (s t+1 , a t , r t , s t ). Then, the electronic device can store the sample data corresponding to the first audio frame in the experience replay buffer module, so that the policy update module in the electronic device can update the model parameters of the audio equalization model in real time or offline based on the sample data in the experience replay buffer module. Among them, the policy update module can be used to update the audio equalization model. That is, as shown in t+1 the flowchart shown in Figure 4 .

[0153] It should be noted that the experience replay buffer module can be sorted in a queue to sequentially put historical data, so as to realize the efficient use of historical interaction data. The number of historical interaction data that can be stored in the experience replay buffer module is not limited in this embodiment of the present application. For example, the experience replay buffer module can record 1M historical interaction data.

[0154] Step 204: The electronic device updates the model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and the second audio frame parameters related to the second audio frame, and obtains an updated audio equalization model.

[0155] In the embodiments of the present application, the second audio frame may be the audio frame next to the first audio frame in the first audio signal.

[0156] In some embodiments of the present application, after the electronic device obtains the gain of the first audio frame through the audio equalization model, the electronic device can obtain the spectral amplitude of the second audio frame to obtain the above second audio frame parameters.

[0157] In some embodiments of the present application, the above-mentioned second audio frame parameters include: the spectral amplitude of the second audio frame, the spectral amplitudes of the first N audio frames before the second audio frame, and the gain of the previous audio frame of the second audio frame.

[0158] It should be noted that the spectral amplitudes of the first N audio frames before the second audio frame can be understood as: the spectral amplitudes of the first N - 1 audio frames before the first audio frame and the spectral amplitude of the first audio frame; the gain of the previous audio frame of the second audio frame can be understood as the gain of the first audio frame.

[0159] In some embodiments of the present application, the above-mentioned audio equalization model can adopt Policy Gradient, Q-learning, or an Actor-Critic-based method to adjust the parameters in the audio equalization model. The embodiments of the present application do not limit this.

[0160] Among them, the above-mentioned Actor-Critic-based method can include but is not limited to any one of the following: Proximal Policy Optimization (PPO), SAC, etc.

[0161] In some embodiments of the present application, the above-mentioned audio equalization model includes: an agent network, at least two first policy update networks, and at least two second policy update networks. One second policy update network corresponds to one first policy update network. Combining Figure 1 , as Figure 5 shown, the above-mentioned step 204 can be specifically implemented by the following steps 204a to 204e.

[0162] Step 204a: The electronic device outputs the gain of the second audio frame through the agent network based on the second audio frame parameters.

[0163] In some embodiments of the present application, the above-mentioned audio equalization model can be a model constructed based on the SAC algorithm.

[0164] It should be noted that the networks maintained by the above-mentioned SAC algorithm can include but are not limited to: an Actor network, a Critic network, and a target Critic network. Among them, the Critic network and the target network can be deployed in an offline environment such as a Graphics Processing Unit (GPU) cluster.

[0165] Exemplarily, the audio equalization model of the embodiments of the present application may include: an agent network, that is, an Actor network maintained by an SAC algorithm; two first policy update networks, that is, two Critic networks maintained by an SAC algorithm; two second measurement update networks, that is, two target Critic networks maintained by an SAC algorithm. Among them, each Critic network corresponds to a target Critic network, and the network structure of the target Critic network is the same as that of the Critic network, that is, a second policy update network corresponds to a first policy update network.

[0166] In some embodiments of the present application, the above-mentioned Critic network may include two independent Q networks, such as Q1 and Q2. The Q network can be used to estimate the cumulative return under a given state-action pair, and its parameters can be denoted as θ1 and θ2 respectively.

[0167] In some embodiments of the present application, the above-mentioned Critic network may output value parameters based on the audio frame parameters and the gain of the audio frame. That is, the second value parameter corresponding to the first policy update network in the subsequent steps, such as value and

[0168] It should be noted that using the two Q networks above can reduce the bias of overestimation of the Q value in subsequent calculations. The Q value can be called the action value function, which is used to represent the expected return that can be obtained by taking a specific action under a given state.

[0169] In some embodiments of the present application, the above-mentioned target Critic network may output action value parameters based on the audio frame parameters and the gain of the audio frame. That is, the first value parameter corresponding to the second policy update network in the subsequent steps, for example, action value and action value Its parameters can be denoted as and

[0170] In some embodiments of the present application, the electronic device may input the second audio frame parameters into the audio equalization model, so that the audio equalization model can encode the second audio frame parameters into a feature vector through the fully connected hidden layer of the agent network; and output the action probability parameters corresponding to the feature vector of the second audio frame parameters through the fully connected output layer of the agent network; and further, the audio equalization model may perform sampling processing on the action probability parameters to obtain the gain of the second audio frame.

[0171] It should be noted that the above-mentioned agent network outputs the gain of the second audio frame based on the second audio frame parameters, which is similar to the solution in step 201 above where the audio equalization model outputs the gain of the first audio frame based on the first audio frame parameters. For specific details, please refer to the relevant description in step 201 above, and the embodiments of the present application will not repeat it here.

[0172] Step 204b: The electronic device updates the network through the second policy, and based on the first reward value, the first value parameter, and the action probability parameter of the second audio frame, outputs a target value.

[0173] In the embodiments of the present application, the above-mentioned first value parameter is a parameter obtained by the second policy update network based on the second audio frame parameters and the gain of the second audio frame, that is, a parameter obtained by the target critic network based on the second audio frame parameters and the gain of the second audio frame

[0174] In some embodiments of the present application, during the offline training phase, the target Critic network can be based on the above-mentioned first reward value r t 、the above-mentioned first value parameter and the action probability parameter π of the above-mentioned second audio frame φ (a t+1 |s t+1 ), and through the following formula (1), calculate a target value y t to supervise and update the Critic network so that the Q value it outputs approaches the true future return expectation.

[0175]

[0176] Among them, γ is the discount factor, which can be used to balance the agent's emphasis on the current reward r t and future rewards. A higher γ can encourage long-term exploration, while a lower γ may lead to the policy converging to the local optimum faster; α is the temperature parameter, which can be used to balance the weights of the reward term and the entropy term. The larger the value, the more it encourages retaining randomness;

[0177] The first reward term can represent the target network's estimate of the future return, and calculates the long-term benefit that can be obtained by the equilibrium strategy when taking action a t+1 at the next moment s t+1 , that is, the above-mentioned first value parameter; at the same time, using and the minimum value of can reduce the bias of overestimation of the Q value;

[0178] The second entropy term -αlogπ φ (a t+1 |s t+1) can represent the reward for random exploration of the current policy, so that the Q-value evaluation takes into account the incentive for exploration; π φ (a t+1 |s t+1 ) represents the probability distribution of the action a output by the current Actor network in the state s t+1 , which can be transformed into the entropy of the current policy through logarithmic operation. t+1

[0179] Step 204c: The electronic device updates the network parameters of the first policy update network based on the target value and the second value parameter.

[0180] In the embodiments of the present application, the above-mentioned second value parameter is a parameter obtained by the first policy update network based on the first audio frame parameter and the gain of the first audio frame.

[0181] In some embodiments of the present application, after calculating the above-mentioned target value y t , the electronic device can define a loss function for each Critic network, that is Then, the gradient descent method can be used to update the parameters θ1 and θ2 in the Critic network, such as

[0182] In some embodiments of the present application, the electronic device can update θ1 and θ2 by minimizing this mean square error.

[0183] Step 204d: The electronic device outputs the updated second value parameter through the updated first policy update network.

[0184] In some embodiments of the present application, after the electronic device updates the parameters θ1 and θ2 in the Critic network, the Critic network can calculate the updated and based on the updated parameters θ1 and θ2, that is, the above-mentioned updated second value parameter.

[0185] Step 204e: The electronic device updates the network parameters of the agent network of the audio equalization model based on the updated second value parameter to obtain an updated audio equalization model.

[0186] In some embodiments of the present application, the electronic device can update the policy of the Actor network based on the updated Critic network through the following formula (2). Then, the electronic device can update the parameter φ using gradient descent, such as φ←φ - η π ▽ φ J π (φ), to obtain an updated Actor network, that is, to obtain an updated audio equalization model.

[0187]

[0188] where π φ (a t |s t ) represents the probability distribution of the action a output by the current Actor network in the state s t and can be transformed into the entropy α log π of the current policy through logarithmic operation t (a φ |s t ). t )

[0189] The first reward term calculates the benefit that can be obtained by the equilibrium strategy when taking the action a at the moment s t , that is, the second value parameter after the above update; at the same time, using t the minimum value of and can reduce the bias of overestimation of the Q value

[0190] In some embodiments of the present application, the above steps of updating the audio equalization model can be continuously performed in an offline environment until convergence by cyclic iteration, or the cyclic iteration reaches a preset number of rounds

[0191] In some embodiments of the present application, the above update of the audio equalization model can be performed by the electronic device in an offline environment such as a GPU cluster, and then the electronic device can synchronize the updated parameter φ to the CPU side deployed in the online environment so that the online inference can also use a better strategy

[0192] In some embodiments of the present application, the above synchronization time can be determined based on the time preset by the electronic device or the user or the number of update iterations

[0193] Exemplarily, in combination with Figure 4 , as Figure 6 shown, the target Critic network can obtain sample data (s t , a t , r t , s t+1 ) from the experience replay buffer module, and obtain the action a t+1 corresponding to the next audio frame to calculate the target value y t . Then, the Critic network can be based on the target value y tPerform gradient descent update to obtain the updated Critic network. The updated Critic network can calculate the Q value and transmit it to the Actor network for gradient descent update, and the update of the equilibrium model is implemented in the offline environment GPU. Further, the updated Actor network in the offline environment can synchronize parameters to the Actor network in the online environment CPU, so as to implement the update of the Actor network in the online environment.

[0194] In the embodiment of the present application, the electronic device can determine the first reward value based on the adjusted first audio frame, so as to update the audio equilibrium model by combining the first audio frame parameters and the gain of the first audio frame. Therefore, the adjusted audio equilibrium model can further meet the user's demand for audio signal equalization adjustment, and thus improve the effect of the electronic device performing audio equalization adjustment.

[0195] In the model training method provided in the embodiment of the present application, since the electronic device inputs the parameters of the first audio frame and the historical frame of the first audio frame into the audio equilibrium model to obtain the gain of the first audio frame, the gain of the first audio frame is obtained based on the characteristics of the audio frame input in real time, that is, the gain can more accurately adjust the first audio frame, so as to improve the playback effect of the adjusted first audio frame. Then, since the electronic device can determine the first reward value based on the adjusted first audio frame, and update the audio equilibrium model by combining the parameters corresponding to the first audio frame and the next audio frame of the first audio frame, the adjusted audio equilibrium model can further meet the user's demand for audio signal equalization adjustment, and thus improve the effect of the electronic device performing audio equalization adjustment.

[0196] In some embodiments of the present application, after the above step 204c, the model training method provided in the embodiment of the present application may further include the following step 301.

[0197] Step 301: The electronic device updates the network parameters of the second policy update network based on the updated network parameters of the first policy update network through the second policy update network.

[0198] In some embodiments of the present application, after the electronic device updates the above-mentioned Critic network, the parameters and in the target Critic network can be softly updated, such as where τ is a small positive number, which can ensure that the target Critic network smoothly follows the Critic network.

[0199] In this way, the target Critic network can ensure the smoothness of future value prediction through soft update, and can cooperate with the Critic network to make the training of the SAC algorithm more stable and efficient.

[0200] In some embodiments of the present application, the model training method provided by the embodiments of the present application may further include the following steps 401 to 404.

[0201] Step 401, the electronic device inputs the first audio frame parameters and the gain of the first audio frame into the first policy update network.

[0202] It can be understood that the electronic device may input the first audio frame parameters s t and the gain a of the first audio frame t into the Critic network.

[0203] Step 402, the electronic device extracts the spectral amplitude of the first audio frame and the local features of the spectral amplitudes of the first N audio frames of the first audio frame through the convolutional layer of the first policy update network, and encodes the local features into a second feature vector.

[0204] In some embodiments of the present application, the electronic device may input the spectral amplitude of the first audio frame and the spectral amplitudes of the first N audio frames of the first audio frame included in the first audio frame parameters, that is, F t , into the Critic network of the above SAC algorithm to obtain the above local features.

[0205] In some embodiments of the present application, the above Critic network may use a combined network structure of CNN+MLP.

[0206] It should be noted that in the CNN+MLP combined network structure, usually the CNN network structure is used as a feature extractor to extract the feature representation of the input data, and then these features are input into the MLP network structure for further processing and classification, so as to make full use of the feature extraction ability of the CNN network structure and the classification ability of the MLP network structure through the combined network structure to improve the performance of the model.

[0207] In some embodiments of the present application, the electronic device may regard F t as 2D data to extract the above local features through the convolutional layer of the CNN network structure, such as a two-dimensional convolutional kernel.

[0208] Exemplarily, assuming F t ∈R 4×512 , the electronic device may regard this F t as a (4,512) matrix to extract local features through a two-dimensional convolutional kernel and encode them into a 1024-dimensional feature vector, that is, the above second feature vector.

[0209] Step 403: The electronic device updates the network through the first policy, splices and fuses the second feature vector, the gain of the previous audio frame of the first audio frame, and the gain of the first audio frame to obtain a third feature vector.

[0210] In some embodiments of the present application, the above splicing and fusion process may be the front and back splicing of values.

[0211] Exemplarily, when the above second feature vector is a vector of 1024, the gain G of the previous audio frame of the first audio frame t-1 is a vector of 512, and the gain a of the first audio frame t is a vector of 512, the electronic device can obtain a third feature vector V through the splicing and fusion process, and the dimension of the third feature vector V can be 1024 + 512 + 512 = 2048.

[0212] Step 404: The electronic device updates the fully connected layer of the network through the first policy, and maps the third feature vector to a second value parameter.

[0213] In the embodiments of the present application, the above second value parameter is used to characterize the value of equalizing the first audio frame by using the gain of the first audio frame.

[0214] In some embodiments of the present application, the above Critic network may include two independent Q networks, such as Q1 and Q2. The Q network can be used to estimate the cumulative return under a given state-action pair, and its parameters can be denoted as θ1 and θ2 respectively.

[0215] In some embodiments of the present application, the above Critic network may map and output a value parameter based on the above third feature vector such as value and

[0216] It should be noted that using the above two Q networks can reduce the bias of overestimating the Q value in subsequent calculations. The Q value can be called the action value function, which is used to represent the expected return that can be obtained by taking a specific action under a given state.

[0217] It should be noted that the scheme in which the second policy update module obtains the first value parameter, that is, the target Critic network outputs the action value parameter is similar to the above steps 401 to 404, and the embodiments of the present application will not elaborate here.

[0218] The embodiments of the present application provide a model training method, Figure 7 showing a flowchart of a model inference method provided by the embodiments of the present application. AsFigure 7 As shown, the model inference method provided by the embodiments of the present application may include the following steps 501 to 504.

[0219] Step 501: When a third audio frame of a second audio signal is collected, obtain third audio frame parameters related to the third audio frame.

[0220] In the embodiments of the present application, the above-mentioned third audio frame may be an audio frame that needs to be processed in real time.

[0221] In the embodiments of the present application, the above-mentioned third audio frame parameters may include: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, and the gain of the previous audio frame of the third audio frame.

[0222] Wherein, N is a positive integer.

[0223] Step 502: The electronic device inputs the third audio frame parameters into the audio equalization model for audio equalization processing, and outputs the gain of the third audio frame.

[0224] Step 503: The electronic device adjusts the spectral amplitude of the third audio frame by using the gain of the third audio frame to obtain an adjusted third audio frame.

[0225] Step 504: The electronic device plays the adjusted third audio frame through a speaker.

[0226] In some embodiments of the present application, the steps of the above steps 501 to 504 are similar to the above steps 201 and 202, and "the electronic device plays the adjusted first audio frame" in step 203a. The embodiments of the present application will not be elaborated herein.

[0227] Exemplarily, as Figure 8 shown, the electronic device may input the second audio frame parameter s of the second audio frame collected in the physical environment or the simulation environment t into the Actor network in the CPU, that is, the above-mentioned updated audio equalization model, so as to output the gain a corresponding to the second audio frame t , and adjust the spectral amplitude of the second audio frame, and then further play it through a speaker.

[0228] In the model inference method provided in the embodiments of the present application, since the electronic device can obtain audio frame parameters related to the collected audio frame based on the collected audio frame and input them into the audio equalization model to obtain the gain corresponding to the audio frame, the gain of the audio frame is output in real time based on the characteristics of the audio frame. That is, the gain can more accurately adjust the audio frame in real time. Thus, it can improve the convenience of the electronic device to perform equalization adjustment on the audio, and at the same time improve the effect of the electronic device to perform audio equalization adjustment.

[0229] In some embodiments of the present application, the above audio equalization model includes an experience replay buffer module. The model inference method provided in the embodiments of the present application may further include the following steps 601 to 603.

[0230] Step 601: The electronic device determines a second reward value based on the adjusted third audio frame.

[0231] It should be noted that the scheme for the electronic device to determine the second reward value is similar to the scheme for the electronic device to determine the first reward value in the above step 203 and its related steps, and the embodiments of the present application will not elaborate here.

[0232] Step 602: The electronic device obtains fourth audio frame parameters corresponding to the fourth audio frame.

[0233] In the embodiments of the present application, the above fourth audio frame is the next audio frame of the third audio frame in the second audio signal.

[0234] It should be noted that the scheme for the electronic device to obtain the fourth audio frame parameters corresponding to the fourth audio frame can refer to the scheme for the electronic device to obtain the first audio frame parameters or the second audio frame parameters in the above embodiments, and the embodiments of the present application will not elaborate here.

[0235] Step 603: The electronic device caches the third audio frame parameters, the gain of the third audio frame, the second reward value, and the fourth audio frame parameters through the experience replay buffer module.

[0236] In some embodiments of the present application, after the electronic device outputs the gain of the third audio frame for the third audio frame based on the audio equalization model, the environment can obtain the reward value corresponding to the returned gain and start to obtain the audio frame parameters corresponding to the next audio frame of the third audio frame, that is, the fourth audio frame. It can be understood that the audio equalization model and the environment have an interaction, which can be recorded as a set of sample data. Then, the electronic device can store the sample data corresponding to the third audio frame in the experience replay buffer module. Further, the policy update module of the electronic device can subsequently update the model parameters of the audio equalization model in real time or offline based on the sample data in the experience replay buffer module. Among them, the policy update module can be used to update the audio equalization model.

[0237] It should be noted that the experience replay buffer module can be sorted by queue to sequentially put historical data, so as to achieve efficient utilization of historical interaction data. The number of historical interaction data that can be stored in the experience replay buffer module is not limited in this embodiment of the present application. For example, the experience replay buffer module can record 1M pieces of historical interaction data.

[0238] For each scenario applicable to the embodiments of the present application, combined with the above implementation solutions of the embodiments of the present application, specific examples are given below to illustrate the implementation process in each scenario of the embodiments of the present application.

[0239] Example 1, offline update phase, as Figure 9 shown.

[0240] When the electronic device obtains the audio frame of the audio signal based on the experience replay buffer module, the electronic device can first convert the obtained audio frame into a frequency domain form for representation through FFT. Then, the electronic device can input the audio frame into the audio equalization model, that is, the Actor network of the SAC algorithm, to obtain the gain corresponding to the audio frame. Thus, the electronic device can adjust the spectral amplitude of the audio frame based on the gain output by the Actor network, and convert the adjusted audio frame from the frequency domain form to the time domain form through IFFT. Thus, the electronic device can play the adjusted audio frame in the physical environment or the simulation environment through the speaker system, and can measure the sound pressure signal in the environment through the microphone to calculate the reward feedback based on the measured information. Then, the electronic device can calculate the target value y t through the target Critic network of the SAC algorithm, and perform gradient descent update through the Critic network of the SAC algorithm based on the target value y t to obtain the updated Critic network. The updated Critic network can calculate the Q value and transmit it to the Actor network for gradient descent update to implement the update of the audio equalization model in the offline environment GPU. Further, the updated Actor network in the offline environment can synchronize parameters to the Actor network in the online environment CPU, so as to implement the update of the Actor network in the online environment. In addition, the updated Critic network can also perform soft update on the parameters of the target Critic network.

[0241] Example 2, online inference phase, as Figure 10 shown.

[0242] When the electronic device obtains the audio frame of the audio signal, the electronic device can first convert the obtained audio frame into a frequency-domain form for representation based on FFT. Then, the electronic device can input the audio frame into the audio equalization model, that is, the Actor network of the SAC algorithm, to obtain the gain corresponding to the audio frame. Thus, the electronic device can adjust the spectral amplitude of the audio frame based on the gain output by the Actor network, and convert the adjusted audio frame from the frequency-domain form to the time-domain form through IFFT. Thus, the electronic device can play the adjusted audio frame in a physical environment or a simulation environment through a speaker system. Then, the electronic device can record the played audio frame to calculate the reward value corresponding to the audio frame, and obtain the next audio frame of the audio frame and the audio frame parameters related to the next audio frame. Thus, the electronic device can combine the audio frame parameters, gain, reward value of the audio frame and the audio frame parameters of the next audio frame into a set of sample data and store it in the experience replay cache module.

[0243] It should be noted that the above-mentioned various method embodiments, or various possible implementation manners in each method embodiment, can be executed independently, or, on the premise of no contradiction, can also be executed in combination with each other, which can be specifically determined according to actual usage requirements, and the embodiments of the present application do not limit this.

[0244] It should be noted that for the model training method provided by the embodiments of the present application, the execution subject can be a model training device. In the embodiments of the present application, taking the model training device executing the model training method as an example, the model training device provided by the embodiments of the present application is described.

[0245] Figure 11 A possible structural schematic diagram of the model training device involved in the embodiments of the present application is shown. As Figure 11 shown, the model training device 70 may include: a processing module 71, an adjustment module 72, a determination module 73, and an update module 74;

[0246] Among them, the processing module 71 is configured to input the first audio frame parameters related to the first audio frame of the first audio signal into the audio equalization model for audio equalization processing, and output the gain of the first audio frame. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, and N is a positive integer;

[0247] The adjustment module 72 is configured to adjust the spectral amplitude of the first audio frame by using the gain of the first audio frame to obtain the adjusted first audio frame;

[0248] The determination module 73 is configured to determine the first reward value based on the first audio frame adjusted by the adjustment module 72;

[0249] An update module 74, configured to update model parameters of an audio equalization model based on first audio frame parameters, a gain of a first audio frame, and a first reward value determined by a determination module 73, so as to obtain an updated audio equalization model.

[0250] In the model training apparatus provided in the embodiments of the present application, since the model training apparatus inputs parameters of a first audio frame and a historical frame of the first audio frame into an audio equalization model to obtain a gain of the first audio frame, the gain of the first audio frame is obtained based on features of a real-time input audio frame, that is, the gain can more accurately adjust the first audio frame, so as to improve the playing effect of the adjusted first audio frame. Then, since the model training apparatus can determine a reward value based on the adjusted first audio frame, and update the audio equalization model by combining parameters corresponding to the first audio frame and the next audio frame of the first audio frame, the adjusted audio equalization model can further meet the user's demand for equalization adjustment of an audio signal. In this way, the effect of the model training apparatus performing audio equalization adjustment can be improved.

[0251] In a possible implementation manner, the above audio equalization model includes an agent network; the processing module 71 includes: an input sub-module, an encoding sub-module, an output sub-module, and a sampling sub-module; the input sub-module is specifically configured to input first audio frame parameters into the audio equalization model; the encoding sub-module is specifically configured to encode the first audio frame parameters into a first feature vector through a fully connected hidden layer of the agent network; the output sub-module is specifically configured to output action probability parameters corresponding to the first feature vector through a fully connected output layer of the agent network; the sampling sub-module is specifically configured to perform sampling processing on the action probability parameters corresponding to the first feature vector to obtain a gain of the first audio frame.

[0252] In a possible implementation manner, the above determination module 73 includes: a playing sub-module, a measuring sub-module, and a determining sub-module; the playing sub-module is specifically configured to play the adjusted first audio frame through a speaker of an electronic device; the measuring sub-module is specifically configured to measure an output of the speaker to obtain a sound pressure signal and a distortion energy value; the determining sub-module is specifically configured to determine a first reward value based on the sound pressure signal, the distortion energy value, a spectrum difference value, and a gain change value, where the first spectrum difference value is used to characterize a difference between a spectrum of the adjusted first audio frame and a spectrum of a preset audio frame, and the first gain change value is used to characterize a change amplitude between a gain of the first audio frame and a gain of the previous audio frame of the first audio frame.

[0253] In a possible implementation, the above audio equalization model includes: an agent network, at least two first policy update networks, and at least two second policy update networks, where one second policy update network corresponds to one first policy update network; the processing module 71 is specifically configured to output the gain of the second audio frame based on the second audio frame parameter through the agent network; the update module 74 includes: an output sub-module and an update sub-module; the output sub-module is specifically configured to output a target value through the second policy update network based on the first reward value, the first value parameter, and the action probability parameter of the second audio frame, where the first value parameter is a parameter obtained by the second policy update network based on the second audio frame parameter and the gain of the second audio frame; the update sub-module is specifically configured to update the network parameters of the first policy update network based on the target value and the second value parameter, where the second value parameter is a parameter obtained by the first policy update network based on the first audio frame parameter and the gain of the first audio frame; the output sub-module is further configured to output the updated second value parameter through the updated first policy update network; the update sub-module is further configured to update the network parameters of the agent network of the audio equalization model based on the updated second value parameter to obtain an updated audio equalization model.

[0254] In a possible implementation, the above update module 74 is further configured to, after updating the network parameters of the first policy update network based on the target value, update the network parameters of the second policy update network through the second policy update network based on the updated network parameters of the first policy update network.

[0255] In a possible implementation, the model training device 70 provided in the embodiments of the present application may further include: an input module; the input module is configured to input the first audio frame parameter and the gain of the first audio frame into the first policy update network; the processing module 71 is further configured to extract the spectral amplitude of the first audio frame and the local features of the spectral amplitudes of the first N audio frames before the first audio frame through the convolutional layer of the first policy update network, and encode the local features into a second feature vector; and, through the first policy update network, splice and fuse the second feature vector, the gain of the audio frame before the first audio frame, and the gain of the first audio frame to obtain a third feature vector; and, map the third feature vector to a second value parameter through the fully connected layer of the first policy update network, where the second value parameter is used to characterize the value of equalizing the first audio frame by using the gain of the first audio frame.

[0256] Figure 12 FIG. shows a possible structural schematic diagram of the model inference device involved in the embodiments of the present application. As Figure 12 shown, the model inference device 80 may include: an acquisition module 81, a processing module 82, an adjustment module 83, and a playback module 84;

[0257] Among them, an acquisition module 81 is configured to obtain third audio frame parameters related to a third audio frame when the third audio frame of a second audio signal is acquired. The third audio frame is an audio frame that needs to be processed in real time. The third audio frame parameters include: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer;

[0258] A processing module 82 is configured to input the third audio frame parameters obtained by the acquisition module 81 into an audio equalization model for audio equalization processing, and output the gain of the third audio frame;

[0259] An adjustment module 83 is configured to adjust the spectral amplitude of the third audio frame by using the gain of the third audio frame output by the processing module 82 to obtain an adjusted third audio frame;

[0260] A playback module 84 is configured to play the adjusted third audio frame through a speaker.

[0261] In the model inference device provided in the embodiment of the present application, since the model inference device can, based on the acquired audio frame, obtain audio frame parameters related to the audio frame and input them into the audio equalization model to obtain the gain corresponding to the audio frame, the gain of the audio frame is output in real time based on the characteristics of the audio frame. That is, the gain can more accurately adjust the audio frame in real time. In this way, the convenience of the model inference device for performing equalization adjustment on audio can be improved, and at the same time, the effect of the model inference device for performing audio equalization adjustment can be improved.

[0262] In a possible implementation manner, the above audio equalization model includes an experience replay buffer module; the model inference device 80 provided in the embodiment of the present application further includes: a determination module and a buffer module;

[0263] The determination module is configured to determine a second reward value based on the adjusted third audio frame;

[0264] The acquisition module 81 is further configured to obtain fourth audio frame parameters corresponding to a fourth audio frame, and the fourth audio frame is the next audio frame of the third audio frame in the second audio signal;

[0265] The buffer module is configured to cache the third audio frame parameters, the gain of the third audio frame, the second reward value, and the fourth audio frame parameters through the experience replay buffer module.

[0266] The model training device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0267] The model training device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0268] The model training device provided in the embodiments of the present application can implement each process implemented in the above method embodiments. To avoid repetition, it will not be elaborated here.

[0269] Optionally, as Figure 13 shown, the embodiments of the present application further provide an electronic device 90, including a processor 91 and a memory 92. A program or instruction that can run on the processor 91 is stored on the memory 92. When the program or instruction is executed by the processor 91, each step of the above model training method embodiments is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0270] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0271] Figure 14 Schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0272] The electronic device 100 includes, but is not limited to, components such as a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110.

[0273] Those skilled in the art can understand that the electronic device 100 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 110 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 14 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than those shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0274] Among them, taking the electronic device in the model training stage of the electronic device 100 as an example:

[0275] The processor 110 is configured to input first audio frame parameters related to a first audio frame in a first audio signal into an audio equalization model for audio equalization processing, and output a gain of the first audio frame. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N previous audio frames of the first audio frame, the gain of the previous audio frame of the first audio frame, where N is a positive integer; and adjust the spectral amplitude of the first audio frame by using the gain of the first audio frame to obtain an adjusted first audio frame; and, based on the adjusted first audio frame, determine a first reward value; and, based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and second audio frame parameters related to a second audio frame, update the model parameters of the audio equalization model to obtain an updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal.

[0276] In the electronic device provided in the embodiments of the present application, since the electronic device inputs parameters of a first audio frame and historical frames of the first audio frame into an audio equalization model to obtain the gain of the first audio frame, the gain of the first audio frame is obtained based on the characteristics of the audio frame input in real time, that is, the gain can more accurately adjust the first audio frame, so as to improve the playback effect of the adjusted first audio frame. Then, since the electronic device can determine a reward value based on the adjusted first audio frame, and update the model of the audio equalization model by combining the parameters corresponding to the first audio frame and the next audio frame of the first audio frame, the adjusted audio equalization model can further meet the user's demand for equalization adjustment of the audio signal. In this way, the effect of the electronic device performing audio equalization adjustment can be improved.

[0277] Optionally, the above audio equalization model includes an agent network; a processor 110, specifically configured to input the first audio frame parameters into the audio equalization model; encode the first audio frame parameters into a first feature vector through the fully connected hidden layer of the agent network; and output the action probability parameters corresponding to the first feature vector through the fully connected output layer of the agent network; and perform sampling processing on the action probability parameters corresponding to the first feature vector to obtain the gain of the first audio frame.

[0278] Optionally, the audio output unit 103 is configured to play the adjusted first audio frame through the speaker of the electronic device; the input unit 104 is configured to measure the output of the speaker to obtain a sound pressure signal and a distortion energy value; the processor 110 is specifically configured to determine a first reward value based on the sound pressure signal, the distortion energy value, the spectral difference value, and the gain change value, where the first spectral difference value is used to characterize the difference between the spectrum of the adjusted first audio frame and the spectrum of the preset audio frame, and the first gain change value is used to characterize the change amplitude between the gain of the first audio frame and the gain of the previous audio frame of the first audio frame.

[0279] Optionally, the above audio equalization model includes an agent network, at least two first policy update networks, and at least two second policy update networks, where one second policy update network corresponds to one first policy update network; the processor 110 is specifically configured to output the gain of the second audio frame based on the second audio frame parameters through the agent network; and output a target value through the second policy update network based on the first reward value, the first value parameter, and the action probability parameters of the second audio frame, where the first value parameter is a parameter obtained by the second policy update network based on the second audio frame parameters and the gain of the second audio frame; and update the network parameters of the first policy update network based on the target value and the second value parameter, where the second value parameter is a parameter obtained by the first policy update network based on the first audio frame parameters and the gain of the first audio frame; and output an updated second value parameter through the updated first policy update network; and update the network parameters of the agent network of the audio equalization model based on the updated second value parameter to obtain an updated audio equalization model.

[0280] Optionally, the processor 110 is further configured to, after updating the network parameters of the first policy update network based on the target value, update the network parameters of the second policy update network through the second policy update network based on the updated network parameters of the first policy update network.

[0281] Optionally, the processor 110 is further configured to input the first audio frame parameter and the gain of the first audio frame into the first policy update network; extract the spectral amplitude of the first audio frame and the local features of the spectral amplitudes of the first N audio frames of the first audio frame through the convolutional layer of the first policy update network, and encode the local features into a second feature vector; and, through the first policy update network, splice and fuse the second feature vector, the gain of the previous audio frame of the first audio frame, and the gain of the first audio frame to obtain a third feature vector; and, through the fully connected layer of the first policy update network, map the third feature vector to a second value parameter, where the second value parameter is used to represent the value of performing equalization processing on the first audio frame using the gain of the first audio frame.

[0282] Taking the electronic device 100 as an example of the electronic device in the model inference stage:

[0283] The processor 110 is configured to, when the third audio frame of the second audio signal is acquired, obtain a third audio frame parameter related to the third audio frame, where the third audio frame is an audio frame to be processed in real time, and the third audio frame parameter includes: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames of the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer; input the third audio frame parameter into the audio equalization model for audio equalization processing, and output the gain of the third audio frame; and adjust the spectral amplitude of the third audio frame using the gain of the third audio frame to obtain an adjusted third audio frame; the audio output unit 103 is configured to play the adjusted third audio frame through the speaker.

[0284] In the electronic device provided in the embodiments of the present application, since the electronic device can obtain the audio frame parameter related to the audio frame based on the acquired audio frame and input it into the audio equalization model to obtain the gain corresponding to the audio frame, the gain of the audio frame is output in real time based on the features of the audio frame, that is, the gain can more accurately adjust the audio frame in real time. Thus, the convenience of the electronic device for performing equalization adjustment on the audio can be improved, and at the same time, the effect of the electronic device for performing audio equalization adjustment can be improved.

[0285] Optionally, the above audio equalization model includes an experience replay buffer module; the processor 110 is further configured to determine a second reward value based on the adjusted third audio frame; obtain a fourth audio frame parameter corresponding to the fourth audio frame, where the fourth audio frame is the next audio frame of the third audio frame in the second audio signal; and cache the third audio frame parameter, the gain of the third audio frame, the second reward value, and the fourth audio frame parameter through the experience replay buffer module.

[0286] The electronic device provided by the embodiment of the present application can implement each process implemented by the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0287] For the beneficial effects of various implementation manners in this embodiment, reference may be specifically made to the beneficial effects of the corresponding implementation manners in the above method embodiment. To avoid repetition, it will not be elaborated here.

[0288] It should be understood that in the embodiment of the present application, the input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The graphics processor 1041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also referred to as a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. The other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0289] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0290] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 110 either.

[0291] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0292] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0293] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0294] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0295] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0296] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described method may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0297] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0298] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A model training method, characterized in that, The method includes: Inputting first audio frame parameters related to a first audio frame in a first audio signal into an audio equalization model for audio equalization processing to output a gain of the first audio frame, where the first audio frame parameters include: a spectral amplitude of the first audio frame, spectral amplitudes of the first N audio frames before the first audio frame, and a gain of the audio frame before the first audio frame, and N is a positive integer; Adjusting the spectral amplitude of the first audio frame by using the gain of the first audio frame to obtain the adjusted first audio frame; Determining a first reward value based on the adjusted first audio frame; Updating model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and second audio frame parameters related to a second audio frame to obtain the updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal.

2. The method according to claim 1, wherein The audio equalization model includes an agent network; The inputting first audio frame parameters related to a first audio frame in a first audio signal into an audio equalization model for audio equalization processing to output a gain of the first audio frame includes: Inputting the first audio frame parameters into the audio equalization model; Encoding the first audio frame parameters into a first feature vector through a fully connected hidden layer of the agent network; Outputting action probability parameters corresponding to the first feature vector through a fully connected output layer of the agent network; Performing sampling processing on the action probability parameters corresponding to the first feature vector to obtain the gain of the first audio frame.

3. The method according to claim 1, wherein The determining a first reward value based on the adjusted first audio frame includes: Playing the adjusted first audio frame through a speaker of an electronic device and measuring an output of the speaker to obtain a sound pressure signal and a distortion energy value; Determining the first reward value based on the sound pressure signal, the distortion energy value, a spectral difference value, and a gain change value, where the first spectral difference value is used to represent a difference between a spectrum of the adjusted first audio frame and a spectrum of a preset audio frame, and the first gain change value is used to represent a change amplitude between the gain of the first audio frame and the gain of the audio frame before the first audio frame.

4. The method according to claim 1, wherein The audio equalization model includes: an agent network, at least two first policy update networks, and at least two second policy update networks, and one of the second policy update networks corresponds to one of the first policy update networks; The updating model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value, and second audio frame parameters related to a second audio frame to obtain the updated audio equalization model includes: Outputting a gain of the second audio frame based on the second audio frame parameters through the agent network; Update the network through the second policy update network, and output a target value based on the first reward value, the first value parameter, and the action probability parameter of the second audio frame. The first value parameter is a parameter obtained by the second policy update network based on the second audio frame parameter and the gain of the second audio frame; Update the network parameters of the first policy update network based on the target value and the second value parameter. The second value parameter is a parameter obtained by the first policy update network based on the first audio frame parameter and the gain of the first audio frame; Output the updated second value parameter through the updated first policy update network; Update the network parameters of the agent network of the audio equalization model based on the updated second value parameter to obtain the updated audio equalization model.

5. The method according to claim 4, wherein After updating the network parameters of the first policy update network based on the target value, the method further includes: Update the network parameters of the second policy update network through the second policy update network based on the network parameters of the updated first policy update network.

6. The method according to claim 4, wherein The method further includes: Input the first audio frame parameter and the gain of the first audio frame into the first policy update network; Extract the spectral amplitude of the first audio frame and the local features of the spectral amplitudes of the first N audio frames before the first audio frame through the convolutional layer of the first policy update network, and encode the local features into a second feature vector; Through the first policy update network, splice and fuse the second feature vector, the gain of the previous audio frame of the first audio frame, and the gain of the first audio frame to obtain a third feature vector; Map the third feature vector to the second value parameter through the fully connected layer of the first policy update network. The second value parameter is used to represent the value of equalizing the first audio frame using the gain of the first audio frame.

7. A model inference method, characterized in that, The method includes: When the third audio frame of the second audio signal is collected, obtain the third audio frame parameter related to the third audio frame. The third audio frame is an audio frame that needs to be processed in real time. The third audio frame parameter includes: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer; Input the third audio frame parameter into the audio equalization model for audio equalization processing, and output the gain of the third audio frame; Adjust the spectral amplitude of the third audio frame using the gain of the third audio frame to obtain the adjusted third audio frame; Play the adjusted third audio frame through a speaker.

8. The method according to claim 7, wherein The audio equalization model includes an experience replay buffer module; the method further includes: Determine a second reward value based on the adjusted third audio frame; Obtain the fourth audio frame parameter corresponding to the fourth audio frame. The fourth audio frame is the next audio frame of the third audio frame in the second audio signal; Through the experience replay cache module, cache the third audio frame parameters, the gain of the third audio frame, the second reward value, and the fourth audio frame parameters.

9. A model training device, characterized in that, The model training device includes: a processing module, an adjustment module, a determination module, and an update module; The processing module is configured to input first audio frame parameters related to a first audio frame in a first audio signal into an audio equalization model for audio equalization processing, and output the gain of the first audio frame. The first audio frame parameters include: the spectral amplitude of the first audio frame, the spectral amplitudes of the first N audio frames before the first audio frame, the gain of the previous audio frame of the first audio frame, and N is a positive integer; The adjustment module is configured to adjust the spectral amplitude of the first audio frame by using the gain of the first audio frame output by the processing module to obtain the adjusted first audio frame; The determination module is configured to determine a first reward value based on the first audio frame adjusted by the adjustment module; The update module is configured to update the model parameters of the audio equalization model based on the first audio frame parameters, the gain of the first audio frame, the first reward value determined by the determination module, and second audio frame parameters related to a second audio frame to obtain the updated audio equalization model, where the second audio frame is the next audio frame of the first audio frame in the first audio signal.

10. A model inference device, characterized in that, The model inference device includes: an acquisition module, a processing module, an adjustment module, and a playback module; The acquisition module is configured to, when a third audio frame of a second audio signal is acquired, acquire third audio frame parameters related to the third audio frame. The third audio frame is an audio frame that needs to be processed in real time. The third audio frame parameters include: the spectral amplitude of the third audio frame, the spectral amplitudes of the first N audio frames before the third audio frame, the gain of the previous audio frame of the third audio frame, and N is a positive integer; The processing module is configured to input the third audio frame parameters acquired by the acquisition module into an audio equalization model for audio equalization processing, and output the gain of the third audio frame; The adjustment module is configured to adjust the spectral amplitude of the third audio frame by using the gain of the third audio frame output by the processing module to obtain the adjusted third audio frame; The playback module is configured to play the adjusted third audio frame through a speaker.

11. An electronic device, characterized in that, It includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, it implements the steps of the model training method according to any one of claims 1 to 6, or implements the steps of the model inference method according to claim 7 or 8.

12. A readable storage medium, characterized in that, The program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, it implements the steps of the model training method according to any one of claims 1 to 6, or implements the steps of the model inference method according to claim 7 or 8.