Audio processing method and device, storage medium and program product
By storing the feature extraction vectors of the encoding processing result of the initial audio in the server, and using the audio overscore model to restore high-frequency components, the storage and traffic waste problems caused by initial audio compression are solved, and the audio quality and user experience are improved.
Patent Information
- Application Number
- CN202510372144.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, since the initial audio lacks high-frequency components after compression, meaningless high-frequency data exists in the feature extraction vector, increasing server storage space and traffic consumption.
By storing the encoding processing result feature of the initial audio in the server, and processing the decoded audio using the trained audio overscore model, the high-frequency components are restored, and the restored audio is generated.
Saves server storage space and feature extraction vector traffic sent to the client, improves audio quality and improves user experience.
Smart Images

Figure CN120260523A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to an audio processing method, device, storage medium, and program product. Background Art
[0002] For some music applications or video applications, the playback quality of audio is a very important aspect. Master audio is the original version of audio completed in a recording studio, which represents the highest quality of audio. However, the initial audio stored in the background of the application is usually the audio obtained after compressing the master audio. These compressions result in the loss of high-frequency components in the master audio, thereby reducing the quality of the audio.
[0003] Therefore, some current applications perform upsampling processing on the initial audio to make the frequency of the upsampled audio consistent with the frequency of the master audio, and then perform encoding processing on the upsampled audio to obtain a feature extraction vector, and store the feature extraction vector in the server to reduce the storage space. When a user of the client needs to play the corresponding audio, the server sends the feature extraction vector to the client for decoding to obtain a decoded audio, and the decoded audio is the audio supplemented with high-frequency components.
[0004] However, in the above process, since the high-frequency components are missing in the initial audio, even if the upsampling processing is performed to increase its frequency to the frequency of the master audio, actually the audio does not have energy in the high-frequency components, resulting in a large amount of meaningless high-frequency data in the feature extraction vector, further resulting in an increase in the storage space of the feature extraction vector in the server, and also resulting in an increase in the traffic when the server sends the feature extraction vector to the client, which is a waste of resources. Summary of the Invention
[0005] Embodiments of the present disclosure provide an audio processing method, which can save the storage space of the feature extraction vector in the server and the traffic when the feature extraction vector is sent to the client, thereby avoiding waste of resources. The technical solution is as follows:
[0006] In a first aspect, an audio processing method is provided, and the method includes:
[0007] Generating a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server, where the feature extraction vector is obtained by encoding an initial audio corresponding to the target track;
[0008] Performing audio super-resolution processing on the decoded audio to obtain a restored audio corresponding to the target track; the sampling rate of the restored audio is higher than the sampling rate of the decoded audio.
[0009] In a possible implementation, generating the decoded audio based on the feature extraction vector corresponding to the target track pre-stored in the server includes:
[0010] Sending a request for obtaining the master audio of the target track to the server;
[0011] Receiving the feature extraction vector sent by the server;
[0012] Performing decoding processing on the feature extraction vector to obtain the decoded audio.
[0013] In a possible implementation, generating the decoded audio based on the feature extraction vector corresponding to the target track pre-stored in the server includes:
[0014] Sending a request for obtaining the master audio of the target track to the server;
[0015] Receiving the decoded audio sent by the server, where the decoded audio is obtained by the server performing decoding processing on the feature extraction vector.
[0016] In a possible implementation, performing audio super-resolution processing on the decoded audio to obtain the restored audio corresponding to the target track includes:
[0017] Encoding the spectral features of the decoded audio through the encoding network of the trained audio super-resolution model to obtain a latent vector;
[0018] Decoding the latent vector through the decoding network of the trained audio super-resolution model to obtain high-frequency band spectral features;
[0019] Performing splicing processing on the spectral features of the decoded audio and the high-frequency spectral features to obtain full-band spectral features;
[0020] Determining the restored audio based on the full-band spectral features.
[0021] In a possible implementation, before encoding the spectral features of the decoded audio through the encoding network of the trained audio super-resolution model, the method further includes:
[0022] Performing upsampling processing on the decoded audio to obtain the target audio after upsampling processing;
[0023] Encoding the spectral features of the decoded audio through the encoding network of the trained audio super-resolution model to obtain a latent vector, including:
[0024] The encoding network of the audio super-resolution model completed through training encodes the spectral features of the target audio to obtain a latent vector.
[0025] In a possible implementation, the method further includes:
[0026] Obtain a first sample training set, where the first sample training set includes sample initial audios corresponding to multiple sample tracks;
[0027] Encode the sample initial audio through the encoder of the pre-trained audio processing model to obtain a sample feature extraction vector; decode the sample feature extraction vector through the decoder of the pre-trained audio processing model to obtain a sample decoded audio;
[0028] Train the pre-trained audio processing model based on the sample initial audio and the sample decoded audio until a preset training end condition is met, to obtain a trained audio processing model;
[0029] Encode the initial audio corresponding to the target track through the encoder of the trained audio processing model to obtain a feature extraction vector corresponding to the target track;
[0030] The generating the decoded audio based on the feature extraction vector corresponding to the target track pre-stored in the server includes:
[0031] Decode the feature extraction vector through the decoder of the trained audio processing model to obtain the decoded audio.
[0032] In a possible implementation, the method further includes:
[0033] Obtain a second sample training set, where the second sample training set includes sample master audios corresponding to the multiple sample tracks;
[0034] Perform downsampling processing on each of the sample master audios to obtain the first sample training set.
[0035] In a second aspect, there is provided an audio processing apparatus, the apparatus includes:
[0036] A decoding module, configured to generate a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server, where the feature extraction vector is obtained by encoding an initial audio corresponding to the target track;
[0037] An audio super-resolution module, configured to perform audio super-resolution processing on the decoded audio to obtain a restored audio corresponding to the target track; the sampling rate of the restored audio is higher than the sampling rate of the decoded audio.
[0038] In a possible implementation, the decoding module is configured to:
[0039] Send a request for obtaining the master audio of the target track to the server;
[0040] Receive the feature extraction vector sent by the server;
[0041] Perform decoding processing on the feature extraction vector to obtain the decoded audio.
[0042] In a possible implementation, the decoding module is configured to:
[0043] Send a request for obtaining the master audio of the target track to the server;
[0044] Receive the decoded audio sent by the server, where the decoded audio is obtained by the server after performing decoding processing on the feature extraction vector.
[0045] In a possible implementation, the audio super-resolution module is configured to:
[0046] Perform encoding processing on the spectral features of the decoded audio through the encoding network of the trained audio super-resolution model to obtain a latent vector;
[0047] Perform decoding processing on the latent vector through the decoding network of the trained audio super-resolution model to obtain high-frequency band spectral features;
[0048] Perform splicing processing on the spectral features of the decoded audio and the high-frequency spectral features to obtain full-band spectral features;
[0049] Determine the restored audio based on the full-band spectral features.
[0050] In a possible implementation, the audio super-resolution module is further configured to:
[0051] Perform upsampling processing on the decoded audio to obtain a target audio after upsampling processing;
[0052] The audio super-resolution module is configured to:
[0053] Perform encoding processing on the spectral features of the target audio through the encoding network of the trained audio super-resolution model to obtain a latent vector.
[0054] In a possible implementation, the apparatus further includes a training module, configured to:
[0055] Obtain a first sample training set, where the first sample training set includes sample initial audios corresponding to multiple sample tracks;
[0056] Encode the sample initial audio through the encoder of the pre-trained audio processing model to obtain a sample feature extraction vector; decode the sample feature extraction vector through the decoder of the pre-trained audio processing model to obtain a sample decoded audio;
[0057] Train the pre-trained audio processing model based on the sample initial audio and the sample decoded audio until a preset training end condition is met, to obtain a trained audio processing model;
[0058] Encode the initial audio corresponding to the target track through the encoder of the trained audio processing model to obtain a feature extraction vector corresponding to the target track;
[0059] The decoding module is used for:
[0060] Decode the feature extraction vector through the decoder of the trained audio processing model to obtain the decoded audio.
[0061] In a possible implementation manner, the training module is further used for:
[0062] Obtain a second sample training set, where the second sample training set includes sample master audio corresponding to the multiple sample tracks;
[0063] Perform downsampling processing on each of the sample master audios to obtain the first sample training set.
[0064] In a third aspect, a computer device is provided, which includes a processor and a memory. At least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the operations performed by any one of the above audio processing methods.
[0065] In a fourth aspect, a computer-readable storage medium is provided, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the operations performed by any one of the above audio processing methods.
[0066] In a fifth aspect, a computer program product is provided, in which at least one instruction is included, and the at least one instruction is loaded and executed by the processor to implement the operations performed by any one of the above audio processing methods.
[0067] The beneficial effects brought by the technical solution provided in the embodiments of the present disclosure are as follows: In the solution mentioned in the embodiments of the present disclosure, the data stored in the server is the feature extraction vector obtained by encoding the original audio. Since there is no meaningless audio data in the original audio, there is no meaningless high-frequency data in the feature extraction vector, which saves the storage space of the feature extraction vector in the server and also saves the traffic when the feature extraction vector is sent to the client, thus avoiding waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0070] Figure 2 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0071] Figure 3 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0072] Figure 4 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0073] Figure 5 is a schematic structural diagram of an audio processing device provided by an embodiment of the present disclosure;
[0074] Figure 6 is a block diagram of the structure of a terminal provided by an embodiment of the present disclosure;
[0075] Figure 7 is a block diagram of the structure of a server provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will further describe the embodiments of the present disclosure in detail in conjunction with the accompanying drawings.
[0077] The embodiments of the present disclosure provide an audio processing method, which can be implemented by a computer device. The computer device can be a terminal and a server, etc. The terminal can be a desktop computer, a laptop computer, a tablet computer, a mobile phone, etc.
[0078] The computer device may include a processor, a memory, a communication component, etc.
[0079] The processor can be a central processing unit (CPU). The processor can be used to read instructions and process data. For example, generate decoded audio based on the feature extraction vector corresponding to the target track pre-stored in the server, perform audio super-resolution processing on the decoded audio, and so on.
[0080] The memory can be various volatile memories or non-volatile memories, such as solid state disk (SSD), dynamic random access memory (DRAM), etc. The memory can be used for data storage. For example, store the data of the feature extraction vector corresponding to the target track, store the data of the generated decoded audio, store the data of the restored audio corresponding to the obtained target track, and so on.
[0081] The communication component can be a wired network connector, a wireless fidelity (WiFi) module, a Bluetooth module, a cellular network communication module, etc. The communication component can be used for data transmission with other devices. For example, send a request for obtaining the master audio of the target track to the server, receive the feature extraction vector or the restored audio sent by the server, and so on.
[0082] The target application program is installed in the computer device, such as a music application program, a video application program, etc. The user can open the target application program on the client and play the target track on the target application program. When the client receives a play request or a download request for the target track, the target track can be processed by the audio processing method provided by the embodiments of the present disclosure to obtain its corresponding restored audio, and the restored audio is the audio that makes up for the high-frequency components, thereby improving the audio quality of the target track.
[0083] Figure 1 and Figure 2 are the flowcharts of an audio processing method provided by the embodiments of the present disclosure. Refer to Figure 1 and Figure 2 , this embodiment includes:
[0084] 101. Generate decoded audio based on the feature extraction vector corresponding to the target track pre-stored in the server.
[0085] In implementation, the server in the background of the target application program can store the feature extraction vectors corresponding to each initial audio, where the initial audio is the audio with high-frequency components missing due to the compression of the master audio. For example, the frequency of the master audio is 192 kHz, and the frequency of the initial audio obtained after compressing the master audio is 44100 Hz.
[0086] When the user plays or downloads a target track through a target application on the client, the client can generate decoded audio based on the feature extraction vector corresponding to the target track pre-stored on the server.
[0087] Among them, the feature extraction vector is obtained by encoding the initial audio corresponding to the target track. This encoding process can be implemented by an encoder. The encoder compresses the initial data from a high-dimensional space to a low-dimensional space, thereby obtaining a feature extraction vector with a smaller data volume. In a possible implementation, a trained encoder can be set in the server. Whenever a new initial audio is received in the background, the initial audio can be input into the trained encoder for encoding processing to obtain the feature extraction vector corresponding to the initial audio output by the trained encoder. Then, the server stores the feature extraction vector correspondingly.
[0088] The decoded audio is obtained by decoding the feature extraction vector. This decoding process can be implemented by a trained decoder, which decodes the feature extraction vector in the low-dimensional space back to the high-dimensional space, thereby obtaining the decoded audio. In the embodiments of the present disclosure, the trained decoder can be set in the server or in the client, and it can be set according to actual situations and requirements. Below, these two situations will be introduced in more detail:
[0089] In a possible implementation, the trained decoder can be set in the client. The processing method is: sending a request to obtain the master audio of the target track to the server; receiving the feature extraction vector sent by the server; and performing decoding processing on the feature extraction vector to obtain decoded audio.
[0090] In implementation, when the user wants to play or download the master audio of the target track on the target application, the user can click the play option or download option of the master audio of the corresponding target track, so that the client sends a request to obtain the master audio of the target track to the server. When the server receives the request to obtain the master audio of the target track, it sends the feature extraction vector corresponding to the target track pre-stored to the client. When the client receives the feature extraction vector, it can input it into the trained decoder to perform decoding processing on the feature extraction vector, thereby obtaining the output decoded audio.
[0091] In this processing method, the server only needs to store the feature extraction vectors corresponding to each initial audio. Compared with the initial audio after upsampling processing in the related art, there is no meaningless audio data in the initial audio in the embodiments of the present disclosure. Therefore, there is no meaningless high-frequency data in the feature extraction vectors corresponding to the initial audio, saving the storage space of the feature extraction vectors in the server and also saving the traffic when the feature extraction vectors are sent to the client, thus avoiding waste of resources.
[0092] In another possible implementation manner, the trained decoder can be set in the server, and its processing method is: sending the master audio of the target track to the server; receiving the decoded audio sent by the server, where the decoded audio is obtained by the server after decoding the feature extraction vector.
[0093] In implementation, when a user wants to play or download the master audio of a target track on a target application, the user can click the play option or download option of the corresponding master audio of the target track, so that the client sends a request for obtaining the master audio of the target track to the server. When the server receives the request for obtaining the master audio of the target track, it inputs the feature extraction vector corresponding to the target track stored in advance into the trained decoder for decoding processing to obtain the output decoded audio. Then, the server can send the obtained decoded audio to the client.
[0094] In this processing method, the server only needs to store the feature extraction vectors corresponding to each initial audio. Compared with the initial audio after upsampling processing in the related art, there is no meaningless audio data in the initial audio in the embodiments of the present disclosure. Therefore, there is no meaningless high-frequency data in the feature extraction vectors corresponding to the initial audio, saving the storage space of the feature extraction vectors in the server. Moreover, the data sent to the client is the decoded audio that has not been upsampled, saving traffic, thus avoiding waste of resources.
[0095] It can be understood that both of the above two processing methods complete the distribution of the decoded audio through the trained encoder and the trained decoder. In this process, the frequency of the decoded audio can be the same as that of the initial audio, so that the decoded audio output by the trained decoder is as identical as possible to the initial audio.
[0096] The present disclosure embodiments do not specifically limit the size relationship between the data volume of the feature extraction vector and the data volume of the decoded audio, which can be set according to requirements and the types of machine learning models applied by the decoder and the encoder.
[0097] 102. Perform audio super-resolution processing on the decoded audio to obtain the restored audio corresponding to the target track.
[0098] In implementation, a trained audio super-resolution model is set in the client. When the client obtains the decoded audio through the above step 101, it can perform audio super-resolution processing on the decoded audio based on the trained audio super-resolution model to compensate for the missing high-frequency components in the decoded audio, thereby obtaining a restored audio. The restored audio is a high-quality audio that is as close as possible to the master audio after this series of processes. The frequency of the restored audio can be the same as that of the master audio, thus greatly improving the similarity between the two, and further improving the audio quality of the restored audio.
[0099] After obtaining the restored audio, the client can play or download it, so that the user can obtain high-quality audio, thereby improving the user experience.
[0100] In a possible implementation, referring to Figure 3 , the method of performing audio super-resolution processing on the decoded audio based on the trained audio super-resolution model can be: encoding the spectral features of the decoded audio through the encoding network of the trained audio super-resolution model to obtain a latent vector; decoding the latent vector through the decoding network of the trained audio super-resolution model to obtain high-frequency band spectral features; splicing the spectral features of the decoded audio and the high-frequency spectral features to obtain full-band spectral features; determining the restored audio based on the full-band spectral features.
[0101] In implementation, the audio super-resolution model can include an encoding network and a decoding network. After obtaining the decoded audio, feature extraction processing can be performed on the decoded audio first to obtain the spectral features of the decoded audio, and then the spectral features of the decoded audio are input into the trained audio super-resolution model. First, the encoding network in the trained audio super-resolution model encodes the spectral features of the decoded audio to obtain an output latent vector. Then, the latent vector is input into the decoding network in the trained audio super-resolution model for decoding processing to obtain an output high-frequency band spectral feature. The high-frequency band spectral feature is the missing high-frequency component predicted by the trained audio super-resolution model for the decoded audio.
[0102] Since the decoded audio is an audio lacking high-frequency components, the spectral features of the decoded audio belong to low-frequency band spectral features. Therefore, after obtaining the high-frequency band spectral features, the high-frequency band spectral features can be spliced with the spectral features of the decoded audio to obtain full-band spectral features. It can be understood that during splicing, corresponding splicing can be performed in the order of each audio frame.
[0103] After obtaining the full-band spectral features, corresponding time-domain waveforms can be obtained through inverse Fourier transform and other processes to obtain the restored audio.
[0104] Furthermore, referring toFigure 4 Before encoding the spectral features of the decoded audio by the encoding network of the audio super-resolution model completed through training, the following processing can also be performed: Upsample the decoded audio to obtain the target audio after upsampling. Correspondingly, the method for the encoding network of the audio super-resolution model completed through training to obtain the latent vector is: Encode the spectral features of the target audio by the encoding network of the audio super-resolution model completed through training to obtain the latent vector.
[0105] In implementation, after obtaining the decoded audio, the decoded audio can be upsampled first to obtain the target audio, so as to narrow the gap between the frequency of the target audio and the frequency of the master audio of the target track, thereby facilitating the design of the audio super-resolution model.
[0106] After obtaining the target audio, feature extraction processing can be performed on the target audio to obtain the spectral features of the target audio, and then the spectral features of the target audio are input into the audio super-resolution model completed through training. The encoding network in the audio super-resolution model completed through training encodes the spectral features of the target audio to obtain the latent vector, and the decoding network in the audio super-resolution model completed through training decodes the latent vector to obtain the high-frequency band spectral features.
[0107] In a possible implementation manner, the frequency of the target audio after upsampling can be the same as the frequency of the master audio corresponding to the target audio, thereby facilitating the design of the audio super-resolution model more.
[0108] In a possible implementation manner, the encoder and decoder described in step 101 above can be the encoder and decoder included in a trained audio processing model. Before step 101, the pre-trained audio processing model can also be trained to obtain the trained audio processing model. The training method can be: Obtain a first sample training set, where the first sample training set includes sample initial audios corresponding to multiple sample tracks; Encode the sample initial audios by the encoder of the pre-trained audio processing model to obtain sample feature extraction vectors; Decode the sample feature extraction vectors by the decoder of the pre-trained audio processing model to obtain sample decoded audios; Train the pre-trained audio processing model based on the sample initial audios and the sample decoded audios until a preset training end condition is met to obtain the trained audio processing model.
[0109] In implementation, the sample initial audio corresponding to the sample track can be obtained from the first sample training set, and the sample initial audio can be an audio lacking high-frequency components. The obtained sample initial audio is input into a pre-trained audio processing model. The encoder in the pre-trained audio processing model encodes the sample initial audio to obtain a sample feature extraction vector, and the decoder in the pre-trained audio processing model decodes the sample feature extraction vector to obtain a sample decoded audio.
[0110] Then, the sample initial audio and the sample decoded audio are input into a loss function to obtain a loss value. Based on this loss value, the encoder and decoder in the pre-trained audio processing model are adjusted. Then, it is determined whether the loss value meets the preset training end condition. If it meets, the trained audio processing model is obtained, that is, the trained encoder and the trained decoder are obtained. If it does not meet, a new sample initial audio is obtained from the first sample training set, and the pre-trained audio processing model is trained again until the preset training end condition is met.
[0111] In a possible implementation manner, the method for obtaining the first sample training set can be: obtaining a second sample training set, where the second sample training set includes sample master audio corresponding to multiple sample tracks; performing downsampling processing on each sample master audio respectively to obtain the first sample training set.
[0112] In implementation, downsampling processing can be performed on the sample master audio corresponding to multiple sample tracks, so that it lacks high-frequency components, thereby obtaining the sample initial audio corresponding to each sample track, that is, obtaining the first sample training set.
[0113] In a possible implementation manner, the following is an example of the preset training end condition:
[0114] First, the number of times of adjusting the parameters of the pre-trained audio processing model reaches the parameter adjustment times threshold. In implementation, the staff can preset the parameter adjustment times threshold in advance. When the parameter adjustment times reach the parameter adjustment times threshold, the training can be stopped, and the encoder obtained after the last parameter adjustment is determined as the trained encoder, and the decoder obtained after the last parameter adjustment is determined as the trained decoder. The parameter adjustment times threshold can be any reasonable value, such as 300 or 350, etc. The embodiments of the present disclosure do not limit this.
[0115] Second, the loss values obtained from continuous preset number of trainings are all less than the preset loss value threshold. Both the preset number and the preset loss value threshold can be any reasonable values. For example, the preset number can be 4 or 5, etc., and the preset loss value threshold can be 0.06, etc. The embodiments of the present disclosure do not limit this.
[0116] Thirdly, the number of times of parameter tuning reaches the parameter tuning times threshold, and the loss values obtained from consecutive preset number of training times are all less than the preset loss value threshold.
[0117] The preset training end condition can be any one of the above three, or other end conditions, and the embodiments of the present disclosure do not limit this.
[0118] In the embodiments of the present disclosure, after obtaining the trained encoder and the trained decoder, the pre-trained audio super-resolution model can also be trained to obtain the trained audio super-resolution model. The structure of the audio super-resolution model in the embodiments of the present disclosure can be a generative model in Generative Adversarial Networks (GAN), and its training method can be: obtaining a third sample training set, where the third sample training set includes a plurality of second sample master audio; performing downsampling processing on the second sample master audio to obtain the second sample initial audio corresponding to the second sample master audio; inputting the second sample initial audio into the trained encoder to obtain the second sample feature extraction vector, and inputting the second sample feature extraction vector into the trained decoder to obtain the second sample decoded audio; encoding the spectral features of the second sample decoded audio through the encoding network in the pre-trained audio super-resolution model to obtain the sample hidden vector corresponding to the second sample target audio; decoding the sample hidden vector through the decoding network in the pre-trained audio super-resolution model to obtain the predicted high-frequency band spectral features; splicing the spectral features of the second sample decoded audio and the predicted high-frequency band spectral features to obtain the predicted full-frequency band spectral features; determining the predicted restored audio based on the predicted full-frequency band spectral features; inputting the predicted restored audio and the second sample master audio into the pre-trained discriminant model to obtain the eigenvalue; training the pre-trained audio super-resolution model and the pre-trained discriminant model based on the eigenvalue and the preset standard eigenvalue, and if the second preset training end condition is satisfied, the trained audio super-resolution model and the trained discriminant model are obtained.
[0119] In implementation, after training the pre-trained audio processing model to obtain the trained audio processing model, the pre-trained audio super-resolution model can be trained.
[0120] When training a pre-trained audio super-resolution model, the second sample master audio can be obtained from the third sample training set and downsampled to obtain the second sample initial audio with missing high-frequency components. The second sample initial audio is input into the trained encoder to obtain the second sample feature extraction vector, and then the second sample feature extraction vector is input into the trained decoder to obtain the output second sample decoded audio. Then, the spectral features of the obtained second sample decoded audio can be input into the pre-trained audio super-resolution model. The encoding network in the pre-trained audio super-resolution model encodes the spectral features of the second sample decoded audio to obtain the sample latent vector, and the decoding network in the pre-trained audio super-resolution model decodes the sample latent vector to obtain the predicted high-frequency band spectral features. The spectral features of the second sample decoded audio and the predicted high-frequency band spectral features are concatenated to obtain the predicted full-frequency band spectral features, and based on the predicted full-frequency band spectral features, the predicted restored audio is determined.
[0121] Then, the predicted restored audio and the second sample master audio are input into the pre-trained discriminant model to obtain the eigenvalue. The eigenvalue and the preset standard eigenvalue can be input into the second loss function to obtain the second loss value. Based on this second loss value, the pre-trained audio super-resolution model and the pre-trained discriminant model are tuned. Then, it is determined whether the second loss value satisfies the second preset training end condition. If it satisfies, the trained audio super-resolution model and the trained discriminant model are obtained. If it does not satisfy, a new second sample master audio is obtained from the second sample training set, and the pre-trained audio super-resolution model and the pre-trained discriminant model are trained again until the second preset training end condition is satisfied. Among them, the preset standard eigenvalue can be set according to the design. For example, in the embodiments of the present disclosure, the preset standard eigenvalue can be 0.5, that is, the discriminant model cannot determine whether the predicted restored audio is the real audio, so as to achieve the Nash equilibrium.
[0122] Through the above multiple trainings, the high-frequency band spectral features output by the trained audio super-resolution model can gradually tend to be the same as the high-frequency components in the second sample master audio, thereby completing the training of the audio super-resolution model.
[0123] After obtaining the trained audio super-resolution model through the above training, the audio super-resolution processing in step 102 above can be performed through the trained audio super-resolution model.
[0124] In the embodiments of the present disclosure, the setting of the second preset training end condition can be the same as one of the preset training end conditions described above, or it can also be other settings. The embodiments of the present disclosure do not limit this.
[0125] An embodiment of the present disclosure provides an audio processing device, which may be the computer device in the above embodiment, such as Figure 5 As shown, the device includes:
[0126] A decoding module 520, configured to generate a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server, where the feature extraction vector is obtained by encoding an initial audio corresponding to the target track;
[0127] An audio super-resolution module 530, configured to perform audio super-resolution processing on the decoded audio to obtain a restored audio corresponding to the target track; the sampling rate of the restored audio is higher than the sampling rate of the decoded audio.
[0128] In a possible implementation manner, the decoding module 520 is configured to:
[0129] Send a request for obtaining the master audio of the target track to the server;
[0130] Receive the feature extraction vector sent by the server;
[0131] Perform decoding processing on the feature extraction vector to obtain the decoded audio.
[0132] In a possible implementation manner, the decoding module 520 is configured to:
[0133] Send a request for obtaining the master audio of the target track to the server;
[0134] Receive the decoded audio sent by the server, where the decoded audio is obtained by the server performing decoding processing on the feature extraction vector.
[0135] In a possible implementation manner, the audio super-resolution module 530 is configured to:
[0136] Encode the spectral features of the decoded audio through an encoding network of a trained audio super-resolution model to obtain a latent vector;
[0137] Decode the latent vector through a decoding network of a trained audio super-resolution model to obtain high-frequency band spectral features;
[0138] Perform splicing processing on the spectral features of the decoded audio and the high-frequency spectral features to obtain full-band spectral features;
[0139] Determine the restored audio based on the full-band spectral features.
[0140] In a possible implementation manner, the audio super-resolution module 530 is further configured to:
[0141] Perform upsampling processing on the decoded audio to obtain the target audio after upsampling processing;
[0142] The audio super-resolution module 530 is used for:
[0143] Encode the spectral features of the target audio through the encoding network of the trained audio super-resolution model to obtain a latent vector.
[0144] In a possible implementation, the device further includes a training module 510, which is used for:
[0145] Obtain a first sample training set, where the first sample training set includes sample initial audios corresponding to multiple sample tracks;
[0146] Encode the sample initial audio through the encoder of the pre-trained audio processing model to obtain a sample feature extraction vector; decode the sample feature extraction vector through the decoder of the pre-trained audio processing model to obtain a sample decoded audio;
[0147] Train the pre-trained audio processing model based on the sample initial audio and the sample decoded audio until a preset training end condition is met, to obtain a trained audio processing model;
[0148] Encode the initial audio corresponding to the target track through the encoder of the trained audio processing model to obtain the feature extraction vector corresponding to the target track;
[0149] The decoding module 520 is used for:
[0150] Decode the feature extraction vector through the decoder of the trained audio processing model to obtain the decoded audio.
[0151] In a possible implementation, the training module 510 is further used for:
[0152] Obtain a second sample training set, where the second sample training set includes sample master audios corresponding to the multiple sample tracks;
[0153] Perform downsampling processing on each of the sample master audios to obtain the first sample training set.
[0154] It should be noted that: when the audio processing device provided in the above embodiment performs audio processing, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device provided in the above embodiment and the embodiment of the audio processing method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0155] Figure 6 FIG. shows a structural block diagram of a terminal 600 provided by an exemplary embodiment of the present application. The terminal may be the computer device in the above embodiment. The terminal 600 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 600 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0156] Generally, the terminal 600 includes: a processor 601 and a memory 602.
[0157] The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0158] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one instruction for being executed by the processor 601 to implement the audio processing method provided in the method embodiments of the present application.
[0159] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 604, a display screen 605, a camera 606, an audio circuit 607, a positioning component 608, and a power supply 609.
[0160] The peripheral device interface 603 may be used to connect at least one peripheral device related to I / O (input / output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 may be implemented on a separate chip or circuit board, and the present embodiment does not limit this.
[0161] The radio frequency circuit 604 is used to receive and transmit RF (radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 604 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (wireless fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (near field communication), and the present application does not limit this.
[0162] The display screen 605 is used to display the UI (user interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input as control signals to the processor 601 for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, which is disposed on the front panel of the terminal 600; in other embodiments, there may be at least two display screens 605, which are respectively disposed on different surfaces of the terminal 600 or are in a foldable design; in still other embodiments, the display screen 605 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal 600. Even further, the display screen 605 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 605 can be prepared using materials such as LCD (liquid crystal display) and OLED (organic light-emitting diode).
[0163] The camera module 606 is used to capture images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (virtual reality) shooting functions, or other combined shooting functions. In some embodiments, the camera module 606 may further include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. A two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0164] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headphone jack.
[0165] The positioning component 608 is used to locate the current geographical location of the terminal 600 to implement navigation or LBS (location-based service). The positioning component 608 may be a positioning component based on GPS (global positioning system), Beidou system, GLONASS system or Galileo system.
[0166] The power supply 609 is used to supply power to each component in the terminal 600. The power supply 609 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 609 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0167] In some embodiments, the terminal 600 further includes one or more sensors 610. The one or more sensors 610 include but are not limited to: an acceleration sensor 611, a gyroscope sensor 612, a pressure sensor 613, a fingerprint sensor 614, an optical sensor 615 and a proximity sensor 616.
[0168] The acceleration sensor 611 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 600. For example, the acceleration sensor 611 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 611. The acceleration sensor 611 can also be used for collecting game or user's motion data.
[0169] The gyroscope sensor 612 can detect the body direction and rotation angle of the terminal 600. The gyroscope sensor 612 can cooperate with the acceleration sensor 611 to collect the 3D actions of the user on the terminal 600. Based on the data collected by the gyroscope sensor 612, the processor 601 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0170] The pressure sensor 613 can be disposed on the side frame of the terminal 600 and / or the lower layer of the display screen 605. When the pressure sensor 613 is disposed on the side frame of the terminal 600, it can detect the holding signal of the user on the terminal 600, and the processor 601 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 613. When the pressure sensor 613 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0171] The fingerprint sensor 614 is used to collect the fingerprints of the user. The processor 601 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 614, or the fingerprint sensor 614 can identify the user's identity according to the collected fingerprints. When the identified user identity is a trusted identity, the processor 601 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 614 can be disposed on the front, back, or side of the terminal 600. When there are physical buttons or manufacturer logos on the terminal 600, the fingerprint sensor 614 can be integrated with the physical buttons or manufacturer logos.
[0172] The optical sensor 615 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera module 606 according to the ambient light intensity collected by the optical sensor 615.
[0173] The proximity sensor 616, also known as the distance sensor, is usually disposed on the front panel of the terminal 600. The proximity sensor 616 is used to collect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the lit state to the off state; when the proximity sensor 616 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the off state to the lit state.
[0174] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the terminal 600, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.
[0175] Figure 7 is a schematic structural diagram of a server provided by an embodiment of the present disclosure. The server 700 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 701 and one or more memories 702. Among them, at least one instruction is stored in the memory 702, and the at least one instruction is loaded and executed by the processor 701 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0176] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in the terminal to complete the audio processing method or the training method of the audio super-resolution model in the above embodiments. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (read-only memory), a RAM (random access memory), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0177] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0178] It should be noted that the information involved in this disclosure (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the "initial audio", "decoded audio", "feature extraction vector", etc. involved in this disclosure are all obtained under full authorization.
[0179] The foregoing are only optional embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An audio processing method, characterized in that, The method includes: Generating a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server, where the feature extraction vector is obtained by encoding an initial audio corresponding to the target track; Performing audio super-resolution processing on the decoded audio to obtain a restored audio corresponding to the target track; the sampling rate of the restored audio is higher than that of the decoded audio.
2. The method according to claim 1, wherein The generating a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server includes: Sending a master audio acquisition request for the target track to the server; Receiving the feature extraction vector sent by the server; Performing decoding processing on the feature extraction vector to obtain the decoded audio.
3. The method according to claim 1, characterized in that, The generating a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server includes: Sending a master audio acquisition request for the target track to the server; Receiving the decoded audio sent by the server, where the decoded audio is obtained by the server performing decoding processing on the feature extraction vector.
4. The method according to claim 1, characterized in that, The performing audio super-resolution processing on the decoded audio to obtain a restored audio corresponding to the target track includes: Encoding the spectral features of the decoded audio through an encoding network of a trained audio super-resolution model to obtain a latent vector; Decoding the latent vector through a decoding network of the trained audio super-resolution model to obtain high-frequency band spectral features; Performing splicing processing on the spectral features of the decoded audio and the high-frequency spectral features to obtain full-band spectral features; Determining the restored audio based on the full-band spectral features.
5. The method according to claim 4, wherein Before encoding the spectral features of the decoded audio through an encoding network of the trained audio super-resolution model, the method further includes: Performing upsampling processing on the decoded audio to obtain a target audio after upsampling processing; The encoding the spectral features of the decoded audio through an encoding network of the trained audio super-resolution model to obtain a latent vector includes: Encoding the spectral features of the target audio through an encoding network of the trained audio super-resolution model to obtain a latent vector.
6. The method according to claim 1, characterized in that, The method further includes: Obtaining a first sample training set, where the first sample training set includes sample initial audios corresponding to multiple sample tracks; Encoding the sample initial audios through an encoder of a pre-trained audio processing model to obtain sample feature extraction vectors; decoding the sample feature extraction vectors through a decoder of the pre-trained audio processing model to obtain sample decoded audios; Training the pre-trained audio processing model based on the sample initial audios and the sample decoded audios until a preset training end condition is satisfied to obtain a trained audio processing model; Encoding the initial audio corresponding to the target track through an encoder of the trained audio processing model to obtain a feature extraction vector corresponding to the target track; The generating a decoded audio based on a feature extraction vector corresponding to a target track pre-stored in a server includes: The decoder of the audio processing model completed through the training decodes the feature extraction vector to obtain the decoded audio.
7. The method according to claim 6, characterized in that, The method further includes: obtaining a second sample training set, where the second sample training set includes the sample master audio corresponding to the multiple sample tracks; performing downsampling processing on each of the sample master audios to obtain the first sample training set.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction is stored in the memory. The at least one instruction is loaded and executed by the processor to implement the operations performed by the audio processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, At least one instruction is stored in the storage medium. The at least one instruction is loaded and executed by the processor to implement the operations performed by the audio processing method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes at least one instruction. The at least one instruction is loaded and executed by the processor to implement the operations performed by the audio processing method according to any one of claims 1 to 7.