Audio processing method, device and electronic device

By using pitch frequency set perceptual domain spectrum information compression and upsampling neural network model decoding during audio encoding, the problem of poor audio encoding and decoding quality at low bit rates is solved, high-quality audio recovery is achieved, and the user experience is improved.

CN113903345BActive Publication Date: 2025-09-26DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111151494.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-09-26
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Existing technologies have difficulty achieving high-quality audio encoding and decoding at low bit rates, resulting in degraded audio quality and poor user experience.

Method used

The audio data is compressed using the pitch frequency set perceptual domain spectrum information, and is decoded using an upsampling neural network model with low computational complexity to achieve audio encoding and decoding at low bit rates.

Benefits of technology

High-quality audio encoding and decoding is achieved at low bit rates, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903345B_ABST
    Figure CN113903345B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio processing method, device, and electronic device. The method includes: obtaining audio data to be processed; wherein the audio data to be processed includes at least one audio frame; extracting audio feature information corresponding to each of the audio frames, and performing quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data; sending the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed, effectively achieving encoding and decoding at a low bit rate and improving user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of audio technology, and in particular to an audio processing method, device, and electronic device. Background Art

[0002] During audio transmission, audio encoding and decoding are often required. That is, the encoder needs to compress and encode the audio to obtain the corresponding encoded audio. After obtaining the encoded audio, the encoded audio is sent to the decoder, so that the decoder can decode the encoded audio and recover the audio.

[0003] Currently, linear prediction analysis (LPA) modeling is commonly used to encode and decode audio data. However, this approach produces a significant amount of residual information, making it difficult to achieve low-bitrate encoding and decoding. Therefore, a method for encoding and decoding audio at low bit rates is urgently needed. Summary of the Invention

[0004] The embodiments of the present disclosure provide an audio processing method, device, and electronic device to implement audio encoding and decoding at a low bit rate.

[0005] In a first aspect, an embodiment of the present disclosure provides an audio processing method, applied to a first device, the method comprising:

[0006] Acquire audio data to be processed; wherein the audio data to be processed includes at least one audio frame;

[0007] Extracting audio feature information corresponding to each of the audio frames, and performing quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data;

[0008] The encoded audio data is sent to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed.

[0009] In a second aspect, an embodiment of the present disclosure provides an audio processing method, applied to a second device, the method comprising:

[0010] Obtaining encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame;

[0011] The target upsampling network model is used to decode the encoded audio data to obtain the audio data to be processed.

[0012] In a third aspect, an embodiment of the present disclosure provides an audio processing device, applied to a first device, the device comprising:

[0013] A first processing module is configured to obtain audio data to be processed; wherein the audio data to be processed includes at least one audio frame;

[0014] The first processing module is further configured to extract audio feature information corresponding to each of the audio frames, and perform quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data;

[0015] The first transceiver module is configured to send the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed.

[0016] In a fourth aspect, an embodiment of the present disclosure provides an audio processing device, applied to a second device, the device comprising:

[0017] a second transceiver module, configured to obtain the encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame;

[0018] The second processing module is used to decode the encoded audio data using a target upsampling network model to obtain the audio data to be processed.

[0019] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: at least one processor and a memory.

[0020] The memory stores computer-executable instructions.

[0021] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method described in the first aspect and various possible designs of the first aspect.

[0022] In a sixth aspect, an embodiment of the present disclosure provides an electronic device, comprising: at least one processor and a memory.

[0023] The memory stores computer-executable instructions.

[0024] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method described in the second aspect and various possible designs of the second aspect.

[0025] In the seventh aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which stores computer execution instructions. When the processor executes the computer execution instructions, it implements the audio processing method described in the first aspect and various possible designs of the first aspect.

[0026] In an eighth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the audio processing method described in the second aspect and various possible designs of the second aspect is implemented.

[0027] In a ninth aspect, an embodiment of the present disclosure provides a computer program product, comprising a computer program. When the computer program is executed by a processor, the audio processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0028] In a tenth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the audio processing method described in the second aspect and various possible designs of the second aspect.

[0029] The audio processing method, device, and electronic device provided in this embodiment indicate that, upon obtaining audio data to be processed, the audio data to be processed needs to be encoded so that the audio data to be processed can be transmitted to a second device. The method then extracts audio feature information corresponding to each audio frame in the audio data to be processed to achieve preliminary compression of the audio data to be processed. After obtaining the audio feature information corresponding to each audio frame, the audio feature information corresponding to each audio frame is quantized and compressed, i.e., further compressed, to obtain more compressed encoded audio, i.e., encoded audio data. The encoded audio data is then transmitted to the second device, so that the second device uses a target upsampling network to decode the encoded audio data, i.e., restore the encoded and compressed audio data to obtain the corresponding audio data to be processed. Since the audio data to be processed is compressed twice when it is encoded, the encoded audio data is more compressed than the original audio data to be processed. Even if the current bit rate is low, the encoded audio data can be successfully sent to the second device. After receiving the encoded audio data, the second device uses the target upsampling network to decode the encoded audio data, and can quickly restore the audio data with high quality, ensuring the quality of the decoded audio data, thereby effectively realizing encoding and decoding at a low bit rate and improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0031] Figure 1 A schematic diagram of a scenario of the audio processing method provided in an embodiment of the present disclosure;

[0032] Figure 2 Schematic diagram of the process of the audio processing method provided in the embodiment of the present disclosure Figure 1 ;

[0033] Figure 3 Schematic diagram of the process of the audio processing method provided in the embodiment of the present disclosure Figure 2 ;

[0034] Figure 4 Schematic diagram of the process of the audio processing method provided in the embodiment of the present disclosure Figure 3 ;

[0035] Figure 5 Schematic diagram of the process of the audio processing method provided in the embodiment of the present disclosure Figure 4 ;

[0036] Figure 6 A schematic diagram of the structure of an upsampling neural network model provided in an embodiment of the present disclosure;

[0037] Figure 7 Schematic diagram of the structure of the residual network provided by the embodiment of the present disclosure

[0038] Figure 8 The structural framework of the audio processing device provided by the embodiment of the present disclosure Figure 1 ;

[0039] Figure 9 The structural framework of the audio processing device provided by the embodiment of the present disclosure Figure 2 ;

[0040] Figure 10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0042] In the prior art, when encoding and decoding audio data, linear prediction analysis modeling is generally used to encode and decode audio data. However, since the use of linear prediction analysis modeling for encoding and decoding will produce a lot of residual information, it is difficult to achieve encoding and decoding below 6kbps. When achieving encoding and decoding below 6kbps, frequency domain-based methods (such as codec2, melp, etc.) are generally used for encoding and decoding. However, due to the large loss caused by compression, the audio quality obtained by decoding is poor, resulting in sound distortion, loss of details, etc., which makes it difficult for users to understand, affecting the user experience.

[0043] Therefore, to address the above issues, the technical concept of the present invention is to compress audio data using perceptual domain spectrum information of the pitch frequency set during encoding, which facilitates the decoding end to use the compressed audio data to restore authentic, intelligible audio. During decoding, a neural network model based on upsampling is used with low computational complexity to generate audio segments at a time and recover the corresponding audio data, achieving encoding and decoding at low bit rates. This improves the audio quality of encoding and decoding at low bit rates while achieving high-speed processing and enhancing the user experience.

[0044] Figure 1 A schematic diagram of a scenario of an audio processing method provided by an embodiment of the present invention, such as Figure 1 As shown, when the first device 101 needs to transmit the audio data to be processed to the second device 102, the first device 101 encodes the audio data to be processed to obtain encoded audio data that is convenient for transmission, and sends the encoded audio data to the second device 102. After receiving the encoded audio data, the second device 102 decompresses it, that is, decodes it, to restore it to the audio data to be processed.

[0045] The first device 101 may be a computer device (e.g., a desktop computer, a laptop computer, an all-in-one computer, etc.), a mobile terminal (e.g., a mobile phone, a tablet computer, etc.), or the like. The second device 102 may also be a computer device (e.g., a desktop computer, a laptop computer, an all-in-one computer, etc.), a mobile terminal (e.g., a mobile phone, a tablet computer, etc.), or the like.

[0046] Optionally, the first device 101 and the second device 102 may be the same device or different devices, which is not limited herein.

[0047] refer to Figure 2 , Figure 2 Schematic diagram of the audio processing method provided in the embodiment of the present disclosure Figure 1 The method of this embodiment can be applied to Figure 1 On the first device shown, the audio processing method includes:

[0048] S201: Acquire audio data to be processed, wherein the audio data to be processed includes at least one audio frame.

[0049] S202: extracting audio feature information corresponding to each audio frame, and performing quantization compression based on the audio feature information corresponding to each audio frame to obtain encoded audio data.

[0050] In an embodiment of the present disclosure, when a first device needs to transmit audio data to a second device, the first device uses the audio data as unprocessed audio data, where the unprocessed audio data is composed of at least one audio frame. Upon obtaining the unprocessed audio data, feature extraction is performed on the unprocessed audio data to obtain audio feature information corresponding to each audio frame in the unprocessed audio data, thereby achieving preliminary audio compression.

[0051] In order to better perform compression, after obtaining the audio feature information corresponding to the audio frame, the audio feature information corresponding to the audio frame is further compressed by utilizing the correlation between the audio frames, that is, quantization compression is performed to obtain encoded audio data.

[0052] The audio frame represents an audio signal of a certain duration, that is, a waveform of a preset duration (eg, 10 ms).

[0053] Optionally, the audio feature information corresponding to the audio frame includes one or more of cepstrum information corresponding to the audio frame, a fundamental frequency corresponding to the audio frame, and a fundamental frequency mutual correlation value corresponding to the audio frame.

[0054] Cepstral information, or cepstral features, contains audio data, specifically the language information within an audio frame, and is widely used in speech recognition scenarios. Pitch frequency includes speaker-related information within an audio frame. Using cepstral features and pitch frequency can effectively recover the speaker and speech content at the decoder, thereby restoring the true voice.

[0055] Furthermore, the process of extracting the cepstrum information corresponding to the audio frame, namely the cepstrum feature, includes the steps of inputting the audio time domain signal frame splicing, windowing, short-time Fourier transform, calculating the frequency band energy, discrete cosine transform, etc.

[0056] Taking a specific application scenario as an example, when the audio data to be processed is 16KHz and the duration of the audio frame is 10ms, the input of an audio frame is x(n), where n=0...160, that is, data of 160 sampling points. When performing signal splicing, the currently input audio frame is spliced ​​with the audio frame input at the previous moment to obtain x(n), where n=-160...160. When performing windowing, the signal x(n) is multiplied by the window function w(n) to obtain the windowed signal x(n)*w(n), n=-160...160. Among them, the window function can be a function such as a Hanning window. When performing Fourier transform, the windowed signal is Fourier transformed to obtain the signal spectrum X(m)=FFT(x(n)*w(n)), and the spectrum energy X is taken. 2 (m). When calculating the band energy, the spectrum is divided into several bands, and the energy of each band is summed. The spectrum is divided into several bands, and the energy of each band is summed, that is, through The frequency band energy B(l) is obtained. When performing discrete cosine transform, the frequency band energy is logarithmized and discrete cosine transform is performed to obtain the cepstrum feature C(K).

[0057] The cepstrum information includes cepstrum parameters corresponding to multiple dimensions, and the number of dimensions corresponding to the cepstrum information can be determined as needed. For example, if the cepstrum information has 18 dimensions, it includes cepstrum parameters corresponding to 18 dimensions; for another example, if the cepstrum information has 80 dimensions, it includes cepstrum parameters corresponding to 80 dimensions.

[0058] Furthermore, when determining the fundamental frequency corresponding to the audio frame, an autocorrelation algorithm can be used for determination. The specific determination process is similar to the existing fundamental frequency determination process, that is, the autocorrelation result between the audio time domain signal corresponding to the current audio frame and the audio time domain signal corresponding to the audio frame in the previous period of time (for example, the audio time domain signal corresponding to the audio frame at the previous moment) is calculated, and the fundamental frequency corresponding to the current audio frame is obtained through the maximum value of the autocorrelation, that is, the time delay corresponding to the maximum autocorrelation value.

[0059] In addition, when calculating the fundamental frequency corresponding to the audio frame, it is necessary to determine the maximum autocorrelation value, and use the maximum autocorrelation value as the fundamental frequency cross-correlation value corresponding to the audio frame.

[0060] In any embodiment, when extracting audio feature information corresponding to an audio frame, the required audio features to be extracted may be determined based on the bit rate. The specific process is as follows: obtaining a current bit rate. If the current bit rate is greater than a preset bit rate, extracting cepstrum information corresponding to each audio frame. If the current bit rate is less than or equal to the preset bit rate, advancing the cepstrum information and pitch frequency corresponding to each audio frame.

[0061] Specifically, when the current bit rate is greater than the preset bit rate, it indicates that the current bit rate is high. To ensure audio quality, the audio data to be processed can be compressed using only the cepstral features, and the cepstral information corresponding to each audio frame in the audio data to be processed is extracted. When the current bit rate is less than or equal to the preset bit rate, it indicates that the current bit rate is low. To further compress the audio data to be processed, the fundamental frequency and cepstral features can be used to compress the audio data to be processed, and the cepstral information and fundamental frequency corresponding to each audio frame in the audio data to be processed are extracted.

[0062] In addition, when the current bit rate is greater than the preset bit rate, the cepstrum information and pitch frequency corresponding to the audio frame may also be extracted to further compress the audio data to be processed using the cepstrum information and pitch frequency.

[0063] S203: Send the encoded audio data to the second device, so that the second device uses the target upsampling network model to decode the encoded audio data to obtain audio data to be processed.

[0064] In an embodiment of the present disclosure, after the audio data to be processed is compressed, that is, after the encoded audio data is obtained, the encoded audio data is sent to the second device, so that when the second device receives the encoded audio data, it uses the target upsampling network model to decode the encoded audio data, that is, performs restoration processing to obtain the corresponding audio data.

[0065] In the embodiment of the present disclosure, after determining the audio feature information corresponding to the audio frame, in order to meet the needs of low bit rate scenarios, the audio feature information is further compressed by utilizing the correlation between audio frames, that is, quantization compression is performed to achieve maximum compression of the audio data, which is convenient for transmission in low bit rate scenarios.

[0066] As can be seen from the above description, when the audio data to be processed is obtained, it indicates that the audio data to be processed needs to be encoded so that the audio data to be processed can be transmitted to the second device. Then, the audio feature information corresponding to each audio frame in the audio data to be processed is extracted to achieve preliminary compression of the audio data to be processed. After obtaining the audio feature information corresponding to each audio frame, the audio feature information corresponding to each audio frame is quantized and compressed, that is, further compressed to obtain encoded audio data that is more convenient for transmission. The encoded audio data is transmitted to the second device so that the second device uses the target upsampling network to decode the encoded audio data, that is, the encoded and compressed audio data is restored to obtain the corresponding audio data to be processed. Since the audio data to be processed is compressed twice when it is encoded, the encoded audio data is more compressed than the original audio data to be processed. Even if the current bit rate is low, the encoded audio data can be successfully sent to the second device. After receiving the encoded audio data, the second device uses the target upsampling network to decode the encoded audio data, and can quickly restore the audio data with high quality, ensuring the quality of the decoded audio data, thereby effectively realizing encoding and decoding at a low bit rate and improving user satisfaction.

[0067] refer to Figure 3 , Figure 3 Schematic diagram of the audio processing method provided in the embodiment of the present disclosure Figure 2 This embodiment describes in detail the process of further compressing the corresponding audio feature information by utilizing the association between audio frames after obtaining the audio feature information corresponding to the audio frame. The audio processing method includes:

[0068] S301: Acquire audio data to be processed, wherein the audio data to be processed includes at least one audio frame.

[0069] S302: Merge audio feature information corresponding to audio frames to obtain at least one audio package, where the audio package includes audio feature information corresponding to a preset number of audio frames.

[0070] In the disclosed embodiment, the audio feature information corresponding to the preset number of audio frames is merged to obtain a corresponding audio package. The audio package includes the audio feature information corresponding to the preset number of audio frames. For example, if the preset number is 4 and the duration of the audio frame is 10ms, the audio package includes the audio feature information corresponding to 40ms of audio signal.

[0071] S303: For each audio packet, perform quantization compression on the audio packet according to audio feature information corresponding to the audio frames in the audio packet to obtain an encoded audio packet.

[0072] In the embodiment of the present disclosure, after obtaining the audio package, the audio package is further compressed using the correlation between the audio feature information corresponding to the audio frames in the audio package to obtain an encoded audio package, that is, to obtain encoded audio data.

[0073] In an embodiment of the present disclosure, optionally, when quantizing and compressing an audio package based on audio feature information corresponding to an audio frame in the audio package, the difference between cepstral information corresponding to a first audio frame and cepstral information of a second audio frame in the audio package is calculated to obtain a cepstral difference corresponding to the second audio frame. The cepstral information corresponding to the first audio frame is the cepstral information corresponding to an audio frame in the audio package, and the cepstral information corresponding to the second audio frame is the cepstral information corresponding to any audio frame in the audio package other than the cepstral information corresponding to the first audio frame. Vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information, and an encoded audio package is generated based on the quantized cepstral information and the cepstral difference corresponding to the second audio frame.

[0074] Specifically, for each audio packet, cepstral information corresponding to the first audio frame is obtained from the cepstral information corresponding to the audio frames in the audio packet, and the cepstral information corresponding to the remaining audio frames is used as the cepstral information corresponding to the second audio frame. For the cepstral information corresponding to each second audio frame, the cepstral information corresponding to the second audio frame and the cepstral information corresponding to the first audio frame are calculated to obtain a cepstral difference corresponding to the second audio frame. When the second device performs decoding, it can use the cepstral difference corresponding to the second audio frame and the cepstral information corresponding to the first audio frame to restore the cepstral information corresponding to the second audio frame.

[0075] After determining the cepstral information corresponding to the first audio frame in the audio packet, compression is continued, i.e., vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information corresponding to the audio packet. After obtaining the quantized cepstral information corresponding to the audio packet, an encoded audio packet is determined using cepstral differences between the quantized cepstral information corresponding to the audio packet and cepstral differences corresponding to each second audio frame corresponding to the audio packet.

[0076] The audio package includes cepstral information corresponding to a preset number of audio frames, and cepstral information corresponding to one audio frame is arbitrarily selected from the cepstral information corresponding to the preset number of audio frames, and the cepstral information corresponding to the selected audio frame is used as the cepstral information corresponding to the first audio frame, or the cepstral information corresponding to the designated audio frame is used as the cepstral information corresponding to the first audio frame. For example, the preset number is 4, and the cepstral information corresponding to the designated audio frame is the cepstral information corresponding to the first audio frame in the audio package, that is, the cepstral information corresponding to the first audio frame is the cepstral information corresponding to the first audio frame in the audio package, and accordingly, the cepstral information corresponding to the remaining three audio frames are all cepstral information corresponding to the second audio frame.

[0077] When calculating the cepstral information corresponding to the second audio frame and the cepstral information corresponding to the first audio frame, a difference is performed based on the dimensions. That is, for each dimension corresponding to the cepstral information, the difference between the cepstral parameters corresponding to the dimension in the cepstral information corresponding to the second audio frame and the cepstral parameters corresponding to the dimension in the cepstral information corresponding to the first audio frame is calculated to obtain the cepstral difference corresponding to the dimension. For example, the difference between the 0th dimension corresponding to the second audio frame and the 0th dimension corresponding to the first audio frame is calculated to obtain the cepstral difference corresponding to the 0th dimension.

[0078] Further, optionally, when vector quantizing the cepstral information corresponding to the first audio frame, the cepstral parameters corresponding to the target dimension in the cepstral information corresponding to the first audio frame are obtained, and the cepstral parameters corresponding to the target dimension are vector quantized to obtain quantized cepstral information.

[0079] Specifically, the cepstrum parameters corresponding to the target dimension are obtained from the cepstrum information corresponding to the first audio frame in the audio package, and the cepstrum parameters corresponding to the target dimension are vector quantized based on the vector quantization technology to obtain the quantized cepstrum information corresponding to the audio package.

[0080] The target dimensions include one or more dimensions, which can be set by relevant personnel based on actual circumstances. For example, the target dimensions are 17 dimensions excluding the 0th dimension. For another example, when the audio feature information only includes cepstrum information, the target dimensions can be all dimensions, that is, vector quantization is performed on all cepstrum parameters in the cepstrum information.

[0081] Optionally, when the extracted audio feature information also includes the fundamental frequency, the fundamental frequency needs to be compressed during quantization compression. The specific process is as follows:

[0082] Cepstrum parameters other than the cepstrum parameters corresponding to the target dimension in the cepstrum information corresponding to the first audio frame are obtained, and are determined as target cepstrum parameters.

[0083] An average value of pitch frequencies corresponding to audio frames in the audio package is obtained, and is determined as an average pitch frequency corresponding to the audio package.

[0084] The pitch slope corresponding to the audio packet is determined according to the pitch frequency corresponding to the audio frame in the audio packet.

[0085] The encoded audio packet is determined according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0086] Specifically, the average of the pitch frequencies corresponding to all audio frames in the audio package is calculated to obtain the average pitch frequency corresponding to the audio package. Since the pitch slope corresponding to the audio package, that is, the pitch frequency corresponding to the audio frames in the audio package, changes linearly over time, the pitch slope can be determined using the pitch frequencies corresponding to the audio frames in the audio package. That is, a linear fit is performed on the pitch frequencies corresponding to the audio frames in the audio package to obtain the pitch slope corresponding to the audio package. For example, if the coefficient corresponding to the pitch slope is 1, then the pitch slope is 1.

[0087] Optionally, after determining the average pitch frequency, pitch slope, target cepstral parameter, quantized cepstral information corresponding to the audio package, and the cepstral difference corresponding to the second audio frame, the encoded audio package is determined using the average pitch frequency, pitch slope, target cepstral parameter, quantized cepstral information corresponding to the audio package, and the cepstral difference corresponding to the second audio frame. The specific process includes:

[0088] The fundamental frequency cross-correlation value, average fundamental frequency, fundamental pitch slope and target cepstrum parameters are scalar quantized to obtain comprehensive feature information.

[0089] The encoded audio packet is obtained according to the comprehensive feature information, the quantized cepstral information and the cepstral difference corresponding to the second audio frame.

[0090] Specifically, after determining the fundamental frequency mutual correlation value, average fundamental frequency, pitch slope, and target cepstral parameter corresponding to the audio packet, scalar quantization is used to quantize the fundamental frequency mutual correlation value, average fundamental frequency, pitch slope, and target cepstral parameter into a single value, thereby obtaining comprehensive feature information. A compressed packet is generated containing the comprehensive feature information corresponding to the audio packet, the quantized feature information, and the cepstral difference values ​​corresponding to each second audio frame corresponding to the audio packet, thereby obtaining an encoded audio packet.

[0091] In the disclosed embodiment, in order to improve the audio quality obtained by decoding, more important target cepstral parameters can be obtained from the cepstral information corresponding to the audio frame, and scalar quantized with parameters such as the fundamental frequency cross-correlation value, average fundamental frequency, and fundamental frequency slope.

[0092] In addition, it should be noted that other encoding methods can be used when encoding the audio data to be processed. Relevant personnel can use them according to actual needs, and they are not described in detail here. For example, a method of linear prediction coefficient + fundamental frequency is used; another example is a method of fundamental frequency + energy of each harmonic.

[0093] S303: Send the encoded audio data to the second device, so that the second device uses the target upsampling network model to decode the encoded audio data to obtain audio data to be processed.

[0094] In the disclosed embodiment, when compressing audio data, the fundamental frequency is combined with the cepstrum information, i.e., the perceptual domain frequency information, for compression, which is beneficial for the second device, i.e., the decoding end, to restore real and reliable audio and ensure the quality of the audio obtained by decoding, thereby effectively improving the audio quality obtained at a low bit rate.

[0095] refer to Figure 4 , Figure 4 Schematic diagram of the audio processing method provided in the embodiment of the present disclosure Figure 3 The method of this embodiment can be applied in Figure 1 On the second device shown, the audio processing method includes:

[0096] S401: Obtain encoded audio data sent by the first device, where the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame.

[0097] S402: Using a target upsampling network model, the encoded audio data is decoded to obtain audio data to be processed.

[0098] In an embodiment of the present disclosure, after receiving the encoded audio data sent by the first device, the second device uses the trained upsampling network model, i.e., the target upsampling network model, to decode the encoded audio data to obtain corresponding audio data to be processed, i.e., restore it to the corresponding audio data.

[0099] The encoded audio data includes encoded audio packets.

[0100] In the disclosed embodiment, the use of a neural network for decoding can improve the decoded audio quality, and because when using an autoregressive network model for decoding, the output of each audio sampling point requires the information of all previous sampling points, resulting in a slower running speed, that is, a slower decoding speed, which cannot meet the needs of real-time decoding of audio data. Therefore, in order to improve the decoding speed while ensuring the decoded audio quality, a neural network model based on upsampling is used to decode the encoded audio data, which can achieve high-speed processing while improving the decoded audio quality.

[0101] In the embodiment of the present disclosure, when decoding is required, the target upsampling network model is used to decode the encoded audio data, that is, the encoded and compressed audio data is restored with high quality. Since the target upsampling network model does not rely on future information and does not introduce delay when restoring the audio, the decoding efficiency is greatly improved while maintaining the quality of the decoded audio, thereby meeting the decoding requirements at low bit rates.

[0102] refer to Figure 5 , Figure 5 Schematic diagram of the audio processing method provided in the embodiment of the present disclosure Figure 4 This embodiment describes in detail the process of decoding the encoded audio data using the target upsampling network model. The audio processing method includes:

[0103] S501: Obtain encoded audio data sent by the first device, where the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame.

[0104] S502: Perform feature recovery processing on the encoded audio data to obtain audio feature information corresponding to at least one audio frame.

[0105] In an embodiment of the present disclosure, for each encoded audio packet in the encoded audio data, feature recovery processing is performed on the encoded audio packet, that is, the compressed bit stream is restored to obtain audio feature information corresponding to a corresponding number of audio frames, for example, audio feature information corresponding to a preset number of audio frames.

[0106] Optionally, when the extracted audio feature information corresponding to the audio frame includes cepstral information, pitch frequency, and pitch mutual correlation value, codeword search technology is used to decode the quantized cepstral information corresponding to the encoded audio package to obtain cepstral parameters corresponding to the target dimension corresponding to the first audio frame in the audio package. Furthermore, codeword search technology is used to decode the comprehensive feature information corresponding to the audio package to obtain the average pitch frequency, pitch slope, pitch mutual correlation value corresponding to the audio package, and target cepstral parameters corresponding to the first audio frame.

[0107] When determining the cepstral parameters and target cepstral parameters corresponding to the target dimension for the first audio frame, the cepstral parameters are combined to obtain cepstral information corresponding to the first audio frame. For each second audio frame corresponding to the audio package, the cepstral information corresponding to the second audio frame is obtained based on the cepstral difference value corresponding to the second audio frame and the cepstral information corresponding to the first audio frame.

[0108] After determining the pitch slope and average pitch frequency corresponding to the audio packet, during decoding, the pitch frequency corresponding to each audio frame in the audio packet can be determined based on the pitch slope and average pitch frequency, i.e., determined by pitch = t * β * main_pitch, where pitch is the pitch frequency corresponding to the audio frame, β is the pitch slope, and main_pitch is the average pitch frequency. For example, if the duration corresponding to the audio frame is 10ms, then when determining the pitch frequency corresponding to the first audio frame, t = 10ms, and the pitch frequency corresponding to the first audio frame is determined to be 10 * β * main_pitch.

[0109] After obtaining the cepstrum information, fundamental frequency, and fundamental frequency cross-correlation value corresponding to the audio frame, the audio feature information corresponding to the audio frame can be obtained.

[0110] Optionally, when the audio feature information corresponding to the extracted audio frame only includes cepstral information, indicating that the encoded audio package includes quantized cepstral information, the quantized cepstral information is decoded using codeword search technology, that is, restored to the cepstral information corresponding to the first audio frame in the audio package. For each second audio frame corresponding to the audio package, the cepstral information corresponding to the second audio frame is obtained based on the cepstral difference corresponding to the second audio frame and the cepstral information corresponding to the first audio frame. After determining the cepstral information corresponding to the first audio frame and the cepstral information corresponding to the second audio frame in the audio package, the audio feature information corresponding to each audio frame in the audio package is obtained.

[0111] S503: Inputting the audio feature information corresponding to each audio frame into the target upsampling network model respectively, so that the target upsampling network model performs upsampling processing on the audio feature information corresponding to each audio frame to obtain an audio waveform corresponding to each audio frame.

[0112] In an embodiment of the present disclosure, after obtaining the audio feature information corresponding to the audio frame, it is input into the target upsampling network model so that the target upsampling network model performs relevant upsampling processing to obtain the audio waveform corresponding to the audio, that is, the time domain waveform.

[0113] Taking a specific application scenario as an example, the sampling rate is 16KHz, the duration of the audio frame is 10ms, and the audio feature information corresponding to the audio frame includes 18-dimensional cepstrum information, fundamental frequency, and fundamental frequency cross-correlation value. Then, 10ms, that is, the feature information corresponding to the audio frame (20*1, including 18-dimensional cepstrum information, one-dimensional fundamental frequency, and one-dimensional fundamental frequency cross-correlation value) is input into the target upsampling network model, and the target upsampling network model outputs 160 audio sampling points (1*169). Therefore, during the processing of the target upsampling network model, upsampling operations need to be performed in the time dimension, such as Figure 6As shown, the target upsampling network model is divided into three layers of network stacking, and the upsampling multiples of each layer are ×8, ×5, and ×4 respectively. After each upsampling, a residual network (such as Figure 7 The up-sampling result is adjusted to make it closer to the actual audio.

[0114] In an embodiment of the present disclosure, optionally, before using the target upsampling network model for decoding, it is necessary to first train the initial upsampling network model to obtain an upsampling network model that can be accurately decoded, that is, to obtain a target upsampling network model that can be decoded with high quality. The training process is: obtaining audio feature information corresponding to the sample audio, and training the initial upsampling network model according to the audio feature information corresponding to the sample audio and a preset loss function to obtain a target upsampling network model, wherein the preset loss function includes a spectral loss function and / or a discriminator loss function.

[0115] Specifically, using the spectrum as a loss function can make the generated spectrum close to the spectrum corresponding to the actual audio data, and using a discriminator can determine whether the generated audio is real audio. When the initial upsampling network model is trained using a preset loss function and audio feature information corresponding to the sample audio, a corresponding loss value is calculated. When the loss value is greater than the preset loss value, it indicates that the trained upsampling network model does not meet the requirements and further training is required, and the initial upsampling network model is continued to be trained. When the loss value is less than or equal to the preset loss value, it indicates that the trained upsampling network model meets the requirements, and the trained upsampling network model is used as the target network model.

[0116] Among them, when the spectrum or discriminator is used as the loss function, that is, the preset loss function includes the spectrum loss function or the discriminator loss function, correspondingly, the loss value includes the spectrum loss value or the discriminator loss value, then when the spectrum loss value is greater than the first preset loss value or the discriminator loss value is greater than the second preset loss value, continue to train the initial upsampling network model.

[0117] When the spectrum and discriminator are used as the loss function, that is, the preset loss function includes the spectrum loss function and the discriminator loss function, and correspondingly, the loss value includes the spectrum loss value and the discriminator loss value, then when the spectrum loss value is greater than the first preset loss value and the discriminator loss value is greater than the second preset loss value, continue to train the initial upsampling network model.

[0118] In addition, optionally, the upsampling neural network can also output sub-band audio (for example, two sub-bands of high frequency band and low frequency band), which are then combined into complete audio through filters. For example, when the duration of the audio frame is 10ms and the sampling rate is 16KHz, when the upsampling neural network outputs two sub-bands, the output is 2 segments of 80 audio sampling points (2*80), and the sub-band filter is used to synthesize the output audio into 160 audio sampling points.

[0119] In the embodiment of the present disclosure, when the spectrum and the discriminator are used as loss functions at the same time, the generated audio is closer to the real audio, and the generated audio can be prevented from having noise, so that the generated audio is similar to the original audio and smoother, thereby improving user satisfaction.

[0120] S504: Obtain the audio waveform corresponding to each audio frame output by the target upsampling network model, and determine it as the audio data to be processed.

[0121] In an embodiment of the present disclosure, after determining each audio frame in the audio data to be processed, that is, the audio waveform corresponding to each audio frame, using the target upsampling network model, the audio waveforms corresponding to each audio frame are combined to obtain the audio data to be processed.

[0122] In any embodiment, optionally, harmonic enhancement processing is performed on the audio data to be processed.

[0123] Specifically, in order to avoid the decoded audio data, that is, the decoded audio data to be processed, from causing damage to the hearing sense, the decoded audio data to be processed can be harmonically enhanced, that is, the fundamental frequency and its multiple frequency energy are enhanced through a filter. The time domain form of the filter is: y(n) = αx(n) + (1-α)y(nT), where y is the enhanced audio signal, α is the preset enhancement coefficient, and T is the frequency point interval corresponding to the fundamental frequency, that is, the interval time between sampling points.

[0124] The value range of α is 0 to 1.

[0125] In the disclosed embodiments, the upsampling network model takes audio feature information as input and outputs the time-domain waveform of an audio frame, saving bandwidth otherwise occupied by residual or phase information required in traditional codecs. Furthermore, the upsampling network model is highly efficient, using audio frames as the generation unit, and can run in real time.

[0126] In an embodiment of the present disclosure, when encoding the audio data to be processed, the audio feature information corresponding to each audio frame in the audio data to be processed is extracted, and the audio feature information is quantized and compressed to obtain the corresponding encoded audio data. When decoding the encoded audio data, the feature information is recovered to obtain the audio feature information corresponding to each audio frame, and the audio feature information corresponding to each audio frame is input into a neural network model based on upsampling, that is, a target upsampling network model, to obtain the waveform corresponding to each audio frame, thereby realizing low-latency audio encoding and decoding at a low bit rate. When the user network bandwidth is insufficient, only a very small amount of traffic is required to achieve a smooth voice call, and when the user network is stuck, redundant information with low bandwidth can be sent to perform audio recovery for packet loss scenarios, thereby improving the user experience.

[0127] In the disclosed embodiment, since the input of the model is audio features and the output is an audio waveform, when training the upsampling network model, due to the difficulty in convergence, the real audio waveform is not used as the objective function. Instead, the frequency and / or discriminator is used as the loss function, so that the audio waveform output by the upsampling network model is close to the original audio waveform, thereby ensuring the quality of the decoded audio.

[0128] Corresponding to the above Figures 2 to 3 The audio processing method of the embodiment, Figure 8 The structural framework of the audio processing device provided by the embodiment of the present disclosure Figure 1 , the audio processing device is applied to the first device. For the convenience of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 8 The audio processing device 80 includes: a first processing module 801 and a first transceiver module 802.

[0129] The first processing module is configured to obtain audio data to be processed; wherein the audio data to be processed includes at least one audio frame;

[0130] The first processing module is further configured to extract audio feature information corresponding to each of the audio frames, and perform quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data;

[0131] The first transceiver module is configured to send the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed.

[0132] In one embodiment of the present disclosure, the first processing module is further configured to:

[0133] Merging the audio feature information corresponding to the audio frames to obtain at least one audio package, wherein the audio package includes audio feature information corresponding to a preset number of audio frames;

[0134] For each audio packet, the audio packet is quantized and compressed according to audio feature information corresponding to audio frames in the audio packet to obtain an encoded audio packet.

[0135] In one embodiment of the present disclosure, the audio feature information corresponding to the audio frame includes cepstrum information corresponding to the audio frame;

[0136] The first processing module is further configured to:

[0137] Calculating a difference between cepstral information corresponding to a first audio frame and cepstral information of a second audio frame in the audio package to obtain a cepstral difference corresponding to the second audio frame; wherein the cepstral information corresponding to the first audio frame is cepstral information corresponding to an audio frame in the audio package, and the cepstral information corresponding to the second audio frame is cepstral information corresponding to any audio frame in the audio package except the cepstral information corresponding to the first audio frame;

[0138] Vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information, and an encoded audio packet is generated according to a difference between the quantized cepstral information and the cepstral information corresponding to the second audio frame.

[0139] In one embodiment of the present disclosure, the first processing module is further configured to:

[0140] Obtain cepstral parameters corresponding to a target dimension in the cepstral information corresponding to the first audio frame, and perform vector quantization on the cepstral parameters corresponding to the target dimension to obtain quantized cepstral information.

[0141] In one embodiment of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency corresponding to the audio frame;

[0142] The first processing module is further configured to:

[0143] Obtaining cepstral parameters other than the cepstral parameters corresponding to the target dimension from the cepstral information corresponding to the first audio frame, and determining the cepstral parameters as target cepstral parameters;

[0144] Obtaining an average value of pitch frequencies corresponding to audio frames in the audio package, and determining the average value as the average pitch frequency corresponding to the audio package;

[0145] Determining a pitch slope corresponding to the audio package according to a pitch frequency corresponding to an audio frame in the audio package;

[0146] The encoded audio packet is determined according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0147] In one embodiment of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency mutual correlation value corresponding to the audio frame;

[0148] The first processing module is further configured to:

[0149] Performing scalar quantization on the fundamental frequency mutual correlation value, the average fundamental frequency, the fundamental pitch slope, and the target cepstrum parameter to obtain comprehensive feature information;

[0150] The encoded audio packet is obtained according to the comprehensive feature information, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0151] In one embodiment of the present disclosure, the first processing module is further configured to:

[0152] Get the current bit rate;

[0153] If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames;

[0154] If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the pitch frequency corresponding to each audio frame are advanced.

[0155] The device provided in this embodiment can be used to perform the above Figure 2 and Figure 3 The technical solution of the method embodiment has similar implementation principles and technical effects, and will not be repeated here in this embodiment.

[0156] Corresponding to the above Figures 4 and 5 The audio processing method of the embodiment, Figure 9 The structural framework of the audio processing device provided by the embodiment of the present disclosure Figure 2 , the audio processing device is applied to the second device. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 10 The audio processing device 100 includes: a second transceiver module 901 and a second processing module 902 .

[0157] The second transceiver module is configured to obtain the encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame;

[0158] The second processing module is used to decode the encoded audio data using a target upsampling network model to obtain the audio data to be processed.

[0159] In one embodiment of the present disclosure, the second processing module is further configured to:

[0160] Performing feature recovery processing on the encoded audio data to obtain audio feature information corresponding to at least one audio frame;

[0161] Inputting the audio feature information corresponding to each audio frame into the target upsampling network model respectively, so that the target upsampling network model performs upsampling processing on the audio feature information corresponding to each audio frame to obtain an audio waveform corresponding to each audio frame;

[0162] The audio waveform corresponding to each audio frame output by the target upsampling network model is obtained and determined as the audio data to be processed.

[0163] In one embodiment of the present disclosure, the second processing module is further configured to:

[0164] Obtain audio feature information corresponding to the sample audio, and train an initial upsampling network model based on the audio feature information corresponding to the sample audio and a preset loss function to obtain the target upsampling network model, wherein the preset loss function includes a spectral loss function and / or a discriminator loss function.

[0165] In one embodiment of the present disclosure, the second processing module is further configured to:

[0166] Perform harmonic enhancement processing on the audio data to be processed.

[0167] The device provided in this embodiment can be used to perform the above Figure 4 and Figure 5 The technical solution of the method embodiment has similar implementation principles and technical effects, and will not be repeated here in this embodiment.

[0168] refer to Figure 10, which shows a schematic structural diagram of an electronic device 1000 suitable for implementing the embodiments of the present disclosure. The electronic device 1000 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0169] like Figure 10 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1009 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0170] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1009 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0171] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1009, or installed from the storage device 1009, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0172] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0173] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0174] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0175] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the audio processing method described above is implemented.

[0176] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0177] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0178] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0179] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0180] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an audio processing method, applied to a first device, the method comprising:

[0182] Acquire audio data to be processed; wherein the audio data to be processed includes at least one audio frame;

[0183] Extracting audio feature information corresponding to each of the audio frames, and performing quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data;

[0184] The encoded audio data is sent to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed.

[0185] According to one or more embodiments of the present disclosure, the step of performing quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data includes:

[0186] Merging the audio feature information corresponding to the audio frames to obtain at least one audio package, wherein the audio package includes audio feature information corresponding to a preset number of audio frames;

[0187] For each audio packet, the audio packet is quantized and compressed according to audio feature information corresponding to audio frames in the audio packet to obtain an encoded audio packet.

[0188] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame includes cepstrum information corresponding to the audio frame;

[0189] The step of quantizing and compressing the audio packet according to the audio feature information corresponding to the audio frame in the audio packet to obtain an encoded audio packet includes:

[0190] Calculating a difference between cepstral information corresponding to a first audio frame and cepstral information of a second audio frame in the audio package to obtain a cepstral difference corresponding to the second audio frame; wherein the cepstral information corresponding to the first audio frame is cepstral information corresponding to an audio frame in the audio package, and the cepstral information corresponding to the second audio frame is cepstral information corresponding to any audio frame in the audio package except the cepstral information corresponding to the first audio frame;

[0191] Vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information, and an encoded audio packet is generated according to a difference between the quantized cepstral information and the cepstral information corresponding to the second audio frame.

[0192] According to one or more embodiments of the present disclosure, performing vector quantization on the cepstrum information corresponding to the first audio frame to obtain quantized cepstrum information includes:

[0193] Obtain cepstrum parameters corresponding to a target dimension in the cepstrum information corresponding to the first audio frame, and perform vector quantization on the cepstrum parameters corresponding to the target dimension to obtain quantized cepstrum information.

[0194] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency corresponding to the audio frame;

[0195] The generating an encoded audio packet according to the quantized cepstral information and the cepstral difference corresponding to the second audio frame includes:

[0196] Obtaining cepstral parameters other than the cepstral parameters corresponding to the target dimension from the cepstral information corresponding to the first audio frame, and determining the cepstral parameters as target cepstral parameters;

[0197] Obtaining an average value of pitch frequencies corresponding to audio frames in the audio package, and determining the average value as the average pitch frequency corresponding to the audio package;

[0198] Determining a pitch slope corresponding to the audio package according to a pitch frequency corresponding to an audio frame in the audio package;

[0199] The encoded audio packet is determined according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0200] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency mutual correlation value corresponding to the audio frame;

[0201] The determining the encoded audio packet according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame includes:

[0202] Performing scalar quantization on the fundamental frequency mutual correlation value, the average fundamental frequency, the fundamental pitch slope, and the target cepstrum parameter to obtain comprehensive feature information;

[0203] The encoded audio packet is obtained according to the comprehensive feature information, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0204] According to one or more embodiments of the present disclosure, extracting audio feature information corresponding to each of the audio frames includes:

[0205] Get the current bit rate;

[0206] If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames;

[0207] If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the pitch frequency corresponding to each audio frame are advanced.

[0208] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an audio processing method, applied to a second device, the method comprising:

[0209] Obtaining encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame;

[0210] The target upsampling network model is used to decode the encoded audio data to obtain the audio data to be processed.

[0211] According to one or more embodiments of the present disclosure, the adopting a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed includes:

[0212] Performing feature recovery processing on the encoded audio data to obtain audio feature information corresponding to at least one audio frame;

[0213] Inputting the audio feature information corresponding to each audio frame into the target upsampling network model respectively, so that the target upsampling network model performs upsampling processing on the audio feature information corresponding to each audio frame to obtain an audio waveform corresponding to each audio frame;

[0214] The audio waveform corresponding to each audio frame output by the target upsampling network model is obtained and determined as the audio data to be processed.

[0215] According to one or more embodiments of the present disclosure, the method further includes:

[0216] Obtain audio feature information corresponding to the sample audio, and train an initial upsampling network model based on the audio feature information corresponding to the sample audio and a preset loss function to obtain the target upsampling network model, wherein the preset loss function includes a spectral loss function and / or a discriminator loss function.

[0217] According to one or more embodiments of the present disclosure, the method further includes:

[0218] Perform harmonic enhancement processing on the audio data to be processed.

[0219] In a third aspect, according to one or more embodiments of the present disclosure, an audio processing device is provided, applied to a first device, the audio processing device comprising:

[0220] A first processing module is configured to obtain audio data to be processed; wherein the audio data to be processed includes at least one audio frame;

[0221] The first processing module is further configured to extract audio feature information corresponding to each of the audio frames, and perform quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data;

[0222] The first transceiver module is configured to send the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed.

[0223] According to one or more embodiments of the present disclosure, the first processing module is further configured to:

[0224] Merging the audio feature information corresponding to the audio frames to obtain at least one audio package, wherein the audio package includes audio feature information corresponding to a preset number of audio frames;

[0225] For each audio packet, the audio packet is quantized and compressed according to audio feature information corresponding to audio frames in the audio packet to obtain an encoded audio packet.

[0226] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame includes cepstrum information corresponding to the audio frame;

[0227] The first processing module is further configured to:

[0228] Calculating a difference between cepstral information corresponding to a first audio frame and cepstral information of a second audio frame in the audio package to obtain a cepstral difference corresponding to the second audio frame; wherein the cepstral information corresponding to the first audio frame is cepstral information corresponding to an audio frame in the audio package, and the cepstral information corresponding to the second audio frame is cepstral information corresponding to any audio frame in the audio package except the cepstral information corresponding to the first audio frame;

[0229] Vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information, and an encoded audio packet is generated according to a difference between the quantized cepstral information and the cepstral information corresponding to the second audio frame.

[0230] According to one or more embodiments of the present disclosure, the first processing module is further configured to:

[0231] Obtain cepstrum parameters corresponding to a target dimension in the cepstrum information corresponding to the first audio frame, and perform vector quantization on the cepstrum parameters corresponding to the target dimension to obtain quantized cepstrum information.

[0232] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency corresponding to the audio frame;

[0233] The first processing module is further configured to:

[0234] Obtaining cepstral parameters other than the cepstral parameters corresponding to the target dimension from the cepstral information corresponding to the first audio frame, and determining the cepstral parameters as target cepstral parameters;

[0235] Obtaining an average value of pitch frequencies corresponding to audio frames in the audio package, and determining the average value as the average pitch frequency corresponding to the audio package;

[0236] Determining a pitch slope corresponding to the audio package according to a pitch frequency corresponding to an audio frame in the audio package;

[0237] The encoded audio packet is determined according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0238] According to one or more embodiments of the present disclosure, the audio feature information corresponding to the audio frame further includes a fundamental frequency mutual correlation value corresponding to the audio frame;

[0239] The first processing module is further configured to:

[0240] Performing scalar quantization on the fundamental frequency mutual correlation value, the average fundamental frequency, the fundamental pitch slope, and the target cepstrum parameter to obtain comprehensive feature information;

[0241] The encoded audio packet is obtained according to the comprehensive feature information, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

[0242] According to one or more embodiments of the present disclosure, the first processing module is further configured to:

[0243] Get the current bit rate;

[0244] If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames;

[0245] If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the pitch frequency corresponding to each audio frame are advanced.

[0246] In a fourth aspect, according to one or more embodiments of the present disclosure, an audio processing device is provided, applied to a second device, the audio processing device comprising:

[0247] a second transceiver module, configured to obtain the encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame;

[0248] The second processing module is used to decode the encoded audio data using a target upsampling network model to obtain the audio data to be processed.

[0249] According to one or more embodiments of the present disclosure, the second processing module is further configured to:

[0250] Performing feature recovery processing on the encoded audio data to obtain audio feature information corresponding to at least one audio frame;

[0251] Inputting the audio feature information corresponding to each audio frame into the target upsampling network model respectively, so that the target upsampling network model performs upsampling processing on the audio feature information corresponding to each audio frame to obtain an audio waveform corresponding to each audio frame;

[0252] The audio waveform corresponding to each audio frame output by the target upsampling network model is obtained and determined as the audio data to be processed.

[0253] According to one or more embodiments of the present disclosure, the second processing module is further configured to:

[0254] Obtain audio feature information corresponding to the sample audio, and train an initial upsampling network model based on the audio feature information corresponding to the sample audio and a preset loss function to obtain the target upsampling network model, wherein the preset loss function includes a spectral loss function and / or a discriminator loss function.

[0255] In one embodiment of the present disclosure, the second processing module is further configured to:

[0256] Perform harmonic enhancement processing on the audio data to be processed.

[0257] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0258] The memory stores computer-executable instructions;

[0259] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method described in the first aspect and various possible designs of the first aspect.

[0260] In a sixth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0261] The memory stores computer-executable instructions;

[0262] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method described in the second aspect and various possible designs of the second aspect.

[0263] In the seventh aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the audio processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0264] In an eighth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer execution instructions. When the processor executes the computer execution instructions, the audio processing method described in the second aspect and various possible designs of the second aspect is implemented.

[0265] In a ninth aspect, an embodiment of the present disclosure provides a computer program product, comprising a computer program. When the computer program is executed by a processor, the audio processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0266] In a tenth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the audio processing method described in the second aspect and various possible designs of the second aspect.

[0267] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0268] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0269] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An audio processing method, characterized in that: include: Acquire audio data to be processed; wherein the audio data to be processed includes at least one audio frame; Extracting audio feature information corresponding to each of the audio frames, and performing quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data; Sending the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed; The extracting audio feature information corresponding to each of the audio frames includes: Get the current bit rate; If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames; If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the fundamental frequency corresponding to each audio frame are extracted.

2. The method according to claim 1, characterized in that The quantizing and compressing the audio feature information corresponding to each of the audio frames to obtain encoded audio data includes: Merging the audio feature information corresponding to the audio frames to obtain at least one audio package, wherein the audio package includes audio feature information corresponding to a preset number of audio frames; For each audio packet, the audio packet is quantized and compressed according to audio feature information corresponding to audio frames in the audio packet to obtain an encoded audio packet.

3. The method according to claim 2, characterized in that The audio feature information corresponding to the audio frame includes cepstrum information corresponding to the audio frame; The step of quantizing and compressing the audio packet according to the audio feature information corresponding to the audio frame in the audio packet to obtain an encoded audio packet includes: Calculating a difference between cepstral information corresponding to a first audio frame and cepstral information of a second audio frame in the audio package to obtain a cepstral difference corresponding to the second audio frame; wherein the cepstral information corresponding to the first audio frame is cepstral information corresponding to an audio frame in the audio package, and the cepstral information corresponding to the second audio frame is cepstral information corresponding to any audio frame in the audio package except the cepstral information corresponding to the first audio frame; Vector quantization is performed on the cepstral information corresponding to the first audio frame to obtain quantized cepstral information, and an encoded audio packet is generated according to a difference between the quantized cepstral information and the cepstral information corresponding to the second audio frame.

4. The method according to claim 3, characterized in that The performing vector quantization on the cepstrum information corresponding to the first audio frame to obtain quantized cepstrum information includes: Obtain cepstral parameters corresponding to a target dimension in the cepstral information corresponding to the first audio frame, and perform vector quantization on the cepstral parameters corresponding to the target dimension to obtain quantized cepstral information.

5. The method according to claim 4, characterized in that The audio feature information corresponding to the audio frame also includes the fundamental frequency corresponding to the audio frame; The generating an encoded audio packet according to the quantized cepstral information and the cepstral difference corresponding to the second audio frame includes: Obtaining cepstral parameters other than the cepstral parameters corresponding to the target dimension from the cepstral information corresponding to the first audio frame, and determining the cepstral parameters as target cepstral parameters; Obtaining an average value of pitch frequencies corresponding to audio frames in the audio package, and determining the average value as the average pitch frequency corresponding to the audio package; Determining a pitch slope corresponding to the audio package according to a pitch frequency corresponding to an audio frame in the audio package; The encoded audio packet is determined according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

6. The method according to claim 5, characterized in that The audio feature information corresponding to the audio frame also includes a fundamental frequency mutual correlation value corresponding to the audio frame; The determining the encoded audio packet according to the average pitch frequency, the pitch slope, the target cepstral parameter, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame includes: Performing scalar quantization on the fundamental frequency mutual correlation value, the average fundamental frequency, the fundamental pitch slope, and the target cepstrum parameter to obtain comprehensive feature information; The encoded audio packet is obtained according to the comprehensive feature information, the quantized cepstral information, and the cepstral difference corresponding to the second audio frame.

7. An audio processing method, characterized in that: include: Obtaining encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame; Using a target upsampling network model, decoding the encoded audio data to obtain the audio data to be processed; The extracting audio feature information corresponding to each audio frame in the audio data to be processed includes: Get the current bit rate; If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames; If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the fundamental frequency corresponding to each audio frame are extracted.

8. The method according to claim 7, characterized in that The adopting the target upsampling network model to decode the encoded audio data to obtain the audio data to be processed includes: Performing feature recovery processing on the encoded audio data to obtain audio feature information corresponding to at least one audio frame; Inputting the audio feature information corresponding to each audio frame into the target upsampling network model respectively, so that the target upsampling network model performs upsampling processing on the audio feature information corresponding to each audio frame to obtain an audio waveform corresponding to each audio frame; The audio waveform corresponding to each audio frame output by the target upsampling network model is obtained and determined as the audio data to be processed.

9. The method according to claim 7, characterized in that The method further comprises: Obtain audio feature information corresponding to the sample audio, and train an initial upsampling network model based on the audio feature information corresponding to the sample audio and a preset loss function to obtain the target upsampling network model, wherein the preset loss function includes a spectral loss function and / or a discriminator loss function.

10. The method according to any one of claims 7 to 9, characterized in that The method further comprises: Perform harmonic enhancement processing on the audio data to be processed.

11. An audio processing device, characterized in that: include: A first processing module is configured to obtain audio data to be processed; wherein the audio data to be processed includes at least one audio frame; The first processing module is further configured to extract audio feature information corresponding to each of the audio frames, and perform quantization compression based on the audio feature information corresponding to each of the audio frames to obtain encoded audio data; A first transceiver module is configured to send the encoded audio data to a second device, so that the second device uses a target upsampling network model to decode the encoded audio data to obtain the audio data to be processed; The first processing module is specifically configured to: Get the current bit rate; If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames; If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the fundamental frequency corresponding to each audio frame are extracted.

12. An audio processing device, characterized in that: include: a second transceiver module, configured to obtain the encoded audio data sent by the first device, wherein the encoded audio data is obtained by the first device extracting audio feature information corresponding to each audio frame in the audio data to be processed and performing quantization compression based on the audio feature information corresponding to each audio frame; A second processing module is configured to decode the encoded audio data using a target upsampling network model to obtain the audio data to be processed; The extracting audio feature information corresponding to each audio frame in the audio data to be processed includes: Get the current bit rate; If the current bit rate is greater than the preset bit rate, extracting cepstrum information corresponding to each of the audio frames; If the current bit rate is less than or equal to the preset bit rate, the cepstrum information and the fundamental frequency corresponding to each audio frame are extracted.

13. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method according to any one of claims 1 to 6.

14. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the audio processing method according to any one of claims 7 to 10.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the audio processing method according to any one of claims 1 to 6 is implemented.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the audio processing method according to any one of claims 7 to 10 is implemented.

17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the audio processing method according to any one of claims 1 to 6 is implemented.

18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the audio processing method according to any one of claims 7 to 10 is implemented.

Citation Information

Patent Citations

  • Systems and methods for energy efficient and low power distributed automatic speech recognition on wearable devices

    CN108694938A

  • Voice generation method and device

    CN110930976A