Audio processing method and apparatus, audio processing model training method and apparatus, and device
By introducing a long short-time quantization network into the audio processing model, the problem of low processing efficiency caused by traditional quantizers using the same bitrate for each audio frame is solved, achieving more efficient audio data processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2026-04-02
AI Technical Summary
In traditional end-to-end neural network codecs, the quantizer quantizes each audio frame at the same bit rate, resulting in low efficiency in audio data processing.
Long-short-time quantization (LSQ) networks are used to perform LQ and SQ processing on audio coding features. By combining LQ and SQ features, the bit rate of audio quantization features is reduced, thereby improving coding efficiency and transmission efficiency.
It improves the encoding and transmission efficiency of the audio processing model and enhances the processing effect of audio data.
Smart Images

Figure CN2025102954_02042026_PF_FP_ABST
Abstract
Description
Audio processing methods, training methods for audio processing models, devices and equipment
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411365294.8, filed on September 27, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to an audio processing method, an audio processing model training method, an apparatus, and a device. Background Technology
[0004] Traditional end-to-end neural network codecs typically consist of encoders, quantizers, and decoders, and can be used to compress audio data into vector format and then reconstruct the audio data through decoding. However, traditional quantizers use the same bitrate for each audio frame, resulting in low processing efficiency for audio data. Summary of the Invention
[0005] This disclosure provides an audio processing method, an audio processing model training method, an apparatus, and a device to improve the processing efficiency of audio data.
[0006] In a first aspect, embodiments of this disclosure provide an audio processing method, the method comprising:
[0007] The audio input to be processed is encoded by the encoder in the pre-trained audio processing model to obtain audio encoding features;
[0008] The audio coding features are input into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0009] The audio quantization features are input into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features.
[0010] The audio dequantization features are input into the decoder in the audio processing model for decoding, and the target reconstructed audio is output.
[0011] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed.
[0012] Secondly, embodiments of this disclosure also provide a method for training an audio processing model, the method comprising:
[0013] The training audio is input into the encoder of the untrained audio processing model to encode the audio and obtain audio coding features.
[0014] The audio coding features are input into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0015] The audio quantization features are input into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features.
[0016] The audio dequantization features are input into the decoder in the audio processing model for decoding, and the predicted reconstructed audio is output.
[0017] Based on the training audio and the predicted reconstructed audio, a target loss value is determined, and based on the target loss value, the untrained audio processing model is iteratively trained to obtain a trained audio processing model.
[0018] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the training audio.
[0019] Thirdly, embodiments of this disclosure also provide an audio processing apparatus, the apparatus comprising:
[0020] The audio input module is used to encode the pre-trained audio processing model's encoder to obtain audio encoding features.
[0021] The audio quantization feature determination module is used to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0022] An audio dequantization feature determination module is used to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features;
[0023] The target reconstructed audio determination module is used to input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the target reconstructed audio.
[0024] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed.
[0025] Fourthly, embodiments of this disclosure also provide a training apparatus for an audio processing model, the apparatus comprising:
[0026] The training audio input module is used to encode the encoder in the untrained audio processing model by training the audio input to obtain audio coding features.
[0027] The audio quantization feature determination module is used to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0028] An audio dequantization feature determination module is used to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features;
[0029] The predictive reconstructed audio determination module is used to input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the predicted reconstructed audio.
[0030] An audio processing model training module is used to determine a target loss value based on the training audio and the predicted reconstructed audio, and to iteratively train the untrained audio processing model based on the target loss value to obtain a trained audio processing model.
[0031] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the training audio.
[0032] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0033] At least one processor; and
[0034] A memory communicatively connected to the at least one processor; wherein,
[0035] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio processing method and / or the training method of the audio processing model described in any embodiment of the present disclosure.
[0036] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute and implement the audio processing method and / or the training method of the audio processing model described in any embodiment of this disclosure.
[0037] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the audio processing method and / or the training method of the audio processing model described in any embodiment of this disclosure. Attached Figure Description
[0038] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0039] Figure 1 is a schematic diagram of the framework of an end-to-end neural network encoder and decoder;
[0040] Figure 2 is a schematic diagram of frame quantization in an end-to-end neural network codec;
[0041] Figure 3 is a flowchart of an audio processing method provided in an embodiment of this disclosure;
[0042] Figure 4 is a schematic diagram of the architecture of a long short-time quantization network provided in an embodiment of this disclosure;
[0043] Figure 5 is a schematic diagram of the architecture of a long short-time dequantization network provided in an embodiment of this disclosure;
[0044] Figure 6 is a schematic diagram showing the comparison between an audio coding feature and an STFT spectrum provided in an embodiment of this disclosure;
[0045] Figure 7 is a flowchart of another audio processing method provided in an embodiment of this disclosure;
[0046] Figure 8 is a schematic diagram of another long-short-time quantization network architecture provided in an embodiment of this disclosure;
[0047] Figure 9 is a schematic diagram of the combined architecture of a long-term feature extraction module, a short-term feature extraction module, and a feature fusion module provided in an embodiment of this disclosure;
[0048] Figure 10 is a schematic diagram of another combined architecture of a long-term feature extraction module, a short-term feature extraction module, and a feature fusion module provided in an embodiment of the present disclosure;
[0049] Figure 11 is a schematic diagram of the structure of a decoder provided in an embodiment of this disclosure;
[0050] Figure 12 is a schematic diagram of another decoder provided in an embodiment of this disclosure;
[0051] Figure 13 is a flowchart of a training method for an audio processing model provided in an embodiment of this disclosure;
[0052] Figure 14 is a flowchart of another training method for an audio processing model provided in an embodiment of this disclosure;
[0053] Figure 15 is a schematic diagram of the architecture of a discriminator provided in an embodiment of this disclosure;
[0054] Figure 16 is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present disclosure;
[0055] Figure 17 is a schematic diagram of the structure of a training device for an audio processing model provided in an embodiment of this disclosure;
[0056] Figure 18 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0057] Figure 1 is a schematic diagram of an end-to-end neural network codec framework. Specifically, the raw audio passes through the encoding end, where AI (Artificial Intelligence) technology is used to encode the features of the raw audio, and the codebook in VQ (Vector Quantization) technology is used for quantization processing to output the bitstream. After transmission through the channel, the decoding end receives the bitstream, uses the codebook in VQ technology for dequantization processing, and uses AI technology for decoding and reconstruction to output the reconstructed audio.
[0058] In traditional end-to-end neural network codecs, the encoder encodes multiple audio frames divided with a fixed encoding frame length, and the quantizer quantizes and dequantizes the encoding result of each audio frame using the same bit rate.
[0059] Figure 2 illustrates the frame-by-frame quantization of an end-to-end neural network codec. Specifically, taking an input audio sampling rate of 16kHz as an example, and assuming the encoder's frame length is 20ms, each audio frame contains 320 sampling points with a dimension of 1×320. After encoding each audio frame, audio coding features are obtained, with the feature dimension of the encoding result corresponding to each audio frame being 512×1. The quantization network at the encoding end of the quantizer quantizes the encoding result of each audio frame in the audio coding features using a bitrate of 6kbps, obtaining audio quantized features. These audio quantized features are then transmitted to the dequantization network at the decoding end of the quantizer (not shown in Figure 2). The audio quantized features sequentially pass through the dequantization network and the decoder to obtain the output reconstructed audio.
[0060] Because traditional quantizers use the same bitrate to quantize each audio frame, the quantization results corresponding to multiple audio frames in the audio quantization feature contain repeated quantization information, which leads to poor coding efficiency and transmission efficiency of the audio processing model, and consequently, low processing efficiency of audio data.
[0061] To address the aforementioned technical problems, embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0062] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0063] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0064] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0065] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0066] Figure 3 is a flowchart of an audio processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations involving audio reconstruction processing. The method can be executed by an audio processing device, which can be implemented in hardware and / or software. This audio processing device can be configured in an electronic device, typically a mobile terminal or tablet computer. As shown in Figure 3, the method includes:
[0067] S110. Encode the encoder in the pre-trained audio processing model to obtain audio coding features from the input audio to be processed.
[0068] For example, audio can be categorized by content, such as music, waveform audio, or speech; and by format, such as WAV, MP3, WMA, or FLAC. There are no restrictions on the audio content or format here; specific settings can be customized according to actual needs.
[0069] In this embodiment, the audio processing model includes a cascaded encoder, quantizer, and decoder.
[0070] The encoder is used to compress the audio to be processed in the temporal dimension and expand it to the feature dimension. For example, the encoder includes, but is not limited to, encoders based on fully convolutional networks, encoders based on recurrent neuron layers, encoders based on Opus, or encoders based on SoundStream, etc. The encoder used is not limited here, and can be customized according to actual needs.
[0071] S120. Input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain audio quantization features.
[0072] In this embodiment, the quantizer includes a long short-time quantization network disposed at the encoding end and a long short-time dequantization network disposed at the decoding end. The long short-time quantization network is used to perform long-time dimension compression quantization and short-time dimension compression quantization on the audio encoded features according to the codebook, and transmits the resulting audio quantized features to the long short-time dequantization network at a specified bit rate.
[0073] In this embodiment, the audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames in the long-time quantization features is less than the number of feature frames in the short-time quantization features. The long-time quantization features characterize the quantization results of the common features corresponding to the audio segments obtained from the audio segmentation to be processed. The number of audio segments can be one or more, and each audio segment corresponds to at least two audio frames.
[0074] In an optional embodiment, the long-short-time quantization network includes a long-time feature extraction module and a long-short-time quantization layer. Accordingly, the step of inputting the audio encoded features into the long-short-time quantization network of the quantizer in the audio processing model for long-short-time quantization processing to obtain audio quantization features includes: inputting the audio encoded features into the long-time feature extraction module for long-time feature extraction to obtain audio long-time features; and inputting the audio encoded features and the audio long-time features into the long-short-time quantization layer for long-short-time quantization processing to obtain audio quantization features.
[0075] Specifically, the audio long-term features belong to the multi-frame level features corresponding to the audio coding features, and the audio long-term features represent the common features corresponding to the audio segments obtained by dividing the audio to be processed.
[0076] In an optional embodiment, the step of inputting the audio coding features into the long-term feature extraction module for long-term feature extraction to obtain audio long-term features includes: performing frame-by-frame processing on the audio coding features according to a preset quantization step size to obtain at least one frame-by-frame coding feature; performing long-term feature extraction on each frame-by-frame coding feature to obtain frame-by-frame long-term features; and determining the audio long-term features based on at least one frame-by-frame long-term feature.
[0077] For example, an audio coding feature X is represented as [T,D], where T represents the number of feature frames and D represents the feature dimension. Correspondingly, multiple frame-based coding features can be represented as [T / S,S,D], and the number of features when the frame length is T / S.
[0078] In one optional embodiment, the preset quantization step size is less than the number of feature frames of the audio coding feature. If the preset quantization step size is equal to the number of feature frames of the audio coding feature, it indicates that the audio long-time feature is a time-invariant feature, representing the common features corresponding to the entire audio to be processed, but the time-invariant feature cannot meet the latency requirements in the audio reconstruction scenario.
[0079] In one optional embodiment, the preset quantization step size is any divisor of the encoder's output frame rate other than 1. For example, assuming the encoder's downsampling rate M = 320, for the audio to be processed with a sampling rate fs = 16kHz, the number of feature frames corresponding to each 1-second audio segment is T = fs / M = 50 frames. That is, the encoder's output frame rate is 50 frames / s, then the preset quantization step size S can be 2 frames, 5 frames, 10 frames, 25 frames, or 50 frames.
[0080] In one optional embodiment, the long-term feature extraction module is either a one-dimensional convolutional module or an average pooling module. Specifically, when the long-term feature extraction module is a one-dimensional convolutional module, the convolutional kernel of the long-term feature extraction module is 1 × a preset quantization stride, and the frame-by-frame long-term features represent the convolution result of the frame-by-frame encoded features; when the long-term feature extraction module is an average pooling module, the frame-by-frame long-term features represent the average value of the frame-by-frame encoded features. The average pooling module has lower complexity than the one-dimensional convolutional module.
[0081] Specifically, at least one segmented long-time feature is concatenated to obtain the audio long-time feature. The number of feature frames in the audio long-time feature is the same as the number of features in the segmented coded feature, and is the ratio between the number of feature frames in the audio coded feature and the preset quantization step size.
[0082] In an optional embodiment, the long-short-time quantization layer includes a long-time quantization module, a short-time feature extraction module, and a short-time quantization module. Correspondingly, the step of inputting the audio encoding features and the audio long-time features into the long-short-time quantization layer for long-short-time quantization processing to obtain audio quantization features includes: inputting the audio long-time features into the long-time quantization module for quantization processing to obtain long-time quantization features; inputting the audio encoding features and the audio long-time features into the short-time feature extraction module for short-time feature extraction to obtain audio short-time features; and inputting the audio short-time features into the short-time quantization module for quantization processing to obtain short-time quantization features.
[0083] Figure 4 is a schematic diagram of the architecture of a long-short-time quantization network provided in an embodiment of this disclosure. Specifically, the audio encoded features X are input into the long-time feature extraction module Extractor to obtain the output audio long-time features X. L The audio coding feature X and the audio long-time feature X L The input is fed into the short-time feature extraction module EncSyn to obtain the output audio short-time features X. S X L The input is fed into the long-time quantization module to obtain the output long-time quantization features. and audio short-time features X S The input is fed into the short-time quantization module to obtain the output short-time quantization features.
[0084] In this embodiment, the short-time audio feature is the feature obtained by removing the long-time audio feature from the audio coding features. See Figure 4, audio long-time feature X. L It can be represented as X L =f Extractor (X), short-time audio features X S It can be represented as X S =f EncSyn (X,X L ).
[0085] Specifically, audio short-time features are subframe-level features corresponding to audio coding features, and they represent the detailed features of each audio frame in the audio to be processed. The number of feature frames for audio short-time features is the same as the number of feature frames for audio coding features.
[0086] In one optional embodiment, the short-term feature extraction module is either a one-dimensional convolutional module or a difference module. Specifically, when the short-term feature extraction module is a one-dimensional convolutional module, the kernel of the short-term feature extraction module is 1, the number of input channels is 2N, and the number of output channels is 1N, where N is any positive integer. The difference module has lower complexity than the one-dimensional convolutional module.
[0087] For example, the quantization algorithms used by the long-time quantization module or the short-time quantization module include, but are not limited to, residual vector quantization algorithm, grouped vector quantization algorithm or scalar quantization algorithm. The quantization algorithm used by the long-time quantization module or the short-time quantization module is not limited here, and can be customized according to actual needs.
[0088] When the preset quantization step size is a divisor of the encoder's output frame rate, assuming that in the long-short-time quantization network, the long-time quantization module has N... q1 The codebook has N1 codes in each vector layer, and there are N short-time quantization modules. q2 A long short time quantization (LST) network has a vector layer, with N2 codebooks in each layer. The encoder's output frame rate is F, and the preset quantization stride is Stride. The bitrate of the LST network satisfies the following formula:
[0089] S130. Input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain audio dequantization features.
[0090] Specifically, the long-short-time dequantization network is used to perform long-time and short-time dequantization processing on audio quantization features based on the codebook.
[0091] In an optional embodiment, the long-short-time dequantization network includes a long-time dequantization module, a short-time dequantization module, and a feature fusion module. Correspondingly, the step of inputting the audio quantization features into the long-short-time dequantization network in the quantizer for long-short-time dequantization processing to obtain audio dequantization features includes: inputting the long-time quantization features into the long-time dequantization module for dequantization processing to obtain long-time dequantization features; inputting the short-time quantization features into the short-time dequantization module for dequantization processing to obtain short-time dequantization features; and inputting the long-time dequantization features and the short-time dequantization features into the feature fusion module for fusion processing to obtain audio dequantization features.
[0092] Specifically, the long-time dequantization module performs long-time dimension dequantization processing on long-time quantized features, while the short-time dequantization module performs short-time dimension dequantization processing on short-time quantized features. The long-time dequantized features represent approximate estimates of the long-time audio features, and the short-time dequantization features represent approximate estimates of the short-time audio features. Quantizers with different quantization and dequantization performance are used, so that the long-time dequantized features (or short-time dequantized features) may be completely identical to the long-time audio features (or short-time audio features), or they may not be completely identical.
[0093] Figure 5 is a schematic diagram of the architecture of a long-short-time dequantization network provided in an embodiment of this disclosure. Specifically, the long-time quantization features... The input is fed into the long-time dequantization module to obtain the output long-time dequantization feature X′. L Quantize short-time features The input is fed into the short-time dequantization module to obtain the output short-time dequantization feature X′. S Quantize the long-time solution features X′ L and short-time solution quantization feature X′ S The input is fed into the feature fusion module DecSyn to obtain the output audio dequantization feature X′. Referring to Figure 5, the audio dequantization feature X′ can be expressed as X′=f DecSyn (X′ L ,X′ S ).
[0094] In one optional embodiment, when the short-time feature extraction module is a one-dimensional convolution module, the feature fusion module is also a one-dimensional convolution module; when the short-time feature extraction module is a difference module, the feature fusion module is an addition module. Specifically, the one-dimensional convolution module has a kernel of 1, 2N input channels, and 1N output channels, where N is any positive integer.
[0095] S140. Input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the target reconstructed audio.
[0096] The decoder is used to perform decoding and reconstruction based on the audio dequantization features. In one optional embodiment, the network architecture of the decoder is symmetrically configured with respect to the encoder in the audio processing model. Exemplary decoders include, but are not limited to, decoders based on fully convolutional networks, decoders based on Opus, or decoders based on SoundStream, etc. The decoder used is not limited here, and can be customized according to actual needs.
[0097] Figure 6 is a schematic diagram comparing an audio coding feature with an STFT spectrum provided in an embodiment of this disclosure. Specifically, the dark gray area at the top of Figure 6 represents the visualized audio coding feature obtained by the encoder after processing the input audio, and the black area at the bottom represents the STFT (short-time Fourier transform) spectrum obtained by performing a short-time Fourier transform on the input audio. As can be seen from the four (but not limited to four) light gray line areas marked in Figure 6, the audio coding feature exhibits significant consistency within some feature frames and aligns with the STFT spectrum in the frequency dimension. In other words, the audio coding feature shares common characteristics within some feature frames.
[0098] From the perspective of the time dimension of the audio to be processed, there is usually redundant feature information in the audio, such as pronunciation timbre and ambient sound. This feature information usually does not change or changes very little over a period of time.
[0099] The technical solution of this embodiment sets the quantizer in the audio processing model to include a long short-time quantization network and a long short-time dequantization network. The long short-time quantization network is used to perform long short-time quantization processing on the audio coding features obtained by the encoder to obtain long-time quantization features and short-time quantization features. The long short-time dequantization network and the decoder determine the target reconstructed audio based on the long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed. This solves the problem of traditional quantizers using the same bit rate for quantization, reduces the bit rate required for audio quantization features, thereby improving the coding efficiency and transmission efficiency of the audio processing model, and thus improving the processing efficiency of audio data.
[0100] Figure 7 is a flowchart of another audio processing method provided in an embodiment of this disclosure. This embodiment further refines the "long-short-time quantization layer" in the above embodiments. Specifically, the long-short-time quantization layer includes a long-time quantization module, a long-time dequantization module, a short-time feature extraction module, and a short-time quantization module. As shown in Figure 7, the method includes:
[0101] S210. Encode the encoder in the pre-trained audio processing model to obtain audio coding features from the audio input to be processed.
[0102] In this embodiment, S210 corresponds to or is similar to S110 shown in Figure 3 of the above embodiment, and will not be described again in this embodiment.
[0103] S220. Input the audio encoding features into the long-term feature extraction module to extract long-term features and obtain audio long-term features.
[0104] S230. Input the audio long-time features into the long-time quantization module for quantization processing to obtain long-time quantization features.
[0105] In this embodiment, S220-S230 corresponds to or is similar to the content regarding "long-term quantization features" in the above embodiments, and will not be repeated here.
[0106] S240. Input the long-time quantization feature into the long-time dequantization module for dequantization processing to obtain the long-time dequantization feature.
[0107] Specifically, the long-time dequantization module is used to perform long-time dimension dequantization processing on the long-time quantization features. The long-time dequantization features represent the approximate estimation results of the long-time audio features. This embodiment uses an example where the long-time dequantization features are not completely identical to the long-time audio features for explanation.
[0108] S250. Input the audio encoding features and the long-time dequantization features into the short-time feature extraction module to extract short-time features and obtain audio short-time features.
[0109] In this embodiment, the short-time audio feature is the audio coding feature after removing the long-time dequantization feature.
[0110] S260. Input the audio short-time features into the short-time quantization module for quantization processing to obtain short-time quantization features.
[0111] Figure 8 is a schematic diagram of another long-short-time quantization network architecture provided in an embodiment of this disclosure. Specifically, the audio encoded features X are input into the long-short-time feature extraction module Extractor to obtain the output audio long-time features X. L X L The input is fed into the long-time quantization module to obtain the output long-time quantization features. Long-term quantization features The input is fed into the long-time dequantization module to obtain the output long-time dequantization feature X′. L The audio coding features X and the long-time dequantization features X′ are combined. LThe input is fed into the short-time feature extraction module EncSyn to obtain the output audio short-time features X. S and the short-time audio features X S The input is fed into the short-time quantization module to obtain the output short-time quantization features. See Figure 8, audio short-time feature X S It can be represented as X S =f EncSyn (X,X′ L ).
[0112] S270. Input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain audio dequantization features.
[0113] In an optional embodiment, both the short-term feature extraction module EncSyn and the feature fusion module DecSyn are one-dimensional convolutional modules. Figure 9 is a schematic diagram of the combined architecture of a long-term feature extraction module, a short-term feature extraction module, and a feature fusion module provided in an embodiment of this disclosure. Figure 9 takes the long-term feature extraction module Extractor as a one-dimensional convolutional module, the number of feature frames of the audio coding features as 8 frames, and the preset quantization step size as 4 frames as an example. Specifically, the long-term feature extraction module performs long-term feature extraction on the frame coding features (X1-X4 or X5-X8) composed of every 4 audio frames, and obtains 2 frame long-term features, namely X1-X4-X5-X8 ... L1 and X L2 Using the frame-by-frame long-time feature X L1 For example, the short-time feature extraction module EncSyn extracts features based on frame coding features X1-X4 and frame long-time features X. L1 The corresponding frame-by-frame long-time dequantization feature X′ L1 Convolution processing is performed to obtain the frame-by-frame short-time features X. S1 -X S4 Correspondingly, the feature fusion module DecSyn dequantizes the feature X′ based on the frame-by-frame long-time dequantization. L1 and frame-by-frame short-time features X S1 -X S4 The corresponding frame-by-frame short-time dequantization feature X′ S1 -X′ S4 Perform convolution processing to obtain the frame dequantization features X′1-X′4 corresponding to the frame coding features X1-X4.
[0114] In another optional embodiment, the short-term feature extraction module EncSyn is a difference module, and the feature fusion module DecSyn is a summation module. Figure 10 is a schematic diagram of another combined architecture of a long-term feature extraction module, a short-term feature extraction module, and a feature fusion module provided in an embodiment of this disclosure. Figure 10 takes the long-term feature extraction module as a one-dimensional convolution module, the feature dimension of the audio coding features as 8 frames, and the preset quantization step size as 4 frames as an example. Similarly, taking the frame-segmented long-term feature X... L1 For example, the short-time feature extraction module EncSyn extracts features based on frame coding features X1-X4 and frame long-time features X. L1 The corresponding frame-by-frame long-time dequantization feature X′ L1 Perform interpolation to obtain the frame-by-frame short-time features X. S1 -X S4 Correspondingly, the feature fusion module DecSyn dequantizes the feature X′ based on the frame-by-frame long-time dequantization. L1 and frame-by-frame short-time features X S1 -X S4 The corresponding frame-by-frame short-time dequantization feature X′ S1 -X′ S4 Summation is performed to obtain the frame dequantization features X′1-X′4 corresponding to the frame coding features X1-X4.
[0115] Referring to Figures 9 and 10 above, taking the frame coding features X1-X4 as an example, the audio quantization data packets transmitted to the long short-time dequantization network contain...
[0116] S280. Input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the target reconstructed audio.
[0117] In this embodiment, S280 corresponds to or is similar to S140 shown in Figure 3 of the above embodiment, and will not be described again in this embodiment.
[0118] Because the long-time quantization module uses a limited number of codebooks to quantize long-time audio features, some feature information is lost. For example, suppose the long-time audio feature is 6.5 and the long-time quantization feature is 'a', but the long-time dequantized feature obtained by dequantizing the long-time quantization feature 'a' may be 6. Therefore, the long-time audio feature loses 0.5 after quantization and dequantization.
[0119] The technical solution of this embodiment includes a long-time dequantization module in the long-short-time quantization layer. This module dequantizes the long-time quantization features to obtain long-time dequantized features. A short-time feature extraction module then extracts short-time audio features from the audio coding features and the long-time dequantized features. Taking the example above, assuming the audio coding feature is 10, in this embodiment, the audio short-time feature is 4 instead of 3.5. This achieves the goal of retaining the feature information lost by the long-time quantization module in the audio short-time features, further improving the accuracy of the audio quantization features and thus further ensuring the processing effect of the audio data.
[0120] Based on the above embodiments, optionally, the decoder includes a convolutional network, a decoding network, and an output network in series. The decoding network includes at least one decoding module in series. The decoding module includes a convolutional activation function layer, a transposed convolutional layer, and a residual structure in series. The residual structure includes at least one residual unit in series.
[0121] Figure 11 is a schematic diagram of a decoder structure provided in an embodiment of this disclosure. Figure 11 uses an example where the decoding network contains four decoding modules and the residual structure contains three residual units. Specifically, "k" in Figure 11 represents the convolution kernel, also known as a filter, which is a matrix used to perform convolution operations on the input data. The size of the convolution kernel determines the neighborhood range considered by the convolution operation. Here, it is a one-dimensional convolution with a kernel size of 1×7. "N" or "n" in Figure 11 represents the number of output channels of the convolution. The number of input channels of the convolution is ignored in Figure 11. It can be understood that the number of input channels of the convolution is equal to the number of output channels of the previous layer. "Dilation" in Figure 11 represents the dilation rate, which is used to control the spacing between elements in the convolution kernel. For example, when the dilation rate is 1, the convolution kernel elements are closely adjacent; when the dilation rate is greater than 1, there will be gaps between the convolution kernel elements, thereby expanding the receptive field without increasing the size of the convolution kernel. "S" or "stride" in Figure 11 represents the stride, which refers to the step size by which the convolution kernel moves on the input data. For example, a stride of 1 means the convolutional kernel moves one element at a time, and a stride of 2 means the convolutional kernel moves two elements at a time. A larger stride results in a smaller output feature map size. In Figure 11, "C" represents a hyperparameter in the decoder.
[0122] In one optional embodiment, the residual unit includes a first activation function layer, a first convolutional layer, a second activation function layer, and a residual output layer connected in series. The number of output channels in the first activation function layer is the same as the number of output channels in the second activation function layer.
[0123] In another optional embodiment, the residual unit includes a first activation function layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a residual output layer connected in series. The first and third convolutional layers have the same number of output channels, while the second convolutional layer has more output channels than the first convolutional layer. Correspondingly, the second activation function layer has more output channels than the first activation function layer.
[0124] For example, when the number of output channels of the first convolutional layer is 1N, the number of output channels of the second convolutional layer is 2N or 3N. In an optional embodiment, the number of output channels of the second convolutional layer is four times the number of output channels of the first convolutional layer.
[0125] The advantage of setting up a second convolutional layer for dimensionality increase, an activation function layer, and a third convolutional layer for dimensionality reduction is that it performs intra-unit feature expansion on the residual units, enabling them to learn higher-dimensional features, thereby further improving the processing effect of audio data.
[0126] In neural networks, the activation function is a crucial component. Its role is to introduce non-linear characteristics into the network. Without activation functions, a neural network merely performs linear combinations of input data, resulting in very limited expressive power and making it difficult to solve complex non-linear problems.
[0127] In one optional embodiment, the decoder employs an activation function such as the Sigmoid function, Tanh function, ReLU (Rectified Linear Unit) function, Leaky ReLU function, ELU function, or Snake function. The Sigmoid function compresses the input value to a range of 0 to 1. The Tanh function compresses the input value to a range of -1 to 1. The ReLU function outputs 0 when the input value is less than 0 and outputs the same value when the input value is greater than 0. The Leaky ReLU function is an improved version of the ReLU function, outputting a slope value related to the input value when the input value is less than 0. The ELU function outputs the same value when the input value is greater than 0 and outputs an exponential value related to the input value when the input value is less than or equal to 0.
[0128] For example, the Snake function satisfies the following formula:
[0129] Where x represents the input value and α represents the learnable model parameters.
[0130] Based on the above embodiments, optionally, the decoder satisfies at least one of the following conditions: 1) it uses the SnakeBeta activation function; 2) at least one of the convolutional structures in the convolutional network, the transposed convolutional layer, and the convolutional layer in the residual unit uses grouped convolution; wherein, the number of groups in the convolutional structure is the common divisor of the number of input channels and the number of output channels of the convolutional structure; 3) the residual structure contains 4 residual units, and the dilation rates of each residual unit are 1, 3, 9, and 27, respectively.
[0131] In one specific embodiment, for example, the SnakeBeta activation function satisfies the following formula:
[0132] Where x represents the input value, and α and β both represent learnable model parameters.
[0133] The advantage of setting the SnakeBeta activation function is that it further improves the processing effect of audio data.
[0134] In one specific embodiment, the number of groups corresponding to different convolutional structures can be the same or different. Specifically, in grouped convolution, the number of channels in the input feature map is grouped according to the number of groups, and the channels in each group are convolved with the corresponding convolutional kernel. The number of groups determines the parallelism and computational efficiency of the convolution.
[0135] The advantages of setting up grouped convolution are twofold. First, it reduces the computational cost of convolution operations, thereby lowering the computational cost and enabling the decoder to run more efficiently on devices with limited computing resources, thus further improving the processing efficiency of audio data. Second, it can play a regularization role to a certain extent, preventing overfitting of the audio processing model.
[0136] The advantage of setting up four residual units is that by increasing the number of layers in the residual structure, the receptive field of the decoder can be increased, thereby further improving the processing effect of audio data.
[0137] Figure 12 is a schematic diagram of another decoder structure provided in an embodiment of this disclosure. Specifically, "group" or "G" in Figure 12 represents the number of groups. Figure 12 takes C=16 as an example and sets G=8. Referring to Figure 12, the decoder uses the Snake Beta activation function, and all convolutional structures use grouped convolution. The residual structures contain residual units with dilation rates of 1, 3, 9, and 27, respectively. The number of output channels of the second convolutional layer is four times the number of output channels of the first convolutional layer.
[0138] In audio processing scenarios, the encoding end (including the encoder and quantization network in the quantizer) of the audio processing model can be deployed on a server, while the decoding end (including the decoder and dequantization network in the quantizer) is typically deployed on a mobile device. Since mobile devices have relatively limited computing power, the decoder complexity is usually required to be low. However, traditional decoders experience a rapid decline in decoding quality as complexity decreases. This embodiment primarily optimizes the decoder architecture to mitigate the degradation in decoding quality caused by lower complexity. In other words, this embodiment minimizes the decoder complexity while maintaining decoding quality.
[0139] Since the server has relatively sufficient computing power, this embodiment does not provide a detailed explanation of the encoder architecture optimization. However, it is understood that the above-mentioned decoder architecture optimizations can be applied partially or entirely to the encoder. When all of the above-mentioned decoder architecture optimizations are applied to the encoder, the audio processing model is a symmetrical end-to-end neural network encoder-decoder.
[0140] Figure 13 is a flowchart of a training method for an audio processing model provided in an embodiment of this disclosure. This embodiment is applicable to training an audio processing model for audio reconstruction. The method can be executed by a training device for the audio processing model, which can be implemented in hardware and / or software. This training device can be configured in an electronic device, typically a mobile terminal or tablet computer. As shown in Figure 13, the method includes:
[0141] S310. Input the training audio into the encoder of the untrained audio processing model and encode it to obtain audio coding features.
[0142] For example, training audio can be categorized based on sound content, such as music, waveform audio, or speech; and based on audio format, such as WAV, MP3, WMA, or FLAC. There are no restrictions on the sound content or format of the training audio here; specific settings can be customized according to actual needs.
[0143] In this embodiment, the untrained audio processing model includes a cascaded encoder, quantizer, and decoder.
[0144] The encoder is used to compress the training audio in the temporal dimension and expand it to the feature dimension. For example, the encoder includes, but is not limited to, encoders based on fully convolutional networks, encoders based on recurrent neuron layers, encoders based on Opus, or encoders based on SoundStream, etc. The encoder used is not limited here, and can be customized according to actual needs.
[0145] S320. Input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain audio quantization features.
[0146] In this embodiment, the quantizer includes a long short-time quantization network disposed at the encoding end and a long short-time dequantization network disposed at the decoding end. The long short-time quantization network is used to perform long-time dimension compression quantization and short-time dimension compression quantization on the audio encoded features according to the codebook, and transmits the resulting audio quantized features to the long short-time dequantization network at a specified bit rate.
[0147] In this embodiment, the audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames for the long-time quantization features is less than the number of feature frames for the short-time quantization features. The long-time quantization features characterize the quantization results of the common features corresponding to the audio segments obtained from the training audio segmentation. The number of audio segments can be one or more, and each audio segment corresponds to at least two audio frames.
[0148] The technical content of the "long short-time quantization network" in this embodiment is the same as or similar to the technical content of the "long short-time quantization network" in the audio processing method provided in any of the above embodiments, and will not be described again in this embodiment.
[0149] S330. Input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain audio dequantization features.
[0150] Specifically, the long-short-time dequantization network is used to perform long-time and short-time dequantization processing on audio quantization features based on the codebook.
[0151] The technical content of the "long short-time dequantization network" in this embodiment is the same as or similar to the technical content of the "long short-time dequantization network" in the audio processing method provided in any of the above embodiments, and will not be described again in this embodiment.
[0152] S340. Input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the predicted reconstructed audio.
[0153] The technical content of the "decoder" in this embodiment is the same as or similar to the technical content of the "decoder" in the audio processing method provided in any of the above embodiments, and will not be repeated here.
[0154] S350. Based on the training audio and the predicted reconstructed audio, determine the target loss value, and based on the target loss value, iteratively train the untrained audio processing model to obtain the trained audio processing model.
[0155] In one optional embodiment, the loss function corresponding to the target loss value, for example, includes, but is not limited to, the squared loss function, the logarithmic loss function, the exponential loss function, the mean squared error loss function, the logistic regression loss function, the Huber loss function, the cross-entropy loss function, the Kullback-Leibler divergence loss function, and the spectral loss function, etc. For example, the spectral loss function can be the STFT spectral loss function or the multi-scale mel-reconstruction spectral loss function.
[0156] In one specific embodiment, determining the target loss value based on the training audio and the predicted reconstructed audio includes: when the loss function corresponding to the target loss value includes the STFT spectral loss function, performing Fast Fourier Transform on the training audio and the predicted reconstructed audio respectively to obtain the training STFT spectrum and the predicted STFT spectrum; and determining the STFT spectral loss value based on the training STFT spectrum and the predicted STFT spectrum.
[0157] In this embodiment, for example, the training STFT spectrum and the predicted STFT spectrum can be the amplitude spectrum or complex spectrum of the STFT spectrum.
[0158] In another specific embodiment, determining the target loss value based on the training audio and the predicted reconstructed audio includes: when the loss function corresponding to the target loss value includes a multi-scale Mel reconstruction spectrum loss function, performing Fast Fourier Transform on the training audio and the predicted reconstructed audio respectively to obtain the training STFT spectrum and the predicted STFT spectrum; using a Mel filter bank to filter the training STFT spectrum and the predicted STFT spectrum respectively to obtain training filtered data and predictive filtered data; performing logarithmic processing on the training filtered data and the predictive filtered data respectively to obtain training Mel spectrum data and predictive Mel spectrum data; and determining the multi-scale Mel reconstruction spectrum loss value based on the training Mel spectrum data and the predicted Mel spectrum data.
[0159] The Mel filter bank contains multiple Mel filters, each with a different window length to generate Mel spectra at different resolutions. For example, the Mel filter bank contains seven Mel filters with window lengths of 32, 64, 128, 256, 512, 1024, and 2048, respectively, with corresponding step sizes of 4, 16, 32, 64, 128, 256, and 512.
[0160] In this embodiment, both the training Mel spectrum data and the predicted Mel spectrum data are logarithmic Mel spectra.
[0161] For example, the multi-scale Mel reconstruction spectral loss value L mel Satisfy the following formula:
[0162] Where x represents the training audio, and G(x) represents the predicted reconstructed audio. This represents the f-th logarithmic Mel spectrum value in the t-th audio frame corresponding to the s-th Mel filter.
[0163] Based on the above embodiments, optionally, the target loss value is the multi-scale Mel reconstruction spectral loss value L. mel The result is a weighted sum of the feature matching loss function value, the codebook loss value, and the commitment loss value. The feature matching loss function value represents the difference between the output features of each network architecture in the audio processing model based on the training audio and the output features based on the predicted reconstructed audio. The codebook loss value is trained to make the codebook in the quantizer approximate the audio coding features output by the encoder. The commitment loss value is trained to make the audio coding features output by the encoder approximate the codebook in the quantizer.
[0164] For example, the multi-scale Mel reconstruction spectral loss value L mel The loss weights corresponding to the feature matching loss function value, codebook loss value, and commitment loss value can be 15, 2, 1, and 0.25, respectively. There is no limit to the loss weight corresponding to each loss value here, and it can be customized according to actual needs.
[0165] In an optional embodiment, determining the target loss value based on the training audio and the predicted reconstructed audio includes: performing transient feature detection on the training audio to obtain transient audio and non-transient audio; determining a transient enhancement spectrum loss value based on the transient audio, non-transient audio, and the predicted reconstructed audio; and determining the target loss value based on the transient enhancement spectrum loss value; wherein the loss weight corresponding to the transient audio is higher than the loss weight corresponding to the non-transient audio.
[0166] Specifically, a transient signal refers to a signal that changes rapidly over a short period of time, has abundant high-frequency components, and distinct start and end points. Examples of transient signals include percussion instruments, the beginning of piano notes, and explosions.
[0167] In this embodiment, the transient enhancement spectral loss value is the transient enhancement STFT spectral loss value or the transient enhancement Mel spectrum loss value.
[0168] Taking the loss weights corresponding to transient audio and non-transient audio as 2 and 1 respectively, and the transient enhanced spectral loss value as the transient enhanced Mel spectrum loss value as an example, the transient enhanced Mel spectrum loss value L mel-transient Satisfy the following formula:
[0169] Where, λ t Satisfy the following formula:
[0170] Where x∈transient represents transient audio. This indicates non-transient audio.
[0171] In audio processing scenarios, transient audio is crucial for the realism of reconstructed audio. Proper handling of transient audio can make the predicted and reconstructed audio sound clearer, more vivid, and more lifelike, while neglecting transient audio can lead to a blurry or unimpressive reconstructed audio. Setting a transient enhancement spectral loss value can improve the reconstruction quality of the audio processing model.
[0172] In an optional embodiment, determining the target loss value based on the training audio and the predicted reconstructed audio includes: acquiring training spectrum data corresponding to the training audio and predicted spectrum data corresponding to the predicted reconstructed audio; comparing the training spectrum data and the predicted spectrum data to obtain first spectrum data and second spectrum data corresponding to the training spectrum data; determining an energy suppression spectrum loss value based on the first spectrum data, the second spectrum data, and the predicted spectrum data; and determining the target loss value based on the energy suppression spectrum loss value.
[0173] In this embodiment, each training spectrum value in the first spectrum data is less than the corresponding predicted spectrum value in the predicted spectrum data, each training spectrum value in the second spectrum data is greater than or equal to the corresponding predicted spectrum value in the predicted spectrum data, and the loss weight corresponding to the first spectrum data is higher than the loss weight corresponding to the second spectrum data.
[0174] In this embodiment, the energy-suppressed spectral loss value is either the energy-suppressed STFT spectral loss value or the energy-suppressed Mel spectral loss value. Specifically, when the energy-suppressed spectral loss value is the energy-suppressed STFT spectral loss value, the training spectral data and the predicted spectral data are the amplitude spectrum or complex spectrum of the STFT spectrum; when the energy-suppressed spectral loss value is the energy-suppressed Mel spectral loss value, the training spectral data are the training Mel spectral data, the predicted spectral data are the predicted Mel spectral data, the first spectral data are the first Mel spectral data, and the second spectral data are the second Mel spectral data.
[0175] Taking the loss weights corresponding to the first Mel data and the second Mel data as 2 and 1 respectively, and the energy-suppressed spectral loss value as the energy-suppressed Mel spectral loss value as an example, the energy-suppressed Mel spectral loss value L mel-energy Satisfy the following formula:
[0176] Where, λ e Satisfy the following formula:
[0177] in, This indicates that the logarithmic Mel-spectrum value in the training Mel-spectrum data belongs to the first Mel-spectrum data. This indicates that the logarithmic Mel value in the training Mel spectrum data belongs to the second Mel spectrum data.
[0178] When the predicted reconstructed audio contains excess reconstruction energy, it sounds noisy and is easily detected by the human ear. The benefit of setting an energy suppression spectral loss value is that it penalizes the excess reconstruction energy contained in the predicted reconstructed audio, thereby reducing the likelihood that the predicted reconstructed audio will contain more energy than the training audio.
[0179] Based on the above embodiments, optionally, the multi-scale Mel reconstruction spectrum loss value L is determined according to the transient enhanced spectrum Mel spectrum loss value and the energy suppressed Mel spectrum loss value. mel .
[0180] Specifically, the multi-scale Mel reconstruction spectral loss value L in any of the above embodiments mel Satisfying the formula:
[0181] L mel =L mel-transient +L mel-energy .
[0182] The technical solution of this embodiment solves the problem of traditional quantizers using the same bit rate by training an audio processing model equipped with a long short-time quantization network and a long short-time dequantization network, thereby reducing the computational load of the audio processing model and improving the training efficiency of the audio processing model.
[0183] Figure 14 is a flowchart of another training method for an audio processing model provided in an embodiment of this disclosure. This embodiment further refines the step of "determining the target loss value based on the training audio and the predicted reconstructed audio" in the above embodiment. As shown in Figure 14, the method includes:
[0184] S410. Input the training audio into the encoder of the untrained audio processing model and encode it to obtain audio coding features.
[0185] S420. Input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain audio quantization features.
[0186] S430. Input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain audio dequantization features.
[0187] S440. Input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the predicted reconstructed audio.
[0188] In this embodiment, S410-S440 corresponds to or is similar to S310-S340 shown in Figure 13 of the above embodiment, and will not be described again in this embodiment.
[0189] S450. Input the predicted reconstructed audio into the discriminator to obtain the output predicted classification result.
[0190] The prediction classification result indicates whether the predicted reconstructed audio is a real video or a generated video. Specifically, if the predicted reconstructed audio has significant feature differences from the real audio, the discriminator considers the predicted reconstructed audio to be a generated video; if the predicted reconstructed audio does not have significant feature differences from the real audio, the discriminator considers the predicted reconstructed audio to be a real video.
[0191] In one alternative embodiment, the discriminator is a multi-period waveform discriminator (MPD) or a multi-resolution spectrogram discriminator (MRSD).
[0192] Among them, the multi-period waveform discriminator is used to model discontinuous sampling points. For example, prime numbers 2, 3, 5, 7, and 11 are set as different periods. The input audio is recombined into a two-dimensional signal according to the period, and then the two-dimensional signal after periodic resampling is processed separately by convolution.
[0193] The multi-resolution spectral discriminator analyzes the characteristics of the spectrogram at different resolutions to determine the difference between the spectrum of the predicted reconstructed audio and the spectrum of the real audio. Multi-resolution analysis helps capture spectral details and patterns at different scales, thus providing a more comprehensive and accurate assessment of the quality of the predicted reconstructed audio. For example, at low resolution, the focus can be on overall spectral trends and large structures, while at high resolution, more subtle frequency variations and local features can be captured.
[0194] In one optional embodiment, the discriminator includes resolution spectrum networks corresponding to at least two window lengths, each resolution spectrum network comprising a spectrum conversion module, a parallel frequency band feature extraction layer corresponding to at least two frequency bands, a feature concatenation layer, a global convolutional network, and a discriminative output layer. Correspondingly, the step of inputting the predicted reconstructed audio into the discriminator to obtain the output predicted classification result includes: inputting the predicted reconstructed audio into each resolution spectrum network to obtain the output resolution classification result; and determining the predicted classification result based on at least two resolution classification results.
[0195] Specifically, for each resolution spectral network, the predicted and reconstructed audio is input into the spectral conversion module for spectral conversion to obtain an audio spectrogram; the audio spectrogram and the corresponding audio frequency bands for each frequency band are input into the corresponding frequency band feature extraction layer for frequency band feature extraction to obtain audio frequency band features; each audio frequency band feature is input into the feature concatenation layer for concatenation processing to obtain global spectral features; the global spectral features are input into the global convolutional network for global feature extraction to obtain global audio features; the global audio features are input into the discriminative output layer for one-dimensional convolution processing to obtain the resolution classification result.
[0196] Specifically, the audio spectrogram can be an amplitude spectrum or a complex spectrum. In one specific embodiment, the spectrum is converted to obtain a complex spectrum as the audio spectrogram.
[0197] In one alternative embodiment, the global convolutional network includes at least one convolutional network, which includes two-dimensional convolutional layers and activation function layers.
[0198] In one alternative embodiment, the frequency band feature extraction layer consists of multiple convolutional networks connected in series.
[0199] For example, in a global convolutional network or frequency band feature extraction layer, the activation function used in the activation function layer includes, but is not limited to, the Sigmoid function, Tanh function, ReLU function, Leaky ReLU function, ELU function, or Snake function. There is no limitation on the activation function used in the activation function layer here, and it can be customized according to actual needs.
[0200] Figure 15 is a schematic diagram of the architecture of a discriminator provided in an embodiment of this disclosure. Figure 15 uses a resolution spectrum network corresponding to a window length of 512 as an example. Specifically, "STFT1" represents the spectrum conversion module, and frequency bands 1 to 5 represent 0-0.8kHz, 0.8-2kHz, 2-4kHz, 4-6kHz, and 6-8kHz, respectively. Here, the number of frequency band feature extraction layers and the frequency band corresponding to each frequency band feature extraction layer are not limited, and can be customized according to actual needs.
[0201] Referring to Figure 15, the network structure within the dashed box in Figure 15 is a global convolutional network. Figure 15 illustrates this example using a global convolutional network containing two convolutional networks. In the two-dimensional convolutional layer, "32" represents the number of output channels, "[3,9]" represents the convolutional kernel, and "[1,2]" represents the stride.
[0202] Dividing the audio spectrogram into multiple frequency bands for feature extraction emphasizes different harmonic modes in each band, allowing the discriminator to learn the corresponding feature information for each band. However, this also leads to harmonic mode mismatch issues between adjacent frequency bands, resulting in poor sound quality of the reconstructed audio output by the trained audio processing model. This embodiment addresses this by setting a global convolutional network after the feature concatenation layer, enabling the discriminator to model the entire spectrogram. This solves the harmonic mode mismatch problem between adjacent frequency bands, thereby improving the classification performance of the discriminator.
[0203] S460. Determine the target loss value based on the training audio, the predicted reconstructed audio, and the predicted classification result.
[0204] Based on the above embodiments, optionally, the target loss value also includes a weighted result of the adversarial loss value. For example, the weight of the adversarial loss value can be 1.
[0205] S470. Based on the target loss value, iteratively train the untrained audio processing model to obtain a trained audio processing model.
[0206] Based on the above embodiments, optionally, according to the target loss value, iteratively train the untrained audio processing model to obtain a trained audio processing model, including: iteratively training the untrained audio processing model according to the target loss value to obtain a pre-trained audio processing model; performing channel pruning on the pre-trained audio processing model and using the processed audio processing model as the untrained audio processing model; returning to the step of inputting the training audio into the encoder in the untrained audio processing model for encoding processing to obtain audio encoding features; until the training termination condition is met, using the pre-trained audio processing model in the current iteration as the trained audio processing model.
[0207] Taking the audio processing model in the audio processing method provided in the above embodiment as an example, channel pruning is a pruning process for "C" in the audio processing model.
[0208] For example, training termination conditions include, but are not limited to, the number of pruning iterations reaching a threshold, the number of pruned channels meeting a channel number range, and the training audio being the last training data.
[0209] In an optional embodiment, before using the pre-trained audio processing model from the current iteration as the trained audio processing model, the method further includes: fine-tuning the pre-trained audio processing model from the current iteration. This setup has the advantage of reducing the performance loss that channel pruning might cause.
[0210] The advantage of setting up pruning training is that it removes unimportant or redundant channel parameters from the audio processing model. While ensuring the processing performance of the audio processing model, it minimizes the size and computational cost of the trained audio processing model, thereby further improving the processing efficiency of the audio processing model.
[0211] In the training scenario of audio processing models, the discriminator has the following technical effects: 1) Evaluating generation quality: The discriminator can judge the similarity between the predicted reconstructed audio and the training audio, providing a quantitative evaluation of generation quality for the training process. 2) Providing gradient feedback: Based on the prediction classification results of the predicted reconstructed audio, it provides gradient information to the audio processing model, guiding the model to optimize its parameters towards generating more realistic audio. 3) Promoting convergence: Through continuous adversarial training with the discriminator, the audio processing model can gradually improve, making the reconstructed audio more consistent with the distribution of real audio, thereby promoting the convergence of the entire training process. 4) Preventing pattern collapse: The existence of the discriminator helps prevent the audio processing model from falling into simple repetitive patterns or generating overly similar audio, encouraging the audio processing model to produce diverse and high-quality outputs. 5) Enhancing generalization ability: In the interaction with the discriminator, the audio processing model can learn different types and styles of audio features, thereby improving its generalization ability to various input conditions.
[0212] The following are embodiments of the audio processing apparatus provided in this disclosure. This apparatus and the audio processing method described above belong to the same inventive concept. For details not described in detail in the embodiments of the audio processing apparatus, please refer to the content of the audio processing method in the above embodiments.
[0213] Figure 16 is a schematic diagram of an audio processing apparatus provided in an embodiment of the present disclosure. As shown in Figure 16, the apparatus includes: an audio input module 510 to be processed, an audio quantization feature determination module 520, an audio dequantization feature determination module 530, and a target reconstructed audio determination module 540.
[0214] The audio input module 510 is used to encode the encoder in the pre-trained audio processing model of the audio input to obtain audio encoding features.
[0215] The audio quantization feature determination module 520 is used to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0216] The audio dequantization feature determination module 530 is used to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features.
[0217] The target reconstructed audio determination module 540 is used to input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the target reconstructed audio.
[0218] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed.
[0219] The technical solution of this embodiment sets the quantizer in the audio processing model to include a long short-time quantization network and a long short-time dequantization network. The long short-time quantization network is used to perform long short-time quantization processing on the audio coding features obtained by the encoder to obtain long-time quantization features and short-time quantization features. The long short-time dequantization network and the decoder determine the target reconstructed audio based on the long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed. This solves the problem of traditional quantizers using the same bit rate for quantization, reduces the bit rate required for audio quantization features, and thus improves the processing efficiency of audio data.
[0220] In an optional embodiment, the long-short-time quantization network includes a long-time feature extraction module and a long-short-time quantization layer. Correspondingly, the audio quantization feature determination module 520 includes:
[0221] The audio long-term feature determination unit is used to input the audio encoded features into the long-term feature extraction module to extract long-term features and obtain audio long-term features.
[0222] The audio quantization feature determination unit is used to input the audio coding features and the audio long-time features into the long-short-time quantization layer for long-short-time quantization processing to obtain the audio quantization features.
[0223] In one optional embodiment, the audio long-term feature determination unit is specifically used for:
[0224] Based on a preset quantization step size, the audio coding features are processed into frames to obtain at least one frame-based coding feature.
[0225] Long-term features are extracted for each frame coding feature to obtain frame long-term features, and audio long-term features are determined based on at least one frame long-term feature.
[0226] In an optional embodiment, the long-short-time quantization layer includes a long-time quantization module, a long-time dequantization module, a short-time feature extraction module, and a short-time quantization module. Correspondingly, the audio quantization feature determination unit is specifically used for:
[0227] The audio long-time features are input into the long-time quantization module for quantization processing to obtain long-time quantized features;
[0228] The long-time quantization features are input into the long-time dequantization module for dequantization processing to obtain long-time dequantization features;
[0229] The audio encoding features and the long-time dequantization features are input into the short-time feature extraction module to extract short-time features and obtain audio short-time features.
[0230] The audio short-time features are input into the short-time quantization module for quantization processing to obtain short-time quantized features.
[0231] In an optional embodiment, the long-short-time dequantization network includes a long-time dequantization module, a short-time dequantization module, and a feature fusion module. Correspondingly, the audio dequantization feature determination module 530 is specifically used for:
[0232] The long-time quantization features are input into the long-time dequantization module for dequantization processing to obtain long-time dequantization features;
[0233] The short-time quantization features are input into the short-time dequantization module for dequantization processing to obtain short-time dequantization features;
[0234] The long-time dequantization features and the short-time dequantization features are input into the feature fusion module for fusion processing to obtain audio dequantization features.
[0235] In one alternative embodiment, the decoder includes a convolutional network, a decoding network, and an output network in series. The decoding network includes at least one decoding module in series. The decoding module includes a convolutional activation function layer, a transposed convolutional layer, and a residual structure in series. The residual structure includes at least one residual unit in series.
[0236] In one optional embodiment, the residual unit includes a first activation function layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a residual output layer connected in series, wherein the first convolutional layer and the third convolutional layer have the same number of output channels, and the second convolutional layer has more output channels than the first convolutional layer.
[0237] In an alternative embodiment, the decoder satisfies at least one of the following conditions:
[0238] 1) The Snake Beta activation function is used;
[0239] 2) At least one of the convolutional structures in the convolutional network, the transposed convolutional layer, and the convolutional layer in the residual unit adopts grouped convolution; wherein, the number of groups in the convolutional structure is the common divisor of the number of input channels and the number of output channels of the convolutional structure;
[0240] 3) The residual structure contains 4 residual units, and the expansion rates of each residual unit are 1, 3, 9 and 27, respectively.
[0241] The audio processing apparatus provided in this disclosure can execute the audio processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0242] The following are embodiments of the audio processing model training apparatus provided in this disclosure. This apparatus and the audio processing model training method described above belong to the same inventive concept. For details not described in detail in the embodiments of the audio processing model training apparatus, please refer to the content of the audio processing model training method described in the above embodiments.
[0243] Figure 17 is a schematic diagram of the structure of a training device for an audio processing model provided in an embodiment of the present disclosure. As shown in Figure 17, the device includes: a training audio input module 610, an audio quantization feature determination module 620, an audio dequantization feature determination module 630, a prediction and reconstruction audio determination module 640, and an audio processing model training module 650.
[0244] Among them, the training audio input module 610 is used to encode the encoder in the audio processing model that has not been fully trained, and obtain audio coding features.
[0245] The audio quantization feature determination module 620 is used to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features.
[0246] The audio dequantization feature determination module 630 is used to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features.
[0247] The predictive reconstructed audio determination module 640 is used to input the audio dequantization features into the decoder in the audio processing model for decoding processing and output the predicted reconstructed audio.
[0248] The audio processing model training module 650 is used to determine a target loss value based on the training audio and the predicted reconstructed audio, and to iteratively train the untrained audio processing model based on the target loss value to obtain a trained audio processing model.
[0249] The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the training audio.
[0250] The technical solution of this embodiment solves the problem of traditional quantizers using the same bit rate by training an audio processing model equipped with a long short-time quantization network and a long short-time dequantization network, thereby reducing the computational load of the audio processing model and improving the training efficiency of the audio processing model.
[0251] In an optional embodiment, the audio processing model training module 650 includes:
[0252] The first target loss value determination unit is used to input the predicted reconstructed audio into the discriminator to obtain the output predicted classification result;
[0253] The target loss value is determined based on the training audio, the predicted reconstructed audio, and the predicted classification result.
[0254] In an optional embodiment, the discriminator includes a resolution spectrum network corresponding to at least two window lengths, wherein the resolution spectrum network includes a spectrum conversion module, a parallel frequency band feature extraction layer corresponding to at least two frequency bands, a feature concatenation layer, a global convolutional network, and a discriminative output layer; and a first target loss value determination unit, specifically used for:
[0255] The predicted and reconstructed audio is input into each resolution spectral network to obtain the output resolution classification result;
[0256] The predicted classification result is determined based on at least two resolution classification results.
[0257] In an optional embodiment, the audio processing model training module 650 includes:
[0258] The second target loss value determination unit is used to perform transient feature detection on the training audio to obtain transient audio and non-transient audio;
[0259] The transient enhancement spectral loss value is determined based on the transient audio, non-transient audio, and the predicted reconstructed audio.
[0260] Based on the transient enhancement spectrum loss value, determine the target loss value;
[0261] The loss weight corresponding to the transient audio is higher than the loss weight corresponding to the non-transient audio.
[0262] In an optional embodiment, the audio processing model training module 650 includes:
[0263] The third target loss value determination unit is used to obtain the training spectrum data corresponding to the training audio and the prediction spectrum data corresponding to the predicted reconstructed audio.
[0264] The training spectrum data and the predicted spectrum data are compared to obtain the first spectrum data and the second spectrum data corresponding to the training spectrum data.
[0265] Based on the first spectrum data, the second spectrum data, and the predicted spectrum data, determine the energy suppression spectrum loss value;
[0266] Based on the energy suppression spectrum loss value, determine the target loss value;
[0267] Wherein, each training spectrum value in the first spectrum data is less than the corresponding predicted spectrum value in the predicted spectrum data, each training spectrum value in the second spectrum data is greater than or equal to the corresponding predicted spectrum value in the predicted spectrum data, and the loss weight corresponding to the first spectrum data is higher than the loss weight corresponding to the second spectrum data.
[0268] In an optional embodiment, the audio processing model training module 650 includes:
[0269] An audio processing model training unit is used to iteratively train an untrained audio processing model based on the target loss value to obtain a pre-trained audio processing model.
[0270] The pre-trained audio processing model is pruned, and the resulting audio processing model is used as the untrained audio processing model.
[0271] Return to the step of encoding the trained audio input into the encoder of the untrained audio processing model to obtain the audio encoded features;
[0272] The audio processing model that has been pre-trained during the current iteration is used as the audio processing model that has been trained until the training termination condition is met.
[0273] The training apparatus for the audio processing model provided in this disclosure can execute the training method for the audio processing model provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0274] Referring now to FIG18, a schematic diagram of the structure of an electronic device (e.g., a terminal device or a server) 700 suitable for implementing embodiments of the present disclosure is shown. The terminal device in embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG18 is merely an example and should not impose any limitation on the functionality and scope of use of embodiments of the present disclosure.
[0275] As shown in Figure 18, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0276] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 18 shows electronic device 700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0277] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0278] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0279] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0280] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0281] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire at least two Internet Protocol (IP) addresses; send a node evaluation request including the at least two IP addresses to a node evaluation device, wherein the node evaluation device selects an IP address from the at least two IP addresses and returns it; and receive the IP address returned by the node evaluation device; wherein the acquired IP address indicates an edge node in a content delivery network.
[0282] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive a node evaluation request including at least two Internet Protocol (IP) addresses; select an IP address from the at least two IP addresses; and return the selected IP address; wherein the received IP address indicates an edge node in the content delivery network.
[0283] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0284] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0285] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0286] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0287] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0288] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0289] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0290] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An audio processing method, comprising: The audio input to be processed is encoded by the encoder in the pre-trained audio processing model to obtain audio encoding features; The audio coding features are input into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features. The audio quantization features are input into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features. The audio dequantization features are input into the decoder in the audio processing model for decoding, and the target reconstructed audio is output. The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed.
2. The method according to claim 1, wherein, The long-short-time quantization network includes a long-time feature extraction module and a long-short-time quantization layer. Correspondingly, the long-short-time quantization network, which inputs the audio encoded features into the quantizer of the audio processing model, performs long-short-time quantization processing to obtain audio quantized features, including: The audio encoding features are input into the long-term feature extraction module for long-term feature extraction to obtain the audio long-term features; The audio coding features and the audio long-time features are input into the long-short-time quantization layer for long-short-time quantization processing to obtain the audio quantization features.
3. The method according to claim 2, wherein, The step of inputting the audio encoding features into the long-term feature extraction module for long-term feature extraction to obtain audio long-term features includes: Based on a preset quantization step size, the audio coding features are processed into frames to obtain at least one frame-based coding feature. Long-term features are extracted for each frame coding feature to obtain frame long-term features, and the audio long-term features are determined based on at least one frame long-term feature.
4. The method according to claim 2, wherein, The long-short-time quantization layer includes a long-time quantization module, a long-time dequantization module, a short-time feature extraction module, and a short-time quantization module. Correspondingly, the step of inputting the audio encoded features and the audio long-time features into the long-short-time quantization layer for long-short-time quantization processing to obtain audio quantization features includes: The audio long-time features are input into the long-time quantization module for quantization processing to obtain long-time quantized features; The long-time quantization features are input into the long-time dequantization module for dequantization processing to obtain long-time dequantization features; The audio encoding features and the long-time dequantization features are input into the short-time feature extraction module to extract short-time features and obtain audio short-time features. The audio short-time features are input into the short-time quantization module for quantization processing to obtain short-time quantized features.
5. The method according to any one of claims 1-4, wherein, The long-short-time dequantization network includes a long-time dequantization module, a short-time dequantization module, and a feature fusion module. Correspondingly, the step of inputting the audio quantization features into the long-short-time dequantization network in the quantizer for long-short-time dequantization processing to obtain the audio dequantization features includes: The long-time quantization features are input into the long-time dequantization module for dequantization processing to obtain long-time dequantization features; The short-time quantization features are input into the short-time dequantization module for dequantization processing to obtain short-time dequantization features; The long-time dequantization features and the short-time dequantization features are input into the feature fusion module for fusion processing to obtain the audio dequantization features.
6. The method according to any one of claims 1-5, wherein, The decoder includes a convolutional network, a decoding network, and an output network connected in series. The decoding network includes at least one decoding module connected in series. The decoding module includes a convolutional activation function layer, a transposed convolutional layer, and a residual structure connected in series. The residual structure includes at least one residual unit connected in series.
7. The method according to claim 6, wherein, The residual unit includes a first activation function layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a residual output layer connected in series. The first convolutional layer and the third convolutional layer have the same number of output channels, and the second convolutional layer has more output channels than the first convolutional layer.
8. The method according to claim 6 or 7, wherein, The decoder satisfies at least one of the following conditions: 1) The Snake Beta activation function is used; 2) At least one of the convolutional structures in the convolutional network, the transposed convolutional layer, and the convolutional layer in the residual unit adopts grouped convolution; wherein, the number of groups in the convolutional structure is the common divisor of the number of input channels and the number of output channels of the convolutional structure; 3) The residual structure contains 4 residual units, and the expansion rates of each residual unit are 1, 3, 9 and 27, respectively.
9. A training method for an audio processing model, comprising: The training audio is input into the encoder of the untrained audio processing model to encode the audio and obtain audio coding features. The audio coding features are input into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features. The audio quantization features are input into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features. The audio dequantization features are input into the decoder in the audio processing model for decoding, and the predicted reconstructed audio is output. Based on the training audio and the predicted reconstructed audio, a target loss value is determined, and based on the target loss value, the untrained audio processing model is iteratively trained to obtain a trained audio processing model. The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the training audio.
10. The method according to claim 9, wherein, The step of determining the target loss value based on the training audio and the predicted reconstructed audio includes: The predicted reconstructed audio is input into the discriminator to obtain the output predicted classification result; The target loss value is determined based on the training audio, the predicted reconstructed audio, and the predicted classification result.
11. The method according to claim 10, wherein, The discriminator includes a resolution spectrum network corresponding to at least two window lengths, and the resolution spectrum network includes a spectrum conversion module, a parallel frequency band feature extraction layer corresponding to at least two frequency bands, a feature concatenation layer, a global convolutional network, and a discriminant output layer. The step of inputting the predicted reconstructed audio into the discriminator to obtain the output prediction classification result includes: The predicted and reconstructed audio is input into each resolution spectral network to obtain the output resolution classification result; The predicted classification result is determined based on at least two resolution classification results.
12. The method according to claim 9, wherein, The step of determining the target loss value based on the training audio and the predicted reconstructed audio includes: Transient feature detection is performed on the training audio to obtain transient and non-transient audio. The transient enhancement spectral loss value is determined based on the transient audio, non-transient audio, and the predicted reconstructed audio. The target loss value is determined based on the transient enhancement spectrum loss value; The loss weight corresponding to the transient audio is higher than the loss weight corresponding to the non-transient audio.
13. The method according to claim 9, wherein, The step of determining the target loss value based on the training audio and the predicted reconstructed audio includes: Obtain the training spectrum data corresponding to the training audio and the prediction spectrum data corresponding to the predicted reconstructed audio; The training spectrum data and the predicted spectrum data are compared to obtain the first spectrum data and the second spectrum data corresponding to the training spectrum data. Based on the first spectrum data, the second spectrum data, and the predicted spectrum data, determine the energy suppression spectrum loss value; The target loss value is determined based on the energy suppression spectrum loss value; Wherein, each training spectrum value in the first spectrum data is less than the corresponding predicted spectrum value in the predicted spectrum data, each training spectrum value in the second spectrum data is greater than or equal to the corresponding predicted spectrum value in the predicted spectrum data, and the loss weight corresponding to the first spectrum data is higher than the loss weight corresponding to the second spectrum data.
14. The method according to any one of claims 9-13, wherein, The step of iteratively training the untrained audio processing model to obtain a trained audio processing model based on the target loss value includes: Based on the target loss value, the untrained audio processing model is iteratively trained to obtain a pre-trained audio processing model. The pre-trained audio processing model is pruned, and the resulting audio processing model is used as the untrained audio processing model. Return to the step of encoding the trained audio input into the encoder of the untrained audio processing model to obtain the audio encoded features; Until the training termination condition is met, the pre-trained audio processing model completed in the current iteration is taken as the trained audio processing model.
15. An audio processing apparatus, comprising: The audio input module is configured to encode the pre-trained audio processing model's encoder to obtain audio encoding features. The audio quantization feature determination module is configured to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features. The audio dequantization feature determination module is configured to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features; The target reconstructed audio determination module is configured to input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the target reconstructed audio. The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the audio to be processed.
16. A training device for an audio processing model, comprising: The training audio input module is configured to encode the encoder in the audio processing model that has not been fully trained, thereby obtaining audio encoded features. The audio quantization feature determination module is configured to input the audio encoding features into the long short-time quantization network of the quantizer in the audio processing model for long short-time quantization processing to obtain the audio quantization features. The audio dequantization feature determination module is configured to input the audio quantization features into the long short-time dequantization network in the quantizer for long short-time dequantization processing to obtain the audio dequantization features; The predictive reconstructed audio determination module is configured to input the audio dequantization features into the decoder in the audio processing model for decoding processing, and output the predicted reconstructed audio. The audio processing model training module is configured to determine a target loss value based on the training audio and the predicted reconstructed audio, and to iteratively train the untrained audio processing model based on the target loss value to obtain a trained audio processing model. The audio quantization features include long-time quantization features and short-time quantization features. The number of feature frames of the long-time quantization features is less than the number of feature frames of the short-time quantization features. The long-time quantization features represent the quantization results of the common features corresponding to the audio segments obtained by dividing the training audio.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio processing method of any one of claims 1-8 and / or the training method of the audio processing model of any one of claims 9-14.
18. A computer-readable storage medium storing computer instructions, said computer instructions being configured to cause a processor to execute and implement the audio processing method of any one of claims 1-8 and / or the training method of the audio processing model of any one of claims 9-14.
19. A computer program product comprising a computer program that, when executed by a processor, implements the audio processing method according to any one of claims 1-8 and / or the training method of the audio processing model according to any one of claims 9-14.
Citation Information
Patent Citations
Audio event detection method and device
CN102486920A
Multi-description speech coding and decoding methods and systems based on linear predictive residual classified quantization
CN108109629A
Audio scene recognition method and device based on long-term and short time feature extraction
CN108305616A
Audio signal encoding device and its encoding program
JP2004301972A
Perceptually-based loss functions for audio encoding and decoding based on machine learning
US20210082444A1