Audio compression method and apparatus
By determining the quantization layer based on the complexity of audio segments and combining it with generative adversarial training, the audio compression model is optimized. This solves the problems of information loss and storage cost in the processing of complex sound details in existing audio compression models, achieving the best balance between sound quality and compression rate.
Patent Information
- Application Number
- CN202411852007.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Existing audio compression models are prone to information loss when compressing complex sound details, and a uniform compression rate cannot balance sound quality and storage costs.
By determining the complexity of the audio segment to be compressed, an adaptive quantization layer is selected for quantization processing. Combined with generative adversarial training and a discriminative model, the audio compression model is optimized to adapt to audio segments of different complexities.
While ensuring sound quality, the compression rate can be flexibly adjusted to adapt to audio segments of varying complexity, achieving the best balance between sound quality and compression rate, and reducing storage space and transmission time.
Smart Images

Figure CN119785806B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to an audio compression method and apparatus. Background Technology
[0002] Audio compression refers to the process of reducing the storage space occupied by audio data through specific encoding techniques, while maintaining the original audio quality as much as possible without significant damage.
[0003] The goal is to construct audio compression models using unsupervised learning methods and generate low-bitrate compressed audio based on these models. However, audio compression models apply a uniform compression rate to the input audio, which may result in the loss of some audio information for input audio containing complex sound details. Summary of the Invention
[0004] This invention provides an audio compression method and apparatus to address the deficiencies in the prior art.
[0005] This invention provides an audio compression method, comprising the following steps:
[0006] Determine the complexity of the audio segment to be compressed;
[0007] Based on the audio compression model, the complexity of the audio segment to be compressed is applied to compress the audio segment to obtain compressed audio;
[0008] The audio compression model includes multiple quantization layers. The audio compression model is used to encode the audio segment to be compressed to obtain coded features, then quantize the coded features using a target quantization layer, and decode the quantized coded features to obtain decoded features. The compressed audio is determined based on the decoded features. The target quantization layer is at least one quantization layer selected from the multiple quantization layers based on the complexity of the audio segment to be compressed.
[0009] According to an audio compression method provided by the present invention, the audio compression model is trained based on the following steps:
[0010] The original audio segments of the sample are encoded to obtain the sample encoding features;
[0011] Based on the complexity of the original audio segment of the sample, a sample quantization layer is determined from the plurality of quantization layers to quantize the encoded features of the sample;
[0012] The sample encoding features are quantized based on the sample quantization layer to obtain sample quantized features;
[0013] The sample quantization features are decoded to obtain sample decoding features, and the sample compressed audio corresponding to the original audio segment of the sample is generated based on the sample decoding features;
[0014] Based on the difference between the compressed audio sample and the original audio segment of the sample, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0015] According to an audio compression method provided by the present invention, the method involves training an initial generation model of the audio compression model based on the difference between the sample compressed audio and the original sample audio segment to obtain the audio compression model, comprising:
[0016] Based on the differences between the sample decoding features and the sample encoding features, as well as the differences between the sample compressed audio and the sample original audio segment, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0017] According to an audio compression method provided by the present invention, the audio compression model is obtained by performing generative adversarial training in conjunction with a discriminative model;
[0018] The process of training an initial generation model for the audio compression model based on the difference between the compressed audio sample and the original audio segment to obtain the audio compression model includes:
[0019] The compressed audio sample is input into the initial discrimination model to obtain the first classification result output by the initial discrimination model.
[0020] The original audio segment of the sample is input into the initial discrimination model to obtain the second classification result output by the initial discrimination model;
[0021] Based on the differences between the first classification result and the classification label corresponding to the compressed audio sample, the differences between the second classification result and the classification label corresponding to the original audio segment of the sample, and the differences between the compressed audio sample and the original audio segment of the sample, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0022] According to an audio compression method provided by the present invention, the initial discrimination model includes multiple branches, each branch is used to extract audio features of the input audio, and to determine the classification result based on the audio features of the input audio, wherein the scale of the audio features extracted by each branch is different;
[0023] The audio compression model is obtained by jointly training the initial generation model and the initial discrimination model based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, including:
[0024] Based on the differences between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples, the differences between the first classification result and the corresponding classification label of the compressed audio samples, the differences between the second classification result and the corresponding classification label of the original audio segments, and the differences between the compressed audio samples and the original audio segments, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0025] According to an audio compression method provided by the present invention, determining the complexity of the audio segment to be compressed includes:
[0026] The complexity of the audio segment to be compressed is determined based on the audio complexity classification model.
[0027] The audio complexity classification model is trained based on multiple sample audio segments and the complexity labels of each sample audio segment.
[0028] According to an audio compression method provided by the present invention, the complexity label of each sample audio segment is determined based on the following steps:
[0029] Determine the scores of multiple different quality metrics for each audio segment in the samples;
[0030] Based on preset weights, the scores of multiple different quality indicators for each sample audio segment are weighted and summed to determine the quality score of each sample audio segment.
[0031] The complexity label of each audio sample is determined based on its quality score and the total number of quantization layers.
[0032] The present invention also provides an audio compression device, comprising the following modules:
[0033] The determination unit is used to determine the complexity of the audio segment to be compressed;
[0034] The compression unit is used to compress the audio segment to be compressed based on the audio compression model and by applying the complexity of the audio segment to be compressed, so as to obtain compressed audio.
[0035] The audio compression model includes multiple quantization layers. The audio compression model is used to encode the audio segment to be compressed to obtain coded features, then quantize the coded features using a target quantization layer, and decode the quantized coded features to obtain decoded features. The compressed audio is determined based on the decoded features. The target quantization layer is at least one quantization layer selected from the multiple quantization layers based on the complexity of the audio segment to be compressed.
[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the audio compression methods described above.
[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio compression method as described above.
[0038] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the audio compression methods described above.
[0039] The audio compression method and apparatus provided by this invention select one or more suitable quantization layers from multiple quantization layers for quantization processing based on the complexity of the audio segment to be compressed. This not only ensures that high-complexity audio segments retain more audio quality details during quantization, avoiding excessive audio quality loss, but also allows low-complexity audio segments to achieve higher compression ratios by selecting higher quantization layers. Thus, while maintaining acceptable audio quality, it significantly reduces storage space and transmission time, improving the overall efficiency and practicality of audio compression. In other words, the audio compression model in this invention can flexibly adjust the compression ratio while ensuring audio quality, enabling the audio compression model to adapt to audio segments of different complexities and achieve the optimal balance between audio quality and compression ratio. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the audio compression method provided by the present invention.
[0042] Figure 2 This is a flowchart illustrating the audio compression model training method provided by the present invention.
[0043] Figure 3 This is a flowchart illustrating another audio compression method provided by the present invention.
[0044] Figure 4 This is a schematic diagram of the audio compression device provided by the present invention.
[0045] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0047] Currently, most audio compression models employ an encoding / decoding model. This model typically includes an encoder, a residual vector quantizer (RVQ), and a decoder. Both the encoder and decoder are fully convolutional structures. The encoder maps the input raw audio to a sequence of embeddings at a lower sampling rate. The RVQ quantizes the embedding sequence, and the decoder reconstructs the original audio based on the quantized embedding sequence, resulting in compressed audio. The training objective of the encoding / decoding model is to minimize the difference between the compressed and original audio. However, using a uniform compression rate to compress the original audio can lead to several problems. Using a consistently high compression rate may result in the loss of high-frequency details in the original audio, and speech may become blurred due to over-compression. Conversely, using a consistently low compression rate, while preserving more audio information, significantly increases storage and transmission costs.
[0048] In response, this invention provides an audio compression method. Figure 1 This is a flowchart illustrating the audio compression method provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110 and 120.
[0049] Step 110: Determine the complexity of the audio segment to be compressed.
[0050] Specifically, the audio segment to be compressed is the audio segment that needs to be compressed. The audio segment to be compressed can be collected by a sound pickup device, input by the user, or obtained from the Internet through web crawling. This embodiment of the invention does not make specific limitations on this.
[0051] The complexity of an audio segment to be compressed refers to the degree of variation in the audio segment in the time and frequency domains, including characteristics such as dynamic range, frequency components, sonic details, and redundancy. Audio segments with high complexity typically contain more sonic details and variations, while audio segments with low complexity are relatively smooth and simple.
[0052] For highly complex audio segments, more detail is typically retained in the compressed audio to maintain sound quality. Therefore, highly complex audio segments can use a lower compression rate to preserve this detail. Conversely, less complex audio segments usually retain less detail in the compressed audio. Therefore, less complex audio segments can use a higher compression rate to reduce storage and transmission costs.
[0053] Therefore, different compression requirements correspond to different levels of complexity in the audio segments to be compressed. Thus, this embodiment of the invention first determines the complexity of the audio segment to be compressed and then adapts a corresponding compression strategy based on that complexity, thereby achieving efficient and high-quality audio compression. The complexity of the audio segment to be compressed can be obtained through time-domain and frequency-domain analysis of the audio signal of the segment.
[0054] Step 120: Based on the audio compression model, apply the complexity of the audio segment to be compressed to compress the audio segment and obtain compressed audio;
[0055] The audio compression model includes multiple quantization layers. After encoding the audio segment to be compressed to obtain coded features, the audio compression model uses a target quantization layer to quantize the coded features and decodes the quantized coded features to obtain decoded features. The compressed audio is determined based on the decoded features. The target quantization layer is at least one quantization layer selected from multiple quantization layers based on the complexity of the audio segment to be compressed.
[0056] Specifically, the quantization layer is used to convert continuous sample values in the audio segment to be compressed into discrete values. The quantization layer is a key parameter in the audio compression model, determining the precision and stride of the quantization of the encoded features. Higher quantization layers (i.e., later quantization layers) typically have larger quantization strides and lower quantization precision, resulting in more information being discarded during quantization and achieving a higher compression ratio. Conversely, lower quantization layers (i.e., earlier quantization layers) have smaller quantization strides and higher quantization precision, preserving more audio quality details, but at a correspondingly lower compression ratio.
[0057] For highly complex audio segments to be compressed (such as music segments containing rich details and dynamic range), a lower quantization layer can be selected to retain more audio details. Although the compression ratio will be slightly lower, the fidelity of the sound quality is guaranteed. For less complex audio segments (such as voice calls or simple background music), a higher quantization layer can be selected to significantly improve the compression ratio while maintaining an acceptable level of sound quality, thus reducing the use of storage space or transmission bandwidth.
[0058] Based on this, after inputting the audio segment to be compressed and its complexity into the audio compression model, the audio compression model selects at least one quantization layer from multiple quantization layers as the target quantization layer based on the complexity. As the complexity increases, the number of corresponding target quantization layers decreases.
[0059] Audio compression models typically select quantization layers sequentially (from front to back) because lower quantization layers retain more sonic detail, while higher quantization layers achieve higher compression ratios. By selecting quantization layers sequentially, the audio compression model can gradually adjust the compression ratio while maintaining sound quality. To reduce compression (i.e., improve sound quality), increase the number of quantization layers selected (i.e., select more lower quantization layers); conversely, to increase compression (i.e., reduce sound quality), decrease the number of quantization layers selected (i.e., select fewer higher quantization layers).
[0060] For example, for a highly complex audio segment 1, the first m quantization layers can be selected as the target quantization layers; for a less complex audio segment 2, the first n quantization layers can be selected as the target quantization layers, where m <n。
[0061] In summary, when selecting quantization layers, audio compression models can balance the relationship between sound quality and compression ratio based on the complexity of the audio segment to be compressed. As complexity increases, the audio compression model tends to choose lower quantization layers to preserve sound quality. Simultaneously, by sequentially selecting quantization layers and adjusting their number, the audio compression model can flexibly adjust the compression ratio while maintaining sound quality. This allows the audio compression model to adapt to audio segments of varying complexity and achieve the optimal balance between sound quality and compression ratio.
[0062] As an optional embodiment, after determining the target quantization layer corresponding to the audio segment to be compressed, the audio segment to be compressed can be encoded to obtain coded features, and the target quantization layer can quantize the coded features to convert continuous coded feature values into discrete, finite numbers of values, thereby reducing the number of bits required to represent the audio segment to be compressed and thus reducing the amount of data.
[0063] After quantization, the audio compression model decodes the quantized coded features. Decoding is the reverse process of encoding; it converts the quantized coded features back into a format or representation that more closely resembles the audio segment to be compressed. In this process, the decoder attempts to recover lost information from the coded features to approximate the quality of the audio segment as closely as possible. Once decoding is complete, based on the decoded features, the audio compression model can determine the final compressed audio.
[0064] In this process, determining the target quantization layer can be performed synchronously with encoding the audio segment to be compressed. After determining the target quantization layer, the encoded features can be quantized based on the target quantization layer.
[0065] Before compressing an audio segment using an audio compression model and applying its complexity, the audio compression model can be pre-trained. This can be achieved through the following steps: First, collect a large number of sample original audio segments and their complexities, and manually label them to determine their corresponding compressed audio tags. Then, train the initial model based on the sample original audio segments, their complexities, and their corresponding compressed audio tags to obtain the audio compression model.
[0066] Furthermore, based on the audio compression model, the complexity of the audio segment to be compressed is applied, and before compressing the audio segment to obtain the compressed audio, the audio segment to be compressed can also be resampled.
[0067] The audio compression method provided in this invention selects one or more suitable quantization layers from multiple quantization layers based on the complexity of the audio segment to be compressed. This not only ensures that high-complexity audio segments retain more audio quality details during quantization, avoiding excessive audio quality loss, but also allows low-complexity audio segments to achieve higher compression ratios by selecting higher quantization layers. This significantly reduces storage space and transmission time while maintaining acceptable audio quality, improving the overall efficiency and practicality of audio compression. In other words, the audio compression model in this invention can flexibly adjust the compression ratio while ensuring audio quality, enabling it to adapt to audio segments of varying complexity and achieve an optimal balance between audio quality and compression ratio.
[0068] Based on the above embodiments, the audio compression model is trained using the following steps:
[0069] The original audio segments of the sample are encoded to obtain the sample encoding features;
[0070] The sample quantization layer used to quantize the encoded features of the sample is determined from multiple quantization layers based on the complexity of the original audio segment of the sample.
[0071] The sample encoded features are quantized based on the sample quantization layer to obtain the sample quantized features;
[0072] The sample quantization features are decoded to obtain sample decoding features, and the sample compressed audio corresponding to the original audio segment of the sample is generated based on the sample decoding features;
[0073] Based on the differences between the compressed audio samples and the original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0074] Specifically, the original audio segments of a sample refer to the original audio data used to train the audio compression model, which can be representative segments selected from a large amount of audio data. Sample encoded features are the features obtained after encoding the original audio segments of the samples.
[0075] Furthermore, the complexity of the original audio segment refers to the degree of variation of the original audio segment in the time and frequency domains, including characteristics such as dynamic range, frequency components, sonic details, and redundancy. Audio segments with high complexity typically contain more sonic details and variations, while those with low complexity are relatively smooth and simple. The sample quantization layer refers to the quantization layer selected based on the complexity of the original audio segment, used to quantize the encoded features of the sample to obtain the quantized features.
[0076] After obtaining the sample quantization features, the sample quantization features are decoded to obtain the sample decoding features. Based on the sample decoding features, the sample compressed audio corresponding to the original audio segment is generated. The sample compressed audio is smaller than the original audio segment, but the training objective is to require that the sample compressed audio retain as little of the main information of the original audio segment as possible.
[0077] In this regard, considering that the difference between the compressed audio sample and the original audio sample segment is used to characterize the integrity and fidelity of the original audio sample segment retained by the compressed audio sample, the smaller the difference between the two, the more successfully the compressed audio sample retains most of the important information of the original audio sample segment, thus achieving efficient compression and good sound quality preservation. Optionally, the difference between the compressed audio sample and the original audio sample segment can usually be measured by quantifying sound quality loss indicators (such as signal-to-noise ratio, distortion, etc.).
[0078] To address this, embodiments of the present invention can determine the training loss value based on the difference between the compressed audio sample and the original audio sample, and update the parameters of the initial generation model based on the training loss value to obtain the audio compression model. The initial generation model can be understood as the starting point for training the audio compression model; it is an untrained model whose parameters are randomly generated, and the process of training the initial generation model is the process of optimizing its parameters.
[0079] Based on any of the above embodiments, the initial generation model of the audio compression model is trained based on the difference between the compressed audio sample and the original audio sample to obtain the audio compression model, including:
[0080] Based on the differences between sample decoding features and sample encoding features, as well as the differences between sample compressed audio and sample original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0081] Specifically, information loss is unavoidable during audio compression. Information loss can occur in both the encoding and decoding stages (the process of converting the original audio into a compressed format) and the decoding stage (the process of restoring the compressed format back to the original audio). Considering only the differences between sample compressed audio and sample original audio segments only reflects the overall difference between the decoded audio and the original audio, and cannot accurately capture the specific information loss during the encoding process.
[0082] For example, during the encoding process, information loss may occur in the sample encoded features, while the features obtained during decoding are based on the sample quantized features for feature reconstruction. If the difference between the decoded sample compressed audio and the original sample audio segment is small, it indicates that the decoding process retains important features of the original audio segment, thus enabling the reconstruction of a sample compressed audio with minimal difference from the original sample audio segment based on the sample decoded features. However, if the difference between the sample encoded features and the sample decoded features is large, it indicates that information loss occurred during the encoding process, meaning that the sample encoded features lost some important features of the original sample audio segment.
[0083] To address this, embodiments of the present invention train the initial generation model of the audio compression model based on the differences between sample decoding features and sample encoding features, as well as the differences between sample compressed audio and sample original audio segments. In other words, during the training process, embodiments of the present invention not only consider the overall differences between the decoded audio and the original audio, but also consider the information loss during the encoding and decoding processes. This allows for more precise optimization of the compression algorithm, enabling the trained audio compression model to produce compressed audio with better sound quality and smaller data volume.
[0084] Among them, the sample encoding feature is obtained by encoding the original audio segment of the sample, and thus the sample encoding feature retains the main features of the original audio segment of the sample. The sample decoding feature is the feature reconstructed based on the sample quantization feature. If the difference between the sample decoding feature and the sample encoding feature is small, it indicates that the main features of the original audio segment of the sample are preserved as much as possible during the encoding and decoding process.
[0085] Therefore, it can be seen that by combining the differences between sample decoding features and sample encoding features, as well as the differences between sample compressed audio and sample original audio segments for training, the performance of the audio compression model can be comprehensively improved, making it more efficient, accurate and reliable in practical applications.
[0086] Based on any of the above embodiments, the audio compression model is obtained by performing generative adversarial training in conjunction with the discriminative model;
[0087] Based on the differences between the compressed audio samples and the original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model, which includes:
[0088] The compressed audio sample is input into the initial discriminant model to obtain the first classification result output by the initial discriminant model;
[0089] The original audio segment of the sample is input into the initial discrimination model to obtain the second classification result output by the initial discrimination model;
[0090] Based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0091] Specifically, the audio compression model is obtained by performing generative adversarial training in conjunction with the discriminative model. In other words, the audio compression model can be regarded as a generative model. Its training goal is to make the generated sample compressed audio approximate the original audio segment, that is, to retain the main features of the original audio segment as much as possible, while reducing the amount of data.
[0092] In generative adversarial training, the discriminative model continuously challenges the audio compression model, forcing it to generate more diverse and indistinguishable audio signals. This helps to enhance the audio compression model's adaptability to different types of audio data, thereby improving its robustness and generalization ability.
[0093] Without introducing a discriminative model, audio compression models may easily fall into overfitting, that is, they become overly dependent on training data and cannot handle untrained audio segments well. This weakens the generalization ability of the audio compression model and makes it unable to adapt to the diverse needs of practical applications.
[0094] To address this, this invention introduces a discriminative model for generative adversarial training to improve the quality, robustness, and generalization ability of the audio compression model, and to promote innovation in audio compression models. The specific training method is as follows:
[0095] The compressed audio sample is input into the initial discriminant model to obtain the first classification result output by the initial discriminant model. Similarly, the original audio segment of the sample is input into the initial discriminant model to obtain the second classification result output by the initial discriminant model. The initial discriminant model can be understood as the starting point for training the discriminant model, and its parameters are randomized.
[0096] Based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0097] The first classification result characterizes the initial discriminant model's predicted classification of the compressed audio sample. The smaller the difference between the first classification result and the corresponding classification label of the compressed audio sample, the more accurately the initial discriminant model can identify the compressed audio sample. The second classification result characterizes the initial discriminant model's predicted classification of the original audio segment. The smaller the difference between the second classification result and the corresponding classification label of the original audio segment, the more accurately the initial discriminant model can identify the original audio segment. The difference between the compressed audio sample and the original audio segment is used to characterize whether the initial generative model can generate compressed audio samples that are as close as possible to the original audio segments. That is, the training goal of the initial generative model is to generate compressed audio samples that are as close as possible to the original audio segments to deceive the discriminant model. The training goal of the initial discriminant model is to accurately distinguish between the original audio segments and the compressed audio samples. By alternately optimizing the initial generative model and the initial discriminant model, the audio compression model and the discriminant model are obtained.
[0098] Optionally, based on the difference between the first classification result and the corresponding classification label of the compressed audio sample, and the difference between the second classification result and the corresponding classification label of the original audio segment, the discriminator loss value can be determined. Based on the difference between the compressed audio sample and the original audio segment, the generator loss value can be determined. Based on the discriminator loss value and the generator loss value, the adversarial training loss value can be determined. Based on the adversarial training loss value, the parameters of the initial generator model and the initial discriminator model are updated to obtain the audio compression model and the discriminator model.
[0099] Based on any of the above embodiments, the initial discrimination model includes multiple branches, each branch is used to extract audio features of the input audio, and the classification result is determined based on the audio features of the input audio. The scale of the audio features extracted by each branch is different.
[0100] Based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the sample decoding features and the sample encoding features, the initial generation model and the initial discriminant model are jointly trained to obtain an audio compression model, including:
[0101] Based on the differences between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples, the differences between the first classification result and the corresponding classification label of the compressed audio samples, the differences between the second classification result and the corresponding classification label of the original audio segments, and the differences between the compressed audio samples and the original audio segments, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0102] Specifically, considering that audio usually contains a variety of audio features, such as frequency, amplitude, timbre, etc., these audio features may differ at different scales. That is, for the same audio, the audio features at different scales contain different audio detail information. Since the audio detail information contains key acoustic features and subtle changes in sound events, capturing this audio detail information can help the initial discrimination model more accurately identify the original audio segments and compressed audio samples.
[0103] When the initial discriminative model can accurately identify subtle differences between the original audio clip and the compressed audio clip, it can provide more valuable feedback to the initial generative model. The initial generative model can then use this feedback to fine-tune its generated compressed audio clips, making them closer to the quality of the original audio clip and reducing distortion or loss introduced by compression.
[0104] Furthermore, the initial discrimination model comprises multiple branches, each extracting audio features at a different scale, meaning each branch focuses on a different frequency range of the input audio. In other words, some branches can be configured to focus on high-frequency components, while others focus on low- or mid-frequency components, thus helping to preserve details in the high-frequency range while reducing unnecessary smoothing.
[0105] Based on this, the initial discrimination model introduced in this embodiment of the invention includes multiple branches. Each branch is used to extract audio features from the input audio and determine the classification result based on the audio features of the input audio. The scale of the audio features extracted by each branch is different. For example, the initial discrimination model may include m identical network branches. Each network branch is used to perform Short-Time Fourier Transform (STFT) on the input audio features at different scales to obtain the corresponding audio features. Optionally, m can be equal to 5, and the window length scale of the STFT corresponding to each network branch includes 2048, 1024, 512, 256, and 128. The input audio of each network branch can first pass through a 2D convolutional layer (3*8 kernel, 32 channels), followed by 2D convolutions with increasing dilation rates in the time dimension (dilation rates of 1, 2, and 4 respectively), and finally two 2D convolutional layers to provide the final prediction result.
[0106] Based on the introduction of multiple branches, the initial generation model and the initial discrimination model are jointly trained based on the differences between the audio features of the original audio segments extracted by each branch and the audio features of the compressed audio samples, the differences between the first classification result and the corresponding classification label of the compressed audio samples, the differences between the second classification result and the corresponding classification label of the original audio segments, and the differences between the compressed audio samples and the original audio segments, to obtain the audio compression model.
[0107] The difference between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples is used to evaluate whether the compressed audio samples generated by the initial generation model closely approximate the original audio segments from the perspective of audio details. The smaller the difference between the two, the more likely the initial generation model can generate compressed audio samples that closely approximate the original audio segments.
[0108] Based on any of the above embodiments, determining the complexity of the audio segment to be compressed includes:
[0109] The complexity of the audio segment to be compressed is determined based on the audio complexity classification model.
[0110] The audio complexity classification model is trained based on multiple sample audio segments and the complexity labels of each sample audio segment.
[0111] Specifically, the audio complexity classification model can be a lightweight neural network model (such as MobileNet), thereby reducing the number of parameters and computational complexity. Before determining the complexity of the audio segment to be compressed based on the audio complexity classification model, the audio complexity classification model can be pre-trained. This can be achieved by performing the following steps: First, collect a large number of sample audio segments and manually label the complexity of each sample audio segment. Then, train the initial model based on multiple sample audio segments and their complexity labels to obtain the audio complexity classification model. The sample audio segments here can be the same as or different from the original sample audio segments mentioned above; this embodiment of the invention does not specifically limit this. The audio complexity classification model can be a gating unit.
[0112] Optionally, the sample audio segments can include speech segments, music segments, ambient sound segments, etc. Speech segments include the DAPS dataset, Common Voice dataset, and VCTK dataset; music segments include the MUSDB dataset and Jamendo dataset; and ambient sound segments include the AudioSet dataset. Furthermore, sample audio segments can also be obtained by segmenting larger audio files, such as splitting the audio file into segments of 15ms each to obtain multiple sample audio segments.
[0113] Before training based on sample audio segments, each sample audio segment can be normalized to the range of -24 dB LUFS. This ensures that the loudness of different sample audio segments remains consistent, thus avoiding model training bias caused by loudness differences. Furthermore, the normalized sample audio segments can be resampled to 24 kHz, which reduces data dimensionality and computational cost, accelerating the model training process and reducing resource consumption. The collected sample audio segments can be divided into training, validation, and test sets. The test set can be a mixture of data from the DAPS, CommonVoice, VCTK, MUSDB, Jamendo, and AudioSet datasets, selected in a certain proportion.
[0114] Based on any of the above embodiments, the complexity label of each sample audio segment is determined based on the following steps:
[0115] Determine the scores of multiple different quality metrics for each audio segment in the samples;
[0116] Based on preset weights, the scores of multiple different quality indicators for each sample audio segment are weighted and summed to determine the quality score of each sample audio segment.
[0117] The complexity label of each audio sample is determined based on its quality score and the total number of quantization layers.
[0118] Specifically, different quality index scores are scores obtained by evaluating the quality of sample audio segments in different dimensions, such as clarity, purity, dynamic range, and timbre. Optionally, quality indicators may include PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligence), MOS (Mean Opinion Score), and WER (Word Error Rate).
[0119] After determining the scores for each quality indicator, the scores of multiple different quality indicators for each sample audio segment are summed based on preset weights to determine the quality score of each sample audio segment. The preset weights can be determined based on the importance of each quality indicator; the higher the importance, the greater the corresponding weight.
[0120] As an optional implementation, the quality score distribution of each sample audio segment can be statistically analyzed, and the complexity level of each sample audio segment can be divided according to the quality score distribution and the total number of quantization layers. For example, if the total number of quantization layers is 8, then 8 complexity levels can be divided according to the quality score distribution, with each complexity level corresponding to a corresponding quality score range. Thus, the complexity label of each sample audio segment can be determined based on its quality score.
[0121] Based on any of the above embodiments Figure 2 This is a flowchart illustrating the audio compression model training method provided by the present invention, as shown below. Figure 2As shown, the initial generative model of the audio compression model includes an encoder, multiple quantization layers (such as Q1, Q2, Q3, Q4, ...), and a decoder. The audio compression model is obtained through generative adversarial training in conjunction with the discriminative model. The encoder can consist of one one-dimensional convolutional layer and four convolutional blocks. The kernel size of the one-dimensional convolutional layer is 7, and the number of channels is 32. Each convolutional block consists of one residual unit and one stride convolutional downsampling layer, with the kernel size being twice the stride. The residual unit contains two convolutions with a kernel size of 3 and one skip connection, finally connected to a one-dimensional convolutional layer with a kernel size of 7. The strides of the four convolutional blocks are 2, 4, 5, and 8, respectively, to achieve effective feature extraction. The decoder adopts a mirror configuration of the encoder, replacing the convolutional layers with deconvolutional layers to reconstruct the audio signal. The quantization layer can use RVQ (Residual Vector Quantizer) to achieve effective quantization of audio features. The initial discriminant model consists of multiple branches, each used to extract audio features from the input audio and determine the classification result based on the audio features of the input audio. The scale of the audio features extracted by each branch is different.
[0122] First, the original audio segments of the samples are obtained and input into an encoder for encoding to obtain the sample encoding features. Simultaneously, the original audio segments are input into an audio complexity classification model to obtain the complexity of the original audio segments.
[0123] Next, based on the complexity of the original audio segment, the sample quantization layer is determined from multiple quantization layers to quantize the encoded features of the sample. The encoded features of the sample are then input into the sample quantization layer for quantization to obtain the quantized features of the sample. For example, if the total number of quantization layers is 8, the complexity is normalized to 8 threshold ranges. If the complexity of the original audio segment is 5, it belongs to the threshold range of [5,6]. Therefore, the first 5 RVQ layers are used as the sample quantization layers.
[0124] After obtaining the sample quantization features, the sample quantization features are input into the decoder for decoding to obtain the sample decoding features, and the sample compressed audio corresponding to the original audio segment of the sample is generated based on the sample decoding features.
[0125] Furthermore, the compressed audio samples are input into the initial discriminant model to obtain the audio features of the compressed audio samples extracted by each branch of the initial discriminant model, as well as the first classification result output by the initial discriminant model. Simultaneously, the original audio segments of the samples are input into the initial discriminant model to obtain the audio features of the original audio segments of the samples extracted by each branch of the initial discriminant model, as well as the second classification result output by the initial discriminant model.
[0126] Based on the differences between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples, the multi-scale spectral reconstruction loss value is determined; based on the differences between the first classification result and the corresponding classification label of the compressed audio samples, and the differences between the second classification result and the corresponding classification label of the original audio segments, the discrimination loss value is determined; based on the differences between the sample decoding features and the sample encoding features, the feature matching loss value is determined; based on the differences between the compressed audio samples and the original audio segments, the generation loss value is determined.
[0127] Based on the multi-scale spectral reconstruction loss, discrimination loss, feature matching loss, and generation loss, the training loss value is determined, and the initial generation model and the initial discrimination model are jointly trained based on the training loss value to obtain the audio compression model and the discrimination model.
[0128] The audio complexity classification model is trained based on multiple sample audio segments and their complexity labels. The complexity label of each sample audio segment is determined by the following steps: determining multiple different quality index scores for each sample audio segment; weighting and summing the multiple different quality index scores for each sample audio segment based on preset weights to determine the quality score of each sample audio segment; and determining the complexity label of each sample audio segment based on its quality score and the total number of quantization layers.
[0129] Figure 3 This is a flowchart illustrating another audio compression method provided by the present invention, as shown below. Figure 3 As shown, the audio segment to be compressed is resampled, and the resampled audio segment to be compressed is input into the audio compression model. The audio compression model compresses the audio segment to be compressed based on the complexity of the audio segment to be compressed, and obtains the compressed audio.
[0130] The audio compression apparatus provided by the present invention will be described below. The audio compression apparatus described below can be referred to in correspondence with the audio compression method described above.
[0131] Based on any of the above embodiments Figure 4 This is a schematic diagram of the audio compression device provided by the present invention, as shown below. Figure 4 As shown, the device includes:
[0132] The determination unit 410 is used to determine the complexity of the audio segment to be compressed;
[0133] Compression unit 420 is used to compress the audio segment to be compressed based on the audio compression model and the complexity of the audio segment to be compressed, so as to obtain compressed audio.
[0134] The audio compression model includes multiple quantization layers. After encoding the audio segment to be compressed to obtain coded features, the audio compression model uses a target quantization layer to quantize the coded features and decodes the quantized coded features to obtain decoded features. The compressed audio is determined based on the decoded features. The target quantization layer is at least one quantization layer selected from multiple quantization layers based on the complexity of the audio segment to be compressed.
[0135] Based on any of the above embodiments, the audio compression model is trained using the following steps:
[0136] The original audio segments of the sample are encoded to obtain the sample encoding features;
[0137] The sample quantization layer used to quantize the encoded features of the sample is determined from multiple quantization layers based on the complexity of the original audio segment of the sample.
[0138] The sample encoded features are quantized based on the sample quantization layer to obtain the sample quantized features;
[0139] The sample quantization features are decoded to obtain sample decoding features, and the sample compressed audio corresponding to the original audio segment of the sample is generated based on the sample decoding features;
[0140] Based on the differences between the compressed audio samples and the original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0141] Based on any of the above embodiments, the initial generation model of the audio compression model is trained based on the difference between the compressed audio sample and the original audio sample to obtain the audio compression model, including:
[0142] Based on the differences between sample decoding features and sample encoding features, as well as the differences between sample compressed audio and sample original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model.
[0143] Based on any of the above embodiments, the audio compression model is obtained by performing generative adversarial training in conjunction with the discriminative model;
[0144] Based on the differences between the compressed audio samples and the original audio segments, the initial generation model of the audio compression model is trained to obtain the audio compression model, which includes:
[0145] The compressed audio sample is input into the initial discriminant model to obtain the first classification result output by the initial discriminant model;
[0146] The original audio segment of the sample is input into the initial discrimination model to obtain the second classification result output by the initial discrimination model;
[0147] Based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0148] Based on any of the above embodiments, the initial discrimination model includes multiple branches, each branch is used to extract audio features of the input audio, and the classification result is determined based on the audio features of the input audio. The scale of the audio features extracted by each branch is different.
[0149] Based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, the initial generation model and the initial discriminant model are jointly trained to obtain the audio compression model, including:
[0150] Based on the differences between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples, the differences between the first classification result and the corresponding classification label of the compressed audio samples, the differences between the second classification result and the corresponding classification label of the original audio segments, and the differences between the compressed audio samples and the original audio segments, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
[0151] Based on any of the above embodiments, determining the complexity of the audio segment to be compressed includes:
[0152] The complexity of the audio segment to be compressed is determined based on the audio complexity classification model.
[0153] The audio complexity classification model is trained based on multiple sample audio segments and the complexity labels of each sample audio segment.
[0154] Based on any of the above embodiments, the complexity label of each sample audio segment is determined based on the following steps:
[0155] Determine the scores of multiple different quality metrics for each audio segment in the samples;
[0156] Based on preset weights, the scores of multiple different quality indicators for each sample audio segment are weighted and summed to determine the quality score of each sample audio segment.
[0157] The complexity label of each audio sample is determined based on its quality score and the total number of quantization layers.
[0158] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute an audio compression method, which includes: determining the complexity of an audio segment to be compressed; compressing the audio segment to be compressed based on the complexity of the audio segment to be compressed using an audio compression model to obtain compressed audio; the audio compression model includes multiple quantization layers, which are used to encode the audio segment to be compressed to obtain encoded features, then quantize the encoded features using a target quantization layer, and decode the quantized encoded features to obtain decoded features, and determine the compressed audio based on the decoded features; the target quantization layer is at least one quantization layer selected from the multiple quantization layers based on the complexity of the audio segment to be compressed.
[0159] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0160] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the audio compression method provided by the above methods. The method includes: determining the complexity of an audio segment to be compressed; compressing the audio segment to be compressed based on the complexity of the audio segment to be compressed using an audio compression model to obtain compressed audio; the audio compression model includes multiple quantization layers, which are used to encode the audio segment to be compressed to obtain encoded features, quantize the encoded features using a target quantization layer, and decode the quantized encoded features to obtain decoded features, and determine the compressed audio based on the decoded features; the target quantization layer is at least one quantization layer selected from the multiple quantization layers based on the complexity of the audio segment to be compressed.
[0161] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the audio compression method provided by the methods described above. The method includes: determining the complexity of an audio segment to be compressed; compressing the audio segment to be compressed based on an audio compression model, applying the complexity of the audio segment to be compressed, to obtain compressed audio; the audio compression model includes multiple quantization layers, which are used to encode the audio segment to be compressed to obtain encoded features, then quantize the encoded features using a target quantization layer, and decode the quantized encoded features to obtain decoded features, and determine the compressed audio based on the decoded features; the target quantization layer is at least one quantization layer selected from the multiple quantization layers based on the complexity of the audio segment to be compressed.
[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0163] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio compression method, characterized in that, include: Determine the complexity of the audio segment to be compressed; Based on the audio compression model, the complexity of the audio segment to be compressed is applied to compress the audio segment to obtain compressed audio; The audio compression model includes multiple quantization layers. The audio compression model is used to encode the audio segment to be compressed to obtain coded features, then quantize the coded features using a target quantization layer, and then decode the quantized coded features to obtain decoded features. The compressed audio is then determined based on the decoded features. The target quantization layer is at least one quantization layer selected from the plurality of quantization layers based on the complexity of the audio segment to be compressed; The audio compression model is obtained by using the original audio segments as samples and performing generative adversarial training on a joint discriminative model. The audio compression model is trained based on the following steps: The compressed audio sample is input into the initial discrimination model to obtain the first classification result output by the initial discrimination model; the compressed audio sample is the audio obtained by compressing the original audio segment of the sample by the initial generation model of the audio compression model; The original audio segment of the sample is input into the initial discrimination model to obtain the second classification result output by the initial discrimination model; Based on the differences between the first classification result and the classification label corresponding to the compressed audio sample, the differences between the second classification result and the classification label corresponding to the original audio segment of the sample, and the differences between the compressed audio sample and the original audio segment of the sample, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
2. The audio compression method according to claim 1, characterized in that, The audio compression model is trained based on the following steps: The original audio segments of the sample are encoded to obtain the sample encoding features; Based on the complexity of the original audio segment of the sample, a sample quantization layer is determined from the plurality of quantization layers to quantize the encoded features of the sample; The sample encoding features are quantized based on the sample quantization layer to obtain sample quantized features; The sample quantization features are decoded to obtain sample decoding features, and the sample compressed audio corresponding to the original audio segment of the sample is generated based on the sample decoding features; Based on the difference between the compressed audio sample and the original audio segment of the sample, the initial generation model of the audio compression model is trained to obtain the audio compression model.
3. The audio compression method according to claim 2, characterized in that, The process of training an initial generation model for the audio compression model based on the difference between the compressed audio sample and the original audio segment to obtain the audio compression model includes: Based on the differences between the sample decoding features and the sample encoding features, as well as the differences between the sample compressed audio and the sample original audio segment, the initial generation model of the audio compression model is trained to obtain the audio compression model.
4. The audio compression method according to claim 3, characterized in that, The initial discrimination model includes multiple branches, each branch is used to extract audio features of the input audio and determine the classification result based on the audio features of the input audio. The scale of the audio features extracted by each branch is different. The audio compression model is obtained by jointly training the initial generation model and the initial discrimination model based on the differences between the first classification result and the corresponding classification label of the compressed audio sample, the differences between the second classification result and the corresponding classification label of the original audio segment, and the differences between the compressed audio sample and the original audio segment, including: Based on the differences between the audio features of the original audio segments extracted from each branch and the audio features of the compressed audio samples, the differences between the first classification result and the corresponding classification label of the compressed audio samples, the differences between the second classification result and the corresponding classification label of the original audio segments, and the differences between the compressed audio samples and the original audio segments, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
5. The audio compression method according to any one of claims 1 to 4, characterized in that, Determining the complexity of the audio segment to be compressed includes: The complexity of the audio segment to be compressed is determined based on the audio complexity classification model. The audio complexity classification model is trained based on multiple sample audio segments and the complexity labels of each sample audio segment.
6. The audio compression method according to claim 5, characterized in that, The complexity label for each sample audio segment is determined based on the following steps: Determine the scores of multiple different quality metrics for each audio segment in the samples; Based on preset weights, the scores of multiple different quality indicators for each sample audio segment are weighted and summed to determine the quality score of each sample audio segment. The complexity label of each audio sample is determined based on its quality score and the total number of quantization layers.
7. An audio compression device, characterized in that, include: The determination unit is used to determine the complexity of the audio segment to be compressed; The compression unit is used to compress the audio segment to be compressed based on the audio compression model and by applying the complexity of the audio segment to be compressed, so as to obtain compressed audio. The audio compression model includes multiple quantization layers. The audio compression model is used to encode the audio segment to be compressed to obtain coded features, then quantize the coded features using a target quantization layer, and then decode the quantized coded features to obtain decoded features. The compressed audio is then determined based on the decoded features. The target quantization layer is at least one quantization layer selected from the plurality of quantization layers based on the complexity of the audio segment to be compressed; The audio compression model is obtained by using the original audio segments as samples and performing generative adversarial training on a joint discriminative model. The audio compression model is trained based on the following steps: The compressed audio sample is input into the initial discrimination model to obtain the first classification result output by the initial discrimination model; the compressed audio sample is the audio obtained by compressing the original audio segment of the sample by the initial generation model of the audio compression model; The original audio segment of the sample is input into the initial discrimination model to obtain the second classification result output by the initial discrimination model; Based on the differences between the first classification result and the classification label corresponding to the compressed audio sample, the differences between the second classification result and the classification label corresponding to the original audio segment of the sample, and the differences between the compressed audio sample and the original audio segment of the sample, the initial generation model and the initial discrimination model are jointly trained to obtain the audio compression model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the audio compression method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the audio compression method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio compression method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Audio coding and decoding method and device based on diffusion model, storage medium and equipment
CN117577121A
Audio cloud storage method and device
CN118486315A