Audio data processing method and apparatus, electronic device and storage medium
By acquiring the Mel spectrum of audio data and using a residual denoising diffusion model to predict high-frequency features, the problem of high-frequency component loss in audio data in existing technologies is solved, thereby improving audio quality and listening experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-05-30
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies, in scenarios such as voice calls, video conferencing, and live video streaming, reduce the audio sampling rate and compress audio data, resulting in the loss of high-frequency components, which affects audio quality and the user's listening experience.
By acquiring the Mel spectrum data of the initial audio data, the high-frequency features are predicted using the residual denoising diffusion model in the band extension module to generate enhanced Mel spectrum data. The audio restoration module then processes the data to generate optimized audio data, thereby improving audio resolution.
The high-frequency characteristics of the audio data were restored, improving audio quality and user listening experience, and generating higher-resolution optimized audio data.
Smart Images

Figure CN2025098688_07052026_PF_FP_ABST
Abstract
Description
Audio data processing methods, devices, electronic equipment and storage media
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411546262.8, filed on October 31, 2024, entitled "Audio Data Processing Method, Apparatus, Electronic Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of artificial intelligence technology, and in particular to an audio data processing method, apparatus, electronic device, and storage medium. Background Technology
[0004] Currently, in various scenarios such as voice calls, video conferencing, and live video streaming, audio data is usually compressed to form low-fidelity audio that retains only low-frequency components for subsequent encoding and decoding processing and data transmission, thereby reducing the overhead of equipment and network resources. Summary of the Invention
[0005] This disclosure provides an audio data processing method, apparatus, electronic device, and storage medium.
[0006] In a first aspect, embodiments of this disclosure provide an audio data processing method, including:
[0007] The process involves: acquiring Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution; processing the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data, wherein the bandwidth extension module includes a residual denoising diffusion model for predicting corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, thereby generating enhanced Mel spectrum data with the low-frequency features and the high-frequency features; and processing the enhanced Mel spectrum data based on an audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0008] Secondly, embodiments of this disclosure provide an audio data processing apparatus, comprising:
[0009] An acquisition unit is used to acquire Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution;
[0010] The first processing unit is used to process the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data to generate enhanced Mel spectrum data with the low-frequency features and the high-frequency features.
[0011] The second processing unit is used to process the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0012] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0013] The memory stores computer-executed instructions;
[0014] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the audio data processing method as described in the first aspect and various possible designs of the first aspect.
[0015] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the audio data processing method described in the first aspect and various possible designs of the first aspect.
[0016] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the audio data processing method described in the first aspect and various possible designs of the first aspect.
[0017] The audio data processing method, apparatus, electronic device, and storage medium provided in this embodiment acquire Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution; process the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data, wherein the bandwidth extension module includes a residual denoising diffusion model for predicting corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, thereby generating enhanced Mel spectrum data containing the low-frequency features and the high-frequency features; process the enhanced Mel spectrum data based on an audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, the second audio resolution being greater than the first audio resolution. By converting the initial audio data into Mel spectrum data, the high-frequency features of the Mel spectrum data are extended using a bandwidth extension module including a residual denoising diffusion model to form enhanced Mel spectrum data containing high-frequency features, and then the enhanced Mel spectrum data is restored to generate optimized audio data with a higher audio resolution. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 is an application scenario diagram of the audio data processing method provided in the embodiments of this disclosure;
[0020] Figure 2 is a schematic flowchart of the audio data processing method provided in an embodiment of this disclosure;
[0021] Figure 3 is a schematic diagram of a process for generating optimized audio data according to an embodiment of this disclosure;
[0022] Figure 4 is a flowchart of the specific implementation of step S102 in the embodiment shown in Figure 2;
[0023] Figure 5 is a flowchart of the specific implementation of step S1022 in the embodiment shown in Figure 4;
[0024] Figure 6 is a schematic diagram of a process for reconstructing enhanced Mel spectrum data using a residual denoising diffusion model according to an embodiment of this disclosure;
[0025] Figure 7 is a schematic flowchart of the audio data processing method provided in an embodiment of this disclosure;
[0026] Figure 8 is a schematic diagram of a first network structure provided in an embodiment of this disclosure;
[0027] Figure 9 is a flowchart of the specific implementation of step S205 in the embodiment shown in Figure 7;
[0028] Figure 10 is a structural block diagram of the audio data processing apparatus provided in an embodiment of this disclosure;
[0029] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure;
[0030] Figure 12 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0033] The application scenarios of the embodiments of this disclosure are explained below:
[0034] Figure 1 illustrates an application scenario of the audio data processing method provided in this embodiment. This method can be applied to applications (APPs) with audio acquisition and playback functions, such as video conferencing applications, live video streaming applications, and music applications. More specifically, it can be applied to applications involving spread spectrum of audio data. The executing entity in this embodiment can be a terminal device running the aforementioned application with audio acquisition and playback functions, a server deploying the server-side component of the aforementioned application, or other electronic devices performing similar functions. When the executing entity is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application. When the executing entity is a server, the server-side component of the aforementioned application with audio acquisition and playback functions can run partially or entirely on the server, executing the method provided in this embodiment on the server side, while the terminal device runs the client-side component of the application. Communication between the server and the terminal device is based on server-client communication, enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.
[0035] In some embodiments, the terminal device or server can implement the audio data processing method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run, or mini-programs embedded in any APP, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin; the specific implementation can be configured as needed. Furthermore, in implementing the audio data processing method provided in this application, the terminal device can execute the method by running computer-executable instructions or computer programs set locally, or by calling computer-executable instructions or computer programs set in an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.
[0036] Referring to Figure 1, when a voice call is made between terminal device D1 and terminal device D2, terminal device D1 first samples the user's voice to obtain the corresponding audio data. Then, after encoding and other processing steps, the audio data is sent to terminal device D2. After decoding and other processing steps on terminal device D2, it is played back, thereby achieving the purpose of playing the user's voice from terminal device D1 on terminal device D2.
[0037] Existing solutions often lead to a decrease in audio quality, affecting the user's listening experience. In application scenarios such as those shown in Figure 1, to reduce the overhead of device and network resources, audio is typically acquired at a low sampling rate on the terminal device D1 side, or the acquired audio data is compressed to obtain low-fidelity audio data with a smaller file size and lower audio resolution. This reduces the cost of subsequent data processing and transmission. However, these existing solutions result in the loss of high-frequency components in the audio data, reducing the realism and clarity of the sound, thus affecting audio quality and the user's listening experience.
[0038] This disclosure provides an audio data processing method to solve the above-mentioned problems.
[0039] Referring to Figure 2, which is a schematic flowchart of an audio data processing method provided in this embodiment, the method can be applied to a terminal device or a server. The audio data processing method includes:
[0040] Step S101: Obtain the Mel spectrum data corresponding to the initial audio data, wherein the initial audio data has a first audio resolution.
[0041] For example, referring to the application scenario diagram shown in Figure 1, this embodiment takes a terminal device as the execution subject to introduce the implementation process of the provided audio data processing method. The terminal device can be terminal device D1 shown in Figure 1, i.e., the audio data acquisition end, or terminal device D2 shown in Figure 1, i.e., the audio data playback end. Specifically, firstly, the terminal device acquires the Mel-spectrum data corresponding to the initial audio data. The Mel-spectrum is a linear transformation of the logarithmic energy spectrum based on a nonlinear Mel scale of sound frequency. It represents the short-time power spectrum of sound and can be used to represent the spectrum of short-term audio information. It can be generated by performing a short-time Fourier transform (STFT) on the initial audio data in the time domain, followed by a nonlinear scale transformation of the logarithmic energy spectrum. The principle of the Mel-spectrum is based on the logarithmic spectrum represented by a nonlinear Mel scale and its linear cosine transformation. On the Mel scale, the frequency is distributed logarithmically, which simulates the perceptual characteristics of the human ear for different frequencies, i.e., high resolution for low frequencies and high resolution for high frequencies.
[0042] In one possible implementation, the Mel spectrum data corresponding to the initial audio data is generated on other external devices, such as an external data server. The terminal device can directly receive the Mel spectrum data sent by the data server. In another possible implementation, after the terminal device obtains the initial audio data, it processes the initial audio data to generate the corresponding Mel spectrum data. The specific implementation method can be set as needed.
[0043] It should be noted that the initial audio data has a first audio resolution, which is smaller than the second audio resolution of the optimized audio data finally generated in this embodiment. In other words, the initial audio data can be considered audio data with a lower audio resolution. In other descriptions with the same meaning, this could mean that the initial audio data has lower clarity, fewer high-frequency components, or a lower sampling frequency. This is essentially due to the use of a lower audio sampling rate or downsampling of the original audio data. Because of the lower audio resolution, it has a relatively smaller data volume. Correspondingly, the Mel-spectrum data obtained after transforming the initial audio data also has the above characteristics, namely, lower audio resolution in the time domain and less obvious (missing) high-frequency features in the frequency domain.
[0044] Step S102: Process the Mel spectrum data using the band extension module to generate enhanced Mel spectrum data. The band extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, so as to generate enhanced Mel spectrum data with both low-frequency and high-frequency features.
[0045] For example, further, after obtaining the Mel spectrum data, the Mel spectrum data is input into the band extension module. After processing by the band extension module, the corresponding enhanced Mel spectrum data is output. As described earlier, the initial audio data corresponding to the Mel spectrum data has low audio resolution, therefore the high-frequency features corresponding to the generated Mel spectrum data are missing (or very small), thus failing to represent audio details. After processing by the band extension module, which uses the Residual Denoising Diffusion Models (RDDM) set in the module to spread the Mel spectrum data, it utilizes the low-frequency features of the Mel spectrum data to generate corresponding high-frequency features, thereby generating enhanced Mel spectrum data with both low-frequency and high-frequency features.
[0046] In this embodiment, the low-frequency and high-frequency characteristics of Mel-spectrum data are relative concepts. The boundary between them is determined based on the sampling frequency corresponding to the initial audio data. According to the Nyquist sampling theorem, the highest effective frequency of the audio data is half of the sampling frequency Fs, i.e., Fs / 2. Therefore, in one implementation, the low-frequency characteristics of Mel-spectrum data refer to the characteristics of frequency components with frequencies less than or equal to the highest effective frequency (Fs / 2), i.e., the true frequency characteristics; while the high-frequency characteristics refer to the characteristics of frequency components with frequencies greater than the highest effective frequency (Fs / 2), i.e., the predicted frequency characteristics.
[0047] Figure 3 is a schematic diagram of a process for generating optimized audio data according to an embodiment of this disclosure. As shown in Figure 3, after obtaining the Mel spectrum data corresponding to the initial audio data, the Mel spectrum data is input into a band extension module for processing. The maximum frequency of the initial Mel spectrum data is N Hz, and the total duration is T seconds. For example, N is 2048, meaning the maximum frequency is 2048 Hz, and T is 10, meaning the total duration is 10 seconds. After inputting the Mel spectrum data into the band extension module, the band extension module predicts based on the frequency characteristics below 2048 Hz, extending the frequency characteristics of the 2048 Hz-4096 Hz band (maximum frequency is 2N), i.e., high-frequency characteristics, thereby generating enhanced Mel spectrum data with both low-frequency and high-frequency characteristics. As shown in the figure, the generated enhanced Mel spectrum data is equivalent to the Mel spectrum data of the input band extension module. It primarily contains frequency characteristics in the 2048Hz-4096Hz band, meaning it includes more high-frequency features. Therefore, the optimized audio data obtained based on this enhanced Mel spectrum data has better sound detail and higher sound quality, improving the user's listening experience. The aforementioned step of extending high-frequency features is implemented by the residual denoising diffusion model in the band extension module. The residual denoising diffusion model is a model used in image processing in existing technologies. In this embodiment, the residual denoising diffusion model is used to process the Mel spectrum data to achieve the purpose of band extension.
[0048] Furthermore, in one possible implementation, the residual denoising diffusion model includes a first network structure and a second network structure. The first network structure estimates residual information, which characterizes the difference between the noisy Mel-spectrum data and the target Mel-spectrum data with high-frequency features. The second network structure estimates noise information, which characterizes the noise components in the current Mel-spectrum data. Accordingly, as shown in Figure 4, the specific implementation of step S102 includes:
[0049] Step S1021: Add Gaussian noise to the Mel spectrum data to generate model input data;
[0050] Step S1022: Iteratively process the model input data through the first network structure and the second network structure to reconstruct the model input data into enhanced Mel spectrum data.
[0051] For example, firstly, the residual denoising diffusion model in this embodiment includes two independent network structures, a first network structure and a second network structure, used to simulate residual diffusion and noise diffusion. During inference, the first network structure can be used to estimate residual information, i.e., information representing the difference between the noisy Mel spectrum data and the target Mel spectrum data with high-frequency characteristics. The second network structure can be used to estimate noise information, i.e., information representing the noise component in the current Mel spectrum data. The process of processing the Mel spectrum data through the residual denoising diffusion model is the reverse process of the diffusion process. Therefore, it is first necessary to add Gaussian noise to the Mel spectrum data to generate noisy Mel spectrum data, i.e., model input data. Then, the model input data is input into the residual denoising diffusion model in the band extension module. The first and second network structures in the residual denoising diffusion model are used to iteratively process the model input data in multiple rounds until it is reconstructed into enhanced Mel spectrum data. More specifically, exemplarily, the first network structure and the second network structure are configured in parallel. The first network structure estimates the residual information and inputs it into the second network structure. The second network structure combines the residual information to estimate the noise information and removes the noise information from the current model input data (Mel spectrum data containing noise), i.e., the reverse process of noise diffusion, to obtain updated Mel spectrum data containing less noise. This process is repeated until the target number of iterations or other preset stopping conditions are reached, thus obtaining enhanced Mel spectrum data. The ability of the first and second network structures to estimate residual information and noise information is obtained by pre-training the residual denoising diffusion model. The specific implementation method for training the residual denoising diffusion model will not be described in this embodiment.
[0052] In one possible implementation, as shown in Figure 5, the specific implementation of step S1022 includes:
[0053] Step S1022-1: Estimate the residual information of the input data of the model through the first network structure, and estimate the noise information of the input data of the model through the second network structure;
[0054] Step S1022-2: Remove residual information and noise information from the model input data to obtain updated model input data;
[0055] Step S1022-3: If the loop stopping condition is not met, return to step S1022-1.
[0056] Step S1022-4: If the loop stopping condition is met, then based on the updated model input data, the enhanced Mel spectrum data is obtained.
[0057] Figure 6 is a schematic diagram illustrating the process of reconstructing enhanced Mel spectrum data using a residual denoising diffusion model according to an embodiment of this disclosure. The process is described below with reference to Figure 6. Exemplarily, the first network structure processes the noisy model input data to estimate the noise information corresponding to the model input data, while the second network structure processes the noisy model input data to estimate the residual information of the model input data. Then, the residual information and noise information are removed from the model input data to obtain updated model input data. The noise information is used to guide the denoising of the noisy model input data, while the residual information guides the model input data to converge towards enhanced Mel spectrum data. After multiple rounds of updating the model input data based on the noise information and residual information, the model input data is gradually reconstructed into enhanced Mel spectrum data containing high-frequency features. For example, after M iterations, the loop stops, and the updated model input data generated in the current round is used as the enhanced Mel spectrum data. For example, after obtaining updated model input data for testing, if the updated model input data exhibits specific spectral characteristics, the loop termination condition is considered met. In this case, the updated model input data generated in the current round is used as the enhanced Mel spectrum data. Otherwise, the loop steps described above are repeated.
[0058] In this embodiment, the model input data containing Gaussian noise is processed in multiple rounds using a residual denoising diffusion model to gradually remove the noise from the model input data. The residual information is then used to guide the generation of high-frequency features, ultimately resulting in enhanced Mel spectrum data containing both low-frequency and high-frequency features. This achieves the spread spectrum of the Mel spectrum data and ultimately improves the audio resolution of the audio data.
[0059] Step S103: Process the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0060] For example, after obtaining the enhanced Mel spectrum data, a pre-set audio restoration module is used to process the enhanced Mel spectrum data. This audio restoration module contains specific data processing logic that restores the Mel spectrum data to time-domain data. By processing the enhanced Mel spectrum data through the audio restoration module, corresponding optimized audio data is obtained. Since the optimized audio data is obtained based on the enhanced Mel spectrum data containing high-frequency features, it has higher frequency characteristics, i.e., a second audio resolution greater than the first audio resolution. Subsequently, depending on the specific technical scenario, the optimized audio data can be played. Because it has a higher audio resolution than the initial audio data, it offers improved sound details and quality, resulting in a more realistic sound.
[0061] The audio restoration module may include inverse transformation processing logic for the Mel spectrum, that is, by performing inverse Fourier transform and other processing steps on the Mel spectrum, it is converted into a time-domain signal, thereby generating optimized audio data. In another possible implementation, the audio restoration module includes a pre-trained generative adversarial network (GAN), which can generate the corresponding time-domain signal based on the Mel spectrum. The specific implementation principle and training process are not described in this embodiment.
[0062] In this embodiment, Mel spectrum data corresponding to initial audio data is acquired, wherein the initial audio data has a first audio resolution. A bandwidth extension module is used to process the Mel spectrum data to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, used to predict corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, thereby generating enhanced Mel spectrum data with both low-frequency and high-frequency features. An audio restoration module processes the enhanced Mel spectrum data to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution. By converting the initial audio data into Mel spectrum data, the high-frequency features of the Mel spectrum data are extended using a bandwidth extension module including a residual denoising diffusion model to form enhanced Mel spectrum data containing high-frequency features. Then, the enhanced Mel spectrum data is restored to generate optimized audio data with a higher audio resolution, thereby achieving spectrum spreading of the initial audio data. This results in optimized audio data with improved sound details and texture, enhancing the user's listening experience.
[0063] Referring to Figure 7, which is a schematic flowchart of the audio data processing method provided in this embodiment, this embodiment further refines steps S102-S103 based on the embodiment shown in Figure 2. The audio data processing method includes:
[0064] Step S201: Obtain the Mel spectrum data corresponding to the initial audio data, wherein the initial audio data has a first audio resolution.
[0065] Step S202: Add Gaussian noise to the Mel spectrum data to generate model input data.
[0066] Step S203: Obtain the first hyperparameter of the first network structure and the second hyperparameter of the second network structure. The first hyperparameter is used to characterize the residual diffusion rate corresponding to the first network structure, and the second hyperparameter is used to indicate the noise diffusion rate corresponding to the second network structure.
[0067] Step S204: Iteratively process the model input data using the first network structure and the corresponding first hyperparameters, and the second network structure and the corresponding second hyperparameters, to reconstruct the model input data into enhanced Mel spectrum data.
[0068] For example, in one possible implementation, both the first and second network structures are U-shaped network (Unet) structures. One is used to estimate residual information, i.e., the difference between the current noisy model input data and the enhanced Mel spectrum data with high-frequency features, while the other is used to estimate noise information, i.e., the noise level of the current noisy model input data. After training the first and second network structures, they can be equipped with the corresponding abilities to estimate residual information and noise information. Simultaneously, the first and second network structures have corresponding hyperparameters, namely, a first hyperparameter of the first network structure and a second hyperparameter of the second network structure. Both the first and second hyperparameters can be sets of multiple independent hyperparameters, represented in the form of a hyperparameter table. The first and second hyperparameters are configured before or after model training. The first hyperparameter characterizes the residual propagation rate corresponding to the first network structure, and the second hyperparameter indicates the noise propagation rate corresponding to the second network structure. By adjusting the first and second hyperparameters and the corresponding number of iterations, different audio data spread spectrum tasks can be applied. By using a first network structure and its corresponding first hyperparameters, and a second network structure and its corresponding second hyperparameters, the model input data is iteratively processed to achieve more accurate control over the denoising process of the model input data, resulting in more precise enhanced Mel spectrum data.
[0069] In one possible implementation, the process of iteratively processing the model input data to reconstruct the model input data into enhanced Mel spectrum data can be represented by the following equation (1):
[0070] Among them, I t This is the current model input data, I t-1This is the updated model input data. and These are the first and second hyperparameters corresponding to the current diffusion node. and These are the first and second hyperparameters of the previous diffusion node, respectively; I res Residual information, ∈ θ This represents noise information. The first and second hyperparameters corresponding to the aforementioned diffusion nodes are both suppression values. The above formula can be used to update the model input data in multiple rounds until it is reconstructed into enhanced Mel spectrum data.
[0071] Further, exemplarily, the first network structure is implemented using a Diffusion Transformer (DiT) architecture based on a multi-head attention mechanism. Figure 8 is a schematic diagram of the structure of a first network structure provided in an embodiment of this disclosure. As shown in Figure 8, the first network structure sequentially includes a multi-head attention layer, a one-dimensional convolutional layer (Conv1d), a Gaussian error linear unit layer (GELU), a random deactivation layer (Dropout), and a one-dimensional convolutional layer (Conv1d). The above-mentioned first network structure is implemented based on the Diffusion Transformer architecture. After training the above-mentioned first network structure, the estimation of residual information can be achieved.
[0072] Step S205: The enhanced Mel spectrum data is processed by the generator in the generative adversarial network to obtain optimized audio data.
[0073] For example, a generative adversarial network consists of a generator and a discriminator. After adversarial training, the generator can convert enhanced Mel-spectral data into corresponding temporal data, i.e., optimize audio data. For example, the generator includes a mapping layer, a transpose layer, an upsampling layer, and a convolutional layer arranged sequentially. As shown in Figure 9, the specific implementation of step S205 includes:
[0074] Step S2051: The enhanced Mel spectrum data is mapped in a high-dimensional space through a mapping layer to obtain enhanced spectrum features.
[0075] Step S2052: Perform a one-dimensional transposed convolution operation on the enhanced spectral features through the transposed layer, and then use the upsampling layer to upsample to obtain the upsampled features.
[0076] Step S2053: Perform one-dimensional convolution operation on the upsampled features through a convolutional structure layer to obtain optimized audio data.
[0077] For example, the mapping layer, or one-dimensional convolutional layer, performs a high-dimensional spatial mapping on the enhanced Mel spectral data through one-dimensional convolution, obtaining feature data with higher dimensions, i.e., enhanced spectral features. Next, a transposed layer performs a one-dimensional transposed convolution operation on the enhanced spectral features. The transposed layer is implemented based on transposed one-dimensional convolution. Then, the transposed result is upsampled, for example, by increasing the number of data points through interpolation, to obtain upsampled features. This step can be repeated several times until the desired upsampled features are obtained. Finally, a convolutional structure layer processes the upsampled features. This convolutional structure layer can also be implemented based on one-dimensional convolution, mapping the upsampled features back to the corresponding data, i.e., optimizing the audio data.
[0078] Further, optionally, before obtaining the Mel spectrum data corresponding to the initial audio data in this embodiment, the method further includes:
[0079] Step S200: Train the residual denoising diffusion model in the band extension module.
[0080] Specifically, exemplarily, step S200 is implemented in the following ways:
[0081] Step S2001: Obtain sample audio data with high-frequency features, and generate initial sample data based on the sample audio data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data.
[0082] Step S2002: Add Gaussian noise to the initial sample data to form noisy audio data, and use the noisy audio data and sample audio data as training samples to train the initial model and obtain the residual denoising diffusion model.
[0083] For example, based on the previous embodiments, the residual denoising diffusion model is a pre-trained model. The trained residual denoising diffusion model can convert Mel spectrum data into enhanced Mel spectrum data, thereby achieving spectrum expansion. In one possible implementation, before using the residual denoising diffusion model to perform spectrum expansion on the Mel spectrum data, this embodiment trains the residual denoising diffusion model. Specifically, firstly, sample audio data with high-frequency characteristics is acquired, i.e., audio data with high sampling rate and high audio resolution. Then, the sample audio data is used to generate corresponding initial sample data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data, for example, by low-pass filtering the sample audio data or by downsampling the sample audio data. Then, Gaussian noise is added to the initial sample data to form noisy audio data, and the noisy audio data and the corresponding sample audio data are combined to form training samples, wherein the noisy audio data corresponds to the model input and the sample audio data corresponds to the model output. The initial model is trained sequentially to enable it to generate enhanced Mel spectrum data based on the Mel spectrum data.
[0084] In this embodiment, the implementation of step S201 is the same as that of step S101 in the embodiment shown in FIG2 of this disclosure, and will not be described in detail here.
[0085] Corresponding to the audio data processing method in the above embodiments, Figure 10 is a structural block diagram of the audio data processing apparatus provided in this disclosure embodiment. The method described in the above embodiments can be executed by this audio data processing apparatus, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.
[0086] For ease of illustration, only the parts relevant to embodiments of this disclosure are shown. Referring to FIG10, the audio data processing apparatus 3 includes:
[0087] The acquisition unit 31 is used to acquire Mel spectrum data corresponding to the initial audio data, wherein the initial audio data has a first audio resolution;
[0088] The first processing unit 32 is used to process Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data to generate enhanced Mel spectrum data with both low-frequency and high-frequency features.
[0089] The second processing unit 33 is used to process the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0090] According to one or more embodiments of this disclosure, the residual denoising diffusion model includes a first network structure and a second network structure, wherein the first network structure estimates residual information, which is used to characterize the difference between noisy Mel spectrum data and target Mel spectrum data with high-frequency characteristics; the second network structure estimates noise information, which is used to characterize the noise component in the current Mel spectrum data; and the first processing unit 32, when processing the Mel spectrum data using the bandwidth extension module to generate enhanced Mel spectrum data, specifically performs the following: adding Gaussian noise to the Mel spectrum data to generate model input data; and iteratively processing the model input data through the first network structure and the second network structure to reconstruct the model input data into enhanced Mel spectrum data.
[0091] According to one or more embodiments of this disclosure, when the first processing unit 32 iteratively processes the model input data through a first network structure and a second network structure to reconstruct the model input data into enhanced Mel spectrum data, it is specifically configured to: estimate the residual information of the model input data through the first network structure and estimate the noise information of the model input data through the second network structure; remove the residual information and noise information from the model input data to obtain updated model input data; if the loop stopping condition is not met, return to the steps of estimating the residual information of the model input data through the first network structure and estimating the noise information of the model input data through the second network structure; if the loop stopping condition is met, obtain enhanced Mel spectrum data based on the updated model input data.
[0092] According to one or more embodiments of this disclosure, a first network structure has a first hyperparameter, which is used to characterize the residual diffusion rate corresponding to the first network structure; a second network structure has a second hyperparameter, which is used to indicate the noise diffusion rate corresponding to the second network structure; the first processing unit 32 is further configured to: acquire the first hyperparameter and the second hyperparameter; when the first processing unit 32 iteratively processes the model input data through the first network structure and the second network structure to reconstruct the model input data into enhanced Mel spectrum data, it is specifically configured to: iteratively process the model input data through the first network structure and the corresponding first hyperparameter, and the second network structure and the corresponding second hyperparameter, to reconstruct the model input data into enhanced Mel spectrum data.
[0093] According to one or more embodiments of this disclosure, the first network structure is implemented through a diffusion transformer architecture based on a multi-head attention mechanism.
[0094] According to one or more embodiments of this disclosure, the audio restoration module includes a generative adversarial network; the second processing unit 33 is specifically used to: process the enhanced Mel spectrum data through a generator in the generative adversarial network to obtain optimized audio data.
[0095] According to one or more embodiments of this disclosure, the generator includes a mapping layer, a transpose layer, an upsampling layer, and a convolutional structure layer arranged sequentially. When the second processing unit 33 processes the enhanced Mel spectrum data through the generator in the generative adversarial network to obtain optimized audio data, it is specifically used to: map the enhanced Mel spectrum data into a high-dimensional space through the mapping layer to obtain enhanced spectral features; perform a one-dimensional transpose convolution operation on the enhanced spectral features through the upsampling layer and then upsample them to obtain upsampled features; and perform a one-dimensional convolution operation on the upsampled features through the convolutional structure layer to obtain optimized audio data.
[0096] According to one or more embodiments of this disclosure, the first processing unit 32 is further configured to: acquire sample audio data with high-frequency features, and generate initial sample data based on the sample audio data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data; add Gaussian noise to the initial sample data to form noisy audio data, and use the noisy audio data and the sample audio data as training samples to train the initial model to obtain a residual denoising diffusion model.
[0097] The acquisition unit 31, the first processing unit 32, and the second processing unit 33 are connected in sequence. The audio data processing device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0098] Figure 11 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 11, the electronic device 4 includes:
[0099] Processor 41, and memory 42 communicatively connected to processor 41;
[0100] Memory 42 stores instructions executed by the computer;
[0101] The processor 41 executes computer execution instructions stored in the memory 42 to implement the audio data processing method in the embodiments shown in Figures 2-9.
[0102] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0103] The relevant explanations can be understood by referring to the descriptions and effects of the steps in the embodiments corresponding to Figures 2-9, which will not be elaborated on here.
[0104] This disclosure provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, these instructions are used to implement the audio data processing method provided in any of the embodiments corresponding to Figures 2-9 of this disclosure.
[0105] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the audio data processing method provided in any of the embodiments corresponding to Figures 2-9 of this disclosure.
[0106] To implement the above embodiments, this disclosure also provides an electronic device.
[0107] Referring to Figure 12, a schematic diagram of the structure of an electronic device 900 suitable for implementing embodiments of the present disclosure is shown. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 12 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.
[0108] As shown in Figure 12, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0109] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 shows electronic device 900 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0110] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0111] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0112] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0113] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0114] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0116] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.
[0117] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] In a first aspect, according to one or more embodiments of this disclosure, an audio data processing method is provided, comprising:
[0120] The process involves: acquiring Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution; processing the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data, wherein the bandwidth extension module includes a residual denoising diffusion model for predicting corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, thereby generating enhanced Mel spectrum data with the low-frequency features and the high-frequency features; and processing the enhanced Mel spectrum data based on an audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0121] According to one or more embodiments of this disclosure, the residual denoising diffusion model includes a first network structure and a second network structure, wherein the first network structure estimates residual information, which is used to characterize the difference between noisy Mel-spectrum data and target Mel-spectrum data with high-frequency features; the second network structure estimates noise information, which is used to characterize the noise component in the current Mel-spectrum data; the step of processing the Mel-spectrum data using a bandwidth extension module to generate enhanced Mel-spectrum data includes: adding Gaussian noise to the Mel-spectrum data to generate model input data; and iteratively processing the model input data through the first network structure and the second network structure to reconstruct the enhanced Mel-spectrum data.
[0122] According to one or more embodiments of this disclosure, the model input data is iteratively processed through a first network structure and a second network structure to reconstruct the enhanced Mel spectrum data, including: estimating residual information of the model input data through the first network structure and estimating noise information of the model input data through the second network structure; removing the residual information and the noise information from the model input data to obtain updated model input data; if the loop stopping condition is not met, returning to the steps of estimating the residual information of the model input data through the first network structure and estimating the noise information of the model input data through the second network structure; if the loop stopping condition is met, obtaining the enhanced Mel spectrum data based on the updated model input data.
[0123] According to one or more embodiments of this disclosure, the first network structure has a first hyperparameter, which characterizes the residual propagation rate corresponding to the first network structure; the second network structure has a second hyperparameter, which indicates the noise propagation rate corresponding to the second network structure; the method further includes: acquiring the first hyperparameter and the second hyperparameter; the step of iteratively processing the model input data through the first network structure and the second network structure to reconstruct the model input data into the enhanced Mel spectrum data includes: iteratively processing the model input data through the first network structure and the corresponding first hyperparameter, and the second network structure and the corresponding second hyperparameter, to reconstruct the model input data into the enhanced Mel spectrum data.
[0124] According to one or more embodiments of this disclosure, the first network structure is implemented using a diffusion converter architecture based on a multi-head attention mechanism.
[0125] According to one or more embodiments of this disclosure, the audio restoration module includes a generative adversarial network; processing the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data includes: processing the enhanced Mel spectrum data through a generator in the generative adversarial network to obtain the optimized audio data.
[0126] According to one or more embodiments of this disclosure, the generator includes a mapping layer, a transpose layer, an upsampling layer, and a convolutional structure layer arranged sequentially. The generator in the generative adversarial network processes the enhanced Mel-spectrum data to obtain the optimized audio data, including: mapping the enhanced Mel-spectrum data to a high-dimensional space through the mapping layer to obtain enhanced spectral features; performing a one-dimensional transpose convolution operation on the enhanced spectral features through the upsampling layer and then upsampling to obtain upsampled features; and performing a one-dimensional convolution operation on the upsampled features through the convolutional structure layer to obtain the optimized audio data.
[0127] According to one or more embodiments of this disclosure, the method further includes: acquiring sample audio data with high-frequency features, and generating initial sample data based on the sample audio data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data; adding Gaussian noise to the initial sample data to form noisy audio data, and using the noisy audio data and the sample audio data as training samples to train the initial model to obtain the residual denoising diffusion model.
[0128] Secondly, according to one or more embodiments of the present disclosure, an audio data processing apparatus is provided, comprising:
[0129] An acquisition unit is used to acquire Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution;
[0130] The first processing unit is used to process the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data to generate enhanced Mel spectrum data with the low-frequency features and the high-frequency features.
[0131] The second processing unit is used to process the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
[0132] According to one or more embodiments of this disclosure, the residual denoising diffusion model includes a first network structure and a second network structure, wherein the first network structure estimates residual information, which is used to characterize the difference between noisy Mel spectrum data and target Mel spectrum data with high-frequency features; the second network structure estimates noise information, which is used to characterize the noise components in the current Mel spectrum data; and when the first processing unit processes the Mel spectrum data using the bandwidth extension module to generate enhanced Mel spectrum data, it is specifically used to: add Gaussian noise to the Mel spectrum data to generate model input data; and iteratively process the model input data through the first network structure and the second network structure to reconstruct the enhanced Mel spectrum data.
[0133] According to one or more embodiments of this disclosure, when the first processing unit iteratively processes the model input data through the first network structure and the second network structure to reconstruct the enhanced Mel spectrum data, it is specifically configured to: estimate the residual information of the model input data through the first network structure and estimate the noise information of the model input data through the second network structure; remove the residual information and the noise information from the model input data to obtain updated model input data; if the loop stopping condition is not met, return to the steps of estimating the residual information of the model input data through the first network structure and estimating the noise information of the model input data through the second network structure; if the loop stopping condition is met, obtain the enhanced Mel spectrum data based on the updated model input data.
[0134] According to one or more embodiments of this disclosure, the first network structure has a first hyperparameter, which characterizes the residual diffusion rate corresponding to the first network structure; the second network structure has a second hyperparameter, which indicates the noise diffusion rate corresponding to the second network structure; the first processing unit is further configured to: acquire the first hyperparameter and the second hyperparameter; when the first processing unit iteratively processes the model input data through the first network structure and the second network structure to reconstruct the model input data into the enhanced Mel spectrum data, it is specifically configured to: iteratively process the model input data through the first network structure and the corresponding first hyperparameter, and the second network structure and the corresponding second hyperparameter, to reconstruct the model input data into the enhanced Mel spectrum data.
[0135] According to one or more embodiments of this disclosure, the first network structure is implemented using a diffusion converter architecture based on a multi-head attention mechanism.
[0136] According to one or more embodiments of this disclosure, the audio restoration module includes a generative adversarial network; the second processing unit is specifically used to: process the enhanced Mel spectrum data through a generator in the generative adversarial network to obtain the optimized audio data.
[0137] According to one or more embodiments of this disclosure, the generator includes a mapping layer, a transpose layer, an upsampling layer, and a convolutional structure layer arranged sequentially. When the second processing unit processes the enhanced Mel-spectrum data through the generator in the generative adversarial network to obtain the optimized audio data, it specifically performs the following: mapping the enhanced Mel-spectrum data in a high-dimensional space through the mapping layer to obtain enhanced spectral features; performing a one-dimensional transpose convolution operation on the enhanced spectral features through the upsampling layer and then upsampling to obtain upsampled features; and performing a one-dimensional convolution operation on the upsampled features through the convolutional structure layer to obtain the optimized audio data.
[0138] According to one or more embodiments of this disclosure, the first processing unit is further configured to: acquire sample audio data with high-frequency features, and generate initial sample data based on the sample audio data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data; add Gaussian noise to the initial sample data to form noisy audio data, and use the noisy audio data and the sample audio data as training samples to train the initial model to obtain the residual denoising diffusion model.
[0139] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0140] The memory stores computer-executed instructions;
[0141] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the audio data processing method as described in the first aspect and various possible designs of the first aspect.
[0142] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, which, when executed by a processor, implement the audio data processing method described in the first aspect and various possible designs of the first aspect.
[0143] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the audio data processing method as described in the first aspect and various possible designs of the first aspect.
[0144] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0145] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0146] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An audio data processing method, comprising: Obtain Mel spectrum data corresponding to the initial audio data, wherein the initial audio data has a first audio resolution; The Mel spectrum data is processed using a bandwidth extension module to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data, so as to generate enhanced Mel spectrum data with the low-frequency features and the high-frequency features. The enhanced Mel spectrum data is processed by the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
2. The method according to claim 1, wherein the residual denoising diffusion model comprises a first network structure and a second network structure, wherein, The first network structure estimates residual information, which is used to characterize the difference between noisy Mel spectrum data and target Mel spectrum data with high-frequency features; the second network structure estimates noise information, which is used to characterize the noise component in the current Mel spectrum data. The process of using the bandwidth extension module to process the Mel spectrum data and generate enhanced Mel spectrum data includes: Gaussian noise is added to the Mel spectrum data to generate model input data; The model input data is iteratively processed through the first network structure and the second network structure to reconstruct the enhanced Mel spectrum data.
3. The method according to claim 2, wherein iteratively processing the model input data through the first network structure and the second network structure to reconstruct the enhanced Mel spectrum data comprises: The residual information of the model input data is estimated using the first network structure, and the noise information of the model input data is estimated using the second network structure. The residual information and the noise information are removed from the model input data to obtain the updated model input data; If the loop termination condition is not met, return to the steps of estimating the residual information of the model input data through the first network structure and estimating the noise information of the model input data through the second network structure; If the loop termination condition is met, the enhanced Mel spectrum data is obtained based on the updated model input data.
4. The method according to claim 2, wherein the first network structure has a first hyperparameter, the first hyperparameter being used to characterize the residual diffusion rate corresponding to the first network structure; The second network structure has a second hyperparameter, which is used to indicate the noise propagation rate corresponding to the second network structure; The method further includes: obtaining the first hyperparameter and the second hyperparameter; The step of iteratively processing the model input data through the first network structure and the second network structure to reconstruct the enhanced Mel spectrum data includes: The model input data is iteratively processed using the first network structure and the corresponding first hyperparameter, and the second network structure and the corresponding second hyperparameter, to reconstruct the enhanced Mel spectrum data.
5. The method according to claim 2, wherein the first network structure is implemented by a diffusion transformer architecture based on a multi-head attention mechanism.
6. The method according to claim 1, wherein the audio restoration module includes a generative adversarial network; Based on the processing of the enhanced Mel spectrum data by the audio restoration module, optimized audio data is obtained, including: The enhanced Mel spectrum data is processed by the generator in the generative adversarial network to obtain the optimized audio data.
7. The method according to claim 6, wherein the generator comprises a mapping layer, a transpose layer, an upsampling layer, and a convolutional structure layer arranged sequentially, and processes the enhanced Mel-spectrum data through the generator in the generative adversarial network to obtain the optimized audio data, including: The enhanced Mel spectrum data is mapped in a high-dimensional space through the mapping layer to obtain enhanced spectrum features; The enhanced spectral features are subjected to a one-dimensional transposed convolution operation through the upsampling layer, and then upsampled to obtain upsampled features. The optimized audio data is obtained by performing a one-dimensional convolution operation on the upsampled features through the convolutional structure layer.
8. The method according to claim 1, wherein the method further comprises: Acquire sample audio data with high-frequency features, and generate initial sample data based on the sample audio data, wherein the initial sample data is sound data generated using the low-frequency components of the sample audio data; Gaussian noise is added to the initial sample data to form noisy audio data, and the noisy audio data and the sample audio data are used as training samples to train the initial model to obtain the residual denoising diffusion model.
9. An audio data processing apparatus, comprising: An acquisition unit is used to acquire Mel spectrum data corresponding to initial audio data, wherein the initial audio data has a first audio resolution; The first processing unit is used to process the Mel spectrum data using a bandwidth extension module to generate enhanced Mel spectrum data. The bandwidth extension module includes a residual denoising diffusion model, which is used to predict the corresponding high-frequency features based on the low-frequency features of the Mel spectrum data to generate enhanced Mel spectrum data with the low-frequency features and the high-frequency features. The second processing unit is used to process the enhanced Mel spectrum data based on the audio restoration module to obtain optimized audio data, wherein the optimized audio data has a second audio resolution, which is greater than the first audio resolution.
10. An electronic device, comprising: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the audio data processing method as described in any one of claims 1 to 8.
11. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the audio data processing method as described in any one of claims 1 to 8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the audio data processing method as described in any one of claims 1 to 8.