Audio processing methods and apparatuses, model training method and apparatus, storage medium and electronic device

By combining the conditional network model and the generative network model, the contradiction between transmission efficiency and quality in audio data compression processing is resolved, and high-quality audio data transmission and recovery at low bit rates is achieved.

WO2025213848A1PCT designated stage Publication Date: 2025-10-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140548
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2024-12-19
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

In the existing technology, audio data cannot be compressed to achieve both low bit rate and high playback quality, resulting in a contradiction between transmission efficiency and audio quality that is difficult to resolve.

Method used

A conditional network model is used to convert audio data into conditional data, and the audio data is restored at the receiving end by generating a network model. Quantization coding and super-prior models are combined to optimize data transmission, reducing the amount of transmitted data while maintaining high quality.

Benefits of technology

It achieves the goal of significantly reducing the amount of data in the transmission process while ensuring the quality of audio data, improving transmission efficiency, and improving the quality of restored audio data through compensation processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140548_16102025_PF_FP_ABST
    Figure CN2024140548_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Audio processing methods and apparatuses, a model training method and apparatus, a storage medium, and an electronic device. An audio processing method comprises: acquiring audio data to be processed (S110); on the basis of a preset conditional network model, processing said audio data to obtain conditional data corresponding to said audio data (S120); and transmitting the conditional data corresponding to said audio data to a receiving end, such that the receiving end performs audio data recovery on the conditional data corresponding to said audio data (S130). Conditional data is used to represent audio data to be processed, the data volume of the conditional data is small, and the transmission speed is high. When the conditional data is transmitted to the receiving end, the receiving end can perform audio data recovery processing on the conditional data to obtain high-fidelity audio data, thereby achieving high-quality audio data transmission while reducing the transmission volume.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, model training method, device, storage medium and electronic device

[0001] This application claims priority to Chinese Patent Application No. 202410433170.2, filed on April 10, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to an audio processing method, a model training method, a device, a storage medium and an electronic device. BACKGROUND

[0003] Digital storage and transmission of audio data are very important in the process of saving and transmitting audio signals. Generally, compression processing is performed on audio data before storage or transmission. The purpose of compression processing is to reduce the storage and transmission amount of audio data.

[0004] The low code rate of compression processing of audio data and the playback quality of audio data are mutually exclusive. In general, compressing audio data to a lower code rate means lower storage and transmission amount, but at the same time, it also leads to a decline in voice quality when playing audio data. SUMMARY

[0005] Embodiments of the present disclosure provide an audio processing method, a model training method, a device, a storage medium and an electronic device to reduce the amount of data in the transmission process while ensuring the quality of audio data.

[0006] In a first aspect, embodiments of the present disclosure provide an audio processing method, comprising:

[0007] obtaining to-be-processed audio data;

[0008] processing the to-be-processed audio data based on a pre-set conditional network model to obtain conditional data corresponding to the to-be-processed audio data;

[0009] transmitting the conditional data corresponding to the to-be-processed audio data to a receiving end to enable the receiving end to perform audio data recovery based on the conditional data corresponding to the to-be-processed audio data.

[0010] In a second aspect, embodiments of the present disclosure also provide an audio processing method, comprising:

[0011] receiving conditional information transmitted by a sending end;

[0012] obtaining noise data, processing the conditional information and the noise data based on a pre-trained generation network model to obtain target audio data.

[0013] In a third aspect, the embodiments of the present disclosure further provide a model training method, comprising:

[0014] obtaining sample audio data, and obtaining a conditional network model to be trained and a generative network model to be trained;

[0015] inputting the sample audio data into the conditional network model to be trained to obtain first audio data and conditional data output by the conditional network model to be trained;

[0016] performing noise adding processing on the sample audio data to obtain noise-added audio data, inputting the conditional data and the noise-added audio data into the generative network model to be trained to obtain generative data output by the generative network model to be trained;

[0017] generating a first loss function based on the first audio data and the sample audio data, and generating a second loss function based on the generative data;

[0018] performing parameter adjustment on the conditional network model to be trained based on the first loss function and / or the second loss function, and performing parameter adjustment on the generative network model to be trained based on the second loss function, until a trained conditional network model and a trained generative network model are obtained.

[0019] In a fourth aspect, the embodiments of the present disclosure further provide an audio processing apparatus, comprising:

[0020] an audio data obtaining module configured to obtain audio data to be processed;

[0021] a conditional data generating module configured to process the audio data to be processed based on a pre-set conditional network model to obtain conditional data corresponding to the audio data to be processed;

[0022] a conditional data transmitting module configured to transmit the conditional data corresponding to the audio data to be processed to a receiving end, so that the receiving end performs audio data recovery based on the conditional data corresponding to the audio data to be processed.

[0023] In a fifth aspect, the embodiments of the present disclosure further provide an audio processing apparatus, comprising:

[0024] a conditional information receiving module configured to receive conditional information transmitted by a sending end;

[0025] a target audio data generating module configured to obtain noise data, and process the conditional information and the noise data based on a pre-trained generative network model to obtain target audio data.

[0026] In a sixth aspect, the embodiments of the present disclosure further provide a model training module, comprising:

[0027] a sample obtaining module, configured to obtain sample audio data, and obtain a conditional network model to be trained and a generative network model to be trained;

[0028] a processing module, configured to input the sample audio data into the conditional network model to be trained, to obtain first audio data and conditional data output by the conditional network model to be trained; perform noise adding processing on the sample audio data, to obtain noise-added audio data; input the conditional data and the noise-added audio data into the generative network model to be trained, to obtain generative data output by the generative network model to be trained;

[0029] a loss function generating module, configured to generate a first loss function based on the first audio data and the sample audio data, and generate a second loss function based on the generative data;

[0030] a model parameter adjusting module, configured to perform parameter adjustment on the conditional network model to be trained based on the first loss function and / or the second loss function, and perform parameter adjustment on the generative network model to be trained based on the second loss function.

[0031] In a seventh aspect, an electronic device is provided, and the electronic device includes:

[0032] one or more processors;

[0033] a storage device configured to store one or more programs,

[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement one or more of the audio processing method and the model training method provided by any of the embodiments of the present disclosure.

[0035] In an eighth aspect, a storage medium containing computer executable instructions is provided, and the computer executable instructions, when executed by a computer processor, are configured to perform one or more of the audio processing method and the model training method provided by any of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0036] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and the sizes of the components and elements are not necessarily drawn to scale.

[0037] FIG. 1 is a flow diagram of an audio processing method provided by an embodiment of the present disclosure;

[0038] FIG. 2 is a structural schematic diagram of a conditional network model according to an embodiment of the present disclosure;

[0039] FIG. 3 is a flowchart of an audio processing method according to an embodiment of the present disclosure;

[0040] FIG. 4 is a structural schematic diagram of a generating network model according to an embodiment of the present disclosure;

[0041] FIG. 5 is a structural schematic diagram of a noise processing module according to an embodiment of the present disclosure;

[0042] FIG. 6 is a flowchart of a model training method according to an embodiment of the present disclosure;

[0043] FIG. 7 is a flowchart of a model training method according to an embodiment of the present disclosure;

[0044] FIG. 8 is a structural schematic diagram of an audio processing apparatus according to an embodiment of the present disclosure;

[0045] FIG. 9 is a structural schematic diagram of an audio processing apparatus according to an embodiment of the present disclosure;

[0046] FIG. 10 is a structural schematic diagram of a model training apparatus according to an embodiment of the present disclosure; and

[0047] FIG. 11 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thoroughly and completely understood as a whole. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0049] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0050] As used herein, the term "comprises" and its variations are open-ended, meaning "includes but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions of other terms will be given in the following description.

[0051] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the sequence or interdependence of the functions performed by these devices, modules or units.

[0052] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless otherwise explicitly indicated in the context.

[0053] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0054] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0055] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware performing the operation of the technical solutions of the present disclosure according to the prompt information.

[0056] As an optional but non-limiting implementation, in response to receiving the active request of the user, the user is sent prompt information, for example, in the form of a pop-up window, which can present the prompt information in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0057] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0058] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.

[0059] The transmission of audio data is from a sending end to a receiving end. Exemplarily, the audio data transmission can include uplink transmission and downlink transmission. For example, in the uplink transmission, the sending end is an electronic device with an audio signal collection function, such as a mobile phone, a sound pickup, etc., and the receiving end can be an edge server, etc. The sending end uploads the collected audio data to the receiving end. In the downlink transmission, the sending end can be an edge server, and the receiving end can be an electronic device with an audio playback function, such as a mobile phone, a loudspeaker, etc., to realize pulling of audio data from the edge server. In some embodiments, the sending end can be an electronic device with an audio signal collection function, and the receiving end can be an electronic device with an audio playback function. It can be understood that any electronic device can act as a sending end or a receiving end in different scenarios. For example, a mobile terminal such as a mobile phone has an audio collection function and an audio playback function, and can act as a sending end to send audio data to other electronic devices, or can act as a receiving end to receive audio signals sent by other electronic devices. In some embodiments, the sending end and the receiving end can be electronic devices with a data transmission function, which have a data sending function or a data receiving function, which are not limited here.

[0060] In the transmission process of audio data, the compression processing cannot balance the sound quality and transmission efficiency, and due to the high code rate and large data volume of audio data, in order to avoid that the compressed audio data obtained through compression processing cannot be played, the data volume of the compressed audio data retrieved is limited. To solve the above problems, the present embodiment provides an audio processing method, as shown in FIG. 1, which is a flowchart of an audio processing method provided by the present embodiment. The present embodiment is applicable to the case that the sending end performs conditional conversion on the audio data to be processed instead of compression processing. The method can be executed by an audio processing device, which can be realized in the form of software and / or hardware, and can be integrated in the sending end. Alternatively, the method can be realized by an electronic device, which can be a mobile terminal such as a mobile phone, a tablet computer, an audio collector, a PC terminal or a server, etc.

[0061] As shown in FIG. 1, the method includes:

[0062] S110, obtaining audio data to be processed.

[0063] S120, processing the audio data to be processed based on a pre-set conditional network model to obtain conditional data corresponding to the audio data to be processed.

[0064] S130, transmitting the conditional data corresponding to the audio data to be processed to a receiving end, so that the receiving end performs audio data recovery based on the conditional data corresponding to the audio data to be processed.

[0065] In this embodiment, the audio data to be processed can be offline audio data or real-time online audio data. For example, the offline audio data can be a song, music, recording data, etc., and the real-time online audio data can be online conference audio data, online telephone audio data, live audio data, etc.

[0066] It can be understood that the audio data to be processed can be complete audio data to be transmitted or a data packet in the complete audio data to be transmitted, and the data packet can include partial audio data. Taking offline audio data as an example, the complete audio data to be transmitted is divided into multiple audio segments, and the partial audio data of each audio segment is transmitted in the form of a data packet. The partial audio data of each audio segment is sequentially taken as the audio data to be processed to implement the audio processing method of this embodiment, and the multiple audio segments are sequentially transmitted. Taking real-time online audio data as an example, real-time audio data is acquired, the acquired real-time audio data is segmented, the partial audio data of each audio segment is transmitted in the form of a data packet, and the sequentially acquired partial audio data of each audio segment is taken as the audio data to be processed.

[0067] In this embodiment, the audio data to be processed is input into the conditional network model through the pre-set conditional network model to obtain the conditional data corresponding to the audio data to be processed. The conditional data can be a data type representing the audio data to be processed. Optionally, the conditional data can be vector data.

[0068] Compared with compressed audio data, the conditional data has a smaller data volume. The audio data to be processed is converted into corresponding conditional data, and the audio data to be processed is represented by the conditional data. Correspondingly, the transmission of the audio data to be processed is realized through the transmission of the conditional data, the data volume is reduced, and the transmission efficiency is improved.

[0069] It should be noted that the receiving end has the function of restoring audio data, and can obtain target audio data based on the condition data transmitted by the sending end. Optionally, the sending end and the receiving end establish a communication connection, and determine whether the receiving end has the function of restoring audio data, for example, obtain the function information (support or not support the function of restoring audio data) of the client. If the receiving end does not have the function of restoring audio data, the compressed audio data is transmitted to the receiving end by compressing the to-be-processed audio data. Taking the uplink transmission as an example, a function request is sent to the receiving end (for example, an edge server), the receiving end feeds back first function information when it has the function of restoring audio data, and the receiving end feeds back second function information when it does not have the function of restoring audio data, and the sending end calls the conditional network model to execute the audio processing method in the embodiment based on the first function information. Taking the downlink transmission as an example, the sending end receives the audio data request of the requesting end (for example, a client), and the audio data request can include a function field, for example, the function field is 1, indicating that the receiving end has the function of restoring audio data, and the function field is 0, indicating that the receiving end does not have the function of restoring audio data. The sending end determines whether to convert the to-be-processed audio data into condition data by analyzing the function field.

[0070] On the basis of the above embodiment, the conditional network model can be a neural network model. Optionally, the conditional network model can be a diffusion model.

[0071] In some embodiments, the conditional network model includes a first encoding module, a first connection module and a first decoding module connected in sequence, wherein the first decoding module includes a plurality of convolution blocks, and at least some of the convolution blocks in the first decoding module output local condition data respectively, and each of the convolution blocks outputs local condition data to splice the condition data.

[0072] The to-be-processed audio data is input into the conditional network model, and the first encoding module, the first connection module, and the first decoding module are sequentially processed with the input to-be-processed audio data to obtain local conditional data output by at least a local convolution block in the first decoding module. The local conditional data output by at least a local convolution block is spliced to obtain the conditional data corresponding to the to-be-processed audio data. It can be understood that the to-be-processed audio data is a frequency domain signal, and a modified discrete cosine transform (MDCT) is performed on the to-be-processed audio data. The data obtained by the transformation is input into the conditional network model. The first encoding data includes a plurality of convolution blocks, which are used for encoding processing of the input to-be-processed audio data. The first connection module can include a plurality of convolution blocks, which are used for recursive processing of the encoded data output by the first encoding module, and input the output data into the first decoding module for decoding processing. The number of convolution blocks included in the first encoding module, the first connection module, and the first decoding module is not limited.

[0073] Taking the first decoding module including N convolution blocks as an example, the first N-1 convolution blocks output local conditional data, and the local conditional data output by the first N-1 convolution blocks is spliced in order to obtain the conditional data.

[0074] Referring to FIG. 2, FIG. 2 is a structural schematic diagram of a conditional network model provided by an embodiment of the present disclosure. In FIG. 2, the first encoding module includes five convolution blocks connected in sequence, the first connection module includes two convolution blocks connected in sequence, and the first decoding module includes five convolution blocks connected in sequence. The first four convolution blocks in the first decoding module sequentially output local conditional data C0, C1, C2, and C3. The local conditional data C0, C1, C2, and C3 are spliced in the output order to obtain the conditional data C corresponding to the to-be-processed audio data. The fifth convolution block in the first decoding module outputs processed audio data, which can be ignored in the inference process.

[0075] In the first encoding module, a normalization layer GDN and a convolution layer are further arranged between adjacent convolution blocks. The normalization layer GDN replaces the PReLU layer to avoid the problem of a large amount of data being clipped. Each convolution layer in the first encoding module passes through a coordinator network layer to make the length of the encoded data the same, and finally accumulates to generate a hidden space representation, which is input into the first connection module.

[0076] In the above embodiment, the convolution block can include multiple convolution layers and multiple normalization layers. The structure of the convolution block is not limited here and can be built according to the construction requirements of the model. Optionally, the convolution block can include a down-sampling layer, a first convolution layer, a feature linear modulation layer, a second convolution layer, a third convolution layer, and an up-sampling layer connected in sequence, and each convolution layer is followed by a normalization layer. The output end of the down-sampling layer is jump-connected to the input end of the up-sampling layer. The first convolution layer, the second convolution layer, and the third convolution layer include one convolution layer with a kernel width of 5 and two convolution layers with a kernel width of 3. For example, the kernel width of the first convolution layer is 5, and the kernel widths of the second convolution layer and the third convolution layer are 3.

[0077] In the above embodiment, the conditional network model can be independently trained. For example, a conditional network model to be trained is obtained, sample audio data is obtained, the sample audio data is input into the conditional network model to be trained, conditional data and processed audio data are obtained, a loss function 1 is generated based on the sample audio data and the processed audio data, and the model parameters of the conditional network model to be trained are adjusted based on the loss function 1. It can be understood that the conditional data output by the conditional network model to be trained affects the quality of the target audio data recovered by the receiving end. The target audio data or quality data of the target audio data of the receiving end is obtained, a loss function 2 is generated based on the target audio data and the sample audio data, or a loss function 3 is generated based on the quality data of the target audio data, and the model parameters of the conditional network model to be trained are adjusted based on one or more of the loss function 1, the loss function 2, and the loss function 3. The above training process is iteratively performed, and the conditional network model trained iteratively is obtained when a training end condition is met. The training end condition can be a preset training times, a training accuracy, and a loss function reaching a convergence state.

[0078] The technical solution provided by the embodiments of the present disclosure converts the audio data to be processed into conditional data through a conditional network model, and represents the audio data to be processed through the conditional data. The data amount of the conditional data is small, and the transmission speed is fast. When the conditional data is transmitted to the receiving end, the receiving end can perform audio data recovery processing through the conditional data to obtain high-fidelity audio data, and the reduction of the transmission amount and the high-quality audio data are realized.

[0079] On the basis of the above-mentioned embodiments, in order to further reduce the amount of data to be transmitted, the conditional data corresponding to the audio data to be processed is subjected to first quantization and encoding processing to obtain first conditional encoding data; the first conditional encoding data is transmitted to the receiving end, so that the receiving end recovers the audio data based on the first conditional encoding data. The amount of data of the first conditional encoding data is less than the amount of data of the conditional data. By using the first conditional encoding data to represent the audio data to be processed, the amount of data is further reduced compared with the conditional data. The first conditional encoding data is transmitted to the receiving end, and the amount of data transmitted in the transmission process is further reduced, thereby improving the transmission efficiency. It can be understood that the receiving end has the function of recovering the audio data based on the first conditional encoding data.

[0080] Optionally, the first quantization and encoding processing is vector quantization and encoding processing, for example, it can be RVQ (Residual vector Quantization) encoding processing. The first quantization and encoding processing is performed based on a pre-trained first encoding model, and accordingly, the first encoding model can be an RVQ encoding model. Optionally, the first encoding model is a three-layer model, and the codebook size of the model is 1024.

[0081] On the basis of the above-mentioned embodiments, in the process of subjecting the conditional data to the first quantization and encoding processing, there is data damage, and accordingly, the data quality of the target audio data achieved by recovering the audio data based on the first conditional encoding data is lower than the data quality of the target audio data achieved by recovering the audio data based on the conditional data.

[0082] In view of the above problems, the damage data caused by the first quantization and encoding processing is determined, and the first conditional encoding data is compensated by the damage data to reduce the influence of the first quantization and encoding processing on the audio data recovery.

[0083] Optionally, residual conditional data is determined based on the conditional data and the first conditional encoding data; the residual conditional data is subjected to second quantization and encoding processing to obtain second conditional encoding data; and the first conditional encoding data and the second conditional encoding data are transmitted to the receiving end, so that the receiving end recovers the audio data based on the first conditional encoding data and the second conditional encoding data.

[0084] The residual condition data can be understood as damage data caused by the first quantization encoding processing. The second condition encoding data obtained by performing the second quantization encoding processing on the residual condition data has a smaller data quantity than the residual condition data. The second condition encoding data is used as compensation data of the first condition encoding data, and the second condition encoding data and the first condition encoding data form target transmission data. The second condition encoding data and the first condition encoding data are transmitted to the receiving end, so that the data quality of the target audio data obtained by the receiving end for audio data recovery can be improved.

[0085] Optionally, the second quantization encoding processing is scalar quantization encoding processing. The second quantization encoding processing is performed based on a pre-trained second encoding model. For example, the second quantization encoding processing can be entropy encoding. Through entropy encoding, discrete numerical values are compressed to a bit rate close to entropy representation. Through the context information of the encoding symbol, a suitable probability can be selected to reduce the redundancy between symbols and improve the encoding efficiency. Optionally, the second encoding model can be an EC (Entropy Coding) model.

[0086] In the above embodiment, the condition data is generated by a condition network model, and the condition data has certain feature redundancy. In this embodiment, an extra random variable (i.e., encoding auxiliary information) is introduced by introducing a hyper-prior model to capture the redundancy in the condition data, so as to assist the second quantization encoding processing and improve the accuracy of the second condition encoding data output by the second quantization encoding processing, and further improve the data quality of the target audio data recovered by the receiving end.

[0087] Optionally, the residual condition data is input into a hyper-prior module to obtain encoding auxiliary information. The residual condition data is subjected to the second quantization encoding processing based on the encoding auxiliary information to obtain second condition encoding data. Specifically, the output end of the hyper-prior module is connected to the input end of the second encoding model. The encoding auxiliary information is input into the second encoding model as prior information. The second encoding model performs the second quantization encoding processing on the residual condition data based on the encoding auxiliary information to obtain the second condition encoding data.

[0088] In the embodiment of the present disclosure, the residual condition data of the condition data and the first condition encoding data is subjected to the second quantization encoding processing to obtain the second condition encoding data. The second condition encoding data is used as compensation data of the first condition encoding data. The first condition encoding data and the second condition encoding data form target transmission data. The target transmission data more accurately represents the audio data to be processed, so that the receiving end can recover high-fidelity target audio data based on the first condition encoding data and the second condition encoding data, improve the data quality of the target audio data, and achieve the effect of small transmission quantity and high audio quality.

[0089] FIG. 3 is a flow diagram of an audio processing method provided by an embodiment of the present disclosure, which is applicable to a case where a receiving end recovers audio data based on condition information transmitted by a sending end. The method can be executed by an audio processing apparatus integrated in the receiving end, which can be implemented in the form of software and / or hardware, and can be implemented by an electronic device such as a mobile terminal, a PC terminal, or a server.

[0090] Referring to FIG. 3, the method specifically includes the following steps.

[0091] S210, receiving condition information transmitted by the sending end.

[0092] S220, obtaining noise data, and processing the condition information and the noise data based on a pre-trained generation network model to obtain target audio data.

[0093] The condition information can be condition data generated by a condition network model in the sending end, can also be first condition encoding data obtained by performing first quantization encoding processing on the condition data, and can further include the first condition encoding data and second condition encoding data. The specific content of the condition information is determined according to the data transmitted by the sending end.

[0094] The receiving end is configured with a pre-trained generation network model, and the condition information and the noise data are processed by the generation network model to recover the audio data, so as to obtain the target audio data. It can be understood that the electronic device can be used as the sending end and the receiving end in different scenarios, and the electronic device can be configured with the condition network model and the generation network model.

[0095] In some embodiments, the generation network model can be a neural network model, for example, the generation network model can be a generator in a generative adversarial network. The condition information is taken as the output of the generation network model, and the target audio data output by the generation network model is obtained, or the condition information and the noise data are taken as the output of the generation network model, and the target audio data output by the generation network model is obtained. The noise data can be Gaussian noise.

[0096] In some embodiments, the generation network model can be a diffusion model, for example, a forward diffusion process and a reverse diffusion process of the diffusion model obtain the generation network model. For example, the forward diffusion process of the generation network model can simulate a noise adding process of the audio data, and the reverse diffusion process of the generation network model can simulate a noise reduction process of the audio data. The low-noise high-fidelity target audio data is obtained by the reverse diffusion process of the generation network model.

[0097] Optionally, the pre-trained generative network model is used to process the condition information and the noise data to obtain target audio data, including: inputting the condition information and the noise data into the generative network model to obtain generated data output by the generative network model; and obtaining the target audio data based on the generated data and the noise data.

[0098] In some embodiments, the generated data output by the generative network model can be audio feature data, and the noise data is processed by the generated data to obtain the target audio data; in some embodiments, the generated data output by the generative network model can be score data, for example, data between 0 and 1, used to represent the processing effect of the generative network model on the condition information and the noise data. The noise data is processed by the score data to obtain the target audio data. Specifically, a denoising model can be pre-set, and the generated data and the noise data are input into the denoising model to obtain the target audio data, where the denoising model can be a mathematical model.

[0099] In some embodiments, the target audio data is obtained based on the generated data and the noise data by iteratively processing the noise data to obtain the target audio data.

[0100] Optionally, the target audio data is obtained based on the generated data and the noise data, including: obtaining intermediate audio data based on the generated data and the noise data; inputting the intermediate audio data and the condition information into the generative network model to obtain new generated data; obtaining new intermediate audio data based on the new generated data and the intermediate audio data; and determining the new intermediate audio data as the target audio data when the new generated data or the number of iterations meets a preset condition.

[0101] The process of generating intermediate audio data can be understood as a process of processing the noise data by the generated data once, and the intermediate audio data is noise-containing audio data obtained by processing the noise data. The intermediate audio data is used as audio data to be denoised in the next iteration process. The intermediate audio data and the condition information are input into the generative network model to obtain new generated data, which can be used to represent the processing effect of the generative network model on the condition information and the intermediate audio data, and can be used to process the intermediate audio data. The intermediate audio data is processed by the new generated data to obtain new intermediate audio data, where the data quality of the new intermediate audio data is higher than that of the intermediate audio data before denoising, that is, the noise content in the new intermediate audio data is reduced.

[0102] The new intermediate audio data is taken as audio data to be denoised in the next iteration process, and the next denoising processing is performed, and the iteration is executed. In each iteration process, the new generated data is judged. Taking the score data as an example, the score data is compared with the preset threshold. The score data is compared with the preset threshold, which indicates that the iteration end condition is reached, and the iteration process is ended. The new intermediate audio data generated in the current iteration is determined as the target audio data. It is known that in each iteration process, whether the current iteration number meets the preset iteration number is determined. If it is met, the iteration process is ended, and the new intermediate audio data generated in the current iteration is determined as the target audio data.

[0103] In each iteration process, the process of obtaining the intermediate audio data can be represented by the following formula:

[0104] Among them, Xt-1 is the audio data before denoising, Xt is the audio data after denoising, Yt is the generated data, σ is the standard deviation of the noise signal, C is the conditional information, and α and β are hyperparameters. The denoising process represented by the above formula is the reverse diffusion process of the diffusion model, so t n represents the previous iteration, t n-1 represents the next iteration. Exemplarily, t n =N represents the first iteration process, t n =0 represents the last iteration process. When t n =N, Xt-1 is the noise data before denoising, and correspondingly, Xt is the intermediate audio data. For example Xt is the t n th intermediate audio data, Xt+1 is the new intermediate audio data obtained by denoising the t n th intermediate audio data.

[0105] The reverse denoising process is iteratively executed until t n =0, the iteration process is ended, and X0 is determined as the target audio data.

[0106] In the above embodiment, the generation network model comprises a second encoding module, a second connection module and a second decoding module connected in sequence, the second encoding module and the second decoding module each comprise a preset number of convolutional blocks, and the convolutional blocks having a corresponding relationship in the second encoding module and the second decoding module are jump connected. The i th convolutional block in the second encoding module is jump connected with the N+1-i th convolutional block in the second decoding module, and N is a preset number.

[0107] The noise data is input to the second encoding module, the noise data is encoded based on the second encoding module, and each convolution block in the second encoding module outputs an encoded feature V, which is output to the next convolution block and the corresponding convolution block in the second decoding module. The noise data input to the second encoding module can be standard deviation data of the noise signal, which is input to the first convolution block of the second encoding module. The input information of each convolution block in the second encoding module further includes noise representation data extracted from the noise data based on the noise processing module. The second connection module can include a plurality of convolution blocks, and the input information of each convolution block in the second connection module further includes noise representation data extracted based on the noise data. The second connection module recursively processes the encoded features output by the second encoding module tree and inputs the output data to the second decoding module for decoding processing.

[0108] The condition information includes a plurality of local condition data spliced, the plurality of local condition data in the condition information is split to obtain a plurality of local condition data, and the plurality of local condition data is input to a plurality of convolution blocks of the second decoding module in turn. The input information of the plurality of convolution blocks of the second decoding module further includes noise representation data extracted based on the noise data.

[0109] For example, referring to FIG. 4, FIG. 4 is a structural schematic diagram of a network model generated according to an embodiment of the present disclosure. The second encoding module in FIG. 4 includes five convolution blocks connected in turn, the second connection module includes two convolution blocks connected in turn, and the second decoding module includes five convolution blocks connected in turn. The structure of each convolution block is not limited here.

[0110] The extraction process of the noise representation data can be converting the noise data to a latent space through random Fourier transform processing to obtain a latent data representation of the noise data, for example, [cos(2πf1logσ), …, cos(2πfNlogσ), sin(2πf1logσ), …, sin(2πfNlogσ)], where σ is the standard deviation data of the noise signal, and f is the frequency.

[0111] The noise processing module extracts features from the latent data representation to obtain noise representation data. For example, referring to FIG. 5, FIG. 5 is a structural schematic diagram of a noise processing module according to an embodiment of the present disclosure. The three repeated linear layers and the activation function layer in the noise processing module realize feature extraction of the latent data representation to obtain noise representation data g.

[0112] In the above embodiment, the training process of the generation network model can be to obtain the condition information of the sample audio data, which can be output by the sending end. It can be explained that, in the case that the conditional network model and the generation network model are configured in the same electronic device, the condition information of the sample audio data can be generated based on the conditional network model in the electronic device. The condition information of the sample audio data is input into the generation network model to be trained to obtain predicted video data, and a loss function A is generated based on the predicted video data and the sample audio data, or a loss function B is generated based on the data quality of the predicted video data, and the model parameters of the generation network model to be trained are conditioned based on the loss function A and / or the loss function B. In the above training process, when the end condition is met, the trained generation network model is obtained.

[0113] The technical scheme provided in the embodiment, when the condition information transmitted by the sending end is received at the receiving end, restores the audio data based on the generation network model to obtain the target audio data, and realizes the restoration of high-quality target audio data from low-transmission condition information, which takes into account the small amount of transmission data and the quality of the audio data.

[0114] FIG. 6 is a flowchart of a model training method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of jointly training a conditional network model and a generation network model. The method can be executed by a model training device, which can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server, etc. The trained conditional network model can be configured in the sending end, and the trained generation network model can be configured in the receiving end, or the trained conditional network model and the trained generation network model can be configured in the same electronic device to switch the sending end and the receiving end in different scenarios.

[0115] Referring to FIG. 6, the method includes the following steps:

[0116] S310, obtaining sample audio data, and obtaining a conditional network model to be trained and a generation network model to be trained.

[0117] S320, inputting the sample audio data into the conditional network model to be trained to obtain first audio data and condition data output by the conditional network model to be trained.

[0118] S330, performing noise adding processing on the sample audio data to obtain noise-added audio data, inputting the condition data and the noise-added audio data into the generation network model to be trained to obtain generation data output by the generation network model to be trained.

[0119] S340, generate a first loss function based on the first audio data and the sample audio data, and generate a second loss function based on the generated data.

[0120] S350, perform parameter adjustment on the conditional network model to be trained based on the first loss function and / or the second loss function, and perform parameter adjustment on the generation network model to be trained based on the second loss function, until a trained conditional network model and a trained generation network model are obtained.

[0121] In the embodiment, the conditional network model and the generation network model are jointly trained, instead of being trained independently, so that the conditional network model and the generation network model can be trained synchronously, and the training efficiency of the conditional network model and the generation network model is improved. Meanwhile, the joint training of the conditional network model and the generation network model does not require determining the label data of the sample audio data, and simplifies the processing of the sample audio data.

[0122] The sample audio data X is input into the conditional network model to be trained, to obtain first audio data X' and conditional data. The sample audio data X is subjected to noise addition processing, for example, Gaussian noise data is added to the sample audio data X, to obtain noise-added audio data. The noise-added audio data and the conditional data are input into the generation network model to be trained, to obtain generated data.

[0123] The first loss function is formed based on the sample audio data X and the first audio data X', and the first loss function can be a cross-entropy function, for example. The function type of the first loss function is not limited herein. The generated data can be score data, and the second loss function is generated based on the generated data. The second loss function can be negatively correlated with the generated data, that is, the greater the generated data, the smaller the loss value of the second loss function.

[0124] The conditional network model to be trained is adjusted based on the first loss function and / or the second loss function, for example, the first loss function and the second loss function can be weighted to obtain a first target loss function, and the conditional network model to be trained is adjusted based on the first target loss function. The generation network model to be trained is adjusted based on the second loss function. The above training process is iteratively performed, to obtain a trained conditional network model and a trained generation network model.

[0125] In the above embodiment, the encoding model can also be included in the sending end, to encode the conditional data and reduce the data amount of the conditional data. Correspondingly, the training of the encoding model can also be included in the above training process, to obtain a trained encoding model.

[0126] Optionally, the training process of the encoding model comprises: inputting the conditional data into the encoding model to be trained to obtain encoded data; inputting the encoded data and the noisy audio data into the generation network model to be trained to obtain generated data output by the generation network model to be trained; generating a third loss function based on the encoded data and the conditional data; and adjusting parameters of the encoding model to be trained based on the third loss function and / or the second loss function.

[0127] The third loss function is generated based on input information (i.e., the conditional data) and output information (i.e., the encoded data) of the encoding model. The third loss function can be a cross-entropy function, and the type of the third loss function is not limited herein. The encoded data is an influencing factor of the generated data, and the second loss function generated based on the generated data can be used to adjust parameters of the encoding model.

[0128] The third loss function and / or the second loss function can be weighted to obtain a second target loss function, and the encoding model to be trained is adjusted based on the second target loss function. The above adjustment process is iteratively performed, and the trained encoding model is obtained when a training end condition is met.

[0129] In the above embodiment, the encoding model comprises a first quantization encoding module and a second quantization encoding model. Correspondingly, the training of the encoding model comprises the training of the first quantization encoding module and the second quantization encoding model. Optionally, the inputting of the conditional data into the encoding model to be trained to obtain the encoded data comprises: inputting the conditional data into the first quantization encoding module to be trained to obtain first encoded data; and inputting the conditional data and the first encoded data into the second quantization encoding module to be trained to obtain second encoded data.

[0130] Correspondingly, the third loss function corresponding to the first quantization encoding module is generated based on the conditional data and the first encoded data. For example, the third loss function can be a cross-entropy function of the conditional data and the first encoded data. The third loss function corresponding to the second quantization encoding module is generated based on the residual conditional data and the second encoded data. For example, the third loss function can be a cross-entropy function of the residual conditional data and the second encoded data.

[0131] In the above embodiment, the conditional data output by the conditional network model is used as an influencing factor of the encoding model and the generation network model in the training process. In the case that the conditional data is inaccurate, the trained encoding model and the generation network model cannot be obtained. To solve the above problem, the joint training process of the conditional network model, the encoding model and the generation network model is divided into two stages, the conditional network model is trained preferentially, and the encoding model and the generation network model are trained based on the trained conditional network model, so as to avoid the influence of the conditional data output by the untrained conditional network model on the training process of the encoding model and the generation network model.

[0132] Optionally, in the first training stage, the conditional network model to be trained and the generation network model to be trained are trained to obtain a trained conditional network model and an intermediate generation network model; in the second training stage, the model parameters of the trained conditional network model are fixed, and the encoding model to be trained and the intermediate generation network model are trained to obtain a trained generation network model and a trained encoding model. For example, refer to FIG. 7, which is a flowchart of a model training method provided by an embodiment of the present disclosure.

[0133] In the first training stage, the sample audio data X is input into the conditional network model to be trained to obtain the first audio data X' and the conditional data output by the conditional network model to be trained. The sample audio data X is subjected to noise adding processing to obtain noise-added audio data, the conditional data and the noise-added audio data are input into the generation network model to be trained to obtain the generation data output by the generation network model to be trained. The first loss function is generated based on the first audio data X' and the sample audio data X, and the second loss function is generated based on the generation data. The parameter adjustment of the conditional network model to be trained is performed based on the first loss function and / or the second loss function, and the parameter adjustment of the generation network model to be trained is performed based on the second loss function. The above training process is repeated until the conditional network model meets the training end condition, the first training stage is ended, and the trained conditional network model and the intermediate generation network model are obtained.

[0134] In the second training stage, the sample audio data X is input into the trained conditional network model to obtain conditional data output by the conditional network model to be trained; the conditional data is input into the encoding model to be trained to obtain encoding data. The sample audio data X is subjected to noise adding processing to obtain noise-added audio data, and the encoding data and the noise-added audio data are input into the intermediate generation network model to obtain generated data output by the intermediate generation network model; a third loss function is generated based on the encoding data and the conditional data, and a second loss function is generated based on the generated data; the encoding model to be trained is subjected to parameter adjustment based on the third loss function and / or the second loss function, and the intermediate generation network model is subjected to parameter adjustment based on the second loss function. The above training process is cycled until a training condition is met, and the trained encoding model and the generation network model are obtained.

[0135] The technical scheme provided in the embodiment can jointly train the conditional network model, the encoding model and the generation network model, and can simultaneously train the conditional network model, the encoding model and the generation network model, thereby accelerating the training efficiency of the conditional network model, the encoding model and the generation network model.

[0136] FIG. 8 is a structural schematic diagram of an audio processing device provided in an embodiment of the present disclosure, as shown in FIG. 8, the device comprises an audio data acquisition module 410, a conditional data generation module 420 and a conditional data transmission module 430.

[0137] The audio data acquisition module 410 is configured to acquire audio data to be processed.

[0138] The conditional data generation module 420 is configured to process the audio data to be processed based on a pre-set conditional network model to obtain conditional data corresponding to the audio data to be processed.

[0139] The conditional data transmission module 430 is configured to transmit the conditional data corresponding to the audio data to be processed to a receiving end, so that the receiving end performs audio data recovery based on the conditional data corresponding to the audio data to be processed.

[0140] The technical scheme provided in the embodiment of the present disclosure converts the audio data to be processed into conditional data by using the conditional network model, and represents the audio data to be processed by using the conditional data. The data amount of the conditional data is small, and the transmission speed is fast. When the conditional data is transmitted to the receiving end, the receiving end can perform audio data recovery processing by using the conditional data to obtain high-fidelity audio data, thereby achieving the consideration of both reducing the transmission amount and high-quality audio data.

[0141] On the basis of the above-mentioned embodiment, the device can further comprise:

[0142] The first encoding module is configured to perform first quantization encoding processing on the condition data corresponding to the audio data to be processed to obtain first condition encoding data.

[0143] The condition data transmission module 430 is configured to transmit the first condition encoding data to the receiving end, so that the receiving end performs audio data recovery based on the first condition encoding data.

[0144] On the basis of the above-mentioned embodiments, optionally, the apparatus further comprises:

[0145] The second encoding module is configured to determine residual condition data based on the condition data and the first condition encoding data, perform second quantization encoding processing on the residual condition data to obtain second condition encoding data.

[0146] The condition data transmission module 430 is configured to transmit the first condition encoding data and the second condition encoding data to the receiving end, so that the receiving end performs audio data recovery based on the first condition encoding data and the second condition encoding data.

[0147] On the basis of the above-mentioned embodiments, optionally, the apparatus further comprises:

[0148] The hyper-prior processing module is configured to input the residual condition data into a hyper-prior module to obtain encoding auxiliary information.

[0149] The second encoding module is configured to perform second quantization encoding processing on the residual condition data based on the encoding auxiliary information to obtain second condition encoding data.

[0150] On the basis of the above-mentioned embodiments, optionally, the first quantization encoding processing is vector quantization encoding processing, and the second quantization encoding processing is scalar quantization encoding processing; the first quantization encoding processing is performed based on a pre-trained first encoding model, and the second quantization encoding processing is performed based on a pre-trained second encoding model.

[0151] On the basis of the above-mentioned embodiments, optionally, the condition network model comprises a first encoding module, a first connection module and a first decoding module connected in sequence, wherein the first decoding module comprises a plurality of convolution blocks, at least some of the convolution blocks in the first decoding module output local condition data respectively, and the convolution blocks output local condition data respectively to splice the condition data.

[0152] The audio processing apparatus provided in the embodiments of the present disclosure can perform the audio processing method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0153] It is worth noting that each unit and module included in the above apparatus is only divided according to functional logic, and is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual distinction, and does not serve to limit the protection scope of the embodiments of the present disclosure.

[0154] FIG. 9 is a structural schematic diagram of an audio processing apparatus provided by an embodiment of the present disclosure, as shown in FIG. 9, the apparatus includes a condition information receiving module 510 and a target audio data generating module 520.

[0155] The condition information receiving module 510 is configured to receive condition information transmitted by a sending end.

[0156] The target audio data generating module 520 is configured to obtain noise data, process the condition information and the noise data based on a pre-trained generation network model, and obtain target audio data.

[0157] The technical scheme provided by the embodiments of the present disclosure, when the condition information transmitted by the sending end is received at the receiving end, the condition information and the noise data are processed based on the generation network model to recover the audio data, and the target audio data is obtained, which realizes the recovery of high-quality target audio data from low-transmission condition information, and takes into account the small amount of transmission data and the quality of the audio data.

[0158] On the basis of the above-mentioned embodiments, optionally, the target audio data generating module 520 is configured to:

[0159] input the condition information and the noise data into the generation network model to obtain generated data output by the generation network model; and obtain the target audio data based on the generated data and the noise data.

[0160] Optionally, the target audio data generating module 520 is further configured to:

[0161] obtain intermediate audio data based on the generated data and the noise data; input the intermediate audio data and the condition information into the generation network model to obtain new generated data; obtain new intermediate audio data based on the new generated data and the intermediate audio data; and determine the new intermediate audio data as the target audio data when the new generated data or the number of iterations meets a preset condition.

[0162] On the basis of the above-mentioned embodiments, optionally, the generation network model includes a second encoding module, a second connection module and a second decoding module connected in sequence, and each of the second encoding module and the second decoding module includes a preset number of convolution blocks, and the convolution blocks having a corresponding relationship in the second encoding module and the second decoding module are connected in a skip connection manner.

[0163] The input information of the second encoding module includes the noise data; the conditional information includes a plurality of local conditional data corresponding to a plurality of convolution blocks input into the second decoding module; and the input information of each convolution block in the second encoding module and the second decoding module and the second connection module further includes noise representation data extracted based on the noise data.

[0164] The audio processing apparatus provided by the embodiments of the present disclosure can execute the audio processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0165] It is worth noting that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual distinction, and does not limit the protection scope of the embodiments of the present disclosure.

[0166] FIG. 10 is a structural schematic diagram of a model training apparatus provided by an embodiment of the present disclosure, as shown in FIG. 10, the apparatus includes a sample acquisition module 610, a processing module 620, a loss function generation module 630, and a model parameter adjustment module 640.

[0167] The sample acquisition module 610 is configured to acquire sample audio data, and acquire a conditional network model to be trained and a generative network model to be trained.

[0168] The processing module 620 is configured to input the sample audio data into the conditional network model to be trained, to obtain first audio data and conditional data output by the conditional network model to be trained; perform noise adding processing on the sample audio data to obtain noise-added audio data, input the conditional data and the noise-added audio data into the generative network model to be trained, and obtain generative data output by the generative network model to be trained.

[0169] The loss function generation module 630 is configured to generate a first loss function based on the first audio data and the sample audio data, and generate a second loss function based on the generative data.

[0170] The model parameter adjustment module 640 is configured to perform parameter adjustment on the conditional network model to be trained based on the first loss function and / or the second loss function, and perform parameter adjustment on the generative network model to be trained based on the second loss function.

[0171] The technical solution provided by the embodiments of the present disclosure trains the conditional network model and the generation network model jointly, replaces the independent training mode of the conditional network model and the generation network model, and synchronously trains the conditional network model and the generation network model to accelerate the training efficiency of the conditional network model and the generation network model.

[0172] On the basis of the above-mentioned embodiments, optionally, the processing module 620 is further configured to input the conditional data into the to-be-trained encoding model to obtain encoding data; input the encoding data and the noisy audio data into the to-be-trained generation network model to obtain generation data output by the to-be-trained generation network model.

[0173] The loss function generation module 630 is further configured to generate a third loss function based on the encoding data and the conditional data.

[0174] The model parameter adjustment module 640 is further configured to adjust the parameters of the to-be-trained encoding model based on the third loss function and / or the second loss function.

[0175] Optionally, the encoding model comprises a first quantization encoding module and a second quantization encoding model.

[0176] The processing module 620 is further configured to input the conditional data into the to-be-trained first quantization encoding module to obtain first encoding data; input the conditional data and the first encoding data into the to-be-trained second quantization encoding module as residual conditional data to obtain second encoding data.

[0177] Optionally, a third loss function corresponding to the first quantization encoding module is generated based on the conditional data and the first encoding data; and a third loss function corresponding to the second quantization encoding module is generated based on the residual conditional data and the second encoding data.

[0178] Optionally, the model parameter adjustment module 640 is further configured to, in a first training stage, train the to-be-trained conditional network model and the to-be-trained generation network model to obtain a trained conditional network model and an intermediate generation network model; and in a second training stage, fix the model parameters of the trained conditional network model, train the to-be-trained encoding model and the intermediate generation network model to obtain a trained generation network model and a trained encoding model.

[0179] The model training apparatus provided by the embodiments of the present disclosure can execute the model training method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0180] It is noted that the units and modules included in the above apparatus are only divided according to function logic, and are not limited to the above division, as long as the corresponding functions can be implemented; in addition, the specific names of each functional unit are only for easy mutual distinction, and do not serve to limit the protection scope of the embodiments of the present disclosure.

[0181] FIG. 11 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Referring to FIG. 11, a structural schematic diagram of an electronic device (e.g., a terminal device or a server in FIG. 11) 500 suitable for implementing an embodiment of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (e.g., a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 11 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0182] As shown in FIG. 11, the electronic device 500 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0183] Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 11 shows the electronic device 500 with various devices, it is understood that all the shown devices are not required to be implemented or provided. More or less devices can be alternatively implemented or provided.

[0184] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication apparatus 509, or installed from the storage apparatus 508, or installed from the ROM 502. When the computer program is executed by the processing apparatus 501, the above-mentioned functions defined in the methods of embodiments of the present disclosure are executed.

[0185] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0186] The electronic device provided by the embodiments of the present disclosure and the audio processing method or model training method provided by the above-mentioned embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiments, and the present embodiment has the same beneficial effects as the above-mentioned embodiments.

[0187] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the audio processing method or model training method provided by the above-mentioned embodiments.

[0188] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used or used in conjunction with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a RF (radio frequency) or the like, or any suitable combination of the foregoing.

[0189] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0190] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and not be assembled into the electronic device.

[0191] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device:

[0192] The computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: acquire to-be-processed audio data; process the to-be-processed audio data based on a pre-set conditional network model to obtain conditional data corresponding to the to-be-processed audio data; and transmit the conditional data corresponding to the to-be-processed audio data to a receiving end, so that the receiving end performs audio data recovery based on the conditional data corresponding to the to-be-processed audio data.

[0193] Alternatively, the computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: receive conditional information transmitted by a sending end; acquire noise data, and process the conditional information and the noise data based on a pre-trained generation network model to obtain target audio data.

[0194] Alternatively, the computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: acquire sample audio data, and acquire a to-be-trained conditional network model and a to-be-trained generation network model; input the sample audio data into the to-be-trained conditional network model to obtain first audio data and conditional data output by the to-be-trained conditional network model; perform noise adding processing on the sample audio data to obtain noise-added audio data, input the conditional data and the noise-added audio data into the to-be-trained generation network model to obtain generated data output by the to-be-trained generation network model; generate a first loss function based on the first audio data and the sample audio data, and generate a second loss function based on the generated data; adjust parameters of the to-be-trained conditional network model based on the first loss function and / or the second loss function, and adjust parameters of the to-be-trained generation network model based on the second loss function, until a trained conditional network model and a trained generation network model are obtained.

[0195] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0196] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0197] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination thereof. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first obtaining unit can also be described as "a unit for obtaining at least two Internet protocol addresses".

[0198] The functions described in this specification can be performed at least in part by one or more hardware logic components. For example, non-limiting examples of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), program-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0199] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0200] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0201] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0202] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. An audio processing method, comprising: Get the audio data to be processed; Processing the audio data to be processed based on a preset conditional network model to obtain conditional data corresponding to the audio data to be processed; The conditional data corresponding to the audio data to be processed is transmitted to the receiving end, so that the receiving end recovers the audio data based on the conditional data corresponding to the audio data to be processed.

2. The method according to claim 1, further comprising: performing a first quantization encoding process on the conditional data corresponding to the audio data to be processed to obtain first conditional encoded data; The first conditionally coded data is transmitted to the receiving end, so that the receiving end recovers the audio data based on the first conditionally coded data.

3. The method according to claim 2, further comprising: determining residual conditional data based on the conditional data and the first conditionally coded data; performing a second quantization coding process on the residual conditional data to obtain second conditional coded data; The first conditionally coded data and the second conditionally coded data are transmitted to the receiving end, so that the receiving end recovers the audio data based on the first conditionally coded data and the second conditionally coded data.

4. The method according to claim 3, further comprising: Inputting the residual conditional data into a super priori module to obtain coding auxiliary information; A second quantization encoding process is performed on the residual conditional data based on the encoding auxiliary information to obtain second conditional encoded data.

5. The method according to claim 3 or 4, wherein: The first quantization encoding process is a vector quantization encoding process, and the second quantization encoding process is a scalar quantization encoding process; The first quantization encoding process is performed based on a pre-trained first encoding model, and the second quantization encoding process is performed based on a pre-trained second encoding model.

6. The method according to any one of claims 1 to 5, wherein: The conditional network model includes a first encoding module, a first connection module and a first decoding module connected in sequence, wherein the first decoding module includes multiple convolution blocks, at least the local convolution blocks in the first decoding module respectively output local conditional data, and each of the convolution blocks respectively outputs local conditional data and splices them to obtain the conditional data.

7. An audio processing method, comprising: receiving condition information transmitted by the sending end; Noise data is acquired, and the condition information and the noise data are processed based on a pre-trained generative network model to obtain target audio data.

8. The method according to claim 7, wherein: The pre-trained generative network model is used to process the condition information and the noise data to obtain target audio data, including: Inputting the condition information and the noise data into the generative network model to obtain generated data output by the generative network model; The target audio data is obtained based on the generated data and the noise data.

9. The method according to claim 8, wherein The obtaining the target audio data based on the generated data and the noise data includes: obtaining intermediate audio data based on the generated data and the noise data; Inputting the intermediate audio data and the condition information into the generative network model to obtain new generated data; obtaining new intermediate audio data based on the newly generated data and the intermediate audio data; When the newly generated data or the number of iterations meets a preset condition, the new intermediate audio data is determined as the target audio data.

10. The method according to any one of claims 7 to 9, wherein: The generation network model includes a second encoding module, a second connection module, and a second decoding module connected in sequence, wherein the second encoding module and the second decoding module respectively include a preset number of convolution blocks, and the convolution blocks with corresponding relationships in the second encoding module and the second decoding module are skip-connected; The input information of the second encoding module includes the noise data; The condition information includes a plurality of local condition data, and the plurality of local condition data are correspondingly input into a plurality of convolution blocks of the second decoding module; Each convolution block in the second encoding module and the second decoding module and the input information of the second connection module also includes noise representation data extracted based on noise data.

11. A model training method comprising: Obtaining sample audio data, as well as a conditional network model to be trained and a generative network model to be trained; Inputting the sample audio data into the conditional network model to be trained to obtain first audio data and conditional data output by the conditional network model to be trained; Noising the sample audio data to obtain noisy audio data, inputting the conditional data and the noisy audio data into the generative network model to be trained to obtain generated data output by the generative network model to be trained; generating a first loss function based on the first audio data and the sample audio data, and generating a second loss function based on the generated data; The parameters of the conditional network model to be trained are adjusted based on the first loss function and / or the second loss function, and the parameters of the generative network model to be trained are adjusted based on the second loss function until a trained conditional network model and a trained generative network model are obtained.

12. The method according to claim 11, further comprising: Inputting the conditional data into the encoding model to be trained to obtain encoding data; Inputting the encoded data and the noisy frequency data into the generative network model to be trained to obtain generated data output by the generative network model to be trained; generating a third loss function based on the encoded data and the conditional data; Parameters of the encoding model to be trained are adjusted based on the third loss function and / or the second loss function.

13. The method according to claim 12, wherein: The coding model includes a first quantization coding module and a second quantization coding model; The step of inputting the conditional data into the coding model to be trained to obtain the coding data comprises: Inputting the conditional data into a first quantization encoding module to be trained to obtain first encoded data; Inputting residual conditional data of the conditional data and the first coded data into the second quantization coding module to be trained to obtain second coded data; A third loss function corresponding to the first quantization encoding module is generated based on the conditional data and the first encoded data; The third loss function corresponding to the second quantization encoding module is generated based on the residual condition data and the second encoded data.

14. The method according to claim 12 or 13, further comprising: In a first training phase, the conditional network model to be trained and the generative network model to be trained are trained to obtain a trained conditional network model and an intermediate generative network model; In the second training stage, the model parameters of the trained conditional network model are fixed, and the coding model to be trained and the intermediate generation network model are trained to obtain a trained generation network model and a trained coding model.

15. An audio processing device, comprising: An audio data acquisition module is configured to acquire audio data to be processed; a conditional data generating module, configured to process the audio data to be processed based on a preset conditional network model to obtain conditional data corresponding to the audio data to be processed; The conditional data transmission module is configured to transmit the conditional data corresponding to the audio data to be processed to a receiving end, so that the receiving end recovers the audio data based on the conditional data corresponding to the audio data to be processed.

16. An audio processing device, comprising: a condition information receiving module, configured to receive condition information transmitted by a sending end; The target audio data generation module is configured to obtain noise data, process the condition information and the noise data based on a pre-trained generation network model, and obtain target audio data.

17. A model training module, comprising: A sample acquisition module is configured to acquire sample audio data, and acquire a conditional network model to be trained and a generative network model to be trained; a processing module configured to input the sample audio data into the conditional network model to be trained, and obtain the first audio data and conditional data output by the conditional network model to be trained; Noising the sample audio data to obtain noisy audio data, inputting the conditional data and the noisy audio data into the generative network model to be trained to obtain generated data output by the generative network model to be trained; a loss function generating module, configured to generate a first loss function based on the first audio data and the sample audio data, and generate a second loss function based on the generated data; The model parameter adjustment module is configured to adjust the parameters of the conditional network model to be trained based on the first loss function and / or the second loss function, and to adjust the parameters of the generative network model to be trained based on the second loss function.

18. An electronic device comprising: one or more processors; a storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement one or more of the audio processing method described in any one of claims 1-6, the audio data method described in any one of claims 7-10, and the model training method described in any one of claims 11-14.

19. A storage medium containing computer-executable instructions, wherein: When executed by a computer processor, the computer executable instructions are used to execute one or more of the audio processing method described in any one of claims 1-6, the audio data method described in any one of claims 7-10, and the model training method described in any one of claims 11-14.

Citation Information

Patent Citations

  • Method and device for realizing vocoder based on variational auto-encoder

    CN111724809A

  • Audio encoding method and device and audio decoding method and device

    CN112259110A

  • Audio coding and decoding method and device, storage medium and computer equipment

    CN116504254A

  • Audio signal recovery method and device, electronic equipment and readable storage medium

    CN116705040A

  • Voice data processing method and device, electronic equipment and storage medium

    CN117095686A