Audio coding and decoding system training method and apparatus, audio coding method and apparatus, audio decoding method and apparatus, electronic device, computer-readable storage medium, and computer program product
By fixing the parameters of the first decoding network and only updating the second encoding network, the problems of long training cycles and high upgrade costs of audio encoding and decoding are solved, and efficient encoding and decoding are achieved to meet user needs and ensure forward compatibility.
Patent Information
- Application Number
- PCT/CN2024/124001
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-10-10
- Publication Date
- 2025-07-03
AI Technical Summary
The existing audio codec system has a long training cycle and high upgrade cost, which cannot meet user needs. The deep learning-based codec system has a high complexity, which limits its promotion in real-time audio and video applications.
By fixing the parameters of the first decoding network, only the second encoding network to be trained is updated, the training cycle is shortened and the upgrade cost is reduced. At the same time, neural network technology is used for encoding and decoding, reducing the complexity of the algorithm.
It shortens the training cycle of the audio codec system, reduces the upgrade cost, ensures forward compatibility, and improves coding efficiency and quality.
Smart Images

Figure CN2024124001_03072025_PF_FP_ABST
Abstract
Description
Audio encoding and decoding system training method, audio encoding method, audio decoding method, device, electronic device, computer-readable storage medium and computer program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] The embodiments of this application are based on the Chinese patent application with application number 202311832234.8 and application date December 26, 2023, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into the embodiments of this application as a reference. Technical Field
[0003] The present application relates to artificial intelligence technology, and in particular to a training method, an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for an audio codec system. Background Art
[0004] Audio coding and decoding technology is a key application in the field of artificial intelligence and a core technology for communication services, including remote audio and video calls. Simply put, speech coding technology aims to transmit as much voice information as possible using less network bandwidth. From the perspective of Shannon's information theory, speech coding is a form of source coding. The goal of source coding is to compress the data volume as much as possible at the encoding end, removing redundancy, while enabling lossless (or near-lossless) recovery at the decoding end.
[0005] In the related art, the training cycle of the encoding network and the decoding network in the audio codec system is very long, the upgrade cost is too high, and it cannot meet user needs.
[0006] Summary of the Invention
[0007] Embodiments of the present application provide a training method, an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for an audio codec system, which can shorten the training cycle of the audio codec system.
[0008] The technical solution of the embodiment of the present application is implemented as follows:
[0009] The present invention provides a method for training an audio codec system, which is applied to an electronic device and includes:
[0010] Acquire a first audio codec system, wherein the first audio codec system includes a first encoding network and a first decoding network;
[0011] In response to a configuration request for the first encoding network, generating a second encoding network to be trained corresponding to the first encoding network;
[0012] encoding the first audio sample based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample, and decoding the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample;
[0013] The parameters of the second encoding network to be trained are updated based on the reconstructed audio samples to obtain a trained second encoding network.
[0014] The present application provides an audio encoding method, which is applied to an electronic device and includes:
[0015] Get audio signal;
[0016] Invoking a trained second encoding network in an audio codec system to perform network coding processing on the audio signal to obtain a second encoding feature of the audio signal, wherein the audio codec system includes the trained second encoding network and the first decoding network;
[0017] performing signal encoding processing on the second encoding feature of the audio signal to obtain a second audio code stream of the audio signal;
[0018] In which, both the first audio code stream and the second audio code stream can be decoded by the first decoding network to obtain a reconstructed audio signal corresponding to the audio signal. The first audio code stream is an audio code stream obtained after the audio signal is processed by the first encoding network, and the trained second encoding network is obtained by training the first encoding network using the training method of the above-mentioned audio codec system.
[0019] The present invention provides an audio decoding method for an electronic device, including:
[0020] Get audio stream;
[0021] Performing signal decoding processing on the audio code stream to obtain a coding feature estimation value corresponding to the audio code stream;
[0022] Invoking a first decoding network in an audio codec system to perform network decoding processing on the coding feature estimation value to obtain a reconstructed audio signal corresponding to the audio code stream;
[0023] In which, the audio codec system includes a trained second encoding network and the first decoding network, the audio code stream is obtained after the audio signal is processed by the trained second encoding network or the first encoding network, and the trained second encoding network is obtained by training the first encoding network through the training method of the above-mentioned audio codec system.
[0024] The present invention provides a training device for an audio codec system, including:
[0025] A first acquisition module is configured to acquire a first audio codec system, wherein the first audio codec system includes a first encoding network and a first decoding network;
[0026] a determining module configured to generate, in response to a configuration request for the first encoding network, a second encoding network to be trained corresponding to the first encoding network;
[0027] a training module configured to encode the first audio sample based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample, and decode the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample;
[0028] The parameters of the second encoding network to be trained are updated based on the reconstructed audio samples to obtain a trained second encoding network.
[0029] The present invention provides an audio encoding device, including:
[0030] A second acquisition module is configured to acquire an audio signal;
[0031] an encoding module configured to call a trained second encoding network in an audio codec system to perform network coding on the audio signal to obtain a second encoding feature of the audio signal, wherein the audio codec system includes the trained second encoding network and the first decoding network;
[0032] a signal encoding module configured to perform signal encoding processing on the second encoding feature of the audio signal to obtain a second audio code stream of the audio signal;
[0033] In which, both the first audio code stream and the second audio code stream can be decoded by the first decoding network to obtain a reconstructed audio signal corresponding to the audio signal. The first audio code stream is an audio code stream obtained after the audio signal is processed by the first encoding network, and the trained second encoding network is obtained by training the first encoding network using the training method of the audio codec system.
[0034] The present invention provides an audio decoding device, including:
[0035] A third acquisition module is configured to acquire an audio stream;
[0036] a signal decoding module configured to perform signal decoding processing on the audio code stream to obtain an estimated coding feature value corresponding to the audio code stream;
[0037] a decoding module configured to call a first decoding network in the audio codec system to decode the coding feature estimate to obtain a reconstructed audio signal corresponding to the audio code stream;
[0038] In which, the audio codec system includes a trained second encoding network and the first decoding network, the audio code stream is obtained after the audio signal is processed by the trained second encoding network or the first encoding network, and the trained second encoding network is obtained by training the first encoding network using the training method of the audio codec system.
[0039] An embodiment of the present application provides an electronic device, comprising:
[0040] Memory for storing computer programs or computer-executable instructions;
[0041] The processor is configured to implement the audio codec system training method, audio encoding method, or audio decoding method provided in the embodiment of the present application when executing the computer program or computer executable instructions stored in the memory.
[0042] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implements the training method of the audio codec system, or the audio encoding method or audio decoding method provided in the embodiment of the present application.
[0043] An embodiment of the present application provides a computer program product, including computer-executable instructions, which, when executed by a processor, implement the training method of the audio codec system, or the audio encoding method or audio decoding method provided in the embodiment of the present application.
[0044] The embodiments of the present application have the following beneficial effects:
[0045] Since the parameters of the first decoding network remain unchanged when training the audio codec system, only the parameters of the second encoding network to be trained are updated, thereby saving the training process for the first decoding network. Compared with the related art of simultaneously training the first decoding network and the second encoding network to be trained, the training cycle of the audio codec system can be shortened, thereby reducing the upgrade cost of the encoding network in the audio codec system to meet the actual application needs of users; and, by fixing the parameters of the first decoding network unchanged and training the second encoding network, it can be ensured that the code stream processed by the trained second encoding network can be correctly decoded by the first decoding network. Since the code stream processed by the first encoding network can also be correctly decoded by the first decoding network, the forward compatibility of the audio codec system is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is a schematic diagram showing a comparison of spectrums at different bit rates provided in an embodiment of the present application;
[0047] FIG2A is a schematic diagram of a training platform of an audio codec system provided in an embodiment of the present application;
[0048] FIG2B is a schematic diagram of the architecture of an audio codec system provided in an embodiment of the present application;
[0049] FIG3A is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0050] FIG3B is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0051] FIG3C is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0052] 4A-4H are flowcharts of a training method for an audio codec system according to an embodiment of the present application;
[0053] FIG5 is a schematic diagram of a flow chart of an audio encoding method provided in an embodiment of the present application;
[0054] FIG6A is a schematic diagram of a flow chart of an audio decoding method provided in an embodiment of the present application;
[0055] FIG6B is a schematic diagram of a voice communication link provided in an embodiment of the present application;
[0056] FIG7 is a flow chart of a training method for an audio codec system according to an embodiment of the present application;
[0057] FIG8A is a schematic diagram of a common convolutional network provided in an embodiment of the present application;
[0058] FIG8B is a schematic diagram of a dilated convolutional network provided in an embodiment of the present application;
[0059] FIG9 is a schematic diagram of a first coding network provided in an embodiment of the present application;
[0060] FIG10A is a schematic diagram of a residual block structure used in a coding block provided in an embodiment of the present application;
[0061] FIG10B is a schematic diagram of the residual unit structure provided in an embodiment of the present application;
[0062] FIG11 is a schematic diagram of a first decoding network provided in an embodiment of the present application;
[0063] FIG12 is a schematic diagram of a second coding network provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0065] In the following description, the terms "first\second" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0066] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0067] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0069] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0070] 1) Neural Network (NN): A mathematical model that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing. This network relies on the complexity of the system to process information by adjusting the connections between its numerous nodes.
[0071] 2) Deep Learning (DL): A new research direction in the field of machine learning (ML), deep learning studies the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. Its ultimate goal is to enable machines to have human-like analytical learning capabilities and to recognize data such as text, images, and sounds.
[0072] 3) Quantization: This is the process of approximating a signal's continuous values (or a large number of discrete values) to a finite number (or a small number) of discrete values. Quantization includes vector quantization (VQ) and scalar quantization.
[0073] Vector quantization is an effective lossy compression technique, based on Shannon's rate-distortion theory. The basic principle of vector quantization is to replace the input vector with the index of the codeword in the code table that best matches it (also known as the quantization value) for transmission and storage, while decoding requires only a simple table lookup. For example, several scalar data items form a vector space, which is then divided into several small regions. During quantization, the vectors falling into the small regions are replaced with the corresponding indices of the input vectors.
[0074] Scalar quantization is the quantization of scalars, i.e. one-dimensional vector quantization, which divides the dynamic range into several small intervals, each of which has a representative value (i.e., index). When the input signal falls into a certain interval, the input signal is quantized to the representative value.
[0075] 4) Entropy Coding: This is a lossless coding method that uses the entropy principle to prevent any information loss during the encoding process. It is also a key module in lossy coding and is located at the end of the encoder. Entropy coding includes Shannon coding, Huffman coding, Exp-Golomb coding, and arithmetic coding.
[0076] 5) Quadrature Mirror Filters (QMF): This is a filter pair that combines analysis and synthesis. The QMF analysis filter decomposes the subband signal to reduce the signal bandwidth so that each subband signal can be processed smoothly through its own channel. The QMF synthesis filter synthesizes the subband signals recovered by the decoder, for example, through zero-value interpolation and bandpass filtering to reconstruct the original audio signal.
[0077] Speech coding technology aims to transmit as much voice information as possible while minimizing network bandwidth. Speech codecs can achieve compression ratios of over 10 times. This means that 10MB of speech data can be compressed by the codec to only 1MB, significantly reducing the bandwidth required to transmit the information. For example, for a wideband speech signal with a sampling rate of 16,000Hz, using a 16-bit sampling depth (the level of detail in recording speech intensity), the bit rate (the amount of data transmitted per unit time) of the uncompressed version is 256kbps. Using speech coding technology, even with lossy encoding, the quality of the reconstructed speech signal within a bit rate range of 10-20kbps can be close to that of the uncompressed version, and may even be perceived as indistinguishable. For services requiring even higher sampling rates, such as ultra-wideband speech at 32,000Hz, the bit rate must be at least 30kbps.
[0078] To ensure smooth communication within communication systems, the industry deploys standard voice codec protocols, such as those from international and domestic standards organizations like ITU-T, 3GPP, IETF, AVS, and CCSA, as well as standards like G.711, G.722, the AMR series, EVS, and OPUS. Figure 1 shows a schematic diagram comparing spectra at different bit rates, demonstrating the relationship between compression bit rate and quality. Curve 101 shows the spectrum of the original speech (i.e., the uncompressed signal); Curve 102 shows the spectrum of the OPUS encoder at a bit rate of 20 kbps; and Curve 103 shows the spectrum of the OPUS encoder at a bit rate of 6 kbps. As shown in Figure 1, as the encoding bit rate increases, the compressed signal becomes closer to the original signal.
[0079] The principles of speech coding are roughly as follows: speech coding can directly encode speech waveform samples sample by sample; or, based on the principle of human vocalization, relevant low-dimensional features are extracted, the encoding end encodes the features, and the decoding end reconstructs the speech signal based on these parameters.
[0080] The above coding principles are all derived from voice signal modeling, that is, compression methods based on signal processing, which cannot guarantee the encoding quality of audio. In response to this, the embodiments of the present application use technology based on deep learning to improve the encoding and decoding efficiency while ensuring voice quality. At present, technology based on deep learning can indeed bring low bit rate and high quality effects. However, this method also has the following two typical problems:
[0081] 1) This codec method is relatively complex. For some real-time audio and video applications, excessive complexity can limit the widespread adoption of audio applications. Furthermore, the training cycle for deep learning-based codec systems is long, ranging from days to weeks, resulting in high iteration costs.
[0082] 2) In order to ensure the forward compatibility principle in the communication system, the old version of the encoder and decoder can only be replaced first, and then the encoder and decoder can be updated, which makes the engineering transformation difficult.
[0083] To address the above issues, embodiments of the present application provide a training method, an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for an audio codec system. The following describes exemplary applications of the electronic device provided by embodiments of the present application. The electronic device provided by embodiments of the present application can be implemented as a terminal device, as a server, or collaboratively implemented by a terminal device and a server.
[0084] The following describes exemplary applications of electronic devices provided by embodiments of the present application. The electronic devices provided by embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices, and in-vehicle devices), smart phones, smart speakers, smart watches, smart TVs, and in-vehicle terminals. The electronic devices provided by embodiments of the present application can be implemented as independent physical servers, or as server clusters or distributed systems composed of multiple physical servers, or as cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms.
[0085] See Figure 2A, which is a schematic diagram of an application scenario of a training platform 10A for an audio codec system provided in an embodiment of the present application. The terminal device 100 is connected to the server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0086] The terminal device 100 (running a client, such as an audio client, a car client, etc.) can be used to obtain training requests for the audio codec system. For example, when the user loads the first audio codec system (including the first encoding network and the first decoding network) through the terminal 100, the terminal 100 automatically obtains the first audio codec system and automatically generates a training request for the audio codec system.
[0087] In some embodiments, a training plug-in for an audio codec system may be embedded in the client running in the terminal device 100 to implement a training method for the audio codec system locally on the client. For example, after the terminal device 100 obtains a training request for the audio codec system, it calls the training plug-in for the audio codec system to implement the training method for the audio codec system. The method includes first obtaining a first audio codec system, wherein the first audio codec system includes a first encoding network and a first decoding network; in response to a configuration request for the first encoding network, generating a second encoding network to be trained corresponding to the first encoding network; encoding the first audio sample based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample; and decoding the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample; and updating the parameters of the second encoding network to be trained based on the reconstructed audio sample to obtain a trained second encoding network. Since the parameters of the first decoding network remain unchanged when training the second encoding network, the training process for the first decoding network is saved. Compared with the related art of simultaneously training the first decoding network and the second encoding network to be trained, the training cycle of the audio codec system can be shortened, thereby reducing the upgrade cost of the encoding network in the audio codec system to meet the actual application needs of users.
[0088] In some embodiments, after the terminal device 100 obtains a training request for the audio codec system, it calls the training interface of the audio codec system of the server 200 (which can be provided in the form of a cloud service, i.e., the training service of the audio codec system), and the server 200 calls the training plug-in of the audio codec system to implement the training method of the audio codec system. First, a first audio codec system is obtained, wherein the first audio codec system includes a first encoding network and a first decoding network; in response to the configuration request for the first encoding network, a second encoding network to be trained corresponding to the first encoding network is generated; based on the second encoding network to be trained, the first audio sample is encoded to obtain an audio code stream sample of the first audio sample, and the audio code stream sample is decoded based on the first decoding network to obtain a reconstructed audio sample of the first audio sample; based on the reconstructed audio sample, the parameters of the second encoding network to be trained are updated to obtain a trained second encoding network. Since the parameters of the first decoding network remain unchanged when training the second encoding network, the training process for the first decoding network is saved. Compared with the related art of simultaneously training the first decoding network and the second encoding network to be trained, the training cycle of the audio codec system can be shortened, thereby reducing the upgrade cost of the encoding network in the audio codec system to meet the actual application needs of users.
[0089] In summary, after the second encoding network is trained in the above manner, the trained second encoding network and the first decoding network can be put into online application, that is, the trained second encoding network and the first decoding network are respectively integrated into a terminal device to implement the audio encoding method and the audio decoding method respectively.
[0090] For example, see Figure 2B, which is a schematic diagram of the architecture of the audio codec system 10B provided in an embodiment of the present application (i.e., an online audio codec system including a second encoding network and a first decoding network). The audio decoding system 10B includes: a server 200, a network 300, a terminal device 400 (i.e., an encoding end) and a terminal device 500 (i.e., a decoding end), wherein the network 300 can be a local area network, a wide area network, or a combination of the two.
[0091] In some embodiments, a client 410 runs on the terminal device 400. The client 410 can be various types of clients, such as an instant messaging client, a web conferencing client, a live broadcast client, a browser, etc. In response to an audio collection instruction triggered by a sender (such as the initiator of a web conferencing session, a host, or the initiator of a voice call), the client 410 calls the microphone of the terminal device 400 to collect audio signals, and performs audio encoding processing on the collected audio signals to obtain an audio stream.
[0092] For example, the client 410 calls the audio encoding method provided in the embodiment of the present application to encode the collected audio signal, that is, calls the trained second encoding network to encode the audio signal to obtain the second encoding feature of the audio signal; performs signal encoding processing on the second encoding feature of the audio signal to obtain the second audio code stream of the audio signal.
[0093] The client 410 can send the audio code stream package to the server 200 via the network 300, so that the server 200 sends the audio code stream package to the terminal device 500 associated with the recipient (such as a participant, audience, or recipient of a voice call in a web conference).
[0094] After receiving the audio code stream package sent by the server 200, the client 510 (such as an instant messaging client, a web conferencing client, a live broadcast client, a browser, etc.) running on the terminal device 500 can perform audio decoding processing on the audio code stream package to obtain a reconstructed audio signal, thereby realizing audio communication.
[0095] For example, the client 510 calls the audio decoding method provided in the embodiment of the present application to decode the received audio code stream encapsulation, that is, performs signal decoding processing on the audio code stream to obtain the coding feature estimation value corresponding to the audio code stream; calls the first decoding network to decode the coding feature estimation value to obtain the reconstructed audio signal corresponding to the audio code stream.
[0096] It should be noted that the server 200 shown in Figure 2A and the server 200 shown in Figure 2B can be the same server or different servers. The terminal device 100 shown in Figure 2A and the terminal device 400 or terminal device 500 shown in Figure 2B can be the same terminal device or different terminal devices.
[0097] In some embodiments, the terminal or server can implement the training method, audio encoding method or audio decoding method of the audio codec system provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a live broadcast application or an instant messaging application; it can also be a small program embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0098] 3A , which is a schematic diagram of the structure of an electronic device 500 provided in an embodiment of the present application. Taking the electronic device 500 as a server as an example, the electronic device 500 shown in FIG3A includes: at least one processor 520, a memory 550, at least one network interface 530, and a user interface 540. The various components in the electronic device 500 are coupled together via a bus system 550. It will be understood that the bus system 550 is used to implement connection and communication between these components. In addition to including a data bus, the bus system 550 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in FIG3A , various buses are labeled as the bus system 550.
[0099] The processor 520 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0100] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 520.
[0101] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0102] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0103] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0104] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 530. Exemplary network interfaces 530 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
[0105] In some embodiments, the training device of the audio codec system provided in the embodiments of the present application can be implemented in software. Figure 3A shows a training device 555 of the audio codec system stored in the memory 550, which can be software in the form of programs and plug-ins, including the following software modules: a first acquisition module 5551, a determination module 5552, and a training module 5553, wherein the first acquisition module 5551, the determination module 5552, and the training module 5553 are used to implement the training function of the audio codec system. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented.
[0106] Referring to FIG. 3B , FIG. 3B is a schematic diagram of the structure of an electronic device 600 provided in an embodiment of the present application. For example, the electronic device 600 is a terminal device. FIG. 3B shows the electronic device 600 including at least one processor 620, a memory 650, at least one network interface 630, and a user interface 640. The various components in the electronic device 600 are coupled together via a bus system 650. The memory 650 includes an operating system 651 and a network communication module 652. It should be noted that the functions of the structure in FIG. 3B are similar to those of the structure in FIG. 3A . The audio encoding device provided in an embodiment of the present application can be implemented in software. FIG. 3B shows an audio encoding device 655 stored in the memory 650. This device can be software in the form of a program or plug-in, and includes the following software modules: a second acquisition module 6551, an encoding module 6552, and a signal encoding module 6553. The second acquisition module 6551, the encoding module 6552, and the signal encoding module 6553 are used to implement audio encoding functions. These modules are logically connected and can be arbitrarily combined or further separated according to the functions implemented.
[0107] Referring to FIG. 3C , FIG. 3C is a schematic diagram of the structure of an electronic device 700 provided in an embodiment of the present application. For example, the electronic device 700 is a terminal device. The electronic device 700 shown in FIG. 3B includes at least one processor 720, a memory 750, at least one network interface 730, and a user interface 740. The various components in the electronic device 700 are coupled together via a bus system 750. The memory 750 includes an operating system 751 and a network communication module 752. It should be noted that the functions of the structure in FIG. 3C are similar to those of the structure in FIG. 3A . The audio decoding device provided in an embodiment of the present application can be implemented in software. FIG. 3B shows an audio decoding device 755 stored in the memory 750. This device can be software in the form of a program or plug-in, and includes the following software modules: a third acquisition module 7551, a signal decoding module 7552, and a decoding module 7553. The third acquisition module 7551, the signal decoding module 7552, and the decoding module 7553 are used to implement audio decoding functionality. These modules are logically connected and can be arbitrarily combined or further separated based on the functionality implemented.
[0108] As previously mentioned, the audio codec system training method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG4A , FIG4A is a flow chart of the audio codec system training method provided in the embodiments of the present application. The following description will be made in conjunction with steps 1 to 4 shown in FIG4A .
[0109] In step 1, a first audio codec system is obtained, wherein the first audio codec system includes a first encoding network and a first decoding network.
[0110] Among them, the first encoding network and the first decoding network are obtained by training the second audio sample. Here, the first encoding network and the first decoding network in the first audio codec system are matched, that is, the first decoding network can correctly decode the audio code stream encoded by the first encoding network. The first encoding network and the first decoding network are trained neural networks and can be applied to online services. For example, the first audio codec system is an audio codec system that has been applied online, that is, the first encoding network is an encoding network applied online, which can encode the received audio signal to obtain an audio code stream; the first decoding network is a decoding network applied online, which can decode the received audio code stream to obtain a reconstructed audio signal, thereby restoring the audio signal.
[0111] In step 2, in response to a configuration request for the first encoding network, a second encoding network to be trained corresponding to the first encoding network is generated.
[0112] The configuration request for the first coding network is used to instruct modification of the configuration data of the first coding network to generate a second coding network to be trained corresponding to the first coding network, that is, the second coding network to be trained is the coding network obtained after updating the first coding network.
[0113] Here, when the first encoding network in the first audio codec system needs to be updated, the configuration data of the first encoding network is modified, and a configuration request for the first encoding network is automatically generated. In response to the configuration request for the first encoding network, a second encoding network to be trained corresponding to the first encoding network is determined.
[0114] In some embodiments, the configuration request includes a homogeneous network configuration request for the first coding network, wherein the homogeneous network configuration request indicates that the network structure of the second coding network to be trained obtained by updating the first coding network is the same as the network structure of the first coding network, that is, the second coding network to be trained is an isomorphic network of the first coding network constructed based on the homogeneous network configuration request. The homogeneous network configuration request is used to indicate the modification of at least one of the following first configuration data: the second audio sample for training the first audio codec system, and the training strategy for training the first audio codec system. Among them, the training strategy is a strategy for training a neural network (such as the first coding network, the first decoding network), which cannot change the structure of the neural network, such as an unsupervised learning strategy, a supervised learning strategy, a training process strategy (such as the number of iterations, training batches, learning rate), etc. By reasonably selecting and combining these strategies, the training effect of the neural network can be effectively improved, so that it can show better performance and generalization ability in practical applications.
[0115] Supervised learning strategies use labeled data for training and adjust the neural network's parameters by comparing its output with the true labels. Unsupervised learning strategies do not use labeled data for training and adjust the neural network's parameters by discovering patterns and structures in the data. The number of iterations in the training process strategy represents the number of updates the neural network makes on the entire training dataset. Choosing an appropriate number of iterations can effectively improve the training performance of the neural network. Excessive iterations can lead to overfitting, while too few iterations can lead to underfitting. The batch size in the training process strategy represents the number of samples used in each iteration. Choosing an appropriate number of iterations can effectively improve the training performance of the neural network. Small batches can reduce memory requirements, improve computational efficiency, and help stabilize gradient updates, but can also lead to training instability. Large batches can slow down training. The learning rate in the training process strategy represents a hyperparameter that controls the step size for model parameter updates. Setting an appropriate learning rate ensures rapid convergence and gradually reducing the learning rate as training progresses improves training stability.
[0116] Here, the first encoding network and the first decoding network are trained through the second audio sample. When the isomorphic network configuration request is used to indicate the replacement of the second audio sample, the first audio sample is the second audio sample after replacement based on the isomorphic network configuration request, that is, the first audio sample is different from the second audio sample. For example, the first audio sample is a speech sample and the second audio sample is a song sample. The first encoding network trained through the speech sample is used to encode the speech signal and is suitable for speech scenarios; the second encoding network trained through the song sample is used to encode the song signal and is suitable for singing scenarios.
[0117] Here, when the homogeneous network configuration request is used to instruct modification of the training strategy for training the first audio codec system, the modified training strategy is used to train the second coding network to be trained in combination with the first audio sample to obtain the trained second coding network, that is, the second coding network is trained using the first audio sample and the modified training strategy to obtain the trained second coding network. The training strategy used to train the second coding network is a modified training strategy of the training strategy for training the first coding network, that is, the training strategy for training the second coding network is different from the training strategy for training the first coding network, for example, the training strategy for training the first coding network is an unsupervised learning strategy, and the training strategy for training the second coding network is a supervised learning strategy.
[0118] In some embodiments, the configuration request includes a heterogeneous network configuration request for the first coding network, the heterogeneous network configuration request being used to instruct modification of at least one of the following second configuration data: the network structure of the first coding network and the parameter values of the first coding network. Step 2 can be implemented by modifying the second configuration data of the first coding network based on the heterogeneous network configuration request to obtain a second coding network to be trained. The second coding network to be trained is a heterogeneous network of the first coding network.
[0119] For example, when the heterogeneous network configuration request is used to indicate the modification of the network structure of the first coding network, the network structure of the first coding network is modified based on the heterogeneous network configuration request to obtain the second coding network to be trained, wherein the network structure of the second coding network to be trained is different from the network structure of the first coding network. The network structure of the second coding network to be trained can be more complex than the network structure of the first coding network, and the network structure of the second coding network to be trained can also be simpler than the network structure of the first coding network.
[0120] For example, when the heterogeneous network configuration request is used to indicate the modification of the parameter amount of the first coding network, the parameter amount of the first coding network is modified based on the heterogeneous network configuration request to obtain the second coding network to be trained, wherein the parameter amount of the second coding network to be trained is different from the parameter amount of the first coding network. The parameter amount of the second coding network to be trained can be more than the parameter amount of the first coding network, for example, the variable dimension output by a certain convolutional layer in the second coding network to be trained is greater than the variable dimension output by the convolutional layer corresponding to the same layer in the second coding network; the parameter amount of the second coding network to be trained can be less than the parameter amount of the first coding network, for example, the variable dimension output by a certain convolutional layer in the second coding network to be trained is less than the variable dimension output by the convolutional layer corresponding to the same layer in the second coding network.
[0121] It should be noted that the configuration request in the embodiment of the present application is not limited to the above-mentioned heterogeneous network configuration request or the above-mentioned homogeneous network configuration request, and the configuration request may include the above-mentioned heterogeneous network configuration request and the above-mentioned homogeneous network configuration request.
[0122] In some embodiments, a second audio codec system, ie, a new audio codec system, is constructed based on the second encoding network to be trained and the first decoding network.
[0123] Here, since the first coding network needs to be updated, based on the configuration request for the first coding network, the second coding network to be trained corresponding to the first coding network is determined, and the first audio codec system is retained. The first audio codec system can still be put online for use. Based on the second coding network to be trained and the first decoding network, a new audio codec system, namely the second audio codec system, is constructed so that the new audio codec system can be put online for use later. It should be noted that after the second audio codec system is put online for use, the first audio codec system can be taken offline, that is, the latest audio codec system is adopted; after the second audio codec system is put online for use, the first audio codec system can also be retained, that is, both audio codec systems can be put online for use, and the two audio codec systems can be used in different scenarios. For example, if the first audio codec system online is used to transmit voice signals, then the first audio codec system is suitable for voice scenarios; if the second audio codec system online is used to transmit song signals, then the second audio codec system is suitable for singing scenarios.
[0124] In step 3, the first audio sample is encoded based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample, and the audio code stream sample is decoded based on the first decoding network to obtain a reconstructed audio sample of the first audio sample, and the parameters of the second encoding network to be trained are updated based on the reconstructed audio sample to obtain a trained second encoding network.
[0125] 4B , which is a flow chart of a training method for an audio codec system according to an embodiment of the present application, shows that step 3 can be implemented by following steps 41-45:
[0126] In step 41, a network coding process is performed on the first audio sample through a second coding network to be trained to obtain coding features of the first audio sample.
[0127] As shown in step 2 above, the second coding network to be trained can be a homogeneous network of the first coding network. The second coding network to be trained is trained using training samples or training strategies that are different from those of the first coding network, so that the trained second coding network and the first coding network are applicable to different scenarios. The second coding network to be trained can also be a heterogeneous network of the first coding network. The network structure of the second coding network to be trained can be more complex than that of the first coding network, that is, the complexity of the second coding network is increased relative to the first coding network, and the coding effect can be improved by increasing the computing power of the coding network. The network structure of the second coding network to be trained can also be simpler than that of the first coding network, that is, the structure of the first coding network is reduced to obtain a simpler second coding network, so as to reduce the complexity of the coding network and adapt to lightweight applications.
[0128] The following is an example of increasing the complexity of the second encoding network relative to the first encoding network (such as the second encoding network to be trained includes the network structure of the first encoding network and a one-dimensional convolutional layer).
[0129] 4C , which is a flow chart of a training method for an audio codec system according to an embodiment of the present application, shows that step 41 can be implemented by following the steps 411 - 412 :
[0130] In step 411, network coding is performed on the first audio sample using the network structure of the first coding network included in the second coding network to be trained to obtain initial coding features of the first audio sample.
[0131] 4D , which is a flowchart of a training method for an audio codec system according to an embodiment of the present application, shows that step 411 can be implemented by following the steps 4111-4112:
[0132] In step 4111, feature extraction processing is performed on the first audio sample using the network structure of the first encoding network included in the second encoding network to be trained to obtain audio features of the first audio sample.
[0133] In some embodiments, step 4111 can be implemented in the following manner: performing causal convolution processing on the first audio sample through the network structure of the first encoding network included in the second encoding network to be trained to obtain causal convolution features; and performing pooling processing on the causal convolution features to obtain audio features of the first audio sample.
[0134] In the field of audio codecs, neural network (NN) operations such as causal convolution and pooling play a crucial role in processing first audio samples (i.e., samples of an audio signal) and extracting features from the audio signal. In audio codecs, causal convolution can be used to extract local features from audio signals. By applying a convolution kernel (a learnable filter), convolution can be performed on the time dimension of the audio signal to capture patterns and resonances. Causal convolution can extract both time-domain and frequency-domain features from the audio signal, which can be used for tasks such as noise reduction, feature extraction, and signal separation. Pooling reduces the time dimension of the audio signal, thereby reducing data complexity and computational effort. Pooling samples a local region of the input signal and aggregates the information in that region, such as the maximum or average value, to generate a more compact feature representation. In audio signal processing, pooling can help improve the robustness and generalization of the network and reduce the risk of overfitting. In the field of audio codecs, operations such as convolution and pooling can be used to implement tasks such as feature extraction, encoding, and decoding of audio signals by constructing appropriate neural network structures. These operations help improve the efficiency and quality of audio signal processing and expand the application scope of audio codec technology in fields such as audio processing, speech recognition, and music generation.
[0135] In step 4112, residual processing is performed on the audio features of the first audio sample by using at least one residual unit in the first encoding network included in the second encoding network to be trained to obtain initial encoding features of the first audio sample.
[0136] In neural network models, residual units refer to a special structure used to build residual networks (Residual Networks, ResNets). Residual units are designed to address the problems of vanishing and exploding gradients during deep neural network training, and to help the network better learn features. Residual units introduce skip connections, which directly add the input to the output rather than simply passing it between layers. This skip connection allows the network to learn the residual function—the difference between the input and output—rather than directly learning the mapping relationship. This design makes the network easier to optimize and also helps alleviate the vanishing gradient problem.
[0137] Here, by performing residual processing on the audio features on the encoding side, based on the characteristics of residual processing, while ensuring comprehensive learning of the audio features, it is also possible to better utilize the shallow feature information of the audio features and avoid missing the shallow feature information of the audio features.
[0138] Based on the characteristics of the residual unit, the residual processing in step 4112 is used to calculate the residual of the audio feature on the encoding side, and determine the residual of the audio feature as the initial encoding feature for subsequent signal encoding. For example, the residual of the audio feature is obtained by adding the audio feature to the output of the residual unit, that is, the audio feature is used as the input of the residual unit, and after the audio feature is processed by the residual unit, the output of the residual unit is obtained, and through the jump connection characteristics of the residual unit, the input of the residual unit and the output of the residual unit are added to obtain the residual of the audio feature.
[0139] Referring to FIG. 4E , FIG. 4E is a flowchart of a training method for an audio codec system provided in an embodiment of the present application. FIG. 4E shows that step 4112 can be implemented through steps 41121 - 41122 .
[0140] In step 41121, feature residual processing is performed on the audio features of the first audio sample by using at least one residual unit in the first encoding network included in the second encoding network to be trained to obtain residual features of the first audio sample.
[0141] The feature residual processing in step 41121 is used to calculate the residual of the audio feature, and determine the residual of the audio feature as the residual feature of the first audio sample for subsequent feature encoding.
[0142] In some embodiments, when the at least one residual unit is a single residual unit, step 41121 may be implemented by performing a single residual processing on the audio features of the first audio sample using a single residual unit in a first encoding network included in the second encoding network to be trained, thereby obtaining a residual feature of the first audio sample. The single residual processing of the single residual unit is used to calculate a single residual corresponding to the first audio sample on the encoding side.
[0143] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 41121 can be implemented in the following manner: performing residual processing on the audio features of the first audio sample through the first residual unit of the plurality of cascaded residual units, wherein the one-time residual processing of the first residual unit is used to calculate the residual of the first audio sample once, and the residual of the first audio sample is determined as the residual result of the first residual unit; the residual result output by the first residual unit is output to the subsequent cascaded residual unit, and the residual processing and the output of the residual result are continued through the subsequent cascaded residual units, wherein the one-time residual processing of the subsequent cascaded residual unit is used to calculate the residual of the residual result input to the subsequent cascaded residual unit; and the residual result output by the last residual unit is used as the residual feature of the first audio sample.
[0144] In some embodiments, the processing process of the residual unit is as follows: the kth residual unit of multiple cascaded residual units performs the following processing: convolution processing is performed on the input of the kth residual unit to obtain the convolution result of the kth residual unit; the convolution result of the kth residual unit is added to the input of the kth residual unit to obtain the residual result output by the kth residual unit, wherein k is a positive integer that increases in sequence, 1≤k≤J, J is the number of residual units, when k is 1, the input of the kth residual unit is the audio feature of the first audio sample, when k is not 1, the input of the kth residual unit is the residual result output by the k-1th residual unit. That is, performing a residual processing on the audio features of the first audio sample through the first residual unit of multiple cascaded residual units can be achieved in the following way: performing the following processing through the first residual unit of multiple cascaded residual units: performing convolution processing on the audio features of the first audio sample to obtain the convolution result of the first residual unit; adding the convolution result of the first residual unit to the audio features of the first audio sample to obtain the residual result output by the first residual unit. Continuing to perform residual processing once through subsequent cascaded residual units and outputting the residual results can be achieved in the following manner: performing the following processing on the jth residual unit of multiple cascaded residual units: performing convolution processing on the residual result output by the j-1th residual unit to obtain the convolution result of the jth residual unit; adding the convolution results of the jth residual units and the residual result output by the j-1th residual unit to obtain the residual result output by the jth residual unit; outputting the residual result output by the jth residual unit to the j+1th residual unit; wherein j is a positive integer that increases successively, 1<j<J, and J is the number of residual units.
[0145] Continuing with the above embodiment, each residual unit includes a dilated convolution operator; performing the following processing on the kth residual unit of a plurality of cascaded residual units: performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit can be achieved in the following manner: performing the following processing on the kth residual unit of a plurality of cascaded residual units: performing dilated convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit. That is, the dilated convolution operator included in the first residual unit is used to perform dilated convolution processing on the audio features of the first audio sample to obtain the dilated convolution result of the first residual unit. The jth residual unit of multiple cascaded residual units performs the following processing: the residual result output by the j-1th residual unit is subjected to dilated convolution processing through the dilated convolution operator included in the jth residual unit to obtain the dilated convolution result of the jth residual unit, where j is a positive integer that increases in sequence, 1<j≤J, and J is the number of residual units. It should be noted that each residual unit contains a dilated convolution operator with a specified dilation rate. Using a dilated convolution operator with a progressive dilation rate is equivalent to using different receptive fields to extract input features at different resolutions, which can better perform a relatively comprehensive analysis of the data. After each residual unit is convolved with the dilated convolution operator with the dilation rate, it is added to the shallow features from the jump connection (i.e., the input of each residual unit), thereby directly utilizing the shallow feature information, so that the network can fully utilize the shallow feature information during the learning process.
[0146] Continuing with the above embodiment, each residual unit includes not only a dilated convolution operator but also at least one causal convolution operator. After performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit, causal convolution processing is performed on the obtained dilated convolution result using the at least one causal convolution operator included in the kth residual unit, and the obtained causal convolution result is used as the convolution result of the kth residual unit. That is, causal convolution processing is performed on the dilated convolution result of the first residual unit using the at least one causal convolution operator included in the first residual unit, and the obtained causal convolution result is used as the convolution result output by the first residual unit. The dilated convolution operator included in the j-th residual unit is used to perform dilated convolution processing on the residual result output by the j-1-th residual unit. After obtaining the dilated convolution result of the j-th residual unit, the dilated convolution result of the j-th residual unit is causally convolved by at least one causal convolution operator included in the j-th residual unit, and the causal convolution result of the j-th residual unit is used as the convolution result of the j-th residual unit. It should be noted that each residual unit also includes at least one causal convolution operator, which continues to extract local information of the features input to the causal convolution operator through the causal convolution operator.
[0147] In neural network models, causal convolution is a special type of convolution when processing time series data (audio signals are a type of time series data). It ensures that the output of the neural network depends only on the current and previous time steps, thereby maintaining temporal causality. In practical applications, causal convolution can adjust the size of the convolution kernel to ensure that the convolution kernel does not span the area before the current time step. This can effectively capture long-term dependencies in the time series while avoiding the problem of vanishing or exploding gradients caused by confusion caused by future information. Causal convolution is particularly important in fields such as natural language processing, speech recognition, and time series prediction because it follows the temporal order of the data, avoids confusion with past information, and can effectively process and predict long time series data. In tasks such as speech recognition and time series prediction, causal convolution has demonstrated superior performance due to its ability to maintain temporal order.
[0148] In some embodiments, when grouped convolution is applied to the dilated convolution operator included in the residual unit, dilated convolution processing is performed on the audio features, which can be achieved by: grouping the input channels of the audio features of the first audio sample to obtain multiple groups, wherein each group includes the first elements (i.e., first eigenvalues) corresponding to at least two channels in the audio features of the first audio sample; and performing dilated convolution processing on the first elements in each group. When grouped convolution is applied to the causal convolution operator included in the residual unit, causal convolution processing is performed on the obtained dilated convolution result, which can be achieved by: grouping the input channels of the dilated convolution result to obtain multiple groups, wherein each group includes the second elements (i.e., second eigenvalues) corresponding to at least two channels in the dilated convolution result; and performing causal convolution processing on the second elements in each group.
[0149] For example, grouped convolution can be applied to the convolution operator of the residual unit (including the dilated convolution operator and the causal convolution operator). Grouped convolution divides the input channels into multiple groups for convolution operations, and only the input channels and output channels within each group are associated. It should be noted that when the input channels are divided into multiple groups, the corresponding output channels are also divided into multiple groups. That is, the number of input channel groups is the same as the number of output channel groups. This ensures that after the convolution within a group, only the input channels and output channels within each group are associated. Here, assume that the input channel input to a feature of a convolution operator has 4 input channels and 4 output channels. If the number of groups is 1, each input channel is associated with 4 output channels. If the number of groups is 2, the 4 input channels are first divided into two groups, 0-1 and 2-3. Within each group, the input channels are associated with the output channels within the group. For example, input channels 0-1 in the first group are associated with output channels 0-1, and input channels 2-3 in the second group are associated with output channels 2-3. As shown in Figure 6A, when the grouped convolution scheme is not used, each input channel is associated with four output channels; as shown in Figure 6B, when the grouped convolution scheme is not used, the 0th output channel is only associated with the 0th-1st input channels, not with the 2nd-3rd input channels, and the 2nd output channel is only associated with the 2nd-3rd input channels, not with the 0th-1st input channels. This comparison shows that the introduction of grouped convolution can prevent any input channel from being associated with all output channels, reduce the number of connections, and reduce complexity.
[0150] Following step 41121, in step 41122, feature encoding is performed on the residual features to obtain initial encoding features of the first audio sample.
[0151] In some embodiments, step 41122 can be implemented by: performing convolution processing on the residual feature to obtain a convolution feature, wherein the number of channels of the convolution feature is greater than the number of channels of the residual feature; and performing pooling processing on the convolution feature to obtain an initial encoding feature of the first audio sample.
[0152] In some embodiments, the second coding network to be trained includes a first coding network including multiple cascaded coding blocks, each coding block including at least one residual unit and a feature coding block; step 4112 is implemented by multiple cascaded coding blocks, and step 4112 can be implemented in the following manner: performing residual processing on the audio features through at least one residual unit in the multiple cascaded coding blocks to obtain the residual features of the first audio sample; performing feature coding processing on the residual features through the feature coding blocks in the multiple cascaded coding blocks to obtain the initial coding features of the first audio sample.
[0153] In some embodiments, residual processing is performed on audio features by at least one residual unit in multiple cascaded coding blocks to obtain residual features of an audio signal. This can be achieved in the following manner: residual processing is performed on the audio features of a first audio sample by at least one residual unit in a first coding block of multiple cascaded coding blocks, and the residual result output by at least one residual unit in the first coding block is output to a feature coding block in the first coding block; residual processing is performed on the coding result output by a feature coding block in an i-1th coding block by at least one residual unit in an i-th coding block of multiple cascaded coding blocks, and the residual result output by at least one residual unit in the i-th coding block is output to a feature coding block in an i-th coding block; and the residual result output by at least one residual unit in the last coding block is used as the residual feature of the first audio sample; wherein i is a positive integer that increases successively, 1<i≤I, and I is the number of coding blocks. Obtaining the initial encoding features of the first audio sample by performing feature encoding processing on the residual features using a feature encoding block in a plurality of cascaded encoding blocks can be achieved in the following manner: performing feature encoding processing on the residual features of the first audio sample using a feature encoding block in a last encoding block in the plurality of cascaded encoding blocks to obtain the initial encoding features of the first audio sample. The encoding features are obtained by performing the following processing on the last encoding block in the plurality of cascaded encoding blocks: performing convolution processing on the residual features to obtain convolution features, wherein the number of channels of the convolution features is greater than the number of channels of the residual features; and performing pooling processing on the convolution features to obtain the initial encoding features of the first audio sample.
[0154] In step 412, the initial coding features are convolved using a one-dimensional convolutional layer included in the second coding network to be trained to obtain coding features of the first audio sample.
[0155] The main difference between the second encoding network to be trained and the first encoding network is that a one-dimensional convolutional layer is added to the end of the first encoding network. The input and output dimensions of this one-dimensional convolutional layer are the same. By increasing the number of layers in the encoding network, the feature extraction capability of the encoding network can be improved. Similarly, the number of layers in the encoding network can be further increased, including by adding one or more residual units to the one-dimensional convolutional layer.
[0156] In step 42, signal encoding processing is performed on the encoding feature of the first audio sample to obtain an audio code stream sample of the first audio sample.
[0157] In some embodiments, step 42 may be implemented by: performing quantization processing on the coding feature of the first audio sample to obtain an index value of the coding feature; and performing entropy coding processing on the index value of the coding feature to obtain an audio code stream sample of the first audio sample.
[0158] It should be noted that, in step 3, “encoding the first audio sample based on the second encoding network to be trained to obtain an encoding result of the first audio sample” is implemented through the above steps 41 and 42.
[0159] In step 43, signal decoding processing is performed on the audio code stream sample of the first audio sample to obtain a coding feature estimation value corresponding to the audio code stream sample.
[0160] In some embodiments, step 44 can be implemented by: performing entropy decoding on the audio code stream sample of the first audio sample to obtain an index value corresponding to the audio code stream sample; and performing inverse quantization on the index value corresponding to the audio code stream sample to obtain a coding feature estimation value corresponding to the audio code stream sample.
[0161] Inverse quantization is achieved by querying a quantization table, which is a mapping table generated by quantization during the encoding process. For example, for a received audio stream sample, entropy decoding is first performed, and then an estimate of the feature vector is obtained by querying a quantization table (i.e., inverse quantization, which is a mapping table generated by quantization during the encoding process). This is the estimated value of the coding feature corresponding to the audio stream sample.
[0162] In step 44, the coding feature estimation value is subjected to network decoding processing through the first decoding network to obtain a reconstructed audio sample corresponding to the audio code stream sample.
[0163] It should be noted that network decoding is the inverse of network encoding. Therefore, the values generated during network decoding are estimates relative to the values generated during network encoding. For example, the coding features generated during network decoding are estimates relative to the coding features generated during network encoding, i.e., estimated coding features.
[0164] Referring to FIG. 4F , FIG. 4F is a flowchart of a training method for an audio codec system provided in an embodiment of the present application. FIG. 4F shows that step 44 can be implemented through steps 441 - 442 .
[0165] In step 441, residual processing is performed on the coding feature estimation value corresponding to the audio stream sample by at least one residual unit in the first decoding network to obtain the audio feature estimation value corresponding to the audio stream sample.
[0166] Based on the characteristics of the residual unit, at least one residual unit is used to perform residual processing on the coding feature estimation value corresponding to the audio code stream sample for calculating the residual of the coding feature estimation value on the decoding side. For example, the residual of the coding feature estimation value is obtained by adding the coding feature estimation value to the output of the residual unit on the decoding side, that is, the coding feature estimation value is used as the input of the residual unit, and after the coding feature estimation value is processed by the residual unit on the decoding side, the output of the residual unit is obtained, and through the jump connection characteristics of the residual unit, the input of the residual unit on the decoding side and the output of the residual unit are added to obtain the residual of the coding feature estimation value.
[0167] 4G , which is a flowchart of a training method for an audio codec system provided by an embodiment of the present application. FIG4G shows that step 441 can be implemented through steps 4411-4412:
[0168] In step 4411, feature decoding processing is performed on the coding feature estimation value corresponding to the audio code stream sample through the first decoding network to obtain the residual feature estimation value corresponding to the audio code stream sample.
[0169] For example, feature decoding is the inverse process of feature encoding, and feature decoding processing is performed on the encoded feature estimate to obtain a residual feature estimate (an estimate) corresponding to the audio stream sample. In an embodiment of the present application, a first decoding network can be called to perform feature decoding processing on the encoded feature estimate corresponding to the audio stream sample through the first decoding network to obtain a residual feature estimate corresponding to the audio stream sample.
[0170] In some embodiments, step 4411 can be implemented by: performing convolution processing on the coding feature estimation value corresponding to the audio code stream sample to obtain a convolution feature, wherein the number of channels of the convolution feature is less than the number of channels of the coding feature estimation value corresponding to the audio code stream sample; and performing upsampling processing on the convolution feature to obtain a residual feature estimation value corresponding to the audio code stream sample.
[0171] In the field of audio codecs, upsampling is used to increase the resolution of feature maps (i.e., convolutional features) to more accurately reconstruct audio signals. Upsampling involves interpolation or other forms of upsampling techniques to generate higher-precision feature maps, which helps to better restore the original details and characteristics of the audio signal during the decoding process. By using neural network techniques such as convolution, pooling, and upsampling in audio decoding, useful features can be effectively extracted, computational complexity can be reduced, and the original content of the audio signal can be more accurately reconstructed. These technologies are of great significance for improving the performance and efficiency of audio decoding and help promote the development and application of audio codec technology.
[0172] Of course, before step 4411, causal convolution can also be performed on the coding feature estimation value corresponding to the audio code stream sample to obtain the coding feature estimation value after causal convolution, and step 4411 can be executed based on the coding feature estimation value after causal convolution, that is, feature decoding processing is performed on the coding feature estimation value after causal convolution to obtain the residual feature estimation value corresponding to the audio code stream sample.
[0173] In step 4412, feature residual processing is performed on the residual feature estimation value corresponding to the audio stream sample by at least one residual unit in the first decoding network to obtain the audio feature estimation value corresponding to the audio stream sample.
[0174] Here, residual processing is performed on the residual feature estimation values corresponding to the audio code stream samples to ensure that the residual feature estimation values are fully learned while better utilizing the shallow feature information of the residual feature estimation values to avoid missing the shallow feature information.
[0175] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 4412 may be implemented by performing a residual process on the residual feature estimation value through one residual unit to obtain an audio feature estimation value corresponding to the audio code stream sample.
[0176] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 4412 can be implemented in the following manner: performing a residual processing on the residual feature estimation value through the first residual unit of the plurality of cascaded residual units; outputting the residual result output by the first residual unit to the subsequent cascaded residual units, and continuing the residual processing and outputting the residual result through the subsequent cascaded residual units; and using the residual result output by the last residual unit as the audio feature estimation value corresponding to the audio code stream sample.
[0177] In some embodiments, the processing process of the residual unit is as follows: the kth residual unit of multiple cascaded residual units performs the following processing: convolution processing is performed on the input of the kth residual unit to obtain the convolution result of the kth residual unit; the convolution result of the kth residual unit is added to the input of the kth residual unit to obtain the residual result output by the kth residual unit, wherein k is a positive integer that increases in sequence, 1≤k≤J, J is the number of residual units, when k is 1, the input of the kth residual unit is the residual feature estimation value, when k is not 1, the input of the kth residual unit is the residual result output by the k-1th residual unit. That is, residual processing of the residual feature estimate can be achieved in the following manner: the first residual unit of multiple cascaded residual units performs the following processing: convolution processing is performed on the residual feature estimate to obtain the convolution result of the first residual unit; the convolution result of the first residual unit is added to the residual feature estimate to obtain the residual result output by the first residual unit. Continuing the residual processing and outputting the residual results through subsequent cascaded residual units can be achieved in the following manner: performing the following processing through the j-th residual unit of multiple cascaded residual units: performing the following processing through the j-th residual unit of multiple cascaded residual units: performing convolution processing on the residual result output by the j-1-th residual unit to obtain the convolution result of the j-th residual unit; adding the convolution results of the j-th residual units and the residual result output by the j-1-th residual unit to obtain the residual result output by the j-th residual unit; outputting the residual result output by the j-th residual unit to the j+1-th residual unit; wherein j is a positive integer that increases successively, 1<j<J, and J is the number of residual units.
[0178] Continuing with the above embodiment, each residual unit includes a dilated convolution operator; performing the following processing on the kth residual unit of multiple cascaded residual units: performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit can be achieved in the following manner: performing the following processing on the kth residual unit of multiple cascaded residual units: performing dilated convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit. That is, the dilated convolution operator included in the first residual unit is used to perform dilated convolution processing on the residual features to obtain the dilated convolution result of the first residual unit. The following processing is performed on the jth residual unit of multiple cascaded residual units: the residual result output by the j-1th residual unit is subjected to dilated convolution processing by the dilated convolution operator included in the jth residual unit to obtain the dilated convolution result of the jth residual unit, where j is a positive integer that increases in sequence, 1<j≤J, and J is the number of residual units.
[0179] Continuing with the above embodiment, each residual unit includes not only a dilated convolution operator but also at least one causal convolution operator. After performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit, causal convolution processing is performed on the obtained dilated convolution result using the at least one causal convolution operator included in the kth residual unit, and the obtained causal convolution result is used as the convolution result of the kth residual unit. That is, causal convolution processing is performed on the dilated convolution result of the first residual unit using the at least one causal convolution operator included in the first residual unit, and the obtained causal convolution result is used as the convolution result output by the first residual unit. The residual result output by the j-1th residual unit is subjected to dilated convolution processing through the dilated convolution operator included in the jth residual unit. After obtaining the dilated convolution result of the jth residual unit, the dilated convolution result of the jth residual unit is subjected to causal convolution processing through at least one causal convolution operator included in the jth residual unit, and the causal convolution result of the jth residual unit is used as the convolution result of the jth residual unit.
[0180] In some embodiments, when grouped convolution is applied to the dilated convolution operator included in the residual unit, dilated convolution processing is performed on the audio features, which can be achieved by: grouping the input channels of the residual feature estimate to obtain multiple groups, wherein each group includes the first elements (i.e., first eigenvalues) corresponding to at least two channels in the residual feature estimate; and performing dilated convolution processing on the first elements in each group. When grouped convolution is applied to the causal convolution operator included in the residual unit, causal convolution processing is performed on the obtained dilated convolution result, which can be achieved by: grouping the input channels of the dilated convolution result to obtain multiple groups, wherein each group includes the second elements (i.e., second eigenvalues) corresponding to at least two channels in the dilated convolution result; and performing causal convolution processing on the second elements in each group.
[0181] In some embodiments, the first decoding network includes multiple cascaded decoding blocks, each decoding block includes a feature decoding block and at least one residual unit; step 441 can be implemented in the following manner: through the feature decoding blocks in the multiple cascaded decoding blocks, feature decoding processing is performed on the coding feature estimation values corresponding to the audio code stream samples to obtain residual feature estimation values corresponding to the audio code stream samples; correspondingly, residual processing is performed on the residual feature estimation values corresponding to the audio code stream samples through at least one residual unit in the multiple cascaded decoding blocks to obtain audio feature estimation values corresponding to the audio code stream samples.
[0182] In some embodiments, feature decoding blocks in a plurality of cascaded decoding blocks perform feature decoding processing on coding feature estimates corresponding to audio stream samples to obtain residual feature estimates corresponding to the audio stream samples. This can be achieved in the following manner: feature decoding blocks in a first decoding block of the plurality of cascaded decoding blocks perform feature decoding processing on coding feature estimates corresponding to the audio stream samples, and output the decoding result output by the feature decoding block in the first decoding block to at least one residual unit in the first decoding block; feature decoding blocks in an i-th decoding block of the plurality of cascaded decoding blocks perform feature decoding processing on the residual result output by at least one residual unit in an i-1-th decoding block, and output the decoding result output by the feature decoding block in the i-th decoding block to at least one residual unit in the i-th decoding block; and use the decoding result output by the feature decoding block in the last decoding block as the residual feature estimate corresponding to the audio stream samples. Wherein, i is a successively increasing positive integer, 1<i≤I, and I is the number of decoding blocks. The decoding result output by the feature decoding block in the first decoding block is obtained by performing the following processing on the feature decoding block in the first decoding block of the multiple cascaded decoding blocks: performing convolution processing on the estimated encoding feature values corresponding to the audio code stream samples to obtain convolution features, wherein the number of channels of the convolution features is less than the number of channels of the encoding features; and performing upsampling processing on the convolution features to obtain the decoding result output by the feature decoding block in the first decoding block. The decoding result output by the feature decoding block in the i-th decoding block is obtained by performing the following processing on the feature decoding block in the i-th decoding block: performing convolution processing on the residual result output by at least one residual unit in the i-1-th decoding block to obtain convolution features, wherein the number of channels of the convolution features is less than the number of channels of the residual result output by the at least one residual unit; and performing upsampling processing on the convolution features to obtain the decoding result output by the feature decoding block in the i-th decoding block.
[0183] In some embodiments, residual processing is performed on the residual feature estimation value corresponding to the audio code stream sample through at least one residual unit in multiple cascaded decoding blocks, and the audio feature estimation value corresponding to the audio code stream sample is obtained. This can be achieved in the following manner: residual processing is performed on the residual feature estimation value corresponding to the audio code stream sample through at least one residual unit in the last decoding block of multiple cascaded decoding blocks, and the audio feature estimation value corresponding to the audio code stream sample is obtained.
[0184] Following the above step 441, in step 442, feature reconstruction processing is performed on the audio feature estimation value corresponding to the audio stream sample to obtain a reconstructed audio sample corresponding to the audio stream sample.
[0185] Here, feature reconstruction is the inverse process of feature extraction. The audio feature estimation value is upgraded through feature reconstruction to achieve the function of data decompression.
[0186] In some embodiments, step 442 may be implemented by upsampling the audio feature estimation values corresponding to the audio stream samples to obtain upsampled features; and performing causal convolution on the upsampled features to obtain reconstructed audio samples corresponding to the audio stream samples.
[0187] In step 45, the parameters of the second encoding network to be trained are updated based on the reconstructed audio samples to obtain a trained second encoding network.
[0188] It should be noted that before applying the second coding network, it is necessary to train the second coding network to be trained, and then put the trained second coding network into application. For example, updating the parameters of the second coding network to be trained based on the reconstructed audio sample can be achieved in the following way: after determining the value of the loss function of the second coding network based on the reconstructed audio sample and the first audio sample, it can be judged whether the value of the loss function exceeds the preset threshold. When the value of the loss function exceeds the preset threshold, the error signal of the second coding network is determined based on the loss function, the error information is back-propagated in the second coding network, and the model parameters of each layer are updated during the propagation process. Among them, the embodiments of the present application are not limited to the form of the loss function, for example, it can be a cross entropy loss function, an L2 loss function, etc.
[0189] Here, we explain backpropagation. Training sample data is input into the input layer of a neural network model, passing through the hidden layers, and finally reaching the output layer to output the result. This is the forward propagation process of the neural network model. Since there is an error between the output of the neural network model and the actual result, the error between the output and the actual value is calculated and propagated backward from the output layer to the hidden layers until it reaches the input layer. During the backpropagation process, the values of the model parameters are adjusted based on the error. Specifically, a loss function is constructed based on the error between the output and the actual value, and the partial derivatives of the loss function with respect to the model parameters are calculated layer by layer to generate the gradient of the loss function with respect to the model parameters of each layer. Since the direction of the gradient indicates the direction of error expansion, the gradient of the model parameters is negated and summed with the original parameters of each layer. The resulting sum is used as the updated model parameters of each layer, thereby reducing the error caused by the model parameters. This process is iterated continuously until convergence. The second encoding network is a neural network model.
[0190] It should be noted that, when the first audio sample involved in the above steps 41 - 45 is an audio signal sample, the audio code stream sample is a full-frequency code stream sample of the audio signal sample.
[0191] Referring to FIG. 4H , FIG. 4H is a flow chart of a training method for an audio codec system provided in an embodiment of the present application. When the first audio sample is a low-frequency sub-band signal obtained by sub-band decomposing an audio signal sample, and the audio bitstream sample is a low-frequency bitstream sample corresponding to the audio signal sample, FIG. 4H shows that step 3 can be implemented through steps 4-1 to 4-7:
[0192] In step 4-1, the audio signal samples are decomposed into sub-bands to obtain low-frequency sub-band signals of the audio signal samples.
[0193] Here, when the audio signal sample is an ultra-wideband signal, the audio signal sample is sub-band decomposed to obtain a low-frequency sub-band signal and a high-frequency sub-band signal of the audio signal sample. It should be noted that the embodiment of the present application does not limit the frequency bands of the low-frequency sub-band signal and the high-frequency sub-band signal, that is, the low-frequency sub-band signal and the high-frequency sub-band signal obtained by decomposition can be two sub-band signals obtained by evenly dividing the frequency band of the audio signal, or can also be two sub-band signals obtained by unevenly dividing the frequency band of the audio signal. For example, if the effective bandwidth of the audio signal sample x is 0-16kHz, then the low-frequency sub-band signal x LB and high frequency subband signal x HB The effective bandwidths are 0-8kHz and 8-16kHz respectively, and the low-frequency sub-band signal x LB and high frequency subband signal x HB The effective bandwidth can also be 0-6kHz and 6-16kHz, respectively. In addition, the embodiment of the present application does not limit the number of divided frequency bands, that is, the frequency band of the audio signal can be evenly / non-evenly divided to obtain two sub-band signals, or the frequency band of the audio signal can be evenly / non-evenly divided to obtain more than two sub-band signals, for example, 3, 4, or more sub-band signals.
[0194] In step 4-2, the low-frequency sub-band signal is subjected to network coding processing by the second coding network to be trained to obtain the low-frequency coding features of the low-frequency sub-band signal.
[0195] Here, the processing process of step 4-2 is similar to the processing process of step 41. The processing object of step 4-2 is the low-frequency sub-band signal, while the processing object of step 41 is the audio signal sample.
[0196] It should be noted that embodiments of the present application can perform high-frequency analysis processing on high-frequency subband signals to obtain high-frequency coding features of the high-frequency subband signals. Since low-frequency subband signals have a greater impact on audio coding than high-frequency subband signals, differentiated signal processing is performed on low-frequency subband signals and high-frequency subband signals, so that the feature dimension of the high-frequency coding features is lower than the feature dimension of the low-frequency coding features. The high-frequency analysis processing is used to reduce the dimensionality of the high-frequency subband signals to achieve data compression. The high-frequency coding features are features that characterize the high-frequency subband signals, and the feature dimension of the high-frequency coding features is smaller than the feature dimension of the high-frequency subband signals.
[0197] In some embodiments, high-frequency analysis and processing can be implemented in the following manner: calling a third coding network to encode the high-frequency sub-band signal to obtain the high-frequency coding features of the high-frequency sub-band signal, wherein the number of channels of the third coding network is less than the number of channels of the second coding network. The third coding network can be a trained coding network or a coding network to be trained. Here, another third coding network structure similar to the second coding network is introduced to generate a low-dimensional feature vector (i.e., the high-frequency coding features of the high-frequency sub-band signal). Compared with the low-frequency sub-band signal, the high-frequency sub-band signal is relatively less important to quality. Therefore, the third coding network structure for the high-frequency sub-band signal does not need to be as complex as the second coding network.
[0198] In some embodiments, since the high-frequency sub-band signal is relatively less important to quality than the low-frequency sub-band signal, the high-frequency sub-band signal can be compressed by another method, namely, band expansion (recovering a broadband speech signal from a band-limited narrow-band speech signal) to quickly compress the high-frequency sub-band signal and extract the high-frequency coding features of the high-frequency sub-band signal.
[0199] In some embodiments, high-frequency analysis processing can be achieved in the following manner: performing frequency domain transformation processing based on multiple sample points included in the high-frequency sub-band signal to obtain transformation coefficients corresponding to the multiple sample points; dividing the transformation coefficients corresponding to the multiple sample points into multiple sub-bands; averaging the transformation coefficients included in each sub-band to obtain the average energy corresponding to each sub-band, and using the average energy as the sub-band spectrum envelope corresponding to each sub-band; determining the sub-band spectrum envelopes corresponding to the multiple sub-bands as the high-frequency coding features of the high-frequency sub-band signal.
[0200] It should be noted that the frequency domain transformation methods in the embodiments of the present application include modified discrete cosine transform (MDCT), discrete cosine transform (DCT), fast Fourier transform (FFT), etc., and the embodiments of the present application are not limited to the frequency domain transformation method. The averaging processing in the embodiments of the present application includes arithmetic mean and geometric mean, and the embodiments of the present application are not limited to the averaging processing method.
[0201] In some embodiments, frequency domain transform processing is performed based on multiple sample points included in the high-frequency sub-band signal to obtain transform coefficients corresponding to the multiple sample points, including: obtaining a reference high-frequency sub-band signal of a reference audio signal, wherein the reference audio signal is an audio signal adjacent to the audio signal; based on the multiple sample points included in the reference high-frequency sub-band signal and the multiple sample points included in the high-frequency sub-band signal, discrete cosine transform processing is performed on the multiple sample points included in the high-frequency sub-band signal to obtain transform coefficients corresponding to the multiple sample points included in the high-frequency sub-band signal.
[0202] In some embodiments, the process of performing geometric averaging on the transform coefficients included in each sub-band is as follows: determining the sum of the squares of the transform coefficients corresponding to the sample points included in each sub-band; and determining the ratio of the sum of the squares to the number of sample points included in the sub-band to obtain the average energy corresponding to each sub-band.
[0203] As an example, for a high frequency subband signal x comprising 320 points HB (n), call the modified discrete cosine transform (MDCT) to generate 320-point MDCT coefficients (i.e., transform coefficients corresponding to multiple sample points included in the high-frequency subband signal). Specifically, if there is a 50% overlap, the high-frequency data of the n+1th frame (i.e., the reference audio signal) can be merged (concatenated) with the high-frequency data of the nth frame (i.e., the audio signal), and the 640-point MDCT can be calculated to obtain the MDCT coefficients of the first 320 points.
[0204] The 320-point MDCT coefficients are divided into N subbands (i.e., the transform coefficients corresponding to multiple sample points are divided into multiple subbands). A subband here is a group of multiple adjacent MDCT coefficients. The 320-point MDCT coefficients can be divided into 8 subbands. For example, the 320 points can be evenly distributed, that is, the number of points included in each subband is the same. Of course, the embodiment of the present application cannot perform non-uniform division of the 320 points, such as a subband with a lower frequency including fewer MDCT coefficients (higher frequency resolution) and a subband with a higher frequency including more MDCT coefficients (lower frequency resolution).
[0205] According to Nyquist's sampling theorem (to recover the original signal from a sampled signal without distortion, the sampling frequency must be greater than twice the original signal's highest frequency. If the sampling frequency is less than twice the highest frequency, the signal spectrum will exhibit aliasing, while if the sampling frequency is greater than twice the highest frequency, the signal spectrum will exhibit no aliasing), the 320-point MDCT coefficients above represent a spectrum between 8 and 16 kHz. However, ultra-wideband voice communication does not necessarily require a spectrum extending to 16 kHz. For example, if the spectrum is set to 14 kHz, only the MDCT coefficients of the first 240 points need to be considered, and the number of subbands can be controlled to 6.
[0206] For each sub-band, the average energy of all MDCT coefficients in the current sub-band is calculated (i.e., the transform coefficients included in each sub-band are averaged) as the sub-band spectrum envelope (the spectrum envelope is a smooth curve passing through the main peak points of the spectrum). For example, if the MDCT coefficients included in the current sub-band are x(n), n = 1, 2, ..., 40, then the average energy Y is calculated by the geometric mean = ((x(1) 2 +x(2) 2 +…+x(40) 2 ) / 40). When the 320-point MDCT coefficients are divided into 8 sub-bands, 8 sub-band spectrum envelopes can be obtained. These 8 sub-band spectrum envelopes are the eigenvectors F of the generated high-frequency sub-band signal. HB (n), i.e., high-frequency coding features.
[0207] In step 4-3, signal coding processing is performed on the low-frequency coding features of the low-frequency sub-band signal to obtain low-frequency code stream samples of the low-frequency sub-band signal.
[0208] Here, the processing process of step 4-3 is similar to the processing process of step 42. The processing object of step 4-3 is the coding feature of the low-frequency sub-band signal, while the processing object of step 42 is the coding feature of the audio signal sample.
[0209] In step 4-4, signal decoding processing is performed on the low-frequency code stream samples of the low-frequency sub-band signal to obtain low-frequency coding feature estimation values corresponding to the low-frequency code stream samples.
[0210] Here, the processing of step 4-4 is similar to the processing of step 43.
[0211] In some embodiments, the embodiments of the present application can also perform signal encoding processing on the high-frequency coding features of the high-frequency sub-band signal, and perform signal decoding processing on the high-frequency code stream samples of the obtained high-frequency sub-band signal to obtain the high-frequency coding feature estimation value corresponding to the high-frequency code stream sample.
[0212] In step 4-5, the low-frequency coding feature estimation value is decoded by the first decoding network to obtain the low-frequency sub-band signal estimation value corresponding to the low-frequency code stream sample.
[0213] Here, the processing process of step 4-5 is similar to the processing process of step 44. The processing object of step 4-5 is the low-frequency coding feature estimation value, while the processing object of step 44 is the coding feature estimation value.
[0214] In some embodiments, the embodiments of the present application may further perform high-frequency reconstruction processing on the high-frequency coding feature estimation values corresponding to the high-frequency code stream samples to obtain high-frequency sub-band signal estimation values corresponding to the high-frequency code stream samples.
[0215] High-frequency reconstruction and high-frequency analysis are inverse processes. For example, when the encoder uses the encoding network to encode the high-frequency subband signal to obtain the high-frequency coding features, the decoder uses the decoding network to decode the high-frequency coding feature estimates to obtain the corresponding high-frequency subband signal estimates.
[0216] In some embodiments, when the encoding end performs frequency band expansion processing on the high-frequency sub-band signal to obtain high-frequency coding features, the decoding end performs inverse frequency band expansion processing on the high-frequency coding feature estimation value to obtain the high-frequency sub-band signal estimation value corresponding to the high-frequency code stream sample.
[0217] In some embodiments, the high-frequency coding feature estimate is subjected to inverse processing of frequency band expansion to obtain a high-frequency sub-band signal estimate corresponding to the high-frequency code stream sample, including: performing frequency domain transform processing based on multiple sample points included in the low-frequency sub-band signal estimate to obtain transform coefficients corresponding to the multiple sample points; performing spectrum replication processing on the transform coefficients of the latter half of the transform coefficients corresponding to the multiple sample points to obtain reference transform coefficients of the reference high-frequency sub-band signal; performing gain processing on the reference transform coefficients of the reference high-frequency sub-band signal based on the sub-band spectrum envelope corresponding to the high-frequency coding feature estimate to obtain the gained reference transform coefficients; and performing inverse frequency domain transform processing on the gained reference transform coefficients to obtain the corresponding high-frequency sub-band signal estimate.
[0218] It should be noted that the frequency domain transformation methods of the embodiments of the present application include modified discrete cosine transform (MDCT), discrete cosine transform (DCT), fast Fourier transform (FFT), etc. The embodiments of the present application are not limited to the frequency domain transformation method.
[0219] In some embodiments, based on the subband spectral envelope corresponding to the high-frequency coding feature estimation value, the reference transform coefficient of the reference high-frequency subband signal is gain-processed to obtain the gained reference transform coefficient, including: based on the subband spectral envelope corresponding to the high-frequency coding feature estimation value, the reference transform coefficient of the reference high-frequency subband signal is divided into multiple subbands; for any subband among the multiple subbands, the following processing is performed: determining a first average energy corresponding to the subband in the subband spectral envelope, and determining a second average energy corresponding to the subband; determining a gain factor based on the ratio of the first average energy to the second average energy; multiplying the gain factor by each reference transform coefficient included in the subband to obtain the gained reference transform coefficient.
[0220] In step 4-6, sub-band synthesis processing is performed based on the low-frequency sub-band signal estimation value to obtain reconstructed audio samples.
[0221] For example, subband synthesis processing is the inverse process of subband decomposition processing. The decoding end performs subband synthesis processing on the low-frequency subband signal estimation value and the high-frequency subband signal estimation value to restore the audio signal, where the reconstructed audio sample is the restored reconstructed signal.
[0222] In some embodiments, subband synthesis processing is performed on the low-frequency subband signal estimation value and the high-frequency subband signal estimation value, including: upsampling the low-frequency subband signal estimation value to obtain a low-pass filtered signal; upsampling the high-frequency subband signal estimation value to obtain a high-frequency filtered signal; and filtering synthesis processing is performed on the low-pass filtered signal and the high-frequency filtered signal to obtain a reconstructed audio sample.
[0223] In step 4-7, the parameters of the second encoding network to be trained are updated based on the reconstructed audio samples to obtain the trained second encoding network.
[0224] Here, the processing of steps 4-7 is similar to the processing of step 45.
[0225] In some embodiments, the generative adversarial network includes a generative model and a discriminative model, the generative model is a second audio codec system, and the second audio codec system includes a second encoding network to be trained and a first decoding network; the second encoding network to be trained in step 3 is obtained by training in the following manner: based on the generative model and the discriminative model in the generative adversarial network, the following training tasks are performed alternately: based on the first audio sample, the generative model is trained so that the generative model generates a reconstructed audio sample based on the audio sample; based on the first audio sample and the reconstructed audio sample, the discriminative model is trained so that the discriminative model distinguishes between the audio sample and the reconstructed audio sample; wherein, when training the generative model, the parameters of the discriminative model are fixed unchanged; when training the discriminative model, the parameters of the generative model are fixed unchanged.
[0226] Here, the parameters of the fixed discriminant model remain unchanged, and the method for training the generative model is similar to the above step 3, that is, the generative model is regarded as a first audio codec system including a second encoding network to be trained.
[0227] As previously mentioned, the audio encoding method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG. 5 , FIG. 5 is a schematic flow chart of the audio encoding method provided in the embodiments of the present application. The audio encoding function is implemented by the audio encoding method, and the following description is provided in conjunction with steps 101 to 103 shown in FIG. 5 .
[0228] In step 101, an audio signal is acquired.
[0229] As an example of obtaining an audio signal, the encoding end responds to the audio collection instruction triggered by the sender (such as the initiator of the online conference, the host, the initiator of the voice call, etc.), and calls the microphone of the terminal device of the encoding end to collect the audio signal to obtain the audio signal (also called the input signal).
[0230] In step 102, a trained second encoding network in an audio codec system is called to encode an audio signal to obtain a second encoding feature of the audio signal, wherein the audio codec system includes a trained second encoding network and a first decoding network.
[0231] Here, the process of encoding the audio signal through the trained second encoding network in step 102 is similar to the above step 41, except that the processing object in step 102 is the audio signal, while the processing object in step 41 is the first audio sample.
[0232] It should be noted that when the audio signal is an ultra-wideband signal and needs to be decomposed into sub-bands, the audio signal can be decomposed into sub-bands first, and then the sub-band signal can be decomposed. The processing process is similar to the process shown in FIG. 4H above.
[0233] Following step 102 , in step 103 , signal encoding processing is performed on the second encoding feature of the audio signal to obtain a second audio code stream of the audio signal.
[0234] Among them, both the first audio code stream and the second audio code stream can be decoded by the first decoding network to obtain reconstructed audio signals corresponding to the audio signals. The first audio code stream is the audio code stream obtained after the audio signal is processed by the first encoding network, and the trained second encoding network is obtained by training the first encoding network through the training method of the above-mentioned audio codec system.
[0235] In some embodiments, step 103 can be implemented by: quantizing the second coding feature of the audio signal to obtain an index value of the second coding feature; and entropy coding the index value of the second coding feature to obtain a second audio code stream of the audio signal.
[0236] As previously mentioned, the audio decoding method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG6A , FIG6A is a schematic flow chart of the audio decoding method provided in the embodiments of the present application. The audio decoding function is implemented by the audio decoding method. The audio decoding method is the inverse process of the audio encoding method described above, and will be described below in conjunction with the steps shown in FIG6A .
[0237] In step 201, an audio stream is obtained.
[0238] As an example, after the audio code stream is encoded by the audio encoding method shown in Figure 5, the audio code stream is transmitted to the decoding end. After the decoding end receives the audio code stream, it performs audio decoding processing on the audio code stream to reconstruct the reconstructed audio signal.
[0239] In step 202, signal decoding processing is performed on the audio code stream to obtain an estimated value of a coding feature corresponding to the audio code stream.
[0240] It should be noted that signal decoding is the inverse process of signal encoding. Since the decoding process of the received code stream is the inverse of the encoding process, the values generated during the decoding process are estimates relative to the values during the encoding process. For example, the estimated coding feature value generated during the decoding process is an estimate of the coding feature value during the encoding process.
[0241] In step 203, the first decoding network in the audio codec system is called to decode the coding feature estimation value to obtain a reconstructed audio signal corresponding to the audio code stream.
[0242] It should be noted that the decoding process is the inverse of the encoding process. Since the decoding process of the received bitstream is the inverse of the encoding process, the values generated during the decoding process are estimates relative to the values during the encoding process. For example, the reconstructed audio signal generated during the decoding process is an estimate of the audio signal during the encoding process.
[0243] Here, in step 203, the first decoding network in the audio codec system is called, and the process of decoding the coding feature estimation value is similar to the above-mentioned step 44, except that the processing object in step 203 is the coding feature estimation value of the audio code stream, while the processing object in step 44 is the coding feature estimation value of the audio code stream sample.
[0244] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0245] The embodiments of the present application can be applied to various audio scenarios, such as voice calls, instant messaging, etc. The following description will be made using a voice call as an example:
[0246] In related technologies, the principles of speech coding are roughly as follows: speech coding can directly encode speech waveform samples sample by sample; or, based on the principle of human vocalization, relevant low-dimensional features are extracted, the encoding end encodes the features, and the decoding end reconstructs the speech signal based on these parameters.
[0247] The above coding principles are all derived from voice signal modeling, that is, compression methods based on signal processing, which cannot guarantee the encoding quality of audio. In response to this, the embodiments of the present application use technology based on deep learning to improve the encoding and decoding efficiency while ensuring voice quality. At present, technology based on deep learning can indeed bring low bit rate and high quality effects. However, this method also has the following two typical problems:
[0248] 1) This codec method is relatively complex. For some real-time audio and video applications, excessive complexity can limit the widespread adoption of audio applications. Furthermore, the training cycle for deep learning-based codec systems is long, ranging from days to weeks, resulting in high iteration costs.
[0249] 2) In order to ensure the forward compatibility principle in the communication system, the old version of the encoder and decoder can only be replaced first, and then the encoder and decoder can be updated, which makes the engineering transformation difficult.
[0250] In order to solve the above problems, the embodiment of the present application provides a speech encoding and decoding method (i.e., an audio encoding method and an audio decoding method). Based on the characteristics of the audio signal, after processing based on the neural network (NN, Neural Network) technology, a feature vector with a lower dimension than the input is obtained. Among them, the neural network adopts an operation similar to "blocking" inside, which can reduce the complexity of the algorithm and improve the encoding effect. The decoding end decodes the received code stream to obtain the feature vector, calls the inverse process of the corresponding encoding end, and completes the reconstruction of the signal. In particular, the neural network of the decoding end adopts an operation similar to "blocking" inside to reduce the complexity of the algorithm. On the basis of the above, in order to ensure the forward compatibility of the encoding and decoding system, the embodiment of the present application proposes a pre-trained decoder mechanism (i.e., a training method for the audio encoding and decoding system). According to the above scheme, a coding network (denoted as the first coding network) and a decoding network (denoted as the first decoding network) are first trained; then, when the parameters of the first decoding network are fixed, the coding network configuration is modified and retrained to obtain the second coding network. In particular, whether the feature vectors processed and compressed by the first coding network or the second coding network are called, high-quality speech can be restored. It's important to note that by fixing the parameters of the decoding network and allowing for various changes to the encoding network, forward compatibility risks are mitigated. Therefore, online services can be updated without impacting the encoding network. Furthermore, changes to the encoding network can both improve encoding performance and enable lightweight modifications to the encoding network, reducing encoding complexity.
[0251] The embodiment of the present application can be applied to the voice communication link shown in Figure 6B. Taking the Voice over Internet Protocol (VoIP) conference system based on the Internet Protocol as an example, the voice codec technology involved in the embodiment of the present application is deployed in the encoding and decoding parts to solve the basic function of voice compression. The encoder is deployed on the uplink client 601, and the decoder is deployed on the downlink client 602. The voice is collected through the uplink client and pre-processed, enhanced, and encoded. The encoded code stream is transmitted to the downlink client 602 via the network. The downlink client 602 performs decoding, enhancement, and other processing to play back the decoded voice on the downlink client 602.
[0252] To ensure forward compatibility (i.e., compatibility between the new encoder and the existing encoder), a transcoder must be deployed in the system's backend (i.e., server) to ensure interoperability between the new encoder and the existing encoder. For example, if the transmitter (uplink client) is the new NN encoder and the receiver (downlink client) is the Public Switched Telephone Network (PSTN) (G.722), the backend must execute the NN decoder to generate the voice signal and then call the G.722 encoder to generate a specific bitstream to implement the transcoding function. This allows the receiver to correctly decode the specific bitstream.
[0253] Before specifically introducing the training method of the audio codec system provided in the embodiment of the present application, the following first introduces the dilated convolutional network and model training.
[0254] Referring to Figures 8A and 8B, Figure 8A is a schematic diagram of a normal convolution (e.g., causal convolution) network provided in an embodiment of the present application, and Figure 8B is a schematic diagram of a dilated convolution network provided in an embodiment of the present application. Relative to ordinary convolution networks, dilated convolution can increase the receptive field while keeping the size of the feature map unchanged, and can also avoid errors caused by upsampling and downsampling. Although the convolution kernel sizes (Kernel Size) shown in Figures 8A and 8B are both 3×3; however, the receptive field 801 of the ordinary convolution shown in Figure 8A is only 3, while the receptive field 802 of the dilated convolution shown in Figure 8B reaches 5. That is to say, for a convolution kernel of size 3×3, the receptive field of the ordinary convolution shown in Figure 8A is 3, and the dilation rate (the number of intervals between points in the convolution kernel) is 1; while the receptive field of the dilated convolution shown in Figure 8B is 5, and the dilation rate is 2.
[0255] The convolution kernel can also be moved on a plane similar to Figure 8A or Figure 8B. This involves the concept of stride rate. For example, each time the convolution kernel shifts 1 square, the corresponding stride rate is 1.
[0256] There's also the concept of convolution channels, which refers to the number of convolution kernel parameters used to perform the convolution analysis. In theory, a greater number of channels provides a more comprehensive signal analysis and higher accuracy; however, a higher number of channels also increases complexity. For example, a 1×320 tensor can be convolved using 24 channels, resulting in a 24×320 tensor as the output.
[0257] It should be noted that the size of the dilated convolution kernel (for example, for speech signals, the size of the convolution kernel can be set to 1×3), the expansion rate, the shift rate and the number of channels can be defined by yourself according to actual application needs. The embodiments of this application do not make specific restrictions on this.
[0258] Regarding model training, the parameters of the encoding network and decoding network are adjusted simultaneously during the training process. After the encoding network and decoding network converge, the trained encoding network and the trained decoding network can be integrated into the audio codec system to achieve high-quality compression. However, during the iteration process of the audio codec system, on the one hand, it is hoped that the compression quality will be improved through system updates, and on the other hand, the bitstream must be forward compatible, that is, the bitstream compressed by the new version of the encoder can be accurately decoded by the old version of the decoder and generate high-quality speech. For a neural network-based codec system, both the new encoding network and the old encoding network are required to be compatible with the same old decoding network, that is, both can achieve high-quality compression effects.
[0259] Among them, the new and old coding networks (i.e., the new coding network and the old coding network) mainly include the following two situations:
[0260] First, the parameters of the old and new encoding network structures are exactly the same, but a new encoding network is trained for a specific purpose. A typical scenario is to improve the encoding performance of the encoding network for a specific language or for different audio signals.
[0261] Second, the new and old encoding network structures are completely different. For example, based on the original encoding network structure, the number of network layers can be increased or the network structure or parameter count of specific modules can be changed. For example, increasing the complexity of the encoding network can improve encoding performance by increasing computing power; the encoding network structure can also be tailored to reduce encoder complexity and adapt to lightweight applications.
[0262] The model training process can be summarized as follows: Based on the prepared training data and the pre-designed network structure, loss function, and optimizer, the training data is input into the model, the gradient is calculated using the backpropagation algorithm, and the model parameters are updated using the optimizer; the validation dataset is input into the model, the loss function value and prediction accuracy are calculated to evaluate the model's performance; furthermore, based on the validation results, the model's hyperparameters, such as the learning rate and regularization coefficient, are adjusted to achieve the best training effect, and the model is saved; finally, the test dataset is input into the model, and the prediction results and accuracy are calculated to evaluate the model's generalization ability. After the above training process is completed, the model can be deployed and applied, and the trained model can be deployed into real-world applications to predict new input data.
[0263] In the field of speech coding and decoding, end-to-end neural networks are primarily involved. The neural network is divided into two parts: the encoding network and the decoding network. The network structure of each part is predefined. The encoding network maps the input time-domain signal into a low-dimensional feature vector using its nonlinear prediction capabilities. The feature vector is smaller than the input time-domain signal dimension to achieve compression. For example, after processing the input time-domain signal through the encoding network, a low-dimensional feature vector is obtained, where each component of each dimension is [-1, 1] to facilitate normalization. All components of this low-dimensional feature vector are then quantized at a specified compression rate to obtain a quantized low-dimensional feature vector. The decoding network uses the quantized low-dimensional feature vector and its nonlinear prediction capabilities to predict a time-domain signal with the same dimension as the input. Ideally, the predicted time-domain signal should approximate the input time-domain signal. Within a specified compression rate (which determines the quantization accuracy of the low-dimensional feature vector), a smaller prediction error indicates better quality for the reconstructed speech signal. Therefore, end-to-end neural network training is based on the model training and deployment methods described above. Based on the training data, a combination of an encoding network and a decoding network is trained. This combination can compress the input time-domain signal through the encoding network and restore the speech signal through the decoding network.
[0264] The following describes the training method for the audio codec system provided in the embodiment of the present application.
[0265] The training method of the audio codec system provided in the embodiment of the present application is implemented through the scalable neural network training platform as shown in Figure 7.
[0266] Through the scalable neural network training platform, an encoding network (denoted as a first encoding network) and a decoding network (denoted as a first decoding network) are first trained.
[0267] Then, based on the first encoding network and the first decoding network, a first audio encoding and decoding system is constructed. For the input audio signal x(n) of the nth frame, the first encoding network is called to obtain a low-dimensional feature vector F1(n). The dimension of the feature vector F1(n) is smaller than the dimension of the input audio signal to reduce the amount of data. For example, for each frame x(n), the neural network (encoding part) is called to generate a lower-dimensional feature vector F1(n). The embodiment of the present application does not limit other NN structures, such as autoencoder (Autoencoder), fully connected (FC, Full-Connection) network, long short-term memory (LSTM, Long Short-Term Memory) network, convolutional neural network (CNN, Convolutional Neural Network) + LSTM, etc. Among them, the use of similar "blocking" operations inside the neural network can reduce the complexity of the algorithm and improve the encoding effect. For F1(n), code tables with different quantization precisions are used for quantization and encoding to achieve multi-rate encoding and decoding effects.
[0268] For the high frequency sub-band signal x obtained by sub-band decomposition of the input audio signal HB (n), considering that high frequencies are not as important to quality as low frequencies, the high frequency subband signal x HB (n) Other schemes can be used to extract the feature vector F HB (n). For example, the frequency band extension technology based on speech signal analysis can realize the generation of high-frequency sub-band signals with only a small number of bits; it can also use the same NN structure as the low-frequency sub-band signal or a more streamlined network (for example, the output feature vector is smaller than the low-frequency feature vector F LB (n) smaller).
[0269] The eigenvector corresponding to the subband signal (i.e. F LB (n) and F HB (n)) performs vector quantization or scalar quantization, and performs entropy coding on the quantized value, and transmits the encoded code stream (low-frequency code stream and high-frequency code stream) to the decoding end.
[0270] For the estimated value F1′(n) of the feature vector obtained by decoding, the first decoding network is called to generate the estimated value x1′(n) of the audio signal to achieve the audio decoding effect.
[0271] Based on the above-mentioned first encoding network and first decoding network, the parameters of the first decoding network are fixed unchanged, and the scalable neural network training platform is called to retrain a second encoding network.
[0272] A second audio codec system is constructed based on the second encoding network and the first decoding network. For the input audio signal x(n) of frame n, the second encoding network is invoked to obtain a low-dimensional feature vector F2(n). The dimension of feature vector F2(n) is smaller than that of the input audio signal to reduce the data size. Feature vector F2(n) is vector- or scalar-quantized, and the quantized index value is entropy-encoded and transmitted to the decoder.
[0273] For the estimated value F2′(n) of the feature vector obtained by decoding, the first decoding network is called to generate an estimated value x′2(n) of the audio signal to complete the decoding.
[0274] Through the above steps, the same decoder (implemented by the decoding network) can correspond to different encoders (implemented by the encoding network), and one or more audio codec systems (i.e., at least one audio codec system) can be combined. These one or more codec systems can perform voice compression on the same voice input. Because these one or more codec systems use the same decoder, forward compatibility of bitstreams in voice communications is achieved, while also leaving significant room for encoder scalability.
[0275] The following describes in detail the audio encoding method, audio decoding method, and audio coding method of the audio coding system provided in the embodiments of the present application.
[0276] In some embodiments, a speech signal with a sampling rate of Fs = 16000 Hz is used as an example (it should be noted that the method provided in the embodiments of the present application is also applicable to scenarios with other sampling rates, including but not limited to: 8000 Hz, 32000 Hz, and 48000 Hz). At the same time, assuming that the frame length is set to 20 ms, therefore, for Fs = 16000 Hz, each frame contains 320 sample points.
[0277] The encoding end and the decoding end are described in detail below with reference to the first encoding network and the first decoding network shown in FIG7 .
[0278] The process of the first encoding network and the first decoding network is as follows:
[0279] For an audio signal (mono signal) with a sampling rate of Fs=16000 Hz, the input signal of the nth frame includes 320 sample points, which are recorded as input signal x(n).
[0280] Step 11: Encode through the first encoding network.
[0281] Based on the input signal x(n), the first coding network is called to generate a lower-dimensional feature vector F1(n). It should be noted that the dimension of x(n) is 320 and the dimension of F1(n) is 56. In terms of data volume, the first coding network plays a role of "dimensionality reduction" and realizes the function of data compression. Among them, the embodiments of the present application are not limited to the dimension of F1(n), and can also be other dimensions smaller than x(n).
[0282] Referring to the network structure diagram of the first coding network shown in FIG9 , the process of data compression performed by the first coding network is described in detail below:
[0283] First, a 16-channel causal convolution is called to expand the input tensor (i.e., vector) to a 16×320 tensor.
[0284] Then, the 16×320 tensor is preprocessed. For example, after performing a convolution operation on the 16×320 tensor, a pooling operation with a factor of 2 is performed, and the activation function can be PReLU to generate a 16×160 tensor.
[0285] Next, four encoding blocks with different downsampling factors (Down_factor) are cascaded. Each encoding block contains a residual block, a convolutional layer, and a pooling layer. Each residual block includes five residual units (RUs) based on dilated convolution (the input and output feature dimensions of the RUs do not change). A convolutional layer is used to double the number of input channels, and the activation function can be PReLU to ensure data volume and avoid data loss. The pooling layer is a pooling operation with a Down_factor to complete downsampling and achieve data compression. Here, the Down_factor of the four encoding blocks is set to 2, 4, 4, and 5, respectively. Therefore, the number of output channels of the four encoding blocks is set to 32, 64, 128, and 256, respectively. After processing by the four encoding blocks, the input 16×160 tensor is converted into 32×80, 64×20, 128×5, and 256×1 tensors, respectively. The embodiment of the present application is not limited to the number of coding blocks, and can be any positive integer such as 2, 3, 4, 5, etc. In addition, the embodiment of the present application is not limited to the number of residual units in a coding block, and can be any positive integer such as 2, 3, 4, 5, 6, etc. The number of residual units in multiple coding blocks can be the same or different, for example, one coding block contains 4 residual units and another coding block contains 5 residual units.
[0286] Here we further introduce the residual unit. The residual unit refers to a module in a deep neural network. By introducing cross-layer connections in the neural network, the neural network is easier to optimize during the training process, avoiding problems such as gradient disappearance or gradient explosion. The core idea is to perform residual learning on the input within the module, that is, to bypass a part of the layer through a direct path and pass the input information directly to the output, so that the network can better utilize shallow feature information during the learning process. Figure 10A is a schematic diagram of the residual block structure used in the encoding block in the first encoding network. This residual block includes 4 residual units based on dilated convolution, and each residual unit contains a dilated convolution block with a specified dilation rate (Dilation rate), that is, each dilated convolution block contains a convolution operator with a specified dilation rate (such as Dilation rate = 3). In an embodiment of the present application, the use of dilated convolution blocks with 5 progressive dilation rates is equivalent to using different receptive fields to extract input features at different resolutions, which can better perform a relatively comprehensive analysis of the data. After the residual is processed by 5 dilation-rate-specified dilation-rate dilated convolution blocks, it is added to the input from the jump connection to obtain the output of the residual block, and the output is sent to the convolution layer connected to the residual block.
[0287] Here, any residual unit in Figure 10A is further described, as shown in Figure 10B. For any residual unit, it contains at least one dilated convolution with a specified expansion rate (used to expand the receptive field), and PReLU can be used as the activation function; in addition, one or more causal convolutions (used to extract local information) can be cascaded, and PReLU can be used as the activation function. The convolution kernel size of the dilated convolution with the above-mentioned specified expansion rate can be 3, 5, 7, 9, etc., and the convolution kernel size of the above-mentioned causal convolution can be 1, 3, etc. This embodiment of the present application does not limit the convolution kernel size of the dilated convolution or causal convolution with the above-mentioned specified expansion rate. In addition, the causal convolution or dilated convolution in the embodiment of the present application can also be implemented by other specific convolution units with similar or equivalent functions.
[0288] Furthermore, for residual units, a grouped convolution algorithm is introduced to reduce algorithmic complexity. Grouped convolution divides the input channels into multiple groups for convolution operations, where only the input channels within each group are associated with the output channels. Here, assume there are 16 input channels and 32 output channels. If the number of groups is 1, each input channel is associated with all 32 output channels. If the number of groups is 2, the 16 input channels are first divided into two groups, 0-7 and 8-15. Within each group, the input channels are associated with the output channels within that group. For example, input channels 0-7 in the first group are associated with output channels 0-15, while input channels 8-15 in the second group are associated with output channels 16-31. For example, output channel 0 is associated only with input channels 0-7, not with input channels 8-15. Output channel 25 is associated only with input channels 8-15, not with input channels 0-7. From this comparison, it can be seen that the introduction of grouped convolution can avoid the association between any input channel and all output channels, reduce the number of connections, and reduce complexity. Of course, since the larger the number of groups, the smaller the correlation between the input channel and the output channel, which will also affect the encoding effect, it is not the case that the larger the number of groups, the better. In the embodiment of the present application, the hole convolution contained in the 5 residual blocks corresponding to the 4 coding blocks can use different grouping number configurations. The specific grouping number configuration is shown in Table 1.
[0289] Table 1. Grouping configurations used by residual units in different coding blocks
[0290] Finally, the 256×1 tensor undergoes similar preprocessing causal convolution to output a 56-dimensional feature vector F1(n). According to the calculation of the first encoding network, each value in the 56-dimensional feature vector F1(n) is between [-1, 1].
[0291] Step 12: quantization encoding.
[0292] For the low-dimensional feature vector F1(n), scalar quantization (each component is quantized separately) and entropy coding methods can be performed. In addition, the embodiments of the present application do not limit the technical combination of vector quantization (combining multiple adjacent components into a vector for joint quantization) and entropy coding.
[0293] Here we explain the quantization coding of the low-dimensional feature vector F1(n). According to the above description, after the input signal is processed by the first coding network, a 56-dimensional feature vector F1(n) is obtained. The embodiment of the present application provides a method based on scalar quantization and entropy coding, including: 1) For each dimension in F1(n), the interval [-1,1] is divided into 11 equal parts to form a codebook containing 11 elements, and each dimension value is quantized into one of the 11 elements; 2) According to the Shannon entropy theorem, for a codebook containing 11 elements and uniformly distributed, the entropy (average bit) is Therefore, the average bit rate corresponding to each frame of the low-frequency subband signal is 193.76 bits; 3) For every 20ms framing method, there are 50 frames in 1 second, so the average bit rate is 9.69kbps (i.e., bit rate mode). According to the entropy coding theory, probability distribution statistics can be performed on each of the above dimensions to generate 56 code tables. Generally, each dimension is non-uniformly distributed, so the actual bit rate is around 9.69kbps, or less than 9.69kbps. Generally, the higher the bit rate mode, the more bits the encoder will use for encoding, and the corresponding reconstructed speech quality will be higher. It should be noted that the embodiment of the present application is not limited to the bit rate mode, and can be 8.88kbps, 6.50kbps, etc.
[0294] In summary, after quantization coding, a code stream can be generated. In the range of 5-10 kbps, the embodiment of the present application can achieve high-quality compression for 16 kHz broadband signals.
[0295] Step 13: Quantization decoding.
[0296] Quantization decoding is the inverse process of quantization encoding. For the received bit stream (including high-frequency bit stream and low-frequency bit stream), entropy decoding is first performed, and the estimated value F′1(n) of the low-dimensional feature vector is obtained by looking up the quantization table.
[0297] Step 14: Decoding is performed through the first decoding network.
[0298] First, based on the estimated value F′1(n) of the low-dimensional feature vector, the first decoding network shown in Figure 11 is called to generate an estimated value x′1(n) of the audio signal, that is, to reconstruct the audio signal. Among them, the first decoding network is similar to the first encoding network, such as causal convolution, and the post-processing structure is similar to the pre-processing structure in the first encoding network. The decoding block structure is symmetrical with the encoding block on the encoding side. The encoding block on the encoding side first performs a dilated convolution and then pooling to complete downsampling. The decoding block on the decoding side first performs pooling to complete upsampling and then performs a dilated convolution. The specific process of the first decoding network is as follows:
[0299] First, a causal convolution is called, which can transform the input tensor F′ into LB(n), a tensor expanded from 56×1 to 256×1.
[0300] Next, cascade four decoding blocks with different upsampling factors (Up_factor). Each decoding block contains a convolution layer, an upsampling module, and a residual block, wherein a convolution layer is used to halve the number of input channels; an upsampling module contains a specific Up_factor for completing upsampling; and a residual block includes five residual units (Residual Unit) based on hole convolution. The Up_factor of the four decoding blocks is set to 5, 4, 4, and 2, respectively. Therefore, the number of output channels of the four decoding blocks is set to 128, 64, 32, and 16, respectively. After processing by the four decoding blocks, the 256×1 tensor is converted into tensors of 128×5, 64×20, 32×80, and 16×160, respectively. Among them, the embodiment of the present application is not limited to the number of decoding blocks, and can be any positive integer such as 2, 3, 4, or 5.
[0301] Here, for the upsampling module containing a specific Up_factor, a Repeat operation can be used to complete the upsampling operation by repeated filling, thereby saving complexity.
[0302] The configuration of the five residual units based on dilated convolution at the decoder is similar to that of the residual units at the encoder, including but not limited to the internal structure of the residual units, convolution kernel size, and dilation rate. Table 2 shows the number of groups used for dilated convolution in the decoder block. The decoder block uses a larger number of groups, 2, to associate more input and output channels, improving speech reconstruction quality.
[0303] Table 2. Grouping configurations used by residual units in different decoding blocks
[0304] Then, the 16×160 tensor output by the cascaded decoding block is post-processed. For example, a Repeat operation with a factor of 2 is performed on the 16×160 tensor output by the cascaded decoding block to complete upsampling, and then a convolution operation is performed and an activation function is used for the PReLU operation to generate a 16×320 tensor.
[0305] Finally, a causal convolution is called to convert the input 16×320 tensor into a 1×320 tensor to reconstruct the input signal.
[0306] The following describes the training process of the first encoding network and the first decoding network:
[0307] The embodiment of the present application can collect data and jointly train the relevant networks of the encoding and decoding ends to obtain optimal parameters. The user only needs to prepare the data and set the corresponding network structure. After the training is completed in the background, the trained model can be put into use.
[0308] It's important to note that adversarial learning training can be used in end-to-end neural network encoding and decoding systems. Adversarial learning training works by pitting a generative model (generator) against a discriminator (discriminator) to improve the performance of the generative model. Specifically, the generative model attempts to generate realistic samples to fool the discriminator, while the discriminator attempts to discern the differences between real and generated samples. This adversarial process continues iteratively until the quality of samples generated by the generative model is sufficiently high.
[0309] When the adversarial learning training mechanism is used to train the first encoding network and the first decoding network, the first encoding network and the first decoding network can be used as generative models, and the performance of the first encoding network and the first decoding network can be improved by using the adversarial learning training mechanism to train the first encoding network and the first decoding network.
[0310] As described above, after the first encoding network and the first decoding network are trained, the first audio codec system is obtained. In actual applications, there is still a need for further version updates, such as: fine-tuning the effect for specific applications (for example, increasing the adaptability of a certain language); further improving voice quality by increasing the complexity of the encoding network; and serving some lightweight applications by reducing the complexity of the encoding network.
[0311] The aforementioned update requirements require a crucial prerequisite: the decoder of an older version can correctly decode the bitstream from the encoder of a newer version. For an audio codec system based on an end-to-end neural network, the decoder network parameters must be fixed before the encoder network is retrained.
[0312] Similar to the training process for the first encoding network and the first decoding network, the second audio codec system (including the second encoding network and the first decoding network) of the end-to-end neural network can use an adversarial learning training mechanism. The adversarial learning training principle is to improve the performance of the generative model by pitting the generative model against the discriminative model.
[0313] The difference from the training process of the first encoding network and the first decoding network is that the training process of the second encoding network and the first decoding network is based on the model (including parameters) that has been trained for the first encoding network and the first decoding network. The training process of the second encoding network and the first decoding network is as follows:
[0314] 1) Load the trained model parameters of the first encoding network and the first decoding network.
[0315] 2) In the scalable neural network training platform, the configuration of the second encoding network is used to replace the first encoding network to modify the configuration of the encoding network. The configuration includes but is not limited to the model structure, loss function, optimizer, etc.
[0316] Among them, the second coding network is flexibly set up, that is, compared with the first coding network, the model complexity can be increased to improve the quality, or the model complexity can be reduced to be suitable for lightweight applications, etc. The above-mentioned modification of the coding network configuration includes but is not limited to: 1) "isomorphic" configuration, the network structure and parameter quantity of the first coding network and the second coding network are exactly the same, that is, the second coding network is an isomorphic network of the first coding network, and the second coding network is obtained by retraining by modifying the training data or training strategy; 2) "heterogeneous" configuration, the network structure and parameter quantity of the first coding network and the second coding network change (such as the second coding network has more layers than the first coding network, or the intermediate variable dimensions of some layers of the second coding network are more), that is, the second coding network is a heterogeneous network of the first coding network. In addition, changes in training data or training strategies in the "isomorphic" configuration can also be applied to scenarios with "heterogeneous" configurations.
[0317] It should be noted that in this embodiment of the present application, the gradient update in the first decoding network needs to be set to false. By setting the gradient update in the first decoding network to false, when training the second encoding network and the first decoding network end-to-end, only the parameters of the first decoding network can be fixed, so the first decoding network does not participate in the gradient calculation. In this way, when retraining the second encoding network and the first decoding network, any parameter updates are only performed in the second encoding network.
[0318] The training data used to train the second encoding network and the first decoding network may be consistent with or inconsistent with the training data used to train the first encoding network and the first decoding network.
[0319] The above configuration yields a second encoding network and a first decoding network, i.e., a second audio codec system. This ensures forward compatibility of the codestream: compression is performed based on the feature vector F1(n) generated by the first encoding network, and compression is performed based on the feature vector F2(n) generated by the second encoding network. After quantization and decoding, feature vectors F1(n) and F2(n) can be decoded by the same first decoding network to generate a speech signal.
[0320] After the second encoding network and the first decoding network are trained, they can be put into use. The following describes the application of the second encoding network and the first decoding network:
[0321] The second encoding network, shown in Figure 12, differs from the first in that a one-dimensional convolutional layer is added to the end of the first encoding network. The input variables of this one-dimensional convolutional layer have a dimension of [56, 1], and the output variable F2(n) has a dimension of [56, 1]. Increasing the number of encoding network layers improves the network's feature extraction capabilities. Similarly, the number of encoding network layers can be further increased, for example, by adding one or more residual units to the one-dimensional convolutional layer.
[0322] The quantization encoding of F2(n) can be performed in the same manner as the quantization encoding of F1(n). In this way, the operation at the encoding end is completed.
[0323] On the decoding side, consistent with the above embodiment, decoding obtains an estimate of the low-dimensional feature vector F′2(n), and the first decoding network is invoked to generate an estimate of the input signal x′2(n). The decoding side operates in the same manner as in the above embodiment. Specifically, the decoding side reuses the first decoding network (with identical parameters). This resolves the forward compatibility issue.
[0324] In summary, the training method, audio encoding method, and audio decoding method of the audio codec system provided in the embodiments of the present application significantly improve the audio quality while ensuring acceptable complexity compared to signal processing schemes through the organic combination of signal decomposition, signal processing technology and deep neural networks.
[0325] So far, the audio encoding method or audio decoding method provided by the embodiment of the present application has been described in combination with the exemplary application and implementation of the terminal device provided by the embodiment of the present application. The embodiment of the present application also provides an audio encoding device and an audio decoding device. In actual applications, the functional modules in the audio encoding device and the audio decoding device can be implemented by the hardware resources of the electronic device (such as terminal equipment, server or server cluster), such as computing resources such as processors, communication resources (such as for supporting various communication modes such as optical cables and cellular), and memories. Figure 3A shows a training device 555 of an audio codec system stored in a memory 550, Figure 3B shows an audio encoding device 655 stored in a memory 650, and Figure 3C shows an audio encoding device 755 stored in a memory 750. It can be software in the form of programs and plug-ins, for example, software modules designed in programming languages such as C / C++ and Java, application software designed in programming languages such as C / C++ and Java, or dedicated software modules in large software systems, application program interfaces, plug-ins, cloud services, etc. The following examples illustrate different implementation methods.
[0326] The training device 555 for the audio codec system includes a series of modules, including a first acquisition module 5551, a determination module 5552, a construction module 5553, and a training module 5554. The following further describes how the various modules in the training device 555 for the audio codec system provided in the embodiment of the present application cooperate to implement the training scheme for the audio codec system.
[0327] The first acquisition module 5551 is configured to acquire a first audio codec system, wherein the first audio codec system includes a first encoding network and a first decoding network; the determination module 5552 is configured to generate a second encoding network to be trained corresponding to the first encoding network in response to a configuration request for the first encoding network; the training module 5553 is configured to encode the first audio sample based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample, and decode the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample; and update the parameters of the second encoding network to be trained based on the reconstructed audio sample to obtain a trained second encoding network.
[0328] In some embodiments, the configuration request includes a homogeneous network configuration request for the first encoding network, and the homogeneous network configuration request is used to indicate the modification of at least one of the following first configuration data: second audio samples for training the first audio codec system, and a training strategy for training the first audio codec system.
[0329] In some embodiments, when the homogeneous network configuration request is used to indicate modification of the second audio sample, the first audio sample is the second audio sample modified based on the homogeneous network configuration request; when the homogeneous network configuration request is used to indicate modification of the training strategy, the modified training strategy is used to train the second coding network to be trained in combination with the first audio sample to obtain the trained second coding network.
[0330] In some embodiments, the configuration request includes a heterogeneous network configuration request for the first coding network, and the heterogeneous network configuration request is used to indicate modification of at least one of the following second configuration data: a network structure of the first coding network, and a parameter quantity of the first coding network.
[0331] In some embodiments, the determination module 5552 is further configured to modify the second configuration data of the first encoding network based on the heterogeneous network configuration request to obtain the second encoding network to be trained.
[0332] In some embodiments, the generative adversarial network includes a generative model and a discriminative model, wherein the generative model is a second audio codec system, and the second audio codec system includes the second encoding network to be trained and the first decoding network; the training module 5553 is also configured to alternately perform the following training tasks based on the generative model and the discriminative model in the generative adversarial network: based on the first audio sample, train the generative model, wherein the trained generative model is used to generate a reconstructed audio sample based on the audio sample; based on the second audio sample and the reconstructed audio sample, train the discriminative model, wherein the trained discriminative model is used to distinguish between the audio sample and the reconstructed audio sample; wherein, when training the generative model, the parameters of the discriminative model are fixed unchanged; when training the discriminative model, the parameters of the generative model are fixed unchanged; and the second encoding network in the trained generative model is used as the trained second encoding network.
[0333] In some embodiments, the training module 5553 is further configured to perform network coding processing on the first audio sample through the second coding network to be trained to obtain the coding features of the first audio sample; and perform signal coding processing on the coding features of the first audio sample to obtain an audio code stream sample of the first audio sample.
[0334] In some embodiments, the second encoding network to be trained includes the network structure and one-dimensional convolution layer of the first encoding network; the training module 5553 is further configured to perform network coding processing on the first audio sample through the network structure of the first encoding network included in the second encoding network to be trained, to obtain the initial encoding features of the first audio sample; and perform convolution processing on the initial encoding features through the one-dimensional convolution layer included in the second encoding network to be trained, to obtain the encoding features of the first audio sample.
[0335] In some embodiments, the training module 5553 is further configured to perform the following processing through the network structure of the first encoding network included in the second encoding network to be trained: performing feature extraction processing on the first audio sample to obtain audio features of the first audio sample; and using at least one residual unit in the network structure to perform residual processing on the audio features to obtain initial encoding features of the first audio sample.
[0336] In some embodiments, the training module 5553 is further configured to perform signal decoding processing on the audio code stream sample of the first audio sample to obtain a coding feature estimation value corresponding to the audio code stream sample; and perform network decoding processing on the coding feature estimation value through the first decoding network to obtain the reconstructed audio sample.
[0337] In some embodiments, the training module 5553 is further configured to perform the following processing through the first decoding network: performing residual processing on the coding feature estimation value using at least one residual unit included in the first decoding network to obtain an audio feature estimation value corresponding to the audio code stream sample; performing feature reconstruction processing on the audio feature estimation value to obtain a reconstructed audio sample corresponding to the audio code stream sample.
[0338] In some embodiments, when the first audio sample is a low-frequency sub-band signal obtained by sub-band decomposing an audio signal sample, the audio code stream sample is a low-frequency code stream sample corresponding to the audio signal sample; when the first audio sample is the audio signal sample, the audio code stream sample is a full-frequency code stream sample corresponding to the audio signal sample.
[0339] The audio encoding device 655 includes a series of modules, including a second acquisition module 6551, an encoding module 6552, and a signal encoding module 6553. The following further describes how the various modules in the audio encoding device 655 provided in the embodiment of the present application cooperate to implement the audio encoding solution.
[0340] The second acquisition module 6551 is configured to acquire an audio signal; the encoding module 6552 is configured to call the trained second encoding network in the audio codec system to perform network coding processing on the audio signal to obtain a second encoding feature of the audio signal, wherein the audio codec system includes the trained second encoding network and the first decoding network; the signal encoding module 6553 is configured to perform signal coding processing on the second encoding feature of the audio signal to obtain a second audio code stream of the audio signal; wherein both the first audio code stream and the second audio code stream can be decoded by the first decoding network to obtain a reconstructed audio signal corresponding to the audio signal, the first audio code stream is the audio code stream obtained after the audio signal is processed by the first encoding network, and the trained second encoding network is trained for the first encoding network using the training method of the audio codec system.
[0341] The audio decoding device 755 includes a series of modules, including a third acquisition module 7551, a signal decoding module 7552, and a decoding module 7553. The following further describes how the various modules in the audio decoding device 555 provided in the embodiment of the present application cooperate to implement the audio decoding solution.
[0342] The third acquisition module 7551 is configured to obtain an audio code stream; the signal decoding module 7552 is configured to perform signal decoding processing on the audio code stream to obtain a coding feature estimation value corresponding to the audio code stream; the decoding module 7553 is configured to call the first decoding network in the audio codec system, perform network decoding processing on the coding feature estimation value, and obtain a reconstructed audio signal corresponding to the audio code stream; wherein the audio codec system includes a trained second coding network and the first decoding network, the audio code stream is obtained after the audio signal is processed by the trained second coding network or the first coding network, and the trained second coding network is obtained by training the first coding network using the training method of the audio codec system.
[0343] The embodiments of the present application provide a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer program or computer-executable instructions, causing the electronic device to perform the training method, audio encoding method, and audio decoding method of the audio codec system described in the embodiments of the present application.
[0344] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the training method, audio encoding method, and audio decoding method of the audio codec system provided in the embodiment of the present application, for example, the training method of the audio codec system shown in Figure 4A.
[0345] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.
[0346] In some embodiments, computer executable instructions (executable instructions for short) may be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as standalone programs or as modules, components, subroutines or other units suitable for use in a computing environment.
[0347] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0348] As an example, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0349] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0350] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A training method for an audio encoding and decoding system, applied to an electronic device, the method comprising: Obtaining a first audio encoding and decoding system, wherein the first audio encoding and decoding system includes a first encoding network and a first decoding network; Responding to a configuration request for the first encoding network, generating a second encoding network to be trained corresponding to the first encoding network; Encoding a first audio sample based on the second encoding network to be trained, obtaining an audio bitstream sample of the first audio sample, and decoding the audio bitstream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample; Updating parameters of the second encoding network to be trained based on the reconstructed audio sample to obtain a trained second encoding network.
2. The method according to claim 1, wherein The configuration request includes a homogeneous network configuration request for the first encoding network, and the homogeneous network configuration request is used to indicate modifying at least one of the following first configuration data: a second audio sample for training the first audio encoding and decoding system, a training strategy for training the first audio encoding and decoding system.
3. The method according to claim 2, wherein When the homogeneous network configuration request is used to indicate modifying the second audio sample, the first audio sample is the second audio sample modified based on the homogeneous network configuration request; When the homogeneous network configuration request is used to indicate modifying the training strategy, the modified training strategy is used to train the second encoding network to be trained in combination with the first audio sample to obtain the trained second encoding network.
4. The method according to any one of claims 1-3, wherein The configuration request includes a heterogeneous network configuration request for the first encoding network, and the heterogeneous network configuration request is used to indicate modifying at least one of the following second configuration data: the network structure of the first encoding network, the number of parameters of the first encoding network.
5. The method according to claim 4, wherein The step of responding to a configuration request for the first encoding network and generating a second encoding network to be trained corresponding to the first encoding network includes: Modifying the second configuration data of the first encoding network based on the heterogeneous network configuration request to obtain the second encoding network to be trained.
6. The method according to any one of claims 1-5, wherein A generative adversarial network includes a generator model and a discriminator model, the generator model is a second audio encoding and decoding system, and the second audio encoding and decoding system includes the second encoding network to be trained and the first decoding network; The second encoding network to be trained is obtained by the following method: Based on the generator model and the discriminator model in the generative adversarial network, alternately performing the following training tasks: Training the generator model based on the first audio sample, wherein the trained generator model is used to generate a reconstructed audio sample based on an audio sample; Training the discriminator model based on the second audio sample and the reconstructed audio sample, wherein the trained discriminator model is used to distinguish between the audio sample and the reconstructed audio sample; Among them, when training the generation model, the parameters of the discriminant model are fixed; when training the discriminant model, the parameters of the generation model are fixed; The second encoding network in the trained generation model is used as the trained second encoding network.
7. The method according to any one of claims 1-6, wherein, The encoding the first audio sample based on the second encoding network to be trained to obtain an encoding result of the first audio sample includes: Performing network encoding processing on the first audio sample through the second encoding network to be trained to obtain The encoding features of the first audio sample; Performing signal encoding processing on the encoding features of the first audio sample to obtain an audio code stream sample of the first audio sample.
8. The method according to claim 7, wherein The second encoding network to be trained includes the network structure of the first encoding network and a one-dimensional convolutional layer; The encoding the first audio sample through the second encoding network to be trained to obtain the encoding features of the first audio sample includes: Performing network encoding processing on the first audio sample through the network structure of the first encoding network included in the second encoding network to be trained to obtain initial encoding features of the first audio sample; Performing convolutional processing on the initial encoding features through the one-dimensional convolutional layer included in the second encoding network to be trained to obtain the encoding features of the first audio sample.
9. The method according to claim 8, wherein, The encoding the first audio sample through the network structure of the first encoding network included in the second encoding network to be trained to obtain the initial encoding features of the first audio sample includes: Performing the following processing through the network structure of the first encoding network included in the second encoding network to be trained: Performing feature extraction processing on the first audio sample to obtain audio features of the first audio sample; Performing residual processing on the audio features by using at least one residual unit in the network structure to obtain the initial encoding features of the first audio sample.
10. The method according to any one of claims 1-9, wherein, The decoding the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample includes: Performing signal decoding processing on the audio code stream sample of the first audio sample to obtain an estimated value of the encoding features corresponding to the audio code stream sample; Performing network decoding processing on the estimated value of the encoding features through the first decoding network to obtain the reconstructed audio sample.
11. The method according to claim 10, wherein, The decoding the estimated value of the encoding features through the first decoding network to obtain the reconstructed audio sample includes: Performing the following processing through the first decoding network: Performing residual processing on the estimated value of the encoding features by using at least one residual unit included in the first decoding network to obtain an estimated value of the audio features corresponding to the audio code stream sample; Performing feature reconstruction processing on the estimated value of the audio features to obtain the reconstructed audio sample.
12. The method according to any one of claims 1-11, wherein When the first audio sample is a low-frequency subband signal obtained by performing subband decomposition on an audio signal sample, the audio code stream sample is the low-frequency code stream sample corresponding to the audio signal sample; When the second audio sample is the audio signal sample, the audio code stream sample is the full-frequency code stream sample corresponding to the audio signal sample.
13. An audio encoding method, applied to an electronic device, the method comprising: Obtaining an audio signal; Invoking a trained second encoding network in an audio codec system to perform network encoding processing on the audio signal to obtain a second encoding feature of the audio signal, wherein the audio codec system includes the trained second encoding network and a first decoding network; Performing signal encoding processing on the second encoding feature of the audio signal to obtain a second audio code stream of the audio signal; Wherein both a first audio code stream and the second audio code stream can decode a reconstructed audio signal corresponding to the audio signal through the first decoding network, the first audio code stream is an audio code stream obtained after the audio signal is processed by a first encoding network, and the trained second encoding network is trained for the first encoding network by the training method of the audio codec system according to any one of claims 1-12.
14. An audio decoding method, applied to an electronic device, the method comprising: Obtaining an audio code stream; Performing signal decoding processing on the audio code stream to obtain an estimated value of the encoding feature corresponding to the audio code stream; Invoking a first decoding network in an audio codec system to perform network decoding processing on the estimated value of the encoding feature to obtain a reconstructed audio signal corresponding to the audio code stream; Wherein the audio codec system includes a trained second encoding network and the first decoding network, the audio code stream is obtained after the audio signal is processed by the trained second encoding network or the first encoding network, and the trained second encoding network is trained for the first encoding network by the training method of the audio codec system according to any one of claims 1-12.
15. A training device for an audio codec system, the device comprising: A first obtaining module configured to obtain a first audio codec system, wherein the first audio codec system includes a first encoding network and a first decoding network; A determining module configured to generate a second encoding network to be trained corresponding to the first encoding network in response to a configuration request for the first encoding network; A training module configured to perform encoding processing on a first audio sample based on the second encoding network to be trained to obtain an audio code stream sample of the first audio sample, and perform decoding processing on the audio code stream sample based on the first decoding network to obtain a reconstructed audio sample of the first audio sample; Updating parameters of the second encoding network to be trained based on the reconstructed audio sample to obtain a trained second encoding network.
16. An electronic device, the electronic device comprising: A memory for storing a computer program or computer-executable instructions; A processor, when executing a computer program or computer-executable instructions stored in the memory, implements the training method of the audio codec system according to any one of claims 1 to 12, or the audio encoding method according to claim 13, or the audio decoding method according to claim 14.
17. A computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the training method of the audio codec system according to any one of claims 1 to 12, or the audio encoding method according to claim 13, or the audio decoding method according to claim 14.
18. A computer program product includes computer-executable instructions, which, when executed by a processor, implement the training method of the audio codec system according to any one of claims 1 to 12, or the audio encoding method according to claim 13, or the audio decoding method according to claim 14.
Citation Information
Patent Citations
Voice processing method, device, system and equipment and storage medium
CN114842857A
System and method for training audio codec
CN116011556A
Training method, coding method, decoding method and device of audio coding and decoding system
CN117831548A
Self-supervised pitch estimation
US20220343896A1
Cited By
Controllable causal double-path speech enhancement method and system
CN120510860A
Voice coding and decoding method, device, equipment and medium
CN121054006A