Audio coding and decoding method and device, equipment and storage medium

By determining the structure delay of the audio data at the encoding end and determining the forward audio data, the problem of uncontrollable delay of the audio data encoding and decoding structure in the prior art is solved, and efficient audio encoding and decoding is achieved.

CN120220703APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311805291.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the existing end-to-end audio encoding and decoding schemes process the length of the data to be encoded is less than the minimum input length, the delay of the encoding and decoding structure of the audio data is uncontrollable, which in turn affects the encoding and decoding performance.

Method used

The forward audio data is determined by determining the current audio data to be encoded and the target structure delay at the encoding end, and based on the structural delay corresponding to the target structure delay and the current audio data, the forward audio data is determined. Then, based on the current audio data and forward audio data, the input data of the encoded network, including the current audio data and forward audio data, is determined, and then encoded through the encoded network to obtain an audio code stream.

Benefits of technology

It realizes that while ensuring the quality of audio data encoding and decoding, the structure delay is controlled and the audio encoding and decoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220703A_ABST
    Figure CN120220703A_ABST
Patent Text Reader

Abstract

The invention provides an audio encoding and decoding method, device and equipment and a storage medium, the method can be applied to the fields of artificial intelligence, audio encoding and decoding and the like, and the method comprises the following steps: obtaining a first audio code stream which is obtained by encoding input data through an encoding network, the input data comprises to-be-coded current audio data and forward audio data of the current audio data, and the forward audio data is determined based on target structure time delay and structure time delay corresponding to the current audio data; determining an input feature vector of a decoding network based on the first audio code stream; and decoding the input feature vector through a decoding network to obtain reconstruction data of the current audio data. The input audio data of the coding network comprises the current audio data and the forward audio data of the current audio data, so that the decoding quality can be improved, and the structure time delay can be controlled, thereby improving the decoding performance of the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to an audio encoding and decoding method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of deep learning technologies, deep learning technologies have been widely applied to signal processing technologies in different dimensions, such as audio, images, and videos. Taking audio signals as an example, in an end-to-end audio encoding and decoding scheme based on deep learning, an encoder maps an audio signal to a coded vector through an encoding network, and further generates a corresponding binary bitstream file through quantization technology. A decoder obtains the transmitted information by reading the binary bitstream file, obtains the corresponding coded vector through inverse quantization technology, and then uses the coded vector as the input of the decoding network to decode and obtain the final reconstructed audio signal.

[0003] In current end-to-end audio encoding and decoding schemes, the encoding network has a minimum input length requirement. When the length of the data to be encoded is less than the minimum input length, data padding is required. However, the current padding method makes the structural delay of the audio data encoding and decoding uncontrollable, thereby resulting in poor audio encoding and decoding performance. Summary of the Invention

[0004] The present application provides an audio encoding and decoding method, apparatus, device, and storage medium, which can control the structural delay while ensuring the encoding and decoding quality of audio data, thereby improving the audio encoding and decoding efficiency.

[0005] In a first aspect, the present application provides an audio decoding method, including:

[0006] Obtain a first audio bitstream, where the first audio bitstream is obtained by encoding input data through an encoding network, and the input data includes current audio data to be encoded and forward audio data of the current audio data, and the forward audio data is determined based on a target structural delay and a structural delay corresponding to the current audio data;

[0007] Based on the first audio bitstream, determine an input feature vector of a decoding network;

[0008] Decode the input feature vector through a decoding network to obtain reconstructed data of the current audio data.

[0009] In a second aspect, the present application provides an audio encoding method, including:

[0010] Determine current audio data to be encoded and a target structural delay;

[0011] Determine the structural delay corresponding to the current audio data, and determine the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data;

[0012] Based on the current audio data and the forward audio data, determine the input data of the encoding network, where the input data includes the current audio data and the forward audio data;

[0013] Encode the input data through the encoding network to obtain a first audio bitstream.

[0014] In a third aspect, the present application provides an audio decoding device, including:

[0015] An acquisition unit, configured to acquire a first audio bitstream, where the first audio bitstream is obtained by encoding input data through an encoding network, the input data includes the current audio data to be encoded and the forward audio data of the current audio data, and the forward audio data is determined based on a target structural delay and the structural delay corresponding to the current audio data;

[0016] An input determination unit, configured to determine an input feature vector of the decoding network based on the first audio bitstream;

[0017] A decoding unit, configured to decode the input feature vector through the decoding network to obtain the reconstructed data of the current audio data.

[0018] In some embodiments, the input determination unit is specifically configured to perform inverse quantization on the first audio bitstream to obtain a first feature vector; determine the input feature vector of the decoding network based on the scale size of the first feature vector and the scale size of the input feature vector of the decoding network.

[0019] In some embodiments, the input determination unit is specifically configured to determine the scale difference between the scale size of the input feature vector of the decoding network and the scale size of the first feature vector; based on the scale difference, determine a second feature vector from the feature vectors of the backward audio data of the current audio data; and obtain the input feature vector of the decoding network based on the first feature vector and the second feature vector.

[0020] In some embodiments, the structural delay corresponding to the current audio data is determined based on the sampling rate and the number of samples included in the current audio data.

[0021] In some embodiments, the structural delay corresponding to the current audio data is the ratio of the number of samples included in the current audio data to the sampling rate.

[0022] In some embodiments, the forward audio data is determined based on the structural delay corresponding to the forward audio data, and the structural delay corresponding to the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data.

[0023] In some embodiments, the structural delay corresponding to the forward audio data is the difference between the target structural delay and the structural delay corresponding to the current audio data.

[0024] In some embodiments, the forward audio data is determined based on the number of samples included in the forward audio data, and the number of samples included in the forward audio data is determined based on the structural delay corresponding to the forward audio data and the sampling rate.

[0025] In some embodiments, the number of samples included in the forward audio data is the product of the structural delay corresponding to the forward audio data and the sampling rate; alternatively, the number of samples included in the forward audio data is determined based on a first number of samples and the total number of samples of the input data of the encoding network, and the first number of samples is the product of the structural delay corresponding to the forward audio data and the sampling rate.

[0026] In some embodiments, if the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is the first number of samples; if the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data.

[0027] In some embodiments, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, then the input data of the encoding network includes the current audio data and the forward audio data; if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then the input data of the encoding network includes the current audio data, the backward audio data of the current audio data, and the forward audio data.

[0028] In some embodiments, the backward audio data is obtained based on the number of samples included in the backward audio data, and the number of samples included in the backward audio data is determined based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data.

[0029] In some embodiments, the number of samples included in the backward audio data is the total number of samples of the input data of the encoding network minus the number of samples included in the current audio data and the number of samples included in the forward audio data.

[0030] In some embodiments, the current audio data is determined based on the duration of the audio frame corresponding to the encoding network.

[0031] In some embodiments, the duration of the target structural delay is greater than the duration of the audio frame.

[0032] In some embodiments, the target structural delay is input by an object; or, the target structural delay is obtained by the object selecting the first candidate structural delay from N candidate structural delays displayed, and the duration of each candidate structural delay is greater than the duration of the audio frame, where N is a positive integer.

[0033] Fourthly, the present application provides an audio encoding device, including:

[0034] An acquisition unit, configured to determine current audio data to be encoded and a target structural delay;

[0035] A delay determination unit, configured to determine the structural delay corresponding to the current audio data, and determine the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data;

[0036] An input determination unit, configured to determine the input data of the encoding network based on the current audio data and the forward audio data, where the input data includes the current audio data and the forward audio data;

[0037] An encoding unit, configured to encode the input data through an encoding network to obtain a first audio bitstream.

[0038] In some embodiments, the delay determination unit is configured to obtain the sampling rate corresponding to the encoding network; and determine the structural delay corresponding to the current audio data based on the sampling rate and the number of samples included in the current audio data.

[0039] In some embodiments, the delay determination unit is configured to determine the ratio of the number of samples included in the current audio data to the sampling rate as the structural delay corresponding to the current audio data.

[0040] In some embodiments, the delay determination unit is configured to determine the structural delay corresponding to the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data; and determine the forward audio data based on the structural delay corresponding to the forward audio data.

[0041] In some embodiments, a delay determination unit is configured to determine the structural delay corresponding to the forward audio data as the difference between the structural delay of the target structure and the structural delay corresponding to the current audio data.

[0042] In some embodiments, a delay determination unit is configured to determine the number of samples included in the forward audio data based on the structural delay corresponding to the forward audio data and the sampling rate; and determine the forward audio data based on the number of samples included in the forward audio data.

[0043] In some embodiments, a delay determination unit is configured to determine the number of samples included in the forward audio data as the product of the structural delay corresponding to the forward audio data and the sampling rate; or determine the product of the structural delay corresponding to the forward audio data and the sampling rate as the first number of samples, and determine the number of samples included in the forward audio data based on the first number of samples and the total number of samples of the input data of the encoding network.

[0044] In some embodiments, a delay determination unit is configured to, when the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, determine the first number of samples as the number of samples included in the forward audio data; when the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, determine the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data as the number of samples included in the forward audio data.

[0045] In some embodiments, an input determination unit is configured to, when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, determine the current audio data and the forward audio data as the input data of the encoding network; when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, obtain the backward audio data of the current audio data, and determine the current audio data, the backward audio data and the forward audio data as the input data of the encoding network.

[0046] In some embodiments, an input determination unit is configured to determine the number of samples included in the backward audio data based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data; and obtain the backward audio data based on the number of samples included in the backward audio data.

[0047] In some embodiments, an input determination unit is configured to subtract the number of samples included in the current audio data and the number of samples included in the forward audio data from the total number of samples of the input data of the encoding network, to obtain the number of samples included in the backward audio data.

[0048] In some embodiments, an acquisition unit is specifically configured to acquire an audio frame corresponding to the encoding network; and determine the current audio data based on the duration of the audio frame.

[0049] In some embodiments, the duration of the target structural delay is greater than the duration of the audio frame.

[0050] In some embodiments, an acquisition unit is specifically configured to acquire the target structural delay input by an object; or display N candidate structural delays corresponding to the encoding network, and in response to a selection operation of the object on a first candidate structural delay among the N candidate structural delays, determine the first candidate structural delay as the target structural delay, where the duration of each candidate structural delay is greater than the duration of the audio frame, and N is a positive integer.

[0051] In a fifth aspect, an electronic device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.

[0052] In a sixth aspect, a chip is provided for implementing the method in any one of the first aspect to the second aspect or its various implementation manners above. Specifically, the chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method in any one of the first aspect to the second aspect or its various implementation manners above.

[0053] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program enables a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.

[0054] In an eighth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions enable a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.

[0055] In a ninth aspect, a computer program is provided, which when running on a computer, enables the computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners above.

[0056] In summary, in the present application, the encoding end determines the current audio data to be encoded and the target structural delay, and determines the structural delay corresponding to the current audio data. Then, based on the target structural delay and the structural delay corresponding to the current audio data, the forward audio data is determined. Next, the encoding end determines the input audio data of the encoding network based on the current audio data and the forward audio data. The input audio data includes the current audio data and the forward audio data. Finally, the input audio data is encoded by the encoding network to obtain the first audio bitstream. That is to say, in the embodiments of the present application, in order to ensure the encoding and decoding effect, the input audio data of the encoding network includes the current audio data and the forward audio data of the current audio data. In this way, compared with the zero-padding scheme, the encoding and decoding quality can be improved. At the same time, in the embodiments of the present application, the forward audio data is determined based on the structural delay corresponding to the current audio data and the target structural delay. In this way, the control of the structural delay can be realized, and excessive structural delay can be avoided. Furthermore, while ensuring the audio encoding quality, the control of the structural delay is realized, thereby improving the encoding performance of the audio data. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0058] Figure 1 It is a schematic block diagram of an audio encoding and decoding system according to an embodiment of the present application;

[0059] Figure 2 It is a schematic diagram of an end-to-end audio encoding and decoding system based on deep learning according to an embodiment of the present application;

[0060] Figure 3 It is a network structure block diagram of an encoder-decoder constructed based on a convolutional neural network in an embodiment of the present application;

[0061] Figure 4 It is a schematic diagram of input data supplementation according to an embodiment of the present application;

[0062] Figure 5 It is a schematic flowchart of an audio encoding method provided by an embodiment of the present application;

[0063] Figure 6A It is a schematic diagram of one type of input data of the encoding network;

[0064] Figure 6B It is a schematic diagram of another type of input data of the encoding network;

[0065] Figure 6CAnother schematic diagram of the input data for the encoding network;

[0066] Figure 7 A related encoding schematic diagram for the network using causal convolution;

[0067] Figure 8A A schematic diagram of encoding;

[0068] Figure 8B Another schematic diagram of encoding;

[0069] Figure 9 Schematic flowchart of the audio decoding method provided by an embodiment of the present application;

[0070] Figure 10 Schematic block diagram of the audio encoding device provided by an embodiment of the present application;

[0071] Figure 11 Schematic block diagram of the audio decoding device provided by an embodiment of the present application;

[0072] Figure 12 Schematic block diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0073] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0074] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.

[0075] The technical solution proposed in this application can be applied to technical fields such as artificial intelligence and audio coding and decoding, and is used to control the structural delay of audio coding and decoding while ensuring the performance of audio and video coding and decoding, thereby improving the audio coding and decoding effect.

[0076] The following introduces the relevant concepts involved in the embodiments of this application.

[0077] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0078] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0079] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behavior to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0080] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0081] This embodiment of the application mainly introduces the application of artificial intelligence technology in audio coding and decoding technology.

[0082] Audio coding and decoding: The audio coding process compresses audio into smaller data, and the decoding process restores the smaller data to audio. The smaller encoded data is used for network transmission and occupies less bandwidth.

[0083] Audio sampling rate: The audio sampling rate describes the number of data contained within a unit of time (1 second). For example, an 8k sampling rate contains 8000 sampling points, and each sampling point corresponds to a short integer.

[0084] Codebook: A collection of multiple vectors, and the same codebook is stored on both the encoder and decoder sides.

[0085] Quantization: Find the vector in the codebook that is closest to the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.

[0086] Quantizer: The quantizer is responsible for the quantization work and is responsible for updating the vectors within the codebook.

[0087] Audio frame: Represents the minimum voice duration for a single transmission in the network.

[0088] Short-Time Fourier Transform: STFT. Divide a long-time signal into several shorter equal-length signals, and then calculate the Fourier transform of each shorter segment separately. It is usually used to depict the changes in the frequency domain and time domain and is an important tool in time-frequency analysis.

[0089] The audio coding and decoding method provided by this embodiment of the application can be applied to the fields of audio coding and decoding, hardware audio coding and decoding, dedicated circuit video coding and decoding, real-time audio coding and decoding, etc. For example, the solution of this application can be combined with the audio video coding standard (AVS for short), such as the H.264 / audio video coding (AVC for short) standard. Or, the solution of this application can be combined with other proprietary or industry standards. It should be understood that the technology of this application is not limited to any specific coding and decoding standard or technology.

[0090] The audio coding and decoding method provided by this embodiment of the application can be applied to any end-to-end audio coding and decoding scheme based on deep learning.

[0091] For ease of understanding, first combine Figure 1 to introduce the audio coding and decoding system involved in this embodiment of the application.

[0092] Figure 1Schematic block diagram of an audio encoding and decoding system according to an embodiment of the present application. It should be noted that Figure 1 This is only an example. The audio encoding and decoding system of the embodiments of the present application includes but is not limited to Figure 1 as shown. As Figure 1 shown, the audio encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) audio data to generate a bitstream, and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded audio data.

[0093] The encoding device 110 of the embodiments of the present application can be understood as a device with audio encoding function, and the decoding device 120 can be understood as a device with audio decoding function. That is, the embodiments of the present application include a wider range of devices for the encoding device 110 and the decoding device 120, such as smart phones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, televisions, cameras, playback devices, digital media players, audio game consoles, in-vehicle computers, etc.

[0094] In some embodiments, the encoding device 110 can transmit the encoded audio data (such as a bitstream) to the decoding device 120 via a channel 130. The channel 130 can include one or more media and / or devices capable of transmitting the encoded audio data from the encoding device 110 to the decoding device 120.

[0095] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to directly transmit the encoded audio data to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded audio data according to a communication standard and transmit the modulated audio data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media can also include wired communication media, such as one or more physical transmission lines.

[0096] In another example, the channel 130 includes a storage medium that can store the audio data encoded by the encoding device 110. The storage medium includes a variety of locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded audio data from the storage medium.

[0097] In another example, the channel 130 may include a storage server that can store the audio data encoded by the encoding device 110. In this example, the decoding device 120 may download the stored encoded audio data from the storage server. Optionally, the storage server can store the encoded audio data and transmit the encoded audio data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.

[0098] In some embodiments, the encoding device 110 includes an audio encoder 112 and an output interface 113. Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0099] In some embodiments, in addition to the audio encoder 112 and the input interface 113, the encoding device 110 may further include an audio source 111.

[0100] The audio source 111 may include at least one of an audio acquisition device (e.g., a microphone), an audio archive, an audio input interface, and a computer voice system. Among them, the audio input interface is used to receive audio data from an audio content provider, and the computer voice system is used to generate audio data.

[0101] The audio encoder 112 encodes the audio data from the audio source 111 to generate a bitstream. The bitstream contains the encoded information of the audio data in the form of a bitstream. The encoded information may include the encoded audio data and associated data. The associated data may include quantization parameters and other syntax structures, etc. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0102] The audio encoder 112 directly transmits the encoded audio data to the decoding device 120 via the output interface 113. The encoded audio data may also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.

[0103] In some embodiments, the decoding device 120 includes an input interface 121 and an audio decoder 122.

[0104] In some embodiments, in addition to the input interface 121 and the audio decoder 122, the decoding device 120 may further include a playback device 123.

[0105] Among them, the input interface 121 includes a receiver and / or a modem. The input interface 121 may receive the encoded audio data through the channel 130.

[0106] The audio decoder 122 is used to decode the encoded audio data to obtain the decoded audio data and transmit the decoded audio data to the playback device 123.

[0107] The playback device 123 plays the decoded audio data. The playback device 123 can be integrated with the decoding device 120 or be external to the decoding device 120. The playback device 123 can include various playback devices.

[0108] In addition, Figure 1 merely for example, the technical solution of the embodiments of the present application is not limited to Figure 1 , for example, the technology of the present application can also be applied to unilateral audio encoding or unilateral audio decoding.

[0109] Figure 2 It is a schematic diagram of an end-to-end audio codec system based on deep learning involved in the embodiments of the present application. As Figure 2 shown, the audio codec system of the embodiments of the present application includes: an encoding network 210, a quantization module 211, an inverse quantization module 212, and a decoding network 213.

[0110] During encoding, the encoding end (also referred to as the sending end) will first input the input audio data into the encoding network 210 for non-linear transformation to obtain the encoding vector of the input audio data (also referred to as the embedding sequence or latent variable, etc.). Then, the quantization module 211 quantizes the encoding vector of the audio data to obtain the quantization result of the encoding vector. For example, a residual-based vector quantizer is used to select the corresponding quantization parameters according to the target bit rate. Finally, the quantized encoding vector is encoded and converted into a binary bitstream.

[0111] During decoding, the decoding end (also referred to as the receiving end) first recovers the quantization result of the encoding vector from the bitstream, and then further recovers the encoding vector through the inverse quantization module 212 and inputs it into the decoding network 213 for non-linear transformation to obtain the reconstructed audio data.

[0112] Figure 3 It is a network structure block diagram of an encoder-decoder constructed based on a convolutional neural network in an embodiment of the present application.

[0113] As Figure 3 shown, the network structure of the encoder-decoder includes an encoding network 310 and a decoding network 320, where the encoding network 310 can be implemented as software such as Figure 1 shown in the audio-video encoding device 110, and the decoding network 320 can be implemented as software such as Figure 1 shown in the audio-video decoding device 120. In some embodiments, the encoding network 310 is also referred to as the encoder 310, and the decoding network 320 is also referred to as the decoder 320.

[0114] At the data sending end, the audio data can be encoded and compressed through the encoding network 310. In an embodiment of the present application, the encoding network 310 may include an input layer 311, one or more encoding modules 312, and an output layer 313.

[0115] Exemplarily, the input layer 311 and the output layer 313 may be convolutional layers constructed based on one-dimensional convolutional kernels. Between the input layer 311 and the output layer 313, a plurality of (e.g., 4) encoding modules (EncoderBlock) 312 are connected in sequence. Each encoding module 312 includes a plurality of residual (ResidualUnit) modules, and each residual module contains a plurality of convolutional layers.

[0116] For example, in the input stage of the encoder, the original audio data to be encoded is sampled, and a vector with a channel number of c and a dimension of w can be obtained; this vector is input into the input layer 311, and after convolutional processing, a feature vector with a channel number of 32c and a dimension of w can be obtained. In some alternative embodiments, to improve the encoding efficiency, the encoding network 310 can encode a batch of audio vectors simultaneously.

[0117] In the downsampling stage of the encoder, the first encoding module reduces the vector dimension to 1 / 2 and doubles the channel number, obtaining a feature vector with a channel number of 64c and a dimension of 1 / 2w; the second encoding module reduces the vector dimension to 1 / 4 and doubles the channel number, obtaining a feature vector with a channel number of 128c and a dimension of 1 / 8w; the third encoding module reduces the vector dimension to 1 / 5 and doubles the channel number, obtaining a feature vector with a channel number of 256c and a dimension of 1 / 40w; the fourth encoding module reduces the vector dimension to 1 / 8 and doubles the channel number, obtaining a feature vector with a channel number of 512c and a dimension of 1 / 320w.

[0118] In the output stage of the encoder, the output layer 313 performs convolutional processing on the feature vector with a channel number of 512c and a dimension of 1 / 320w to obtain an encoded vector with a channel number of 1 and a dimension of K.

[0119] The encoded vector is input into the quantizer 330, the codebook index corresponding to the encoded vector can be queried in the codebook, and the codebook index is encoded to obtain a binary code stream, and then the binary code stream is sent to the data receiving end.

[0120] The data receiving end decodes the received binary code stream to obtain the codebook index, and performs inverse quantization based on the codebook index to obtain the reconstructed encoded vector. Finally, the reconstructed encoded vector is decoded by the decoding network 320 to obtain the restored audio data.

[0121] In one embodiment of the present application, the decoding network 320 may include an input layer 321, one or more decoding modules 322, and an output layer 323. Each decoding module 322 includes a plurality of ResidualUnit modules, and each residual module includes a plurality of convolutional layers.

[0122] After the data receiving end decodes the bitstream to obtain the codebook index, it can first query the codebook vector corresponding to the codebook index in the codebook through the quantizer 320, and then obtain the encoded vector for audio data reconstruction based on the codebook vector. For example, the reconstructed encoded vector can be a vector with 1 channel and dimension K. In some alternative embodiments, to improve the decoding efficiency, the data receiving end can decode a batch of codebook vectors simultaneously.

[0123] In the input stage of the decoder, the reconstructed encoded vector is input to the input layer 321, and after convolutional processing, a feature vector with 512c channels and a dimension of 1 / 320w can be obtained.

[0124] In the decoding stage of the decoder, the first decoding module increases the vector dimension by 8 times and reduces the number of channels by 2 times, obtaining a feature vector with 256c channels and a dimension of 1 / 40w; the second decoding module increases the vector dimension by 5 times and reduces the number of channels by 2 times, obtaining a feature vector with 128c channels and a dimension of 1 / 8w; the third decoding module increases the vector dimension by 4 times and reduces the number of channels by 2 times, obtaining a feature vector with 64c channels and a dimension of 1 / 2w; the fourth decoding module increases the vector dimension by 2 times and reduces the number of channels by 2 times, obtaining a feature vector with 32 channels and a dimension of w.

[0125] In the output stage of the decoder, after the output layer 323 performs convolutional processing on the feature vector with 32 channels and a dimension of w, the reconstructed audio data with 1 channel and a dimension of w is restored.

[0126] As can be seen from the above, in the end-to-end audio (or other one-dimensional signal) encoding and decoding scheme based on deep learning, the encoder maps the audio signal to a latent variable through the encoding network, and further generates a corresponding binary bitstream file through quantization; the decoding end obtains the transmitted information by reading the binary bitstream file, and obtains the corresponding latent variable through the inverse quantization technology. Then, the latent variable is used as the input of the decoding network, and the final reconstructed audio signal is decoded.

[0127] Traditional codecs can only process input data of a specific length, such as one frame with a length of 20 ms. If the length of the input data is less than 20 ms, then zero-padding operations need to be performed on the input data until the data length reaches the preset frame length. On the contrary, if the length of the input data is greater than 20 ms, then the input data will be sliced so that the length of each slice except the last one is the preset frame length, and zero-padding operations will be performed on the last frame as needed. After that, each frame of data will be used as the basic coding unit for encoding and decoding operations.

[0128] Deep learning-based audio codecs support variable-length input data. At the same time, there is also a minimum effective input length to ensure that the coded bitstream is not affected by additional zero-padding operations. Generally speaking, this minimum effective input length is jointly determined by a specific encoding network and decoding network.

[0129] For the real-time audio transmission scenario of neural network-based codecs, in one example, as Figure 4 shown, the network uses conventional convolution, and the data to be encoded is placed in the middle of the input, and zero-padding is used in a symmetric manner at both ends to reach the minimum effective length of the network. However, when zero-padding is used, coding noise will be introduced, which will reduce the encoding and decoding effect of the audio data. To improve the audio coding effect, in some embodiments of the present application, the coding noise can be reduced by supplementing the uncoded data. However, to meet the minimum effective input length requirement, the length of the forward data to be encoded used is a fixed value, and using more forward data to be encoded will result in a longer structural delay, which will further lead to an uncontrollable structural delay and poor audio coding and decoding performance.

[0130] To solve the above technical problems, an audio encoding and decoding method with controllable structural delay is proposed in an embodiment of the present application. Specifically, the encoding end determines the current audio data to be encoded and the target structural delay, and determines the structural delay corresponding to the current audio data. Then, based on the target structural delay and the structural delay corresponding to the current audio data, the forward audio data is determined. Next, the encoding end determines the input audio data of the encoding network based on the current audio data and the forward audio data. The input audio data includes the current audio data and the forward audio data. Finally, the input audio data is encoded by the encoding network to obtain the first audio code stream. Correspondingly, the decoding end obtains the first audio code stream, and determines the input code stream of the decoding network based on the first audio code stream. Then, the input code stream is decoded by the decoding network to obtain the reconstructed data of the current audio data. That is to say, in order to ensure the encoding and decoding effect in the embodiment of the present application, the input audio data of the encoding network includes the current audio data and the forward audio data of the current audio data. In this way, compared with the zero-padding scheme, the encoding and decoding quality can be improved. At the same time, in the embodiment of the present application, the forward audio data is determined based on the structural delay corresponding to the current audio data and the target structural delay. In this way, the control of the structural delay can be realized, and excessive structural delay can be avoided. Furthermore, while ensuring the audio encoding and decoding quality, the control of the structural delay is realized, thereby improving the encoding and decoding performance of the audio data.

[0131] The technical solutions of the embodiments of the present application are described in detail below through some embodiments. These embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0132] Since the present application mainly relates to the improvement of the encoding end, the audio encoding method provided in the embodiments of the present application is introduced first with the encoding end as an example.

[0133] Figure 5 It is a schematic flowchart of the audio encoding method provided in an embodiment of the present application. The execution subject of the embodiment of the present application is a device with audio encoding function, such as an audio encoding device. In some embodiments, the audio encoding device may be Figure 1 the encoding device in. For the convenience of description, the embodiment of the present application is described by taking the execution subject as the encoding device as an example.

[0134] As Figure 5 shown, the audio encoding method of the embodiment of the present application includes:

[0135] S101. Determine the current audio data to be encoded and the target structural delay.

[0136] In the embodiments of the present application, the encoding network has a requirement for the total number of samples of the input data, which can be understood as the minimum total number of samples of the input data of the encoding network. When encoding the current audio data, if the number of samples included in the current audio data does not meet the requirement for the total number of samples of the input data of the encoding network, then it is necessary to supplement the samples of the current audio data so that the number of samples included in the supplemented current audio data is not less than the total number of samples of the input data of the encoding network.

[0137] It should be noted that the total number of samples of the input data of different encoding networks may be different. In the embodiments of the present application, after the encoding network is selected, the total number of samples of the input data of the encoding network can also be obtained. Among them, the total number of samples of the input data of the encoding network is determined based on information such as the number of network layers of the encoding network, the convolution operations between the layers, and the size of the output features of the encoding network. That is to say, the total number of samples of the input data of the encoding network can ensure that after being processed by the encoding network, the size of the output feature vector meets the preset output size.

[0138] In some related technologies, by padding zeros at one or both ends of the current audio data, the requirement for the total number of samples of the input data of the encoding network is met. However, zero-padding will introduce input noise, reduce the accuracy of the input data of the encoding network, and further make the feature information of the current audio data extracted by the encoding network inaccurate. When performing subsequent encoding processing based on this inaccurate feature information, the encoding quality of the current audio data will be reduced.

[0139] To solve this technical problem, in the embodiments of the present application, when encoding the current audio data, if the number of samples included in the current audio data does not meet the requirement for the total number of samples of the input data of the encoding network, then real unencoded audio data is supplemented to complete the input of the encoding network. This can reduce input noise and thus improve the encoding quality of the current audio data.

[0140] Furthermore, since the unencoded audio data needs to be loaded in a real-time scenario due to not being encoded, it will introduce structural latency. In the embodiments of the present application, in order to control the structural latency, when encoding the current audio data, the target structural latency is also obtained. The target structural latency can be understood as the maximum structural latency allowed when encoding the current audio data.

[0141] The following are at least one of the possible structural latencies that may exist when encoding and decoding the current audio data in the embodiments of the present application:

[0142] The buffering time of the current audio data;

[0143] The loading time of the forward unencoded data of the current audio data;

[0144] The running time required to encode the current audio data;

[0145] The network transmission duration;

[0146] The running time required to decode the current audio data,

[0147] After decoding the current audio data, the duration of additional buffered data required for filtering processing.

[0148] In some embodiments of the present application, the structural delay at the encoding end involved in the embodiments of the present application mainly includes: the buffering time of the current audio data and the loading time of the forward uncoded data of the current audio data.

[0149] In the embodiments of the present application, the encoding structural delay can be controlled by controlling the length of the current audio data (i.e., the number of samples included) and the length of the forward audio data (i.e., the number of samples included).

[0150] In the embodiments of the present application, the current audio data to be encoded can be a segment of audio data to be encoded. For example, in a real-time scenario, the current audio data to be encoded can be understood as a segment of real-time audio data generated within a current period of time (e.g., 20 ms). For another example, in a non-real-time scenario, the current audio data to be encoded can be a segment of audio data to be encoded in an audio file to be encoded.

[0151] The embodiments of the present application do not limit the specific size of the current audio data. For example, the current audio data to be encoded can include one or more audio frames. Optionally, the current audio data can also include incomplete audio frames, such as including 1 / 2 of an audio frame or 1 / 4 of an audio frame, etc.

[0152] In the embodiments of the present application, an audio frame can be understood as a data segment with a specified time length obtained after frame division processing and windowing processing of the original audio data.

[0153] The embodiments of the present application do not limit the specific acquisition method of the original audio data.

[0154] In some examples, the original audio can be the voice collected by the terminal.

[0155] In some examples, the original audio data can be the sound signal collected in a network voice call or video call scenario.

[0156] In some examples, the original audio data can be the sound signal collected in a live broadcast scenario, or the sound signal collected in an online karaoke scenario, or the sound signal collected in a voice broadcast scenario.

[0157] In some examples, the original audio data may be audio data obtained from a storage resource. For example, the original audio data may be stored speech, music, video, etc.

[0158] In one possible implementation, when the embodiments of the present application perform audio frame partitioning on the original audio data, a preset duration can be set for partitioning. For example, the original audio data is partitioned into an audio frame every 10 ms or 20 ms.

[0159] In some embodiments of the present application, the size of the current audio data is a preset value. For example, the current audio data is audio data with a duration of 20 ms.

[0160] In some embodiments of the present application, the size of the current audio data can be specified by the user. For example, the user specifies that the duration of the current audio data is 30 ms.

[0161] In an audio encoding and decoding scheme based on a neural network, after the encoding network training is completed, the audio frames corresponding to the encoding network are also determined. That is to say, different encoding networks can correspond to different audio frames, and of course, they can also correspond to the same audio frame, which is specifically determined based on the network structure and related parameters of the encoding network.

[0162] In some embodiments of the present application, the size of the current audio data is related to the audio frames corresponding to the encoding network. At this time, the above S101 includes the following steps S101-A1 to S101-A2:

[0163] S101-A1. Obtain the audio frames corresponding to the encoding network;

[0164] S101-A2. Determine the current audio data based on the duration of the audio frames.

[0165] In this implementation, before encoding the audio data, the encoding end first selects an encoding network, and then the audio frames corresponding to the encoding network can be obtained. In one example, if the encoding network supports multiple audio frames, the multiple audio frames can be displayed to the user, and the user's selection operation on one of the multiple audio frames is received. Then, the audio frame selected by the user is determined as the audio frame corresponding to the encoding network. Then, based on this audio frame, the current audio data is determined.

[0166] In one example, the current audio data includes one audio frame, and the duration of one audio frame is 20 ms. At this time, the audio data within 20 ms to be encoded currently can be determined as the current audio data.

[0167] In one example, when the current audio data includes multiple audio frames, assuming there are 2 audio frames, the duration of one audio frame is 20 ms. At this time, the audio data within 40 ms to be currently encoded can be determined as the current audio data.

[0168] In this implementation manner, when determining the current audio data to be encoded based on the duration of the audio frames corresponding to the encoding network, then after the encoding network is selected, the size of the current audio data is also determined.

[0169] In the embodiments of the present application, in order to implement the control of the structural delay, the encoding end also determines a target structural delay, which can be understood as the maximum structural delay allowed when encoding the current audio data.

[0170] It should be noted that in the embodiments of the present application, if the size of the current audio data is an audio frame corresponding to the encoding network, then the duration of the above-mentioned target structural delay is greater than the duration of this audio frame. This is because the encoding structural delay of the present application mainly includes the buffering time of the current audio data and the loading time of the forward unencoded data of the current audio data. The target structural delay is the sum of the buffering time of the current audio data and the loading time of the forward unencoded data of the current audio data. If the current audio data includes one audio frame, then the buffering time of the current audio data is equal to the duration of this audio frame. At this time, it can be deduced that the target structural delay is greater than the duration of this audio frame. For example, when the duration of the audio frame corresponding to the encoding network is 20 ms, then the duration of this target structural delay is greater than 20 ms, for example, the duration of the target structural delay is 30 ms or 40 ms.

[0171] The embodiments of the present application do not limit the specific manner in which the encoding end determines the target structural delay.

[0172] In one possible implementation manner, the target structural delay is a preset value corresponding to different application scenarios. For example, if the embodiments of the present application are applied to real-time scenario 1, and the structural delay required by this real-time scenario 1 is t1 ms, therefore, this t1 ms is determined as the target structural delay. For another example, if the embodiments of the present application are applied to real-time scenario 2, and the structural delay required by this real-time scenario 2 is t2 ms, therefore, this t2 ms is determined as the target structural delay.

[0173] In one possible implementation manner, the target structural delay is input by an object. That is to say, when the encoding end encodes a certain audio data, the object can indicate the target structural delay. In this way, when the encoding end encodes the current audio data, it can obtain this target structural delay and perform encoding based on this target structural delay.

[0174] In a possible implementation, the target structural delay is selected by the object from N candidate structural delays, where N is a positive integer. That is, the encoding end displays the N candidate structural delays corresponding to the encoding network, and the object selects the first candidate structural delay among the displayed N candidate structural delays. In response to the object's selection operation of the first candidate structural delay among the N candidate structural delays, the encoding end determines the first candidate structural delay as the target structural delay. It should be noted that the duration of each of the N candidate structural delays is greater than the duration of the audio frame corresponding to the encoding network, so that the object can select the required target structural delay from these N candidate structural delays, and it can be ensured that the selected target structural delay is less than the duration of the audio frame, thereby ensuring the accuracy of the determination of the target structural delay.

[0175] In the embodiment of the present application, after the encoding end determines the current audio data to be encoded and the target structural delay, it performs the steps of S102 as follows.

[0176] S102: Determine the structural delay corresponding to the current audio data, and determine the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data.

[0177] In the embodiment of the present application, when the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, it is necessary to determine the forward uncoded data of the current audio data, and then use the forward uncoded data to fill in the current audio data. However, too long forward uncoded data will introduce a large structural delay. Therefore, in the embodiment of the present application, in order to control the structural delay, after the encoding end determines the current audio data to be encoded and the target structural delay based on the above steps, it determines the forward audio data corresponding to the current audio data based on the current audio data and the target structural delay, so as to avoid selecting too long forward audio data, resulting in a large structural delay.

[0178] In the embodiment of the present application, the structural delay of the encoding end mainly includes the structural delay corresponding to the current audio data and the structural delay corresponding to the forward audio data. After the encoding end determines the current audio data to be encoded based on the above steps, it can determine the structural delay corresponding to the current audio data, and then based on the target structural delay and the structural delay corresponding to the current audio data, it can determine the forward audio data.

[0179] Next, the specific process of determining the structural delay corresponding to the current audio data will be introduced.

[0180] The embodiment of the present application does not limit the specific manner in which the encoding end determines the structural delay corresponding to the current audio data.

[0181] In a possible implementation, if the current audio data includes an audio frame and the duration of an audio frame is 20 ms, then this 20 ms can be determined as the structural delay corresponding to the current audio data.

[0182] In a possible implementation, the encoding end can determine the structural delay corresponding to the current audio data through the following steps S102-A1 and S102-A2:

[0183] S102-A1: Obtain the sampling rate corresponding to the encoding network;

[0184] S102-A2: Based on the sampling rate and the number of samples included in the current audio data, determine the structural delay corresponding to the current audio data.

[0185] The sampling rate (Fs) refers to the number of samplings per second, usually measured in Hz (Hertz). For example, a 16 kHz signal means sampling 16,000 times per second, that is, the signal per second consists of 16,000 points.

[0186] The sampling duration (T) refers to the duration of the audio, usually expressed in seconds.

[0187] The relationship between the sampling duration and the sampling rate can be calculated by the following formula (1):

[0188] Sampling duration = Number of sampling points / Sampling rate (1)

[0189] Among them, the number of sampling points (N) refers to the total number of samples in the audio, which is equal to the sampling rate multiplied by the duration. For example, if the sampling rate of an audio file is 44.1 kHz (that is, 44,100 samples are collected per second) and the duration is 10 seconds, then the number of sampling points is 44,100 x 10 = 441,000. According to the above formula, the duration can be calculated as 441,000 / 44,100 = 10 seconds.

[0190] In the embodiments of the present application, different encoding networks can correspond to different sampling rates. After the encoding end selects an encoding network, the sampling rate corresponding to this encoding network is also determined. In this way, the encoding end can determine the structural delay corresponding to the current audio data based on the number of samples included in the current audio data and this sampling rate.

[0191] The embodiments of the present application do not limit the specific manner in which the encoding end determines the structural delay corresponding to the current audio data based on the number of samples included in the current audio data and this sampling rate.

[0192] In a possible implementation, the encoding end may determine the sampling duration corresponding to the current audio data based on the number of samples included in the current audio data and the sampling rate, and then determine the sampling duration as the structural delay corresponding to the current audio data. That is to say, the encoding end may determine the ratio of the number of samples included in the current audio data to the sampling rate as the structural delay corresponding to the current audio data.

[0193] In a possible implementation, the encoding end may determine the sampling duration corresponding to the current audio data based on the number of samples included in the current audio data and the sampling rate, and then correct the sampling duration, and determine the corrected duration as the structural delay corresponding to the current audio data. For example, when the caching speed (or loading speed) of the current audio data is greater than the sampling speed, a certain value is subtracted from or a certain coefficient is multiplied by the sampling duration corresponding to the current audio data to obtain the structural delay corresponding to the current audio data.

[0194] In the embodiment of the present application, after the encoding end determines the structural delay corresponding to the current audio data based on the above steps, based on the target structural delay and the structural delay corresponding to the current audio data, the forward uncoded data of the current audio data is determined. This can ensure that the sum of the structural delays corresponding to the forward uncoded data and the current audio data is not greater than the target structural delay, and thus effectively control the structural delay while ensuring the encoding quality.

[0195] The embodiment of the present application does not limit the specific manner in which the encoding end determines the forward uncoded data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data.

[0196] In some embodiments, if the current audio data is an audio frame with a duration of 20 ms, and assuming the target structural delay is 30 ms, then the 10-ms uncoded audio data forward of the current audio data may be determined as the forward audio data.

[0197] In some embodiments, the above S102-A2 includes the following steps of S102-A21 and S102-A22:

[0198] S102-A21: Determine the structural delay corresponding to the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data;

[0199] S102-A22: Determine the forward audio data based on the structural delay corresponding to the forward audio data.

[0200] In this implementation, when determining the forward audio data, the encoding end may first determine the structural delay corresponding to the forward uncoded data based on the target structural delay and the structural delay corresponding to the current audio data.

[0201] For example, the difference between the structural delay of the target structure and the structural delay corresponding to the current audio data is determined as the structural delay corresponding to the forward audio data.

[0202] For another example, the difference between the structural delay of the target structure and the structural delay corresponding to the current audio data is determined. If the difference between the structural delay of the target structure and the structural delay corresponding to the current audio data is too small, for example, less than a preset value, then the preset shortest structural delay can be determined as the structural delay corresponding to the forward uncoded data, which can prevent the forward uncoded data from being too small and affecting the coding effect.

[0203] After the encoding end determines the structural delay corresponding to the forward uncoded data based on the above steps, the forward audio data can be determined based on this structural delay.

[0204] The embodiments of the present application do not limit the specific manner in which the encoding end determines the forward audio data based on the structural delay corresponding to the forward audio data.

[0205] In some embodiments, if the duration of the structural delay corresponding to the forward audio data is the same as the duration of an audio frame corresponding to the encoding network, then an uncoded audio frame in the forward direction of the current audio data can be determined as the forward audio data.

[0206] In some embodiments, the above S102-A22 includes the following steps of S102-A221 and S102-A222:

[0207] S102-A221: Determine the number of samples included in the forward audio data based on the structural delay corresponding to the forward audio data and the sampling rate;

[0208] S102-A222: Determine the forward audio data based on the number of samples included in the forward audio data.

[0209] In this implementation manner, after the encoding end determines the structural delay corresponding to the forward audio data, the number of samples included in the forward audio data is determined based on the structural delay corresponding to the forward audio data and the sampling rate.

[0210] In the embodiments of the present application, the manner of determining the forward audio data based on the number of samples included in the forward audio data in the above S102-A222 includes at least the following several types:

[0211] Method 1: The product of the structural delay corresponding to the forward audio data and the sampling rate is determined as the number of samples included in the forward audio data.

[0212] In the first implementation method, after determining the number of samples included in the forward audio data, the encoding end directly determines the product of the structural delay corresponding to the forward audio data and the sampling rate as the number of samples included in the forward audio data.

[0213] Exemplarily, the encoding end determines the number of samples included in the forward audio data based on the following formula (2):

[0214] N3 = (T - N2 / Fs) * Fs (2)

[0215] Where N3 is the number of samples included in the forward audio data, N2 is the number of samples included in the current audio data, Fs is the sampling rate, T is the target structural delay, N2 / Fs is the structural delay corresponding to the current audio data, and T - N2 / Fs is the structural delay corresponding to the forward audio data.

[0216] In the second method, the encoding end determines the product of the structural delay corresponding to the forward audio data and the sampling rate to obtain the first number of samples, and determines the number of samples included in the forward audio data based on the first number of samples and the total number of samples of the input data of the encoding network.

[0217] In the second method, the encoding end first determines the product of the structural delay corresponding to the forward audio data and the sampling rate as the first number of samples, and then compares the first number of samples with the total number of samples of the input data of the encoding network to determine the number of samples included in the forward audio data.

[0218] For example, if the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, then the first number of samples is determined as the number of samples included in the forward audio data.

[0219] For another example, if the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data is determined as the number of samples included in the forward audio data.

[0220] That is to say, in the second method, by controlling the sum of the number of samples included in the current audio data and the number of samples included in the forward audio data to be not greater than the total number of samples of the input data of the encoding network, it is possible to prevent the forward audio data from being too large, thereby ensuring the encoding effect.

[0221] After the encoding end determines the forward audio data of the current audio based on the above steps, it executes the following step S103.

[0222] S103. Determine the input audio data of the encoding network based on the current audio data and the forward audio data.

[0223] In the embodiments of the present application, in order to control the structural delay, the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data. This can ensure that the structural delay during the encoding of the current audio data, that is, the sum of the structural delay corresponding to the current audio data and the structural delay corresponding to the forward audio data, does not exceed the target structural delay, thereby realizing the control of the structural delay in the audio encoding process and improving the encoding performance of the audio data.

[0224] As can be seen from the above, in the embodiments of the present application, in order to control the structural delay, the encoding end determines the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data. At this time, the sum of the number of samples included in the current audio data and the forward audio data may or may not be the same as the total number of samples of the input data of the encoding network. Based on this, the encoding end determines the input audio data of the encoding network based on the current audio data and the forward audio data, including at least the following situations:

[0225] Situation 1, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, then the current audio data and the forward audio data are determined as the input audio data of the encoding network.

[0226] Exemplarily, assume that the number of samples included in the forward audio data is N3, the number of samples included in the current audio data is N2, and the total number of samples of the input data of the encoding network is N. As Figure 6A shown, if the sum of the number of samples N2 included in the current audio data and the number of samples N3 included in the forward audio data is equal to the total number of samples N of the input data of the encoding network, that is, N2 + N3 = N, at this time, the current audio data and the forward audio data can be determined as the input audio data of the encoding network. That is to say, in this situation 1, the input audio data of the encoding network only includes the current audio data and the forward audio data.

[0227] Situation 2, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then the current audio data and the forward audio data are determined as the input audio data of the encoding network.

[0228] Exemplarily, assume that the number of samples included in the forward audio data is N3, the number of samples included in the current audio data is N2, and the total number of samples of the input data of the encoding network is N. As Figure 6BAs shown, if the sum of the number of samples N2 included in the current audio data and the number of samples N3 included in the forward audio data is greater than the total number of samples N of the input data of the encoding network, that is, N2 + N3 > N, at this time, the current audio data and the forward audio data can be determined as the input audio data of the encoding network. That is to say, in this case 2, the input audio data of the encoding network only includes the current audio data and the forward audio data.

[0229] Case 3, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then the backward audio data of the current audio data is obtained, and the current audio data, the backward audio data, and the forward audio data are determined as the input audio data of the encoding network.

[0230] In this case 3, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then the input of the encoding network needs to be continuously supplemented. In this embodiment, in order to ensure the encoding quality, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then the backward audio data of the current audio data is obtained to supplement the input.

[0231] The embodiments of the present application do not limit the specific manner of obtaining the backward audio data of the current audio data.

[0232] In a possible implementation manner, the backward audio data includes at least one audio frame. For example, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then at least one encoded audio frame backward of the current audio data is selected as the backward audio data of the current audio data.

[0233] In a possible implementation manner, the encoding end determines the number of samples included in the backward audio data based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data, and then obtains the backward audio data based on the number of samples included in the backward audio data.

[0234] For example, the encoding end determines the difference between the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data as the number of samples included in the backward audio data.

[0235] Exemplarily, the encoding end determines the number of samples included in the backward audio data through the following formula (3):

[0236] N1 = N - N2 - N3 (3)

[0237] Wherein, N1 is the number of samples included in the backward audio data, N is the total number of samples of the input data of the encoding network, N2 is the number of samples included in the current audio data, and N3 is the number of samples included in the forward audio data.

[0238] For another example, the encoding end determines the difference between the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data. If the difference is less than a preset value, the preset number of samples is determined as the number of samples included in the backward audio data, which can prevent the amount of data of the selected backward audio data from being too small.

[0239] In one example, assume that the number of samples included in the forward audio data is N3, the number of samples included in the current audio data is N2, the number of samples included in the backward audio data is N1, and the total number of samples of the input data of the encoding network is N. As Figure 6C shown, if the sum of the number of samples N2 included in the current audio data, the number of samples N3 included in the forward audio data, and the number of samples N1 included in the backward audio data is equal to the total number of samples N of the input data of the encoding network, that is, N1 + N2 + N3 = N, at this time, the backward audio data, the current audio data, and the forward audio data can be determined as the input audio data of the encoding network. That is to say, in this case 3, the input audio data of the encoding network includes not only the current audio data and the forward audio data, but also the backward audio data of the current audio data.

[0240] In this case 3, when encoding the current audio data, if the number of samples included in the current audio data is less than the number of samples included in the encoding network, the forward audio data of the current audio data can be determined based on the above-mentioned target structure delay and the current audio data, and the forward audio data can be used as the context information of the current audio data. At the same time, the encoding end determines the backward audio data of the current audio data based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data, and the forward audio data can be used as the context information of the current audio data. In this way, when encoding the current audio data, the context information of the current audio data is considered, thereby improving the encoding quality of the current audio data. Figure 7 A schematic diagram of a related encoding for the network to use causal convolution, as Figure 7 shown, the current audio data to be encoded is placed on the far right of the input, and the left side is filled with the encoded data to meet the requirements of the minimum effective length of the encoding network. However Figure 7 in the encoding method shown, only the backward audio data is used to provide the context during encoding, so that the encoding network can only extract a sub-optimal feature representation of the current audio data, thereby affecting the encoding quality of the audio signal. And Figure 6CIn the embodiment of the present application shown, when encoding the current audio data, the encoding end not only uses the backward audio data to provide the context for the current audio data, but also uses the forward audio data to provide the following context for the current audio data, so that the encoding network can extract the optimal feature representation of the current audio data, thereby improving the encoding quality of the audio signal.

[0241] S104. Encode the input audio data through an encoding network to obtain a first audio bitstream.

[0242] In the embodiment of the present application, after determining the input audio data of the encoding network based on the above steps, the encoding end inputs the input audio data into the encoding network for encoding to obtain the feature information corresponding to the input audio data, and then performs quantization and encoding processing on the feature information to obtain an audio bitstream. For the sake of convenience of description, this audio bitstream is denoted as the first audio bitstream.

[0243] In some embodiments of the present application, if the input audio data of the encoding network determined by the encoding end includes the current audio data and the forward audio data of the current audio data, for example, it is placed in the input buffer in the order of the current audio data and the forward audio data and encoded to obtain a first audio bitstream. For example Figure 8A As shown, the encoding network encodes the input audio data composed of the current audio data and the forward audio data to obtain the feature information of the current audio data and the feature information of the forward audio data. Then, the quantizer performs quantization processing on the feature information of the current audio data and the feature information of the forward audio data to obtain the codebook index corresponding to each feature information, and encodes the codebook index to obtain a binary bitstream.

[0244] In some embodiments of the present application, if the input audio data of the encoding network determined by the encoding end includes the current audio data, the forward audio data of the current audio data, and the backward audio data of the current audio data, for example, it is placed in the input buffer in the order of the backward audio data, the current audio data, and the forward audio data and encoded to obtain a first audio bitstream. For example Figure 8B As shown, the encoding network encodes the input audio data composed of the current audio data, the forward audio data, and the backward audio data to obtain the feature information of the current audio data, the feature information of the forward audio data, and the feature information of the backward audio data. Then, the quantizer performs quantization processing on the feature information of the current audio data, the feature information of the forward audio data, and the feature information of the backward audio data to obtain the codebook index corresponding to each feature information, and encodes the codebook index to obtain a binary bitstream. For example, it is placed in the input buffer in the order of the backward audio data, the current audio data, and the forward audio data and encoded to obtain a first audio bitstream.

[0245] The audio encoding method provided by the embodiments of the present application determines the current audio data to be encoded and the target structural delay, and determines the structural delay corresponding to the current audio data. Then, based on the target structural delay and the structural delay corresponding to the current audio data, the forward audio data is determined. Next, the encoding end determines the input audio data of the encoding network based on the current audio data and the forward audio data. The input audio data includes the current audio data and the forward audio data. Finally, the input audio data is encoded by the encoding network to obtain the first audio bitstream. That is to say, in order to ensure the encoding and decoding effect, the input audio data of the encoding network in the embodiments of the present application includes the current audio data and the forward audio data of the current audio data. In this way, compared with the zero-padding scheme, the encoding and decoding quality can be improved. At the same time, in the embodiments of the present application, the forward audio data is determined based on the structural delay corresponding to the current audio data and the target structural delay, so that the control of the structural delay can be realized, and excessive structural delay can be avoided. Furthermore, while ensuring the audio encoding quality, the control of the structural delay is realized, thereby improving the encoding performance of the audio data.

[0246] The above describes the audio encoding method involved in the embodiments of the present application. Below, taking the decoding end as an example, the audio decoding method provided by the embodiments of the present application will be introduced.

[0247] Figure 9 It is a schematic flowchart of the audio decoding method provided by an embodiment of the present application. The execution subject of the embodiments of the present application can be a device with specific audio decoding functions, such as an audio decoding device. In some embodiments, the audio decoding device can be Figure 1 the decoding device in. For the convenience of description, the embodiments of the present application will be described by taking the execution subject as the decoding device as an example.

[0248] As Figure 9 shown, the audio decoding method of the embodiments of the present application includes the following steps:

[0249] S201. Obtain the first audio bitstream.

[0250] Among them, the first audio bitstream is obtained by encoding the input audio data through an encoding network. The input audio data includes the current audio data and the forward audio data of the current audio data, and the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data.

[0251] For the specific generation process of the above first audio bitstream, reference can be made to the specific introduction in the above encoding embodiments, which will not be elaborated here.

[0252] In some embodiments, the current audio data is determined based on the duration of the audio frame corresponding to the encoding network.

[0253] In some embodiments, the duration of the target structure latency is greater than the duration of the audio frame.

[0254] In one example, the target structure latency is input by an object.

[0255] In one example, the target structure latency is obtained by an object selecting the first candidate structure latency from among N candidate structure latencies displayed, and the duration of each candidate structure latency is greater than the duration of the audio frame, where N is a positive integer.

[0256] In some embodiments, the structure latency corresponding to the current audio data is determined based on the sampling rate and the number of samples included in the current audio data.

[0257] In one example, the structure latency corresponding to the current audio data is the ratio of the number of samples included in the current audio data to the sampling rate.

[0258] Wherein, for the specific process of determining the structure latency corresponding to the current audio data, reference may be made to the relevant description of S102 above, and details will not be elaborated here.

[0259] In some embodiments, the forward audio data is determined based on the structure latency corresponding to the forward audio data, and the structure latency corresponding to the forward audio data is determined based on the target structure latency and the structure latency corresponding to the current audio data.

[0260] In one example, the structure latency corresponding to the forward audio data is the difference between the target structure latency and the structure latency corresponding to the current audio data.

[0261] In some embodiments, the forward audio data is determined based on the number of samples included in the forward audio data, and the number of samples included in the forward audio data is determined based on the structure latency corresponding to the forward audio data and the sampling rate.

[0262] Exemplarily, the number of samples included in the forward audio data is the product of the structure latency corresponding to the forward audio data and the sampling rate.

[0263] Exemplarily, the number of samples included in the forward audio data is determined based on the first number of samples and the total number of samples of the input data of the coding network, and the first number of samples is the product of the structure latency corresponding to the forward audio data and the sampling rate.

[0264] In one example, if the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the coding network, then the number of samples included in the forward audio data is the first number of samples.

[0265] In another example, when the sum of the number of samples in the first sample and the number of samples included in the current audio data is greater than the total number of samples in the input data of the encoding network, the number of samples included in the forward audio data is the difference between the total number of samples in the input data of the encoding network and the number of samples included in the current audio data.

[0266] The specific process of determining the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data can refer to the relevant description of S102 above, and will not be elaborated here.

[0267] In some embodiments, when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples in the input data of the encoding network, the input audio data of the encoding network includes the current audio data and the forward audio data.

[0268] In some embodiments, when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples in the input data of the encoding network, the input audio data of the encoding network includes the current audio data, the backward audio data of the current audio data, and the forward audio data.

[0269] Exemplarily, the backward audio data is obtained based on the number of samples included in the backward audio data, and the number of samples included in the backward audio data is determined based on the total number of samples in the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data.

[0270] For example, the number of samples included in the backward audio data is the total number of samples in the input data of the encoding network minus the number of samples included in the current audio data and the number of samples included in the forward audio data.

[0271] The specific process of determining the input audio data of the encoding network based on the current audio data and the forward audio data can refer to the relevant description of S103 above, and will not be elaborated here.

[0272] S202. Determine the input feature vector of the decoding network based on the first audio bitstream.

[0273] From the above Figure 3It can be seen that after the encoding end obtains the current audio data to be encoded, it determines the input data of the encoding network based on the current audio data and the forward audio data of the current audio data, and encodes the input data through the encoding network to obtain the feature vector corresponding to the input data. Then, the feature vector is quantized to obtain the quantization result, and the quantization result is entropy encoded to obtain the first audio bitstream. Then, the encoding end transmits the first audio bitstream to the decoding end. After obtaining the first audio bitstream, the decoding end first performs inverse quantization on the first audio bitstream to obtain the feature vector, and records the feature vector as the first feature vector. The first feature vector can be understood as the reconstructed value of the feature vector output by the encoding network. Then, the decoding end performs decoding processing on the first feature vector through the decoding network to obtain the reconstructed data of the current audio data.

[0274] The embodiments of the present application do not limit the specific manner in which the decoding end determines the input feature vector of the decoding network based on the first audio bitstream.

[0275] In some embodiments, the decoding end performs inverse quantization on the first audio bitstream to obtain the first feature vector, and further uses the first feature vector as the input feature vector of the decoding network.

[0276] In some embodiments, the above S202 includes the following steps of S202-A and S202-B:

[0277] S202-A: Perform inverse quantization on the first audio bitstream to obtain the first feature vector;

[0278] S202-B: Determine the input feature vector of the decoding network based on the scale size of the first feature vector and the scale size of the input feature vector of the decoding network

[0279] In the embodiments of the present application, the input data of the decoding network is a feature vector, and the decoding network has requirements for the scale of the input feature vector. Based on this, in this implementation manner, after the decoding end performs inverse quantization on the first audio bitstream to obtain the first feature vector, it is also necessary to obtain the scale size of the input feature vector of the decoding network, and then determine the input feature vector of the decoding network based on the scale size of the first feature vector and the scale size of the input feature vector of the decoding network.

[0280] The scale of the feature vector in the embodiments of the present application includes at least one of the size and the number of channels of the feature vector, where the size of the feature vector includes the width W and the height H of the feature vector.

[0281] In some embodiments, when the scale of the first feature vector is smaller than the scale of the input feature vector of the decoding network, the width and height of the first feature vector are supplemented so that the size of the supplemented first feature vector is the same as the size of the input feature vector of the decoding network. Then, the supplemented first feature vector is determined as the input feature vector of the decoding network. For example, when the size of the first feature vector is a 1×3 matrix and the size of the input feature vector of the decoding network is 1×5, 2 elements can be supplemented after the first feature vector, and these 2 supplemented elements can be obtained by replicating the last element or the last 2 elements of the first feature vector.

[0282] In some embodiments, the decoding end determines the scale difference between the scale size of the input feature vector of the decoding network and the scale size of the first feature vector. Then, based on this scale difference, a second feature vector is determined from the feature vectors of the backward audio data of the current audio data. In this way, the input feature vector of the decoding network can be obtained based on the first feature vector and the second feature vector.

[0283] In this implementation manner, when the scale of the first feature vector is smaller than the size of the input feature vector of the decoding network, a second feature vector is selected from the feature vectors of the decoded audio data to supplement the first feature vector, and the supplemented first feature vector is determined as the input feature vector of the decoding network.

[0284] For example, assume that the size of the first feature vector is 1×m1 and the size of the input feature vector of the decoding network is 1×m2. When m1 is smaller than m2, a second feature vector with a size of 1×(m2 - m1) is selected from the feature vectors of the backward audio data (i.e., the decoded audio data) of the current audio data. Then, the second feature vector with a size of 1×(m2 - m1) is supplemented to the first feature vector with a size of 1×m1 to obtain an input feature vector with a size of 1×m2.

[0285] In some examples, assume the number of channels c1 of the first feature vector and the number of channels c2 of the input feature vector of the decoding network. When c1 is smaller than c2, a second feature vector with c2 - c1 channels is selected from the feature vectors of the backward audio data (i.e., the decoded audio data) of the current audio data. Then, the second feature vector with c2 - c1 channels is concatenated with the first feature vector with c1 channels in channels to obtain an input feature vector with c2 channels.

[0286] After the decoding end determines the input feature vector of the decoding network based on the above steps, it executes the following step S203.

[0287] S203. Decode the input feature vector through the decoding network to obtain the reconstructed data of the current audio data.

[0288] In the embodiment of the present application, the input bitstream includes a first audio bitstream, which is obtained by encoding input audio data through an encoding network. The input audio data includes current audio data and forward audio data of the current audio data, and the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data. That is to say, in order to ensure the decoding effect in the embodiment of the present application, the input audio data of the encoding network includes the current audio data and the forward audio data of the current audio data. In this way, compared with the zero-padding scheme, the decoding quality of the audio data can be improved. At the same time, in the embodiment of the present application, the forward audio data is determined based on the structural delay corresponding to the current audio data and the target structural delay, so that the control of the structural delay can be realized, and excessive structural delay can be avoided. Furthermore, while ensuring the audio decoding quality, the control of the structural delay can be realized, thereby improving the decoding performance of the audio data.

[0289] In the embodiment of the present application, the input feature information of the decoding network includes not only the first feature information corresponding to the first audio bitstream, but also the second feature vector of the backward audio data. In this way, it can be ensured that the scale size of the feature vector input to the decoding network meets the scale size requirement of the input feature vector of the decoding network, thereby ensuring the decoding performance of the decoding network, realizing the accurate reconstruction of the current audio data, and further improving the decoding effect of the audio data.

[0290] As can be seen from the above, the first audio bitstream in the embodiment of the present application is encoded based on the current audio data and the forward audio data of the current audio data. In this way, when the decoding end decodes the first audio bitstream, it can obtain not only the reconstructed data of the current audio data, but also the reconstructed data of the forward audio data of the current audio data.

[0291] For the audio decoding method in the embodiment of the present application, the decoding end obtains a first audio bitstream, which is obtained by encoding input data through an encoding network. The input data includes the current audio data to be encoded and the forward audio data of the current audio data, and the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data; based on the first audio bitstream, the input feature vector of the decoding network is determined; the input feature vector is decoded through the decoding network to obtain the reconstructed data of the current audio data. That is to say, in order to ensure the decoding effect in the embodiment of the present application, the input audio data of the encoding network includes the current audio data and the forward audio data of the current audio data. In this way, compared with the zero-padding scheme, the decoding quality can be improved. At the same time, in the embodiment of the present application, the forward audio data is determined based on the structural delay corresponding to the current audio data and the target structural delay, so that the control of the structural delay can be realized, and excessive structural delay can be avoided. Furthermore, while ensuring the audio decoding quality, the control of the structural delay can be realized, thereby improving the decoding performance of the audio data.

[0292] In combination with the above Figures 5 to 9 , embodiments of the audio encoding and decoding method of the present application have been described in detail. In combination with Figures 10 to 11 , embodiments of the device of the present application will be described in detail.

[0293] Figure 10 FIG. is a schematic block diagram of an audio decoding device provided by an embodiment of the present application. The device 10 can be applied to a decoding device.

[0294] As Figure 10 shown, the audio decoding device 10 includes:

[0295] An obtaining unit 11, configured to obtain a first audio bitstream, where the first audio bitstream is obtained by encoding input data through an encoding network, and the input data includes current audio data to be encoded and forward audio data of the current audio data, and the forward audio data is determined based on a target structural delay and a structural delay corresponding to the current audio data;

[0296] An input determining unit 12, configured to determine an input feature vector of a decoding network based on the first audio bitstream;

[0297] A decoding unit 13, configured to decode the input feature vector through a decoding network to obtain reconstructed data of the current audio data.

[0298] In some embodiments, the input determining unit 12 is specifically configured to perform inverse quantization on the first audio bitstream to obtain a first feature vector; and determine the input feature vector of the decoding network based on a scale size of the first feature vector and a scale size of the input feature vector of the decoding network.

[0299] In some embodiments, the input determining unit 12 is specifically configured to determine a scale difference between a scale size of the input feature vector of the decoding network and a scale size of the first feature vector; determine a second feature vector from feature vectors of backward audio data of the current audio data based on the scale difference; and obtain the input feature vector of the decoding network based on the first feature vector and the second feature vector.

[0300] In some embodiments, the structural delay corresponding to the current audio data is determined based on the sampling rate and the number of samples included in the current audio data.

[0301] In some embodiments, the structural delay corresponding to the current audio data is a ratio of the number of samples included in the current audio data to the sampling rate.

[0302] In some embodiments, the forward audio data is determined based on the structural delay corresponding to the forward audio data, and the structural delay corresponding to the forward audio data is determined based on the target structural delay and the structural delay corresponding to the current audio data.

[0303] In some embodiments, the structural delay corresponding to the forward audio data is the difference between the target structural delay and the structural delay corresponding to the current audio data.

[0304] In some embodiments, the forward audio data is determined based on the number of samples included in the forward audio data, and the number of samples included in the forward audio data is determined based on the structural delay corresponding to the forward audio data and the sampling rate.

[0305] In some embodiments, the number of samples included in the forward audio data is the product of the structural delay corresponding to the forward audio data and the sampling rate; alternatively, the number of samples included in the forward audio data is determined based on a first number of samples and the total number of samples of the input data of the encoding network, and the first number of samples is the product of the structural delay corresponding to the forward audio data and the sampling rate.

[0306] In some embodiments, if the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is the first number of samples; if the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data.

[0307] In some embodiments, if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, then the input data of the encoding network includes the current audio data and the forward audio data; if the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then the input data of the encoding network includes the current audio data, the backward audio data of the current audio data, and the forward audio data.

[0308] In some embodiments, the backward audio data is obtained based on the number of samples included in the backward audio data, and the number of samples included in the backward audio data is determined based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data.

[0309] In some embodiments, the number of samples included in the backward audio data is the total number of samples of the input data of the encoding network minus the number of samples included in the current audio data and the number of samples included in the forward audio data.

[0310] In some embodiments, the current audio data is determined based on the duration of the audio frame corresponding to the encoding network.

[0311] In some embodiments, the duration of the target structural delay is greater than the duration of the audio frame.

[0312] In some embodiments, the target structural delay is input by an object; or, the target structural delay is obtained by the object selecting an operation on a first candidate structural delay among N candidate structural delays displayed, and the duration of each candidate structural delay is greater than the duration of the audio frame, where N is a positive integer.

[0313] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 10 The illustrated device can execute the embodiments of the above audio decoding method, and the foregoing and other operations and / or functions of each module in the device respectively implement the above method embodiments. For the sake of brevity, it will not be elaborated here.

[0314] Figure 11 It is a schematic block diagram of an audio encoding device provided by an embodiment of the present application. The device 20 can be applied to an encoding device.

[0315] As Figure 11 shown, the audio encoding device 20 includes:

[0316] An obtaining unit 21, configured to determine current audio data to be encoded and a target structural delay;

[0317] A delay determining unit 22, configured to determine the structural delay corresponding to the current audio data, and determine the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data;

[0318] An input determining unit 23, configured to determine the input data of the encoding network based on the current audio data and the forward audio data, where the input data includes the current audio data and the forward audio data;

[0319] An encoding unit 24, configured to encode the input data through an encoding network to obtain a first audio bitstream.

[0320] In some embodiments, the latency determination unit 22 is configured to obtain the sampling rate corresponding to the encoding network; and determine the structural latency corresponding to the current audio data based on the sampling rate and the number of samples included in the current audio data.

[0321] In some embodiments, the latency determination unit 22 is configured to determine the ratio of the number of samples included in the current audio data to the sampling rate as the structural latency corresponding to the current audio data.

[0322] In some embodiments, the latency determination unit 22 is configured to determine the structural latency corresponding to the forward audio data based on the target structural latency and the structural latency corresponding to the current audio data; and determine the forward audio data based on the structural latency corresponding to the forward audio data.

[0323] In some embodiments, the latency determination unit 22 is configured to determine the difference between the target structural latency and the structural latency corresponding to the current audio data as the structural latency corresponding to the forward audio data.

[0324] In some embodiments, the latency determination unit 22 is configured to determine the number of samples included in the forward audio data based on the structural latency corresponding to the forward audio data and the sampling rate; and determine the forward audio data based on the number of samples included in the forward audio data.

[0325] In some embodiments, the latency determination unit 22 is configured to determine the product of the structural latency corresponding to the forward audio data and the sampling rate as the number of samples included in the forward audio data; or determine the product of the structural latency corresponding to the forward audio data and the sampling rate as the first number of samples, and determine the number of samples included in the forward audio data based on the first number of samples and the total number of samples of the input data of the encoding network.

[0326] In some embodiments, the latency determination unit 22 is configured to, when the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, determine the first number of samples as the number of samples included in the forward audio data; and when the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, determine the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data as the number of samples included in the forward audio data.

[0327] In some embodiments, the input determination unit 23 is configured to determine the current audio data and the forward audio data as the input data of the encoding network when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network; when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, obtain the backward audio data of the current audio data, and determine the current audio data, the backward audio data, and the forward audio data as the input data of the encoding network.

[0328] In some embodiments, the input determination unit 23 is configured to determine the number of samples included in the backward audio data based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data; and obtain the backward audio data based on the number of samples included in the backward audio data.

[0329] In some embodiments, the input determination unit 23 is configured to subtract the number of samples included in the current audio data and the number of samples included in the forward audio data from the total number of samples of the input data of the encoding network to obtain the number of samples included in the backward audio data.

[0330] In some embodiments, the acquisition unit 21 is specifically configured to acquire the audio frame corresponding to the encoding network; and determine the current audio data based on the duration of the audio frame.

[0331] In some embodiments, the duration of the target structural delay is greater than the duration of the audio frame.

[0332] In some embodiments, the acquisition unit 21 is specifically configured to acquire the target structural delay input by the object; or, display N candidate structural delays corresponding to the encoding network, and in response to a selection operation of the object on a first candidate structural delay among the N candidate structural delays, determine the first candidate structural delay as the target structural delay, the duration of each candidate structural delay is greater than the duration of the audio frame, and N is a positive integer.

[0333] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 11 The illustrated apparatus can execute the embodiments of the above audio encoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, they are not elaborated here.

[0334] In the above, the apparatus according to the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules. Specifically, each step of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0335] Figure 12 is a schematic block diagram of an electronic device provided by an embodiment of the present application, Figure 12 The electronic device can be the above-mentioned encoding device or a decoding device.

[0336] Such as Figure 12 As shown, the electronic device 30 may include:

[0337] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.

[0338] For example, the processor 32 can be used to execute the steps in the above method 200 according to the instructions in the computer program 33.

[0339] In some embodiments of the present application, the processor 32 may include, but is not limited to:

[0340] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0341] In some embodiments of the present application, the memory 31 includes, but is not limited to:

[0342] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0343] In some embodiments of the present application, the computer program 33 can be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.

[0344] As Figure 12 shown, the electronic device 30 may further include:

[0345] A transceiver 34, which can be connected to the processor 32 or the memory 31.

[0346] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 can include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas can be one or more.

[0347] It should be understood that the various components in the computing device 30 are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0348] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon. When the computer program is executed by a computer, the computer is enabled to execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product containing instructions. When the instructions are executed by a computer, the computer is enabled to execute the method of the above method embodiment.

[0349] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method of the above method embodiment.

[0350] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0351] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0352] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.

[0353] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0354] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An audio decoding method, characterized in that, Including: Obtain a first audio bitstream, which is obtained by encoding input data through an encoding network. The input data includes current audio data to be encoded and forward audio data of the current audio data. The forward audio data is determined based on a target structural delay and a structural delay corresponding to the current audio data. Based on the first audio bitstream, determine an input feature vector of a decoding network. Decode the input feature vector through the decoding network to obtain reconstructed data of the current audio data.

2. The method according to claim 1, wherein The determining the input feature vector of the decoding network based on the first audio bitstream includes: Perform inverse quantization on the first audio bitstream to obtain a first feature vector. Based on a scale size of the first feature vector and a scale size of the input feature vector of the decoding network, determine the input feature vector of the decoding network.

3. The method according to claim 2, characterized in that The determining the input feature vector of the decoding network based on the scale size of the first feature vector and the scale size of the input feature vector of the decoding network includes: Determine a scale difference between a scale size of the input feature vector of the decoding network and the scale size of the first feature vector. Based on the scale difference, determine a second feature vector from feature vectors of backward audio data of the current audio data. Based on the first feature vector and the second feature vector, obtain the input feature vector of the decoding network.

4. The method according to claim 1, wherein The structural delay corresponding to the current audio data is a ratio of the number of samples included in the current audio data to a sampling rate. The forward audio data is determined based on a structural delay corresponding to the forward audio data. The structural delay corresponding to the forward audio data is a difference between the target structural delay and the structural delay corresponding to the current audio data.

5. The method according to claim 4, characterized in that, The forward audio data is determined based on the number of samples included in the forward audio data. The number of samples included in the forward audio data is determined based on the structural delay corresponding to the forward audio data and the sampling rate. The structural delay corresponding to the forward audio data is a difference between the target structural delay and the structural delay corresponding to the current audio data.

6. The method according to claim 5, wherein The number of samples included in the forward audio data is a product of the structural delay corresponding to the forward audio data and the sampling rate. Or, The number of samples included in the forward audio data is determined based on a first number of samples and a total number of samples of the input data of the encoding network. The first number of samples is a product of the structural delay corresponding to the forward audio data and the sampling rate. Wherein, if a sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is the first number of samples. If a sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then the number of samples included in the forward audio data is a difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data.

7. The method according to any one of claims 1-6, characterized in that when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, the input data of the encoding network includes the current audio data and the forward audio data; when the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, the input data of the encoding network includes the current audio data, the backward audio data of the current audio data, and the forward audio data.

8. The method according to claim 7, wherein The backward audio data is obtained based on the number of samples included in the backward audio data, and the number of samples included in the backward audio data is determined based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data.

9. The method according to any one of claims 1 to 6, characterized in that, The current audio data is determined based on the duration of the audio frame corresponding to the encoding network, and the duration of the target structural delay is greater than the duration of the audio frame.

10. The method according to claim 1, wherein The target structural delay is input by the object; or, the target structural delay is obtained by the object selecting the first candidate structural delay among N candidate structural delays displayed, and the duration of each candidate structural delay is greater than the duration of the audio frame, where N is a positive integer.

11. An audio encoding method, characterized in that, Comprising: determining the current audio data to be encoded and the target structural delay; determining the structural delay corresponding to the current audio data based on the sampling rate corresponding to the encoding network and the number of samples included in the current audio data, and determining the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data; determining the input data of the encoding network based on the current audio data and the forward audio data, where the input data includes the current audio data and the forward audio data; encoding the input data through the encoding network to obtain a first audio code stream.

12. The method according to claim 11, wherein The determining the structural delay corresponding to the current audio data based on the sampling rate and the number of samples included in the current audio data includes: determining the ratio of the number of samples included in the current audio data to the sampling rate as the structural delay corresponding to the current audio data; The determining the forward audio data of the current audio data based on the target structural delay and the structural delay corresponding to the current audio data includes: determining the structural delay corresponding to the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data; determining the forward audio data based on the structural delay corresponding to the forward audio data.

13. The method according to claim 12, wherein The determining the structural delay corresponding to the forward audio data based on the target structural delay and the structural delay corresponding to the current audio data includes: determining the difference between the target structural delay and the structural delay corresponding to the current audio data as the structural delay corresponding to the forward audio data; Determining the forward audio data based on the structural delay corresponding to the forward audio data includes: Determining the number of samples included in the forward audio data based on the structural delay corresponding to the forward audio data and the sampling rate; Determining the forward audio data based on the number of samples included in the forward audio data.

14. The method according to claim 13, wherein The determining the number of samples included in the forward audio data based on the structural delay corresponding to the forward audio data and the sampling rate includes: Determining the product of the structural delay corresponding to the forward audio data and the sampling rate as the number of samples included in the forward audio data; or Determining the product of the structural delay corresponding to the forward audio data and the sampling rate as the first number of samples; If the sum of the first number of samples and the number of samples included in the current audio data is less than or equal to the total number of samples of the input data of the encoding network, then determining the first number of samples as the number of samples included in the forward audio data; If the sum of the first number of samples and the number of samples included in the current audio data is greater than the total number of samples of the input data of the encoding network, then determining the difference between the total number of samples of the input data of the encoding network and the number of samples included in the current audio data as the number of samples included in the forward audio data.

15. The method according to any one of claims 12 - 14, characterized in that The determining the input data of the encoding network based on the current audio data and the forward audio data includes: If the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is equal to the total number of samples of the input data of the encoding network, then determining the current audio data and the forward audio data as the input data of the encoding network; If the sum of the number of samples included in the forward audio data and the number of samples included in the current audio data is less than the total number of samples of the input data of the encoding network, then obtaining the backward audio data of the current audio data, and determining the current audio data, the backward audio data and the forward audio data as the input data of the encoding network.

16. The method according to claim 15, wherein The obtaining the backward audio data of the current audio data includes: Determining the number of samples included in the backward audio data based on the total number of samples of the input data of the encoding network, the number of samples included in the current audio data, and the number of samples included in the forward audio data; Obtaining the backward audio data based on the number of samples included in the backward audio data.

17. The method according to any one of claims 12 - 14, characterized in that, The determining the current audio data includes: Obtaining the audio frame corresponding to the encoding network; Determining the current audio data based on the duration of the audio frame.

18. The method according to any one of claims 12 - 14, characterized in that, When the duration of the target structural delay is greater than the duration of the audio frame, the determining the target structural delay includes: Obtaining the target structural delay input by the object; or Displaying N candidate structural delays corresponding to the encoding network, and in response to a selection operation of the object on a first candidate structural delay among the N candidate structural delays, determining the first candidate structural delay as the target structural delay, and the duration of each candidate structural delay is greater than the duration of the audio frame, where N is a positive integer.

19. An electronic device, comprising a processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 10 or 11 to 18 above.

20. A computer-readable storage medium, characterized in that, For storing a computer program; The computer program causes a computer to execute the method according to any one of claims 1 to 10 or 11 to 18 above.