Audio encoding and decoding method, device and equipment
By dividing the audio data into multiple audio blocks and performing inverse transformation or transformation of signal strength on each audio block, the problem of signal strength fluctuations in audio encoding and decoding is solved, and the encoding and decoding effect is improved.
Patent Information
- Application Number
- CN202311550914.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, there are problems of signal strength fluctuations in the audio encoding and decoding process, resulting in unsatisfactory encoding and decoding effects.
By dividing the audio data to be encoded or decoded into a plurality of audio blocks, each audio block includes 2 or more audio frames, the signal strength of each audio block is inversely transformed or transformed, and then encoded or decoded.
It realizes smooth processing of signal strength, improves the encoding and decoding effect of audio data, and reduces the fluctuations of signal strength.
Smart Images

Figure CN120020946A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to an audio encoding and decoding method, apparatus, and device. Background Art
[0002] With the rapid development of deep learning technologies, deep learning technologies have been widely applied to signal processing technologies in different dimensions, such as audio, images, and videos. Taking audio signals as an example, in an end-to-end audio encoding and decoding scheme based on deep learning, an encoder maps an audio signal to a coded vector through an encoding network, and further generates a corresponding binary bitstream file through quantization technology. A decoding end obtains a quantization result by reading the binary bitstream file, and performs inverse quantization on the quantization result through inverse quantization technology to obtain a reconstructed coded vector. Then, the reconstructed coded vector is used as the input of a decoding network, and the final reconstructed audio signal is decoded.
[0003] In order to improve the audio encoding and decoding effect, before encoding audio data, the encoding end first performs signal strength transformation processing, such as loudness normalization processing. After the decoding network at the corresponding decoding end reconstructs the audio data, an inverse transformation is performed on the signal strength of the audio data, such as inverse normalization processing on the loudness of the reconstructed audio data. However, current signal strength transformation methods have the problem of signal strength fluctuations, which makes the audio encoding and decoding effect unsatisfactory. Summary of the Invention
[0004] The present application provides an audio encoding and decoding method, apparatus, device, and storage medium, which can make the transformation of signal strength smoother, thereby improving the audio encoding and decoding effect.
[0005] In a first aspect, the present application provides an audio decoding method, including:
[0006] Decoding a bitstream to obtain a currently reconstructed audio block by a decoding network, where the currently reconstructed audio block is one of M audio blocks obtained by partitioning target audio data, M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0007] Performing an inverse transformation on the signal strength of the currently reconstructed audio block to obtain a reconstructed value of the currently reconstructed audio block.
[0008] In a second aspect, the present application provides an audio encoding method, including:
[0009] Performing block partitioning on target audio data to be encoded to obtain M audio blocks, M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0010] For the current audio block among the M audio blocks, perform an inverse transform on the signal strength of the current audio block to obtain the transformed current audio block;
[0011] Encode the transformed current audio block to obtain a bitstream.
[0012] In a third aspect, the present application provides an audio decoding device, including:
[0013] A decoding unit, configured to decode the bitstream to obtain the current audio block reconstructed by the decoding network. The current audio block is one of the M audio blocks obtained by partitioning the target audio data. M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0014] A transformation unit, configured to perform an inverse transform on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block.
[0015] In some embodiments, the transformation unit is specifically configured to determine the inverse transform value of the signal strength of the current audio block; based on the inverse transform value, perform an inverse transform on the signal strength of the reconstructed current audio block to obtain the current audio block.
[0016] In some embodiments, the transformation unit specifically determines the inverse transform value of the signal strength of the current audio block based on the reconstructed current audio block.
[0017] In some embodiments, the signal strength of the current audio frame includes at least one of the amplitude, energy, and loudness of the current audio block.
[0018] In some embodiments, if the signal strength of the current audio frame includes the loudness of the current audio block, and the inverse transform value includes a loudness inverse gain value, then the transformation unit is specifically configured to decode the bitstream to obtain the first reference loudness value of the current audio block. The first reference loudness value is calculated based on the original audio data of the current audio block; determine the second reference loudness value of the current audio block based on the reconstructed current audio block; determine the loudness inverse gain value based on the first reference loudness value and the second reference loudness value of the current audio block.
[0019] In some embodiments, the transformation unit is specifically configured to, for each audio frame included in the current audio block, determine the second reference loudness value of the audio frame based on the reconstructed audio data after loudness normalization of the audio frame; determine the second reference loudness value of the current audio block based on the second reference loudness value of each audio frame.
[0020] In some embodiments, the transformation unit is specifically configured to subtract the first reference loudness value of the current audio block from the second reference loudness value of the current audio block to obtain a first difference value; and obtain the loudness anti-gain value based on the first difference value and a first preset value.
[0021] In some embodiments, the transformation unit is specifically configured to multiply the first difference value by the first preset value to obtain a first product; and perform a preset operation on the first product to obtain the loudness anti-gain value.
[0022] In some embodiments, the transformation unit is specifically configured to multiply each element in the reconstructed current audio block by the loudness anti-gain value to obtain the reconstructed value of the current audio block.
[0023] In some embodiments, the transformation unit is specifically configured to determine the length information of the current audio block; and decode the code stream corresponding to the current audio block based on the length information of the current audio block to obtain the current audio block after loudness normalization.
[0024] In some embodiments, the transformation unit is specifically configured to decode the code stream to obtain the length information of the current audio frame.
[0025] In some embodiments, the M audio blocks are obtained by performing block division on multiple audio frames included in the target audio data based on the audio block division length determined according to the length of the target audio data.
[0026] Fourthly, the present application provides an audio encoding device, including:
[0027] A block division unit, configured to perform block division on target audio data to be encoded to obtain M audio blocks, where M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0028] A transformation unit, configured to transform the signal strength of the current audio block among the M audio blocks to obtain a transformed current audio block;
[0029] An encoding unit, configured to encode the transformed current audio block to obtain a code stream.
[0030] In some embodiments, the transformation unit is specifically configured to determine a transformation value of the signal strength of the current audio block based on the original audio data of the current audio block; and transform the signal strength of the current audio block based on the transformation value to obtain the transformed current audio block.
[0031] In some embodiments, the signal strength of the current audio block includes at least one of the amplitude, energy, and loudness of the current audio block.
[0032] In some embodiments, if the signal strength of the current audio block includes the loudness of the current audio block and the transformation value includes a loudness gain value, the transformation unit is specifically configured to determine a first reference loudness value of the current audio block based on the original audio data of the current audio block; determine a normalized target loudness value corresponding to the current audio block; and determine the loudness gain value of the current audio block based on the normalized target loudness value and the first reference loudness value of the current audio block.
[0033] In some embodiments, the transformation unit is specifically configured to, for each audio frame included in the current audio block, determine a first reference loudness value of the audio frame based on the original audio data of the audio frame; and determine a first reference loudness value of the current audio block based on the first reference loudness value of each audio frame.
[0034] In some embodiments, the transformation unit is specifically configured to determine the loudness value of the target audio data; and determine a normalized target loudness value corresponding to the current audio block based on the loudness value of the target audio data.
[0035] In some embodiments, the transformation unit is specifically configured to subtract the normalized target loudness value from the first reference loudness value of the current audio block to obtain a second difference value; and obtain the loudness gain value based on the second difference value and a second preset value.
[0036] In some embodiments, the transformation unit is specifically configured to multiply the second difference value by the second preset value to obtain a second product; and perform a preset operation on the second product to obtain the loudness gain value.
[0037] In some embodiments, the encoding unit is further configured to write the first reference loudness value of the current audio block into the code stream.
[0038] In some embodiments, the transformation unit is specifically configured to multiply the loudness value of each element of the current audio block by the loudness gain value to obtain the transformed current audio block.
[0039] In some embodiments, the encoding unit is further configured to write the length information of the current audio block into the code stream.
[0040] In a fifth aspect, an electronic device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in any one of the first aspect to the second aspect or its various implementation manners described above.
[0041] In some embodiments, the block partitioning unit is specifically configured to determine the audio block partitioning length based on the length of the target audio data; and perform block partitioning on a plurality of audio frames included in the target audio data based on the audio block partitioning length to obtain the M audio blocks.
[0042] In a sixth aspect, there is provided a chip for implementing the method in any one of the first aspect to the second aspect or its various implementation manners. Specifically, the chip includes: a processor configured to call and run a computer program from a memory, so that a device installed with the chip executes the method in any one of the first aspect to the second aspect or its various implementation manners.
[0043] In a seventh aspect, there is provided a computer-readable storage medium for storing a computer program, where the computer program causes a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0044] In an eighth aspect, there is provided a computer program product including computer program instructions, where the computer program instructions cause a computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0045] In a ninth aspect, there is provided a computer program, which when running on a computer, causes the computer to execute the method in any one of the first aspect to the second aspect or its various implementation manners.
[0046] In summary, in this application, by decoding a bitstream, a current audio block reconstructed by a decoding network is obtained. The current audio block is one of the M audio blocks obtained by performing block partitioning on target audio data, where M is a positive integer greater than 1, and an audio block includes two or more audio frames. Then, an inverse transform is performed on the signal strength of the reconstructed current audio block to obtain the reconstruction value of the current audio block. That is to say, in the embodiments of this application, two or more audio frames are divided into one audio block, and then signal strength transformation processing is performed in units of audio blocks, so that smooth processing of audio signal strength can be achieved, such as achieving smooth transformation of loudness, amplitude, energy, etc., thereby improving the decoding effect of audio data. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1Schematic block diagram of an audio encoding and decoding system according to an embodiment of the present application;
[0049] Figure 2 Schematic diagram of an end-to-end audio encoding and decoding system based on deep learning according to an embodiment of the present application;
[0050] Figure 3 Network structure block diagram of an encoder-decoder constructed based on a convolutional neural network in an embodiment of the present application;
[0051] Figure 4 Schematic diagram of an end-to-end audio encoding and decoding system based on deep learning according to an embodiment of the present application;
[0052] Figure 5 Schematic flow diagram of an audio decoding method provided in an embodiment of the present application;
[0053] Figure 6 Schematic diagram of an inverse quantization;
[0054] Figure 7 Schematic flow diagram of an audio encoding method provided in an embodiment of the present application;
[0055] Figure 8 Schematic diagram of a quantization;
[0056] Figure 9 Schematic block diagram of an audio decoding device provided in an embodiment of the present application;
[0057] Figure 10 Schematic block diagram of an audio encoding device provided in an embodiment of the present application;
[0058] Figure 11 Schematic block diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0060] It should be noted that in the description of the present application, the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.
[0061] The technical solution proposed in the present application can be applied to technical fields such as artificial intelligence and audio coding and decoding, and is used to reduce the complexity of audio coding and decoding while ensuring the performance of audio and video coding and decoding, thereby improving the efficiency of audio coding and decoding.
[0062] The following introduces the related concepts involved in the embodiments of the present application.
[0063] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0064] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0065] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0066] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0067] This application example mainly introduces the application of artificial intelligence technology in audio encoding and decoding technology.
[0068] Audio encoding and decoding: The audio encoding process compresses audio into smaller data, and the decoding process restores the smaller data to audio. The encoded smaller data is used for network transmission and occupies less bandwidth.
[0069] Audio sampling rate: The audio sampling rate describes the number of data contained in a unit of time (1 second). For example, an 8k sampling rate contains 8000 sampling points, and each sampling point corresponds to a short integer.
[0070] Codebook: A collection of multiple vectors, and the same codebook is saved on both the encoder and decoder sides.
[0071] Quantization: Find the vector in the codebook that is closest to the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.
[0072] Quantizer: The quantizer is responsible for the quantization work and is responsible for updating the vectors in the codebook.
[0073] Audio frame: Represents the minimum voice duration of a single transmission in the network.
[0074] Short Time Fourier Transform: STFT. Divide a long-time signal into several shorter equal-length signals, and then calculate the Fourier transform of each shorter segment separately. It is usually used to depict the changes in the frequency domain and time domain and is an important tool in time-frequency analysis.
[0075] The audio encoding and decoding method provided by the embodiments of the present application can be applied to the fields of audio encoding and decoding, hardware audio encoding and decoding, dedicated circuit video encoding and decoding, real-time audio encoding and decoding, etc. For example, the solution of the present application can be combined with the audio video coding standard (AVS for short), for example, the H.264 / audio video coding (AVC) standard. Alternatively, the solution of the present application can be combined with other proprietary or industry standards. It should be understood that the technology of the present application is not limited to any specific encoding and decoding standard or technology.
[0076] The audio encoding and decoding method provided by the embodiments of the present application can be applied to any end-to-end audio encoding and decoding solution based on deep learning.
[0077] For ease of understanding, first, in combination with Figure 1 the audio encoding and decoding system involved in the embodiments of the present application will be introduced.
[0078] Figure 1 is a schematic block diagram of an audio encoding and decoding system involved in the embodiments of the present application. It should be noted that Figure 1 is only an example. The audio encoding and decoding system of the embodiments of the present application includes but is not limited to Figure 1 as shown. As Figure 1 shown, the audio encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) audio data to generate a bitstream and transmit the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded audio data.
[0079] The encoding device 110 of the embodiments of the present application can be understood as a device with audio encoding function, and the decoding device 120 can be understood as a device with audio decoding function. That is, the embodiments of the present application include a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, TVs, cameras, playback devices, digital media players, audio game consoles, in-vehicle computers, etc.
[0080] In some embodiments, the encoding device 110 can transmit the encoded audio data (such as the bitstream) to the decoding device 120 via the channel 130. The channel 130 can include one or more media and / or devices capable of transmitting the encoded audio data from the encoding device 110 to the decoding device 120.
[0081] In one example, channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded audio data directly to the decoding device 120 in real time. In this example, the encoding device 110 may modulate the encoded audio data according to a communication standard and transmit the modulated audio data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.
[0082] In another example, channel 130 includes a storage medium that can store the audio data encoded by the encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded audio data from the storage medium.
[0083] In another example, channel 130 may include a storage server that can store the audio data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded audio data from the storage server. Optionally, the storage server can store the encoded audio data and can transmit the encoded audio data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.
[0084] In some embodiments, the encoding device 110 includes an audio encoder 112 and an output interface 113. Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0085] In some embodiments, in addition to including the audio encoder 112 and the input interface 113, the encoding device 110 may further include an audio source 111.
[0086] The audio source 111 may include at least one of an audio acquisition device (e.g., a microphone), an audio archive, an audio input interface, and a computer voice system. Among them, the audio input interface is used to receive audio data from an audio content provider, and the computer voice system is used to generate audio data.
[0087] The audio encoder 112 encodes the audio data from the audio source 111 to generate a bitstream. The bitstream contains the encoded information of the audio data in the form of a bitstream. The encoded information may include the encoded audio data and associated data. The associated data may include quantization parameters and other syntax structures, etc. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0088] The audio encoder 112 directly transmits the encoded audio data to the decoding device 120 via the output interface 113. The encoded audio data can also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.
[0089] In some embodiments, the decoding device 120 includes an input interface 121 and an audio decoder 122.
[0090] In some embodiments, in addition to the input interface 121 and the audio decoder 122, the decoding device 120 may further include a playback device 123.
[0091] Among them, the input interface 121 includes a receiver and / or a modem. The input interface 121 can receive the encoded audio data through the channel 130.
[0092] The audio decoder 122 is used to decode the encoded audio data to obtain the decoded audio data and transmit the decoded audio data to the playback device 123.
[0093] The playback device 123 plays the decoded audio data. The playback device 123 can be integrated with the decoding device 120 or outside the decoding device 120. The playback device 123 can include various playback devices.
[0094] In addition, Figure 1 For example only, the technical solutions of the embodiments of the present application are not limited to Figure 1 , for example, the technology of the present application can also be applied to unilateral audio encoding or unilateral audio decoding.
[0095] Figure 2 It is a schematic diagram of an end-to-end audio encoding and decoding system based on deep learning involved in the embodiments of the present application. As Figure 2 shown, the audio encoding and decoding system of the embodiments of the present application includes: an encoding network 210, a quantization module 211, an inverse quantization module 212, and a decoding network 213.
[0096] During encoding, the encoding end (also called the sending end) will first input the input audio data into the encoding network 210 for non-linear transformation to obtain the encoding vector of the input audio data (also called the embedding sequence or latent variable, etc.). Then, the quantization module 211 quantizes the encoding vector of the audio data to obtain the quantization result of the encoding vector. For example, a residual-based vector quantizer is used to select the corresponding quantization parameters according to the target bit rate. Finally, the quantized encoding vector is encoded and converted into a binary bitstream.
[0097] During decoding, the decoding end (also known as the receiving end) first recovers the quantization result of the encoded vector from the bitstream, and then further recovers the encoded vector through the inverse quantization module 212 and inputs it into the decoding network 213 for non-linear transformation to obtain the reconstructed audio data.
[0098] Figure 3 It is a network structure block diagram of an encoder-decoder constructed based on a convolutional neural network in an embodiment of the present application.
[0099] As Figure 3 shown, the network structure of the encoder-decoder includes an encoding network 310 and a decoding network 320. Among them, the encoding network 310 can be implemented as software such as Figure 1 shown in the audio-video encoding device 110, and the decoding network 320 can be implemented as software such as Figure 1 shown in the audio-video decoding device 120. In some embodiments, the encoding network 310 is also referred to as the encoder 310, and the decoding network 320 is also referred to as the decoder 320.
[0100] At the data sending end, the audio data can be encoded and compressed through the encoding network 310. In an embodiment of the present application, the encoding network 310 may include an input layer 311, one or more encoding modules 312, and an output layer 313.
[0101] Exemplarily, the input layer 311 and the output layer 313 may be convolutional layers constructed based on one-dimensional convolutional kernels. A plurality of (for example, 4) encoding modules (EncoderBlock) 312 are sequentially connected between the input layer 311 and the output layer 313. Each encoding module 312 includes a plurality of residual (ResidualUnit) modules, and each residual module contains a plurality of convolutional layers.
[0102] For example, in the input stage of the encoder, data sampling is performed on the original audio data to be encoded, and a vector with a channel number of c and a dimension of w can be obtained; this vector is input into the input layer 311, and after convolutional processing, a feature vector with a channel number of 32c and a dimension of w can be obtained. In some optional implementation manners, in order to improve the encoding efficiency, the encoding network 310 can simultaneously perform encoding processing on a batch of audio vectors.
[0103] In the downsampling stage of the encoder, the first encoding module reduces the vector dimension to 1 / 2 and doubles the number of channels, obtaining a feature vector with 64c channels and a dimension of 1 / 2w; the second encoding module reduces the vector dimension to 1 / 4 and doubles the number of channels, obtaining a feature vector with 128c channels and a dimension of 1 / 8w; the third encoding module reduces the vector dimension to 1 / 6 and doubles the number of channels, obtaining a feature vector with 256c channels and a dimension of 1 / 48w; the fourth encoding module reduces the vector dimension to 1 / 8 and doubles the number of channels, obtaining a feature vector with 512c channels and a dimension of 1 / 384w.
[0104] In the output stage of the encoder, the output layer 313 performs convolution processing on the feature vector with 512c channels and a dimension of 1 / 384w to obtain an encoded vector with 1 channel and a dimension of K.
[0105] The encoded vector is input into the quantizer 330, and the codebook index corresponding to the encoded vector can be queried in the codebook, and the codebook index is encoded to obtain a binary code stream, and then the binary code stream is sent to the data receiving end.
[0106] The data receiving end receives the binary code stream, decodes it to obtain the codebook index, and performs inverse quantization based on the codebook index to obtain the reconstructed encoded vector. Finally, the reconstructed encoded vector is decoded by the decoding network 320 to obtain the restored audio data.
[0107] In an embodiment of the present application, the decoding network 320 may include an input layer 321, one or more decoding modules 322, and an output layer 323. Each decoding module 322 includes a plurality of ResidualUnit modules, and each residual module contains a plurality of convolutional layers.
[0108] After the data receiving end decodes the code stream to obtain the codebook index, it can first query the codebook vector corresponding to the codebook index in the codebook through the quantizer 320, and then obtain the encoded vector for audio data reconstruction based on the codebook vector. For example, the reconstructed encoded vector may be a vector with 1 channel and a dimension of K. In some alternative embodiments, to improve the decoding efficiency, the data receiving end may decode a batch of codebook vectors simultaneously.
[0109] In the input stage of the decoder, the reconstructed encoded vector is input into the input layer 321, and after convolution processing, a feature vector with 512c channels and a dimension of 1 / 384w can be obtained.
[0110] In the decoding stage of the decoder, the first decoding module increases the vector dimension by 8 times and reduces the number of channels by 2 times, obtaining a feature vector with 256 channels and a dimension of 1 / 48w; the second decoding module increases the vector dimension by 6 times and reduces the number of channels by 2 times, obtaining a feature vector with 128 channels and a dimension of 1 / 8w; the third decoding module increases the vector dimension by 4 times and reduces the number of channels by 2 times, obtaining a feature vector with 64 channels and a dimension of 1 / 2w; the fourth decoding module increases the vector dimension by 2 times and reduces the number of channels by 2 times, obtaining a feature vector with 32 channels and a dimension of w.
[0111] In the output stage of the decoder, after the output layer 323 performs convolution processing on the feature vector with 32 channels and a dimension of w, the reconstructed audio data with 1 channel and a dimension of w is restored.
[0112] Loudness, also known as volume, is the strength of the sound felt by the human ear. It is a subjective perception of the sound size by humans. The size of loudness depends on the amplitude of the sound at the receiving end. For the same sound source, the farther the amplitude propagates, the smaller the loudness; when the propagation distance is fixed, the greater the amplitude of the sound source, the greater the loudness. The size of loudness is closely related to sound intensity, but the change of loudness with sound intensity is not a simple linear relationship, but close to a logarithmic relationship. When the frequency of the sound and the waveform of the sound wave change, the human perception of the loudness size will also change.
[0113] In the radio and television industry, when viewers switch between different channels or when a TV drama switches to an advertisement or an advertisement switches to a TV drama, they will find that there is an obvious change in volume. For example, when a TV drama switches to an advertisement, the sound of the advertisement is too loud or too small, and the volume needs to be adjusted frequently, resulting in a poor user experience. Therefore, it is necessary to normalize the loudness of the audio data.
[0114] Based on this, in some scenarios, in order to improve the encoding and decoding effect of audio data, before encoding the audio data, the encoding end first transforms the signal intensity of the audio data, such as normalizing the loudness of the audio data. Exemplarily, before inputting the audio data to be encoded into the Figure 3 encoding network, the encoding end first transforms the signal intensity of the audio data, such as normalizing the loudness. Correspondingly, the decoding end performs an inverse transformation on the signal intensity of the audio data reconstructed by the Figure 3 decoding network, such as inverse normalizing the loudness, so as to obtain the reconstructed value of the audio data.
[0115] In one example, the end-to-end audio encoding and decoding system provided by the embodiments of the present application is as Figure 4As shown in the figure, the encoding device 410 includes: a signal strength inverse transformation module 411, an encoding network 412, a quantizer 413, and an encoding module 414. Correspondingly, the decoding device 420 includes: a decoding module 421, an inverse quantizer 422, a decoding network 423, and a signal strength transformation module 424.
[0116] During encoding, at the encoding end, the signal strength inverse transformation module 411 first performs an inverse transformation on the signal strength of the audio data to be encoded, for example, performs loudness normalization processing, to obtain the audio data after signal strength transformation. Then, the audio data after signal strength transformation is input into the encoding network 412 for non-linear conversion to obtain the encoding vector of the audio data. Next, the quantizer 413 performs quantization processing on the encoding vector of the audio data to obtain the quantization result of the audio data. Finally, the encoding module 414 encodes the quantization result of the audio data to obtain a bitstream.
[0117] During decoding, at the decoding end, the decoding module 421 first decodes the bitstream to obtain the quantization result of the audio data, and then the inverse quantizer 422 performs inverse quantization processing on the quantization result of the audio data to obtain the decoding vector of the audio data. Next, the decoding network 423 performs non-linear transformation on the decoding vector of the audio data to obtain the reconstructed value of the audio data after signal strength transformation. Finally, the signal strength transformation module 424 performs an inverse transformation on the signal strength of the reconstructed value of the audio data after signal strength transformation to obtain the reconstructed value of the audio data.
[0118] In some embodiments, the signal strength transformation is uniformly performed on an entire audio file, which is applicable to the offline encoding scenario and not applicable to the real-time scenario (Real-Time Communication, abbreviated as RTC) scenario.
[0119] In some embodiments, in the real-time scenario, by default, a single audio frame is used as the unit of signal strength transformation, and the signal strength transformation is performed on each audio frame, which easily leads to large fluctuations in signal strength.
[0120] To solve the above technical problems, in the encoding end of the embodiments of the present application, the target audio data to be encoded is divided into multiple audio blocks, each audio block including two or more audio frames. Then, the signal strength of each audio block is transformed, and then the transformed audio block is encoded to obtain a bitstream. In this way, the decoding end decodes the bitstream to obtain the currently reconstructed audio block by the decoding network, and performs an inverse transformation on the signal strength of the reconstructed currently audio block. For example, an inverse transformation is performed on the signal strength of the reconstructed currently audio block by using a transformation method opposite to that of the encoding end to obtain the reconstructed value of the currently audio block. That is to say, in the embodiments of the present application, two or more audio frames are divided into one audio block, and then the signal strength is transformed and processed in units of audio blocks. This can not only be applied to real-time scenarios, but also achieve smooth processing of the audio signal strength, thereby improving the encoding and decoding effect of audio data.
[0121] The technical solutions of the embodiments of the present application will be described in detail through some embodiments below. These embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0122] First, taking the decoding end as an example, the audio decoding method provided by the embodiments of the present application will be introduced.
[0123] Figure 5 It is a schematic flowchart of the audio decoding method provided by an embodiment of the present application. The execution subject of the embodiments of the present application is a device with audio decoding function, such as an audio decoding device. In some embodiments, the audio decoding device may be the Figure 1 decoding device in. For ease of description, the embodiments of the present application will be described by taking the execution subject as the decoding device as an example.
[0124] As Figure 5 shown, the audio decoding method of the embodiments of the present application includes:
[0125] S101. Decode the bitstream to obtain the currently reconstructed audio block by the decoding network.
[0126] Among them, the currently audio block is one of the M audio blocks obtained by block-dividing the target audio data, M is a positive integer greater than 1, and the audio block includes two or more audio frames.
[0127] In the embodiments of the present application, the target audio data to be decoded can be an audio data segment of any length.
[0128] In one example, the target audio data to be decoded includes multiple audio frames.
[0129] In the embodiments of the present application, an audio frame can be understood as a data segment with a specified time length obtained by performing frame division and windowing processing on the original audio data.
[0130] The embodiments of the present application do not limit the specific acquisition method of the original audio data.
[0131] In some examples, the original audio data can be the voice collected by the terminal.
[0132] In some examples, the original audio data can be the sound signal collected in the scenario of a network voice call or a video call.
[0133] In some examples, the original audio data can be the sound signal collected in the live broadcast scenario, or the sound signal collected in the online singing scenario, or the sound signal collected in the voice broadcast scenario.
[0134] In some examples, the original audio data can be the audio data obtained from the storage resource. For example, the original audio data can be the stored voice, music, video, etc.
[0135] In a possible implementation manner, when the embodiments of the present application perform audio frame division on the original audio data, a preset duration can be set for division. For example, every 10 ms of the original audio data in the original audio data is divided into an audio frame.
[0136] In order to store and transmit the audio data over a long distance, it is necessary to perform audio encoding on the obtained original audio data to reduce the size of the audio data, thereby reducing the storage space of the audio data or reducing the traffic bandwidth consumed by the long-distance transmission.
[0137] In the audio encoding process, first, the obtained target audio data is divided into blocks to obtain M audio blocks. Each of the M audio blocks includes two or more audio frames. For the current audio block among the M audio blocks, the signal strength of the current audio block is transformed to obtain the current audio block after the signal strength transformation. Then, a non-linear transformation is performed on the transformed current audio block. For example, the current audio block is upsampled and feature-transformed through an encoding network to obtain the encoding vector of the current audio block. Then, a quantizer is used to quantize the encoding vector of the current audio block, and finally, the quantized encoding vector is encoded to obtain a bitstream.
[0138] In the embodiments of the present application, the transformation of the signal strength of the current audio block includes the transformation of one or more of the dimensions describing the signal strength, such as the amplitude, energy, and loudness of the current audio frame.
[0139] In the embodiments of the present application, the number of audio frames included in each of the M audio blocks may be the same or different. That is to say, the length of each audio block may be exactly the same, or completely different, or not completely the same. The embodiments of the present application do not limit this and can be determined based on actual needs.
[0140] In some embodiments, the M audio blocks are obtained by dividing every K consecutive audio frames among the multiple audio frames included in the target audio data into one audio block, where K is a positive integer greater than 1. For example, when K = 3 and the target audio data includes 21 audio frames, every 3 consecutive audio frames of the target audio data can be divided into one audio block, resulting in 7 audio blocks. For example, the first audio frame, the second audio frame, and the third audio frame in the target audio data are divided into one audio block, the fourth audio frame, the fifth audio frame, and the sixth audio frame are divided into one audio block, and so on. Another example is that when K = 3 and the target audio data includes 14 audio frames, then every 3 consecutive audio frames among the first 12 audio frames of the target audio data can be divided into one audio block, resulting in 4 audio blocks, and the last 2 audio frames among these 14 audio frames are divided into one audio block, for a total of 5 audio blocks. Or, every 3 consecutive audio frames among the first 9 audio frames of these 14 audio frames are divided into one audio block, resulting in 3 audio blocks, and finally the last 5 audio frames of these 14 audio frames are divided into one audio block, for a total of 4 audio blocks. Of course, according to actual needs, other division methods may also be included, and the embodiments of the present application do not limit this.
[0141] In some embodiments, the M audio blocks are obtained by performing block division on the multiple audio frames included in the target audio data based on the audio block division length determined according to the length of the target audio data. Among them, the greater the length of the target audio data, the greater the corresponding audio block division length. The specific division process can refer to the relevant description at the encoding end.
[0142] It should be noted that in the encoding process of the embodiments of the present application, in addition to performing loudness normalization in units of audio blocks during signal strength transformation, in other stages of encoding (such as the non-linear transformation and quantization stages), non-linear transformation and quantization processing can be performed in units of audio blocks or in units of audio frames. The embodiments of the present application do not limit this.
[0143] Correspondingly, at the decoding end, in addition to performing signal strength transformation in units of audio blocks during signal strength transformation, in other stages of decoding (such as the inverse quantization stage and the non-linear transformation stage), inverse quantization and non-linear transformation can be performed in units of audio blocks or in units of audio frames. The embodiments of the present application do not limit this.
[0144] In the embodiments of the present application, the decoding end decodes each of the M audio blocks one by one. In the embodiments of the present application, the processes of decoding each of the M audio blocks by the decoding end are basically the same. For the convenience of description, the current audio block among the M audio frames is taken as an example for illustration here.
[0145] In some embodiments, when the decoding end decodes with an audio block as the decoding unit, for the current audio block, the decoding end first obtains the bitstream of the current audio block, decodes the bitstream of the current audio block to obtain the quantization result of the current audio block. Then, the quantization result of the current audio block is inverse-quantized to obtain the decoding vector of the current audio block, and the decoding vector of the current audio block is input into the decoding network for decoding to obtain the current audio block after signal strength transformation.
[0146] In some embodiments, when the decoding end decodes with an audio frame as the decoding unit, for each audio frame included in the current audio block, the decoding end first decodes the bitstream to obtain the quantization result of the audio frame. Then, the quantization result of the audio frame is inverse-quantized to obtain the decoding vector of the audio frame, and the decoding vector of the audio frame is input into the decoding network for decoding to obtain the reconstructed audio frame. Referring to this method, the reconstructed audio frames of each audio frame in the current audio block can be decoded, and then the reconstructed current audio block can be obtained.
[0147] In the embodiments of the present application, since the inverse transformation process of signal strength is performed with an audio block as the unit. Therefore, when the decoding end decodes the current audio block, it is first necessary to determine the length information of the current audio block. That is to say, the above S101 includes the following steps of S101-A and S101-B:
[0148] S101-A. Determine the length information of the current audio block;
[0149] S101-B. Based on the length information of the current audio block, decode the bitstream corresponding to the current audio block to obtain the current audio block after loudness normalization.
[0150] In the embodiments of the present application, when decoding the current audio block, it is first necessary to determine the length information of the current audio block, and the length information of the current audio block can be understood as the information indicating the number of audio frames included in the current audio block.
[0151] The embodiments of the present application do not limit the specific manner in which the decoding end determines the length information of the current audio block.
[0152] In a possible implementation, the decoding end and the encoding end adopt a default length as the length of the audio block. For example, the decoding end and the encoding end default that the audio block includes 3 audio frames or 4 audio frames, etc. In this way, the decoding end determines the default length as the length information of the current audio block.
[0153] In a possible implementation, the encoding end determines the block division length of the audio block. For example, the encoding end determines that every 3 (of course, other values such as 3, 5, etc.) audio frames of the target audio data are divided into one audio block. In this way, the encoding end writes the length information of the current audio block into the code stream. At this time, the decoding end can obtain the length information of the current audio frame by decoding the code stream.
[0154] In an example, if the lengths of each of the M audio blocks are the same, the encoding end can write the length information of the audio block once in the code stream. For example, the encoding end writes the length information of the audio block in the syntax information of the code stream. Correspondingly, the decoding end can decode the length information once and obtain the length information of each of the M audio blocks. For example, the decoding end decodes the syntax information in the code stream to obtain the length information of each of the M audio blocks.
[0155] In an example, the encoding end can write the length information of each of the M audio blocks into the code stream. Correspondingly, the decoding end can obtain the length information of each of the M audio blocks by decoding the code stream.
[0156] In an example, the encoding end can write the length information of the audio blocks with inconsistent lengths among the M audio blocks into the code stream. In this way, the decoding end can obtain the length information of each of the audio blocks with different lengths among the M audio blocks by decoding the code stream.
[0157] After obtaining the length information of the current audio block based on the above steps, the decoding end determines the code stream of the current audio block from the code stream based on the length information of the current audio block, and then decodes the code stream of the current audio block. For example, it performs inverse quantization and non-linear transformation processing to obtain the reconstructed current audio block.
[0158] The quantization and non-linear transformation processes in the decoding process are introduced below.
[0159] In some embodiments, if the quantizer used in the embodiments of the present application is a residual-based vector quantizer, the quantization result corresponding to the current audio block decoded at the decoding end includes the codebook index corresponding to each quantization layer of the quantizer and the residual vector corresponding to the last quantization layer. At this time, when the decoding end performs inverse quantization on the quantization result, first, based on the codebook index corresponding to each quantization layer, the corresponding vector is queried in the corresponding codebook, and then the vectors obtained by the query are added to the residual vector corresponding to the last quantization layer to obtain the decoded vector of the current audio block (also referred to as the reconstructed encoded vector).
[0160] Exemplarily, as Figure 6 shown, assuming that the residual-based vector quantizer includes 3 quantization layers, the decoding end decodes the code stream to obtain the residual vector corresponding to the last quantization layer of the quantizer and the codebook indexes index0, index1, and index2 corresponding to each of the 3 quantization layers. In this way, the decoding end first queries the codebook vector corresponding to the codebook index index2 in the codebook corresponding to the 3rd quantization layer, and then determines the reconstructed value y3' of the output vector of the 3rd vector layer as the codebook vector. The sum value of the vector y3' and the residual vector is determined as the reconstructed value of the residual vector Next, the codebook vector corresponding to the codebook index index1 is queried in the codebook corresponding to the 2nd quantization layer, and then the codebook vector is determined as the reconstructed value y2' of the output vector of the 2nd vector layer. The sum value of the output vector y2' and the reconstructed value of the residual vector is determined as the reconstructed value of the residual vector Next, the codebook vector corresponding to the codebook index index0 is queried in the codebook corresponding to the 1st quantization layer, and then the codebook vector is determined as the reconstructed value y1' of the output vector of the 1st vector layer. The sum value of the vector y1' and the reconstructed value of the residual vector is determined as the reconstructed encoded vector y0' of the audio data. Next, the codebook vector corresponding to the codebook index index0 is queried in the codebook corresponding to the 1st quantization layer, and then the codebook vector is determined as the reconstructed value y1' of the output vector of the 1st vector layer. The sum value of the vector y1' and the reconstructed value of the residual vector is determined as the reconstructed encoded vector y0' of the audio data.
[0161] In some embodiments, the number of quantization layers included in the above Figure 6 shown residual-based vector quantizer can be adaptively selected based on the statistical value of the effective audio information of the current audio block to be processed.
[0162] After obtaining the decoded vector of the current audio block based on the above steps, the decoding end performs non-linear transformation processing on the decoded vector of the current audio block through a decoding network, such as performing upsampling and feature transformation processing, to obtain the reconstructed current audio block.
[0163] For example, through the above Figure 3The decoding network shown performs a non - linear transformation on the decoding vector of the current audio block to obtain the reconstructed current audio block.
[0164] Based on the above steps, after the decoding end decodes the bitstream and obtains the current audio block reconstructed by the decoding network, it executes the steps of S102 as follows.
[0165] S102: Perform an inverse transformation on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block.
[0166] In the embodiment of the present application, the target audio data is divided into M audio blocks, and each audio block in the M audio blocks includes two or more audio frames. For the current audio block in the M audio blocks, during decoding, first, based on the above steps, the current audio block reconstructed by the decoding network is determined, and then the inverse transformation process of the signal strength of the entire current audio block is performed, so that the signal strength of the current audio frame is smoother, thereby improving the decoding effect of the audio data.
[0167] The embodiment of the present application does not limit the specific manner in which the decoding end performs an inverse transformation on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block.
[0168] In some embodiments, the above S102 includes the steps of S102 - A and S102 - B as follows:
[0169] S102 - A: Determine the inverse transformation value of the signal strength of the current audio block;
[0170] S102 - B: Based on the inverse transformation value, perform an inverse transformation on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block.
[0171] In this implementation manner, when the decoding end performs an inverse transformation on the signal strength of the current audio block reconstructed by the decoding network, it first determines the inverse transformation value of the signal strength of the current audio block.
[0172] In the embodiment of the present application, the decoding end can determine the inverse transformation value of the signal strength of the current audio block at least through the following several methods:
[0173] In Method 1, when the encoding end transforms the signal strength of the current audio block, it determines the transformation value of the signal strength of the current audio block. Then, while transforming the signal strength of the current audio block based on the transformation value of the signal strength of the current audio block, it can also write the determined transformation value of the signal strength of the current audio block into the bitstream. In this way, the decoding end decodes the bitstream to obtain the transformation value of the signal strength of the current audio block, and then determines the inverse transformation value of the signal strength of the current audio block based on the transformation value of the signal strength of the current audio block. For example, the decoding end determines the reciprocal of the transformation value of the signal strength of the current audio block as the inverse transformation value of the signal strength of the current audio block.
[0174] In the embodiments of the present application, the inverse transformation can be understood as the inverse process of the transformation.
[0175] In Method 2, the decoding end determines the inverse transformation value of the signal strength of the current audio block through the following steps of S102 - A1:
[0176] S102 - A1: Based on the reconstructed current audio block, determine the inverse transformation value of the signal strength of the current audio block.
[0177] In this Method 2, the decoding end determines the inverse transformation value of the signal strength of the current audio block based on the reconstructed audio data of the current audio block reconstructed by the above decoding network.
[0178] The embodiments of the present application do not limit the specific method for the decoding end to determine the inverse transformation value of the signal strength of the current audio block based on the reconstructed current audio block.
[0179] In a possible implementation, based on a preset transformation calculation formula, the signal strength of the reconstructed audio data of the reconstructed current audio block is calculated to obtain the first signal strength value of the current audio block. Then, based on the first signal strength value of the current audio block, the inverse transformation value of the current audio block is determined. For example, a correspondence table between different signal strength values and inverse transformation values of different signal strengths is set up, and this correspondence table can be obtained through calculation, experiment, or experience. Then, the inverse transformation value of the signal strength corresponding to the first signal strength value in this correspondence is determined as the inverse transformation value of the signal strength of the current audio block.
[0180] In some embodiments, the signal strength of the current audio block includes at least one of the amplitude, energy, and loudness of the current audio block. That is to say, in the embodiments of the present application, the transformation of the signal strength of the current audio block includes the transformation of at least one of the amplitude, energy, and loudness of the current audio block. Correspondingly, the inverse transformation of the signal strength of the current audio block includes the inverse transformation of at least one of the amplitude, energy, and loudness of the current audio block.
[0181] In a possible implementation, if the signal strength of the current audio frame includes the loudness of the current audio block and the inverse transform value includes the loudness inverse gain value, then the above S102-A1 includes the following steps of S102-A11 to S102-A13:
[0182] S102-A11. Decode the bitstream to obtain a first reference loudness value of the current audio block, where the first reference loudness value is calculated based on the original audio data of the current audio block;
[0183] S102-A12. Determine a second reference loudness value of the current audio block based on the reconstructed current audio block;
[0184] S102-A13. Determine the loudness inverse gain value based on the first reference loudness value and the second reference loudness value of the current audio block.
[0185] In the embodiments of the present application, if the transformation of the signal strength of the current audio block at the encoding end includes loudness normalization processing, then when the encoding end performs loudness normalization processing on the loudness of the current audio block, it calculates the loudness based on the original audio data of the current audio block to obtain a first reference loudness value of the current audio block, and then determines the loudness gain value of the current audio block based on the first reference loudness value of the current audio block. Then, the loudness normalization processing is performed on the current audio block using the loudness gain value of the current audio block. At the same time, the encoding end writes the calculated first reference loudness value of the current audio block into the bitstream. In this way, the decoding end obtains the first reference loudness value of the current audio block by decoding the bitstream.
[0186] It should be noted that there is no sequence preference between the above S102-A12 and the above S102-A11, that is, S102-A12 can be executed before the above S102-A11, or after S102-A11, or executed synchronously with the above S102-A11.
[0187] In the embodiments of the present application, the decoding end calculates the loudness based on the reconstructed audio data of the reconstructed current audio block to obtain a second reference loudness value of the current audio block.
[0188] The embodiments of the present application do not limit the specific manner in which the decoding end determines the second reference loudness value of the current audio block based on the reconstructed audio data of the reconstructed current audio block.
[0189] In a possible implementation, the decoding end uses a preset loudness calculation method, such as the Moore and Zwicker loudness calculation model, to calculate the loudness based on the reconstructed audio data of the reconstructed current audio block to obtain a loudness value, and then records the loudness value as the second reference loudness value of the current audio block.
[0190] In a possible implementation, the decoding end can determine the second reference loudness value of the current audio block through the following steps S102-A12-a and S102-A12-b:
[0191] S102-A12-a. For each audio frame included in the current audio block, determine the second reference loudness value of the audio frame based on the reconstructed audio data after loudness normalization of the audio frame;
[0192] S102-A12-b. Based on the second reference loudness value of each audio frame, determine the second reference loudness value of the current audio block.
[0193] In the embodiments of the present application, in order to improve the smoothness of loudness, multiple audio frames are divided into one audio block for overall loudness normalization processing. That is to say, the current audio block in the embodiments of the present application includes multiple audio frames, for example, includes 2 or more than 2 audio frames. When the decoding end calculates the second reference loudness value of the current audio block, it determines the second reference loudness value of each audio frame included in the current audio block by performing loudness calculation on each audio frame included in the current audio block, and then determines the second reference loudness value of the current audio block based on the second reference loudness value of each audio frame included in the current audio block.
[0194] The embodiments of the present application do not limit the specific manner in which the decoding end determines the second reference loudness value of each audio frame in the current audio block. For example, for each audio frame included in the current audio block, the decoding end samples a preset loudness calculation method and performs loudness calculation based on the reconstructed audio data after loudness normalization of the audio frame to obtain the second reference loudness value of the audio frame.
[0195] After the decoding end determines the second reference loudness value of each audio frame in the current audio block, it determines the second reference loudness value of the current audio block based on the second reference loudness value of each audio frame. For example, if the current audio block includes 3 audio frames, the average value or weighted average value of the second reference loudness values of these 3 audio frames is determined as the second reference loudness value of the current audio block. For another example, a preset data processing method is sampled to process the second reference loudness values of the 3 audio frames included in the current audio block to obtain the second reference loudness value of the current audio block.
[0196] After the decoding determines the first reference loudness value and the second reference loudness value of the current audio block based on the above steps, it executes the above step S102-A13 to determine the loudness anti-gain value based on the first reference loudness value of the current audio block and the second reference loudness value of the current audio block.
[0197] As can be seen from the above, the first reference loudness value of the current audio block can be understood as the loudness value before the loudness of the current audio block is normalized, and the second reference loudness value can be understood as the loudness value after the loudness of the current audio block is normalized. In this way, the decoding end can determine the loudness anti-gain value of the current audio block based on the first reference loudness value and the second reference loudness value of the current audio block.
[0198] The embodiments of the present application do not limit the specific manner of determining the loudness anti-gain value based on the first reference loudness value of the current audio block and the second reference loudness value of the current audio block.
[0199] In one example, the decoding end determines the difference between the first reference loudness value and the second reference loudness value of the current audio block as the loudness anti-gain value of the current audio block.
[0200] In one example, the decoding end determines the loudness gain value of the current audio block through the following steps of S102-A13-a1 and S102-A13-a2:
[0201] S102-A13-a1: Subtract the first reference loudness value of the current audio block from the second reference loudness value of the current audio block to obtain a first difference;
[0202] S102-A13-a2: Based on the first difference and the first preset value, obtain the loudness anti-gain value of the current audio block.
[0203] In this implementation manner, the decoding end subtracts the first reference loudness value of the current audio block from the second reference loudness value of the current audio block to obtain a difference, denoted as the first difference. Then, based on the first difference and the first preset value, the loudness anti-gain value of the current audio block is determined.
[0204] The embodiments of the present application do not limit the specific manner of determining the loudness anti-gain value of the current audio block based on the first difference and the first preset value.
[0205] For example, the decoding end determines the product of the first difference and the first preset value as the loudness anti-gain value of the current audio block.
[0206] For another example, the decoding end multiplies the first difference by the first preset value to obtain a first product, and then performs a preset operation on the first product to obtain the loudness anti-gain value of the current audio block.
[0207] The embodiments of the present application do not limit the specific manner of the preset operation.
[0208] In one example, the above preset operation is a logarithmic operation. That is, the decoding end performs a logarithmic operation on the first product obtained by multiplying the first difference and the first preset value to obtain the loudness anti-gain value of the current audio block.
[0209] Exemplarily, the decoding end can determine the loudness anti-gain value of the current audio block through the following formula (1):
[0210] gain_1 = exp((ref_db - decode_db) * GAIN_FACTOR_1) (1)
[0211] Wherein, gain_1 is the loudness anti-gain value of the current audio frame, exp is the exponential operation with e as the base, ref_db is the first reference loudness value of the current audio block, decode_db is the second reference loudness value of the current audio block, (ref_db - decode_db) is the first difference, GAIN_FACTOR_1 is the first preset value, which is a constant, and (ref_db - decode_db) * GAIN_FACTOR_1 is the first product.
[0212] It should be noted that in addition to using the above method to determine the loudness anti-gain value of the current audio block, the decoding end can also use other methods to determine the loudness anti-gain value of the current audio block.
[0213] After the decoding end determines the inverse transform value of the current audio block based on the above steps, it executes the steps of S102-B above, and based on the inverse transform value of the current audio block, performs an inverse transform process on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block.
[0214] In the embodiments of the present application, the process of the decoding end performing an inverse transform on the signal strength of the reconstructed current audio block based on the inverse transform value of the current audio block can be understood as the inverse process of the encoding end performing a transform on the signal strength of the audio data of the current audio block based on the transform value of the current audio block.
[0215] In some embodiments, if in the embodiments of the present application, the inverse transform of the signal strength of the reconstructed current audio block includes performing an inverse normalization operation on the loudness of the reconstructed current audio frame, then the above decoding end can multiply each element in the reconstructed current audio block by the determined loudness anti-gain value to obtain the reconstructed value of the current audio block.
[0216] Exemplarily, the decoding end uses the following formula (2) to obtain the reconstructed value of the current audio block:
[0217] decode_audio_data_normalize = decode_audio_data * gain_1 (2)
[0218] Among them, decode_audio_data_normalize is the reconstructed value of the current audio block, and decode_audio_data is the current audio block reconstructed by the above decoding network.
[0219] The above embodiments describe the process of the decoding end determining the reconstructed audio data of the current audio block after signal strength transformation, and performing inverse signal strength transformation on the reconstructed audio data of the current audio block after signal strength transformation. The decoding end can refer to the above method to determine the reconstructed value of each audio block among the M audio blocks obtained by dividing the target audio data, and finally splice the reconstructed values of the M audio blocks after inverse signal strength transformation to obtain the complete reconstructed audio data of the target audio data.
[0220] In the audio decoding method provided by the embodiments of the present application, the decoding end obtains the current audio block reconstructed by the decoding network by decoding the code stream. The current audio block is one of the M audio blocks obtained by block-dividing the target audio data, where M is a positive integer greater than 1, and the audio block includes two or more audio frames. Then, inverse signal strength transformation processing is performed on the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block. That is to say, the embodiments of the present application divide two or more audio frames into one audio block, and then perform inverse signal strength transformation processing in units of audio blocks, which can achieve smooth processing of audio signal strength and thus improve the decoding effect of audio data.
[0221] The above introduces the audio decoding method involved in the embodiments of the present application. Next, taking the encoding end as an example, the audio encoding method provided by the embodiments of the present application will be introduced.
[0222] Figure 7 It is a schematic flowchart of the audio encoding method provided by an embodiment of the present application. The execution subject of the embodiments of the present application can be a device with specific audio encoding functions, such as an audio encoding device. In some embodiments, the audio encoding device can be Figure 1 the encoding device in. For ease of description, the embodiments of the present application will be described by taking the execution subject as the encoding device as an example.
[0223] As Figure 7 shown, the audio encoding method of the embodiments of the present application includes the following steps:
[0224] S201. Perform block division on the target audio data to be encoded to obtain M audio blocks.
[0225] Among them, M is a positive integer greater than 1, and the audio block includes two or more audio frames.
[0226] In the embodiments of the present application, the target audio data to be encoded can be an audio data segment of any length.
[0227] In one example, the target audio data to be encoded includes multiple audio frames.
[0228] In the embodiments of the present application, an audio frame can be understood as a data segment with a specified time length obtained after frame division and windowing processing of the original audio data.
[0229] The embodiments of the present application do not limit the specific acquisition method of the original audio data.
[0230] In some examples, the original audio data can be the voice collected by the terminal.
[0231] In some examples, the original audio data can be the sound signal collected in the scenario of a network voice call or a video call.
[0232] In some examples, the original audio data can be the sound signal collected in the live broadcast scenario, or the sound signal collected in the online singing scenario, or the sound signal collected in the voice broadcast scenario.
[0233] In some examples, the original audio data can be the audio data obtained from the storage resource. For example, the original audio data can be the stored voice, music, video, etc.
[0234] In a possible implementation manner, when the embodiments of the present application perform audio frame division on the original audio data, a preset duration can be set for division. For example, every 10 ms of the original audio data in the original audio data is divided into an audio frame.
[0235] The target audio data in the embodiments of the present application can be understood as any type of audio data to be encoded. For example, it can be a segment of audio data to be encoded in the offline state, or a segment of audio data to be encoded in the real-time scenario.
[0236] In order to store and transmit the audio data over a long distance, it is necessary to perform audio encoding on the obtained original audio data to reduce the size of the audio data, thereby reducing the storage space of the audio data or reducing the traffic bandwidth consumed by the long-distance transmission.
[0237] During the audio encoding process, for the encoding effect of the audio data, before the encoding end inputs the target audio data into the encoding network for non-linear transformation, the signal intensity of the target audio data is first transformed. However, in the related signal intensity transformation method, there are signal intensity fluctuations and unevenness, resulting in a poor encoding effect of the audio data.
[0238] To solve the above technical problems, in the embodiments of the present application, the target audio data to be encoded is divided into M audio blocks, and each audio block includes multiple audio frames. For example, each audio block includes 2 or more audio frames. When performing the transformation processing of the signal strength, the audio frames included in one audio block are processed as a whole for the transformation processing of the signal strength. In this way, the problem of loudness fluctuations can be prevented, the smooth processing of the loudness can be achieved, and it can be applied to real-time scenarios, thereby improving the encoding effect of the audio data.
[0239] In the embodiments of the present application, the number of audio frames included in each of the M audio blocks may be the same or different. That is to say, the length of each audio frame may be exactly the same, or completely different, or not completely the same. The embodiments of the present application do not limit this and can be determined based on actual needs.
[0240] In some embodiments, if the lengths of the audio blocks are the same, that is, the number of included audio frames is the same, the encoding end can adopt a unified audio block division length to divide the audio frames included in the target audio data into M audio blocks of the same size, that is, divided into M groups. For example, every K consecutive audio frames among the multiple audio frames included in the target audio data are divided into one audio block, where K is a positive integer greater than 1. For example, K = 3, and the target audio data includes 21 audio frames. In this way, every 3 consecutive audio frames of the target audio data can be divided into one audio block, obtaining 7 audio blocks. For example, the first audio frame, the second audio frame, and the third audio frame in the target audio data are divided into one audio block, the fourth audio frame, the fifth audio frame, and the sixth audio frame are divided into one audio block, and so on. For another example, if K = 3 and the target audio data includes 14 audio frames, then the first 12 audio frames of the target audio data can be divided into one audio block for every 3 consecutive audio frames, obtaining 4 audio blocks, and the last 2 audio frames among these 14 audio frames are divided into one audio block, a total of 5 audio blocks are divided. Or, the first 9 audio frames of these 14 audio frames are divided into one audio block for every 3 consecutive audio frames, obtaining 3 audio blocks, and finally the last 5 audio frames of these 14 audio frames are divided into one audio block, a total of 4 audio blocks are obtained. Of course, according to actual needs, other division methods may also be included, and the embodiments of the present application do not limit this.
[0241] The embodiments of the present application do not limit the determination method of the above audio block division length.
[0242] In some embodiments, the above audio block division length is a preset value. For example, it is preset to divide 3 or 4 audio frames into one audio block.
[0243] In some embodiments, the audio block division length is determined based on the length and / or signal strength of the target audio data. For example, different lengths and / or different signal strengths of the audio data have corresponding relationships with different block division lengths. Therefore, the block division length corresponding to the target audio data can be selected from this corresponding relationship based on the length and / or signal strength of the target audio data, and then the target audio data is divided based on this block division length to obtain M audio blocks.
[0244] Exemplarily, assume that the corresponding relationship between different lengths of audio data and different block division lengths is shown in Table 1:
[0245] Table 1
[0246] Length of audio data Block division length A1 3 A2 4 …… ……
[0247] For example, assume that the length of the target audio data is A2. Referring to Table 1 above, the block division length corresponding to the length A2 of the audio data is 4. Therefore, at the encoding end, according to the block division length of 4, every 4 consecutive audio frames in the target audio data are divided into one audio block to obtain M audio blocks.
[0248] In some embodiments, after the encoding end determines the lengths of the M audio blocks based on the above steps, the length information of the M audio blocks is written into the bitstream so that the decoding end can perform audio data decoding on the M audio blocks based on the length information of the audio blocks.
[0249] In one example, when the lengths of each of the M audio blocks are the same, the encoding end can write the length information of the audio blocks once into the bitstream. For example, the encoding end writes the length information of the audio blocks into the syntax information of the bitstream.
[0250] In one example, the encoding end can write the length information of each of the M audio blocks into the bitstream.
[0251] In one example, the encoding end can write the length information of the audio blocks with inconsistent lengths among the M audio blocks into the bitstream.
[0252] After the encoding end divides the target audio data to be encoded into M audio blocks based on the above steps, the following steps of S202 are executed.
[0253] S202: For the current audio block among the M audio blocks, transform the signal strength of the current audio block to obtain the transformed current audio block.
[0254] In the embodiments of the present application, the encoding end transforms the signal strength of each of the M audio blocks in the same way. For the sake of description, the current audio block among the M audio blocks is taken as an example for illustration here. The current audio block can be understood as any one of the M audio blocks.
[0255] As can be seen from the above, the current audio block in the embodiment of the present application includes multiple audio frames. When the signal strength changes, the signal strength of an entire current audio block is processed for transformation, making the signal strength smoother, thereby improving the encoding effect of audio data.
[0256] In some embodiments, the transformation of the signal strength of the current audio block includes transforming one or more of the dimensions describing the signal strength, such as the amplitude, energy, loudness, etc. of the current audio block.
[0257] The embodiment of the present application places no restrictions on the specific manner in which the encoding end transforms the signal strength of the current audio block to obtain the current audio block with the signal strength transformed.
[0258] In some embodiments, the encoding end may adopt a preset signal strength calculation method to calculate the average signal strength of the current audio block, and then, based on the average signal strength of the current audio block, perform normalization processing on the signal strength of the current audio block to obtain the current audio block with the signal strength transformed.
[0259] In some embodiments, the above S202 includes the following steps of S202-A and S202-B:
[0260] S202-A: Determine the transformation value of the signal strength of the current audio block based on the original audio data of the current audio block;
[0261] S202-B: Based on the transformation value, transform the signal strength of the current audio block pair to obtain the current audio block after transformation.
[0262] In this implementation manner, when the encoding end transforms the signal strength of the current audio block, it first determines the transformation value of the signal strength of the current audio block based on the original audio data of the current audio block, and then, based on the transformation value of the signal strength, transforms the signal strength of the audio data block to obtain the current audio block after transformation.
[0263] In some embodiments, the transformation value of the signal strength can be understood as the transformation value of the signal strength, for example, including the amplitude transformation value, the loudness transformation value (also known as the loudness gain value), the energy transformation value, etc.
[0264] The embodiment of the present application places no restrictions on the specific manner in which the encoding end determines the loudness gain value of the current audio block based on the original audio data of the current audio block.
[0265] In a possible implementation, based on a preset calculation formula for transformation values, signal strength values are calculated for the original audio data of the current audio block to obtain the second signal strength value of the current audio block. Then, based on the second signal strength value of the current audio block, the transformation value of the signal strength of the current audio block is determined. For example, a correspondence table between different signal strength values and different signal strength transformation values is set up, and this correspondence table can be obtained through calculation, experiment, or experience. Furthermore, the transformation value corresponding to the second signal strength value in this correspondence is determined as the transformation value of the signal strength of the current audio block.
[0266] In some embodiments, the signal strength of the current audio block includes at least one of the amplitude, energy, and loudness of the current audio block.
[0267] In a possible implementation, if the signal strength of the current audio block includes the loudness of the current audio block and the transformation value includes a loudness gain value, then the above S202-A includes the following steps S202-A1 to S202-A3:
[0268] S202-A1. Based on the original audio data of the current audio block, determine the first reference loudness value of the current audio block;
[0269] S202-A2. Determine the normalized target loudness value corresponding to the current audio block;
[0270] S202-A3. Based on the normalized target loudness value and the first reference loudness value of the current audio block, determine the loudness gain value of the current audio block.
[0271] In this implementation, when the encoding end performs loudness normalization processing on the current audio block, it first calculates the loudness based on the original audio data of the current audio block to obtain the first reference loudness value of the current audio block, and determines the normalized target loudness value corresponding to the current audio block. Furthermore, based on the first reference loudness value and the normalized target loudness value of the current audio block, the loudness gain value of the current audio block is determined.
[0272] It should be noted that there is no sequence preference between the above S202-A1 and the above S202-A2. That is to say, S202-A1 can be executed before the above S202-A2, or after S202-A2, or executed synchronously with the above S202-A2.
[0273] The embodiments of the present application do not limit the specific manner in which the encoding end determines the first reference loudness value of the current audio block based on the original audio data of the current audio block.
[0274] In a possible implementation, the encoding end adopts a preset loudness calculation method, such as the Moore and Zwicker loudness calculation model, to calculate the loudness based on the original audio data of the current audio block, obtaining a loudness value, and then recording this loudness value as the first reference loudness value of the current audio block.
[0275] In a possible implementation, the encoding end can determine the first reference loudness value of the current audio block through the following steps S202-A1-a and S202-A1-b:
[0276] S202-A12-a: For each audio frame included in the current audio block, determine the first reference loudness value of the audio frame based on the original audio data of the audio frame;
[0277] S202-A12-b: Based on the first reference loudness value of each audio frame, determine the first reference loudness value of the current audio block.
[0278] In the embodiments of the present application, in order to improve the smoothness of the loudness, multiple audio frames are divided into an audio block for overall loudness normalization processing. That is to say, the current audio block in the embodiments of the present application includes multiple audio frames, for example, including 2 or more audio frames. When the encoding end calculates the first reference loudness value of the current audio block, it determines the first reference loudness value of each audio frame included in the current audio block by calculating the loudness of each audio frame included in the current audio block, and then determines the first reference loudness value of the current audio block based on the first reference loudness value of each audio frame included in the current audio block.
[0279] The embodiments of the present application do not limit the specific manner in which the encoding end determines the first reference loudness value of each audio frame in the current audio block. For example, for each audio frame included in the current audio block, the encoding end adopts a preset loudness calculation method to calculate the loudness based on the original audio data of the audio frame, obtaining the first reference loudness value of the audio frame.
[0280] Exemplarily, the encoding end determines the effective loudness value of the audio frame based on a preset effective loudness calculation method and based on the original audio data of the audio frame, and then determines this effective loudness value as the first reference loudness value of the audio frame. The embodiments of the present application do not limit the specific manner of calculating the first reference loudness value of the audio frame.
[0281] After the encoding end determines the first reference loudness value of each audio frame in the current audio block, based on the first reference loudness value of each audio frame, it determines the first reference loudness value of the current audio block. For example, if the current audio block includes 3 audio frames, the average value or weighted average value of the first reference loudness values of these 3 audio frames is determined as the first reference loudness value of the current audio block. For another example, by sampling a preset data processing method to process the first reference loudness values of the 3 audio frames included in the current audio block, the first reference loudness value of the current audio block is obtained.
[0282] Further, the encoding end determines the normalized target loudness value corresponding to the current audio block.
[0283] In some embodiments, the above normalized target loudness value is a preset value. That is to say, in the embodiments of the present application, for the audio data to be encoded, the normalized target loudness values corresponding to the audio blocks included in each audio data are all the same. For example, they are all -16 dB.
[0284] In some embodiments, when the loudness of the audio data to be encoded changes, different normalized target loudness values can be set. Based on this, the above S202-A2 includes the following steps of S202-A21 and S202-A22:
[0285] S202-A21: Determine the loudness value of the target audio data;
[0286] S202-A22: Based on the loudness value of the target audio data, determine the normalized target loudness value corresponding to the current audio block.
[0287] In this implementation manner, the normalized target loudness values corresponding to audio data with different loudnesses. Based on this, the encoding end can calculate the loudness value of the target audio data to be encoded. For example, by using a preset loudness calculation method, the loudness value of the target audio data is determined. Then, based on the loudness value of the target audio data, the normalized target loudness value corresponding to the target audio data is determined. Furthermore, the normalized target loudness value corresponding to the target audio data is determined as the normalized target loudness value corresponding to each of the M audio blocks obtained by dividing the target audio data. In this way, the normalized target loudness value corresponding to the current audio block can be obtained.
[0288] In one example, in the embodiments of the present application, different loudness value intervals correspond to different normalized target loudness values. Exemplarily, as shown in Table 2:
[0289] Table 2
[0290] Loudness value range Normalized target loudness value [x1, x4] C1 [x5, x10] C2 …… ……
[0291] As shown in Table 2, the normalized target loudness value corresponding to the loudness value range [x1, x4] is C1, the normalized target loudness value corresponding to the loudness value range [x5, x10] is C2, and so on. In this way, the encoding end can look up Table 2 based on the determined loudness value of the target audio data, and obtain the normalized target loudness value corresponding to the current audio block. Suppose the calculated loudness value of the target audio data at the encoding end is x6, and x6 is in the loudness value range [x5, x10]. The normalized target loudness value corresponding to the loudness value range [x5, x10] in Table 2 is C2. Therefore, it can be determined that the normalized target loudness value corresponding to the current audio block is C2.
[0292] In one example, in the embodiments of the present application, different loudness values correspond to different normalized target loudness values. Exemplarily, as shown in Table 3:
[0293] Table 3
[0294] Loudness value Normalized target loudness value x1 C1 x2 C2 x3 C3 …… ……
[0295] As shown in Table 3, the normalized target loudness value corresponding to the loudness value x1 is C1, the normalized target loudness value corresponding to the loudness value x2 is C2, and so on. In this way, the encoding end can look up Table 3 based on the determined loudness value of the target audio data, and obtain the normalized target loudness value corresponding to the current audio block. Suppose the calculated loudness value of the target audio data at the encoding end is x3, and the normalized target loudness value corresponding to the loudness value x3 in Table 3 is C3. Therefore, it can be determined that the normalized target loudness value corresponding to the current audio block is C3.
[0296] After determining the first reference loudness value and the normalized target loudness value of the current audio block based on the above steps, the encoding end executes the steps of S202-A3 above, and determines the loudness gain value of the current audio block based on the first reference loudness value and the normalized target loudness value of the current audio block.
[0297] As can be seen from the above, the first reference loudness value of the current audio block can be understood as the loudness value before the loudness of the current audio block is not normalized, and the normalized target loudness value can be understood as the target loudness value after the loudness of the current audio block is normalized. In this way, the encoding end can determine the loudness gain value of the current audio block based on the first reference loudness value and the normalized target loudness value of the current audio block.
[0298] The embodiments of the present application do not limit the specific manner of determining the loudness gain value of the current audio block based on the first reference loudness value of the current audio block and the normalized target loudness value corresponding to the current audio block.
[0299] In one example, the encoding end determines the difference between the first reference loudness value of the current audio block and the normalized target loudness value corresponding to the current audio block as the loudness gain value of the current audio block.
[0300] In one example, the encoding end determines the loudness gain value of the current audio block through the following steps of S202 - A31 and S202 - A32:
[0301] S202 - A31: Subtract the normalized target loudness value corresponding to the current audio block from the first reference loudness value of the current audio block to obtain a second difference;
[0302] S202 - A32: Based on the second difference and a second preset value, obtain the loudness gain value of the current audio block.
[0303] In this implementation, the encoding end subtracts the normalized target loudness value corresponding to the current audio block from the first reference loudness value of the current audio block to obtain a difference, denoted as the second difference. Then, based on this second difference and the second preset value, the loudness gain value of the current audio block is determined.
[0304] The embodiments of the present application do not limit the specific manner in which the encoding end determines the loudness gain value of the current audio block based on the second difference and the second preset value.
[0305] For example, the encoding end determines the product of the second difference and the second preset value as the loudness gain value of the current audio block.
[0306] For another example, the encoding end multiplies the second difference by the second preset value to obtain a second product, and then performs a preset operation on this second product to obtain the loudness gain value of the current audio block.
[0307] The embodiments of the present application do not limit the specific manner of the preset operation.
[0308] In one example, the above - mentioned preset operation is a logarithmic operation. That is, the encoding end performs a logarithmic operation on the second product obtained by multiplying the second difference and the second preset value to obtain the loudness gain value of the current audio block.
[0309] Exemplarily, the encoding end can determine the loudness gain value of the current audio block through the following formula (3):
[0310] gain _2= exp((tar_db - ref_db) * GAIN_FACTOR_2) (3)
[0311] Among them, gain_2 is the loudness gain value of the current audio frame, exp is the exponential operation with e as the base, tar_db is the normalized target loudness value corresponding to the current audio block, ref_db is the first reference loudness value of the current audio block, (tar_db - ref_db) is the second difference, GAIN_FACTOR_2 is the second preset value, which is a constant. Optionally, the above second preset value may be equal to the above first preset value. (tar_db - ref_db) * GAIN_FACTOR_2 is the second product.
[0312] It should be noted that in addition to using the above method to determine the loudness gain value of the current audio block, the encoding end can also use other methods to determine the loudness gain value of the current audio block. The embodiments of the present application do not limit this.
[0313] In some embodiments, the encoding end also writes the above determined first reference loudness value of the current audio block into the code stream. In this way, the decoding end obtains the first reference loudness value of the current audio block by decoding the code stream, and then determines the loudness anti-gain value of the current audio block based on the first reference loudness value.
[0314] In some embodiments, the encoding end can also directly write the above determined loudness gain value of the current audio block into the code stream. In this way, the decoding end can obtain the loudness gain value of the current audio block by decoding the code stream, and then determine the loudness anti-gain value of the current audio block based on the loudness gain value of the current audio block.
[0315] After the encoding end determines the transformation value of the signal strength of the current audio block based on the above steps, it executes the steps of S202-B above, and transforms the signal strength of the original audio data of the current audio block based on the transformation value of the current audio block to obtain the transformed current audio block.
[0316] In some embodiments, if the transformation of the signal strength of the current audio block includes normalizing the loudness of the current audio frame, the encoding end multiplies each element in the audio data of the current audio block by the above calculated loudness gain value respectively to obtain the current audio block after loudness normalization.
[0317] Exemplarily, the encoding end uses the following formula (4) to obtain the current audio block after loudness normalization:
[0318] d audio_data_normalize = audio_data * gain_2 (4)
[0319] Among them, d audio_data_normalize is the current audio block after loudness normalization, and audio_data is the original audio data of the current audio block.
[0320] The above embodiments introduce the specific process of the encoding end transforming the signal strength of the original audio data of the current audio block to obtain the transformed current audio block. The encoding end can refer to the above method to transform the signal strength of each of the M audio blocks obtained by dividing the target audio data, so as to obtain each audio block after the signal strength transformation.
[0321] Based on the above steps, after the encoding end transforms the signal strength of the current audio block to obtain the transformed current audio block, it executes the following step S203.
[0322] S203: Encode the transformed current audio block to obtain a bitstream.
[0323] In the embodiments of the present application, after the encoding end obtains the transformed current audio block based on the above steps, it inputs the transformed current audio block into an encoding network for non-linear transformation, such as downsampling and feature transformation processing, to obtain the encoding vector of the current audio block. Then, the encoding vector of the current audio block is quantized to obtain the quantization result of the current audio block, and then the quantization result is encoded to obtain a bitstream.
[0324] The embodiments of the present application do not limit the specific network structure of the encoding network. For example Figure 3 As shown, the encoding network includes 1 input layer, 4 encoding feature processing modules, and 1 output layer. The encoding end processes the features of the transformed current audio block through the input layer to obtain the feature information 1 of the transformed current audio block. Then, the feature information 1 is respectively input into the first encoding feature processing module for processing, such as downsampling and / or feature transformation processing, to obtain the feature information 2 of the transformed current audio block. Then, the feature information 2 is input into the second encoding feature processing module for downsampling and / or feature transformation processing, to obtain the feature information 3 of the transformed current audio block. Then, the feature information 3 is input into the third encoding feature processing module for downsampling and / or feature transformation processing, to obtain the feature information 4 of the transformed current audio block. Finally, the feature information 4 is input into the output layer for processing, to obtain the feature information 5 of the transformed current audio block. Finally, the feature information 5 is input into the output layer, and the feature vector 5 is converted into an encoding vector with a preset dimension (such as k) and a preset number of channels (such as 1).
[0325] Then, the encoding end quantizes the encoding vector of the current audio block to obtain a quantization result, and encodes the quantization result to obtain a bitstream.
[0326] The embodiments of the present application do not limit the specific manner in which the encoding end quantizes the encoding vector of the current audio block.
[0327] In some embodiments, the encoding end samples a preset quantizer to quantize the encoding vector of the current audio block, and obtains the quantization result of the current audio block.
[0328] In some embodiments, when quantizing the encoding vector of the current audio block, the encoding end may determine the statistical value of the effective audio information of the current audio block, and then, based on the statistical value of the effective audio information of the current audio block, select the target quantizer of the current audio block, and then use the target quantizer to quantize the encoding vector of the current audio block to obtain a quantization result, and encode the quantization result to obtain a bitstream.
[0329] In some embodiments, if the quantizer adopted in the embodiments of the present application is a Figure 4 residual-based vector quantizer as shown, the encoding end inputs the encoding vector of the current audio block into the first quantization layer of the target quantizer for quantization to obtain a first quantization vector and the codebook index corresponding to the first quantization layer; based on the encoding vector and the first quantization vector of the current audio block, obtain a first residual vector; input the first residual vector into the second quantization layer of the target quantizer for quantization, and repeat the iteration to obtain the residual vector corresponding to the last quantization layer of the target quantizer and the codebook index corresponding to each quantization layer in the target quantizer; encode the residual vector corresponding to the last quantization layer of the target quantizer and the codebook index corresponding to each quantization layer in the target quantizer to obtain a bitstream.
[0330] Exemplarily, as Figure 8 shown, assuming that the target quantizer includes 3 quantization layers, the encoding end inputs the encoding vector y0 of the current audio block into the first quantization layer of the target quantizer for quantization, searches for the codebook vector matching the encoding vector y0 among the codebook vectors included in the codebook 1 corresponding to the first quantization layer as the output vector y1 of the first quantization layer, and at the same time, records the codebook index index0 of the vector y1 in the codebook 1 corresponding to the first quantization layer. Take the difference between the encoding vector y0 and the vector y1 as the residual vector Then input the residual vector into the second quantization layer of the target quantizer for quantization, search for the codebook vector matching the residual vector among the codebook vectors included in the codebook 2 corresponding to the second quantization layer as the output vector y2 of the second quantization layer, and at the same time, records the codebook index index1 of the vector y2 in the codebook 2 corresponding to the second quantization layer. Then, take the difference between the residual vector and the vector y2 as the residual vector Then input the residual vector into the third quantization layer of the target quantizer for quantization, search for the codebook vector matching the residual vector The matched codebook vector is used as the output vector y3 of the third quantization layer. Meanwhile, the codebook index index2 corresponding to the vector y3 in the codebook 3 of the third quantization layer is recorded. Then, the difference between the residual vector and the vector y3 is used as the residual vector
[0331] As can be seen from the above, in the embodiment of the present application, the target quantizer is used to quantize the encoding vector of the current audio block, and the final quantization result obtained at least includes the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook index corresponding to each quantization layer. For example, as Figure 8 shown, the final quantization result includes the residual vector corresponding to the last quantization layer of the target quantizer and the codebook indexes index0, index1, and index2 corresponding to each quantization layer.
[0332] Then, the encoding end encodes the above quantization result to obtain a bitstream. For example, the encoding end encodes the residual vector corresponding to the last quantization layer of the target quantizer and the codebook indexes index0, index1, and index2 corresponding to each quantization layer to form a binary bitstream.
[0333] Taking the encoding process of the current audio block in the M audio blocks divided from the target audio data as an example for illustration, the encoding processes of other audio blocks can refer to the encoding process of the current audio block above, so as to realize the complete encoding of the target audio data.
[0334] In the audio encoding method provided by the embodiment of the present application, the encoding end performs block division on the target audio data to be encoded to obtain M audio blocks, where M is a positive integer greater than 1, and the audio block includes two or more audio frames; for the current audio block in the M audio blocks, the signal strength of the current audio block is transformed to obtain the current audio block after signal strength transformation; the transformed current audio block is encoded to obtain a bitstream. That is to say, in the embodiment of the present application, two or more audio frames are divided into one audio block, and then the signal strength is transformed in units of audio blocks, so that the smoothing processing of the audio signal strength can be realized. Furthermore, encoding the audio block after signal strength smoothing processing can improve the encoding effect of the audio data.
[0335] Above, in combination with Figures 5 to 8 , the embodiments of the audio encoding and decoding method of the present application are described in detail. Below, in combination with Figures 9 to 10 , the embodiments of the device of the present application are described in detail.
[0336] Figure 9It is a schematic block diagram of an audio decoding device provided by an embodiment of the present application. The device 10 can be applied to a decoding device.
[0337] As Figure 9 shown, the audio decoding device 10 includes:
[0338] A decoding unit 11, configured to decode a bitstream to obtain a currently reconstructed audio block of the decoding network, where the currently reconstructed audio block is one of M audio blocks obtained by partitioning target audio data, M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0339] An inverse transformation unit 12, configured to perform an inverse transformation on the signal strength of the reconstructed currently reconstructed audio block to obtain a reconstructed value of the currently reconstructed audio block.
[0340] In some embodiments, the transformation unit 12 is specifically configured to determine an inverse transformation value of the signal strength of the currently reconstructed audio block; based on the inverse transformation value, perform an inverse transformation on the signal strength of the reconstructed currently reconstructed audio block to obtain the currently reconstructed audio block.
[0341] In some embodiments, the transformation unit 12 is specifically configured to determine an inverse transformation value of the currently reconstructed audio block based on the currently reconstructed audio block.
[0342] In some embodiments, the signal strength of the currently reconstructed audio block includes at least one of the amplitude, energy, and loudness of the currently reconstructed audio block.
[0343] In some embodiments, if the signal strength of the currently reconstructed audio frame includes the loudness of the currently reconstructed audio block, and the inverse transformation value includes a loudness inverse gain value, then the transformation unit 12 is specifically configured to decode the bitstream to obtain a first reference loudness value of the currently reconstructed audio block, where the first reference loudness value is calculated based on the original audio data of the currently reconstructed audio block; determine a second reference loudness value of the currently reconstructed audio block based on the currently reconstructed audio block; determine the loudness inverse gain value based on the first reference loudness value and the second reference loudness value of the currently reconstructed audio block.
[0344] In some embodiments, the transformation unit 12 is specifically configured to, for each audio frame included in the currently reconstructed audio block, determine a second reference loudness value of the audio frame based on the reconstructed audio data after loudness normalization of the audio frame; determine a second reference loudness value of the currently reconstructed audio block based on the second reference loudness value of each audio frame.
[0345] In some embodiments, the transformation unit 12 is specifically configured to subtract the first reference loudness value of the current audio block from the second reference loudness value of the current audio block to obtain a first difference value; and obtain the loudness anti-gain value based on the first difference value and a first preset value.
[0346] In some embodiments, the transformation unit 12 is specifically configured to multiply the first difference value by the first preset value to obtain a first product; and perform a preset operation on the first product to obtain the loudness anti-gain value.
[0347] In some embodiments, the transformation unit 12 is specifically configured to multiply each element in the reconstructed current audio block by the loudness anti-gain value to obtain the reconstructed value of the current audio block.
[0348] In some embodiments, the transformation unit 12 is specifically configured to determine the length information of the current audio block; and decode the code stream corresponding to the current audio block based on the length information of the current audio block to obtain the reconstructed current audio block.
[0349] In some embodiments, the transformation unit 12 is specifically configured to decode the code stream to obtain the length information of the current audio frame.
[0350] In some embodiments, the M audio blocks are obtained by performing block division on a plurality of audio frames included in the target audio data based on the audio block division length determined according to the length of the target audio data.
[0351] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 9 the illustrated apparatus can execute the embodiments of the above audio decoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, they will not be elaborated here.
[0352] Figure 10 is a schematic block diagram of an audio encoding apparatus provided in an embodiment of the present application. The apparatus 20 can be applied to an encoding device.
[0353] As Figure 10 shown, the audio encoding apparatus 20 includes:
[0354] a block division unit 21, configured to perform block division on target audio data to be encoded to obtain M audio blocks, where M is a positive integer greater than 1, and the audio block includes two or more audio frames;
[0355] A transform unit 22, configured to transform the signal strength of the current audio block among the M audio blocks to obtain a transformed current audio block;
[0356] An encoding unit 23, configured to encode the transformed current audio block to obtain a bitstream.
[0357] In some embodiments, the transform unit 22 is specifically configured to determine a transform value of the signal strength of the current audio block based on the original audio data of the current audio block; and transform the signal strength of the current audio block based on the transform value to obtain the transformed current audio block.
[0358] In some embodiments, the signal strength of the current audio block includes at least one of the amplitude, energy, and loudness of the current audio block.
[0359] In some embodiments, if the signal strength of the current audio block includes the loudness of the current audio block and the transform value includes a loudness gain value, the transform unit 22 is specifically configured to determine a first reference loudness value of the current audio block based on the original audio data of the current audio block; determine a normalized target loudness value corresponding to the current audio block; and determine the loudness gain value of the current audio block based on the normalized target loudness value and the first reference loudness value of the current audio block.
[0360] In some embodiments, the transform unit 22 is specifically configured to, for each audio frame included in the current audio block, determine a first reference loudness value of the audio frame based on the original audio data of the audio frame; and determine the first reference loudness value of the current audio block based on the first reference loudness values of each audio frame.
[0361] In some embodiments, the transform unit 22 is specifically configured to determine a loudness value of the target audio data; and determine a normalized target loudness value corresponding to the current audio block based on the loudness value of the target audio data.
[0362] In some embodiments, the transform unit 22 is specifically configured to subtract the normalized target loudness value from the first reference loudness value of the current audio block to obtain a second difference; and obtain the loudness gain value based on the second difference and a second preset value.
[0363] In some embodiments, the transform unit 22 is specifically configured to multiply the second difference by the second preset value to obtain a second product; and perform a preset operation on the second product to obtain the loudness gain value.
[0364] In some embodiments, the encoding unit 23 is further configured to write the first reference loudness value of the current audio block into the bitstream.
[0365] In some embodiments, the transformation unit 22 is specifically configured to multiply the loudness value of each element of the current audio block by the loudness gain value to obtain the transformed current audio block.
[0366] In some embodiments, the encoding unit 23 is further configured to write the length information of the current audio block into the bitstream.
[0367] In some embodiments, the block partitioning unit 21 is specifically configured to determine the audio block partitioning length based on the length of the target audio data; and perform block partitioning on a plurality of audio frames included in the target audio data based on the audio block partitioning length to obtain the M audio blocks.
[0368] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, details are not described herein again. Specifically, Figure 10 The illustrated apparatus can execute the embodiments of the above audio encoding method, and the foregoing and other operations and / or functions of each module in the apparatus are respectively for implementing the above method embodiments. For the sake of brevity, details are not described herein again.
[0369] In the foregoing, the apparatus of the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or can be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0370] Figure 11 is a schematic block diagram of an electronic device provided by an embodiment of the present application. Figure 11 The electronic device can be the above encoding device or a decoding device.
[0371] Such as Figure 11 As shown, the electronic device 30 may include:
[0372] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transfer the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.
[0373] For example, the processor 32 can be used to execute the steps in the above method 200 according to the instructions in the computer program 33.
[0374] In some embodiments of the present application, the processor 32 may include but is not limited to:
[0375] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.
[0376] In some embodiments of the present application, the memory 31 includes but is not limited to:
[0377] A volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0378] In some embodiments of the present application, the computer program 33 may be divided into one or more modules. The one or more modules are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.
[0379] As Figure 11 shown, the electronic device 30 may further include:
[0380] A transceiver 34, which may be connected to the processor 32 or the memory 31.
[0381] Wherein, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.
[0382] It should be understood that the various components in the electronic device 30 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.
[0383] According to one aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or rather, the embodiments of the present application further provide a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.
[0384] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.
[0385] In other words, when implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0386] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0387] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.
[0388] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0389] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An audio decoding method, characterized in that: include: Decoding the bitstream to obtain a current audio block reconstructed by a decoding network, wherein the current audio block is one of M audio blocks obtained by dividing the target audio data into blocks, where M is a positive integer greater than 1, and the audio block includes 2 or more audio frames; Perform an inverse transformation on the signal strength of the reconstructed current audio block to obtain a reconstructed value of the current audio block.
2. The method according to claim 1, characterized in that The inverse transformation of the signal strength of the reconstructed current audio block to obtain the reconstructed value of the current audio block includes: Determine, based on the reconstructed current audio block, an inverse transformation value of a signal strength of the current audio block, where the signal strength of the current audio block includes at least one of an amplitude, energy, and loudness of the current audio block; Based on the inverse transformation value, an inverse transformation is performed on the signal strength of the reconstructed current audio block to obtain a reconstructed value of the current audio block.
3. The method according to claim 2, characterized in that If the signal strength of the current audio frame includes the loudness of the current audio block, and the inverse transformation value includes an inverse loudness gain value, then determining the inverse transformation value of the signal strength of the current audio block based on the reconstructed current audio block includes: Decoding the bitstream to obtain a first reference loudness value of the current audio block, where the first reference loudness value is calculated based on original audio data of the current audio block; Determining a second reference loudness value of the current audio block based on the reconstructed current audio block; The loudness inverse gain value is determined based on a first reference loudness value of the current audio block and a second reference loudness value of the current audio block.
4. The method according to claim 3, characterized in that The determining, based on the reconstructed current audio block, a second reference loudness value of the current audio block comprises: For each audio frame included in the current audio block, determining a second reference loudness value of the audio frame based on reconstructed audio data after loudness normalization of the audio frame; A second reference loudness value of the current audio block is determined based on the second reference loudness value of each audio frame.
5. The method according to claim 3, characterized in that: The determining the loudness inverse gain value based on the first reference loudness value of the current audio block and the second reference loudness value of the current audio block includes: Subtracting a first reference loudness value of the current audio block from a second reference loudness value of the current audio block to obtain a first difference value; The loudness inverse gain value is obtained based on the first difference and a first preset value.
6. The method according to claim 5, characterized in that The step of obtaining the loudness inverse gain value based on the first difference and a first preset value includes: multiplying the first difference by the first preset value to obtain a first product; A preset operation is performed on the first product to obtain the loudness inverse gain value.
7. The method according to any one of claims 3 to 6, characterized in that: The performing inverse transformation on the signal strength of the reconstructed current audio block based on the inverse transformation value to obtain the reconstructed value of the current audio block includes: Each element in the reconstructed current audio block is multiplied by the inverse loudness gain value to obtain a reconstructed value of the current audio block.
8. The method according to claim 1, characterized in that The decoding bitstream is used to obtain a current audio block reconstructed by a decoding network, including: Decoding the bit stream to determine the length information of the current audio block; Based on the length information of the current audio block, a bit stream corresponding to the current audio block is decoded to obtain the reconstructed current audio block.
9. The method according to claim 1, characterized in that: The M audio blocks are obtained by dividing a plurality of audio frames included in the target audio data into blocks based on an audio block division length determined based on a length of the target audio data.
10. An audio encoding method, characterized in that: include: Divide the target audio data to be encoded into blocks to obtain M audio blocks, where M is a positive integer greater than 1, and the audio block includes 2 or more audio frames; For a current audio block among the M audio blocks, transform a signal strength of the current audio block to obtain a transformed current audio block; The transformed current audio block is encoded to obtain a bit stream.
11. The method according to claim 10, characterized in that The transforming the signal strength of the current audio block to obtain a transformed current audio block includes: determining, based on original audio data of the current audio block, a transformed value of a signal strength of the current audio block, where the signal strength of the current audio block includes at least one of an amplitude, energy, and loudness of the current audio block; Based on the transformation value, the signal strength of the current audio block is transformed to obtain the transformed current audio block.
12. The method according to claim 11, characterized in that If the signal strength of the current audio block includes the loudness of the current audio block, and the transformation value includes a loudness gain value, then determining the transformation value of the signal strength of the current audio block based on the original audio data of the current audio block includes: Determine a first reference loudness value of the current audio block based on original audio data of the current audio block; Determining a normalized target loudness value corresponding to the current audio block; A loudness gain value of the current audio block is determined based on the normalized target loudness value and a first reference loudness value of the current audio block.
13. The method according to claim 12, characterized in that The determining, based on original audio data of the current audio block, a first reference loudness value of the current audio block includes: For each audio frame included in the current audio block, determining a first reference loudness value of the audio frame based on original audio data of the audio frame; Based on the first reference loudness value of each audio frame, a first reference loudness value of the current audio block is determined.
14. The method according to claim 12, characterized in that The determining a normalized target loudness value corresponding to the current audio block includes: Determining a loudness value of the target audio data; Based on the loudness value of the target audio data, a normalized target loudness value corresponding to the current audio block is determined.
15. The method according to claim 12, characterized in that The determining, based on the normalized target loudness value and the first reference loudness value of the current audio block, a loudness gain value of the current audio block comprises: Subtracting the normalized target loudness value from the first reference loudness value of the current audio block to obtain a second difference value; The loudness gain value is obtained based on the second difference and a second preset value.
16. The method according to claim 15, characterized in that The step of obtaining the loudness gain value based on the second difference and the second preset value includes: multiplying the second difference by the second preset value to obtain a second product; A preset operation is performed on the second product to obtain the loudness gain value.
17. The method according to claim 10, characterized in that The target audio data to be encoded is divided into blocks to obtain M audio blocks, including: Based on the length of the target audio data, determine the audio block division length; Based on the audio block division length, multiple audio frames included in the target audio data are divided into blocks to obtain the M audio blocks.
18. An audio decoding device, characterized in that: include: A decoding unit, configured to decode a bit stream to obtain a current audio block reconstructed by a decoding network, wherein the current audio block is one of M audio blocks obtained by dividing the target audio data into blocks, wherein M is a positive integer greater than 1, and the audio block includes two or more audio frames; The transform unit is configured to perform an inverse transform on the signal strength of the reconstructed current audio block to obtain a reconstructed value of the current audio block.
19. An audio encoding device, characterized in that: include: A block division unit, configured to divide the target audio data to be encoded into blocks to obtain M audio blocks, where M is a positive integer greater than 1, and the audio block includes 2 or more audio frames; a transforming unit, configured to transform, for a current audio block among the M audio blocks, a signal strength of the current audio block to obtain a transformed current audio block; The encoding unit is used to encode the transformed current audio block to obtain a code stream.
20. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1 to 9 or 10 to 17.