Audio encoding and decoding method, device and equipment
By counting the effective audio information of the audio data and selecting the target quantizer for quantization, the problem of high complexity of audio encoding and decoding in the prior art is solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202311472254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The existing end-to-end audio codec solutions have high encoding and decoding complexity, resulting in low efficiency.
By counting the effective audio information of the audio data, selecting the target quantizer, quantizing the encoded vector of the audio data, reducing the complexity of the quantization operation.
It improves the encoding and decoding efficiency of audio data, reduces complexity, and improves the encoding and decoding performance.
Smart Images

Figure CN119943067A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to an audio encoding and decoding method, apparatus, and device. Background Art
[0002] With the rapid development of deep learning technology, deep learning technology has been widely used in the processing technology of signals of different dimensions, such as audio, image and video. Taking audio signals as an example, in the end-to-end audio codec solution based on deep learning, the encoder maps the audio signal into a coding vector through the encoding network, and further generates the corresponding binary code stream file through quantization technology. The decoding end obtains the transmitted information by reading the binary code stream file, and obtains the corresponding coding vector through the inverse quantization technology, and then uses the coding vector as the input of the decoding network to decode and obtain the final reconstructed audio signal.
[0003] However, the current end-to-end audio coding and decoding solutions have high coding and decoding complexity, resulting in low audio coding and decoding efficiency. Summary of the invention
[0004] The present application provides an audio coding and decoding method, apparatus, and device, which can reduce the complexity of audio coding and decoding and improve the efficiency of audio coding and decoding.
[0005] In a first aspect, the present application provides an audio decoding method, comprising:
[0006] Decoding a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data by a target quantizer, wherein the coding vector is obtained by performing a nonlinear transformation on the audio data, and the target quantizer is selected based on a statistical value of effective audio information of the audio data;
[0007] De-quantizing the quantization result to obtain a reconstructed coding vector of the audio data;
[0008] The reconstructed coding vector is decoded to obtain a reconstructed value of the audio data.
[0009] In a second aspect, the present application provides an audio encoding method, comprising:
[0010] Acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data;
[0011] Performing statistics on the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and selecting a target quantizer based on the statistical value of the effective audio information;
[0012] The target quantizer is used to quantize the coding vector to obtain a quantization result, and the quantization result is encoded to obtain a code stream.
[0013] In a third aspect, the present application provides an audio decoding device, comprising:
[0014] A decoding unit, configured to decode a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data by a target quantizer, wherein the coding vector is obtained by performing a nonlinear transformation on the audio data, and the target quantizer is selected based on a statistical value of effective audio information of the audio data;
[0015] A dequantization unit, used for dequantizing the quantization result to obtain a reconstructed coding vector of the audio data;
[0016] The reconstruction unit is used to decode the reconstructed coding vector to obtain a reconstructed value of the audio data.
[0017] In some embodiments, the effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
[0018] In some embodiments, the target quantizer is determined based on a target quantization level of a quantizer corresponding to the audio data, and the target quantization level is determined based on the statistical value of the effective audio information.
[0019] In some embodiments, if the effective audio information statistical value is less than a first preset value, the target quantization layer number is the preset first quantization layer number; or, the target quantization layer number is the second quantization layer number corresponding to the effective audio information statistical value in the first corresponding relationship, and the first corresponding relationship includes the relationship between different effective audio information statistical value intervals and different quantization layer numbers; or, the target quantization layer number is determined based on the second corresponding relationship and the effective audio information statistical value, and the second corresponding relationship is the relationship between different effective audio information statistical values and different quantization layer numbers.
[0020] In some embodiments, the target quantizer is composed of quantization layers of the target number of quantization layers selected from a total number of quantization layers of a preset quantizer.
[0021] In some embodiments, among the target number of quantization layers selected from the total number of quantization layers of the preset quantizer, at least two quantization layers are non-adjacent quantization layers.
[0022] In some embodiments, the decoding unit is also used to decode the bit stream to obtain indication information of the target quantizer; the dequantization unit is specifically used to determine the target quantizer based on the indication information; and use the target quantizer to dequantize the quantization result to obtain a reconstructed coding vector of the audio data.
[0023] In some embodiments, the indication information includes index information of the quantization layer included in the target quantizer, and the inverse quantization unit is specifically used to select the corresponding quantization layer from the total quantization layers of the preset quantizer based on the index information of the quantization layer to form the target quantizer.
[0024] In some embodiments, the inverse quantization unit is specifically used to query in the codebook corresponding to each quantization layer of the target quantizer based on the codebook subscript corresponding to each quantization layer of the target quantizer included in the quantization result, to obtain the quantization vector of each quantization layer of the target quantizer; based on the residual vector corresponding to the last quantization layer of the target quantizer in the quantization result, and the quantization vector of each quantization layer of the target quantizer, obtain the reconstructed coding vector of the audio data.
[0025] In a fourth aspect, the present application provides an audio encoding device, including:
[0026] An acquisition unit, used to acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data;
[0027] a determination unit, configured to perform statistics on the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and select a target quantizer based on the statistical value of the effective audio information;
[0028] The quantization unit is used to quantize the coding vector using the target quantizer to obtain a quantization result, and encode the quantization result to obtain a code stream.
[0029] In some embodiments, the effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
[0030] In some embodiments, if the effective audio information statistics of the audio data include the energy value of the audio data, the determination unit is specifically used to determine the energy value of each audio frame in the audio data; determine the energy value of the audio data based on the energy value of each audio frame; and determine the effective audio information statistics of the audio data based on the energy value of the audio data.
[0031] In some embodiments, the determination unit is specifically configured to determine, for an i-th audio frame in the audio data, a sum of squares of sample data included in the i-th audio frame as an energy value of the i-th audio frame, where i is a positive integer.
[0032] In some embodiments, if the effective audio information statistics of the audio data include the spectrum envelope statistics of the audio data, the determination unit is specifically used to determine the spectrum envelope statistics of each audio frame in the audio data; determine the spectrum envelope statistics of the audio data based on the spectrum envelope statistics of each audio frame; and determine the effective audio information statistics of the audio data based on the spectrum envelope statistics of the audio data.
[0033] In some embodiments, the determination unit is specifically used to convert the i-th audio frame in the audio data from a time domain signal to a frequency domain signal, where i is a positive integer; and perform statistics on the spectrum envelope of the frequency domain signal of the i-th audio frame to obtain the spectrum envelope statistics of the i-th audio frame.
[0034] In some embodiments, if the effective audio information statistics of the audio data include the zero-crossing rate of the audio data, the determination unit is specifically used to determine the zero-crossing rate of each audio frame in the audio data; determine the zero-crossing rate of the audio data based on the zero-crossing rate of each audio frame; and determine the effective audio information statistics of the audio data based on the zero-crossing rate of the audio data.
[0035] In some embodiments, the determination unit is specifically used to determine the target quantization level of the quantizer corresponding to the audio data based on the statistical value of the effective audio information; and select the target quantizer based on the target quantization level.
[0036] In some embodiments, the determination unit is specifically used to determine a preset first quantization layer number as the target quantization layer number if the effective audio information statistical value is less than a first preset value; or, obtain a first correspondence between different effective audio information statistical value intervals and different quantization layer numbers, and determine the second quantization layer number corresponding to the effective audio information statistical value in the first correspondence as the target quantization layer number; or, obtain a second correspondence between different effective audio information statistical values and different quantization layer numbers, and determine the target quantization layer number based on the second correspondence and the effective audio information statistical value.
[0037] In some embodiments, the determination unit is specifically configured to select quantization layers of the target number of quantization layers from a total number of quantization layers of a preset quantizer to form the target quantizer.
[0038] In some embodiments, among the quantization layers of the target number of quantization layers selected from the total quantization layers of the preset quantizer, at least two quantization layers are non-adjacent quantization layers.
[0039] In some embodiments, the quantization unit is specifically used to input the encoding vector of the audio data into the first quantization layer of the target quantizer for quantization, and obtain a first quantization vector and a codebook subscript corresponding to the first quantization layer; based on the encoding vector of the audio data and the first quantization vector, obtain a first residual vector; input the first residual vector into the second quantization layer of the target quantizer for quantization, repeat iterations, and obtain the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook subscript corresponding to each quantization layer in the target quantizer; encode the residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer to obtain the code stream.
[0040] In some embodiments, the bitstream also includes indication information for indicating the target quantizer.
[0041] In some embodiments, the indication information includes index information of a quantization layer included in the target quantizer.
[0042] In a fifth aspect, an electronic device is provided, comprising a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in any one of the first to second aspects or their implementations.
[0043] In a sixth aspect, a chip is provided for implementing the method in any one of the first to second aspects or their respective implementations. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method in any one of the first to second aspects or their respective implementations.
[0044] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method in any one of the first to second aspects above or in each of their implementations.
[0045] In an eighth aspect, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions enable a computer to execute the method in any one of the first to second aspects above or in each of their implementations.
[0046] In a ninth aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or in each of their implementations.
[0047] In summary, the present application obtains the effective audio information statistics of the audio data by counting the effective audio information of the audio data, and then selects the target quantizer based on the effective audio information statistics, and finally uses the target quantizer to quantize the encoding vector of the audio data. In other words, the embodiment of the present application adaptively selects the most suitable quantizer based on the characteristics of the input audio data (i.e., the effective audio information statistics), avoids multiple quantization operations on the data of the low audio signal, thereby reducing the complexity of the quantization operation and improving the encoding and decoding efficiency of the audio data. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 A schematic block diagram of an audio codec system according to an embodiment of the present application;
[0050] Figure 2 A schematic diagram of an end-to-end audio codec system based on deep learning involved in an embodiment of the present application;
[0051] Figure 3 A network structure block diagram of a codec constructed based on a convolutional neural network in one embodiment of the present application;
[0052] Figure 4 A schematic diagram of a residual-based vector quantizer according to an embodiment of the present application;
[0053] Figure 5 A flowchart of an audio encoding method provided in one embodiment of the present application;
[0054] Figure 6 is a schematic diagram of a coding network;
[0055] Figure 7 A schematic diagram of selecting a target quantizer;
[0056] Figure 8 A schematic diagram of quantization using a target quantizer;
[0057] Fig. 9 A flowchart of an audio decoding method provided in an embodiment of the present application;
[0058] Fig.10 It is a schematic diagram of dequantization;
[0059] Fig.11 is a schematic block diagram of an audio encoding device provided by an embodiment of the present application;
[0060] Fig.12 is a schematic block diagram of an audio decoding device provided by an embodiment of the present application;
[0061] Fig.13 It is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0063] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In an embodiment of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units does not have to be limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "multiple" refers to two or more than two.
[0064] The technical solution proposed in this application can be applied to technical fields such as artificial intelligence and audio coding and decoding, and can be used to reduce the complexity of audio coding and decoding while ensuring the performance of audio and video coding and decoding, thereby improving the efficiency of audio coding and decoding.
[0065] The following is an introduction to the relevant concepts involved in the embodiments of the present application.
[0066] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0067] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0068] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0069] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0070] The embodiments of this application mainly introduce the application of artificial intelligence technology in audio coding and decoding technology.
[0071] Audio encoding and decoding: The audio encoding process is to compress the audio into smaller data, and the decoding process is to restore the smaller data to audio. The encoded smaller data is used for network transmission and occupies less bandwidth.
[0072] Audio sampling rate: The audio sampling rate describes the number of data contained in a unit of time (1 second). For example, an 8k sampling rate contains 8000 sampling points, and each sampling point corresponds to a short integer.
[0073] Codebook: A collection of multiple vectors. The encoder and decoder both store the same codebook.
[0074] Quantization: Find the closest vector in the codebook for the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.
[0075] Quantizer: The quantizer is responsible for quantization and updating the vectors in the codebook.
[0076] Audio frame: Indicates the minimum duration of voice for a single transmission in the network.
[0077] Short Time Fourier Transform: STFT. It divides a long signal into several shorter signals of equal length, and then calculates the Fourier transform of each shorter segment. It is usually used to describe the changes in the frequency domain and time domain, and is an important tool in time-frequency analysis.
[0078] The audio coding method provided in the embodiments of the present application can be applied to the field of audio coding, the field of hardware audio coding, the field of dedicated circuit video coding, the field of real-time audio coding, etc. For example, the scheme of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard. Alternatively, the scheme of the present application can be combined with other exclusive or industry standards. It should be understood that the technology of the present application is not limited to any specific coding standard or technology.
[0079] The audio coding and decoding method provided in the embodiments of the present application can be applied to any end-to-end audio coding and decoding solution based on deep learning.
[0080] To facilitate understanding, first combine Figure 1 The audio codec system involved in the embodiment of the present application is introduced.
[0081] Figure 1 This is a schematic block diagram of an audio codec system involved in an embodiment of the present application. It should be noted that: Figure 1 This is just an example. The audio codec system of the embodiment of the present application includes but is not limited to Figure 1 As shown. Figure 1As shown, the audio codec system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compression) the audio data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded audio data.
[0082] The encoding device 110 of the embodiment of the present application can be understood as a device with an audio encoding function, and the decoding device 120 can be understood as a device with an audio decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, playback devices, digital media players, audio game consoles, car computers, etc.
[0083] In some embodiments, the encoding device 110 may transmit the encoded audio data (eg, a code stream) to the decoding device 120 via the channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded audio data from the encoding device 110 to the decoding device 120.
[0084] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded audio data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded audio data according to the communication standard and transmit the modulated audio data to the decoding device 120. The communication medium includes a wireless communication medium, such as a radio frequency spectrum, and optionally, the communication medium may also include a wired communication medium, such as one or more physical transmission lines.
[0085] In another example, the channel 130 includes a storage medium, which can store the audio data encoded by the encoding device 110. The storage medium includes a variety of locally accessible data storage media, such as optical disks, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded audio data from the storage medium.
[0086] In another example, the channel 130 may include a storage server that can store the audio data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded audio data from the storage server. Optionally, the storage server can store the encoded audio data and can transmit the encoded audio data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.
[0087] In some embodiments, the encoding device 110 includes an audio encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0088] In some embodiments, the encoding device 110 may further include an audio source 111 in addition to the audio encoder 112 and the input interface 113 .
[0089] The audio source 111 may include at least one of an audio acquisition device (eg, a microphone), an audio archive, an audio input interface, and a computer voice system, wherein the audio input interface is used to receive audio data from an audio content provider, and the computer voice system is used to generate audio data.
[0090] The audio encoder 112 encodes the audio data from the audio source 111 to generate a bitstream. The bitstream contains the encoding information of the audio data in the form of a bitstream. The encoding information may include the encoded audio data and associated data. The associated data may include quantization parameters and other syntax structures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0091] The audio encoder 112 transmits the encoded audio data directly to the decoding device 120 via the output interface 113. The encoded audio data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.
[0092] In some embodiments, the decoding device 120 includes an input interface 121 and an audio decoder 122 .
[0093] In some embodiments, the decoding device 120 may include a playback device 123 in addition to the input interface 121 and the audio decoder 122 .
[0094] The input interface 121 includes a receiver and / or a modem. The input interface 121 can receive the encoded audio data through the channel 130 .
[0095] The audio decoder 122 is used to decode the encoded audio data to obtain decoded audio data, and transmit the decoded audio data to the playback device 123 .
[0096] The playback device 123 plays the decoded audio data. The playback device 123 may be integrated with the decoding device 120 or may be external to the decoding device 120. The playback device 123 may include a variety of playback devices.
[0097] also, Figure 1 This is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1For example, the technology of the present application can also be applied to single-sided audio encoding or single-sided audio decoding.
[0098] Figure 2 Schematic diagram of an end-to-end audio codec system based on deep learning involved in an embodiment of the present application. Figure 2 As shown, the audio codec system of the embodiment of the present application includes: an encoding network 210, a quantization module 211, an inverse quantization module 212 and a decoding network 213.
[0099] During encoding, the encoding end (also called the transmitting end) will first input the input audio data into the encoding network 210 for nonlinear transformation to obtain the encoding vector (also called embedded sequence or hidden variable, etc.) of the input audio data. Then, the encoding vector of the audio data is quantized by the quantization module 211 to obtain the quantization result of the encoding vector. For example, a residual-based vector quantizer is used to select the corresponding quantization parameter according to the target bit rate. Finally, the quantized encoding vector is encoded and converted into a binary code stream.
[0100] During decoding, the decoding end (also called the receiving end) first recovers the quantization result of the coding vector from the bit stream, and then further recovers the coding vector through the inverse quantization module 212, and inputs it into the decoding network 213 for nonlinear transformation to obtain reconstructed audio data.
[0101] Figure 3 This is a block diagram of the network structure of a codec built based on a convolutional neural network in one embodiment of the present application.
[0102] like Figure 3 As shown, the network structure of the codec includes an encoding network 310 and a decoding network 320, wherein the encoding network 310 can be implemented as software as shown in FIG. Figure 1 The audio and video encoding device 110 and the decoding network 320 shown in FIG. 3 can be implemented as software as shown in FIG. Figure 1 The audio and video decoding device 120 is shown. In some embodiments, the encoding network 310 is also referred to as an encoder 310 , and the decoding network 320 is also referred to as a decoder 320 .
[0103] At the data transmission end, the audio data may be encoded and compressed through the encoding network 310. In an embodiment of the present application, the encoding network 310 may include an input layer 311, one or more encoding modules 312 and an output layer 313.
[0104] Exemplarily, the input layer 311 and the output layer 313 may be convolutional layers constructed based on a one-dimensional convolution kernel, and multiple (e.g., 4) encoding modules (EncoderBlock) 312 are sequentially connected between the input layer 311 and the output layer 313. Each encoding module 312 includes multiple residual (ResidualUnit) modules, and each residual module includes multiple convolutional layers.
[0105] For example, at the input stage of the encoder, the original audio data to be encoded is sampled to obtain a vector with c channels and w dimensions; the vector is input to the input layer 311, and after convolution processing, a feature vector with 32c channels and w dimensions can be obtained. In some optional implementations, in order to improve encoding efficiency, the encoding network 310 can encode a batch of audio vectors at the same time.
[0106] In the downsampling stage of the encoder, the first encoding module reduces the vector dimension to 1 / 2 and increases the number of channels by 2 times, obtaining a feature vector with 64 channels and 1 / 2w dimension; the second encoding module reduces the vector dimension to 1 / 4 and increases the number of channels by 2 times, obtaining a feature vector with 128 channels and 1 / 8w dimension; the third encoding module reduces the vector dimension to 1 / 5 and increases the number of channels by 2 times, obtaining a feature vector with 256 channels and 1 / 40w dimension; the fourth encoding module reduces the vector dimension to 1 / 8 and increases the number of channels by 2 times, obtaining a feature vector with 512 channels and 1 / 320w dimension.
[0107] In the output stage of the encoder, the output layer 313 performs convolution processing on the feature vector with a channel number of 512c and a dimension of 1 / 320w to obtain a coding vector with a channel number of 1 and a dimension of K.
[0108] The coding vector is input to the quantizer 330, and the codebook index corresponding to the coding vector can be queried in the codebook, and the codebook index is encoded to obtain a binary code stream, which is then sent to the data receiving end.
[0109] The data receiving end decodes the received binary code stream to obtain a codebook index, and performs inverse quantization based on the codebook index to obtain a reconstructed coding vector. Finally, the reconstructed coding vector is decoded through a decoding network 320 to obtain restored audio data.
[0110] In one embodiment of the present application, the decoding network 320 may include an input layer 321, one or more decoding modules 322, and an output layer 323. Each decoding module 322 includes a plurality of residual units, and each residual unit includes a plurality of convolutional layers.
[0111] After the data receiving end decodes the code stream to obtain the codebook index, the data receiving end may first query the codebook vector corresponding to the codebook index in the codebook through the quantizer 320, and then obtain the encoding vector reconstructed by the audio data based on the codebook vector. For example, the reconstructed encoding vector may be a vector with a channel number of 1 and a dimension of K. In some optional implementations, in order to improve decoding efficiency, the data receiving end may decode a batch of codebook vectors at the same time.
[0112] In the input stage of the decoder, the reconstructed coding vector is input to the input layer 321, and after convolution processing, a feature vector with a channel number of 512c and a dimension of 1 / 320w can be obtained.
[0113] In the decoding stage of the decoder, the first decoding module increases the vector dimension to 8 times and reduces the number of channels by 2 times, obtaining a feature vector with 256 channels and 1 / 40w dimension; the second decoding module increases the vector dimension to 5 times and reduces the number of channels by 2 times, obtaining a feature vector with 128 channels and 1 / 8w dimension; the third decoding module increases the vector dimension to 4 times and reduces the number of channels by 2 times, obtaining a feature vector with 64 channels and 1 / 2w dimension; the fourth decoding module increases the vector dimension to 2 times and reduces the number of channels by 2 times, obtaining a feature vector with 32 channels and w dimension.
[0114] In the output stage of the decoder, the output layer 323 performs convolution processing on the feature vector with 32 channels and a dimension of w, and restores the reconstructed audio data with 1 channel and a dimension of w.
[0115] In some embodiments, in order to improve the audio coding effect, when quantizing the coding vector of the audio data, a residual-based vector quantizer is used, that is, the above-mentioned Figure 3 The quantizer 330 in is a residual-based vector quantizer.
[0116] In one example, if Figure 4 As shown in the figure, the residual-based vector quantizer contains N vector quantization layers (VQ), each of which maintains a codebook of length L. The length of each codeword is M, that is, a codebook includes L codebook vectors, and the length of each codebook vector is M. Its specific operation is: the encoding vector x∈R S×D As input (S is the number of frames, D is the dimension of the feature). Assume that the i-th frame feature of x is recorded as y0∈R 1×D , will first pass through the first vector quantization layer, and query the codebook corresponding to the first vector quantization layer with y0∈R 1×D The codebook vector closest to the nearest distance obtains the corresponding vector quantization result And the corresponding codebook index index_i0, and calculate the quantized residual As the input of the next vector quantization layer. Then iterate N-1 times the same quantization and residual calculation operations until all N vector quantization layers are traversed, and output the codebook index index_ij corresponding to each layer, j = 1, ... N, and write it into the bit stream after binarization (such as 7 can be represented as 0 ... 111, the total number of bits is log2 (L)). The quantization process ends after processing all frames of x, and outputs the bit stream to be transmitted.
[0117] The decoder decodes the bitstream and reads the corresponding codebook index for each frame. ij ,i=1,…S,j=1,…N, recover the codebook vector from the codebook. Specifically, for the i-th frame, assume that the information recovered at each layer is Then sum the results of each layer to get the reconstructed coding vector of the i-th frame
[0118] In the current end-to-end audio coding and decoding scheme, when quantizing the encoding vector of audio data, the characteristics of the audio data itself are not taken into consideration. Instead, the same quantization operation is used uniformly for quantization, which introduces additional computational complexity during quantization and consumes additional bit rate, thereby making the audio data coding and decoding complexity high and the coding and decoding efficiency low.
[0119] In order to solve the above technical problems, the embodiment of the present application proposes an adaptive quantization method for end-to-end audio compression. Specifically, the effective audio information of the audio data is counted to obtain the effective audio information statistics of the audio data, and then based on the effective audio information statistics, a target quantizer is selected, and finally the target quantizer is used to quantize the encoding vector of the audio data. In other words, the embodiment of the present application adaptively selects the most suitable quantizer based on the characteristics of the input audio data (i.e., the effective audio information statistics), avoids multiple quantization operations on the data of the low audio signal, and thus can reduce the complexity of the quantization operation, can improve the encoding efficiency of the audio data, and effectively control the complexity of decoding.
[0120] The technical solutions of the embodiments of the present application are described in detail below through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0121] First, taking the encoding end as an example, the audio encoding method provided in the embodiment of the present application is introduced.
[0122] Figure 51 is a flow chart of an audio encoding method provided in an embodiment of the present application. The execution subject of the embodiment of the present application is a device having an audio encoding function, such as an audio encoding device. In some embodiments, the audio encoding device can be Figure 1 For ease of description, the embodiment of the present application is described by taking the execution subject as an encoding device as an example.
[0123] like Figure 5 As shown, the audio encoding method of the embodiment of the present application includes:
[0124] S101. Acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data.
[0125] In the embodiment of the present application, the audio data to be decoded may be a segment of audio data of any length.
[0126] In one example, the audio data to be decoded includes one or more audio frames. In one example, the audio data may also include an incomplete audio frame, such as 1 / 2 of an audio frame or 1 / 4 of an audio frame.
[0127] In the embodiment of the present application, the audio frame can be understood as a data segment with a specified time length obtained after framing and windowing the original audio data.
[0128] The embodiment of the present application does not limit the specific method of obtaining the original audio data.
[0129] In some examples, the original audio may be speech captured by the terminal.
[0130] In some examples, the original audio data may be a sound signal collected in an Internet voice call or video call scenario.
[0131] In some examples, the original audio data may be a sound signal collected in a live broadcast scenario, a sound signal collected in an online karaoke scenario, or a sound signal collected in a voice broadcast scenario.
[0132] In some examples, the original audio data may be audio data acquired from a storage resource, such as stored voice, music, video, etc.
[0133] In a possible implementation, when dividing the original audio data into audio frames, the embodiment of the present application may set a preset time length for division, for example, dividing every 10 ms of original audio in the original audio data into an audio frame.
[0134] In order to store and transmit audio data over long distances, the acquired original audio data needs to be encoded to reduce the size of the audio data, thereby reducing the storage space of the audio data or reducing the traffic bandwidth consumed by long-distance transmission.
[0135] In the related art, during the audio encoding process, the acquired audio data is firstly subjected to a nonlinear transformation, for example, the audio data is subjected to a nonlinear transformation through a coding network to obtain a coding vector of the audio data. Next, a unified quantizer is used to quantize the coding vector of the audio data, and finally the quantized coding vector is encoded to obtain a code stream. However, the effective audio information included in different audio data may be different. For example, some audio frames in an audio sequence include effective audio signals, while some audio frames do not include effective audio signals, such as audio frames in a noisy or silent state, which do not include effective audio signals. In the related art, when a unified quantizer is used to quantize all audio data, the characteristics of the audio data itself are not considered. For example, when the same quantizer is used to quantize some noise frames or silent frames as audio frames including effective speech signals, the quantization complexity will be increased, the quantization time will be increased, and the coding efficiency of the audio will be reduced.
[0136] In order to solve the above technical problems, when quantizing the coding vector of audio data, the embodiment of the present application adaptively selects the most suitable quantizer based on the characteristics of the audio data (i.e., the statistical value of the effective audio information), so as to avoid the quantization complexity caused by multiple quantization operations on the data of the low audio signal, thereby reducing the complexity of the quantization operation and improving the coding efficiency of the audio data.
[0137] In some embodiments, during quantization, the audio frames in the audio sequence may be quantized frame by frame, that is, during each quantization, the coding vector of an audio frame is quantized.
[0138] In some embodiments, during quantization, multiple audio frames in an audio sequence may be uniformly quantized, that is, during each quantization, the encoding vectors of multiple audio frames are quantized.
[0139] The embodiment of the present application does not limit the specific method in which the encoding device performs nonlinear transformation on the audio data to obtain the encoding vector of the audio data.
[0140] In some embodiments, the encoding device downsamples the audio data to obtain the encoding vector of the audio data. For example, the audio data is downsampled once or multiple times to obtain the encoding vector of the audio data.
[0141] In some embodiments, the encoding device downsamples and nonlinearly transforms the audio data through a convolutional layer-based encoding network to obtain an encoding vector of the audio data.
[0142] The embodiment of the present application does not limit the specific network structure of the encoding network based on the convolutional layer.
[0143] In one possible implementation, Figure 6 As shown, the convolutional layer-based encoding network in the embodiment of the present application includes an input layer, multiple encoding modules and an output layer, each encoding module is composed of multiple residual units, and each residual unit is composed of one or more convolutional layers. In the embodiment of the present application, the above encoding module is mainly used to downsample the input feature information and increase the number of channels. For example, for Figure 6 A coding module in the above coding module further downsamples the feature information output by the previous coding module to reduce the dimension of the feature information and increases the number of channels of the feature information, for example, doubling the number of channels to increase the number of channels of the feature information. Then, the feature information with reduced dimension and increased number of channels is input into the next coding module, and finally the feature information output by the output layer is obtained, and the feature information output by the output layer is recorded as the coding vector of the audio data.
[0144] For example, assume that the convolutional layer-based encoding network includes 1 input layer, 4 encoding modules, and 1 output layer. For each audio frame in the audio data, the audio frame is sampled to obtain a vector with c channels and 8000 dimensions. Then, the vector is input to the input layer, and after convolution processing, a feature vector with 32 channels and 8000 dimensions can be obtained. Then, the feature vector with 32 channels and 8000 dimensions is input to the first encoding module for dimensionality reduction processing, reducing the vector dimension to 1 / 2, and increasing the number of channels by 2 times, obtaining a feature vector with 64 channels and 4000 dimensions. The feature vector with 64 channels and 4000 dimensions is input to the second encoding module, reducing the vector dimension to 1 / 4, and increasing the number of channels by 2 times, obtaining a feature vector with 128 channels and 1000 dimensions. The feature vector with 128 channels and 1000 dimensions is input into the third encoding module, the vector dimension is reduced to 1 / 5, and the number of channels is doubled to obtain a feature vector with 256 channels and 200 dimensions. The feature vector with 256 channels and 200 dimensions is input into the fourth encoding module, the vector dimension is reduced to 1 / 8, and the number of channels is doubled to obtain a feature vector with 512 channels and 25 dimensions. Finally, the feature vector with 512 channels and 25 dimensions is input into the output layer for convolution processing to obtain a coding vector with 1 channel and D dimensions.
[0145] Based on the above steps, the encoding vector x∈R of the audio data can be obtained N×D , where D is the dimension of the encoding vector and N is the number of audio frames.
[0146] In some embodiments, the encoding end may also use other methods to perform nonlinear transformation on the audio data to obtain the encoding vector of the audio data, which is not limited in this embodiment of the present application.
[0147] S102: Count the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and select a target quantizer based on the statistical value of the effective audio information.
[0148] It should be noted that the embodiment of the present application does not limit the specific execution order of performing nonlinear transformation on the audio data in the above S102 and the above S101 to obtain the encoding vector of the audio data. In other words, S102 can be executed before determining the encoding vector of the audio data, or after determining the encoding vector of the audio data, or synchronously with determining the encoding vector of the audio data.
[0149] That is, in some embodiments, after the encoder obtains the audio data to be encoded, the audio data may be first nonlinearly transformed to obtain the encoding vector of the audio data. Then, the effective audio information of the audio data is counted to obtain the effective audio information statistics of the audio data, and the target quantizer is selected based on the effective audio information statistics.
[0150] In some embodiments, after the encoder obtains the audio data to be encoded, it can first count the effective audio information of the audio data to obtain the effective audio information statistics of the audio data, and select the target quantizer based on the effective audio information statistics. Then, the audio data is nonlinearly transformed to obtain the encoding vector of the audio data.
[0151] In some embodiments, after the encoding end obtains the audio data to be encoded, it can perform nonlinear transformation on the audio data to obtain the encoding vector of the audio data, and at the same time, perform statistics on the effective audio information of the audio data to obtain the effective audio information statistical value of the audio data, and select the target quantizer based on the effective audio information statistical value.
[0152] When quantizing the coding vectors in the audio sequence, the related technology selects a unified quantization operation for quantization, which increases the quantization complexity and quantization time for noise frames and silent frames without effective audio signals, thereby reducing the coding efficiency.
[0153] In the embodiment of the present application, based on the above steps, after obtaining the audio data, the effective audio information of the audio data is counted to obtain the statistical value of the effective audio information of the audio data, and then based on the statistical value of the effective audio information of the audio data, a target quantizer is selected, and then the target quantizer is used to quantize the coding vector of the audio data. For example, if the statistical value of the effective audio information of the audio data is less than or equal to a certain preset value, it means that the audio data is low audio data, such as a noise frame or a silent frame. At this time, a quantizer including fewer vector quantization layers can be used to quantize the coding vector of the audio data, which can reduce the quantization complexity. If the statistical value of the effective audio information of the audio data is greater than a certain preset value, it means that the audio data is high audio data, such as a speech frame. At this time, a quantizer including more vector quantization layers can be used to quantize the coding vector of the audio data, which can ensure the quantization effect. It can be seen from this that the embodiment of the present application can reduce the quantization complexity while ensuring the quantization performance through the statistical value of the effective audio information of the audio data, thereby improving the coding efficiency of the audio.
[0154] The specific process of determining the effective audio information statistics of audio data is introduced below.
[0155] In some embodiments, the effective audio information of the audio data is counted by a pre-trained classification neural network model. For example, the audio data is input into the classification neural network model, and the classification neural network model predicts whether the audio data is a low audio signal or a high audio signal, wherein the low audio signal indicates that the audio data includes less effective audio information (for example, the audio data is noise data or silent data), and then the first value (for example, 0) is determined as the effective audio information statistical value of the audio data. The high audio signal indicates that the audio data includes more effective audio information, and then the second value (for example, 1) is determined as the effective audio information statistical value of the audio data. That is, through the classification neural network model, the audio data is divided into a low audio signal or a high audio signal, and then the effective audio information statistical value of the audio data is determined to be either a first value (for example, 0) or a second value (for example, 1). Finally, based on the effective audio information statistical value of the audio data, a target quantizer is selected for the audio data. For example, if the audio data is a low audio signal, that is, the effective audio information statistical value of the audio data is a first value (for example, 0), a quantizer with fewer quantization layers is selected as the target quantizer of the audio data. If the audio data is a high-frequency audio signal, that is, the effective audio information statistic of the audio data is a second value (eg, 1), a quantizer with a larger number of quantization levels is selected as the target quantizer of the audio data.
[0156] In some embodiments, the effective audio information statistics of the audio data include at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data. It should be noted that in addition to including at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, the effective audio information statistics of the audio data may also include other statistics, and the embodiments of the present application are not limited to this. The following takes the effective audio information statistics of the audio data including at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data as an example to introduce the process of determining the effective audio information statistics of the audio data. Exemplarily, at least the following situations may be included:
[0157] In case 1, if the effective audio information statistics of the audio data include the energy value of the audio data, the encoder can determine the energy value of the audio data, and then determine the effective audio information statistics of the audio data based on the energy value of the audio data.
[0158] In one implementation of situation 1, the encoder determines the energy value of the audio data as a whole, for example, by taking the sample data included in the audio data as a whole and calculating an energy value as the energy value of the audio data.
[0159] In one implementation of situation 1, the above S102 counts the effective audio information of the audio data to obtain the effective audio information statistics of the audio data, including the following steps S102-A1 to S102-A3:
[0160] S102-A1, determining the energy value of each audio frame in the audio data;
[0161] S102-A2, determining the energy value of the audio data based on the energy value of each audio frame;
[0162] S102-A3. Determine a statistical value of effective audio information of the audio data based on the energy value of the audio data.
[0163] In this embodiment, the energy value of the audio data can be calculated to determine whether the audio data is silent data or noise data, or valid audio data. This is because the energy value of a silent signal or a noise signal is low, while the energy of an audio signal is high. Therefore, the statistical value of valid audio information of the audio data can be determined by calculating the energy value.
[0164] Specifically, the encoding end determines the energy value of each audio frame in the audio data, and then determines the energy value of the audio data based on the energy value of each audio frame. Finally, based on the energy value of the audio data, the effective audio information statistics of the audio data are determined.
[0165] In the embodiment of the present application, the specific process of determining the energy value of each audio frame in the audio data is basically the same. For the convenience of description, the process of determining the energy value of the i-th audio frame in the audio data is taken as an example for explanation.
[0166] The embodiment of the present application does not limit the specific method of determining the energy value of the ith audio frame. For example, the energy value of the ith audio frame is determined based on the sample data included in the ith audio frame.
[0167] In a possible implementation, the square sum of the sample data included in the ith audio frame is determined as the energy value of the ith audio frame. For example, assuming that the number of samples included in the ith audio frame is Framesize, the square sum of the sample data included in the ith frame is determined as the energy value E of the ith audio frame. i .
[0168] Exemplarily, the encoder determines the energy value of the i-th audio frame by the following formula (1):
[0169]
[0170] Among them, E i is the energy value of the i-th audio frame, and x(t) is the sample data of a sample in the i-th audio frame.
[0171] In a possible implementation manner, the sum of sample data included in the i-th audio frame is determined as the energy value of the i-th audio frame.
[0172] Based on the above steps, the terminal device can determine the energy value of each audio frame in the audio data.
[0173] In some embodiments, if the audio data to be encoded includes an audio frame, the effective audio information statistics of the audio data are determined based on the determined energy value of the audio frame. For example, the energy value of the audio frame is determined as the effective audio information statistics of the audio data.
[0174] In some embodiments, if the audio data to be encoded includes multiple audio frames, the encoding end determines the energy value of each audio frame in the audio data based on the above steps, and then executes the above S102-A2 steps to determine the energy value of the audio data based on the energy value of each audio frame in the audio data.
[0175] The embodiment of the present application does not limit the specific implementation method of the above S102-A2.
[0176] In one example, the encoding end determines the energy value of the audio data by summing the energy values of the audio frames included in the audio data. For example, if the audio data to be encoded includes 3 audio frames, the energy value of the first audio frame, the energy value of the second audio frame, and the energy value of the third audio frame are summed to determine the energy value of the audio data.
[0177] In one example, the encoding end determines the energy value of the audio data by averaging the energy values of the audio frames included in the audio data. For example, if the audio data to be encoded includes 3 audio frames, the energy value of the first audio frame, the energy value of the second audio frame, and the energy value of the third audio frame divided by 3 is determined as the energy value of the audio data.
[0178] After the encoder determines the energy value of the audio data based on the above steps, it executes the above steps S102-A3 to determine the effective audio information statistics of the audio data based on the energy value of the audio data.
[0179] In some embodiments, if the effective audio information statistics of the audio data only include the energy value of the audio data, the energy value of the audio data determined above can be determined as the effective audio information statistics of the audio data, and then based on the energy value of the audio data, the target quantizer of the audio data is selected. For example, if the energy value of the audio data is less than a certain threshold T (for example, T=0.06), the audio data is determined to be low-energy audio data, generally a noise signal or a silent signal, with very little audio information. At this time, in order to reduce the quantization complexity, a quantizer with fewer quantization layers is selected as the target quantizer of the audio data. Among them, the specific process of selecting the target quantizer refers to the specific introduction of the following embodiment, which will not be repeated here.
[0180] Case 2: If the effective audio information statistics of the audio data include the spectrum envelope statistics of the audio data, the encoding end can determine the effective audio information statistics of the audio data by determining the spectrum envelope statistics of the audio data, and then based on the spectrum envelope statistics of the audio data.
[0181] In one implementation of situation 2, the encoder determines the spectrum envelope statistics of the audio data as a whole, for example, by determining the spectrum envelope statistics of the audio data as a whole based on the sample data included in the audio data.
[0182] In one implementation of situation 2, the above S102 counts the effective audio information of the audio data to obtain the effective audio information statistics of the audio data, including the following steps S102-B1 to S102-B3:
[0183] S102-B1, determining a spectrum envelope statistic of each audio frame in the audio data;
[0184] S102-B2, determining a spectrum envelope statistic of the audio data based on the spectrum envelope statistic of each audio frame;
[0185] S102-B3. Determine effective audio information statistics of the audio data based on the spectrum envelope statistics of the audio data.
[0186] The spectrum envelope is a curve formed by connecting the highest points of the amplitude of different frequencies. It is called the spectrum envelope. The spectrum is a collection of many different frequencies, forming a very wide frequency range, and the amplitudes of different frequencies may be different. Speech is a complex multi-frequency signal, and each frequency component has a different amplitude. When they are arranged according to the size of the frequency, the curve formed by the top is called the speech spectrum envelope. The shape of the envelope changes with the sound produced. The sound waves generated by the vibration of the vocal cords will resonate when passing through the vocal tract composed of the oral cavity, nasal cavity, etc. As a result of the resonance, certain areas of the spectrum will be strengthened. Therefore, the shape of the spectrum envelope varies from person to person. But generally speaking, it has several peaks and troughs. The speech spectrum envelope is closely related to the semantic information and personality information of the speech signal.
[0187] Based on this, in an embodiment of the present application, the spectral envelope statistics of the audio data can be calculated to determine whether the audio data is silent data or noise data, or valid audio data. This is because the frequency of the silent signal or the noise signal is lower, while the frequency of the audio signal is higher. Therefore, the valid audio information statistics of the audio data can be determined by calculating the spectral envelope statistics.
[0188] Specifically, the encoding end determines the spectrum envelope statistics of each audio frame in the audio data, and then determines the spectrum envelope statistics of the audio data based on the spectrum envelope statistics of each audio data. Finally, based on the spectrum envelope statistics of the audio data, the effective audio information statistics of the audio data are determined.
[0189] In the embodiment of the present application, the specific process of determining the spectrum envelope statistics of each audio frame in the audio data is basically the same. For the convenience of description, the process of determining the spectrum envelope statistics of the i-th audio frame in the audio data is taken as an example for explanation.
[0190] The embodiment of the present application does not limit the specific method of determining the spectrum envelope statistical value of the i-th audio frame.
[0191] In a possible implementation, the encoder converts the i-th audio frame from a time domain signal to a frequency domain signal, where i is a positive integer, and then performs statistics on the spectrum envelope of the frequency domain signal of the i-th audio frame to obtain a spectrum envelope statistical value of the i-th audio frame.
[0192] For example, the encoding end uses short-time Fourier transform (SFT) to convert the i-th audio frame from a time domain signal to a frequency domain signal, and constructs the spectrum envelope of the i-th audio frame based on the frequency domain signal of the i-th audio frame. Then, based on the spectrum envelope of the i-th audio frame, the envelope of the frequency band is counted to obtain the spectrum envelope statistics of the i-th audio frame.
[0193] Based on the above steps, the terminal device can determine the spectrum envelope statistics of each audio frame in the audio data.
[0194] In some embodiments, if the audio data includes an audio frame, the effective audio information statistics of the audio data are determined based on the determined spectrum envelope statistics of the audio frame. For example, the spectrum envelope statistics of the audio frame are determined as the effective audio information statistics of the audio data.
[0195] In some embodiments, if the audio data includes multiple audio frames, the encoding end determines the spectrum envelope statistics of each audio frame in the audio data based on the above steps, and then executes the above S102-B2 steps to determine the spectrum envelope statistics of the audio data based on the spectrum envelope statistics of each audio frame in the audio data.
[0196] The embodiment of the present application does not limit the specific implementation method of the above S102-B2.
[0197] In one example, the encoding end determines the spectrum envelope statistics of the audio data by summing the spectrum envelope statistics of each audio frame included in the audio data. For example, if the audio data includes three audio frames, the spectrum envelope statistics of the first audio frame, the spectrum envelope statistics of the second audio frame, and the spectrum envelope statistics of the third audio frame are summed to determine the spectrum envelope statistics of the audio data.
[0198] In one example, the encoding end determines the spectral envelope statistics of the audio data by taking the average value of the spectral envelope statistics of each audio frame included in the audio data. For example, if the audio data includes three audio frames, the sum of the spectral envelope statistics of the first audio frame, the spectral envelope statistics of the second audio frame, and the spectral envelope statistics of the third audio frame divided by 3 is determined as the spectral envelope statistics of the audio data.
[0199] After the encoder determines the spectrum envelope statistics of the audio data based on the above steps, it executes the above steps S102-B3 to determine the effective audio information statistics of the audio data based on the spectrum envelope statistics of the audio data.
[0200] In some embodiments, if the effective audio information statistics of the audio data only include the spectrum envelope statistics of the audio data, the sum of the number of signals in the spectrum envelope of the audio data whose frequency is greater than a preset frequency (for example, 15 Hz) can be determined as the effective audio information statistics of the audio data.
[0201] In some embodiments, if the effective audio information statistics of the audio data only include the spectrum envelope statistics of the audio data, the effective audio information statistics of the audio data can be determined based on the spectrum envelope statistics of the audio data determined above and the number of samples of the audio data.
[0202] For example, if the spectral envelope statistics of the above-mentioned audio data are the sum of the spectral envelope statistics of each audio frame in the audio data, then the ratio of the sum of the number of signals (also called samples) in the spectral envelope statistics of the audio data with a frequency greater than a preset frequency (for example, 15 Hz) to the sum of the total number of samples of the audio data can be determined as the effective audio information statistics of the audio data.
[0203] For another example, if the spectral envelope statistic of the above-mentioned audio data is the average value of the spectral envelope statistic of each audio frame included in the audio data, then the ratio of the sum of the number of signals (also called samples) with a frequency greater than a preset frequency (for example, 15 Hz) in the average value of the spectral envelope statistic of each audio frame included in the audio data to the average value of the number of samples of the audio data can be determined as the effective audio information statistic of the audio data.
[0204] Based on the above steps, after the encoding end determines the effective audio information statistics of the audio data based on the spectral envelope statistics of the audio data, it selects the target quantizer of the audio data based on the effective audio information statistics of the audio data. For example, if the proportion of signals less than a preset frequency (e.g., 15Hz) in the spectral envelope statistics of the audio data is above a preset proportion (e.g., 60%), the audio data is determined to be a low audio signal, generally noise data or silent data, with very little audio information. At this time, in order to reduce the quantization complexity, a quantizer with fewer quantization layers is selected as the target quantizer of the audio data. Among them, the specific process of selecting the target quantizer refers to the specific introduction of the following embodiment, which will not be repeated here.
[0205] Case 3: if the effective audio information statistics of the audio data include the zero-crossing rate of the audio data, the encoder can determine the effective audio information statistics of the audio data by determining the zero-crossing rate of the audio data, and then determine the effective audio information statistics of the audio data based on the zero-crossing rate of the audio data.
[0206] In one implementation of situation 3, the encoder determines the zero-crossing rate of the audio data as a whole. For example, the sample data included in the audio data is taken as a whole to determine a zero-crossing rate as a spectral envelope statistic of the audio data.
[0207] In one implementation of situation 3, the above S102 counts the effective audio information of the audio data to obtain the effective audio information statistics of the audio data, including the following steps S102-C1 to S102-C3:
[0208] S102-C1, determining the zero-crossing rate of each audio frame in the audio data;
[0209] S102-C2, determining a zero-crossing rate of the audio data based on the zero-crossing rate of each audio frame;
[0210] S102-C3. Determine a statistical value of effective audio information of the audio data based on the zero-crossing rate of the audio data.
[0211] The zero-crossing rate (ZCR) is the number of times a frame of speech time domain signal crosses 0 (time axis), that is, the number of times the speech signal passes through the zero point (from positive to negative or from negative to positive). When the ratio of the short-term zero-crossing rate of an audio frame to the number of samples framesize included in the frame (i.e., the number of zero-crossing times / framesize) is lower than a preset value (e.g., 30%, which is acceptable), the audio frame is identified as a low audio frame.
[0212] Based on this, in an embodiment of the present application, the zero-crossing rate of the audio data can be calculated to determine whether the audio data is silent data or noise data, or valid audio data.
[0213] Specifically, the encoding end determines the zero-crossing rate of each audio frame in the audio data, and then determines the zero-crossing rate of the audio data based on the zero-crossing rate of each audio frame. Finally, based on the zero-crossing rate of the audio data, the effective audio information statistics of the audio data are determined.
[0214] In the embodiment of the present application, the specific process of determining the zero-crossing rate of each audio frame in the audio data is basically the same. For the convenience of description, the process of determining the zero-crossing rate of the i-th audio frame in the audio data is taken as an example for explanation.
[0215] The embodiment of the present application does not limit the specific method of determining the zero-crossing rate of the i-th audio frame.
[0216] In a possible implementation, the encoder determines the number of times that sample data (ie, speech signal) included in the i-th audio frame crosses the zero time axis as the zero-crossing rate of the i-th audio frame.
[0217] Based on the above steps, the terminal device can determine the zero-crossing rate of each audio frame in the audio data.
[0218] In some embodiments, if the audio data includes an audio frame, the effective audio information statistics of the audio data are determined based on the determined zero-crossing rate of the audio frame. For example, the zero-crossing rate of the audio frame is determined as the effective audio information statistics of the audio data.
[0219] In some embodiments, if the audio data includes multiple audio frames, the encoding end determines the zero-crossing rate of each audio frame in the audio data based on the above steps, and then executes the above S102-C2 steps to determine the zero-crossing rate of the audio data based on the zero-crossing rate of each audio frame in the audio data.
[0220] The embodiment of the present application does not limit the specific implementation method of the above S102-C2.
[0221] In one example, the encoding end determines the zero-crossing rate of the audio data by summing the zero-crossing rates of the audio data. For example, when the audio data includes three audio frames, the sum of the zero-crossing rate of the first audio frame, the zero-crossing rate of the second audio frame, and the zero-crossing rate of the third audio frame is determined as the zero-crossing rate of the audio data.
[0222] In one example, the encoder determines the zero-crossing rate of the audio data by averaging the zero-crossing rates of the audio frames included in the audio data. For example, when the audio data includes three audio frames, the sum of the zero-crossing rate of the first audio frame, the zero-crossing rate of the second audio frame, and the zero-crossing rate of the third audio frame divided by 3 is determined as the zero-crossing rate of the audio data.
[0223] After the encoder determines the zero-crossing rate of the audio data based on the above steps, it executes the above steps S102-C3 to determine the effective audio information statistics of the audio data based on the zero-crossing rate of the audio data.
[0224] In some embodiments, if the effective audio information statistics of the audio data only include the zero-crossing rate of the audio data, the zero-crossing rate of the audio data determined above may be determined as the effective audio information statistics of the audio data.
[0225] In some embodiments, if the effective audio information statistics of the audio data only include the zero-crossing rate of the audio data, the effective audio information statistics of the audio data can be determined based on the zero-crossing rate of the audio data and the number of samples of the audio data determined above.
[0226] For example, if the zero-crossing rate of the audio data is the sum of the zero-crossing rates of each audio frame in the audio data, the ratio of the sum of the zero-crossing rates of the audio data to the sum of the number of samples of the audio data can be determined as the effective audio information statistic of the audio data.
[0227] For another example, if the zero-crossing rate of the audio data is the average value of the zero-crossing rate of the audio data, the ratio of the average value of the zero-crossing rate of the audio data to the average value of the number of samples of the audio data can be determined as the effective audio information statistic of the audio data.
[0228] Based on the above steps, the encoding end determines the effective audio information statistics of the audio data based on the zero-crossing rate of the audio data, and then selects the target quantizer of the audio data based on the effective audio information statistics of the audio data. For example, if the ratio of the zero-crossing rate of the audio data to the number of samples of the audio data is less than a preset value (for example, 30%), the audio data is determined to be a low audio signal, generally a noise frame or a silent frame, with very little audio information. At this time, in order to reduce the quantization complexity, a quantizer with fewer quantization layers is selected as the target quantizer of the audio data. Among them, the specific process of selecting the target quantizer refers to the specific introduction of the following embodiment, which will not be repeated here.
[0229] The above describes the specific process of determining the effective audio information statistics of the audio data by the encoder. The following describes the process of selecting a target quantizer for the audio data based on the effective audio information statistics of the audio data by the encoder.
[0230] The embodiment of the present application does not limit the specific manner in which the encoding end selects a target quantizer for the audio data based on the statistical value of the effective audio information of the audio data.
[0231] In some embodiments, the embodiments of the present application set different quantizers for different effective audio information statistics. Exemplarily, as shown in Table 1, different effective audio information statistics correspond to different quantizers:
[0232] Table 1
[0233] Valid audio information statistics [a1, a2) Quantizer 1 Valid audio information statistics [a2, a3) Quantizer 2 …… ……
[0234] As shown in Table 1, different effective audio information statistical value intervals correspond to different quantizers, where quantizer 1 and quantizer 2 are different, for example, quantizer 1 and quantizer 2 include different numbers of quantization layers. Based on this, the encoding end can query the target quantizer of the audio data obtained in Table 1 based on the effective audio information statistical value of the audio data calculated in the steps. For example, the effective audio information statistical value of the audio data calculated above is a2, and quantizer 1 can be determined as the target quantizer of the audio data.
[0235] In some embodiments, the encoder may select a target quantizer through the following steps S102-D1 and S102-D2:
[0236] S102-D1, determining a target quantization level number of a quantizer corresponding to the audio data based on the effective audio information statistics;
[0237] S102-D2. Select a target quantizer based on the target number of quantization layers.
[0238] In this embodiment, the encoding end can also adopt this method to determine the target quantizer of the audio data. Specifically, the encoding end determines the target quantization layer number of the quantizer corresponding to the audio data based on the game audio information statistics of the audio data determined above. Then, based on the target quantization layer number, the quantization layer included in the preset quantizer is deleted, and the deleted quantizer is determined as the target quantizer of the audio data.
[0239] The following is an introduction to the specific method for the encoder to determine the target quantization level of the quantizer corresponding to the audio data based on the effective audio information statistics of the audio data. It should be noted that in the embodiment of the present application, the specific method for the encoder to determine the target quantization level of the quantizer corresponding to the audio data based on the effective audio information statistics of the audio data includes but is not limited to the following:
[0240] In a first approach, if the statistical value of the effective audio information of the audio data is less than a first preset value, the preset first quantization level number is determined as the target quantization level number.
[0241] In the first method, the quantizer is divided into two categories, one is a quantizer with a first quantization layer number, and the other is a quantizer with a third quantization layer number. Optionally, the third quantization layer number is greater than the first preset quantization layer number. In this way, when determining the target quantizer of the audio data, based on the above steps, the effective audio information statistics of the audio data are determined, and then the effective audio information statistics of the audio data are compared with the first preset value. If the effective audio information statistics of the audio data are less than the first preset value, the preset first quantization layer number is determined as the target quantization layer number corresponding to the audio data. If the effective audio information statistics of the audio data are greater than or equal to the first preset value, the preset third quantization layer number is determined as the target quantization layer number corresponding to the audio data.
[0242] In one example, the corresponding relationship between the effective audio information statistical value and the target quantization level is shown in Table 2:
[0243] Table 2
[0244]
[0245]
[0246] As shown in Table 2, in this implementation, the encoder compares the effective audio information statistic B calculated based on the above steps with the first preset value C. If the effective audio information statistic B is less than the first preset value C, the preset first quantization layer number is determined as the target quantization layer number corresponding to the audio data. If the effective audio information statistic B is greater than or equal to the first preset value C, the preset third quantization layer number is determined as the target quantization layer number corresponding to the audio data.
[0247] In some embodiments, the third number of quantization layers may be the number of quantization layers included in a preset quantizer. That is, in an embodiment of the present application, when the coding vector of the audio data is quantized using a preset quantizer, firstly, based on the above steps, the effective audio information statistic B of the audio data is determined, and then the effective audio information statistic B is compared with the first preset value C. If the effective audio information statistic B of the audio data is less than the first preset value C, it means that the audio data is a low-frequency signal. In order to reduce the quantization complexity, fewer quantization layers are used to quantize the coding vector of the audio data, and then the first number of quantization layers are selected from the quantization layers included in the preset quantizer, and the coding vector of the audio data is quantized to reduce the quantization complexity. Since fewer quantization layers mean fewer codebook indexes to be encoded (because one quantization layer corresponds to one codebook index to be encoded), the information to be encoded is reduced, thereby improving the coding efficiency of the audio data.
[0248] If the effective audio information statistic B of the audio data is greater than or equal to the first preset value C, it means that the audio data is not a low audio signal. In order to ensure the encoding performance, the preset quantizer is used to quantize the encoding vector of the audio data.
[0249] Method 2: In the embodiment of the present application, different effective audio information statistical value intervals correspond to different quantization levels.
[0250] Exemplarily, the first correspondence between different valid audio information statistical value intervals and different quantization levels is shown in Table 3:
[0251] Table 3
[0252] Valid audio information statistics The number of quantization layers of the quantizer [a1,a2) A1 [a2,a3) A2 …… ……
[0253] As shown in Table 3, different effective audio information statistical value intervals correspond to different quantization layer numbers.
[0254] Based on this, when the encoding end determines the target quantization layer number of the quantizer corresponding to the audio data based on the effective audio information statistics, it first obtains the first correspondence between different effective audio information statistical value intervals and different quantization layer numbers, and determines the second quantization layer number corresponding to the effective audio information statistics in the first correspondence as the target quantization layer number.
[0255] For example, based on the above steps, the encoding end calculates the effective audio information statistic of the audio data as a2, and queries in the above Table 3 to obtain the quantization level corresponding to the effective audio information statistic a2 as A2, and then determines A2 as the target quantization level of the target quantizer of the audio data.
[0256] As can be seen from the above, in some embodiments, the effective audio information statistics of the audio data include at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero crossing rate of the audio data. The first correspondence between different effective audio information statistics intervals and different quantization levels shown in Table 3 includes the correspondence between at least one of the energy value interval of the audio data, the spectrum envelope statistics interval of the audio data, and the zero crossing rate interval of the audio data, and different quantization levels.
[0257] In one example, if the effective audio information statistics of the audio data include any one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, then the first corresponding relationship includes the corresponding relationship between any one of the energy value interval, the spectrum envelope statistics interval, and the zero-crossing rate interval and different quantization levels. For example, it includes the corresponding relationship between different energy value intervals and different quantization levels, or includes the corresponding relationship between different spectrum envelope statistics intervals and different quantization levels, or includes the corresponding relationship between different zero-crossing rate intervals and different quantization levels.
[0258] In one example, if the effective audio information statistics of the audio data include any two of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, then the first corresponding relationship includes the corresponding relationship between any two of the energy value interval, the spectrum envelope statistics interval, and the zero-crossing rate interval and different quantization layers. Or include the corresponding relationship between different energy value intervals and spectrum envelope statistics intervals and different quantization layers, in which case the effective audio information statistics interval in Table 3 includes 2 parameters (energy value interval, spectrum envelope statistics interval). Or include the corresponding relationship between different energy value intervals and zero-crossing rate intervals and different quantization layers, in which case the effective audio information statistics interval in Table 3 includes 2 parameters (energy value interval, zero-crossing rate interval). Or the corresponding relationship between different zero-crossing rate intervals and spectrum envelope statistics intervals and different quantization layers, in which case the effective audio information statistics in Table 3 include 2 parameters (zero-crossing rate interval, spectrum envelope statistics interval).
[0259] In one example, if the effective audio information statistics of the audio data include the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero crossing rate of the audio data, then the first correspondence includes the correspondence between the energy value interval, the spectrum envelope statistics interval, and the zero crossing rate interval and different quantization layers. At this time, the effective audio information statistics interval in Table 4 includes 3 parameters (energy value interval, spectrum envelope statistics interval, zero crossing rate interval).
[0260] Method three, in the embodiment of the present application, different effective audio information statistical values correspond to different quantization levels.
[0261] Exemplarily, the second corresponding relationship between different effective audio information statistical values and different quantization levels is shown in Table 4:
[0262] Table 4
[0263] Valid audio information statistics The number of quantization layers of the quantizer a1 A1 a2 A2 a3 A3 …… ……
[0264] As shown in Table 4, different effective audio information statistics correspond to different quantization levels.
[0265] Based on this, when the encoding end determines the target quantization layer number of the quantizer corresponding to the audio data based on the effective audio information statistics, it first obtains the second correspondence between different effective audio information statistics and different quantization layers, and determines the fourth quantization layer number corresponding to the effective audio information statistics in the second correspondence as the target quantization layer number.
[0266] For example, based on the above steps, the encoding end calculates the effective audio information statistic of the audio data as a3, and queries in the above Table 3 to obtain the quantization level corresponding to the effective audio information statistic a3 as A3, and then determines A3 as the target quantization level of the target quantizer of the audio data.
[0267] In one possible implementation, the second corresponding relationship between different effective audio information statistics and different quantization levels is a linear or nonlinear mathematical relationship. Therefore, based on the linear or nonlinear mathematical relationship, the target quantization level corresponding to the effective audio information statistics of the audio data can be determined.
[0268] For example, the second corresponding relationship between different effective audio information statistics and different quantization levels is shown in formula (2):
[0269] F=f(x) (2)
[0270] Wherein, x is the effective audio information statistic, f() is the second correspondence between the effective audio information statistic and the number of quantization layers, and F is the number of quantization layers. In this way, after the encoder determines the effective audio information statistic of the audio data based on the above steps, it substitutes the effective audio information statistic of the audio data into the above formula (2) to calculate the target quantization layer number corresponding to the audio data.
[0271] As can be seen from the above, in some embodiments, the effective audio information statistics of the audio data include at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data. The second correspondence between different effective audio information statistics and different quantization levels shown in Table 4 includes the correspondence between at least one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, and different quantization levels.
[0272] In one example, if the effective audio information statistics of the audio data include any one of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, then the second corresponding relationship includes the corresponding relationship between any one of the energy value, the spectrum envelope statistics, and the zero-crossing rate and different numbers of quantization layers. For example, it includes the corresponding relationship between different energy values and different numbers of quantization layers, or includes the corresponding relationship between different spectrum envelope statistics and different numbers of quantization layers, or includes the corresponding relationship between different zero-crossing rates and different numbers of quantization layers.
[0273] In one example, if the effective audio information statistics of the audio data include any two of the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero-crossing rate of the audio data, then the second corresponding relationship includes the corresponding relationship between any two of the energy value, the spectrum envelope statistics, and the zero-crossing rate and different numbers of quantization layers. For example, including the corresponding relationship between different energy values and spectrum envelope statistics and different numbers of quantization layers, in this case, the effective audio information statistics in Table 4 include 2 parameters (energy value, spectrum envelope statistics). Or including the corresponding relationship between different energy values and zero-crossing rates and different numbers of quantization layers, in this case, the effective audio information statistics in Table 4 include 2 parameters (energy value, zero-crossing rate). Or including the corresponding relationship between different zero-crossing rates and spectrum envelope statistics and different numbers of quantization layers, in this case, the effective audio information statistics in Table 4 include 2 parameters (zero-crossing rate, spectrum envelope statistics).
[0274] In one example, if the effective audio information statistics of the audio data include the energy value of the audio data, the spectrum envelope statistics of the audio data, and the zero crossing rate of the audio data, then the second corresponding relationship includes the energy value, the spectrum envelope statistics, and the zero crossing rate and the corresponding relationship between different quantization levels. At this time, the effective audio information statistics in Table 4 include 3 parameters (energy value, spectrum envelope statistics, zero crossing rate).
[0275] After the encoder determines the target number of quantization levels of the quantizer corresponding to the audio data based on the above steps, it executes the above steps S102-D2 to select a target quantizer based on the target number of quantization levels.
[0276] The embodiment of the present application does not limit the specific manner in which the encoding end selects the target quantizer based on the target number of quantization layers.
[0277] In one possible implementation, the encoding end in the embodiment of the present application can obtain multiple quantizers with different quantization layers, so that the encoding end can select a quantizer with a target quantization layer number from the multiple quantizers without quantization layers based on the target quantization layer number determined above as the target quantizer for the audio data.
[0278] In a possible implementation, the encoding end includes a preset quantizer, so that the encoding end can select quantization layers of a target number of quantization layers from the total quantization layers of the preset quantizer to form a target quantizer for audio data.
[0279] For example, Figure 7As shown, the preset quantizer includes 8 quantization layers, and the target quantization layer number corresponding to the audio data determined by the encoder based on the above steps is 3, so 3 quantization layers are selected from the 8 quantization layers included in the preset quantizer to form the target quantizer of the audio data. For example, the encoder randomly selects 3 quantization layers from the 8 quantization layers included in the preset quantizer, that is, quantization layer 1, quantization layer 3 and quantization layer 5 to form the target quantizer of the audio data.
[0280] In order to ensure the quantization effect, in some embodiments, at least two quantization layers are non-adjacent quantization layers among the quantization layers of the target number of quantization layers selected from the total quantization layers of the preset quantizer. For example, three non-adjacent quantization layers are randomly selected from the eight quantization layers included in the preset quantizer to form the target quantizer.
[0281] After the encoder determines the target quantizer of the audio data based on the above steps, it executes the following step S103.
[0282] S103: Use a target quantizer to quantize the coding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a bit stream.
[0283] In an embodiment of the present application, when the encoding end quantizes the encoding vector of the audio data, it determines the effective audio information statistics of the audio data, and then selects the target quantizer of the audio data based on the effective audio information statistics of the audio data, and then uses the target quantizer to quantize the encoding vector of the audio data to obtain a quantization result, and encodes the quantization result to obtain a bit stream. In other words, the embodiment of the present application adaptively selects the most suitable quantizer based on the characteristics of the input audio data (i.e., the effective audio information statistics), avoids multiple quantization operations on the data of the low audio signal, and then can reduce the complexity of the quantization operation, can improve the encoding efficiency of the audio data, and effectively control the complexity of decoding.
[0284] The specific process of quantizing the coding vector of audio data using the target quantizer is introduced below.
[0285] The target quantizer of the embodiment of the present application includes a target quantization layer, for example, the target quantizer includes K quantization layers. In this way, the coding vector of the audio data is iteratively quantized using the K quantization layers included in the target quantizer, and finally the quantization result of the audio data is obtained. For example, the coding vector of the audio data is input into the first quantization layer of the target quantizer for quantization, and the first quantization vector and the codebook subscript corresponding to the first quantization layer are obtained; based on the coding vector of the audio data and the first quantization vector, the first residual vector is obtained; the first residual vector is input into the second quantization layer of the target quantizer for quantization, and it is repeated to obtain the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook subscript corresponding to each quantization layer in the target quantizer; the residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer are encoded to obtain a code stream.
[0286] Exemplarily, taking the i-th audio frame in the audio data as an example, first, the coding vector of the i-th audio frame is input into the first quantization layer in the target quantizer for quantization, and a vector 1 matching the coding vector of the i-th audio frame is searched in the vectors included in the codebook corresponding to the first quantization layer as the first quantization vector of the first quantization layer, and the codebook index of the vector 1 in the codebook corresponding to the first quantization layer is recorded. Then, the difference between the coding vector of the i-th audio frame and the first quantization vector is input into the second quantization layer of the target quantizer for quantization, and a vector 2 matching the vector 1 is searched in the vectors included in the codebook corresponding to the second quantization layer as the second quantization vector of the second quantization layer, and the codebook index of the vector 2 in the codebook corresponding to the second quantization layer is recorded. Then, the difference between the first quantization vector of the first quantization layer and the second quantization vector of the second quantization layer is input into the third quantization layer of the target quantizer for quantization, and so on, the residual vector corresponding to the last quantization layer of the target quantizer and the codebook index of the quantization result of each quantization layer in the target quantizer in the corresponding codebook can be obtained.
[0287] For example, assume that the target quantizer includes three quantization layers, such as Figure 8 In this way, the encoder inputs the encoding vector y0 of the i-th audio frame into the first quantization layer of the target quantizer for quantization, searches for a vector matching the encoding vector y0 in the vectors included in the codebook 1 corresponding to the first quantization layer, and uses it as the output vector y1 of the first quantization layer. At the same time, the codebook index index0 of the vector y1 in the codebook 1 corresponding to the first quantization layer is recorded. The difference between the encoding vector y0 and the vector y1 is used as the residual vector Then the residual vector The input is quantized in the second quantization layer of the target quantizer, and the residual vector is searched in the vector included in the codebook 2 corresponding to the second quantization layer. The matched vector is used as the output vector y2 of the second quantization layer. At the same time, the codebook index index1 of the vector y2 in the codebook 2 corresponding to the second quantization layer is recorded. Then, the residual vector The difference with vector y2 is used as the residual vector Then the residual vector The input is quantized in the third quantization layer of the target quantizer, and the residual vector is searched in the vector included in the codebook 3 corresponding to the third quantization layer. The matched vector is used as the output vector y3 of the third quantization layer. At the same time, the codebook index index2 of the vector y3 in the codebook 3 corresponding to the third quantization layer is recorded. Then, the residual vector The difference with vector y3 is used as the residual vector
[0288] As can be seen from the above, in the embodiment of the present application, the target quantizer is used to quantize the coding vector of the audio data, and the final quantization result obtained includes at least the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook index corresponding to each quantization layer. Figure 8 As shown, the final quantization result includes the residual vector corresponding to the last quantization layer of the target quantizer And the codebook indexes index0, index1 and index2 corresponding to each quantization layer.
[0289] Next, the encoder encodes the above quantization results to obtain a bitstream. For example, the encoder encodes the residual vector corresponding to the last quantization layer of the target quantizer And the codebook indexes index0, index1 and index2 corresponding to each quantization layer are encoded to form a binary code stream.
[0290] As can be seen from the above, in the embodiment of the present application, the encoding end selects a target quantizer for the audio data based on the effective audio information statistics of the audio data. Therefore, in order to ensure that the decoding end can accurately decode the bit stream of the audio data, the bit stream in the embodiment of the present application also includes indication information for indicating the target quantizer.
[0291] The embodiment of the present application does not limit the specific form of the indication information.
[0292] In a possible implementation, if when determining the target quantizer as described above, the encoder selects from a plurality of preset quantizers with different quantization layers based on the statistical value of the effective audio information of the audio data, then the indication information may include index information of the target quantizer in the plurality of quantizers with different quantization layers. In this way, the decoder can determine the target quantizer from the quantizers with different quantization layers based on the index information, and then accurately dequantize the quantization result of the audio data based on the target quantizer to obtain the reconstructed value of the coding vector of the audio data, and then decode based on the reconstructed value of the coding vector of the audio data to obtain the reconstructed value of the audio data, thereby achieving accurate decoding of the audio data.
[0293] In a possible implementation, if when determining the target quantizer as above, the encoder selects a target number of quantization layers from the total quantization layers included in the preset quantizer, then the indication information may include index information of the quantization layers included in the target quantizer, and the index information is used to indicate the sequence number of the quantization layer in the quantization layers included in the preset quantizer. In this way, the decoder can select the target quantization layer from the quantization layers included in the preset quantizer based on the index information of the quantization layer included in the target quantizer, obtain the target quantizer, and then accurately dequantize the quantization result of the audio data based on the target quantizer.
[0294] In the audio encoding method provided by the embodiment of the present application, the encoding end counts the effective audio information of the audio data to be encoded, obtains the effective audio information statistics of the audio data, and then selects the target quantizer based on the effective audio information statistics, and finally uses the target quantizer to quantize the encoding vector of the audio data. In other words, the embodiment of the present application adaptively selects the most suitable quantizer based on the characteristics of the input audio data (i.e., the effective audio information statistics), avoids multiple quantization operations on the data of the low audio signal, thereby reducing the complexity of the quantization operation and improving the encoding efficiency of the audio data.
[0295] The above describes the audio encoding method involved in the embodiment of the present application. The following describes the audio decoding method provided by the embodiment of the present application by taking the decoding end as an example.
[0296] Fig. 9 1 is a flow chart of an audio decoding method provided by an embodiment of the present application. The execution subject of the embodiment of the present application may be a device with a specific audio decoding function, such as an audio decoding device. In some embodiments, the audio decoding device may be Figure 1 For the convenience of description, the present application embodiment is described by taking the execution subject as a decoding device as an example.
[0297] like Fig. 9 As shown, the audio decoding method of the embodiment of the present application includes the following steps:
[0298] S201, decoding a bit stream to obtain a quantization result corresponding to the audio data.
[0299] The quantization result is obtained by quantizing the coding vector of the audio data through a target quantizer. For the specific quantization process, refer to the relevant description of S103 above.
[0300] The encoding vector of the audio data is obtained by performing nonlinear transformation on the audio data. The specific extraction process of the encoding vector refers to the relevant description of S101 above.
[0301] The target quantizer is selected based on the statistical value of the effective audio information of the audio data. The specific selection process refers to the relevant description of S102 above.
[0302] Exemplarily, the effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
[0303] In some embodiments, the target quantizer is determined based on a target quantization level of a quantizer corresponding to the audio data, and the target quantization level is determined based on a statistical value of effective audio information.
[0304] In an example, if the statistical value of the effective audio information is less than a first preset value, the target quantization level is a preset first quantization level.
[0305] In an example, the target quantization level is the second quantization level corresponding to the effective audio information statistical value in the first corresponding relationship, and the first corresponding relationship includes the relationship between different effective audio information statistical value intervals and different quantization levels.
[0306] In an example, the target quantization level is determined based on a second corresponding relationship and a statistical value of effective audio information, where the second corresponding relationship is a relationship between different statistical values of effective audio information and different quantization levels.
[0307] In some embodiments, the target quantizer is composed of a target number of quantization layers selected from a total number of quantization layers of a preset quantizer.
[0308] Exemplarily, among the target number of quantization layers selected from the total number of quantization layers of the preset quantizer, at least two quantization layers are non-adjacent quantization layers.
[0309] In an embodiment of the present application, when encoding audio data, the encoding end obtains the encoding vector of the audio data, and determines the effective audio information statistics of the audio data, and then selects a target quantizer for the audio data based on the effective audio information statistics, and then uses the target quantizer to quantize the encoding vector of the audio data to obtain the quantization result of the audio data, and finally encodes the quantization result to obtain a code stream. Correspondingly, after the decoding end obtains the code stream, it decodes the code stream to obtain the quantization result of the audio data, and then dequantizes the quantization result to obtain the reconstructed encoding vector of the audio data (also referred to as the reconstructed value of the encoding vector), and finally decodes the reconstructed encoding vector of the audio data to obtain the reconstructed value of the audio data. That is, in an embodiment of the present application, the encoding end adaptively selects the most suitable quantizer based on the effective audio information statistics of the input audio data, avoids multiple quantization operations on the data of the low audio signal, and then can reduce the complexity of the quantization operation, and can improve the encoding efficiency of the audio data. At the same time, when the data of the low audio signal is quantized using fewer quantization layers, the codebook index to be encoded can be reduced, which can not only reduce the encoded data at the encoding end, but also reduce the decoded data at the decoding end, thereby improving the decoding performance of the audio data.
[0310] According to the above encoding process, in some embodiments, the above quantization result includes the codebook index corresponding to each quantization layer and the residual vector corresponding to the last quantization layer when the target quantizer quantizes the encoding vector of the audio data. In this way, the decoding end decodes the code stream of the audio data to obtain the codebook index corresponding to each quantization layer and the residual vector corresponding to the last quantization layer.
[0311] S202: Dequantize the quantization result to obtain a reconstructed coding vector of the audio data.
[0312] In an embodiment of the present application, if the decoding end decodes the code stream of the audio data, and the obtained quantization result includes the codebook index corresponding to each quantization layer and the residual vector corresponding to the last quantization layer, the decoding end queries the corresponding codebook based on the codebook index corresponding to each quantization layer to obtain the corresponding vector, and then accumulates the vectors obtained by the query with the residual vector corresponding to the last quantization layer to obtain the reconstructed coding vector of the audio data.
[0313] For example, Fig.10 As shown, assuming that the target quantizer includes 3 layers, the decoder decodes the bitstream and obtains the residual vector corresponding to the last quantization layer of the target quantizer And the codebook indexes index0, index1 and index2 corresponding to each quantization layer. In this way, firstly query the vector corresponding to the codebook index index2 in the codebook corresponding to the third quantization layer, and then determine the vector as the reconstructed value y3' of the output vector of the third vector layer, and compare the vector y3' with the residual vector The sum of the values is determined as the residual vector The reconstruction value Next, the vector corresponding to the codebook index index1 is queried in the codebook corresponding to the second quantization layer, and the vector is determined as the reconstructed value y2' of the output vector of the second vector layer, and the output vector y2' is compared with the reconstructed value of the residual vector The sum of the values is determined as the residual vector The reconstruction value Next, the vector corresponding to the codebook index index0 is queried in the codebook corresponding to the first quantization layer, and the vector is determined as the reconstructed value y1' of the output vector of the first vector layer, and the vector y1' is compared with the reconstructed value of the residual vector The sum of is determined as the reconstructed value y0′ of the encoding vector y0.
[0314] As can be seen from the above, when the quantization result of the audio data is dequantized, the decoding end needs to query the vector corresponding to the codebook index from the codebook corresponding to the quantization layer of the target quantizer. In other words, when the quantization result of the audio data is dequantized, the target quantizer needs to be determined. Based on this, the above S202 includes the following steps S202-A1 to S202-A3:
[0315] S202-A1, decoding the bit stream to obtain indication information of the target quantizer;
[0316] S202-A2, determining a target quantizer based on the indication information;
[0317] S202-A3. Use the target quantizer to dequantize the quantization result to obtain the reconstructed coding vectors of N audio data.
[0318] In this implementation, the code stream also includes indication information of the target quantizer, and the indication information is used to indicate the target quantizer. At this time, when the decoder decodes the code stream, in addition to obtaining the quantization result of the audio data, the indication information of the target quantizer is also included. Then, the decoder determines the target quantizer based on the indication information.
[0319] The embodiment of the present application does not limit the specific manner in which the decoding end determines the target quantizer based on the indication information.
[0320] In a possible implementation, if the encoding end determines the target quantizer based on the effective audio information statistics of the audio data and selects from a plurality of preset quantizers with different quantization layers, the indication information may include index information of the target quantizer in the plurality of quantizers with different quantization layers. In this way, the decoding end may select the corresponding quantization layer from the total quantization layers of the preset quantizer based on the index information of the quantization layer to form the target quantizer.
[0321] In a possible implementation, if the encoding end determines the target quantizer and selects the target number of quantization layers from the total quantization layers included in the preset quantizer, the indication information may include index information of the vector quantization layers included in the target quantizer. In this way, the decoding end can select the target quantization layer from the quantization layers included in the preset quantizer based on the index information of the quantization layer included in the target quantizer to obtain the target quantizer.
[0322] After the decoder determines the target quantizer based on the above steps, it uses the target quantizer to dequantize the quantization result to obtain the reconstructed coding vector of the audio data. For example, the decoder searches the codebook corresponding to each quantization layer of the target quantizer based on the codebook subscript corresponding to each quantization layer of the target quantizer included in the quantization result to obtain the quantization vector of each quantization layer of the target quantizer; based on the residual vector corresponding to the last quantization layer of the target quantizer in the quantization result and the quantization vector of each quantization layer of the target quantizer, the reconstructed coding vector of the audio data is obtained.
[0323] For example, the preset quantizer includes 8 quantization layers. The encoder selects 3 quantization layers from the 8 quantization layers to form the target quantizer. At the same time, the encoder also writes the indexes of the selected 3 quantization layers into the bitstream. For example, the encoder writes the residual vector corresponding to the last quantization layer of the target quantizer to In addition to encoding the codebook index index0, index1 and index2 corresponding to each quantization layer, the indexes index0', index1' and index2' of these three quantization layers are also encoded. In this way, the decoder decodes the bitstream and obtains the residual vector corresponding to the last quantization layer of the target quantizer. As well as the codebook index index0, index1 and index2 corresponding to each quantization layer, the indexes index0', index1' and index2' of the three quantization layers can also be obtained. Then, the decoder selects three quantization layers from the quantization layers included in the preset quantizer based on the indexes index0', index1' and index2' of the three quantization layers to form the target quantizer. Fig.10As shown, based on the codebooks corresponding to the three quantization layers and the codebook indexes index0, index1 and index2 corresponding to each quantization layer, the output vectors corresponding to the three quantization layers can be queried, and finally the residual vector corresponding to the last quantization layer of the decoded target quantizer can be obtained. The output vectors corresponding to the three quantization layers are added together to obtain the reconstructed coding vector of the audio data.
[0324] S203. Decode the reconstructed coding vector to obtain a reconstructed value of the audio data.
[0325] After obtaining the reconstructed coding vector of the audio data based on the above steps, the decoding end decodes the reconstructed coding vector to obtain the reconstructed value of the audio data. For example, the reconstructed coding vector of the audio data is input into the decoding network for decoding to obtain the reconstructed value of the audio data.
[0326] The audio decoding method of the embodiment of the present application decodes the bit stream to obtain the quantization result corresponding to the audio data, the quantization result is obtained by quantizing the coding vector of the audio data through a target quantizer, the coding vector of the audio data is obtained by nonlinearly transforming the audio data, and the target quantizer is selected based on the effective audio information statistics of the audio data; the quantization result is dequantized to obtain the reconstructed coding vector of the audio data; the reconstructed coding vector is decoded to obtain the reconstructed value of the audio data. That is, in the embodiment of the present application, the encoding end adaptively selects the most suitable quantizer based on the effective audio information statistics of the input audio data, avoids multiple quantization operations on the data of the low audio signal, and can thereby reduce the codebook index to be decoded, thereby improving the decoding performance of the audio data.
[0327] Combination of the above Figures 5 to 10 , describes in detail the audio encoding and decoding method embodiment of the present application, and the following is combined with Figure 11 to Figure 12 , describe in detail the device embodiments of the present application.
[0328] Fig.11 1 is a schematic block diagram of an audio decoding apparatus provided in an embodiment of the present application. The apparatus 10 can be applied to a decoding device.
[0329] like Fig.11 As shown, the audio decoding device 10 includes:
[0330] A decoding unit 11 is used to decode a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data by a target quantizer, wherein the coding vector of the audio data is obtained by performing a nonlinear transformation on the audio data, and the target quantizer is selected based on a statistical value of effective audio information of the audio data;
[0331] A dequantization unit 12, configured to dequantize the quantization result to obtain a reconstructed coding vector of the audio data;
[0332] The reconstruction unit 13 is used to decode the reconstructed coding vector to obtain the reconstructed value of the audio data.
[0333] In some embodiments, the effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
[0334] In some embodiments, the target quantizer is determined based on a target quantization level of a quantizer corresponding to the audio data, and the target quantization level is determined based on the statistical value of the effective audio information.
[0335] In some embodiments, if the effective audio information statistical value is less than a first preset value, the target quantization layer number is the preset first quantization layer number; or, the target quantization layer number is the second quantization layer number corresponding to the effective audio information statistical value in the first corresponding relationship, and the first corresponding relationship includes the relationship between different effective audio information statistical value intervals and different quantization layer numbers; or, the target quantization layer number is determined based on the second corresponding relationship and the effective audio information statistical value, and the second corresponding relationship is the relationship between different effective audio information statistical values and different quantization layer numbers.
[0336] In some embodiments, the target quantizer is composed of quantization layers of the target number of quantization layers selected from a total number of quantization layers of a preset quantizer.
[0337] In some embodiments, among the target number of quantization layers selected from the total number of quantization layers of the preset quantizer, at least two quantization layers are non-adjacent quantization layers.
[0338] In some embodiments, the decoding unit 11 is also used to decode the bit stream to obtain indication information of the target quantizer; the dequantization unit 12 is specifically used to determine the target quantizer based on the indication information; and use the target quantizer to dequantize the quantization result to obtain a reconstructed coding vector of the audio data.
[0339] In some embodiments, the indication information includes index information of the quantization layer included in the target quantizer, and the inverse quantization unit 12 is specifically used to select a corresponding quantization layer from the total quantization layers of a preset quantizer based on the index information of the quantization layer to form the target quantizer.
[0340] In some embodiments, the inverse quantization unit 12 is specifically used to query in the codebook corresponding to each quantization layer of the target quantizer based on the codebook subscript corresponding to each quantization layer of the target quantizer included in the quantization result, to obtain the quantization vector of each quantization layer of the target quantizer; based on the residual vector corresponding to the last quantization layer of the target quantizer in the quantization result, and the quantization vector of each quantization layer of the target quantizer, obtain the reconstructed coding vector of the audio data.
[0341] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Fig.11 The device shown can execute the above-mentioned audio decoding method embodiment, and the above-mentioned and other operations and / or functions of each module in the device are respectively for realizing the above-mentioned method embodiment, which will not be described in detail here for the sake of brevity.
[0342] Fig.12 20 is a schematic block diagram of an audio encoding device provided in an embodiment of the present application. The device 20 can be applied to an encoding device.
[0343] like Fig.12 As shown, the audio encoding device 20 includes:
[0344] An acquisition unit 21 is used to acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data;
[0345] A determination unit 22, configured to perform statistics on the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and select a target quantizer based on the statistical value of the effective audio information;
[0346] The quantization unit 23 is used to quantize the coding vector of the audio data using the target quantizer to obtain a quantization result, and encode the quantization result to obtain a code stream.
[0347] In some embodiments, the effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
[0348] In some embodiments, if the effective audio information statistics of the audio data include the energy value of the audio data, the determination unit 22 is specifically used to determine the energy value of each audio frame in the audio data; determine the energy value of the audio data based on the energy value of each audio frame; and determine the effective audio information statistics of the audio data based on the energy value of the audio data.
[0349] In some embodiments, the determination unit 22 is specifically configured to determine, for an i-th audio frame in the audio data, a sum of squares of sample data included in the i-th audio frame as an energy value of the i-th audio frame, where i is a positive integer.
[0350] In some embodiments, if the effective audio information statistics of the audio data include the spectrum envelope statistics of the audio data, the determination unit 22 is specifically used to determine the spectrum envelope statistics of each audio frame in the audio data; determine the spectrum envelope statistics of the audio data based on the spectrum envelope statistics of each audio frame; and determine the effective audio information statistics of the audio data based on the spectrum envelope statistics of the audio data.
[0351] In some embodiments, the determination unit 22 is specifically used to convert the i-th audio frame in the audio data from a time domain signal to a frequency domain signal, where i is a positive integer; and perform statistics on the spectrum envelope of the frequency domain signal of the i-th audio frame to obtain the spectrum envelope statistics of the i-th audio frame.
[0352] In some embodiments, if the effective audio information statistics of the audio data include the zero-crossing rate of the audio data, the determination unit 22 is specifically used to determine the zero-crossing rate of each audio frame in the audio data; determine the zero-crossing rate of the audio data based on the zero-crossing rate of each audio frame; and determine the effective audio information statistics of the audio data based on the zero-crossing rate of the audio data.
[0353] In some embodiments, the determination unit 22 is specifically configured to determine a target quantization level number of a quantizer corresponding to the audio data based on the statistical value of the effective audio information; and select the target quantizer based on the target quantization level number.
[0354] In some embodiments, the determination unit 22 is specifically used to determine a preset first quantization layer number as the target quantization layer number if the effective audio information statistical value is less than a first preset value; or, obtain a first correspondence between different effective audio information statistical value intervals and different quantization layer numbers, and determine the second quantization layer number corresponding to the effective audio information statistical value in the first correspondence as the target quantization layer number; or, obtain a second correspondence between different effective audio information statistical values and different quantization layer numbers, and determine the target quantization layer number based on the second correspondence and the effective audio information statistical value.
[0355] In some embodiments, the determination unit 22 is specifically configured to select quantization layers of the target number of quantization layers from a total number of quantization layers of a preset quantizer to form the target quantizer.
[0356] In some embodiments, among the quantization layers of the target number of quantization layers selected from the total quantization layers of the preset quantizer, at least two quantization layers are non-adjacent quantization layers.
[0357] In some embodiments, the quantization unit 23 is specifically used to input the encoding vector of the audio data into the first quantization layer of the target quantizer for quantization, and obtain a first quantization vector and a codebook subscript corresponding to the first quantization layer; based on the encoding vector of the audio data and the first quantization vector, obtain a first residual vector; input the first residual vector into the second quantization layer of the target quantizer for quantization, repeat iterations, and obtain the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook subscript corresponding to each quantization layer in the target quantizer; encode the residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer to obtain the code stream.
[0358] In some embodiments, the bitstream also includes indication information for indicating the target quantizer.
[0359] In some embodiments, the indication information includes index information of a quantization layer included in the target quantizer.
[0360] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Fig.12 The device shown can execute the above-mentioned audio encoding method embodiment, and the above-mentioned and other operations and / or functions of each module in the device are respectively for realizing the above-mentioned method embodiment, which will not be described in detail here for the sake of brevity.
[0361] The above describes the device of the embodiment of the present application from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software modules in the decoding processor to perform. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.
[0362] Fig.13is a schematic block diagram of an electronic device provided in an embodiment of the present application, Fig.13 The electronic device may be the above-mentioned encoding device or the decoding device.
[0363] like Fig.13 As shown, the electronic device 30 may include:
[0364] The memory 31 and the processor 32, the memory 31 is used to store the computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiment of the present application.
[0365] For example, the processor 32 may be configured to execute the steps in the method 200 according to the instructions in the computer program 33 .
[0366] In some embodiments of the present application, the processor 32 may include but is not limited to:
[0367] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0368] In some embodiments of the present application, the memory 31 includes but is not limited to:
[0369] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0370] In some embodiments of the present application, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording pages provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.
[0371] like Fig.13 As shown, the electronic device 30 may further include:
[0372] The transceiver 34 may be connected to the processor 32 or the memory 31 .
[0373] The processor 32 may control the transceiver 34 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.
[0374] It should be understood that the various components in the computing device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0375] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.
[0376] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.
[0377] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a solid state drive (solid state disk, SSD)), etc.
[0378] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0379] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0380] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0381] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An audio decoding method, characterized in that: include: Decoding a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data by a target quantizer, wherein the coding vector is obtained by performing a nonlinear transformation on the audio data, and the target quantizer is selected based on a statistical value of effective audio information of the audio data; De-quantizing the quantization result to obtain a reconstructed coding vector of the audio data; The reconstructed coding vector is decoded to obtain a reconstructed value of the audio data.
2. The method according to claim 1, characterized in that: The effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
3. The method according to claim 1 or 2, characterized in that: The target quantizer is determined based on a target quantization level of a quantizer corresponding to the audio data, and the target quantization level is determined based on a statistical value of the effective audio information.
4. The method according to claim 3, characterized in that If the effective audio information statistical value is less than a first preset value, the target quantization layer number is a preset first quantization layer number; or, The target quantization level is the second quantization level corresponding to the effective audio information statistical value in the first corresponding relationship, wherein the first corresponding relationship includes the relationship between different effective audio information statistical value intervals and different quantization levels; or, The target number of quantization levels is determined based on a second corresponding relationship and the effective audio information statistical value, where the second corresponding relationship is a relationship between different effective audio information statistical values and different numbers of quantization levels.
5. The method according to claim 3, characterized in that: The target quantizer is composed of quantization layers of the target number of quantization layers selected from the total quantization layers of the preset quantizer, and at least two of the quantization layers are non-adjacent quantization layers.
6. The method according to claim 1 or 2, characterized in that: The step of inverse quantizing the quantization result to obtain a reconstructed coding vector of the audio data includes: Decoding the bitstream to obtain indication information of the target quantizer, wherein the indication information includes index information of a quantization layer included in the target quantizer; Based on the indication information, a corresponding quantization layer is selected from a total quantization layer of a preset quantizer to determine the target quantizer; The target quantizer is used to inversely quantize the quantization result to obtain a reconstructed coding vector of the audio data.
7. The method according to claim 6, characterized in that The step of using the target quantizer to dequantize the quantization result to obtain a reconstructed coding vector of the audio data includes: Based on the codebook subscript corresponding to each quantization layer of the target quantizer included in the quantization result, query in the codebook corresponding to each quantization layer of the target quantizer to obtain a quantization vector of each quantization layer of the target quantizer; Based on the residual vector corresponding to the last quantization layer of the target quantizer in the quantization result and the quantization vector of each quantization layer of the target quantizer, a reconstructed coding vector of the audio data is obtained.
8. An audio encoding method, characterized in that: include: Acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data; Performing statistics on the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and selecting a target quantizer based on the statistical value of the effective audio information; The target quantizer is used to quantize the coding vector to obtain a quantization result, and the quantization result is encoded to obtain a code stream.
9. The method according to claim 8, characterized in that The effective audio information statistics of the audio data include at least one of an energy value of the audio data, a spectrum envelope statistics of the audio data, and a zero-crossing rate of the audio data.
10. The method according to claim 9, characterized in that If the effective audio information statistics of the audio data include the energy value of the audio data, then the effective audio information of the audio data is counted to obtain the effective audio information statistics of the audio data, including: Determining an energy value of each audio frame in the audio data; Determining an energy value of the audio data based on the energy value of each audio frame; Based on the energy value of the audio data, a statistical value of effective audio information of the audio data is determined.
11. The method according to claim 10, characterized in that The determining the energy value of each audio frame in the audio data comprises: For an i-th audio frame in the audio data, a sum of squares of sample data included in the i-th audio frame is determined as an energy value of the i-th audio frame, where i is a positive integer.
12. The method according to claim 9, characterized in that If the effective audio information statistics of the audio data include the spectrum envelope statistics of the audio data, then the effective audio information of the audio data is counted to obtain the effective audio information statistics of the audio data, including: Determining a spectral envelope statistic of each audio frame in the audio data; Determine a spectral envelope statistic of the audio data based on a spectral envelope statistic of each audio frame; Based on the spectrum envelope statistics of the audio data, the effective audio information statistics of the audio data are determined.
13. The method according to claim 12, characterized in that The determining of the spectrum envelope statistics of each frame of audio data included in the audio data comprises: For an i-th audio frame in the audio data, convert the i-th audio frame from a time domain signal to a frequency domain signal, where i is a positive integer; Statistics are performed on the frequency spectrum envelope of the frequency domain signal of the i-th audio frame to obtain a statistical value of the frequency spectrum envelope of the i-th audio frame.
14. The method according to claim 9, characterized in that If the effective audio information statistics of the audio data include the zero-crossing rate of the audio data, then the effective audio information of the audio data is counted to obtain the effective audio information statistics of the audio data, including: Determining a zero crossing rate of each audio frame in the audio data; Determining a zero-crossing rate of the audio data based on a zero-crossing rate of each audio frame; Based on the zero-crossing rate of the audio data, a statistical value of effective audio information of the audio data is determined.
15. The method according to any one of claims 8 to 14, characterized in that: The step of selecting a target quantizer based on the effective audio information statistics comprises: Determining a target quantization level of a quantizer corresponding to the audio data based on the effective audio information statistics; Based on the target number of quantization levels, the target quantizer is selected.
16. The method according to claim 15, characterized in that The step of selecting a target quantizer based on the effective audio information statistics comprises: If the effective audio information statistical value is less than a first preset value, the preset first quantization layer number is determined as the target quantization layer number; or, Acquire a first correspondence between different valid audio information statistical value intervals and different quantization levels, and determine the second quantization level corresponding to the valid audio information statistical value in the first correspondence as the target quantization level; or, A second correspondence between different effective audio information statistics and different quantization levels is obtained, and the target quantization level is determined based on the second correspondence and the effective audio information statistics.
17. The method according to any one of claims 8 to 14, characterized in that: The step of quantizing the coding vector using the target quantizer to obtain a quantization result, and encoding the quantization result to obtain a code stream includes: Inputting the encoding vector into a first quantization layer of the target quantizer for quantization, and obtaining a first quantization vector and a codebook subscript corresponding to the first quantization layer; Obtaining a first residual vector based on the encoding vector and the first quantization vector; Input the first residual vector into the second quantization layer of the target quantizer for quantization, repeat the iteration, and obtain the residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer; The residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer are encoded to obtain the code stream.
18. An audio decoding device, characterized in that: include: A decoding unit, configured to decode a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data by a target quantizer, wherein the coding vector is obtained by performing a nonlinear transformation on the audio data, and the target quantizer is selected based on a statistical value of effective audio information of the audio data; A dequantization unit, used for dequantizing the quantization result to obtain a reconstructed coding vector of the audio data; The reconstruction unit is used to decode the reconstructed coding vector to obtain a reconstructed value of the audio data.
19. An audio encoding device, characterized in that: include: An acquisition unit, used to acquire audio data to be encoded, and perform nonlinear transformation on the audio data to obtain an encoding vector of the audio data; a determination unit, configured to perform statistics on the effective audio information of the audio data to obtain a statistical value of the effective audio information of the audio data, and select a target quantizer based on the statistical value of the effective audio information; The quantization unit is used to quantize the coding vector using the target quantizer to obtain a quantization result, and encode the quantization result to obtain a code stream.
20. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1 to 7 or 8 to 18.