Audio signal encoding apparatus and decoding apparatus, and encoding method and decoding method
Patent Information
- Application Number
- CN202111171436.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-10-28
- Filing Date
- 2015-07-03
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2035-07-03
AI Technical Summary
但是,该途径(approach)与电波和频带的有效利用相反
[0033] The encoding and decoding devices according to the present invention can reduce the overall bit rate and can also encode and decode high-quality audio signals.
Smart Images

Figure CN114023341B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on July 3, 2015, with application number 201580015301.4 and entitled "Audio Signal Encoding Device, Audio Signal Decoding Device, Audio Signal Encoding Method and Audio Signal Decoding Method". Technical Field
[0002] This invention relates to encoding and decoding techniques for improving the sound quality of acoustic signals, such as speech or music signals. Background Technology
[0003] Encoding techniques that compress audio signals at low bit rates are crucial for the efficient use of radio waves in mobile communications. Furthermore, expectations for improved voice call quality have been rising in recent years, with a desire for highly immersive call services. To achieve such services, wide-bandwidth audio signals could be encoded at high bit rates. However, this approach contradicts the efficient utilization of radio waves and frequency bands.
[0004] Here, as an example, we discuss the audio signal encoding technology used in the G.719 standard (Non-Patent Document 1).
[0005] In the G.719 standard, when encoding audio signals, a specified number of bits are allocated to the spectrum after frequency conversion of the audio signal. Specifically, the spectrum is divided into sub-bands with a specified bandwidth, and the units (the necessary number of bits) used for sequential quantization via dot vector quantization, starting from the sub-band with the higher energy, are allocated as follows.
[0006] (1) Allocate 1 unit to the subband with the highest energy from the entire subband.
[0007] Since 1 bit is allocated for each spectrum, for example, if the number of spectrum samples in a subband is 8, then 1 unit is 8 bits (also, the maximum number of bits that can be allocated for each spectrum is 9 bits, for example, if the number of spectrum samples in a subframe is 8, then it can be allocated up to 72 bits).
[0008] (2) Allocating 1 unit of subband will reduce the quantization subband energy by 2 levels (6dB). If the number of bits allocated to a subband of 1 unit exceeds the maximum value (9 bits), it will be removed from the quantization object in the next subsequent cycle.
[0009] (3) Return to (1) above and repeat the same process.
[0010] Figure 6This represents the subband energy within each subband. The horizontal axis represents frequency, and the vertical axis represents amplitude on a logarithmic scale. In the graph, subband energy is represented by horizontal lines instead of dots, with the width of each line representing the bandwidth of the subband.
[0011] Figure 7 , Figure 8 These are graphs showing examples of bit allocation results for each subband when using the coding method specified in the G.719 standard. The horizontal axis of each graph represents frequency, and the vertical axis represents the number of bits allocated. Furthermore, Figure 7 This is for a bit rate of 128k bit / s. Figure 8 This is for a bit rate of 64k bit / s.
[0012] At 128k bit / s, there is an abundance of available bit assets, so for many subbands (spectrums), the maximum available bit size is 9 bits, which can guarantee high-quality audio signals.
[0013] In contrast, at 64k bit / s, although no subband with the maximum value of 9 bits is allocated, there is also no subband without allocated bits. This can be said to suppress the degradation of audio signal quality and also take into account the effective use of radio waves and frequency bands.
[0014] Existing technical documents
[0015] Patent documents
[0016] Patent Document 1: Japanese Patent Publication No. 2013-534328; Patent Document 2: International Publication No. 2005 / 027095
[0017] Non-patent literature
[0018] Non-patent literature 1: ITU-T Standard G.719, 2008 Summary of the Invention
[0019] However, there is a need to achieve further efficient utilization of radio waves and frequency bands. Here, when using the method described above in the G.719 standard to encode an audio signal with a sampling frequency of around 32kHz at a low bit rate of around 20kbps or less, there is a problem that the number of units (bits) used to quantize all subbands cannot be guaranteed.
[0020] Figure 9 This diagram illustrates an example of bit allocation for each subband when using the encoding method defined by the G.719 standard at 20k bit / s. As a result, the high-frequency band may not be allocated bits, and consequently, the auditoryally important low-frequency band may also be excluded. Consequently, the spectrum in this subband cannot be encoded, leading to a significant degradation in the quality of the audio signal.
[0021] In contrast, a method that dynamically changes the bit allocation method is also considered (Patent Document 1).
[0022] However, changing the bit allocation method with a single encoding method (quantization method) without changing the encoding method (quantization method) has its limitations in dealing with the degradation of audio signal quality.
[0023] This invention provides encoding and decoding techniques for reducing the overall bit rate and achieving high-quality audio signals.
[0024] The audio signal encoding device of the present invention includes: a time-frequency conversion unit that converts the input audio signal to the frequency domain and generates a spectrum, divides the spectrum into sub-bands of each defined frequency band, and outputs the sub-band spectrum; a sub-band energy quantization unit that quantizes the sub-band energy for each sub-band; a pitch calculation unit that analyzes the tonality of the sub-band spectrum and outputs the analysis result; a bit allocation unit that, based on the tonality analysis result and the quantized sub-band energy, selects a second sub-band quantized by the second quantization unit from the sub-bands and determines the number of first bits allocated to the first sub-band quantized by the first quantization unit; and a multiplexing unit that multiplexes and outputs information containing the encoding information output from the first quantization unit and the second quantization unit, the quantized sub-band energy, and the tonality analysis result. The first quantization unit performs pulse encoding on the sub-band spectrum contained in the first sub-band using bits composed of the number of first bits, and the second quantization unit encodes the sub-band spectrum contained in the second sub-band using a pitch filter.
[0025] The audio signal encoding apparatus of the present invention includes: a time-frequency conversion unit that generates a spectrum by performing a frequency domain conversion on an input audio signal, divides the spectrum into sub-bands of a predetermined frequency band, and outputs the sub-band spectrum, wherein the input audio signal includes a music signal, a speech signal, or a signal that is a mixture of a music signal and a speech signal; a sub-band energy quantization unit that quantizes the sub-band energy of each sub-band; a pitch calculation unit that analyzes the pitch of the sub-band spectrum and outputs the analysis result; a bit allocation unit that, based on the pitch analysis result and the quantized sub-band energy, selects a second sub-band from the sub-bands that is quantized by a second quantization unit, and determines the number of first bits to be allocated to the first sub-band that is quantized by a first quantization unit; and a multiplexing unit that takes the encoded information output from the first quantization unit and the second quantization unit, the quantized sub-band energy, and the pitch analysis... The result is multiplexed into information and multiplexed information is output, wherein the first quantization unit encodes the subband spectrum within the subband spectrum contained in the first subband by using the first number of bits; and the second quantization unit encodes the subband spectrum within the subband spectrum contained in the second subband by using an encoding method, the encoding method comprising: identifying the second subband to which the second quantization unit intends to perform quantization; searching for subbands or frequency bands, wherein the normalized subband spectrum of the identified second subband and the quantized spectrum of the subband or the frequency band have the greatest correlation; and using the position of the subband or the frequency band to generate hysteresis information of the second subband; encoding the hysteresis information to generate encoded information output from the second quantization unit; wherein the hysteresis information is the absolute position of the subband or the frequency band, or the hysteresis information is the relative position of the subband or the frequency band, or the hysteresis information is the number of the subband.
[0026] The audio signal decoding apparatus for decoding encoded information of the present invention includes: a separation unit (201) for separating encoded information into first encoded information, second encoded information, quantized sub-band energy obtained by quantizing the energy of each sub-band in the sub-band, and an analysis result of the pitch of each sub-band calculated in the sub-band; and a bit allocation unit (203) for selecting a second sub-band from the sub-bands to be decoded by the second decoding unit (205) based on the pitch analysis result and the quantized sub-band energy, and determining a first sub-band to be allocated to the sub-bands to be decoded by the first decoding unit (204). The first bit number; and the frequency-time conversion unit (207), which generates and outputs an output audio signal by performing a time-domain conversion on the spectrum output from the second decoding unit (205), the first decoding unit (204) generates a first decoded spectrum by decoding the first encoded information using the first bit number, the second decoding unit (205) generates a second decoded spectrum by decoding the second encoded information, and the second decoding unit generates a regenerated spectrum by combining the second decoded spectrum and the first decoded spectrum, wherein the output audio signal includes a music signal, a speech signal, or a signal that is a mixture of music and speech signals.
[0027] The terminal device of the present invention includes: an audio signal encoding device as described above; and an antenna for transmitting encoded information.
[0028] The terminal device of the present invention includes: an antenna that receives encoded information and outputs it to a separation unit (201); and an audio signal decoding device as described above.
[0029] The audio signal encoding method of the present invention includes: generating a spectrum by performing a frequency domain conversion on an input audio signal, wherein the input audio signal includes a music signal, a speech signal, or a signal that is a mixture of music and speech signals; dividing the spectrum into sub-bands of a predetermined frequency band and outputting the sub-band spectrum; quantizing the sub-band energy of each sub-band; analyzing the tonality of the sub-band spectrum and outputting the analysis result; selecting a second sub-band from the sub-bands based on the tonality analysis result and the quantized sub-band energy; determining the number of first bits to be allocated to the first sub-band among the sub-bands; generating first encoded information by encoding the sub-band spectrum contained in the first sub-band using the number of first bits; and encoding the sub-band spectrum using an encoding method. The second sub-band is used to generate second coded information from the sub-band spectrum within the sub-band spectrum contained in the second sub-band, wherein the encoding method includes: identifying the second sub-band to which the second quantization unit is to perform quantization; searching for sub-bands or frequency bands, wherein the normalized sub-band spectrum of the identified second sub-band and the quantized spectrum of the sub-band or the frequency band have the greatest correlation; generating hysteresis information of the second sub-band using the position of the sub-band or the frequency band, wherein the hysteresis information is the absolute position of the sub-band or the frequency band, or the hysteresis information is the number of the sub-band; encoding the hysteresis information to generate the second coded information; and multiplexing the first coded information and the second coded information together and outputting them.
[0030] The present invention discloses a method for decoding audio signals of decoded coded information. The method includes: separating coded information into first coded information, second coded information, quantized sub-band energy obtained by quantizing the energy of each sub-band, and an analysis result of the pitch of each sub-band; selecting a second sub-band from the sub-bands based on the analysis result of the pitch and the quantized sub-band energy; determining a first number of bits to be allocated to the first sub-band; generating a first decoded spectrum by decoding the first coded information using the first number of bits; generating a second decoded spectrum by decoding the second coded information; generating a regenerated spectrum by combining the second decoded spectrum and the first decoded spectrum; and generating and outputting an output audio signal by performing a time-domain conversion on the regenerated spectrum, wherein the output audio signal includes a music signal, a speech signal, or a signal that is a mixture of a music signal and a speech signal.
[0031] The present invention provides a computer program product having program code for the method described above.
[0032] Furthermore, these general and specific methods can also be implemented through systems, methods, integrated circuits, or computer programs, or through any combination of systems, devices, methods, integrated circuits, and computer programs.
[0033] The encoding and decoding devices according to the present invention can reduce the overall bit rate and can also encode and decode high-quality audio signals. Attached Figure Description
[0034] Figure 1 This is a structural diagram of the encoding device in Embodiment 1 of the present invention.
[0035] Figure 2 This is a detailed structural diagram of the bit allocation unit of the encoding device in Embodiment 1 of the present invention.
[0036] Figure 3 This is an explanatory diagram illustrating the operation of the encoding device in Embodiment 1 of the present invention.
[0037] Figure 4 This is a structural diagram of the decoding device in Embodiment 2 of the present invention.
[0038] Figure 5 This is a detailed structural diagram of the bit allocation unit of the decoding device in Embodiment 2 of the present invention.
[0039] Figure 6 This is an explanatory diagram illustrating the subband energy in existing coding devices.
[0040] Figure 7 This is an explanatory diagram illustrating the bit allocation results of subbands in a prior art encoding device.
[0041] Figure 8 This is an explanatory diagram illustrating the bit allocation results of subbands in a prior art encoding device.
[0042] Figure 9 This is an explanatory diagram illustrating the bit allocation results of subbands in a prior art encoding device. Detailed Implementation
[0043] Hereinafter, the structure and operation of embodiments of the present invention will be described with reference to the accompanying drawings. Furthermore, the input signal to the encoding device of the present invention and the output signal from the decoding device, i.e., the audio signal, are concepts including speech signals, wider-bandwidth music signals, and signals that combine them.
[0044] In this invention, "input audio signal" refers to a combination of music and speech signals, or a signal that includes a mixture of both. Furthermore, "quantized subband energy" is the sum or average of the energy of the subband spectrum within a subband, i.e., the energy obtained by quantizing the subband energy. Subband energy can be calculated, for example, as the sum of squares of the subband spectra within a subband. "Pitchiness" refers to the degree to which the peak value of the spectrum is established in a specific frequency component; its analysis results can be expressed numerically or symbolically. "Pulse coding" refers to coding that approximates the spectrum using pulses.
[0045] "Relatively low" refers to being lower than other subbands, such as being lower than the average of all subbands or being lower than a specified value. "High-frequency subband" refers to a subband located on the high-frequency side among multiple subbands.
[0046] Furthermore, the terms first (spectrum) quantization unit, second (spectrum) quantization unit, first (spectrum) decoding unit, second (spectrum) decoding unit, first subband, second subband, third subband, fourth subband, first number of bits, second number of bits, third number of bits, and fourth number of bits, as described in the embodiments and claims, respectively represent categories and do not represent order.
[0047] (Implementation Method 1) Figure 1 This is a block diagram illustrating the structure and operation of the audio signal encoding device 100 according to Embodiment 1. Figure 1 The audio signal encoding device 100 shown includes a time-frequency conversion unit 101, a sub-band energy quantization unit 102, a pitch calculation unit 103, a bit allocation unit 104, a normalization unit 105, a first spectrum quantization unit 106, a second spectrum quantization unit 107, and a multiplexing unit 108. Furthermore, an antenna A is connected to the multiplexing unit 108. Moreover, combining the audio signal encoding device 100 and the antenna A constitutes a terminal device or a base station device.
[0048] The time-frequency conversion unit 101 converts the input audio signal in the time domain to the frequency domain, generating the spectrum of the input audio signal (hereinafter referred to as "spectrum"). As an example of time-frequency conversion, MDCT (Improved Discrete Cosine Transform) can be cited, but it is not limited to this. For example, DCT (Discrete Cosine Transform), DFT (Discrete Fourier Transform), Fourier Transform, etc. can also be used.
[0049] Furthermore, the time-frequency conversion unit 101 divides the spectrum into predetermined frequency bands, or sub-bands. These predetermined frequency bands can be either equally spaced or, for example, wider in high-frequency bands and narrower in low-frequency bands.
[0050] Then, the time-frequency conversion unit 101 outputs the spectrum divided into each sub-band as the sub-band spectrum to the sub-band energy quantization unit 102, the pitch calculation unit 103, and the normalization unit 105.
[0051] The subband energy quantization unit 102 calculates the energy of the subband spectrum (i.e., the subband energy) for each subband, quantizes it, and obtains the quantized subband energy. Specifically, the subband energy can be calculated by the sum of squares of the subband spectra within the subband, but it is not limited to this. For example, the subband energy can be calculated by integrating the amplitude of the subband spectrum for each subband. Furthermore, when averaging the subband energy, the sum of squares is divided by the number of spectra within the subband (subband width). Then, the subband energy obtained in this way is quantized according to a specified pitch width.
[0052] Then, the obtained quantized subband energy is output to the normalization unit 105 and the bit allocation unit 104, and the encoded quantized subband energy after encoding is output to the multiplexing unit 108.
[0053] The pitch calculation unit 103 analyzes the sub-band spectrum contained in each sub-band to determine the pitch characteristic. Pitch characteristic refers to the degree to which the peak of the spectrum exists in a specific frequency component; it is a concept of peak characteristic, which implies the presence of a significant peak. For example, it can be quantitatively calculated as the ratio of the amplitude of the average spectrum within the target sub-band to the amplitude of the largest spectrum within that sub-band. If this value exceeds a predetermined threshold, the spectrum of that sub-band is defined as having pitch characteristic (peak characteristic). In this embodiment, a "1" is generated as a peak / pitch flag when the value exceeds the predetermined threshold, and a "0" is generated as a peak / pitch flag when the value is below the predetermined threshold. This is output as the analysis result to the bit allocation unit 104 and the multiplexing unit 108. Of course, the above ratio can also be directly output as the analysis result.
[0054] The meaning of the pitch calculation unit is as follows.
[0055] At low bit rates, for efficient quantization of subbands where the spectrum energy is dispersed across the entire subband spectrum (e.g., using the low-frequency spectrum to represent the high-frequency spectrum), the use of pitch filter-based methods (i.e., methods that use the low-frequency spectrum to represent the high-frequency spectrum) is effective. Therefore, the energy dispersion within a subband is determined by a measure of the peak / tonality of the spectrum (such as the ratio of peak power to average power), and subbands with low peak / tonality become the targets for pitch filter-based quantization.
[0056] Bit allocation unit 104, referring to the quantized subband energy and peak / pitch flags of each subband, allocates bits from the bit assets, which represent the total number of bits available for encoding, for the subband spectrum in each subband. Specifically, it calculates and determines the number of bits allocated to the subband quantized by the first spectrum quantization unit, i.e., the first subband, and outputs it to the first spectrum quantization unit 106 as bit allocation information. Furthermore, it selects and determines the subband quantized by the second spectrum quantization unit 107, i.e., the second subband, and outputs it to the second spectrum quantization unit 107 as the quantization mode.
[0057] The details of the structure and operation of the bit allocation unit 104 will be described later.
[0058] Furthermore, in this embodiment, the bit allocation unit 104 references the peak / tone flag and the quantized subband energy of each subband in an order that is arbitrary.
[0059] Furthermore, the second sub-band that becomes the object of quantization in the second spectral quantization unit 107 can be the entire frequency band as a candidate. However, generally, the frequency bands with lower energy and lower tonality of the quantized sub-band are mainly high-frequency bands. Therefore, it is also possible to select only the sub-bands that exist in specific high-frequency bands. For example, it is possible to select only 4 or 5 sub-bands of the high-frequency band as objects.
[0060] Alternatively, typically, the low-frequency band has a higher tonality, while the high-frequency band has a lower tonality. Therefore, the high-frequency sub-bands of the audio signal essentially become the object of quantization based on the pitch filter. Thus, it is also possible to select sub-bands based on tonality, with the entire high-frequency side becoming the object of pitch filter quantization, and only transmitting the signal of that sub-band as the quantization mode.
[0061] Normalization unit 105 generates normalized subband spectra by normalizing (dividing) the spectra of each subband with respect to the input quantized subband energy. This normalizes the differences in amplitude between subbands. Normalization unit 105 then outputs the normalized subband spectra to the first spectral quantization unit 106 and the second spectral quantization unit 107.
[0062] Furthermore, the normalization unit 105 has an arbitrary structure.
[0063] Furthermore, the normalization unit 105 is a single structure in this embodiment, but it can also be configured as two units in front of the first spectrum quantization unit 106 and the second spectrum quantization unit 107 respectively.
[0064] The first spectrum quantization unit 106 is an example of a first quantization unit. It uses bits consisting of the first number of bits allocated by the bit allocation unit 104 to quantize the subband spectrum of the input normalized subband spectrum that belongs to the first subband to be quantized by the first spectrum quantization unit 106. Then, the quantization result is output as a quantized spectrum to the second spectrum quantization unit 107, and the quantized spectrum is encoded and the generated first encoded information is output to the multiplexing unit 108.
[0065] The first spectral quantization unit 106 uses a pulse coding unit, but examples of pulse coding units include a dot vector quantization unit that performs dot vector quantization and a pulse coding unit that performs pulse coding that approximates the subband spectrum with a few pulses. That is, any quantization unit can be used as long as it is a quantization method suitable for quantizing a high-tonality spectrum or a method that performs quantization with a few pulses.
[0066] Furthermore, at very low bit rates, it is possible to expect better sound quality by using pulse coding quantization, which approximates the subband spectrum with fewer pulses than lattice vector quantization.
[0067] The second spectral quantization unit 107 is an example of a second quantization unit, for example, which can employ an extended band (prediction model of the pitch filter) quantization method.
[0068] Here, the pitch filter is a processing block that performs the processing represented by Equation 1 below.
[0069] (1)
[0070] Generally, a pitch filter refers to a filter that enhances the pitch period (T) of the signal on the time axis (enhancing the pitch component on the frequency axis). When the number of taps is 1, for a discrete signal x[i], it is, for example, a digital filter represented by Equation 1. However, in this embodiment, the pitch filter is defined as a processing block that performs the processing represented by Equation 1, and it is not necessary to enhance the pitch of the signal on the time axis.
[0071] In this embodiment, the pitch filter (the processing block represented by Equation 1) is applied to quantize the MDCT coefficient sequence Mq[i]. Specifically, in Equation 1, x[i] = 0 (i ≥ K, where K is the lower frequency limit of the MDCT coefficients to be encoded), y[i] = Mq[i] (i < K), and y[i] is calculated (K ≤ i ≤ K', where K' is the upper frequency limit of the MDCT coefficients to be encoded). The T with the smallest error between the MDCT coefficients Mt[i] to be encoded and the calculated y[i] is encoded as hysteresis information. Spectral coding based on such a pitch filter is disclosed in Patent Document 2, etc.
[0072] The second spectrum quantization unit 107 determines the second sub-band (normalized sub-band spectrum) to be quantized by the second spectrum quantization unit 107 with reference to the quantization mode. Thus, K and K' are determined. Then, the sub-band or frequency band with the largest correlation between the normalized sub-band spectrum (Mt[i], equivalent to K≤i≤K') and the quantized spectrum (Mq[i], equivalent to i<K) in the determined second sub-band (frequency K~K') is searched, and this position is generated as hysteresis information (equivalent to T). Hysteresis information, for example, may include the absolute or relative position of the sub-band or frequency band, or the sub-band number. Then, the second spectrum quantization unit 107 encodes the hysteresis information and outputs it as second encoded information to the multiplexing unit 108.
[0073] Furthermore, in this embodiment, the encoded quantized subband energy is multiplexed and transmitted by the multiplexing unit 108, and gain can be generated on the decoding unit side, so the gain is not encoded. However, the gain can also be encoded and transmitted. At that time, the gain between the subbands that have the largest quantized spectrum related to the second subband to be quantized is calculated, and the second spectrum quantization unit 107 encodes the hysteresis information and gain, and outputs it as the second encoded information to the multiplexing unit 108.
[0074] Furthermore, generally speaking, the bandwidth of high-frequency subbands is set wider than that of low-frequency subbands. However, for a portion of the copied low-frequency subband, due to its lower energy, there may be instances where the subband is not quantized by the lattice vector. In such cases, that subband can be treated as zero-spectrum, or noise can be added to avoid abrupt changes in the spectrum between subbands.
[0075] The multiplexing unit 108 multiplexes the quantized subband energy, the first coding information, the second coding information, and the peak / tone flag, and outputs them as coding information to antenna A.
[0076] Then, antenna A sends encoded information to the audio signal decoding device. The encoded information reaches the audio signal decoding device via various nodes or base stations.
[0077] Next, the details of bit allocation unit 104 will be explained.
[0078] Figure 2 This is a block diagram showing the detailed structure and operation of the bit allocation unit 104 of the audio signal encoding device 100 in Embodiment 1. Figure 2 The bit allocation unit 104 shown consists of a bit library 111, a bit library 112, a bit allocation calculation unit 113, and a quantization mode determination unit 114.
[0079] The bit library 111 refers to the output of the tone calculation unit 103, namely the peak / tone flag, to ensure the number of bits required for the second spectral quantization performed by the second spectral quantization unit 107 when the peak / tone flag is "0".
[0080] In this embodiment, the number of bits required for encoding hysteresis information is ensured based on the pitch filter. Furthermore, the ensured number of bits is removed from the total number of bits available for quantization, i.e., the bit assets, and the remaining bit assets are output to the bit library 112. Additionally, bit assets are supplied from the subband energy quantization unit 102, meaning that bits remaining after removing the number of bits required for variable-length encoding of the quantized subband energy can be used for quantization (encoding) of the first spectrum quantization unit 106, the second spectrum quantization unit 107, and the peak / pitch flag. The subband energy quantization unit 102 is not limited to generating bit asset information.
[0081] Bit library 112 ensures the number of bits used for the peak / pitch flag. For example, in this embodiment, the peak / pitch flag is transmitted with 5 sub-bands of the high-frequency band, so bit library 112 ensures 5 bits.
[0082] Then, bit library 112 removes the number of bits secured by bit library 112 from the bit assets input from bit library 111, and outputs the resulting number of bits to the bit allocation calculation unit 113 in the adaptive bit allocation unit. Furthermore, the total number of bits secured in bit library 111 and bit library 112 is the third number of bits. Additionally, subbands with a peak / pitch flag of zero correspond to the third subband.
[0083] Furthermore, the order of bit library 111 and bit library 112 can be interchanged. Additionally, in this embodiment, the bit libraries are divided into blocks 111 and 112, but the bit libraries can also be processed simultaneously within a single block. Alternatively, these operations can be performed within the bit allocation calculation unit 113.
[0084] The bit allocation calculation unit 113 calculates the bit allocation for each subband quantized by the first spectrum quantization unit 106. Specifically, firstly, the number of bits output from the bit library 112 is allocated to each subband based on the quantized subband energy. The allocation method, as described in the prior art section, determines auditory importance based on the magnitude of the quantized subband energy, and prioritizes bit allocation for subbands deemed important. As a result, no bits are allocated to subbands with quantized subband energy of zero, below zero, or a specified value.
[0085] Furthermore, referring to the peak / pitch flag input during allocation, sub-bands with a peak / pitch flag of "0" (the 3rd sub-band) are excluded from the bit allocation targets. That is, only sub-bands with higher peak characteristics (here, sub-bands with a peak / pitch flag set to "1") are allocated bits as target sub-bands. Then, the sub-bands to be allocated bits (the 1st sub-band) are determined, and the number of bits allocated to each sub-band is combined and set as allocation bit information, which is first output to the quantization mode determination unit 114.
[0086] The quantization mode determination unit 114 receives the allocated bit information and peak / tone flag output from the bit allocation calculation unit 113. Then, in the case of a high-frequency subband with high tone (the quantization target of the first spectrum quantization unit 106) that has not been bit-allocated, this subband is redefined as a subband quantized by the second spectrum quantization unit 107 (the fourth subband). The number of bits necessary for quantization by the second spectrum quantization unit (the fourth number of bits) is output to the bit allocation calculation unit 113 in order to subtract from the allocated bit information. That is, the number of bits necessary for quantization by the second spectrum quantization unit 107 is allocated to this frequency band, and the allocated number of bits (the fourth number of bits) is output. Alternatively, the number of bits equivalent to the allocated number can be subtracted from the bit assets available to the first spectrum quantization unit 106 and output to the bit allocation calculation unit 113.
[0087] Furthermore, the quantization mode determination unit 114 determines the sub-band quantized by the second spectrum quantization unit 107 and outputs it to the second spectrum quantization unit 107 as the quantization mode. Specifically, the high-frequency sub-band with a low pitch (peak / pitch flag "0") (the third sub-band) and the high-frequency sub-band without allocated bits (the fourth sub-band) are determined as the sub-bands quantized by the second spectrum quantization unit 107 (the second sub-band) and output as the quantization mode.
[0088] In the bit allocation calculation unit 113, the bit assets are updated again by subtracting the number of bits received from the quantization mode determination unit 114 (the fourth bit number) from the number of bits (bit assets) input from the bit library 112, and the bit allocation for the subband quantized by the first spectrum quantization unit 106 is calculated again. If the updated bit assets are received from the quantization mode determination unit, the updated bit assets are used to calculate the bit allocation for the subband quantized by the first spectrum quantization unit 106 again. Finally, the first bit number is the value obtained by subtracting the third and fourth bit numbers from the total number of bits (bit assets).
[0089] Then, the recalculated number of bits (the first number of bits) and the information of the subband quantized by the first spectrum quantization unit 106 (the first subband) are used as the allocation bit information and output to the first spectrum quantization unit 106 this time.
[0090] Furthermore, since the bit allocation calculation unit 113 calculates the bit allocation result in the first step, the allocated bit information can be directly output to the first spectrum quantization unit 106 if it is not necessary to perform the bit allocation and other recalculations in any subband.
[0091] Figure 3 This is a flowchart illustrating the operation of the audio signal encoding device 100 in Embodiment 1, specifically, a flowchart illustrating the operation of the bit allocation unit 104.
[0092] First, the bit allocation unit 104 obtains the quantized subband energy from the subband energy quantization unit 102 (S1).
[0093] Next, the bit allocation unit 104 obtains the peak / pitch flag in the high frequency band from the tone calculation unit 103 (S2).
[0094] Then, bit allocation unit 104 determines the sub-band (third sub-band) to be quantized by second spectrum quantization unit 107 based on peak / tone flag, and ensures the bits (third bit number) used for quantization by second spectrum quantization unit 107 in bit library 111 and bit library 112 (S3).
[0095] In the bit allocation calculation unit 113, based on the quantized subband energy, the bit allocation unit 104 determines the number of bits allocated to the subband that is the object of quantization of the first spectrum quantization unit 106 (S4).
[0096] In the quantization mode determination unit 114, the bit allocation unit 104 checks the allocated bits for the high-frequency subband determined by the bit allocation calculation unit 113, and, as needed, determines the subband (second subband) to be quantized by the second spectrum quantization unit 107 again, and updates the bit assets for the first subband quantization unit 106 (S5).
[0097] Then, finally, the bit allocation unit 104 uses the updated bit assets in the bit allocation calculation unit 113 again to calculate the bit allocation (first bit number) to the first spectrum quantization unit 106 again (S6).
[0098] The audio signal encoding device according to this embodiment can reduce the overall bit rate and achieve high-quality audio signal encoding.
[0099] In particular, according to Figure 2 , Figure 3Its structure and operation prevent unquantized (bit allocation as "0") subbands from occurring in high-frequency bands with exceptionally wide subband widths, enabling the maximum number of subbands quantized by the first quantization unit to be allocated by bit. Therefore, it can extract optimal performance within a limited bit rate and achieve adaptive bit allocation.
[0100] (Implementation Method 2) Figure 4 This is a block diagram illustrating the structure and operation of the audio signal decoding device 200 according to Embodiment 2. Figure 4 The audio signal decoding device 200 shown comprises a separation unit 201, a sub-band energy decoding unit 202, a bit allocation unit 203, a first decoding unit 204, a second decoding unit 205, a denormalization unit 206, and a frequency-to-time conversion unit 207. Furthermore, an antenna A is connected to the separation unit 201. Moreover, combining the audio signal decoding device 200 and the antenna A constitutes a terminal device or a base station device.
[0101] Separation unit 201 receives the encoded information received by antenna A and separates the encoded quantized subband energy, the first encoded information, the second encoded information, and the peak / tone flag. Then, the encoded quantized subband energy is output to subband energy decoding unit 202, the first encoded information is output to first decoding unit 204, the second encoded information is output to second decoding unit 205, and the peak / tone flag is output to bit allocation unit 203.
[0102] Subband energy decoding unit 202 decodes the encoded quantized subband energy, generates decoded quantized subband energy, and outputs it to bit allocation unit 203 and denormalization unit 206.
[0103] Bit allocation unit 203 determines the allocation of bits in the first decoding unit 204 and the second decoding unit 205 by referring to the decoding quantization subband energy of each subband and the peak / pitch flag. Specifically, it determines the number of bits allocated (first bit number) and the subband in which the allocated bits are located when the first encoded information is decoded by the first decoding unit 204, and outputs this as bit allocation information. At the same time, it determines and selects the subband to be decoded (second subband) of the second encoded information decoded by the second decoding unit 205, and outputs it to the second decoding unit 205 as the quantization mode.
[0104] like Figure 5 As shown, the bit allocation unit 203 has the same structure and operation as the bit allocation unit 104 described on the encoding device side, so for details of its operation, please refer to the description of the bit allocation unit 104 on the encoding device side.
[0105] The first decoding unit 204 uses the number of the first bits represented in the allocated bit information to decode the first encoded information, generate the first decoded spectrum, and output it to the second decoding unit 205.
[0106] The second decoding unit 205 uses the first decoding spectrum to decode the second encoded information for the sub-band determined in the quantization mode, generates the second decoding spectrum, and combines the second decoding spectrum with the first decoding spectrum to generate and output the regenerated spectrum.
[0107] The normalization unit 206 refers to the energy of the decoded quantized subband, adjusts the amplitude (gain) of the regenerated spectrum, and outputs it to the frequency-time conversion unit 207.
[0108] The frequency-to-time conversion unit 207 converts the regenerated spectrum in the frequency domain into an output acoustic signal in the time domain and outputs it. As an example of frequency-to-time conversion, the inverse conversion of the frequency-to-time conversion can be listed.
[0109] The audio signal decoding device according to this embodiment can reduce the overall bit rate and achieve high-quality audio signal decoding.
[0110] (Summary) The audio signal encoding device and audio signal decoding device of the present invention have been described above in Embodiments 1 and 2. The encoding device and decoding device of the present invention can also be in the form of semi-finished products and components, such as motherboards and semiconductor elements, or in the form of finished products such as terminal devices and base station devices. When the encoding device and decoding device of the present invention are in the form of semi-finished products and components, they can be combined with antennas, DA / AD converters, amplification units, speakers, and microphones to become finished products.
[0111] Furthermore, Figure 1 , Figure 2 , Figure 4 , Figure 5 The block diagram illustrates the structure and operation (method) of hardware designed for special purposes, and includes the implementation by installing a program in general-purpose hardware to perform the operation (method) of the present invention, which is then executed by a processor. Examples of electronic computers used as general-purpose hardware include various mobile information terminals such as personal computers, smartphones, and mobile phones.
[0112] Furthermore, dedicated hardware is not limited to finished products (consumer electronics) such as mobile phones and landlines, but also includes semi-finished products and components such as motherboards and semiconductor elements.
[0113] According to embodiments of this disclosure, at least the following audio signal encoding device, audio signal decoding device, audio signal encoding method, and audio signal decoding method are disclosed.
[0114] An audio signal encoding apparatus according to an embodiment of the present disclosure includes: a time-frequency conversion unit that converts an input audio signal to the frequency domain and generates a spectrum, divides the spectrum into sub-bands of each defined frequency band, and outputs the sub-band spectrum; a sub-band energy quantization unit that quantizes the sub-band energy for each sub-band; a pitch calculation unit that analyzes the tonality of the sub-band spectrum and outputs the analysis result; a bit allocation unit that, based on the tonality analysis result and the quantized sub-band energy, selects a second sub-band quantized by a second quantization unit from the sub-bands and determines the number of first bits allocated to the first sub-band quantized by a first quantization unit; and a multiplexing unit that multiplexes and outputs information including encoding information output from the first quantization unit and the second quantization unit, the quantized sub-band energy, and the tonality analysis result, wherein the first quantization unit pulse-encodes the sub-band spectrum contained in the first sub-band using bits composed of the number of first bits; and the second quantization unit encodes the sub-band spectrum contained in the second sub-band using a pitch filter.
[0115] According to an embodiment of the audio signal encoding apparatus of the present disclosure, the bit allocation unit selects the second sub-band from the sub-band of the high-frequency band.
[0116] According to an embodiment of the audio signal encoding apparatus of the present disclosure, the bit allocation unit selects the sub-band whose pitch is lower than a predetermined threshold as the second sub-band.
[0117] According to an embodiment of the audio signal encoding apparatus of the present disclosure, the bit allocation unit selects the sub-band whose quantization sub-band energy is zero or lower than a predetermined value as the second sub-band.
[0118] According to an embodiment of the audio signal encoding apparatus of this disclosure, the bit allocation unit determines the first bit number as the number of bits obtained by subtracting the second bit number allocated to the second subband from the total number of bits available for quantization.
[0119] According to an embodiment of the audio signal encoding apparatus of this disclosure, the bit allocation unit calculates the number of third bits allocated to the third sub-band selected from the total number of bits based on the analysis result of the tone. When allocating the number of bits after subtracting the third bit from the total number of bits to the first sub-band based on the quantized sub-band energy, the sub-band without allocated bits is selected as the fourth sub-band. The number of fourth bits allocated when the fourth sub-band is encoded by the second quantization unit is calculated. The third and fourth sub-bands are reselected as the second sub-band quantized by the second quantization unit. The number of bits after subtracting the third and fourth bit from the total number of bits is determined as the number of first bits allocated to the first sub-band quantized by the first quantization unit.
[0120] According to an embodiment of the audio signal encoding apparatus of the present disclosure, the analysis result of the pitch calculation unit is output as a flag indicating whether the pitch is higher than a predetermined threshold.
[0121] According to an embodiment of the present disclosure, an audio signal decoding apparatus decodes encoded information output from an audio signal encoding apparatus, comprising: a separation unit that separates the encoded information into first encoded information, second encoded information, quantized sub-band energy quantized by the energy obtained for each sub-band, and an analysis result of the tonality calculated for each sub-band; a bit allocation unit that, based on the tonality analysis result and the quantized sub-band energy, selects the second sub-band decoded by the second decoding unit from the sub-bands and determines the number of first bits allocated to the first sub-band decoded by the first decoding unit; and a frequency-time conversion unit that converts the spectrum output from the second decoding unit to the time domain, generates and outputs an output audio signal, wherein the first decoding unit generates a first decoded spectrum by decoding the first encoded information using bits composed of the number of first bits, the second decoding unit decodes the second encoded information to generate a second decoded spectrum, and generates a regenerated spectrum by decoding using the second decoded spectrum and the first decoded spectrum.
[0122] A terminal device according to an embodiment of the present disclosure includes: an audio signal encoding device according to an embodiment of the present disclosure; and an antenna for transmitting the encoded information.
[0123] A base station apparatus according to an embodiment of the present disclosure includes: an audio signal encoding device according to an embodiment of the present disclosure; and an antenna for transmitting the encoded information.
[0124] A terminal device according to an embodiment of the present disclosure includes: an antenna that receives the encoded information and outputs it to the separation unit; and an audio signal decoding device according to an embodiment of the present disclosure.
[0125] A base station apparatus according to an embodiment of the present disclosure includes: an antenna that receives the encoded information and outputs it to the separation unit; and an audio signal decoding apparatus according to an embodiment of the present disclosure.
[0126] An audio signal encoding method according to an embodiment of the present disclosure includes the following steps: converting an input audio signal to the frequency domain and generating a spectrum; dividing the spectrum into sub-bands of each defined frequency band and outputting the sub-band spectrum; quantizing the sub-band energy of each sub-band; analyzing the tonality of the sub-band spectrum and outputting the analysis result; selecting a second sub-band from the sub-bands based on the tonality analysis result and the quantized sub-band energy; determining the number of first bits allocated to the first sub-band; encoding the sub-band spectrum contained in the first sub-band using bits composed of the number of first bits to generate first encoding information; encoding the sub-band spectrum contained in the second sub-band using a pitch filter to generate second encoding information; multiplexing the first encoding information and the second encoding information and outputting them.
[0127] An audio signal decoding method according to an embodiment of the present disclosure for decoding encoded information output from an audio signal encoding device includes the following steps: separating the encoded information into first encoded information, second encoded information, quantized sub-band energy (quantized by calculating the energy of each sub-band), and an analysis result of the tonal characteristics calculated for each sub-band; selecting a second sub-band from the sub-bands based on the tonal characteristics analysis result and the quantized sub-band energy; determining a first number of bits allocated to the first sub-band; decoding the first encoded information using bits composed of the first number of bits and generating a first decoded spectrum; decoding the second encoded information and generating a second decoded spectrum; decoding using the second decoded spectrum and the first decoded spectrum and generating a regenerated spectrum; converting the regenerated spectrum to the time domain, generating an output audio signal, and outputting it.
[0128] Industrial applicability
[0129] The audio signal encoding device and audio signal decoding device of the present invention can be applied to equipment and components related to the recording, transmission and reproduction of audio signals.
[0130] Label Explanation
[0131] 100 Audio Signal Encoding Device
[0132] 101 Time-to-Frequency Conversion Unit
[0133] 102 Sub-band Energy Quantization Units
[0134] 103 Pitch Calculation Units
[0135] 104-bit allocation unit
[0136] 105 normalized units
[0137] 106 First Spectrum Quantization Unit
[0138] 107 Second Spectrum Quantization Unit
[0139] 108 reuse units
[0140] 111 Bit Library
[0141] 112-bit library
[0142] 113-bit allocation calculation unit
[0143] 114 Quantization Mode Determination Unit
[0144] 200 Audio Signal Decoding Device
[0145] 201 Separation Unit
[0146] 202 Sub-band Energy Decoding Units
[0147] 203-bit allocation unit
[0148] 204 Decoding Unit 1
[0149] 205 Decoding Unit 2
[0150] 206 Solution normalization unit
[0151] 207 Frequency-to-Time Conversion Unit
[0152] 211 Bit Library
[0153] 212-bit library
[0154] 213-bit allocation calculation unit
[0155] 214 Quantization Mode Determination Unit
Claims
1. An audio signal encoding device, comprising: The time-frequency conversion unit generates a spectrum by performing a frequency domain conversion on the input audio signal, divides the spectrum into sub-bands of a specified frequency band, and outputs the sub-band spectrum. The input audio signal includes a music signal, a speech signal, or a signal that is a mixture of music and speech signals. Subband energy quantization unit, quantizes the subband energy for each subband; The pitch calculation unit analyzes the tonal characteristics of the sub-band spectrum and outputs the analysis results; A bit allocation unit, based on the analysis results of the tone and the energy of the quantized subband, selects a second subband from the subbands to be quantized by the second quantization unit, and determines the number of first bits to be allocated to the first subband from the subbands to be quantized by the first quantization unit; and The multiplexing unit multiplexes the encoded information output from the first quantization unit and the second quantization unit, the quantization sub-band energy, and the tone analysis results into information and outputs the multiplexed information, wherein... The first quantization unit encodes the subband spectrum within the subband spectrum contained in the first subband by using the first number of bits; and The second quantization unit encodes the sub-band spectrum within the sub-band spectrum contained in the second sub-band using an encoding method, the encoding method including: Identify the second sub-band to which the second quantization unit will perform quantization. The search identifies sub-bands or frequency bands where the normalized sub-band spectrum of the identified second sub-band has the highest correlation with the quantized spectrum of the sub-band or frequency band. The hysteresis information of the second sub-band is generated using the position of the sub-band or the frequency band. The hysteresis information is encoded to generate the encoded information output from the second quantization unit; Wherein, the hysteresis information is the absolute position of the sub-band or the frequency band, or the hysteresis information is the relative position of the sub-band or the frequency band, or the hysteresis information is the number of the sub-band.
2. The audio signal encoding device as described in claim 1, The bit allocation unit selects the second sub-band from the sub-bands in the high frequency range.
3. The audio signal encoding device as described in claim 2, The bit allocation unit selects the sub-band whose pitch is below a predetermined threshold as the second sub-band.
4. The audio signal encoding device as described in claim 2, The bit allocation unit selects the subband whose quantization subband energy is equal to zero or lower than a specified value as the second subband.
5. The audio signal encoding device as described in claim 1, The bit allocation unit determines the number of the first bit by subtracting the number of the second bit to be allocated to the second subband from the total number of bits available for quantization.
6. The audio signal encoding device as described in claim 5, The bit allocation unit calculates the number of 3rd bits to be allocated to the 3rd subband from the total number of bits, the 3rd subband being selected from the subbands based on the tonal analysis results; when the number of bits obtained by subtracting the 3rd bit from the total number of bits is allocated to the 1st subband based on the quantized subband energy, the subbands with unallocated bits are selected as the 4th subband, and the number of 4th bits allocated when the 4th subband is encoded by the 2nd quantization unit is calculated, and The third and fourth subbands are selected as additional second subbands to be quantized by the second quantization unit, and the number of bits obtained by subtracting the third and fourth bits from the total number of bits is determined as the first number of bits to be allocated to the first subband to be quantized by the first quantization unit.
7. The audio signal encoding device as described in claim 1, The analysis results from the pitch calculation unit are output as a flag indicating whether the pitch is higher than a predetermined threshold.
8. The audio signal encoding device as claimed in claim 1, wherein the device is configured to: Obtain the quantized subband energy. Obtain peak / pitch indicators in the high-frequency range. Identify the subband to which the second quantization unit will perform quantization and retain the bits to be used in the quantization of the second quantization unit. The number of bits to be allocated to the subband that will be quantized by the first quantization unit is determined based on the quantization subband energy. Check the number of bits allocated to the subband in the high-frequency range, identify the second subband to which the second quantization unit will perform quantization as needed, and update the bit budget for the first quantization unit. The bit allocation for the first quantization unit is recalculated using the updated bit budget.
9. An audio signal decoding device for decoding encoded information, the audio signal decoding device comprising: The separation unit separates the encoded information into first encoded information, second encoded information, quantized subband energy obtained by quantizing the energy of each subband in the subband, and the analysis result of the tonality calculated for each subband in the subband, wherein the second encoded information includes hysteresis information; The bit allocation unit, based on the analysis results of the tone and the quantized subband energy, selects a second subband from the subbands to be decoded by the second decoding unit, and determines the number of first bits to be allocated to the first subband from the subbands to be decoded by the first decoding unit. as well as The frequency-time conversion unit generates and outputs an audio signal by performing a time-domain conversion on the regenerated spectrum output from the second decoding unit. The first decoding unit generates a first decoded spectrum by decoding the first encoded information using the first number of bits, and The second decoding unit generates a second decoded spectrum by decoding the second encoded information including the hysteresis information, and the second decoding unit generates the regenerated spectrum by combining the second decoded spectrum and the first decoded spectrum. The output audio signal includes music signal, voice signal, or a mixture of music signal and voice signal.
10. A terminal device, comprising: The audio signal encoding device according to claim 1; as well as Antenna for transmitting encoded information.
11. A terminal device, comprising: It receives the encoded information and outputs it to the antenna of the separation unit; And the audio signal decoding device as described in claim 9.
12. A method for encoding audio signals, comprising: A spectrum is generated by performing a frequency domain conversion on an input audio signal, wherein the input audio signal includes a music signal, a speech signal, or a mixture of music and speech signals; The spectrum is divided into sub-bands of a specified frequency band and the sub-band spectrum is output. Quantize the subband energy for each of the subbands; Analyze the tonal characteristics of the subband spectrum and output the analysis results; Based on the analysis results of the tonality and the quantized subband energy, a second subband is selected from the subbands; Determine the number of bits to be allocated to the first subband of the first subband; The first encoded information is generated by encoding the subband spectrum within the subband spectrum contained in the first subband using the first number of bits; The second coded information is generated by encoding the subband spectrum within the subband spectrum contained in the second subband using an encoding method, wherein the encoding method includes: Identify the second sub-band to which the second quantization unit needs to be quantized. The search identifies sub-bands or frequency bands where the normalized sub-band spectrum of the identified second sub-band has the highest correlation with the quantized spectrum of the sub-band or frequency band. The hysteresis information of the second sub-band is generated using the position of the sub-band or the frequency band, wherein the hysteresis information is the absolute position of the sub-band or the frequency band, or the hysteresis information is the sub-band number. Encode the hysteresis information to generate the second encoded information; and The first encoded information and the second encoded information are multiplexed together and output.
13. A method for decoding audio signals containing decoded information, the audio signal decoding method comprising: The encoded information is separated into first encoded information, second encoded information, quantized subband energy obtained by quantizing the energy of each subband in the subband, and analysis results of the tonality calculated for each subband in the subband, wherein the second encoded information includes hysteresis information; Based on the analysis results of the tonality and the quantized subband energy, a second subband is selected from the subbands; Determine the number of bits to be allocated to the first subband of the first subband; The first decoded spectrum is generated by decoding the first encoded information using the first number of bits; A second decoded spectrum is generated by decoding the second encoded information including the hysteresis information, and a regenerated spectrum is generated by combining the second decoded spectrum and the first decoded spectrum; and An output audio signal is generated and output by performing a time-domain conversion on the regenerated spectrum, wherein the output audio signal includes a music signal, a speech signal, or a mixture of music and speech signals.
14. A computer program product having program code for performing the method of claim 12 or claim 13.
Citation Information
Patent Citations
Systems, methods, apparatus and computer-readable media for dynamic bit allocation
JP2013534328A
Encoder apparatus and decoder apparatus
WO2005027095A1
Speech audio encoding device, speech audio decoding device, speech audio encoding method, and speech audio decoding method
CN104737227A
Acoustic signal encoding device, acoustic signal decoding device, method for encoding acoustic signal, and method for decoding acoustic signal
CN106133831A