Decoding apparatus, decoding method, computer-readable recording medium, and computer program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2018-12-03
- Publication Date
- 2026-08-07
AI Technical Summary
[0016]根据编码装置以及解码装置,能够进行编码以及解码,使得摩擦音等音信号在听觉上的劣化变少。
Smart Images

Figure CN117351969B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application filed on December 3, 2018, with application number 201880086667.4, entitled "Decoding device, encoding device, method and procedure thereof", and filed by Nippon Telegraph and Telephone Corporation. Technical Field
[0002] This invention relates to signal processing techniques, such as audio signal encoding techniques, and to techniques for encoding or decoding sample strings derived from the spectrum of an audio signal. Background Technology
[0003] In audio signal compression encoding, to improve compression efficiency, the audio signal has traditionally been represented as a spectrum string, and this spectrum string is encoded by bit allocation that takes into account auditory importance. This bit allocation, considering auditory importance, is achieved by prioritizing bit allocation to samples corresponding to low frequencies in the spectrum string. As a result, sometimes a structure is adopted where no bits are allocated to samples corresponding to high frequencies in the spectrum string, and direct information about these high-frequency samples is not encoded in the encoding device. Since the decoder in the corresponding encoding device sets the values of high-frequency samples in the spectrum string to 0 to obtain the decoded sound, a band-spreading technique as described in Non-Patent Document 1 is sometimes used. This technique involves adjusting the amplitude of the low-frequency sample string while copying it, and outputting the decoded result of the high-frequency sample string as the result of the high-frequency sample string. This is based on the fact that humans have low sensitivity to high frequencies when listening to audio, and do not experience discomfort if they can hear low-frequency overtones. By allocating the bits saved in the high-frequency band to the low-frequency band, information more important to human auditory characteristics can be represented with high precision. Thus, the encoding method for audio signals is usually designed to allocate more bits to the low-frequency spectrum.
[0004] Existing technical documents
[0005] Non-patent literature
[0006] Non-patent literature 1: M. Arora, J. Lee, and S. Park, “High Quality Blind Bandwidth Extension of Audio for Portable Player Applications,” AES 120th Convention, Paris, France, 2006. Summary of the Invention
[0007] The problem that the invention aims to solve
[0008] According to the bandwidth extension technology of Non-Patent Document 1, for most sounds in natural speech, the bandwidth extension sound obtained from the decoded sound obtained by the decoding device can be obtained with minimal degradation in auditory quality. However, there are also sounds in natural speech, such as the fricative sounds in human speech, where energy is concentrated at high frequencies and there is basically no energy at low frequencies. If such sound signals are encoded in the encoding device with the bit allocation as described above, especially under low bit rate conditions, the decoded sound obtained from the decoding device has a large distortion in the main frequency components of the sound. If the bandwidth extension technology of Non-Patent Document 1 is used to obtain the bandwidth extension sound from the decoded sound, there is a problem of auditory degradation of the bandwidth extension sound.
[0009] Therefore, the object of the present invention is to provide an encoding apparatus that performs compression encoding on the encoding side based on bandwidth expansion on the decoding side, a decoding apparatus that performs decoding in conjunction with bandwidth expansion on the decoding side, methods and procedures thereof, so that the auditory degradation of sound signals such as fricatives is reduced.
[0010] Methods for solving problems
[0011] One aspect of the decoding apparatus of the present invention includes: a decoding unit that decodes a spectrum code of a frame unit within a predetermined time interval, wherein the spectrum code is a spectrum code for which no bits are allocated to a portion of the high-domain side, to obtain a sample string in the frequency domain; a bandwidth extension unit that, by configuring samples of K samples contained in the sample string in the frequency domain obtained by the decoding unit on the high-domain side compared to the sample string in the frequency domain obtained by the decoding unit, to obtain a decoded extended spectrum sequence, wherein K is an integer of 2 or more; and a fricative adjustment release unit that, when the information indicating whether it is an input fricative tone indicates that it is a fricative tone, obtains a result in which all or part of the low-domain side frequency sample string in the decoded extended spectrum sequence obtained by the bandwidth extension unit that is located on the low-domain side compared to a predetermined frequency, and the same number of the high-domain side frequency sample strings in the decoded extended spectrum sequence obtained by the bandwidth extension unit that are located on the high-domain side compared to a predetermined frequency, are swapped as the spectrum sequence of the decoded tone signal; otherwise, the decoded extended spectrum sequence obtained by the bandwidth extension unit is directly used as the spectrum sequence of the decoded tone signal.
[0012] One aspect of the present invention is a decoding apparatus that decodes a spectrum code of a frame unit within a predetermined time interval to obtain a spectrum sequence of a decoded audio signal. The apparatus includes: a decoding unit that, when information indicating whether an input fricative tone is present indicates that the input tone is fricative, decodes the spectrum code to obtain a frequency domain spectrum sequence by not allocating bits to a portion of the low-domain side; otherwise, decodes the spectrum code to obtain a frequency domain spectrum sequence by not allocating bits to a portion of the high-domain side; and a fricative tone corresponding frequency band extension unit that, when information indicating whether an input tone is fricative indicates that the input tone is fricative, extends the frequency domain spectrum sequence obtained by the decoding unit to a lower-domain side to obtain a spectrum sequence of the decoded audio signal; otherwise, extends the frequency domain spectrum sequence obtained by the decoding unit to a higher-domain side to obtain a spectrum sequence of the decoded audio signal.
[0013] An encoding apparatus according to one aspect of the present invention includes an encoding unit that encodes a sample string of frequencies corresponding to an audio signal within a frame unit of a predetermined time interval into a spectral code by encoding a portion of the high-domain side without allocating bits. The encoding apparatus includes: a fricative tone determination unit that determines whether the audio signal is a fricative tone; and a fricative tone adjustment unit that, if the fricative tone determination unit determines that the audio signal is a fricative tone, obtains an adjusted spectral sequence by swapping all or a portion of the low-domain spectral sequence of the audio signal's spectral sequence located in the low-domain side compared to the predetermined frequency with all or a portion of the same number of high-domain spectral sequences located in the high-domain side compared to the predetermined frequency. Otherwise, the spectral sequence corresponding to the audio signal is directly obtained as the adjusted spectral sequence. The encoding unit encodes the adjusted spectrum sequence obtained by the friction tone adjustment unit as a sample string of frequencies corresponding to the sound signal to obtain a spectrum code. The encoding device also includes a band-spreading gain encoding unit, which stores multiple codes and gain candidate vectors corresponding to each code. Each gain candidate vector contains K gain candidate values. The code corresponding to the gain candidate vector is obtained as a band-spreading gain code and output. The gain candidate vector is the gain candidate vector whose error is minimized by multiplying the K adjusted spectra in the coded part of the adjusted spectrum sequence with bits allocated by the encoding part and the K gain candidate values contained in the gain candidate vector, and the sequence of the absolute values of the K adjusted spectra in the coded part of the adjusted spectrum sequence without bits allocated by the encoding part. Here, K is an integer greater than 2.
[0014] One aspect of the present invention provides a decoding method that decodes a spectral code of a frame unit within a defined time interval to obtain a spectral sequence of a decoded audio signal. The method includes: a decoding step where, if information indicating whether an input fricative tone is present indicates that the input tone is fricative, bits are not allocated to a portion of the lower domain side of the spectral code, and the spectral code is decoded to obtain a frequency domain spectral sequence; otherwise, bits are not allocated to a portion of the higher domain side of the spectral code, and the spectral code is decoded to obtain a frequency domain spectral sequence; and a fricative corresponding frequency band extension step where, if information indicating whether an input tone is fricative indicates that the input tone is fricative, the frequency domain spectral sequence obtained in the decoding step is extended to a lower domain side to obtain the spectral sequence of the decoded audio signal; otherwise, the frequency domain spectral sequence obtained in the decoding step is extended to a higher domain side to obtain the spectral sequence of the decoded audio signal.
[0015] The effects of the invention
[0016] The encoding and decoding devices enable encoding and decoding, thereby reducing the auditory degradation of sound signals such as fricatives. Attached Figure Description
[0017] Figure 1 This is a block diagram illustrating an example of an encoding device according to the first embodiment.
[0018] Figure 2 This is a flowchart illustrating an example of the encoding method of the first embodiment.
[0019] Figure 3 This is a block diagram illustrating an example of a decoding device according to the first embodiment.
[0020] Figure 4 This is a flowchart illustrating an example of the decoding method of the first embodiment.
[0021] Figure 5 This is a diagram used to illustrate an example of friction noise adjustment.
[0022] Figure 6 This is a diagram used to illustrate an example of friction noise adjustment.
[0023] Figure 7 This is a diagram used to illustrate an example of friction noise adjustment.
[0024] Figure 8 This is a diagram used to illustrate an example of friction noise adjustment.
[0025] Figure 9 This is a block diagram illustrating an example of an encoding device according to the second embodiment.
[0026] Figure 10 This is a flowchart illustrating an example of the encoding method of the second embodiment.
[0027] Figure 11 This is a block diagram illustrating an example of a decoding device according to the second embodiment.
[0028] Figure 12 This is a flowchart illustrating an example of the decoding method of the second embodiment.
[0029] Figure 13 This diagram illustrates an example of bandwidth extension processing and friction tone adjustment release processing.
[0030] Figure 14 This diagram illustrates an example of bandwidth extension processing and friction tone adjustment release processing. Detailed Implementation
[0031] <First Implementation>
[0032] The first embodiment is a prerequisite for the second embodiment, which is an embodiment of the present invention.
[0033] The system of the first embodiment includes an encoding device and a decoding device. The encoding device encodes a time-domain audio signal input in frames of a predetermined time length to obtain a code, and outputs it. The code output by the encoding device is input to the decoding device. The decoding device decodes the input code and outputs a time-domain audio signal in frames. The audio signal input to the encoding device is, for example, a speech signal or sound signal obtained by picking up speech or music through a microphone and performing an analog-to-digital (A / D) conversion. Furthermore, the audio signal output by the decoding device can be heard by reproducing it through a loudspeaker, for example, after being converted by an analog-to-digital (A / D) converter.
[0034] Encoding device
[0035] Reference Figure 1 The processing procedure of the encoding device in the first embodiment will be explained. For example... Figure 1 As illustrated, the encoding apparatus of the first embodiment includes: a frequency domain transformation unit 11, a fricative sound determination unit 12, a fricative sound adjustment unit 13, an encoding unit 14, and a multiplexing unit 15. A time-domain audio signal input to the encoding apparatus is input to the frequency domain transformation unit 11. The encoding apparatus performs processing in each unit on a frame unit of a predetermined time length. The encoding method of the first embodiment performs the following through each unit of the encoding apparatus: Figure 2 The process is implemented by steps S11 to S15 as illustrated in the example.
[0036] Furthermore, the encoding device can be configured to input a frequency-domain audio signal instead of a time-domain audio signal. In this configuration, the encoding device may not include a frequency-domain transformation unit 11; it is sufficient to input a frequency-domain audio signal of a predetermined time-length frame unit to the fricative sound determination unit 12 and the fricative sound adjustment unit 13.
[0037] [Frequency Domain Transformation Unit 11]
[0038] The time-domain audio signal input to the encoding device is input into the frequency domain transformation unit 11. The frequency domain transformation unit 11 transforms the input time-domain audio signal into an N-point frequency domain spectrum sequence X0,…,X in frame units of a predetermined time length, for example, by modified discrete cosine transform (MDCT). N-1 Then output (step S11). N is a positive integer, such as N=32, etc. Moreover, the subscripts added to X are numbered sequentially starting from the lowest frequency spectrum. As a transformation method to the frequency domain, various well-known transformation methods other than MDCT can be used (e.g., Discrete Fourier Transform, Short-Time Fourier Transform, etc.).
[0039] The frequency domain transformation unit 11 outputs the transformed spectral sequence to the fricative sound determination unit 12 and the fricative sound adjustment unit 13. Furthermore, the frequency domain transformation unit 11 can also apply filtering and compressing processing to the transformed spectral sequence for auditory weighting, using the filtered or compressed sequence as the spectral sequence X0,…,X. N-1 Output.
[0040] [Friction noise determination unit 12 (friction noise determination device)]
[0041] In the friction sound determination unit 12, for example, the spectrum sequence X0,…,X output by the input frequency domain transformation unit 11 is... N-1 The fricative sound determination unit 12 uses the input spectrum sequence X0,…,X in frames. N-1 The system determines whether the sound signal is a fricative tone and outputs the determination result as fricative tone determination information to the fricative tone adjustment unit 13 and the multiplexing unit 15 (step S12). For example, 1 bit of information can be used as the fricative tone determination information. That is, the fricative tone determination unit 12 outputs a bit "1" as fricative tone determination information in frames when the sound signal is a fricative tone, and outputs a bit "0" as fricative tone determination information when the sound signal in that frame is not a fricative tone.
[0042] The fricative sound determination unit 12, for example, calculates the input spectrum sequence X0,…,X N-1The average energy of the samples located on the higher domain side relative to the input spectral sequence X0,…,X N-1 The index that the larger the ratio of the average energy of the samples located on the lower domain side is, the larger the value is used as the index for the sound of fricative sound in that frame. The fricative sound determination unit 12 determines that the sound is fricative sound if the calculated index is greater than or above a predetermined threshold, and determines that the sound is not fricative sound if the calculated index is less than or below the predetermined threshold.
[0043] If integer values greater than 1 and less than N-1 are designated as MA, and integer values greater than MA and less than N are designated as MB, then the fricative sound determination unit 12, for example, uses the spectrum sequence X0,…,X N-1 The sample numbers below MA are X0, ..., X MA Let the sample be located on the lower domain side, and let the spectral sequence X0,…,X be... N-1 Samples with a sample number of MB or higher, i.e., X MB ,…,X N-1 Let X0,…,X be samples located on the higher domain side. MA The average of the absolute values and the average of the sums of squares of all or a portion of the sample values is set as the low-domain average energy, and X is... MB ,…,X N-1 The average of the absolute values and the average of the sums of squares of all or part of the sample values is set as the high-domain average energy. The value obtained by dividing the high-domain average energy by the low-domain average energy is used as an index of the fricative tone.
[0044] Furthermore, an integer value MA can be set such that samples from the low-domain side, which are the objects of the calculation of the low-domain average energy in the friction sound determination unit 12, are included in the low-domain spectral sequence of the friction sound adjustment unit 13 described later. That is, the integer value MA used in the friction sound determination unit 12 can be set to a value less than the integer value M of the friction sound adjustment unit 13 described later. Moreover, an integer value MB can be set such that samples from the high-domain side, which are the objects of the calculation of the high-domain average energy in the friction sound determination unit 12, are included in the high-domain spectral sequence of the friction sound adjustment unit 13 described later. That is, the integer value MB used in the friction sound determination unit 12 can be set to a value greater than or equal to the integer value M of the friction sound adjustment unit 13 described later.
[0045] The samples X0,…,X located on the lower domain side are... MA When a portion of the sample values are used in the calculation of the aforementioned indicators, they can be derived from X0,…,X MA The side with the lowest frequency is used to calculate the above indicators using one or more sample values. That is, α can be set as a positive integer less than MA, and X0,…,X αThe average of the absolute values and the average of the sums of squares of the sample values is set as the low-domain average energy. The value of α is predetermined based on prior experiments, such that if X0,…,X… α For sounds other than those with fricative properties, the spectrum can be changed to the range that is normally possible.
[0046] In the encoding process of the encoding unit 14 described later, due to the constraint of the maximum number of bits obtained in the encoding process, sometimes no bits are allocated to several samples starting from the highest frequency in the adjusted spectrum sequence. In this case, sometimes regardless of whether the spectrum adjustment process in the friction tone adjustment unit 13 described later is performed or not, no bits are allocated to β samples (β is a positive integer) starting from the highest frequency in the spectrum sequence. In such a case, X can be... MB ,…,X N-1 The X values of the first β samples, starting from the highest frequency, were removed. MB ,…,X N-1-β Used for calculating the aforementioned indicators. That is, X can be... MB ,…,X N-1-β The average of the absolute values and the average of the sums of squares of the sample values is set as the high-domain average energy. Moreover, the value of β can be predetermined in relation to the encoding processing performed by the pre-designed encoding unit 14 and the adjustment processing performed by the fricative tone adjustment unit 13.
[0047] Figure 5 and Figure 6 This is an example of the friction tone adjustment unit 13 described later, with N=32 and M=20. In these examples, X0,…,X in the spectral sequence… 19 Set as a low-domain side spectral sequence, X in the spectral sequence 20 ,…,X 31 The spectral sequence is set to the high-domain side. Therefore, the fricative sound determination unit 12 sets MA to a value less than 20, for example, 19, sets MB to a value greater than 20, for example, 20, and sets X0,…,X 19 The average of the absolute values and the average of the sums of squares of all or a portion of the sample values is set as the low-domain average energy, and X is... 20 ,…,X 31 The average of the absolute values or the average of the sums of squares of all or a portion of the sample values can be set as the high-domain average energy. Here, when α=8, the friction sound determination unit 12 sets the average of the absolute values or the average of the sums of squares of the sample values of X0, ..., X8 as the low-domain average energy. And here, when β=4, the friction sound determination unit 12 sets X... 20 ,…,X 27The average of the absolute values and the average of the sums of squares of the sample values can be set as the high-domain average energy.
[0048] Moreover, such as Figure 1 As indicated by the dashed lines, the fricative tone determination unit 12 inputs not the spectral sequence output by the frequency domain transformation unit 11, but a time-domain audio signal input to the encoding device. Using the input time-domain audio signal, it determines whether the audio signal of a frame is a fricative tone. This determination can be performed, for example, by calculating the number of zero-crossings of the input time-domain audio signal as an index for whether the frame is a fricative tone. If the calculated index is greater than or above a predetermined threshold, the tone is determined to be a fricative tone; otherwise, if the calculated index is below or below the predetermined threshold, the tone is determined not to be a fricative tone.
[0049] [Friction adjustment section 13]
[0050] The frequency spectrum sequence X0,…,X is input to the frequency domain transformation unit 11 in the friction noise adjustment unit 13. N-1 The friction tone determination unit 12 outputs friction tone determination information. The friction tone adjustment unit 13 adjusts the input spectrum sequence X0,…,X in frames, when the input friction tone determination information indicates that the tone is a friction tone. N-1 The adjusted spectrum sequence Y0,…,Y is obtained by performing the following spectrum adjustment process. N-1 The resulting adjusted spectral sequence Y0,…,Y N-1 The output is sent to the encoding unit 14, where the spectral sequence X0,…,X is processed if the fricative determination information indicates that the sound is not a fricative. N-1 Directly used as the adjusted spectral sequence Y0,…,Y N-1 Output to encoding unit 14 (step S13).
[0051] If we define M as an integer value greater than 1 and less than N, then for example, if we define the spectrum sequence X0,…,X… N-1 The samples with sample numbers less than M are X0,…,X M-1 The sample group is set as the low-domain side spectrum sequence, and the spectrum sequence X0,…,X is... N-1 Samples with sample number M or higher, i.e., X M ,…,X N-1 If the sample group is set as a high-domain side spectrum sequence, then when the fricative sound determination information indicates a fricative sound, the adjustment process performed by the fricative sound adjustment unit 13 is as follows: that is, the low-domain side spectrum sequence X0,…,X is obtained. M-1 All or part of the samples, and the same number of high-domain side spectral sequences X M ,…,XN-1 The result of swapping all or part of the samples is used as the adjusted spectral sequence Y0,…,Y N-1 The following describes the adjustment process performed by the friction noise adjustment unit 13. Various adjustments can be performed by the friction noise adjustment unit 13, including those described below, but the specific adjustments are predetermined.
[0052] [Example 1 of the adjustment process performed by the friction noise adjustment unit 13]
[0053] When the fricative sound determination information indicates that the sound is fricative, for example, the fricative sound adjustment unit 13 obtains the adjusted spectrum sequence Y0,…,Y by performing the following steps 1-1 to 1-6. N-1 Furthermore, steps 1-1 to 1-6 below are divided into 6 steps to easily understand the operation of the friction noise adjustment unit 13. However, performing steps 1-1 to 1-6 separately is just one example. The friction noise adjustment unit 13 can also perform the equivalent processing of steps 1-1 to 1-6 in one step by changing the arrangement of elements or replacing the index.
[0054] Step 1-1: Convert the spectral sequence X0,…,X N-1 The sample group of samples with sample numbers less than M is set as the low-domain side spectrum sequence X0,…,X M-1 The spectral sequence X0,…,X N-1 The sample group with sample number M or above is set as the high-domain side spectrum sequence X. M ,…,X N-1 .
[0055] Step 1-2: Extract the low-domain side spectrum sequence X0,…,X obtained in step 1-1. M-1 The C samples (where C is a positive integer) contained therein are used as the adjustment targets for the higher domain side.
[0056] Step 1-3: Extract the high-domain side spectrum sequence X obtained in step 1-1. M ,…,X N-1 The C samples contained therein are used as the adjustment targets for the lower domain side.
[0057] Steps 1-4: Obtain the sample positions of the adjustment target samples from the low-domain side spectrum sequence extracted in Step 1-2, and configure the results of the adjustment target samples extracted from the high-domain side spectrum sequence in Step 1-3 as the low-domain side adjusted spectrum sequence Y0,…,Y M-1 .
[0058] Steps 1-5: Obtain the sample positions of the adjustment target samples from the high-domain side spectrum sequence extracted in Step 1-3, and configure the results of the adjustment target samples extracted from the low-domain side spectrum sequence in Step 1-2 as the high-domain side adjusted spectrum sequence Y. M ,…,Y N-1 .
[0059] Steps 1-6: The adjusted spectrum sequence Y0,…,Y obtained in Steps 1-4 is processed using the low-domain side. M-1 The high-domain side adjusted spectral sequence Y obtained in steps 1-5 M ,…,Y N-1 By combining these sequences, we obtain the adjusted spectral sequence Y0,…,Y N-1 .
[0060] exist Figure 5 The examples shown are steps 1-1 to 1-6 when N=32, M=20, and C=8. The fricative sound adjustment unit 13 first processes the spectral sequence X0,…,X… 31 In X0,…,X 19 Set as the low-domain side spectral sequence, and set X 20 ,…,X 31 Set as the high-domain side spectrum sequence (step 1-1). The friction tone adjustment unit 13 extracts the low-domain side spectrum sequence X0,…,X 19 The eight samples X2, ..., X9 included are used as the adjustment targets for the higher domain side (steps 1-2). The fricative sound adjustment unit 13 extracts the higher domain side spectral sequence X. 20 ,…,X 31 The 8 samples X included 20 ,…,X 27 As the target sample for adjustment towards the lower domain side (steps 1-3), the friction tone adjustment unit 13 obtains the sample positions where X2, ..., X9 exist in the lower domain side spectral sequence and configures X... 20 ,…,X 27 The result is the adjusted spectral sequence Y0,…,Y on the low-domain side. 19 (Steps 1-4). The friction sound adjustment unit 13 obtains the presence of X in the high-domain side spectral sequence. 20 ,…,X 27 The sample positions were configured with the results of X2,…,X9, as the adjusted spectral sequence Y on the high-domain side. 20 ,…,Y 31 (Steps 1-5). The friction sound adjustment unit 13 adjusts the low-domain side spectral sequence Y0,…,Y 19 And the high-domain side has adjusted the spectral sequence Y 20 ,…,Y 31 By combining these sequences, we obtain the adjusted spectral sequence Y0,…,Y 31(Steps 1-6)
[0061] [Example 2 of the adjustment process performed by the friction noise adjustment unit 13]
[0062] Furthermore, the friction noise adjustment unit 13 can also replace the above-described steps 1-4 and perform the following steps 1-4'.
[0063] Step 1-4': In step 1-2, the remaining samples from the low-domain side spectral sequence that were used to adjust the samples towards the high-domain side are squeezed towards the low-domain side. The empty sample positions on the high-domain side are then filled with the adjustment samples taken from the high-domain side spectral sequence in step 1-3, and the result is taken as the adjusted low-domain side spectral sequence Y0,…,Y M-1 .
[0064] The friction tone adjustment unit 13 performs step 1-4' instead of step 1-4. In the subsequent encoding unit 14, samples with lower corresponding frequencies are encoded to increase the importance of hearing.
[0065] Thus, when the fricative sound determination unit 12 determines that the sound is a fricative sound, the fricative sound adjustment unit 13 can also be configured to construct an adjusted spectrum sequence by using a low-domain side adjusted spectrum sequence and a high-domain side adjusted spectrum sequence, including a portion of the samples in the low-domain side spectrum sequence in the high-domain side adjusted spectrum sequence, placing the remaining samples in the low-domain side of the low-domain side adjusted spectrum sequence, placing a portion of the samples in the high-domain side spectrum sequence in the high-domain side of the low-domain side adjusted spectrum sequence, and including the remaining samples in the high-domain side spectrum sequence in the high-domain side adjusted spectrum sequence, thereby obtaining an adjusted spectrum sequence.
[0066] [Example 3 of the adjustment process performed by the friction noise adjustment unit 13]
[0067] Similarly, the friction noise adjustment unit 13 can replace the above-described steps 1-5 and perform the following steps 1-5'.
[0068] Step 1-5': After extracting the adjustment target samples from the high-domain side spectrum sequence in Step 1-3, the remaining samples are squeezed towards the low-domain side. The empty high-domain side sample positions are then filled with the adjustment target samples extracted from the low-domain side spectrum sequence in Step 1-2, and the result is taken as the high-domain side adjusted spectrum sequence Y. M ,…,Y N-1 .
[0069] Step 1-5' is performed by replacing step 1-5 with the friction tone adjustment unit 13. In the subsequent encoding unit 14, the auditory importance of the sample that was originally located on the low domain side can be increased and the sample that was originally located on the high domain side can be encoded.
[0070] Figure 6 The example shown is when N=32, M=20, and C=8, step 1-4' is performed instead of step 1-1 to step 1-6, and step 1-5' is performed instead of step 1-5. The friction sound adjustment unit 13 first processes the spectrum sequence X0,…,X… 31 In X0,…,X 19 Set as the low-domain side spectral sequence, and set X 20 ,…,X 31 Set as the high-domain side spectrum sequence (step 1-1). The friction tone adjustment unit 13 extracts the low-domain side spectrum sequence X0,…,X 19 The eight samples X2, ..., X9 included are used as the adjustment targets for the higher domain side (steps 1-2). The fricative tone adjustment unit 13 extracts the higher domain side spectral sequence X. 20 ,…,X 31 The 8 samples X included 20 ,…,X 27 As a sample for adjustment towards the lower domain side (steps 1-3), the friction tone adjustment unit 13 will use X from the lower domain side spectral sequence. 10 ,…,X 19 Squeezing towards the lower domain side, X after squeezing towards the lower domain side 10 ,…,X 19 High-domain side configuration X 20 ,…,X 27 The result is obtained as the adjusted spectral sequence Y0,…,Y on the low-domain side. 19 (Steps 1-4'). The friction sound adjustment unit 13 will adjust the X in the high-domain side spectrum sequence. 28 ,…,X 31 Squeezing towards the lower domain side, X after squeezing towards the lower domain side 28 ,…,X 31 The high-domain side configurations X2,…,X9 are used to obtain the results as the high-domain side adjusted spectral sequence Y. 20 ,…,Y 31 (Steps 1-5'). The friction sound adjustment unit 13 adjusts the low-domain side spectral sequence Y0,…,Y 19 And the high-domain side has adjusted the spectral sequence Y 20 ,…,Y 31 By combining these sequences, we obtain the adjusted spectral sequence Y0,…,Y 31 (Steps 1-6)
[0071] In this way, if the fricative sound determination unit 12 determines that the sound is a fricative sound, the fricative sound adjustment unit 13 constructs an adjusted spectrum sequence by using the adjusted spectrum sequence on the low-domain side and the adjusted spectrum sequence on the high-domain side. A portion of the samples in the low-domain side spectrum sequence are placed on the high-domain side of the adjusted spectrum sequence, the remaining samples in the low-domain side spectrum sequence are included in the adjusted spectrum sequence on the low-domain side, a portion of the samples in the high-domain side spectrum sequence are included in the adjusted spectrum sequence on the low-domain side, and the remaining samples in the high-domain side spectrum sequence are placed on the low-domain side of the adjusted spectrum sequence on the high-domain side, thereby obtaining an adjusted spectrum sequence.
[0072] [Example 4 of the adjustment process performed by the friction noise adjustment unit 13]
[0073] Furthermore, it is desirable that the friction tone adjustment unit 13, in the adjustment target samples from the low-domain side spectrum sequence to the high-domain side in steps 1-2 described above, does not include one or more samples starting from the lowest frequency. This is because low-frequency samples contribute to the continuity of the signal waveform between frames and should be encoded with more bits allocated in the encoding unit 14. That is, when γ is set to a positive integer, X from the low-domain side spectrum sequence... γ ,…,X M-1 Select C samples to adjust, for example, X γ ,…,X γ+C-1 The target sample can be set as the target sample. Furthermore, increasing the value of γ increases the continuity of the signal waveform between frames, but relatively reduces the number of bits allocated to other samples in the encoding section 14, thus lowering the auditory quality of the decoded audio within the frame. Therefore, considering these factors, the value of γ can be determined through prior experiments.
[0074] In the above Figure 5 and Figure 6 In the example, let γ = 2, such that the adjustment target sample from the low-domain side spectrum sequence to the high-domain side does not include the two samples X0 and X1 starting from the lowest frequency in the low-domain side spectrum sequence.
[0075] In other words, if the fricative sound determination unit 12 determines that the sound is a fricative sound, the fricative sound adjustment unit 13 obtains the result of swapping a portion of the high-domain side in the low-domain side spectrum sequence and all or part of the same number of high-domain side spectrum sequences, as the adjusted spectrum sequence.
[0076] [Example 5 of the adjustment process performed by the friction noise adjustment unit 13]
[0077] In the encoding process of the encoding unit 14 described later, due to the constraint of the maximum number of bits obtained in the encoding process, sometimes no bits are allocated to several samples starting from the highest frequency in the adjusted spectral sequence. In this case, for the high-domain side spectral sequence X... M ,…,X N-1 One or more samples starting from the highest frequency in the sequence can be excluded from encoding, and the high-domain side spectrum sequence X can be used instead. M ,…,X N-1 The remaining samples located on the lower domain side are designated as encoding targets. Therefore, in this case, the friction tone adjustment unit 13 ensures that the adjustment target samples from the higher domain side spectrum sequence to the lower domain side in steps 1-3 above do not include one or more samples from the higher domain side spectrum sequence starting from the highest frequency.
[0078] In the above Figure 5 and Figure 6 In the example, this means that the four samples X, starting from the highest frequency in the high-domain side spectrum sequence, are not included. 28 ,…,X 31 It is included in the adjustment object sample from the high-domain side spectral sequence to the low-domain side.
[0079] In other words, if the fricative sound determination unit 12 determines that the sound is a fricative sound, the fricative sound adjustment unit 13 obtains an adjusted spectral sequence by swapping all or part of the low-domain side spectral sequence and a part of the high-domain side spectral sequence located in the low-domain side.
[0080] [Encoding section 14]
[0081] The adjusted spectrum sequence Y0,…,Y is input into the encoding unit 14 and output from the friction tone adjustment unit 13. N-1 The encoding unit 14 uses a method that prioritizes allocating bits to samples with smaller sample numbers on a frame-by-frame basis, for example, using the same method as in Non-Patent Document 1, to encode the input adjusted spectral sequence Y0,…,Y0. N-1 The spectrum code is obtained by encoding and then output to the multiplexing unit 15 (step S14).
[0082] Here, the method for prioritizing the allocation of bits to samples with smaller sample numbers is, for example, the following: The adjusted spectral sequence Y0,…,Y N-1The sequence is divided into multiple sub-sequences. For sub-sequences with smaller sample numbers, each sample within the sub-sequence is divided by the smallest gain value. The integer values of the division results are then encoded using variable-length or fixed-length codes, or vector quantized, to obtain the code corresponding to the adjusted spectral sequence, i.e., the spectral code. In this case, for sub-sequences with larger sample numbers, a code corresponding to that sub-sequence may not be obtained. That is, bits may not be allocated to sub-sequences with larger sample numbers.
[0083] For the adjusted spectral sequence Y0,…,Y N-1 The portion of the sequence with smaller sample numbers is encoded by dividing the values of the samples within that portion of the sequence by the gain of the smaller values, resulting in larger integer values. Therefore, each integer value is allocated more bits for encoding. On the other hand, for the adjusted spectral sequence Y0,…,Y… N-1 The portion of the sequence with the largest sample number is encoded by dividing the value of each sample in the partial sequence by the gain of the largest value, resulting in smaller integer values. Therefore, each integer value is encoded using fewer bits. Most of the integer values obtained by dividing the value of each sample in the partial sequence by the gain of the largest value are 0.
[0084] Moreover, such as Figure 1 As shown by the dashed line, if the fricative tone adjustment unit 13 and the encoding unit 14 are set as the fricative tone corresponding encoding unit 17, it can be said that when the fricative tone determination unit 12 determines that it is a fricative tone, the fricative tone corresponding encoding unit 17 encodes the spectrum sequence by encoding the bits preferentially allocated on the high domain side to obtain the spectrum code. In other cases, the fricative tone corresponding encoding unit 17 encodes the spectrum sequence by encoding the bits preferentially allocated on the low domain side to obtain the spectrum code.
[0085] [Reuse Section 15]
[0086] The multiplexing unit 15 receives the friction tone determination information output by the friction tone determination unit 12 and the spectrum code output by the encoding unit 14. The multiplexing unit 15 outputs a code, frame by frame, obtained by concatenating the code corresponding to the input friction tone determination information and the spectrum code (step S15). If the friction tone determination information output by the friction tone determination unit 12 is a 1-bit information, the friction tone determination information output by the friction tone determination unit 12 and input to the multiplexing unit 15 is itself set to the code corresponding to the friction tone determination information.
[0087] Decoding Device
[0088] Reference Figure 3 This explains the processing procedure of the decoding device in the first embodiment. For example... Figure 3As illustrated, the decoding apparatus of the first embodiment includes a multiplexing separation unit 21, a decoding unit 22, a fricative adjustment release unit 23, and a time-domain transformation unit 24. The code output by the encoding unit is input into the decoding apparatus. The code input to the decoding apparatus is input to the multiplexing separation unit 21. The decoding apparatus performs processing in each unit for a predetermined frame unit. The decoding method of the first embodiment performs the following through each unit of the decoding apparatus: Figure 4 The process is implemented by steps S21 to S24 as illustrated in the example.
[0089] [Reuse Separation Section 21]
[0090] The code is input to the encoding device in the multiplexing separation unit 21. The multiplexing separation unit 21 separates the input code into a code corresponding to the fricative sound determination information and a spectrum code in frame units, outputs the fricative sound determination information obtained from the code corresponding to the fricative sound determination information to the fricative sound adjustment release unit 23, and outputs the spectrum code to the decoding unit 22 (step S21).
[0091] When the friction tone determination information is set as 1 bit information, it is sufficient to set the code itself corresponding to the friction tone determination information input to the multiplexing separation unit 21 as the friction tone determination information.
[0092] [Decoding Section 22]
[0093] The spectrum code output by the multiplexing and demultiplexing unit 21 is input into the decoding unit 22. The decoding unit 22 decodes the input spectrum code frame by frame using a decoding method corresponding to the encoding method performed by the encoding unit 14 of the encoding apparatus to obtain the decoded adjusted spectrum sequence ^Y0,…,^Y. N-1 The resulting decoded adjusted spectral sequence ^Y0,…,^Y N-1 Output to the friction noise adjustment and release unit 23 (step S22).
[0094] When the encoding unit 14 of the encoding device decodes the spectrum code using a decoding method corresponding to the encoding method described above, the decoding unit 22 decodes the spectrum code to obtain an integer value string. It then combines multiple sample value sequences obtained by multiplying the gain of the smaller sample number portion of the sequence by the integer value to obtain the decoded adjusted spectrum sequence ^Y0,…,^Y. N-1 When a portion of the sequence with a large sample number is not bit-allocated in the encoding device, the value of the decoded adjusted spectrum corresponding to that portion of the sequence is set to 0. Furthermore, for samples with integer values of 0, multiplying by the gain also results in a value of 0, so the value of the decoded adjusted spectrum becomes 0. In other words, for a portion of the sequence with a large sample number, most integer values are 0, and the value of the decoded adjusted spectrum is mostly 0.
[0095] In this way, the decoding unit 22 decodes the spectrum code of the frame unit within the specified time interval, and the spectrum code that prioritizes the allocation of bits to the low domain side, to obtain a sample string in the frequency domain corresponding to the decoded audio signal (decoding the adjusted spectrum sequence).
[0096] [Friction noise adjustment and release section 23]
[0097] The friction sound determination information output from the multiplexing and demultiplexing unit 21 and the decoded adjusted spectrum sequence ^Y0,…,^Y output from the decoding unit 22 are input into the friction sound adjustment release unit 23. N-1 The fricative adjustment and de-escalation unit 23, on a frame-by-frame basis, adjusts the input decoded spectral sequence ^Y0,…,^Y when the input fricative determination information indicates that the tone is fricative. N-1 The following adjustment and removal process is performed to obtain the decoded spectrum sequence ^X0,…,^X N-1 The resulting decoded spectrum sequence ^X0,^X1,…,^X N-1 The output is sent to the time-domain transformation unit 24. If the fricative sound determination information indicates that the sound is not a fricative sound, the adjusted spectral sequence ^Y0,…,^Y is decoded. N-1 Directly used as the decoded spectral sequence ^X0,…,^X N-1 The output is sent to the time domain transformation unit 24 (step S23).
[0098] If we set an integer value greater than 1 and less than N as M, for example, we would decode the adjusted spectrum sequence ^Y0,…,^Y N-1 The samples with sample numbers less than M are ^Y0,…,^Y M-1 The sample group is set as the low-domain side to decode the adjusted spectrum sequence, and the decoded adjusted spectrum sequence ^Y0,…,^Y is... N-1 The samples with sample number M or higher are ^Y M ,…,^Y N-1 If the sample group is set as the high-domain side decoded adjusted spectrum sequence, then when the fricative determination information indicates a fricative tone, the adjustment removal process performed by the fricative adjustment removal unit 23 is as follows: the low-domain side decoded adjusted spectrum sequence ^Y0,…,^Y N-1 All or part of the samples, and the same number of high-domain side decoded adjusted spectral sequences^Y M ,…,^Y N-1 All or part of the samples are swapped, and the swapped result is used as the decoded spectrum sequence ^X0,…,^X N-1 The adjustment release process performed by the friction noise adjustment release unit 23 can include various processes including those exemplified below, but the adjustment release process is predetermined to be the reverse of the adjustment process performed by the friction noise adjustment unit 13 of the corresponding encoding device.
[0099] In other words, if the input information indicating whether a tone is fricative is indeed a fricative tone, the fricative adjustment release unit 23 will swap all or part of the low-domain frequency sample string (low-domain decoded adjusted spectrum sequence) located on the lower domain side of the frequency domain sample string obtained by the decoding unit 22 with all or part of the same number of high-domain frequency sample strings (high-domain decoded adjusted spectrum sequence) located on the higher domain side of the frequency domain sample string obtained by the decoding unit 22, and obtain the swapped result as the spectrum sequence of the decoded tone signal (decoded spectrum sequence). In other cases, the fricative adjustment release unit 23 directly uses the frequency domain sample string (decoded adjusted spectrum sequence) obtained by the decoding unit 22 as the spectrum sequence of the decoded tone signal (decoded spectrum sequence).
[0100] [Example 1 of the adjustment and release process performed by the friction noise adjustment and release unit 23]
[0101] When the fricative sound determination information indicates that the sound is fricative, the fricative sound adjustment release unit 23 obtains the decoded spectrum sequence ^X0,…,^X by performing steps 2-1 to 2-6 as described below. N-1 Furthermore, in order to easily understand the operation of the friction noise adjustment release unit 23, the following steps 2-1 to 2-6 are divided into 6 steps. However, the separation of the following steps 2-1 to 2-6 by the friction noise adjustment release unit 23 is just one example. It is also possible to perform the equivalent processing of steps 2-1 to 2-6 in one step by changing the arrangement elements or replacing the index.
[0102] Step 2-1: Decode the adjusted spectral sequence ^Y0,…,^Y N-1 The sample group of samples with sample numbers less than M is set as the low-domain side decoding adjusted spectrum sequence ^Y0,…,^Y M-1 The decoded spectral sequence ^Y0,…,^Y will be adjusted. N-1 The sample group of samples with sample number M or above is set as the high-domain side decoding adjusted spectrum sequence ^Y M ,…,^Y N-1 .
[0103] Step 2-2: Extract the low-domain side decoded adjusted spectrum sequence ^Y0,…,^Y obtained in Step 2-1. M-1 The C samples (C being a positive integer) contained therein are used as the adjustment targets for the higher domain side.
[0104] Step 2-3: Extract the high-domain side decoded adjusted spectrum sequence ^Y obtained in step 2-1. M ,…,^Y N-1The C samples contained therein are used as the adjustment targets for the lower domain side.
[0105] Step 2-4: Obtain the sample positions of the adjusted target samples from the low-domain decoded adjusted spectrum sequence extracted in Step 2-2, and configure the results of the adjusted target samples from the high-domain decoded adjusted spectrum sequence extracted in Step 2-3, as the low-domain decoded spectrum sequence ^X0,…,^X M-1 .
[0106] Step 2-5: Obtain the sample positions of the adjusted target samples from the high-domain decoded adjusted spectrum sequence extracted in Step 2-3, and configure the results of the adjusted target samples extracted from the low-domain decoded adjusted spectrum sequence in Step 2-2 as the high-domain decoded spectrum sequence^X. M ,…,^X N-1 .
[0107] Step 2-6: Decode the low-domain side spectral sequence ^X0,…,^X obtained in step 2-4. M-1 , and the high-domain side decoded spectrum sequence ^X obtained in steps 2-5 M ,…,^X N-1 By combining these sequences, we obtain the decoded spectral sequence ^X0,…,^X N-1 .
[0108] Figure 7 This example illustrates steps 2-1 to 2-6 when N=32, M=20, and C=8. The fricative tone adjustment release unit 23 first decodes the adjusted spectral sequence ^Y0,…,^Y... 31 In ^Y0,…,^Y 19 Set the low-domain side decoding to an adjusted spectral sequence, and set ^Y 20 ,…,^Y 31 Set the high-domain side decoded adjusted spectrum sequence (step 2-1). The fricative tone adjustment release unit 23 extracts the low-domain side decoded adjusted spectrum sequence ^Y0,…,^Y. 19 The eight samples ^Y2, ..., ^Y9 contained therein are used as adjustment targets for the higher domain side (step 2-2). The fricative tone adjustment release unit 23 extracts the decoded adjusted spectrum sequence ^Y from the higher domain side. 20 ,…,^Y 31 The 8 samples included^Y 20 ,…,^Y 27 As the adjustment target sample for the lower domain side (steps 2-3), the fricative tone adjustment release unit 23 obtains the sample positions where ^Y2, ..., ^Y9 exist in the decoded adjusted spectrum sequence on the lower domain side and configures ^Y. 20 ,…,^Y 27The result is the low-domain side decoded spectral sequence ^X0,…,^X 19 (Steps 2-4). The friction tone adjustment release unit 23 obtains the ^Y present in the decoded adjusted spectrum sequence on the high-domain side. 20 ,…,^Y 27 The sample positions are configured with the results of ^Y2,…,^Y9, which serve as the high-domain side decoded spectral sequence ^X. 20 ,…,^X 31 (Steps 2-5). The friction tone adjustment and release unit 23 decodes the low-domain side spectrum sequence ^X0,…,^X 19 Decoding the spectral sequence ^X on the high-domain side 20 ,…,^X 31 By combining these sequences, we obtain the decoded spectral sequence ^X0,…,^X 31 (Steps 2-6)
[0109] [Example 2 of the adjustment and release process performed by the friction noise adjustment and release unit 23]
[0110] When the friction noise adjustment unit 13 of the encoding device performs step 1-4' instead of step 1-4, the friction noise adjustment release unit 23 performs step 2-4' instead of the above-mentioned step 2-4.
[0111] Step 2-4': After extracting the adjustment target samples from the low-domain side decoded and adjusted spectrum sequence in Step 2-2, the remaining samples are squeezed towards both the low-domain and high-domain sides. The sample positions in the empty gaps are then filled with the adjustment target samples extracted from the high-domain side decoded and adjusted spectrum sequence in Step 2-3, resulting in the configured low-domain side decoded spectrum sequence ^X0,…,^X. M-1 .
[0112] [Example 3 of the adjustment process performed by the friction noise adjustment unit 13]
[0113] When the friction noise adjustment unit 13 of the encoding device performs step 1-5' instead of step 1-5, the friction noise adjustment release unit 23 performs step 2-5' instead of the above-mentioned step 2-5.
[0114] Step 2-5': After extracting the adjustment target samples from the high-domain side decoded and adjusted spectrum sequence in Step 2-3, the remaining samples are squeezed towards the high-domain side. The empty sample positions on the low-domain side are then configured as the high-domain side decoded spectrum sequence obtained in Step 2-2, where the adjustment target samples were extracted from the low-domain side decoded and adjusted spectrum sequence. M ,…,^X N-1 .
[0115] Figure 8This example illustrates how, when N=32, M=20, and C=8, step 2-4' is performed by replacing step 2-4 in steps 2-1 to 2-6, and step 2-5' is performed by replacing step 2-5. The fricative tone adjustment release unit 23 first decodes the adjusted spectrum sequence ^Y0,…,^Y. 31 In ^Y0,…,^Y 19 Set the low-domain side decoding to an adjusted spectral sequence, and set ^Y 20 ,…,^Y 31 Set the high-domain side decoded adjusted spectrum sequence (step 2-1). The fricative tone adjustment release unit 23 extracts the low-domain side decoded adjusted spectrum sequence ^Y0,…,^Y. 19 The 8 samples included^Y 12 ,…,^Y 19 As the sample to be adjusted towards the higher domain (step 2-2), the friction tone adjustment release unit 23 extracts the decoded adjusted spectrum sequence from the higher domain. 20 ,…,^Y 31 The 8 samples included^Y 24 ,…,^Y 31 As the sample to be adjusted towards the lower domain (steps 2-3), the friction tone adjustment release unit 23 squeezes ^Y0, ^Y1 in the decoded and adjusted spectrum sequence on the lower domain side towards the lower domain side, and squeezes ^Y2, ..., ^Y... 11 Squeezing towards the higher domain side, and configuring ^Y in the gaps that are left open. 24 ,…,^Y 31 The configured result is used as the low-domain side decoding spectrum sequence ^X0,…,^X 19 (Steps 2-4'). The friction tone adjustment release unit 23 decodes the ^Y in the adjusted spectrum sequence on the high-domain side. 20 ,…,^Y 23 Squeezing towards the higher domain side, and then ^Y after squeezing towards the higher domain side 20 ,…,^Y 23 Low-domain side configuration ^Y 12 ,…,^Y 19 The configured result is used as the high-domain side decoding spectrum sequence ^X 20 ,…,^X 31 (Steps 2-5'). The friction tone adjustment and release unit 23 decodes the low-domain side spectrum sequence ^X0,…,^X 19 and high-domain side decoding spectral sequence ^X 20 ,…,^X 31 By combining these sequences, we obtain the decoded spectral sequence ^X0,…,^X 31 (Steps 2-6)
[0116] [Example 4 of the adjustment and release process performed by the friction noise adjustment and release unit 23]
[0117] In step 1-2, if the friction tone adjustment unit 13 of the encoding device does not include one or more samples starting from the lowest frequency in the adjustment target samples from the low-domain side spectrum sequence to the high-domain side, the friction tone adjustment release unit 23 in step 2-2 makes it so that the adjustment target samples from the low-domain side decoded and adjusted spectrum sequence to the high-domain side do not include one or more samples starting from the lowest frequency.
[0118] [Example 5 of the adjustment and release process performed by the friction noise adjustment and release unit 23]
[0119] In steps 1-3, if the friction tone adjustment unit 13 of the encoding device does not include one or more samples starting from the highest frequency in the adjustment target samples from the high-domain side spectrum sequence to the low-domain side, the friction tone adjustment release unit 23 in steps 2-3 ensures that the friction tone adjustment target samples from the high-domain side decoded and adjusted spectrum sequence to the low-domain side do not include one or more samples starting from the highest frequency.
[0120] Moreover, such as Figure 3 As shown by the dashed line, if the decoding unit 22 and the fricative adjustment release unit 23 are set as the fricative corresponding decoding unit 26, then when the information indicating whether it is an input fricative tone indicates that it is a fricative tone, the fricative corresponding decoding unit 26 preferentially allocates bits in the high domain side of the spectrum code and decodes the spectrum code to obtain a spectrum sequence (decoded spectrum sequence). In other cases, the fricative corresponding decoding unit 26 preferentially allocates bits in the low domain side of the spectrum code and decodes the spectrum code to obtain a spectrum sequence (decoded spectrum sequence).
[0121] [Time Domain Transformation Unit 24]
[0122] The decoded spectrum sequence ^X0,…,^X is input to the friction tone adjustment and release unit 23 and output in the time-domain transformation unit 24. N-1 The time-domain transformation unit 24 uses a time-domain transformation method corresponding to the frequency-domain transformation method performed by the frequency-domain transformation unit 11 of the encoding apparatus, such as inverse MDCT, to transform the decoded spectrum sequence ^X0,…,^X per frame. N-1 The signal is transformed into a time domain signal to obtain a frame-unit audio signal (decoded audio signal) and output (step S24).
[0123] Furthermore, when the frequency domain transformation unit 11 of the encoding device applies auditory weighted filtering and companding processing to the spectrum sequence obtained by the transformation, the time domain transformation unit 24 transforms the result of the inverse filtering or inverse companding processing corresponding to these processing on the decoded spectrum sequence into a time domain signal and outputs the decoded audio signal obtained therefrom.
[0124] Furthermore, the decoding device can also be configured to output a frequency-domain decoded audio signal instead of a time-domain decoded audio signal. In this configuration, the decoding device may not include a time-domain transformation unit 24, and the decoded spectrum sequence of the frame unit obtained by the friction tone adjustment release unit 23 may be linked sequentially according to time intervals to output the frequency-domain decoded audio signal.
[0125] Effects and benefits
[0126] The encoding and decoding apparatus according to the first embodiment, by adding a fricative adjustment processing and a corresponding fricative adjustment release processing to the encoding processing and decoding processing structure designed to allocate more bits to the low-frequency spectrum as in the conventional way, can compress and encode even sound signals containing fricatives, thereby reducing auditory degradation.
[0127] Conventional techniques that can compress and encode audio signals, even those containing fricative sounds, to minimize auditory degradation include encoding / decoding techniques that prioritize bit allocation to high-energy subbands. However, these techniques require sending bit allocation information for each subband from the encoding side to the decoding side. In contrast, the encoding and decoding apparatus according to the first embodiment can perform compression encoding by sending only 1 bit of fricative sound determination information from the encoding side to the decoding side, thus minimizing auditory degradation even for audio signals containing fricative sounds.
[0128] <Modifications of the First Embodiment>
[0129] The variation of the first embodiment differs from the first embodiment only in the fricative sound determination unit 12 included in the encoding device. The other structures of the encoding device and the decoding device are the same as those of the first embodiment. Hereinafter, the operation of the fricative sound determination unit 12, which differs from the first embodiment, and the resulting effects on the encoding and decoding devices will be explained.
[0130] [Fricative Sound Judgment Section 12]
[0131] The friction sound determination unit 12 of the modified embodiment of the first embodiment has a comparison result storage unit (not shown).
[0132] The fricative sound determination unit 12 calculates the input spectrum sequence X0,…,X in each frame. N-1 The average energy of the samples located on the high-domain side relative to the input spectral sequence X0,…,X N-1 The larger the ratio of the average energy of the samples located on the lower domain side, the larger its value becomes. This serves as an indicator that the frame is a fricative sound. The comparison result information indicates whether the calculated index is greater than a predetermined threshold or is above the threshold.
[0133] The comparison result storage unit stores the comparison result information in an amount equivalent to a predetermined number of past frames. That is, the friction sound determination unit 12 stores the comparison result information calculated from the spectral sequence of the frame in a new frame-by-frame manner in the comparison result storage unit, and deletes the oldest stored comparison result information.
[0134] The fricative sound determination unit 12 uses comparison result information calculated from the spectrum sequence of the current frame and comparison result information of a predetermined number of past frames stored in the comparison result storage unit. If more than half of the comparison result information or more than half of the comparison result information indicates that it is greater than or above a predetermined threshold, it is determined to be a fricative sound. If not, it is determined not to be a fricative sound. The determination result is output as fricative sound determination information to the fricative sound adjustment unit 13 and the multiplexing unit 15.
[0135] Thus, it is also possible that, in multiple frames containing the frame, if the index of the ratio of the average energy of the spectrum on the high-domain side of the spectrum to the average energy of the spectrum on the low-domain side of the audio signal is greater than a predetermined threshold, or if the number of frames with a value greater than or equal to the threshold is greater than the number of frames that are not such, or if the number of frames that are not such is greater than or equal to the threshold, then the frame is determined to be a fricative sound.
[0136] The information used for determining the friction sound can be, for example, 1 bit of information, or the average of the absolute values and the average of the sum of squares of all or a portion of the sample values can be used as the average energy, which is the same as the friction sound determination unit 12 in the first embodiment.
[0137] Effects and benefits
[0138] If the processing in the encoding and decoding apparatus of the first embodiment is performed, for frames that undergo adjustment processing and adjustment de-processing, a decoded tone with less coding distortion in the high-domain components and more coding distortion in the low-domain components is obtained; for frames that do not undergo adjustment processing and adjustment de-processing, a decoded tone with more coding distortion in the high-domain components and less coding distortion in the low-domain components is obtained. Therefore, there is a possibility of waveform discontinuity in the decoded tone occurring at the boundary between frames that undergo adjustment processing and adjustment de-processing and frames that do not undergo adjustment processing and adjustment de-processing. That is, if the determination result of the friction sound determination unit 12 switches frequently, the waveform discontinuity of the decoded tone occurs frequently, and this discontinuity is perceived, which may lead to a deterioration in auditory quality. Compared with the encoding apparatus of the first embodiment, the encoding apparatus of the modified example of the first embodiment can suppress the frequent switching of the determination result of the friction sound determination unit 12, suppress the frequency of occurrence of waveform discontinuity in the decoded tone, and suppress the deterioration in auditory quality caused by the perceived discontinuity.
[0139] In the friction sound determination unit 12 of the modified embodiment of the first embodiment, although increasing the number of comparison result information used in the determination can suppress frequent switching of the determination result of the friction sound determination unit 12 and suppress the frequency of discontinuity in the waveform of the decoded sound, the number of comparison result information used in the determination needs to be determined by considering the trade-off between the degradation of auditory quality caused by perceived discontinuity and the auditory quality of each frame of decoded sound. For example, when the frame length is 3 ms, the number of comparison result information used in the determination can be set to 16.
[0140] <Second Implementation>
[0141] The system of the second embodiment of the present invention also includes an encoding device and a decoding device, just like the system of the first embodiment.
[0142] The second embodiment differs from the first embodiment in that the spectrum without allocated bits in the encoding device is restored in the decoding device; that is, the bandwidth is extended in the decoding device. In the second embodiment, the decoding device extends the bandwidth by decoding the adjusted spectrum sequence after the spectrum has been altered based on the fricative tone determination information. Regarding the spectrum without allocated bits in the encoding device, the time intervals of non-fricative tones are contained in the high domain, and the time intervals of fricative tones are contained in the low domain. Therefore, in the second embodiment, for the time intervals of non-fricative tones, the high domain spectrum is reproduced by copying the low domain spectrum, thereby extending the bandwidth; for the time intervals of fricative tones, the low domain spectrum is reproduced by copying the high domain spectrum, thereby extending the bandwidth.
[0143] In the second embodiment, the spectrum is copied by multiplying the spectrum of the source being copied by a gain. Therefore, in addition to the processing performed by the encoding device of the first embodiment, the encoding device of the second embodiment also calculates the gain used by the decoding device of the second embodiment and outputs a code corresponding to the calculated gain.
[0144] Encoding device
[0145] Reference Figure 9 The processing procedure of the encoding device in the second embodiment will be explained. For example... Figure 9 As illustrated, the encoding device of the second embodiment includes: a frequency domain transformation unit 11, a friction tone determination unit 12, a friction tone adjustment unit 13, an encoding unit 14, a bandwidth extension gain encoding unit 16, and a multiplexing unit 15. Figure 9 The encoding device of the second embodiment and Figure 1 The difference in the encoding device is that it has a band-spreading gain encoding unit 16, and the code output by the multiplexing unit 15 also includes the band-spreading gain code output by the band-spreading gain encoding unit 16. The other structures of the encoding device in the second embodiment, namely the frequency domain transformation unit 11, the fricative sound determination unit 12, the fricative sound adjustment unit 13, and the encoding unit 14, operate the same as those in the encoding device of the first embodiment, so only the key parts of the operation will be described below.
[0146] In the encoding apparatus, a time-domain audio signal is input in frame units of a predetermined time length. The time-domain audio signal input to the encoding apparatus is then input to the frequency domain transformation unit 11. The encoding apparatus performs processing in frame units of a predetermined time length in each component. The encoding method of the second embodiment performs the following through each component of the encoding apparatus: Figure 10 The process is implemented by steps S11 to S16 as illustrated in the example.
[0147] [Frequency Domain Transformation Unit 11]
[0148] Frequency domain transformation unit 11 transforms the time-domain audio signal input to the encoding device into an N-point frequency domain spectral sequence X0,…,X in frame units. N-1 Output afterwards (step S11).
[0149] [Fricative Sound Judgment Section 12]
[0150] The friction sound determination unit 12 uses the frequency domain transformation unit 11 to obtain the spectrum sequence X0,…,X in frame units. N-1Alternatively, the audio signal input to the time domain of the encoding device is used to determine whether the audio signal is a fricative tone, and the determination result is output as fricative tone determination information (step S12). The fricative tone determination unit 12 of the encoding device in the first embodiment outputs the fricative tone determination information to the fricative tone adjustment unit 13 and the multiplexing unit 15, but the fricative tone determination unit 12 of the encoding device in the second embodiment outputs the fricative tone determination information to the fricative tone adjustment unit 13 and the multiplexing unit 15, and also to the bandwidth spread gain encoding unit 16. Moreover, the fricative tone determination unit 12 of the encoding device in the second embodiment can also perform the same operation as the fricative tone determination unit 12 of the encoding device in the modified example of the first embodiment.
[0151] In other words, if the ratio of the average energy of the spectrum on the high-domain side to the average energy of the spectrum on the low-domain side in a certain frame's spectral sequence is greater than a predetermined threshold or is above the threshold, the fricative sound determination unit 12 determines that the sound signal is a fricative sound.
[0152] Furthermore, if, in multiple frames containing a certain frame, the index in which the ratio of the average energy of the spectrum on the high-domain side to the average energy of the spectrum on the low-domain side in the spectrum sequence is greater than a predetermined threshold, or if the number of frames exceeding the threshold is greater than the number of frames that are not such, or if the number of frames that are not such is greater than the number of frames that are not such, then the fricative sound determination unit 12 determines that the sound signal is a fricative sound.
[0153] [Friction adjustment section 13]
[0154] The fricative adjustment unit 13 adjusts the frequency spectrum sequence X0,…,X obtained by the frequency domain transformation unit 11 in frames, when the fricative determination information obtained by the fricative determination unit 12 indicates that the sound is a fricative. N-1 Perform spectrum adjustment processing to obtain the adjusted spectrum sequence Y0,…,Y N-1 The resulting adjusted spectral sequence Y0,…,Y N-1 The output is sent to the encoding unit 14. If the fricative determination information obtained by the fricative determination unit 12 indicates that the sound is not a fricative, the frequency spectrum sequence X0,…,X obtained by the frequency domain transformation unit 11 is then processed. N-1 Directly used as the adjusted spectral sequence Y0,…,Y N-1 Output to encoding unit 14 (step S13).
[0155] The frequency spectrum adjustment process performed by the friction sound adjustment unit 13 is as follows: the frequency spectrum sequence X0,…,X is adjusted. N-1 The low-domain side spectral sequence X0,…,X M-1 All or part of the samples, and the same number of spectral sequences X0, X…, X N-1 The high-domain side spectrum sequence XM ,…,X N-1 All or part of the samples are swapped, and the swapped result is used as the adjusted spectral sequence Y0,…,Y. N-1 .
[0156] In other words, when the fricative sound determination unit 12 determines that the sound is a fricative sound, the fricative sound adjustment unit 13 swaps all or part of the low-domain side spectrum sequence that is located in the lower domain compared to the specified frequency in the spectrum sequence of the sound signal, and all or part of the high-domain side spectrum sequence that is located in the higher domain compared to the specified frequency in the same number of spectrum sequences, and obtains the swapped result as the adjusted spectrum sequence. In other cases, the fricative sound adjustment unit 13 directly obtains the spectrum sequence corresponding to the sound signal as the adjusted spectrum sequence.
[0157] [Encoding section 14]
[0158] Encoding unit 14, in frame units, allocates bits preferentially to samples with smaller sample numbers, and processes the adjusted spectrum sequence Y0,…,Y obtained by friction tone adjustment unit 13. N-1 The spectrum code is obtained by encoding and then output to the multiplexing unit 15 (step S14).
[0159] The method of prioritizing bit allocation for samples with smaller sample numbers in the encoding unit 14 of the encoding apparatus in the first embodiment can be either a method of allocating bits for all samples in the adjusted spectrum sequence, or a method of not allocating bits for a portion of samples with larger sample numbers. In contrast, the method of prioritizing bit allocation for samples with smaller sample numbers in the encoding unit 14 of the encoding apparatus in the second embodiment is limited to a method of not allocating bits for a portion of the adjusted spectrum with larger sample numbers in the adjusted spectrum sequence. Furthermore, this bit allocation method is predetermined and stored in the encoding unit 14, and also stored in the band spread gain encoding unit 16 described later.
[0160] Encoding unit 14, for example, for the adjusted spectral sequence Y0,…,Y N-1 The K largest (K≦N / 2) adjusted spectra of the N adjusted spectra are Y. N-K ,…,Y N-1 Without allocating bits, start with the NK adjusted spectra Y0,…,Y0 from the side with the smaller remaining sample number. N-K-1 Allocate bits to the adjusted spectral sequence Y0,…,Y N-1 The spectrum code is obtained by encoding and then output to the multiplexing unit 15. That is, the encoding unit 14 essentially only encodes the adjusted spectrum sequence Y0,…,Y N-1 NK adjusted spectra Y0,…,Y, starting from the smaller sample number in the N adjusted spectra.N-K-1 The spectrum code is obtained by encoding.
[0161] [Bandwidth spread gain coding unit 16]
[0162] The adjusted spectrum sequence Y0,…,Y is at least input to the friction tone adjustment unit 13 output by the frequency band extension gain encoding unit 16. N-1 The band-spreading gain coding unit 16, in frame units, at least according to the input adjusted spectrum sequence Y0,…,Y N-1 The bandwidth spread gain code is obtained as described below, and the obtained bandwidth spread gain code is output to the multiplexing unit 15 (step S16).
[0163] The adjusted spectral sequence Y0,…,Y is set to be input only into the bandwidth spread gain coding section 16. N-1 In the case of the structure, for example as in Example 1 below, the band spread gain coding unit 16, in frame units, according to the input adjusted spectrum sequence Y0,…,Y N-1 The obtained bandwidth spread gain code is output to the multiplexing unit 15.
[0164] Furthermore, it can also be configured in the band-spreading gain coding section 16, in addition to the input adjusted spectrum sequence Y0,…,Y N-1 The structure also includes the friction sound determination information output by the friction sound determination unit 12. In this structure, for example, as in Example 2 below, the band spread gain coding unit 16, frame by frame, determines the friction sound determination information based on the input adjusted spectrum sequence Y0,…,Y0. N-1 The frequency band spread gain code is obtained from the friction sound determination information and then output to the multiplexing unit 15.
[0165] In the storage unit 161 of the band-spreading gain coding unit 16, multiple groups are pre-stored, each consisting of a candidate gain vector and a code capable of determining the candidate gain vector. Each candidate gain vector is composed of a number of candidate gain values from multiple samples. The band-spreading gain coding unit 16 obtains and outputs the code corresponding to the candidate gain vector as a band-spreading gain code on a frame-by-frame basis. The candidate gain vector is the one whose absolute value is the sum of the absolute value of the difference between the absolute value of the product of the adjusted spectrum value allocated by the coding unit 14 (with bits allocated) and the absolute value of the adjusted spectrum value not allocated by the coding unit 14 (without bits allocated). Alternatively, a square value or similar value can be used instead of the absolute value.
[0166] The following explains how the adjusted spectrum of bits allocated by the coding unit 14 is derived from the adjusted spectrum sequence Y0,…,Y N-1 The NK adjusted spectra Y0,…,Y starting from the sample number smaller in the sample numberN-K-1 The adjusted spectrum without allocated bits in the encoding section 14 is derived from the adjusted spectrum sequence Y0,…,Y N-1 The K adjusted spectra Y starting from the side with the larger sample number N-K ,…,Y N-1 Examples of such situations.
[0167] [Example 1 of the band-spreading gain coding unit 16]
[0168] In this example, we assume that J groups of gain candidate vectors and codes are stored in storage unit 161, and each gain candidate vector is composed of gain candidate values of K samples. Hereinafter, let's assume that the J gain candidate vectors are respectively designated as G... j (j=0,…,J-1), will be compared with the gain candidate vector G j Let the code for each of the (j=0,…,J-1) be C. Gj (j=0,…,J-1), each gain candidate vector G j Given K candidate gain values g j,k The explanation will be based on the structure (k=0,…,K-1).
[0169] The gain candidate vector G is output by the band-spreading gain coding unit 16 and stored in the storage unit 161. j E is obtained from (j=0,…,J-1) by the following equation (1). j The candidate vector for minimum gain G j The corresponding code C Gj As a frequency band extension gain code C G .
[0170] ···(1)
[0171] In other words, the band-spreading gain coding unit 16 obtains the code corresponding to the gain candidate vector as the band-spreading gain code and outputs it, wherein the gain candidate vector is the adjusted spectrum Y0,…,Y0 ... N-K-1 The K adjusted spectra Y starting from the side with the larger sample number N-2K ,…,Y N-K-1 The gain candidate value g that constitutes the gain candidate vector j,0 ,…,g j,K-1 The absolute value of the product of the two products |Y N-2K g j,0 |,…,|Y N-K-1 g j,K | Adjusted spectrum Y with no bits allocated to the encoding section 14 N-K ,…,Y N-1 Their respective absolute values |Y N-K |,…,|YN-1 The absolute value of the difference ||Y N-2K g j,0 |-|Y N-K ||,…,||Y N-K-1 g j,K |-|Y N-1 The sum of ||E j This is the candidate vector for the minimum gain.
[0172] [Example 2 of the band-spreading gain coding unit 16]
[0173] In this example, the storage unit 161 stores J groups of gain candidate vectors and codes, similar to Example 1. However, unlike Example 1, it is assumed that both gain candidate vectors for fricative sounds and gain candidate vectors for non-fricative sounds are stored as gain candidate vectors. That is, the storage unit 161 stores J groups of gain candidate vectors for fricative sounds and gain candidate vectors for non-fricative sounds and codes, where each gain candidate vector for fricative sounds and each gain candidate vector for non-fricative sounds is composed of gain candidate values of K samples. Hereinafter, the J gain candidate vectors for fricative sounds will be designated as G1. j (j=0,…,J-1), the J non-frictional tones are respectively set as G2 by the gain candidate vectors. j (j=0,…,J-1), with the friction tone using the gain candidate vector G1 j Each corresponding to (j=0,…,J-1) and associated with the non-friction tone using the gain candidate vector G2 j Let the code for each of the (j=0,…,J-1) be C. Gj (j=0,…,J-1) will be used for explanation. Furthermore, let each fricative tone be represented by a gain candidate vector G1. j The quantity of K samples, i.e., K candidate gain values g1 j,k (k=0,…,K-1) is used to construct each non-frictional tone, and the gain candidate vector G2 is used for each tone. j The quantity of K samples, i.e., K candidate gain values g2 j,k The explanation will be based on the structure (k=0,…,K-1).
[0174] When the input fricative sound determination information indicates that the sound is fricative, the band-spreading gain coding unit 16 uses the gain candidate vector G1 to encode the fricative sound stored in the storage unit 161. j (j=0,…,J-1) is set as the candidate gain vector G. j (j=0,…,J-1), when the input fricative determination information indicates that the tone is not a fricative tone, the band-spreading gain coding unit 16 uses the gain candidate vector G2 to encode the non-fricative tone stored in the storage unit 161. j (j=0,…,J-1) is set as the candidate gain vector G.j (j=0,…,J-1), will be compared with the gain candidate vector G j E in (j=0,…,J-1) obtained by equation (1) above j The candidate vector for minimum gain G j The corresponding band spread gain code C Gj As a frequency band extension gain code C G Output.
[0175] In other words, when the input fricative tone determination information indicates a fricative tone, the band-spreading gain coding unit 16 sets the fricative tone stored in the storage unit 161 as a gain candidate vector using the gain candidate vector. When the input fricative tone determination information indicates a non-fricative tone, the band-spreading gain coding unit 16 sets the non-fricative tone stored in the storage unit 161 as a gain candidate vector using the gain candidate vector, obtains the code corresponding to the gain candidate vector as the band-spreading gain code, and outputs it. The gain candidate vector is the adjusted spectrum Y0,…,Y0 ... N-K-1 The K adjusted spectra Y starting from the side with the larger sample number N-2K ,…,Y N-K-1 The gain candidate value g that constitutes the gain candidate vector j,0 ,…,g j,K-1 The absolute value of the product of the two products |Y N-2K g j,0 |,…,|Y N-K-1 g j,K-1 | Adjusted spectrum Y with no bits allocated to the encoding section 14 N-K ,…,Y N-1 Their respective absolute values |Y N-K |,…,|Y N-1 The absolute value of the difference ||Y N-2K g j,0 |-|Y N-K ||,…,||Y N-K-1 g j,K-1 |-|Y N-1 The sum of ||E j This is the candidate vector for the minimum gain.
[0176] Alternatively, the band-spreading gain coding unit 16 may store multiple codes, a gain candidate vector for fricatives corresponding to each code, and a gain candidate vector for non-fricatives corresponding to each code. If the fricatives determination unit 12 determines that the sound is a fricative, the band-spreading gain coding unit 16 uses the gain candidate vector for fricatives as the gain candidate vector. Otherwise, the band-spreading gain coding unit 16 uses the gain candidate vector for non-fricatives as the gain candidate vector.
[0177] [Example 1 of Examples 1 and 2 of the Bandwidth Spread Gain Coding Unit 16]
[0178] In Examples 1 and 2 above, the adjusted spectrum of the object of the multiplication operation, which is set as the gain candidate value, is set as the adjusted spectrum Y0,…,Y0 that has been allocated bits from the coding unit 14. N-K-1 The K adjusted spectra Y starting from the side with the larger sample number N-2K ,…,Y N-K-1 However, the adjusted spectrum of the object of the multiplication operation set as the gain candidate value is only the adjusted spectrum Y0,…,Y0 of the bits allocated by the coding section 14. N-K-1 The K adjusted spectra corresponding to the K predetermined sample numbers are obtained.
[0179] [Example 2 of Example 1 and Example 2 of Bandwidth Spread Gain Coding Unit 16]
[0180] In Examples 1 and 2 above, Y is in ascending order of the value of k in equation (1). N-2K+k, g j,k, Y N-K+k They are related, but any kind of relationship is acceptable as long as it is predetermined.
[0181] [A specific example of the bandwidth spread gain coding unit 16]
[0182] This section describes a specific example of the band-spreading gain coding unit 16 when N=32 and K=12. This specific example corresponds to a variation 2 of example 2 of the band-spreading gain coding unit 16. Figure 13 and Figure 14 The following is an example of the bandwidth extension unit 25 and the friction tone adjustment release unit 23 of the decoding device described later when N=32 and K=12.
[0183] Figure 13 This is an example of a case where the fricative tone determination information indicates a tone that is not fricative. As described later, the bandwidth extension unit 25 of the decoding device performs the process of setting the 8th to 19th decoded adjusted spectra as copy sources, obtaining the values of the decoded adjusted spectra of these copy sources multiplied by the bandwidth extension gain, and using them as the 20th to 31st decoded extended spectra in sample number order. Therefore, when the input fricative tone determination information indicates a tone that is not fricative, the bandwidth extension gain encoding unit 16 sets the non-fricative tone stored in the storage unit 161 as a gain candidate vector, and obtains the code corresponding to the gain candidate vector as the bandwidth extension gain code, wherein the gain candidate vector is the adjusted spectrum Y0,…,Y0 that has been allocated bits from the encoding unit 14. 19 The 12 adjusted spectra Y8, ..., Y starting from the side with the larger sample number 19The gain candidate values g that constitute the gain candidate vector j,0 ,…,g j,11 The absolute value of the product of the two products |Y8g j,0 |,…,|Y 19 g j,11 | Adjusted spectrum Y with no bits allocated to the encoding section 14 20 ,…,Y 31 Their respective absolute values |Y 20 |,…,|Y 31 |Absolute value of the difference||Y8g j,0 |-|Y 20 ||,…,||Y 19 g j,11 |-|Y 31 The sum of ||E j This is the candidate vector for the minimum gain.
[0184] Figure 14 This is an example of a case where the fricative tone determination information indicates a fricative tone. The bandwidth extension unit 25 of the decoding device performs the following processing as described later: The 8th to 19th decoded adjusted spectra are set as copy sources. The values of the decoded adjusted spectra of these copy sources are multiplied by the bandwidth extension gain to obtain a result where the sample numbers following the 16th to 19th are in the order of the 8th to 15th sample numbers, which is used as the decoded extended spectra of the 20th to 31st samples. Therefore, when the input fricative tone determination information indicates a fricative tone, the bandwidth extension gain encoding unit 16 sets the fricative tone stored in the storage unit 161 using a gain candidate vector as a gain candidate vector, and obtains a code corresponding to the gain candidate vector as a bandwidth extension gain code. The gain candidate vector is the adjusted spectrum Y0,…,Y0 ... 19 The 12 adjusted spectra starting from the larger of the sample numbers in the sample are Y8,…,Y 19 The gain candidate value g that constitutes the gain candidate vector j,0 ,…,g j,11 The absolute value of the product of the two products |Y8g j,0 |,…,|Y 19 g j,11 | Adjusted spectrum Y with no bits allocated to the encoding section 14 24 ,…,Y 31 ,Y 20 ,…,Y 23 Their respective absolute values |Y 24 |,…,|Y 31 |,|Y 20 |,…,|Y 23 |Absolute value of the difference||Y8g j,0 |-|Y24 ||,…,||Y 15 g j,7 |-|Y 31 ||,||Y 16 g j,8 |-|Y 20 ||,…,||Y 19 g j,11 |-|Y 23 The sum of ||E j This is the candidate vector for the minimum gain.
[0185] Thus, the bandwidth-spreading gain coding unit 16 stores multiple codes and gain candidate vectors corresponding to each code. Each gain candidate vector contains K gain candidate values (K is an integer greater than or equal to 2). The bandwidth-spreading gain coding unit 16 obtains the code corresponding to the gain candidate vector as the bandwidth-spreading gain code and outputs it. The gain candidate vector is the gain candidate vector whose error is minimized by multiplying the K values of the adjusted spectrum in the adjusted spectrum sequence (which has K bits allocated by the coding unit 14) with the K gain candidate values contained in the gain candidate vector, and by multiplying the K values of the adjusted spectrum in the adjusted spectrum sequence (which has K bits not allocated by the coding unit 14) with the K values of the adjusted spectrum.
[0186] The operation of the bandwidth extension gain encoding unit 16 corresponds to the operation of the bandwidth extension unit 25 and the friction tone adjustment release unit 23 of the decoding device. Figure 8 In the example, the friction tone adjustment release unit 23 of the decoding device sets the 20th to 23rd decoded spread spectrum (sample number smaller on the side of the 20th to 31st decoded spread spectrum) as the decoded spectrum with sample numbers from 28th to 31st, and sets the 24th to 31st decoded spread spectrum (sample number larger on the side of the 20th to 31st decoded spread spectrum) as the decoded spectrum with sample numbers from 2nd to 9th. The bandwidth extension unit 25 of the decoding device takes into account the frequency of the decoded spectrum obtained by the operation of the friction tone adjustment release unit 23 and performs... Figure 14 The action.
[0187] That is, regardless of whether the fricative sound determination information indicates a fricative sound or not, the bandwidth extension unit 25 of the decoding device performs frequency matching processing with the frequency range in the decoded spectrum. Therefore, the bandwidth extension gain encoding unit 16 also performs operations corresponding to those of the bandwidth extension unit 25.
[0188] [Reuse Section 15]
[0189] The multiplexing unit 15 receives the friction sound determination information output by the friction sound determination unit 12, the spectrum code output by the encoding unit 14, and the bandwidth spread gain code output by the bandwidth spread gain encoding unit 16. The multiplexing unit 15 outputs a code obtained by concatenating the code corresponding to the input friction sound determination information, the spectrum code, and the bandwidth spread gain code (step S15).
[0190] Decoding Device
[0191] Reference Figure 11 The processing procedure of the decoding device in the second embodiment will be explained. Figure 11 As illustrated in the example, the decoding device of the second embodiment includes a multiplexing separation unit 21, a decoding unit 22, a bandwidth extension unit 25, a friction tone adjustment release unit 23, and a time domain transformation unit 24. Figure 11 The decoding device of the second embodiment and Figure 3 The first embodiment of the decoding device differs in that it has a bandwidth extension unit 25, and the multiplexing separation unit 21 also obtains a bandwidth extension gain code from the input code. The other structures of the second embodiment's decoding device, namely the operation of the decoding unit 22, the fricative tone adjustment release unit 23, and the time-domain transformation unit 24, are the same as those of the first embodiment's decoding device; therefore, only the essential parts of the operation will be described below.
[0192] The code output by the encoding device is input into the decoding device. The code input to the decoding device is input to the multiplexing separation unit 21. The decoding device performs frame unit processing of a predetermined time length in each component. The decoding method of the second embodiment performs the following through each component of the decoding device: Figure 12 The process is implemented by steps S21 to S25 as illustrated in the example.
[0193] [Reuse Separation Section 21]
[0194] The multiplexing separation unit 21 separates the input code into a code corresponding to the friction tone determination information, a frequency band extension gain code, and a spectrum code. It outputs the friction tone determination information obtained from the code corresponding to the friction tone determination information to the friction tone adjustment release unit 23 and the frequency band extension unit 25, outputs the frequency band extension gain code to the frequency band extension unit 25, and outputs the spectrum code to the decoding unit 22 (step S21).
[0195] [Decoding Section 22]
[0196] The decoding unit 22 decodes the input spectrum code in frame units by performing decoding processing corresponding to the encoding processing performed by the encoding unit 14 of the encoding device, and outputs the decoded adjusted spectrum sequence (step S22).
[0197] As described above, since the encoding unit 14 of the encoding apparatus in the second embodiment performs encoding processing that does not allocate bits to samples with a large portion of the sample numbers, even if the spectrum code is decoded, the values of the decoded adjusted spectrum for these sample numbers cannot be obtained. In the case of the encoding unit 14 described above, the decoding unit 22 decodes the spectrum code to obtain NK decoded adjusted spectra starting from the smaller sample number: ^Y0,...,^Y N-K-1 The decoding has adjusted the spectral sequence.
[0198] Furthermore, the value of the decoded adjusted spectrum for sample numbers for which no bits are allocated in the encoding unit 14 can also be set to 0. That is, in the case of the encoding unit 14 described above, the decoding unit 22 can also decode the spectrum code and decode the K decoded adjusted spectra of the sample numbers starting from the larger one. N-K ,…,^Y N-1 Each value is set to 0, resulting in the decoded adjusted spectrum sequence ^Y0,…,^Y N-1 .
[0199] In this way, the decoding unit 22 decodes the spectrum code of the frame unit within the specified time interval, and the spectrum code that has not been allocated bits to a part of the high domain side, to obtain the sample string in the frequency domain (decoding the adjusted spectrum sequence).
[0200] However, as will be described later, if the input information indicating whether a tone is fricative indicates that it is a fricative tone, the fricative tone adjustment release unit 23 obtains the result of swapping all or part of the low-domain frequency sample string located on the lower domain side compared to the specified frequency in the decoded extended spectrum sequence (spectrum sequence based on the decoded adjusted spectrum sequence) obtained by the band extension unit 25 (described later) with all or part of the high-domain frequency sample string located on the higher domain side compared to the specified frequency in the decoded extended spectrum sequence obtained by the band extension unit 25, and uses this result as the spectrum sequence of the decoded tone signal. In other cases, the fricative tone adjustment release unit 23 directly uses the decoded extended spectrum sequence obtained by the band extension unit 25 as the spectrum sequence of the decoded tone signal. That is, if the input information indicating whether a sound is a fricative indicates that it is a fricative, the decoding unit 22 is configured not to allocate bits to a portion of the low-domain side of the spectrum code, and decodes the spectrum code to obtain a frequency domain spectrum sequence (decodes an adjusted spectrum sequence). In other cases, the decoding unit 22 is configured not to allocate bits to a portion of the high-domain side of the spectrum code, and decodes the spectrum code to obtain a frequency domain spectrum sequence (decodes an adjusted spectrum sequence).
[0201] The decoding unit 22 of the decoding device in the first embodiment outputs the obtained decoded adjusted spectrum sequence to the friction tone adjustment release unit 23, but the decoding unit 22 of the decoding device in the second embodiment outputs the obtained decoded adjusted spectrum sequence to the bandwidth extension unit 25.
[0202] [Bandwidth extension section 25]
[0203] In the bandwidth extension section 25, at least the bandwidth extension gain code output from the multiplexing separation section 21 and the decoded adjusted spectrum sequence output from the decoding section 22 are input. The bandwidth extension section 25 obtains the decoded extended spectrum sequence frame by frame, based at least on the input bandwidth extension gain code and the decoded adjusted spectrum sequence, as described below. ~ Y0,…, ~ Y N-1 The resulting decoded spread spectrum sequence ~ Y0,…, ~ Y N-1 Output to the friction noise adjustment and release unit 23 (step S25).
[0204] In the case where only the band-spreading gain code and the decoded adjusted spectrum sequence are input to the band-spreading unit 25, as in Example 1 below, the band-spreading unit 25 obtains the decoded spread spectrum sequence frame by frame based on the input band-spreading gain code and the decoded adjusted spectrum sequence. ~ Y0,…, ~ Y N-1 The resulting decoded spread spectrum sequence ~ Y0,…, ~ Y N-1 Output to the friction noise adjustment and release unit 23.
[0205] Furthermore, the band extension unit 25 can also be configured to receive, in addition to the input band extension gain code and the decoded adjusted spectrum sequence, the friction tone determination information output by the multiplexing separation unit 21. With this configuration, for example as in Example 2 below, the band extension unit 25 obtains the decoded extended spectrum sequence frame by frame based on the input band extension gain code, the decoded adjusted spectrum sequence, and the friction tone determination information. ~ Y0,…, ~ Y N-1 The resulting decoded spread spectrum sequence ~ Y0,…, ~ Y N-1 Output to the friction noise adjustment and release unit 23.
[0206] In the storage unit 251 of the band extension unit 25, the same as that stored in the storage unit 161 of the band extension gain coding unit 16 of the coding device, multiple groups are pre-stored, each consisting of a gain candidate vector (which is a candidate for gain vector) and a code that can determine the gain candidate vector. Each gain candidate vector consists of multiple sample values of gain candidate values. The band extension unit 25 obtains a sequence consisting of the result of multiplying each sample value of the copy source with each band extension gain contained in the gain candidate vector determined by the code corresponding to the band extension gain code, and setting the result as the decoded extended spectrum corresponding to the adjusted spectrum that has not been allocated bits in the coding unit 14 of the coding device, and the result of directly setting the decoded adjusted spectrum obtained by decoding the spectrum code as the decoded extended spectrum. The copy source is all or part of the decoded adjusted spectrum obtained by decoding the spectrum code (the decoded adjusted spectrum corresponding to the adjusted spectrum that has been allocated bits in the coding unit 14 of the coding device).
[0207] The following explains how the adjusted spectrum of bits allocated by the coding unit 14 is derived from the adjusted spectrum sequence Y0,…,Y N-1 The NK adjusted spectra Y0,…,Y starting from the sample number smaller in the sample number N-K-1 The adjusted spectrum without allocated bits in the coding section 14 is derived from the adjusted spectrum sequence Y0,…,Y N-1 The K adjusted spectra Y starting from the side with the larger sample number N-K ,…,Y N-1 Examples of this situation include: that is, illustrating how decoding the spectral code yields the decoded adjusted spectral sequence ^Y0,…,^Y. N-K-1 Examples of such cases. [Example 1 of the bandwidth extension section 25]
[0208] In this example, let J groups of gain candidate vectors and codes be stored in storage unit 251, where each gain candidate vector is composed of gain candidate values equivalent to K samples. Hereinafter, let the J gain candidate vectors be denoted as G... j (j=0,…,J-1), will be compared with the gain candidate vector G j The code corresponding to each of (j=0,…,J-1) is set to C. Gj (j=0,…,J-1), each gain candidate vector G j Let g be the quantity of K samples, i.e., K candidate gain values. j,k The explanation will be based on the structure (k=0,…,K-1).
[0209] The bandwidth extension unit 25 will decode the adjusted spectrum ^Y0,…,^Y N-K-1 Directly set it to NK decoded spread spectra starting from the smaller sample number of the decoded spread spectrum sequence. ~ Y0,…, ~Y N-K-1 The bandwidth extension unit 25 also stores the gain candidate vector G from the storage unit 251. j In (j=0,…,J-1), obtain the code C corresponding to the input. Gj The K gain candidate values contained in the gain candidate vector with equal band-spreading gain codes are used as the band-spreading gain g0,…,g K-1 The bandwidth extension unit 25 further extends the decoded adjusted spectrum ^Y0,…,^Y N-K-1 The K decoders starting with the larger sample number have had their spectrum adjusted. N-2K ,…,^Y N-K-1 and bandwidth spread gain g0,…,g K-1 The value of ^Y after multiplying them separately N-2K g0,…,^Y N-K-1 g K-1 Let K be the K decoded spread spectra starting from the side with the larger sample number in the decoded spread spectrum sequence. ~ Y N-K ,…, ~ Y N-1 .
[0210] [Example 2 of the bandwidth extension section 25]
[0211] In this example, J groups of gain candidate vectors and codes are stored in storage unit 251, similar to Example 1. However, unlike Example 1, two types of gain candidate vectors are stored: one for fricative sounds and one for non-fricative sounds. That is, J groups of gain candidate vectors for fricative sounds, one for non-fricative sounds, and codes are stored in storage unit 251, and each gain candidate vector for fricative sounds and each gain candidate vector for non-fricative sounds is composed of the gain candidate values of K samples. Hereinafter, the J gain candidate vectors for fricative sounds will be designated as G1. j (j=0,…,J-1), let the J non-frictional tones be represented by the gain candidate vectors G2 respectively. j (j=0,…,J-1), will be compared with the friction tone using the gain candidate vector G1 j Each corresponding to (j=0,…,J-1) and associated with the non-friction tone using the gain candidate vector G2 j Let the code for each of the (j=0,…,J-1) be C. Gj (j=0,…,J-1) are used for explanation. Furthermore, each fricative tone is represented by a gain candidate vector G1. j Let g1 be the quantity of K samples, i.e., K candidate gain values. j,k (k=0,…,K-1) is used to construct each non-frictional tone, and the gain candidate vector G2 is used for each tone. j Let g2 be the quantity of K samples, i.e., K candidate gain values. j,kThe structure (k=0,…,K-1) is used to illustrate this.
[0212] The bandwidth extension unit 25 will decode the adjusted spectrum ^Y0,…,^Y N-K-1 Directly set it to NK decoded spread spectra starting from the smaller sample number of the decoded spread spectrum sequence. ~ Y0,…, ~ Y N-K-1 Furthermore, when the input fricative sound determination information indicates that the sound is fricative, the bandwidth extension unit 25 uses the fricative sound gain candidate vector G1 stored in the storage unit 251. j (j=0,…,J-1) is set as the candidate gain vector G. j (j=0,…,J-1), when the input fricative sound determination information indicates that the sound is not a fricative sound, the bandwidth extension unit 25 uses the gain candidate vector G2 to store the non-fricative sound in the storage unit 251. j (j=0,…,J-1) is set as the candidate gain vector G. j (j=0,…,J-1), thus obtaining the gain candidate vector G. j The symbol C corresponding to the input is in (j=0,…,J-1). Gj The K gain candidate values contained in the gain candidate vector with equal band-spreading gain codes are used as the band-spreading gain g0,…,g K-1 The bandwidth extension unit 25 further extends the decoded adjusted spectrum ^Y0,…,^Y N-K-1 The K decoders starting with the larger sample number have had their spectrum adjusted. N-2K ,…,^Y N-K-1 With bandwidth spread gain g0,…,g K-1 The value of ^Y after multiplying them separately N-2K g0,…,^Y N-K-1 g K-1 Let K be the decoded spread spectra starting from the side with the larger sample number in the decoded spread spectrum sequence. ~ Y N-K ,…, ~ Y N-1 .
[0213] [Example 1 of the variations of Example 1 and Example 2 of the bandwidth extension section 25]
[0214] In Examples 1 and 2 above, let the decoded adjusted spectrum of the object of the multiplication operation of the band spread gain be denoted as the decoded adjusted spectrum ^Y0,…,^Y obtained by decoding the spectrum code. N-K-1 The K adjusted spectra starting from the side with the larger sample number ^Y N-2K ,…,^Y N-K-1However, the decoded adjusted spectrum of the object of the multiplication operation denoted by the band-spreading gain is simply the decoded adjusted spectrum ^Y0,…,^Y obtained by decoding the spectral code. N-K-1 The K decoders corresponding to the pre-determined K sample numbers have already had their spectra adjusted.
[0215] [Example 2 of the variations of Example 1 and Example 2 of the bandwidth extension section 25]
[0216] In Examples 1 and 2 above, decoding with k values from smallest to largest has adjusted the spectrum^Y. N-2K+k The bandwidth extension gain g increases with the value of k from smallest to largest. k Multiplying them together gives the decoded spread spectrum of k in ascending order. ~ Y N-K+k That is, perform associations with k values from smallest to largest, but any association is acceptable as long as it is a pre-determined association.
[0217] [Specific example of bandwidth extension 25]
[0218] A specific example of the bandwidth extension section 25 with N=32 and K=12 will be explained. This specific example corresponds to variation 2 of example 2 of the bandwidth extension section 25. Figure 13 and Figure 14 This is an example of the processing of the frequency band extension section 25 and the friction sound adjustment release section 23 when N=32 and K=12.
[0219] Figure 13 This is an example of a fricative sound determination information indicating a sound that is not a fricative sound. The bandwidth extension unit 25 adjusts the spectrum ^Y0,…,^Y obtained through decoding the spectrum code. 19 Set directly to decode spread spectrum ~ Y0,…, ~ Y 19 The bandwidth extension section 25 also obtains the code C corresponding to the input. Gj The 12 gain candidate values contained in the gain candidate vector with equal band-spreading gain codes are used as the band-spreading gain g0,…,g 11 The bandwidth extension unit 25 further extends the decoded adjusted spectrum ^Y0,…,^Y 19 The first 12 decoders, starting with the sample number with the larger value, have had their spectra adjusted ^Y8,…,^Y 19 With bandwidth spread gain g0,…,g 11 The values after multiplying them separately are ^Y8g0,…,^Y 19 g 11 Let K be the decoded spread spectra starting from the side with the larger sample number in the decoded spread spectrum sequence. ~ Y 20 ,…, ~ Y31 .
[0220] Figure 14 This is an example of a fricative sound indicating that the sound is fricative. The bandwidth extension unit 25 adjusts the spectrum ^Y0,…,^Y obtained from decoding the spectrum code. 19 Set directly to decode spread spectrum ~ Y0,…, ~ Y 19 The bandwidth extension section 25 also obtains the code C corresponding to the input. Gj The 12 gain candidate values contained in the gain candidate vector with equal band-spreading gain codes are used as the band-spreading gain g0,…,g 11 The bandwidth extension unit 25 further extends the decoded adjusted spectrum ^Y0,…,^Y 19 The first 12 decoders, starting with the sample number with the larger value, have had their spectra adjusted ^Y8,…,^Y 19 With bandwidth spread gain g0,…,g 11 The values after multiplying them separately are ^Y8g0,…,^Y 19 g 11 Let K be the K decoded spread spectra starting from the side with the larger sample number in the decoded spread spectrum sequence. ~ Y 24 ,…, ~ Y 31, ~ Y 20 ,…, ~ Y 23 That is, the bandwidth extension section 25 performs the following processing: it decodes the adjusted spectrum from the 8th to the 19th frequency band ^Y8,…,^Y 19 Set as a copy source, and adjust the spectrum ^Y8,…,^Y of the decoded copies of these sources. 19 The values of the bandwidth spread gain g0,…,g 11 The value after multiplication is ^Y8g0,…,^Y 19 g 11 The decoded spread spectrum is arranged in the order corresponding to the 16th to 19th sample numbers of the decoded adjusted spectrum. ~ Y 20 =^Y 16 g8,…, ~ Y 23 =^Y 19 g 11 The following is the decoded extended spectrum corresponding to the sample numbers from the 8th to the 15th of the decoded adjusted spectrum. ~ Y 24 =^Y8g0,…, ~ Y 31 =^Y 15The result of g7's sequence becomes the decoded spread spectrum from the 20th to the 31st. ~ Y 20 ,…, ~ Y 31 .
[0221] The operation of the frequency band extension unit 25 corresponds to the operation of the friction noise adjustment release unit 23. Figure 8 In the example, the friction tone adjustment release unit 23 decodes the 20th to 31st extended spectrum. ~ Y 20 ,…, ~ Y 31 The 20th to 23rd decoded spread spectrum on the side with the smaller sample number. ~ Y 20 ,…, ~ Y 23 Let ^X be the decoding spectrum from sample number 28 to 31. 28 ,…,^X 31 Decode the 20th to 31st extended spectrum ~ Y 20 ,…, ~ Y 31 The 24th to 31st decoded spread spectrum on the side with the larger sample number. ~ Y 24 ,…, ~ Y 31 Let the decoded spectrum be ^X2, ...,^X8 from the 2nd to the 9th sample number. The bandwidth extension unit 25 considers the frequency range of the decoded spectrum obtained through the operation of the friction tone adjustment release unit 23, and performs... Figure 14 The operation involves matching the frequency range of the decoding device 25 with the frequency range in the decoding spectrum, regardless of whether the fricative sound determination information indicates a fricative sound or not.
[0222] In this way, the bandwidth extension unit 25 configures samples based on K (K is an integer greater than 2) of the frequency domain sample string (decoded adjusted spectrum sequence) obtained by decoding the spectrum code by the decoding unit 22 on the high domain side, thereby obtaining the decoded extended spectrum sequence.
[0223] More specifically, for example, the band extension unit 25 obtains a group of K band extension gains by decoding the band extension gain code. On the higher domain side of the sample string (decoded adjusted spectrum sequence) in the frequency domain obtained by decoding the spectrum code obtained by decoding unit 22, K samples contained in the sample string in the frequency domain obtained by decoding unit 22 are multiplied by K band extension gains, thereby obtaining a decoded extended spectrum sequence.
[0224] Furthermore, assuming that the bandwidth extension unit 25 stores multiple codes, friction tone gain candidate vectors corresponding to each code, and non-friction tone gain candidate vectors corresponding to each code, and assuming that each friction tone gain candidate vector and non-friction tone gain candidate vector contains K gain candidate values, the bandwidth extension unit 25 can decode the bandwidth extension gain code to obtain a group of K bandwidth extension gains as follows: if the information indicating whether it is an input friction tone indicates that it is a friction tone, the K gain candidate values contained in the friction tone gain candidate vector whose corresponding code is the same as the bandwidth extension gain code are set as a group of K bandwidth extension gains; otherwise, the K gain candidate values contained in the non-friction tone gain candidate vector whose corresponding code is the same as the bandwidth extension gain code are set as a group of K bandwidth extension gains.
[0225] [Friction noise adjustment and release section 23]
[0226] The friction noise determination information output by the multiplexing separation unit 21 and the decoded spread spectrum sequence output by the bandwidth extension unit 25 are input into the friction noise adjustment and release unit 23. ~ Y0,…, ~ Y N-1 When the fricative determination information input in frames indicates that a sound is a fricative, the fricative adjustment release unit 23 decodes the input spread spectrum sequence. ~ Y0,…, ~ Y N-1 After adjustment and removal processing, the decoded spectrum sequence ^X0,…,^X is obtained. N-1 The resulting decoded spectrum sequence ^X0,…,^X N-1 The output is sent to the time-domain transformation unit 24. If the fricative sound determination information indicates that the sound is not a fricative sound, the fricative sound adjustment release unit 23 will decode the spread spectrum sequence. ~ Y0,…, ~ Y N-1 Directly used as the decoded spectral sequence ^X0,…,^X N-1 The output is sent to the time domain transformation unit 24 (step S23).
[0227] The adjustment and release process performed by the friction tone adjustment and release unit 23 is to process the decoded spread spectrum sequence. ~ Y0,…, ~ Y N-1 The friction tone adjustment release unit 23 of the decoding device of the first embodiment decodes the adjusted spectrum sequence ^Y0,…,^Y. N-1 The same processing is performed. That is, if an integer value greater than 1 and less than N is set as M, then, for example, when decoding the spread spectrum sequence... ~ Y0,…, ~ Y N-1 Samples with sample numbers less than M are... ~ Y0,…, ~ Y M-1 The sample group is set as the low-domain side decoding spread spectrum sequence, and the decoding spread spectrum sequence is... ~ Y0,…, ~ Y N-1 The samples with sample number M or higher are... ~ Y M ,…, ~ Y N-1 If the sample group is set as the high-domain side decoded spread spectrum sequence, then when the fricative determination information indicates a fricative tone, the adjustment and release process performed by the fricative adjustment and release unit 23 is as follows: The low-domain side decoded spread spectrum sequence is obtained. ~ Y0,…, ~ Y N-1 All or part of the samples, and the same number of high-domain side decoded spread spectrum sequences ~ Y M ,…, ~ Y N-1 The results of all or part of the samples are used as the decoded spectral sequence ^X0,…,^X N-1 .
[0228] In other words, if the information indicating whether it is an input fricative tone indicates that it is a fricative tone, the fricative tone adjustment release unit 23 obtains all or part of the low-domain frequency sample string that is located on the lower domain side compared to the specified frequency in the decoded extended spectrum sequence obtained by the bandwidth extension unit 25, and the same number of all or part of the high-domain frequency sample string that is located on the higher domain side compared to the specified frequency in the decoded extended spectrum sequence obtained by the bandwidth extension unit 25, and uses this as the spectrum sequence (decoded spectrum sequence) of the decoded tone signal. In other cases, the fricative tone adjustment release unit 23 directly uses the decoded extended spectrum sequence obtained by the bandwidth extension unit 25 as the spectrum sequence (decoded spectrum sequence) of the decoded tone signal.
[0229] Moreover, such as Figure 11As shown by the dashed line, if the frequency band extension unit 25 and the fricative tone adjustment release unit 23 are set as the fricative tone corresponding frequency band extension unit 27, then when the information indicating whether it is an input fricative tone indicates that it is a fricative tone, the fricative tone corresponding frequency band extension unit 27 extends the frequency domain spectrum sequence (decoded adjusted spectrum sequence) obtained by the decoding unit 22 to the lower domain side to obtain the spectrum sequence of the decoded tone signal (decoded spectrum sequence). In other cases, the fricative tone corresponding frequency band extension unit 27 extends the frequency domain spectrum sequence obtained by the decoding unit 22 to the higher domain side to obtain the spectrum sequence of the decoded tone signal (decoded spectrum sequence).
[0230] [Time Domain Transformation Unit 24]
[0231] The time-domain transformation unit 24, for each frame, uses a time-domain transformation method corresponding to the frequency-domain transformation method performed by the frequency-domain transformation unit 11 of the encoding apparatus to transform the decoded spectrum sequence ^X0,…,^X. N-1 The signal is transformed into a time domain signal to obtain a frame-unit audio signal (decoded audio signal) and output (step S24).
[0232] Effects and benefits
[0233] The encoding and decoding apparatus according to the second embodiment, like the encoding and decoding apparatus of the first embodiment, performs fricative tone adjustment processing and fricative tone adjustment de-processing, prioritizing the allocation of bits in the higher domain during the time interval of the fricative tone and prioritizing the allocation of bits in the lower domain during the time interval where such a time interval does not exist, thereby reducing auditory degradation even for sound signals containing fricative tones.
[0234] The encoding and decoding apparatus according to the second embodiment further utilizes a bandwidth expansion gain. In the time interval of a fricative tone, the bandwidth is expanded by replicating the high-domain spectrum and reproducing the low-domain spectrum. Conversely, in time intervals where this is not the case, the bandwidth is expanded by replicating the low-domain spectrum and reproducing the high-domain spectrum. Therefore, even with sound signals containing fricative tones, auditory degradation can be further reduced compared to the first embodiment. In this case, by using a bandwidth expansion gain based on the amplitude of the spectrum to maintain frequency order during replication, the original spectrum outline is reproduced as closely as possible, improving auditory quality.
[0235] Furthermore, if the friction sound determination unit 12 of the modified example of the first embodiment is used as the friction sound determination unit 12 of the encoding device of the second embodiment, compared with the structure of using the friction sound determination unit 12 of the first embodiment as the friction sound determination unit 12 of the encoding device of the second embodiment, it is possible to suppress the frequent switching of the determination result of the friction sound determination unit 12, suppress the frequency of the occurrence of discontinuity in the waveform of the decoded sound, and suppress the degradation of auditory quality caused by the perception of such discontinuity.
[0236] [Program and recording medium]
[0237] Alternatively, the encoding device, decoding device, and friction sound detection device can be implemented using a computer. In this case, a program describes the processing content of the functions that each encoding device, decoding device, and friction sound detection device should have. Then, by executing this program on the computer, the various encoding devices, decoding devices, and friction sound detection devices are implemented on the computer.
[0238] The program describing this processing can be recorded on a computer-readable recording medium. Such a computer-readable recording medium includes, for example, any medium such as a magnetic recording device, optical disc, optical-magnetic recording medium, semiconductor memory, etc.
[0239] Furthermore, the processing of each part can be constructed by executing a prescribed program on a computer, or at least a part of these processes can be implemented by hardware.
[0240] Furthermore, it goes without saying that appropriate modifications can be made without departing from the spirit of the invention.
Claims
1. A decoding apparatus for decoding the spectral code of a frame unit within a specified time interval to obtain a spectral sequence of a decoded audio signal, comprising: The decoding unit, when the information indicating whether it is an input fricative tone indicates that it is a fricative tone, assumes that no bits are allocated to a portion of the lower domain side of the spectral code, and decodes the spectral code to obtain a frequency domain spectral sequence; otherwise, assumes that no bits are allocated to a portion of the higher domain side of the spectral code, and decodes the spectral code to obtain a frequency domain spectral sequence; and The fricative tone corresponding frequency band extension section, when the information indicating whether it is an input fricative tone indicates that it is a fricative tone, extends the frequency domain spectrum sequence obtained by the decoding section to the lower domain side to obtain the spectrum sequence of the decoded tone signal; otherwise, it extends the frequency domain spectrum sequence obtained by the decoding section to the higher domain side to obtain the spectrum sequence of the decoded tone signal.
2. A decoding method, comprising decoding the spectral code of a frame unit within a specified time interval to obtain a spectral sequence of a decoded audio signal, including: The decoding step, in the case where the information indicating whether it is an input fricative tone indicates that it is a fricative tone, assumes that no bits are allocated to a portion of the lower domain side of the spectral code, and decodes the spectral code to obtain a frequency domain spectral sequence; otherwise, assumes that no bits are allocated to a portion of the higher domain side of the spectral code, and decodes the spectral code to obtain a frequency domain spectral sequence; and In the fricative tone corresponding frequency band extension step, if the information indicating whether it is an input fricative tone indicates that it is a fricative tone, the frequency domain spectrum sequence obtained in the decoding step is extended to the lower domain side to obtain the spectrum sequence of the decoded tone signal. Otherwise, the frequency domain spectrum sequence obtained in the decoding step is extended to the higher domain side to obtain the spectrum sequence of the decoded tone signal.
3. A computer-readable recording medium containing a program for enabling a computer to function as various parts of the decoding apparatus of claim 1.
4. A computer program product comprising a computer program for enabling a computer to function as a component of the decoding apparatus of claim 1.
Citation Information
Patent Citations
Decoding apparatus, decoding method and communication terminals and base station apparatus
CN101656074A
Encoding device, decoding device, methods therefor, program, and recording medium
CN107210042A