Decoding device, decoding method, recording medium, and computer program product
By using the shape parameter η of the generalized Gaussian distribution in the frequency domain for encoding and decoding, the problem of undetermined parameter η in the prior art is solved, the efficiency of encoding and decoding is improved, and a more accurate timing signal feature representation is achieved.
Patent Information
- Application Number
- CN202111170288.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-04-13
- Filing Date
- 2016-01-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2036-01-27
AI Technical Summary
In the prior art, the encoding and decoding processing structure based on parameter η has not been determined, resulting in low encoding and decoding efficiency, and it is difficult to effectively use parameter η to represent the characteristics of the timing signal.
The timing signal is encoded and decoded in the frequency domain. By setting the parameter η as a positive number, using the shape parameters of the generalized Gaussian distribution as an approximation of the whitening spectrum sequence, one or a variable parameter of the multiple parameters η is selected for encoding and decoding.
Appropriate encoding and decoding processing based on parameter η is realized, the encoding and decoding efficiency is improved, and the characteristics of timing signals can be represented more accurately.
Smart Images

Figure CN113921021B_ABST
Abstract
Description
[0001] This application is a divisional application of the following patent application: the application date is January 27, 2016, the application number is 201680007279.3, and the invention name is "Encoding device, encoding method and recording medium". Technical Field
[0002] The present invention relates to a technology for encoding or decoding a time series signal such as a sound signal. Background Art
[0003] As parameters that represent the characteristics of a time-series signal such as an audio signal, parameters such as LSP are known (for example, refer to Non-Patent Document 1).
[0004] Since LSPs contain multiple values, they are sometimes difficult to use directly for audio classification or segment estimation. For example, since LSPs contain multiple values, threshold-based processing using LSPs is not simple.
[0005] Furthermore, although not publicly known, the inventors have proposed a parameter η. This parameter η is a shape parameter that determines the probability distribution to which the encoding target belongs in an encoding scheme that performs arithmetic coding on quantized values of frequency-domain coefficients using a linear prediction envelope, such as that used in the 3GPP EVS (Enhanced Voice Services) standard. Parameter η is correlated with the distribution of the encoding target, and appropriately determining parameter η enables efficient encoding and decoding.
[0006] Furthermore, the parameter η can be an indicator of the characteristics of a time series signal. Therefore, although it is not publicly known, it is conceivable to determine an appropriate encoding or decoding structure based on the parameter η and perform the encoding or decoding process with the determined structure.
[0007]
Prior art literature
[0008]
Non-patent literature
[0009] [Non-patent document 1] Takehiro Moriya, "Indispensable technology for high-voltage compression sound symbolization: Linear Sequence (LSP)", NTT Technology Instruments, September 2014, P.58-60 Summary of the Invention
[0010] Problems to be solved by the invention
[0011] However, there is no known technology for determining an appropriate encoding or decoding process structure based on the parameter η and performing the encoding or decoding process with the determined structure.
[0012] The object of the present invention is to provide a structure for determining appropriate encoding or decoding processing based on parameter η, an encoding device, a decoding device, their methods, a program, and a recording medium for performing encoding or decoding processing of the determined structure.
[0013] Means for solving problems
[0014] According to one embodiment of the present invention, a coding device encodes a time series signal for each predetermined time interval in the frequency domain, wherein a parameter η is set to a positive number, and the parameter η corresponding to the time series signal is used as a shape parameter of a generalized Gaussian distribution that approximates a histogram of a whitened spectral sequence. Any one of a plurality of parameters η can be selected for each predetermined time interval, or the parameter η is variable. The above-mentioned whitened spectral sequence is a sequence obtained by dividing a frequency domain sample string by a spectral envelope estimated by treating the ηth power of the absolute value of the frequency domain sample string corresponding to the time series signal as a power spectrum. The coding device includes: a coding unit that encodes the time series signal for each predetermined time interval through a coding process having a structure determined at least based on the parameter η for each predetermined time interval.
[0015] According to one embodiment of the present invention, a coding device encodes a time series signal for each predetermined time interval in the frequency domain, wherein a parameter η is set to a positive number, and any one of a plurality of parameters η can be selected for each predetermined time interval or the parameter η is variable. The coding device includes: a coding unit that encodes a frequency domain sample string corresponding to the time series signal for each predetermined time interval by changing the bit allocation or substantially changing the bit allocation based on the value of the spectral envelope estimated by treating the ηth power of the absolute value of the frequency domain sample string corresponding to the time series signal as an estimate of the spectral envelope of the power spectrum, and outputs a parameter code representing the parameter η corresponding to the output code.
[0016] According to one aspect of the present invention, a decoding device is provided in which a parameter η is set to a positive number, and a parameter code representing the parameter η is used as a code representing a shape parameter of a generalized Gaussian distribution that approximates a histogram of a whitened spectrum sequence, wherein the whitened spectrum sequence is a sequence obtained by dividing a frequency domain sample string by a spectrum envelope estimated by treating the ηth power of the absolute value of the frequency domain sample string corresponding to the parameter η as a power spectrum. The decoding device includes: a parameter code decoding unit that decodes an input parameter code to obtain the parameter η; a determination unit that determines a structure of a decoding process based on at least the obtained parameter η; and a decoding unit that decodes the input code through a decoding process having the determined structure.
[0017] According to one aspect of the present invention, a decoding device obtains a frequency domain sample string corresponding to a time series signal by decoding in the frequency domain, wherein the decoding device includes: a parameter code decoding unit that decodes an input parameter code to obtain a parameter η; a linear prediction coefficient decoding unit that obtains coefficients that can be converted into linear prediction coefficients by decoding an input linear prediction coefficient code; a non-smoothed spectrum envelope sequence generation unit that uses the obtained parameter η to obtain a non-smoothed spectrum envelope sequence, wherein the non-smoothed spectrum envelope sequence is a sequence of amplitude spectrum envelopes corresponding to coefficients that can be converted into linear prediction coefficients raised to the power of 1 / η; and a decoding unit that decodes the input integer signal code according to a bit allocation that is changed or substantially changed based on the non-smoothed spectrum envelope sequence, thereby obtaining a frequency domain sample string corresponding to the time series signal.
[0018] Effects of the Invention
[0019] Based on the parameter η, the structure of an appropriate encoding process or decoding process can be determined, and the encoding process or decoding process of the determined structure can be performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a block diagram for explaining an example of a conventional encoding device.
[0021] Figure 2 This is a block diagram for explaining an example of a conventional encoding unit.
[0022] Figure 3 This is a diagram used to illustrate the generalized Gaussian distribution.
[0023] Figure 4 This is a block diagram for explaining an example of an encoding device.
[0024] Figure 5 This is a flowchart for explaining an example of an encoding method.
[0025] Figure 6 This is a block diagram for explaining an example of an encoding unit.
[0026] Figure 7 This is a block diagram for explaining an example of an encoding unit.
[0027] Figure 8 This is a flowchart for explaining an example of processing of the encoding unit.
[0028] Figure 9 This is a block diagram for explaining an example of a decoding device.
[0029] Figure 10 This is a flowchart for explaining an example of a decoding method.
[0030] Figure 11This is a flowchart for explaining an example of processing of the decoding unit.
[0031] Figure 12 This is a block diagram for explaining an example of an encoding device.
[0032] Figure 13 This is a flowchart for explaining an example of an encoding method.
[0033] Figure 14 This is a block diagram for explaining an example of a parameter determination unit.
[0034] Figure 15 This is a flowchart for explaining an example of a parameter determination method.
[0035] Figure 16 It is a histogram used to illustrate the technical background.
[0036] Figure 17 This is a block diagram for explaining an example of an encoding device.
[0037] Figure 18 This is a flowchart for explaining an example of an encoding method.
[0038] Figure 19 This is a block diagram for explaining an example of a decoding device.
[0039] Figure 20 This is a flowchart for explaining an example of a decoding method.
[0040] Figure 21 This is a block diagram for explaining an example of a parameter determination unit.
[0041] Figure 22 This is a flowchart for explaining an example of a parameter determination method.
[0042] Figure 23 This is a diagram used to illustrate the generalized Gaussian distribution. DETAILED DESCRIPTION [Technical Background]
[0044] Known methods for encoding low-bit audio signals (e.g., approximately 10 kbit / s to 20 kbit / s) include adaptive coding of orthogonal transform coefficients in the frequency domain, such as the DFT (Discrete Fourier Transform) and the MDCT (Modified Discrete Cosine Transform). For example, the MPEG USAC (Unified Speech and Audio Coding) standard technology includes a TCX (Transform Coded Excitation) coding mode, in which MDCT coefficients are normalized and quantized for each frame before variable-length coding (e.g., see Reference 1).
[0045] [Reference 1] M.Neuendorf, et al., "MPEG Unified Speech and Audio Coding-The ISO / MPEG Standard for High-Efficiency Audio Coding of all Content Types", AES132 nd Convention,Budapest,Hungary,2012.
[0046] Figure 1 The following describes an example of the structure of a coding device based on the existing TCX. Figure 1 's various departments.
[0047] <Frequency Domain Converter 11>
[0048] An audio signal, which is a time-series signal in the time domain, is input to the frequency domain conversion unit 11. The audio signal is, for example, a speech signal or an acoustic signal.
[0049] The frequency domain converter 11 converts the inputted time domain audio signal into an N-point MDCT coefficient sequence X(0), X(1), ..., X(N-1) in the frequency domain in units of frames of a predetermined time length, where N is a positive integer.
[0050] The converted MDCT coefficient sequence X(0), X(1), . . . , X(N−1) is output to the envelope normalization unit 15 .
[0051] <Linear prediction analysis unit 12>
[0052] The linear prediction analysis unit 12 receives an audio signal as a time-series signal in the time domain.
[0053] The linear prediction analysis unit 12 performs linear prediction analysis on the input audio signal in units of frames to generate linear prediction coefficients α1, α2, ..., α p Furthermore, the linear prediction analysis unit 12 generates linear prediction coefficients α1, α2, ..., α p Encode to generate linear prediction coefficient code. An example of linear prediction coefficient code is and linear prediction coefficients α1, α2, ..., α p The code corresponding to the quantized value string of the corresponding LSP (Line Spectrum Pairs) parameter string is the LSP code. p is an integer greater than or equal to 2.
[0054] Furthermore, the linear prediction analysis unit 12 generates quantized linear prediction coefficients ^α1, ^α2, ..., ^α corresponding to the generated linear prediction coefficient code.p .
[0055] The generated quantized linear prediction coefficients ^α1, ^α2, …, ^α p The linear prediction coefficient codes are output to the smoothed amplitude spectrum envelope sequence generator 14 and the non-smoothed amplitude spectrum envelope sequence generator 13. The generated linear prediction coefficient codes are output to the decoding device.
[0056] In linear prediction analysis, for example, a method is used in which the autocorrelation of the input audio signal is calculated frame by frame, and the Levinson-Durbin algorithm is applied using the calculated autocorrelation to obtain linear prediction coefficients. Alternatively, a method is also used in which the MDCT coefficient sequence calculated by the frequency domain conversion unit 11 is input to the linear prediction analysis unit 12, and the Levinson-Durbin algorithm is applied to the sequence of square values of each coefficient in the MDCT coefficient sequence obtained by inverse Fourier transform to obtain the linear prediction coefficients.
[0057] <Smoothed Amplitude Spectrum Envelope Sequence Generator 14>
[0058] The smoothed amplitude spectrum envelope sequence generator 14 receives the quantized linear prediction coefficients ^α1, ^α2, ..., ^α generated by the linear prediction analyzer 12. p .
[0059] The smoothed amplitude spectrum envelope sequence generator 14 uses the quantized linear prediction coefficients ^α1, ^α2, ..., ^α p , generating a smoothed amplitude spectrum envelope sequence defined by the following equation (B1): γ (0),^W γ (1),…,^W γ (N-1). Let · be a real number, exp(·) be an exponential function with the Napier number as the base, and j be an imaginary unit. γ is a positive constant less than 1 and is a coefficient that reduces the amplitude unevenness of the amplitude spectrum envelope sequence ^W(0), ^W(1), …, ^W(N-1) defined by the following equation (B2). In other words, it is a coefficient that smoothes the amplitude spectrum envelope sequence.
[0060]
Number 1
[0061]
[0062] The generated smoothed amplitude spectrum envelope sequence ^W γ (0),^W γ (1),…,^W γ (N−1) is output to the envelope normalization unit 15 and the variance parameter determination unit 163 of the encoding unit 16 .
[0063] <Non-smoothed Amplitude Spectrum Envelope Sequence Generator 13>
[0064] The non-smoothed amplitude spectrum envelope sequence generator 13 receives the quantized linear prediction coefficients ^α1, ^α2, ..., ^α generated by the linear prediction analyzer 12. p .
[0065] The non-smoothed amplitude spectrum envelope sequence generation unit 13 uses the quantized linear prediction coefficients ^α1, ^α2, ..., ^α p , generating the non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1),…, ^W(N-1) defined by the above formula (B2).
[0066] The generated non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), . . . , ^W(N−1) is output to the variance parameter determination unit 163 of the encoding unit 16 .
[0067] <Envelope Normalization Section 15>
[0068] The envelope normalization unit 15 receives as input the MDCT coefficient sequence X(0), X(1), ..., X(N-1) generated by the frequency domain conversion unit 11 and the smoothed amplitude spectrum envelope sequence output by the smoothed amplitude spectrum envelope sequence generation unit 14. γ (0),^W γ (1),…,^W γ (N-1).
[0069] The envelope normalization unit 15 normalizes the amplitude spectrum envelope sequence by using the values of the smoothed amplitude spectrum envelope sequence. γ (k) Normalize each coefficient X(k) of the MDCT coefficient string to generate a normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1). That is, X N (k) = X(k) / ^W γ (k)[k=0,1,…,N-1].
[0070] The generated normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1) is output to the encoding unit 16.
[0071] Here, in order to realize quantization with reduced distortion as perceived by the auditory sense, the envelope normalization unit 15 uses a sequence of attenuated amplitude spectrum envelopes, that is, a smoothed amplitude spectrum envelope sequence. γ (0),^W γ (1),…,^Wγ (N-1), normalize the MDCT coefficient string X(0), X(1), ..., X(N-1) in units of frames.
[0072] <Encoding unit 16>
[0073] The normalized MDCT coefficient sequence X generated by the envelope normalization unit 15 is input to the encoding unit 16. N (0),X N (1),…,X N (N-1), the smoothed amplitude spectrum envelope sequence outputted by the smoothed amplitude spectrum envelope sequence generator 14 γ (0),^W γ (1),…,^W γ (N-1), the non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), ..., ^W(N-1) output by the non-smoothed amplitude spectrum envelope sequence generation unit 13.
[0074] The encoding unit 16 generates the normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1) corresponding code.
[0075] and the normalized MDCT coefficient string X is generated N (0),X N (1),…,X N The code corresponding to (N-1) is output to the decoding device.
[0076] The normalized MDCT coefficient string X N (0),X N (1),…,X N Each coefficient of (N-1) is divided by the gain (global gain) g, and the result is a quantized integer value sequence, that is, a quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The code obtained by encoding (N-1) is defined as an integer signal code. In the technique of Non-Patent Document 1, the encoder 16 determines a gain g, which is a value as large as possible and equal to the number of bits of the integer signal code and a pre-allocated number of bits, namely, the allocated number of bits B. The encoder 16 then generates a gain code corresponding to the determined gain g and an integer signal code corresponding to the determined gain g.
[0077] The generated gain code and integer signal code are used as the normalized MDCT coefficient string X N (0),X N (1),…,X N(N-1) corresponding codes are output to the decoding device.
[0078] [Specific Example of Encoding Processing by the Encoding Unit 16]
[0079] A specific example of the encoding process performed by the encoding unit 16 will be described.
[0080] Figure 2 16 shows a specific example of the structure of the encoding unit 16. Figure 2 As shown, the coding unit 16 includes, for example, a gain acquisition unit 161, a quantization unit 162, a variance parameter determination unit 168, an arithmetic coding unit 169, a gain coding unit 165, a determination unit 166, and a gain update unit 167. Figure 2 's various departments.
[0081] <Gain Acquisition Unit 161>
[0082] The gain acquisition unit 161 obtains the normalized MDCT coefficient sequence X from the input. N (0),X N (1),…,X N (N-1) A global gain g is determined and outputted so that the number of bits of the integer signal code is equal to or less than the pre-allocated number of bits, i.e., the allocated number of bits B, and is as large as possible. The global gain g obtained by the gain acquisition unit 161 becomes the initial value of the global gain used by the quantization unit 162.
[0083] <Quantization Unit 162>
[0084] The quantization unit 162 obtains the normalized MDCT coefficient string X to be input. N (0),X N (1),…,X N The sequence of integer parts of the results of dividing each coefficient of (N-1) by the global gain g obtained by the gain acquisition unit 161 or the gain update unit 167 is the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1), and output.
[0085] Here, the global gain g used by the quantization unit 162 during the first execution is the global gain g obtained by the gain acquisition unit 161, that is, the initial value of the global gain. Furthermore, the global gain g used by the quantization unit 162 during the second and subsequent executions is the global gain g obtained by the gain update unit 167, that is, the updated value of the global gain.
[0086] <Variance Parameter Determination Unit 163>
[0087] The variance parameter determination unit 163 determines the variance parameter based on the input non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), ..., ^W(N-1) and the input smoothed amplitude spectrum envelope sequence ^W γ (0),^W γ (1),…,^W γ (N-1), the variance parameter for each frequency is obtained by the following formula (B3) And output.
[0088]
Number 2
[0089]
[0090] <Arithmetic Coding Unit 164>
[0091] The arithmetic coding unit 164 uses the variance parameter obtained by the variance parameter determination unit 163 The quantized normalized coefficient sequence X obtained by the quantization unit 162 is Q (0),X Q (1),…,X Q (N-1) performs arithmetic coding to obtain an integer signal code, and outputs the integer signal code and the number of consumed bits C, which is the number of bits of the integer signal code. This arithmetic coding performs an optimal bit allocation when the quantized and normalized coefficient sequence at each frequency k (=0, ..., N-1) follows a Laplace distribution related to the following probability variable X, for example, as shown by the following equation.
[0092]
Number 3
[0093]
[0094] <Determination Unit 166>
[0095] When the number of times the gain is updated is a predetermined number, the determination unit 166 outputs an integer signal code and outputs an instruction signal to the gain coding unit 165 for encoding the global gain g obtained by the gain updating unit 167. When the number of times the gain is updated is less than the predetermined number, the determination unit 166 outputs the number of consumed bits C measured by the arithmetic coding unit 164 to the gain updating unit 167.
[0096] <Gain Update Unit 167>
[0097] When the number of consumed bits C measured by the arithmetic coding unit 164 is greater than the number of allocated bits B, the gain update unit 167 updates the value of the global gain g to a larger value and outputs it. When the number of consumed bits C is less than the number of allocated bits B, the gain update unit 167 updates the value of the global gain g to a smaller value and outputs the updated value of the global gain g.
[0098] <Gain Encoding Unit 165>
[0099] The gain coding unit 165 encodes the global gain g obtained by the gain updating unit 167 in accordance with the instruction signal output by the determination unit 166 to obtain a gain code and outputs the code.
[0100] The integer signal code output by the determination unit 166 and the gain code output by the gain encoding unit 165 are output to the decoding apparatus as codes corresponding to the normalized MDCT coefficient sequence.
[0101] As described above, in conventional TCX coding, the MDCT coefficient sequence is normalized using a smoothed amplitude spectrum envelope sequence that weakens the non-smoothed amplitude spectrum envelope, and then the normalized MDCT coefficient sequence is encoded. This coding method is adopted in the aforementioned MPEG-4USAC and other standards.
[0102] Conventional encoding devices use arithmetic coding to optimize bit allocation for the Laplace distribution. Furthermore, because arithmetic coding utilizes information about the concavity and convexity of the spectral envelope, a variance parameter corresponding to the variance of the aforementioned Laplace distribution is generated based on the envelope value. However, the probability distributions to which encoding targets belong vary, and they do not uniformly conform to the Laplace distribution. Therefore, applying the same bit allocation to encoding targets belonging to distributions excluded from the hypothesis could reduce compression efficiency. Furthermore, when introducing other distributions, improving efficiency is difficult unless variance parameters for those distributions are also generated and information about the concavity and convexity of the spectral envelope is accurately incorporated, as in conventional encoding devices.
[0103] In addition, compared to normalization based on a non-smoothed amplitude spectrum envelope sequence, normalization based on a smoothed amplitude spectrum envelope does not whiten the MDCT sequence X(0), X(1), ..., X(N-1). Specifically, compared to normalization based on a non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), ..., ^W(N-1) to obtain a normalized sequence X(0) / ^W(0), X(1) / ^W(1), ..., X(N-1) / ^W(N-1), normalization based on a smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), ..., ^W(N-1) to obtain a normalized sequence X(0), X(1), ..., X(N-1) / ^W(N-1). γ (0),^W γ (1),…,^W γ (N-1) is normalized to obtain the normalized MDCT coefficient string X N (0) = X(0) / ^W γ (0),X N (1) = X(1) / ^W γ(1),…,X N (N-1)=X(N-1) / ^W γ (N-1), only ^W(0) / ^W γ (0),^W(1) / ^W γ (1),…,^W(N-1) / ^W γ Therefore, if it is assumed that the MDCT coefficient sequence X(0), X(1), ..., X(N-1) is normalized by the non-smoothed amplitude spectrum envelope sequence ^W(0), ^W(1), ..., ^W(N-1), and the convexity and concavity of the envelope of the normalized sequence X(0) / ^W(0), X(1) / ^W(1), ..., X(N-1) / ^W(N-1) is flattened to a degree suitable for encoding in the encoding unit 16, then the normalized MDCT coefficient sequence X(0), X(1), ..., X(N-1) input to the encoding unit 16 is flattened. N (0),X N (1),…,X N (N-1), leaving ^W(0) / ^W γ (0),^W(1) / ^W γ (1),…,^W(N-1) / ^W γ (N-1) sequence (hereinafter, normalized amplitude spectrum envelope sequence) N (0),^W N (1),…,^W N (N-1)) represents the concavity and convexity of the envelope.
[0104] Figure 16 Indicates the concavity and convexity of the envelope of the normalized MDCT sequence ^W(0) / ^W γ (0),^W(1) / ^W γ (1),…,^W(N-1) / ^W γ The frequency of occurrence of the values of each coefficient contained in the normalized MDCT coefficient sequence when (N-1) takes each value. The curve of envelope: 0.2-0.3 represents the concavity and convexity of the envelope of the normalized MDCT sequence ^W(k) / ^W γ (k) is greater than 0.2 and less than 0.3, and corresponds to the normalized MDCT coefficient X of sample k. N The frequency of the value of (k). Envelope: The curve of 0.3-0.4 represents the concavity and convexity of the envelope of the normalized MDCT sequence ^W(k) / ^W γ (k) is greater than 0.3 and less than 0.4, and corresponds to the normalized MDCT coefficient X of sample k. N The frequency of the value of (k). Envelope: The curve of 0.4-0.5 represents the concavity and convexity of the envelope of the normalized MDCT sequence ^W(k) / ^W γ(k) is greater than 0.4 and less than 0.5, and corresponds to the normalized MDCT coefficient X of sample k. N The frequency of the value of (k).
[0105] look Figure 16 As can be seen, the average value of each coefficient included in the normalized MDCT coefficient sequence is approximately zero, but the variance is correlated with the envelope value. Specifically, the greater the concavity of the envelope of the normalized MDCT sequence, the wider the foot of the curve representing frequency. Therefore, it can be seen that the variance of the normalized MDCT coefficients is correlated to be large. To achieve more efficient compression, encoding is performed that exploits this correlation. Specifically, encoding is performed that modifies or substantially modifies the bit allocation of each coefficient in the frequency domain coefficient sequence to be encoded based on the spectral envelope.
[0106] For this purpose, for example, after the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) When arithmetic coding is performed, a variance parameter determined based on the spectrum envelope is used.
[0107] In addition, when there is diversity in the probability distribution to which the coding object belongs, if the optimal bit allocation for the coding object assuming that it belongs to a certain probability distribution (for example, Laplace distribution) is performed on the coding object belonging to a probability distribution excluded from the assumption, the compression efficiency may decrease.
[0108] Therefore, as the probability distribution to which the encoding target belongs, a distribution capable of expressing various probability distributions, that is, a generalized Gaussian distribution represented by the following equation is used.
[0109]
Number 4
[0110]
[0111] The generalized Gaussian distribution is obtained by changing the parameter η (>0) as the shape parameter, such as Figure 3 As shown in FIG, when η = 1, it is a Laplace distribution, and when η = 2, it is a Gaussian distribution. Thus, various distributions can be expressed. η is a predetermined number greater than 0. The value of η can be predetermined, or selected or changed in each frame as a predetermined time interval. In addition, the above formula The variance parameter is the value corresponding to the distribution, and the information of the concavity and convexity of the spectrum envelope is incorporated into the variance parameter. The quantized normalized coefficient X at each frequency k is Q (k) constituted in compliance with It becomes the optimal arithmetic code when , and encoding is performed using arithmetic codes based on this structure.
[0112] For example, we further introduce the energy σ in addition to the prediction residual 2 In addition to the information of the global gain g, the information of the distribution is also used, for example, by the following formula (A1) to calculate the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The variance parameters of the coefficients of (N-1).
[0113]
Number 5
[0114]
[0115] Where σ is σ 2 The square root of .
[0116] Specifically, the Levinson-Durbin algorithm is performed on a sequence of values obtained by inverse Fourier transforming the absolute values of the MDCT coefficients raised to the power of n, and the quantized linear prediction coefficients ^α1, ^α2, ..., ^α are replaced. p The linear prediction coefficients thus obtained are quantized using β1, ^β2, …, ^β p , the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1),…, ^H(N-1) and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1) is respectively obtained by the following formula (A2) and formula (A3):
[0117]
Number 6
[0118]
[0119] And find out, and divide each coefficient of the obtained non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) by the corresponding smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ The normalized amplitude spectrum envelope sequence ^H is obtained by using the coefficients of (N-1) N (0) = ^H(0) / ^H γ (0),^H N (1) = ^H(1) / ^H γ (1),…,^H N (N-1)=^H(N-1) / ^H γ (N-1), the variance parameter is calculated according to the above formula (A1) based on the normalized amplitude spectrum envelope sequence and the global gain g.
[0120] Here, σ in formula (A1) 2 / η / g is a value closely related to entropy, and if the bit rate is constant, the value of each frame will change little. Therefore, as σ 2 / η When a fixed value is used, there is no need to add new information for the method of the present invention.
[0121] The above technology is based on the quantization and normalization of the coefficient sequence X Q (0),X Q (1),…,X Q (N-1) is a technique for minimizing the code length when performing arithmetic coding. The following describes the derivation of the above technique.
[0122] If the quantization is done sufficiently carefully, the quantization normalization coefficient X Q (k) respectively through the variance parameter The code length when encoding is performed using arithmetic codes with a generalized Gaussian distribution with a shape parameter η is
[0123]
Number 7
[0124]
[0125] In order to reduce the code length, it is considered to obtain the variance parameter sequence based on the linear prediction coefficients that have been quantized and encoded. The above formula (A4) can be rewritten as
[0126]
Number 8
[0127]
[0128] where ln is the logarithm with Napier's number as base, C is a constant for the variance parameter, and D IS (X|Y) is the Itakura-Saito distance from X to Y
[0129]
Number 9
[0130]
[0131] That is, the problem of minimizing the code length L of the variance parameter sequence is reduced to and |X Q (k)| η Here, if the variance parameter sequence and linear prediction coefficients β1,β2,…,β p , the energy of the prediction residual σ 2If the corresponding relationship determines one, it is possible to establish an optimization problem for finding the linear prediction coefficient that minimizes the code length, but in order to use the existing high-speed solution, the correspondence is established here as follows.
[0132]
Number 10
[0133]
[0134] If the influence of quantization is ignored, the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) can use the MDCT sequence X(0), X(1), ..., X(N-1) and the smoothed amplitude spectrum envelope ^H γ (0),^H γ (1),…,^H γ (N-1), global gain g and are expressed as X Q (k) = X(k) / (g^H γ (k)), so the term that depends on the variance parameter of formula (A5) is expressed by formula (A6), as
[0135]
Number 11
[0136]
[0137] As shown, the Itakura-Saito distance between the absolute value of the MDCT coefficient sequence and the fully polar spectral envelope is expressed. Conventional linear prediction analysis, namely, applying the Levinson-Durbin algorithm to a sequence obtained by inverse Fourier transforming the power spectrum, is known to find linear prediction coefficients that minimize the Itakura-Saito distance between the power spectrum and the fully polar spectral envelope. Therefore, the aforementioned code length minimization problem can be optimally solved by applying the Levinson-Durbin algorithm to a sequence obtained by inverse Fourier transforming the nth power of the amplitude spectrum, i.e., the nth power of the absolute value of the MDCT coefficient sequence, in the same manner as conventional methods.
[0138] [First embodiment]
[0139] (coding)
[0140] Figure 4 FIG. 1 shows a configuration example of an encoding device according to the first embodiment. Figure 4 As shown, the encoding device of the first embodiment includes, for example, a frequency domain conversion unit 21, a linear prediction analysis unit 22, a non-smoothed amplitude spectrum envelope sequence generation unit 23, a smoothed amplitude spectrum envelope sequence generation unit 24, an envelope normalization unit 25, an encoding unit 26 and a parameter determination unit 27. Figure 5 An example of each process of the encoding method according to the first embodiment implemented by the encoding device is shown.
[0141] The following describes Figure 4 's various departments.
[0142] <Parameter determination unit 27>
[0143] In the first embodiment, the parameter determination unit 27 can select any one of the plurality of parameters η for each predetermined time interval.
[0144] Assume that a plurality of parameters η are stored as candidate parameters η in the parameter determination unit 27. The parameter determination unit 27 sequentially reads one parameter η from the plurality of parameters and outputs it to the linear prediction analysis unit 22, the unsmoothed amplitude spectrum envelope sequence generation unit 23, and the encoding unit 26 (step A0).
[0145] The frequency domain conversion unit 21, linear prediction analysis unit 22, unsmoothed amplitude spectrum envelope sequence generation unit 23, smoothed amplitude spectrum envelope sequence generation unit 24, envelope normalization unit 25, and encoding unit 26 perform the processing of, for example, steps A1 to A6 described below, based on the parameters η sequentially read by the parameter determination unit 27, to generate a code for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval. Generally, given a given parameter η, two or more codes may be obtained for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval. In this case, the code for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval is a code that combines these two or more codes. In this example, the code combines the linear prediction coefficient code, the gain code, and the integer signal code. Thus, a code for each parameter η is obtained for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval.
[0146] After processing in step A6, the parameter determination unit 27 selects a code from the codes obtained for each parameter n for the frequency domain sample sequence corresponding to the time series signal in the same predetermined time interval, and determines the parameter n corresponding to the selected code (step A7). The determined parameter n becomes the parameter n for the frequency domain sample sequence corresponding to the time series signal in the same predetermined time interval. The parameter determination unit 27 then outputs a code representing the selected code and the determined parameter n to the decoding device. The details of the processing in step A7 performed by the parameter determination unit 27 will be described later.
[0147] In the following, it is assumed that the parameter determination unit 27 reads one parameter η and processes the read parameter η.
[0148] <Frequency Domain Converter 21>
[0149] An audio signal, which is a time-series signal in the time domain, is input to the frequency domain conversion unit 21. Examples of the audio signal are a digital voice signal or a digital acoustic signal.
[0150] The frequency domain converter 21 converts the inputted time domain audio signal into an N-point MDCT coefficient sequence X(0), X(1), ..., X(N-1) in the frequency domain in units of frames of a predetermined time length (step A1). N is a positive integer.
[0151] The obtained MDCT coefficient strings X(0), X(1), . . . , X(N−1) are output to the linear prediction analysis unit 22 and the envelope normalization unit 25 .
[0152] Unless otherwise specified, subsequent processing is performed in units of frames.
[0153] In this way, the frequency domain converter 21 obtains a frequency domain sample sequence, which is, for example, an MDCT coefficient sequence corresponding to the audio signal.
[0154] <Linear prediction analysis unit 22>
[0155] The MDCT coefficient sequence X(0), X(1), . . . , X(N-1) obtained by the frequency domain conversion unit 21 is input to the linear prediction analysis unit 22.
[0156] The linear prediction analysis unit 22 uses the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to perform linear prediction analysis on ~R(0), ~R(1), ..., ~R(N-1) defined by the following equation (A7) to generate linear prediction coefficients β1, β2, ..., β p , and generate linear prediction coefficients β1,β2,…,β p The linear prediction coefficient code and the quantized linear prediction coefficients corresponding to the linear prediction coefficient code, that is, the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p (Step A2).
[0157]
Number 12
[0158]
[0159] The generated quantized linear prediction coefficients ^β1,^β2,…,^β p It is output to the non-smoothed spectrum envelope sequence generator 23 and the smoothed amplitude spectrum envelope sequence generator 24. In addition, the energy σ of the prediction residual is calculated during the linear prediction analysis process. 2 At this time, the energy of the calculated prediction residual σ 2 The variance parameter is output to the variance parameter determination unit 268 of the encoding unit 26 .
[0160] In addition, the generated linear prediction coefficient code is sent to the parameter determination unit 27.
[0161] Specifically, the linear prediction analysis unit 22 first performs an operation equivalent to the inverse Fourier transform, which treats the nth power of the absolute value of the MDCT coefficient string X(0), X(1), ..., X(N-1) as a power spectrum, i.e., the operation of formula (A7), thereby obtaining a time domain signal string corresponding to the nth power of the absolute value of the MDCT coefficient string X(0), X(1), ..., X(N-1), i.e., a pseudo-correlation function signal string ~R(0), ~R(1), ..., ~R(N-1). Furthermore, the linear prediction analysis unit 22 uses the obtained pseudo-correlation function signal string ~R(0), ~R(1), ..., ~R(N-1) to perform linear prediction analysis and generate linear prediction coefficients β1, β2, ..., β p Then, the linear prediction analysis unit 22 generates the linear prediction coefficients β1, β2, ..., β p Encode to obtain the linear prediction coefficient code and the quantized linear prediction coefficients ^β1, ^β2, ..., ^β corresponding to the linear prediction coefficient code p .
[0162] Linear prediction coefficients β1,β2,…,β p It is a linear prediction coefficient corresponding to the time domain signal when the nth power of the absolute value of the MDCT coefficient string X(0), X(1), ..., X(N-1) is regarded as the power spectrum.
[0163] The linear prediction analysis unit 22 generates the linear prediction coefficient code using, for example, existing coding techniques. Examples of existing coding techniques include coding techniques that use the code corresponding to the linear prediction coefficient itself as the linear prediction coefficient code, coding techniques that convert the linear prediction coefficient into LSP parameters and use the code corresponding to the LSP parameters as the linear prediction coefficient code, and coding techniques that convert the linear prediction coefficient into PARCOR coefficients and use the code corresponding to the PARCOR coefficients as the linear prediction coefficient code. For example, coding techniques that use the code corresponding to the linear prediction coefficient itself as the linear prediction coefficient code include techniques that predetermine multiple quantized linear prediction coefficient candidates, store each candidate in association with a linear prediction coefficient code, and determine any one of the candidates as the quantized linear prediction coefficient for the generated linear prediction coefficient, thereby obtaining the quantized linear prediction coefficient and the linear prediction coefficient code. For example, a coding technology that sets the code corresponding to the linear prediction coefficient itself as the linear prediction coefficient code is the following technology: a plurality of quantized linear prediction coefficient candidates are determined in advance, each candidate is stored in advance corresponding to the linear prediction coefficient code, and any one of the candidates is determined as the quantized linear prediction coefficient for the generated linear prediction coefficient, thereby obtaining the quantized linear prediction coefficient and the linear prediction coefficient code.
[0164] In this way, the linear prediction analysis unit 22 performs linear prediction analysis using, for example, a pseudo-correlation function signal string obtained by performing inverse Fourier transform of the absolute value of the frequency domain sample string as the MDCT coefficient string to the power of η as the power spectrum, to generate coefficients that can be converted into linear prediction coefficients.
[0165] <Non-smoothed Amplitude Spectrum Envelope Sequence Generator 23>
[0166] The non-smoothed amplitude spectrum envelope sequence generation unit 23 receives the quantized linear prediction coefficients ^β1, ^β2, ..., ^β generated by the linear prediction analysis unit 22. p .
[0167] The non-smoothed amplitude spectrum envelope sequence generation unit 23 generates and quantizes linear prediction coefficients ^β1, ^β2, ..., ^β p The corresponding sequence of amplitude spectrum envelopes is the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) (step A3).
[0168] The generated non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), . . . , ^H(N−1) is output to the encoding unit 26 .
[0169] The non-smoothed amplitude spectrum envelope sequence generation unit 23 uses the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p , as the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1), the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1) defined by formula (A2) is generated.
[0170]
Number 13
[0171]
[0172] In this manner, the unsmoothed amplitude spectrum envelope sequence generator 23 estimates the spectrum envelope by obtaining a sequence of amplitude spectrum envelopes corresponding to the coefficients generated by the linear prediction analysis unit 22 and convertible into linear prediction coefficients, i.e., an unsmoothed spectrum envelope sequence, which is a sequence obtained by raising the sequence of amplitude spectrum envelopes to the power of 1 / n. Here, c is an arbitrary number, and a sequence consisting of a plurality of values raised to the power of c is a sequence consisting of values obtained by raising each of the plurality of values to the power of c. For example, a sequence obtained by raising the sequence of the amplitude spectrum envelope to the power of 1 / n is a sequence consisting of values obtained by raising each coefficient of the amplitude spectrum envelope to the power of 1 / n.
[0173] The 1 / n-th power processing performed by the non-smoothed amplitude spectrum envelope sequence generator 23 is caused by the processing of treating the n-th power of the absolute value of the frequency domain sample sequence as a power spectrum performed by the linear prediction analysis unit 22. In other words, the 1 / n-th power processing performed by the non-smoothed amplitude spectrum envelope sequence generator 23 is performed to return the value raised to the n-th power by the processing of treating the n-th power of the absolute value of the frequency domain sample sequence as a power spectrum performed by the linear prediction analysis unit 22 to the original value.
[0174] <Smoothed Amplitude Spectrum Envelope Sequence Generator 24>
[0175] The smoothed amplitude spectrum envelope sequence generator 24 receives the quantized linear prediction coefficients ^β1, ^β2, ..., ^β generated by the linear prediction analyzer 22. p .
[0176] The smoothed amplitude spectrum envelope sequence generation unit 24 generates the attenuated and quantized linear prediction coefficients ^β1, ^β2, ..., ^β p The corresponding sequence of amplitude convex and concave of the amplitude spectrum envelope sequence is the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1)(Step A4).
[0177] The generated smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N−1) is output to the envelope normalization unit 25 and the encoding unit 26 .
[0178] The smoothed amplitude spectrum envelope sequence generator 24 uses the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p and correction coefficient γ, as the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1), generate the smoothed amplitude spectrum envelope sequence defined by equation (A3) ^H γ (0),^H γ (1),…,^H γ (N-1).
[0179]
Number 14
[0180]
[0181] Here, the correction coefficient γ is a predetermined constant less than 1, and is a coefficient that reduces the unevenness of the amplitude of the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1). In other words, it is a coefficient that smoothes the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1).
[0182] <Envelope Normalization Section 25>
[0183] The envelope normalization unit 25 receives as input the MDCT coefficient sequence X(0), X(1), ..., X(N-1) obtained by the frequency domain conversion unit 21 and the smoothed amplitude spectrum envelope sequence ^H generated by the smoothed amplitude spectrum envelope generation unit 24. γ (0),^H γ (1),…,^H γ (N-1).
[0184] The envelope normalization unit 25 normalizes the amplitude spectrum of the corresponding smoothed amplitude spectrum envelope sequence by using the corresponding smoothed amplitude spectrum envelope sequence. γ (0),^H γ (1),…,^H γ The values of (N-1) are used to normalize the coefficients of the MDCT coefficient string X(0), X(1), ..., X(N-1), thereby generating a normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1)(Step A5).
[0185] The generated normalized MDCT coefficient string is output to the encoding unit 26 .
[0186] The envelope normalization unit 25 sets k=0, 1, ..., N-1, for example, and divides each coefficient X(k) of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) by the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1) values, thereby generating a normalized MDCT coefficient string X N (0),X N (1),…,X N The coefficients X of (N-1) N (k). That is, let k = 0, 1, ..., N-1, X N (k) = X(k) / ^H γ (k).
[0187] <Encoding Unit 26>
[0188] The normalized MDCT coefficient sequence X generated by the envelope normalization unit 25 is input to the encoding unit 26. N (0),X N (1),…,X N (N-1), the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the unsmoothed amplitude spectrum envelope generation unit 23, and the smoothed amplitude spectrum envelope sequence ^H generated by the smoothed amplitude spectrum envelope generation unit 24. γ (0),^H γ (1),…,^H γ (N-1) and the energy σ of the average residual calculated by the linear prediction analysis unit 22 2 .
[0189] The encoding unit 26 performs, for example, Figure 8 The processing of steps A61 to A65 is performed to perform encoding (step A6).
[0190] The encoding unit 26 obtains the normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1) corresponding global gain g (step A61), find the normalized MDCT coefficient string X N (0),X N (1),…,X N The result of dividing each coefficient of (N-1) by the global gain g is a quantized integer value sequence, that is, the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) (step A62), according to the global gain g and the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1) and the energy of the average residual σ 2 , and the quantized normalized coefficient sequence X is obtained by formula (A1) Q (0),X Q (1),…,X Q The variance parameters corresponding to the coefficients of (N-1) (Step A63), using the variance parameter The quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q(N-1) performs arithmetic coding to obtain an integer signal code (step A64), and obtains a gain code corresponding to the global gain g (step A65).
[0191]
Number 15
[0192]
[0193] Here, the normalized amplitude spectrum envelope sequence in the above formula (A1) is N (0),^H N (1),…,^H N The value of the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) is divided by the corresponding smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ The values obtained from the respective values of (N-1) are values obtained by the following formula (A8).
[0194]
Number 16
[0195]
[0196] The generated integer signal code and gain code are output to the parameter determination unit 27 as codes corresponding to the normalized MDCT coefficient sequence.
[0197] Through steps A61 to A65, the encoding unit 26 implements the following functions: determining the number of bits of the integer signal code to be a pre-allocated number of bits, that is, a global gain g that is less than the allocated number of bits B and as large as possible, and generating a gain code corresponding to the determined global gain g and an integer signal code corresponding to the determined global gain g.
[0198] Among the steps A61 to A65 performed by the encoding unit 26, the characteristic process is step A63, which is to normalize the coefficient sequence X by the global gain g and the quantization. Q (0),X Q (1),…,X Q The encoding process itself of encoding each of (N-1) to obtain a code corresponding to the normalized MDCT coefficient string includes various known techniques including the technique described in Non-Patent Document 1. Two specific examples of the encoding process performed by the encoding unit 26 are described below.
[0199] [Specific Example 1 of Encoding Processing by the Encoding Unit 26]
[0200] As a specific example 1 of the encoding process performed by the encoding unit 26, an example in which no loop process is included will be described.
[0201] Figure 6 The configuration example of the encoding unit 26 of the specific example 1 is shown in FIG. Figure 6 As shown, the encoding unit 26 of the specific example 1 includes, for example, a gain acquisition unit 261, a quantization unit 262, a variance parameter determination unit 268, an arithmetic encoding unit 269, and a gain encoding unit 265. Figure 6 's various departments.
[0202] <Gain Acquisition Unit 261>
[0203] The gain acquisition unit 261 receives the normalized MDCT coefficient sequence X generated by the envelope normalization unit 25. N (0),X N (1),…,X N (N-1).
[0204] The gain acquisition unit 261 obtains the normalized MDCT coefficient sequence X N (0),X N (1),…,X N In (N-1), the global gain g is determined so that the number of bits of the integer signal code is less than the pre-allocated number of bits, that is, the allocated number of bits B and is as large as possible, and is output (step S261). The gain acquisition unit 261, for example, takes the normalized MDCT coefficient string X N (0),X N (1),…,X N The square root of the total energy of (N-1) and the multiplication value of a constant having a negative correlation with the number of allocated bits B are obtained as the global gain g and output. Alternatively, the gain acquisition unit 261 may obtain the normalized MDCT coefficient sequence X N (0),X N (1),…,X N The relationship between the total energy of (N-1), the number of allocated bits B, and the global gain g is tabulated in advance, and the global gain g is obtained and outputted by referring to the table.
[0205] In this way, the gain acquisition unit 261 obtains a gain for dividing all samples of the normalized MDCT coefficient sequence, that is, the normalized frequency domain sample sequence, for example.
[0206] The obtained global gain g is output to the quantization unit 262 and the variance parameter determination unit 268 .
[0207] <Quantization Unit 262>
[0208] The normalized MDCT coefficient sequence X generated by the envelope normalization unit 25 is input to the quantization unit 262. N (0),X N (1),…,X N (N−1) and the global gain g obtained by the gain acquisition unit 261 .
[0209] The quantization unit 262 obtains the normalized MDCT coefficient string X N (0),X N (1),…,X N The sequence of the integer parts of the results of dividing each coefficient of (N-1) by the global gain g is the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1), and output (step S262).
[0210] In this manner, the quantization unit 262 divides each sample of the normalized MDCT coefficient sequence, that is, the normalized frequency domain sample sequence, by a gain, and quantizes the sample to obtain a quantized normalized coefficient sequence.
[0211] The obtained quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) is output to the arithmetic coding unit 269.
[0212] <Variance Parameter Determination Unit 268>
[0213] The variance parameter determination unit 268 receives as input the parameter η read by the parameter determination unit 27, the global gain g obtained by the gain acquisition unit 261, the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the unsmoothed amplitude spectrum envelope generation unit 23, and the smoothed amplitude spectrum envelope sequence ^H generated by the smoothed amplitude spectrum envelope generation unit 24. γ (0),^H γ (1),…,^H γ (N-1) and the energy σ of the prediction residual obtained by the linear prediction analysis unit 22 2 .
[0214] The variance parameter determination unit 268 determines the variance parameter based on the global gain g, the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1), and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1) and the energy of the prediction residual σ 2 , through the above formula (A1) and formula (A8), we get the variance parameter sequence The variance parameters are calculated and output (step S268).
[0215] The obtained variance parameter sequence It is output to the arithmetic coding unit 269.
[0216] <Arithmetic Coding Unit 269>
[0217] The arithmetic coding unit 269 receives as input the parameter η read by the parameter determination unit 27 and the quantized normalized coefficient sequence X obtained by the quantization unit 262. Q (0),X Q (1),…,X Q (N-1) and the variance parameter sequence obtained by the variance parameter determination unit 268
[0218] The arithmetic coding unit 269 uses the variance parameter sequence The variance parameters are used as the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The variance parameters corresponding to the coefficients of (N-1) are used to normalize the quantized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) performs arithmetic coding to obtain an integer signal code, and outputs it (step S269).
[0219] The arithmetic coding unit 269 performs arithmetic coding on the quantized and normalized coefficient sequence X. Q (0),X Q (1),…,X Q The coefficients of (N-1) follow the generalized Gaussian distribution When the bit allocation is optimal, encoding is performed by arithmetic coding based on the bit allocation performed.
[0220] The obtained integer signal code is output to the parameter determination unit 27 .
[0221] It is also possible to cross the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q At this time, from formula (A1) and formula (A8), it can be seen that the variance parameter sequence The various variance parameters are based on the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1),…, ^H(N-1), so it can be said that the arithmetic coding unit 269 performs coding in which the bit allocation is substantially changed based on the estimated spectrum envelope (non-smoothed amplitude spectrum envelope).
[0222] <Gain Encoding Unit 265>
[0223] The global gain g obtained by the gain acquisition unit 261 is input to the gain encoding unit 265 .
[0224] The gain coding unit 265 codes the global gain g to obtain a gain code, and outputs the obtained gain code (step S265 ).
[0225] The generated integer signal code and gain code are output to the parameter determination unit 27 as codes corresponding to the normalized MDCT coefficient sequence.
[0226] Steps S261, S262, S268, S269, and S265 of this specific example 1 correspond to the above-mentioned steps A61, A62, A63, A64, and A65, respectively.
[0227] [Specific Example 2 of the Encoding Processing by the Encoding Unit 26]
[0228] As a specific example 2 of the encoding process performed by the encoding unit 26, an example including a loop process will be described.
[0229] Figure 7 The configuration example of the encoding unit 26 of the specific example 2 is shown in FIG. Figure 7 As shown, the encoding unit 26 of the specific example 2 includes, for example, a gain acquisition unit 261, a quantization unit 262, a variance parameter determination unit 268, an arithmetic encoding unit 269, a gain encoding unit 265, a determination unit 266, and a gain update unit 267. Figure 7 's various departments.
[0230] <Gain Acquisition Unit 261>
[0231] The gain acquisition unit 261 receives the normalized MDCT coefficient sequence X generated by the envelope normalization unit 25. N (0),X N (1),…,X N (N-1).
[0232] The gain acquisition unit 261 obtains the normalized MDCT coefficient sequence X N (0),X N (1),…,X N In (N-1), the global gain g is determined so that the number of bits of the integer signal code is less than the pre-allocated number of bits, that is, the allocated number of bits B and is as large as possible, and is output (step S261). The gain acquisition unit 261, for example, takes the normalized MDCT coefficient string X N (0),X N (1),…,X N A multiplication value of the square root of the total energy of (N-1) and a constant having a negative correlation with the number of allocated bits B is obtained as a global gain g and outputted.
[0233] The obtained global gain g is output to the quantization unit 262 and the variance parameter determination unit 268 .
[0234] The global gain g obtained by the gain acquisition unit 261 becomes an initial value of the global gain used in the quantization unit 262 and the variance parameter determination unit 268 .
[0235] <Quantization Unit 262>
[0236] The normalized MDCT coefficient sequence X generated by the envelope normalization unit 25 is input to the quantization unit 262. N (0),X N (1),…,X N (N−1) and the global gain g obtained by the gain acquisition unit 261 or the gain update unit 267 .
[0237] The quantization unit 262 obtains the normalized MDCT coefficient string X N (0),X N (1),…,X N The sequence of the integer parts of the results of dividing each coefficient of (N-1) by the global gain g is the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1), and output (step S262).
[0238] Here, the global gain g used by the quantization unit 262 during the first execution is the global gain g obtained by the gain acquisition unit 261, that is, the initial value of the global gain. Furthermore, the global gain g used by the quantization unit 262 during the second and subsequent executions is the global gain g obtained by the gain update unit 267, that is, the updated value of the global gain.
[0239] The obtained quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) is output to the arithmetic coding unit 269.
[0240] <Variance Parameter Determination Unit 268>
[0241] The variance parameter determination unit 268 receives as input the parameter η read by the parameter determination unit 27, the global gain g obtained by the gain acquisition unit 261 or the gain update unit 267, the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the unsmoothed amplitude spectrum envelope generation unit 23, and the smoothed amplitude spectrum envelope sequence ^H generated by the smoothed amplitude spectrum envelope generation unit 24. γ (0),^H γ (1),…,^H γ (N-1) and the energy σ of the prediction residual obtained by the linear prediction analysis unit 22 2 .
[0242] The variance parameter determination unit 268 determines the variance parameter based on the global gain g, the unsmoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1), and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1) and the energy of the prediction residual σ 2 , through the above formula (A1) and formula (A8), we get the variance parameter sequence The variance parameters are calculated and output (step S268).
[0243] Here, the global gain g used by the variance parameter determination unit 268 during the first execution is the global gain g obtained by the gain acquisition unit 261, that is, the initial value of the global gain. Furthermore, the global gain g used by the variance parameter determination unit 268 during the second and subsequent executions is the global gain g obtained by the gain update unit 267, that is, the updated value of the global gain.
[0244] The obtained variance parameter sequence It is output to the arithmetic coding unit 269.
[0245] <Arithmetic Coding Unit 269>
[0246] The arithmetic coding unit 269 receives as input the parameter η read by the parameter determination unit 27 and the quantized normalized coefficient sequence X obtained by the quantization unit 262. Q (0),X Q (1),…,X Q (N-1) and the variance parameter sequence obtained by the variance parameter determination unit 268
[0247] The arithmetic coding unit 269 uses the variance parameter sequence The variance parameters are used as the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The variance parameters corresponding to the coefficients of (N-1) are used to normalize the quantized coefficient sequence X Q (0),X Q (1),…,X Q (N-1) performs arithmetic coding to obtain an integer signal code and the number of consumed bits C, which is the number of bits of the integer signal code, and outputs them (step S269).
[0248] The arithmetic coding unit 269 forms a quantized and normalized coefficient sequence X during arithmetic coding. Q (0),X Q (1),…,X QThe coefficients of (N-1) follow the generalized Gaussian distribution The optimal arithmetic code is used for encoding based on the arithmetic code structure. As a result, the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The expected value of the bit allocation of each coefficient of (N-1) is obtained by the variance parameter sequence and was decided.
[0249] The obtained integer signal code and the number of consumed bits C are output to the determination unit 266 .
[0250] It is also possible to cross the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q At this time, from formula (A1) and formula (A8), it can be seen that due to the variance parameter sequence The various variance parameters are based on the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1),…, ^H(N-1), so it can be said that the arithmetic coding unit 269 performs coding in which the bit allocation is substantially changed based on the estimated spectrum envelope (non-smoothed amplitude spectrum envelope).
[0251] <Determination Unit 266>
[0252] The integer signal code obtained by the arithmetic coding unit 269 is input to the determination unit 266 .
[0253] When the number of times the gain is updated is a predetermined number, the determination unit 266 outputs an integer signal code and outputs an instruction signal to the gain coding unit 265 for encoding the global gain g obtained by the gain updating unit 267. When the number of times the gain is updated is less than the predetermined number, the number of consumed bits C measured by the arithmetic coding unit 264 is output to the gain updating unit 267 (step S266).
[0254] <Gain Update Unit 267>
[0255] The gain updating unit 267 receives input of the number C of consumed bits measured by the arithmetic coding unit 264 .
[0256] When the number of consumed bits C is greater than the number of allocated bits B, the gain update unit 267 updates the value of the global gain g to a larger value and outputs it; when the number of consumed bits C is less than the number of allocated bits B, the gain update unit 267 updates the value of the global gain g to a smaller value and outputs the updated value of the global gain g (step S267).
[0257] The updated global gain g obtained by the gain updating unit 267 is output to the quantization unit 262 and the gain encoding unit 265 .
[0258] <Gain Encoding Unit 265>
[0259] The gain encoding unit 265 receives input of the output instruction from the determination unit 266 and the global gain g obtained by the gain updating unit 267 .
[0260] The gain coding unit 265 codes the global gain g according to the instruction signal to obtain a gain code and outputs the obtained gain code (step 265 ).
[0261] The integer signal code output by the determination unit 266 and the gain code output by the gain encoding unit 265 are output to the parameter determination unit 27 as codes corresponding to the normalized MDCT coefficient sequence.
[0262] That is, in this specific example 2, the last step S267 corresponds to the above-mentioned step A61, and steps S262, S263, S264, and S265 correspond to the above-mentioned steps A62, A63, A64, and A65, respectively.
[0263] Note that Specific Example 2 of the encoding process performed by the encoding unit 26 is described in further detail in International Publication No. WO2014 / 054556 and the like.
[0264] [Modification of the Encoding Unit 26]
[0265] The encoding unit 26 may perform encoding in which the bit allocation is changed based on the estimated spectrum envelope (non-smoothed amplitude spectrum envelope) by performing the following processing, for example.
[0266] The encoding unit 26 first obtains the normalized MDCT coefficient string X N (0),X N (1),…,X N (N-1) corresponding to the global gain g, to find the normalized MDCT coefficient string X N (0),X N (1),…,X N The result of dividing each coefficient of (N-1) by the global gain g is a quantized integer value sequence, that is, the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q (N-1).
[0267] Assume that the quantized normalized coefficient sequence X Q (0),X Q (1),…,X Q The quantization bits corresponding to the coefficients of (N-1) areQ (k) is consistent with the distribution range, and the range can be determined based on the estimated value of the envelope. It is also possible to encode the estimated value of the envelope for each of a plurality of samples, but the encoding unit 26 can use, for example, the value of the normalized amplitude spectrum envelope sequence based on linear prediction as shown in the following equation (A9): N (k) to determine X Q (k) Scope.
[0268]
Number 17
[0269]
[0270] In the case of X in a certain k Q (k) When quantizing, in order to Q The square error of (k) is set to the minimum, which can be based on
[0271]
Number 18
[0272] The number of bits to be allocated is set to b(k)
[0273]
Number 19
[0274]
[0275] B is a predetermined positive integer. At this time, the encoding unit 26 may perform a process of readjusting b(k) by rounding b(k) to an integer or setting b(k)=0 when b(k) is less than 0.
[0276] Furthermore, the encoding unit 26 can also determine the number of allocated bits by aggregating a plurality of samples instead of allocating bits for each sample, and can also perform vector quantization for aggregating a plurality of samples instead of scalar quantization for each sample.
[0277] If the X of sample k Q The number of quantization bits b(k) of (k) is provided above, and each sample is encoded, then X Q (k) can be -2 b(k)-1 to 2 b(k)-1 This 2 b(k) The encoding unit 26 encodes each sample with b(k) bits to obtain an integer signal code.
[0278] The generated integer signal code is output to the decoding device. Q The b(k)-bit integer signal codes corresponding to (k) are sequentially output to the decoding device starting from k=0.
[0279] Assume X Q (k) Exceeding -2 above b(k)-1 to 2b(k)-1 In the case of a range of , it is replaced by the maximum value or the minimum value.
[0280] If g is too small, quantization distortion will occur in the replacement. If g is too large, the quantization error will increase. Q The range of (k) is too small compared to b(k), and information cannot be used effectively. Therefore, g can be optimized.
[0281] The encoding unit 26 encodes the global gain g to obtain a gain code and outputs the obtained gain code.
[0282] As shown in this modified example of the encoding unit 26, the encoding unit 26 can perform encoding other than arithmetic coding.
[0283] <Parameter determination unit 27>
[0284] Through the processing of steps A1 to A6, the code generated for each parameter η for the frequency domain sample string corresponding to the timing signal of the same predetermined time interval (in this example, the linear prediction coefficient code, gain code and integer signal code) is input into the parameter determination unit 27.
[0285] The parameter determination unit 27 selects a code from the codes obtained for each parameter n for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval, and determines the parameter n corresponding to the selected code (step A7). The determined parameter n becomes the parameter n for the frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval. The parameter determination unit 27 then outputs the selected code and the parameter code representing the determined parameter n to the decoding device. The code is selected based on at least one of the code size and the coding distortion corresponding to the code. For example, the code with the smallest code size or the code with the smallest coding distortion is selected.
[0286] Here, coding distortion refers to the error between the frequency domain sample sequence obtained from the input signal and the frequency domain sample sequence obtained by locally decoding the generated code. The encoding device may include a coding distortion calculation unit for calculating coding distortion. This coding distortion calculation unit includes a decoding unit that performs the same processing as the decoding device described below, and locally decodes the code generated by the decoding unit. The coding distortion calculation unit then calculates the error between the frequency domain sample sequence obtained from the input signal and the frequency domain sample sequence obtained by local decoding, and sets this error as coding distortion.
[0287] (decoding)
[0288] Figure 9 : shows an example of the structure of a decoding device corresponding to the encoding device. Figure 9As shown, the decoding device of the first embodiment includes, for example, a linear prediction coefficient decoding unit 31, a non-smoothed amplitude spectrum envelope sequence generation unit 32, a smoothed amplitude spectrum envelope sequence generation unit 33, a decoding unit 34, an envelope denormalization unit 35, a time domain conversion unit 36, and a parameter decoding unit 37. Figure 10 An example of each process of the decoding method according to the first embodiment implemented by the decoding device is shown.
[0289] The decoding device receives as input at least the parameter code output by the encoding device, a code corresponding to the normalized MDCT coefficient string, and a linear prediction coefficient code.
[0290] The following describes Figure 9 's various departments.
[0291] <Parameter decoding unit 37>
[0292] The parameter code output from the encoding device is input to the parameter decoding unit 37 .
[0293] The parameter decoding unit 37 decodes the parameter code to obtain a decoding parameter η. The obtained decoding parameter η is output to the unsmoothed amplitude spectrum envelope sequence generation unit 32, the smoothed amplitude spectrum envelope sequence generation unit 33, and the decoding unit 34. The parameter decoding unit 37 stores a plurality of decoding parameter η candidates. The parameter decoding unit 37 obtains a candidate decoding parameter η corresponding to the parameter code as the decoding parameter η. The plurality of decoding parameters η stored in the parameter decoding unit 37 are the same as the plurality of parameters η stored in the parameter determination unit 27 of the encoding device.
[0294] <Linear Prediction Coefficient Decoding Unit 31>
[0295] The linear prediction coefficient code outputted from the encoding device is inputted to the linear prediction coefficient decoding unit 31 .
[0296] The linear prediction coefficient decoding unit 31 decodes the input linear prediction coefficient code for each frame using, for example, a conventional decoding technique, to obtain decoded linear prediction coefficients ^β1, ^β2, ..., ^β p (Step B1).
[0297] The obtained decoded linear prediction coefficients ^β1,^β2,…,^β p The signal is output to the non-smoothed amplitude spectrum envelope sequence generation unit 32 and the non-smoothed amplitude spectrum envelope sequence generation unit 33 .
[0298] Here, conventional decoding techniques include, for example, techniques that, when the linear prediction coefficient code corresponds to a quantized linear prediction coefficient, yield decoded linear prediction coefficients identical to the quantized linear prediction coefficients obtained by decoding the linear prediction coefficient code; and techniques that, when the linear prediction coefficient code corresponds to a quantized LSP parameter, yield decoded LSP parameters identical to the quantized LSP parameters obtained by decoding the linear prediction coefficient code. Furthermore, as is well known, linear prediction coefficients and LSP parameters can be converted into each other; conversion between decoded linear prediction coefficients and decoded LSP parameters can be performed based on the input linear prediction coefficient code and information required for subsequent processing. The above-described techniques, including the decoding of the linear prediction coefficient code and the conversion performed as needed, are referred to as "decoding based on conventional decoding techniques."
[0299] In this way, the linear prediction coefficient decoding unit 31 decodes the input linear prediction coefficient code to generate coefficients that can be converted into linear prediction coefficients corresponding to the pseudo-correlation function signal string, which is obtained by performing an inverse Fourier transform of the ηth power of the absolute value of the frequency domain sample string corresponding to the time series signal as the power spectrum.
[0300] <Non-smoothed Amplitude Spectrum Envelope Sequence Generator 32>
[0301] The non-smoothed amplitude spectrum envelope sequence generator 32 receives as input the decoding parameter η obtained by the parameter decoder 37 and the decoded linear prediction coefficients ^β1, ^β2, ..., ^β obtained by the linear prediction coefficient decoder 31. p .
[0302] The non-smoothed amplitude spectrum envelope sequence generator 32 generates and decodes the linear prediction coefficients ^β1, ^β2, ..., ^β using the above equation (A2). p The corresponding sequence of amplitude spectrum envelopes is the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) (step B2).
[0303] The generated non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), . . . , ^H(N−1) is output to the decoding unit 34 .
[0304] In this way, the non-smoothed amplitude spectrum envelope sequence generator 32 obtains a sequence obtained by raising the sequence of amplitude spectrum envelopes corresponding to the coefficients convertible into linear prediction coefficients generated by the linear prediction coefficient decoder 31 to the power of 1 / η, i.e., a non-smoothed spectrum envelope sequence.
[0305] <Smoothed Amplitude Spectrum Envelope Sequence Generator 33>
[0306] The smoothed amplitude spectrum envelope sequence generator 33 receives as input the decoding parameter η obtained by the parameter decoder 37 and the decoded linear prediction coefficients ^β1, ^β2, ..., ^β obtained by the linear prediction coefficient decoder 31. p .
[0307] The smoothed amplitude spectrum envelope sequence generator 33 generates the weakened and decoded linear prediction coefficients ^β1, ^β2, ..., ^β by using the above formula A(3). p The corresponding sequence of amplitude convexity of the amplitude spectrum envelope sequence is the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1)(Step B3).
[0308] The generated smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N−1) is output to the decoding unit 34 and the envelope denormalization unit 35 .
[0309] <Decoding Unit 34>
[0310] The decoding unit 34 receives as input the decoding parameter η obtained by the parameter decoding unit 37, the code corresponding to the normalized MDCT coefficient string output by the encoding device, the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the non-smoothed amplitude spectrum envelope generation unit 32, and the smoothed amplitude spectrum envelope sequence ^H generated by the smoothed amplitude spectrum envelope generation unit 33. γ (0),^H γ (1),…,^H γ (N-1).
[0311] The decoding unit 34 includes a variance parameter determination unit 342 .
[0312] The decoding unit 34 performs, for example, Figure 11 The decoding is performed by performing the processing from step B41 to step B44 shown in FIG. 4 (step B4). That is, the decoding unit 34 decodes the gain code included in the code corresponding to the input normalized MDCT coefficient string for each frame to obtain the global gain g (step B41). The variance parameter determination unit 342 of the decoding unit 34 determines the global gain g based on the global gain g and the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1), through the above formula (A1), the variance parameter sequence is obtained The decoding unit 34 calculates the variance parameters of the sequence of variance parameters. The structure of arithmetic decoding corresponding to each variance parameter is used to perform arithmetic decoding on the integer signal code contained in the code corresponding to the normalized MDCT coefficient string to obtain a decoded normalized coefficient sequence ^X Q (0),^X Q (1),…,^X Q (N-1) (step B43), decode the normalized coefficient sequence ^X Q (0),^X Q (1),…,^X Q Each coefficient of (N-1) is multiplied by the global gain g to generate the decoded normalized MDCT coefficient string ^X N (0),^X N (1),…,^X N (N-1) (Step B44) In this way, the decoding unit 34 can decode the input integer signal code based on the bit allocation that is substantially changed based on the non-smoothed spectrum envelope sequence.
[0313] In addition, when encoding is performed by the process described in [Variation of Encoding Unit 26], the decoding unit 34 performs the following process, for example. The decoding unit 34 decodes the gain code included in the code corresponding to the input normalized MDCT coefficient string for each frame to obtain the global gain g. The variance parameter determination unit 342 of the decoding unit 34 determines the global gain g based on the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) and the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1), through the above formula (A9), the variance parameter sequence is obtained The decoding unit 34 can be based on the variance parameter sequence The variance parameters of Calculate b(k) using formula (A10), and set X Q The value of (k) is decoded in sequence by its bit number b(k) to obtain the decoded normalized coefficient sequence ^X Q (0),^X Q (1),…,^X Q (N-1), decoded normalized coefficient sequence ^X Q (0),^X Q (1),…,^X Q Each coefficient of (N-1) is multiplied by the global gain g to generate the decoded normalized MDCT coefficient string ^X N (0),^X N (1),…,^XN (N-1) In this way, the decoding unit 34 can decode the input integer signal code according to the bit allocation that changes based on the non-smoothed spectrum envelope sequence.
[0314] The generated decoded normalized MDCT coefficient string ^X N (0),^X N (1),…,^X N (N−1) is output to the envelope denormalization unit 35 .
[0315] <Envelope Denormalization Section 35>
[0316] The smoothed amplitude spectrum envelope sequence generated by the smoothed amplitude spectrum envelope generation unit 33 is input to the envelope denormalization unit 35. γ (0),^H γ (1),…,^H γ (N-1) and the decoded normalized MDCT coefficient string ^X generated by the decoding unit 34 N (0),^X N (1),…,^X N (N-1).
[0317] The envelope denormalization unit 35 uses the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ (N-1), decode the normalized MDCT coefficient string ^X N (0),^X N (1),…,^X N (N-1) is inversely normalized to generate a decoded MDCT coefficient string ^X(0), ^X(1), ..., ^X(N-1) (step B5).
[0318] The generated decoded MDCT coefficient strings ^X(0), ^X(1), . . . , ^X(N−1) are output to the time domain conversion unit 36 .
[0319] For example, the envelope denormalization unit 35 sets k=0, 1, ..., N-1, and normalizes the decoded MDCT coefficient string ^X N (0),^X N (1),…,^X N The coefficients of (N-1)^X N (k) multiplied by the smoothed amplitude spectrum envelope sequence ^H γ (0),^H γ (1),…,^H γ Each envelope value of (N-1)^H γ(k) and generate the decoded MDCT coefficient string ^X(0),^X(1),...,^X(N-1). That is, let k = 0, 1,..., N-1, ^X(k) = ^X N (k)×^H γ (k).
[0320] <Time Domain Converter 36>
[0321] The decoded MDCT coefficient sequence ^X(0), ^X(1), . . . , ^X(N-1) generated by the envelope denormalization unit 35 is input to the time domain conversion unit 36 .
[0322] The time domain converter 36 converts the decoded MDCT coefficient sequence ^X(0), ^X(1), ..., ^X(N-1) obtained by the envelope denormalizer 35 into the time domain for each frame to obtain a frame-based audio signal (decoded audio signal) (step B6).
[0323] In this way, the decoding device obtains a time series signal through decoding in the frequency domain.
[0324] [Second embodiment]
[0325] The encoding device and method of the first embodiment are an encoding device and method that encode each of a plurality of parameters η to generate a code, select an optimal code from the codes generated for each parameter η, and output the selected code and a parameter code corresponding to the selected code.
[0326] In contrast, the encoding device and method of the second embodiment first determine parameter n by parameter determination unit 27, then perform encoding based on the determined parameter n to generate and output a code. In the second embodiment, parameter n is made variable by parameter determination unit 27 for each predetermined time interval. Here, making parameter n variable for each predetermined time interval means that parameter n can also change if the predetermined time interval changes, and it is assumed that the value of parameter n does not change within the same time interval.
[0327] The following description will focus on the parts that are different from the first embodiment, and duplicate descriptions of the parts that are the same as the first embodiment will be omitted.
[0328] (coding)
[0329] Figure 12 FIG. 4 shows a configuration example of an encoding device according to the second embodiment. Figure 12 As shown, the encoding device includes, for example, a frequency domain conversion unit 21, a linear prediction analysis unit 22, a non-smoothed amplitude spectrum envelope sequence generation unit 23, a smoothed amplitude spectrum envelope sequence generation unit 24, an envelope normalization unit 25, an encoding unit 26, and a parameter determination unit 27'. Figure 13An example of each process of the encoding method implemented by the encoding device is shown.
[0330] The following describes Figure 12 's various departments.
[0331] <Parameter determination unit 27'>
[0332] The parameter determination unit 27' receives as input a time-domain audio signal as a time-series signal. Examples of the audio signal include a digital voice signal or a digital sound signal.
[0333] The parameter determination unit 27 ′ determines the parameter η through a process described later based on the input timing signal (step A7 ′).
[0334] η determined by the parameter determination unit 27 ′ is output to the linear prediction analysis unit 22 , the non-smoothed amplitude spectrum envelope estimation unit 23 , the smoothed amplitude spectrum envelope estimation unit 24 , and the encoding unit 26 .
[0335] Furthermore, the parameter determination unit 27′ generates a parameter code by encoding the determined n. The generated parameter code is sent to the decoding device.
[0336] The parameter determination unit 27 ′ will be described in detail later.
[0337] The frequency domain conversion unit 21, linear prediction analysis unit 22, unsmoothed amplitude spectrum envelope sequence generation unit 23, smoothed amplitude spectrum envelope sequence generation unit 24, envelope normalization unit 25, and encoding unit 26 generate a code (steps A1 to A6) based on the parameter η determined by the parameter determination unit 27' through the same process as in the first embodiment. In this example, the code is a combination of a linear prediction coefficient code, a gain code, and an integer signal code. The generated code is transmitted to the decoding device.
[0338] Figure 14 27 ' shows a configuration example of the parameter determination unit. Figure 14 As shown, the parameter determination unit 27' includes, for example, a frequency domain conversion unit 41, a spectrum envelope estimation unit 42, a whitening spectrum sequence generation unit 43, and a parameter acquisition unit 44. The spectrum envelope estimation unit 42 includes, for example, a linear prediction analysis unit 421 and a non-smoothed amplitude spectrum envelope sequence generation unit 422. For example, Figure 2 An example of each process of the parameter determination method implemented by the parameter determination unit 27 ′ is shown.
[0339] The following describes Figure 14 's various departments.
[0340] <Frequency Domain Converter 41>
[0341] A time-domain audio signal, which is a time-series signal, is input to the frequency domain converter 41. Examples of the audio signal are a digital voice signal or a digital sound signal.
[0342] The frequency domain converter 41 converts the inputted time domain audio signal into an N-point MDCT coefficient sequence X(0), X(1), ..., X(N-1) in the frequency domain in units of frames of a predetermined time length, where N is a positive integer.
[0343] The obtained MDCT coefficient sequence X(0), X(1), . . . , X(N−1) is output to the spectrum envelope estimation unit 42 and the whitened spectrum sequence generation unit 43 .
[0344] Unless otherwise specified, subsequent processing is performed in units of frames.
[0345] In this way, the frequency domain converter 41 obtains a frequency domain sample sequence, for example, an MDCT coefficient sequence corresponding to the audio signal (step C41 ).
[0346] <Spectrum Envelope Estimation Unit 42>
[0347] The MDCT coefficient sequence X(0), X(1), . . . , X(N-1) obtained by the frequency domain conversion section 41 is input to the spectrum envelope estimation section 42.
[0348] The spectrum envelope estimation unit 42 estimates the spectrum envelope using the η0 power of the absolute value of the frequency domain sample sequence corresponding to the time series signal as the power spectrum based on the parameter η0 determined by a predetermined method (step C42).
[0349] The estimated spectrum envelope is output to the whitening spectrum sequence generation unit 43 .
[0350] The spectrum envelope estimation unit 42 generates a non-smoothed amplitude spectrum envelope sequence through processing by, for example, a linear prediction analysis unit 421 and a non-smoothed amplitude spectrum envelope sequence generation unit 422 described below, thereby estimating the spectrum envelope.
[0351] Assume that parameter η0 is determined by a predetermined method. For example, η0 is set to a predetermined number greater than 0. For example, η0 = 1. Alternatively, η obtained in a frame earlier than the frame for which the current parameter η is desired to be obtained may be used. The frame earlier than the frame for which the current parameter η is desired to be obtained (hereinafter referred to as the current frame) is, for example, a frame earlier than the current frame and a frame near the current frame. The frame near the current frame is, for example, the frame immediately before the current frame.
[0352] <Linear prediction analysis unit 421>
[0353] The MDCT coefficient sequence X(0), X(1), . . . , X(N-1) obtained by the frequency domain conversion unit 41 is input to the linear prediction analysis unit 421.
[0354] The linear prediction analysis unit 421 uses the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to perform linear prediction analysis on ~R(0), ~R(1), ..., ~R(N-1) defined by the following equation (C1) to generate linear prediction coefficients β1, β2, ..., β p , and generate linear prediction coefficients β1,β2,…,β p The linear prediction coefficient code and the quantized linear prediction coefficients corresponding to the linear prediction coefficient code, that is, the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p .
[0355]
Number 20
[0356]
[0357] The generated quantized linear prediction coefficients ^β1,^β2,…,^β p The signal is output to the non-smoothed spectrum envelope sequence generation unit 422 .
[0358] Specifically, the linear prediction analysis unit 421 first obtains a time domain signal sequence corresponding to the power of the absolute value of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to the power of η0 by performing an operation equivalent to the inverse Fourier transform, i.e., the operation of formula (C1), using the absolute value of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to the power of η0. In addition, the linear prediction analysis unit 421 uses the obtained pseudo-correlation function signal sequence ~R(0), ~R(1), ..., ~R(N-1) to perform linear prediction analysis and generate linear prediction coefficients β1, β2, ..., β p Furthermore, the linear prediction analysis unit 421 generates linear prediction coefficients β1, β2, ..., β p Encode to obtain the linear prediction coefficient code and the quantized linear prediction coefficients ^β1, ^β2, ..., ^β corresponding to the linear prediction coefficient code p .
[0359] Linear prediction coefficients β1,β2,…,β p It is the linear prediction coefficient corresponding to the time domain signal when the η0 power of the absolute value of the MDCT coefficient string X(0), X(1), ..., X(N-1) is regarded as the power spectrum.
[0360] The linear prediction analysis unit 421 generates the linear prediction coefficient code using, for example, existing coding techniques. Examples of existing coding techniques include: a coding technique that uses the code corresponding to the linear prediction coefficient itself as the linear prediction coefficient code; a coding technique that converts the linear prediction coefficient into an LSP parameter and uses the code corresponding to the LSP parameter as the linear prediction coefficient code; and a coding technique that converts the linear prediction coefficient into a PARCOR coefficient and uses the code corresponding to the PARCOR coefficient as the linear prediction coefficient code.
[0361] In this way, the linear prediction analysis unit 421 performs linear prediction analysis using, for example, a pseudo-correlation function signal string obtained by performing inverse Fourier transform of the MDCT coefficient string, i.e., the absolute value of the frequency domain sample string to the power of η, as the power spectrum, to generate coefficients that can be converted into linear prediction coefficients (step C421).
[0362] <Non-smoothed Amplitude Spectrum Envelope Sequence Generator 422>
[0363] The non-smoothed amplitude spectrum envelope sequence generation unit 422 receives the quantized linear prediction coefficients ^β1, ^β2, ..., ^β generated by the linear prediction analysis unit 421. p .
[0364] The non-smoothed amplitude spectrum envelope sequence generation unit 422 generates and quantizes the linear prediction coefficients ^β1, ^β2, ..., ^β p The corresponding sequence of amplitude spectrum envelopes is the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1),…, ^H(N-1).
[0365] The generated non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), . . . , ^H(N−1) is output to the whitening spectrum sequence generation unit 43 .
[0366] The non-smoothed amplitude spectrum envelope sequence generator 422 uses the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p , as the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1), the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1) defined by formula (C2) is generated.
[0367]
Number 21
[0368]
[0369] In this way, the non-smoothed amplitude spectrum envelope sequence generation unit 422 estimates the spectrum envelope by obtaining a sequence of the amplitude spectrum envelope corresponding to the pseudo-correlation function signal string raised to the power of 1 / η0, i.e., a non-smoothed spectrum envelope sequence, based on the coefficients that can be converted into linear prediction coefficients generated by the linear prediction analysis unit 421 (step C422).
[0370] <Whitening spectrum sequence generation unit 43>
[0371] The whitening spectrum sequence generation unit 43 receives input of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) obtained by the frequency domain conversion unit 41 and the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the non-smoothed amplitude spectrum envelope generation unit 422.
[0372] The whitened spectrum sequence generator 43 generates a whitened spectrum sequence X by dividing each coefficient of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) by each value of the corresponding non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1). W (0),X W (1),…,X W (N-1).
[0373] The generated whitened spectrum sequence X W (0),X W (1),…,X W (N-1) is output to the parameter acquisition unit 44 .
[0374] The whitened spectrum sequence generator 43 generates a whitened spectrum sequence X by dividing each coefficient X(k) of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) by each value ^H(k) of the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1), for example, with k=0, 1, ..., N-1. W (0),X W (1),…,X W Each value X of (N-1) W (k). That is, let k = 0, 1, ..., N-1, X W (k) = X(k) / ^H(k).
[0375] In this way, the whitened spectrum sequence generator 43 obtains a whitened spectrum sequence obtained by dividing, for example, the MDCT coefficient sequence, that is, the frequency domain sample sequence, by, for example, the spectrum envelope sequence, that is, the non-smoothed amplitude spectrum envelope sequence (step C43 ).
[0376] <Parameter Acquisition Unit 44>
[0377] The parameter acquisition unit 44 receives the whitened spectrum sequence X generated by the whitened spectrum sequence generation unit 43 as input. W (0),X W (1),…,X W (N-1).
[0378] The parameter acquisition unit 44 obtains the generalized Gaussian distribution with the parameter η as the shape parameter for the whitened spectrum sequence X. W (0),X W (1),…,X W In other words, the parameter acquisition unit 44 determines that the parameter η is a generalized Gaussian distribution whose shape parameter is close to the whitened spectrum sequence X. W (0),X W (1),…,X W The parameter η of the distribution of the (N-1) histogram.
[0379] The generalized Gaussian distribution with parameter η as the shape parameter is defined as follows: γ is the gamma function.
[0380]
Number 22
[0381]
[0382] The generalized Gaussian distribution is obtained by changing the shape parameter η, as Figure 3 As shown, various distributions can be expressed, such as a Laplace distribution when η=1 and a Gaussian distribution when η=2. is the parameter corresponding to the variance.
[0383] Here, η obtained by the parameter acquisition unit 44 is defined by, for example, the following formula (C3). -1 It is the inverse function of function F. This formula is derived by the so-called moment method.
[0384]
Number 23
[0385]
[0386] In the inverse function F -1 When the parameter acquisition unit 44 is explicitly defined, the inverse function F -1 Enter m1 / ((m2) 1 / 2 ) value, the parameter η can be obtained.
[0387] In the inverse function F -1In the case where it is not explicitly defined, in order to calculate the value of η defined in Equation (C3), for example, the parameter acquisition unit 44 can obtain the parameter η by the first method or the second method described below.
[0388] Describe the first method for obtaining the parameter η. In the first method, the parameter acquisition unit 44 calculates m1 / ((m2) 1 / 2 ) based on the whitened spectrum sequence, refers to a pair of multiple different ηs prepared in advance and F(η) corresponding to η, and obtains the η corresponding to the F(η) closest to the calculated m1 / ((m2) 1 / 2 ).
[0389] Pairs of multiple different ηs prepared in advance and F(η) corresponding to η are stored in the storage unit 441 of the parameter acquisition unit 44 in advance. The parameter acquisition unit 44 refers to the storage unit 441, finds the F(η) closest to the calculated m1 / ((m2) 1 / 2 ), and reads and outputs the η corresponding to the found F(η) from the storage unit 441.
[0390] The F(η) closest to the calculated m1 / ((m2) 1 / 2 ) means the F(η) for which the absolute value of the difference from the calculated m1 / ((m2) 1 / 2 ) becomes the smallest.
[0391] Describe the second method for obtaining the parameter η. In the second method, the inverse function F -1 's approximate curve function is, for example, expressed by the following Equation (C3'), ~ F -1 , and the parameter acquisition unit 44 calculates m1 / ((m2) 1 / 2 ) based on the whitened spectrum sequence, and obtains η by calculating the output value when the calculated m1 / ((m2) ~ F -1 is input to the approximate curve function 1 / 2 .
[0392] In addition, the η obtained by the parameter acquisition unit 44 can be defined by generalizing Equation (C3) by using positive integers q1 and q(where q1 < q2) determined in advance not as in Equation (C3) but as in Equation (C3").
[0393]
Equation 24
[0394]
[0395] In addition, when η is defined by the formula (C3"), η can also be obtained by the same method as when η is defined by the formula (C3). That is, the parameter acquisition unit 44 calculates m based on the q1 moment of the whitened spectrum sequence. q1 and m as its q2 moment q2 The value of m q1 / ((m q2 ) q1 / q2 ) After that, for example, similarly to the first and second methods described above, it is possible to refer to a plurality of different η prepared in advance and the F'(η) corresponding to η to obtain the m closest to the calculated value. q1 / ((m q2 ) q1 / q2 ) of F'(η) corresponding to η, or the inverse function F' -1 The approximate curve function is as ~F' -1 , calculated on the approximate curve function ~F -1 Enter the calculated m q1 / ((m q2 ) q1 / q2 ) and find η.
[0396] In this way, it can be said that η is based on two different moments m of different order. q1 ,m q2 For example, we can use two different moments m of different orders to calculate the value of q1 ,m q2 η is obtained by the ratio of the value of the lower-order moment or its value (hereinafter referred to as the former) to the value of the higher-order moment or its value (hereinafter referred to as the latter), the value based on the ratio, or the value obtained by dividing the former by the latter. For example, if the moment is m and Q is a predetermined real number, it is m Q In addition, these values can be input into the approximate curve function ~F -1 And find η. As above, the approximate curve function ~F' -1 It is sufficient to use a monotonically increasing function whose output is positive in the domain of definition used.
[0397] The parameter determination unit 27' may determine the parameter η through a loop process. That is, the parameter determination unit 27' may further perform the process of the spectrum envelope estimation unit 42, the whitened spectrum sequence generation unit 43, and the parameter acquisition unit 44 one or more times, setting the parameter η determined by the parameter acquisition unit 44 to the parameter η0 determined by a predetermined method.
[0398] At this time, for example, as in Figure 14As shown by the middle dashed line, the parameter η obtained by the parameter acquisition unit 44 is output to the spectrum envelope estimation unit 42. The spectrum envelope estimation unit 42 uses the η obtained by the parameter acquisition unit 44 as the parameter η0 and performs the same processing as described above to estimate the spectrum envelope. The whitened spectrum sequence generation unit 43 performs the same processing as described above based on the newly estimated spectrum envelope to generate a whitened spectrum sequence. The parameter acquisition unit 44 performs the same processing as described above based on the newly generated whitened spectrum sequence to obtain the parameter η.
[0399] For example, the processing of the spectrum envelope estimation unit 42, the whitening spectrum sequence generation unit 43, and the parameter acquisition unit 44 may be further performed a predetermined number of times, τ. τ is a predetermined positive integer, for example, τ=1 or τ=2.
[0400] The spectrum envelope estimation unit 42 may repeat the processes of the spectrum envelope estimation unit 42 , the whitening spectrum sequence generation unit 43 , and the parameter acquisition unit 44 until the absolute value of the difference between the parameter η obtained this time and the parameter η obtained previously becomes equal to or smaller than a predetermined threshold.
[0401] (decoding)
[0402] Since the decoding device and method of the second embodiment are the same as those of the first embodiment, repeated description will be omitted.
[0403] [[Modification of the Second Embodiment]]
[0404] In addition, as long as the configuration of the encoding process can be determined based on at least the parameter η, the encoding process may be any process, and an encoding process other than the encoding process of the encoding unit 26 may be used.
[0405] A modification of the second embodiment will be described below, in which the encoding process is not limited to the encoding process performed by the encoding unit 26 .
[0406] (coding)
[0407] An example of an encoding device and method according to a modified example of the second embodiment will be described.
[0408] like Figure 17 As shown, the encoding device of the modified example of the second embodiment includes, for example, a parameter determination unit 27', an acoustic feature extraction unit 521, a determination unit 522, and an encoding unit 523. Figure 18 The encoding method is realized by performing the various processes illustrated in FIG.
[0409] Hereinafter, each unit of the encoding device will be described.
[0410] <Parameter determination unit 27'>
[0411] The parameter determination unit 27' receives as input a time-domain audio signal in units of frames, which is a time-series signal. Examples of the audio signal include a digital voice signal or a digital sound signal.
[0412] The parameter determination unit 27' determines the parameter η based on the input time series signal through the processing described below (step FE1). The parameter determination unit 27' performs processing for each frame of a predetermined time length. In other words, the parameter η is determined for each frame.
[0413] The parameter η determined by the parameter determination unit 27 ′ is output to the determination unit 522 .
[0414] Figure 21 27 ' shows a configuration example of the parameter determination unit. Figure 21 As shown, the parameter determination unit 27' includes, for example, a frequency domain conversion unit 41, a spectrum envelope estimation unit 42, a whitening spectrum sequence generation unit 43, and a parameter acquisition unit 44. The spectrum envelope estimation unit 42 includes, for example, a linear prediction analysis unit 421 and a non-smoothed amplitude spectrum envelope sequence generation unit 422. For example, Figure 22 An example of each process of the parameter determination method implemented by the parameter determination unit 27 ′ is shown.
[0415] The following describes Figure 21 's various departments.
[0416] <Frequency Domain Converter 41>
[0417] The frequency domain conversion unit 41 receives input of a time-domain audio signal as a time-series signal.
[0418] The frequency domain converter 41 converts the inputted time domain audio signal into an N-point MDCT coefficient sequence X(0), X(1), ..., X(N-1) in the frequency domain in units of frames of a predetermined time length, where N is a positive integer.
[0419] The obtained MDCT coefficient sequence X(0), X(1), . . . , X(N−1) is output to the spectrum envelope estimation unit 42 and the whitened spectrum sequence generation unit 43 .
[0420] Unless otherwise specified, subsequent processing is performed in units of frames.
[0421] In this way, the frequency domain converter 41 obtains a frequency domain sample sequence, for example, an MDCT coefficient sequence corresponding to the time series signal (step C41 ).
[0422] <Spectrum Envelope Estimation Unit 42>
[0423] The MDCT coefficient sequence X(0), X(1), . . . , X(N-1) obtained by the frequency domain conversion section 41 is input to the spectrum envelope estimation section 42.
[0424] The spectrum envelope estimation unit 42 estimates the spectrum envelope using the η0 power of the absolute value of the frequency domain sample sequence corresponding to the time series signal as the power spectrum based on the parameter η0 determined by a predetermined method (step C42).
[0425] The estimated spectrum envelope is output to the whitening spectrum sequence generation unit 43 .
[0426] The spectrum envelope estimation unit 42 generates a non-smoothed amplitude spectrum envelope sequence through processing by, for example, a linear prediction analysis unit 421 and a non-smoothed amplitude spectrum envelope sequence generation unit 422 described below, thereby estimating the spectrum envelope.
[0427] Assume that parameter η0 is determined by a predetermined method. For example, η0 is set to a predetermined number greater than 0. For example, η0 = 1. Alternatively, η obtained in a frame earlier than the frame for which the current parameter η is desired to be obtained may be used. The frame earlier than the frame for which the current parameter η is desired to be obtained (hereinafter referred to as the current frame) is, for example, a frame earlier than the current frame and a frame near the current frame. The frame near the current frame is, for example, the frame immediately before the current frame.
[0428] <Linear prediction analysis unit 421>
[0429] The MDCT coefficient sequence X(0), X(1), . . . , X(N-1) obtained by the frequency domain conversion unit 41 is input to the linear prediction analysis unit 421.
[0430] The linear prediction analysis unit 421 uses the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to perform linear prediction analysis on ~R(0), ~R(1), ..., ~R(N-1) defined by the following equation (C1) to generate linear prediction coefficients β1, β2, ..., β p , and generate linear prediction coefficients β1,β2,…,β p The linear prediction coefficient code and the quantized linear prediction coefficients corresponding to the linear prediction coefficient code, that is, the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p .
[0431]
Number 25
[0432]
[0433] The generated quantized linear prediction coefficients ^β1,^β2,…,^β p The signal is output to the non-smoothed spectrum envelope sequence generation unit 422 .
[0434] Specifically, the linear prediction analysis unit 421 first obtains a time domain signal sequence corresponding to the power of the absolute value of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to the power of η0 by performing an operation equivalent to the inverse Fourier transform, i.e., the operation of formula (C1), using the absolute value of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) to the power of η0. In addition, the linear prediction analysis unit 421 uses the obtained pseudo-correlation function signal sequence ~R(0), ~R(1), ..., ~R(N-1) to perform linear prediction analysis and generate linear prediction coefficients β1, β2, ..., β p Furthermore, the linear prediction analysis unit 421 generates linear prediction coefficients β1, β2, ..., β p Encode to obtain the linear prediction coefficient code and the quantized linear prediction coefficients ^β1, ^β2, ..., ^β corresponding to the linear prediction coefficient code p .
[0435] Linear prediction coefficients β1,β2,…,β p It is the linear prediction coefficient corresponding to the time domain signal when the η0 power of the absolute value of the MDCT coefficient string X(0), X(1), ..., X(N-1) is regarded as the power spectrum.
[0436] The linear prediction analysis unit 421 generates the linear prediction coefficient code using, for example, existing coding techniques. Examples of existing coding techniques include: a coding technique that uses the code corresponding to the linear prediction coefficient itself as the linear prediction coefficient code; a coding technique that converts the linear prediction coefficient into an LSP parameter and uses the code corresponding to the LSP parameter as the linear prediction coefficient code; and a coding technique that converts the linear prediction coefficient into a PARCOR coefficient and uses the code corresponding to the PARCOR coefficient as the linear prediction coefficient code.
[0437] In this way, the linear prediction analysis unit 421 performs linear prediction analysis using, for example, a pseudo-correlation function signal string obtained by performing inverse Fourier transform of the MDCT coefficient string, i.e., the absolute value of the frequency domain sample string to the power of η, to generate a linear prediction coefficient (step C421).
[0438] <Non-smoothed Amplitude Spectrum Envelope Sequence Generator 422>
[0439] The non-smoothed amplitude spectrum envelope sequence generation unit 422 receives the quantized linear prediction coefficients ^β1, ^β2, ..., ^β generated by the linear prediction analysis unit 421. p .
[0440] The non-smoothed amplitude spectrum envelope sequence generation unit 422 generates and quantizes the linear prediction coefficients ^β1, ^β2, ..., ^βp The corresponding sequence of amplitude spectrum envelopes is the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1),…, ^H(N-1).
[0441] The generated non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), . . . , ^H(N−1) is output to the whitening spectrum sequence generation unit 43 .
[0442] The non-smoothed amplitude spectrum envelope sequence generator 422 uses the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p , as the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1), the non-smoothed amplitude spectrum envelope sequence ^H(0),^H(1),…,^H(N-1) defined by formula (C2) is generated.
[0443]
Number 26
[0444]
[0445] In this way, the non-smoothed amplitude spectrum envelope sequence generation unit 422 estimates the spectrum envelope by obtaining a sequence of the amplitude spectrum envelope corresponding to the pseudo-correlation function signal string raised to the power of 1 / η0, i.e., a non-smoothed spectrum envelope sequence, based on the coefficients that can be converted into linear prediction coefficients generated by the linear prediction analysis unit 421 (step C422).
[0446] In addition, the non-smoothed spectrum envelope sequence generation unit 422 may replace the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p The linear prediction coefficients β1, β2, ..., β generated by the linear prediction analysis unit 421 are used. p , and obtain the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1). At this time, the linear prediction analysis unit 421 may not obtain the quantized linear prediction coefficients ^β1, ^β2, ..., ^β p processing.
[0447] <Whitening spectrum sequence generation unit 43>
[0448] The whitening spectrum sequence generation unit 43 receives input of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) obtained by the frequency domain conversion unit 41 and the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1) generated by the non-smoothed amplitude spectrum envelope generation unit 422.
[0449] The whitened spectrum sequence generator 43 generates a whitened spectrum sequence X by dividing each coefficient of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) by each value of the corresponding non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1). W (0),X W (1),…,X W (N-1).
[0450] The generated whitened spectrum sequence X W (0),X W (1),…,X W (N-1) is output to the parameter acquisition unit 44 .
[0451] The whitened spectrum sequence generator 43 generates a whitened spectrum sequence X by dividing each coefficient X(k) of the MDCT coefficient sequence X(0), X(1), ..., X(N-1) by each value ^H(k) of the non-smoothed amplitude spectrum envelope sequence ^H(0), ^H(1), ..., ^H(N-1), for example, with k=0, 1, ..., N-1. W (0),X W (1),…,X W Each value X of (N-1) W (k). That is, let k = 0, 1, ..., N-1, X W (k) = X(k) / ^H(k).
[0452] In this way, the whitened spectrum sequence generator 43 obtains a whitened spectrum sequence obtained by dividing, for example, the MDCT coefficient sequence, that is, the frequency domain sample sequence, by, for example, the spectrum envelope sequence, that is, the non-smoothed amplitude spectrum envelope sequence (step C43 ).
[0453] <Parameter Acquisition Unit 44>
[0454] The parameter acquisition unit 44 receives the whitened spectrum sequence X generated by the whitened spectrum sequence generation unit 43 as input. W (0),X W (1),…,X W (N-1).
[0455] The parameter acquisition unit 44 obtains the generalized Gaussian distribution with the parameter η as the shape parameter for the whitened spectrum sequence X. W (0),X W (1),…,X W In other words, the parameter acquisition unit 44 determines that the parameter η is a generalized Gaussian distribution whose shape parameter is close to the whitened spectrum sequence X. W (0),X W (1),…,X WThe parameter η of the distribution of the (N-1) histogram.
[0456] The generalized Gaussian distribution with parameter η as the shape parameter is defined as follows: γ is the gamma function.
[0457]
Number 27
[0458]
[0459] The generalized Gaussian distribution is obtained by changing the shape parameter η, as Figure 23 As shown, various distributions can be expressed, such as a Laplace distribution when η=1 and a Gaussian distribution when η=2. is the parameter corresponding to the variance.
[0460] Here, η obtained by the parameter acquisition unit 44 is defined by, for example, the following formula (C3). -1 is the inverse function of function F. This formula is derived by the so-called method of moments.
[0461]
Number 28
[0462]
[0463] In the inverse function F -1 When the parameter acquisition unit 44 is explicitly defined, the inverse function F -1 Enter m1 / ((m2) 1 / 2 ) value, the parameter η can be obtained.
[0464] In the inverse function F -1 When it is not explicitly defined, the parameter acquisition unit 44 can obtain the parameter η by, for example, the first method or the second method described below in order to calculate the value of η defined in the formula (C3).
[0465] A first method for obtaining the parameter η will be described. In the first method, the parameter acquisition unit 44 calculates m1 / ((m2)) based on the whitened spectrum sequence. 1 / 2 ), refer to the pre-prepared different multiple η and F (η) corresponding to η, and obtain the m1 / ((m2)) closest to the calculated 1 / 2 ) corresponds to the η of F(η).
[0466] The pairs of different η and F(η) corresponding to η prepared in advance are stored in the storage unit 441 of the parameter acquisition unit 44. The parameter acquisition unit 44 refers to the storage unit 441 and finds the closest m1 / ((m2)) to the calculated value. 1 / 2 ) of F(η), and reads η corresponding to the found F(η) from the storage unit 441 and outputs it.
[0467] The F(η) closest to the calculated m1 / ((m2) 1 / 2 ) refers to the F(η) for which the absolute value of the difference from the calculated m1 / ((m2) 1 / 2 ) becomes the smallest.
[0468] Describe the second method for obtaining the parameter η. In the second method, the approximate curve function of the inverse function F -1 is, for example, expressed by the following formula (C3’) ~ F -1 , and the parameter acquisition unit 44 calculates m1 / ((m2) 1 / 2 ) based on the whitened frequency spectrum sequence, and obtains η by calculating the output value when the calculated m1 / ((m2) ~ F -1 ) is input to the approximate curve function 1 / 2 .
[0469] In addition, η obtained by the parameter acquisition unit 44 can be defined by generalizing formula (C3) by using positive integers q1 and q2 (where q1 < q2) determined in advance not as in formula (C3) but as in formula (C3”).
[0470]
Equation 29
[0471] <00q1 ,m q2 For example, we can use two different moments m of different orders to calculate the value of q1 ,m q2 η is obtained by the ratio of the value of the lower-order moment or its value (hereinafter referred to as the former) to the value of the higher-order moment or its value (hereinafter referred to as the latter), the value based on the ratio, or the value obtained by dividing the former by the latter. For example, if the moment is m and Q is a predetermined real number, it is m Q In addition, these values can be input into the approximate curve function ~F -1 And find η. As above, the approximate curve function ~F' -1 It is sufficient to use a monotonically increasing function whose output is positive in the domain of definition used.
[0474] The parameter determination unit 27' may determine the parameter η through a loop process. That is, the parameter determination unit 27' may further perform the process of the spectrum envelope estimation unit 42, the whitened spectrum sequence generation unit 43, and the parameter acquisition unit 44 one or more times, setting the parameter η determined by the parameter acquisition unit 44 to the parameter η0 determined by a predetermined method.
[0475] At this time, for example, as in Figure 21 As shown by the middle dashed line, the parameter η obtained by the parameter acquisition unit 44 is output to the spectrum envelope estimation unit 42. The spectrum envelope estimation unit 42 uses the η obtained by the parameter acquisition unit 44 as the parameter η0 and performs the same processing as described above to estimate the spectrum envelope. The whitened spectrum sequence generation unit 43 performs the same processing as described above based on the newly estimated spectrum envelope to generate a whitened spectrum sequence. The parameter acquisition unit 44 performs the same processing as described above based on the newly generated whitened spectrum sequence to obtain the parameter η.
[0476] For example, the processing of the spectrum envelope estimation unit 42, the whitening spectrum sequence generation unit 43, and the parameter acquisition unit 44 may be further performed a predetermined number of times, τ. τ is a predetermined positive integer, for example, τ=1 or τ=2.
[0477] The spectrum envelope estimation unit 42 may repeat the processes of the spectrum envelope estimation unit 42 , the whitening spectrum sequence generation unit 43 , and the parameter acquisition unit 44 until the absolute value of the difference between the parameter η obtained this time and the parameter η obtained previously becomes equal to or smaller than a predetermined threshold.
[0478] <Acoustic Feature Extraction Unit 521>
[0479] The acoustic feature extraction unit 521 receives input of a time-domain audio signal in units of frames, which is a time-series signal.
[0480] The acoustic feature extraction unit 521 calculates an index representing the magnitude of the sound of the time-series signal as an acoustic feature (step FE2). The calculated index representing the magnitude of the sound is output to the determination unit 522. Furthermore, the acoustic feature extraction unit 521 generates an acoustic feature code corresponding to the acoustic feature and outputs it to the decoding device.
[0481] The index indicating the magnitude of the sound of the time series signal may be any index as long as it indicates the magnitude of the sound of the time series signal. The index indicating the magnitude of the sound of the time series signal is, for example, the energy of the time series signal.
[0482] In this example, since the determination unit 522 described below determines the structure of the encoding process based not only on the parameter η but also on the index indicating the size of the sound, the acoustic feature extraction unit 521 calculates the index indicating the size of the sound. However, if the determination unit 522 only uses the parameter η to determine the structure of the encoding process and does not use the index indicating the size of the sound, the acoustic feature extraction unit 521 does not need to calculate the index indicating the size of the sound.
[0483] <Determination Unit 522>
[0484] The determination unit 522 receives the parameter η determined by the parameter determination unit 27' and the index representing the magnitude of the sound of the time series signal calculated by the acoustic feature extraction unit 521. Furthermore, the sound signal is input as a frame unit of the time series signal as needed.
[0485] The determination unit 522 determines the structure of the encoding process based on at least the parameter n (step FE3), generates a determination code capable of determining the structure of the encoding process, and outputs it to the decoding device. In addition, the information about the structure of the encoding process determined by the determination unit 522 is output to the encoding unit 523.
[0486] The determination unit 522 may determine the configuration of the encoding process based only on the parameter η, or may determine the configuration of the encoding process based on the parameter η and other parameters.
[0487] The coding process structure may be a coding method such as TCX (Transform Coded Excitation) or ACELP (Algebraic Code Excited Linear Prediction), and may also include the frame length, which is the unit of time processing in a particular coding method, the number of bits allocated to the code, the order of coefficients that can be converted into linear prediction coefficients, and the values of any parameters used in the coding process. Specifically, the parameter η can be used to appropriately determine the frame length, which is the unit of time processing in a particular coding method, the number of bits allocated to the code, the order of coefficients that can be converted into linear prediction coefficients, and the values of any parameters used in the coding process.
[0488] In addition, refer to Figure 12 as well as Figure 13 The encoding device and method of the second embodiment described above determine the value of the parameter used in the encoding process based on the parameter η. Therefore, it can be said that referring to Figure 12 as well as Figure 13 The encoding device and method of the second embodiment described above are an example of a modified example of the second embodiment in which the structure of the encoding process is determined based on the parameter η.
[0489] The code that can identify the structure of the coding process can be any code as long as it can identify the structure of the coding process. For example, the code that can identify the structure of the coding process is a flag based on a predetermined bit string: as the structure of the coding process, it is "11" when TCX with a long frame length is determined, "100" when TCX with a short frame length is determined, "101" when ACELP is determined, and "0" when, for example, a coding process is determined to transmit only the noise level or to determine the iso-low bits. The code that can identify the structure of the coding process can be, for example, a parameter code representing the parameter n.
[0490] The determination code that can determine the configuration of the encoding process also determines the configuration of the corresponding decoding process if the configuration of the encoding process is determined by the determination code. Therefore, it can be said that the determination code can also determine the configuration of the decoding process.
[0491] Hereinafter, first, an example will be described in which the encoding process is determined based on the parameter η and an index representing the volume of the sound of the time series signal.
[0492] The determination unit 522 compares the index indicating the magnitude of the sound of the time series signal with a predetermined threshold value C e In addition, the parameter η is compared with the predetermined threshold C ηAs an indicator of the size of the sound of the time series signal, for example, when the average amplitude (the square root of the average energy per sample) is used, let C e = Maximum amplitude value * (1 / 128). For example, if the precision is 16 bits, the maximum amplitude value is 32768, so let C e =256. In addition, for example, let C η =1.
[0493] If the index representing the size of the sound of the time series signal is greater than or equal to the predetermined threshold C e And the parameter η<predetermined threshold C η , since the time series signal is likely to be music composed mainly of wind instruments or string instruments (hereinafter referred to as sustained music), the determination unit 522 determines to perform encoding suitable for sustained music. Encoding suitable for sustained music is, for example, TCX encoding with a long frame length, specifically, TCX encoding with a frame length of 1024 pixels.
[0494] If the index representing the size of the sound of the time series signal is greater than or equal to the predetermined threshold C e And the parameter η≥predetermined threshold C η , then the time series signal is likely to be music with speech or percussion instruments with large time variations as the main body.
[0495] At this time, the determination unit 522 divides the input time series signal into four subframes as needed, and measures the energy of the time series signal in each subframe. The determination unit 522 divides the arithmetic mean of the energy of the four subframes by the geometric mean to obtain the value F = ((1 / 4)Σ4 subframe energy) / ((Π subframe energy) 1 / 4 ) is the predetermined threshold C F If the above is true, the time series signal is likely to be music with large temporal variations. In this case, the determination unit 522 decides to perform encoding processing suitable for music with large temporal variations. The encoding processing suitable for music with large temporal variations is, for example, TCX encoding processing with a short frame length, specifically, TCX encoding processing with a frame length of 256 points. For example, let C E =1.5.
[0496] If the value F is less than the predetermined threshold C F , then the time series signal is likely to be speech. In this case, the determination unit 522 decides to perform a coding process suitable for speech. Coding processes suitable for speech are speech coding processes such as ACELP and CELP (Code Excited Linear Prediction).
[0497] If the index representing the size of the sound of the time series signal is less than the predetermined threshold value Ce And the parameter η≥predetermined threshold C η , then the time series signal is likely to be a silent interval. A silent interval does not mean a completely silent interval, but rather a period in which background sound or surrounding noise is present despite the absence of the target sound. In this case, determination unit 522 determines that the time series signal is a silent interval.
[0498] If the index representing the size of the sound of the time series signal is less than the predetermined threshold value C e And the parameter η<predetermined threshold C η , the time series signal is likely to be low-volume, continuous music, i.e., background music (hereinafter referred to as characteristic background music such as BGM). In this case, the determination unit 522 determines to perform encoding processing suitable for characteristic background music such as BGM. An encoding processing suitable for characteristic background music such as BGM is, for example, TCX encoding processing with a short frame length, specifically, TCX encoding processing with 256-bit frames.
[0499] Furthermore, the determination unit 522 may determine the configuration of the encoding process based not only on the parameter η but also on at least one of the temporal variation of the index representing the sound level of the time series signal, the spectral shape, the temporal variation of the spectral shape, and the degree of periodicity of the fundamental pitch. Furthermore, when using at least one of the temporal variation of the index representing the sound level of the time series signal, the spectral shape, the temporal variation of the spectral shape, and the degree of periodicity of the fundamental pitch, the acoustic feature extraction unit 521 calculates the acoustic feature used by the determination unit 522 from the temporal variation of the index representing the sound level of the time series signal, the spectral shape, the temporal variation of the spectral shape, and the degree of periodicity of the fundamental pitch, and outputs it to the determination unit 522. Furthermore, the acoustic feature extraction unit 521 generates an acoustic feature code corresponding to the calculated acoustic feature code and outputs it to the decoding device.
[0500] The following describes respectively (1) a case where the structure of the coding process is determined based on the temporal variation of the parameter η and an indicator representing the size of the sound of the time series signal; (2) a case where the structure of the coding process is determined based on the parameter η and the spectral shape of the time series signal; (3) a case where the structure of the coding process is determined based on the temporal variation of the parameter η and the spectral shape of the time series signal; and (4) a case where the structure of the coding process is determined based on the parameter η and the periodicity of the fundamental frequency of the time series signal.
[0501] (1) In the case of determining a structure for encoding processing based on a parameter η and temporal variation of an index representing the size of the sound of a time series signal, the determination unit 522 determines whether the temporal variation of the index representing the size of the sound of the time series signal is large, and further determines whether the parameter η is large.
[0502] Whether the temporal variation of the index representing the magnitude of the sound of the time series signal is large can be determined based on, for example, a predetermined threshold value C E That is, if the temporal variation of the index representing the size of the sound of the time series signal is greater than or equal to the predetermined threshold value C E ', it can be determined that the temporal variation of the index representing the size of the sound of the time series signal is large. Otherwise, it can be determined that the temporal variation of the index representing the size of the sound of the time series signal is small.
[0503] Whether the parameter η is large can be determined based on a predetermined threshold value C η That is, if the parameter η ≥ the predetermined threshold C η , then it can be determined that the parameter η is large, otherwise, it can be determined that the parameter η is small.
[0504] If the temporal variation of the index representing the size of the sound of the time series signal is large and the parameter is large, the time series signal is likely to be speech. In this case, the determination unit 522 decides to perform coding processing suitable for speech. For example, using the value F obtained by dividing the arithmetic mean of the energy of the four subframes constituting the time series signal by the geometric mean, F = ((1 / 4)Σ4 subframe energy) / ((Π subframe energy) 1 / 4 ), set C E '=1.5.
[0505] If the temporal variation of the index representing the volume of the time series signal is large and the parameter is small, the time series signal is likely to be music with large temporal variation. In this case, the determination unit 522 determines to perform encoding processing suitable for music with large temporal variation.
[0506] If the temporal variation of the index representing the volume of the time-series signal is small and the parameter η is large, the time-series signal is likely to be a silent interval. In this case, the determination unit 522 determines that the time-series signal is a silent interval.
[0507] If the temporal variation of the index representing the size of the sound of the time series signal is small and the parameter η is small, it is highly likely that the music is composed mainly of wind instruments or string instruments with sustained sounds. In this case, the determination unit 522 decides to perform encoding processing suitable for sustained music.
[0508] (2) When the configuration of the encoding process is determined based on the parameter η and the spectral shape of the time series signal, the determination unit 522 determines whether the spectral shape of the time series signal is flat and whether the parameter η is large.
[0509] Whether the spectrum shape of the time series signal is flat can be determined based on a predetermined threshold E V For example, if the absolute value of the first PARCOR coefficient corresponding to the timing signal is less than the predetermined threshold EV (For example, E V =0.7), it can be determined that the spectrum shape of the time series signal is flat; otherwise, it can be determined that the spectrum shape of the time series signal is not flat.
[0510] If the spectrum of the time series signal is flat and the parameter η is large, the time series signal is likely to be a silent interval. In this case, the determination unit 522 determines that the time series signal is a silent interval.
[0511] If the spectrum of the time series signal is flat and the parameter η is small, the time series signal is likely to be music with large temporal variations. In this case, the determination unit 522 decides to perform encoding processing suitable for music with large temporal variations.
[0512] If the spectral shape of the time series signal is not flat and the parameter η is large, the time series signal is likely to be speech. In this case, the determination unit 522 decides to perform encoding processing suitable for speech.
[0513] If the spectrum shape of the time series signal is not flat and the parameter η is small, it is highly likely that the signal is composed mainly of sustained sounds of wind instruments or string instruments. In this case, the determination unit 522 decides to perform encoding processing suitable for sustained music.
[0514] (3) When the configuration of the encoding process is determined based on the parameter η and the temporal variation of the spectral shape of the time series signal, the determination unit 522 determines whether the temporal variation of the spectral shape of the time series signal is large and whether the parameter η is large.
[0515] Whether the temporal variation of the spectrum shape of the time series signal is flat can be determined based on a predetermined threshold value E V For example, if the arithmetic mean of the absolute values of the first PARCOR coefficients of the four subframes constituting the time series signal is divided by the geometric mean, the value F V =((1 / 4)Σ4 subframes’ absolute value of the first PARCOR coefficient) / ((Π the absolute value of the first PARCOR coefficient) 1 / 4 ) is the predetermined threshold E V '(For example, E V '=1.2) or above, it can be determined that the temporal variation of the spectral shape of the time series signal is large. Otherwise, it can be determined that the temporal variation of the spectral shape of the time series signal is small.
[0516] If the temporal variation of the spectral shape of the time series signal is large and the parameter η is large, the time series signal is likely to be speech. In this case, the determination unit 522 decides to perform encoding processing suitable for speech.
[0517] If the temporal variation of the spectral shape of the time series signal is large and the parameter η is small, the time series signal is likely to be music with large temporal variation. In this case, the determination unit 522 decides to perform encoding processing suitable for music with large temporal variation.
[0518] If the temporal variation of the spectral shape of the time series signal is small and the parameter η is large, the time series signal is likely to be a silent interval. In this case, the determination unit 522 determines that the time series signal is a silent interval.
[0519] If the temporal variation of the spectral shape of the time series signal is small and the parameter η is small, the signal is likely to be music composed mainly of wind instruments or string instruments with sustained sounds. In this case, the determination unit 522 decides to perform encoding processing suitable for sustained music.
[0520] (4) When the configuration of the encoding process is determined based on the parameter η and the periodicity of the pitch of the time series signal, the determination unit 522 determines whether the periodicity of the pitch of the time series signal is large and whether the parameter η is large.
[0521] Whether the periodicity of the fundamental pitch of the time series signal is large can be determined based on a predetermined threshold value C P That is, if the periodicity of the fundamental tone of the time series signal is greater than or equal to the predetermined threshold C P If the periodicity of the fundamental pitch is large, then the periodicity of the fundamental pitch of the time series signal can be determined to be small. As the periodicity of the fundamental pitch, for example, the normalized correlation function of the sequence that deviates from the fundamental pitch period by τ samples is used.
[0522]
Number 30
[0523]
[0524] (where x(i) is the sample value of the time series and N is the number of samples in the frame), let C P =0.8.
[0525] If the periodicity of the fundamental pitch is large and the parameter η is large, the time series signal is likely to be speech. In this case, the determination unit 522 decides to perform encoding processing suitable for speech.
[0526] If the periodicity of the fundamental tone is large and the parameter η is small, the probability of the music being composed mainly of sustained sounds of wind instruments or string instruments is high. In this case, the determination unit 522 decides to perform encoding processing suitable for sustained music.
[0527] When the periodicity of the fundamental pitch is small and the parameter η is large, the time series signal is likely to be a silent interval. In this case, the determination unit 522 determines that the time series signal is a silent interval.
[0528] If the periodicity of the fundamental pitch is small and the parameter η is small, the time series signal is likely to be music with large temporal variations. In this case, the determination unit 522 decides to perform encoding processing suitable for music with large temporal variations.
[0529] <Encoding Unit 523>
[0530] The encoding unit 523 receives input of the audio signal in frame units as a time-series signal and information on the configuration of the encoding process determined by the determination unit 522 .
[0531] The coding unit 523 encodes the input time series signal through a coding process with a predetermined structure to generate a code (step FE4 ). The generated code is output to the decoding device.
[0532] When a coding process suitable for continuous music is determined, for example, TCX (Transform Coded Excitation) coding with a long frame length is performed, specifically, TCX coding of 1024-point frames is performed. In this case, a code representing a fixed value of η (e.g., η = 0.8) may be output to the decoding device as a parameter code, rather than a code representing the parameter η determined by the parameter determination unit 27'.
[0533] When an encoding process suitable for music with large temporal variations is determined, for example, a TCX encoding process with a short frame length, specifically, a TCX encoding process with a frame length of 256 points is performed.
[0534] When a coding process suitable for characteristic background sound such as BGM is determined, for example, TCX coding with a short frame length, specifically, TCX coding with 256 points, is performed. Furthermore, in this case, a code representing a fixed value of η (e.g., η = 0.8) may be output to the decoding device as a parameter code, rather than a code representing the parameter η determined by the parameter determination unit 27'.
[0535] When a coding process suitable for speech is determined, speech coding processes such as ACELP (Algebraic Code Excited Linear Prediction) and CELP (Code Excited Linear Prediction) are performed.
[0536] When it is determined that the time series signal is a silent interval, the encoding unit 523 does not encode the input time series signal, but performs processing using, for example, (i) the first method or (ii) the second method described below.
[0537] (i) First method
[0538] The encoder 523 transmits information indicating that the interval is silent to the decoding device. The information indicating that the interval is silent is transmitted as a low-order bit, such as 1 bit. After transmitting the information indicating that the interval is silent, the encoder 523 may not transmit the information indicating that the interval is silent again while the determination unit 522 determines that the time-series signal being processed is a silent interval.
[0539] (ii) Second method
[0540] The encoding unit 523 transmits information indicating whether it is a silent period, the shape of the spectrum envelope of the time series signal, and information on the amplitude of the time series signal to the decoding device.
[0541] (decoding)
[0542] An example of a decoding device and method will be described.
[0543] like Figure 19 As shown, the decoding device includes, for example, a determination code decoding unit 525, an acoustic feature code decoding unit 526, a determination unit 527, and a decoding unit 528. Each unit of the decoding device performs Figure 20 The decoding method is realized by performing the various processes illustrated in FIG.
[0544] The following describes each unit of the decoding device.
[0545] <Determination Code Decoding Unit 525>
[0546] The identification code output from the encoding device is input to the identification code decoding unit 525 .
[0547] The identification code decoding unit 525 decodes the identification code and acquires information on the configuration of the encoding process (step FD1 ). The acquired information on the configuration of the encoding process is output to the identification unit 527 .
[0548] When the determination code is a parameter code, the determination code decoding unit 525 decodes the parameter code to obtain a parameter η, and outputs the obtained parameter η to the determination unit 527 as information on the configuration of the encoding process.
[0549] <Acoustic Feature Code Decoding Unit 526>
[0550] The acoustic feature code decoder 526 receives input of the acoustic feature code output from the encoding device.
[0551] The acoustic feature code decoding unit 526 decodes the acoustic feature code to obtain an acoustic feature, which is at least one of an indicator indicating the sound level of the time series signal, a temporal variation in the indicator indicating the sound level, a spectral shape, a temporal variation in the spectral shape, and a degree of periodicity of the fundamental pitch (step FD2). The obtained acoustic feature is output to the determination unit 527.
[0552] On the encoding side, the configuration of the encoding process is determined based only on the parameter η, and when the acoustic feature value and the acoustic feature value code are not generated, the acoustic feature value code decoding unit 526 does not perform any processing.
[0553] <Determination Unit 527>
[0554] The determination unit 527 receives input of information on the configuration of the encoding process obtained by the determination code decoding unit 525. The determination unit 527 also receives input of the acoustic feature quantity obtained by the acoustic feature quantity code decoding unit 526 as needed.
[0555] The determination unit 527 determines the structure of the decoding process based on the information regarding the encoding process structure (step FD3). For example, the determination unit 527 determines the structure of the decoding process corresponding to the encoding process structure determined by the information regarding the encoding process structure. Alternatively, the determination unit 527 may determine the structure of the decoding process based on the information regarding the encoding process structure and the acoustic feature values, as needed. The determined information regarding the decoding process structure is output to the decoding unit 528.
[0556] The following describes an example in which the parameter η is input as information about the structure of the encoding process, and at least one of the following acoustic feature quantities is input: an indicator representing the sound volume of a time series signal, a temporal variation in the indicator representing the sound volume, a spectral shape, a temporal variation in the spectral shape, and a degree of periodicity of the fundamental frequency.
[0557] In this case, it is assumed that a judgment criterion similar to the judgment criterion for specifying the configuration of the encoding process performed by the determination unit 522 of the encoding device is previously determined for the determination unit 527 of the decoding device. Based on this judgment criterion, the determination unit 527 uses the parameter η and the acoustic feature amount to determine the configuration of the decoding process corresponding to the configuration of the encoding process determined by the determination unit 522.
[0558] The determination criteria for the specific configuration of the encoding process performed by the determination unit 522 of the encoding device have been described in (Encoding), and therefore, repeated description will be omitted here.
[0559] For example, as the decoding processing structure, one of a decoding process suitable for continuous music, a decoding process suitable for music with large temporal variations, a decoding process suitable for characteristic background sounds such as background music, and a decoding process suitable for speech is determined. Alternatively, the determination unit 527 determines that the time series signal is a silent interval.
[0560] <Decoding Unit 528>
[0561] The decoding unit 528 receives input of the code output by the encoding device and the information on the configuration of the decoding process determined by the determination unit 527 .
[0562] The decoding unit 528 obtains a frame-based audio signal as a time-series signal through decoding processing with a predetermined structure (step FD4 ).
[0563] When a decoding process suitable for continuous music is determined, for example, a TCX (Transform Coded Excitation) decoding process having a long frame length is performed, specifically, a TCX decoding process for a frame having 1024 dots is performed.
[0564] When a decoding process suitable for music with large temporal variations is determined, for example, a TCX decoding process with a short frame length, specifically, a TCX decoding process with a frame length of 256 dots, is performed.
[0565] When a decoding process suitable for characteristic background sound such as BGM is determined, for example, a TCX decoding process with a short frame length, specifically, a TCX decoding process with a 256-point frame length is performed.
[0566] When a decoding process suitable for speech is determined, speech decoding processes such as ACELP (Algebraic Code Excited Linear Prediction) and CELP (Code Excited Linear Prediction) are performed.
[0567] When the decoding device receives information indicating a silent interval or when the determination unit 527 determines that the time series signal is a silent interval, the decoding unit 528 performs processing using, for example, (i) the first method or (ii) the second method described below.
[0568] (i) First method
[0569] This corresponds to (i) the first method on the encoding side.
[0570] The decoding unit 528 generates predetermined noise.
[0571] (ii) Second method
[0572] Decoding unit 528 uses the information about the shape of the spectral envelope of the time series signal and the amplitude of the time series signal, received along with the information indicating the silent interval, to transform the noise into a predetermined noise and output it. The noise transformation method can be based on existing methods used in EVS (Enhanced Voice Service) and other services.
[0573] In this way, the decoding unit 528 can generate noise when acquiring information indicating that it is a silent period.
[0574] [Modifications, etc.]
[0575] If the linear prediction analysis unit 22 and the non-smoothed amplitude spectrum envelope sequence generation unit 23 are considered as a single spectrum envelope estimation unit 2A, then the spectrum envelope estimation unit 2A estimates the spectrum envelope (non-smoothed amplitude spectrum envelope sequence) by using the nth power of the absolute value of the frequency domain sample sequence corresponding to the time series signal, for example, the MDCT coefficient sequence, as the power spectrum. Here, "using the nth power" means using the spectrum of the nth power when the power spectrum is generally used.
[0576] In this case, it can be said that the linear prediction analysis unit 22 of the spectrum envelope estimation unit 2A performs linear prediction analysis using, for example, a pseudo-correlation function signal sequence obtained by inverse Fourier transforming the absolute values of the frequency domain sample sequence, which is the MDCT coefficient sequence, raised to the power of n, as the power spectrum, thereby obtaining coefficients that can be converted into linear prediction coefficients. Furthermore, it can be said that the unsmoothed amplitude spectrum envelope sequence generation unit 23 of the spectrum envelope estimation unit 2A estimates the spectrum envelope by obtaining a sequence, i.e., an unsmoothed spectrum envelope sequence, obtained by raising the sequence of amplitude spectrum envelopes corresponding to the coefficients that can be converted into linear prediction coefficients, obtained by the linear prediction analysis unit 22, to the power of 1 / n.
[0577] In addition, if the smoothed amplitude spectrum envelope sequence generation unit 24, the envelope normalization unit 25 and the encoding unit 26 are understood as a single encoding unit 2B, it can be said that the encoding unit 2B performs encoding on each coefficient of the frequency domain sample string, such as the MDCT coefficient string corresponding to the time series signal, in such a manner that the bit allocation is changed or the bit allocation is substantially changed based on the spectrum envelope (non-smoothed amplitude spectrum envelope sequence) estimated by the spectrum envelope estimation unit 2A.
[0578] If the decoding unit 34 and the envelope denormalization unit 35 are understood as a decoding unit 3A, it can be said that the decoding unit 3A decodes the input integer signal code according to the bit allocation that changes based on the non-smoothed spectrum envelope sequence or the bit allocation that actually changes, thereby obtaining a frequency domain sample string corresponding to the time series signal.
[0579] If the encoder 2B performs encoding that changes the bit allocation based on the spectral envelope (non-smoothed amplitude-spectral envelope sequence) or substantially changes the bit allocation, encoding processing other than the arithmetic coding described above can be performed. In this case, the decoder 3A performs decoding processing corresponding to the encoding processing performed by the encoder 2B.
[0580] For example, the encoder 2B may perform Golomb-Rice encoding on the frequency domain sample sequence using Rice parameters determined based on the spectral envelope (unsmoothed amplitude-spectral envelope sequence). In this case, the decoder 3A may perform Golomb-Rice decoding using Rice parameters determined based on the spectral envelope (unsmoothed amplitude-spectral envelope sequence).
[0581] In the first embodiment, the encoding device does not need to complete the encoding process when determining parameter η. In other words, parameter determination unit 27 can determine parameter η based on the estimated code quantity. In this case, encoding unit 2B uses each of multiple parameters η to obtain an estimated code quantity for a frequency domain sample sequence corresponding to the time series signal of the same predetermined time interval, obtained through the same encoding process as described above. Based on the obtained estimated code quantities, parameter determination unit 27 selects one of the multiple parameters η. For example, it selects the parameter η with the smallest estimated code quantity. Encoding unit 2B performs the same encoding process as described above using the selected parameter η, obtains a code, and outputs it.
[0582] The encoding device may also include Figure 4 or Figure 12 The segmentation unit 28 is shown as a dotted line in the figure. Based on the frequency domain sample string, for example, the MDCT coefficient string generated by the frequency domain conversion unit 21, the segmentation unit 28 generates a first frequency domain sample string consisting of samples corresponding to the periodic component of the frequency domain sample string and a second frequency domain sample string consisting of samples other than the samples corresponding to the periodic component of the frequency domain sample string, and outputs information indicating the samples corresponding to the periodic component as auxiliary information to the decoding device.
[0583] In other words, the first frequency domain sample string is a sample string consisting of samples corresponding to the mountain portion of the frequency domain sample string, and the second frequency domain sample string is a sample string consisting of samples corresponding to the valley portion of the frequency domain sample string.
[0584] For example, a sample string consisting of one or more consecutive samples including samples corresponding to the periodicity or fundamental frequency of the time series signal corresponding to the frequency domain sample string in the frequency domain sample string, and all or part of one or more consecutive samples including samples corresponding to integer multiples of the periodicity or fundamental frequency of the time series signal corresponding to the frequency domain sample string in the frequency domain sample string is generated as a first frequency domain sample string, and a sample string consisting of samples not included in the first frequency domain sample string in the frequency domain sample string is generated as a second frequency domain sample string. The first and second frequency domain sample strings can be generated using the method described in International Publication No. WO2012 / 046685.
[0585] The linear prediction analysis unit 22, the unsmoothed amplitude spectrum envelope sequence generation unit 23, the smoothed amplitude spectrum envelope sequence generation unit 24, the envelope normalization unit 25, the encoding unit 26, and the parameter determination unit 27 perform the encoding process described in the first embodiment or the second embodiment on each of the first frequency domain sample sequence and the second frequency domain sample sequence to generate a code. Specifically, for example, when arithmetic coding is performed, a parameter code, a linear prediction coefficient code, an integer signal code, and a gain code are generated for the first frequency domain sample sequence, and a parameter code, a linear prediction coefficient code, an integer signal code, and a gain code are generated for the second frequency domain sample sequence.
[0586] In this way, by encoding each of the first frequency domain sample string and the second frequency domain sample string, encoding can be performed more efficiently.
[0587] In this case, the decoding device may further include Figure 9 The combining unit 38 is shown by a dotted line in FIG. The decoding device performs the decoding process described in the first embodiment or the second embodiment based on the code corresponding to the first frequency domain sample string (for example, a parameter code, a linear prediction coefficient code, an integer signal code, and a gain code) to obtain a decoded first frequency domain sample string. In addition, the decoding device performs the decoding process described in the first embodiment or the second embodiment based on the code corresponding to the second frequency domain sample string (for example, a parameter code, a linear prediction coefficient code, an integer signal code, and a gain code) to obtain a decoded second frequency domain sample string. The combining unit 38 uses the input auxiliary information to appropriately combine the decoded first frequency domain sample string and the decoded second frequency domain sample string to obtain, for example, a decoded frequency domain sample string such as a decoded MDCT coefficient string ^X(0), ^X(1), ..., ^X(N-1). The time domain conversion unit converts the decoded frequency domain sample string into the time domain to obtain a time series signal. Combining using auxiliary information can be performed using the method described in International Publication WO2012 / 046685.
[0588] In addition, when the bit rate is low or when the amount of code needs to be further reduced, only the first frequency domain sample string can be encoded in the encoding device, and only the code corresponding to the first frequency domain sample string can be generated, and the code corresponding to the second frequency domain sample string cannot be generated. In the decoding device, the first frequency domain sample string obtained from the code and the second frequency domain sample string with the sample value set to 0 can be used to obtain the decoded frequency domain sample string.
[0589] Furthermore, the linear prediction analysis unit 22, the unsmoothed amplitude spectrum envelope sequence generation unit 23, the smoothed amplitude spectrum envelope sequence generation unit 24, the envelope normalization unit 25, the encoding unit 26, and the parameter determination unit 27 can perform the encoding process described in the first embodiment or the second embodiment on the sample sequence obtained by combining the first frequency domain sample sequence and the second frequency domain sample sequence, thereby generating a code. For example, when arithmetic coding is performed, a parameter code, a linear prediction coefficient code, an integer signal code, and a gain code corresponding to the sample sequence are generated.
[0590] By encoding the entire column of sample strings in this way, encoding can be performed more efficiently.
[0591] At this time, the decoding device performs the decoding process described in the first or second embodiment to obtain a decoded, aligned sample sequence. Using the input side information, the decoded, aligned sample sequence is aligned according to a rule corresponding to the rule used by the encoding device to generate the first and second frequency domain sample sequences. For example, a decoded frequency domain sample sequence is obtained as the decoded MDCT coefficient sequence ^X(0), ^X(1), ..., ^X(N-1). The time domain converter 36 converts the decoded frequency domain sample sequence into the time domain to obtain a time series signal. Alignment using side information can be performed using the method described in International Publication WO2012 / 046685.
[0592] Furthermore, the encoding device may select any of the following methods for each frame: (1) a method of encoding a frequency domain sample string to generate a code; (2) a method of encoding each of the first frequency domain sample string and the second frequency domain sample string to generate a code; (3) a method of encoding only the first frequency domain sample string to generate a code; (4) a method of encoding a sample string obtained by combining the first frequency domain sample string and the second frequency domain sample string, i.e., a sequenced sample string, to generate a code. In this case, the encoding device also outputs a code indicating which of the methods (1) to (4) has been selected, and the decoding device performs decoding processing corresponding to any of the above methods based on the code input for each frame.
[0593] Furthermore, the parameter determination unit 27 of the encoding device and the parameter decoding unit 37 of the decoding device may store candidates for the parameter η corresponding to each of the above-mentioned methods (1) to (4). Similarly, the linear prediction analysis unit 22 of the encoding device and the linear prediction coefficient decoding unit 31 of the decoding device may store candidates for the quantized linear prediction coefficient and candidates for the decoded linear prediction coefficient corresponding to each of the above-mentioned methods (1) to (4).
[0594] The non-smoothed amplitude spectrum envelope sequence generation unit 23 and the non-smoothed amplitude spectrum envelope sequence generation unit 422 can, for example, transform the spectrum envelope sequence (non-smoothed amplitude spectrum envelope sequence) based on the periodic component of the frequency domain sample sequence, which is the MDCT coefficient sequence ^X(0), ^X(1), ..., ^X(N-1), to generate a periodic synthetic envelope sequence. Similarly, the non-smoothed amplitude spectrum envelope sequence generation unit 32 can, for example, transform the spectrum envelope sequence (non-smoothed amplitude spectrum envelope sequence) based on the periodic component of the decoded frequency domain sample sequence, which is the decoded MDCT coefficient sequence ^X(0), ^X(1), ..., ^X(N-1), to generate a periodic synthetic envelope sequence. In this case, the variance parameter determination unit 268, the decoding unit 34, and the whitening spectrum sequence generation unit 43 of the encoding unit 26 use the periodic synthetic envelope sequence instead of the spectrum envelope sequence (non-smoothed amplitude spectrum envelope sequence) and perform the same processing as described above. Since the periodic synthetic envelope sequence has good approximation accuracy near the peak value caused by the pitch period of the time series signal, the coding efficiency can be improved by using the periodic synthetic envelope sequence.
[0595] For example, a sequence in which the values of at least integer multiples of the period of the frequency domain sample sequence and samples near integer multiples of the period in the spectrum envelope sequence change more significantly as the period of the frequency domain sample sequence increases is defined as a periodic synthetic envelope sequence. Alternatively, a sequence in which the values of at least integer multiples of the period of the frequency domain sample sequence and samples near integer multiples of the period in the spectrum envelope sequence change more significantly as the degree of periodicity of the time series signal increases is defined as a periodic synthetic envelope sequence. Alternatively, a sequence in which the values of many samples near integer multiples of the period of the frequency domain sample sequence in the spectrum envelope sequence change more significantly as the period of the frequency domain sample sequence increases is defined as a periodic synthetic envelope sequence.
[0596] Furthermore, N and U can be set as positive integers, T can be set as the interval of the periodic component of the frequency domain sample string, L can be set as the number of digits below the decimal point of the interval T, v can be set as an integer greater than 1, floor(·) can be set as a function that discards the digits below the decimal point and returns an integer value, Round(·) can be set as a function that rounds the first decimal place and returns an integer value, and T' can be set as T×2 L, let ^H[0],…,^H[N-1] be the spectrum envelope sequence, let δ be the value that determines the mixing ratio of the spectrum envelope ^H[n] and the periodic envelope P[k], about (U×T') / 2 L -v-1≦k≦(U×T') / 2 L An integer k in the range of +v-1, such as
[0597]
Number 31
[0598] or
[0599]
[0600] in,
[0601] h=2.8·(1.125-exp(-0.07·T′ / 2 L )),
[0602] PD=0.5·(2.6-exp(-0.05·T′ / 2 L ))
[0603] The periodic envelope sequence P[1], ..., P[N] is obtained in this way, and the periodic comprehensive envelope sequence defined by the following formula is obtained using the obtained periodic envelope sequence P[1], ..., P[N]. M [1],…,^H M [N]h and PD may be predetermined values other than the above examples.
[0604]
Number 32
[0605]
[0606] The value δ, which determines the mixing ratio between the spectral envelope ^H[n] and the periodic envelope P[k], can be predetermined in the encoding and decoding devices, or a code representing the information about δ determined by the encoding device can be generated and output to the decoding device. In the latter case, the decoding device determines δ by decoding the input code representing δ. Using the determined δ, the unsmoothed amplitude spectrum envelope sequence generator 32 of the decoding device can determine a periodic integrated envelope sequence identical to the periodic integrated envelope sequence generated by the encoding device.
[0607] If Figure 12 If the spectrum envelope estimation unit 2A, encoding unit 2B, frequency domain conversion unit 21 and segmentation unit 28 are understood as a coding unit 2C, it can be said that the coding unit 2C encodes the time series signal of each predetermined time interval through a coding process of a structure determined at least based on the parameter η of each predetermined time interval.
[0608] In addition, if Figure 17 If the sound feature extraction unit 521, determination unit 522 and encoding unit 523 are understood as a coding unit 2D, it can be said that the coding unit 2D encodes the time series signal of each predetermined time interval through a coding process of a structure determined at least based on the parameter η of each predetermined time interval.
[0609] In this way, it can be considered that the encoding unit 2C and the encoding unit 2D perform the same processing.
[0610] The processes described above are not only executed in time series in the order described but may also be executed in parallel or individually according to the processing capability of the device executing the processes or as needed.
[0611] Furthermore, various processes in each method or device can be implemented by a computer. In this case, the processing content of each method or device is described in a program. Then, by executing the program on a computer, various processes in each method or device are implemented on the computer.
[0612] The program describing the processing contents can be recorded on a computer-readable recording medium. Examples of the computer-readable recording medium include a magnetic recording device, an optical disc, a magneto-optical recording medium, and a semiconductor memory.
[0613] The program may be distributed, for example, by selling, transferring, or lending a removable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be distributed by storing the program in a storage device of a server computer and transferring the program from the server computer to other computers via a network.
[0614] A computer that executes such a program, for example, first temporarily stores a program recorded in a removable recording medium or a program forwarded from a server computer in its own storage device. Furthermore, when executing a process, the computer reads the program stored in its own recording device and executes the process according to the read program. Furthermore, as another embodiment of the program, the computer can directly read the program from the removable recording medium and execute the process according to the program. Furthermore, each time a program is forwarded from a server computer to the computer, the computer can sequentially execute the process according to the received program. Furthermore, a structure can also be provided in which the program is not forwarded from the server computer to the computer, and the above-mentioned process is executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only based on its execution instructions and results. Furthermore, the program includes information for the processing of the electronic computer and data that complies with the program (data that is not a direct instruction to the computer but has the nature of specifying the computer's processing, etc.).
[0615] Furthermore, although each device is configured by executing a predetermined program on a computer, at least a part of the processing contents may be realized by hardware.
Claims
1. A decoding device, which obtains a frequency domain sample string corresponding to a time series signal by decoding in the frequency domain, wherein: The decoding device comprises: The parameter code decoding unit obtains a candidate for the decoding parameter η corresponding to the input parameter code from a plurality of candidates for the decoding parameter η as the decoding parameter η; A linear prediction coefficient decoding unit decodes the input linear prediction coefficient code to obtain a decoded linear prediction coefficient; A non-smoothed spectrum envelope sequence generating unit, using the decoding parameter η obtained above to obtain a non-smoothed spectrum envelope sequence through equation A2; and The decoding unit decodes the input integer signal code according to the bit allocation that is changed or substantially changed based on the non-smoothed spectrum envelope sequence, thereby obtaining a frequency domain sample string corresponding to the time series signal. N is a positive integer, · is a real number, exp(·) is an exponential function with Napier number as base, j is an imaginary unit, ^β1,^β2,…,^β p is the above-mentioned decoded linear prediction coefficient, ^H(0), ^H(1), ..., ^H(N-1) is the above-mentioned non-smoothed spectrum envelope sequence, p is an integer greater than 2, and formula A2 is the following formula 2. A decoding method, wherein a frequency domain sample string corresponding to a time series signal is obtained by decoding in the frequency domain, wherein: The decoding method comprises: a parameter code decoding step of obtaining, from a plurality of candidates for decoding parameter η, a candidate for decoding parameter η corresponding to the input parameter code as decoding parameter η; a linear prediction coefficient decoding step, decoding the input linear prediction coefficient code to obtain a decoded linear prediction coefficient; a step of generating a non-smoothed spectrum envelope sequence, using the decoding parameter η obtained above to obtain the non-smoothed spectrum envelope sequence through formula A2; and A decoding step of decoding the input integer signal code according to the bit allocation that is changed or substantially changed based on the non-smoothed spectrum envelope sequence, thereby obtaining a frequency domain sample string corresponding to the time series signal. N is a positive integer, · is a real number, exp(·) is an exponential function with Napier number as base, j is an imaginary unit, ^β1,^β2,…,^β p is the above-mentioned decoded linear prediction coefficient, ^H(0), ^H(1), ..., ^H(N-1) is the above-mentioned non-smoothed spectrum envelope sequence, p is an integer greater than 2, and formula A2 is the following formula 3. A computer-readable recording medium recording a program for causing a computer to function as each unit of the decoding device according to claim 1. 4 . A computer program product comprising a computer program for causing a computer to function as each unit of the decoding device according to claim 1 .
Citation Information
Patent Citations
Coding method, decoding method, coding device, decoding device, program, and recording medium
WO2012046685A1
Coding method, coding device, program, and recording medium
WO2014054556A1
Intensified audio-frequency coding-decoding device and method
CN1677493A
Audio signal classifier
US20110016077A1