Coding device, coding method, decoding device, decoding method, and program

By selecting Hoffman encoding or arithmetic encoding according to the change in the transform window length in the encoding device, the problem of low encoding efficiency in the prior art is solved, and higher encoding efficiency and lower decoding complexity are achieved, and efficient encoding of various sound materials is suitable.

CN112400203BActive Publication Date: 2025-08-19SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980039838.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-21
Filing Date
2019-06-07
Publication Date
2025-08-19
Estimated Expiration
2039-06-07

AI Technical Summary

Technical Problem

The existing encoding technology is difficult to achieve higher encoding efficiency and higher compression efficiency when sending a variety of sound materials, especially in 7.1 surround sound reproduction or enhanced rendering reproduction of 3D audio, requiring higher speed audio channel decoding.

Method used

By using the transform window length change in the encoding device, Hoffman encoding and arithmetic encoding are selectively performed, the spectrum information is encoded, and a suitable encoding scheme is selected according to the change in the transform window length, such as using Hoffman encoding when the transform window length becomes smaller and larger, and in other cases arithmetic encoding is used.

Benefits of technology

Improves encoding efficiency, reduces the computational complexity in the decoding process, and achieves higher compression efficiency in different types of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112400203B_ABST
    Figure CN112400203B_ABST
Patent Text Reader

Abstract

The present technology relates to an encoder and encoding method, a decoder and decoding method, and a program that can improve coding efficiency. The encoder includes a time-frequency conversion unit for performing time-frequency conversion on an audio signal using a conversion window; and an encoding unit for performing Huffman coding on spectral information obtained through time-frequency conversion when the conversion window length switches from a short conversion window length to a long conversion window length, and performing arithmetic coding on the spectral information when the conversion window length does not switch from a short conversion window length to a long conversion window length. The present technology can be applied to both the encoder and the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an encoding device, an encoding method, a decoding device, a decoding method, and a program, and particularly to an encoding device, an encoding method, a decoding device, a decoding method, and a program capable of improving encoding efficiency. Background Art

[0002] As methods for encoding audio signals, known methods include encoding according to the MPEG (Moving Picture Experts Group)-2 AAC (Advanced Audio Coding) standard, the MPEG-4 AAC standard, the MPEG-D USAC (Uniform Speech and Audio Coding) standard, and the MPEG-H 3D Audio standard using the MPEG-D USAC standard as a core encoder, etc., which are international standards (e.g., refer to NPLs 1 and 2).

[0003] [Citation List]

[0004] [Non-patent literature]

[0005] [NPL 1]

[0006] INTERNATIONAL STANDARD ISO / IEC 14496-3Fourth edition2009-09-01Information technology-coding of audio-visual objects-part3:Audio

[0007] [NPL 2]

[0008] INTERNATIONAL STANDARD ISO / IEC 23003-3Frist edition2012-04-01Information technology-coding of audio-visual objects-part3:Unified speechand audio coding Summary of the Invention

[0009] [Technical Issues]

[0010] At the same time, in order to transmit a variety of sound materials (objects) achieved through reproduction with enhanced presentation compared to conventional 7.1 surround sound reproduction or "3D audio," it is necessary to use an encoding technology that can decode more audio channels at a higher speed and with higher compression efficiency. In other words, it is necessary to improve encoding efficiency.

[0011] The present technology has been achieved in view of this situation, and an object of the present technology is to achieve improvement in encoding efficiency.

[0012] [Solution to the problem]

[0013] An encoding device according to a first aspect of the present technology includes: a time-frequency transform section that performs time-frequency transform on an audio signal using a transform window; and an encoding section that performs Huffman encoding on spectrum information obtained by time-frequency transform in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length, and performs arithmetic encoding on the spectrum information in a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length.

[0014] The encoding method or program according to the first aspect of the present technology includes: performing time-frequency transform on an audio signal using a transform window; performing Huffman coding on spectral information obtained by the time-frequency transform in a case where the transform window length of the transform window is changed from a small transform window length to a large transform window length; and performing arithmetic coding on the spectral information in a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length.

[0015] According to a first aspect of the present technology, time-frequency transform is performed on an audio signal using a transform window; in a case where the transform window length of the transform window is changed from a small transform window length to a large transform window length, Huffman coding is performed on spectral information obtained by the time-frequency transform; and in a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length, arithmetic coding is performed on the spectral information.

[0016] According to the second aspect of the present technology, a decoding device includes: a demultiplexing unit that demultiplexes a coded bit stream and extracts transform window information indicating the type of transform window used in the time-frequency transform of an audio signal and coded data about spectrum information obtained by the time-frequency transform from the coded bit stream; and a decoding unit that decodes the coded data by a decoding scheme corresponding to Huffman coding, in a case where the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length.

[0017] The decoding method or program according to the second aspect of the present technology includes the following steps: demultiplexing a coded bit stream, and extracting transform window information indicating the type of transform window used in the time-frequency transform of an audio signal and encoded data about spectrum information obtained by the time-frequency transform from the coded bit stream; and in a case where the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length, decoding the coded data by a decoding scheme corresponding to Huffman coding.

[0018] According to a second aspect of the present technology, a coded bit stream is demultiplexed, transform window information indicating the type of transform window used in the time-frequency transform of an audio signal and coded data about spectrum information obtained by the time-frequency transform are extracted from the coded bit stream; and in a case where the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the coded data is decoded by a decoding scheme corresponding to Huffman coding.

[0019] [Advantageous Effects of the Invention]

[0020] According to the first and second aspects of the present technology, encoding efficiency can be improved.

[0021] It should be noted that the advantages are not always limited to those described herein, but may be any advantageous effects described in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] [ Figure 1 ]

[0023] Figure 1 This is an explanatory diagram of MPEG-4 AAC encoding.

[0024] [ Figure 2 ]

[0025] Figure 2 This is an explanatory diagram of the types of conversion windows in MPEG-4 AAC.

[0026] [ Figure 3 ]

[0027] Figure 3 This diagram explains MPEG-D USAC encoding.

[0028] [ Figure 4 ]

[0029] Figure 4 This is an explanatory diagram of the types of transformation windows in MPEG-D USAC.

[0030] [ Figure 5 ]

[0031] Figure 5 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0032] [ Figure 6 ]

[0033] Figure 6 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0034] [ Figure 7 ]

[0035] Figure 7 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0036] [ Figure 8 ]

[0037] Figure 8 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0038] [ Figure 9 ]

[0039] Figure 9 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0040] [ Figure 10 ]

[0041] Figure 10 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0042] [ Figure 11 ]

[0043] Figure 11 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0044] [ Figure 12 ]

[0045] Figure 12 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0046] [ Figure 13 ]

[0047] Figure 13 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0048] [ Figure 14 ]

[0049] Figure 14 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0050] [ Figure 15 ]

[0051] Figure 15 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0052] [ Figure 16 ]

[0053] Figure 16 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0054] [ Figure 17 ]

[0055] Figure 17 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0056] [ Figure 18 ]

[0057] Figure 18 This diagram illustrates the coding efficiency of Huffman coding and arithmetic coding.

[0058] [ Figure 19 ]

[0059] Figure 19 is a diagram describing an embodiment of a configuration of an encoding device.

[0060] [ Figure 20 ]

[0061] Figure 20 is a flowchart showing the encoding process.

[0062] [ Figure 21 ]

[0063] Figure 21 is a diagram for describing an embodiment of a configuration of a decoding device.

[0064] [ Figure 22 ]

[0065] Figure 22 is a flowchart showing the decoding process.

[0066] [ Figure 23 ]

[0067] Figure 23 This is an explanatory diagram of the encoding efficiency according to this technology.

[0068] [ Figure 24 ]

[0069] Figure 24 This is an explanatory diagram of the encoding efficiency according to this technology.

[0070] [ Figure 25 ]

[0071] Figure 25 is a diagram describing an embodiment of the syntax of a channel stream.

[0072] [ Figure 26 ]

[0073] Figure 26 is a diagram describing an embodiment of the syntax of ics_info.

[0074] [ Figure 27 ]

[0075] Figure 27 is a flowchart showing the encoding process.

[0076] [ Figure 28 ]

[0077] Figure 28 is a flowchart showing the decoding process.

[0078] [ Figure 29 ]

[0079] Figure 29 is a flowchart showing the encoding process.

[0080] [ Figure 30 ]

[0081] Figure 30 is a diagram describing an example of a configuration of a computer. DETAILED DESCRIPTION

[0082] Embodiments to which the present technology is applied will be described below with reference to the drawings.

[0083] <First embodiment>

[0084] <This technology>

[0085] First, an overview of the present technology will be described. Although the signal to be encoded may be any type of signal such as an audio signal and an image signal, the present technology will be described below by taking, for example, a case where the object to be encoded is an audio signal.

[0086] For example, Figure 1 As described in , in MPEG-4 AAC, an audio signal is encoded.

[0087] In other words, when the encoding (decoding) process starts, first, time-frequency transform is performed on the audio signal using MDCT (Modified Discrete Cosine Transform).

[0088] Next, the MDCT coefficient (ie, spectrum information obtained through MDCT) is quantized for each scale factor band, and the quantized MDCT coefficient is obtained as a quantization result.

[0089] Here, the scale factor band refers to a band obtained by combining a plurality of sub-bands having a predetermined bandwidth, that is, an analysis power of a QMF (Quadrature Mirror Filter) analysis filter.

[0090] When quantized MDCT coefficients are obtained through quantization, Huffman coding is applied to each section in which the quantized NDCT coefficients and Huffman codebook information are encoded using the same Huffman codebook. In other words, Huffman coding is performed. Note that a section refers to a band obtained by combining multiple scale factor bands.

[0091] The Huffman code is output as coded data on the audio signal, that is, the Huffman-coded quantized MDCT coefficients and Huffman codebook information obtained as described above.

[0092] Furthermore, it is known that, in time-frequency transform, selecting an appropriate transform window according to the characteristics of a normally processed audio signal can enable compression of the audio signal with higher sound quality than using a single transform window.

[0093] For example, it is known that a transform window with a small transform window length is suitable for a music signal with a strong attack characteristic accompanied by sudden time changes (aggressive music signal), and a transform window with a large transform window length is suitable for a music signal with a strong stationary characteristic not accompanied by sudden time changes (stationary music signal).

[0094] Specifically, in MPEG4 AAC, for example, Figure 2 As described in , MDCT is performed while appropriately changing to a suitable window sequence among four window sequences.

[0095] exist Figure 2 In

[0045] , "window_sequence" represents a window sequence. Here, the window sequence represents the type of the transformation window, that is, the window type.

[0096] Specifically, in MPEG4 AAC, one type may be selected as a window sequence (ie, window type) from among four types of transform windows, namely, ONLY_LONG_SEQUENCE, LONG_START_SEQUENCE, EIGHT_SHORT_SEQUENCE, and LONG_STOP_SEQUENCE.

[0097] Moreover, in Figure 2 In the example, "num_windows" indicates the number of transform windows used when performing MDCT using a transform window of each window type, and the shape of the transform window is shown in each "look" box. Specifically, in each "look" box, the horizontal direction indicates Figure 2 The time direction in , and the vertical direction represents the size of the transform window at each sampling position, that is, Figure 2 The size of the coefficient by which each sample in is multiplied.

[0098] In MPEG4 AAC, when MDCT is performed on an audio signal, ONLY_LONG_SEQUENCE is selected for a frame having a strong stationary characteristic. The transform window represented by this ONLY_LONG_SEQUENCE is a transform window having a transform window length of 2048 samples.

[0099] Furthermore, for a frame having a strong aggressive characteristic, EIGHT_SHORT_SEQUENCE is selected. The transform window represented by this EIGHT_SHORT_SEQUENCE is eight transform windows divided in the time direction, and the transform window length of each divided transform window is 256 samples.

[0100] The transform window represented by EIGHT_SHORT_SEQUENCE is smaller in transform window length than other transform windows such as the transform window represented by LONG_STOP_SEQUENCE.

[0101] For a frame in which the window_sequence is transformed from ONLY_LONG_SEQUENCE to EIGHT_SHORT_SEQUENCE, LONG_START_SEQUENCE is selected. The transform window indicated by this LONG_START_SEQUENCE is a transform window having a transform window length of 2048 samples.

[0102] For frames where the window_sequence transitions from EIGHT_SHORT_SEQUENCE to ONLY_LONG_SEQUENCE, LONG_STOP_SEQUENCE is selected.

[0103] In other words, when the transform window length of the transform window changes from a small transform window length to a large transform window length, LONG_STOP_SEQUENCE is selected. The transform window represented by LONG_STOP_SEQUENCE is a transform window having a transform window length of 2048 samples.

[0104] It should be noted that the details of the transform window used in MPEG4 AAC are described in detail in, for example, "INTERNATIONAL STANDARD ISO / IEC 14496-3 Fourth edition 2009-09-01 Information technology-coding of audio-visual objects-part 3: Audio".

[0105] On the other hand, Figure 3 As described in , in MPEG-D USAC, the audio signal is encoded.

[0106] In other words, similarly to the case of MPEG-4 AAC, when the encoding (decoding) process starts, first, time-frequency transform is performed on the audio signal using MDCT.

[0107] Then, the MDCT coefficients obtained through the time-frequency transform are quantized according to the scale factor bands, and the quantized MDCT coefficients are obtained as quantization results.

[0108] Furthermore, context-dependent arithmetic coding is performed on the quantized MDCT coefficients, and the arithmetically coded quantized MDCT coefficients are output as coded data on the audio signal.

[0109] In context-based arithmetic coding, a plurality of occurrence probability tables are prepared, in each of which a short code is assigned to an input bit sequence with a high occurrence probability and a long code is assigned to the input bit sequence with a low occurrence probability.

[0110] Furthermore, an effective occurrence probability table is selected based on the encoding results (context) of the previously quantized MDCT coefficients that are close in time and frequency to the quantized MDCT coefficients to be encoded. In other words, the occurrence probability table is appropriately changed based on the correlation between the quantized MDCT coefficients that are close in time and frequency. Furthermore, the quantized MDCT coefficients are encoded using the selected occurrence probability table.

[0111] In context-based arithmetic coding, performing encoding by selecting a valid occurrence probability table from among a plurality of occurrence probability tables makes it possible to achieve higher encoding efficiency.

[0112] Furthermore, unlike Huffman coding, arithmetic coding does not require codebook information to be transmitted. Therefore, compared to Huffman coding, arithmetic coding can reduce the amount of code corresponding to the codebook information.

[0113] It should be noted that Figure 4 As described in , in MPEG-D USAC, MDCT is performed while appropriately changing to a suitable window sequence among five window sequences.

[0114] exist Figure 4 In FIG. 1 , “Window” represents a window sequence, “num_windows” represents the number of transform windows used when performing MDCT using a transform window of each window type, and the shape of the transform window is shown in each “Window Shape” box.

[0115] In MPEG-D USAC, one type can be selected as a window sequence from among five types of transform windows, namely, ONLY_LONG_SEQUENCE, LONG_START_SEQUENCE, EIGHT_SHORT_SEQUENCE, LONG_STOP_SEQUENCE, and STOP_START_SEQUENCE.

[0116] Specifically, among window_sequences, that is, among window types, ONLY_LONG_SEQUENCE, LONG_START_SEQUENCE, EIGHT_SHORT_SEQUENCE, and LONG_STOP_SEQUENCE are the same as in the case of MPEG4AAC.

[0117] In MPEG-D USAC, in addition to these four window types, STOP_START_SEQUENCE is further prepared.

[0118] For frames where the window_sequence transitions from LONG_STOP_SEQUENCE to LONG_START_SEQUENCE, STOP_START_SEQUENCE is selected.

[0119] The transform window indicated by this STOP_START_SEQUENCE is a transform window having a transform window length of 2048 samples.

[0120] It should be noted that the details of MPEG-D USAC are described in, for example, "INTERNATIONAL STANDARD ISO / IEC 23003-3 Frist edition 2012-04-01 Information technology-coding of audio-visual objects-part 3: Unified speech and audio coding".

[0121] It should also be noted that MPEG4 AAC is abbreviated as "AAC" and MPEG-D USAC is abbreviated as "USAC".

[0122] The above-described comparison between AAC and USAC indicates that context-based arithmetic coding is adopted in the current USAC, which is considered to be higher in compression efficiency (coding efficiency) than Huffman coding adopted in AAC.

[0123] However, for all audio signals, context-based arithmetic coding is not always greater (higher) than Huffman coding in terms of compression efficiency.

[0124] In arithmetic coding based on the USAC context, the code is shorter and the coding efficiency tends to be higher than that of AAC Huffman coding for a stationary music signal; however, for an aggressive music signal, the code becomes longer and the coding efficiency tends to be lower.

[0125] Figures 5 to 18 This embodiment is described. It should be noted that Figures 5 to 18 In FIG, the horizontal axis represents time, that is, a frame of an audio signal, and the vertical axis represents the number of encoded bits (the number of necessary bits) or the difference in the number of necessary bits when encoding an audio signal (the number of different bits). Specifically, one frame contains 1024 samples.

[0126] Figure 5 The number of necessary bits required when MDCT and quantization are performed on a still music signal used as an audio signal and AAC Huffman coding is performed on the quantized MDCT coefficients after quantization, and the number of necessary bits required when USAC arithmetic coding is performed on the same quantized MDCT coefficients after quantization are described.

[0127] In this embodiment, the dotted line L11 represents the number of necessary bits in the USAC arithmetic coding of each frame, and the dotted line L12 represents the number of necessary bits in the AAC Huffman coding of each frame. In this embodiment, it should be understood that in most frames, the USAC arithmetic coding has fewer necessary bits than the AAC Huffman coding.

[0128] Further, Figure 6 Described Figure 5 It should be noted that the same reference symbols are used to represent Figure 5 The part corresponding to Figure 6 and omit its description.

[0129] from Figure 6 As is apparent from the portion described in , the number of necessary bits between AAC Huffman coding and USAC arithmetic coding is approximately 100 to 150 bits, and USAC arithmetic coding is greater (higher) in coding efficiency than AAC Huffman coding.

[0130] Figure 7 Describes the difference between the number of necessary bits in AAC Huffman coding and the number of necessary bits in USAC arithmetic coding, that is, Figure 5 Describes the different number of bits in each frame.

[0131] exist Figure 7 , the horizontal axis represents frame (time), and the vertical axis represents the number of different bits. Note that here, the number of different bits is obtained by subtracting the number of necessary bits for AAC Huffman coding from the number of necessary bits for USAC arithmetic coding.

[0132] from Figure 7As is apparent from the above, in the case where the audio signal is a stationary music signal, that is, the audio signal has a stationary characteristic, the different bit numbers appear as negative values in most frames. In other words, it should be understood that in most frames, USAC arithmetic coding requires fewer bits than AAC Huffman coding.

[0133] Therefore, in a case where the audio signal to be encoded is a stationary signal, selecting arithmetic coding as the encoding scheme makes it possible to obtain higher encoding efficiency.

[0134] Moreover, in the MDCT process, a window sequence, that is, the type of window sequence, is selected in each frame. Figure 2 The four window sequences described in Figure 7 When the graph of different bit numbers described in the above is divided into four graphs, the four graphs are Figures 8 to 11 The diagram described in .

[0135] In other words, Figure 8 Indicates the number of different bits in each frame, where ONLY_LONG_SEQUENCE is selected as Figure 7 A sequence of windows of different bit numbers in the described frame.

[0136] same, Figure 9 Describes the different number of bits per frame, where LONG_START_SEQUENCE is selected as Figure 7 A sequence of windows of different bit numbers in the described frame. Figure 10 Describes the different number of bits per frame, where EIGHT_SHORT_SEQUENCE is selected as Figure 7 A sequence of windows of different bit numbers in the described frame.

[0137] Further, Figure 11 Describes the different number of bits per frame, where LONG_STOP_SEQUENCE is selected as Figure 7 A sequence of windows of different bit numbers in the described frame.

[0138] It should be noted that the horizontal axes each represent a frame (time) and the vertical axes each represent Figures 8 to 11 The number of different bits in .

[0139] from Figures 8 to 11It is obvious that in most frames, ONLY_LONG_SEQUENCE is selected because the audio signal is a stationary music signal. In addition, it is obvious that for the remaining selected LONG_START_SEQUENCE, EIGHT_SHORT_SEQUENCE, and LONG_STOP_SEQUENCE, there are fewer frames.

[0140] like Figure 11 As described in , in the case where LONG_STOP_SEQUENCE is selected, the number of different bits is a positive value; thus, in most frames, the coding efficiency of AAC Huffman coding is higher. Needless to say, as Figure 7 As described in , it should be understood that USAC arithmetic coding as a whole is higher in coding efficiency than AAC Huffman coding.

[0141] on the other hand, Figures 12 to 18 Respectively Figures 5 to 11 correspond to each other, and each represents a necessary bit number or a different bit number in a case where the audio signal is an aggressive music signal.

[0142] In other words, Figure 12 The necessary number of bits required when performing MDCT and quantization on an aggressive music signal serving as an audio signal and performing AAC Huffman coding on the quantized MDCT coefficients after quantization, and the necessary number of bits required when performing USAC arithmetic coding on the same quantized MDCT coefficients after quantization are described.

[0143] In the present embodiment, a dotted line L31 indicates the necessary number of bits in USAC arithmetic coding of each frame, and a dotted line L32 indicates the necessary number of bits in AAC Huffman coding of each frame.

[0144] In this embodiment, the number of necessary bits for USAC arithmetic coding is less than that for AAC Huffman coding in most frames. However, in the case of aggressive music signals, the number of necessary bits for AAC Huffman coding is greater than that for USAC arithmetic coding in the case of static music signals.

[0145] Further, Figure 13 Described Figure 12 It should be noted that the same reference symbols are used to represent Figure 12 The part corresponding to Figure 13 and its description will be omitted.

[0146] from Figure 13It should be understood from the description that in several frames, AAC Huffman coding is less in terms of the necessary number of bits than USAC arithmetic coding.

[0147] Figure 14 describes the difference between the number of necessary bits in AAC Huffman coding and the number of necessary bits in USAC arithmetic coding, that is, Figure 12 Describes the different number of bits in each frame.

[0148] exist Figure 14 , the horizontal axis represents frame (time), and the vertical axis represents different bit numbers. It should be noted that here, the different bit numbers are obtained by subtracting the necessary bit number for AAC Huffman coding from the necessary bit number for USAC arithmetic coding.

[0149] from Figure 14 As is apparent from FIG, in the case where the audio signal is an aggressive music signal, that is, the audio signal has an aggressive characteristic, the different bit numbers appear as negative values in most frames.

[0150] However, compared to the case where the audio signal is a stationary music signal, in the case where the audio signal is an aggressive music signal, the number of frames in which the different bit numbers show positive values is larger. In other words, it should be understood that in the case where the audio signal is an aggressive music signal, AAC Huffman coding requires less bits than USAC arithmetic coding in most frames.

[0151] Furthermore, in the MDCT process, a window sequence, i.e., the type of window sequence, is selected in each frame. Figure 2 The four window sequences described in Figure 14 When the different bit numbers described in the figure are divided into four graphs, the four graphs are Figures 15 to 18 The diagram described in .

[0152] In other words, Figure 15 Indicates the number of different bits in each frame, where ONLY_LONG_SEQUENCE is selected as Figure 14 A sequence of windows of varying bit counts describing a frame.

[0153] same, Figure 16 Describes the number of different bits in each frame, where LONG_START_SEQUENCE is selected as Figure 14 A sequence of windows of varying bit counts describing a frame. Figure 17 Describes the number of different bits in each frame, where EIGHT_SHORT_SEQUENCE is selected as Figure 14 A sequence of windows of varying bit counts describing a frame.

[0154] Further, Figure 18 Describes the number of different bits in each frame, where LONG_STOP_SEQUENCE is selected as Figure 14 A sequence of windows of varying bit counts describing a frame.

[0155] It should be noted that the horizontal axes each represent a frame (time), and the vertical axes each represent Figures 15 to 18 The number of different bits in .

[0156] from Figures 15 to 18 It is apparent that, in the case where the audio signal is an aggressive music signal, the ratio of selecting EIGHT_SHORT_SEQUENCE, LONG_START_SEQUENCE, or LONG_STOP_SEQUENCE as the window sequence is higher than that in the case where the audio signal is a stationary music signal.

[0157] Further, it should be understood that, similar to the case of a stationary music signal, in most frames, even if the audio signal is an aggressive music signal, in the case where ONLY_LONG_SEQUENCE, LONG_START_SEQUENCE, or EIGHT_SHORT_SEQUENCE is selected, USAC arithmetic coding is higher in coding efficiency than AAC Huffman coding.

[0158] However, it should be understood that in the case where LONG_STOP_SEQUENCE is selected, AAC Huffman coding is smaller in the number of necessary bits and higher in coding efficiency than USAC arithmetic coding in most frames.

[0159] This is because context correlation in USAC arithmetic coding is reduced due to the conversion between a frame with a strong attack characteristic and a frame with a strong stationary characteristic, and an inefficient occurrence probability table is selected.

[0160] It should be noted that since the quantized MDCT coefficients are encoded using a transform window divided into eight in the time direction, the number of bits (code amount) required in the USAC arithmetic coding for each frame in which EIGHT_SHORT_SEQUENCE is selected is not large. In other words, the encoding of the quantized MDCT coefficients is performed eight times to correspond to the eight divided transform windows, each having a transform window length of 256 samples in the time direction; thereby, the degree of reduction in context correlation is dispersed and alleviated.

[0161] As described above, in a case where an audio signal has an aggressive characteristic, specifically, at the time of transformation from a frame using a transform window with a small transform window length to a frame using a transform window with a large transform window length, in a frame, that is, in each frame in which LONG_STOP_SEQUENCE is selected, USAC arithmetic coding is lower in coding efficiency (compression efficiency) than AAC Huffman coding.

[0162] Furthermore, an increase in the code length of arithmetic coding naturally leads to an increase in computational complexity at the time of decoding.

[0163] Further, arithmetic coding has the characteristic that decoding cannot be performed without unifying all signs of one quantized MDCT coefficient, and requires greater computational complexity than Huffman coding due to a large amount of computational processing occurring per bit.

[0164] Therefore, in order to solve the problem, the present technology aims to improve encoding efficiency and reduce computational complexity in a decoding process by appropriately selecting an encoding scheme when encoding an audio signal.

[0165] Specifically, for example, in the case of a transformation from a frame in which time-frequency transformation is performed using a transform window having a small transform window length to a frame in which time-frequency transformation is performed using a transform window having a larger transform window length than the former frame, the quantized spectral information is Huffman encoded in a codec using a time-frequency transform similar to USAC.

[0166] For example, in the case of USAC, in each frame in which LONG_STOP_SEQUENCE is selected, Huffman coding is selected as the coding scheme.

[0167] Further, for other frames, that is, frames except for each frame when transforming from a small transform window length to a large transform window length, Huffman coding or arithmetic coding is selected as the encoding scheme.

[0168] At this time, including a decision flag for identifying the selected coding scheme in the coded bitstream as needed can enable the decoding side to identify the selected coding scheme. In other words, indicating a change in the decision flag or decoding scheme in the decoder syntax can enable the decoding side to appropriately change the decoding scheme.

[0169] <Configuration Example of Encoding Device>

[0170] Subsequently, specific embodiments of encoding and decoding devices to which this technology is applied will be described. It should be noted that embodiments for performing encoding and decoding based on MPEG-D USAC will be described later. However, any other codec may be used as long as the time-frequency transform information is encoded by appropriately changing the transform window length and selecting any of a plurality of encoding schemes including context-based arithmetic coding.

[0171] Figure 19 is a diagram describing an embodiment of a configuration of an encoding device to which the present technology is applied.

[0172] Figure 19 The encoding device 11 described in has a time-frequency transform section 21 , a normalization section 22 , a quantization section 23 , a coding scheme selection section 24 , an encoding section 25 , a bit control section 26 , and a multiplexing section 27 .

[0173] The time-frequency transform section 21 selects a transform window for each frame of the supplied audio signal and performs time-frequency transform on the audio signal using the selected transform window.

[0174] Furthermore, the time-frequency transform section 21 supplies spectrum information obtained by the time-frequency transform to the normalization section 22 and supplies transform window information indicating the type of transform window (window sequence) selected for each frame to the encoding scheme selection section 24 and the multiplexing section 27 .

[0175] For example, the time-frequency transform section 21 performs MDCT as the time-frequency transform and obtains MDCT coefficients as the spectrum information. For example, the description will be continued considering a case where the spectrum information is the MDCT coefficients.

[0176] The normalization section 22 normalizes the MDCT coefficient supplied from the time-frequency transform section 21 based on the normalization parameter supplied from the bit control section 26 , and supplies the normalized MDCT coefficient obtained as a result of the normalization to the quantization section 23 , and supplies the parameter associated with the normalization to the multiplexing section 27 .

[0177] The quantization section 23 quantizes the normalized MDCT coefficient supplied from the normalization section 22 and supplies the quantized MDCT coefficient obtained as a result of the quantization to the encoding scheme selection section 24 .

[0178] The coding scheme selection section 24 selects a coding scheme based on the transform window information supplied from the time-frequency transform section 21 and supplies the quantized MDCT coefficient supplied from the quantization section 23 to a module in the coding section 25 according to the selection result of the coding scheme.

[0179] The encoding section 25 encodes the quantized MDCT coefficients supplied from the encoding scheme selection section 24 by the encoding scheme selected (specified) by the encoding scheme selection section 24. The encoding section 25 has a Huffman encoding section 31 and an arithmetic encoding section 32.

[0180] In the case where the quantized MDCT coefficients are supplied from the encoding scheme selection section 24, the Huffman encoding section 31 encodes the quantized MDCT coefficients by the Huffman encoding scheme. In other words, the quantized MDCT coefficients are subjected to Huffman encoding.

[0181] The Huffman coding unit 31 supplies the MDCT-coded data obtained through Huffman coding and Huffman codebook information to the bit control unit 26. Here, the Huffman codebook information indicates the Huffman codebook used during Huffman coding. Furthermore, the Huffman codebook information supplied to the bit control unit 26 is subjected to Huffman coding.

[0182] In the case where the quantized MDCT coefficients are supplied from the encoding scheme selection section 24, the arithmetic encoding section 32 encodes the quantized MDCT coefficients by the arithmetic encoding scheme. In other words, the quantized MDCT coefficients are subjected to context-based arithmetic encoding.

[0183] The arithmetic coding section 32 supplies the MDCT-coded data obtained by the arithmetic coding to the bit control section 26 .

[0184] When the MDCT encoded data and the Huffman codebook information are supplied from the Huffman encoding section 31 to the bit control section 26 or when the MDCT encoded data are supplied from the arithmetic encoding section 32 to the bit control section 26 , the bit control section 26 determines the bit amount and the sound quality.

[0185] In other words, the bit control section 26 determines whether the bit amount (code amount) of the MDCT encoded data is within the target bit amount to be used, and determines whether the sound quality of the sound based on the MDCT encoded data is within an allowable range.

[0186] In a case where the bit amount and the like of the MDCT encoded data are within the target bit amount to be used and the sound quality is within the allowable range, the bit control section 26 supplies the supplied MDCT encoded data and the like to the multiplexing section 27 .

[0187] On the contrary, in a case where the bit amount of MDCT encoded data, etc. is not within the target bit amount to be used or the sound quality is not within the allowable range, the bit control section 26 resets the parameters supplied to the normalization section 22 and supplies the reset parameters to the normalization section 22 to complete the encoding again.

[0188] The multiplexing section 27 multiplexes the MDCT encoded data and Huffman codebook information supplied from the bit control section 26, the transform window information supplied from the time-frequency transform section 21, and the parameters supplied from the normalization section 22, and outputs an encoded bit stream obtained as a result of the multiplexing.

[0189] <Description of Encoding Process>

[0190] Next, the operation performed by the encoding device 11 will be described. In other words, reference will be made to Figure 20 The flowchart in describes the encoding process performed by the encoding device 11. It should be noted that this encoding process is performed on each frame of the audio signal.

[0191] In step S11 , the time-frequency transform section 21 performs time-frequency transform on the supplied frames of the audio signal.

[0192] In other words, the time-frequency conversion unit 21 determines the aggressive or static characteristics of the frame to be processed of the audio signal based on, for example, the MDCT coefficients that are close to the MDCT coefficients in time and frequency, or the size and variation of the audio signal. In other words, the time-frequency conversion unit 21 determines whether the audio signal has an aggressive or static characteristic based on the size and variation of the MDCT coefficients, the size and variation of the audio signal, and the like.

[0193] The time-frequency transform section 21 selects a transform window for the frame to be processed based on the result of the determination of the attack characteristic or the stationary characteristic, the result of the selection of the transform window for the frame immediately preceding the frame to be processed, etc., and performs time-frequency transform on the frame to be processed of the audio signal using the selected transform window. The time-frequency transform section 21 supplies the MDCT coefficients obtained by the time-frequency transform to the normalization section 22 and supplies transform window information indicating the type of the selected transform window to the encoding scheme selection section 24 and the multiplexing section 27.

[0194] In step S12, the normalization section 22 normalizes the MDCT parameters supplied from the time-frequency transform section 21 based on the parameters supplied from the bit control section 26, and supplies the normalized MDCT coefficients obtained as a result of the normalization to the quantization section 23, and supplies the parameters associated with the normalization to the multiplexing section 27.

[0195] In step S13 , the quantization section 23 quantizes the normalized MDCT coefficient supplied from the normalization section 22 and supplies the quantized MDCT coefficient obtained as a result of the quantization to the encoding scheme selection section 24 .

[0196] In step S14 , the encoding scheme selection section 24 determines whether the type of the transform window is LONG_STOP_SEQUENCE, that is, the window sequence indicated by the transform window information supplied from the time-frequency transform section 21 .

[0197] In the case where it is judged in step S14 that the window sequence is LONG_STOP_SEQUENCE, the encoding scheme selection section 24 supplies the quantized MDCT coefficient supplied from the quantization section 23 to the Huffman encoding section 31 , and then the process proceeds to step S15 .

[0198] The frame for which LONG_STOP_SEQUENCE is selected is a frame at the time of transition from a frame having a strong attack characteristic and a small transform window length (ie, EIGHT_SHORT_SEQUENCE) to a frame having a strong stationary characteristic and a large transform window length (ie, ONLY_LONG_SEQUENCE).

[0199] For example, as referenced Figure 18 As described, in the case of the frame whose transform window length is changed from a small transform window length to a large transform window length, that is, the frame in which LONG_STOP_SEQUENCE is selected, Huffman coding is higher in encoding efficiency than arithmetic coding.

[0200] Therefore, when encoding this frame, the Huffman coding scheme is selected as the encoding scheme. In other words, similar to MPEG4 AAC, the quantized MDCT coefficients and Huffman codebook information are encoded using the same Huffman codebook using the Huffman code for each part.

[0201] In step S15 , the Huffman encoding section 31 performs Huffman encoding on the quantized MDCT coefficient supplied from the encoding scheme selection section 24 using the Huffman codebook information, and supplies the MDCT encoded data and the Huffman codebook information to the bit control section 26 .

[0202] The bit control section 26 determines the target code amount and sound quality to be used based on the MDCT encoded data and Huffman codebook information supplied from the Huffman encoding section 31. The encoding device 11 repeatedly performs a series of processes including parameter resetting, normalization, quantization, and Huffman encoding until MDCT encoded data and Huffman codebook information of the target bit amount and target quality are obtained.

[0203] Further, when the MDCT encoded data and the Huffman codebook information of the target bit amount and the target quality are obtained, the bit control section 26 supplies the MDCT encoded data and the Huffman codebook information to the multiplexing section 27 , and the process proceeds to step S17 .

[0204] On the other hand, if it is determined in step S14 that the window sequence is not LONG_STOP_SEQUENCE, that is, if the transform window length is not changed from a small transform window length to a large transform window length, the process proceeds to step S16. In this case, the encoding scheme selection unit 24 supplies the quantized MDCT coefficients supplied from the quantization unit 23 to the arithmetic encoding unit 32.

[0205] In step S16, the arithmetic coding section 32 performs context-based arithmetic coding on the quantized MDCT coefficients supplied from the coding scheme selection section 24 and supplies MDCT coded data to the bit control section 26. In other words, the quantized MDCT coefficients are subjected to arithmetic coding.

[0206] The bit control section 26 determines the bit amount to be used and the sound quality based on the MDCT encoded data supplied from the arithmetic encoding section 32. The encoding device 11 repeatedly performs processing including parameter resetting, normalization, quantization, and arithmetic encoding until MDCT encoded data of a target bit amount and target quality is obtained.

[0207] Further, when the MDCT encoded data of the target bit amount and the target quality is obtained, the bit control section 26 supplies the MDCT encoded data to the multiplexing section 27 , and the process proceeds to step S17 .

[0208] When the process of step S15 or S16 is executed, the process of step S17 is executed.

[0209] In other words, in step S17 , the multiplexing section 27 performs multiplexing to generate an encoded bit stream and transmits (outputs) the obtained encoded bit stream to a decoding device or the like.

[0210] For example, in the case where the processing of step S15 is performed, the multiplexing section 27 multiplexes the MDCT encoded data and Huffman codebook information supplied from the bit control section 26, the transform window information supplied from the time-frequency transform section 21, and the parameters supplied from the normalization section 22 and generates an encoded bit stream.

[0211] Further, for example, in a case where the processing of step S16 is performed, the multiplexing section 27 multiplexes the MDCT encoded data supplied from the bit control section 26, the transform window information supplied from the time-frequency transform section 21, and the parameters supplied from the normalization section 22 and generates an encoded bit stream.

[0212] When the encoded bit stream obtained in this way is output, the encoding process ends.

[0213] As described so far, the encoding device 11 selects an encoding scheme according to the type of transform window used in time-frequency transform. This makes it possible to select an appropriate encoding scheme for each frame and improve encoding efficiency.

[0214] <Configuration Example of Decoding Device>

[0215] Subsequently, a decoding device that receives the encoded bit stream output from the encoding device 11 and performs decoding will be described.

[0216] For example, the configuration of the decoding device is as follows Figure 21 As described in.

[0217] Figure 21 The decoding device 71 described in has an acquisition section 81 , a demultiplexing section 82 , a decoding scheme selection section 83 , a decoding section 84 , an inverse quantization section 85 , and an inverse time-frequency transform section 86 .

[0218] The acquisition section 81 acquires an encoded bit stream by receiving the encoded bit stream supplied from the encoding device 11 and supplies the encoded bit stream to the demultiplexing section 82 .

[0219] The demultiplexing section 82 demultiplexes the encoded bit stream supplied from the acquisition section 81 and supplies the MDCT encoded data and Huffman codebook information obtained by demultiplexing to the decoding scheme selection section 83. Furthermore, the demultiplexing section 82 supplies the parameter associated with normalization and obtained by demultiplexing to the inverse quantization section 85 and supplies the transform window information obtained by demultiplexing to the decoding scheme selection section 83 and the inverse time-frequency transform section 86.

[0220] The decoding scheme selection section 83 selects a decoding scheme based on the transform window information supplied from the demultiplexing section 82 , and supplies the MDCT encoded data and the like supplied from the demultiplexing section 82 to a module in the decoding section 84 according to the selection result of the decoding scheme.

[0221] The decoding section 84 decodes the MDCT-encoded data and the like supplied from the decoding scheme selection section 83. The decoding section 84 has a Huffman decoding section 91 and an arithmetic decoding section 92.

[0222] In the case where MDCT encoded data and Huffman codebook information are supplied from the decoding scheme selection section 83, the Huffman decoding section 91 decodes the MDCT encoded data using the Huffman codebook information by a decoding scheme corresponding to Huffman encoding, and supplies the quantized MDCT coefficients obtained as a result of the decoding to the inverse quantization section 85.

[0223] In a case where the MDCT encoded data is supplied from the decoding scheme selection section 83 , the arithmetic decoding section 92 decodes the MDCT encoded data by a decoding scheme corresponding to arithmetic coding and supplies quantized MDCT coefficients obtained as a result of the decoding to the inverse quantization section 85 .

[0224] The inverse quantization section 85 inversely quantizes the quantized MDCT coefficients supplied from the Huffman decoding section 91 or the arithmetic decoding section 92 using the parameters supplied from the demultiplexing section 82, and supplies the MDCT coefficients obtained as a result of the inverse quantization to the inverse time-frequency transform section 86. More specifically, the inverse quantization section 85 obtains the MDCT coefficients by, for example, multiplying the values obtained by inversely quantizing the quantized MDCT coefficients by the parameters supplied from the demultiplexing section 82.

[0225] The inverse time-frequency transform section 86 performs inverse time-frequency transform on the MDCT coefficient supplied from the inverse quantization section 85 based on the transform window information supplied from the demultiplexing section 82 and outputs the output audio signal (i.e., the time signal obtained as a result of the inverse time-frequency transform) to the subsequent stage.

[0226] <Description of decoding process>

[0227] Next, the operation performed by the decoding device 71 will be described. In other words, reference will be made to Figure 22 The flowchart of describes the decoding process performed by the decoding device 71. It should be noted that this decoding process starts when the acquisition section 81 receives an encoded bit stream corresponding to one frame.

[0228] In step S41, the demultiplexing section 82 demultiplexes the coded bit stream supplied from the acquisition section 81 and supplies the MDCT coded data and the like obtained by demultiplexing to the decoding scheme selection section 83 and the like. In other words, the MDCT coded data, transform window information, and various parameters are extracted from the coded bit stream.

[0229] In this case, when the audio signal (MDCT coefficient) is subjected to Huffman coding, MDCT coded data and Huffman codebook information are extracted from the coded bit stream. Conversely, when the audio signal is subjected to arithmetic coding, MDCT coded data are extracted from the coded bit stream.

[0230] Further, the demultiplexing section 82 supplies the parameters associated with normalization and obtained by demultiplexing to the inverse quantization section 85 , and supplies the transform window information obtained by demultiplexing to the decoding scheme selection section 83 and the time-frequency inverse transform section 86 .

[0231] In step S42 , the decoding scheme selection section 83 determines whether the type of the transform window indicated by the transform window information supplied from the demultiplexing section 82 is LONG_STOP_SEQUENCE.

[0232] In the case where it is determined in step S42 that the type of the transform window is LONG_STOP_SEQUENCE, the decoding scheme selection section 83 supplies the MDCT encoded data and Huffman codebook information supplied from the demultiplexing section 82 to the Huffman decoding section 91 , and the process proceeds to step S43 .

[0233] In this case, the frame to be processed is the frame at the time of changing from a frame with a small transform window length to a frame with a large transform window length. In other words, the transform window indicated by the transform window information is the transform window selected when changing from a small transform window length to a large transform window length. Thus, the decoding scheme selection unit 83 selects the decoding scheme corresponding to Huffman coding as the decoding scheme.

[0234] In step S43, the Huffman decoding section 91 decodes the MDCT encoded data and Huffman codebook information (ie, Huffman code) supplied from the decoding scheme selection section 83. Specifically, the Huffman decoding section 91 obtains quantized MDCT coefficients based on the Huffman codebook information and the MDCT encoded data.

[0235] The Huffman decoding section 91 supplies the quantized MDCT coefficient obtained by decoding to the inverse quantization section 85 , and then the process proceeds to step S45 .

[0236] In contrast, in the case where it is determined in step S42 that the type of the transform window is not LONG_STOP_SEQUENCE, the decoding scheme selection section 83 supplies the MDCT encoded data supplied from the demultiplexing section 82 to the arithmetic decoding section 92 , and the process proceeds to step S44 .

[0237] In this case, the frame to be processed is not the frame at the time of changing from a frame with a small transform window length to a frame with a large transform window length. In other words, the transform window indicated by the transform window information is not the transform window selected when changing from a small transform window length to a large transform window length. Therefore, the decoding scheme selection unit 83 selects the decoding scheme corresponding to arithmetic coding as the decoding scheme.

[0238] In step S44 , the arithmetic decoding section 92 decodes the MDCT-encoded data (ie, the arithmetic code) supplied from the decoding scheme selection section 83 .

[0239] The arithmetic decoding section 92 supplies the quantized MDCT coefficient obtained by decoding the MDCT-encoded data to the inverse quantization section 85 , and then the process proceeds to step S45 .

[0240] When the process of step S43 or S44 is executed, the process of step S45 is executed.

[0241] In step S45 , the inverse quantization section 85 inversely quantizes the quantized MDCT coefficients supplied from the Huffman decoding section 91 or the arithmetic decoding section 92 using the parameters supplied from the demultiplexing section 82 , and supplies the MDCT coefficients obtained as a result of the demultiplexing to the inverse time-frequency transform section 86 .

[0242] In step S46 , the time-frequency inverse transform section 86 performs time-frequency inverse transform on the MDCT coefficient supplied from the inverse quantization section 85 based on the transform window information supplied from the demultiplexing section 82 and outputs the output audio signal obtained as a result of the time-frequency inverse transform to the subsequent stage.

[0243] When the output audio signal is output, the decoding process ends.

[0244] As described so far, the decoding device 71 selects a decoding scheme based on the transform window information obtained by demultiplexing the encoded bit stream and performs decoding using the selected decoding scheme. Specifically, when the transform window type is LONG_STOP_SEQUENCE, a decoding scheme corresponding to Huffman coding is selected; otherwise, a decoding scheme corresponding to arithmetic coding is selected. This improves encoding efficiency on the encoding side and reduces the throughput (computational complexity) of the decoding process on the decoding side.

[0245] Meanwhile, a scheme in which Huffman coding is performed on a frame in which LONG_STOP_SEQUENCE is selected and arithmetic coding is performed on a frame in which a transform window type other than LONG_STOP_SEQUENCE is selected is referred to as a “hybrid coding scheme.” According to this hybrid coding scheme, encoding efficiency can be improved and throughput in the decoding process can be reduced.

[0246] For example, as Figure 5 The situation described in Figure 23 A graph depicting the difference in the number of necessary bits between the case where Huffman coding is used for a frame for which a LONG_STOP_SEQUENCE compliant with USAC is selected (i.e., the case where encoding is performed by a hybrid coding scheme) and the case where AAC Huffman coding is always used when encoding the same still music signal.

[0247] It should be noted that Figure 23 , the horizontal axis represents frame (time) and the vertical axis represents the number of different bits. The number of different bits mentioned here is obtained by subtracting the number of necessary bits for AAC Huffman coding from the number of necessary bits for the hybrid coding scheme.

[0248] Figure 23 The different number of bits per frame described in Figure 7 The different bit numbers described in . Figure 23 and Figure 7 , that is, the comparison between the case where encoding is performed by the hybrid encoding scheme and the case where arithmetic coding is always performed indicates that although Figure 23 The embodiment in is higher in coding efficiency, however, the difference in coding efficiency is not that great.

[0249] On the contrary, as Figure 12 The situation described in Figure 24 The difference in the number of necessary bits between the case where Huffman coding is used for a frame for which a LONG_STOP_SEQUENCE compliant with USAC is selected (i.e., the case where encoding is performed by a hybrid coding scheme) and the case where AAC Huffman coding is always used when encoding the same aggressive music signal is described.

[0250] It should be noted that Figure 24 , the horizontal axis represents frame (time), and the vertical axis represents different bit numbers. The different bit numbers mentioned here are obtained by subtracting the necessary bit number of AAC Huffman coding from the necessary bit number of the hybrid coding scheme.

[0251] Figure 24 The different number of bits per frame described in Figure 14 The different bit numbers described in . Figure 24 and Figure 14 Comparison between, that is, a case where encoding is performed by a hybrid encoding scheme and a case where arithmetic coding is always performed represents that, Figure 24 The embodiment in is minimal in terms of the number of different bits. In other words, the comparison represents Figure 24 The embodiment in has a greater improvement in coding efficiency.

[0252] Furthermore, by using a hybrid coding scheme, arithmetic coding is not used for each frame for which LONG_STOP_SEQUENCE is selected, but Huffman coding is used, so that the throughput in the decoding process of the frame can also be reduced.

[0253] <Second embodiment>

[0254] <Encoding Scheme Selection>

[0255] Meanwhile, it has been described so far that arithmetic coding is always selected as the encoding scheme for each frame in which a transform window type other than LONG_STOP_SEQUENCE is selected. However, when selecting an encoding scheme, it is preferable to consider not only encoding efficiency (compression efficiency) but also tolerances such as throughput and sound quality.

[0256] Thus, for example, for a frame in which a transform window type other than LONG_STOP_SEQUENCE is selected, Huffman coding or arithmetic coding may also be selected.

[0257] In this case, for example, a decision flag indicating selection of Huffman coding or arithmetic coding as the coding scheme during the encoding process is stored in the encoded bit stream.

[0258] For example, it is assumed here that a value of “1” of the determination flag indicates selection of the Huffman coding scheme and a value of “0” of the determination flag indicates selection of the arithmetic coding scheme.

[0259] In the case where the frame is a frame for which a transform window type other than LONG_STOP_SEQUENCE is selected, that is, in the case where the transform window length is not changed from a small transform window length to a large transform window length, the determination flag can be regarded as selection information indicating the encoding scheme selected for the frame to be processed. In other words, the determination flag can be regarded as selection information indicating the selection result of the encoding scheme.

[0260] It should be noted that for this frame, since the Huffman coding scheme is always selected, the decision flag is not included in the encoded bit stream for the frame where LONG_STOP_SEQUENCE is selected.

[0261] For example, Figure 25 , in the case where a determination flag is stored in a coded bit stream as needed, a syntax of a channel stream corresponding to one frame of an audio signal in a predetermined channel of the coded bit stream may be a syntax based on MPEG-D USAC.

[0262] exist Figure 25 In the described embodiment, the portion indicated by the arrow Q11, that is, the portion of the characters "ics_info()", represents ics_info in which information associated with the transform window and the like is stored.

[0263] In addition, the portion of the character "section_data()" indicated by the arrow Q12 represents section_data. Huffman codebook information and the like are stored in this section_data. Further, Figure 25 The characters "ac_spectral_data" in represent MDCT-encoded data.

[0264] Also, for example, the syntax of the part ics_info represented by the characters "ics_info()" is as follows Figure 26 As described in.

[0265] exist Figure 26In the described embodiment, the portion of characters "window_sequence" represents transformation window information, that is, a window sequence, and the portion of characters "window_shape" represents the shape of the transformation window.

[0266] Furthermore, the character "huffman_coding_flag" portion represents a determination flag.

[0267] Here, if the conversion window information stored in the portion of the character "window_sequence" indicates LONG_STOP_SEQUENCE, the judgment flag is not stored in ics_info. On the contrary, if the conversion window information indicates a type other than LONG_STOP_SEQUENCE, the judgment flag is stored in ics_info.

[0268] Therefore, in Figure 25 In the described embodiment, in which the Figure 26 The conversion window information in the portion of the character "window_sequence" of the character "window_sequence" indicates a type other than LONG_STOP_SEQUENCE and a judgment flag having a value of "1" is stored in Figure 26 In the case of the character "huffman_coding_flag" in the section, the Huffman codebook information and the like are stored in section_data. Further, Figure 26 When the conversion window information in the portion of the character "window_sequence" indicates LONG_STOP_SEQUENCE, Huffman codebook information and the like are also stored in section_data.

[0269] <Description of Encoding Process>

[0270] as Figure 25 and Figure 26 In the embodiment described in, for example, in the case where the judgment flag is stored in the encoded bit stream as needed, the encoding device 11 performs Figure 27 The following will refer to the encoding process described in Figure 27 The flowchart in describes the encoding process performed by the encoding device 11.

[0271] It should be noted that because the processing Figure 20 The processing of steps S11 to S15 in are similar, so the description of the processing of steps S71 to S75 will be omitted.

[0272] When it is determined in step S74 that the window sequence is not LONG_STOP_SEQUENCE, the encoding scheme selection unit 24 determines in step S76 whether or not to perform arithmetic coding.

[0273] For example, the encoding scheme selection section 24 determines whether to perform arithmetic encoding based on designation information supplied from a high-order control device.

[0274] Here, for example, the designated information refers to information indicating a coding scheme designated by a content producer, etc. For example, when the window sequence is not a frame of a LONG_STOP_SEQUENCE, the content producer can designate Huffman coding or arithmetic coding as the coding scheme for each frame.

[0275] In this case, when the encoding scheme indicated by the designation information is arithmetic coding, the encoding scheme selection unit 24 determines in step S76 to perform arithmetic coding. On the other hand, when the encoding scheme indicated by the designation information is Huffman coding, the encoding scheme selection unit 24 determines in step S76 not to perform arithmetic coding.

[0276] Further, the encoding scheme selection section 24 may select an encoding scheme in step S76 based on resources of the decoding device 71 and the encoding device 11 , ie, throughput, bit rate of the audio signal to be encoded, whether real-time characteristics are required, and the like.

[0277] Specifically, for example, in a case where the bit rate of the audio signal is high and sufficient sound quality can be ensured, the encoding scheme selection section 24 may select Huffman encoding having a lower throughput and determine not to perform arithmetic encoding in step S76 .

[0278] Moreover, for example, in a case where real-time characteristics are required, the decoding device 71 has fewer resources, and it is important to immediately perform encoding and decoding processing with low throughput that is superior to sound quality, the encoding scheme selection unit 24 may select Huffman coding and determine in step S76 not to perform arithmetic coding.

[0279] In a case where real-time characteristics are required or the decoding side thus has fewer resources, selecting Huffman coding as the coding scheme makes it possible to perform processing (operation) at a higher speed than when arithmetic coding is always performed.

[0280] It should be noted that regarding the resources of the decoding device 71, before starting encoding processing, etc., it is sufficient to obtain in advance from the decoding device 71 only the computing processing capability of the device in which the decoding device 71 is set, information indicating the memory capability, etc. as resource information about the decoding device 71.

[0281] If it is determined in step S76 that arithmetic coding is to be performed, the coding scheme selection section 24 supplies the quantized MDCT coefficients supplied from the quantization section 23 to the arithmetic coding section 32, and then performs the processing of step S77. In other words, in step S77, context-based arithmetic coding is performed on the quantized MDCT coefficients.

[0282] It should be noted that because the processing Figure 20 , so the description of the process of step S77 will be omitted. When the process of step S77 is executed, the process proceeds to step S79.

[0283] In contrast, in the case where it is determined not to perform arithmetic coding, that is, in step S76 , it is determined to perform Huffman coding, the coding scheme selection section 24 supplies the quantized MDCT coefficient supplied from the quantization section 23 to the Huffman coding section 31 , and the process proceeds to step S78 .

[0284] In step S78, similar processing to step S75 is performed, and MDCT encoded data and Huffman codebook information obtained as a result of the processing are supplied from the Huffman encoding section 31 to the bit control section 26. When the processing of step S78 is performed, the processing proceeds to step S79.

[0285] When the process of step S77 or S78 is executed, the bit control unit 26 generates a determination flag in step S79.

[0286] For example, in the case where the processing of step S77 (ie, arithmetic coding) is performed, the bit control section 26 generates a determination flag having a value of “0” and supplies the generated determination flag to the multiplexing section 27 together with the MDCT encoded data supplied from the arithmetic coding section 32 .

[0287] Further, for example, in a case where the processing of step S78 (i.e., Huffman encoding) is performed, the bit control section 26 generates a judgment flag having a value of "1" and supplies the generated judgment flag to the multiplexing section 27 together with the MDCT encoded data and Huffman codebook information supplied from the Huffman encoding section 31.

[0288] When the process of step S79 is executed, the process proceeds to step S80.

[0289] When the processing of step S75 or S79 is performed, the multiplexing section 27 performs multiplexing in step S80 to generate a coded bit stream and sends the obtained coded bit stream to the decoding device 71. It should be noted that in step S80, the same processing as in step S75 or S79 is basically performed. Figure 20 The same processing is performed as in step S17.

[0290] For example, when the process of step S75 is executed, the multiplexing unit 27 generates a coded bit stream in which the MDCT coded data, Huffman codebook information, transform window information, and parameters from the normalization unit 22 are stored. This coded bit stream does not include a determination flag.

[0291] Further, for example, in the case where the process of step S78 is performed, the multiplexing section 27 generates a coded bit stream in which the determination flag from the normalization section 22, MDCT coded data, Huffman codebook information, transform window information, and parameters are stored.

[0292] Also, for example, in the case where the process of step S77 is performed, the multiplexing section 27 generates an encoded bit stream in which the determination flag from the normalization section 22, the MDCT encoded data, the transform window information, and the parameters are stored.

[0293] When the encoded bit stream is generated and output in this way, the encoding process ends.

[0294] As described so far, for frames whose window sequence is not LONG_STOP_SEQUENCE, the encoding device 11 selects Huffman coding or arithmetic coding and performs encoding using the selected encoding scheme. This makes it possible to select an appropriate encoding scheme for each frame, improve encoding efficiency, and achieve encoding with a higher degree of freedom.

[0295] <Description of decoding process>

[0296] Further, in which the encoding device 11 performs reference Figure 27 In the case of the encoding process described, the decoding device 71 performs Figure 28 The decoding process described in .

[0297] The following will refer to Figure 28 The flowchart of FIG. 7 describes the decoding process performed by the decoding device 71. It should be noted that since the process is Figure 22 Since the processing of steps S41 to S43 in the above is similar, the description of the processing of steps S121 to S123 will be omitted. However, it should be noted that in the case where the judgment flag is extracted from the encoded bit stream by demultiplexing in step S121, the judgment flag is supplied from the demultiplexing section 82 to the decoding scheme selection section 83.

[0298] In the case where it is determined in step S122 that the window sequence is not LONG_STOP_SEQUENCE, the decoding scheme selection section 83 determines in step S124 whether the MDCT encoded data is arithmetic coding based on the determination flag supplied from the demultiplexing section 82. In other words, the decoding scheme selection section 83 determines whether the encoding scheme of the MDCT encoded data is arithmetic coding.

[0299] For example, in the case where the value of the judgment flag is "1", the decoding scheme selection section 83 judges that the MDCT-coded data is not an arithmetic code, that is, a Huffman code, and in the case where the value of the judgment flag is "0", the decoding scheme selection section 83 judges that the MDCT-coded data is an arithmetic code. In this way, the decoding scheme selection section 83 selects Huffman coding or arithmetic coding of the decoding scheme corresponding to the coding scheme indicated by the judgment flag.

[0300] In the case where it is determined in step S124 that the MDCT coded data is not an arithmetic code, that is, a Huffman code, the decoding scheme selection section 83 supplies the MDCT coded data and the Huffman codebook information supplied from the demultiplexing section 82 to the Huffman decoding section 91, and the process proceeds to step S123. Then, in step S123, the Huffman code is decoded.

[0301] In contrast, in the case where it is determined in step S124 that the MDCT encoded data is an arithmetic code, the decoding scheme selection section 83 supplies the MDCT encoded data supplied from the demultiplexing section 82 to the arithmetic decoding section 92 , and the process proceeds to step S125 .

[0302] In step S125, the MDCT coded data (ie, arithmetic code) is decoded by a decoding scheme corresponding to arithmetic coding. Figure 22 The process is similar to step S44 in , so the description of the process of step S125 will be omitted.

[0303] When the processing of step S123 or S125 is executed, then, the processing of steps S126 and S127 is executed, and the decoding processing ends. Figure 22 Steps S45 and S46 in are similar, so the description of steps S126 and S127 will be omitted.

[0304] As described so far, the decoding device 71 selects a decoding scheme based on the transform window information and the judgment flag and performs decoding. Specifically, even for a frame whose window sequence is not a LONG_STOP_SEQUENCE, since the correct decoding scheme can be selected by referring to the judgment flag, not only can the encoding efficiency be improved and the throughput on the decoding side be reduced, but also encoding and decoding can be achieved with a higher degree of freedom.

[0305] <Third embodiment>

[0306] <Description of Encoding Process>

[0307] Alternatively, for a frame whose window sequence is not LONG_STOP_SEQUENCE, in the case where Huffman coding or arithmetic coding is selected, a coding scheme with a smaller number of necessary bits may be selected.

[0308] For example, in a case where the throughput of the decoding device 71 or the encoding device 11 has tolerance and coding efficiency (compression efficiency) is prioritized, the necessary number of bits for Huffman coding and arithmetic coding can be calculated, and for frames whose window sequence is not LONG_STOP_SEQUENCE, a coding scheme with a smaller number of necessary bits can be selected.

[0309] In this case, for example, the encoding device 11 performs Figure 29 In other words, the following will refer to the encoding process described in Figure 29 The flowchart of is used to describe the encoding process performed by the encoding device 11.

[0310] It should be noted that because the processing Figure 20 The processing of steps S11 to S15 in are similar, so the description of the processing of steps S151 to S155 will be omitted.

[0311] In the case where it is determined in step S154 that the window sequence is not LONG_STOP_SEQUENCE, the coding scheme selection section 24 supplies the quantized MDCT coefficients supplied from the quantization section 23 to the Huffman coding section 31 and the arithmetic coding section 32, and the process proceeds to step S156. In this case, the timing of selecting (adopting) the coding scheme in step S154 has not yet been determined.

[0312] In step S156, the arithmetic coding section 32 performs context-based arithmetic coding on the quantized MDCT coefficients supplied from the coding scheme selection section 24 and supplies MDCT coded data obtained as a result of the coding to the bit control section 26. Figure 20 Similar processing is performed in step S16.

[0313] In step S157, the Huffman encoding section 31 performs Huffman encoding on the quantized MDCT coefficient supplied from the encoding scheme selection section 24 and supplies MDCT encoded data obtained as a result of the encoding and Huffman codebook information to the bit control section 26. In step S157, processing similar to step S1155 is performed.

[0314] In step S158 , the bit control section 26 compares the bit number of the MDCT encoded data and the Huffman codebook information supplied from the Huffman encoding section 31 with the bit number of the MDCT encoded data supplied from the arithmetic encoding section 32 and selects a coding scheme.

[0315] In other words, in a case where the bit number (code amount) of MDCT encoded data obtained by Huffman encoding and Huffman codebook information are smaller than the bit number of MDCT encoded data obtained by arithmetic encoding, the bit control section 26 selects Huffman encoding as the encoding scheme.

[0316] In this case, the bit control section 26 supplies the MDCT encoded data obtained by the Huffman encoding and the Huffman codebook information to the multiplexing section 27 .

[0317] In contrast, in a case where the bit number of MDCT encoded data obtained by arithmetic encoding is equal to or smaller than the bit number of MDCT encoded data obtained by Huffman encoding and the Huffman codebook information, the bit control section 26 selects arithmetic coding as the encoding scheme.

[0318] In this case, the bit control section 26 supplies the MDCT encoded data obtained by the arithmetic encoding to the multiplexing section 27 .

[0319] In this way, the actual number of bits (code amount) of Huffman coding is compared with the actual number of bits of arithmetic coding, that is, the necessary number of bits in these coding schemes is compared with each other, so that the coding scheme with the smaller necessary number of bits can be selected. Basically, in this case, Huffman coding or arithmetic coding is selected as the coding scheme based on the necessary number of bits at the time of Huffman coding and the necessary number of bits at the time of arithmetic coding, and encoding is performed using the selected coding scheme.

[0320] In step S159 , the bit control section 26 generates a determination flag according to the selection result of the encoding scheme in step S158 and supplies the generated determination flag to the multiplexing section 27 .

[0321] For example, the bit control section 26 generates a determination flag having a value of “1” in the case where Huffman coding is selected as the coding scheme, and generates a determination flag having a value of “0” in the case where arithmetic coding is selected as the coding scheme.

[0322] When the determination flag is generated in this manner, the process proceeds to step S160.

[0323] When the process of step S159 is executed or the process of step S155 is executed, the process of step S160 is executed and the encoding process ends. Figure 27 The process is similar to step S80 in , so the description of the process of step S160 will be omitted.

[0324] As described so far, for a frame whose window sequence is not a LONG_STOP_SEQUENCE, the encoding device 11 selects a coding scheme with a smaller number of necessary bits from Huffman coding and arithmetic coding, and generates a coded bit stream containing MDCT-coded data encoded using the selected coding scheme. In this way, it is possible to select an appropriate coding scheme for each frame, improve coding efficiency, and achieve coding with a higher degree of freedom.

[0325] Further, reference is performed therein Figure 29 In the case of the encoding process described, the decoding device 71 performs the reference Figure 28 The decoding process described.

[0326] As described so far, according to the present technology, by appropriately selecting a coding scheme, it is possible to improve coding efficiency (compression efficiency) and reduce throughput during decoding, compared to the case where only arithmetic coding is used.

[0327] Furthermore, in the second and third embodiments, for frames whose window sequence is not a LONG_STOP_SEQUENCE, even in situations where, for example, the bit rate of the audio signal is high and the sound quality is sufficiently high, or where throughput is more important than sound quality, a suitable encoding scheme can be selected. This allows encoding and decoding to be achieved with a higher degree of freedom. In other words, for example, the throughput during decoding can be controlled more flexibly.

[0328] <Configuration Example of Computer>

[0329] Meanwhile, the series of processes described above can be executed by hardware or by software. In the case of executing the series of processes by software, a program configuring the software is installed in a computer. Here, the types of computers include computers integrated into dedicated hardware, computers that can perform various functions by installing various programs in the computer (for example, a general-purpose personal computer), etc.

[0330] Figure 30 This is a block diagram illustrating an embodiment of a configuration of hardware of a computer that causes a program to execute the series of processing described above.

[0331] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are connected to one another via a bus 504 .

[0332] The input / output interface 505 is also connected to the bus 504 . An input section 506 , an output section 507 , a recording section 508 , a communication section 509 , and a drive 510 are connected to the input / output interface 505 .

[0333] The input unit 506 includes a keyboard, a mouse, a microphone, an imaging element, and the like. The output unit 507 includes a display, a speaker, and the like. The recording unit 508 includes a hard disk, a nonvolatile memory, and the like. The communication unit 509 includes a network interface and the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0334] In the computer of the configuration described above, the CPU 501 loads a program recorded in, for example, the recording section 508 to the RAM 503 via the input / output interface 505 and the bus 504 and executes the program, thereby performing the series of processes described above.

[0335] For example, the program executed by the computer (CPU 501) can be provided by recording the program in a removable recording medium 511 serving as a package medium, etc. Alternatively, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or a digital satellite service.

[0336] In the computer, the program can be installed into the recording section 508 via the input / output interface 505 by loading the removable recording medium 511 into the drive 510. Alternatively, the communication section 509 can receive the program via a wired or wireless transmission medium and install the program into the recording section 508. In another alternative, the program can be installed in advance into the ROM 502 or the recording section 508.

[0337] The program executed by the computer may be a program that executes processing in time series in the order described in this specification or may be a program that executes processing in parallel or at necessary timing such as a calling timing.

[0338] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various changes can be made without departing from the essence of the present technology.

[0339] For example, the present technology can adopt a cloud computing configuration in which a plurality of devices processes one function in a shared or collaborative manner over a network.

[0340] Furthermore, each step described in the above flowchart can be executed not only by one device but also by multiple devices in a shared manner.

[0341] Furthermore, when one step includes a plurality of types of processing, the plurality of types of processing included in one step can be executed not only by one device but also by a plurality of devices in a shared manner.

[0342] Further, the present technology is configured as follows.

[0343] (1) An encoding device comprising:

[0344] a time-frequency transform unit that performs time-frequency transform on the audio signal using a transform window; and

[0345] The encoding unit performs Huffman encoding on spectrum information obtained by time-frequency transform in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length, and performs arithmetic encoding on the spectrum information in a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length.

[0346] (2) The encoding device according to (1), further comprising:

[0347] The multiplexing unit multiplexes the encoded data on the spectrum information and the transform window information indicating the type of the transform window used in the time-frequency transform to generate an encoded bit stream.

[0348] (3) The encoding device according to (1) or (2), wherein

[0349] In a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length, the encoding section encodes the spectrum information by a coding scheme, which is Huffman coding or arithmetic coding.

[0350] (4) The encoding device according to (3), wherein

[0351] The encoding section encodes the spectrum information by an encoding scheme selected based on the necessary number of bits in the encoding process, the bit rate of the audio signal, resource information on the decoding side, or designation information on the encoding scheme.

[0352] (5) The encoding device according to (3) or (4), wherein

[0353] In a case where a transform window length of the transform window is not changed from a small transform window length to a large transform window length, the multiplexing section multiplexes selection information of a coding scheme indicating spectrum information, the encoded data, and the transform window information to generate an encoded bit stream.

[0354] (6) A coding method comprising:

[0355] By encoding device,

[0356] performing a time-frequency transform on the audio signal using a transform window;

[0357] performing Huffman encoding on spectrum information obtained by time-frequency transform in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length; and

[0358] In a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length, arithmetic encoding is performed on the spectrum information.

[0359] (7) A program for causing a computer to execute a process comprising the following steps:

[0360] performing a time-frequency transform on the audio signal using a transform window;

[0361] performing Huffman encoding on spectrum information obtained by time-frequency transform in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length; and

[0362] In a case where the transform window length of the transform window is not changed from a small transform window length to a large transform window length, arithmetic encoding is performed on the spectrum information.

[0363] (8) A decoding device comprising:

[0364] a demultiplexing section that demultiplexes the encoded bit stream and extracts, from the encoded bit stream, transform window information indicating the type of transform window used in time-frequency transform of the audio signal and encoded data on spectrum information obtained by the time-frequency transform; and

[0365] The decoding section decodes the encoded data by a decoding scheme corresponding to Huffman coding when the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length.

[0366] (9) The decoding device according to (8), wherein

[0367] In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the decoding section decodes the encoded data by a decoding scheme corresponding to arithmetic coding.

[0368] (10) The decoding device according to (8), wherein

[0369] In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the decoding unit decodes the encoded data by a decoding scheme corresponding to the encoding scheme, the encoding scheme is Huffman encoding or arithmetic coding, and the encoding scheme is indicated by the selection information extracted from the encoded bit stream.

[0370] (11) A decoding method comprising:

[0371] By decoding device,

[0372] demultiplexing the encoded bit stream and extracting, from the encoded bit stream, transform window information indicating the type of transform window used in time-frequency transform of the audio signal and encoded data on spectrum information obtained by the time-frequency transform; and

[0373] In a case where the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the encoded data is decoded by a decoding scheme corresponding to Huffman coding.

[0374] (12) A program for causing a computer to execute a process comprising the following steps:

[0375] demultiplexing the encoded bit stream and extracting, from the encoded bit stream, transform window information indicating the type of transform window used in time-frequency transform of the audio signal and encoded data on spectrum information obtained by the time-frequency transform; and

[0376] In a case where the transform window indicated by the transform window information is a transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the encoded data is decoded by a decoding scheme corresponding to Huffman coding.

[0377] [Reference Number List]

[0378] 11 encoding device, 21 time-frequency conversion unit, 24 encoding scheme selection unit, 26 bit control unit, 27 multiplexing unit, 31 Huffman encoding unit, 32 arithmetic encoding unit, 71 decoding device, 81 acquisition unit, 82 demultiplexing unit, 83 decoding scheme selection unit, 91 Huffman decoding unit, 92 arithmetic decoding unit.

Claims

1. An encoding device, comprising: a time-frequency transform unit that performs time-frequency transform on the audio signal using a transform window to obtain MDCT coefficients as spectrum information; a normalization unit, normalizing the MDCT coefficients based on a normalization parameter to obtain normalized MDCT coefficients; a quantization unit, configured to quantize the normalized MDCT coefficients to obtain quantized MDCT coefficients; an encoding unit that performs Huffman encoding on the quantized MDCT coefficients when a transform window length of the transform window is changed from a small transform window length to a large transform window length, and performs arithmetic encoding and Huffman encoding on the quantized MDCT coefficients when the transform window length of the transform window is not changed from the small transform window length to the large transform window length; a coding scheme selection unit that supplies the quantized MDCT coefficients to the Huffman coding unit and the arithmetic coding unit without changing the transform window length from a small transform window length to a large transform window length, a bit control unit that compares the actual number of bits of Huffman coding with the actual number of bits of arithmetic coding to select a coding scheme with a smaller number of necessary bits between the Huffman coding and the arithmetic coding; and a multiplexing unit that multiplexes the encoded data on the spectrum information and the transform window information indicating the type of the transform window used in the time-frequency transform to generate an encoded bit stream, wherein The multiplexing unit multiplexes the selection information of the coding scheme indicating the spectrum information, the encoded data, and the transform window information to generate a coded bit stream without changing the transform window length of the transform window from the small transform window length to the large transform window length.

2. The encoding device according to claim 1, wherein The encoding section encodes the spectrum information by an encoding scheme selected based on a necessary number of bits in an encoding process, a bit rate of the audio signal, resource information on a decoding side, or designation information on the encoding scheme.

3. A coding method comprising: By encoding device, Performing a time-frequency transform on the audio signal using a transform window to obtain MDCT coefficients as spectral information; Normalizing the MDCT coefficients based on a normalization parameter to obtain normalized MDCT coefficients; quantizing the normalized MDCT coefficients to obtain quantized MDCT coefficients; performing Huffman coding on the quantized MDCT coefficients in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length; performing arithmetic coding and Huffman coding on the quantized MDCT coefficients without changing the transform window length of the transform window from the small transform window length to the large transform window length; as well as supplying the quantized MDCT coefficients to a Huffman coding section and an arithmetic coding section without changing from a small transform window length to a large transform window length, comparing the actual number of bits of Huffman coding with the actual number of bits of arithmetic coding to select a coding scheme with a smaller number of necessary bits between the Huffman coding and the arithmetic coding, and The coded data on the spectrum information and the transform window information indicating the type of the transform window used in the time-frequency transform are multiplexed to generate a coded bit stream, wherein: Without changing the transform window length of the transform window from the small transform window length to the large transform window length, selection information of the encoding scheme indicating the spectrum information, encoding data, and transform window information are multiplexed to generate an encoded bit stream.

4. A storage medium comprising a program which, when executed by a computer including the storage medium, causes the computer to execute a process comprising the following steps: Performing a time-frequency transform on the audio signal using a transform window to obtain MDCT coefficients as spectral information; Normalizing the MDCT coefficients based on a normalization parameter to obtain normalized MDCT coefficients; quantizing the normalized MDCT coefficients to obtain quantized MDCT coefficients; performing Huffman coding on the quantized MDCT coefficients in a case where a transform window length of the transform window is changed from a small transform window length to a large transform window length; performing arithmetic coding and Huffman coding on the quantized MDCT coefficients without changing the transform window length of the transform window from the small transform window length to the large transform window length; and supplying the quantized MDCT coefficients to a Huffman coding section and an arithmetic coding section without changing from a small transform window length to a large transform window length, and, comparing the actual number of bits of Huffman coding with the actual number of bits of arithmetic coding to select a coding scheme with a smaller number of necessary bits between the Huffman coding and the arithmetic coding, and The coded data on the spectrum information and the transform window information indicating the type of the transform window used in the time-frequency transform are multiplexed to generate a coded bit stream, wherein: Without changing the transform window length of the transform window from the small transform window length to the large transform window length, selection information of the encoding scheme indicating the spectrum information, encoding data, and transform window information are multiplexed to generate an encoded bit stream.

5. A decoding device comprising: a demultiplexing section that demultiplexes a coded bit stream and extracts, from the coded bit stream, transform window information indicating the type of transform window used in time-frequency transform of an audio signal and coded data on MDCT coefficients as spectral information obtained by the time-frequency transform; and a decoding unit that decodes the encoded data using a decoding scheme corresponding to Huffman coding when the transform window indicated by the transform window information is the transform window selected when the transform window length is changed from a small transform window length to a large transform window length, Without changing from a small transform window length to a large transform window length, the quantized MDCT coefficients are supplied to the Huffman coding section and the arithmetic coding section, wherein the MDCT coefficients are normalized based on a normalization parameter to obtain normalized MDCT coefficients, and the normalized MDCT coefficients are quantized to obtain the quantized MDCT coefficients, and wherein, The selection between the Huffman coding and the arithmetic coding is based on comparing the actual number of bits of the Huffman coding with the actual number of bits of the arithmetic coding, and selecting the coding scheme with the smaller number of necessary bits, wherein In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the decoding section decodes the encoded data by a decoding scheme corresponding to arithmetic coding, and wherein, In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the decoding unit decodes the encoded data by a decoding scheme corresponding to a coding scheme, which is the Huffman coding or arithmetic coding, and the coding scheme is indicated by selection information extracted from the encoded bit stream.

6. A decoding method comprising: By decoding device, demultiplexing a coded bit stream and extracting, from the coded bit stream, transform window information indicating the type of transform window used in time-frequency transform of an audio signal and coded data on MDCT coefficients as spectral information obtained by the time-frequency transform; and In a case where the transform window indicated by the transform window information is the transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the encoded data is decoded by a decoding scheme corresponding to Huffman coding, wherein Without changing from a small transform window length to a large transform window length, the quantized MDCT coefficients are supplied to the Huffman coding section and the arithmetic coding section, wherein the MDCT coefficients are normalized based on a normalization parameter to obtain normalized MDCT coefficients, and the normalized MDCT coefficients are quantized to obtain the quantized MDCT coefficients, and wherein, The selection between the Huffman coding and the arithmetic coding is based on comparing the actual number of bits of the Huffman coding with the actual number of bits of the arithmetic coding, and selecting the coding scheme with the smaller number of necessary bits, wherein In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the encoded data is decoded by a decoding scheme corresponding to arithmetic coding, and wherein, In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the encoded data is decoded by a decoding scheme corresponding to a coding scheme, the coding scheme being the Huffman coding or arithmetic coding, and the coding scheme being indicated by selection information extracted from the encoded bit stream.

7. A storage medium comprising a program which, when executed by a computer comprising the storage medium, causes the computer to execute a process comprising the following steps: demultiplexing a coded bit stream and extracting, from the coded bit stream, transform window information indicating the type of transform window used in time-frequency transform of an audio signal and coded data on MDCT coefficients as spectral information obtained by the time-frequency transform; and In a case where the transform window indicated by the transform window information is the transform window selected when the transform window length is changed from a small transform window length to a large transform window length, the encoded data is decoded by a decoding scheme corresponding to Huffman coding, wherein Without changing from a small transform window length to a large transform window length, the quantized MDCT coefficients are supplied to the Huffman coding section and the arithmetic coding section, wherein The MDCT coefficients are normalized based on a normalization parameter to obtain normalized MDCT coefficients, and the normalized MDCT coefficients are quantized to obtain quantized MDCT coefficients, and wherein, The selection between the Huffman coding and the arithmetic coding is based on comparing the actual number of bits of the Huffman coding with the actual number of bits of the arithmetic coding, and selecting the coding scheme with the smaller number of necessary bits, wherein In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the encoded data is decoded by a decoding scheme corresponding to arithmetic coding, and wherein, In a case where the transform window indicated by the transform window information is not the transform window selected when the transform window length is changed from the small transform window length to the large transform window length, the encoded data is decoded by a decoding scheme corresponding to a coding scheme, the coding scheme being the Huffman coding or arithmetic coding, and the coding scheme being indicated by selection information extracted from the encoded bit stream.

Citation Information

Patent Citations

  • Audio Encoder and Decoder

    US20130282383A1