Encoding device and method, decoding device and method, and program

The encoding device improves 3D Audio coding efficiency by prioritizing audio signals and inserting silence data to meet real-time constraints, enabling effective transmission of multiple audio objects with minimal sound quality loss.

JP7827065B2Active Publication Date: 2026-03-10SONY GROUP CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing 3D Audio coding technologies, such as MPEG-H 3D Audio, require improved coding efficiency to efficiently transmit and stream live performances in real-time, as they handle multiple audio objects with metadata for position and gain, necessitating a more efficient coding technique.

Method used

An encoding device and method that generates priority information based on audio signals and metadata, performs time-frequency transformation, and quantizes MDCT coefficients in order of priority, with additional encoding for high-priority signals, and inserts pre-generated silence data if real-time limits are not met.

Benefits of technology

Enhances coding efficiency while maintaining real-time operation, allowing for the transmission of a larger number of audio objects with minimized sound quality degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827065000001
    Figure 0007827065000001
  • Figure 0007827065000002
    Figure 0007827065000002
  • Figure 0007827065000003
    Figure 0007827065000003
Patent Text Reader

Abstract

The present technology pertains to an encoding device and method, a decoding device and method, and a program, with which it is possible to improve encoding efficiency in a state in which real-time operation is maintained. The encoding device is provided with: a priority degree information generation unit for generating, on the basis of an audio signal and / or metadata for an audio signal, priority degree information indicating the priority degree of the audio signal; a time frequency conversion unit for performing time frequency conversion on the audio signal and generating an MDCT coefficient; and a bit allocation unit for performing, on a plurality of audio signals, quantization of the MDCT coefficient of the audio signal, in order from the audio signal having the highest priority degree indicated by the priority degree signal. This technology is applicable to an encoding device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to an encoding device and method, a decoding device and method, and a program, and in particular to an encoding device and method, a decoding device and method, and a program that enable improvement in encoding efficiency while maintaining real-time operation. [Background technology]

[0002] Conventionally, encoding technologies such as the international standard MPEG (Moving Picture Experts Group)-D USAC (Unified Speech and Audio Coding) standard and the MPEG-H 3D Audio standard, which uses the MPEG-D USAC standard as a core coder, have been known (see, for example, Non-Patent Documents 1 to 3). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] ISO / IEC 23003-3, MPEG-D USAC [Non-patent document 2] ISO / IEC 23008-3, MPEG-H 3D Audio [Non-patent document 3] ISO / IEC 23008-3:2015 / AMENDMENT3, MPEG-H 3D Audio Phase 2 Summary of the Invention [Problem to be solved by the invention]

[0004] 3D Audio, which is handled by standards such as MPEG-H 3D Audio, has metadata for each object, such as horizontal and vertical angles that indicate the position of the sound material (object), distance, and gain for the object, and can reproduce the three-dimensional direction, distance, and spread of sound.As a result, 3D Audio enables audio playback with a more realistic feel than conventional stereo playback.

[0005] However, in order to transmit the data of the many objects realized by 3D Audio, a coding technique that can compress and decode more audio channels efficiently and quickly is required. In other words, there is a demand for improved coding efficiency.

[0006] Furthermore, in order to live stream live performances and concerts using 3D Audio, it is necessary to improve coding efficiency while achieving real-time performance.

[0007] The present technology has been made in view of such circumstances, and makes it possible to improve coding efficiency while maintaining real-time operation. [Means for solving the problem]

[0008] An encoding device according to a first aspect of the present technology includes: a priority information generation unit that generates priority information indicating a priority of an audio signal based on at least one of an audio signal and metadata of the audio signal; a time-frequency transformation unit that performs a time-frequency transformation on the audio signal to generate MDCT coefficients; and a bit allocation unit that quantizes the MDCT coefficients of a plurality of the audio signals in order from the audio signal with the highest priority indicated by the priority information.

[0009] An encoding method or program according to a first aspect of the present technology includes the steps of generating priority information indicating a priority of an audio signal based on at least one of an audio signal and metadata of the audio signal, performing a time-frequency transform on the audio signal, generating MDCT coefficients, and quantizing the MDCT coefficients of a plurality of the audio signals in order from the audio signal with the highest priority indicated by the priority information.

[0010] In a first aspect of the present technology, priority information indicating a priority of an audio signal is generated based on at least one of an audio signal and metadata of the audio signal, a time-frequency transform is performed on the audio signal to generate MDCT coefficients, and the MDCT coefficients of a plurality of the audio signals are quantized in order from the audio signal with the highest priority indicated by the priority information.

[0011] A decoding device according to a second aspect of the present technology includes a decoding unit that acquires encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in order of priority, starting from the audio signals with the highest priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals, and decodes the encoded audio signals.

[0012] A decoding method or program according to a second aspect of the present technology includes the steps of: obtaining encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in order of priority, starting from the audio signals with the highest priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals; and decoding the encoded audio signals.

[0013] In a second aspect of the present technology, for a plurality of audio signals, encoded audio signals are obtained by quantizing MDCT coefficients of the audio signals in order of priority, starting from the audio signals with the highest priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals, and the encoded audio signals are decoded.

[0014] A coding device according to a third aspect of the present technology includes: a coding unit that codes an audio signal and generates an encoded audio signal; a buffer that holds a bitstream consisting of the encoded audio signal for each frame; and an insertion unit that, if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, inserts pre-generated encoded silence data into the bitstream as the encoded audio signal for the frame to be processed.

[0015] An encoding method or program according to a third aspect of the present technology includes the steps of: encoding an audio signal to generate an encoded audio signal; storing a bitstream consisting of the encoded audio signal for each frame in a buffer; and, if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, inserting pre-generated encoded silence data into the bitstream as the encoded audio signal for the frame to be processed.

[0016] In a third aspect of the present technology, an audio signal is encoded to generate an encoded audio signal, a bitstream consisting of the encoded audio signal for each frame is stored in a buffer, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, pre-generated encoded silence data is inserted into the bitstream as the encoded audio signal for the frame to be processed.

[0017] A decoding device according to a fourth aspect of the present technology includes a decoding unit that encodes an audio signal to generate an encoded audio signal, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, acquires the bitstream obtained by inserting pre-generated encoded silence data into a bitstream consisting of the encoded audio signals for each frame as the encoded audio signal for the frame to be processed, and decodes the encoded audio signal.

[0018] A decoding method or program according to a fourth aspect of the present technology includes a step of encoding an audio signal to generate an encoded audio signal, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, acquiring the bitstream obtained by inserting pre-generated encoded silence data as the encoded audio signal for the frame to be processed into a bitstream consisting of the encoded audio signals for each frame, and decoding the encoded audio signal.

[0019] In a fourth aspect of the present technology, an audio signal is encoded to generate an encoded audio signal, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, a bitstream obtained by inserting pre-generated encoded silence data as the encoded audio signal for the frame to be processed into a bitstream consisting of the encoded audio signals for each frame is obtained, and the encoded audio signal is decoded.

[0020] A coding device according to a fifth aspect of the present technology includes: a time-frequency transform unit that performs time-frequency transform on an audio signal of an object to generate MDCT coefficients; a psychoacoustic parameter calculation unit that calculates psychoacoustic parameters based on the MDCT coefficients and setting information related to a masking threshold for the object; and a bit allocation unit that performs bit allocation processing based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients.

[0021] An encoding method or program according to a fifth aspect of the present technology includes the steps of performing a time-frequency transform on an audio signal of an object to generate MDCT coefficients, calculating psychoacoustic parameters based on the MDCT coefficients and setting information related to a masking threshold for the object, and performing bit allocation processing based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients.

[0022] In a fifth aspect of the present technology, a time-frequency transform is performed on an audio signal of an object to generate MDCT coefficients, psychoacoustic parameters are calculated based on the MDCT coefficients and setting information related to a masking threshold for the object, and bit allocation processing is performed based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients. [Brief explanation of the drawings]

[0023] [Figure 1] FIG. 1 illustrates an example of the configuration of an encoder. [Figure 2] FIG. 2 illustrates an example of the configuration of an object audio encoding unit. [Figure 3] 10 is a flowchart illustrating an encoding process. [Figure 4] 10 is a flowchart illustrating a bit allocation process. [Figure 5] FIG. 10 is a diagram illustrating an example of the syntax of metadata Config. [Figure 6] FIG. 10 is a diagram illustrating an example of the configuration of a decoder. [Figure 7] FIG. 10 is a diagram illustrating a configuration example of an unpacking / decoding unit. [Figure 8] 10 is a flowchart illustrating a decoding process. [Figure 9] 10 is a flowchart illustrating a selective decoding process. [Figure 10] FIG. 2 illustrates an example of the configuration of an object audio encoding unit. [Figure 11] FIG. 1 illustrates an example of the configuration of a content distribution system. [Figure 12] FIG. 10 is a diagram illustrating an example of input data. [Figure 13] FIG. 10 is a diagram illustrating the calculation of a context. [Figure 14] FIG. 1 illustrates an example of the configuration of an encoder. [Figure 15] FIG. 2 illustrates an example of the configuration of an object audio encoding unit. [Figure 16] FIG. 2 illustrates an example of the configuration of an initialization unit. [Figure 17] 10A and 10B are diagrams illustrating an example of progress information and a determination of whether or not processing can be completed. [Figure 18] FIG. 1 is a diagram illustrating an example of a bit stream made up of coded data. [Figure 19] FIG. 10 is a diagram illustrating an example of the syntax of encoded data. [Figure 20] FIG. 10 is a diagram illustrating an example of extended data. [Figure 21] FIG. 2 is a diagram illustrating segment data. [Figure 22] FIG. 10 is a diagram illustrating an example of the configuration of AudioPreRoll(). [Figure 23] 10 is a flowchart illustrating an initialization process. [Figure 24] 10 is a flowchart illustrating an encoding process. [Figure 25] 10 is a flowchart illustrating an encoded Mute data insertion process. [Figure 26] FIG. 10 is a diagram illustrating a configuration example of an unpacking / decoding unit. [Figure 27] 10 is a flowchart illustrating a decoding process. [Figure 28] FIG. 1 illustrates an example of the configuration of an encoder. [Figure 29] FIG. 2 illustrates an example of the configuration of an object audio encoding unit. [Figure 30] 10 is a flowchart illustrating an encoding process. [Figure 31] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0025] First Embodiment About this technology This technology performs encoding processing that takes into account the importance of objects (audio), thereby improving encoding efficiency while maintaining real-time operation and increasing the number of objects that can be transmitted.

[0026] For example, to realize live streaming, the encoding process must be performed in real time. In other words, if f frames of audio are to be transmitted per second, the encoding of one frame and the output of the bitstream must be completed within 1 / f seconds.

[0027] In order to achieve the goal of performing the encoding process in real time, the following approach is effective.

[0028] · The encoding process is carried out in stages. First, the minimum necessary encoding is completed, and then additional encoding processing is performed with improved encoding efficiency. If the additional encoding processing is not completed after a predetermined time limit has elapsed, the processing is terminated at that point and the result of the encoding processing at the previous stage is output. Furthermore, if the minimum required encoding has not been completed when the specified time limit has elapsed, the process is terminated and a bitstream of Mute data prepared in advance is output.

[0029] When audio signals of multiple channels or multiple objects are played back simultaneously, the sounds played back by those audio signals include sounds that are important compared to other sounds and sounds that are not so important. For example, an unimportant sound is a sound that does not cause a listener to feel uncomfortable even if a particular sound is not played back from the entire sound.

[0030] If additional coding processing that improves coding efficiency is performed using a processing order that does not take into account the importance of the audio, i.e., the importance of the channel or object, processing may be terminated even if the audio is important, resulting in deterioration of sound quality.

[0031] Therefore, this technology performs additional encoding processing that increases the encoding efficiency in order of audio importance, thereby making it possible to improve the encoding efficiency of the entire content while maintaining real-time operation.

[0032] In this way, the more important the audio, the more additional encoding processing is completed, while the less important the audio is, the less additional encoding processing is completed and only the minimum necessary encoding is performed, thereby improving the encoding efficiency of the entire content and increasing the number of objects that can be transmitted.

[0033] As described above, in the present technology, when encoding the audio signals of each channel constituting a multi-channel signal and the audio signals of objects, additional encoding processes with improved encoding efficiency are performed in descending order of priority for the audio signals of each channel and each object, thereby improving the encoding efficiency of the entire content in real-time processing.

[0034] Note that the following describes the case where the audio signal of an object is coded according to the MPEG-H standard, but similar processing is performed when coding is performed according to the MPEG-H standard including the audio signal of a channel, or when coding is performed using other methods.

[0035] <Encoder configuration example> FIG. 1 is a diagram showing an example of the configuration of an embodiment of an encoder to which the present technology is applied.

[0036] The encoder 11 shown in FIG. 1 is made up of a signal processing device such as a computer that functions as an encoder (encoding device).

[0037] In the example shown in Fig. 1, audio signals of N objects and metadata of these N objects are input to the encoder 11 and encoded according to the MPEG-H standard. Note that #0 to #N-1 in Fig. 1 represent object numbers indicating the N objects.

[0038] The encoder 11 includes an object metadata encoding unit 21, an object audio encoding unit 22, and a packing unit .

[0039] The object metadata encoding unit 21 encodes the metadata of each of the N objects supplied thereto in accordance with the MPEG-H standard, and supplies the resulting encoded metadata to a packing unit 23 .

[0040] For example, the metadata of an object includes object position information indicating the position of the object in three-dimensional space, a priority value indicating the priority (degree of importance) of the object, and a gain value indicating a gain for gain compensation of the audio signal of the object. In particular, in this example, the metadata includes at least the priority value.

[0041] Here, the object position information includes, for example, a horizontal angle (Azimuth), a vertical angle (Elevation), and a distance (Radius).

[0042] The horizontal and vertical angles are the horizontal and vertical angles that indicate the position of the object as seen from the reference listening position in three-dimensional space. The radius indicates the distance from the reference listening position to the object, indicating the position of the object in three-dimensional space. This object position information can be said to be information that indicates the sound source position of the sound based on the object's audio signal.

[0043] Additionally, the metadata of an object may include parameters for spread processing to widen the sound image of the object.

[0044] The object audio encoding unit 22 encodes the audio signal of each of the N supplied objects in accordance with the MPEG-H standard based on the Priority value included in the metadata of each supplied object, and supplies the resulting encoded audio signal to a packing unit 23.

[0045] The packing unit 23 packs the encoded metadata supplied from the object metadata encoding unit 21 and the encoded audio signal supplied from the object audio encoding unit 22, and outputs the resulting encoded bitstream.

[0046] <Configuration example of object audio encoding unit> The object audio encoding unit 22 is configured as shown in FIG. 2, for example.

[0047] In the example of FIG. 2, the object audio encoding unit 22 includes a priority information generation unit 51 , a time-frequency conversion unit 52 , an auditory psycho-parameter calculation unit 53 , a bit allocation unit 54 , and an encoding unit 55 .

[0048] The priority information generation unit 51 generates priority information indicating the priority of each object, i.e., the priority of the audio signal, based on at least one of the audio signal of each supplied object and the Priority value included in the metadata of each supplied object, and supplies the generated priority information to the bit allocation unit 54.

[0049] For example, the priority information generation unit 51 analyzes the level of priority of the audio signal of an object based on the sound pressure and spectral shape of the audio signal, correlation of spectral shapes between audio signals of a plurality of objects or channels, etc. Then, the priority information generation unit 51 generates priority information based on the analysis results.

[0050] Furthermore, for example, the metadata of an MPEG-H object contains a Priority value, which is a parameter indicating the priority of the object, as a 3-bit integer ranging from 0 to 7, and the larger the Priority value, the higher the priority of the object.

[0051] The priority value may be set intentionally by the content creator, or it may be set automatically by the application that generates the metadata by analyzing the audio signal of each object. Also, without the intention of the content creator or analysis of the audio signal, the priority value may be set to a fixed value such as the highest priority "7" as the application default.

[0052] Therefore, when the priority information generating unit 51 generates priority information for an object (audio signal), it may use only the analysis results of the audio signal without using the priority value, or it may use both the priority value and the analysis results.

[0053] For example, when both the priority value and the analysis result are used, even if the analysis result of the audio signal is the same, an object having a larger (higher) priority value can be given a higher priority.

[0054] The time-frequency transform unit 52 performs time-frequency transform using MDCT (Modified Discrete Cosine Transform) on the audio signal of each supplied object.

[0055] The time-frequency transform unit 52 supplies the MDCT coefficients, which are frequency spectrum information of each object obtained by the time-frequency transform, to the bit allocation unit 54 .

[0056] The psychoacoustic parameter calculation unit 53 calculates psychoacoustic parameters for taking into account human auditory characteristics (auditory masking) based on the supplied audio signal of each object, and supplies the calculated psychoacoustic parameters to the bit allocation unit .

[0057] The bit allocation unit 54 performs bit allocation processing based on the priority information supplied from the priority information generation unit 51, the MDCT coefficients supplied from the time-frequency conversion unit 52, and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53.

[0058] In the bit allocation process, the quantization bits and quantization noise for each scale factor band are calculated and evaluated, and bit allocation is performed based on a psychoacoustic model. Then, based on the bit allocation results, the MDCT coefficients for each scale factor band are quantized to obtain quantized MDCT coefficients.

[0059] The bit allocation unit 54 supplies the quantized MDCT coefficients for each scale factor band of each object obtained in this manner to the encoding unit 55 as the quantization results of each object, more specifically, as the quantization results of the MDCT coefficients of each object.

[0060] Here, the scale factor band is a band (frequency band) obtained by bundling a plurality of sub-bands (here, MDCT resolution) of a predetermined bandwidth based on the characteristics of human hearing.

[0061] By the bit allocation process described above, some of the quantization bits of the scale factor bands where the quantization noise generated in the quantization of the MDCT coefficients is masked and not perceptible are allocated (diverted) to the scale factor bands where the quantization noise is more perceptible. This suppresses deterioration of sound quality overall and enables efficient quantization. In other words, it is possible to improve coding efficiency.

[0062] For an object for which the quantized MDCT coefficients could not be obtained within the time limit for real-time processing, the bit allocation unit 54 supplies the encoding unit 55 with pre-prepared Mute data as the quantization result of that object.

[0063] The mute data is zero data indicating the value "0" of the MDCT coefficient of each scale factor band. More specifically, the quantized value of the mute data, i.e., the quantized MDCT coefficient of MDCT coefficient "0", is output to the encoding unit 55. Note that although the mute data is output to the encoding unit 55 here, mute information indicating whether the quantization result (quantized MDCT coefficient) is mute data or not may be supplied to the encoding unit 55 instead of supplying the mute data. In this case, the encoding unit 55 switches between performing normal encoding processing or directly encoding the quantized MDCT coefficient of MDCT coefficient "0" according to the mute information. Furthermore, instead of encoding the quantized MDCT coefficient of MDCT coefficient "0", it may use encoded data of MDCT coefficient "0" prepared in advance.

[0064] Furthermore, the bit allocation unit 54 supplies Mute information indicating whether or not the quantization result (quantized MDCT coefficient) is Mute data, for example, for each object, to the packing unit 23. The packing unit 23 stores the Mute information supplied from the bit allocation unit 54 in an ancillary area or the like of the encoded bit stream.

[0065] The encoding unit 55 encodes the quantized MDCT coefficients for each scale factor band of each object supplied from the bit allocation unit 54, and supplies the resulting encoded audio signal to the packing unit 23.

[0066] <Description of the encoding process> Next, a description will be given of the operation of the encoder 11. That is, the encoding process by the encoder 11 will be described below with reference to the flowchart in FIG.

[0067] In step S11, the object metadata encoding unit 21 encodes the metadata of each object supplied thereto, and supplies the resulting encoded metadata to the packing unit 23.

[0068] In step S12, the priority information generation unit 51 generates priority information for each object based on at least one of the audio signal of each supplied object and the Priority value of the metadata of each supplied object, and supplies it to the bit allocation unit 54.

[0069] In step S13, the time-frequency transform unit 52 performs time-frequency transform using MDCT on the audio signal of each supplied object, and supplies the resulting MDCT coefficients for each scale factor band to the bit allocation unit .

[0070] In step S14, the psychoacoustic parameter calculation unit 53 calculates psychoacoustic parameters based on the audio signal of each object supplied, and supplies the calculated psychoacoustic parameters to the bit allocation unit .

[0071] In step S15, the bit allocation unit 54 performs bit allocation processing based on the priority information supplied from the priority information generation unit 51, the MDCT coefficients supplied from the time-frequency conversion unit 52, and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53.

[0072] The bit allocation unit 54 supplies the quantized MDCT coefficients obtained by the bit allocation process to the encoding unit 55, and also supplies Mute information to the packing unit 23. The bit allocation process will be described in detail later.

[0073] In step S16, the encoding unit 55 encodes the quantized MDCT coefficients supplied from the bit allocation unit 54, and supplies the resulting encoded audio signal to the packing unit 23.

[0074] For example, the encoding unit 55 performs context-based arithmetic coding on the quantized MDCT coefficients, and outputs the coded quantized MDCT coefficients as a coded audio signal to the packing unit 23. Note that the coding method is not limited to arithmetic coding. For example, coding may be performed using Huffman coding or another coding method.

[0075] In step S17 , the packing unit 23 packs the encoded metadata supplied from the object metadata encoding unit 21 and the encoded audio signal supplied from the encoding unit 55 .

[0076] At this time, the packing unit 23 stores the mute information supplied from the bit allocation unit 54 in an ancillary area or the like of the coded bit stream.

[0077] Then, the packing unit 23 outputs the coded bit stream obtained by packing, and the coding process ends.

[0078] In this way, the encoder 11 generates priority information based on the audio signal and priority value of the object, and performs bit allocation processing using the priority information. This improves the encoding efficiency of the entire content in real-time processing, and enables the transmission of data for a larger number of objects.

[0079] <Explanation of Bit Allocation Processing> Next, the bit allocation process corresponding to the process in step S15 in FIG. 3 will be described with reference to the flowchart in FIG.

[0080] In step S41, based on the priority information supplied from the priority information generating unit 51, the bit allocating unit 54 sets the order of processing of each object (processing order) in descending order of priority indicated by the priority information.

[0081] In this example, of the total N objects, the processing order of the object with the highest priority is set to "0," and the processing order of the object with the lowest priority is set to "N-1." Note that the processing order setting is not limited to this; for example, the processing order of the object with the highest priority may be set to "1," and the processing order of the object with the lowest priority may be set to "N," or priorities may be represented by symbols other than numbers.

[0082] Thereafter, the minimum necessary quantization process, that is, the minimum necessary encoding process is performed on objects in order from the highest priority object.

[0083] That is, in step S42, the bit allocation unit 54 sets the processing target ID indicating the object to be processed to "0".

[0084] The value of this processing target ID is updated by incrementing it by 1 starting from "0." If the value of the processing target ID is n, the object indicated by that processing target ID is the nth object in the processing order set in step S41.

[0085] Therefore, in the bit allocation section 54, each object is processed in the processing order set in step S41.

[0086] In step S43, the bit allocation unit 54 determines whether the value of the processing target ID is less than N or not.

[0087] If it is determined in step S43 that the value of the processing target ID is less than N, that is, if quantization processing has not yet been performed on all objects, the processing of step S44 is performed.

[0088] That is, in step S44, the bit allocation unit 54 performs the minimum necessary quantization process on the MDCT coefficients for each scale factor band of the object to be processed indicated by the processing object ID.

[0089] Here, the minimum necessary quantization process is the first quantization process that is performed before the bit allocation loop process.

[0090] Specifically, the bit allocation unit 54 calculates and evaluates the quantization bits and quantization noise for each scale factor band based on the psychoacoustic parameters and MDCT coefficients, thereby determining the target number of bits (quantization bit count) for the quantized MDCT coefficients for each scale factor band.

[0091] The bit allocation unit 54 quantizes the MDCT coefficients for each scale factor band so that the quantized MDCT coefficients for each scale factor band become data within the target number of quantization bits, thereby obtaining quantized MDCT coefficients.

[0092] Furthermore, the bit allocation unit 54 generates and stores Mute information indicating that the quantization result is not Mute data for the object to be processed.

[0093] In step S45, the bit allocation unit 54 determines whether or not the time is within a predetermined time limit for real-time processing.

[0094] For example, if a predetermined time has elapsed since the bit allocation process started, it is determined that the time limit has not been reached.

[0095] This time limit is a threshold value set (determined) by the bit allocation unit 54, taking into consideration the processing time required by the encoding unit 55 and packing unit 23 downstream of the bit allocation unit 54, for example, so that the encoded bit stream can be output (distributed) in real time, i.e., so that the encoding process can be performed in real time.

[0096] Furthermore, this time limit may be dynamically changed based on the results of the bit allocation process so far, such as the values ​​of the quantized MDCT coefficients of the object obtained by the bit allocation unit 54 so far.

[0097] If it is determined in step S45 that the time is within the time limit, then the process proceeds to step S46.

[0098] In step S46, the bit allocation unit 54 saves (holds) the quantized MDCT coefficients obtained by the process in step S44 as the quantization result of the object to be processed, and adds "1" to the value of the object ID to be processed, thereby setting a new object that has not yet undergone the minimum necessary quantization process as the next object to be processed.

[0099] After the process of step S46 is performed, the process returns to step S43, and the above-described process is repeated. That is, the minimum necessary quantization process is performed on the new object to be processed.

[0100] In this way, in steps S43 to S46, the minimum necessary quantization process is performed on each object in descending order of priority, thereby improving coding efficiency.

[0101] If it is determined in step S45 that the time limit has not been reached, that is, if the time limit has been reached, the minimum necessary quantization process for each object is terminated, and the process then proceeds to step S47. In other words, in this case, for objects that were not selected as processing targets, the process is terminated without completing the minimum necessary quantization process.

[0102] In step S47, the bit allocation unit 54 saves (retains) the quantization values ​​of the Mute data prepared in advance for objects that were not processed in the above-mentioned steps S43 to S46, i.e., objects for which the minimum necessary quantization processing has not been completed, as the quantization results for each of those objects.

[0103] That is, in step S47, for an object for which the minimum necessary quantization process has not been completed, the quantized value of the Mute data is used as the quantization result of that object.

[0104] Furthermore, for objects for which the minimum necessary quantization process has not been completed, the bit allocation unit 54 generates and holds Mute information indicating that the quantization result is Mute data.

[0105] After the process of step S47 is performed, the process proceeds to step S54.

[0106] If it is determined in step S43 that the value of the processing target ID is not less than N, that is, if the minimum necessary quantization process has been completed for all objects within the time limit, the process proceeds to step S48.

[0107] In step S48, the bit allocation unit 54 sets the processing target ID indicating the object to be processed to "0." As a result, the objects are again designated as processing targets in descending order of priority, and subsequent processing is carried out.

[0108] In step S49, the bit allocation unit 54 determines whether the value of the processing target ID is less than N or not.

[0109] If it is determined in step S49 that the value of the processing target ID is less than N, that is, if the additional quantization process (additional encoding process) has not yet been performed on all objects, the process proceeds to step S50.

[0110] In step S50, the bit allocation unit 54 performs an additional quantization process, i.e., an additional bit allocation loop process, once on the MDCT coefficients for each scale factor band of the object to be processed indicated by the processing object ID, and updates and saves the quantization results as necessary.

[0111] Specifically, the bit allocation unit 54 recalculates and reevaluates the quantization bits and quantization noise for each scale factor band based on the psychoacoustic parameters and the quantized MDCT coefficients, which are the quantization results for each scale factor band of the object obtained by the previous processes, such as the minimum necessary quantization process, etc. As a result, a new target number of quantization bits for the quantized MDCT coefficients is determined for each scale factor band.

[0112] The bit allocation unit 54 re-quantizes the MDCT coefficients for each scale factor band so that the quantized MDCT coefficients for each scale factor band become data within the target number of quantization bits, thereby obtaining quantized MDCT coefficients.

[0113] If the process of step S50 results in quantized MDCT coefficients of higher quality with less quantization noise than the quantized MDCT coefficients held as the quantization results of the object, the bit allocation unit 54 replaces the quantized MDCT coefficients held so far with the newly obtained quantized MDCT coefficients and stores them. In other words, the held quantized MDCT coefficients are updated.

[0114] In step S51, the bit allocation unit 54 determines whether or not the time is within a predetermined time limit for real-time processing.

[0115] For example, in step S51, similarly to step S45, if a predetermined time has elapsed since the start of the bit allocation process, it is determined that the time is not within the time limit.

[0116] The time limit in step S51 may be the same as that in step S45, or, as described above, may be dynamically changed depending on the results of the bit allocation processing performed so far, i.e., the minimum necessary quantization processing and additional bit allocation loop processing.

[0117] If it is determined in step S51 that the time is within the time limit, there is still time remaining until the time limit, so the process proceeds to step S52.

[0118] In step S52, the bit allocation unit 54 determines whether or not the loop process of the additional quantization process, that is, the additional bit allocation loop process, has ended.

[0119] For example, in step S52, it is determined that the loop processing has ended when the additional bit allocation loop processing has been repeated a predetermined number of times, or when the difference in quantization noise between the most recent two additional bit allocation loop processing operations is below a threshold value.

[0120] If it is determined in step S52 that the loop processing has not yet ended, the process returns to step S50, and the above-described processing is repeated.

[0121] On the other hand, if it is determined in step S52 that the loop processing has ended, the processing of step S53 is carried out.

[0122] In step S53, the bit allocation unit 54 saves (holds) the quantized MDCT coefficients updated in step S50 as the final quantization result of the object to be processed, and adds "1" to the value of the object ID to be processed, thereby making a new object that has not yet undergone additional quantization processing the next object to be processed.

[0123] After the process of step S53 is performed, the process returns to step S49, and the above-described process is repeated, that is, an additional quantization process is performed on the new object to be processed.

[0124] In this way, in steps S49 to S53, additional quantization processing is performed on each object in descending order of priority, thereby further improving coding efficiency.

[0125] If it is determined in step S51 that the time is not within the time limit, that is, if the time limit has been reached, the additional quantization process for each object is discontinued, and the process then proceeds to step S54.

[0126] In other words, in this case, the minimum necessary quantization process is completed for some objects, but the process is terminated with the additional quantization process remaining incomplete, so that for some objects, the results of the minimum necessary quantization process are output as the final quantized MDCT coefficients.

[0127] However, in steps S49 to S53, processing is performed in descending order of priority, so the object whose processing is aborted is an object with a relatively low priority. That is, for the object with a high priority, high-quality quantized MDCT coefficients are obtained, so degradation of sound quality can be minimized.

[0128] Furthermore, if it is determined in step S49 that the value of the processing target ID is not less than N, that is, if the additional quantization process is completed for all objects within the time limit, the process proceeds to step S54.

[0129] If the process of step S47 is performed, if it is determined in step S49 that the value of the processing target ID is not less than N, or if it is determined in step S51 that the time limit has not been reached, the process of step S54 is performed.

[0130] In step S 54 , the bit allocation unit 54 outputs the quantized MDCT coefficients held as the quantization results for each object, that is, the stored quantized MDCT coefficients, to the encoding unit 55 .

[0131] At this time, for objects for which the minimum necessary quantization process has not been completed, the quantized value of the Mute data held as the quantization result is output to the encoding unit 55.

[0132] Furthermore, the bit allocation unit 54 supplies the mute information of each object to the packing unit 23, and the bit allocation process ends.

[0133] When the Mute information is supplied to the packing unit 23, in step S17 of FIG. 3 described above, the packing unit 23 stores the Mute information in the coded bitstream.

[0134] The mute information is flag information having a value of "0" or "1".

[0135] Specifically, for example, if all quantized MDCT coefficients in a frame to be coded for an object are 0, that is, if the quantization result is Mute data, the value of the Mute information is set to "1." On the other hand, if the quantization result is not Mute data, the value of the Mute information is set to "0."

[0136] Such mute information is described, for example, in the metadata of an object, an ancillary area of ​​an encoded bitstream, etc. Note that the mute information is not limited to flag information, and may include alphabets, other symbols, or a character string such as "MUTE."

[0137] As an example, FIG. 5 shows a syntax example in which Mute information is added to ObjectMetadataConfig() of MPEG-H.

[0138] In the example of FIG. 5, the Mute information "mutedObjectFlag[o]" is stored in the metadata Config for the number of objects (num_objects).

[0139] As described above, if the quantized MDCT coefficients of an object are all "0", "1" is set as the Mute information (mutedObjectFlag[o]), and "0" is set otherwise.

[0140] By describing such Mute information, the decoding side can use 0 data (zero data) as the IMDCT output instead of performing IMDCT on objects whose Mute information is "1", thereby realizing faster decoding processing.

[0141] In this way, the bit allocation unit 54 performs the minimum necessary quantization process and additional quantization process on objects in order of priority.

[0142] In this way, the higher the priority of an object, the more additional quantization processing (additional bit allocation loop processing) can be completed, improving the coding efficiency of the entire content even in real-time processing, and allowing the data of more objects to be transmitted.

[0143] In the above, the case has been described in which priority information is input to the bit allocation unit 54 and the time-frequency conversion unit 52 performs time-frequency conversion on all objects. However, priority information may also be supplied to the time-frequency conversion unit 52, for example.

[0144] In such a case, the time-frequency transform unit 52 does not perform time-frequency transform on objects with low priority indicated by the priority information, and replaces all MDCT coefficients of each scale factor band with 0 data (zero data) and supplies them to the bit allocation unit 54.

[0145] By doing so, compared to the configuration shown in FIG. 2, it is possible to further reduce the processing time and amount of processing for low-priority objects, and ensure more processing time for high-priority objects.

[0146] <Decoder configuration example> Next, a decoder that receives (acquires) the coded bitstream output from the encoder 11 shown in FIG. 1 and decodes the coded metadata and coded audio signal will be described.

[0147] Such a decoder may be configured, for example, as shown in FIG.

[0148] The decoder 81 shown in FIG. 6 includes an unpacking / decoding unit 91, a rendering unit 92, and a mixing unit 93.

[0149] The unpacking / decoding unit 91 obtains the coded bit stream output from the encoder 11, and unpacks and decodes the coded bit stream.

[0150] The unpacking / decoding unit 91 supplies the audio signal of each object obtained by unpacking and decoding and the metadata of each object to the rendering unit 92. At this time, the unpacking / decoding unit 91 decodes the encoded audio signal of each object in accordance with Mute information included in the encoded bitstream.

[0151] The rendering unit 92 generates M-channel audio signals based on the audio signals of each object supplied from the unpacking / decoding unit 91 and the object position information included in the metadata of each object, and supplies the M-channel audio signals to the mixing unit 93. At this time, the rendering unit 92 generates the M-channel audio signals so that the sound image of each object is localized at the position indicated by the object position information of that object.

[0152] The mixing unit 93 supplies the audio signals of each channel supplied from the rendering unit 92 to external speakers corresponding to each channel, and reproduces the audio.

[0153] In addition, if the encoded bitstream contains encoded audio signals for each channel, the mixing unit 93 performs weighted addition on the audio signals for each channel supplied from the unpacking / decoding unit 91 and the audio signals for each channel supplied from the rendering unit 92 for each channel to generate the final audio signals for each channel.

[0154] <Configuration example of unpacking / decoding unit> Moreover, the unpacking / decoding unit 91 of the decoder 81 shown in FIG. 6 is configured in more detail as shown in FIG.

[0155] The unpacking / decoding unit 91 shown in FIG. 7 includes a Mute information acquisition unit 121, an object audio signal acquisition unit 122, an object audio signal decoding unit 123, an output selection unit 124, a zero value output unit 125, and an IMDCT unit 126.

[0156] The mute information acquisition unit 121 acquires mute information of the audio signal of each object from the supplied coded bitstream and supplies the information to the output selection unit 124 .

[0157] The mute information acquisition unit 121 also acquires and decodes encoded metadata of each object from the supplied encoded bitstream, and supplies the resulting metadata to the rendering unit 92. The mute information acquisition unit 121 also supplies the supplied encoded bitstream to the object audio signal acquisition unit 122.

[0158] The object audio signal acquisition unit 122 acquires the coded audio signal of each object from the coded bitstream supplied from the mute information acquisition unit 121 , and supplies the coded audio signal to the object audio signal decoding unit 123 .

[0159] The object audio signal decoding unit 123 decodes the coded audio signal of each object supplied from the object audio signal acquisition unit 122, and supplies the resulting MDCT coefficients to the output selection unit .

[0160] The output selection unit 124 selectively switches the output destination of the MDCT coefficients of each object supplied from the object audio signal decoding unit 123 based on the mute information of each object supplied from the mute information acquisition unit 121 .

[0161] Specifically, when the value of the Mute information for a predetermined object is “1”, that is, when the quantization result is Mute data, the output selection unit 124 sets the MDCT coefficient of the object to 0 and supplies the 0-value output unit 125. In other words, zero data is supplied to the 0-value output unit 125.

[0162] On the other hand, if the value of the Mute information for a specific object is “0”, that is, if the quantization result is not Mute data, the output selection unit 124 supplies the MDCT coefficients of that object supplied from the object audio signal decoding unit 123 to the IMDCT unit 126.

[0163] The zero value output unit 125 generates an audio signal based on the MDCT coefficients (zero data) supplied from the output selection unit 124, and supplies the signal to the rendering unit 92. In this case, since the MDCT coefficients are zero, a silent audio signal is generated.

[0164] The IMDCT unit 126 performs IMDCT on the MDCT coefficients supplied from the output selection unit 124 to generate an audio signal, and supplies the audio signal to the rendering unit 92 .

[0165] <Description of Decryption Process> Next, the operation of the decoder 81 will be described.

[0166] When the decoder 81 receives a coded bit stream for one frame from the encoder 11, it performs a decoding process to generate an audio signal and outputs it to a speaker. The decoding process performed by the decoder 81 will be described below with reference to the flowchart in FIG.

[0167] In step S81, the unpacking / decoding unit 91 acquires (receives) the coded bit stream transmitted from the encoder 11.

[0168] In step S82, the unpacking / decoding unit 91 performs selective decoding processing.

[0169] The selective decoding process, which will be described in detail later, involves selectively decoding the encoded audio signal of each object based on the mute information. The resulting audio signal of each object is then supplied to the rendering unit 92. The metadata of each object obtained from the encoded bitstream is also supplied to the rendering unit 92.

[0170] In step S83, the rendering unit 92 renders the audio signal of each object based on the audio signal of each object supplied from the unpacking / decoding unit 91 and the object position information included in the metadata of each object.

[0171] For example, the rendering unit 92 generates audio signals for each channel using VBAP (Vector Based Amplitude Panning) based on the object position information so that the sound image of each object is localized at the position indicated by the object position information, and supplies the generated signals to the mixing unit 93. Note that the rendering method is not limited to VBAP, and other formats may also be used. As described above, the object position information consists of, for example, a horizontal angle (Azimuth), a vertical angle (Elevation), and a distance (Radius), but may also be expressed, for example, by Cartesian coordinates (X, Y, Z).

[0172] In step S84, the mixing unit 93 supplies the audio signals of each channel supplied from the rendering unit 92 to the speakers corresponding to those channels to reproduce sound. Once the audio signals of each channel have been supplied to the speakers, the decoding process ends.

[0173] In this way, the decoder 81 obtains the mute information from the coded bitstream and decodes the coded audio signal of each object according to the mute information.

[0174] <Description of Selective Decoding Process> Next, the selective decoding process corresponding to the process in step S82 in FIG. 8 will be described with reference to the flowchart in FIG.

[0175] In step S111, the mute information acquisition unit 121 acquires mute information of the audio signal of each object from the supplied coded bitstream and supplies the information to the output selection unit .

[0176] In addition, the Mute information acquisition unit 121 acquires and decodes the encoded metadata of each object from the encoded bitstream, supplies the resulting metadata to the rendering unit 92, and supplies the encoded bitstream to the object audio signal acquisition unit 122.

[0177] In step S112, the object audio signal acquisition unit 122 sets the object number of the object to be processed to 0 and holds it.

[0178] In step S113, the object audio signal acquisition unit 122 determines whether the object number it holds is less than the number N of objects.

[0179] If it is determined in step S113 that the object number is less than N, then in step S114, the object audio signal decoding unit 123 decodes the coded audio signal of the object to be processed.

[0180] That is, the object audio signal acquisition unit 122 acquires the encoded audio signal of the object to be processed from the encoded bitstream supplied from the mute information acquisition unit 121 , and supplies the acquired encoded audio signal to the object audio signal decoding unit 123 .

[0181] The object audio signal decoding unit 123 then decodes the coded audio signal supplied from the object audio signal acquisition unit 122, and supplies the resulting MDCT coefficients to the output selection unit .

[0182] In step S115, the output selection unit 124 determines whether the value of the Mute information of the object to be processed supplied from the Mute information acquisition unit 121 is "0" or not.

[0183] If it is determined in step S115 that the value of the Mute information is “0”, the output selection unit 124 supplies the MDCT coefficients of the object to be processed, supplied from the object audio signal decoding unit 123, to the IMDCT unit 126, and the process proceeds to step S116.

[0184] In step S116, the IMDCT unit 126 performs IMDCT on the MDCT coefficients supplied from the output selection unit 124 to generate an audio signal of the object to be processed, and supplies the audio signal to the rendering unit 92. Once the audio signal has been generated, the process then proceeds to step S117.

[0185] On the other hand, if it is determined in step S115 that the value of the Mute information is not “0”, that is, the value of the Mute information is “1”, the output selection unit 124 sets the MDCT coefficients to 0 and supplies them to the zero value output unit 125.

[0186] The zero value output unit 125 generates an audio signal of the object to be processed from the MDCT coefficients that are zero and are supplied from the output selection unit 124, and supplies the signal to the rendering unit 92. Therefore, the zero value output unit 125 does not actually perform any processing for generating an audio signal, such as IMDCT.

[0187] It should be noted that the audio signal generated by the zero value output unit 125 is a silent signal. Once the audio signal is generated, the process then proceeds to step S117.

[0188] If it is determined in step S115 that the value of the Mute information is not "0" or if an audio signal is generated in step S116, then in step S117, the object audio signal acquisition unit 122 adds 1 to the object number it holds and updates the object number of the object being processed.

[0189] Once the object number is updated, the process returns to step S113, and the above-described process is repeated, i.e., an audio signal for a new object to be processed is generated.

[0190] Also, if it is determined in step S113 that the object number of the object to be processed is not less than N, the selective decoding process ends because audio signals have been obtained for all objects, and then the process proceeds to step S83 in Figure 8.

[0191] In this way, the decoder 81 decodes the encoded audio signal while determining, for each object in the frame to be processed, based on the mute information, whether or not to decode the encoded audio signal.

[0192] That is, the decoder 81 decodes only the necessary encoded audio signals according to the mute information of each audio signal. This not only reduces the amount of calculation required for decoding while minimizing deterioration in the sound quality of the sound reproduced by the audio signals, but also reduces the amount of calculation required for subsequent processing, such as processing in the rendering unit 92.

[0193] Second Embodiment <Configuration example of object audio encoding unit> Furthermore, the first embodiment described above is an example of delivering fixed viewpoint 3D audio content (audio signals), in which case the user's listening position is a fixed position.

[0194] However, in MPEG-I free viewpoint 3D audio, the user's listening position is not fixed, but can move to any position. Therefore, the priority of each object also changes depending on the relationship between the user's listening position and the object's position (positional relationship).

[0195] Therefore, when the content (audio signal) to be distributed is free viewpoint 3D audio, priority information may be generated taking into account the audio signal of the object, the priority value of the metadata, the object position information, and the listening position information indicating the user's listening position.

[0196] In such a case, the object audio encoding unit 22 of the encoder 11 may be configured as shown in Fig. 10. In Fig. 10, the same reference numerals are used to designate parts corresponding to those in Fig. 2, and their description will be omitted where appropriate.

[0197] The object audio encoding unit 22 shown in FIG. 10 includes a priority information generating unit 51, a time-frequency transforming unit 52, an auditory psycho-parameter calculating unit 53, a bit allocating unit 54, and an encoding unit 55.

[0198] The configuration of the object audio encoding unit 22 in FIG. 10 is basically the same as the configuration shown in FIG. 2, but differs from the example shown in FIG. 2 in that object position information and listening position information are also supplied to the priority information generation unit 51 in addition to the priority value.

[0199] That is, in the example of Figure 10, the priority information generation unit 51 is supplied with the audio signal of each object, the priority value and object position information contained in the metadata of each object, and listening position information indicating the user's listening position in three-dimensional space.

[0200] For example, the listening position information is received (acquired) by the encoder 11 from the decoder 81 to which the content is distributed.

[0201] Furthermore, since the content here is free viewpoint 3D Audio, the object position information included in the metadata is, for example, the position of the sound source in three-dimensional space, i.e., coordinate information indicating the absolute position of the object, etc. However, without being limited to this, the object position information may also be coordinate information indicating the relative position of the object.

[0202] The priority information generation unit 51 generates priority information based on at least one of the audio signal of each object, the priority value of each object, and the object position information and listening position information (metadata and listening position information) of each object, and supplies it to the bit allocation unit 54.

[0203] For example, compared to when the object is close to the user (listener), the volume of the object tends to decrease and the priority of the object tends to decrease as the distance between the object and the user increases.

[0204] Therefore, for example, the priority information generating unit 51 may adjust the priority calculated based on the audio signal and Priority value of the object using a low-order nonlinear function that decreases the priority as the distance between the object and the user's listening position increases, and the priority information indicating the adjusted priority may be used as the final priority information. In this way, priority information that is more subjective can be obtained.

[0205] Even when the object audio encoding unit 22 has the configuration shown in FIG. 10, the encoder 11 performs the encoding process described with reference to FIG.

[0206] However, in step S12, the priority information is generated using object position information and listening position information as well as the object position information and listening position information as needed. That is, the priority information is generated based on the audio signal, the Priority value, and at least one of the object position information and the listening position information.

[0207] Third Embodiment <Configuration example of a content distribution system> However, even if real-time processing limiting processes are implemented to improve encoding efficiency in live streaming of live performances or concerts as in the first embodiment, the processing load on the hardware implementing the encoder may suddenly increase due to interrupts from the OS (Operating System). In such cases, the number of objects whose processing is not completed within the real-time processing limit may increase, causing a sense of discomfort to the listener. In other words, the sound quality may deteriorate.

[0208] Therefore, in order to suppress the occurrence of such audible discomfort, i.e., deterioration in sound quality, multiple input data with different numbers of objects may be prepared by pre-rendering, and each input data may be encoded using separate hardware.

[0209] In this case, for example, among the coded bitstreams for which no restriction processing for real-time processing has been performed, the coded bitstream with the largest number of objects is output to the decoder 81. Therefore, even if one of the multiple pieces of hardware has a hardware that has experienced a sudden increase in processing load due to an OS interrupt or the like, it is possible to suppress the occurrence of an audible sense of discomfort.

[0210] In this way, when a plurality of input data are prepared in advance, a content distribution system for distributing content is configured as shown in FIG. 11, for example.

[0211] The content distribution system shown in FIG. 11 includes encoders 201-1 to 201-3 and an output unit 202.

[0212] For example, in a content distribution system, three pieces of input data D1 to D3, each having a different number of objects, are prepared in advance as data for reproducing the same content.

[0213] Here, the input data D1 is data consisting of audio signals and metadata of N objects, and for example, the input data D1 is original data that has not been pre-rendered.

[0214] Furthermore, input data D2 is data consisting of audio signals and metadata for 16 objects, which is fewer than input data D1. For example, input data D2 is data obtained by performing pre-rendering on input data D1.

[0215] Similarly, input data D3 is data consisting of audio signals and metadata for 10 objects, which is fewer than input data D2; for example, input data D3 is data obtained by pre-rendering input data D1.

[0216] Regardless of which of the input data D1 to D3 is used to play back the content (audio), the same sound is basically played back.

[0217] In the content distribution system, input data D1 is supplied (input) to encoder 201-1, input data D2 is supplied to encoder 201-2, and input data D3 is supplied to encoder 201-3.

[0218] The encoders 201-1 to 201-3 are realized by different hardware components such as computers, etc. In other words, the encoders 201-1 to 201-3 are realized by different operating systems.

[0219] The encoder 201-1 performs encoding processing on the supplied input data D1 to generate an encoded bit stream, and supplies the encoded bit stream to the output unit 202.

[0220] Similarly, encoder 201-2 performs an encoding process on the supplied input data D2 to generate an encoded bitstream and supplies it to output unit 202, and encoder 201-3 performs an encoding process on the supplied input data D3 to generate an encoded bitstream and supplies it to output unit 202.

[0221] In the following description, when there is no need to particularly distinguish between the encoders 201-1 to 201-3, they will also be simply referred to as encoders 201.

[0222] Each encoder 201 has the same configuration as the encoder 11 shown in FIG. 1, for example, and generates a coded bit stream by performing the coding process described with reference to FIG.

[0223] Also, although an example in which three encoders 201 are provided in the content distribution system will be described here, the present invention is not limited to this, and two or four or more encoders 201 may be provided.

[0224] The output unit 202 selects one of the coded bitstreams supplied from the plurality of encoders 201 and transmits the selected coded bitstream to the decoder 81.

[0225] For example, the output unit 202 determines whether there is an encoded bitstream among multiple encoded bitstreams that does not contain Mute information with a value of "1", that is, whether there is an encoded bitstream in which the Mute information value of all objects is "0".

[0226] If there is an encoded bitstream that does not contain Mute information with a value of "1", the output unit 202 selects the encoded bitstream that does not contain Mute information with a value of "1" and has the largest number of objects, and transmits it to the decoder 81.

[0227] Furthermore, if there is no coded bitstream that does not contain Mute information with a value of "1", the output unit 202, for example, selects the one with the largest number of objects or the one with the largest number of objects whose Mute information is "0" and transmits it to the decoder 81.

[0228] In this way, by selecting and outputting one of a plurality of coded bitstreams, it is possible to suppress the occurrence of strange sounds to the ear and to achieve high-quality audio reproduction.

[0229] Here, referring to Figure 12, we will explain specific examples of input data D1 to input data D3 when data consisting of metadata and audio signals of N (where N > 16) objects is prepared as original data of the content.

[0230] In this example, the original data is the same for all of the input data D1 to D3, and the number of objects in the data is N.

[0231] In particular, the input data D1 is the original data itself.

[0232] Therefore, the input data D1 consists of metadata and audio signals for the original N objects, and does not include metadata and audio signals for new objects generated by pre-rendering.

[0233] Furthermore, the input data D2 and the input data D3 are data obtained by performing pre-rendering on the original data.

[0234] Specifically, the input data D2 consists of the metadata and audio signals of the four highest priority objects out of the original N objects, and the metadata and audio signals of the 12 new objects generated by pre-rendering.

[0235] The data for the 12 non-original objects included in the input data D2 was generated by pre-rendering based on the data for (N-4) objects that were not included in the input data D2 out of the original N objects.

[0236] Furthermore, in the input data D2, for four objects, the metadata and audio signals of the original objects are not pre-rendered and are included as they are in the input data D2.

[0237] The input data D3 does not include data of the original objects, but is data consisting of metadata and audio signals of 10 new objects generated by pre-rendering.

[0238] The metadata and audio signals of these 10 objects were generated by pre-rendering based on the data of the original N objects.

[0239] As described above, by performing pre-rendering based on the data of the original objects and generating metadata and audio signals for new objects, it is possible to prepare input data with a reduced number of objects.

[0240] Although the original object data here is only input data D1, multiple pieces of original data that have not been pre-rendered may be used as input data, taking into consideration the possibility of sudden OS interrupts, etc. In other words, for example, not only input data D1 but also input data D2 may be original data.

[0241] In this way, even if an OS interrupt or the like occurs suddenly in encoder 201-1 that receives input data D1, degradation of sound quality can be prevented as long as an OS interrupt or the like does not occur in encoder 201-2 that receives input data D2. In other words, encoder 201-2 is likely to obtain an encoded bitstream that does not include Mute information with a value of "1."

[0242] Alternatively, for example, by pre-rendering based on the data of the original objects, a large number of input data with an even smaller number of objects than the input data D3 shown in Fig. 12 may be prepared. Furthermore, the number of object signals (audio signals) and object metadata (metadata) of each of the input data D1, D2, and D3 may be set by the user, or may be dynamically changed depending on the resources of each encoder 201, etc.

[0243] As described above, according to the present technology described in the first to third embodiments, even when all processing in real-time processing cannot be completed within the time limit, the coding efficiency of the entire content can be improved by performing additional bit allocation processing that improves the coding efficiency in descending order of the importance of the audio of objects.

[0244] <Fourth embodiment> <About underflow> As mentioned above, 3D Audio, which is handled by standards such as the MPEG-H 3D Audio standard, has metadata for each object, such as horizontal and vertical angles that indicate the position of the sound material (object), distance, and gain for the object, making it possible to reproduce the three-dimensional direction, distance, and spread of sound.

[0245] In conventional stereo playback, a mixing engineer in the studio would take multi-track data consisting of many sound materials and pan each sound material to the left and right channels in a process called mixdown to obtain a stereo audio signal.

[0246] In contrast, in 3D Audio, individual sound elements called objects are positioned in three-dimensional space, and the positional information of these objects is described as metadata. As a result, in 3D Audio, a large number of objects, or more specifically, the object audio signals of the objects, are encoded before being mixed down.

[0247] However, when encoding a large number of objects in real time, such as for live broadcasting, the transmission device must have high processing power. In other words, if one frame of data cannot be encoded within the specified time, the transmission device will enter an underflow state, meaning that there is no data to transmit, and the transmission process will fail.

[0248] To avoid such underflow, in coding devices that require real-time performance, the bit allocation process, which requires a large amount of computational resources, is controlled so that the process is completed within a specified time.

[0249] In order to keep up with technological advances and reduce costs, today's encoding devices do not use dedicated hardware, but instead often run encoding software on general-purpose hardware such as a PC (Personal Computer) equipped with an OS (Operating System) such as Linux (registered trademark).

[0250] However, in operating systems such as Linux (registered trademark), many system processes other than encoding are executed in parallel, and these system processes are executed as high-priority processes, so they are often executed with priority over the encoding software process. In such cases, in the worst case scenario, the encoding process may not reach the bit allocation process, resulting in an underflow.

[0251] To avoid such underflow, a technique is often used in which silent data (mute data) is encoded and sent when there is no processed data to output.

[0252] Coding standards such as MPEG-D USAC and MPEG-H 3D Audio use context-based arithmetic coding techniques.

[0253] In this context-based arithmetic coding technique, the quantized MDCT coefficients of the previous frame and the current frame are used as contexts, and the appearance frequency table of the quantized MDCT coefficients to be coded is automatically selected based on the contexts, and arithmetic coding is performed.

[0254] Here, a method of calculating a context in context-based arithmetic coding will be described with reference to FIG.

[0255] In FIG. 13, the vertical direction indicates frequency, and the horizontal direction indicates time, that is, frames of the object audio signal.

[0256] Each rectangle or circle represents an MDCT coefficient block of each frequency for each frame, and each MDCT coefficient block includes two MDCT coefficients (quantized MDCT coefficients). In particular, each rectangle represents an MDCT coefficient block that has already been coded, and each circle represents an MDCT coefficient block that has not yet been coded.

[0257] In this example, the MDCT coefficient block BLK11 is the target of coding. At this time, the four MDCT coefficient blocks BLK12 to BLK15 adjacent to the MDCT coefficient block BLK11 are used as the context.

[0258] In particular, MDCT coefficient blocks BLK12 to BLK14 are MDCT coefficient blocks of the same frequency as or adjacent to the frequency of MDCT coefficient block BLK11 in a frame temporally immediately preceding the frame of the MDCT coefficient block BLK11 to be coded.

[0259] The MDCT coefficient block BLK15 is an MDCT coefficient block of a frequency adjacent to the frequency of the MDCT coefficient block BLK11 in the frame of the MDCT coefficient block BLK11 to be coded.

[0260] A context value is calculated based on these MDCT coefficient blocks BLK12 to MDCT coefficient block BLK15, and an occurrence frequency table (arithmetic code frequency table) for encoding the MDCT coefficient block BLK11 to be encoded is selected based on the context value.

[0261] During decoding, variable-length decoding must be performed using the same frequency table as during encoding from the arithmetic code, i.e., the coded quantized MDCT coefficients. Therefore, the exact same calculations must be performed to calculate the context values ​​during encoding and decoding.

[0262] Note that, since the further detailed content of context-based arithmetic coding is not directly related to the present technology, a description thereof will be omitted here.

[0263] However, in the method of encoding and transmitting the mute data described above, calculations are required to encode the mute data itself, which may result in it being impossible to output one frame of data within a specified time.

[0264] Therefore, this technology makes it possible to prevent underflow in software-based encoding devices using an OS such as Linux (registered trademark), even when the encoding method is MPEG-H, which uses context-based arithmetic coding technology.

[0265] In particular, with this technology, even if the encoding process cannot be completed due to other processing loads occurring on the OS, the occurrence of underflow can be prevented by sending pre-prepared encoded mute data.

[0266] <Encoder configuration example> Fig. 14 is a diagram showing an example of the configuration of another embodiment of an encoder to which the present technology is applied. Note that in Fig. 14, parts corresponding to those in Fig. 1 are given the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0267] 14 is, for example, a software-based encoding device using an OS. That is, the encoder 11 is realized by, for example, running encoding software on an information processing device such as a PC using an OS.

[0268] The encoder 11 includes an initialization unit 301, an object metadata encoding unit 21, an object audio encoding unit 22, and a packing unit .

[0269] The initialization unit 301 performs initialization based on initialization information supplied from the OS or the like, such as when the encoder 11 is started up, and generates encoded mute data based on the initialization information and supplies it to the object audio encoding unit 22.

[0270] The coded mute data is data obtained by coding the quantized values ​​of the mute data, i.e., the quantized MDCT coefficients of MDCT coefficients "0". Such coded mute data can be said to be coded silence data obtained by coding the quantized values ​​of the MDCT coefficients of silence data, i.e., the quantized values ​​of the MDCT coefficients of a silent audio signal. Note that, in the following description, it is assumed that context-based arithmetic coding is performed as coding, but the coding is not limited to this and other coding methods may also be used.

[0271] The object audio encoding unit 22 encodes the supplied audio signal of each object (hereinafter also referred to as object audio signal) in accordance with the MPEG-H standard, and supplies the resulting encoded audio signal to the packing unit 23. At this time, the object audio encoding unit 22 appropriately uses the encoded mute data supplied from the initialization unit 301 as the encoded audio signal.

[0272] As in the above-described embodiment, priority information may be calculated in the object audio encoding unit 22 based on the metadata of each object, and the priority information may be used to perform quantization of MDCT coefficients, etc.

[0273] <Configuration example of object audio encoding unit> The object audio encoding unit 22 of the encoder 11 shown in Fig. 14 may be configured, for example, as shown in Fig. 15. In Fig. 15, parts corresponding to those in Fig. 2 are given the same reference numerals, and their description will be omitted where appropriate.

[0274] In the example of Figure 15, the object audio encoding unit 22 includes a time-frequency conversion unit 52, a psychoacoustic parameter calculation unit 53, a bit allocation unit 54, a context processing unit 331, a variable-length encoding unit 332, an output buffer 333, a processing progress monitoring unit 334, a processing completion determination unit 335, and an encoded mute data insertion unit 336.

[0275] The bit allocation unit 54 performs bit allocation processing based on the MDCT coefficients supplied from the time-frequency transform unit 52 and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53. Note that, similar to the above-described embodiment, the bit allocation unit 54 may perform bit allocation processing based on priority information.

[0276] The bit allocator 54 supplies the quantized MDCT coefficients for each scale factor band of each object obtained by the bit allocation process to the context processor 331 and the variable-length encoder 332 .

[0277] Based on the quantized MDCT coefficients supplied from the bit allocation unit 54, the context processing unit 331 determines (selects) an occurrence frequency table that is required when encoding the quantized MDCT coefficients.

[0278] For example, as described with reference to FIG. 13, the context processing unit 331 determines an occurrence frequency table to be used for encoding a quantized MDCT coefficient of interest (MDCT coefficient block) from a representative value of a plurality of quantized MDCT coefficients in the vicinity of the quantized MDCT coefficient of interest.

[0279] The context processing unit 331 supplies the variable length coding unit 332 with an index (hereinafter also referred to as an occurrence frequency table index) indicating the occurrence frequency table of each quantized MDCT coefficient, determined for each quantized MDCT coefficient, more specifically for each MDCT coefficient block.

[0280] The variable-length coding unit 332 refers to the occurrence frequency table indicated by the occurrence frequency table index supplied from the context processing unit 331, and performs variable-length coding on the quantized MDCT coefficients supplied from the bit allocation unit 54, thereby performing lossless compression.

[0281] Specifically, the variable-length coding unit 332 generates an encoded audio signal by performing context-based arithmetic coding as the variable-length coding.

[0282] It should be noted that arithmetic coding is used as a variable-length coding technique in the coding standards disclosed in the above-mentioned Non-Patent Documents 1 to 3. In this technique, it is possible to apply other variable-length coding techniques besides arithmetic coding, such as Huffman coding.

[0283] The variable-length coding unit 332 supplies the coded audio signal obtained by the variable-length coding to the output buffer 333, where it is stored.

[0284] The context processing unit 331 and the variable length coding unit 332 that code the quantized MDCT coefficients correspond to the coding unit 55 of the object audio coding unit 22 shown in FIG.

[0285] The output buffer 333 holds a bit stream consisting of the encoded audio signal for each frame supplied from the variable length coding unit 332, and supplies the held encoded audio signal (bit stream) to the packing unit 23 at an appropriate timing.

[0286] The processing progress monitoring unit 334 monitors the progress of each process performed in the time-frequency conversion unit 52 to the bit allocation unit 54, the context processing unit 331, and the variable-length coding unit 332, and supplies progress information indicating the monitoring results to the processing completion determination unit 335.

[0287] The processing progress monitoring unit 334 instructs the time-frequency conversion unit 52 to the bit allocation unit 54, the context processing unit 331, and the variable-length coding unit 332 to terminate the processing being executed, as appropriate, depending on the judgment result supplied from the processing completion judgment unit 335.

[0288] The processing completion possibility determination unit 335 determines whether the processing of encoding the object audio signal will be completed within a predetermined time based on the progress information supplied from the processing progress monitoring unit 334, and supplies the determination result to the processing progress monitoring unit 334 and the encoded mute data insertion unit 336. More specifically, the determination result is supplied to the encoded mute data insertion unit 336 only when it is determined that the processing will not be completed within the predetermined time.

[0289] The encoded mute data insertion unit 336 inserts encoded mute data prepared (generated) in advance into a bit stream consisting of the encoded audio signal of each frame in the output buffer 333, depending on the judgment result supplied from the processing completion judgment unit 335.

[0290] In this case, the coded mute data is inserted into the bitstream as the coded audio signal of the frame for which it is determined that the processing will not be completed within the predetermined time.

[0291] That is, if it is determined that the processing for a given frame will not be completed within the allotted time, the bit allocation process is aborted, and the coded audio signal for that given frame cannot be obtained. As a result, the coded audio signal for that given frame is not held in the output buffer 333. Therefore, zero data, i.e., coded mute data, which is coded silence data obtained by coding a silent audio signal (silence signal), is inserted into the bitstream as the coded audio signal for that given frame.

[0292] For example, the insertion of encoded mute data may be performed for each object (object audio signal), or when the bit allocation process is terminated, the encoded audio signals of all objects may be encoded mute data.

[0293] <Configuration example of initialization unit> The initialization unit 301 of the encoder 11 shown in FIG. 14 is configured as shown in FIG. 16, for example.

[0294] The initialization unit 301 includes an initialization processing unit 361 and an encoded mute data generation unit 362 .

[0295] Initialization information is supplied to the initialization processing unit 361. For example, the initialization information includes information indicating the number of objects and channels that make up the content to be encoded, i.e., the number of objects and the number of channels.

[0296] The initialization processing unit 361 performs initialization based on the supplied initialization information, and supplies the number of objects indicated by the initialization information, more specifically, object number information indicating the number of objects, to the encoded mute data generation unit 362.

[0297] The coded mute data generation unit 362 generates coded mute data for the number of objects indicated by the object number information supplied from the initialization processing unit 361, and supplies the data to the coded mute data insertion unit 336. That is, the coded mute data generation unit 362 generates coded mute data for each object. Note that the coded mute data for each object is the same data.

[0298] Furthermore, if the encoder 11 also encodes the audio signals of each channel, the encoded mute data generator 362 also generates encoded mute data for the number of channels based on channel number information indicating the number of channels.

[0299] <Processing progress and encoded Mute data> Next, the progress of the processing performed in each section of the encoder 11 and the encoded mute data will be described.

[0300] The processing progress monitoring unit 334 determines the time using a timer supplied from the processor or OS, and generates progress information indicating the degree of progress of processing from the input of one frame of object audio signal to the generation of the encoded audio signal of that frame.

[0301] A specific example of progress information and determination of whether processing is complete or not will now be described with reference to Fig. 17. In Fig. 17, an object audio signal for one frame is assumed to consist of 1024 samples.

[0302] In the example shown in FIG. 17, time t11 indicates the time when the object audio signal of the frame to be processed is supplied to the time-frequency transform unit 52, that is, the time when time-frequency transform of the object audio signal to be processed starts.

[0303] Furthermore, time t12 is the time when a predetermined threshold is reached, and if quantization of the object audio signal, i.e., generation of the quantized MDCT coefficients, is completed by time t12, the coded audio signal of the frame to be processed can be output (transmitted) without delay. In other words, if the process of generating the quantized MDCT coefficients is completed by time t12, no underflow occurs.

[0304] Time t13 is the time when the output of the coded audio signal of the frame to be processed, i.e., the coded bitstream, starts. In this example, the time from time t11 to time t13 is 21 msec.

[0305] The hatched (diagonally lined) rectangular portions indicate the time required to perform processing (hereinafter also referred to as invariant processing) that requires a substantially constant amount of calculation (computational complexity) regardless of the object audio signal, among the processing steps performed to obtain quantized MDCT coefficients from the object audio signal. More specifically, the hatched rectangular portions indicate the time required to complete the invariant processing. For example, time-frequency transformation and calculation of psychoacoustic parameters are invariant processing.

[0306] In contrast, the unhatched rectangular areas indicate the time required for processing (hereinafter also referred to as variable processing) in which the amount of calculation required, i.e., the processing time, varies depending on the object audio signal, among the processes performed to obtain quantized MDCT coefficients from the object audio signal. For example, bit allocation processing is a variable processing.

[0307] The processing progress monitoring unit 334 monitors the progress of processing in the time-frequency conversion unit 52 through the bit allocation unit 54, and monitors the occurrence of interrupt processing in the OS, etc., to identify the time required to complete the invariant processing and the variable processing. Note that the time required to complete the invariant processing and the variable processing changes depending on the occurrence of interrupt processing in the OS, etc.

[0308] For example, the processing progress monitoring unit 334 generates information indicating the time required for the invariant processing to be completed and the time required for the variable processing to be completed as progress information, and supplies this to the processing completion possibility determination unit 335 .

[0309] For example, in the example indicated by arrow Q11, the invariant processing and the variable processing are completed (ended) by time t12 when the threshold is reached, that is, the quantized MDCT coefficients can be obtained by time t12.

[0310] Therefore, the processing completion determination unit 335 supplies the processing progress monitoring unit 334 with a determination result that the processing of encoding the object audio signal will be completed within a predetermined time, that is, by the time when output of the encoded audio signal should start.

[0311] In the example shown by arrow Q12, the invariant process is completed by time t12, but the variable process takes a long time, so the variable process does not complete by time t12. In other words, the completion time of the variable process is slightly after time t12.

[0312] Therefore, the processing completion determination unit 335 supplies a determination result indicating that the processing of encoding the object audio signal will not be completed within a predetermined time to the processing progress monitoring unit 334. More specifically, the processing completion determination unit 335 supplies a determination result indicating that the bit allocation processing needs to be terminated to the processing progress monitoring unit 334.

[0313] In this case, for example, the processing progress monitor 334 instructs the bit allocator 54 to terminate the bit allocation process, more specifically the bit allocation loop process, in response to the determination result supplied from the processing completion determination unit 335.

[0314] This causes the bit allocation loop processing to be terminated in the bit allocation unit 54. However, since the bit allocation unit 54 performs at least the minimum necessary quantization processing, it is possible to obtain quantized MDCT coefficients without causing underflow, although this may result in a decrease in quality.

[0315] Furthermore, in the example shown by arrow Q13, an interrupt occurs in the OS, so the invariant process is not completed by time t12, resulting in an underflow.

[0316] Therefore, the processing completion determination unit 335 supplies the determination result that the processing of encoding the object audio signal will not be completed within the predetermined time to the processing progress monitoring unit 334 and the encoded mute data insertion unit 336. More specifically, the processing completion determination unit 335 supplies the determination result that it is necessary to output encoded mute data to the processing progress monitoring unit 334 and the encoded mute data insertion unit 336.

[0317] In this case, the processes being performed in the time-frequency transform unit 52 through the variable-length coding unit 332 are stopped (terminated), and the coded mute data inserting unit 336 inserts coded mute data.

[0318] Next, the encoded mute data will be explained. Before explaining the encoded mute data, the encoded audio signal will be explained first.

[0319] As described above, the variable-length coding unit 332 supplies the coded audio signal for each frame to the output buffer 333. More specifically, coded data including the coded audio signal is supplied. It is assumed here that the variable-length coding of the quantized MDCT coefficients is performed in accordance with, for example, the MPEG-H 3D Audio standard.

[0320] For example, one frame of coded data includes at least an Indep flag (independence flag), the coded audio signal of the current frame (coded quantized MDCT coefficients), and a preroll frame flag indicating whether or not data related to a preroll frame is present.

[0321] The Indep flag is flag information indicating whether the current frame is a frame that has been coded using prediction or differential coding.

[0322] For example, the value of the Indep flag is "1", i.e., Indep=1, indicates that the current frame is a frame that has been coded without using prediction, differentials, etc. In other words, Indep=1 indicates that the coded audio signal of the current frame is the absolute values ​​of the quantized MDCT coefficients, i.e., the quantized MDCT coefficients are coded as they are.

[0323] Therefore, when playing back an encoded bitstream from the middle, the decoder 81, i.e., the playback device, can start processing (playback) from a frame with Indep = 1. In other words, a frame with Indep = 1 is a randomly accessible frame.

[0324] On the other hand, the value of the Indep flag is "0", i.e., Indep=0, indicates that the current frame is a frame that has been coded using prediction or a difference. In other words, Indep=0 indicates that the coded audio signal of the current frame is a coded differential value between the quantized MDCT coefficients of the current frame and the quantized MDCT coefficients of the frame immediately preceding the current frame. Therefore, a frame with Indep=0 cannot be randomly accessed, i.e., it cannot be used as a random access target.

[0325] The pre-roll frame flag is flag information indicating whether or not the coded data of the current frame includes a coded audio signal of a pre-roll frame.

[0326] For example, if the value of the pre-roll frame flag is "1", the coded data of the current frame includes the coded audio signal (coded quantized MDCT coefficients) of the pre-roll frame.

[0327] In this case, the coded data of the current frame includes an Indep flag, a coded audio signal of the current frame, a pre-roll frame flag, and a coded audio signal of the pre-roll frame.

[0328] On the other hand, if the value of the pre-roll frame flag is "0", the coded data of the current frame does not include the coded audio signal of the pre-roll frame.

[0329] Note that a pre-roll frame is a randomly accessible frame, that is, a frame that is located immediately before a frame with Indep=1.

[0330] Now, with reference to FIG. 18, an example of a bit stream made up of coded data (coded audio signals) of a plurality of frames will be described.

[0331] 18, #x represents the frame number of the object audio signal frame (time frame). Frames without the characters "Indep=1" are considered to be frames with Indep=0.

[0332] For example, "#0" indicates the 0th frame (0th) in the 0 origin, i.e., the first frame, and "#25" indicates the 25th frame. Below, a frame with frame number "#x" will also be referred to as frame #x.

[0333] In Figure 18, the part indicated by arrow Q31 shows the bitstream obtained by the normal encoding process that is performed when the processing completion possibility determination unit 335 determines that the processing will be completed within the specified time.

[0334] In particular, in this example, frame #0 indicated by arrow W11 and frame #25 indicated by arrow W12 are frames with Indep=1, that is, frames that can be accessed randomly.

[0335] For example, if Indep=1 is set for all frames, decoding (playback) can be started from any frame, but since this significantly reduces coding efficiency, coding is generally performed with Indep=1 every few dozen frames. Therefore, in the explanation of Fig. 18, it is assumed that Indep=1 is set every 25 frames.

[0336] Additionally, the text "PreRollFrame (=#24)" written in the frame #25 section indicates that the encoded audio signal of frame #24, which is a preroll frame for frame #25, is stored in the encoded data (bitstream) of frame #25.

[0337] For example, when decoding starts from frame #25, due to the nature of MDCT, the coded audio signal of frame #25 contains only odd function components of the signal (object audio signal). Therefore, if decoding is performed using only the coded audio signal of frame #25, frame #25 cannot be reproduced as complete data, resulting in noise.

[0338] To prevent such noise from occurring, the encoded data of frame #25 contains the encoded audio signal of frame #24, which is a pre-roll frame.

[0339] When decoding starts from frame #25, the encoded audio signal of frame #24, more specifically the even function component of the encoded audio signal, is extracted (taken out) from the encoded data of frame #25 and combined with the odd function component of frame #25.

[0340] This makes it possible to obtain a complete object audio signal as the decoding result of frame #25, and to prevent abnormal sounds from occurring during playback.

[0341] The portion indicated by arrow Q32 shows the bitstream that is obtained when the processing completion determination unit 335 determines that processing will not be completed within the predetermined time in frame #24. That is, the portion indicated by arrow Q32 shows an example in which encoded Mute data is inserted in frame #24.

[0342] In the following description, a frame into which coded mute data is inserted will also be referred to as a mute frame.

[0343] In this example, frame #24 indicated by arrow W13 is a mute frame, and this frame #24 is the frame (pre-roll frame) immediately before frame #25 which is randomly accessible.

[0344] For frame #24, which is a mute frame, coded mute data calculated in advance based on the number of objects at initialization is inserted into the bitstream as the coded audio signal for frame #24. More specifically, coded data including coded mute data is inserted into the bitstream.

[0345] In the coded mute data generation unit 362, coded mute data is generated by arithmetically coding the quantized MDCT coefficients (quantized values ​​of mute data) with MDCT coefficient "0", assuming that frame #24 is a randomly accessible frame, i.e., Indep=1.

[0346] In particular, the coded mute data is generated using only the quantized MDCT coefficients (silence data) for one frame corresponding to the frame to be processed, without using the quantized MDCT coefficients for the frame immediately preceding the frame to be processed. In other words, the coded mute data is generated without using the difference from the immediately preceding frame or the context of the immediately preceding frame.

[0347] This is because at the time of initialization, that is, when the encoded mute data is generated, the data (quantized MDCT coefficients) of frame #23 immediately preceding frame #24 does not exist.

[0348] In this way, if the mute frame is not a randomly accessible frame, the encoded data for the mute frame is generated to include an Indep flag with a value of "1", encoded Mute data as the encoded audio signal of the current frame that is the mute frame, and a pre-roll frame flag with a value of "0".

[0349] In this case, the value of the Indep flag is "1" in the mute frame, but the decoder 81 is configured not to start decoding from that mute frame.

[0350] In this example, the frame #25 that follows the mute frame #24 is a randomly accessible frame, that is, a frame with Indep=1.

[0351] Therefore, the encoded mute data of frame #24, which is a pre-roll frame of frame #25, is stored as an encoded audio signal of the pre-roll frame in the encoded data of frame #25. In this case, for example, the encoded mute data inserting unit 336 inserts (stores) the encoded mute data of frame #24 into the encoded data of frame #25 held in the output buffer 333.

[0352] The portion indicated by arrow Q33 shows an example in which randomly accessible frame #25 is set as a mute frame.

[0353] In frame #25, which is a mute frame, coded data including coded mute data calculated in advance based on the number of objects at initialization is inserted into the bitstream. This coded mute data is obtained by arithmetically coding the quantized MDCT coefficients with MDCT coefficients of "0" assuming Indep=1, as in the example shown by arrow Q32.

[0354] Furthermore, since frame #25 is a randomly accessible frame, the encoded audio signal of the pre-roll frame is also stored in the encoded data of frame #25. In this case, the encoded mute data is used as the encoded audio signal of the pre-roll frame.

[0355] Therefore, if the mute frame is a randomly accessible frame, the encoded data of the mute frame is generated to include an Indep flag with a value of "1", encoded Mute data as the encoded audio signal of the current frame which is a mute frame, a pre-roll frame flag with a value of "1", and encoded Mute data as the encoded audio signal of the pre-roll frame.

[0356] As described above, the encoded mute data insertion unit 336 inserts encoded mute data depending on the type of the current frame, such as whether the current frame to be a mute frame is a pre-roll frame or a randomly accessible frame.

[0357] According to this technology, in a software-based encoding device using an OS such as Linux (registered trademark), it is possible to prevent underflow from occurring even when the encoding method is MPEG-H or the like, which uses context-based arithmetic coding technology.

[0358] In particular, with this technology, it is possible to prevent underflow from occurring even when encoding of an object audio signal is not completed due to other processing loads occurring on the OS, for example.

[0359] <Example of encoded data structure> Next, an example of the structure of encoded data in which encoded audio signals are stored will be described.

[0360] FIG. 19 shows an example of the syntax of the encoded data.

[0361] In this example, "usacIndependencyFlag" represents the Indep flag.

[0362] Furthermore, "mpegh3daSingleChannelElement(usacIndependencyFlag)" represents an object audio signal, more specifically, an encoded audio signal. This encoded audio signal is data of the current frame.

[0363] Furthermore, the encoded data contains extension data indicated by "mpegh3daExtElement(usacIndependencyFlag)".

[0364] This extended data has a structure shown in FIG. 20, for example.

[0365] In the example shown in FIG. 20, segment data indicated by "usacExtElementSegmentData[i]" is stored as appropriate in the extended data.

[0366] The data stored in this segment data and the order in which the data is stored are determined by usacExtElementType, which is config data, as shown in FIG. 21, for example.

[0367] In the example shown in FIG. 21, when usacExtElementType is "ID_EXT_ELE_AUDIOPREROLL", "AudioPreRoll()" is stored in the segment data.

[0368] This "AudioPreRoll()" is data with a configuration shown in FIG. 22, for example.

[0369] In this example, the encoded audio signals of the frames preceding the current frame indicated by "AccessUnit()" are stored by the number indicated by "numPreRollFrames".

[0370] In particular, here, one encoded audio signal indicated by "AccessUnit()" is the encoded audio signal of a pre-roll frame. Also, by increasing the number indicated by "numPreRollFrames", it is possible to store encoded audio signals of frames further ahead (in the past) in time.

[0371] <Explanation of initialization process> Next, the operation of the encoder 11 shown in FIG. 14 will be described.

[0372] First, the initialization process that is performed when the encoder 11 is started will be described with reference to the flowchart of FIG.

[0373] In step S201, the initialization processing unit 361 performs initialization based on the supplied initialization information. For example, the initialization processing unit 361 resets parameters used in the encoding process in each unit of the encoder 11, and resets the output buffer 333.

[0374] Furthermore, the initialization processing unit 361 generates object number information based on the initialization information, and supplies the object number information to the encoded mute data generation unit 362 .

[0375] In step S 202 , the coded mute data generating unit 362 generates coded mute data based on the object number information supplied from the initialization processing unit 361 , and supplies the coded mute data to the coded mute data inserting unit 336 .

[0376] For example, as described with reference to Fig. 18, the coded mute data generation unit 362 generates coded mute data by arithmetically coding the quantized MDCT coefficients with MDCT coefficients of "0" assuming that Indep = 1. Furthermore, coded mute data is generated for the number of objects indicated by the object number information. Once the coded mute data is generated, the initialization process ends.

[0377] In this way, the encoder 11 performs initialization and generates the encoded mute data. By generating the encoded mute data before encoding, the encoded mute data can be inserted as needed when encoding the object audio signal, thereby preventing underflow.

[0378] <Description of the encoding process> After the initialization process is completed, the encoder 11 performs the encoding process and the encoded Mute data insertion process in parallel at any timing. First, the encoding process by the encoder 11 will be described with reference to the flowchart of FIG.

[0379] The processes in steps S231 to S233 are similar to the processes in steps S11, S13, and S14 in FIG. 3, and therefore will not be described further.

[0380] In step S 234 , the bit allocation unit 54 performs bit allocation processing based on the MDCT coefficients supplied from the time-frequency conversion unit 52 and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53 .

[0381] In the bit allocation process, the above-mentioned minimum necessary quantization process and additional bit allocation loop process are performed on the MDCT coefficients for each scale factor band for each object in an arbitrary order.

[0382] The bit allocator 54 supplies the quantized MDCT coefficients obtained by the bit allocation process to the context processor 331 and the variable-length encoder 332 .

[0383] In step S235, the context processing unit 331 selects an occurrence frequency table to be used for encoding the quantized MDCT coefficients based on the quantized MDCT coefficients supplied from the bit allocation unit .

[0384] For example, as described with reference to FIG. 13, the context processing unit 331 calculates a context value for the quantized MDCT coefficient of the current frame to be processed based on the quantized MDCT coefficients of frequencies in the current frame and the frame immediately preceding the current frame that are close to the frequency (scale factor band) of the quantized MDCT coefficient of the current frame.

[0385] Then, the context processing unit 331 selects an occurrence frequency table for encoding the quantized MDCT coefficients to be processed based on the context value, and supplies an occurrence frequency table index indicating the selection result to the variable length coding unit 332 .

[0386] In step S 236 , the variable-length coding unit 332 performs variable-length coding on the quantized MDCT coefficients supplied from the bit allocation unit 54 based on the occurrence frequency table indicated by the occurrence frequency table index supplied from the context processing unit 331 .

[0387] The variable-length coding unit 332 supplies the coded audio signal obtained by variable-length coding, more specifically, coded data including the coded audio signal of the current frame obtained by variable-length coding, to the output buffer 333, where it is stored.

[0388] 18, the variable length coding unit 332 generates coded data including at least the Indep flag, the coded audio signal of the current frame, and the pre-roll frame flag, and stores the coded data in the output buffer 333. As described above, the coded data also includes the coded audio signal of the pre-roll frame as appropriate, depending on the value of the pre-roll frame flag.

[0389] Note that the processes of steps S232 to S236 described above are performed for each object or each frame depending on the result of the process completion determination made by the process completion determination unit 335. That is, depending on the result of the process completion determination, some or all of the processes may not be executed, or the execution of a process may be stopped midway (aborted).

[0390] Furthermore, by the coded mute data insertion process described later, coded mute data is inserted into the bit stream consisting of the coded audio signal (coded data) for each object of each frame held in the output buffer 333 as appropriate.

[0391] The output buffer 333 supplies the stored encoded audio signal (encoded data) to the packing unit 23 at an appropriate timing.

[0392] When the encoded audio signal (encoded data) is supplied for each frame from the output buffer 333 to the packing unit 23, the processing of step S237 is then performed to end the encoding processing, but since the processing of step S237 is the same as the processing of step S17 in Fig. 3, a description thereof will be omitted. In more detail, in step S237, the encoded metadata and the encoded data including the encoded audio signal are packed, and the resulting encoded bitstream is output.

[0393] In this way, the encoder 11 performs variable-length coding, packs the resulting encoded audio signal and encoded metadata, and outputs an encoded bitstream. In this way, object data can be transmitted efficiently.

[0394] <Description of the Encoded Mute Data Insertion Process> Next, an encoded mute data insertion process that is performed simultaneously with the encoding process in the encoder 11 will be described with reference to the flowchart in Fig. 25. For example, the encoded mute data insertion process is performed for each frame of an object audio signal or for each object.

[0395] In step S251, the processing completion determination unit 335 determines whether processing can be completed.

[0396] For example, when the above-described encoding process is started, the processing progress monitoring unit 334 starts monitoring the progress of each process performed in the time-frequency transform unit 52 to the bit allocation unit 54, the context processing unit 331, and the variable-length coding unit 332, and generates progress information. Then, the processing progress monitoring unit 334 supplies the generated progress information to the processing completion determination unit 335.

[0397] The processing completion determination unit 335 then determines whether the processing can be completed based on the progress information supplied from the processing progress monitoring unit 334, and supplies the determination result to the processing progress monitoring unit 334 and the encoded mute data insertion unit 336.

[0398] For example, even if only the minimum necessary quantization process is performed as bit allocation process, if the variable-length coding unit 332 does not complete variable-length coding by the time when packing should start in the packing unit 23, it is determined that the process of coding the object audio signal will not be completed within the predetermined time. Then, the determination result that the process of coding the object audio signal will not be completed within the predetermined time, more specifically, the determination result that coded mute data needs to be output, is supplied to the process progress monitoring unit 334 and the coded mute data inserting unit 336.

[0399] Also, for example, if only the minimum necessary quantization process is performed in the bit allocation process or if the bit allocation loop process is aborted midway, it may be possible to complete the variable-length coding in the variable-length coding unit 332 by the time packing in the packing unit 23 is to start. In such a case, it is determined that the process of coding the object audio signal will not be completed within a predetermined time, but this determination result is not supplied to the coded mute data insertion unit 336 but only to the processing progress monitoring unit 334. More specifically, the determination result that the bit allocation process needs to be aborted is supplied to the processing progress monitoring unit 334.

[0400] The processing progress monitoring unit 334 appropriately controls the execution of processing performed by the time-frequency conversion unit 52 to the bit allocation unit 54, the context processing unit 331, and the variable-length coding unit 332, depending on the judgment result supplied from the processing completion judgment unit 335.

[0401] That is, as described with reference to FIG. 17, the processing progress monitoring unit 334 instructs each processing block from the time-frequency conversion unit 52 to the variable-length coding unit 332 to stop the execution of a process that is about to be performed or to abort a process that is currently being performed, as appropriate.

[0402] Specifically, for example, suppose that a determination result indicating that the process of encoding an object audio signal in a specified frame will not be completed within a specified time, or more specifically, a determination result indicating that encoded mute data needs to be output, is supplied to the processing progress monitoring unit 334.

[0403] In such a case, the processing progress monitoring unit 334 instructs the time-frequency transform unit 52 through the variable-length coding unit 332 to stop the execution of the processing for the predetermined frame or to abort the processing that is currently being performed in those units. Then, in the encoding processing described with reference to Fig. 24, the processing of steps S232 through S236 is stopped or aborted midway.

[0404] Therefore, the variable-length coding unit 332 does not perform variable-length coding of the quantized MDCT coefficients of the predetermined frame, and the coded audio signal (coded data) for the predetermined frame is not supplied from the variable-length coding unit 332 to the output buffer 333.

[0405] Also, for example, suppose that a determination result indicating that the bit allocation process needs to be terminated is supplied to the processing progress monitoring unit 334 at a specific frame. In such a case, the processing progress monitoring unit 334 instructs the bit allocator 54 to perform only the minimum necessary quantization process or to terminate the bit allocation loop process.

[0406] Then, in the encoding process described with reference to FIG. 24, bit allocation processing is performed in step S234 in accordance with an instruction from the processing progress monitor 334.

[0407] In step S252, the coded mute data insertion unit 336 determines whether or not to insert coded mute data, in other words, whether or not the current frame to be processed is a mute frame, based on the determination result supplied from the processing completion determination unit 335.

[0408] For example, in step S252, if the result of the process completion determination is that the process of encoding the object audio signal will not be completed within a predetermined time, or more specifically, that encoded mute data needs to be output, it is determined that encoded mute data should be inserted.

[0409] If it is determined in step S252 that the encoded Mute data is not to be inserted, the process of step S253 is not performed and the encoded Mute data insertion process ends.

[0410] For example, if the processing progress monitoring unit 334 receives a determination result indicating that the bit allocation process needs to be terminated, it determines in step S252 that the encoded Mute data should not be inserted, and the encoded Mute data insertion unit 336 does not insert the encoded Mute data.

[0411] If the current frame to be processed is a randomly accessible frame and the frame immediately preceding the current frame is a mute frame, the coded mute data inserting unit 336 inserts coded mute data of a pre-roll frame.

[0412] That is, for example, as shown by arrow Q32 in Figure 18, the encoded mute data insertion unit 336 inserts encoded mute data as an encoded audio signal of a pre-roll frame into the encoded data of the current frame held in the output buffer 333.

[0413] If it is determined in step S252 that the encoded mute data is to be inserted, in step S253 the encoded mute data inserting unit 336 inserts the encoded mute data into the encoded data of the current frame according to the type of the current frame to be processed.

[0414] More specifically, the encoded Mute data insertion unit 336 generates encoded data for the current frame including an Indep flag with a value of "1", encoded Mute data as the encoded audio signal of the current frame to be processed, and a pre-roll frame flag, as described with reference to Figure 18, for example.

[0415] At this time, if the current frame is a randomly accessible frame, the coded mute data inserting unit 336 also stores coded mute data as a coded audio signal of the pre-roll frame in the coded data of the current frame to be processed.

[0416] The coded mute data inserting unit 336 then inserts the coded data of the current frame into the portion of the bit stream made up of the coded data of each frame held in the output buffer 333, which corresponds to the current frame.

[0417] As described above, if the current frame is a pre-roll frame of the frame immediately following (immediately following) the current frame, encoded mute data is inserted into the encoded data of the next frame as the encoded audio signal of the pre-roll frame at an appropriate timing.

[0418] Furthermore, if the current frame is a mute frame, the variable-length coding unit 332 may generate coded data of the current frame in which no coded audio signal is stored and supply the coded data to the output buffer 333. In such a case, the coded mute data inserting unit 336 inserts coded mute data into the coded data of the current frame stored in the output buffer 333 as coded audio signals of the current frame or pre-roll frame.

[0419] When the coded mute data is inserted into the bitstream held in the output buffer 333, the coded mute data insertion process ends.

[0420] In this way, the encoder 11 inserts the encoded mute data as needed, thereby preventing underflow.

[0421] Note that even when coded mute data is inserted as needed, the bit allocation process may be performed in the order indicated by the priority information in the bit allocation unit 54. In such a case, the bit allocation unit 54 performs the same process as the bit allocation process described with reference to Fig. 4, and inserts coded mute data for objects for which the minimum necessary quantization process has not been completed.

[0422] <Decoder configuration example> 14. The decoder 81, which receives the coded bit stream output by the encoder 11 shown in FIG. 14, has the configuration shown in FIG. 6, for example.

[0423] However, the configuration of the unpacking / decoding unit 91 in the decoder 81 is, for example, the configuration shown in Fig. 26. Note that in Fig. 26, the same reference numerals are used to designate parts corresponding to those in Fig. 7, and descriptions thereof will be omitted as appropriate.

[0424] The unpacking / decoding unit 91 shown in FIG. 26 includes an object audio signal acquisition unit 122, an object audio signal decoding unit 123, and an IMDCT unit 126.

[0425] The object audio signal acquisition unit 122 acquires the encoded audio signal (encoded data) of each object from the supplied encoded bitstream, and supplies the acquired signal to the object audio signal decoding unit 123.

[0426] Furthermore, the object audio signal acquisition unit 122 acquires and decodes the coded metadata of each object from the coded bitstream supplied, and supplies the resulting metadata to the rendering unit 92 .

[0427] <Description of Decryption Process> Next, a description will be given of the operation of the decoder 81. That is, the decoding process performed by the decoder 81 will be described below with reference to the flowchart in FIG.

[0428] In step S271, the unpacking / decoding unit 91 acquires (receives) the coded bit stream transmitted from the encoder 11.

[0429] In step S272, the unpacking / decoding unit 91 decodes the coded bitstream.

[0430] That is, the object audio signal acquisition unit 122 of the unpacking / decoding unit 91 acquires and decodes the coded metadata of each object from the coded bitstream, and supplies the resulting metadata to the rendering unit 92.

[0431] Furthermore, the object audio signal acquisition unit 122 acquires the encoded audio signal (encoded data) of each object from the encoded bitstream, and supplies the acquired signal to the object audio signal decoding unit 123.

[0432] The object audio signal decoding unit 123 then decodes the coded audio signal supplied from the object audio signal acquisition unit 122 and supplies the resulting MDCT coefficients to the IMDCT unit 126 .

[0433] In step S 273 , the IMDCT unit 126 performs IMDCT on the basis of the MDCT coefficients supplied from the object audio signal decoding unit 123 to generate an audio signal for each object, and supplies the generated audio signal to the rendering unit 92 .

[0434] After the IMDCT is performed, the processes of steps S274 and S275 are performed and the decoding process ends. However, since these processes are similar to the processes of steps S83 and S84 in FIG. 8, a description thereof will be omitted.

[0435] In this way, the decoder 81 decodes the coded bitstream and reproduces the audio. In this way, the audio can be reproduced without causing an underflow, i.e., without any interruption.

[0436] Fifth Embodiment <Encoder configuration example> Incidentally, among the objects that make up the content, there are important objects that should not be masked by other objects. Also, even for a single object, among the multiple frequency components contained in the audio signal of the object, there are important frequency components that should not be masked by other objects.

[0437] Therefore, for objects or frequencies that do not want to be masked by other objects, an allowable upper limit (hereinafter also referred to as an allowable masking threshold) of the amount of auditory masking for sounds from all other objects in the three-dimensional space of the object, i.e., the masking threshold (spatial masking threshold), may be set.

[0438] The masking threshold is the threshold of the sound pressure boundary at which a sound becomes inaudible due to masking, and sounds smaller than this threshold are not perceived auditorily. Note that, in the following, frequency masking will be simply referred to as masking, but temporal masking may be used instead of frequency masking, or both frequency masking and temporal masking may be used. Frequency masking is a phenomenon in which, when sounds of multiple frequencies are played simultaneously, a sound of one frequency masks sounds of another frequency, making them less audible. Temporal masking is a phenomenon in which, when a sound is played, sounds played before and after it are masked, making them less audible.

[0439] When setting information indicating such an upper limit (allowable masking threshold) is set, the setting information can be used for bit allocation processing, more specifically, for calculating psychoacoustic parameters.

[0440] The setting information is information about the masking threshold of an important object or frequency that should not be masked by other objects. For example, the setting information includes an object ID indicating an object (audio signal) for which an upper limit is set, information indicating the frequency for which an upper limit is set, and information indicating the set upper limit (allowable masking threshold). That is, for example, the setting information sets an upper limit (allowable masking threshold) for each frequency for each object.

[0441] By using the setting information, it is possible to allocate bits preferentially to objects and frequencies that are important to the content creator, thereby increasing the sound quality ratio compared to other objects and frequencies, thereby improving the sound quality of the entire content and improving coding efficiency.

[0442] Fig. 28 is a diagram showing an example of the configuration of the encoder 11 when using setting information. In Fig. 28, parts corresponding to those in Fig. 1 are given the same reference numerals, and their explanation will be omitted as appropriate.

[0443] The encoder 11 shown in FIG. 28 includes an object metadata encoding unit 21, an object audio encoding unit 22, and a packing unit .

[0444] In this example, unlike the example shown in FIG. 1, the object audio encoding unit 22 is not supplied with the Priority value included in the metadata of the object.

[0445] The object audio encoding unit 22 encodes the audio signals of each of the N supplied objects in accordance with the MPEG-H standard or the like based on the supplied setting information, and supplies the resulting encoded audio signals to a packing unit 23.

[0446] The upper limit indicated by the setting information may be set (input) by the user, or may be set by the object audio encoding unit 22 based on the audio signal.

[0447] Specifically, for example, the object audio encoding unit 22 may perform music analysis based on the audio signal of each object, and set the upper limit value based on the analysis results obtained as a result, such as the genre and melody of the content.

[0448] For example, for a vocal object, important frequency bands of the vocal can be automatically determined based on the analysis results, and an upper limit can be set based on the determination results.

[0449] The upper limit (allowable masking threshold) indicated by the setting information may be set to a common value across all frequencies for one object, or may be set for each frequency for one object. Alternatively, a common upper limit across all frequencies or an upper limit for each frequency may be set for multiple objects.

[0450] <Configuration example of object audio encoding unit> The object audio encoding unit 22 of the encoder 11 shown in Fig. 28 may be configured, for example, as shown in Fig. 29. Note that in Fig. 29, parts corresponding to those in Fig. 2 are given the same reference numerals, and descriptions thereof will be omitted where appropriate.

[0451] In the example shown in FIG. 29, the object audio encoding unit 22 includes a time-frequency transform unit 52, an auditory psycho-parameter calculation unit 53, a bit allocation unit 54, and an encoding unit 55.

[0452] The time-frequency transform unit 52 performs time-frequency transform using MDCT on the audio signal of each object supplied, and supplies the resulting MDCT coefficients to the psychoacoustic parameter calculation unit 53 and the bit allocation unit 54 .

[0453] The psychoacoustic parameter calculation unit 53 calculates psychoacoustic parameters based on the supplied setting information and the MDCT coefficients supplied from the time-frequency conversion unit 52, and supplies the calculated psychoacoustic parameters to the bit allocation unit .

[0454] Here, an example will be described in which the psychoacoustic parameter calculation unit 53 calculates the psychoacoustic parameters based on the setting information and the MDCT coefficients, but the psychoacoustic parameters may also be calculated based on the setting information and the audio signal.

[0455] The bit allocation unit 54 performs bit allocation processing based on the MDCT coefficients supplied from the time-frequency conversion unit 52 and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53 .

[0456] In the bit allocation process, the quantization bits and quantization noise for each scale factor band are calculated and evaluated, and bit allocation is performed based on a psychoacoustic model. Then, based on the bit allocation results, the MDCT coefficients for each scale factor band are quantized to obtain (generate) quantized MDCT coefficients.

[0457] The bit allocation unit 54 supplies the quantized MDCT coefficients for each scale factor band of each object obtained in this manner to the encoding unit 55 as the quantization results of each object, more specifically, as the quantization results of the MDCT coefficients of each object.

[0458] By the bit allocation process described above, some of the quantization bits of the scale factor bands in which the quantization noise generated by the quantization of the MDCT coefficients is masked and not perceived are allocated to the scale factor bands in which the quantization noise is more easily perceived.

[0459] At this time, bits are preferentially allocated to important objects and frequencies (scale factor bands) according to the setting information. In other words, bits are allocated appropriately to objects and frequencies for which upper limits are set according to the upper limits.

[0460] This makes it possible to suppress deterioration of the overall sound quality, especially that of objects and frequencies that the user (content creator) considers important, and to perform efficient quantization, thereby improving coding efficiency.

[0461] In particular, when calculating the quantized MDCT coefficients, a masking threshold (a psychoacoustic parameter) is calculated for each object at each frequency based on the setting information in the psychoacoustic parameter calculation unit 53. Then, during bit allocation processing in the bit allocation unit 54, quantization bits are allocated so that the quantization noise does not exceed the masking threshold.

[0462] For example, when calculating psychoacoustic parameters, parameters are adjusted so that the allowable quantization noise is reduced for frequencies for which an upper limit is set by the setting information, and the psychoacoustic parameters are calculated.

[0463] The amount of parameter adjustment may be varied according to the allowable masking threshold value, i.e., the upper limit value, indicated by the setting information, thereby allowing more bits to be allocated to the relevant frequency.

[0464] The encoding unit 55 encodes the quantized MDCT coefficients for each scale factor band of each object supplied from the bit allocation unit 54, and supplies the resulting encoded audio signal to the packing unit 23.

[0465] <Description of the encoding process> Next, a description will be given of the operation of the encoder 11 configured as shown in Fig. 28. That is, hereinafter, the encoding process by the encoder 11 shown in Fig. 28 will be described with reference to the flowchart in Fig. 30.

[0466] The process of step S301 is the same as the process of step S11 in FIG. 3, and therefore a description thereof will be omitted.

[0467] In step S302, the psychoacoustic parameter calculation unit 53 acquires setting information.

[0468] In step S303, the time-frequency transform unit 52 performs time-frequency transform using MDCT on the audio signal of each supplied object to generate MDCT coefficients for each scale factor band. The time-frequency transform unit 52 supplies the generated MDCT coefficients to the psychoacoustic parameter calculation unit 53 and the bit allocation unit 54.

[0469] In step S304, the psychoacoustic parameter calculation unit 53 calculates psychoacoustic parameters based on the setting information acquired in step S302 and the MDCT coefficients supplied from the time-frequency conversion unit 52, and supplies the calculated psychoacoustic parameters to the bit allocation unit .

[0470] At this time, the psychoacoustic parameter calculation unit 53 calculates psychoacoustic parameters for the objects and frequencies (scale factor bands) indicated by the setting information based on the upper limit values ​​indicated by the setting information so that the allowable quantization noise is small.

[0471] In step S 305 , the bit allocation unit 54 performs bit allocation processing based on the MDCT coefficients supplied from the time-frequency conversion unit 52 and the psychoacoustic parameters supplied from the psychoacoustic parameter calculation unit 53 .

[0472] The bit allocation unit 54 supplies the quantized MDCT coefficients obtained by the bit allocation process to the encoding unit 55 .

[0473] In step S306, the encoding unit 55 encodes the quantized MDCT coefficients supplied from the bit allocation unit 54, and supplies the resulting encoded audio signal to the packing unit 23.

[0474] For example, the coding unit 55 performs context-based arithmetic coding on the quantized MDCT coefficients, and outputs the coded quantized MDCT coefficients as a coded audio signal to the packing unit 23. Note that the coding method is not limited to arithmetic coding, and any other coding method such as Huffman coding or other coding methods may be used.

[0475] In step S307, the packing unit 23 packs the encoded metadata supplied from the object metadata encoding unit 21 and the encoded audio signal supplied from the encoding unit 55, and outputs the encoded bitstream obtained as a result. When the encoded bitstream obtained by packing is output, the encoding process ends.

[0476] In this way, the encoder 11 calculates psychoacoustic parameters based on the setting information and performs bit allocation processing. This allows the content creator to allocate more bits to objects or frequency bands that they want to prioritize, thereby improving encoding efficiency.

[0477] In this embodiment, an example has been described in which priority information is not used in the bit allocation process. However, this is not limiting, and even when priority information is used in the bit allocation process, setting information may be used in calculating the psychoacoustic parameters. In such a case, the setting information is supplied to the psychoacoustic parameter calculation unit 53 of the object audio encoding unit 22 shown in Fig. 2, and the psychoacoustic parameters are calculated using the setting information. Alternatively, the setting information may be supplied to the psychoacoustic parameter calculation unit 53 of the object audio encoding unit 22 shown in Fig. 15, and the setting information may be used in calculating the psychoacoustic parameters.

[0478] <Example of computer configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0479] FIG. 31 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.

[0480] In the computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.

[0481] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0482] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a nonvolatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0483] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.

[0484] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0485] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.

[0486] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0487] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology. For example, although an example in which quantization processing is performed in order from an object with a higher priority has been described as an embodiment of the present technology, quantization processing may be performed in order from an object with a lower priority depending on the use case.

[0488] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0489] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0490] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0491] Furthermore, the present technology can also be configured as follows.

[0492] (1) a priority information generating unit that generates priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; a time-frequency transform unit that performs a time-frequency transform on the audio signal and generates MDCT coefficients; a bit allocation unit that quantizes the MDCT coefficients of the audio signals in order from the audio signal with the highest priority indicated by the priority information; An encoding device comprising: (2) The bit allocation unit performs a minimum necessary quantization process on the MDCT coefficients of the plurality of the audio signals, and also performs an additional quantization process on the MDCT coefficients of the audio signals in order from the audio signal with the highest priority indicated by the priority information, based on the result of the minimum necessary quantization process. The encoding device according to (1). (3) When the bit allocation unit is unable to perform the additional quantization process on all of the audio signals within a predetermined time limit, the bit allocation unit outputs the result of the minimum necessary quantization process as a quantization result of the audio signals for which the additional quantization process has not been completed. The encoding device according to (2). (4) The bit allocation unit performs the minimum necessary quantization process on the audio signals in order of priority indicated by the priority information. The encoding device according to (3). (5) When the bit allocation unit is unable to perform the minimum necessary quantization process on all of the audio signals within the time limit, it outputs a quantized value of zero data as a quantization result of the audio signals for which the minimum necessary quantization process has not been completed. The encoding device according to (4). (6) The bit allocation unit further outputs mute information indicating whether the quantization result of the audio signal is the quantization value of the zero data. (5) The encoding device according to (5). (7) The bit allocation unit determines the time limit based on the processing time required at a stage subsequent to the bit allocation unit. An encoding device according to any one of (3) to (6). (8) The bit allocation unit dynamically changes the time limit based on the result of the minimum necessary quantization process performed so far or the result of the additional quantization process. (7) The encoding device according to (7). (9) The priority information generating unit generates the priority information based on a sound pressure of the audio signal, a spectral shape of the audio signal, or a correlation of the spectral shapes between the plurality of audio signals. An encoding device according to any one of (2) to (8). (10) The metadata includes a pre-generated priority value indicating the priority of the audio signal. An encoding device according to any one of (2) to (9). (11) the metadata includes position information indicating a sound source position of the sound based on the audio signal, The priority information generating unit generates the priority information based on at least the location information and listening position information indicating a listening position of the user. An encoding device according to any one of (2) to (10). (12) The plurality of audio signals includes at least one of the audio signals of an object and the audio signals of a channel. An encoding device according to any one of (2) to (11). (13) further comprising an auditory psychoparameter calculation unit that calculates auditory psychoparameters based on the audio signal; The bit allocation unit performs the minimum necessary quantization process and the additional quantization process based on the psychoacoustic parameters. An encoding device according to any one of (2) to (12). (14) The audio signal processing device further includes an encoding unit that encodes the quantization result of the audio signal output from the bit allocation unit. An encoding device according to any one of (2) to (13). (15) The psychoacoustic parameter calculation unit calculates the psychoacoustic parameter based on the audio signal and setting information related to a masking threshold for the audio signal. The encoding device according to (13). (16) The encoding device generating priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; performing a time-frequency transform on the audio signal to generate MDCT coefficients; For the plurality of audio signals, the MDCT coefficients of the audio signals are quantized in order from the audio signal with the highest priority indicated by the priority information. Encoding method. (17) generating priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; performing a time-frequency transform on the audio signal to generate MDCT coefficients; For the plurality of audio signals, the MDCT coefficients of the audio signals are quantized in order from the audio signal with the highest priority indicated by the priority information. A program that causes a computer to perform a process. (18) a decoding unit that acquires coded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals, and decodes the coded audio signals. Decryption device. (19) The decoding unit further acquires mute information indicating whether the quantization result of the audio signal is a quantization value of zero data, and generates the audio signal based on the MDCT coefficients obtained by the decoding, or generates the audio signal with the MDCT coefficients set to 0, according to the mute information. (18) A decoding device according to (18). (20) The decoding device obtaining encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals; Decoding the encoded audio signal Decryption method. (twenty one) obtaining encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals; Decoding the encoded audio signal A program that causes a computer to perform a process. (twenty two) an encoding unit that encodes an audio signal and generates an encoded audio signal; a buffer for holding a bitstream consisting of the encoded audio signal for each frame; an inserting unit that inserts pre-generated coded silence data into the bitstream as the coded audio signal of the frame to be processed when the coding of the audio signal for the frame to be processed is not completed within a predetermined time; An encoding device comprising: (twenty three) The audio signal processing device further includes a bit allocation unit that quantizes MDCT coefficients of the audio signal, and the encoding unit encodes the quantization result of the MDCT coefficients. The encoding device according to (22). (twenty four) a generator for generating the encoded silence data; The encoding device according to (23). (twenty five) The generating unit generates the encoded silence data by encoding quantized values ​​of MDCT coefficients of silence data. The encoding device according to (24). (26) The generating unit generates the encoded silence data based on only one frame of the silence data. The encoding device according to (24) or (25). (27) the audio signal is an audio signal of a channel or an object, The generating unit generates the encoded silence data based on at least one of the number of channels and the number of objects. An encoding device according to any one of (24) to (26). (28) The insertion unit inserts the encoded silence data according to the type of the frame to be processed. An encoding device according to any one of (22) to (27). (29) When the frame to be processed is a pre-roll frame of a randomly accessible frame, the inserting unit inserts the coded silence data into the bitstream as the coded audio signal of the pre-roll frame for the randomly accessible frame. The encoding device according to (28). (30) When the frame to be processed is a randomly accessible frame, the inserting unit inserts the coded silence data into the bitstream as the coded audio signal of a pre-roll frame for the frame to be processed. The encoding device according to (28) or (29). (31) The insertion unit does not insert the coded silence data if the bit allocation unit performs only the minimum necessary quantization process on the MDCT coefficients or terminates the additional quantization process performed on the MDCT coefficients after the minimum necessary quantization process, and if the coding process of the audio signal is completed within the predetermined time. An encoding device according to any one of (23) to (27). (32) The encoding unit performs variable length encoding on the audio signal. An encoding device according to any one of (22) to (31). (33) The variable length coding is a context-based arithmetic coding. The encoding device according to (32). (34) The encoding device encoding the audio signal to generate an encoded audio signal; storing a bitstream of the encoded audio signal frame by frame in a buffer; If the process of encoding the audio signal for the frame to be processed is not completed within a predetermined time, encoded silence data generated in advance is inserted into the bitstream as the encoded audio signal for the frame to be processed. Encoding method. (35) encoding the audio signal to generate an encoded audio signal; storing a bitstream of the encoded audio signal frame by frame in a buffer; If the process of encoding the audio signal for the frame to be processed is not completed within a predetermined time, encoded silence data generated in advance is inserted into the bitstream as the encoded audio signal for the frame to be processed. A program that causes a computer to perform a process. (36) a decoding unit that encodes an audio signal to generate an encoded audio signal, and, if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, acquires the bit stream obtained by inserting pre-generated encoded silence data into a bit stream made up of the encoded audio signals for each frame as the encoded audio signal for the frame to be processed, and decodes the encoded audio signal. Decryption device. (37) The decoding device An audio signal is encoded to generate an encoded audio signal, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, a bit stream obtained by inserting pre-generated encoded silence data as the encoded audio signal for the frame to be processed into a bit stream made up of the encoded audio signals for each frame is obtained, and the encoded audio signal is decoded. Decryption method. (38) An audio signal is encoded to generate an encoded audio signal, and if the process of encoding the audio signal for a frame to be processed is not completed within a predetermined time, a bit stream obtained by inserting pre-generated encoded silence data as the encoded audio signal for the frame to be processed into a bit stream made up of the encoded audio signals for each frame is obtained, and the encoded audio signal is decoded. A program that causes a computer to perform a process. (39) a time-frequency transform unit that performs a time-frequency transform on the audio signal of the object and generates MDCT coefficients; a psychoacoustic parameter calculation unit that calculates psychoacoustic parameters based on the MDCT coefficients and setting information related to a masking threshold for the object; a bit allocation unit that performs bit allocation processing based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients; An encoding device comprising: (40) The setting information includes information indicating the upper limit of the masking threshold set for each frequency. The encoding device according to (39). (41) The setting information includes information indicating an upper limit of the masking threshold set for one or more of the objects. The encoding device according to (39) or (40). (42) The encoding device Perform a time-frequency transform on the object's audio signal to generate MDCT coefficients; calculating psychoacoustic parameters based on the MDCT coefficients and setting information regarding a masking threshold for the object; performing bit allocation processing based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients; Encoding method. (43) Perform a time-frequency transform on the object's audio signal to generate MDCT coefficients; calculating psychoacoustic parameters based on the MDCT coefficients and setting information regarding a masking threshold for the object; performing bit allocation processing based on the psychoacoustic parameters and the MDCT coefficients to generate quantized MDCT coefficients; A program that causes a computer to execute a process that includes steps. [Explanation of symbols]

[0493] 11 Encoder, 21 Object Metadata Encoding Unit, 22 Object Audio Encoding Unit, 23 Packing Unit, 51 Priority Information Generation Unit, 52 Time-Frequency Transform Unit, 53 Psychoacoustic Parameter Calculation Unit, 54 Bit Allocation Unit, 55 Encoding Unit, 81 Decoder, 91 Unpacking / Decoding Unit, 92 Rendering Unit, 331 Context Processing Unit, 332 Variable-Length Encoding Unit, 333 Output Buffer, 334 Processing Progress Monitoring Unit, 335 Processing Completion Determination Unit, 336 Encoded Mute Data Insertion Unit, 362 Encoded Mute Data Generation Unit

Claims

1. a priority information generating unit that generates priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; a time-frequency transform unit that performs a time-frequency transform on the audio signal and generates MDCT coefficients; a bit allocation unit that quantizes the MDCT coefficients of the audio signals in order from the audio signal with the highest priority indicated by the priority information; Equipped with The bit allocation unit performs a minimum necessary quantization process on the MDCT coefficients of the plurality of the audio signals, and also performs an additional quantization process on the MDCT coefficients of the audio signals in order from the audio signal with the highest priority indicated by the priority information, based on the result of the minimum necessary quantization process. Encoding device.

2. When the bit allocation unit is unable to perform the additional quantization process on all of the audio signals within a predetermined time limit, the bit allocation unit outputs the result of the minimum necessary quantization process as a quantization result of the audio signals for which the additional quantization process has not been completed. The encoding device according to claim 1 .

3. The bit allocation unit performs the minimum necessary quantization process on the audio signals in order of priority indicated by the priority information. The encoding device according to claim 2 .

4. When the bit allocation unit is unable to perform the minimum necessary quantization process on all of the audio signals within the time limit, it outputs a quantized value of zero data as a quantization result of the audio signals for which the minimum necessary quantization process has not been completed. The encoding device according to claim 3 .

5. The bit allocation unit further outputs mute information indicating whether the quantization result of the audio signal is the quantization value of the zero data.

5. The encoding device according to claim 4.

6. The bit allocation unit determines the time limit based on the processing time required at a stage subsequent to the bit allocation unit. The encoding device according to claim 2 .

7. The bit allocation unit dynamically changes the time limit based on the result of the minimum necessary quantization process performed so far or the result of the additional quantization process. The encoding device according to claim 6.

8. The priority information generating unit generates the priority information based on a sound pressure of the audio signal, a spectral shape of the audio signal, or a correlation of the spectral shapes between the plurality of audio signals. The encoding device according to claim 1 .

9. The metadata includes a pre-generated priority value indicating the priority of the audio signal. The encoding device according to claim 1 .

10. the metadata includes position information indicating a sound source position of the sound based on the audio signal, The priority information generating unit generates the priority information based on at least the location information and listening position information indicating a listening position of the user. The encoding device according to claim 1 .

11. The plurality of audio signals includes at least one of the audio signals of an object and the audio signals of a channel. The encoding device according to claim 1 .

12. further comprising an auditory psychoparameter calculation unit that calculates auditory psychoparameters based on the audio signal; The bit allocation unit performs the minimum necessary quantization process and the additional quantization process based on the psychoacoustic parameters. The encoding device according to claim 1 .

13. The audio signal quantization unit may further include a coding unit that codes the quantization result of the audio signal output from the bit allocation unit. The encoding device according to claim 1 .

14. The psychoacoustic parameter calculation unit calculates the psychoacoustic parameter based on the audio signal and setting information related to a masking threshold for the audio signal. The encoding device according to claim 12.

15. The encoding device generating priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; performing a time-frequency transform on the audio signal to generate MDCT coefficients; For the plurality of audio signals, the MDCT coefficients of the audio signals are quantized in order from the audio signal with the highest priority indicated by the priority information. including the steps a minimum necessary quantization process is performed on the MDCT coefficients of the plurality of audio signals, and an additional quantization process is performed on the audio signals in order of priority indicated by the priority information, in which the MDCT coefficients are quantized based on the result of the minimum necessary quantization process; Encoding method.

16. generating priority information indicating a priority of the audio signal based on at least one of an audio signal and metadata of the audio signal; performing a time-frequency transform on the audio signal to generate MDCT coefficients; For the plurality of audio signals, the MDCT coefficients of the audio signals are quantized in order from the audio signal with the highest priority indicated by the priority information. Have the computer execute the process, a minimum necessary quantization process is performed on the MDCT coefficients of the plurality of audio signals, and an additional quantization process is performed on the audio signals in order of priority indicated by the priority information, in which the MDCT coefficients are quantized based on the result of the minimum necessary quantization process; program.

17. a decoding unit configured to obtain encoded audio signals by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals, and to decode the encoded audio signals; The encoded audio signal is obtained by performing a minimum necessary quantization process on the MDCT coefficients of the plurality of audio signals, and performing an additional quantization process on the MDCT coefficients of the plurality of audio signals in order from the audio signal with the highest priority indicated by the priority information, based on the result of the minimum necessary quantization process. Decryption device.

18. The decoding unit further acquires mute information indicating whether the quantization result of the audio signal is a quantization value of zero data, and generates the audio signal based on the MDCT coefficients obtained by the decoding, or generates the audio signal with the MDCT coefficients set to 0, according to the mute information.

18. The decoding device according to claim 17.

19. The decoding device obtaining encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals; decoding the encoded audio signal; The encoded audio signal is obtained by performing a minimum necessary quantization process on the MDCT coefficients of the plurality of audio signals, and performing an additional quantization process on the MDCT coefficients of the plurality of audio signals in order from the audio signal with the highest priority indicated by the priority information, based on the result of the minimum necessary quantization process. Decryption method.

20. obtaining encoded audio signals obtained by quantizing MDCT coefficients of a plurality of audio signals in descending order of priority indicated by priority information generated based on at least one of the audio signals and metadata of the audio signals; Decoding the encoded audio signal Have the computer execute the process, The encoded audio signal is obtained by performing a minimum necessary quantization process on the MDCT coefficients of the plurality of audio signals, and performing an additional quantization process on the MDCT coefficients of the plurality of audio signals in order from the audio signal with the highest priority indicated by the priority information, based on the result of the minimum necessary quantization process. program.

Citation Information

Patent Citations

  • IEC23003-3,

  • IEC23008-3

  • IEC23008-3,

  • Voice encoder and decoder

    JP2000206994A

  • Method and device for audio encoding, and encoding program recording medium

    JP2005148760A