Audio signal coding method, audio signal coding device, program, and recording medium
Patent Information
- Authority / Receiving Office
- IN · IN
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2022-12-19
- Publication Date
- 2026-07-16
AI Technical Summary
The existing embedded encoding/decoding technologies for sound signals with multiple channels and monaural signals result in a larger algorithmic delay for stereo encoding/decoding compared to monaural encoding/decoding, complicating control in communication systems like multipoint control units during conference calls.
A sound signal encoding method that includes a stereo encoding step to obtain a characteristic parameter and a downmix signal, followed by a monaural encoding step using an encoding scheme with overlapping window processing, and an additional encoding step for the overlap section between frames to reduce algorithmic delay.
This method ensures that the algorithmic delay of stereo encoding/decoding is not larger than that of monaural encoding/decoding, facilitating synchronized control and output of stereo and monaural decoded signals in communication systems.
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a technology of embeddedencoding / decoding of a sound signal having a plurality of channelsand a sound signal having one channel.BACKGROUND ART
[0002] As a technology of embedded encoding / decoding of a soundsignal having a plurality of channels and a monaural sound signal,there is a technology of Non-Patent Literature 1. The summary ofthe technology of Non-Patent Literature 1 will be described withan encoding device 500 illustrated in Fig. 5 and a decoding device600 illustrated in Fig. 6. For each frame which is a predeterminedtime section, a stereo encoding unit 510 of the encoding device500 obtains, from a stereo input sound signal which is an inputtedsound signal having a plurality of channels, a stereo code CSrepresenting a characteristic parameter which is a parameterrepresenting a characteristic of difference between the channelsof the stereo input sound signal, and a downmix signal which is asignal obtained by mixing the stereo input sound signal. Amonaural encoding unit 520 of the encoding device 500 encodes thedownmix signal for each frame to obtain a monaural code CM. Amonaural decoding unit 610 of the decoding device 600 decodes themonaural code CM for each frame to obtain a monaural decoded soundsignal which is a decoded signal of the downmix signal. A stereodecoding unit 620 of the decoding device 600 performs, for eachframe, process of obtaining the characteristic parameter which isthe parameter representing the characteristic of the differencebetween the channels by decoding the stereo code CS, and obtaininga stereo decoded sound signal from the monaural decoded soundsignal and the characteristic parameter (so-called upmixprocessing).
[0003] As a monaural encoding / decoding scheme by which a highquality monaural decoded sound signal can be obtained, there is a3GPP EVS standard encoding / decoding scheme described in Non-PatentLiterature 2. By using a high-quality monaural encoding / decodingscheme like that of Non-Patent Literature 2, as the monauralencoding / decoding scheme of Non-Patent Literature 1, there is apossibility that a higher-quality embedded encoding / decoding of asound signal having a plurality of channels and a monaural soundsignal can be realized.PRIOR ART LITERATURENON-PATENT LITERATURE
[0004] Non-Patent Literature 1: Jeroen Breebaart et al.,"Parametric Coding of Stereo Audio", EURASIP Journal on AppliedSignal Processing, pp. 1305-1322, 2005:9.Non-Patent Literature 2: 3GPP, "Codec for Enhanced VoiceServices (EVS); Detailed algorithmic description", TS 26.445.SUMMARY OF THE INVENTIONPROBLEMS TO BE SOLVED BY THE INVENTION
[0005] The upmix processing of Non-Patent Literature 1 is signalprocessing in the frequency domain that includes processing ofapplying a window having overlap between adjacent frames to amonaural decoded sound signal. The monaural encoding / decodingscheme of Non-Patent Literature 2 also includes processing ofapplying a window having overlap between adjacent frames. That is,as for a predetermined range of a boundary part of frames, adecoded sound signal is obtained by combining a signal which isobtained by applying an inclined window in an attenuating shape toa signal obtained by decoding a code of a preceding frame and asignal which is obtained by applying an inclined window in anincreasing shape to a signal obtained by decoding a code of afollowing frame, on both of the decoding side of the stereoencoding / decoding scheme of Non-Patent Literature 1 and thedecoding side of the monaural encoding / decoding scheme of NonPatent Literature 2. From the above, there is a problem that, whena monaural encoding / decoding scheme like that of Non-PatentLiterature 2 is used as a monaural encoding / decoding scheme forembedded encoding / decoding like that of Non-Patent Literature 1, astereo decoded sound signal delays with respect to a monauraldecoded sound signal by an amount corresponding to the window inthe upmix processing, that is, the algorithmic delay of the stereoencoding / decoding is larger than that of the monauralencoding / decoding.
[0006] For example, in a multipoint control unit (MCU) forperforming a conference call at many places, it is common tocontrol to switch that a signal from which place is to beoutputted to which place, for each predetermined time section, andit is difficult to control in a state where a stereo decoded soundsignal delays with respect to a monaural decoded sound signal byan amount corresponding to the window in the upmix processing.Thus, it is assumed that the implementation is such that thecontrol is performed in a state where the stereo decoded soundsignal is delayed by one frame with respect to the monauraldecoded sound signal. That is, in a communication system thatincludes a multipoint control unit, the problem described abovebecomes more prominent, and there is a possibility that thealgorithmic delay of stereo encoding / decoding becomes larger thanthe algorithmic delay of monaural encoding / decoding by one frame.Further, though it becomes possible to control the switching foreach predetermined time section by delaying the stereo decodedsound signal with respect to the monaural decoded sound signal byone frame, there is a possibility that the control about amonaural decoded sound signal from which place and a stereodecoded sound signal from which place are to be combined andoutputted for each time section becomes complicated because themonaural decoded sound signal and the stereo decoded sound signalhave different delays.
[0007] The present invention has been made in view of such aproblem, and its objective is to provide such embeddedencoding / decoding of a sound signal having a plurality of channelsand a monaural sound signal that the algorithmic delay of stereoencoding / decoding is not larger than the algorithmic delay ofmonaural encoding / decoding.MEANS TO SOLVE THE PROBLEMS
[0008] In order to solve the above problem, a sound signalencoding method as one aspect of the present invention is a soundsignal encoding method for encoding an inputted sound signalhaving C channels (C is an integer of 2 or larger) for each frame,the sound signal encoding method comprising: as processing for acurrent frame, a stereo encoding step of obtaining and outputtinga stereo code representing a characteristic parameter which is aparameter representing a characteristic of difference betweenchannels of the sound signal having the C channels; a downmix stepof obtaining a signal by mixing the sound signal having the Cchannels as a downmix signal; and a monaural encoding step ofencoding the downmix signal to obtain and output a monaural code,wherein the monaural encoding step encodes the downmix signal byan encoding scheme that includes processing of applying a windowhaving overlap between frames to obtain the monaural code, and thesound signal encoding method further comprises an additionalencoding step of encoding a part of the downmix signal for asection corresponding to the overlap between the current frame andan immediately following frame to obtain and output an additionalcode.EFFECTS OF THE INVENTION
[0009] According to the present invention, it is possible toprovide such embedded encoding / decoding of a sound signal having aplurality of channels and a monaural sound signal that thealgorithmic delay of stereo encoding / decoding is not larger thanthe algorithmic delay of monaural encoding / decoding.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] [Fig. 1] Fig. 1 is a block diagram showing an example ofan encoding device of each embodiment.[Fig. 2] Fig. 2 is a flowchart showing an example ofprocessing of the encoding device of each embodiment.[Fig. 3] Fig. 3 is a block diagram showing an example of adecoding device of each embodiment.[Fig. 4] Fig. 4 is a flowchart showing an example ofprocessing of the decoding device of each embodiment.[Fig. 5] Fig. 5 is a block diagram showing an example of aconventional encoding device.[Fig. 6] Fig. 6 is a block diagram showing an example of aconventional decoding device.[Fig. 7] Fig. 7 is a diagram schematically showing eachsignal in an encoding device of Non-Patent Literature 2.[Fig. 8] Fig. 8 is a diagram schematically showing eachsignal and algorithmic delay in a decoding device of Non-PatentLiterature 2.[Fig. 9] Fig. 9 is a diagram schematically showing eachsignal and algorithmic delay in a decoding device of Non-PatentLiterature 1 when a monaural encoding / decoding scheme of NonPatent Literature 2 is used as a monaural encoding / decodingscheme.[Fig. 10] Fig. 10 is a diagram schematically showing eachsignal and algorithmic delay in the decoding device of the presentinvention.[Fig. 11] Fig. 11 is a diagram schematically showing eachsignal in the encoding device of the present invention.[Fig. 12] Fig. 12 is a diagram showing an example of afunctional configuration of a computer that realizes each deviceof each embodiment.DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] Before describing each embodiment, each signal andalgorithmic delay in encoding / decoding of the background art and afirst embodiment will be described first, with reference to Figs.7 to 11 schematically illustrating each signal when the framelength is 20 ms. The horizontal axis of each of Figs. 7 to 11 is atime axis. Since description will be made on an example where acurrent frame is processed at time t7 below, descriptions of"past" and "future" are attached to the left end and right end ofthe axis arranged at the top of each diagram, and an upward arrowis attached to the position of t7 which is the time when thecurrent frame is processed. Figs. 7 to 11 schematically show, foreach signal, which time section the signal belongs to, and, when awindow is applied, whether the window is in an increasing shape, aflat shape or an attenuating shape. More specifically, in order tovisually express that combination of a section with the window inthe increasing shape and a section with the window in theattenuating shape yields a signal without windowing, the sectionwith the window in the increasing shape is indicated by atriangular shape that includes a straight line rising to theright, and the section with the window in the attenuating shape isindicated by a triangular shape that includes a straight linefalling to the right in Figs. 7 to 11, because what a windowfunction is strictly like is not important in the descriptionhere. Further, hereinafter, though time of the start of eachsection is identified using words such as "from" or "at and after"in order to avoid a complicated wording expression, the actualstart of each section is the time immediately after the specifiedtime, and the actual start of a digital signal for each section isthe sample immediately after the specified time, as one skilled inthe art could understand.
[0012] Fig. 7 is a diagram schematically showing each signal inan encoding device of Non-Patent Literature 2 that processes thecurrent frame at time t7. It is the signal 1a, which is a monauralsound signal up to t7, that the encoding device of Non-PatentLiterature 2 can use for the processing of the current frame. Inthe processing of the current frame, the encoding device of NonPatent Literature 2 uses the 8.75 ms section from t6 to t7 of thesignal 1a for analysis, as a so-called "look-ahead section", andencodes the signal 1b, which is a signal obtained by applying awindow to a part of the signal 1a for the 23.25 ms section from t1to t6, to obtain and output a monaural code. The shape of thewindow is increasing in the 3.25 ms section from t1 to t2, is flatin the 16.75 ms section from t2 to t5 and is attenuating in the3.25 ms section from t5 to t6. That is, the signal 1b is amonaural sound signal corresponding to the monaural code obtainedin the processing of the current frame. The encoding device ofNon-Patent Literature 2 had already finished similar processing asprocessing of an immediately previous frame at the time when themonaural sound signal up to t3 was inputted, and had alreadyencoded the signal 1c which is a signal obtained by applying awindow in a shape of attenuating in the section from t1 to t2 tothe monaural sound signal for the 23.25 ms section up to t2. Thatis, the signal 1c is a monaural sound signal corresponding to amonaural code obtained in the processing of the immediatelyprevious frame, and the section from t1 to t2 is an overlapsection between the current frame and the immediately previousframe. Further, the encoding device of Non-Patent Literature 2encodes the signal 1d, which is a signal obtained by applying awindow in a shape of increasing in the section from t5 to t6 tothe monaural sound signal for the 23.25 ms section at and aftert5, as processing of an immediately following frame. That is, thesignal 1d is a monaural sound signal corresponding to a monauralcode obtained in the processing of the immediately followingframe, and the section from t5 to t6 is an overlap section betweenthe current frame and the immediately following frame.
[0013] Fig. 8 is a diagram schematically showing each signal in adecoding device of Non-Patent Literature 2 that processes thecurrent frame at time t7 when the monaural code of the currentframe is inputted from the encoding device of Non-PatentLiterature 2. In the processing of the current frame, the decodingdevice of Non-Patent Literature 2 obtains the signal 2a which is adecoded sound signal for the section from t1 to t6, from themonaural code of the current frame. The signal 2a is a decodedsound signal corresponding to the signal 1b and is a signal towhich the window in the shape of increasing in the section from t1to t2, being flat in the section from t2 to t5 and attenuating inthe section from t5 to t6 is applied. The decoding device of NonPatent Literature 2 had already obtained the signal 2b which is adecoded sound signal of the 23.25 ms section up to t2, to whichthe window in the shape of attenuating in the section from t1 tot2 is applied, from the monaural code of the immediately previousframe at time t3 when the monaural code of the immediatelyprevious frame was inputted, as the processing of the immediatelyprevious frame. Further, the decoding device of Non-PatentLiterature 2 obtains the signal 2c which is a decoded sound signalof the 23.25 ms section at and after t5, to which the window inthe shape of increasing in the section from t5 to t6 is applied,from a monaural code of the immediately following frame, asprocessing of the immediately following frame. However, since thesignal 2c has not been obtained at time t7, a complete decodedsound signal is not obtained for the section from t5 to t6 at timet7 though an incomplete decoded sound signal is obtained.Therefore, at time t7, the decoding device of Non-PatentLiterature 2 obtains and outputs the signal 2d, which is amonaural decoded sound signal of the 20 ms section from t1 to t5,by combining the signal 2b obtained in the processing of theimmediately previous frame and the signal 2a obtained in theprocessing of the current frame for the section from t1 to t2 andusing the signal 2a obtained in the processing of the currentframe as it is for the section from t2 to t5. Since the decodingdevice of Non-Patent Literature 2 obtains the decoded soundsignal, whose sections start from t1, at time t7, the algorithmicdelay of the monaural encoding / decoding scheme of Non-PatentLiterature 2 is 32 ms which is a time length from t1 to t7.
[0014] Fig. 9 is a diagram schematically showing each signal inthe decoding device 600 of Non-Patent Literature 1 in a case wherethe monaural decoding unit 610 uses the monaural decoding schemeof Non-Patent Literature 2. At time t7, the stereo decoding unit620 performs stereo decoding processing (upmix processing) of thecurrent frame using the signal 3a, which is a monaural decodedsound signal up to t5 completely obtained by the monaural decodingunit 610. Specifically, the stereo decoding unit 620 uses thesignal 3b, which is a signal of the 23.25 ms section from t0 to t5obtained by applying a window in a shape of increasing in the 3.25ms section from t0 to t1, being flat in the 16.75 ms section fromt1 to t4 and attenuating in the 3.25 ms section from t4 to t5 tothe signal 3a to obtain the signal 3c-i ("i" is a channel number),which is a decoded sound signal from t0 to t5 to which a window inthe same shape as the window for the signal 3b is applied, foreach channel. The stereo decoding unit 620 had already obtainedthe signal 3d-i which is a decoded sound signal of each channelfor the 23.25 ms section up to t1, to which a window in a shape ofattenuating in the section from t0 to t1 is applied, at the pointof time t3, as the processing of the immediately previous frame.Further, the stereo decoding unit 620 obtains the signal 3e-iwhich is a decoded sound signal of each channel for the 23.25 mssection at and after t4, to which the window in the shape ofincreasing in the section from t4 to t5 is applied, as theprocessing of the immediately following frame. However, since thesignal 3e-i is not obtained at time t7, a complete decoded soundsignal is not obtained for the section from t4 to t5 at time t7though an incomplete decoded sound signal is obtained. Therefore,at the time t7, for each channel, the stereo decoding unit 620obtains and outputs the signal 3f-i, which is a complete decodedsound signal for the 20 ms section from t0 to t4, by combining thesignal 3d-i obtained in the processing of the immediately previousframe and the signal 3c-i obtained in the processing of thecurrent frame for the section from t0 to t1 and using the signal3c-i obtained in the processing of the current frame as it is forthe section from t1 to t4. Since the decoding device 600 obtains adecoded sound signal, whose sections start at t0 for each channel,at the point of time t7, the algorithmic delay of the stereoencoding / decoding of Non-Patent Literature 1 using the monauralencoding / decoding scheme of Non-Patent Literature 2 as a monauralencoding / decoding scheme is 35.25 ms which is a time length fromt0 to t7. That is, the algorithmic delay of the stereoencoding / decoding in the embedded encoding / decoding is larger thanthe algorithmic delay of the monaural encoding / decoding.
[0015] Fig. 10 is a diagram schematically showing each signal ina decoding device 200 of the first embodiment described later. Thedecoding device 200 of the first embodiment has a configurationshown in Fig. 3 and includes a monaural decoding unit 210, anadditional decoding unit 230 and a stereo decoding unit 220 thatoperate as described in detail in the first embodiment. At timet7, the stereo decoding unit 220 processes the current frame usingthe signal 4c, which is a completely obtained monaural decodedsound signal up to t6. As stated in the description of Fig. 9, itis the signal 3a which is the monaural decoded sound signal up tot5 that is completely obtained by the monaural decoding unit 210at time t7. Therefore, in the decoding device 200, the additionaldecoding unit 230 decodes an additional code CA to obtain thesignal 4b, which is a monaural decoded sound signal of the 3.25 mssection from t5 to t6 (additional decoding processing), and thestereo decoding unit 220 performs stereo decoding processing(upmix processing) of the current frame using the signal 4cconcatenating the signal 3a which is the monaural decoded soundsignal up to t5 obtained by the monaural decoding unit 210, andthe signal 4b which is the monaural decoded sound signal of thesection from t5 to t6 obtained by the additional decoding unit230. That is, the stereo decoding unit 220 uses the signal 4d ofthe 23.75 ms section from t1 to t6 obtained by applying the windowin the shape of increasing in the 3.25 ms section from t1 to t2,being flat in the 16.75 ms section from t2 to t5 and attenuatingin the 3.25 ms section from t5 to t6 to the signal 4c to obtainthe signal 4e-i, which is a decoded sound signal of the sectionfrom t1 to t6 to which a window in the same shape as the windowfor the signal 4d is applied, for each channel. The stereodecoding unit 220 had already obtained the signal 4f-i which is adecoded sound signal of each channel for the 23.75 ms section upto t2, to which the window in the shape of attenuating in thesection from t1 to t2 is applied, at time t3, as the processing ofthe immediately previous frame. Further, the stereo decoding unit220 obtains the signal 4g-i, which is a decoded sound signal ofeach channel for the 23.75 ms section at and after t5, to whichthe window in the shape of increasing in the section from t5 to t6is applied, as processing of the immediately following frame.However, since the signal 4g-i is not obtained at time t7, acomplete decoded sound signal is not obtained for the section fromt5 to t6 at time t7 though an incomplete decoded sound signal isobtained. Therefore, at time t7, for each channel, the stereodecoding unit 220 obtains and outputs the signal 4h-i, which is acomplete decoded sound signal for the 20 ms section from t1 to t5,by combining the signal 4f-i obtained in the processing of theimmediately previous frame and the signal 4e-i obtained in theprocessing of the current frame for the section from t1 to t2 andusing the signal 4e-i obtained in the processing of the currentframe as it is for the section from t2 to t5. Since the decodingdevice 200 obtains a decoded sound signal, whose sections start att1 for each channel, at the point of time t7, the algorithmicdelay of the stereo encoding / decoding in the embeddedencoding / decoding of the first embodiment is 32 ms which is thetime length from t1 to t7. That is, the algorithmic delay of thestereo encoding / decoding by the embedded encoding / decoding of thefirst embodiment is not larger than the algorithmic delay of themonaural encoding / decoding.
[0016] Fig. 11 is a diagram schematically showing each signal inan encoding device 100 of the first embodiment described later,that is, an encoding device corresponding to the decoding device200 of the first embodiment which is a decoding device for makingeach signal as schematically shown in Fig. 10. The encoding device100 of the first embodiment has a configuration shown in Fig. 1and includes an additional encoding unit 130 that processesencoding the signal 5c which is a part of the signal 1a, which isa monaural sound signal, for the section from t5 to t6, which isan overlap section between the current frame and the immediatelyfollowing frame to obtain the additional code CA, in addition to astereo encoding unit 110 that performs processing similar to thatof the stereo encoding unit 510 of the encoding device 500 and amonaural encoding unit 120 that encodes the signal 1b, which is asignal obtained by applying a window to a part of the signal 1a,which is the monaural sound signal up to t7, for the section fromt1 to t6 to obtain a monaural code CM similarly to the monauralencoding unit 520 of the encoding device 500.
[0017] Hereinafter, the section from t5 to t6, which is theoverlap section between the current frame and the immediatelyfollowing frame, will be called "section X". That is, on theencoding side, the section X is a section for which the monauralencoding unit 120 encodes a monaural sound signal to which awindow is applied, in both of the processing of the current frameand processing of the immediately following frame. Morespecifically, the section X is a section with a predeterminedlength of a sound signal that includes an end of the sound signalthat the monaural encoding unit 120 encodes in the processing ofthe current frame, a section for which the monaural encoding unit120 encodes a sound signal, to which the window in the attenuatingshape is applied, in the processing of the current frame, asection with the predetermined length including the beginning ofthe section encoded by the monaural encoding unit 120 in theprocessing of the immediately following frame, and a section forwhich the monaural encoding unit 120 encodes a sound signal, towhich the window in the increasing shape is applied, in theprocessing of the immediately following frame. Further, on thedecoding side, the section X is a section for which the monauraldecoding unit 210 decodes the monaural code CM to obtain a decodedsound signal to which a window is applied, in both of theprocessing of the current frame and the processing of theimmediately following frame. More specifically, the section X is asection with the predetermined length of a decoded sound signalthat the monaural decoding unit 210 obtains by decoding themonaural code CM in the processing of the current frame, thatincludes the end of the decoded sound signal, a section of thedecoded sound signal that the monaural decoding unit 210 obtainsby decoding the monaural code CM in the processing of the currentframe, to which a window in the attenuating shape is applied, asection with the predetermined length of a decoded sound signalthat the monaural decoding unit 210 obtains by decoding themonaural code CM in the processing of the immediately followingframe, that includes the start of the decoded sound signal, asection of the decoded sound signal that the monaural decodingunit 210 obtains by decoding the monaural code CM in theprocessing of the immediately following frame, to which a windowin the increasing shape is applied, and a section for which themonaural decoding unit 210 obtains a decoded sound signal bycombining the decoded sound signal already obtained by decodingthe monaural code CM in the processing of the current frame andthe decoded sound signal obtained by decoding the monaural code CMin the processing of the immediately following frame, in theprocessing of the immediately following frame.
[0018] Further, hereinafter, the section t1 to t5, which is asection except the section X in a section for which monauralencoding / decoding is performed in the processing of the currentframe, will be called "section Y". That is, the section Y is, onthe encoding side, a part of the section for which the monauralsound signal is encoded by the monaural encoding unit 120 in theprocessing of the current frame except the overlap section betweenthe current frame and the immediately following frame, and on thedecoding side, a part of the section for which the monaural codeCM is decoded by the monaural decoding unit 210 to obtain adecoded sound signal in the processing of the current frame exceptthe overlap section between the current frame and the immediatelyfollowing frame. Since the section Y is a concatenation of asection where the monaural sound signal is represented by themonaural code CM of the current frame and the monaural code CM ofthe immediately previous frame and a section where a monauralsound signal is represented only by the monaural code CM of thecurrent frame, the section Y is a section for which the monauraldecoded sound signal can be completely obtained in processing upto the processing of the current frame.
[0019] <FIRST EMBODIMENT>An encoding device and a decoding device of the firstembodiment will be described.
[0020] <<Encoding device 100>>As shown in Fig. 1, the encoding device 100 of the firstembodiment includes the stereo encoding unit 110, the monauralencoding unit 120 and the additional encoding unit 130. Theencoding device 100 encodes an inputted two-channel stereo timedomain sound signal (a two-channel stereo input sound signal) foreach frame with a predetermined time length, for example, 20 ms toobtain and output a stereo code CS, a monaural code CM and anadditional code CA described later. The two-channel stereo inputsound signal inputted to the encoding device 100 is a digitalvoice or acoustic signal obtained, for example, by picking upsound such as voice and music by each of two microphones, andperforming AD conversion thereof, and consists of an input soundsignal of a left channel which is a first channel and an inputsound signal of a right channel which is a second channel. Codesoutputted by the encoding device 100, that is, the stereo code CS,the monaural code CM and the additional code CA are inputted tothe decoding device 200 described later. The encoding device 100processes steps S111, S121 and S131 illustrated in Fig. 2 for eachframe, that is, each time the above-stated two-channel stereoinput sound signal of the predetermined time length is inputted.In the case of the example described above, the encoding device100 processes steps S111, S121 and S131 for the current frame whena two-channel stereo input sound signal of 20 ms from t3 to t7 isinputted.
[0021] [Stereo encoding unit 110]From the two-channel stereo input sound signal inputted tothe encoding device 100, the stereo encoding unit 110 obtains andoutputs the stereo code CS representing a characteristicparameter, which is a parameter representing a characteristic ofdifference between the inputted sound signals of the two channels,and a downmix signal which is a signal obtained by mixing thesound signals of the two channels (step S111).
[0022] [Example of stereo encoding unit 110]As an example of the stereo encoding unit 110, anoperation of the stereo encoding unit 110 for each frame whentaking information representing strength difference between theinputted sound signals of the two channels for each frequency bandas the characteristic parameter will be described. Note that,though a specific example using a complex DFT (Discrete FourierTransformation) is described below, a well-known method forconversion to the frequency domain other than the complex DFT maybe used. Note that, in the case of converting such a samplesequence, whose number of samples is not a power of two, into thefrequency domain, a well-known technology, such as using a samplesequence with zero stuffing so that the number of samples becomesa power of two, can be used.
[0023] First, the stereo encoding unit 110 performs complex DFTfor each of the inputted sound signals of the two channels toobtain a complex DFT coefficient sequence (step S111-1). Thecomplex DFT coefficient sequence is obtained by applying a windowhaving overlap between frames and using processing inconsideration of symmetry of complex numbers obtained by complexDFT. For example, when the sampling frequency is 32 kHz, theprocessing is performed each time sound signals of the twochannels, each of which has 640 samples corresponding to 20 ms,are inputted; and, for each channel, it is enough to obtain asequence of 372 complex numbers corresponding to the former halfof a sequence of 744 complex numbers to be obtained by performingcomplex DFT for a digital sound signal sample sequence ofsuccessive 744 samples (in the case of the example describedabove, a sample sequence of the section from t1 to t6) as thecomplex DFT coefficient sequence; which 744 samples includes 104samples overlapping with a sample group at the end of theimmediately previous frame (in case of the example describedabove, samples of the section from t1 to t2) and 104 samplesoverlapping with a sample group at the beginning of theimmediately following frame (in the case of the example describedabove, samples of the section from t5 to t6) . Hereinafter, "f"indicates each of integers from 1 to 372; each complex DFTcoefficient of a complex DFT coefficient sequence of the firstchannel is indicated by V1(f); and each complex DFT coefficient ofa complex DFT coefficient sequence of the second channel isindicated by V2(f). Next, from the complex DFT coefficientsequences of the two channels, the stereo encoding unit 110obtains sequences of radiuses of the complex DFT coefficients onthe complex plane (step S111-2). The radius of each complex DFTcoefficient of each channel on the complex plane corresponds tostrength of the sound signal of each channel for each frequencybin. Hereinafter, the radius of the complex DFT coefficient V1(f)of the first channel on the complex plane is indicated by V1r(f),and the radius of the complex DFT coefficient V2(f) of the secondchannel on the complex plane is indicated by V2r(f). Next, thestereo encoding unit 110 obtains an average of ratios of radiusesof one channel and radiuses of the other channel for eachfrequency band, and obtains a sequence of averages as thecharacteristic parameter (step S111-3). The sequence of averagesis the characteristic parameter corresponding to the informationrepresenting the strength difference between the inputted soundsignals of the two channels for each frequency band. For example,in the case of four bands, "f" of which being from 1 to 93, from94 to 186, from 187 to 279 and from 280 to 372, the stereoencoding unit 110 obtain 93 values for each of four bands bydividing the radius V1r(f) of the first channel by the radiusV2r(f) of the second channel, obtains averages thereof as Mr(1),Mr(2), Mr(3) and Mr(4), and obtains a series of average {Mr(1),Mr(2), Mr(3), Mr(4)} as the characteristic parameter.
[0024] Note that the number of bands is only required to be avalue equal to or smaller than the number of frequency bins, andthe same number as the number of frequency bins or 1 may be usedas the number of bands. In the case of using the same value as thenumber of frequency bins as the number of bands, the stereoencoding unit 110 can obtain, for each frequency bin, a value of aratio between a radius of one channel and a radius of the otherchannel and obtain a sequence of the obtained values of ratios asthe characteristic parameter. In the case of using 1 as the numberof bands, the stereo encoding unit 110 can obtain, for eachfrequency bin, a value of a ratio between a radius of one channeland a radius of the other channel and obtain an average of theobtained ratio values for all the bands as the characteristicparameter. Further, in case of adopting multiple bands, the numberof frequency bins to be included in each frequency band isarbitrary. For example, the number of frequency bins to beincluded in a low-frequency band may be smaller than the number offrequency bins to be included in a high-frequency band.
[0025] Further, the stereo encoding unit 110 may use differencebetween a radius of one channel and a radius of the other channel,instead of the ratio between a radius of one channel and a radiusof the other channel. That is, in the case of the exampledescribed above, the stereo encoding unit 110 may use a valueobtained by subtracting the radius V2r(f) of the second channelfrom the radius V1r(f) of the first channel, instead of the valueobtained by dividing the radius V1r(f) of the first channel by theradius V2r(f) of the second channel.
[0026] Furthermore, the stereo encoding unit 110 obtains thestereo code CS which is a code representing the characteristicparameter (step S111-4). The stereo code CS which is a coderepresenting the characteristic parameter can be obtained by awell-known method. For example, the stereo encoding unit 110performs vector quantization of the value sequence obtained atstep S111-3 to obtain a code, and outputs the obtained code as thestereo code CS. Alternatively, for example, the stereo encodingunit 110 performs scalar quantization of each of the valuesincluded in the value sequence obtained at step S111-3 to obtaincodes, and outputs the obtained codes together as the stereo codeCS. Note that, in a case where what is obtained at step S111-3 isone value, the stereo encoding unit 110 can output a code obtainedby scalar quantization of the one value, as the stereo code CS.
[0027] The stereo encoding unit 110 also obtains a downmix signalwhich is a signal obtained by mixing the sound signals of the twochannels of the first channel and the second channel (step S111-5). For example, in the processing of the current frame, thestereo encoding unit 110 obtains a downmix signal, which is amonaural signal obtained by mixing the sound signals of the twochannels, for 20 ms from t3 to t7. The stereo encoding unit 110may mix the sound signals of the two channels in the time domainlike step S111-5A described later or may mix the sound signals ofthe two channels in the frequency domain like step S111-5Bdescribed later. In the case of mixing in the time domain, forexample, the stereo encoding unit 110 obtains a sequence ofaverages of corresponding samples between the sample sequence ofthe sound signal of the first channel and the sample sequence ofthe sound signal of the second channel, as the downmix signalwhich is a monaural signal obtained by mixing the sound signals ofthe two channels (step S111-5A). In the case of mixing in thefrequency domain, for example, the stereo encoding unit 110obtains a complex DFT coefficient sequence applying complex DFT tothe sample sequence of the first channel sound signal, obtains acomplex DFT coefficient sequence applying complex DFT to thesample sequence of the second channel sound signal, obtains aradius average VMr(f) and an angle average VM White Square (f) from each complexDFT coefficient thereof, and obtains a sample sequence applyinginverse complex DFT to a sequence of complex values VM(f) whoseradius is VMr(f) and angle is VM White Square (f) on the complex plane , as thedownmix signal which is a monaural signal obtained by mixing thesound signals of the two channels (step S111-5B).
[0028] Note that, as indicated by two-dot chain lines in Fig. 1,the encoding device 100 may be provided with a downmix unit 150 sothat step S111-5 for obtaining a downmix signal may be processednot within the stereo encoding unit 110 but by the downmix unit150. In this case, the stereo encoding unit 110 obtains andoutputs the stereo code CS representing the characteristicparameter which is a parameter representing the characteristic ofthe difference between the inputted sound signals of the twochannels, from the two-channel stereo input sound signal inputtedto the encoding device 100 (step S111), and the downmix unit 150obtains and outputs the downmix signal, which is a signal obtainedby mixing the sound signals of the two channels, from the twochannel stereo input sound signal inputted to the encoding device100 (step S151). That is, the stereo encoding unit 110 may performsteps S111-1 to step S111-4 described above as step S111, and thedownmix unit 150 may perform step S111-5 described above as stepS151.
[0029] [Monaural encoding unit 120]The downmix signal outputted by the stereo encoding unit110 is inputted to the monaural encoding unit 120. When theencoding device 100 is provided with the downmix unit 150, thedownmix signal outputted by the downmix unit 150 is inputted tothe monaural encoding unit 120. The monaural encoding unit 120encodes the downmix signal by a predetermined encoding scheme toobtain and output the monaural code CM (step S121). As theencoding scheme, an encoding scheme that includes processing ofapplying a window having overlap between frames, for example, likethe 13.2 kbps mode of the 3GPP EVS standard (3GPP TS26.445) ofNon-Patent Literature 2 is used. In the case of the exampledescribed above, in the processing of the current frame, themonaural encoding unit 120 encodes the signal 1b, which is asignal for the section from t1 to t6 obtained by applying a windowin a shape of increasing in the section from t1 to t2 where thecurrent frame and the immediately previous frame overlap,attenuating in the section from t5 to t6 where the current frameand the immediately following frame overlap and being flat in thesection from t2 to t5 between the above sections, to the signal 1awhich is the downmix signal, using the section from t6 to t7 ofthe signal 1a which is the "look-ahead section" for analysisprocessing, to obtain and output the monaural code CM.
[0030] Thus, when the encoding scheme used by the monauralencoding unit 120 includes processing of applying a window havingoverlap and analysis processing using a "look-ahead section", notonly the downmix signal outputted by the stereo encoding unit 110or the downmix unit 150 in the processing of the current frame butalso a downmix signal outputted by the stereo encoding unit 110 orthe downmix unit 150 in frame processing in the past is also usedin encoding processing. Therefore, the monaural encoding unit 120can be provided with a storage not shown to store downmix signalsinputted in frame processing in the past so that the monauralencoding unit 120 can process encoding of the current frame usinga downmix signal stored in the storage, too. Alternatively, thestereo encoding unit 110 or the downmix unit 150 may be providedwith a storage not shown so that the stereo encoding unit 110 orthe downmix unit 150 may output a downmix signal to be used by themonaural encoding unit 120 in encoding processing of the currentframe, including a downmix signal obtained in frame processing inthe past, in the processing of the current frame, and the monauralencoding unit 120 may use the downmix signals inputted from thestereo encoding unit 110 or the downmix unit 150 in the processingof the current frame. Note that storing signals obtained in frameprocessing in the past in a storage not shown and using a signalin the processing of the current frame, like the above processing,are also performed by each unit described later when necessary.Since it is well-known processing in the technological field ofencoding, description thereof will be omitted below in order toavoid redundancy.
[0031] [Additional encoding unit 130]The downmix signal outputted by the stereo encoding unit110 is inputted to the additional encoding unit 130. When theencoding device 100 is provided with the downmix unit 150, thedownmix signal outputted by the downmix unit 150 is inputted tothe additional encoding unit 130. The additional encoding unit 130encodes a part of the inputted downmix signal for the section X toobtain and output the additional code CA (step S131). In the caseof the example described above, the additional encoding unit 130encodes the signal 5c, which is a downmix signal for the sectionfrom t5 to t6, to obtain and output the additional code CA. Forthe encoding, an encoding scheme such as well-known scalarquantization or vector quantization can be used.
[0032] <<Decoding device 200>>As shown in Fig. 3, the decoding device 200 of the firstembodiment includes the monaural decoding unit 210, the additionaldecoding unit 230 and the stereo decoding unit 220. For each framewith the same predetermined time length as in the encoding device100, the decoding device 200 decodes the inputted monaural codeCM, additional code CA and stereo code CS to obtain and output thetwo-channel stereo time-domain sound signal (a two-channel stereodecoded sound signal). The codes inputted to the decoding device200, that is, the monaural code CM, the additional code CA and thestereo code CS are outputted by the encoding device 100. Thedecoding device 200 processes steps S211, S221 and S231illustrated in Fig. 4 for each frame, that is, each time themonaural code CM, the additional code CA and the stereo code CSare inputted at an interval with the predetermined time lengthdescribed above. In the case of the example described above, whenthe monaural code CM, the additional code CA and the stereo codeCS of the current frame are inputted at time t7, 20 ms after t3,when the immediately previous frame was processed, the decodingdevice 200 processes steps S211, S221 and S231 for the currentframe. Note that, as shown by a broken line in Fig. 3, thedecoding device 200 also outputs a monaural decoded sound signalwhich is a monaural time-domain sound signal when necessary.
[0033] [Monaural decoding unit 210]The monaural code CM included among the codes inputted tothe decoding device 200 is inputted to the monaural decoding unit210. The monaural decoding unit 210 obtains and outputs themonaural decoded sound signal for the section Y using the inputtedmonaural code CM (step S211). As a predetermined decoding scheme,a decoding scheme corresponding to the encoding scheme used by themonaural encoding unit 120 of the encoding device 100 is used. Inthe case of the example described above, the monaural decodingunit 210 decodes the monaural code CM of the current frame by thepredetermined decoding scheme to obtain the signal 2a for the23.25 ms section from t1 to t6, to which the window in the shapeof increasing in the 3.25 ms section from t1 to t2, being flat inthe 16.75 ms section from t2 to t5 and attenuating in the 3.25 mssection from t5 to t6 is applied. By combining the signal 2bobtained from the monaural code CM of the immediately previousframe in the processing of the immediately previous frame and thesignal 2a obtained from the monaural code CM of the current framefor the section from t1 to t2 and using the signal 2a obtainedfrom the monaural code CM of the current frame as it is for asection from t2 to t5, the monaural decoding unit 210 obtains andoutputs the signal 2d, which is the monaural decoded sound signalfor 20 ms section from t1 to t5. Note that, since the signal 2afor the section from t5 to t6 obtained from the monaural code CMof the current frame is used as "the signal 2b obtained fromprocessing of an immediately previous frame" in the processing ofthe immediately following frame, the monaural decoding unit 210stores the signal 2a for the section from t5 to t6 obtained fromthe monaural code CM of the current frame into a storage not shownin the monaural decoding unit 210.
[0034] [Additional decoding unit 230]The additional code CA included among the codes inputtedto the decoding device 200 is inputted to the additional decodingunit 230. The additional decoding unit 230 decodes the additionalcode CA to obtain and output an additional decoded signal which isthe monaural decoded sound signal for the section X (step S231).For the decoding, a decoding scheme corresponding to the encodingscheme used by the additional encoding unit 130 is used. In theexample described above, the additional decoding unit 230 decodesthe additional code CA of the current frame to obtain and outputthe signal 4b which is the monaural decoded sound signal for the3.25 ms section from t5 to t6.
[0035] [Stereo decoding unit 220]The monaural decoded sound signal outputted by themonaural decoding unit 210, the additional decoded signaloutputted by the additional decoding unit 230 and the stereo codeCS included among the codes inputted to the decoding device 200are inputted to the stereo decoding unit 220. From the inputtedmonaural decoded sound signal, additional decoded signal andstereo code CS, the stereo decoding unit 220 obtains and outputs astereo decoded sound signal, which is a decoded sound signalhaving the two channels (step S221). More specifically, the stereodecoding unit 220 obtains a decoded downmix signal for a sectionY+X which is a signal obtained by concatenating the monauraldecoded sound signal for the section Y and the additional decodedsignal for the section X (that is, a section obtained byconcatenating the section Y and the section X) (step S221-1), andobtains and outputs the decoded sound signals of the two channelsfrom the decoded downmix signal obtained at step S221-1 by upmixprocessing using the characteristic parameter obtained from thestereo code CS (step S221-2). The upmix processing is processingof obtaining the decoded sound signals of the two channels,regarding the decoded downmix signal as a signal obtained bymixing the decoded sound signals of the two channels and regardingthe characteristic parameter obtained from the stereo code CS asinformation representing the characteristic of difference betweenthe decoded sound signals of the two channels. The same goes foreach embodiment described later. In the case of the exampledescribed above, first, the stereo decoding unit 220 obtains thedecoded downmix signal for the 23.25 ms section from t1 to t6 (thesection from t1 to t6 of the signal 4c) by concatenating themonaural decoded sound signal for the 20 ms section from t1 to t5outputted by the monaural decoding unit 210 (the section from t1to t5 of the signals 2d and 3a) and an additional decoded signal(the signal 4b) for the 3.25 ms section from t5 to t6 outputted bythe additional decoding unit 230. Next, regarding the decodeddownmix signal for the section from t1 to t6 as a signal obtainedby mixing the decoded sound signals of the two channels andregarding the characteristic parameter obtained from the stereocode CS as information representing the characteristic of thedifference between the decoded sound signals of the two channels,the stereo decoding unit 220 obtains and outputs the decoded soundsignals of the two channels for the 20 ms section from t1 to t5(signals 4h-1 and 4h-2).
[0036] [Example of step S221-2 performed by stereo decoding unit220]Step S221-2 performed by the stereo decoding unit 220 whenthe characteristic parameter is information representing thestrength difference between the sound signals of the two channelsfor each frequency band will be described as an example of stepS221-2 performed by the stereo decoding unit 220. First, thestereo decoding unit 220 decodes the inputted stereo code CS toobtain the information representing the strength difference foreach frequency band (S221-21). The stereo decoding unit 220obtains the characteristic parameter from the stereo code CS by ascheme corresponding to the scheme by which the stereo encodingunit 110 of the encoding device 100 obtained the stereo code CSfrom the information representing the strength difference for eachfrequency band. For example, the stereo decoding unit 220 performsvector decoding of the inputted stereo code CS to obtain elementvalues of a vector corresponding to the inputted stereo code CS asinformation representing strength differences for a plurality offrequency bands, respectively. Alternatively, for example, thestereo decoding unit 220 performs scalar decoding of each of codesincluded in the inputted stereo code CS to obtain the informationrepresenting the strength difference for each frequency band. Notethat, in a case where the number of bands is one, the stereodecoding unit 220 performs scalar decoding of the inputted stereocode CS to obtain information representing the strength differencefor the one frequency band, that is, for the whole band.
[0037] Next, regarding the decoded downmix signal as a signalobtained by mixing the decoded sound signals of the two channelsand regarding the characteristic parameter as the informationrepresenting strength difference between the decoded sound signalsof the two channels for each frequency band, the stereo decodingunit 220 obtains and outputs the decoded sound signals of the twochannels from the decoded downmix signal obtained at step S221-1and the characteristic parameter obtained at step S221-21 (stepS220-22). When the stereo encoding unit 110 of the encoding device100 operates in the above-stated specific example using complexDFT, the stereo decoding unit 220 operates at step S221-22 asfollows.
[0038] First, the stereo decoding unit 220 obtains the signal 4dobtained by applying the window in the shape of increasing in the3.25 ms section from t1 to t2, being flat in the 16.75 ms sectionfrom t2 to t5 and attenuating in the 3.25 ms section from t5 to t6to a decoded downmix signal with 744 samples for the 23.25 mssection from t1 to t6 (step S221-221). Next, the stereo decodingunit 220 obtains a sequence of 372 complex numbers of the formerhalf of a sequence of 744 complex numbers to be obtained byperforming complex DFT to the signal 4d as a complex DFTcoefficient sequence (a monaural complex DFT coefficient sequence)(step S221-222). Hereinafter, each complex DFT coefficient of themonaural complex DFT coefficient sequence obtained by the stereodecoding unit 220 is indicated by MQ(f). Next, the stereo decodingunit 220 obtains a radius MQr(f) of each complex DFT coefficienton the complex plane and an angle M QWhite Square (f) of each complex DFTcoefficient on the complex plane from the monaural complex DFTcoefficient sequence (step S221-223). Next, the stereo decodingunit 220 obtains a value by multiplying each radius MQr(f) by asquare root of a corresponding value in the characteristicparameter, as each radius VLQr(f) of the first channel, andobtains a value by dividing each radius MQr(f) by a square root ofa corresponding value in the characteristic parameter, as eachradius VRQr(f) of the second channel (step S221-224). In the caseof the example of the four bands described above, thecorresponding value in the characteristic parameter for eachfrequency bin is Mr(1) when "f" is 1 to 93, Mr(2) when "f" is 94to 186, Mr(3) when "f" is 187 to 279 and Mr(4) when "f" is 280 to372. Note that, when the stereo encoding unit 110 of the encodingdevice 100 uses the difference between the radius of the firstchannel and the radius of the second channel instead of the ratiobetween the radius of the first channel and the radius of thesecond channel, the stereo decoding unit 220 can obtain a value byadding a value obtained by dividing a corresponding value in thecharacteristic parameter by 2 to each radius MQr(f) as each radiusVLQr(f) of the first channel and obtain a value by subtracting thevalue obtained by dividing a corresponding value in thecharacteristic parameter by 2 from each radius MQr(f) as eachradius VRQr(f) of the second channel. Next, the stereo decodingunit 220 performs inverse complex DFT to the sequence of suchcomplex numbers that the radius and angle on the complex plane areVLQr(f) and MQ White Square (f), respectively, to obtain a decoded sound signalof the first channel with the 744 samples for the 23.25 ms sectionfrom t1 to t6 (the signal 4e-1) to which a window is applied, andperforms inverse complex DFT to the sequence of such complexnumbers that the radius and angle on the complex plane are VRQr(f)and MQ White Square (f), respectively, to obtain a decoded sound signal of thesecond channel with the 744 samples for the 23.25 ms section fromt1 to t6 (the signal 4e-2) (step S221-225) to which a window isapplied. The decoded sound signals of the channels to which thewindow is applied obtained at step S221-225 (the signals 4e-1 and4e-2) are signals to which the window in the shape of increasingin the 3.25 ms section from t1 to t2, being flat in the 16.75 mssection from t2 to t5 and attenuating in the 3.25 ms section fromt5 to t6 is applied. Next, the stereo decoding unit 220 obtainsand outputs the decoded sound signals for the 20 ms section fromt1 to t5 (the signal 4h-1 and 4h-2) by combining the signalsobtained at step S221-225 for the immediately previous frame (thesignal 4f-1 and 4f-2) and the signals obtained at step S221-225for the current frame (the signals 4e-1 and 4e-2) for the sectionfrom t1 to t2, respectively, and using the signals obtained atstep S221-225 (the signals 4e-1 and 4e-2) for the current frame asthey are for the section from t2 to t5, for the first and secondchannels, respectively (step S221-226).
[0039] <SECOND EMBODIMENT>Difference between the downmix signal and a locallydecoded signal of monaural encoding for the section Y, which is atime section for which the monaural decoding unit 210 can obtain acomplete monaural decoded sound signal from the monaural code CM,may also be an encoding target of the additional encoding unit130. This embodiment is regarded as a second embodiment, andpoints different from the first embodiment will be described.
[0040] [Monaural encoding unit 120]In addition to encoding the downmix signal by thepredetermined encoding scheme to obtain and output the monauralcode CM, the monaural encoding unit 120 also obtains and outputs asignal obtained by decoding the monaural code CM, that is, amonaural locally decoded signal which is a locally decoded signalof the downmix signal for the section Y (step S122). In the caseof the example described above, in addition to obtaining themonaural code CM of the current frame, the monaural encoding unit120 obtains a locally decoded signal corresponding to the monauralcode CM of the current frame, that is, a locally decoded signal towhich the window in the shape of increasing in the 3.25 ms sectionfrom t1 to t2, being flat in the 16.75 ms section from t2 to t5and attenuating in the 3.25 ms section from t5 to t6 is applied.By combining a locally decoded signal corresponding to themonaural code CM of the immediately previous frame and the locallydecoded signal corresponding to the monaural code CM of thecurrent frame, for the section from t1 to t2, and using thelocally decoded signal corresponding to the monaural code CM ofthe current frame as it is, for the section from t2 to t5, themonaural encoding unit 120 obtains and outputs a locally decodedsignal for the 20 ms section from t1 to t5. As the locally decodedsignal corresponding to the monaural code CM of the immediatelyprevious frame for the section from t1 to t2, a signal stored inthe storage not shown in the monaural encoding unit 120 is used.Since a part of the locally decoded signal corresponding to themonaural code CM of the current frame for the section from t5 tot6 is used as "a locally decoded signal corresponding to themonaural code CM of an immediately previous frame" in theprocessing of the immediately following frame, the monauralencoding unit 120 stores the locally decoded signal for thesection from t5 to t6 obtained from the monaural code CM of thecurrent frame into the storage not shown in the monaural encodingunit 120.
[0041] [Additional encoding unit 130]In addition to the downmix signal, the monaural locallydecoded signal outputted by the monaural encoding unit 120 is alsoinputted to the additional encoding unit 130 as shown by a dashedline in Fig. 1. The additional encoding unit 130 encodes not onlythe downmix signal for the section X which is the encoding targetof the additional encoding unit 130 of the first embodiment, butalso a difference signal between the downmix signal and themonaural locally decoded signal for the section Y (a signalconfigured by subtraction between sample values of correspondingsamples) to obtain and output the additional code CA (step S132).For example, the additional encoding unit 130 can encode each ofthe downmix signal for the section X and the difference signal forthe section Y to obtain codes and obtain a concatenation of theobtained codes as the additional code CA. For the encoding, anencoding scheme similar to that of the additional encoding unit130 of the first embodiment can be used. Further, for example, theadditional encoding unit 130 may encode a signal obtained byconcatenating the difference signal for the section Y and thedownmix signal for the section X to obtain the addition code CA.Further, for example, like [Specific example 1 of additionalencoding unit 130] below, the additional encoding unit 130 canperform first additional encoding that encodes the downmix signalfor the section X to obtain a code (a first additional code CA1)and second additional encoding that encodes a signal obtained byconcatenating the difference signal for the section Y (that is, aquantization error signal of the monaural encoding unit 120) and adifference signal between the downmix signal for the section X andthe locally decoded signal of the first additional encoding (thatis, a quantization error signal of the first additional encoding)to obtain a code (a second additional code CA2), and obtain acombination of the first additional code CA1 and the secondadditional code CA2 as the additional code CA. According to[Specific example 1 of additional encoding unit 130], the secondadditional encoding is performed, with a concatenation of suchsignals that amplitude difference between two sections is assumedto be smaller than that between the difference signal for thesection Y and the downmix signal for the section X as an encodingtarget, and, as for the downmix signal itself, the firstadditional encoding is performed, with only a short time sectionas an encoding target. Therefore, efficient encoding can beexpected.
[0042] [Specific example 1 of additional encoding unit 130]First, the additional encoding unit 130 encodes theinputted downmix signal for the section X to obtain the firstadditional code CA1 (step S132-1; hereinafter also referred to as"first additional encoding") and obtains a locally decoded signalfor the section X corresponding to the first additional code CA1,that is, a locally decoded signal of the first encoding for thesection X (step S132-2). For the first additional encoding, anencoding scheme such as well-known scalar quantization or vectorquantization can be used. Next, the additional encoding unit 130obtains a difference signal (a signal configured by subtractionbetween sample values of corresponding samples) between theinputted downmix signal for the section X and the locally decodedsignal for the section X obtained at step S132-2 (step S132-3).Further, the additional encoding unit 130 obtains a differencesignal (a signal configured by subtraction between sample valuesof corresponding samples) between the downmix signal and themonaural locally decoded signal for the section Y (step S132-4).Next, the additional encoding unit 130 encodes a signal obtainedby concatenating the difference signal for the section Y obtainedat step S132-4 and the difference signal for the section Xobtained at step S132-3 to obtain the second additional code CA2(step S132-5; hereinafter also referred to as "second additionalencoding"). For the second additional encoding, an encoding schemeof collectively encoding a sample sequence obtained byconcatenating a sample sequence of the difference signal for thesection Y obtained at step S132-4 and a sample sequence of thedifference signal for the section X obtained at step S132-3, forexample, an encoding scheme using prediction in the time domain oran encoding scheme adapted to amplitude imbalance in the frequencydomain is used. Next, the additional encoding unit 130 outputs aconcatenation of the first additional code CA1 obtained at stepS132-1 and the second additional code CA2 obtained at step S132-5as the additional code CA (step S132-6).
[0043] Note that the additional encoding unit 130 may target aweighted difference signal for encoding instead of the differencesignal described above. That is, the additional encoding unit 130may code a weighted difference signal (a signal configured byweighted subtraction between sample values of correspondingsamples) between the downmix signal and the monaural locallydecoded signal for the section Y and the downmix signal for thesection X to obtain and output the additional code CA. In the caseof [Specific example 1 of additional encoding unit 130], theadditional encoding unit 130 can obtain the weighted differencesignal (a signal configured by weighted subtraction between samplevalues of corresponding samples) between the downmix signal andthe monaural locally decoded signal for the section Y as theprocessing of step S132-4. Similarly, the additional encoding unit130 may obtain the weighted difference signal (a signal configuredby weighted subtraction between sample values of correspondingsamples) between the inputted downmix signal for the section X andthe locally decoded signal for the section X obtained at stepS132-2 as the processing of step S132-3 of [Specific example 1 ofadditional encoding unit 130]. In these cases, the weight used forgeneration of each weighted difference signal can be encoded by awell-known encoding technology to obtain a code, and the obtainedcode (a code representing the weight) can be included in theadditional code CA. Though the above is the same for eachdifference signal in each embodiment described later, it is wellknown in the technological field of encoding that a weighteddifference signal is targeted by encoding instead of a differencesignal and that, in that case, the weight is also coded.Therefore, in the embodiments described later, each individualdescription will be omitted to avoid redundancy, and onlydescription in which a difference signal and a weighted differencesignal are written together using "or" and description in whichsubtraction and weighted subtraction are written together using"or" will be made.
[0044] [Additional decoding unit 230]The additional decoding unit 230 decodes the additionalcode CA to obtain and output not only the additional decodedsignal for the section X, which is an additional decoded signalobtained by the additional decoding unit 230 of the firstembodiment but also an additional decoded signal for the section Y(step S232). For the decoding, a decoding scheme corresponding tothe encoding scheme used by the additional encoding unit 130 atstep S132 is used. That is, if the additional encoding unit 130uses [Specific example 1 of additional encoding unit 130] at stepS132, the additional decoding unit 230 performs the processing of[Specific example 1 of additional decoding unit 230] describedbelow.
[0045] [Specific example 1 of additional decoding unit 230]First, the additional decoding unit 230 decodes the firstadditional code CA1 included in the additional code CA to obtain afirst decoded signal for the section X (step S232-1; hereinafteralso referred to as "first additional decoding"). For the firstadditional decoding, a decoding scheme corresponding to theencoding scheme used for the first additional encoding by theadditional encoding unit 130 is used. Further, the additionaldecoding unit 230 decodes the second additional code CA2 includedin the additional code CA to obtain a second decoded signal forthe sections Y and X (step S232-2; hereinafter also referred to as"second additional decoding"). For the second additional decoding,a decoding scheme corresponding to the encoding scheme used by theadditional encoding unit 130 for the second additional encoding isused, that is, a decoding scheme that yields, from the code, acollective sample sequence which is a concatenation of a samplesequence of the additional decoded signal for the section Y and asample sequence of the second decoded signal for the section X;for example, a decoding scheme using prediction in the time domainor a decoding scheme adapted to amplitude imbalance in thefrequency domain. Next, the additional decoding unit 230 obtains apart of the second decoded signal obtained at step S232-2 for thesection Y as the additional decoded signal for the section Y,obtains an addition signal (a signal configured by addition ofsample values of corresponding samples) of the first decodedsignal for the section X obtained at step S232-1 and a part of thesecond decoded signal for the section X obtained at step S232-2 asthe additional decoded signal for the section X, and outputs theadditional decoded signals for the sections Y and X (step S232-4).
[0046] Note that, when the additional encoding unit 130 targetsnot a difference signal but a weighted difference signal forencoding, a code representing a weight is also included in theadditional code CA, and thus the additional decoding unit 230 candecode a code of the additional code CA except the coderepresenting the weight to obtain and output an additional decodedsignal, and decode the code representing the weight, which isincluded in the additional code CA, to obtain and output theweight at step S232 described above. In the case of [Specificexample 1 of additional decoding unit 230], a code representing aweight for the section X, which is included in the additional codeCA, is decoded to obtain the weight for the section X; and, atstep S232-4, the additional decoding unit 230 can obtain the partof the second decoded signal obtained at step S232-2 for thesection Y as the additional decoded signal for the section Y,obtain a weighted addition signal (a signal configured by weightedaddition of sample values of corresponding samples) of the firstdecoded signal for the section X obtained at step S232-1 and thepart of the second decoded signal obtained at step S232-2 for thesection X as the additional decoded signal for the section X, andoutput the additional decoded signals for the sections Y and X anda weight for the section Y obtained by decoding a coderepresenting the weight for the section Y, which is included inthe additional code CA. Though the above is the same for additionof signals in each embodiment described later, it is well known inthe technological field of encoding that weighted addition(generation of a weighted sum signal) is performed instead ofaddition (generation of a sum signal) and that, in that case, aweight is also obtained from a code. Therefore, in the embodimentsdescribed later, respective descriptions will be omitted to avoidredundancy, and only description in which addition and weightedaddition are written together using "or" and description in whicha sum signal and a weighted sum signal are written together using"or" will be made.
[0047] [Stereo decoding unit 220]The stereo decoding unit 220 performs steps S222-1 andS222-2 below (step S222). The stereo decoding unit 220 obtains asignal concatenating a sum signal (a signal configured by additionof sample values of corresponding samples) of the monaural decodedsound signal for the section Y and the additional decoded signalfor the section Y and the additional decoded signal for thesection X, as a decoded downmix signal for the section Y+X (stepS222-1) instead of step S221-1 performed by the stereo decodingunit 220 of the first embodiment and obtains and outputs thedecoded sound signals of the two channels from the decoded downmixsignal obtained at step S222-1 by the upmix processing using thecharacteristic parameter obtained from the stereo code CS, usingthe decoded downmix signal obtained at step S222-1 instead of thedecoded downmix signal obtained at step S221-1 (step S222-2).
[0048] Note that, when the additional encoding unit 130 targetsnot a difference signal but a weighted difference signal forencoding, the stereo decoding unit 220 can obtain a signalconcatenating a weighted sum signal (a signal configured byweighted addition of sample values of corresponding samples) ofthe monaural decoded sound signal for the section Y and theadditional decoded signal for the section Y and the additionaldecoded signal for the section X, as the decoded downmix signalfor the section X+Y at step S222-1. For generation of the weightedsum signal (weighted addition of sample values of correspondingsamples) of the monaural decoded sound signal for the section Yand the additional decoded signal for the section Y, the weightfor the section Y outputted by the additional decoding unit 230can be used. Though the above is the same for addition of signalsin each embodiment described later, it is well known in thetechnological field of encoding that weighted addition (generationof a weighted sum signal) is performed instead of addition(generation of a sum signal) and that, in that case, a weight isalso obtained from a code as described in the description of theadditional decoding unit 230. Therefore, in the embodimentsdescribed later, respective description will be omitted to avoidredundancy, and only description in which addition and weightedaddition are written together using "or" and description in whicha sum signal and a weighted sum signal are written together using"or" will be made.
[0049] According to the second embodiment, in addition to thatthe algorithmic delay of stereo encoding / decoding is not largerthan the algorithmic delay of monaural encoding / decoding, it ispossible to cause the quality of a decoded downmix signal used forstereo decoding to be higher than the first embodiment. Therefore,it is possible to cause the sound quality of a decoded soundsignal of each channel obtained by stereo decoding to be high.That is, in the second embodiment, the monaural encodingprocessing by the monaural encoding unit 120 and the additionalencoding processing by the additional encoding unit 130 are usedas encoding processing for high-quality encoding of a downmixsignal; the monaural code CM and the additional code CA areobtained as codes that represents the favorable downmix signal;and the monaural decoding processing by the monaural decoding unit210 and the additional decoding processing by the additionaldecoding unit 230 are used as decoding processing for obtaining ahigh-quality decoded downmix signal. A code amount assigned toeach of the monaural code CM and the additional code CA can bearbitrarily determined according to purposes. In the case ofrealizing higher-quality stereo encoding / decoding in addition tostandard-quality monaural encoding / decoding, a larger code amountcan be assigned to the additional code CA. That is, "monauralcode" and "additional code" are mere convenient names from aviewpoint of stereo encoding / decoding. Since each of the monauralcode CM and the additional code CA is a part of a coderepresenting a downmix signal, one of the codes may be called "afirst downmix code", and the other may be called "a second downmixcode". When it is assumed that a larger code amount is assigned tothe additional code CA, the additional code CA may be called "adownmix code", "a downmix signal code" or the like. The above isthe same in a third embodiment and each embodiment based on thesecond embodiment described after the third embodiment.
[0050] <THIRD EMBODIMENT>There may be a case where it is possible for the stereodecoding unit 220 to obtain the decoded sound signals of the twochannels with a higher sound quality by using a decoded downmixsignal corresponding to a downmix signal obtained by mixing soundsignals of the two channels in the frequency domain, and it ispossible for the monaural decoding unit 210 to obtain a monauraldecoded sound signal with a higher sound quality when the monauralencoding unit 120 encodes a signal obtained by mixing soundsignals of the two channels in the time domain. In such a case, itis preferable that the stereo encoding unit 110 obtains thedownmix signal by mixing the sound signals of the two channelsinputted to the encoding device 100 in the frequency domain, thatthe monaural encoding unit 120 encodes a signal obtained by mixingthe sound signals of the two channels inputted to the encodingdevice 100 in the time domain, and that the additional encodingunit 130 also encodes difference between the signal obtained bymixing the sound signals of the two channels in the frequencydomain and the signal obtained by mixing the sound signals of thetwo channels in the time domain. This embodiment is regarded as athird embodiment, and points different from the second embodimentwill be mainly described.
[0051] [Stereo encoding unit 110]The stereo encoding unit 110 operates as described in thefirst embodiment similarly to the stereo encoding unit 110 of thesecond embodiment. However, as for the processing of obtaining thedownmix signal which is the signal obtained by mixing the soundsignals of the two channels, the processing is performed byprocessing of mixing the sound signals of the two channels in thefrequency domain, for example, like step S111-5B (step S113). Thatis, the stereo encoding unit 110 obtains a downmix signal obtainedby mixing the sound signals of the two channels in the frequencydomain. For example, in the processing of the current frame, thestereo encoding unit 110 can obtain a downmix signal, which is amonaural signal obtained by mixing the sound signals of the twochannels in the frequency domain, for the section from t1 to t6.When the encoding device 100 is also provided with the downmixunit 150, the stereo encoding unit 110 obtains and outputs thestereo code CS representing the characteristic parameter, which isa parameter representing the characteristic of the differencebetween the sound signals of the two channels inputted from thetwo-channel stereo input sound signal inputted to the encodingdevice 100 (step S113), and the downmix unit 150 obtains andoutputs the downmix signal, which is a signal obtained by mixingthe sound signals of the two channels in the frequency domain,from the two-channel stereo input sound signal inputted to theencoding device 100 (step S153).
[0052] [Monaural encoding target signal generation unit 140]The encoding device 100 of the third embodiment alsoincludes a monaural encoding target signal generation unit 140 asshown by a one-dot chain line in Fig. 1. The two-channel stereoinput sound signal inputted to the encoding device 100 is inputtedto the monaural encoding target signal generation unit 140. Themonaural encoding target signal generation unit 140 obtains amonaural encoding target signal, which is a monaural signal, fromthe inputted two-channel stereo input sound signal by processingof mixing the sound signals of the two channels in the time domain(step S143). For example, the monaural encoding target signalgeneration unit 140 obtains a sequence of averages ofcorresponding samples between the sample sequence of the soundsignal of the first channel and the sample sequence of the soundsignal of the second channel, as the monaural encoding targetsignal which is a signal obtained by mixing the sound signals ofthe two channels. That is, the monaural encoding target signalobtained by the monaural encoding target signal generation unit140 is a signal obtained by mixing the sound signals of the twochannels in the time domain. For example, in the processing of thecurrent frame, the monaural encoding target signal generation unit140 can obtain the monaural encoding target signal, which is themonaural signal obtained by mixing the sound signals of the twochannels in the time domain, for 20 ms from t3 to t7.
[0053] [Monaural encoding unit 120]The monaural encoding target signal outputted by themonaural encoding target signal generation unit 140 is inputted tothe monaural encoding unit 120 instead of the downmix signaloutputted by the stereo encoding unit 110 or the downmix unit 150.The monaural encoding unit 120 encodes the monaural encodingtarget signal to obtain and output the monaural code CM (stepS123). For example, in the processing of the current frame, themonaural encoding unit 120 encodes the signal for the section fromt1 to t6 obtained by applying the window in the shape ofincreasing in the section from t1 to t2 where the current frameand the immediately previous frame overlap, attenuating in thesection from t5 to t6 where the current frame and the immediatelyfollowing frame overlap and being flat in the section from t2 tot5 between the above sections, to the monaural encoding targetsignal, using the section from t6 to t7 of the monaural encodingtarget signal which is "a look-ahead section" for analysisprocessing, to obtain and output the monaural code CM.
[0054] [Additional encoding unit 130]Similarly to the additional encoding unit 130 of thesecond embodiment, the additional encoding unit 130 encodes adifference signal or weighted difference signal (a signalconfigured by subtraction or weighted subtraction between samplevalues of corresponding samples) between the downmix signal andthe monaural locally decoded signal for the section Y and thedownmix signal for the section X to obtain and output theadditional code CA (step S133). Here, the downmix signal for thesection Y is a signal obtained by mixing the sound signals of thetwo channels in the frequency domain, and the monaural locallydecoded signal for the section Y is a signal obtained by locallydecoding a signal obtained by mixing the sound signals of the twochannels in the time domain.
[0055] Note that, similarly to the additional encoding unit 130of the second embodiment, the additional encoding unit 130 of thethird embodiment can perform the first additional encoding thatencodes the downmix signal for the section X to obtain a firstadditional code CA and the second additional encoding that encodesa signal obtained by concatenating the difference signal orweighted difference signal for the section Y and the differencesignal or weighted difference signal between the downmix signalfor the section X and the locally decoded signal of the firstadditional encoding to obtain a second additional code CA2, andobtain a concatenation of the first additional code CA1 and thesecond additional code CA2 as the additional code CA, as describedin [Specific example 1 of additional encoding unit 130].
[0056] [Monaural decoding unit 210]Similarly to the monaural decoding unit 210 of the secondembodiment, the monaural decoding unit 210 obtains and outputs themonaural decoded sound signal for the section Y using the monauralcode CM (step S213). However, the monaural decoded sound signalobtained by the monaural decoding unit 120 of the third embodimentis a decoded signal of a signal obtained by mixing sound signalsof the two channels in the time domain.
[0057] [Additional decoding unit 230]Similarly to the additional decoding unit 230 of thesecond embodiment, the additional decoding unit 230 decodes theadditional code CA to obtain and output additional decoded signalsfor the sections Y and X, (step S233). However, the additionaldecoded signal for the section Y includes difference between asignal obtained by mixing sound signals of the two channels in thetime domain and the monaural decoded sound signal, and differencebetween a signal obtained by mixing the sound signals of the twochannels in the frequency domain and the signal obtained by mixingthe sound signals of the two channels in the time domain.
[0058] [Stereo decoding unit 220]The stereo decoding unit 220 performs steps S223-1 andS223-2 below (step S223). Similarly to the additional decodingunit 230 of the second embodiment, the stereo decoding unit 220obtains a signal obtained by concatenating a sum signal orweighted sum signal (a signal configured by addition or weightedaddition of sample values of corresponding samples) of themonaural decoded sound signal for the section Y and an additionaldecoded signal for the section Y and an additional decoded signalfor the section X, as the decoded downmix signal for the sectionY+X (step S223-1), and obtains and outputs the decoded soundsignals of the two channels from the decoded downmix signalobtained at step S223-1 by the upmix processing using thecharacteristic parameter obtained from the stereo code CS (stepS223-2). However, the sum signal for the section Y includes themonaural decoded sound signal obtained by performing monauralencoding / decoding of the signal obtained by mixing the soundsignals of the two channels in the time domain, difference betweenthe signal obtained by mixing the sound signals of the twochannels in the time domain and the monaural decoded sound signal,and difference between the signal obtained by mixing the soundsignals of the two channels in the frequency domain and the signalobtained by mixing the sound signals of the two channels in thetime domain.
[0059] <FOURTH EMBODIMENT>As for the section X, though a correct locally decodedsignal or decoded signal cannot be obtained by the monauralencoding unit 120 or the monaural decoding unit 210 without asignal and codes of the immediately following frame, an incompletelocally decoded signal or decoded signal can be obtained only witha signal and codes of frames up to the current frame. Therefore,each of the first to third embodiments may be changed so that, asfor the section X, the additional encoding unit 130 encodes notthe downmix signal itself but difference between the downmixsignal and a monaural locally decoded signal obtained from thesignal of the frames up to the current frame. This embodiment willbe described as fourth embodiments.
[0060] <<Fourth embodiment A>>First, for a fourth embodiment A, which is a fourthembodiment obtained by changing the second embodiment, descriptionwill be made mainly on points different from the secondembodiment.
[0061] [Monaural encoding unit 120]Similarly to the monaural encoding unit 120 of the secondembodiment, the downmix signal outputted by the stereo encodingunit 110 or the downmix unit 150 is inputted to the monauralencoding unit 120. The monaural encoding unit 120 obtains andoutputs the monaural code CM obtained by encoding the downmixsignal, and a signal obtained by decoding the monaural code CM ofthe frames up to the current frame, that is, a monaural locallydecoded signal, which is a locally decoded signal of the downmixsignal for the section Y+X (step S124). More specifically, inaddition to obtaining the monaural code CM of the current frame,the monaural encoding unit 120 obtains a locally decoded signalcorresponding to the monaural code CM of the current frame, thatis, a locally decoded signal to which the window in the shape ofincreasing in the 3.25 ms section from t1 to t2, being flat in the16.75 ms section from t2 to t5 and attenuating in the 3.25 mssection from t5 to t6 is applied; and, by combining a locallydecoded signal corresponding to the monaural code CM of theimmediately previous frame and the locally decoded signalcorresponding to the monaural code CM of the current frame, forthe section from t1 to t2, and using the locally decoded signalcorresponding to the monaural code CM of the current frame as itis, for the section from t2 to t6, the monaural encoding unit 120obtains and outputs a locally decoded signal for the section of23.25 ms from t1 to t6. However, a locally decoded signal for thesection from t5 to t6 is a locally decoded signal that becomes acomplete locally decoded signal by being combined with a locallydecoded signal to which the window in the increasing shape isapplied, the locally decoded signal being obtained in processingof the immediately following frame, and is an incomplete locallydecoded signal to which the window in the attenuating shape isapplied.
[0062] [Additional encoding unit 130]Similarly to the additional encoding unit 130 of thesecond embodiment, the downmix signal outputted by the stereoencoding unit 110 or the downmix unit 150, and the monaurallocally decoded signal outputted by the monaural encoding unit 120are inputted to the additional encoding unit 130. The additionalencoding unit 130 encodes a difference signal or weighteddifference signal (a signal configured by subtraction or weightedsubtraction between sample values of corresponding samples)between the downmix signal and the monaural locally decoded signalfor the section Y+X to obtain and output the additional code CA(step S134).
[0063] [Monaural decoding unit 210]Similarly to the monaural decoding unit 210 of the secondembodiment, the monaural code CM is inputted to the monauraldecoding unit 210. The monaural decoding unit 210 obtains andoutputs the monaural decoded sound signal for the section Y+Xusing the monaural code CM (step S214). However, the decodedsignal for the section X, that is, the section from t5 to t6 is adecoded signal that becomes a complete decoded signal by beingcombined with a decoded signal to which the window in theincreasing shape is applied, the decoded signal being obtained inprocessing of the immediately following frame, and is anincomplete decoded signal to which the window in the attenuatingshape is applied.
[0064] [Additional decoding unit 230]Similarly to the additional decoding unit 230 of thesecond embodiment, the additional code CA is inputted to theadditional decoding unit 230. The additional decoding unit 230decodes the additional code CA to obtain and output an additionaldecoded signal for the section Y+X (step S234).
[0065] [Stereo decoding unit 220]Similarly to the stereo decoding unit 220 of the secondembodiment, the monaural decoded sound signal outputted by themonaural decoding unit 210, the additional decoded signaloutputted by the additional decoding unit 230 and the stereo codeCS inputted to the decoding device 200 are inputted to the stereodecoding unit 220. The stereo decoding unit 220 obtains a sumsignal or weighted sum signal (a signal configured by addition orweighted addition of sample values of corresponding samples) ofthe monaural decoded sound signal and the additional decodedsignal for the section Y+X as the decoded downmix signal, andobtains and outputs the decoded sound signals of the two channelsfrom the decoded downmix signal by the upmix processing using thecharacteristic parameter obtained from the stereo code CS (stepS224).
[0066] <<Fourth embodiment B>>Note that, by replacing "the downmix signal outputted bythe stereo encoding unit 110 or the downmix unit 150" and "thedownmix signal" in the description of the monaural encoding unit120 of the fourth embodiment A with "the monaural encoding targetsignal outputted by the monaural encoding target signal generationunit 140" and "the monaural encoding target signal", respectively,description mainly on points different from the third embodimentfor a fourth embodiment B, which is a fourth embodiment obtainedby changing the third embodiment, is obtained.
[0067] <<Fourth embodiment C>>Further, by assuming that each of the monaural locallydecoded signal obtained by the monaural encoding unit 120, thedifference signal or weighted difference signal coded by theadditional encoding unit 130 and the additional decoded signalobtained by the additional decoding unit 230 in the description ofthe fourth embodiment A is the signal for the section X, and thestereo decoding unit 220 obtains a signal that is a concatenationof the monaural decoded sound signal for the section Y and a sumsignal or weighted sum signal of the monaural decoded sound signaland the additional decoded signal for the section X as the decodeddownmix signal, a fourth embodiment C, which is a fourthembodiment obtained by changing the first embodiment, is obtained.
[0068] <FIFTH EMBODIMENT>The downmix signal for the section X includes a part thatcan be predicted from the monaural locally decoded signal for thesection Y. Therefore, for the section X, the difference betweenthe downmix signal and a predicted signal from the monaurallocally decoded signal for the section Y may be encoded by theadditional encoding unit 130 in each of the first to fourthembodiments. This embodiment will be described as fifthembodiments.
[0069] <<Fifth embodiment A>>First, regarding fifth embodiments obtained by changingthe second embodiment, the third embodiment, the fourth embodimentA and the fourth embodiment B, respectively, as fifth embodimentA, points different from each of the second embodiment, the thirdembodiment, the fourth embodiment A and the fourth embodiment Bwill be described.
[0070] [Additional encoding unit 130]The additional encoding unit 130 performs steps S135A-1and S135A-2 below (step S135A). First, the additional encodingunit 130 obtains a predicted signal for the monaural locallydecoded signal for the section X from the inputted monaurallocally decoded signal for the section Y or the section Y+X(however, for the section X, an incomplete monaural locallydecoded signal as described above) using a predetermined wellknown prediction technology (step S135A-1). Note that, in the caseof the fifth embodiment obtained by changing the fourth embodimentA or the fifth embodiment obtained by changing the fourthembodiment B, the inputted incomplete monaural locally decodedsignal for the section X is included in the predicted signal forthe section X. Next, the additional encoding unit 130 encodes thedifference signal or weighted difference signal (a signalconfigured by subtraction or weighted subtraction between samplevalues of corresponding samples) between the downmix signal andthe monaural locally decoded signal for the section Y and adifference signal or weighted difference signal (a signalconfigured by subtraction or weighted subtraction between samplevalues of corresponding samples) between the downmix signal forthe section X and the predicted signal obtained at step S135A-1 toobtain and output the additional code CA (step S135A-2). Forexample, a signal obtained by concatenating the difference signalfor the section Y and the difference signal for the section X maybe encoded to obtain the additional code CA. Further, for example,each of the difference signal for the section Y and the differencesignal for the section X may be encoded to obtain a code, and aconcatenation of the obtained codes may be obtained as theadditional code CA. For the encoding, an encoding scheme similarto that of the additional encoding unit 130 of each of the secondembodiment, the third embodiment, the fourth embodiment A and thefourth embodiment B can be used.
[0071] [Stereo decoding unit 220]The stereo decoding unit 220 performs steps S225A-0 toS225A-2 below (step S225A). First, the stereo decoding unit 220obtains the predicted signal for the section X from the monauraldecoded sound signal for the section Y or the section Y+X, usingthe same prediction technology used by the additional encodingunit 130 at step S135 (step S225A-0). Next, the stereo decodingunit 220 obtains a signal obtained by concatenating a sum signalor weighted sum signal (a signal configured by addition orweighted addition of sample values of corresponding samples) ofthe monaural decoded sound signal and the additional decodedsignal for the section Y and a sum signal or weighted sum signal(a signal configured by addition or weighted addition of samplevalues of corresponding samples) of the additional decoded signaland the predicted signal for the section X as the decoded downmixsignal for the section Y+X (step S225A-1). Next, the stereodecoding unit 220 obtains and outputs the decoded sound signals ofthe two channels from the decoded downmix signal obtained at stepS225A-1, by the upmix processing using the characteristicparameter obtained from the stereo code CS (step S225A-2).
[0072] <<Fifth embodiment B>>Next, regarding a fifth embodiment obtained by changingeach of the first embodiment and the fourth embodiment C as afifth embodiment B, points different from each of the firstembodiment and the fourth embodiment C will be described.
[0073] [Monaural encoding unit 120]In addition to the monaural code CM obtained by encodingthe downmix signal, the monaural encoding unit 120 also obtainsand outputs, for the section Y or the section Y+X, a signalobtained by decoding the monaural code CM of the frames up to thecurrent frame, that is, a monaural locally decoded signal which isa locally decoded signal of the inputted downmix signal (stepS125B). However, as described above, the monaural locally decodedsignal for the section X is an incomplete monaural locally decodedsignal.
[0074] [Additional encoding unit 130]The additional encoding unit 130 performs steps S135B-1and S135B-2 below (step S135B). First, the additional encodingunit 130 obtains the predicted signal for the monaural locallydecoded signal for the section X from the inputted monaurallocally decoded signal for the section Y or the section Y+X(however, for the section X, an incomplete monaural locallydecoded signal as described above) using the predetermined wellknown prediction technology (step S135B-1). In the case of thefifth embodiment obtained by changing the fourth embodiment C, theinputted monaural locally decoded signal for the section X isincluded in the predicted signal for the section X. Next, theadditional encoding unit 130 encodes the difference signal orweighted difference signal between the downmix signal for thesection X and the predicted signal (a signal configured bysubtraction or weighted subtraction between sample values ofcorresponding samples) obtained at step S135B-1 to obtain andoutput the additional code CA (step S135B-2). For the encoding,for example, an encoding scheme similar to that of the additionalencoding unit 130 of each of the first embodiment and the fourthembodiment C can be used.
[0075] [Stereo decoding unit 220]The stereo decoding unit 220 performs steps S225B-0 toS225B-2 below (step S225B). First, the stereo decoding unit 220obtains the predicted signal for the section X from the monauraldecoded sound signal for the section Y or the section Y+X, usingthe same prediction technology used by the additional encodingunit 130 (step S225B-0). Next, the stereo decoding unit 220obtains a signal that is a concatenation of the monaural decodedsound signal for the section Y and a sum signal or weighted sumsignal (a signal configured by addition or weighted addition ofsample values of corresponding samples) of the additional decodedsignal and the predicted signal for the section X as the decodeddownmix signal for the section Y+X (step S225B-1). Next, thestereo decoding unit 220 obtains and outputs the decoded soundsignals of the two channels from the decoded downmix signalobtained at step S225B-1, by the upmix processing using thecharacteristic parameter obtained from the stereo code CS (stepS225B-2).
[0076] <SIXTH EMBODIMENT>In the first to fifth embodiments, by the decoding device200 using the additional code CA obtained by the encoding device100, the decoded downmix signal for the section X used by thestereo decoding unit 220 is obtained at least by decoding theadditional code CA. However, without the decoding device 200 usingthe additional code CA, a predicted signal from the monauraldecoded sound signal for the section Y may be used as the decodeddownmix signal for the section X used by the stereo decoding unit220. This embodiment is regarded as a sixth embodiment, and pointsdifferent from the first embodiment will be described.
[0077] <<Encoding device 100>>The encoding device 100 of the sixth embodiment isdifferent from the encoding device 100 of the first embodiment inthat the encoding device 100 of the sixth embodiment does notinclude the additional encoding unit 130, does not code thedownmix signal for the section X, and does not obtain theadditional code CA. That is, the encoding device 100 of the sixthembodiment includes the stereo encoding unit 110 and the monauralencoding unit 120, and the stereo encoding unit 110 and themonaural encoding unit 120 operates as the stereo encoding unit110 and the monaural encoding unit 120 of the first embodiment,respectively.
[0078] <<Decoding device 200>>The decoding device 200 of the sixth embodiment does notinclude the additional decoding unit 230 for decoding theadditional code CA but includes the monaural decoding unit 210 andthe stereo decoding unit 220. Though the monaural decoding unit210 of the sixth embodiment operates as the monaural decoding unit210 of the first embodiment, the monaural decoding unit 210 alsooutputs the monaural decoded sound signal for the section X whenthe stereo decoding unit 220 uses the monaural decoded soundsignal for the section Y+X. Further, the stereo decoding unit 220of the sixth embodiment operates as described below that isdifferent from the operation of the stereo decoding unit 220 ofthe first embodiment.
[0079] [Stereo decoding unit 220]The stereo decoding unit 220 performs steps S226-0 toS226-2 below (step S226). First, the stereo decoding unit 220obtains the predicted signal for the section X from the monauraldecoded sound signal for the section Y or the section Y+X, using apredetermined well-known prediction technology similar to that ofthe fifth embodiment (step S226-0). Next, the stereo decoding unit220 obtains a signal that is a concatenation of the monauraldecoded sound signal for the section Y and the predicted signalfor the section X, as the decoded downmix signal for the sectionY+X (step S226-1) and obtains and outputs the decoded soundsignals of the two channels from the decoded downmix signalobtained at step S226-1 by the upmix processing using thecharacteristic parameter obtained from the stereo code CS (stepS226-2).
[0080] <SEVENTH EMBODIMENT>In each embodiment described above, description has beenmade on an example in which sound signals of two channels arehandled to simplify the description. However, the number ofchannels is not limited thereto and is only required to be two ormore. When the number of channels is indicated by C (C is aninteger of 2 or larger), each embodiment described above can beimplemented, with "two channels" replaced with "C channels" (C isan integer of 2 or larger).
[0081] For example, the encoding device 100 of the first to fifthembodiments can obtain the stereo code CS, the monaural code CMand the additional code CA from inputted sound signals of the Cchannels; the encoding device 100 of the sixth embodiment canobtain the stereo code CS and the monaural code CM from theinputted sound signals of the C channels; the stereo encoding unit110 can obtain and output a code representing informationcorresponding to difference between channels of the inputted soundsignals of the C channels, as the stereo code CS; the stereoencoding unit 110 or the downmix unit 150 can obtain and output asignal obtained by mixing the inputted sound signals of the Cchannels as the downmix signal; and the monaural encoding targetsignal generation unit 140 can obtain and output a signal obtainedby mixing the inputted sound signals of the C channels in the timedomain, as the monaural encoding target signal. For example, theinformation corresponding to the difference between channels ofthe sound signals of the C channels is, for each of C-1 channelsother than a reference channel, information corresponding todifference between the sound signal of the channel and the soundsignal of the reference channel.
[0082] Similarly, the decoding device 200 of the first to fifthembodiments can obtain and output the decoded sound signals of theC channels based on the inputted monaural code CM, additional codeCA and stereo code CS; the decoding device 200 of the sixthembodiment can obtain and output the decoded sound signals of theC channels based on the inputted monaural code CM and stereo codeCS; and the stereo decoding unit 220 can obtain and output thedecoded sound signals of the C channels from the decoded downmixsignal by the upmix processing using the characteristic parameterobtained based on the inputted stereo code CS. More specifically,the stereo decoding unit 220 can obtain and output the decodedsound signals of the C channels, regarding the decoded downmixsignal as a signal obtained by mixing the decoded sound signals ofthe C channels, and regarding the characteristic parameterobtained based on the inputted stereo code CS as informationrepresenting the characteristic of the difference between channelsof the decoded sound signals of the C channels.
[0083] <Program and recording medium>The processing of each unit of each encoding device andeach decoding device described above may be realized by acomputer. In this case, the processing content of a function thateach device should have is written by a program. Then, by loadingthe program onto a storage 1020 of a computer shown in Fig. 12 andcausing a calculation unit 1010, an input unit 1030, an outputunit 1040 and the like to operate, various kinds of processingfunctions of each of the devices described above are realized onthe computer.
[0084] The program in which the processing content is written canbe recorded in a computer-readable recording medium. The computerreadable recording medium is, for example, a non-transitoryrecording medium, specifically, a magnetic recording device, anoptical disk or the like.
[0085] Further, distribution of this program is performed, forexample, by sales, transfer, lending or the like of a portablerecording medium such as a DVD or a CD-ROM in which the program isrecorded. Furthermore, a configuration is also possible in whichthis program is stored in a storage of a server computer, and isdistributed by transfer from the server computer to othercomputers via a network.
[0086] The computer that executes such a program once stores theprogram recorded in a portable recording medium or transferredfrom a server computer into an auxiliary recording unit 1050,which is its own non-transitory storage, first. Then, at the timeof executing processing, the computer loads the program stored inthe auxiliary recording unit 1050, which is its own non-transitorystorage, into the storage 1020 and executes the processingaccording to the loaded program. Further, as another executionform of this program, a computer may directly load the programfrom a portable recording medium into the storage 1020 and executethe processing according to the program. Furthermore, each time aprogram is transferred to the computer from a sever computer, thecomputer may sequentially execute processing according to thereceived program. Further, a configuration is also possible inwhich the above processing is executed by a so-called ASP(application service provider) type service in which, withouttransferring the program to the computer from the server computer,the processing functions are realized only by an instruction toexecute the program and acquisition of a result. Note that it isassumed that, as the program described herein, information whichis provided for processing by an electronic calculator and isequivalent to a program (data or the like which is not a directcommand to the computer but has a nature of specifying processingof the computer) is included.
[0087] Though it is assumed in the description above that thedevice is configured by causing a predetermined program to beexecuted on a computer, at least a part of the processing contentmay be realized by hardware.
[0088] In addition, it goes without saying that it is possible toappropriately make changes within a range not departing from thespirit of the present invention.
Claims
1. A sound signal encoding method for encoding an inputted sound signal of C channels (C is an integer of 2 or larger) for each frame, the sound signal encoding method comprising: as processing for a current frame, a stereo encoding step of obtaining and outputting a stereo code representing a characteristic parameter which is a parameter representing a characteristic of difference between channels of the sound signal having the C channels; a downmix step of obtaining a signal obtained by mixing the sound signal having the C channels as a downmix signal; and a monaural encoding step of encoding the downmix signal to obtain and output a monaural code, wherein the monaural encoding step encodes the downmix signal by an encoding scheme that includes processing of applying a window having overlap between frames to obtain the monaural code, and the sound signal encoding method further comprises an additional encoding step of encoding a part of the downmix signal for a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as "a section X") to obtain and output an additional code.
2. The sound signal encoding method according to claim 1, wherein the monaural encoding step also obtains a monaural locally decoded signal corresponding to the monaural code; and the additional encoding step encodes a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between a part of the downmix signal for a section except the section X (hereinafter referred to as "a section Y") and the monaural locally decoded signal for the section Y, and the downmix signal for the section X to obtain the additional code.
3. The sound signal encoding method according to claim 2, wherein the downmix step obtains a signal obtained by mixing the sound signal having the C channels in a frequency domain as the downmix signal; the sound signal encoding method further comprises a monaural encoding target signal generation step of obtaining a signal by mixing the sound signal having the C channels in a time domain as a monaural encoding target signal; and the monaural encoding step encodes the monaural encoding target signal by the encoding scheme that includes the processing of applying a window having overlap between frames to obtain the monaural code.
4. The sound signal encoding method according to claim 2 or 3, wherein the additional encoding step encodes the section X of the downmix signal to obtain a first additional code and a locally decoded signal for the section X (hereinafter referred to as "a first additional locally decoded signal") corresponding to the first additional code, encodes a signal obtained by concatenating a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the first additional locally decoded signal for the section X, by an encoding scheme of collectively encoding a sample sequence, to obtain a second additional code, and obtain a combination of the first additional code and the second additional code as the additional code.
5. The sound signal encoding method according to claim 2 or 3, wherein the additional encoding step obtains a predicted signal for the section X of the monaural locally decoded signal, from the monaural locally decoded signal for the section Y or the monaural locally decoded signal for the section Y and the section X, and encodes the signal configured by subtraction or weighted subtraction between the sample values of the corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the predicted signal for the section X to obtain the additional code.
6. A sound signal encoding device for encoding an inputted sound signal having C channels (C is an integer of 2 or larger) for each frame, the sound signal encoding device comprising: a stereo encoding unit configured to obtain and output a stereo code representing a characteristic parameter which is a parameter representing a characteristic of difference between channels of the sound signal having the C channels, as processing for a current frame; a downmix unit configured to obtain a signal by mixing the sound signal having the C channels as a downmix signal, as processing for the current frame; and a monaural encoding unit configured to encode the downmix signal to obtain and output a monaural code, as processing for the current frame, wherein the monaural encoding unit encodes the downmix signal by an encoding scheme that includes processing of applying a window having overlap between frames to obtain the monaural code, and the sound signal encoding device further comprises an additional encoding unit configured to encode a part of the downmix signal for a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as "a section X") to obtain and output an additional code.
7. The sound signal encoding device according to claim 6, wherein the monaural encoding unit also obtains a monaural locally decoded signal corresponding to the monaural code; and the additional encoding unit encodes a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between a part of the downmix signal for a section except the section X (hereinafter referred to as "a section Y") and the monaural locally decoded signal for the section Y, and the downmix signal for the section X to obtain the additional code.
8. The sound signal encoding device according to claim 7, wherein the downmix unit obtains a signal obtained by mixing the sound signal having the C channels in a frequency domain as the downmix signal; the sound signal encoding device further comprises a monaural encoding target signal generation unit configured to obtain a signal by mixing the sound signal having the C channels in a time domain as a monaural encoding target signal; and the monaural encoding unit encodes the monaural encoding target signal by the encoding scheme that includes the processing of applying a window having overlap between frames to obtain the monaural code.
9. The sound signal encoding device according to claim 7 or 8, wherein the additional encoding unit encodes the section X of the downmix signal to obtain a first additional code and a locally decoded signal for the section X (hereinafter referred to as "a first additional locally decoded signal"), corresponding to the first additional code, encodes a signal obtained by concatenating a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the first additional locally decoded signal for the section X, by an encoding scheme of collectively encoding a sample sequence, to obtain a second additional code, and obtain a combination of the first additional code and the second additional code as the additional code.
10. The sound signal encoding device according to claim 7 or 8, wherein the additional encoding unit obtains a predicted signal for the section X of the monaural locally decoded signal, from the monaural locally decoded signal for the section Y or the monaural locally decoded signal for the section Y and the section X, and encodes the signal configured by subtraction or weighted subtraction between the sample values of the corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the predicted signal for the section X to obtain the additional code.
11. A program for causing a computer to execute each step of the sound signal encoding method according to any one of claims 1 to 5.
12. A computer-readable recording medium in which a program for causing a computer to execute each step of the sound signal encoding method according to any one of claims 1 to 5 is recorded.