Sound signal decoding method, sound signal decoding device, computer program product, and recording medium

By introducing additional decoding steps and inter-frame overlap window processing technology in the encoding and decoding process, the problem of large delay of stereo encoding/decoding algorithm is solved, and the effect of closeness to mono encoding/decoding delay is achieved, and the control process of the multi-point control device is simplified.

CN115917643BActive Publication Date: 2025-05-02NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080102308.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-24
Publication Date
2025-05-02
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

In the prior art, the delay of the stereo encoding/decoding algorithm is greater than that of the mono encoding/decoding algorithm, which leads to complex control in particular in multi-point control devices.

Method used

By introducing additional decoding steps in the encoding and decoding process, using inter-frame overlapping window processing technology, combined with the processing of stereo code and mono code, embedded encoding/decoding of multiple channel sound signals and mono sound signals is realized.

Benefits of technology

It effectively reduces the delay of the stereo encoding/decoding algorithm, makes it close to the delay of the mono encoding/decoding algorithm, and simplifies the control process of the multi-point control device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115917643B_ABST
    Figure CN115917643B_ABST
Patent Text Reader

Abstract

Provided is embedded decoding of a plurality of channel sound signals and a monophonic sound signal whose algorithm delay of stereo encoding / decoding is not greater than that of monophonic encoding / decoding. A decoding device (200) decodes the code in units of frames to obtain decoded sound signals of a plurality of channels. A monophonic decoding unit (210) decodes the monophonic code in a decoding method including applying a process of overlapping windows between frames to obtain a monophonic decoded sound signal. An additional decoding unit (230) decodes the additional code to obtain an additional decoded signal of an overlapping interval X between a current frame and an immediately subsequent frame. A stereo decoding unit (220) obtains a decoded downmix signal, which is a signal obtained by combining a signal of an interval Y other than interval X in the monophonic decoded sound signal and an additional decoded signal of interval X. A decoded sound signal is obtained from the decoded downmix signal by an upmix process using characteristic parameters obtained from the stereo code and is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for performing embedded encoding / decoding on a plurality of channel sound signals and a single channel sound signal. Background Art

[0002] As a technology for performing embedded encoding / decoding of multi-channel audio signals and monaural audio signals, there is the technology of Non-Patent Document 1. Figure 5 The encoding device 500 and Figure 6 The decoding device 600 illustrated in the figure will be used to describe the outline of the technology of non-patent document 1. The stereo encoding unit 510 of the encoding device 500 obtains a stereo code CS representing a characteristic parameter and a signal obtained by mixing the stereo input sound signal from a sound signal of multiple channels, i.e., a stereo input sound signal, for each predetermined time interval, i.e., each frame. The characteristic parameter is a parameter representing the characteristic of the difference between channels in the stereo input sound signal. The mono encoding unit 520 of the encoding device 500 encodes the downmix signal for each frame to obtain a mono code CM. The mono decoding unit 610 of the decoding device 600 decodes the mono code CM for each frame to obtain a mono decoded sound signal as a decoded signal of the downmix signal. The stereo decoding unit 620 of the decoding device 600 decodes the stereo code CS for each frame to obtain a characteristic parameter as a parameter representing the characteristic of the difference between channels, and performs a process (so-called upmix process) of obtaining a stereo decoded sound signal from the mono decoded sound signal and the characteristic parameter.

[0003] As a monaural encoding / decoding method for obtaining a high-quality monaural decoded audio signal, there is an encoding / decoding method of the 3GPP EVS standard described in non-patent document 2. As the monaural encoding / decoding method of non-patent document 1, if a high-quality monaural encoding / decoding method such as non-patent document 2 is used, it is possible to achieve higher-quality embedded encoding / decoding of multi-channel audio signals and monaural audio signals.

[0004] Prior art literature

[0005] Non-patent literature

[0006] Non-patent literature 1: Jeroen Breebaart et al., "Parametric Coding of StereoAudio", EURASIP Journal on Applied Signal Processing, pp. 1305-1322, 2005: 9.

[0007] Non-patent document 2: 3GPP, "Codec for Enhanced Voice Services (EVS); Detailed algorithmic description", TS26.445. Summary of the invention

[0008] Problems to be solved by the invention

[0009] The upmixing process of Non-Patent Document 1 is a signal processing in the frequency domain including applying a process of overlapping windows between adjacent frames to a mono decoded sound signal. On the other hand, the mono encoding / decoding method of Non-Patent Document 2 also includes a process of applying a process of overlapping windows between adjacent frames. That is, on both the decoding side of the stereo encoding / decoding method of Non-Patent Document 1 and the decoding side of the mono encoding / decoding method of Non-Patent Document 2, a signal of a slanted window in a shape of attenuating the signal obtained by decoding the code of the front frame is synthesized with a signal of a slanted window in a shape of increasing the signal obtained by decoding the code of the rear frame for a predetermined range of the boundary part of the frame, thereby obtaining a decoded sound signal. Therefore, if a mono encoding / decoding method such as Non-Patent Document 2 is used as a mono encoding / decoding method of embedded encoding / decoding such as Non-Patent Document 1, there is a problem that the stereo decoded sound signal is delayed by the window in the upmixing process compared with the mono decoded sound signal, that is, there is a problem that the algorithm delay of the stereo encoding / decoding is larger than that of the mono encoding / decoding.

[0010] For example, in a multipoint control unit (MCU) for conducting a telephone conference at multiple locations, it is usually performed to switch which signal from which location is output to which location for each predetermined time interval. It is assumed that it is difficult to perform control in a state where a stereo decoded audio signal is delayed by a window in the upmixing process compared to a mono decoded audio signal, and it is implemented to perform control in a state where the stereo decoded audio signal is delayed by one frame compared to the mono decoded audio signal. That is, in a communication system including a multipoint control unit, the above-mentioned problem becomes more significant, and there is a possibility that the delay of the algorithm of stereo encoding / decoding is larger by one frame than that of the algorithm of mono encoding / decoding. In addition, if the stereo decoded audio signal is delayed by one frame compared to the mono decoded audio signal, the control itself can be switched for each predetermined time interval, but for each time interval, the control of which mono decoded audio signal from which location is combined with which stereo decoded audio signal from which location and outputted may become complicated due to the difference in delay between the mono decoded audio signal and the stereo decoded audio signal.

[0011] The present invention has been made in view of such a problem, and an object of the present invention is to provide embedded encoding / decoding of audio signals of multiple channels and a monaural audio signal in which the algorithm delay of stereo encoding / decoding is not greater than the algorithm delay of monaural encoding / decoding.

[0012] Means for solving problems

[0013] In order to solve the above-mentioned problems, a sound signal decoding method according to one aspect of the present invention is a sound signal decoding method for decoding an input code in units of frames to obtain a decoded sound signal of C channels (C is an integer greater than or equal to 2), and includes, as processing of a current frame, a mono decoding step of decoding a mono code included in the input code in a decoding method including applying a process of overlapping windows between frames to obtain a mono decoded sound signal; an additional decoding step of decoding an additional code included in the input code to obtain an additional decoded signal, the additional decoded signal being a mono decoded signal of an interval X, which is an overlapping interval between the current frame and the next frame; and a stereo decoding step of obtaining a decoded down-mix signal, the decoded down-mix signal being a signal obtained by connecting a signal of an interval other than the interval X in the mono decoded sound signal and the additional decoded signal of the interval X, and obtaining and outputting a decoded sound signal of C channels from the decoded down-mix signal by an up-mix process using a feature parameter obtained from the stereo code included in the input code.

[0014] Effects of the Invention

[0015] According to the present invention, embedded encoding / decoding of a plurality of channel sound signals and a monaural sound signal can be provided in which the algorithm delay of stereo encoding / decoding is not greater than the algorithm delay of monaural encoding / decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a block diagram showing an example of an encoding device according to each embodiment.

[0017] Figure 2 This is a flowchart showing an example of processing by the encoding device according to each embodiment.

[0018] Figure 3 It is a block diagram showing an example of a decoding device according to each embodiment.

[0019] Figure 4 This is a flowchart showing an example of processing by the decoding device according to each embodiment.

[0020] Figure 5 is a block diagram showing an example of a conventional encoding device.

[0021] Figure 6 This is a block diagram showing an example of a conventional decoding device.

[0022] Figure 7 This is a diagram schematically showing each signal of the encoding device of non-patent document 2.

[0023] Figure 8 This is a diagram schematically showing the delays of various signals and algorithms in the decoding device of Non-Patent Document 2.

[0024] Fig. 9 This is a diagram schematically showing delays of each signal and algorithm of the decoding device of Non-Patent Document 1 when the monaural encoding / decoding method of Non-Patent Document 2 is used as the monaural encoding / decoding method.

[0025] Fig.10 It is a diagram schematically showing delays of various signals and algorithms in the decoding device of the present invention.

[0026] Fig.11 It is a diagram schematically showing each signal of the encoding device of the present invention.

[0027] Fig.12 This is a diagram showing an example of a functional configuration of a computer that realizes each device in each embodiment. DETAILED DESCRIPTION

[0028] Before describing each embodiment, first, with respect to the background technology and the delay of each signal and algorithm for encoding / decoding of the first embodiment, the delay of each signal when the frame length is 20 ms is schematically illustrated. Figures 7 to 11 Provide explanation. Figures 7 to 11 The horizontal axis of each figure is the time axis. In the following, since an example of processing the current frame at time t7 is described, the axis arranged at the top of each figure is marked with "past" on the left and "future" on the right, and an upward arrow is marked at t7, which is the time when the current frame is processed. Figures 7 to 11 In FIG. 1 , for each signal, it is schematically shown which time interval the signal is, and whether the window is in an increasing shape, a flat shape, or a decaying shape when the window is applied. More specifically, it is not important in the description here what the window function is exactly, so in FIG. Figures 7 to 11 In the figure, regarding the intervals of the shape of window increase and the shape of window decrease, in order to visually express the signal that becomes the application window during synthesis, the interval of the shape of window increase is represented by a triangle including a straight line inclined upward to the right, and the interval of the shape of window decrease is represented by a triangle including a straight line inclined downward to the right. In addition, in order to avoid complicated statement expressions, the statements such as "to" and "after" are used to determine the time of the start of each interval, but as can be understood by those skilled in the art, the actual start of each interval is the time after the time just recorded, and the actual start of the digital signal of each interval is the sample after the time just recorded.

[0029] Figure 7 This is a diagram schematically showing various signals of the encoding device of non-patent document 2 that processes the current frame at the time point t7. The encoding device of non-patent document 2 can use signal 1a, which is a monophonic sound signal before t7, for the processing of the current frame. In the encoding device of non-patent document 2, in the processing of the current frame, the 8.75 ms interval from t6 to t7 in signal 1a is used as a so-called "pre-read interval" for analysis, and the signal obtained by applying a window to the signal of the 23.25 ms interval from t1 to t6 in signal 1a, that is, signal 1b, is encoded to obtain a monophonic code and output. The shape of the window is a shape that increases in the interval of 3.25 ms from t1 to t2, is flat in the interval of 16.75 ms from t2 to t5, and is attenuated in the interval of 3.25 ms from t5 to t6. That is, signal 1b is a monophonic sound signal corresponding to the monophonic code obtained by processing the current frame. In the encoding device of non-patent document 2, as the processing of the immediately preceding frame, the same processing is performed at the time point when the monaural sound signal before t3 is input, and the encoding of the signal 1c is completed by applying the window of the shape that attenuates from the interval t1 to t2 to the monaural sound signal of the interval 23.25ms before t2. That is, the signal 1c is a monaural sound signal corresponding to the monaural code obtained by the processing of the immediately preceding frame, and the interval from t1 to t2 is the overlapping interval of the current frame and the immediately preceding frame. In addition, as the processing of the immediately succeeding frame, the encoding device of non-patent document 2 encodes the signal 1d which is a signal obtained by applying the window of the shape that increases from the interval t5 to t6 to the monaural sound signal of the interval 23.25ms after t5. That is, the signal 1d is a monaural sound signal corresponding to the monaural code obtained by the processing of the immediately succeeding frame, and the interval from t5 to t6 is the overlapping interval of the current frame and the immediately succeeding frame.

[0030] Figure 8The figure schematically shows various signals of the decoding device of non-patent document 2 which processes the current frame at the time point t7 when the monaural code of the current frame is input from the encoding device of non-patent document 2. In the processing of the current frame, the decoding device of non-patent document 2 obtains a decoded sound signal, i.e., signal 2a, from the monaural code of the current frame in the interval from t1 to t6. This signal 2a is a decoded sound signal corresponding to signal 1b, and is a signal of an application window of a shape that increases from t1 to t2, is flat from t2 to t5, and is attenuated from t5 to t6. As the processing of the immediately preceding frame, the decoding device of non-patent document 2 obtains a decoded sound signal, i.e., signal 2b, from the application window of a shape that attenuates from t1 to t2 obtained from the monaural code of the immediately preceding frame at the time point t3 when the monaural code of the immediately preceding frame is input. In addition, as processing of the immediately following frame, the decoding device of non-patent document 2 obtains a decoded sound signal, i.e., signal 2c, of the interval of 23.25ms after t5 of the application window of the shape increased from the interval of t5 to t6 from the monaural code of the immediately following frame. However, since signal 2c cannot be obtained at the time point of t7, at the time point of t7, although an incomplete decoded sound signal can be obtained for the interval of t5 to t6, a complete decoded sound signal cannot be obtained. Therefore, at the time point of t7, the decoding device of non-patent document 2 synthesizes signal 2b obtained by the immediately preceding processing and signal 2a obtained by the processing of the current frame for the interval of t1 to t2, and directly uses signal 2a obtained by the processing of the current frame for the interval of t2 to t5, thereby obtaining and outputting a mono decoded sound signal, i.e., signal 2d, of the interval of 20ms from t1 to t5. The decoding device of non-patent document 2 obtains the decoded audio signal of the interval starting from t1 at time point t7, so the algorithm delay of the monaural encoding / decoding method of non-patent document 2 is the time length from t1 to t7, that is, 32 ms.

[0031] Fig. 9This is a diagram schematically showing each signal of the decoding device 600 of Patent Document 1 when the mono decoding unit 610 uses the mono decoding method of Non-Patent Document 2. At time t7, the stereo decoding unit 620 uses the signal 3a which is the mono decoded sound signal until t5 completely obtained by the mono decoding unit 610 to perform stereo decoding processing (upmixing processing) of the current frame. Specifically, the stereo decoding unit 620 adopts a shape that increases in the interval of 3.25ms from t0 to t1, is flat in the interval of 16.75ms from t1 to t4, and attenuates in the interval of 3.25ms from t4 to t5, i.e., signal 3b, for signal 3a, and obtains a decoded sound signal from t0 to t5, i.e., signal 3c-i (i is the channel number), which is a window with the same shape as signal 3b, for each channel. As a process of the immediately preceding frame, the stereo decoding unit 620 has obtained, at time t3, decoded audio signals of each channel in the interval of 23.25 ms from t1 to the application window of the shape of attenuation from t0 to t1. In addition, as a process of the immediately succeeding frame, the stereo decoding unit 620 obtains decoded audio signals of each channel in the interval of 23.25 ms from t4 to t5 in the form of increasing from t4 to t5, namely, signal 3e-i. However, since signal 3e-i cannot be obtained at time t7, at time t7, an incomplete decoded audio signal is obtained for the interval from t4 to t5, but a complete decoded audio signal cannot be obtained. Therefore, at the time point t7, for each channel, for the interval from t0 to t1, the stereo decoding unit 620 synthesizes the signal 3d-i obtained by processing the immediately preceding frame and the signal 3c-i obtained by processing the current frame, and for the interval from t1 to t4, directly uses the signal 3c-i obtained by processing the current frame, thereby obtaining and outputting the complete decoded sound signal, i.e., the signal 3f-i, for the interval of 20ms from t0 to t4. The decoding device 600 obtains the decoded sound signal for the interval starting from t0 of each channel at the time point t7, so the mono encoding / decoding method of non-patent document 2 is used as the mono encoding / decoding method. The algorithm delay of the stereo encoding / decoding of non-patent document 1 as the mono encoding / decoding method is 35.25ms, which is the time length from t0 to t7. That is, the algorithm delay of the stereo encoding / decoding in the embedded encoding / decoding becomes larger than the algorithm delay based on the mono encoding / decoding.

[0032] Fig.10 is a diagram schematically showing various signals of a decoding device 200 according to a first embodiment described later. The decoding device 200 according to the first embodiment is Figure 3The structure shown in FIG. 1 includes a monaural decoding unit 210, an additional decoding unit 230, and a stereo decoding unit 220 that perform the operations described in detail in the first embodiment. At time t7, the stereo decoding unit 220 uses the signal 4c, which is a completely obtained monaural decoded audio signal before t6, to process the current frame. Fig. 9As described above, at the time point t7, the mono decoded sound signal before t5, i.e., signal 3a, is completely obtained by the mono decoder 210. Therefore, in the decoding device 200, the additional decoding unit 230 decodes the additional code CA to obtain a mono decoded sound signal of the interval of 3.25 ms from t5 to t6, i.e., signal 4b (additional decoding process), and the stereo decoder 220 uses the signal 4c obtained by combining the mono decoded sound signal before t5 obtained by the mono decoder 210, i.e., signal 3a, and the mono decoded sound signal of the interval from t5 to t6 obtained by the additional decoding unit 230 to perform stereo decoding process (upmixing process) of the current frame. That is, the stereo decoding unit 220 adopts a signal 4d in the interval of 23.75 ms from t1 to t6, which is a window of 3.25 ms from t1 to t2, in which the signal 4c increases, is flat in the interval of 16.75 ms from t2 to t5, and is attenuated in the interval of 3.25 ms from t5 to t6, and obtains a decoded sound signal, i.e., signal 4e-i, in the interval of t1 to t6, in which the window of the same shape as signal 4d is applied, for each channel. As a process of the immediately preceding frame, the stereo decoding unit 220 has already obtained a signal 4f-i of a decoded sound signal of each channel in the interval of 23.75 ms before t2, in which the window of the shape of attenuation from t1 to t2 is applied, at the time point t3. In addition, as a process of the immediately succeeding frame, the stereo decoding unit 220 obtains a decoded sound signal, i.e., signal 4g-i, in the interval of 23.75 ms after t5, in which the window of the form of increasing from t5 to t6 is applied. However, since the signal 4g-i cannot be obtained at the time point t7, an incomplete decoded sound signal is obtained for the interval from t5 to t6 at the time point t7, but a complete decoded sound signal cannot be obtained. Therefore, at the time point t7, the stereo decoding unit 220 synthesizes the signal 4f-i obtained by processing the immediately preceding frame and the signal 4e-i obtained by processing the current frame for each channel for the interval from t1 to t2, and directly uses the signal 4e-i obtained by processing the current frame for the interval from t2 to t5, thereby obtaining and outputting a complete decoded sound signal, i.e., signal 4h-i, for the interval of 20ms from t1 to t5. The decoding device 200 obtains the decoded sound signal of the interval starting from t1 of each channel at the time point t7, so the algorithm delay of the stereo encoding / decoding in the embedded encoding / decoding of the first embodiment is 32ms in length from t1 to t7. That is, the algorithm delay of the stereo encoding / decoding based on the embedded encoding / decoding of the first embodiment is not greater than the algorithm delay of the monaural encoding / decoding.

[0033] Fig.11 This is for the encoding device 100 of the first embodiment described later, that is, to make each signal Fig.10 FIG. 1 is a diagram schematically showing the encoding device 200 corresponding to the decoding device of the first embodiment of the schematically shown decoding device, and each signal of the encoding device. The encoding device 100 of the first embodiment is Figure 1 The structure shown includes, in addition to: a stereo encoding unit 110 that performs the same processing as the stereo encoding unit 510 of the encoding device 500; and a mono encoding unit 120 that, in the same manner as the mono encoding unit 520 of the encoding device 500, applies a window to the signal of the interval from t1 to t6 in the mono sound signal before t7, i.e., signal 1a, to obtain a mono code CM, and also includes an additional encoding unit 130 that encodes the signal of the interval from t1 to t6 in the signal 1a that is a mono sound signal, i.e., signal 5c, to obtain an additional code CA.

[0034] Hereinafter, the interval from t5 to t6, which is the overlapping interval of the current frame and the immediately following frame, is referred to as "interval X". That is, on the encoding side, interval X is an interval in which the mono encoding unit 120 encodes the mono audio signal of the application window in both the processing of the current frame and the processing of the immediately following frame. More specifically, interval X is an interval of a predetermined length including the end of the audio signal encoded by the mono encoding unit 120 in the processing of the current frame, an interval in which the audio signal of the application window of the shape attenuated by the mono encoding unit 120 in the processing of the current frame is encoded, and an interval of a predetermined length including the beginning of the interval encoded by the mono encoding unit 120 in the processing of the immediately following frame is encoded, and an interval in which the audio signal of the application window of the shape increased by the mono encoding unit 120 in the processing of the immediately following frame is encoded. In addition, on the decoding side, interval X is an interval in which the mono decoding unit 210 decodes the mono code CM in both the processing of the current frame and the processing of the immediately following frame to obtain the decoded audio signal of the application window. In more detail, the interval X is an interval of a predetermined length including the end in the decoded sound signal obtained by the mono decoding unit 210 decoding the mono code CM by processing the current frame, an interval of an application window of an attenuation shape in the decoded sound signal obtained by the mono decoding unit 210 decoding the mono code CM by processing the current frame, an interval of a predetermined length including the start in the decoded sound signal obtained by the mono decoding unit 210 decoding the mono code CM in processing the immediately following frame, an interval of an application window of an increase shape in the decoded sound signal obtained by the mono decoding unit 210 decoding the mono code CM in processing the immediately following frame, and a decoded sound signal obtained by synthesizing the decoded sound signal obtained by decoding the mono code CM by processing the current frame and the decoded sound signal obtained by decoding the mono code CM in processing the immediately following frame.

[0035] Hereinafter, the interval excluding the interval X in the interval to be mono-encoded / decoded in the processing of the current frame, that is, the interval from t1 to t5 is referred to as "interval Y". That is, interval Y is an interval excluding the overlapping interval with the immediately following frame in the interval in which the monaural encoding unit 120 encodes the monaural audio signal in the processing of the current frame, on the encoding side, and is an interval excluding the overlapping interval with the immediately following frame in the interval in which the monaural decoding unit 210 decodes the monaural code CM in the processing of the current frame to obtain a decoded audio signal, on the decoding side. Interval Y is an interval in which a monaural audio signal is represented by the monaural code CM of the current frame and the monaural code CM of the immediately preceding frame, and an interval in which a monaural audio signal is represented by only the monaural code CM of the current frame, and is therefore an interval in which a completely monaural decoded audio signal is obtained until the processing of the current frame.

[0036] <First Embodiment>

[0037] The encoding device and decoding device according to the first embodiment will be described.

[0038] <<Encoding device 100>>

[0039] like Figure 1 As shown, the encoding device 100 of the first embodiment includes a stereo encoding unit 110, a monaural encoding unit 120, and an additional encoding unit 130. The encoding device 100 encodes an input two-channel stereo sound time domain sound signal (two-channel stereo sound input sound signal) in units of frames of a specified time length, such as 20 ms, and obtains and outputs a stereo code CS, a monaural code CM, and an additional code CA described later. The two-channel stereo sound input sound signal input to the encoding device 100 is, for example, a digital sound signal or acoustic signal obtained by collecting sounds such as voices, music, etc. through two microphones and performing AD conversion, and is composed of an input sound signal of a left channel as a first channel and an input sound signal of a right channel as a second channel. The codes output by the encoding device 100, namely, the stereo code CS, the monaural code CM, and the additional code CA, are input to the decoding device 200 described later. The encoding device 100 performs the encoding in units of each frame, that is, each time the two-channel stereo sound input sound signal of the specified time length described above is input, Figure 2 The illustrated processing of step S111, step S121 and step S131. In the above example, when a two-channel stereophonic input sound signal for 20 ms from t3 to t7 is input, the encoding device 100 performs the processing of step S111, step S121 and step S131 on the current frame.

[0040] [Stereo encoding unit 110]

[0041] The stereo encoding unit 110 obtains and outputs a stereo code CS and a downmix signal from a two-channel stereo input sound signal input to the encoding device 100 (step S111). The stereo code CS represents a characteristic parameter, which is a parameter representing a characteristic of the difference between the sound signals of the two channels input, and the downmix signal is a signal obtained by mixing the sound signals of the two channels.

[0042] [Example of the Stereo Encoding Unit 110]

[0043] As an example of the stereo encoding unit 110, the operation of each frame of the stereo encoding unit 110 is described in the case where information representing the intensity difference of each frequency band of the sound signals of the two input channels is used as a feature parameter. In addition, a specific example using complex DFT (Discrete Fourier Transformation) is described below, but a known frequency domain transformation method other than complex DFT can also be used. In addition, when a sample string whose number of samples is not a power of 2 is transformed into the frequency domain, it is sufficient to use a known technique such as a sample string that is zero-filled in a manner that the number of samples becomes a power of 2.

[0044] The stereo encoding unit 110 first performs complex DFT on the input sound signals of the two channels to obtain complex DFT coefficient sequences (step S111-1). The complex DFT coefficient sequence is obtained by applying a window that overlaps between frames and using a process that takes into account the symmetry of the complex numbers obtained by the complex DFT. For example, when the sampling frequency is 32kHz, each time the sound signals of the two channels of 640 samples each are input as samples of 20ms, the process is performed, and for each channel, the sample sequence of the continuous 744-point digital sound signal (the sample sequence of the interval t1 to t6 in the above example) including the 104-point sample overlapping with the last sample group of the immediately preceding frame (the sample in the interval t1 to t2 in the above example) and the 104-point sample overlapping with the first sample group of the immediately following frame (the sample in the interval t5 to t6 in the above example) is subjected to complex DFT and the first half of the 372 complex sequences of the 744 complex sequences obtained are used as the complex DFT sequence. Thereafter, f is an integer between 1 and 372, each complex DFT coefficient of the complex DFT coefficient sequence of the first channel is V1(f), and each complex DFT coefficient of the complex DFT coefficient sequence of the second channel is V2(f). Next, the stereo encoding unit 110 obtains a sequence of radius values ​​on a complex plane based on each complex DFT coefficient from the complex DFT coefficient sequences of the two channels (step S111-2). The radius value on the complex plane of each complex DFT coefficient of each channel is equivalent to the intensity of each frequency point of the sound signal of each channel. Thereafter, the radius value on the complex plane of the complex DFT coefficient V1(f) of the first channel is set to V1r(f), and the radius value on the complex plane of the complex DFT coefficient V2(f) of the second channel is set to V2r(f). Next, the stereo encoding unit 110 obtains the average value of the ratio of the value of the radius of one channel to the value of the radius of the other channel for each frequency band, and obtains a sequence based on the average value as a feature parameter (step S111-3). The sequence of average values ​​is a feature parameter corresponding to information representing the intensity difference of each frequency band of the sound signals of the two channels input. For example, if there are four frequency bands, for each of the four frequency bands f from 1 to 93, 94 to 186, 187 to 279, and 280 to 372, the average values ​​Mr(1), Mr(2), Mr(3), and Mr(4) of 93 values ​​obtained by dividing the value V1r(f) of the radius of the first channel by the value V2r(f) of the radius of the second channel are obtained, and a sequence {Mr(1), Mr(2), Mr(3), and Mr(4)} based on the average value is obtained as a feature parameter.

[0045] In addition, the number of frequency bands can be any value less than the number of frequency points. The same value as the number of frequency points can be used as the number of frequency bands, or 1 can be used. When the same value as the number of frequency points is used as the number of frequency bands, the stereo encoding unit 110 obtains the value of the ratio of the value of the radius of one channel to the value of the radius of the other channel at each frequency point, and obtains a sequence of values ​​based on the obtained ratio as a feature parameter. When 1 is used as the number of frequency bands, the stereo encoding unit 110 only needs to obtain the value of the ratio of the value of the radius of one channel to the value of the radius of the other channel at each frequency point, and obtain the average value of the obtained ratio for all frequency bands as a feature parameter. In addition, when the number of frequency bands is set to multiple, the number of frequencies contained in each frequency band can be arbitrary. For example, the number of frequency points contained in a low-frequency band can be set to be less than the number of frequency points contained in a high-frequency band.

[0046] In addition, the stereo encoding unit 110 may replace the ratio of the radius value of one channel to the radius value of the other channel with the difference between the radius value of one channel and the radius value of the other channel. That is, in the above example, the value obtained by subtracting the radius value V2r(f) of the second channel from the radius value V1r(f) of the first channel may be used instead of the value obtained by dividing the radius value V1r(f) of the first channel by the radius value V2r(f) of the second channel.

[0047] The stereo encoding unit 110 further obtains a code representing the characteristic parameters, namely, a stereo code CS (step S111-4). The code representing the characteristic parameters, namely, the stereo code CS, can be obtained by a known method. For example, the stereo encoding unit 110 performs vector quantization on the sequence of values ​​obtained in step S111-3 to obtain a code, and outputs the obtained code as a stereo code CS. Or, for example, the stereo encoding unit 110 performs scalar quantization on the values ​​included in the sequence of values ​​obtained in step S111-3 to obtain codes, and outputs the code obtained by combining the obtained codes as a stereo code CS. In addition, when a single value is obtained in step S111-3, the stereo encoding unit 110 can output the code obtained by scalar quantizing the single value as a stereo code CS.

[0048] The stereo encoding unit 110 also obtains a signal obtained by mixing the sound signals of the first channel and the second channel, that is, a down-mix signal (step S111-5). For example, in the processing of the current frame, for 20ms from t3 to t7, the stereo encoding unit 110 obtains a monaural signal obtained by mixing the sound signals of the two channels, that is, a down-mix signal. The stereo encoding unit 110 can mix the sound signals of the two channels in the time domain as in step S111-5A described later, or can mix the sound signals of the two channels in the frequency domain as in step S111-5B described later. In the case of mixing in the time domain, for example, the stereo encoding unit 110 obtains a sequence obtained by averaging the corresponding samples of the sample string of the sound signal of the first channel and the sample string of the sound signal of the second channel, which is a monaural signal obtained by mixing the sound signals of the two channels, that is, a down-mix signal (step S111-5A). When mixing is performed in the frequency domain, for example, the stereo encoding unit 110 obtains the complex DFT coefficients of the complex DFT coefficient column obtained by performing complex DFT on the sample string of the sound signal of the first channel, and the radius average VMr(f) and the angle average VMθ(f) of the complex DFT coefficients of the complex DFT coefficient column obtained by performing complex DFT on the sample string of the sound signal of the second channel, and obtains a sample string by performing inverse complex DFT on a sequence of complex VM(f) with a radius VMr(f) and an angle VMθ(f) on the complex plane, to obtain a down-mix signal which is a signal obtained by mixing the sound signals of the two channels into a mono channel (steps S111B to 5B).

[0049] In addition, if Figure 1 As shown by the dot-dash line, the encoding device 100 may include a downmixer 150 so that the processing of step S111-5 for obtaining the downmix signal is performed not in the stereo encoding unit 110 but in the downmixer 150. In this case, the stereo encoding unit 110 obtains and outputs a stereo code CS (step S111), the stereo code CS indicating a feature parameter that is a feature of a difference between two channel sound signals input from a two-channel stereo input sound signal input to the encoding device 100 (step S111), and the downmixer 150 obtains and outputs a signal obtained by mixing the two channel sound signals from the two-channel stereo input sound signal input to the encoding device 100, that is, a downmix signal (step S151). That is, the stereo encoding unit 110 may perform the above-mentioned steps S111-1 to S111-4 as step S111, and the downmixer 150 may perform the above-mentioned step S111-5 as step S151.

[0050] [Monaural encoding unit 120]

[0051] The downmix signal outputted from the stereo encoding unit 110 is inputted to the mono encoding unit 120. When the encoding device 100 includes the downmix unit 150, the downmix signal outputted from the downmix unit 150 is inputted to the mono encoding unit 120. The mono encoding unit 120 encodes the downmix signal using a predetermined encoding method to obtain and output a mono code CM (step S121). As the encoding method, for example, an encoding method including a process of applying a window having overlap between frames, such as the 13.2 kbps mode of the 3GPP EVS standard (3GPP TS26.445) of Non-Patent Document 2, is used. In the above example, the mono encoding unit 120, in the processing of the current frame, uses the shape of the increase of the signal 1a as the downmix signal from the interval from t1 to t2 where the current frame and the immediately previous frame overlap, the shape of the attenuation of the interval from t5 to t6 where the current frame and the immediately subsequent frame overlap, and the signal 1b of the interval from t1 to t6 obtained by applying a flat-shaped window to the interval from t2 to t5 between these intervals. The interval from t6 to t7 of the signal 1a as the "pre-read interval" is also used for analysis and processing, and is encoded to obtain and output the mono code CM.

[0052] In this way, when the encoding method used by the mono encoding unit 120 includes processing with overlapping windows and analysis processing using a "pre-read interval", not only the downmix signal output by the stereo encoding unit 110 or the downmix unit 150 in the processing of the current frame is used for the encoding process, but also the downmix signal output by the stereo encoding unit 110 or the downmix unit 150 in the processing of the past frame is used for the encoding process. Therefore, the mono encoding unit 120 is provided with a storage unit (not shown), and the downmix signal input in the processing of the past frame is pre-stored in the storage unit. The mono encoding unit 120 can use the downmix signal stored in the storage unit to perform the encoding process of the current frame. Alternatively, the stereo encoding unit 110 or the down-mixing unit 150 may include a storage unit (not shown) and the stereo encoding unit 110 or the down-mixing unit 150 may output the down-mix signal used in the encoding process of the current frame by the mono encoding unit 120 including the down-mix signal obtained in the process of the previous frame in the process of the current frame, and the mono encoding unit 120 may use the down-mix signal input from the stereo encoding unit 110 or the down-mixing unit 150 in the process of the current frame. In addition, as required, each unit described later may also perform the above-mentioned process, and the signal obtained in the process of the previous frame may be stored in the storage unit (not shown) and used in the process of the current frame, but this is a well-known process in the technical field of encoding, and therefore, the description thereof will be omitted in order to avoid redundancy.

[0053] [Additional encoding unit 130]

[0054] The downmix signal output by the stereo encoding unit 110 is input to the additional encoding unit 130. When the encoding device 100 includes the downmix unit 150, the downmix signal output by the downmix unit 150 is input to the additional encoding unit 130. The additional encoding unit 130 encodes the downmix signal of the interval X in the input downmix signal, obtains the additional code CA and outputs it (step S131). In the above example, the additional encoding unit 130 encodes the signal 5c which is the downmix signal of the interval from t5 to t6, obtains the additional code CA and outputs it. The encoding can use a known encoding method such as scalar quantization or vector quantization.

[0055] <<Decoding device 200>>

[0056] like Figure 3 As shown, the decoding device 200 of the first embodiment includes a mono decoding unit 210, an additional decoding unit 230, and a stereo decoding unit 220. The decoding device 200 decodes the input mono code CM, additional code CA, and stereo code CS in frame units of the same prescribed time length as the encoding device 100 to obtain and output a binaural stereo time domain sound signal (binaural stereo decoded sound signal). The codes input to the decoding device 200, i.e., the mono code CM, additional code CA, and stereo code CS are output by the encoding device 100. The decoding device 200 performs decoding each time the mono code CM, additional code CA, and stereo code CS are input in frame units, i.e., at intervals of the prescribed time length described above. Figure 4 In the above example, if the mono code CM, additional code CA, and stereo code CS of the current frame are input at time t7, which is 20 ms after t3 when the decoding device 200 performed the processing for the immediately previous frame, then the decoding device 200 performs the processing of steps S211, S221, and S231 for the current frame. Figure 3 As shown by the middle dotted line, the decoding device 200 also outputs a monaural decoded audio signal which is a monaural time-domain audio signal when necessary.

[0057] [Monaural decoding unit 210]

[0058] The monaural code CM included in the code input to the decoding device 200 is input to the monaural decoding unit 210. The monaural decoding unit 210 obtains and outputs a monaural decoded audio signal of the interval Y using the input monaural code CM (step S211). As the predetermined decoding method, a decoding method corresponding to the encoding method used by the monaural encoding unit 120 of the encoding device 100 is used. In the above example, the mono decoding unit 210 decodes the mono code CM of the current frame in a prescribed decoding method, and obtains a signal 2a of the interval t1 to t6 of 23.25 ms with an application window of an increasing shape of 3.25 ms from t1 to t2, a flat shape of 16.75 ms from t2 to t5, and a signal 2b of the interval t1 to t6 of 23.25 ms from t6 with an attenuated shape. For the interval from t1 to t2, the signal 2b obtained from the mono code CM of the immediately preceding frame and the signal 2a obtained from the mono code CM of the current frame are synthesized in the processing of the immediately preceding frame. For the interval from t2 to t5, the signal 2a obtained from the current mono code CM is directly used, thereby obtaining a mono decoded signal of the interval 20 ms from t1 to t5, i.e., signal 2d. In addition, since the signal 2a of the interval from t5 to t6 obtained from the mono code CM of the current frame is used as the "signal 2b obtained in the immediately previous processing" in the processing of the next frame, the mono decoding unit 210 stores the signal 2a of the interval from t5 to t6 obtained from the mono code CM of the current frame in a storage unit not shown in the figure within the mono decoding unit 210.

[0059] [Additional decoding unit 230]

[0060] The additional decoding unit 230 receives the additional code CA included in the code input to the decoding device 200. The additional decoding unit 230 decodes the additional code CA to obtain a decoded audio signal of a monaural section X, i.e., an additional decoded signal, and outputs it (step S231). In decoding, a decoding method corresponding to the encoding method used by the additional encoding unit 130 is used. In the above example, the additional decoding unit 230 decodes the additional code CA of the current frame to obtain a decoded audio signal of a monaural section of 3.25 ms from t5 to t6, i.e., signal 4b, and outputs it.

[0061] [Stereo decoding unit 220]

[0062] The stereo decoding unit 220 is input with the mono decoded audio signal output by the monaural decoding unit 210, the additional decoded signal output by the additional decoding unit 230, and the stereo code CS included in the code input to the decoding device 200. The stereo decoding unit 220 obtains a decoded audio signal of two channels, i.e., a stereo decoded audio signal, based on the input mono decoded audio signal, the additional decoded signal, and the stereo code CS, and outputs it (step S221). More specifically, the stereo decoding unit 220 obtains a decoded down-mixed signal of the interval Y+X (i.e., the interval where the interval Y is connected to the interval X) (step S221-1), and the interval Y+X is obtained by connecting the mono decoded audio signal of the interval Y and the additional decoded signal of the interval X, and obtains a decoded audio signal of two channels from the decoded down-mixed signal obtained in step S221-1 through an up-mixing process using a characteristic parameter obtained from the stereo code CS, and outputs it (step S221-2). Although the same is true in each of the embodiments described below, the upmixing process refers to the following process: the downmix signal is regarded as a signal obtained by mixing the decoded sound signals of the two channels, and the feature parameters obtained from the stereo code CS are regarded as information indicating the characteristics of the difference between the decoded sound signals of the two channels, so as to obtain the decoded sound signals of the two channels. In the above example, first, the stereo decoding unit 220 connects the mono decoded sound signal of the 20ms interval from t1 to t5 output by the mono decoding unit 210 (the interval from t1 to t5 of the signal 2d and the signal 3a) and the additional decoded signal of the 3.25ms interval from t5 to t6 output by the additional decoding unit 230 (the signal 4b), and obtains the decoded downmix signal of the 23.25ms interval from t1 to t6 (the interval from t1 to t6 of the signal 4c). Next, the stereo decoding unit 220 regards the decoded downmix signal in the interval from t1 to t6 as a signal obtained by mixing the decoded sound signals of the two channels, regards the characteristic parameters obtained from the stereo code CS as information representing the differential characteristics of the decoded sound signals of the two channels, obtains the decoded sound signals of the two channels in the interval of 20ms from t1 to t5 (signal 4h-1 and signal 4h-2) and outputs them.

[0063] [Example of Step S221-2 Performed by the Stereo Decoding Unit 220]

[0064] As an example of step S221-2 performed by the stereo decoding unit 220, step S221-2 performed by the stereo decoding unit 220 in the case where the characteristic parameter is information indicating the intensity difference of each frequency band of the sound signals of the two channels will be described. The stereo decoding unit 220 first decodes the input stereo code CS to obtain information indicating the intensity difference of each frequency band (S221-21). The stereo decoding unit 220 obtains the characteristic parameter from the stereo code CS in a manner corresponding to the manner in which the stereo encoding unit 110 of the encoding device 100 obtains the stereo code CS from the information indicating the intensity difference of each frequency band. For example, the stereo decoding unit 220 performs vector decoding on the input stereo code CS, and obtains each element value of the vector corresponding to the input stereo code CS as information indicating the intensity difference of each frequency band of multiple frequency bands. Alternatively, for example, the stereo decoding unit 220 performs scalar decoding on each code included in the input stereo code CS to obtain information indicating the intensity difference of each frequency band. Furthermore, when the number of frequency bands is 1, the stereo decoding unit 220 performs scalar decoding on the input stereo code CS to obtain information indicating the intensity differences of one frequency band, that is, all frequency bands.

[0065] Next, the stereo decoding unit 220 regards the decoded downmix signal obtained in step S221-1 and the feature parameter obtained in step S221-21 as a signal obtained by mixing the decoded audio signals of the two channels, and regards the feature parameter as information indicating the intensity difference of each frequency band of the decoded audio signals of the two channels, and obtains and outputs the decoded audio signals of the two channels (step S220-22). If the stereo encoding unit 110 of the encoding device 100 performs the above-mentioned specific example operation using the complex DFT, step S221-22 of the stereo decoding unit 220 becomes the following operation.

[0066] The stereo decoding unit 220 first obtains the following signal 4d (step S221-221), which is a signal with a window applied to it, with respect to the decoded downmix signal of 744 samples in the interval of 23.25ms from t1 to t6, the shape of increasing in the interval of 3.25ms from t1 to t2, being flat in the interval of 16.75ms from t2 to t5, and being attenuated in the interval of 3.25ms from t5 to t6. Next, the stereo decoding unit 220 obtains a sequence of 372 complex numbers in the first half of the sequence of 744 complex numbers obtained by complex DFT of the signal 4d as a complex DFT coefficient sequence (mono complex DFT coefficient sequence) (steps S221-222). Thereafter, each complex DFT coefficient of the mono complex DFT coefficient sequence obtained by the stereo decoding unit 220 is set to MQ(f). Next, the stereo decoding unit 220 obtains the value of the radius MQr(f) on the complex plane of each complex DFT coefficient and the value of the angle MQθ(f) on the complex plane of each complex DFT coefficient based on the complex DFT coefficient sequence of the monophonic channel (steps S221-223). Next, the stereo decoding unit 220 obtains the value of each radius MQr(f) multiplied by the square root of the corresponding value in the characteristic parameter as the value of each radius VLQr(f) of the first channel, and obtains the value of each radius MQr(f) divided by the square root of the corresponding value in the characteristic parameter as the value of each radius VRQr(f) of the second channel (steps S221-224). Regarding the corresponding values ​​in the characteristic parameters of each frequency point, if the example of the four frequency bands mentioned above is used, f from 1 to 93 is Mr(1), f from 94 to 186 is Mr(2), f from 187 to 279 is Mr(3), and f from 280 to 372 is Mr(4). In addition, when the stereo encoding unit 110 of the encoding device 100 uses the difference between the radius value of the first channel and the radius value of the second channel instead of the ratio between the radius value of the first channel and the radius value of the second channel, the stereo decoding unit 220 can add the value obtained by dividing the corresponding value in the characteristic parameter by 2 to the value of each radius MQr(f) to obtain the value VLQr(f) of each radius of the first channel, and subtract the value obtained by dividing the corresponding value in the characteristic parameter by 2 from the value of each radius MQr(f) to obtain the value VRQr(f) of each radius of the second channel.Next, the stereo decoding unit 220 obtains a decoded sound signal (signal 4e-1) of an application window of the first channel with 744 samples in the interval of 23.25 ms from t1 to t6 by performing inverse complex DFT on a sequence formed by complex numbers with a radius of VLQr(f) and an angle of MQθ(f) on the complex plane, and obtains a sound signal (signal 4e-2) of an application window of the second channel with 744 samples in the interval of 23.25 ms from t1 to t6 by performing inverse complex DFT on a sequence formed by complex numbers with a radius of VRQr(f) and an angle of MQθ(f) on the complex plane (steps S221-225). The decoded sound signals (signal 4e-1 and signal 4e-2) that capture the windows of each channel obtained in steps S221-225 are windowed signals with an increasing shape in the 3.25ms interval from t1 to t2, a flat shape in the 16.75ms interval from t2 to t5, and an attenuated shape in the 3.25ms interval from t5 to t6. Next, for the first channel and the second channel, the stereo decoding unit 220 synthesizes the signals (signal 4f-1, signal 4f-2) obtained in the immediately preceding steps S221-225 and the signals (signal 4e-1, signal 4e-2) obtained in steps S221-225 of the current frame for the interval from t1 to t2, and directly uses the signals (signal 4e-1, signal 4e-2) obtained in steps S221-225 of the current frame for the interval from t2 to t5, thereby obtaining the decoded sound signal (signal 4h-1, signal 4h-2) for the 20ms interval from t1 to t5 and outputs it (steps S221-226).

[0067] <Second Embodiment>

[0068] The difference between the downmix signal and the monaural encoded local decoded signal in the interval Y, which is a time interval in which a complete monaural decoded audio signal is obtained from the monaural code CM in the monaural decoding unit 210, may also be an object of encoding in the additional encoding unit 130. This aspect is taken as a second embodiment, and the differences from the first embodiment will be described.

[0069] [Monaural encoding unit 120]

[0070] In addition to encoding the downmix signal using a predetermined encoding method to obtain and output a mono code CM, the monaural encoding unit 120 also obtains and outputs a signal obtained by decoding the monaural code CM, i.e., a monaural local decoded signal which is a local decoded signal of interval Y of the downmix signal (step S122). In the above example, the monaural coding unit 120 obtains not only the monaural code CM of the current frame but also the local decoded signal corresponding to the monaural code CM of the current frame, that is, the local decoded signal of the applied window having an increasing shape in the interval of 3.25 ms from t1 to t2, a flat shape in the interval of 16.75 ms from t2 to t5, and a decaying shape in the interval of 3.25 ms from t5 to t6. For the interval from t1 to t2, the local decoded signal corresponding to the monaural code CM immediately before and the local decoded signal corresponding to the monaural code CM of the current frame are synthesized, and for the interval from t2 to t5, the local decoded signal corresponding to the monaural code CM of the current frame is used as it is, thereby obtaining and outputting the local decoded signal from t1 to t5. The local decoded signal corresponding to the monaural code CM of the frame immediately before the interval from t1 to t2 uses the signal stored in the storage unit (not shown) in the monaural coding unit 120. The signal in the interval from t5 to t6 among the local decoded signals corresponding to the mono code CM of the current frame is used as the “local decoded signal corresponding to the mono code CM of the immediately preceding frame” in the processing of the immediately subsequent frame, so the mono encoding unit 120 stores the local decoded signal in the interval from t5 to t6 obtained from the mono code CM of the current frame in a storage unit not shown in the figure within the mono encoding unit 120.

[0071] [Additional encoding unit 130]

[0072] In addition to the downmixed signal, such as Figure 1As shown by the dotted line in the middle, the monaural local decoded signal output by the monaural encoding unit 120 is also input to the additional encoding unit 130. The additional encoding unit 130 not only encodes the downmix signal of the interval X that is the encoding object of the additional encoding unit 130 of the first embodiment, but also encodes the difference signal (a signal formed by subtracting the sample values ​​between the corresponding samples) between the downmix signal of the interval Y and the monaural local decoded signal, and obtains and outputs the additional code CA (step S132). For example, the additional encoding unit 130 encodes the downmix signal of the interval X and the difference signal of the interval Y respectively to obtain a code, and obtains the code obtained by connecting the obtained codes as the additional code CA. In the encoding, the same encoding method as the additional encoding unit 130 of the first embodiment can be used. In addition, for example, the additional encoding unit 130 can also encode the signal formed by connecting the difference signal of the interval Y and the downmix signal of the interval X to obtain the additional code CA. In addition, for example, as described in [[Specific Example 1 of the additional encoding unit 130]] below, the additional encoding unit 130 performs first additional encoding and second additional encoding, in which the downmix signal of interval X is encoded to obtain a code (first additional code CA1), and in the second additional encoding, a signal obtained by combining the difference signal of interval Y (i.e., the quantization error signal of the mono encoding unit 120), the downmix signal of interval X, and the difference signal of the first additionally encoded local decoded signal (i.e., the quantization error signal of the first additionally encoded) is encoded to obtain a code (second additional code CA2), and the first additional code CA1 and the second additional code CA2 are combined to obtain the code. According to [[Specific example 1 of the additional encoding unit 130]], a signal obtained by connecting the difference signal of interval Y and the down-mixed signal of interval X, in which the difference in amplitude between the two intervals is smaller than that of the down-mixed signal, is used as the object of encoding for the second additional encoding, and the down-mixed signal itself is used as the object of encoding for the first additional encoding, so that highly efficient encoding can be expected.

[0073] [[Specific example 1 of the additional encoding unit 130]]

[0074] The additional coding unit 130 first encodes the input downlink mixed signal of the interval X to obtain the first additional code CA1 (step S132-1, hereinafter also referred to as the "first additional code"), and obtains the local decoded signal of the interval X corresponding to the first additional code CA1, that is, the local decoded signal of the first additional code of the interval X (step S132-2). In the first additional code, a well-known encoding method such as scalar quantization or vector quantization can be used. Next, the additional coding unit 130 obtains the difference signal (a signal formed by subtracting the sampling values ​​between the corresponding samples) between the input downlink mixed signal of the interval X and the local decoded signal of the interval X obtained in step S132-2 (step S132-3). The additional coding unit 130 also obtains the difference signal (a signal formed by subtracting the sampling values ​​between the corresponding samples) between the downmix signal of the interval Y and the mono local decoded signal (step S132-4). Next, the additional coding unit 130 encodes the signal obtained by connecting the difference signal of the interval Y obtained in step S132-4 and the difference signal of the interval X obtained in step S132-3, thereby obtaining a second additional code CA2 (in step S132-5, hereinafter also referred to as "second additional coding"). In the second additional coding, a coding method is used to encode the sample string of the difference signal of the interval Y obtained in step S132-4 and the sample string of the difference signal of the interval X obtained in step S132-3 together, for example, a coding method using prediction in the time domain or a coding method that adapts to the deviation of the amplitude in the frequency domain is used. Next, the additional coding unit 130 outputs the code obtained by combining the first additional code CA1 obtained in step S132-1 and the second additional code CA2 obtained in step S132-5 as an additional code CA (step S132-6).

[0075] In addition, the additional coding unit 130 may replace the above-mentioned difference signal with a weighted difference signal as the object of coding. That is, the additional coding unit 130 may encode the weighted difference signal (a signal formed by weighted subtraction of sampling values ​​between corresponding samples) between the downmix signal of interval Y and the monaural local decoded signal and the downmix signal of interval X to obtain the additional code CA and output it. In the case of [[Specific example 1 of the additional coding unit 130]], as the processing of step S132-4, the additional coding unit 130 may obtain the weighted difference signal (a signal formed by weighted subtraction of sampling values ​​between corresponding samples) between the downmix signal of interval Y and the monaural local decoded signal. Similarly, as the processing of step S132-3 of [[Specific example 1 of the additional coding unit 130]], the additional coding unit 130 may obtain the weighted difference signal (a signal formed by weighted subtraction of sampling values ​​between corresponding samples) between the input downmix signal of interval X and the local decoded signal of interval X obtained in step S132-2. In these cases, the weights used to generate each weighted difference signal are encoded by a known encoding technique to obtain a code, and the obtained code (code indicating the weight) is included in the additional code CA. These cases are also the same for each difference signal in each embodiment described later, but it is known in the technical field of encoding that the weighted difference signal is used as the object of encoding instead of the difference signal, and the weight is also encoded at this time, so in the embodiment described later, in order to avoid redundancy, a separate description is omitted, and only the description of using "or" to record the difference signal and the weighted difference signal together, and the description of using "or" to record the subtraction operation and the weighted subtraction operation together are given.

[0076] [Additional decoding unit 230]

[0077] The additional decoding unit 230 decodes the additional code CA, and obtains and outputs not only the additional decoding signal obtained by the additional decoding unit 230 of the first embodiment, that is, the additional decoding signal of the interval X, but also obtains and outputs the additional decoding signal of the interval Y (step S232). The decoding method corresponding to the encoding method used by the additional coding unit 130 in step S132 is used. That is, when the additional coding unit 130 uses [[Specific example 1 of the additional coding unit 130]] in step S132, the additional decoding unit 230 performs the processing of [[Specific example 1 of the additional decoding unit 230]] described below.

[0078] [[Specific example 1 of the additional decoding unit 230]]

[0079] The additional decoding unit 230 first decodes the first additional code CA1 included in the additional code CA to obtain a first decoded signal of interval X (step S232-1, hereinafter also referred to as "first additional decoding"). In the first additional decoding, the additional coding unit 130 uses a decoding method corresponding to the coding method used in the first additional coding. In addition, the additional decoding unit 230 decodes the second additional code CA2 included in the additional code CA to obtain second decoded signals of interval Y and interval X (step S232-2, hereinafter also referred to as "second additional decoding"). In the second additional decoding, a decoding method corresponding to the coding method used by the additional coding unit 130 in the second additional coding is used, that is, a decoding method that can obtain a combined sampling string formed by connecting a sampling string of the additional decoded signal from the code to interval Y and a sampling string of the second decoded signal of interval X, for example, a decoding method that applies prediction in the time domain, a decoding method that applies amplitude deviation in the frequency domain. Next, the additional decoding unit 230 obtains the second decoded signal of interval Y in the second decoded signal obtained in step S232-2 as the additional decoded signal of interval Y, obtains the signal obtained by adding the first decoded signal of interval X obtained in step S232-1 and the second decoded signal of interval X in the second decoded signal obtained in step S232-2 (a signal constructed by adding the sampling values ​​between corresponding samples) as the additional decoded signal of interval X, and outputs the additional decoded signals of interval Y and interval X (step S232-4).

[0080] In addition, when the additional coding unit 130 sets the weighted difference signal as the encoding object instead of the difference signal, the additional code CA also includes the code representing the weighting. Therefore, the additional decoding unit 230 decodes the code other than the code representing the weighting in the additional code CA in the above-mentioned step S232 to obtain the additional decoded signal and outputs it, and decodes the code representing the weighting contained in the additional code CA to obtain the weight and output it. In the case of [[Specific example 1 of the additional decoding unit 230]], as long as the weight of the interval X included in the additional code CA is decoded to obtain the weight of the interval X, the additional decoding unit 230 obtains the second decoded signal of the interval Y in the second decoded signal obtained in step S232-2 as the additional decoded signal of the interval Y in step S232-4, and obtains the signal obtained by weighted addition operation of the first decoded signal of the interval X obtained in step S232-1 and the second decoded signal of the interval X in the second decoded signal obtained in step S232-2 (a signal composed of weighted addition operation of sampling values ​​between corresponding samples) as the additional decoded signal of the interval X, and the weighted output of the interval Y obtained by decoding the interval Y, the additional decoded signal of the interval X, and the code representing the weight of the interval Y included in the additional code CA. These are also the same for the addition operations of the signals in each of the embodiments described later. However, in the technical field of coding, it is well known that weighted addition operations (generation of weighted sum signals) are performed instead of addition (generation of sum signals), and weights are obtained according to codes in this case. Therefore, in the embodiments described later, separate descriptions are omitted without becoming redundant, and only the description of using "or" to record the addition and weighted addition operations together, and the description of using "or" to record the sum signal and the weighted sum signal together are performed.

[0081] [Stereo decoding unit 220]

[0082] The stereo decoding unit 220 performs the following steps S22-1 and S222-2 (step S222). The stereo decoding unit 220 replaces step S221-1 performed by the stereo decoding unit 220 of the first embodiment, obtains a decoded down-mix signal of interval Y+X by combining a sum signal (a signal formed by adding sample values ​​between corresponding samples) of the monaural decoded sound signal of interval Y and the additional decoded signal of interval Y with the additional decoded signal of interval X (step S22-1), uses the decoded down-mix signal obtained in step S222-1 instead of the decoded down-mix signal obtained in step S221-1, obtains two-channel decoded sound signals from the decoded down-mix signal obtained in step S222-1 by up-mixing processing using characteristic parameters obtained from the stereo code CS, and outputs the decoded sound signals (step S22-2).

[0083] In addition, when the additional coding unit 130 encodes a weighted difference signal instead of a difference signal, the stereo decoding unit 220 may obtain a signal obtained by combining a weighted sum signal (a signal formed by weighted addition of sampling values ​​between corresponding samples) of the monaural decoded audio signal of the interval Y and the additional decoded signal of the interval Y and the additional decoded signal of the interval X as a decoded down-mix signal of the interval Y+X in step S222-1. In the generation of the weighted sum signal (weighted addition of sampling values ​​between corresponding samples) of the monaural decoded audio signal of the interval Y and the additional decoded signal of the interval Y, the weight of the interval Y output by the additional decoding unit 230 may be used. The same is true for the addition operation of the signal in each of the embodiments described later. As described in the description location of the additional decoding unit 230, in the technical field of coding, a weighted addition operation (generation of a sum signal) is performed instead of an addition operation (generation of a sum signal). It is well known that the weight is obtained from the code at this time. Therefore, in the embodiments described later, in order to avoid redundancy, a separate description is omitted, and only the description of using "or" to record the addition operation and the weighted addition operation together, and the description of using "or" to record the sum signal and the weighted sum signal together are performed.

[0084] According to the second embodiment, in addition to the fact that the algorithm delay of stereo encoding / decoding is not greater than the algorithm delay of mono encoding / decoding, the decoded downmix signal used in stereo decoding can be made higher in quality than in the first embodiment, so the decoded sound signal of each channel obtained by stereo decoding can also be made higher in sound quality. That is, in the second embodiment, the mono encoding process performed by the mono encoding unit 120 and the additional encoding process performed by the additional encoding unit 130 are used as encoding processes for encoding the downmix signal with high quality, the mono code CM and the additional code CA are obtained as codes for well representing the downmix signal, and the mono decoding process performed by the mono decoding unit 210 and the additional decoding process performed by the additional decoding unit 230 are used as decoding processes for obtaining a high-quality decoded downmix signal. The amount of code allocated to each of the mono code CM and the additional code CA can be arbitrarily determined according to the purpose. In addition to the standard quality mono encoding / decoding, in the case of wishing to achieve higher quality stereo encoding / decoding, more code can be allocated to the additional code CA. That is, from the perspective of stereo encoding / decoding, "mono code" and "additional code" are just convenient names. If the mono code CM and the additional code CA are respectively parts of the code representing the down-mix signal, one of them can be called "first down-mix code" and the other can be called "second down-mix code". If it is assumed that more code amount is allocated to the additional code CA, the additional code CA can be called "down-mix code", "down-mix signal code", etc. The above situation is also the same in the third embodiment or each embodiment based on the second embodiment described thereafter.

[0085] <Third Embodiment>

[0086] In the stereo decoding unit 220, a decoded downmix signal corresponding to a downmix signal obtained by mixing the sound signals of the two channels in the frequency domain can be used to obtain a decoded sound signal of the two channels with higher sound quality. In the monaural encoding unit 120, a signal obtained by mixing the sound signals of the two channels in the time domain is encoded, and in the monaural decoding unit 210, a high-quality decoded sound signal of the monaural decoding unit can be obtained. In this case, the stereo encoding unit 110 can mix the sound signals of the two channels input to the encoding device 100 in the frequency domain to obtain a downmix signal, the monaural encoding unit 120 can encode the signal obtained by mixing the sound signals of the two channels input to the encoding device 100 in the time domain, and the additional encoding unit 130 can also encode the difference between the signal obtained by mixing the sound signals of the two channels in the frequency domain and the signal obtained by mixing the sound signals in the time domain. This form is described as a third embodiment, and the difference from the second embodiment is mainly described.

[0087] [Stereo encoding unit 110]

[0088] The stereo encoding unit 110 performs the operation described in the first embodiment in the same manner as the stereo encoding unit 110 of the second embodiment, but the processing of obtaining a signal obtained by mixing the sound signals of the two channels, that is, a down-mix signal, is performed by, for example, mixing the sound signals of the two channels in the frequency domain as in step S111-5B (step S113). That is, the stereo encoding unit 110 obtains a down-mix signal obtained by mixing the sound signals of the two channels in the frequency domain. For example, in the processing of the current frame, the stereo encoding unit 110 may obtain a monaural signal obtained by mixing the sound signals of the two channels in the frequency domain, that is, a down-mix signal, for the interval from t1 to t6. When the encoding device 100 is further provided with a down-mixing unit 150, the stereo encoding unit 110 obtains and outputs a stereo code CS (step S113), wherein the stereo code CS represents a feature parameter, which is a feature of a difference between sound signals of two channels input according to a two-channel stereo input sound signal input to the encoding device 100, and the down-mixing unit 150 obtains and outputs a down-mixed signal (step S153), wherein the down-mixed signal is a signal obtained by mixing sound signals of two channels in the frequency domain according to the two-channel stereo input sound signal input to the encoding device 100.

[0089] [Monaural encoding target signal generation unit 140]

[0090] As in Figure 1As shown by the dot-dash line, the encoding device 100 of the third embodiment further includes a mono encoding object signal generating unit 140. The two-channel stereo input sound signal input to the encoding device 100 is input to the mono encoding object signal generating unit 140. The mono encoding object signal generating unit 140 obtains a mono encoding object signal as a mono signal from the input two-channel stereo input sound signal by mixing the sound signals of two channels in the time domain (step S143). For example, the mono encoding object signal generating unit 140 obtains a signal obtained by mixing the sound signals of two channels, i.e., the mono encoding object signal, based on a sequence of average values ​​between corresponding samples of the sample string of the sound signal of the first channel and the sample string of the sound signal of the second channel. That is, the mono encoding object signal obtained by the mono encoding object signal generating unit 140 is a signal obtained by mixing the sound signals of two channels in the time domain. For example, the monaural encoding target signal generation unit 140 may obtain a monaural encoding target signal which is a monaural signal obtained by mixing two-channel audio signals in the time domain for 20 ms from t3 to t7 in the processing of the current frame.

[0091] [Monaural encoding unit 120]

[0092] The mono encoding unit 120 receives the mono encoding target signal output by the mono encoding target signal generating unit 140 instead of the downmix signal output by the stereo encoding unit 110 or the downmix unit 150. The mono encoding unit 120 encodes the mono encoding target signal to obtain a mono code CM and outputs it (step S123). For example, in the processing of the current frame, the mono encoding unit 120 encodes the mono encoding target signal to obtain a mono code CM by applying a window to the mono encoding target signal, such that the shape of the interval from t1 to t2 where the current frame overlaps with the immediately preceding frame increases, the shape of the interval from t5 to t6 where the current frame overlaps with the immediately following frame decreases, and the shape of the interval from t2 to t5 between these intervals is flat, and also encodes the interval from t6 to t7 of the mono encoding target signal as a "pre-read interval" for analysis processing, and obtains and outputs the mono code CM.

[0093] [Additional encoding unit 130]

[0094] The additional coding unit 130 encodes the difference signal or weighted difference signal (a signal formed by subtracting or weighted subtraction of sample values ​​between corresponding samples) between the downmix signal of interval Y and the monaural local decoded signal and the downmix signal of interval X, and obtains and outputs the additional code CA (step S133). However, the downmix signal of interval Y is a signal obtained by mixing the audio signals of two channels in the frequency domain, and the monaural local decoded signal of interval Y is a signal obtained by locally decoding the signal obtained by mixing the audio signals of two channels in the time domain.

[0095] In addition, the additional coding unit 130 of the third embodiment is similar to the additional coding unit 130 of the second embodiment, as described in [[Specific Example 1 of the additional coding unit 130]], and performs a first additional coding of the first additional code CA obtained by encoding the down-mixed signal of the interval X, and a second additional coding of the second additional code CA2 obtained by encoding the difference signal or the weighted difference signal of the connection interval Y, the difference signal of the down-mixed signal of the interval X, and the local decoded signal of the first additional code, and the code combining the first additional code CA1 and the second additional code CA2 is set as the additional code CA.

[0096] [Monaural decoding unit 210]

[0097] The monaural decoding unit 210 obtains and outputs a monaural decoded audio signal of the interval Y using the monaural code CM, similarly to the monaural decoding unit 210 of the second embodiment (step S213). However, the monaural decoded audio signal obtained by the monaural decoding unit 120 of the third embodiment is a decoded signal of a signal obtained by mixing audio signals of two channels in the time domain.

[0098] [Additional decoding unit 230]

[0099] The additional decoding unit 230 decodes the additional code CA similarly to the additional decoding unit 230 of the second embodiment, obtains and outputs additional decoded signals of interval Y and interval X (step S233). However, the additional decoded signal of interval Y includes the difference between the signal obtained by mixing the audio signals of two channels in the time domain and the mono decoded audio signal, and the difference between the signal obtained by mixing the audio signals of two channels in the frequency domain and the signal obtained by mixing the audio signals of two channels in the time domain.

[0100] [Stereo decoding unit 220]

[0101] The stereo decoding unit 220 performs the following steps S223-1 and S223-2 (step S223). The stereo decoding unit 220 obtains a signal obtained by combining the sum signal of the monaural decoded audio signal of the interval Y and the additional decoded signal of the interval Y or a weighted sum signal (a signal formed by adding or weighted adding the sample values ​​between corresponding samples) and the additional decoded signal of the interval X as a decoded down-mix signal of the interval Y+X (step S223-1), and obtains two-channel decoded audio signals from the decoded down-mix signal obtained in step S223-1 by up-mixing processing using the characteristic parameters obtained from the stereo code CS and outputs them (step S223-2). Among them, the sum signal of interval Y includes: a mono decoded sound signal obtained by mono encoding / decoding a signal obtained by mixing sound signals of two channels in the time domain; the difference between a signal obtained by mixing sound signals of two channels in the time domain and the mono decoded sound signal; and the difference between a signal obtained by mixing sound signals of two channels in the frequency domain and a signal obtained by mixing sound signals of two channels in the time domain.

[0102] <Fourth Embodiment>

[0103] Regarding the interval X, in the mono encoding unit 120 and the mono decoding unit 210, although it is impossible to obtain a correct local decoded signal and a decoded signal without the signal and code of the immediately following frame, an incomplete local decoded signal and a decoded signal can be obtained even if only the signal and code up to the current frame are used. Therefore, the first to third embodiments can also be changed to: for the interval X, instead of the downmix signal itself, the difference between the downmix signal and the monaural local decoded signal obtained from the signal up to the current frame is encoded by the additional encoding unit 130. This method is described as the fourth embodiment.

[0104] <<Fourth Embodiment A>>

[0105] First, a fourth embodiment A which is a fourth embodiment obtained by modifying the second embodiment will be described mainly focusing on points different from the second embodiment.

[0106] [Monaural encoding unit 120]

[0107] Similar to the monaural encoding unit 120 of the second embodiment, the downmix signal output by the stereo encoding unit 110 or the downmix unit 150 is input to the monaural encoding unit 120. The monaural encoding unit 120 obtains a monaural code CM obtained by encoding the downmix signal and a signal obtained by decoding the monaural code CM up to the current frame, that is, a local decoded signal of the downmix signal in the interval Y+X, that is, a monaural local decoded signal, and outputs it (step S124). More specifically, the monaural coding unit 120 obtains the monaural code CM of the current frame and the local decoded signal corresponding to the monaural code CM of the current frame, that is, the local decoded signal of the application window with an increasing shape in the interval of 3.25 ms from t1 to t2, a flat shape in the interval of 16.75 ms from t2 to t5, and an attenuated shape in the interval of 3.25 ms from t5 to t6. For the interval from t1 to t2, the local decoded signal corresponding to the monaural code CM immediately before and the local decoded signal corresponding to the monaural code CM of the current frame are synthesized, and for the interval from t2 to t6, the local decoded signal corresponding to the monaural code CM of the current frame is used directly, thereby obtaining and outputting the local decoded signal from t1 to t2. However, the local decoded signal of the interval from t5 to t6 is a local decoded signal that is a complete local decoded signal by being synthesized with the local decoded signal of the application window with an increasing shape obtained in the processing of the immediately subsequent frame, and is an incomplete local decoded signal of the application window with an attenuated shape.

[0108] [Additional encoding unit 130]

[0109] Similar to the additional coding unit 130 of the second embodiment, the downmix signal output by the stereo coding unit 110 or the downmix unit 150 and the monaural local decoded signal output by the monaural coding unit 120 are input to the additional coding unit 130. The additional coding unit 130 encodes the difference signal or the weighted difference signal (a signal formed by subtracting the sample values ​​between corresponding samples or performing weighted subtraction) between the downmix signal of the interval Y+X and the monaural local decoded signal, obtains an additional code CA, and outputs it (step S134).

[0110] [Monaural decoding unit 210]

[0111] As with the monaural decoding unit 210 of the second embodiment, the monaural code CM is input to the monaural decoding unit 210. The monaural decoding unit 210 obtains and outputs a monaural decoded audio signal of the interval Y+X using the monaural code CM (step S214). However, the decoded signal of the interval X, i.e., the interval from t5 to t6, is a decoded signal that is a complete decoded signal by being synthesized with the decoded signal of the application window of the increasing shape obtained in the processing of the immediately subsequent frame, and is an incomplete decoded signal of the application window of the attenuated shape.

[0112] [Additional decoding unit 230]

[0113] Similar to the additional decoding unit 230 of the second embodiment, the additional code CA is input to the additional decoding unit 230. The additional decoding unit 230 decodes the additional code CA, obtains an additional decoded signal of the interval Y+X, and outputs it (step S234).

[0114] [Stereo decoding unit 220]

[0115] Similar to the stereo decoding unit 220 of the second embodiment, the mono decoded audio signal output by the monaural decoding unit 210, the additional decoded signal output by the additional decoding unit 230, and the stereo code CS input to the decoding device 200 are input to the stereo decoding unit 220. The stereo decoding unit 220 obtains a sum signal or a weighted sum signal (a signal formed by adding or weighted adding sample values ​​between corresponding samples) of the mono decoded audio signal of the interval Y+X and the additional decoded signal as a decoded down-mix signal, and obtains two-channel decoded audio signals from the decoded down-mix signal through up-mix processing using characteristic parameters obtained from the stereo code CS and outputs the decoded audio signals (step S224).

[0116] <<Fourth Embodiment B>>

[0117] In addition, if the "downmix signal output by the stereo encoding unit 110 or the downmix unit 150" and the "downmix signal" of the mono encoding unit 120 in the description of the fourth embodiment A are replaced with the "mono encoding object signal output by the mono encoding object signal generating unit 140" and the "mono encoding object signal", respectively, the description becomes centered on the points different from the third embodiment, that is, the fourth embodiment B, which is a fourth embodiment in which the third embodiment is changed.

[0118] <<Fourth Embodiment C>>

[0119] In addition, if the monaural local decoded signal obtained by the monaural encoding unit 120 in the description of the fourth embodiment, the difference signal or the weighted difference signal encoded by the additional encoding unit 130, and the additional decoded signal obtained by the additional decoding unit 230 are respectively set to interval X, and the stereo decoding unit 220 obtains a signal obtained by connecting the monaural decoded sound signal of interval Y and the sum signal of the monaural decoded sound signal of interval X and the additional decoded signal or the weighted sum signal as a decoded mixed signal, it becomes a fourth embodiment in which the first embodiment is modified, that is, the fourth embodiment C.

[0120] <Fifth Embodiment>

[0121] The downmix signal of the interval X includes a portion that can be predicted from the monaural local decoded signal of the interval Y. Therefore, in each of the first to fourth embodiments, the additional encoder 130 may encode the difference between the downmix signal and the prediction signal from the monaural local decoded signal of the interval Y for the interval X. This method will be described as the fifth embodiment.

[0122] <<Fifth Embodiment A>>

[0123] First, a fifth embodiment obtained by changing each of the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B will be referred to as the fifth embodiment A, and points different from the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B will be described.

[0124] [Additional encoding unit 130]

[0125] The additional coding unit 130 performs the following steps S135A-1 and S135A-2 (step S135A). The additional coding unit 130 first obtains a prediction signal for section X of the monaural local decoded signal from the input monaural local decoded signal of section Y or section Y+X (wherein, as described above, the incomplete monaural local decoded signal of section X) using a predetermined known prediction technique (step S135A-1). In addition, in the case of the fifth embodiment in which the fourth embodiment A is modified or the fifth embodiment in which the fourth embodiment B is modified, the prediction signal for section X includes the input incomplete monaural local decoded signal of section X. Next, the additional coding unit 130 encodes the difference signal or weighted difference signal (a signal formed by subtracting or weighted subtraction of the sample values ​​between corresponding samples) between the downmix signal of the interval Y and the monaural local decoded signal, and the difference signal or weighted difference signal (a signal formed by subtracting or weighted subtraction of the sample values ​​between corresponding samples) between the downmix signal of the interval X and the prediction signal obtained in step S135A-1, obtains and outputs the additional code CA (step S135A-2). For example, the additional code CA may be obtained by encoding a signal obtained by combining the difference signal of the interval Y and the difference signal of the interval X, or by encoding the difference signal of the interval Y and the difference signal of the interval X separately, and obtaining a code obtained by combining the obtained codes as the additional code CA. In the encoding, the same encoding method as that of the additional coding unit 130 in each of the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B may be used.

[0126] [Stereo decoding unit 220]

[0127] The stereo decoding unit 220 proceeds from the following step S225A-0 to step S225A-2 (step S225A). The stereo decoding unit 220 first obtains a prediction signal of the interval X from the mono decoded sound signal of the interval Y or the interval Y+X using the same prediction technique as the prediction technique used by the additional encoding unit 130 in step S135 (step S225A-0). Next, the stereo decoding unit 220 obtains a decoded down-mix signal of the interval Y+X by combining a sum signal or a weighted sum signal (a signal formed by adding or weighted adding the sample values ​​between corresponding samples) of the mono decoded sound signal of the interval Y and the additional decoded signal, and a sum signal or a weighted sum signal (a signal formed by adding or weighted adding the sample values ​​between corresponding samples) of the additional decoded signal of the interval X and the prediction signal (step S225A-1). Next, the stereo decoding unit 220 obtains two-channel decoded audio signals from the decoded downmix signal obtained in step S225A-1 by performing an upmix process using the characteristic parameters obtained from the stereo code CS, and outputs the decoded audio signals (step S225A-2).

[0128] <<Fifth Embodiment B>>

[0129] Next, a fifth embodiment obtained by modifying each of the first embodiment and the fourth embodiment C will be described as a fifth embodiment B, and points different from each of the first embodiment and the fourth embodiment C will be described.

[0130] [Monaural encoding unit 120]

[0131] The mono encoding unit 120 obtains and outputs a signal obtained by decoding the monaural code CM of the interval Y or interval Y+X up to the current frame, i.e., a local decoded signal of the input downmix signal, i.e., a monaural local decoded signal, in addition to the monaural code CM obtained by encoding the downmix signal (step S125B). However, as described above, the monaural local decoded signal of the interval X is an incomplete monaural local decoded signal.

[0132] [Additional encoding unit 130]

[0133] The additional coding unit 130 performs the following steps S135B-1 and S135B-2 (step S135B). The additional coding unit 130 first uses a predetermined known prediction technique to obtain a prediction signal for section X of the monaural local decoded signal from the input monaural local decoded signal of section Y or section Y+X (wherein, as described above, section X is an incomplete monaural local decoded signal) (step S135B-1). In the case of the fifth embodiment in which the fourth embodiment C is changed, the prediction signal for section X includes the input monaural local decoded signal of section X. The additional coding unit 130 then encodes the difference signal or the weighted difference signal (a signal formed by subtracting the sample values ​​between corresponding samples or performing weighted subtraction) between the downlink mixed signal of section X and the prediction signal obtained in step S135B-1 to obtain and output the additional code CA (step S135B-2). For example, the same encoding method as that of the additional encoding unit 130 in each of the first embodiment and the fourth embodiment C may be used for encoding.

[0134] [Stereo decoding unit 220]

[0135] The stereo decoding unit 220 proceeds from the following step S225B-0 to step S225B-2 (step S225B). The stereo decoding unit 220 first uses the same prediction technique as the prediction technique used by the additional encoding unit 130 to obtain a prediction signal for interval X from the mono decoded sound signal of interval Y or interval Y+X (step S225B-0). Next, the stereo decoding unit 220 obtains a signal obtained by connecting the sum signal of the mono decoded sound signal of interval Y, the additional decoded signal of interval X, and the prediction signal, or a weighted sum signal (a signal formed by adding or weighted adding the sample values ​​between corresponding samples) as a decoded down-mix signal of interval Y+X (step S225B-1). Next, the stereo decoding unit 220 obtains two-channel decoded sound signals from the decoded down-mix signal obtained in step S225B-1 by up-mixing processing using the characteristic parameters obtained from the stereo code CS and outputs them (step S225B-2).

[0136] <Sixth Embodiment>

[0137] In the first to fifth embodiments, the decoding device 200 uses the additional code CA obtained by the encoding device 100 to decode at least the additional code CA to obtain a decoded down-mix signal of the section X used in the stereo decoding unit 220. However, the decoding device 200 may use a prediction signal of a monaural decoded audio signal from the section Y as a decoded down-mix signal of the section X used in the stereo decoding unit 220 without using the additional code CA. This embodiment will be referred to as the sixth embodiment, and points different from the first embodiment will be described.

[0138] <<Encoding device 100>>

[0139] The encoding device 100 of the sixth embodiment is different from the encoding device 100 of the first embodiment in that the encoding device 100 does not include the additional encoding unit 130, does not encode the downlink mixed signal of the interval X, and does not obtain the additional code CA. That is, the encoding device 100 of the sixth embodiment includes a stereo encoding unit 110 and a monaural encoding unit 120, and the stereo encoding unit 110 and the monaural encoding unit 120 perform the same operations as the stereo encoding unit 110 and the monaural encoding unit 120 of the first embodiment, respectively.

[0140] <<Decoding device 200>>

[0141] The decoding device 200 of the sixth embodiment does not include the additional decoding unit 230 that decodes the additional code CA, but includes a monaural decoding unit 210 and a stereo decoding unit 220. The monaural decoding unit 210 of the sixth embodiment performs the same operation as the monaural decoding unit 210 of the first embodiment, but when the stereo decoding unit 220 uses the monaural decoded audio signal of the interval Y+X, it also outputs the monaural decoded audio signal of the interval X. In addition, the stereo decoding unit 220 of the sixth embodiment performs the following operation different from the stereo decoding unit 220 of the first embodiment.

[0142] [Stereo decoding unit 220]

[0143] The stereo decoding unit 220 performs the following steps S226-0 to S226-2 (step S226). The stereo decoding unit 220 first obtains a prediction signal for the interval X from the mono decoded sound signal of the interval Y or the interval Y+X using a known prediction technique as specified in the fifth embodiment (step S226-0). Next, the stereo decoding unit 220 obtains a signal obtained by combining the mono decoded sound signal of the interval Y and the prediction signal of the interval X as a decoded down-mix signal of the interval Y+X (step S226-1), and obtains two-channel decoded sound signals from the decoded down-mix signal obtained in step S226-1 through an up-mix process using a feature parameter obtained from the stereo code CS and outputs the decoded sound signals (step S226-2).

[0144] <Seventh Embodiment>

[0145] In the above embodiments, for the sake of simplicity, an example of processing a sound signal of two channels is used for explanation. However, the number of channels is not limited to this, as long as it is 2 or more. If the number of channels is set to C (C is an integer greater than 2), the above embodiments can be implemented by replacing the two channels with C (C is an integer greater than 2) channels.

[0146] For example, the encoding device 100 of the first to fifth embodiments may obtain the stereo code CS, the monaural code CM, and the additional code CA from the input C-channel audio signals, the encoding device 100 of the sixth embodiment may obtain the stereo code CS and the monaural code CM from the input C-channel audio signals, the stereo encoding unit 110 may generate and output a code representing information corresponding to the difference between channels in the input C-channel audio signals as the stereo code CS, the stereo encoding unit 110 or the downmixing unit 150 may output a signal obtained by mixing the input C-channel audio signals as the downmixed signal, and the monaural encoding target signal generating unit 140 may obtain and output a signal obtained by mixing the input C-channel audio signals in the time domain as the encoding target signal. The information corresponding to the difference between channels in the C-channel audio signals is, for example, information corresponding to the difference between the audio signal of the channel and the audio signal of the reference channel for each of the C-1 channels other than the reference channel.

[0147] Similarly, the decoding apparatus 200 of the first to fifth embodiments may obtain and output decoded audio signals of C channels based on the input monaural code CM, the additional code CA, and the stereo code CS, the decoding apparatus 200 of the sixth embodiment may obtain and output decoded audio signals of C channels based on the input monaural code CM and the stereo code CS, and the stereo decoding unit 220 may obtain and output decoded audio signals of C channels from the decoded down-mix signal by up-mixing processing using characteristic parameters obtained based on the input stereo code CS. More specifically, the stereo decoding unit 220 may regard the decoded down-mix signal as a signal mixed with decoded audio signals of C channels, regard the characteristic parameters obtained based on the input stereo code CS as information indicating the characteristics of the difference between channels in the decoded audio signals of C channels, obtain and output decoded audio signals of C channels.

[0148] <Program and Recording Medium>

[0149] The processing of each part of each encoding device and each decoding device can be realized by a computer. In this case, the processing content of the function that each device should have is recorded in a program. Then, the program is read into Fig.12 The storage unit 1020 of the computer shown in the figure operates the arithmetic processing unit 1010, the input unit 1030, the output unit 1040, etc., thereby realizing various processing functions in the above-mentioned devices on the computer.

[0150] The program describing the processing contents can be recorded in a computer-readable recording medium. The computer-readable recording medium is, for example, a non-transitory recording medium, specifically, a magnetic recording device, an optical disk, and the like.

[0151] In addition, the program can be circulated by, for example, selling, transferring, or lending a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program can also be circulated by storing the program in a storage device of a server computer and transmitting the program from the server computer to other computers via a network.

[0152] The computer that executes such a program first temporarily stores the program recorded in a portable recording medium or the program transferred from a server computer in the auxiliary recording unit 1050, which is its own non-temporary storage device, for example. Moreover, when executing the process, the computer reads the program stored in the auxiliary recording unit 1050, which is its own non-temporary storage device, into the storage unit 1020, and executes the process according to the read program. In addition, as another execution method of the program, the computer can directly read the program from the portable recording medium into the storage unit 1020 and execute the process according to the program, and further, each time the program is transferred from the server computer to the computer, the computer can sequentially execute the process according to the received program. In addition, it is also possible to configure that the program is not transferred from the server computer to the computer, but the above-mentioned process is executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only by its execution instruction and result acquisition. In addition, the program in this method includes a program based on the program for the processing of the electronic computer (data that is not a direct instruction to the computer but has the nature of specifying the processing of the computer, etc.).

[0153] In this embodiment, the present device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.

[0154] Furthermore, it is of course possible to make appropriate changes within the scope not departing from the gist of the present invention.

Claims

1. A method for decoding a sound signal, which decodes an input code in units of frames to obtain decoded sound signals of C channels, wherein: C is an integer greater than or equal to 2, and the audio signal decoding method is characterized in that: The current frame processing includes: a mono decoding step, decoding the mono code included in the input code in a decoding method including applying a window process with overlap between frames to obtain a mono decoded sound signal; an additional decoding step of decoding an additional code included in the input code to obtain an additional decoded signal, the additional decoded signal being a mono decoded signal of the overlapping interval X of the current frame and the next frame; and a stereo decoding step, obtaining a decoded downmix signal, wherein the decoded downmix signal is a signal obtained by combining a signal of a section Y other than the section X in the mono decoded sound signal and the additional decoded signal of the section X, By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

2. The sound signal decoding method according to claim 1, characterized in that: In the additional decoding step, the additional code included in the input code is decoded to obtain an additional decoded signal which is a mono decoded signal of the interval Y and the interval X. In the stereo decoding step, A decoded down-mix signal is obtained by combining a signal formed by adding or weighted adding the sample values ​​between the corresponding samples of the monaural decoded audio signal and the additional decoded signal in section Y with the additional decoded signal in section X, By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

3. The sound signal decoding method according to claim 2, characterized in that: In the additional decoding step, The first additional code included in the additional code is decoded to obtain a first additional decoded signal of section X, The second additional code included in the additional code is decoded in a decoding manner of obtaining a combined sampling string from the code to obtain a second additional decoding signal of interval Y and interval X, obtaining the second additional decoded signal of interval Y as the additional decoded signal of interval Y, A signal formed by adding or weighted adding the sample values ​​between corresponding samples of the first additionally decoded signal and the second additionally decoded signal in section X is obtained as the additionally decoded signal in section X.

4. The sound signal decoding method according to claim 2, characterized in that: In the stereo decoding step, A prediction signal for interval X is obtained from the mono decoded sound signal for interval Y or the mono decoded sound signals for interval Y and interval X, a signal obtained by combining a signal formed by adding or weighted adding the sampling values ​​between the corresponding samples of the monaural decoded sound signal and the additional decoded signal in interval Y and a signal formed by adding or weighted adding the sampling values ​​between the corresponding samples of the predicted signal and the additional decoded signal in interval X, i.e., a decoded down-mix signal; By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

5. A sound signal decoding device, which decodes an input code in units of frames to obtain decoded sound signals of C channels, wherein: C is an integer greater than or equal to 2, and the audio signal decoding device is characterized in that: The current frame processing includes: a monaural decoding unit that decodes the monaural code included in the input code by a decoding method including applying a process of overlapping windows between frames to obtain a monaural decoded audio signal; An additional decoding unit decodes the additional code included in the input code to obtain an additional decoded signal, which is a mono decoded signal of the overlapping interval X of the current frame and the next frame; a stereo decoding unit, wherein the decoded downmix signal is obtained by combining a signal of a section Y other than the section X in the mono decoded audio signal and the additional decoded signal of the section X; By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

6. The sound signal decoding device according to claim 5, characterized in that: The additional decoding unit decodes the additional code included in the input code to obtain an additional decoded signal which is a mono decoded signal of the interval Y and the interval X. The stereo decoding unit obtains a decoded down-mix signal which is a signal formed by adding or weighted adding the sample values ​​between the monaural decoded audio signal and the corresponding samples of the additional decoded signal in section Y and the additional decoded signal in section X. By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

7. The sound signal decoding device according to claim 6, characterized in that: The additional decoding unit performs: The first additional code included in the additional code is decoded to obtain a first additional decoded signal of section X, The second additional code included in the additional code is decoded in a decoding manner of obtaining a combined sampling string from the code to obtain a second additional decoding signal of interval Y and interval X, obtaining the second additional decoded signal of interval Y as the additional decoded signal of interval Y, A signal formed by adding or weighted adding the sample values ​​between the corresponding samples of the first additionally decoded signal and the second additionally decoded signal in section X is obtained as the additionally decoded signal in section X.

8. The sound signal decoding device according to claim 6, characterized in that: The stereo decoding unit performs: A prediction signal for interval X is obtained from the mono decoded sound signal for interval Y or the mono decoded sound signals for interval Y and interval X, a signal obtained by combining a signal formed by adding or weighted adding the sampling values ​​between the corresponding samples of the monaural decoded sound signal and the additional decoded signal in interval Y and a signal formed by adding or weighted adding the sampling values ​​between the corresponding samples of the predicted signal and the additional decoded signal in interval X, i.e., a decoded down-mix signal; By performing an upmix process using a feature parameter obtained from a stereo code included in the input code, decoded audio signals of C channels are obtained from the decoded downmix signal and output.

9. A computer program product, comprising a computer program, the computer program being configured to cause a computer to execute the steps of the sound signal decoding method according to any one of claims 1 to 4.

10. A computer-readable recording medium recording a program for causing a computer to execute each step of the sound signal decoding method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Hybrid audio signal encoder, hybrid audio signal decoder, method for encoding audio signal, and method for decoding audio signal

    CN103548080A

  • Audio Encoder For Encoding A Multichannel Signal And Audio Decoder For Decoding An Encoded Audio Signal

    CN107430863A