Method for Encoding Audio Signal, Audio Signal Encoding Device, Computer Program Product, and Recording Medium
Through the stereo encoding method, the encoding processing and feature parameters based on frame units are used to solve the problem of large stereo encoding/decoding delay, and efficient signal control in multi-site teleconferencing systems is realized.
Patent Information
- Application Number
- CN202080102309.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-06-24
AI Technical Summary
In the prior art, the algorithm delay of the stereo encoding/decoding method is greater than that of the mono encoding/decoding method, resulting in complex control in a multi-site teleconferencing system, making it difficult to switch signals within each specified time interval.
The stereo encoding method is adopted, through frame-based encoding processing, including stereo encoding, downmix, mono encoding and additional encoding steps, and the signal encoding and decoding is used to use characteristic parameters to reduce algorithm delay.
The embedded encoding/decoding of multiple channel sound signals and mono sound signals with a stereo encoding/decoding delay of no greater than that of mono encoding/decoding, simplifying signal control in multi-location teleconferencing systems.
Smart Images

Figure CN115917644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for embedded encoding / decoding of audio signals of multiple channels and an audio signal of one channel. Background Art
[0002] As a technique for embedded encoding / decoding of audio signals of multiple channels and a monophonic audio signal, there is the technique of Non-Patent Document 1. The outline of the technique of Non-Patent Document 1 will be described by the encoding device 500 exemplified in Figure 5 and the decoding device 600 exemplified in Figure 6 The stereo encoding unit 510 of the encoding device 500 obtains, for each predetermined time interval, i.e., each frame, a stereo code CS representing characteristic parameters and a down-mixed signal obtained by mixing the stereo input audio signals, from the input audio signals of multiple channels, i.e., the stereo input audio signals. The characteristic parameters are parameters representing the characteristics of the difference between channels in the stereo input audio signals. The monophonic encoding unit 520 of the encoding device 500 encodes the down-mixed signal for each frame to obtain a monophonic code CM. The monophonic decoding unit 610 of the decoding device 600 decodes the monophonic code CM for each frame to obtain a monophonic decoded audio signal as a decoded signal of the down-mixed signal. The stereo decoding unit 620 of the decoding device 600 decodes the stereo code CS for each frame to obtain characteristic parameters as parameters representing the characteristics of the difference between channels, and performs a process (so-called up-mixing process) of obtaining a stereo decoded audio signal from the monophonic decoded audio signal and the characteristic parameters.
[0003] As a monophonic encoding / decoding method for obtaining a high-quality monophonic decoded audio signal, there is the encoding / decoding method of the 3GPP EVS standard described in Non-Patent Document 2. As the monophonic encoding / decoding method of Non-Patent Document 1, if a high-quality monophonic encoding / decoding method such as that of Non-Patent Document 2 is used, it may be possible to achieve higher-quality embedded encoding / decoding of audio signals of multiple channels and a monophonic audio signal.
[0004] Prior Art Documents
[0005] Non-Patent Documents
[0006] Non-Patent Document 1: Jeroen Breebaart et al., "Parametric Coding of Stereo Audio", EURASIP Journal on Applied Signal Processing, pp. 1305-1322, 2005:9.
[0007] Non-Patent Document 2: 3GPP, "Codec for Enhanced Voice Services (EVS); Detailed algorithmic description", TS26.445. Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] The upmix processing in Non-Patent Document 1 is signal processing in the frequency domain that includes processing of applying a window that overlaps between adjacent frames to the mono decoded sound signal. On the other hand, the mono encoding / decoding method in Non-Patent Document 2 also includes processing of applying a window with overlap between adjacent frames. That is, on both the decoding side of the stereo encoding / decoding method in Non-Patent Document 1 and the decoding side of the mono encoding / decoding method in Non-Patent Document 2, for a specified range at the boundary portion of the frame, by synthesizing a signal of an inclined window whose shape attenuates the signal obtained by decoding the code of the front frame and a signal of an inclined window whose shape increases the signal obtained by decoding the code of the rear frame, a decoded sound signal is obtained. Thus, if a mono encoding / decoding method such as that of Non-Patent Document 1, which is an embedded encoding / decoding method, uses a mono encoding / decoding method such as that of Non-Patent Document 2, there is a problem that the stereo decoded sound signal is delayed by the window in the upmix processing compared to the mono decoded sound signal, that is, there is a problem that the algorithm delay of stereo encoding / decoding is larger than that of mono encoding / decoding.
[0010] For example, in a Multipoint Control Unit (MCU) used for a multi-site teleconference, control of switching which signal from which site is output to which site is usually performed for each specified time interval. It is assumed that it is difficult to perform control in a state where the stereo decoded sound signal is delayed by the window in the upmix processing compared to the mono decoded sound signal, and an installation is made to perform control in a state where the stereo decoded sound signal is delayed by 1 frame compared to the mono decoded sound signal. That is, in a communication system including a multipoint control device, the above problem becomes more prominent, and there is a possibility that the algorithm delay of stereo encoding / decoding is 1 frame larger than the algorithm delay of mono encoding / decoding. In addition, if the stereo decoded sound signal is delayed by 1 frame compared to the mono decoded sound signal, the control of switching itself can be performed for each specified time interval, but for each time interval, the control of combining the mono decoded sound signal from which site and the stereo decoded sound signal from which site and outputting them may become complicated due to the different delays between the mono decoded sound signal and the stereo decoded sound signal.
[0011] The present invention has been made in view of such problems, and an object thereof is to provide an embedded encoding / decoding of a sound signal of a plurality of channels and a sound signal of a monophonic channel in which the algorithm delay of stereo encoding / decoding is not greater than the algorithm delay of monophonic encoding / decoding.
[0012] Means for Solving the Problems
[0013] In order to solve the above problems, a sound signal encoding method according to an aspect of the present invention is a sound signal encoding method for encoding an input sound signal of C (C is an integer of 2 or more) channels in units of frames, and includes, as processing of a current frame: a stereo encoding step of obtaining and outputting a stereo code representing characteristic parameters, the characteristic parameters being parameters representing characteristics of differences between channels of sound signals of C channels; a downmixing step of obtaining a signal obtained by mixing sound signals of C channels as a downmixing signal; a monophonic encoding step of encoding the downmixing signal to obtain a monophonic code and outputting it, and in the monophonic encoding step, encoding the downmixing signal in an encoding method including processing of applying a window having an overlap between frames to obtain a monophonic code, and further includes an additional encoding step in which a signal in an overlapping section between the current frame and the immediately subsequent frame in the downmixing signal is encoded to obtain an additional code and output.
[0014] Advantages of the Invention
[0015] According to the present invention, it is possible to provide an embedded encoding / decoding of a sound signal of a plurality of channels and a sound signal of a monophonic channel in which the algorithm delay of stereo encoding / decoding is not greater than the algorithm delay of monophonic encoding / decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a block diagram showing an example of an encoding apparatus according to each embodiment.
[0017] Figure 2 It is a flowchart showing an example of processing of an encoding apparatus according to each embodiment.
[0018] Figure 3 It is a block diagram showing an example of a decoding apparatus according to each embodiment.
[0019] Figure 4 It is a flowchart showing an example of processing of a decoding apparatus according to each embodiment.
[0020] Figure 5 It is a block diagram showing an example of a conventional encoding apparatus.
[0021] Figure 6 It is a block diagram showing an example of a conventional decoding apparatus.
[0022] Figure 7It is a diagram schematically showing each signal of the encoding device of Non-Patent Document 2.
[0023] Figure 8 It is a diagram schematically showing each signal and algorithm delay of the decoding device of Non-Patent Document 2.
[0024] Figure 9 It is a diagram schematically showing each signal and algorithm delay of the decoding device of Non-Patent Document 1 in the case of using the monaural encoding / decoding method of Non-Patent Document 2 as the monaural encoding / decoding method.
[0025] Figure 10 It is a diagram schematically showing each signal and algorithm delay of the decoding device of the present invention.
[0026] Figure 11 It is a diagram schematically showing each signal of the encoding device of the present invention.
[0027] Figure 12 It is a diagram showing an example of the functional structure of a computer that implements each device in each embodiment. Detailed Embodiments
[0028] Before describing each embodiment, first, regarding the background art and each signal and algorithm delay of the encoding / decoding of the first embodiment, from the diagrams that schematically illustrate each signal in the case where the frame length is 20 ms Figures 7 to 11 an explanation will be given. Figures 7 to 11 The horizontal axis of each of the following diagrams is the time axis. Hereinafter, since an example of processing the current frame at the time point of t7 is described, on the axis arranged at the uppermost part of each diagram, a description with the left end marked as "past" and the right end marked as "future" is given, and an upward arrow is marked at t7 which is the time point for processing the current frame. In Figures 7 to 11 , for each signal, it is schematically shown which time interval the signal is in, and whether the window is in an increasing shape, a flat shape, or a decaying shape when the window is applied. More specifically, what the window function is exactly like is not important in the present description. Therefore, in Figures 7 to 11 , regarding the intervals of the window increasing shape and the window decreasing shape, in order to visually represent the signal that becomes the applied window during synthesis, the interval of the window increasing shape is represented by a triangle including a straight line slanting upward to the right, and the interval of the window decreasing shape is represented by a triangle including a straight line slanting downward to the right. Additionally, hereinafter, in order to avoid complicated statement expressions, statements such as "to" and "after" are used to determine the start time of each interval. However, as those skilled in the art can understand, the actual start of each interval is the time point immediately after the recorded time point, and the actual start of the digital signal of each interval is the sample immediately after the recorded time point.
[0029] Figure 7FIG. is a diagram schematically showing signals of an encoding apparatus of Non-Patent Document 2 that processes a current frame at time point t7. The signal that can be used by the encoding apparatus of Non-Patent Document 2 for processing the current frame is Signal 1a, which is a monaural audio signal before t7. In the encoding apparatus of Non-Patent Document 2, in the processing of the current frame, an 8.75 ms interval from t6 to t7 in Signal 1a is used as a so-called "pre-read interval" for analysis, and Signal 1b, which is a signal obtained by applying a window to the signal in the 23.25 ms interval from t1 to t6 in Signal 1a, is encoded to obtain a monaural code and output. The shape of the window is such that it increases in the 3.25 ms interval from t1 to t2, is flat in the 16.75 ms interval from t2 to t5, and decays in the 3.25 ms interval from t5 to t6. That is, this Signal 1b is a monaural audio signal corresponding to the monaural code obtained by processing the current frame. In the encoding apparatus of Non-Patent Document 2, as the processing of the immediately preceding frame, at the time point when a monaural audio signal before t3 is input, the same processing is completed, and the encoding of Signal 1c, which is a signal obtained by applying a window with a decaying shape in the interval from t1 to t2 to the monaural audio signal in the 23.25 ms interval before t2, is completed. That is, this Signal 1c is a monaural audio signal corresponding to the monaural code obtained by processing the immediately preceding frame, and the interval from t1 to t2 is the overlapping interval between the current frame and the immediately preceding frame. Further, as the processing of the immediately succeeding frame, the encoding apparatus of Non-Patent Document 2 encodes Signal 1d, which is a signal obtained by applying a window with an increasing shape in the interval from t5 to t6 to the monaural audio signal in the 23.25 ms interval after t5. That is, this Signal 1d is a monaural audio signal corresponding to the monaural code obtained by processing the immediately succeeding frame, and the interval from t5 to t6 is the overlapping interval between the current frame and the immediately succeeding frame.
[0030] Figure 8FIG. is a diagram showing signals of a decoding apparatus of Non-Patent Document 2 that schematically shows processing of a current frame at time point t7 when a monaural code of the current frame is input from an encoding apparatus of Non-Patent Document 2. In processing of the current frame, the decoding apparatus of Non-Patent Document 2 obtains a decoded audio signal, i.e., signal 2a, in the period from t1 to t6 from the monaural code of the current frame. This signal 2a is a decoded audio signal corresponding to signal 1b, and is a signal of an application window that increases in the period from t1 to t2, is flat in the period from t2 to t5, and decays in the period from t5 to t6. As processing of the immediately preceding frame, at time point t3 when the monaural code of the immediately preceding frame is input, the decoding apparatus of Non-Patent Document 2 has obtained a decoded audio signal, i.e., signal 2b, in a 23.25-ms period up to t2 of an application window that decays in the period from t1 to t2 from the monaural code of the immediately preceding frame. Further, as processing of the immediately succeeding frame, the decoding apparatus of Non-Patent Document 2 obtains a decoded audio signal, i.e., signal 2c, in a 23.25-ms period after t5 of an application window that increases in the period from t5 to t6 from the monaural code of the immediately succeeding frame. However, since signal 2c cannot be obtained at time point t7, at time point t7, although an incomplete decoded audio signal can be obtained for the period from t5 to t6, a complete decoded audio signal cannot be obtained. Therefore, at time point t7, the decoding apparatus of Non-Patent Document 2 synthesizes signal 2b obtained by the immediately preceding processing and signal 2a obtained by processing of the current frame for the period from t1 to t2, and directly uses signal 2a obtained by processing of the current frame for the period from t2 to t5, thereby obtaining and outputting a monaural decoded audio signal, i.e., signal 2d, in a 20-ms period from t1 to t5. Since the decoding apparatus of Non-Patent Document 2 obtains a decoded audio signal in the period starting from t1 at time point t7, the algorithm delay of the monaural encoding / decoding method of Non-Patent Document 2 is the time length from t1 to t7, i.e., 32 ms.
[0031] Figure 9FIG. is a diagram showing signals of the decoding apparatus 600 of Patent Document 1 when the monaural decoding unit 610 uses the monaural decoding method of Non-Patent Document 2. At time t7, the stereo decoding unit 620 performs stereo decoding processing (upmixing processing) of the current frame using the signal 3a as the monaural decoded sound signal until t5 completely obtained by the monaural decoding unit 610. Specifically, the stereo decoding unit 620 applies a window whose shape increases in the 3.25 ms interval from t0 to t1, is flat in the 16.75 ms interval from t1 to t4, and decays in the 3.25 ms interval from t4 to t5 to the signal 3a, and obtains, for each channel, a decoded sound signal, i.e., signal 3c-i, of the window having the same shape as the signal 3b in the 23.25 ms interval from t0 to t5. As processing of the immediately preceding frame, the stereo decoding unit 620 has already obtained, at time t3, the decoded sound signals of each channel, i.e., signal 3d-i, in the 23.25 ms interval up to t1 of the window whose shape decays in the interval from t0 to t1. Further, as processing of the immediately following frame, the stereo decoding unit 620 obtains the decoded sound signals of each channel, i.e., signal 3e-i, in the 23.25 ms interval after t4 of the window having an increasing form in the interval from t4 to t5. However, since the signal 3e-i cannot be obtained at time t7, at time t7, an incomplete decoded sound signal is obtained in the interval from t4 to t5, but a complete decoded sound signal cannot be obtained. Therefore, at time t7, for each channel, in the interval from t0 to t1, the stereo decoding unit 620 synthesizes the signal 3d-i obtained by the processing of the immediately preceding frame and the signal 3c-i obtained by the processing of the current frame, and in the interval from t1 to t4, directly uses the signal 3c-i obtained by the processing of the current frame, thereby obtaining and outputting a complete decoded sound signal, i.e., signal 3f-i, in the 20 ms interval from t0 to t4. Since the decoding apparatus 600 obtains the decoded sound signals in the intervals starting from t0 of each channel at time t7, the algorithm delay of the stereo encoding / decoding of Patent Document 1 using the monaural encoding / decoding method of Non-Patent Document 2 as the monaural encoding / decoding method is 35.25 ms, which is the time length from t0 to t7. That is, the algorithm delay of the stereo encoding / decoding in this embedded encoding / decoding becomes larger than the algorithm delay based on the monaural encoding / decoding.
[0032] Figure 10 FIG. is a diagram schematically showing signals of the decoding apparatus 200 of the first embodiment described later. The decoding apparatus 200 of the first embodiment is Figure 3The structure shown includes a mono decoder 210 that performs the operations described in detail in the first embodiment, an additional decoder 230, and a stereo decoder 220. The stereo decoder 220 processes the current frame at time t7 using the mono decoded sound signal, i.e., signal 4c, which is fully obtained before t6. At Figure 9As described above, at the time point t7, what is completely obtained by the mono decoding unit 210 at the location of the description is the mono decoded sound signal before t5, that is, signal 3a. Therefore, in the decoding device 200, the additional decoding unit 230 decodes the additional code CA to obtain the mono decoded sound signal in the 3.25 ms interval from t5 to t6, that is, signal 4b (additional decoding process). The stereo decoding unit 220 uses the signal 4c obtained by connecting the mono decoded sound signal before t5, that is, signal 3a, obtained by the mono decoding unit 210 and the mono decoded sound signal in the interval from t5 to t6, that is, signal 4b, obtained by the additional decoding unit 230, to perform the stereo decoding process (upmixing process) of the current frame. That is, for signal 4c, the stereo decoding unit 220 adopts the shape that increases in the 3.25 ms interval from t1 to t2, is flat in the 16.75 ms interval from t2 to t5, and decays in the 3.25 ms interval from t5 to t6. For each channel, the decoded sound signal in the interval from t1 to t6 of the application window with the same shape as signal 4d, that is, signal 4e-i, is obtained. As the processing of the immediately preceding frame, at the time point t3, the stereo decoding unit 220 has already obtained the decoded sound signals of each channel in the 23.75 ms interval before t2 of the application window with the shape that decays in the interval from t1 to t2, that is, signal 4f-i. In addition, as the processing of the immediately following frame, the stereo decoding unit 220 obtains the decoded sound signals of each channel in the 23.75 ms interval after t5 of the application window with the increasing form in the interval from t5 to t6, that is, signal 4g-i. However, since signal 4g-i cannot be obtained at the time point t7, an incomplete decoded sound signal is obtained for the interval from t5 to t6 at the time point t7, but a complete decoded sound signal cannot be obtained. Therefore, at the time point t7, for each channel, the stereo decoding unit 220 synthesizes signal 4f-i obtained by the processing of the immediately preceding frame and signal 4e-i obtained by the processing of the current frame for the interval from t1 to t2, and directly uses signal 4e-i obtained by the processing of the current frame for the interval from t2 to t5, so as to obtain and output the complete decoded sound signal in the 20 ms interval from t1 to t5, that is, signal 4h-i. Since the decoding device 200 obtains the decoded sound signals in the interval starting from t1 of each channel at the time point t7, the algorithm delay of the stereo encoding / decoding in the embedded encoding / decoding of the first embodiment is the time length of 32 ms from t1 to t7. That is, the algorithm delay of the stereo encoding / decoding based on the embedded encoding / decoding in the first embodiment is not greater than the algorithm delay based on the mono encoding / decoding.
[0033] Figure 11 It is for the encoding device 100 of the first embodiment described below, so that each signal becomesFigure 10 A diagram schematically showing each signal of the encoding device corresponding to the decoding device 200 of the first embodiment of the decoding device schematically shown. The encoding device 100 of the first embodiment is Figure 1 of the structure shown, and in addition to including: a stereo encoding unit 110 that performs the same processing as the stereo encoding unit 510 of the encoding device 500; and a mono encoding unit 120 that, similarly to the mono encoding unit 520 of the encoding device 500, applies a window to the signal in the interval from t1 to t6 in the mono sound signal 1a before t7, i.e., the signal 1b, and encodes it to obtain a mono code CM, further includes an additional encoding unit 130 that encodes the signal in the overlapping interval from t5 to t6 between the current frame and the next frame in the mono sound signal 1a, i.e., the signal 5c, to obtain an additional code CA.
[0034] Thereafter, the section from t5 to t6, which is the overlapping section between the current frame and the subsequent frame, is referred to as "section X". That is, on the encoding side, section X is the section where the mono encoding unit 120 encodes the mono audio signal of the applied window in both the processing of the current frame and the processing of the subsequent frame. More specifically, section X is the section of a specified length including the terminal in the audio signal encoded by the mono encoding unit 120 in the processing of the current frame, is the section encoding the audio signal of the applied window with the shape attenuated in the processing of the current frame by the mono encoding unit 120, is the section of a specified length including the start end in the section encoded by the mono encoding unit 120 in the processing of the subsequent frame, and is the section encoding the audio signal of the applied window with the shape increased in the processing of the subsequent frame by the mono encoding unit 120. In addition, on the decoding side, section X is the section where the mono decoding unit 210 decodes the mono code CM to obtain the decoded audio signal of the applied window in both the processing of the current frame and the processing of the subsequent frame. More specifically, section X is the section of a specified length including the terminal in the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM through the processing of the current frame, is the section of the applied window with the attenuated shape in the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM through the processing of the current frame. The mono decoding unit 210 is the section of a specified length including the start end in the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM in the processing of the subsequent frame, is the section of the applied window with the increased shape in the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM in the processing of the subsequent frame. The mono decoding unit 210 is the decoded audio signal obtained by synthesizing the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM through the processing of the current frame and the decoded audio signal obtained by the mono decoding unit 210 decoding the mono code CM in the processing of the subsequent frame in the processing of the subsequent frame.
[0035] In addition, hereafter, the section from t1 to t5, which is the section other than section X in the section mono-encoded / decoded in the processing of the current frame, is referred to as "section Y". That is, section Y on the encoding side is the section other than the overlapping section with the subsequent frame in the section where the mono encoding unit 120 encodes the mono audio signal in the processing of the current frame, and on the decoding side is the section other than the overlapping section with the subsequent frame in the section where the mono decoding unit 210 decodes the mono code CM to obtain the decoded audio signal in the processing of the current frame. Section Y is the section formed by connecting the section representing the mono audio signal using the mono code CM of the current frame and the mono code CM of the previous frame and the section representing the mono audio signal using only the mono code CM of the current frame. Therefore, it is the section where the mono decoded audio signal is completely obtained up to the processing of the current frame.
[0036] <First Embodiment>
[0037] The encoding device and the decoding device according to the first embodiment will be described.
[0038] [[Encoding Device 100]]
[0039] As Figure 1 shown, the encoding device 100 according to the first embodiment includes a stereo encoding unit 110, a mono encoding unit 120, and an additional encoding unit 130. The encoding device 100 encodes an input two-channel stereo audio time-domain sound signal (two-channel stereo input sound signal) in units of frames having a predetermined time length of, for example, 20 ms, and obtains and outputs a stereo code CS, a mono code CM, and an additional code CA described later. The two-channel stereo input sound signal input to the encoding device 100 is composed of, for example, a digital sound signal or an audio signal obtained by respectively picking up sounds such as voices and music through two microphones and performing AD conversion, an input sound signal of the left channel as the first channel, and an input sound signal of the right channel as the second channel. The codes output by the encoding device 100, that is, the stereo code CS, the mono code CM, and the additional code CA are input to a decoding device 200 described later. The encoding device 100 performs Figure 2 the processes of step S111, step S121, and step S131 illustrated for each frame, that is, whenever the above-described two-channel stereo input sound signal having the predetermined time length is input. In the above example, when a 20-ms amount of two-channel stereo input sound signal from t3 to t7 is input, the encoding device 100 performs the processes of step S111, step S121, and step S131 for the current frame.
[0040] [[Stereo Encoding Unit 110]]
[0041] The stereo encoding unit 110 obtains and outputs a stereo code CS and a downmix signal (step S111) from the two-channel stereo input sound signal input to the encoding device 100. The stereo code CS represents a characteristic parameter that is a parameter representing the characteristics of the difference between the sound signals of the two input channels, and the downmix signal is a signal obtained by mixing the sound signals of the two channels.
[0042] [[Example of Stereo Encoding Unit 110]]
[0043] As an example of the stereo encoding unit 110, the operation of each frame of the stereo encoding unit 110 will be described in the case where information indicating the intensity difference of each frequency band of the sound signals of the two input channels is used as a characteristic parameter. In addition, a specific example using a complex DFT (Discrete Fourier Transformation) is described below, but a known transformation method to a frequency domain other than the complex DFT may also be used. Further, when transforming a sampling string whose number of samplings is not a power of 2 into the frequency domain, a known technique such as using a sampling string that has been zero-padded so that the number of samplings becomes a power of 2 may be used.
[0044] The stereo encoding unit 110 first performs a complex DFT on the input sound signals of the two channels respectively to obtain a complex DFT coefficient sequence (step S111-1). The complex DFT coefficient sequence is obtained by applying a window with overlap between frames and using a process that takes into account the symmetry of the complex numbers obtained by the complex DFT. For example, when the sampling frequency is 32 kHz, whenever the sound signals of the two channels sampled as a 20-ms quantity of every 640 samples are input, processing is performed. For each channel, a continuous 744-point digital sound signal sampling string (the sampling string in the interval from t1 to t6 in the above example) including 104 points of sampling overlapping with the last sampling group of the immediately preceding frame (the sampling in the interval from t1 to t2 in the above example) and 104 points of sampling overlapping with the first sampling group of the immediately following frame (the sampling in the interval from t5 to t6 in the above example) is subjected to a complex DFT, and the first half of the 372 complex sequences among the 744 complex sequences obtained can be used as the complex DFT sequence. After that, for each integer f from 1 to 372, each complex DFT coefficient of the complex DFT coefficient sequence of the first channel is V1(f), and each complex DFT coefficient of the complex DFT coefficient sequence of the second channel is V2(f). Next, the stereo encoding unit 110 obtains a sequence of values of the radius on the complex plane based on each complex DFT coefficient from the complex DFT coefficient sequences of the two channels (step S111-2). The value of the radius on the complex plane of each complex DFT coefficient of each channel corresponds to the intensity of each frequency point of the sound signal of each channel. After that, the value of the radius on the complex plane of the complex DFT coefficient V1(f) of the first channel is set as V1r(f), and the value of the radius on the complex plane of the complex DFT coefficient V2(f) of the second channel is set as V2r(f). Next, the stereo encoding unit 110 obtains the average value of the ratio of the value of the radius of one channel to the value of the radius of the other channel for each frequency band, and obtains a sequence based on the average value as a characteristic parameter (step S111-3). This sequence of average values is a characteristic parameter corresponding to the information indicating the intensity difference of each frequency band of the input sound signals of the two channels. For example, if it is set to 4 frequency bands, for each of the 4 frequency bands of f from 1 to 93, 94 to 186, 187 to 279, and 280 to 372, the average value Mr(1), Mr(2), Mr(3), Mr(4) of the 93 values obtained by dividing the value of the radius V1r(f) of the first channel by the value of the radius V2r(f) of the second channel is obtained, and a sequence {Mr(1), Mr(2), Mr(3), Mr(4)} based on the average value is obtained as a characteristic parameter.
[0045] In addition, the number of frequency bands may be any value less than or equal to the number of frequency points. The same value as the number of frequency points can be used as the number of frequency bands, or 1 can be used. When the same value as the number of frequency points is used as the number of frequency bands, the stereo encoding unit 110 only needs to obtain the ratio of the radius value of one channel to the radius value of the other channel for each frequency point, and obtain a sequence based on the obtained ratio value as a feature parameter. When 1 is used as the number of frequency bands, the stereo encoding unit 110 only needs to obtain the ratio of the radius value of one channel to the radius value of the other channel for each frequency point, and obtain the average value of the obtained ratio values for all frequency bands as a feature parameter. In addition, when the number of frequency bands is set to multiple, the frequency count included in each frequency band can be arbitrary. For example, the number of frequency points included in a frequency band with a lower frequency can be set to be less than the number of frequency points included in a frequency band with a higher frequency.
[0046] In addition, the stereo encoding unit 110 may also use the difference between the radius value of one channel and the radius value of the other channel instead of the ratio of the radius value of one channel to the radius value of the other channel. That is, in the above example, the value obtained by subtracting the radius value V2r(f) of the second channel from the radius value V1r(f) of the first channel can be used instead of the value obtained by dividing the radius value V1r(f) of the first channel by the radius value V2r(f) of the second channel.
[0047] The stereo encoding unit 110 further obtains a code representing the feature parameter, i.e., the stereo code CS (step S111-4). The code representing the feature parameter, i.e., the stereo code CS, can be obtained by a well-known method. For example, the stereo encoding unit 110 performs vector quantization on the sequence of values obtained in step S111-3 to obtain a code, and outputs the obtained code as the stereo code CS. Or, for example, the stereo encoding unit 110 performs scalar quantization on each value included in the sequence of values obtained in step S111-3 to obtain a code, and outputs the code obtained by combining the obtained codes as the stereo code CS. In addition, when only one value is obtained in step S111-3, the stereo encoding unit 110 only needs to output the code obtained by performing scalar quantization on the one value as the stereo code CS.
[0048] The stereo encoding unit 110 also obtains a signal obtained by mixing the sound signals of the first channel and the second channel, that is, a downmix signal (step S111-5). For example, in the processing of the current frame, for 20 ms from t3 to t7, the stereo encoding unit 110 may obtain a mono signal obtained by mixing the sound signals of the two channels, that is, a downmix signal. The stereo encoding unit 110 may mix the sound signals of the two channels in the time domain as in step S111-5A described later, or may mix the sound signals of the two channels in the frequency domain as in step S111-5B described later. When mixing in the time domain, for example, the stereo encoding unit 110 obtains a sequence obtained by averaging the corresponding samples between the sample string of the sound signal of the first channel and the sample string of the sound signal of the second channel as a mono signal obtained by mixing the sound signals of the two channels, that is, a downmix signal (step S111-5A). When mixing in the frequency domain, for example, the stereo encoding unit 110 obtains the average value VMr(f) of the radii and the average value VMθ(f) of the angles of each complex DFT coefficient of the complex DFT coefficient sequence obtained by performing complex DFT on the sample string of the sound signal of the first channel and each complex DFT coefficient of the complex DFT coefficient sequence obtained by performing complex DFT on the sample string of the sound signal of the second channel, and performs inverse complex DFT on the sequence of complex numbers VM(f) with a radius of VMr(f) and an angle of VMθ(f) on the complex plane to obtain a sample string, and obtains a downmix signal as a signal obtained by mixing the sound signals of the two channels into a mono signal (steps S111B to 5B).
[0049] In addition, as Figure 1 indicated by the dash-dot line, a downmixing unit 150 may also be included in the encoding device 100 so that the process of step S111-5 of obtaining the downmix signal is performed not in the stereo encoding unit 110 but in the downmixing unit 150. In this case, the stereo encoding unit 110 obtains and outputs a stereo code CS (step S111), and the stereo code CS represents a characteristic parameter, which is a characteristic of the difference between the sound signals of the two channels into which the two-channel stereo input sound signal input to the encoding device 100 is input (step S111). The downmixing unit 150 obtains and outputs a signal obtained by mixing the sound signals of the two channels, that is, a downmix signal, from the two-channel stereo input sound signal input to the encoding device 100 (step S151). That is, it may also be that the stereo encoding unit 110 performs the above steps S111-1 to S111-4 as step S111, and the downmixing unit 150 performs the above step S111-5 as step S151.
[0050] [Mono Encoding Unit 120]
[0051] The downmixed signal output from the stereo encoding unit 110 is input to the mono encoding unit 120. When the encoding device 100 includes a downmixing unit 150, the downmixed signal output from the downmixing unit 150 is input to the mono encoding unit 120. The mono encoding unit 120 encodes the downmixed signal using a prescribed encoding method to obtain and output a mono code CM (step S121). As the encoding method, for example, an encoding method including processing of applying a window with overlap between frames, such as the 13.2 kbps mode of the 3GPP EVS standard (3GPP TS26.445) of Non-Patent Document 2, is used. In the case of the above example, in the processing of the current frame, the mono encoding unit 120 uses, as the signal 1b in the interval from t1 to t6 obtained by applying a window with a shape that increases in the interval from t1 to t2 where the current frame overlaps with the immediately preceding frame, a shape that decays in the interval from t5 to t6 where the current frame overlaps with the immediately following frame, and a flat shape in the interval from t2 to t5 between these intervals, the signal 1a as the downmixed signal. Also, the interval from t6 to t7 of the signal 1a in the "pread read interval" is used for analysis processing, encoded, and the mono code CM is obtained and output.
[0052] Thus, when the encoding method used in the mono encoding unit 120 includes processing of applying an overlapping window and analysis processing using a "pread read interval", not only the downmixed signal output in the processing of the current frame by the stereo encoding unit 110 or the downmixing unit 150 is used for the encoding process, but also the downmixed signal output in the processing of past frames by the stereo encoding unit 110 or the downmixing unit 150 is used for the encoding process. Therefore, a storage unit (not shown) is provided in the mono encoding unit 120, and the downmixed signal input in the processing of past frames is stored in advance in the storage unit, and the mono encoding unit 120 can use the downmixed signal stored in the storage unit to perform the encoding process of the current frame. Alternatively, a storage unit (not shown) may be provided in the stereo encoding unit 110 or the downmixing unit 150, and the stereo encoding unit 110 or the downmixing unit 150 outputs, in the processing of the current frame, the downmixed signal used in the encoding process of the current frame by the mono encoding unit 120, including the downmixed signal obtained in the processing of past frames. The mono encoding unit 120 uses the downmixed signal input from the stereo encoding unit 110 or the downmixing unit 150 in the processing of the current frame. In addition, as needed, in each of the units described later, signals obtained in the processing of past frames are stored in a storage unit (not shown) and used in the processing of the current frame, but such processing is well-known in the field of encoding, and thus, for the sake of avoiding redundancy, the description will be omitted hereafter.
[0053] [Additional Encoding Unit 130]
[0054] The downmixed signal output from the stereo encoding unit 110 is input to the additional encoding unit 130. When the encoding device 100 includes a downmixing unit 150, the downmixed signal output from the downmixing unit 150 is input to the additional encoding unit 130. The additional encoding unit 130 encodes the downmixed signal in the section X of the input downmixed signal to obtain an additional code CA and outputs it (step S131). In the above example, the additional encoding unit 130 encodes the signal 5c, which is the downmixed signal in the section from t5 to t6, to obtain an additional code CA and outputs it. Encoding can use a known encoding method such as scalar quantization or vector quantization.
[0055] <<Decoding device 200>>
[0056] As Figure 3 shown, the decoding device 200 of the first embodiment includes a monaural decoding unit 210, an additional decoding unit 230, and a stereo decoding unit 220. The decoding device 200 decodes the input monaural code CM, additional code CA, and stereo code CS in units of frames of the same specified time length as the encoding device 100 to obtain a two-channel stereo time-domain sound signal (two-channel stereo decoded sound signal) and outputs it. The codes input to the decoding device 200, that is, the monaural code CM, additional code CA, and stereo code CS, are output from the encoding device 100. Each time the decoding device 200 is input with the monaural code CM, additional code CA, and stereo code CS at each frame unit, that is, at intervals of the above-specified time length, it performs Figure 4 the processes of step S211, step S221, and step S231 exemplified. In the above example, at the time point t7 when 20 ms has elapsed since t3 when the processing for the immediately preceding frame was performed, if the decoding device 200 is input with the monaural code CM, additional code CA, and stereo code CS of the current frame, it performs the processes of step S211, step S221, and step S231 for the current frame. In addition, as Figure 3 shown by the dashed line, the decoding device 200 also outputs a monaural decoded sound signal, which is a time-domain sound signal of a monaural, if necessary.
[0057] [Monaural decoding unit 210]
[0058] The mono decoding unit 210 is input with the mono code CM included in the code input to the decoding device 200. The mono decoding unit 210 uses the input mono code CM to obtain a mono decoded audio signal for the interval Y and outputs it (step S211). As a prescribed decoding method, a decoding method corresponding to the encoding method used by the mono encoding unit 120 of the encoding device 100 is used. In the above example, the mono decoding unit 210 decodes the mono code CM of the current frame in a prescribed decoding method to obtain a shape that increases in the 3.25 ms interval from t1 to t2, is flat in the 16.75 ms interval from t2 to t5, and decays in the 3.25 ms interval from t5 to t6. For the interval from t1 to t6 of the application window of the signal 2a, for the interval from t1 to t2, in the processing of the immediately preceding frame, the signal 2b obtained by synthesizing the mono code CM of the immediately preceding frame and the signal 2a obtained from the mono code CM of the current frame are combined. For the interval from t2 to t5, the signal 2a obtained directly from the current mono code CM is used to obtain the mono decoded signal for the 20 ms interval from t1 to t5, which is the signal 2d. In addition, the signal 2a in the interval from t5 to t6 obtained from the mono code CM of the current frame is used as the "signal 2b obtained in the immediately preceding processing" in the processing of the immediately following frame. Therefore, the mono decoding unit 210 stores the signal 2a in the interval from t5 to t6 obtained from the mono code CM of the current frame in a storage unit (not shown) within the mono decoding unit 210.
[0059] [Additional decoding unit 230]
[0060] The additional decoding unit 230 is input with the additional code CA included in the code input to the decoding device 200. The additional decoding unit 230 decodes the additional code CA to obtain an additional decoded audio signal for the interval X, which is the additional decoded signal, and outputs it (step S231). In the decoding, a decoding method corresponding to the encoding method used by the additional encoding unit 130 is used. In the above example, the additional decoding unit 230 decodes the additional code CA of the current frame to obtain the mono decoded audio signal in the 3.25 ms interval from t5 to t6, which is the signal 4b, and outputs it.
[0061] [Stereo decoding unit 220]
[0062] The mono decoded audio signal output from the mono decoding unit 210, the additional decoded signal output from the additional decoding unit 230, and the stereo code CS included in the code input to the decoding apparatus 200 are input to the stereo decoding unit 220. The stereo decoding unit 220 obtains two-channel decoded audio signals, i.e., stereo decoded audio signals, based on the input mono decoded audio signal, additional decoded signal, and stereo code CS, and outputs them (step S221). More specifically, the stereo decoding unit 220 obtains a decoded downmix signal for the interval Y+X (i.e., the interval obtained by connecting interval Y and interval X) (step S221-1), where the interval Y+X is obtained by connecting the mono decoded audio signal in interval Y and the additional decoded signal in interval X. Through upmix processing using the characteristic parameters obtained from the stereo code CS, two-channel decoded audio signals are obtained from the decoded downmix signal obtained in step S221-1 and output (step S221-2). Although the same applies to each of the embodiments described later, the upmix processing refers to the following processing: regarding the downmix signal as a signal obtained by mixing two-channel decoded audio signals, and regarding the characteristic parameters obtained from the stereo code CS as information representing the differential characteristics of the two-channel decoded audio signals, to obtain two-channel decoded audio signals. In the above example, first, the stereo decoding unit 220 connects the mono decoded audio signal in the 20 ms interval from t1 to t5 output from the mono decoding unit 210 (the interval from t1 to t5 of signal 2d and signal 3a) and the additional decoded signal in the 3.25 ms interval from t5 to t6 output from the additional decoding unit 230 (signal 4b), to obtain a decoded downmix signal in the 23.25 ms interval from t1 to t6 (the interval from t1 to t6 of signal 4c). Next, the stereo decoding unit 220 regards the decoded downmix signal in the interval from t1 to t6 as a signal obtained by mixing two-channel decoded audio signals, regards the characteristic parameters obtained from the stereo code CS as information representing the differential characteristics of the two-channel decoded audio signals, and obtains two-channel decoded audio signals in the 20 ms interval from t1 to t5 (signal 4h-1 and signal 4h-2) and outputs them.
[0063] [Example of Step S221-2 Performed by Stereo Decoding Unit 220]
[0064] As an example of step S221-2 performed by the stereo decoding unit 220, step S221-2 performed by the stereo decoding unit 220 in the case where the characteristic parameter is information indicating the intensity difference of each frequency band of the sound signals of two channels will be described. The stereo decoding unit 220 first decodes the input stereo code CS to obtain information indicating the intensity difference of each frequency band (S221-21). The stereo decoding unit 220 obtains the characteristic parameter from the stereo code CS in a manner corresponding to the way in which the stereo encoding unit 110 of the encoding device 100 obtains the stereo code CS from the information indicating the intensity difference of each frequency band. For example, the stereo decoding unit 220 performs vector decoding on the input stereo code CS, and obtains the respective element values of the vector corresponding to the input stereo code CS as information indicating the intensity difference of each of a plurality of frequency bands. Or, for example, the stereo decoding unit 220 performs scalar decoding on each code included in the input stereo code CS to obtain information indicating the intensity difference of each frequency band. Further, in the case where the number of frequency bands is 1, the stereo decoding unit 220 performs scalar decoding on the input stereo code CS to obtain information indicating the intensity difference of 1 frequency band, that is, all frequency bands.
[0065] Next, the stereo decoding unit 220 regards the decoded downmix signal obtained in step S221-1 and the characteristic parameter obtained in step S221-21 as a signal obtained by mixing the decoded sound signals of two channels, and regards the characteristic parameter as information indicating the intensity difference of each frequency band of the decoded sound signals of two channels, and obtains and outputs the decoded sound signals of two channels (step S220-22). If the stereo encoding unit 110 of the encoding device 100 performs the operation of the above specific example using the complex DFT, step S221-22 of the stereo decoding unit 220 becomes the following operation.
[0066] The stereo decoding unit 220 first obtains the following signal 4d (step S221-221). The signal 4d is a decoded downmix signal of 744 samples for a 23.25 ms interval from t1 to t6, with an increasing shape in the 3.25 ms interval from t1 to t2, flat in the 16.75 ms interval from t2 to t5, and a decaying shape in the 3.25 ms interval from t5 to t6, which is a windowed signal. Next, the stereo decoding unit 220 obtains a sequence of the first half of 372 complex numbers from the sequence of 744 complex numbers obtained by performing a complex DFT on the signal 4d as the complex DFT coefficient sequence (mono complex DFT coefficient sequence) (step S221-222). After that, each complex DFT coefficient of the mono complex DFT coefficient sequence obtained by the stereo decoding unit 220 is set as MQ(f). Next, the stereo decoding unit 220 obtains the value of the radius on the complex plane of each complex DFT coefficient, MQr(f), and the value of the angle on the complex plane of each complex DFT coefficient, MQθ(f), according to the mono complex DFT coefficient sequence (step S221-223). Next, the stereo decoding unit 220 obtains the value obtained by multiplying the value of each radius MQr(f) by the square root of the corresponding value in the characteristic parameter as the value of each radius VLQr(f) of the first channel, and obtains the value obtained by dividing the value of each radius MQr(f) by the square root of the corresponding value in the characteristic parameter as the value of each radius VRQr(f) of the second channel (step S221-224). If the corresponding value in the characteristic parameter for each frequency point is an example of the above four frequency bands, then for f from 1 to 93 it is Mr(1), for f from 94 to 186 it is Mr(2), for f from 187 to 279 it is Mr(3), and for f from 280 to 372 it is Mr(4). Additionally, when the stereo encoding unit 110 of the encoding device 100 uses the difference between the radius value of the first channel and the radius value of the second channel instead of the ratio of the radius value of the first channel to the radius value of the second channel, the stereo decoding unit 220 can obtain the value obtained by adding the value obtained by dividing the corresponding value in the characteristic parameter by 2 to each radius value MQr(f) as the value of each radius VLQr(f) of the first channel, and obtain the value obtained by subtracting the value obtained by dividing the corresponding value in the characteristic parameter by 2 from each radius value MQr(f) as the value of each radius VRQr(f) of the second channel.Next, the stereo decoding unit 220 obtains a decoded sound signal (signal 4e-1) of the first channel that is the application window of 744 samples in the 23.25 ms interval from t1 to t6 by performing an inverse complex DFT on the sequence formed by the complex numbers with a radius of VLQr(f) and an angle of MQθ(f) on the complex plane, and obtains a sound signal (signal 4e-2) of the application window of the second channel that is 744 samples in the 23.25 ms interval from t1 to t6 by performing an inverse complex DFT on the sequence formed by the complex numbers with a radius of VRQr(f) and an angle of MQθ(f) on the complex plane (steps S221-225). The decoded sound signals (signal 4e-1 and signal 4e-2) that capture the windows of each channel obtained in steps S221-225 are signals with an increasing shape in the 3.25 ms interval from t1 to t2, a flat shape in the 16.75 ms interval from t2 to t5, and a decaying shape in the 3.25 ms interval from t5 to t6. Next, for the first channel and the second channel, the stereo decoding unit 220 synthesizes the signals (signal 4f-1, signal 4f-2) obtained in the immediately preceding steps S221-225 and the signals (signal 4e-1, signal 4e-2) obtained in the current frame's steps S221-225 for the interval from t1 to t2, and directly uses the signals (signal 4e-1, signal 4e-2) obtained in the current frame's steps S221-225 for the interval from t2 to t5, thereby obtaining decoded sound signals (signal 4h-1, signal 4h-2) in the 20 ms interval from t1 to t5 and outputting them (steps S221-226).
[0067] <Second Embodiment>
[0068] The difference between the downmixed signal of the interval Y and the locally decoded signal of the mono encoding can also be an object of encoding in the additional encoding unit 130. The interval Y is the time interval in which the complete mono decoded sound signal is obtained from the mono code CM in the mono decoding unit 210. This form will be described as the second embodiment, and the differences from the first embodiment will be explained.
[0069] [Mono Encoding Unit 120]
[0070] In addition to encoding the downmixed signal by a prescribed encoding method to obtain and output the monaural code CM, the monaural encoding unit 120 also obtains and outputs a signal obtained by decoding the monaural code CM, that is, a monaural partial decoding signal which is a partial decoding signal of the interval Y as the downmixed signal (step S122). In the above example, in addition to obtaining the monaural code CM of the current frame, the monaural encoding unit 120 also obtains a partial decoding signal corresponding to the monaural code CM of the current frame, that is, a partial decoding signal of an application window whose shape increases in the 3.25 ms interval from t1 to t2, is flat in the 16.75 ms interval from t2 to t5, and decays in the 3.25 ms interval from t5 to t6. For the interval from t1 to t2, the partial decoding signal corresponding to the monaural code CM of the immediately preceding frame and the partial decoding signal corresponding to the monaural code CM of the current frame are synthesized. For the interval from t2 to t5, the partial decoding signal corresponding to the monaural code CM of the current frame is directly used, thereby obtaining and outputting the partial decoding signal from t1 to t5. The partial decoding signal corresponding to the monaural code CM of the frame immediately preceding the interval from t1 to t2 uses the signal stored in a storage unit (not shown) within the monaural encoding unit 120. The signal in the interval from t5 to t6 in the partial decoding signal corresponding to the monaural code CM of the current frame is used as the "partial decoding signal corresponding to the monaural code CM of the immediately preceding frame" in the processing of the immediately succeeding frame. Therefore, the monaural encoding unit 120 stores the partial decoding signal in the interval from t5 to t6 obtained from the monaural code CM of the current frame in a storage unit (not shown) within the monaural encoding unit 120.
[0071] [Additional encoding unit 130]
[0072] In addition to the downmixed signal, as Figure 1As shown by the dashed line in the figure, the mono local decoding signal output from the mono encoding unit 120 is also input to the additional encoding unit 130. The additional encoding unit 130 not only encodes the downmix signal in the section X that is the encoding target of the additional encoding unit 130 in the first embodiment, but also encodes the difference signal between the downmix signal in the section Y and the mono local decoding signal (a signal formed by subtracting the sample values between the corresponding samples), and obtains and outputs an additional code CA (step S132). For example, the additional encoding unit 130 can respectively encode the downmix signal in the section X and the difference signal in the section Y to obtain codes, and use the code obtained by concatenating the obtained codes as the additional code CA. The same encoding method as the additional encoding unit 130 in the first embodiment can be used for encoding. In addition, for example, the additional encoding unit 130 can also encode the signal formed by concatenating the difference signal in the section Y and the downmix signal in the section X to obtain the additional code CA. Additionally, for example, as in [[Specific Example 1 of the Additional Encoding Unit 130]] below, the additional encoding unit 130 performs first additional encoding and second additional encoding. In the first additional encoding, the downmix signal in the section X is encoded to obtain a code (first additional code CA1). In the second additional encoding, the signal formed by concatenating the difference signal in the section Y (i.e., the quantization error signal of the mono encoding unit 120), the downmix signal in the section X, and the difference signal between the local decoding signal of the first additional encoding (i.e., the quantization error signal of the first additional encoding) is encoded to obtain a code (second additional code CA2). The combined encoding of the first additional code CA1 and the second additional code CA2 can be used as the additional code CA. According to [[Specific Example 1 of the Additional Encoding Unit 130]], the signal formed by concatenating the difference signal in the section Y and the signal with a smaller amplitude difference between the two sections compared to the downmix signal in the section X is used as the encoding target for the second additional encoding, and the downmix signal itself is used as the encoding target for the short time section for the first additional encoding. Therefore, high-efficiency encoding can be expected.
[0073] [[Specific Example 1 of the Additional Encoding Unit 130]]
[0074] The additional encoding unit 130 first encodes the input downmix signal of the interval X to obtain a first additional code CA1 (step S132-1, hereinafter also referred to as the "first additional code"), and obtains a local decoded signal of the interval X corresponding to the first additional code CA1, that is, a local decoded signal of the first additional code of the interval X (step S132-2). A known encoding method such as scalar quantization or vector quantization can be used for the first additional code. Next, the additional encoding unit 130 obtains a difference signal between the input downmix signal of the interval X and the local decoded signal of the interval X obtained in step S132-2 (a signal formed by subtracting the sampled values between corresponding samples) (step S132-3). The additional encoding unit 130 also obtains a difference signal between the downmix signal of the interval Y and the monaural local decoded signal (a signal formed by subtracting the sampled values between corresponding samples) (step S132-4). Next, the additional encoding unit 130 encodes the signal that combines the difference signal of the interval Y obtained in step S132-4 and the difference signal of the interval X obtained in step S132-3, thereby obtaining a second additional code CA2 (in step S132-5, hereinafter also referred to as "second additional encoding"). In the second additional encoding, an encoding method that encodes the sample string obtained by concatenating the sample string of the difference signal of the interval Y obtained in step S132-4 and the sample string of the difference signal of the interval X obtained in step S132-3 together is used. For example, an encoding method that utilizes prediction in the time domain or an encoding method that adapts to the amplitude deviation in the frequency domain is used. Next, the additional encoding unit 130 outputs the code obtained by combining the first additional code CA1 obtained in step S132-1 and the second additional code CA2 obtained in step S132-5 as the additional code CA (step S132-6).
[0075] In addition, the additional encoding unit 130 may also use the weighted difference signal as the object to be encoded instead of the above-mentioned difference signal. That is, the additional encoding unit 130 may also encode the weighted difference signal (a signal formed by weighted subtraction operation of the sample values between corresponding samples) between the downmixed signal in interval Y and the mono local decoded signal and the downmixed signal in interval X to obtain an additional code CA and output it. If it is [[Specific Example 1 of the Additional Encoding Unit 130]], then as the process of step S132-4, the additional encoding unit 130 only needs to obtain the weighted difference signal (a signal formed by weighted subtraction operation of the sample values between corresponding samples) between the downmixed signal in interval Y and the mono local decoded signal. Similarly, as the process of step S132-3 of [[Specific Example 1 of the Additional Encoding Unit 130]], the additional encoding unit 130 may also obtain the weighted difference signal (a signal formed by weighted subtraction operation of the sample values between corresponding samples) between the input downmixed signal in interval X and the local decoded signal in interval X obtained in step S132-2. In these cases, as long as the weights used to generate each weighted difference signal are encoded by a known encoding technique to obtain a code, and the obtained code (the code representing the weight) is included in the additional code CA. These cases are the same for each difference signal in each of the following embodiments, but using the weighted difference signal as the object to be encoded instead of the difference signal and encoding the weighting at this time are well-known in the field of encoding techniques. Therefore, in the following embodiments, in order to avoid redundancy, separate descriptions are omitted, and only the description of jointly recording the difference signal and the weighted difference signal using "or" and the description of jointly recording the subtraction operation and the weighted subtraction operation using "or" are given.
[0076] [Additional Decoding Unit 230]
[0077] The additional decoding unit 230 decodes the additional code CA, and not only outputs the additional decoded signal in interval X, which is the additional decoded signal obtained by the additional decoding unit 230 in the first embodiment, but also outputs the additional decoded signal in interval Y (step S232). The decoding method used is the decoding method corresponding to the encoding method used by the additional encoding unit 130 in step S132. That is, when the additional encoding unit 130 uses [[Specific Example 1 of the Additional Encoding Unit 130]] in step S132, the additional decoding unit 230 performs the following [[Specific Example 1 of the Additional Decoding Unit 230]] process.
[0078] [[Specific Example 1 of the Additional Decoding Unit 230]]
[0079] The additional decoding unit 230 first decodes the first additional code CA1 included in the additional code CA to obtain a first decoded signal for the interval X (step S232-1, hereinafter also referred to as "first additional decoding"). In the first additional decoding, the additional encoding unit 130 uses a decoding method corresponding to the encoding method used in the first additional code. In addition, the additional decoding unit 230 decodes the second additional code CA2 included in the additional code CA to obtain a second decoded signal for the interval Y and the interval X (step S232-2, hereinafter also referred to as "second additional decoding"). In the second additional decoding, a decoding method corresponding to the encoding method used by the additional encoding unit 130 in the second additional encoding is used, that is, a decoding method capable of obtaining a combined sampled string formed by concatenating the sampled string of the additional decoded signal from the code to the interval Y and the sampled string of the second decoded signal for the interval X. For example, a decoding method applying prediction in the time domain or a decoding method applying the deviation of the amplitude in the frequency domain is used. Next, the additional decoding unit 230 obtains the second decoded signal for the interval Y in the second decoded signal obtained in step S232-2 as the additional decoded signal for the interval Y, and obtains, as the additional decoded signal for the interval X, a signal obtained by adding the first decoded signal for the interval X obtained in step S232-1 and the second decoded signal for the interval X in the second decoded signal obtained in step S232-2 (a signal formed by adding the sampled values between the corresponding samples), and outputs the additional decoded signals for the interval Y and the interval X (step S232-4).
[0080] In addition, when the additional encoding unit 130 encodes not the difference signal but the weighted difference signal, the additional code CA also includes a code representing the weighting. Therefore, the additional decoding unit 230 decodes the code other than the code representing the weighting in the additional code CA in the above step S232 to obtain an additional decoded signal and outputs it, and decodes the code representing the weighting included in the additional code CA to obtain the weighting and outputs it. If it is [[specific example 1 of the additional decoding unit 230]], it is only necessary to decode the code representing the weighting of the interval X included in the additional code CA to obtain the weighting of the interval X. In step S232-4 of the additional decoding unit 230, the second decoded signal of the interval Y in the second decoded signal obtained in step S232-2 is obtained as the additional decoded signal of the interval Y, and the first decoded signal of the interval X obtained in step S232-1 and the second decoded signal of the interval X in the second decoded signal obtained in step S232-2 are subjected to weighted addition operation to obtain a signal (a signal composed of weighted addition operation of the sampled values between the corresponding samples) as the additional decoded signal of the interval X, and the additional decoded signals of the interval Y and the interval X and the weighting of the interval Y obtained by decoding the code representing the weighting of the interval Y included in the additional code CA are output. The same applies to the addition operation of the signals in each of the following embodiments, but in the technical field of encoding, weighted addition operation (generation of a weighted sum signal) is performed instead of addition (generation of a sum signal), and obtaining the weighting from the code at this time is well-known. Therefore, in the following embodiments, a separate description will not be redundantly omitted, and only a description of addition and weighted addition operation together using "or" and a description of the sum signal and the weighted sum signal together using "or" will be given.
[0081] [Stereo decoding unit 220]
[0082] The stereo decoding unit 220 performs the following steps S22-1 and step S222-2 (step S222). Instead of the step S221-1 performed by the stereo decoding unit 220 in the first embodiment, the stereo decoding unit 220 obtains, as the decoded downmix signal of the interval Y+X, a signal obtained by combining the sum signal of the monaural decoded sound signal of the interval Y and the additional decoded signal of the interval Y (a signal composed of addition operation of the sampled values between the corresponding samples) and the additional decoded signal of the interval X (step S22-1). Instead of using the decoded downmix signal obtained in step S221-1, the stereo decoding unit 220 uses the decoded downmix signal obtained in step S222-1, and obtains two-channel decoded sound signals from the decoded downmix signal obtained in step S222-1 through an upmix process using the characteristic parameters obtained from the stereo code CS and outputs them (step S22-2).
[0083] In addition, when the additional encoding unit 130 encodes a weighted difference signal instead of a difference signal, the stereo decoding unit 220 only needs to obtain, in step S222-1, a signal obtained from the weighted sum signal (a signal formed by weighted addition of the sampled values between the corresponding samples) of the mono decoded sound signal in the connection interval Y and the additional decoded signal in the interval Y and the additional decoded signal in the interval X as the decoded downmix signal in the interval Y+X. In the generation of the weighted sum signal (weighted addition of the sampled values between the corresponding samples) of the mono decoded sound signal in the interval Y and the additional decoded signal in the interval Y, the weight of the interval Y output by the additional decoding unit 230 can be used. The same applies to the addition operation of the signals in each of the embodiments described below. As described at the location of the description of the additional decoding unit 230, in the field of encoding, weighted addition operation (generation of a weighted sum signal) is also performed instead of addition operation (generation of a sum signal), and it is well-known to obtain the weight from the code. Therefore, in the embodiments described below, in order not to be redundant, separate descriptions are omitted, and only the description of the addition operation and the weighted addition operation together using "or" and the description of the sum signal and the weighted sum signal together using "or" are provided.
[0084] According to the second embodiment, in addition to the algorithm delay of stereo encoding / decoding not being greater than that of mono encoding / decoding, the decoded downmix signal used in stereo decoding can be made of higher quality than in the first embodiment, so the decoded sound signals of each channel obtained by stereo decoding can also be made of high quality. That is, in the second embodiment, the mono encoding process performed by the mono encoding unit 120 and the additional encoding process performed by the additional encoding unit 130 are used as encoding processes for encoding the downmix signal with high quality, the mono code CM and the additional code CA are obtained as codes for representing the downmix signal well, and the mono decoding process performed by the mono decoding unit 210 and the additional decoding process performed by the additional decoding unit 230 are used as decoding processes for obtaining a high-quality decoded downmix signal. The code amounts allocated to the mono code CM and the additional code CA can be arbitrarily determined according to the use. In addition to standard-quality mono encoding / decoding, when higher-quality stereo encoding / decoding is desired, more code amount can be allocated to the additional code CA. That is, from the perspective of stereo encoding / decoding, "mono code" and "additional code" are just convenient names. If the mono code CM and the additional code CA are respectively a part of the code representing the downmix signal, one of them can be called "first downmix code" and the other can be called "second downmix code". If it is assumed that more code amount is allocated to the additional code CA, the additional code CA can be called "downmix code", "downmix signal code", etc. The above situation is the same in the third embodiment or each of the embodiments based on the second embodiment described later.
[0085] <Third Embodiment>
[0086] In the stereo decoding unit 220, a decoded stereo sound signal of high quality can be obtained more easily using a decoded downmix signal corresponding to a downmix signal obtained by mixing sound signals of two channels in the frequency domain. In the mono encoding unit 120, a high-quality mono decoded sound signal can be obtained in the mono decoding unit 210 for one of the signals obtained by encoding a signal obtained by mixing sound signals of two channels in the time domain. In such a case, the stereo encoding unit 110 can mix the sound signals of two channels input to the encoding apparatus 100 in the frequency domain to obtain a downmix signal, the mono encoding unit 120 encodes the signal obtained by mixing the sound signals of two channels input to the encoding apparatus 100 in the time domain, and the additional encoding unit 130 also encodes the difference between the signal obtained by mixing the sound signals of two channels in the frequency domain and the signal obtained by mixing them in the time domain. This configuration is taken as the third embodiment, and will be described centering on the differences from the second embodiment.
[0087] [Stereo Encoding Unit 110]
[0088] The stereo encoding unit 110 performs the operations described in the first embodiment in the same manner as the stereo encoding unit 110 of the second embodiment. However, the process of obtaining a signal obtained by mixing sound signals of two channels, i.e., a downmix signal, is performed, for example, by mixing the sound signals of two channels in the frequency domain as in step S111-5B (step S113). That is, the stereo encoding unit 110 obtains a downmix signal obtained by mixing sound signals of two channels in the frequency domain. For example, in the processing of the current frame, the stereo encoding unit 110 only needs to obtain a mono signal, i.e., a downmix signal, obtained by mixing sound signals of two channels in the frequency domain for the period from t1 to t6. When the encoding apparatus 100 further includes a downmix unit 150, the stereo encoding unit 110 obtains a stereo code CS and outputs it (step S113). The stereo code CS represents characteristic parameters that are the characteristics of the difference between the sound signals of the two channels input according to the two-channel stereo input sound signal input to the encoding apparatus 100. The downmix unit 150 obtains a downmix signal and outputs it (step S153). The downmix signal is a signal obtained by mixing sound signals of two channels in the frequency domain according to the two-channel stereo input sound signal input to the encoding apparatus 100.
[0089] [Mono Encoding Target Signal Generation Unit 140]
[0090] As in Figure 1As shown by the center dash line, the encoding device 100 of the third embodiment further includes a monaural encoding target signal generation unit 140. In the monaural encoding target signal generation unit 140, a two-channel stereo input sound signal input to the encoding device 100 is input. The monaural encoding target signal generation unit 140 obtains a monaural encoding target signal that is a monaural signal by processing of mixing sound signals of two channels in the time domain from the input two-channel stereo input sound signal (step S143). For example, the monaural encoding target signal generation unit 140 obtains a sequence of averages between corresponding samples of a sampling string of a sound signal based on the first channel and a sampling string of a sound signal of the second channel, and obtains a monaural encoding target signal that is a signal obtained by mixing sound signals of two channels. That is, the monaural encoding target signal obtained by the monaural encoding target signal generation unit 140 is a signal obtained by mixing sound signals of two channels in the time domain. For example, in the processing of the current frame, the monaural encoding target signal generation unit 140 only needs to obtain a monaural signal that is a signal obtained by mixing sound signals of two channels in the time domain for 20 ms from t3 to t7.
[0091] [Monaural Encoding Unit 120]
[0092] In the monaural encoding unit 120, the monaural encoding target signal output by the monaural encoding target signal generation unit 140 is input to replace the down-mixed signal output by the stereo encoding unit 110 or the down-mixing unit 150. The monaural encoding unit 120 encodes the monaural encoding target signal to obtain a monaural code CM and outputs it (step S123). For example, in the processing of the current frame, for the monaural encoding target signal, the monaural encoding unit 120 applies a window to the shape that increases in the interval from t1 to t2 where the current frame overlaps with the immediately preceding frame, the shape that decays in the interval from t5 to t6 where the current frame overlaps with the immediately following frame, and the shape that is flat in the interval from t2 to t5 between these intervals, and obtains a signal in the interval from t1 to t6. The monaural encoding unit 120 also uses the interval from t6 to t7 of the monaural encoding target signal as a "look-ahead interval" for analysis processing to perform encoding, and obtains and outputs the monaural encoding CM.
[0093] [Additional Encoding Unit 130]
[0094] The additional encoding unit 130, similar to the additional encoding unit 130 of the second embodiment, encodes the difference signal or weighted difference signal (a signal formed by performing subtraction or weighted subtraction on the sampled values between corresponding samples) between the downmix signal in section Y and the mono local decoded signal and the downmix signal in section X to obtain and output an additional code CA (step S133). However, the downmix signal in section Y is a signal obtained by mixing the sound signals of two channels in the frequency domain, and the mono local decoded signal in section Y is a signal obtained by locally decoding a signal obtained by mixing the sound signals of two channels in the time domain.
[0095] In addition, the additional encoding unit 130 of the third embodiment, similar to the additional encoding unit 130 of the second embodiment, as described in [[Specific Example 1 of the Additional Encoding Unit 130]], performs a first additional encoding to obtain a first additional code CA1 by encoding the downmix signal in section X, and a second additional encoding to obtain a second additional code CA2 by encoding a signal that combines the difference signal or weighted difference signal in section Y, the downmix signal in section X, and the difference signal or weighted difference signal between the local decoded signal of the first additional code, and sets the code obtained by combining the first additional code CA1 and the second additional code CA2 as the additional code CA.
[0096] [Mono Decoding Unit 210]
[0097] The mono decoding unit 210, similar to the mono decoding unit 210 of the second embodiment, uses the mono code CM to obtain and output the mono decoded sound signal in section Y (step S213). However, the mono decoded sound signal obtained by the mono decoding unit 120 of the third embodiment is a decoded signal of a signal obtained by mixing the sound signals of two channels in the time domain.
[0098] [Additional Decoding Unit 230]
[0099] The additional decoding unit 230, similar to the additional decoding unit 230 of the second embodiment, decodes the additional code CA to obtain and output the additional decoded signals in sections Y and X (step S233). However, the additional decoded signal in section Y includes the difference between the signal obtained by mixing the sound signals of two channels in the time domain and the mono decoded sound signal, and the difference between the signal obtained by mixing the sound signals of two channels in the frequency domain and the signal obtained by mixing the sound signals of two channels in the time domain.
[0100] [Stereo Decoding Unit 220]
[0101] The stereo decoding unit 220 performs the following steps S223-1 and S223-2 (step S223). Similar to the additional decoding unit 230 of the second embodiment, the stereo decoding unit 220 obtains, as the decoded downmix signal for the interval Y+X, the sum signal or weighted sum signal (a signal formed by addition or weighted addition of the sampled values between corresponding samples) of the mono decoded audio signal for the interval Y and the additional decoded signal for the interval Y and the additional decoded signal for the interval X (step S223-1), and obtains and outputs the decoded audio signals for two channels from the decoded downmix signal obtained in step S223-1 through an upmix process using the characteristic parameters obtained from the stereo code CS (step S223-2). Among them, the sum signal for the interval Y includes: the mono decoded audio signal obtained by mono encoding / decoding the signal obtained by time-domain mixing of the audio signals for two channels; the difference between the signal obtained by time-domain mixing of the audio signals for two channels and the mono decoded audio signal; and the difference between the signal obtained by frequency-domain mixing of the audio signals for two channels and the signal obtained by time-domain mixing of the audio signals for two channels.
[0102] <Fourth Embodiment>
[0103] Regarding the interval X, in the mono encoding unit 120 and the mono decoding unit 210, although correct local decoded signals and decoded signals cannot be obtained without the signals and codes of the subsequent frame, incomplete local decoded signals and decoded signals are obtained even with only the signals and codes up to the current frame. Therefore, the first to third embodiments can also be modified as follows: for the interval X, instead of the downmix signal itself, the additional encoding unit 130 encodes the difference between the downmix signal and the mono local decoded signal obtained from the signals up to the current frame. This method will be described as the fourth embodiment.
[0104] <<Fourth Embodiment A>>
[0105] First, regarding the fourth embodiment, i.e., Fourth Embodiment A, which is obtained by modifying the second embodiment, the description will focus on the differences from the second embodiment.
[0106] [Mono Encoding Unit 120]
[0107] Similar to the mono encoding unit 120 of the second embodiment, the mono encoding unit 120 is input with the downmix signal output from the stereo encoding unit 110 or the downmix unit 150. The mono encoding unit 120 obtains the mono code CM obtained by encoding the downmix signal and the signal obtained by decoding the mono code CM up to the current frame, that is, the local decoded signal of the downmix signal in the section Y+X, namely the mono local decoded signal, and outputs it (step S124). More specifically, in addition to obtaining the mono code CM of the current frame, the mono encoding unit 120 also obtains the local decoded signal corresponding to the mono code CM of the current frame, that is, the local decoded signal of the application window whose shape increases in the 3.25 ms section from t1 to t2, is flat in the 16.75 ms section from t2 to t5, and decays in the 3.25 ms section from t5 to t6. Regarding the section from t1 to t2, the local decoded signal corresponding to the mono code CM of the immediately preceding frame and the local decoded signal corresponding to the mono code CM of the current frame are synthesized. Regarding the section from t2 to t6, the local decoded signal corresponding to the mono code CM of the current frame is directly used to obtain the local decoded signal from t1 to t2 and output it. However, the local decoded signal in the section from t5 to t6 is the local decoded signal that becomes a complete local decoded signal by being synthesized with the local decoded signal of the application window with an increasing shape obtained in the processing of the immediately following frame, and is an incomplete local decoded signal of the application window with a decaying shape.
[0108] [Additional encoding unit 130]
[0109] Similar to the additional encoding unit 130 of the second embodiment, the additional encoding unit 130 is input with the downmix signal output from the stereo encoding unit 110 or the downmix unit 150 and the mono local decoded signal output from the mono encoding unit 120. The additional encoding unit 130 encodes the difference signal or the weighted difference signal (a signal formed by subtraction or weighted subtraction of the sample values between corresponding samples) between the downmix signal in the section Y+X and the mono local decoded signal, and obtains the additional code CA and outputs it (step S134).
[0110] [Mono decoding unit 210]
[0111] Similar to the mono decoding unit 210 of the second embodiment, the mono decoding unit 210 is input with the mono code CM. The mono decoding unit 210 uses the mono code CM to obtain the mono decoded audio signal in the section Y+X and outputs it (step S214). However, the decoded signal in the section X, that is, the section from t5 to t6, is the decoded signal that becomes a complete decoded signal by being synthesized with the decoded signal of the application window with an increasing shape obtained in the processing of the immediately following frame, and is an incomplete decoded signal of the application window with a decaying shape.
[0112] [Additional decoding unit 230]
[0113] Similar to the additional decoding unit 230 of the second embodiment, the additional decoding unit 230 is input with the additional code CA. The additional decoding unit 230 decodes the additional code CA, obtains the additional decoding signal for the interval Y + X, and outputs it (step S234).
[0114] [Stereo decoding unit 220]
[0115] Similar to the stereo decoding unit 220 of the second embodiment, the mono decoding sound signal output by the mono decoding unit 210, the additional decoding signal output by the additional decoding unit 230, and the stereo code CS input to the decoding device 200 are input to the stereo decoding unit 220. The stereo decoding unit 220 obtains the sum signal or weighted sum signal (a signal formed by performing addition or weighted addition on the sampled values between corresponding samples) of the mono decoding sound signal and the additional decoding signal for the interval Y + X as the decoded downmix signal, and through the upmix processing using the characteristic parameters obtained from the stereo code CS, obtains the decoded sound signals for two channels from the decoded downmix signal and outputs them (step S224).
[0116] [[Fourth Embodiment B]]
[0117] In addition, if the "downmix signal output by the stereo encoding unit 110 or the downmixing unit 150" and "downmix signal" in the description of the fourth embodiment A are respectively replaced with the "mono encoding target signal output by the mono encoding target signal generation unit 140" and "mono encoding target signal", it becomes a description centered on the points different from the third embodiment of the fourth embodiment, that is, the third embodiment of the fourth embodiment B.
[0118] [[Fourth Embodiment C]]
[0119] Furthermore, if the mono partial decoding signal obtained by the mono encoding unit 120, the difference signal or weighted difference signal encoded by the additional encoding unit 130, and the additional decoding signal obtained by the additional decoding unit 230 in the description of the fourth embodiment are respectively set as the interval X, and the stereo decoding unit 220 obtains the signal obtained by combining the mono decoding sound signal for the interval Y and the signal obtained by adding the mono decoding sound signal for the interval X and the additional decoding signal or weighted sum signal as the decoded mixed signal, it becomes the fourth embodiment, that is, the fourth embodiment C, which modifies the first embodiment.
[0120] [[Fifth Embodiment]]
[0121] The downmixed signal in the interval X contains a part that can be predicted from the monaural local decoded signal according to the interval Y. Therefore, in each of the first to fourth embodiments, for the interval X, the additional encoding unit 130 may also encode the difference between the downmixed signal and the predicted signal of the monaural local decoded signal from the interval Y. This method will be described as the fifth embodiment.
[0122] <<Fifth Embodiment A>>
[0123] First, the fifth embodiment that modifies each of the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B is regarded as the fifth embodiment A, and the points different from each of the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B will be described.
[0124] [Additional Encoding Unit 130]
[0125] The additional encoding unit 130 performs the following steps S135A-1 and S135A-2 (step S135A). The additional encoding unit 130 first uses a prescribed well-known prediction technique to obtain a prediction signal for the interval X of the monaural local decoded signal from the input monaural local decoded signal of the interval Y or the interval Y+X (where, as described above, the monaural local decoded signal of the incomplete interval X). In addition, in the fifth embodiment that modifies the fourth embodiment A or the fifth embodiment that modifies the fourth embodiment B, the incomplete monaural local decoded signal of the input interval X is included in the prediction signal for the interval X. Next, the additional encoding unit 130 encodes the difference signal or weighted difference signal (a signal formed by performing subtraction or weighted subtraction on the sample values between corresponding samples) between the downmixed signal of the interval Y and the monaural local decoded signal, and the difference signal or weighted difference signal (a signal formed by performing subtraction or weighted subtraction on the sample values between corresponding samples) between the downmixed signal of the interval X and the prediction signal obtained in step S135A-1, to obtain an additional code CA and output it (step S135A-2). For example, the signal connecting the difference signal of the interval Y and the difference signal of the interval X may be encoded to obtain the additional code CA. In addition, for example, the difference signal of the interval Y and the difference signal of the interval X may be encoded separately to obtain codes, and the code obtained by connecting the obtained codes may be used as the additional code CA. The same encoding method as that of the additional encoding unit 130 in each of the second embodiment, the third embodiment, the fourth embodiment A, and the fourth embodiment B may be used in the encoding.
[0126] [Stereo Decoding Unit 220]
[0127] The stereo decoding unit 220 performs steps S225A-0 to S225A-2 (step S225A). First, the stereo decoding unit 220 uses the same prediction technique as the prediction technique used by the additional encoding unit 130 in step S135 to obtain a prediction signal for interval X from the mono decoded sound signal in interval Y or interval Y+X (step S225A-0). Next, the stereo decoding unit 220 obtains, as the decoded downmix signal for interval Y+X, a signal that combines the sum signal or weighted sum signal (a signal formed by performing addition or weighted addition on the sample values between corresponding samples) of the mono decoded sound signal in interval Y and the additional decoded signal, and the sum signal or weighted sum signal (a signal formed by performing addition or weighted addition on the sample values between corresponding samples) of the additional decoded signal and the prediction signal in interval X (step S225A-1). Next, the stereo decoding unit 220 obtains and outputs two-channel decoded sound signals from the decoded downmix signal obtained in step S225A-1 through an upmix process using the characteristic parameters obtained from the stereo code CS (step S225A-2).
[0128] <<Fifth Embodiment B>>
[0129] Next, the fifth embodiment, which modifies each of the first and fourth embodiments C, is designated as the fifth embodiment B, and the differences from each of the first and fourth embodiments C will be described.
[0130] [Mono Encoding Unit 120]
[0131] In addition to the mono code CM obtained by encoding the downmix signal, the mono encoding unit 120 also obtains and outputs a signal obtained by decoding the mono code CM for interval Y or interval Y+X up to the current frame, that is, a mono partial decoded signal, which is a partial decoded signal of the input downmix signal. However, as described above, the mono partial decoded signal for interval X is an incomplete mono partial decoded signal.
[0132] [Additional Encoding Unit 130]
[0133] The additional encoding unit 130 performs the following steps S135B-1 and S135B-2 (step S135B). The additional encoding unit 130 first uses a prescribed well-known prediction technique to obtain a prediction signal for the interval X of the mono local decoded signal from the input interval Y or the mono local decoded signal of the interval Y+X (where, as described above, the interval X is an incomplete mono local decoded signal) (step S135B-1). In the case of the fifth embodiment that modifies the fourth embodiment C, the prediction signal for the interval X includes the input mono local decoded signal of the interval X. The additional encoding unit 130 then encodes the difference signal or weighted difference signal (a signal formed by subtraction or weighted subtraction of the sample values between corresponding samples) between the downmixed signal of the interval X and the prediction signal obtained in step S135B-1 to obtain an additional code CA and outputs it (step S135B-2). For example, the same encoding method as that of the additional encoding unit 130 in each of the first embodiment and the fourth embodiment C can be used in the encoding.
[0134] [Stereo decoding unit 220]
[0135] The stereo decoding unit 220 performs step S225B-2 from the following step S225B-0 (step S225B). The stereo decoding unit 220 first uses the same prediction technique as the prediction technique used by the additional encoding unit 130 to obtain a prediction signal for the interval X from the mono decoded sound signal of the interval Y or the interval Y+X (step S225B-0). Next, the stereo decoding unit 220 obtains a signal obtained by connecting the mono decoded sound signal of the interval Y and the sum signal or weighted sum signal (a signal formed by addition or weighted addition of the sample values between corresponding samples) of the additional decoded signal and the prediction signal of the interval X as the decoded downmixed signal of the interval Y+X (step S225B-1). Next, the stereo decoding unit 220 obtains two-channel decoded sound signals from the decoded downmixed signal obtained in step S225B-1 through an upmixing process using the characteristic parameters obtained from the stereo code CS and outputs them (step S225B-2).
[0136] <Sixth Embodiment>
[0137] In the first to fifth embodiments, the decoding device 200 uses the additional code CA obtained by the encoding device 100 to at least decode the additional code CA to obtain the decoded downmix signal of the interval X used in the stereo decoding unit 220. However, the decoding device 200 may also not use the additional code CA, but use the predicted signal of the monaural decoded sound signal from the interval Y as the decoded downmix signal of the interval X used in the stereo decoding unit 220. This embodiment is taken as the sixth embodiment, and the differences from the first embodiment will be described.
[0138] <<Encoding device 100>>
[0139] The encoding device 100 of the sixth embodiment is different from the encoding device 100 of the first embodiment in that it does not include the additional encoding unit 130, does not encode the downmix signal of the interval X, and does not obtain the additional code CA. That is, the encoding device 100 of the sixth embodiment includes a stereo encoding unit 110 and a monaural encoding unit 120, and the stereo encoding unit 110 and the monaural encoding unit 120 perform the same operations as the stereo encoding unit 110 and the monaural encoding unit 120 of the first embodiment, respectively.
[0140] <<Decoding device 200>>
[0141] The decoding device 200 of the sixth embodiment does not include the additional decoding unit 230 for decoding the additional code CA, but includes a monaural decoding unit 210 and a stereo decoding unit 220. The monaural decoding unit 210 of the sixth embodiment performs the same operation as the monaural decoding unit 210 of the first embodiment, but also outputs the monaural decoded sound signal of the interval X when the monaural decoded sound signal of the interval Y+X is used in the stereo decoding unit 220. In addition, the stereo decoding unit 220 of the sixth embodiment performs the following operations different from those of the stereo decoding unit 220 of the first embodiment.
[0142] [Stereo decoding unit 220]
[0143] The stereo decoding unit 220 performs the following steps S226-0 to S226-2 (step S226). The stereo decoding unit 220 first obtains the predicted signal of the interval X according to the monaural decoded sound signal of the interval Y or the interval Y+X by using the same specified well-known prediction technique as in the fifth embodiment (step S226-0). Then, the stereo decoding unit 220 obtains the signal connecting the monaural decoded sound signal of the interval Y and the predicted signal of the interval X as the decoded downmix signal of the interval Y+X (step S226-1), and obtains and outputs the decoded sound signals of two channels from the decoded downmix signal obtained in step S226-1 through the upmix processing using the characteristic parameters obtained from the stereo code CS (step S226-2).
[0144] <Seventh Embodiment>
[0145] In the above-described embodiments, for simplicity of explanation, an example of processing sound signals of two channels has been described. However, the number of channels is not limited to this, and any number of 2 or more is acceptable. If the number of channels is set to C (where C is an integer of 2 or more), then each of the above-described embodiments can be implemented by replacing two channels with C channels (where C is an integer of 2 or more).
[0146] For example, the encoding device 100 of the first to fifth embodiments only needs to obtain the stereo code CS, the mono code CM, and the additional code CA based on the input sound signals of C channels, and the encoding device 100 of the sixth embodiment only needs to obtain the stereo code CS and the mono code CM based on the input sound signals of C channels. The stereo encoding unit 110 generates and outputs, as the stereo code CS, a code representing information of the difference between channels in the input sound signals corresponding to C channels. The stereo encoding unit 110 or the downmixing unit 150 outputs, as a downmixed signal, a signal obtained by mixing the input sound signals of C channels. The mono encoding target signal generation unit 140 only needs to output, as an encoding target signal, a signal obtained by mixing the input sound signals of C channels in the time domain. The information of the difference between channels in the sound signals corresponding to C channels is, for example, information corresponding to the difference between each of the C - 1 channels other than the reference channel and the sound signal of the reference channel and the sound signal of the reference channel.
[0147] Similarly, the decoding device 200 of the first to fifth embodiments only needs to output, as decoded sound signals of C channels, the decoded sound signals of C channels based on the input mono code CM, additional code CA, and stereo code CS, and the decoding device 200 of the sixth embodiment only needs to output, as decoded sound signals of C channels, the decoded sound signals of C channels based on the input mono code CM and stereo code CS. The stereo decoding unit 220 only needs to output, as decoded sound signals of C channels, the decoded sound signals of C channels obtained by upmixing the decoded downmixed signal using the characteristic parameters obtained based on the input stereo code CS. More specifically, the stereo decoding unit 220 can regard the decoded downmixed signal as a signal obtained by mixing the decoded sound signals of C channels, regard the characteristic parameters obtained based on the input stereo code CS as information representing the characteristics of the difference between channels in the decoded sound signals of C channels, and obtain and output the decoded sound signals of C channels.
[0148] <Program and Recording Medium>
[0149] The processing of each part of the above-described encoding apparatuses and decoding apparatuses can be implemented by a computer. In this case, the processing contents of the functions that each apparatus should have are described by a program. Then, this program is read into Figure 12 the storage unit 1020 of the computer shown in FIG. 1, and the arithmetic processing unit 1010, the input unit 1030, the output unit 1040, etc. are operated, whereby various processing functions in the above-described apparatuses are implemented on the computer.
[0150] The program describing this processing content can be recorded on a recording medium readable by a computer. A recording medium readable by a computer is, for example, a non-transitory recording medium. Specifically, it is a magnetic recording device, an optical disc, or the like.
[0151] In addition, the distribution of this program is performed, for example, by selling, transferring, lending a removable recording medium such as a DVD or a CD-ROM on which this program is recorded. Furthermore, it can also be configured to store this program in the storage device of a server computer, and transmit this program from the server computer to other computers via a network, thereby distributing this program.
[0152] A computer that executes such a program first temporarily stores, for example, the program recorded on a removable recording medium or the program transferred from a server computer in an auxiliary recording unit 1050 which is its own non-transitory storage device. And at the time of executing the processing, this computer reads the program stored in the auxiliary recording unit 1050 which is its own non-transitory storage device into the storage unit 1020, and executes the processing according to the read program. Additionally, as another execution mode of this program, the computer can directly read the program from a removable recording medium into the storage unit 1020 and execute the processing according to this program. Furthermore, it can also be configured to execute the processing according to the received program sequentially each time the program is transmitted from the server computer to this computer. Additionally, it can also be configured to execute the above-described processing by a so-called ASP (Application Service Provider) type service that realizes the processing function only through its execution instruction and result acquisition without transmitting the program from the server computer to this computer. Additionally, the program in this mode includes a program based on a program for processing by an electronic computer (data, etc. that are not direct instructions to the computer but have the nature of prescribing the processing of the computer).
[0153] In addition, in this mode, this apparatus is configured by executing a prescribed program on a computer, but at least a part of these processing contents can also be implemented by hardware.
[0154] Furthermore, of course, appropriate changes can be made without departing from the gist of the present invention.
Claims
1. A method for encoding an audio signal, which encodes the input audio signal of C channels in units of frames, where C is an integer greater than or equal to 2, and is characterized in that: As processing of the current frame, it includes: A stereo encoding step of obtaining and outputting a stereo code representing characteristic parameters, where the characteristic parameters are parameters representing the characteristics of the difference between channels of the audio signal of C channels; A downmixing step of obtaining a signal obtained by mixing the audio signals of C channels as a downmixed signal; A mono encoding step of encoding the downmixed signal to obtain a mono code and outputting it; In the mono encoding step, the downmixed signal is encoded to obtain the mono code in an encoding method including processing of applying a window with overlap between frames; It further includes an additional encoding step, in which the signal in the overlapping interval, that is, the signal in interval X, of the current frame and the immediately following frame in the downmixed signal is encoded to obtain an additional code and output it.
2. The method for encoding an audio signal according to claim 1, characterized in that: In the mono encoding step, a mono local decoded signal corresponding to the mono code is also obtained; In the additional encoding step, the signal in the interval other than interval X, that is, the signal in interval Y, of the downmixed signal, the signal formed by subtraction or weighted subtraction of the corresponding samples between the signal in interval Y and the mono local decoded signal, and the downmixed signal in interval X are encoded to obtain an additional code.
3. The method for encoding an audio signal according to claim 2, characterized in that: In the downmixing step, a signal obtained by mixing the audio signals of C channels in the frequency domain is obtained as the downmixed signal; It further includes a mono encoding target signal generation step, in which a signal obtained by mixing the audio signals of C channels in the time domain is obtained as a mono encoding target signal; The mono encoding step encodes the mono encoding target signal in an encoding method including processing of applying a window with overlap between frames to obtain the mono code.
4. The method for encoding an audio signal according to claim 2 or 3, characterized in that: In the additional encoding step, The signal in interval X of the downmixed signal is encoded to obtain a first additional code and a local decoded signal of interval X corresponding to the first additional code, that is, a first additional local decoded signal; The signal obtained by connecting the signal formed by subtraction or weighted subtraction of the corresponding samples between the downmixed signal in interval Y and the mono local decoded signal, and the signal formed by subtraction or weighted subtraction of the corresponding samples between the downmixed signal in interval X and the first additional local decoded signal is encoded in an encoding method of combining and encoding a sampling string to obtain a second additional code; The code obtained by adding the first additional code and the second additional code is used as the additional code.
5. The method for encoding an audio signal according to claim 2 or 3, characterized in that, In the additional encoding step, Based on the mono local decoded signal of interval Y or the mono local decoded signals of interval Y and interval X, a predicted signal of interval X of the mono local decoded signal is obtained, Encode a signal formed by a subtraction operation or a weighted subtraction operation of the sampled values between the corresponding samples of the downmixed signal of interval Y and the mono local decoded signal, and a signal formed by a subtraction operation or a weighted subtraction operation of the sampled values between the corresponding samples of the downmixed signal of interval X and the predicted signal, and obtain an additional code.
6. An audio signal encoding apparatus that encodes an input audio signal of C channels per frame, where C is an integer greater than or equal to 2, characterized in that, As processing of the current frame, it includes: A stereo encoding unit that obtains and outputs a stereo code representing characteristic parameters, where the characteristic parameters are parameters representing the characteristics of the difference between channels of the audio signal of C channels; A downmixing unit that obtains a signal obtained by mixing the audio signals of C channels as a downmixed signal; A mono encoding unit that encodes the downmixed signal to obtain a mono code and outputs it, The mono encoding unit encodes the downmixed signal in an encoding method including processing with a window having an overlap between frames to obtain the mono code, It further includes an additional encoding unit that encodes the overlapping interval of the current frame and the immediately subsequent frame in the downmixed signal, that is, the signal of interval X, to obtain an additional code and outputs it.
7. The audio signal encoding apparatus according to claim 6, characterized in that, The mono encoding unit also obtains a mono local decoded signal corresponding to the mono code, The additional encoding unit encodes a signal formed by a subtraction operation or weighted subtraction of the sampled values between the corresponding samples of the signal of interval Y, which is the interval other than interval X in the downmixed signal, and the mono local decoded signal of interval Y, and the downmixed signal of interval X, to obtain an additional code.
8. The audio signal encoding apparatus according to claim 7, characterized in that, The downmixing unit obtains a signal obtained by mixing the audio signals of C channels in the frequency domain as the downmixed signal, It further includes a mono encoding target signal generation unit that obtains a signal obtained by mixing the audio signals of C channels in the time domain as a mono encoding target signal, The mono encoding unit encodes the mono encoding target signal in an encoding method including processing with a window having an overlap between frames to obtain the mono code.
9. The audio signal encoding apparatus according to claim 7 or 8, characterized in that, The additional encoding unit performs: Encoding the signal of interval X in the downmixed signal to obtain a first additional code and a local decoded signal of interval X corresponding to the first additional code, that is, a first additional local decoded signal, A signal formed by a subtraction operation or a weighted subtraction operation of sampling values between corresponding samples of the downmixed signal in interval Y and the monaural local decoded signal, and a signal obtained by connecting signals formed by a subtraction operation or a weighted subtraction operation of sampling values between corresponding samples of the downmixed signal in interval X and the first additional local decoded signal are encoded using an encoding method that combines and encodes a string of samples to obtain a second additional code. The code obtained by adding the first additional code and the second additional code is used as the additional code.
10. The audio signal encoding apparatus according to claim 7 or 8, characterized in that The additional encoding unit performs: Based on the monaural local decoded signal in interval Y or the monaural local decoded signals in intervals Y and X, a prediction signal for interval X of the monaural local decoded signal is obtained. Encode a signal formed by a subtraction operation or a weighted subtraction operation of sampling values between corresponding samples of the downmixed signal in interval Y and the monaural local decoded signal, and a signal formed by a subtraction operation or a weighted subtraction operation of sampling values between corresponding samples of the downmixed signal in interval X and the prediction signal, and obtain an additional code.
11. A computer program product comprising a program for causing a computer to execute the steps of the audio signal encoding method according to any one of claims 1 to 5.
12. A computer-readable recording medium having recorded thereon a program for causing a computer to execute the steps of the audio signal encoding method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multichannel audio coder and decoder
CN102160113A
Stereo signal encoding device, stereo signal decoding device, stereo signal encoding method, and stereo signal decoding method
CN103180899A