Program for downmixing sound signals
By obtaining the leading relationship and correlation coefficient of the two-channel sound signal for weighted average, an efficient mono signal is generated, which solves the problem of low encoding efficiency in the prior art and realizes efficient mono signal encoding and stereo encoding.
Patent Information
- Application Number
- CN202510686435.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-09
- Filing Date
- 2020-11-04
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art fails to effectively generate mono signals useful for encoding processing from the two-channel sound signals, resulting in inefficient encoding.
By obtaining the leading relationship information and correlation coefficients of the left and right channels, weighted average processing is performed to generate a downmix signal, and combining mono encoding and stereo encoding to generate an efficient mono signal.
It realizes efficient encoding processing from two-channel sound signals to mono signals, improving encoding efficiency and sound quality.
Smart Images

Figure CN120496544A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of November 4, 2020, application number 202080098232.9, and invention name "Sound signal downmixing method, sound signal encoding method, sound signal downmixing device, sound signal encoding device and recording medium". Technical Field
[0002] The present invention relates to a technology for obtaining a monophonic sound signal from a two-channel sound signal in order to encode the sound signal in mono, encode the sound signal by a combination of mono encoding and stereo encoding, perform signal processing on the sound signal in mono, or perform signal processing on a stereo sound signal using a monophonic sound signal. Background Art
[0003] Patent Document 1 discloses a technique for deriving a monaural audio signal from a two-channel audio signal and performing embedded encoding / decoding of the two-channel audio signal and the monaural audio signal. Patent Document 1 discloses a technique in which the input left-channel audio signal and the input right-channel audio signal are averaged for each corresponding sample to obtain a monaural signal, the monaural signal is encoded (mono encoding) to obtain a monaural code, the monaural code is decoded (mono decoding) to obtain a monaural local decoded signal, and the difference (prediction residual signal) between the input audio signal and a prediction signal obtained from the monaural local decoded signal is encoded for each of the left and right channels. In the technology of Patent Document 1, for each channel, a signal obtained by delaying and assigning an amplitude ratio to a monaural local decoded signal is used as a prediction signal. A prediction signal having a delay and amplitude ratio that minimizes the error between the input sound signal and the prediction signal is selected, or a prediction signal having a delay difference and amplitude ratio that maximizes the cross-correlation between the input sound signal and the monaural local decoded signal is used. The prediction signal is subtracted from the input sound signal to obtain a prediction residual signal, which is then used as the target for encoding / decoding, thereby suppressing degradation in the sound quality of the decoded sound signal of each channel.
[0004] Prior art literature
[0005] Patent Literature
[0006] Patent Document 1: WO2006-070751 Summary of the Invention
[0007] Problems to be solved by the invention
[0008] The technology of Patent Document 1 improves the coding efficiency of each channel by optimizing the delay and amplitude ratio assigned to the monaural local decoded signal when obtaining the prediction signal. However, in the technology of Patent Document 1, the monaural local decoded signal is obtained by encoding / decoding a monaural signal obtained by averaging the left and right channel audio signals. In other words, the technology of Patent Document 1 lacks research on how to generate a monaural signal useful for signal processing such as encoding from a two-channel audio signal.
[0009] The present invention aims to provide a technique for obtaining a monaural signal useful for signal processing such as encoding from a two-channel audio signal.
[0010] Means for solving problems
[0011] A sound signal downmixing method according to one embodiment of the present invention obtains a signal obtained by mixing a left-channel input sound signal and a right-channel input sound signal, namely a downmixed signal, and is characterized in that it includes: a left-right relationship information acquisition step, obtaining preceding channel information and a left-right correlation coefficient, wherein the preceding channel information is information indicating which of the left-channel input sound signal and the right-channel input sound signal is preceding, and the left-right correlation coefficient is a correlation coefficient between the left-channel input sound signal and the right-channel input sound signal; and a downmixing step, performing weighted averaging of the left-channel input sound signal and the right-channel input sound signal based on the preceding channel information and the left-right correlation coefficient to obtain a downmixed signal, so that the larger the left-right correlation coefficient γ is, the more the input sound signal of the preceding channel between the left-channel input sound signal and the right-channel input sound signal is included.
[0012] In the audio signal downmixing method according to one embodiment of the present invention, it is characterized in that the sampling number is t, the left channel input audio signal is x L (t), the right channel input sound signal is x R (t), the downmixed signal is x M (t), the left and right correlation coefficient is γ, in the downmixing step, when the preceding channel information indicates that the left channel is the first, for each sample number t, by x M (t)=((1+γ) / 2)×x L (t)+((1-γ) / 2)×x R (t) The downmix signal is obtained. When the preceding channel information indicates that the right channel precedes, for each sample number t, the M (t)=((1-γ) / 2)×x L (t)+((1+γ) / 2)×x R (t) A downmix signal is obtained. When the preceding channel information indicates that no channel is preceded, for each sample number t, x M(t)=(x L (t)+x R (t)) / 2 to obtain the downmixed signal.
[0013] A sound signal downmixing method according to one embodiment of the present invention is characterized in that it includes the above-mentioned sound signal downmixing method as a sound signal downmixing step, and further includes: a mono encoding step of encoding the downmixed signal obtained in the downmixing step to obtain mono encoding; and a stereo encoding step of encoding the left channel input sound signal and the right channel input sound signal to obtain stereo encoding.
[0014] Effects of the Invention
[0015] According to the present invention, a monaural signal useful for signal processing such as encoding can be obtained from a two-channel audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a block diagram showing an example of an encoding device according to the first reference mode and the second embodiment.
[0017] Figure 2 This is a flowchart showing an example of processing by the encoding device of the first reference method.
[0018] Figure 3 This is a block diagram showing an example of a decoding device according to the first reference method.
[0019] Figure 4 This is a flowchart showing an example of processing by the decoding device according to the first reference method.
[0020] Figure 5 This is a flowchart showing an example of processing by the left-channel subtraction gain estimation unit and the right-channel subtraction gain estimation unit according to the first reference method.
[0021] Figure 6 This is a flowchart showing an example of processing by the left-channel subtraction gain estimation unit and the right-channel subtraction gain estimation unit according to the first reference method.
[0022] Figure 7 This is a flowchart showing an example of processing by the left channel subtraction gain decoding unit and the right channel subtraction gain decoding unit according to the first reference aspect.
[0023] Figure 8 This is a flowchart showing an example of processing by the left-channel subtraction gain estimation unit and the right-channel subtraction gain estimation unit according to the first reference method.
[0024] Figure 9 This is a flowchart showing an example of processing by the left-channel subtraction gain estimation unit and the right-channel subtraction gain estimation unit according to the first reference method.
[0025] Figure 10 This is a block diagram showing an example of an encoding device according to the second reference method and the first embodiment.
[0026] Figure 11 This is a flowchart showing an example of processing by the encoding device according to the second reference method.
[0027] Figure 12 This is a block diagram showing an example of a decoding device according to the second reference method.
[0028] Figure 13 This is a flowchart showing an example of processing by the decoding device according to the second reference method.
[0029] Figure 14 This is a flowchart showing an example of processing by the encoding device according to the first embodiment.
[0030] Figure 15 This is a flowchart showing an example of processing by the encoding device according to the second embodiment.
[0031] Figure 16 This is a block diagram showing an example of an encoding device according to the third embodiment.
[0032] Figure 17 This is a flowchart showing an example of processing by the encoding device according to the third embodiment.
[0033] Figure 18 This is a block diagram showing an example of a sound signal coding apparatus according to the fourth embodiment.
[0034] Figure 19 This is a flowchart showing an example of processing by the sound signal coding apparatus according to the fourth embodiment.
[0035] Figure 20 This is a block diagram showing an example of a sound signal processing device according to a fourth embodiment.
[0036] Figure 21 This is a flowchart showing an example of processing by the audio signal processing device according to the fourth embodiment.
[0037] Figure 22 This is a block diagram showing an example of a sound signal downmixing device according to a fourth embodiment.
[0038] Figure 23 This is a flowchart showing an example of processing of the audio signal downmixing device according to the fourth embodiment.
[0039] Figure 24 This is a diagram showing an example of the functional configuration of a computer that realizes each device in the embodiment of the present invention. DETAILED DESCRIPTION
[0040] <First embodiment>
[0041] First, let's explain how to write in a specification. A superscript "^" (like "^x") in a character "x" should normally be written directly above the "x." However, due to limitations on how specifications are written, it may be written as "^x."
[0042] <First Reference Method>
[0043] Before describing the embodiments of the invention, a coding device and a decoding device, each serving as a first reference mode and a second reference mode, will be described as the basis for implementing the second embodiment and the first embodiment. In the specification and claims, a coding device may be referred to as a sound signal coding device, a coding method may be referred to as a sound signal coding method, a decoding device may be referred to as a sound signal decoding device, and a decoding method may be referred to as a sound signal decoding method.
[0044] <<Encoding Device 100>>
[0045] like Figure 1 As shown, an encoding device 100 according to the first reference embodiment includes a downmix unit 110, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel signal subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, and a stereo encoding unit 170. Encoding device 100 encodes an input binaural time-domain audio signal in frames of a predetermined time length, for example, 20 ms, to obtain and output a monaural code CM, a left channel subtraction gain code Cα, a right channel subtraction gain code Cβ, and a stereo code CS, which will be described later. The binaural time-domain audio signal input to the encoding device is, for example, a digital audio signal or audio signal obtained by A / D conversion of sounds such as voices or music picked up by two microphones, respectively. The signal is composed of a left channel input audio signal and a right channel input audio signal. The codes outputted by the encoding device, namely the mono code CM, the left channel subtraction gain code Cα, the right channel subtraction gain code Cβ and the stereo code CS are inputted to the decoding device. The encoding device 100 performs the following steps on each frame: Figure 2 The processing of steps S110 to S170 is illustrated.
[0046] [Downmixing Unit 110]
[0047] The input audio signal of the left channel and the input audio signal of the right channel input to the encoding device 100 are input to the downmixing unit 110. The downmixing unit 110 obtains a downmixed signal, which is a signal obtained by mixing the left channel input audio signal and the right channel input audio signal, from the input audio signal of the left channel and the input audio signal of the right channel, and outputs the downmixed signal (step S110).
[0048] For example, when the number of samples per frame is T, the input audio signal x of the left channel input to the encoding device 100 in units of frames is input to the downmixing unit 110. L (1), x L (2), ..., x L (T) and the right channel input sound signal x R (1), x R (2), ..., x R (T). Here, T is a positive integer. For example, when the frame length is 20 ms and the sampling frequency is 32 kHz, T is 640. The downmixing unit 110 uses a sequence of average values of the sample values of each corresponding sample of the input left channel input sound signal and the input right channel input sound signal as the downmixed signal x. M (1), x M (2), ..., x M (T) is obtained and output. That is, if each sample number is t, then x M (t)=(x L (t)+x R (t)) / 2.
[0049] [Left Channel Subtraction Gain Estimation Unit 120]
[0050] The left channel subtraction gain estimation unit 120 receives the input audio signal x of the left channel input to the encoding device 100. L (1), x L (2), ..., x L (T) and the downmix signal x output by the downmix unit 110 M (1), x M (2), ..., x M (T). The left channel subtraction gain estimation unit 120 obtains and outputs the left channel subtraction gain α and the left channel subtraction gain code Cα representing the left channel subtraction gain α based on the input left channel audio signal and the downmix signal (step S120). The left channel subtraction gain estimation unit 120 obtains the left channel subtraction gain α and the left channel subtraction gain code Cα using a method for obtaining the amplitude ratio g, such as that described in Patent Document 1, a known method exemplified by the method for encoding the amplitude ratio g, or a method based on the newly proposed principle of minimizing quantization error. The principle of minimizing quantization error and the method based on this principle will be described later.
[0051] [Left Channel Signal Subtraction Unit 130]
[0052] The left channel signal subtraction unit 130 receives the left channel input audio signal x input to the encoding device 100. L (1), x L (2), ..., x L (T), the downmix signal x output by the downmixing unit 110 M (1), x M (2), ..., x M (T), the left channel subtraction gain α output by the left channel subtraction gain estimation unit 120. The left channel signal subtraction unit 130 calculates the sample value x of the input audio signal of the left channel for each corresponding sample t. L (t) Subtract the sample value x of the downmixed signal M (t) is multiplied by the left channel subtraction gain α × x M (t) and the value x L (t)-α×x M (t) is the left channel differential signal y L (1), y L (2), ..., y L (T) obtain and output (step S130). That is, y L (t)=x L (t)-α×x M In the encoding device 100, in order to avoid delay or computational processing required to obtain a local decoded signal, the left channel signal subtraction unit 130 uses the unquantized downmix signal x obtained by the downmix unit 110 instead of the monaural encoded local decoded signal, i.e., the quantized downmix signal. M (t) is sufficient. However, when the left channel subtraction gain estimation unit 120 obtains the left channel subtraction gain α by a well-known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing the quantization error, a unit for obtaining a local decoded signal corresponding to the monaural code CM is provided in the subsequent stage of the monaural code unit 160 of the encoding device 100 or in the monaural code unit 160, and in the left channel signal subtraction unit 130, the downmix signal x is replaced with M (1), x M (2), ..., x M (T) As in the conventional encoding apparatus such as Patent Document 1, a quantized downmix signal ^x which is a locally decoded signal for monaural encoding may be used. M (1), ^x M (2), ..., ^x M (T) to get the left channel differential signal.
[0053] [Right Channel Subtraction Gain Estimation Unit 140]
[0054] The right channel subtraction gain estimation unit 140 receives the input audio signal x of the right channel input to the encoding device 100. R (1), x R (2), ..., x R (T) and the downmix signal x output from the downmixing unit 110 M (1), x M (2), ..., x M (T). The right channel subtraction gain estimation unit 140 obtains and outputs the right channel subtraction gain β and the code representing the right channel subtraction gain β, namely, the right channel subtraction gain code Cβ, based on the input right channel input audio signal and the downmix signal (step S140). The right channel subtraction gain estimation unit 140 obtains the right channel subtraction gain β and the right channel subtraction gain code Cβ using a method for obtaining the amplitude ratio g, such as that described in Patent Document 1, a known method exemplified by the method for encoding the amplitude ratio g, or a method based on the newly proposed principle of minimizing quantization error. The principle of minimizing quantization error and the method based on this principle will be described later.
[0055] [Right Channel Signal Subtraction Unit 150]
[0056] The right channel signal subtraction unit 150 receives the right channel input audio signal x input to the encoding device 100. R (1), x R (2), ..., x R (T), the downmix signal x output by the downmixing unit 110 M (1), x M (2), ..., x M (T) and the right channel subtraction gain β output by the right channel subtraction gain estimation unit 140. The right channel signal subtraction unit 150 calculates the sampling value x of the input audio signal of the right channel for each corresponding sample t. R (t) Subtract the sample value x of the downmixed signal M (t) is multiplied by the right channel subtraction gain β, which is the value β×x M (t) and the value x R (t)-β×x M (t) as the right channel differential signal y R (1), y R (2), ..., y R (T) obtain and output (step S150). That is, y R (t)=x R (t)-β×x M(t). In the right channel signal subtraction unit 150, similarly to the left channel signal subtraction unit 130, in order to avoid the delay and computational complexity of obtaining the local decoded signal in the encoding device 100, the unquantized downmix signal x obtained by the downmix unit 110 is used instead of the quantized downmix signal of the monaural encoded local decoded signal. M (t) is sufficient. However, when the right channel subtraction gain estimation unit 140 obtains the right channel subtraction gain β by a well-known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing the quantization error, a unit for obtaining a local decoded signal corresponding to the mono code CM is provided in the subsequent stage of the mono encoding unit 160 of the encoding device 100 or in the mono encoding unit 160. Similarly to the left channel signal subtraction unit 130, in the right channel signal subtraction unit 150, the downmix signal x is replaced with M (1), x M (2), ..., x M (T) As in the conventional encoding apparatus such as Patent Document 1, a quantized downmix signal ^x which is a locally decoded signal of monaural encoding may be used. M (1),^x M (2), ..., ^x M (T) to get the right channel differential signal.
[0057] [Monaural encoding unit 160]
[0058] The downmix signal x output from the downmixing unit 110 is input to the monaural encoding unit 160. M (1), x M (2), ..., x M (T). The mono encoding unit 160 uses a predetermined encoding method to M The input downmix signal is encoded by 1 bit to obtain a mono coded CM and output (step S160). That is, the downmix signal x sampled from the input T is M (1), x M (2), ..., x M (T) gets b M The CM is encoded into a mono channel of 10 bits and outputted. Any encoding method may be used, and an encoding method such as the 3GPP EVS standard may also be used.
[0059] [Stereo encoding unit 170]
[0060] The left channel difference signal y output from the left channel signal subtraction unit 130 is input to the stereo encoding unit 170. L (1), y L (2), ..., y L(T), and the right channel difference signal y output by the right channel signal subtraction unit 150 R (1), y R (2), ..., y R (T). The stereo encoding unit 170 uses a predetermined encoding method to generate a total of b s The input left channel difference signal and right channel difference signal are encoded by 1 bit to obtain a stereo code CS and output (step S170). That is, the left channel difference signal y L (1), y L (2), ..., y L (T) and the input T sampled right channel differential signal y R (1), y R (2),..., y R (T) Get the total b s The stereo code CS of the bits is output. Any encoding method may be used, for example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, a method of encoding the input left channel difference signal and the right channel difference signal independently may be used, or a code obtained by combining all the encodings may be used as the stereo code CS.
[0061] When encoding the input left channel difference signal and right channel difference signal independently, the stereo encoding unit 170 uses b L The left channel differential signal is encoded with b R That is, the stereo encoding unit 170 encodes the right channel difference signal y from the input T sampled left channel difference signal y L (1), y L (2), ..., y L (T) gets b L The left channel differential code CL of the bit is obtained from the right channel differential signal y sampled from the input T R (1), y R (2), ...,y R (T) gets b R The right channel differential code CR of the bit is combined with the left channel differential code CL and the right channel differential code CR as the stereo code CS. L Bit and b R The total number of bits is b S bit.
[0062] When the input left channel difference signal and right channel difference signal are combined and encoded in one encoding method, the stereo encoding unit 170 uses the total bS That is, the stereo encoding unit 170 encodes the left channel difference signal y from the input T sampled left channel difference signal y L (1), y L (2), ..., y L (T) and the input T sampled right channel differential signal y R (1), y R (2), ..., y R (T) gets b S The stereo code CS of the bits is output.
[0063] <<Decoding Device 200>>
[0064] like Figure 3 As shown, the decoding device 200 of the first reference method includes: a mono decoding unit 210, a stereo decoding unit 220, a left channel subtraction gain decoding unit 230, a left channel signal addition unit 240, a right channel subtraction gain decoding unit 250, and a right channel signal addition unit 260. The decoding device 200 decodes the input mono code CM, the left channel subtraction gain code Cα, the right channel subtraction gain code Cβ, and the stereo code CS in frame units of the same time length as the corresponding encoding device 100, obtains and outputs a two-channel stereo time domain decoded sound signal (a left channel decoded sound signal and a right channel decoded sound signal to be described later) in frame units. The decoding device 200 can also be as follows Figure 3 As shown by the dotted line in the middle, a mono decoded audio signal in the time domain (a mono decoded audio signal described later) is output. For example, the decoded audio signal output by the decoding device 200 is reproduced by a speaker by performing DA conversion, so that the decoded audio signal can be received. The decoding device 200 performs D / A conversion on each frame. Figure 4 The processing of steps S210 to S260 is illustrated in FIG.
[0065] [Monaural decoding unit 210]
[0066] The monaural code CM input to the decoding device 200 is input to the monaural decoding unit 210. The monaural decoding unit 210 decodes the input monaural code CM using a predetermined decoding method to obtain a monaural decoded audio signal ^x M (1), ^x M (2), ..., ^x M (T) and output (step S210). As the predetermined decoding method, a decoding method corresponding to the encoding method used in the mono encoding unit 160 of the corresponding encoding device 100 is used. The number of bits of the mono encoded CM is b M .
[0067] [Stereo decoding unit 220]
[0068] The stereo code CS input to the decoding device 200 is input to the stereo decoding unit 220. The stereo decoding unit 220 decodes the input stereo code CS using a predetermined decoding method to obtain a left channel decoded differential signal y. L (1),^y L (2), ..., ^y L (T), right channel decoded differential signal ^y R (1), ^y R (2), ..., ^y R (T) and output (step S220). As the predetermined decoding method, a decoding method corresponding to the encoding method used in the stereo encoding unit 170 of the corresponding encoding device 100 is used. The total number of bits of the stereo encoding CS is b S .
[0069] [Left Channel Subtraction Gain Decoding Unit 230]
[0070] The left channel subtraction gain code Cα input to the decoding device 200 is input to the left channel subtraction gain decoding unit 230. The left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα and obtains the left channel subtraction gain α, which it outputs (step S230). The left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα using a decoding method corresponding to the method used by the left channel subtraction gain estimation unit 120 of the corresponding encoding device 100 to obtain the left channel subtraction gain α. The method by which the left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα and obtains the left channel subtraction gain α when the left channel subtraction gain estimation unit 120 of the corresponding encoding device 100 obtains the left channel subtraction gain α and the left channel subtraction gain code Cα using a method based on the principle of minimizing quantization error will be described later.
[0071] [Left channel signal adding unit 240]
[0072] The left channel signal adding section 240 receives the mono decoded audio signal φx outputted from the mono decoding section 210. M (1), ^x M (2), ..., ^x M (T), the left channel decoded differential signal y output by the stereo decoding unit 220 L (1), ^y L (2), ..., ^y L(T), the left channel subtraction gain α output by the left channel subtraction gain decoding unit 230. The left channel signal adding unit 240 converts the sample value of the left channel decoded differential signal α to the sample value of the left channel decoded differential signal α for each corresponding sample t. L (t), and the sampling value of the mono decoded audio signal ^x M (t) is multiplied by the left channel subtraction gain α×^x M (t) The added value ^y L (t)+α×^x M The sequence of (t) is used as the left channel decoded sound signal ^x L (1), ^x L (2), ..., ^x L (T) obtain and output (step S240). That is, it is ^x L (t)=^y L (t)+α×^x M (t).
[0073] [Right Channel Subtraction Gain Decoding Unit 250]
[0074] The right channel subtraction gain code Cβ input to the decoding device 200 is input to the right channel subtraction gain decoding unit 250. The right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ and obtains the right channel subtraction gain β, which it outputs (step S250). The right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ using a decoding method corresponding to the method used by the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 to obtain the right channel subtraction gain β. The method by which the right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ and obtains the right channel subtraction gain β when the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 obtains the right channel subtraction gain β and the right channel subtraction gain code Cβ using a method based on minimizing quantization error will be described later.
[0075] [Right Channel Signal Adding Unit 260]
[0076] The right channel signal adding section 260 receives the mono decoded audio signal φx outputted from the mono decoding section 210. M (1), ^x M (2), ..., ^x M (T), the right channel decoded differential signal y output by the stereo decoding unit 220 R (1), ^y R (2), ..., ^y R(T), the right channel subtraction gain β output by the right channel subtraction gain decoding unit 250. The right channel signal adding unit 260 adds the sample value of the right channel decoded difference signal to the value of the right channel subtraction gain β for each corresponding sample t. R (t), and the sampling value of the mono decoded audio signal ^x M (t) multiplied by the right channel subtraction gain β M (t) The value obtained by adding ^y R (t)+β×^x M The sequence of (t) is used as the right channel decoded sound signal ^x R (1), ^x R (2), ..., ^x R (T) is performed and output (step S260). R (t)=^y R (t)+β×^x M (t).
[0077] [Principle of minimizing quantization error]
[0078] The principle of minimizing the quantization error is described below. When the left channel difference signal and the right channel difference signal input to the stereo encoding unit 170 are combined and encoded using one encoding method, the number of bits b used in encoding the left channel difference signal is L The number of bits b used in encoding the right channel difference signal R It may not be clearly determined, but the following description assumes that the number of bits used in encoding the left channel differential signal is b. L , the number of bits used in encoding the right channel differential signal is b R In addition, the following description focuses on the left channel, but the same applies to the right channel.
[0079] The above-mentioned encoding device 100 uses b L Bit pair from the left channel input sound signal x L (1), x L (2), ..., x L Each sample value of (T) is subtracted from the downmix signal x M (1), x M (2), ..., x M The left channel differential signal y is composed of the value obtained by multiplying each sample value of (T) by the value of the left channel subtraction gain α. L (1), y L (2), ..., y L (T) is encoded and b M Bit-pair downmix signal x M (1), xM (2), ..., x M (T) is encoded. In addition, the decoding device 200 is based on b L The left channel decoded differential signal is y L (1), ^y L (2), ..., ^y L (hereinafter also referred to as "quantized left channel differential signal") is decoded according to b M Bit encoding for mono decoded audio signal^x M (1), ^x M (2), ..., ^x M After decoding (T) (hereinafter also referred to as "quantized downmix signal"), the quantized downmix signal obtained by decoding is converted into M (1),^x M (2), ..., ^x M The value obtained by multiplying each sample value of (T) by the left channel gain α is added to the quantized left channel difference signal y obtained by decoding. L (1), ^y L (2), ..., ^y L (T) each sampling value, thereby obtaining the left channel decoded audio signal ^x as the left channel decoded audio signal L (1), ^x L (2), ..., ^x L (T) The encoding apparatus 100 and the decoding apparatus 200 should be designed to reduce the energy of the quantization error of the decoded audio signal of the left channel obtained in the above-mentioned process.
[0080] The energy of the quantization error in the decoded signal obtained by encoding / decoding the input signal (hereinafter referred to as "quantization error caused by encoding") is generally proportional to the energy of the input signal and tends to decrease exponentially with the number of bits per sample used in encoding. Therefore, the average energy per sample of the quantization error caused by encoding the left channel difference signal is expressed as a positive number σ. L 2 It can be estimated as shown in the following equation (1-0-1): the average energy per sample of the quantization error generated by encoding the downmix signal is expressed as a positive number σ M 2 It can be estimated as shown in the following formula (1-0-2).
[0081] [Mathematical formula 1]
[0082]
[0083] [Mathematical formula 2]
[0084]
[0085] Here, let the input audio signal x of the left channel be L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T) is the value of each sample value that is approximately considered to be the same sequence. For example, the input audio signal x of the left channel L (1), x L (2), ..., x L (T) and the right channel input signal x R (1), x R (2), ..., x R (T) is equivalent to the case where the sound from the sound source at equal distances from the two microphones is collected in an environment with a lot of background noise or reverberation. Under this condition, the left channel differential signal y L (1), y L (2), ..., y L Each sample value of (T) is related to the downmix signal x M (1), x M (2), ..., x M The value obtained by multiplying each sample value of (T) by (1-α) is equivalent. Therefore, the energy of the left channel difference signal is (1-α) times the energy of the downmix signal. 2 times, so the above σ L 2 The above σ can be used M 2 Replace with (1-α) 2 ×σ M 2 , so the average energy per sample of the quantization error generated by encoding the left channel difference signal can be estimated as shown in the following formula (1-1).
[0086] [Mathematical formula 3]
[0087]
[0088] Furthermore, the average energy per sample of the quantization error of the signal to which the quantized left channel difference signal is added in the decoding device, that is, the average energy per sample of the quantization error of the sequence of values obtained by multiplying each sample value of the quantized downmix signal obtained by decoding by the left channel subtraction gain α, can be estimated as shown in the following formula (1-2).
[0089] [Formula 4]
[0090]
[0091] Assuming that the quantization error resulting from encoding the left channel difference signal and the quantization error of the sequence of values obtained by multiplying each sample value of the decoded quantized downmix signal by the left channel subtraction gain α are uncorrelated, the average energy per sample of the quantization error of the decoded left channel audio signal is estimated by the sum of equations (1-1) and (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error of the decoded left channel audio signal is determined as shown in equation (1-3).
[0092] [Formula 5]
[0093]
[0094] That is, the input sound signal x in the left channel L (1), x L (2), ..., x L (T) and downmix signal xx M (1), x M (2), ..., x M (T) Under the condition that the values of each sample value are approximately considered to be the same sequence, in order to minimize the quantization error of the decoded audio signal of the left channel, the left channel subtraction gain estimation unit 120 can calculate the left channel subtraction gain α by using formula (1-3). The left channel subtraction gain α obtained by formula (1-3) is a value greater than 0 and less than 1. When the number of bits used for two codes is b L with b M When they are equal, it is 0.5, which is the number of bits b used to encode the left channel differential signal. L greater than the number of bits b used to encode the downmix signal M The closer the value is to 0.5, the more bits b are used to encode the downmix signal. M Greater than the number of bits b used to encode the left channel differential signal L The closer it is to 0.5.
[0095] The same is true for the right channel. The input sound signal xR (1), x R (2), ..., x R (T), downmix signal x M (1), x M (2), ..., x M (T) Under the condition that the values of the sample values are approximately considered to be the same sequence, in order to minimize the quantization error of the decoded audio signal of the right channel, the right channel subtraction gain estimation unit 140 can calculate the right channel subtraction gain β using the following formula (1-3-2).
[0096] [Formula 6]
[0097]
[0098] The right channel subtraction gain β obtained by formula (1-3-2) is a value greater than 0 and less than 1. The number of bits used in the two codes is b R with b M When they are equal, it is 0.5, which is the number of bits b used to encode the right channel differential signal. R The more bits b used to encode the downmix signal M The closer the value is to 0.5, the number of bits b used to encode the downmix signal M The more than the number of bits b used to encode the right channel differential signal R The more it is, the closer it is to 0.5.
[0099] Next, the input audio signal x including the left channel L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T) The principle of minimizing the energy of the quantization error of the decoded audio signal of the left channel, including the case where it cannot be regarded as the same sequence, is explained.
[0100] The input sound signal of the left channel is x L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2),..., x M The normalized inner product value r of (T) L It is expressed by the following formula (1-4).
[0101] [Formula 7]
[0102]
[0103] The normalized inner product value r obtained by formula (1-4) L is a real value that is related to the downmix signal x M (1), x M (2), ..., x M Each sample value of (T) is multiplied by the real value r L 'And get the sequence r of sample values L '×x M (1), r L '×x M (2),..., r L '×x M The sequence x obtained by the difference between the sequence of sample values obtained at (T) and each sample value of the input audio signal of the left channel is L (1)-r L '×x M (1), x L (2)-r L '×x M (2), ..., x L (T)-r L '×x M The energy of (T) becomes the smallest real value r L Same value.
[0104] The left channel input sound signal x L (1), x L (2), ..., x L (T) For each sampling number t, it can be decomposed into x L (t)=r L ×x M (t)+(x L (t)- r L ×x M (t)). Here, when x L (t)- r L ×x M The sequence of values of (t) is set as the orthogonal signal x L '(1), x L '(2), ..., x L '(T), according to the decomposition, the sampling values y of the left channel differential signal L (t)=x L (t)-αx M (t) and the downmixed signal x M (1), x M (2), ..., x MEach sample value x of (T) M (t) is multiplied by the normalized inner product value r L and using the left channel subtraction gain α (r L -α) and the value (r L -α)×x M (t) and the sampling values x of the orthogonal signal L '(t) and (r L -α)×x M (t)+x L '(t) equivalent. Orthogonal signal x L '(1), x L '(2), ..., x L '(T) represents the downmix signal x M (1), x M (2), ..., x M (T) is orthogonal, that is, the inner product is 0, so the energy of the left channel difference signal is obtained by multiplying the energy of the downmix signal by (r L -α) 2 Therefore, the energy of the quadrature signal is expressed as the sum of the energy of the quadrature signal and the energy of the quadrature signal. L The average energy per sample of the quantization error generated by encoding the left channel differential signal with a bit is a positive number σ 2 , can be estimated as shown in the following formula (1-5).
[0105] [Formula 8]
[0106]
[0107] Assuming that the quantization error generated by encoding the left channel difference signal and the quantization error of the sequence of values obtained by multiplying each sample value of the decoded quantized downmix signal by the left channel subtraction gain α are uncorrelated, the average energy per sample of the quantization error of the decoded left channel audio signal is estimated by the sum of Equations (1-5) and (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error of the decoded left channel audio signal is determined as shown in Equation (1-6).
[0108] [Formula 9]
[0109]
[0110] That is, in order to minimize the quantization error of the decoded audio signal of the left channel, the left channel subtraction gain estimation unit 120 only needs to calculate the left channel subtraction gain α using equation (1-6). In other words, considering the principle of minimizing the energy of the quantization error, the left channel subtraction gain α should be calculated by normalizing the inner product value r L The value obtained by multiplying the correction factor by the number of bits used for encoding, i.e., b L and b M The correction coefficient is a value greater than 0 and less than 1, and is used to encode the left channel differential signal in the number of bits b. L and the number of bits b used to encode the downmix signal M If the same, it is 0.5, the number of bits b used to encode the left channel differential signal L than the number of bits b used to encode the downmix signal M The more it is, the closer it is to 0.5, the number of bits b used to encode the left channel differential signal L than the number of bits b used to encode the downmix signal M The smaller the value, the closer it is to 0.5.
[0111] The same is true for the right channel. To minimize the quantization error of the decoded audio signal of the right channel, the right channel subtraction gain estimation unit 140 may determine the right channel subtraction gain β using the following equation (1-6-2).
[0112] [Formula 10]
[0113]
[0114] Here, r R is the right channel input sound signal x R (1), x R (2), ..., x R (T) and the downmix signal x M (1),x M (2), ..., x M The normalized inner product value of (T) is expressed by the following formula (1-4-2).
[0115] [Mathematical formula 11]
[0116]
[0117] That is, if the principle of minimizing the energy of the quantization error is considered, the right channel subtraction gain β should use the normalized inner product value r R The value obtained by multiplying the correction factor by the number of bits used for encoding, i.e., b R and bM The correction coefficient is a value greater than 0 and less than 1, and is used to encode the number of bits b of the right channel differential signal. R than the number of bits b used to encode the downmix signal M The larger the value, the closer it is to 0.5. The smaller the number of bits used to encode the right channel difference signal is than the number of bits used to encode the downmix signal, the closer it is to 0.5.
[0118] [Estimation and decoding of subtraction gain based on the principle of minimizing quantization error]
[0119] Specific examples of estimating and decoding the subtraction gain based on the principle of minimizing the quantization error described above will be described. In each example, the left-channel subtraction gain estimation unit 120 and the right-channel subtraction gain estimation unit 140, which estimate the subtraction gain in the encoding device 100, and the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250, which decode the subtraction gain in the decoding device 200, will be described.
[0120] [[Example 1]]
[0121] Example 1 is based on the following principle: L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T) The principle of minimizing the energy of the quantization error of the decoded audio signal of the left channel, including the case where the decoded audio signal of the left channel is not considered to be the same series; and the principle of minimizing the energy of the quantization error of the decoded audio signal of the right channel R (1), x R (2),..., x R (T) and the downmix signal x M (1), x M (2), ..., x M (T) The principle of minimizing the energy of the quantization error in the decoded audio signal of the right channel, regardless of the case where the signals are not considered to be of the same series.
[0122] [Left Channel Subtraction Gain Estimation Unit 120]
[0123] In the left channel subtraction gain estimation unit 120, the left channel subtraction gain candidate α cand (a) and the code Cα corresponding to the candidate cand The group (a) is stored in a plurality of groups (A groups, a=1, ..., A) in advance. The left channel subtraction gain estimation unit 120 performs the following operation: Figure 5The following steps S120-11 to S120-14 are shown.
[0124] The left channel subtraction gain estimation unit 120 first calculates the left channel subtraction gain based on the input audio signal x L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T), the normalized inner product value r of the input audio signal of the left channel of the downmix signal is obtained by formula (1-4) L (Step S120 - 11 ). Then, the left channel subtraction gain estimation unit 120 uses the left channel difference signal y in the stereo encoding unit 170 L (1), y L (2), ..., y L (T) is the number of bits b L , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M , and the number of samples per frame T, the left channel correction coefficient c is obtained by the following formula (1-7): L (Step S120-12).
[0125] [Mathematical formula 12]
[0126]
[0127] Next, the left channel subtraction gain estimation unit 120 obtains the normalized inner product value r obtained in step S120-11. L and the left channel correction coefficient c obtained in step S120-12 L Then, the left channel subtraction gain estimation unit 120 obtains the stored candidate α of the left channel subtraction gain. cand (1), ..., α cand (A) is closest to the multiplication value c obtained in step S120-13 L ×r L Candidate (multiplication value c L ×r L The quantized value of ) is used as the left channel subtraction gain α to obtain the stored code Cα cand (1), ..., Cα candThe code corresponding to the left channel subtraction gain α in (A) is used as the left channel subtraction gain code Cα (step S120 - 14 ).
[0128] In addition, the left channel difference signal y is converted into L (1), y L (2), ..., y L The number of bits b used in encoding (T) L If it is not explicitly determined, the number of bits b of the stereo code CS output by the stereo coding unit 170 is s One-half (i.e., b s / 2) is used as the number of bits b L In addition, the left channel correction coefficient c L It can also be a value other than the value obtained by equation (1-7) itself, but a value greater than 0 and less than 1, and the value is as follows: L (1), y L (2), ..., y L (T) is the number of bits b L and for the downmix signal x M (1), x M (2), ..., x M (T) is the number of bits b M If the same, it is 0.5, the number of bits is b L Bit number b M The more it is, the closer it is to 0 than 0.5, and the number of bits b L Bit number b M The smaller the value, the closer it is to 1 than 0.5. This also applies to the examples described below.
[0129] [Right Channel Subtraction Gain Estimation Unit 140]
[0130] In the right channel subtraction gain estimation unit 140, a plurality of sets (B sets, b=1, . . . , B) of right channel subtraction gain candidate β are stored in advance. cand (b) and the encoding Cβ corresponding to the candidate cand (b) The right channel subtraction gain estimation unit 140 performs the following Figure 5 The following steps S140-11 to S140-14 are shown.
[0131] The right channel subtraction gain estimation unit 140 first calculates the right channel subtraction gain based on the input audio signal x R (1), x R (2), ..., x R (T) and the downmix signal x M (1), xM (2), ..., x M (T), the normalized inner product value r of the input audio signal of the right channel of the downmix signal is obtained by formula (1-4-2) R (Step S140 - 11 ). Then, the right channel subtraction gain estimation unit 140 uses the right channel difference signal y in the stereo encoding unit 170 R (1), y R (2), ..., y R (T) is the number of bits b R , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M , and the number of samples per frame T, the right channel correction coefficient c is obtained by the following formula (1-7-2): R (Step S140-12).
[0132] [Mathematical formula 13]
[0133]
[0134] Next, the right channel subtraction gain estimation unit 140 obtains the normalized inner product value r obtained in step S140-11. R and the right channel correction coefficient c obtained in step S140-12 R Next, the right channel subtraction gain estimation unit 140 obtains the candidate β closest to the stored right channel subtraction gain. cand (1), ..., β cand The multiplication value c obtained in step S140-13 in (B) R ×r R Candidate (multiplication value c R ×r R The quantized value) is used as the right channel subtraction gain β, and the stored code Cβ is obtained. cand (1), ..., Cβ cand The code corresponding to the right channel subtraction gain β in (B) is used as the right channel subtraction gain Cβ (step S140 - 14 ).
[0135] In addition, the right channel difference signal y is converted into R (1), y R (2), ..., y R The number of bits b used in encoding (T) RIf it is not explicitly determined, the number of bits b of the stereo code CS output by the stereo coding unit 170 is s One-half (i.e., b s / 2) is used as the number of bits b R In addition, the right channel correction coefficient c R It can also be a value other than the value obtained by formula (1-7-2) itself, but a value greater than 0 and less than 1, and the value is as follows: R (1), y R (2), ..., y R (T) is the number of bits b R and for the downmix signal x M (1), x M (2), ..., x M (T) (T) The number of bits b of the code M If the same, it is 0.5, the number of bits is b R Bit number b M The more it is, the closer it is to 0 than 0.5, and the number of bits b R Bit number b M The smaller the value, the closer it is to 1 than 0.5. This also applies to the examples described below.
[0136] [Left channel subtraction gain decoding unit 230]
[0137] In the left channel subtraction gain decoding unit 230, similar to the portion stored in the corresponding left channel subtraction gain estimation unit 120 of the encoding device 100, a plurality of sets (A sets, a=1, . . . , A) of candidate left channel subtraction gain α are pre-stored. cand (a) and the encoding Cα corresponding to the candidate cand (a) The left channel subtraction gain decoding unit 230 converts the stored code Cα cand (1), ..., Cα cand A candidate for the left channel subtraction gain corresponding to the input left channel subtraction gain code Cα in (A) is obtained as the left channel subtraction gain α (step S230 - 11 ).
[0138] [Right Channel Subtraction Gain Decoding Unit 250]
[0139] In the right channel subtraction gain decoding unit 250, similar to the portion stored in the corresponding right channel subtraction gain estimation unit 140 of the encoding device 100, a plurality of sets (B sets, b=1, . . . , B) of right channel subtraction gain candidate β are pre-stored. cand (b) and the encoding Cβ corresponding to the candidate cand (b) The right channel subtraction gain decoding unit 250 converts the stored code Cβcand (1), ..., Cβ cand A candidate for the right channel subtraction gain corresponding to the input right channel subtraction gain code Cβ in (B) is obtained as the right channel subtraction gain β (step S250 - 11 ).
[0140] Alternatively, the same subtraction gain candidate may be used for encoding in the left and right channels. Assuming that A and B are the same value, the left channel subtraction gain candidate α stored in the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 may be set to cand (a) and the encoding Cα corresponding to the candidate cand (a), the candidate β of the right channel subtraction gain stored in the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 cand (b) and the encoding Cβ corresponding to the candidate cand The groups of (b) are the same.
[0141] [Variation of Example 1]
[0142] The number of bits b used for encoding the left channel differential signal in the encoding apparatus 100 is L is the number of bits used for decoding the left channel difference signal in the decoding device 200, and is the number of bits used for encoding the downmix signal in the encoding device 100. M The value of is the number of bits used for decoding the downmix signal in the decoding apparatus 200, so the correction coefficient c L The same value can be calculated in the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r L As the object of encoding and decoding, the quantized value of the normalized inner product value is quantized in the encoding device 100 and the decoding device 200. L Multiply by the correction factor c L , the left channel subtraction gain α is obtained. The same is true for the right channel. This method is described as a modification of Example 1.
[0143] [Left Channel Subtraction Gain Estimation Unit 120]
[0144] In the left channel subtraction gain estimation unit 120, a plurality of sets (A sets, a=1, . . . , A) of candidates for normalized inner product values of the left channel are stored in advance. Lcand (a) and the encoding Cα corresponding to the candidate cand (a) group. Figure 6 As shown, the left channel subtraction gain estimation unit 120 performs steps S120-11 and S120-12 described in Example 1, and steps S120-15 and S120-16 described below.
[0145] The left channel subtraction gain estimation unit 120 first performs the same operation as step S120-11 of the left channel subtraction gain estimation unit 120 in Example 1, based on the input left channel input audio signal x L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T), the normalized inner product value r of the input audio signal of the left channel of the downmix signal is obtained by formula (1-4) L (Step S120 - 11 ) Next, the left channel subtraction gain estimation unit 120 obtains a candidate r for the normalized inner product value with the stored left channel. Lcand (1), ..., r Lcand The normalized inner product value r obtained in step S120-11 in (A) L The closest candidate (normalized inner product value r L Quantized value of L , get the stored code Cα cand (1),..., Cα cand The closest candidate in (A) L The corresponding code is used as the left channel subtraction gain code Cα (step S120-15). In addition, the left channel subtraction gain estimation unit 120 uses the left channel difference signal y in the stereo encoding unit 170 in the same manner as step S120-12 of the left channel subtraction gain estimation unit 120 in Example 1. L (1), y L (2), ..., y L (T) is the number of bits b L , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M , and the number of samples per frame T, the left channel correction coefficient c is obtained by formula (1-7) L (Step S120-12). Next, the left channel subtraction gain estimation unit 120 obtains a quantized value of the normalized inner product value obtained in step S120-15. L and the left channel correction coefficient c obtained in step S120-12 L The multiplied value is used as the left channel subtraction gain α (step S120 - 16 ).
[0146] [Right Channel Subtraction Gain Estimation Unit 140]
[0147] In the right channel subtraction gain estimation unit 140, a plurality of sets (B sets, b=1, . . . , B) of candidates for normalized inner product values of the right channel are stored in advance. Rcand (b) and the encoding Cβ corresponding to the candidate cand (b) group. Figure 6 As shown, the right channel subtraction gain estimation unit 140 performs steps S140-11 and S140-12 described in Example 1, and steps S140-15 and S140-16 described below.
[0148] The right channel subtraction gain estimation unit 140 first performs the same operation as step S140-11 of the right channel subtraction gain estimation unit 140 in Example 1, based on the input right channel input audio signal x R (1), x R (2), ..., x R (T) and the downmix signal x M (1), x M (2), ..., x M (T), the normalized inner product value r of the input audio signal of the right channel of the downmix signal is obtained by formula (1-4-2) R (Step S140 - 11 ) Next, the right channel subtraction gain estimation unit 140 obtains a candidate r for the normalized inner product value with the stored right channel. Rcand (1), ..., r Rcand The normalized inner product value r obtained in step S140-11 in (B) R The closest candidate (normalized inner product value r R Quantized value of R , get the stored code Cβ cand (1),..., Cβ cand The closest candidate r in (B) R The corresponding code is used as the right channel subtraction gain code Cβ (step S140-15). Then, the right channel subtraction gain estimation unit 140 uses the right channel difference signal y in the stereo encoding unit 170 in the same manner as step S140-12 of the right channel subtraction gain estimation unit 140 in Example 1. R (1), y R (2), ..., y R (T) is the number of bits b R , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M, and the number of samples per frame T, the right channel correction coefficient c is obtained by formula (1-7-2) R (Step S140-12). Next, the right channel subtraction gain estimation unit 140 obtains a quantized value of the normalized inner product value obtained in step S140-15. R and the right channel correction coefficient c obtained in step S140-12 R The multiplied value is used as the right channel subtraction gain β (step S140 - 16 ).
[0149] [Left channel subtraction gain decoding unit 230]
[0150] In the left channel subtraction gain decoding unit 230, similar to the portion stored in the corresponding left channel subtraction gain estimation unit 120 of the encoding device 100, multiple sets (A sets, a=1, ..., A) of left channel normalized inner product value candidates r are pre-stored. Lcand (a) and the encoding Cα corresponding to the candidate cand (a) The left channel subtraction gain decoding unit 230 performs the following Figure 7 The following steps S230-12 to S230-14 are shown.
[0151] The left channel subtraction gain decoding unit 230 compares the stored code Cα with the cand (1), ..., Cα cand The candidate of the normalized inner product value of the left channel corresponding to the input left channel subtraction gain code Cα in (A) obtains the decoded value ^r as the normalized inner product value of the left channel L (Step S230-12). In addition, the left channel subtraction gain decoding unit 230 uses the left channel decoding differential signal y in the stereo decoding unit 220. L (1), ^y L (2), ..., ^y L (T) The number of bits decoded b L , used in the monaural decoding unit 210 to decode the audio signal in monaural mode M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M , and the number of samples per frame T, the left channel correction coefficient c is obtained by formula (1-7) L (Step S230-13) Next, the left channel subtraction gain decoding unit 230 obtains a decoded value of the normalized inner product value obtained in step S230-12. L and the left channel correction coefficient c obtained in step S230-13 LThe multiplied value is used as the left channel subtraction gain α (step S230 - 14 ).
[0152] In addition, when the stereo code CS is a code obtained by combining the left channel differential code CL and the right channel differential code CR, the left channel decoded differential signal CL is used in the stereo decoding unit 220. L (1), ^y L (2), ..., ^y L (T) The number of decoded bits b L The number of bits of the left channel differential code CL is used in the stereo decoding unit 220 to decode the differential signal ^y for the left channel. L (1), ^y L (2), ..., ^y L (T) The number of decoded bits b L If it is not explicitly determined, the number of bits b of the stereo code CS input to the stereo decoding unit 220 is s One-half (i.e., b s / 2) is used as the number of bits b L In the mono decoding unit 210, the mono decoding audio signal is decoded into M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M Is the number of bits of mono coded CM. Left channel correction coefficient c L It is not the value obtained by equation (1-7) itself, but a value greater than 0 and less than 1, and is the following value: L (1), ^y L (2), ..., ^y L (T) The number of decoded bits b L and for mono decoded sound signal^x M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M If the same, it is 0.5, the number of bits is b L Bit number b M The more it is, the closer it is to 0 than 0.5, and the number of bits b L Bit number b M The smaller it is, the closer it is to 1 than 0.5.
[0153] [Right Channel Subtraction Gain Decoding Unit 250]
[0154] In the right channel subtraction gain decoding unit 250, similar to the portion stored in the corresponding right channel subtraction gain estimation unit 140 of the encoding device 100, multiple sets (B sets, b = 1, ..., B) of right channel normalized inner product value candidates r are pre-stored. Rcand (b) and the encoding Cβ corresponding to the candidate cand (b) The right channel subtraction gain decoding unit 250 performs the following Figure 7 The following steps S250-12 to S250-14 are shown.
[0155] The right channel subtraction gain decoding unit 250 compares the stored code Cβ with the cand (1), ..., Cβ cand The candidate of the normalized inner product value of the right channel corresponding to the input right channel subtraction gain code Cβ in (B) obtains the decoded value ^r as the normalized inner product value of the right channel R (Step S250-12). In addition, the right channel subtraction gain decoding unit 250 uses the differential signal y used for right channel decoding in the stereo decoding unit 220. R (1), ^y R (2), ..., ^y R (T) The number of decoded bits b R , in the mono decoding unit 210 for mono decoding the sound signal ^x M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M , and the number of samples per frame T, the right channel correction coefficient c is obtained by formula (1-7-2) R (Step S250-13) Next, the right channel subtraction gain decoding unit 250 obtains a decoded value of the normalized inner product value obtained in step S250-12. R and the right channel correction coefficient c obtained in step S250-13 R The value obtained by multiplying by R is used as the right channel subtraction gain β (step S250 - 14 ).
[0156] In addition, when the stereo code CS is a code obtained by combining the left channel differential code CL and the right channel differential code CR, the right channel decoded differential signal CL is used in the stereo decoding unit 220. R (1), ^y R (2), ..., ^y R (T) The number of decoded bits b R The number of bits of the right channel differential code CR is used in the stereo decoding unit 220 to decode the differential signal ^y for the right channel. R(1), ^y R (2), ..., ^y R (T) The number of bits decoded b R If it is not explicitly determined, the number of bits b of the stereo code CS input to the stereo decoding unit 220 is s One-half (i.e., b s / 2) is used as the number of bits b R In the mono decoding unit 210, the mono decoding audio signal x is used M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M is the number of bits of mono coded CM. Right channel correction coefficient c R It is not the value obtained by formula (1-7-2) itself, but a value greater than 0 and less than 1, when and is the following value: R (1), ^y R (2), ..., ^y R (T) The number of decoded bits b R and for mono decoded sound signal^x M (1), ^x M (2), ..., ^x M (T) The number of decoded bits b M If the same, it is 0.5, the number of bits is b R Bit number b M The more it is, the closer it is to 0 than 0.5, and the number of bits b R Bit number b M The smaller the value, the closer it is to 1 than 0.5.
[0157] Alternatively, the same normalized inner product value candidate may be used and encoded for the left and right channels. Assuming that A and B are the same value, the normalized inner product value candidate r for the left channel stored in the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 may be the same. Lcand (a) and the encoding Cα corresponding to the candidate cand (a) The candidate r of the normalized inner product value of the right channel stored in the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 Rcand (b) and the encoding Cβ corresponding to the candidate cand The groups of (b) are the same.
[0158] Furthermore, code Cα is essentially a code corresponding to the left channel subtraction gain α. For purposes of matching the terms in the descriptions of encoding device 100 and decoding device 200, it is referred to as left channel subtraction gain coding. However, from the perspective of coding representing a normalized inner product value, it may also be referred to as left channel inner product coding. Similarly, code Cβ may also be referred to as right channel inner product coding.
[0159] [[Example 2]]
[0160] Example 2 will describe an example of using a value obtained by also taking into account input values from past frames as the normalized inner product value. Example 2 does not strictly guarantee intra-frame optimality, namely, minimizing the energy of the quantization error in the decoded audio signal of the left channel and the energy of the quantization error in the decoded audio signal of the right channel. Instead, it reduces the rapid inter-frame fluctuations in the left channel subtraction gain α and the right channel subtraction gain β, thereby reducing the noise generated in the decoded audio signal due to these fluctuations. In other words, Example 2 not only reduces the energy of the quantization error in the decoded audio signal, but also considers the auditory quality of the decoded audio signal.
[0161] Example 2 differs from Example 1 on the encoding side, namely, the left-channel subtraction gain estimation unit 120 and the right-channel subtraction gain estimation unit 140. However, on the decoding side, namely, the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250, the same as Example 1. The following description focuses on the differences between Example 2 and Example 1.
[0162] [Left Channel Subtraction Gain Estimation Unit 120]
[0163] like Figure 8 As shown, the left channel subtraction gain estimation unit 120 performs steps S120-111 to SS120-113 described below, and steps S120-12 to SS120-14 described in Example 1.
[0164] The left channel subtraction gain estimation unit 120 first uses the input left channel input audio signal x L (1), x L (2), ..., x L (T), the input downmix signal x M (1), x M (2), ..., x M (T), and the inner product value E used in the previous frame L (-1), the inner product value E used in the current frame is obtained by the following formula (1-8): L (0) (step S120-111).
[0165] [Formula 14]
[0166]
[0167] Here, ε L is a predetermined value greater than 0 and less than 1, and is stored in advance in the left channel subtraction gain estimation unit 120. In addition, the left channel subtraction gain estimation unit 120 calculates the inner product value E obtained. L (0) is used as the inner product value E used in the previous frame L (-1)" is used in the next frame and stored in the left channel subtraction gain estimation unit 120.
[0168] The left channel subtraction gain estimation unit 120 also uses the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M (-1), the energy E of the downmix signal used in the current frame is obtained by the following formula (1-9): M (0) (step S120-112).
[0169] [Mathematical formula 15]
[0170]
[0171] Here, ε M is a predetermined value greater than 0 and less than 1, and is stored in advance in the left channel subtraction gain estimation unit 120. In addition, the left channel subtraction gain estimation unit 120 calculates the energy E of the downmix signal obtained. M (0) as the energy E of the downmix signal used in the previous frame M (-1)" is used in the next frame and stored in the left channel subtraction gain estimation unit 120.
[0172] Next, the left channel subtraction gain estimation unit 120 uses the inner product value E used in the current frame obtained in step S120-111 L (0) and the energy E of the downmix signal used in the current frame obtained in step S120-112 M (0), the normalized inner product value r is obtained by the following formula (1-10): L (Step S120-113).
[0173] [Mathematical formula 16]
[0174]
[0175] The left channel subtraction gain estimation unit 120 further performs step S120-12, and then replaces the normalized inner product value r obtained in step S120-11 with L , and use the normalized inner product value r obtained in the above steps S120-113 L To perform step S120-13, and then to perform step S120-14.
[0176] In addition, the above ε L and ε M The closer it is to 1, the greater the normalized inner product value r L The more likely it is to include the influence of the left channel input audio signal and the downmixed signal of the past frame, the normalized inner product value r L , through the normalized inner product value r L The obtained left channel subtraction gain α has a smaller variation between frames.
[0177] [Right Channel Subtraction Gain Estimation Unit 140]
[0178] like Figure 8 As shown, the right channel subtraction gain estimation unit 140 performs the following steps S140-111 to S140-113 and steps S140-12 to S140-14 described in Example 1.
[0179] The right channel subtraction gain estimation unit 140 first uses the input right channel input audio signal x R (1), x R (2), ..., x R (T), the input downmix signal x M (1), x M (2), ..., x M (T), and the inner product value E used in the previous frame R (-1), the inner product value E used in the current frame is obtained by the following formula (1-8-2): R (0) (step S140-111).
[0180] [Mathematical formula 17]
[0181]
[0182] Here, ε R is a predetermined value greater than 0 and less than 1, and is stored in advance in the right channel subtraction gain estimation unit 140. In addition, the right channel subtraction gain estimation unit 140 calculates the inner product value E obtained. R (0) is used as the inner product value E used in the previous frame R(-1)” is used in the next frame and stored in the right channel subtraction gain estimation unit 140.
[0183] The right channel subtraction gain estimation unit 140 also uses the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M (-1), the energy E of the downmix signal used in the current frame is obtained by formula (1-9) M (0) (Step S140 - 112 ). The right channel subtraction gain estimation unit 140 is configured to obtain the energy E of the downmix signal. M (0) As the "energy E of the downmix signal used in the previous frame" M (-1)” and is used in the next frame and stored in the right channel subtraction gain estimation unit 140. In addition, the energy E of the downmix signal used in the current frame is also obtained by the left channel subtraction gain estimation unit 120 using equation (1-9). M (0), therefore, only one of step S120-112 performed by the left-channel subtraction gain estimation unit 120 and step S140-112 performed by the right-channel subtraction gain estimation unit 140 may be performed.
[0184] Next, the right channel subtraction gain estimation unit 140 uses the inner product value E used in the current frame obtained in step S140-111 R (0) and the energy E of the downmix signal used in the current frame obtained in step S140-112 M (0), the normalized inner product value r is obtained by the following formula (1-10-2): R (Step S140-113).
[0185] [Mathematical formula 18]
[0186]
[0187] The right channel subtraction gain estimation unit 140 further performs step S140-12, and then replaces the normalized inner product value r obtained in step S140-11 with R , and use the normalized inner product value r obtained in the above step S140-113 R Perform step S140-13, and then perform step S140-14.
[0188] In addition, the above ε R and ε M The closer it is to 1, the greater the normalized inner product value r RThe more likely it is to include the influence of the right channel input audio signal and the downmixed signal of the past frame, the normalized inner product value r R , through the normalized inner product value r R The smaller the inter-frame variation of the obtained right channel subtraction gain β is.
[0189] [Variation of Example 2]
[0190] Regarding Example 2, a modification similar to the modification of Example 1 relative to Example 1 can also be performed. This method will be described as a modification of Example 2. The modification of Example 2 differs from the modification of Example 1 on the encoding side, namely, the left channel subtraction gain estimation unit 120 and the right channel subtraction gain estimation unit 140. However, the modification of Example 2 is the same as the modification of Example 1 on the decoding side, namely, the left channel subtraction gain decoding unit 230 and the right channel subtraction gain decoding unit 250. The differences between the modification of Example 1 and the modification of Example 2 are the same as those of Example 2. Therefore, the modification of Example 2 will be described below with reference to the modification of Example 1 and Example 2 as appropriate.
[0191] [Left Channel Subtraction Gain Estimation Unit 120]
[0192] In the left channel subtraction gain estimation unit 120, similarly to the left channel subtraction gain estimation unit 120 of the modification of Example 1, a plurality of sets (A sets, a=1, . . . , A) of left channel normalized inner product value candidates r are stored in advance. Lcand (a) and the encoding Cα corresponding to the candidate cand (a) group. Figure 9 As shown, the left channel subtraction gain estimation unit 120 performs steps S120-111 to S120-113 similar to Example 2, and steps S120-12, S120-15, and S120-16 similar to the modified example of Example 1. Specifically, the steps are as follows.
[0193] The left channel subtraction gain estimation unit 120 first uses the input left channel input audio signal x L (1), x L (2), ..., x L (T), the input downmix signal x M (1), x M (2), ..., x M (T), and the inner product value E used in the previous frame L (-1), the inner product value E used in the current frame is obtained by formula (1-8) L (0) (step S120-111). The left channel subtraction gain estimation unit 120 also uses the input downmix signal x M (1), x M (2), ..., xM (T) and the energy E of the downmix signal used in the previous frame M (-1), the energy E of the downmix signal used in the current frame is obtained by formula (1-9) M (0) (Step S120-112) Next, the left channel subtraction gain estimation unit 120 uses the inner product value E used in the current frame obtained in step S120-111. L (0) and the energy E of the downmix signal used in the current frame obtained in step S120-112 M (0), the normalized inner product value r is obtained by formula (1-10) L (Step S120 - 113 ) Next, the left channel subtraction gain estimation unit 120 obtains a candidate r for the normalized inner product value with the stored left channel. Lcand (1), ..., r Lcand The normalized inner product value r obtained in step S120-113 in (A) L The closest candidate (normalized inner product value r L Quantized value of L , get the same code Cα as the stored one cand (1), ..., Cα cand The closest candidate r in (A) L The corresponding code is used as the left channel subtraction gain code Cα (step S120-15). In addition, the left channel subtraction gain estimation unit 120 uses the left channel difference signal y in the stereo encoding unit 170 L (1), y L (2), ..., y L (T) is the number of bits b L , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M , and the number of samples per frame T, the left channel correction coefficient c is obtained by formula (1-7) L (Step S120-12). Next, the left channel subtraction gain estimation unit 120 obtains a quantized value of the normalized inner product value obtained in step S120-15. L and the left channel correction coefficient c obtained in step S120-12 L The multiplied value is used as the left channel subtraction gain α (step S120 - 16 ).
[0194] [Right Channel Subtraction Gain Estimation Unit 140]
[0195] In the right channel subtraction gain estimation unit 140, similarly to the right channel subtraction gain estimation unit 140 of the modification of Example 1, a plurality of sets (B sets, b=1, . . . , B) of candidates for normalized inner product values of the right channel are stored in advance. Rcand (b) and the encoding Cβ corresponding to the candidate cand (b) group. Figure 9 As shown, the right channel subtraction gain estimation unit 140 performs steps S140-111 to S140-113 similar to Example 2, and steps S140-12, S140-15, and S140-16 similar to the modified example of Example 1. Specifically, the steps are as follows.
[0196] The right channel subtraction gain estimation unit 140 first uses the input right channel input audio signal x R (1), x R (2), ..., x R (T), the input downmix signal x M (1), x M (2), ..., x M (T), and the inner product value E used in the previous frame R (-1), and through formula (1-8-2), we get the inner product value E used in the current frame. R (0) (step S140-111). The right channel subtraction gain estimation unit 140 also uses the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M (-1), the energy E of the downmix signal used in the current frame is obtained by formula (1-9) M (0) (Step S140-112). Next, the right channel subtraction gain estimation unit 140 uses the inner product value E used in the current frame obtained in step S140-111. R (0) and the energy E of the downmix signal used in the current frame obtained in step S140-112 M (0), the normalized inner product value r is obtained by formula (1-10-2) R (Step S140 - 113 ) Next, the right channel subtraction gain estimation unit 140 obtains a candidate r for the normalized inner product value with the stored right channel. Rcand (1), ..., r Rcand (B) The normalized inner product value r obtained in step S140-113 R The closest candidate (normalized inner product value r R Quantized value of R, get the stored code Cβ cand (1), ..., Cβ cand The closest candidate r in (B) R The corresponding code is used as the right channel subtraction gain code Cβ (step S140-15). In addition, the right channel subtraction gain estimation unit 140 uses the right channel difference signal y in the stereo encoding unit 170 R (1), y R (2), ..., y R (T) is the number of bits b R , used in the mono encoding unit 160 to downmix the signal x M (1), x M (2), ..., x M (T) is the number of bits b M , and the number of samples per frame T, the right channel correction coefficient c is obtained by formula (1-7-2) R (Step S140-12). Next, the right channel subtraction gain estimation unit 140 obtains a quantized value of the normalized inner product value obtained in step S140-15. R and the right channel correction coefficient c obtained in step S140-12 R The multiplied value is used as the right channel subtraction gain β (step S140 - 16 ).
[0197] [[Example 3]]
[0198] For example, if the left-channel input audio signal contains different sounds, such as voices or music, than the right-channel input audio signal, components of the left-channel input audio signal may also contain components of the right-channel input audio signal in the downmix signal. This leads to the following problem: The larger the value used for the left-channel subtraction gain α, the more likely it is that the left-channel decoded audio signal contains sounds from the right-channel input audio signal, which should not be heard. Meanwhile, the larger the value used for the right-channel subtraction gain β, the more likely it is that the right-channel decoded audio signal contains sounds from the left-channel input audio signal, which should not be heard. Therefore, while minimizing the energy of the quantization error in the decoded audio signal is not strictly guaranteed, considering auditory quality, the left-channel subtraction gain α and the right-channel subtraction gain β may be set to values smaller than those obtained in Example 1. Similarly, the left-channel subtraction gain α and the right-channel subtraction gain β may be set to values smaller than those obtained in Example 2.
[0199] Specifically, regarding the left channel, in Examples 1 and 2, the normalized inner product value r L and the left channel correction coefficient c L The multiplication value cL ×r L The quantized value of is set as the left channel subtraction gain α. In Example 3, the normalized inner product value r L , left channel correction coefficient c L and λ, which is a predetermined value greater than 0 and less than 1 L The multiplication value λ L ×c L ×r L The quantized value of is set as the left channel subtraction gain α. Therefore, the multiplication value c can also be set as in Example 1 and Example 2. L ×r L As the object of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, the left channel subtraction gain code Cα represents the multiplication value c L ×r L As the quantized value of the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230, the multiplication value c L ×r L The quantized value of λ L Multiply and get the left channel subtraction gain α. Alternatively, the normalized inner product value r L , left channel correction coefficient c L and the preset value λ L The multiplication value λ L ×c L ×r L The left channel subtraction gain code Cα, which is the target of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, represents the multiplication value λ. L ×c L ×r L quantized value of .
[0200] Similarly, for the right channel, in Examples 1 and 2, the normalized inner product value r R and the right channel correction coefficient c R The multiplication value c R ×r R The quantized value of is set as the right channel subtraction gain β. In contrast, in Example 3, the normalized inner product value r R , right channel correction coefficient c R and λ, which is a predetermined value greater than 0 and less than 1 R The multiplication value λ R ×c R ×r R The quantized value of is set as the right channel subtraction gain β. Therefore, the multiplication value c can also be set as in Example 1 and Example 2. R ×r RAs the object of encoding in the right channel subtraction gain estimation unit 140 and decoding in the right channel subtraction gain decoding unit 250, the right channel subtraction gain code Cβ represents the multiplication value c R ×r R As the quantized value of the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250, the multiplication value c R ×r R The quantized value and λ R Alternatively, the normalized inner product value r R , left channel correction coefficient c R and the predetermined value λ R The multiplication value λ R ×c R ×r R The right channel subtraction gain code Cβ, which is the target of encoding by the right channel subtraction gain estimation unit 140 and decoding by the right channel subtraction gain decoding unit 250, represents the multiplication value λ. R ×c R ×r R In addition, let λ R is L The same value is sufficient.
[0201] [Variation of Example 3]
[0202] As described above, both the encoding device 100 and the decoding device 200 can calculate the same value. Therefore, similarly to the modification of Example 1 and the modification of Example 2, the normalized inner product value r can be set to L represents the target of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230. The left channel subtraction gain code Cα represents the normalized inner product value r L The left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 convert the normalized inner product value r L Quantized value, left channel correction coefficient c L and λ which is a predetermined value greater than 0 and less than 1 L Alternatively, the normalized inner product value r L and λ, which is a predetermined value greater than 0 and less than 1 L The multiplication value λ L ×r L As the object of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, the left channel subtraction gain code Cα represents the multiplication value λ L ×r LAs the quantized value of λ, the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 multiply the value λ by L ×r L The quantized value and the left channel correction coefficient c L Multiply them to get the left channel subtraction gain α.
[0203] The same is true for the right channel, the correction coefficient c R The same value can be calculated in the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r can be calculated similarly to the modification of Example 1 and the modification of Example 2. R As the object of encoding by the right channel subtraction gain estimation unit 140 and decoding by the right channel subtraction gain decoding unit 250, the right channel subtraction gain code Cβ represents the normalized inner product value r R As the quantized value of R Quantized value of the right channel correction coefficient c R and a predetermined value greater than 0 and less than 1, namely λ R Alternatively, the normalized inner product value r R and a predetermined value greater than 0 and less than 1, namely λ R The multiplication value λ R ×r R The right channel subtraction gain code Cβ, which is the target of encoding in the right channel subtraction gain estimation unit 140 and decoding in the right channel subtraction gain decoding unit 250, represents the multiplication value λ. R ×r R The right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 multiply the multiplication value λ R ×r R The quantized value and the right channel correction coefficient c R Multiplying them together gives the right channel subtraction gain β.
[0204] [[Example 4]]
[0205] The auditory quality issue described at the beginning of Example 3 arises when the correlation between the left-channel and right-channel input audio signals is low. This issue rarely arises when the correlation between the left-channel and right-channel input audio signals is high. Therefore, in Example 4, instead of the predetermined value in Example 3, the left-right correlation coefficient γ, which is the correlation coefficient between the left-channel and right-channel input audio signals, is used. As the correlation between the left-channel and right-channel input audio signals increases, priority is given to reducing the energy of the quantization error in the decoded audio signal. Furthermore, as the correlation between the left-channel and right-channel input audio signals decreases, priority is given to suppressing degradation in auditory quality.
[0206] The encoding side of Example 4 is different from Examples 1 and 2, but the decoding side, namely the left channel subtraction gain decoding unit 230 and the right channel subtraction gain decoding unit 250, are the same as Examples 1 and 2. The differences between Example 4 and Examples 1 and 2 are described below.
[0207] [Left-right relationship information estimation unit 180]
[0208] like Figure 1 As indicated by the middle dashed line, the encoding device 100 of Example 4 further includes a left-right correlation information estimation unit 180. The left-right correlation information estimation unit 180 receives inputs of the left channel input audio signal and the right channel input audio signal. The left-right correlation information estimation unit 180 obtains and outputs a left-right correlation coefficient γ based on the received left channel input audio signal and the received right channel input audio signal (step S180).
[0209] The left-right correlation coefficient γ is the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and can also be the sample sequence x of the input audio signal of the left channel. L (1), x L (2), ..., x L (T) and the sample sequence x of the input audio signal of the right channel R (1), x R (2), ..., x R The correlation coefficient γ0 of (T) may also be a correlation coefficient that takes into account the time difference, for example, the correlation coefficient γ between the sample sequence of the input audio signal of the left channel and the sample sequence of the input audio signal of the right channel located later than the sample sequence by τ samples. τ .
[0210] Let τ be information equivalent to the difference between the arrival time from the sound source primarily emitting sound in a space to the left-channel microphone and the arrival time from the sound source to the right-channel microphone (the so-called arrival time difference), when the sound signal obtained by AD-converting the sound collected by the left-channel microphone placed in a certain space is the left-channel input sound signal, and the sound signal obtained by AD-converting the sound collected by the right-channel microphone placed in the same space is the right-channel input sound signal. This information is hereinafter referred to as the left-right time difference. The left-right time difference τ can be obtained by any known method, or by the method described in the left-right relationship information estimation unit 181 of the second reference method. That is, the above-mentioned correlation coefficient γ τ This is information equivalent to the correlation coefficient between the sound signal collected from the sound source to the left channel microphone and the sound signal collected from the sound source to the right channel microphone.
[0211] [Left Channel Subtraction Gain Estimation Unit 120]
[0212] The left channel subtraction gain estimation unit 120 replaces step S120-13 and obtains the normalized inner product value r obtained in step S120-11 or step S120-113. L , the left channel correction coefficient c obtained in step S120-12 L , and the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 (step S120-13"). Next, the left channel subtraction gain estimation unit 120 replaces step S120-14 and obtains the candidate α of the stored left channel subtraction gain. cand (1), ..., α cand The multiplication value γ×c obtained in step S120-13" in (A) L ×r L The closest candidate (multiplication value γ×c L ×r L The quantized value of ) is used as the left channel subtraction gain α to obtain the stored code Cα cand (1), ..., Cα cand The code corresponding to the left channel subtraction gain α in (A) is used as the left channel subtraction gain Cα (step S120 - 14 ”).
[0213] [Right Channel Subtraction Gain Estimation Unit 140]
[0214] The right channel subtraction gain estimation unit 140 replaces step S140-13 and obtains the normalized inner product value r obtained in step S140-11 or step S140-113. R , the right channel correction coefficient c obtained in step S140-12R , and the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 (step S140-13"). Next, the right channel subtraction gain estimation unit 140 replaces step S140-14 and compares the stored candidate β of the right channel subtraction gain to cand (1),..., β cand The multiplication value γ×c obtained in step S140-13" in (B) R ×r R The closest candidate (multiplication value γ×c R ×r R The quantized value of ( ) is obtained as the right channel subtraction gain β, and the stored code Cβ is obtained. cand (1), ..., Cβ cand The code corresponding to the right channel subtraction gain β in (B) is used as the right channel subtraction gain Cβ (step S140 - 14 ”).
[0215] [Variation of Example 4]
[0216] As described above, both the encoding device 100 and the decoding device 200 can calculate the same value. Therefore, the normalized inner product value r L Multiplication value γ×r by left and right correlation coefficient γ L As the object of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, the left channel subtraction gain code Cα represents the multiplication value γ×r L As the quantized value of γ×r is, the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 convert the multiplication value γ×r L The quantized value and the left channel correction coefficient c L Multiply them to get the left channel subtraction gain α.
[0217] The same is true for the right channel, the correction coefficient c R The same value can be calculated in the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r R Multiplication value γ×r by left and right correlation coefficient γ R As the object of encoding by the right channel subtraction gain estimation unit 140 and decoding by the right channel subtraction gain decoding unit 250, the right channel subtraction gain code Cβ represents the multiplication value γ×r R As the quantized value of γ×r is, the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 convert the multiplication value γ×r R The quantized value and the right channel correction coefficient c R Multiplying them together gives the right channel subtraction gain β.
[0218] <Second Reference Method>
[0219] The encoding device and decoding device according to the second reference method will be described.
[0220] <<Encoding Device 101>>
[0221] like Figure 10 As shown, encoding device 101 according to the second reference scheme includes a downmixing unit 110, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel signal subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimation unit 181, and a time shifting unit 191. Encoding device 101 according to the second reference scheme differs from encoding device 100 according to the first reference scheme in that it includes left-right relationship information estimation unit 181 and time shifting unit 191; that left channel signal subtraction gain estimation unit 120, left channel signal subtraction unit 130, right channel signal subtraction gain estimation unit 140, and right channel signal subtraction unit 150 use the signal output by time shifting unit 191 instead of the signal output by downmixing unit 110; and that, in addition to the aforementioned codes, it also outputs a left-right time difference code Cτ, described later. The other structures and operations of the encoding apparatus 101 of the second reference method are the same as those of the encoding apparatus 100 of the first reference method. The encoding apparatus 101 of the second reference method performs the following operations on each frame: Figure 11 The processing of steps S110 to S191 is illustrated. Hereinafter, differences between the encoding apparatus 101 according to the second reference scheme and the encoding apparatus 100 according to the first reference scheme will be described.
[0222] [Left-Right Relationship Information Estimation Unit 181]
[0223] The left-channel input audio signal and the right-channel input audio signal input to the encoding device 101 are input to the left-channel information estimation unit 181. The left-channel information estimation unit 181 obtains a left-right time difference τ and a left-right time difference code Cτ representing the left-right time difference τ from the input left-channel and right-channel input audio signals, and outputs the obtained results (step S181).
[0224] The left-right time difference τ is information equivalent to the difference between the arrival time from the sound source primarily emitting sound in a certain space to the left-channel microphone and the arrival time from that sound source to the right-channel microphone (the so-called arrival time difference), assuming that the sound signal collected by the left-channel microphone placed in a certain space is the left-channel input sound signal and the sound signal collected by the right-channel microphone placed in the same space is the right-channel input sound signal. Furthermore, the left-right time difference τ includes not only the arrival time difference but also information about which microphone arrives first. The left-right time difference τ can take both positive and negative values based on either input sound signal. In other words, the left-right time difference τ indicates whether the same sound signal is contained in the left-channel input sound signal or the right-channel input sound signal. Hereinafter, when the same sound signal is included in the input sound signal of the left channel before the input sound signal of the right channel, it is also referred to as left channel first, and when the same sound signal is included in the input sound signal of the right channel before the input sound signal of the left channel, it is also referred to as right channel first.
[0225] The left-right time difference τ can also be obtained by any known method. For example, the left-right relationship information estimation unit 181 calculates the predetermined τ max to τ min (For example, τ max is a positive number, τ min is a negative number) for each candidate sampling number τ cand Calculate the difference between the sample sequence representing the input audio signal of the left channel and the candidate sample number τ that is shifted backward from the sample sequence cand The value of the correlation magnitude of the sample sequence of the input audio signal of the right channel at the position γ (hereinafter referred to as the correlation value) cand , the correlation value γ cand The maximum number of candidate samples τ cand The left-right time difference τ is obtained. That is, in this example, when the left channel is ahead, the left-right time difference τ is a positive value, and when the right channel is ahead, the left-right time difference τ is a negative value. The absolute value of the left-right time difference τ is a value indicating how far ahead the leading channel is relative to the other channel (the number of samples ahead). For example, when calculating the correlation value γ using only samples within a frame, cand In the case of τ cand When the value is positive, the partial sample sequence x of the input audio signal of the right channel is calculated. R (1+τ cand ), x R (2+τ cand ), ..., x R (T), and the candidate sample number τ that deviates from the partial sample sequencecand A partial sample column x of the input sound signal of the left channel at the position L (1), x L (2), ..., x L (T-τ cand The absolute value of the correlation coefficient of ) is taken as the correlation value γ cand , in τ cand When it is a negative value, calculate the partial sample sequence x of the input sound signal of the left channel L (1-τ cand ), x L (2-τ cand ), ..., x L (T), and the candidate sample number -τ that deviates from the partial sample sequence cand A partial sampling sequence x of the input sound signal of the right channel at the position R (1), x R (2), ..., x R (T+τ cand The absolute value of the correlation coefficient of ) is taken as the correlation value γ cand Of course, in order to calculate the correlation value γ cand Alternatively, one or more samples of the past input sound signal that are continuous with the sample sequence of the input sound signal of the current frame may be used. In this case, the sample sequence of the input sound signal of the past frame may be stored in a storage unit (not shown) within the left-right relationship information estimation unit 181 in a predetermined number of frames.
[0226] Alternatively, for example, instead of the absolute value of the correlation coefficient, the correlation value γ may be calculated using information on the phase of the signal as follows: cand In this example, the left-right relationship information estimation unit 181 first calculates the left channel input audio signal x according to the following equations (3-1) and (3-2). L (1), x L (2), ..., x L (T) and the right channel input sound signal x R (1), x R (2), ..., x R (T) Perform Fourier transform to obtain the spectrum X at each frequency k from 0 to T-1 L (k) and X R (k).
[0227] [Mathematical formula 19]
[0228]
[0229] [Mathematical formula 20]
[0230]
[0231] The left-right relationship information estimation unit 181 uses the obtained spectrum X L (k) and X R (k), the spectrum φ(k) of the phase difference at each frequency k is obtained by the following formula (3-3).
[0232] [Mathematical formula 21]
[0233]
[0234] By performing inverse Fourier transform on the spectrum of the obtained phase difference, as shown in the following formula (3-4), for the phase difference from τ max to τ min The number of candidate samples τ cand , and obtain the phase difference signal ψ(τ cand ).
[0235] [Mathematical formula 22]
[0236]
[0237] The phase difference signal ψ(τ cand ) represents the absolute value of the left channel input sound signal x L (1), x L (2),..., x L (T) and the right channel input sound signal x R (1), x R (2), ..., x R (T) corresponds to a certain correlation of the rationality of the time difference, so the number of candidate samples τ cand The phase difference signal ψ(τ cand ) is taken as the correlation value γ cand The left-right relationship information estimation unit 181 obtains the phase difference signal ψ(τ cand ) is the absolute value of the correlation value γ cand Become the maximum candidate sampling number τ cand Alternatively, the phase difference signal ψ(τ cand ) is taken as the correlation value γ cand , and using, for example, relative to each τ cand Located relative to the phase difference signal ψ(τ cand ) of the absolute value of τ cand The normalized value is the relative difference between the absolute value of the phase difference signal obtained for each of the plurality of candidate sampling numbers before and after. cand, or you can use a predetermined positive number τ range , the average value is obtained by the following formula (3-5), and the obtained average value ψ c (τ cand ) and the phase difference signal ψ(τ cand ), the normalized correlation value obtained by the following formula (3-6) is taken as γ cand use.
[0238] [Mathematical formula 23]
[0239]
[0240] [Mathematical formula 24]
[0241]
[0242] In addition, the normalized correlation value obtained by formula (3-6) is a value greater than or equal to 0 and less than or equal to 1, which represents τ cand As the time difference between the left and right, the more reasonable it is, the closer it is to 1, τ cand The more unreasonable the left and right time difference is, the closer it is to 0.
[0243] Furthermore, the left-right relationship information estimation unit 181 encodes the left-right time difference τ using a predetermined encoding method to obtain a code that can uniquely identify the left-right time difference τ, namely, the left-right time difference code Cτ. As the predetermined encoding method, a well-known encoding method such as scalar quantization can be used. Furthermore, the predetermined number of candidate samples can be obtained from τ. max to τ min The integer values of may also include those in the range from τ max to τ min The fractional value or decimal value between τ max to τ min Any integer value between . In addition, it can be τ max = -τ min , or not. In addition, in the case of a special input audio signal that a certain channel must have first, it can also be τ max and τ min are all positive numbers, or τ max and τ min All are negative numbers.
[0244] Furthermore, when the encoding device 101 estimates the subtraction gain based on the principle of minimizing the quantization error in Example 4 or the modified example of Example 4 described in the first reference mode, the left-right relationship information estimation unit 181 further outputs the correlation value between the sample sequence of the input audio signal of the left channel and the sample sequence of the input audio signal of the right channel at a position shifted backward by the left-right time difference τ, that is, the correlation value for the position from τ to τ. max to τ min The number of candidate samples τ cand The calculated correlation value γ cand The maximum value among them is taken as the left-right correlation coefficient γ (step S180).
[0245] [Time Shift Section 191]
[0246] The downmix signal x output from the downmixing unit 110 is input to the time shift unit 191. M (1), x M (2), ..., x M (T) and the left-right time difference τ output by the left-right relationship information estimation unit 181. When the left-right time difference τ is a positive value (i.e., the left-right time difference τ indicates that the left channel is ahead), the time shift unit 191 shifts the downmix signal x M (1), x M (2), ..., x M (T) is directly output to the left channel subtraction gain estimation unit 120 and the left channel signal subtraction unit 130 (i.e., it is determined to be used in the left channel subtraction gain estimation unit 120 and the left channel signal subtraction unit 130). The downmix signal is delayed by |τ| samples (the number of samples corresponding to the absolute value of the left and right time difference τ, the number of samples corresponding to the magnitude of the left and right time difference τ). M (1-|τ|), x M (2-|τ|), ..., x M (T-|τ|) is the delayed downmix signal x M' (1), x M' (2), ..., x M' (T) is output to the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150 (i.e., it is determined to be used in the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150). When the left-right time difference τ is a negative value (i.e., the left-right time difference τ indicates that the right channel is ahead), the signal x obtained by delaying the downmix signal by |τ| samples is obtained. M (1-|τ|), x M (2-|τ|), ..., x M (T-|τ|) is the delayed downmix signal x M' (1), x M' (2), ..., xM' (T) is output to the left channel subtraction unit 120 and the left channel signal subtraction unit 130 (ie, it is determined to be used in the left channel subtraction unit 120 and the left channel signal subtraction unit 130), and the downmix signal x is converted to M (1), x M (2), ...,x M (T) is directly output to the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150 (i.e., it is determined to be used in the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150). When the left-right time difference τ is 0 (i.e., the left-right time difference τ indicates that neither channel has advanced), the downmix signal x is converted to M (1), x M (2), ..., x M (T) is directly output to the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 (i.e., it is determined to be used in the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150) (step S191). Specifically, for the channel with the shorter arrival time among the left and right channels, the input downmix signal is directly output to the subtraction gain estimation unit and the signal subtraction unit of the channel. For the channel with the longer arrival time among the left and right channels, a signal obtained by delaying the input downmix signal by the absolute value of the left-right time difference τ is output to the subtraction gain estimation unit and the signal subtraction unit of the channel. Furthermore, since the time shift unit 191 uses the downmix signal of the past frame to obtain the delayed downmix signal, a predetermined number of frames of downmix signal input in the past frame is stored in a storage unit (not shown) within the time shift unit 191. Furthermore, if the left channel subtraction gain estimation unit 120 and the right channel subtraction gain estimation unit 140 obtain the left channel subtraction gain α and the right channel subtraction gain β using a known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing quantization error, then a unit that obtains a local decoded signal corresponding to the monaural code CM may be provided in the subsequent stage of the monaural encoding unit 160 of the encoding device 101 or within the monaural encoding unit 160. Alternatively, the time shift unit 191 may use the downmix signal x in place of the downmix signal x. M (1), x M (2), ..., x M (T), using the quantized downmix signal ^x as a mono-encoded local decoded signal M (1), ^x M (2), ..., ^x MIn this case, the time shift unit 191 outputs the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M (T) to replace the downmix signal x M (1), x M (2), ..., x M (T), output delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to replace the delayed downmix signal x M' (1), x M' (2), ..., x M' (T).
[0247] [Left Channel Subtraction Gain Estimation Unit 120, Left Channel Signal Subtraction Unit 130, Right Channel Subtraction Gain Estimation Unit 140, Right Channel Signal Subtraction Unit 150]
[0248] The left channel subtraction gain estimation section 120, the left channel signal subtraction section 130, the right channel subtraction gain estimation section 140, and the right channel signal subtraction section 150 use the downmix signal x input from the time shift section 191. M (1), x M (2), ..., x M (T) or delayed downmix signal x M' (1), x M' (2), ..., x M' (T) The same operation as described in the first reference embodiment is performed, but the downmix signal x outputted from the downmixing unit 110 is replaced with M (1), x M (2), ..., x M (T) (steps S120, S130, S140, S150, S150). That is, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 use the downmix signal x determined by the time shift unit 191. M (1), x M (2), ..., x M (T) or delayed downmix signal x M' (1), x M' (2), ..., x M' (T), and performs the same operation as that described in the first reference mode. In addition, the time shift unit 191 outputs the quantized downmix signal ^x M (1), ^xM (2), ..., ^x M (T) instead of the downmix signal x M (1), x M (2), ..., x M (T), output delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' (T) instead of the delayed downmix signal x M' (1), x M' (2), ..., x M' In the case of (T), the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 use the quantized downmix signal φx input from the time shift unit 191. M (1), ^x M (2), ..., ^x M (T) or delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to carry out the above-mentioned processing.
[0249] <<Decoding Device 201>>
[0250] like Figure 12 As shown, the decoding device 201 of the second reference method includes: a mono decoding unit 210, a stereo decoding unit 220, a left channel subtraction gain decoding unit 230, a left channel signal addition unit 240, a right channel signal subtraction gain decoding unit 250, a right channel signal addition unit 260, a left-right time difference decoding unit 271, and a time shift unit 281. The difference between the decoding device 201 of the second reference method and the decoding device 200 of the first reference method is that in addition to the above-mentioned codes, the left-right time difference code Cτ described later is also input; the left-right time difference decoding unit 271 and the time shift unit 281 are included; and the left channel signal addition unit 240 and the right channel signal addition unit 260 use the signal output by the time shift unit 281 instead of the signal output by the mono decoding unit 210. The other structures and operations of the decoding device 201 of the second reference method are the same as those of the decoding device 200 of the first reference method. The decoding device 201 of the second reference method performs the following operations on each frame: Figure 13 The processing of steps S210 to S281 is illustrated. Hereinafter, the differences between the decoding apparatus 201 according to the second reference aspect and the decoding apparatus 200 according to the first reference aspect will be described.
[0251] [Left-right time difference decoding unit 271]
[0252] The left-right time difference code Cτ input to the decoding device 201 is input to the left-right time difference decoding unit 271. The left-right time difference decoding unit 271 decodes the left-right time difference code Cτ using a predetermined decoding method, obtains the left-right time difference τ, and outputs it (step S271). As the predetermined decoding method, a decoding method corresponding to the encoding method used by the left-right relationship information estimation unit 181 of the corresponding encoding device 101 is used. The left-right time difference τ obtained by the left-right time difference decoding unit 271 is the same value as the left-right time difference τ obtained by the left-right relationship information estimation unit 181 of the corresponding encoding device 101, which is obtained from τ max to τ min Any value in the range of .
[0253] [Time Shift Section 281]
[0254] The time shift unit 281 receives the mono decoded audio signal ^x output from the mono decoding unit 210. M (1), ^x M (2),..., ^x M (T) and the left-right time difference τ output by the left-right time difference decoding unit 271. When the left-right time difference τ is a positive value (i.e., the left-right time difference τ indicates that the left channel is ahead), the time shift unit 281 converts the mono decoded audio signal into M (1), ^x M (2), ..., ^x M (T) is directly output to the left channel signal adding section 240 (ie, it is determined to be used in the left channel signal adding section 240), and the signal ^x obtained by delaying the monaural decoded audio signal by |τ| samples is obtained. M (1-|τ|), ^x M (2-|τ|), ...,^x M (T-|τ|) is the delayed mono decoded sound signal^x M' (1), ^x M' (2), ..., ^x M' (T) is output to the right channel signal adding unit 260 (i.e., it is determined to be used in the right channel signal adding unit 260). When the left-right time difference τ is a negative value (i.e., the left-right time difference τ indicates that the right channel is ahead), the signal ^x obtained by delaying the monaural decoded audio signal by |τ| samples is obtained. M (1-|τ|), ^x M (2-|τ|), ..., ^x M (T-|τ|) is the delayed mono decoded sound signal^x M' (1), ^x M' (2), ..., ^x M'(T) is output to the left channel signal adding section 240 (ie, it is determined to be used in the left channel signal adding section 240), and the mono decoded audio signal is converted to M (1), ^x M (2), ..., ^x M (T) is directly output to the right channel signal adding unit 260 (i.e., it is determined to be used in the right channel signal adding unit 260). When the left and right time difference τ is 0 (i.e., the left and right time difference τ indicates that neither channel has advanced), the mono decoded audio signal is converted to M (1), ^x M (2), ..., ^x M (T) is directly output to the left channel signal adding unit 240 and the right channel signal adding unit 260 (i.e., it is determined to be used by the left channel signal adding unit 240 and the right channel signal adding unit 260) (step S281). Furthermore, since the monaural decoded audio signal of the past frame is used to obtain the delayed monaural decoded audio signal in the time shifting unit 281, a predetermined number of frames of the monaural decoded audio signal input in the past frame is stored in a storage unit (not shown) within the time shifting unit 281.
[0255] [Left Channel Signal Adding Unit 240, Right Channel Signal Adding Unit 260]
[0256] The left channel signal adding section 240 and the right channel signal adding section 260 use the monaural decoded audio signal input from the time shift section 281. M (1), ^x M (2), ..., ^x M (T) or delayed mono decoded sound signal^x M' (1), ^x M' (2),..., ^x M' (T) The same operation as described in the first reference embodiment is performed, but the monaural decoded audio signal ^x outputted by the monaural decoding unit 210 is replaced by M (1), ^x M (2), ..., ^x M (T) (steps S240, S260). That is, the left channel signal adding unit 240 and the right channel signal adding unit 260 use the monaural decoded audio signal ^x determined by the time shift unit 281. M (1),^x M (2), ..., ^x M (T) or delayed mono decoded sound signal^x M' (1), ^x M' (2), ..., ^x M'(T), perform the same operation as described in the first reference method.
[0257] <First embodiment>
[0258] The first embodiment is a modification of the encoding device 101 of the second reference method, in which the downmix signal is generated by taking into account the relationship between the left and right channel input audio signals. The encoding device of the first embodiment is described below. Since the codes generated by the encoding device of the first embodiment can be decoded by the decoding device 201 of the second reference method, the description of the decoding device is omitted.
[0259] <<Encoding Device 102>>
[0260] like Figure 10 As shown, the encoding device 102 of the first embodiment includes a downmixing unit 112, a left channel subtraction gain estimating unit 120, a left channel signal subtracting unit 130, a right channel subtraction gain estimating unit 140, a right channel signal subtracting unit 150, a mono encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimating unit 182, and a time shifting unit 191. The encoding device 102 of the third embodiment differs from the encoding device 101 of the second reference embodiment in that the encoding device 102 includes a left-right relationship information estimating unit 182 instead of the left-right relationship information estimating unit 181, and includes a downmixing unit 112 instead of the downmixing unit 110. Figure 10 As shown by the middle dotted line, the left-right relationship information estimation unit 182 obtains and outputs the left-right correlation coefficient γ and the preceding channel information. The output left-right correlation coefficient γ and the preceding channel information are input to the downmixing unit 112 for use. The other structures and operations of the encoding device 102 of the first embodiment are the same as those of the encoding device 101 of the second reference mode. The encoding device 102 of the first embodiment performs the following operations on each frame: Figure 14 The processing of steps S112 to S191 is illustrated. The following describes the differences between the encoding device 102 of the first embodiment and the encoding device 101 of the second reference mode.
[0261] [Left-Right Relationship Information Estimation Unit 182]
[0262] The left channel input audio signal and the right channel input audio signal input into the encoding device 102 are input to the left-right relationship information estimation unit 182. Based on the input left and right channel input audio signals, the left-right relationship information estimation unit 182 obtains and outputs the left-right time difference τ, the left-right time difference code Cτ representing the left-right time difference τ, the left-right correlation coefficient γ, and the preceding channel information (step S182). The process by which the left-right relationship information estimation unit 182 obtains the left-right time difference τ and the left-right time difference code Cτ is similar to that of the left-right relationship information estimation unit 181 in the second reference embodiment.
[0263] The left-right correlation coefficient γ corresponds to the correlation coefficient between the sound signal collected by the left-channel microphone and the sound signal collected by the right-channel microphone from the sound source, as discussed above in the description of left-right relationship information estimation unit 181 in the second reference embodiment. The preceding channel information corresponds to which microphone the sound from the sound source reaches first. It indicates whether the same sound signal is included first in the left or right channel input sound signal, and thus indicates which of the left and right channels precedes.
[0264] According to the example described above in the description of the left-right relationship information estimation unit 181 in the second reference embodiment, the left-right relationship information estimation unit 182 calculates the correlation value between the sample sequence of the input audio signal of the left channel and the sample sequence of the input audio signal of the right channel at a position that is offset from the sample sequence by the left-right time difference τ, that is, the correlation value for the sample sequence from τ max to τ min The number of candidate samples τ cand The calculated correlation value γ cand The maximum value of is output as the left-right correlation coefficient γ. Furthermore, when the left-right time difference τ is a positive value, the left-right relationship information estimation unit 182 obtains and outputs information indicating that the left channel precedes as the preceding channel information. When the left-right time difference τ is a negative value, the left-right relationship information estimation unit 182 obtains and outputs information indicating that the right channel precedes as the preceding channel information. When the left-right time difference τ is zero, the left-right relationship information estimation unit 182 may obtain and output information indicating that the left channel precedes as the preceding channel information, information indicating that the right channel precedes as the preceding channel information, or information indicating that neither channel precedes as the preceding channel information.
[0265] [Downmixing Unit 112]
[0266] The downmixing unit 112 receives the left channel input audio signal input to the encoding device 102, the right channel input audio signal input to the encoding device 102, the left-right correlation coefficient γ output by the left-right correlation information estimation unit 182, and the preceding channel information output by the left-right correlation information estimation unit 182. The downmixing unit 112 performs a weighted average of the left channel input audio signal and the right channel input audio signal to obtain a downmixed signal, and outputs the downmixed signal so that the larger the left-right correlation coefficient γ, the larger the preceding channel input audio signal of the left channel or the right channel input audio signal (step S112).
[0267] For example, if the absolute value or normalized value of the correlation coefficient is used as the correlation value as in the example described above in the description of the left-right relationship information estimation unit 181 of the second reference method, the obtained left-right correlation coefficient γ is a value greater than or equal to 0 and less than or equal to 1. Therefore, the downmixing unit 112 uses the weight determined by the left-right correlation coefficient γ for each corresponding sample number t to mix the left channel input audio signal x L (t) and the right channel input sound signal x R (t) The value obtained by weighted addition is taken as the downmix signal x M Specifically, when the preceding channel information indicates that the left channel precedes, that is, when the left channel precedes, the downmixing unit 112 sets x M (t) = ((1+γ) / 2)×x L (t)+((1-γ) / 2)×x R (t), when the preceding channel information indicates that the right channel precedes, that is, when the right channel precedes, let x M (t) = ((1-γ) / 2) × x L (t)+((1+γ) / 2)×x R (t), get the downmix signal x M The downmix unit 112 generates a downmix signal in this manner. The smaller the left-right correlation coefficient γ, that is, the smaller the correlation between the left-channel input audio signal and the right-channel input audio signal, the closer the downmix signal is to a signal obtained by averaging the left-channel input audio signal and the right-channel input audio signal. The larger the left-right correlation coefficient γ, that is, the greater the correlation between the left-channel input audio signal and the right-channel input audio signal, the closer the downmix signal is to the input audio signal of the preceding channel between the left-channel input audio signal and the right-channel input audio signal.
[0268] Furthermore, when neither channel is advanced, the downmixing unit 112 may average the left channel input audio signal and the right channel input audio signal so that the left channel input audio signal and the right channel input audio signal are included in the downmixed signal with equal weight, and output the obtained downmixed signal. Therefore, when the advanced channel information indicates that neither channel is advanced, the downmixing unit 112 averages the left channel input audio signal x for each sample number t. L (t) and the input sound signal x of the right channel R (t) The averaged x M (t)=(x L (t)+x R (t)) / 2 is set as the downmix signal x M (t).
[0269] <Second embodiment>
[0270] The encoding device 100 of the first reference embodiment can also be modified to generate a downmix signal by taking into account the relationship between the left-channel input audio signal and the right-channel input audio signal. This embodiment will be described as a second embodiment. Furthermore, since the code generated by the encoding device of the second embodiment can be decoded by the decoding device 200 of the first reference embodiment, the description of the decoding device will be omitted.
[0271] <<Encoding Device 103>>
[0272] like Figure 1 As shown, the encoding device 103 of the second embodiment includes: a downmixing unit 112, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a mono encoding unit 160, a stereo encoding unit 170, and a left-right relationship information estimation unit 183. The encoding device 103 of the second embodiment differs from the encoding device 100 of the first reference embodiment in that the downmixing unit 110 is replaced with the downmixing unit 112. Figure 1 The left-right relationship information estimation unit 183 shown by the dotted line in the middle includes the left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ and the preceding channel information. The output left-right correlation coefficient γ and the preceding channel information are input to the downmixing unit 112 for use. The other structures and operations of the encoding device 103 of the second embodiment are the same as those of the encoding device 100 of the first reference mode. In addition, the operation of the downmixing unit 112 of the encoding device 103 of the second embodiment is the same as the operation of the downmixing unit 112 of the encoding device 102 of the first embodiment. The encoding device 103 of the second embodiment performs the following operations on each frame: Figure 15 The processing of steps S112 to S183 is illustrated. Hereinafter, the differences between the encoding device 103 according to the second embodiment and the encoding device 100 according to the first reference mode and the encoding device 102 according to the first embodiment will be described.
[0273] [Left-Right Relationship Information Estimation Unit 183]
[0274] The left-channel input audio signal and the right-channel input audio signal input to the encoding device 103 are input to the left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains a left-right correlation coefficient γ and preceding channel information based on the input left-channel and right-channel input audio signals, and outputs them (step S183).
[0275] The left-right correlation coefficient γ and preceding channel information obtained and output by the left-right relationship information estimation unit 183 are the same as those described in the first embodiment. Specifically, the left-right relationship information estimation unit 183 is identical to the left-right relationship information estimation unit 182, except that the left-right time difference τ and the left-right time difference code Cτ may not be obtained and output.
[0276] For example, regarding the max to τ min The number of candidate samples τ cand The left-right relationship information estimation unit 183 compares the sample sequence of the input audio signal of the left channel with the sample sequence shifted backward by each candidate sample number τ. cand The correlation value γ of the sample sequence of the input sound signal of the right channel at the position cand The maximum value in is obtained as the left and right correlation coefficient γ and output, and τ when the correlation value is the maximum value cand When the value is positive, the information indicating the left channel is obtained as the preceding channel information and output. When the correlation value is the maximum value, τ cand When the correlation value is the maximum value, the information indicating the right channel is obtained and output as the preceding channel information. cand When it is 0, the left-right relationship information estimation unit 183 may obtain and output information indicating that the left channel precedes as the preceding channel information, may obtain and output information indicating that the right channel precedes as the preceding channel information, or may obtain and output information indicating that neither channel precedes as the preceding channel information.
[0277] <Third embodiment>
[0278] For an encoding device that stereo encodes the input sound signals of each channel instead of the differential signals of each channel, a structure that considers the relationship between the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmix signal can also be adopted. This method is described as a third embodiment.
[0279] <<Encoding Device 104>>
[0280] like Figure 16 As shown, the encoding device 104 of the third embodiment includes: a left-right relationship information estimation unit 183, a downmixing unit 112, a mono encoding unit 160, and a stereo encoding unit 174. The encoding device 104 of the third embodiment performs the following steps on each frame: Figure 17 The processing of step S183, step S112, step S160, and step S174 are illustrated. Hereinafter, the encoding device 104 according to the third embodiment will be described with reference to the description of the second embodiment as appropriate.
[0281] [Left-Right Relationship Information Estimation Unit 183]
[0282] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment. The left-right relationship information estimation unit 183 receives inputs of the left-channel input audio signal and the right-channel input audio signal from the encoding device 104. Based on the received left-channel input audio signal and right-channel input audio signal, the left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the left-channel input audio signal and the right-channel input audio signal, and preceding channel information indicating which of the left-channel and right-channel input audio signals precedes (step S183).
[0283] [Downmixing Unit 112]
[0284] The downmixing unit 112 is the same as the downmixing unit 112 of the second embodiment. Input to the downmixing unit 112 are the left channel input audio signal input to the encoding device 104, the right channel input audio signal input to the encoding device 104, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 183, and the preceding channel information output by the left-right relationship information estimation unit 183. The downmixing unit 112 performs a weighted average of the left and right channel input audio signals to generate a downmixed signal, and outputs the downmixed signal so that the larger the left-right correlation coefficient γ, the larger the preceding channel input audio signal of the left and right channel input audio signals (step S112).
[0285] For example, when the sampling number is set to t and the input audio signal of the left channel is set to x L (t), set the input sound signal of the right channel to x R (t), set the downmix signal to x M At (t), when the preceding channel information indicates that the left channel precedes, the downmixing unit 112 calculates the value of x for each sample number t. M (t)=((1+γ) / 2)×x L (t)+((1-γ) / 2)×x R (t) A downmix signal is obtained. When the preceding channel information indicates that the right channel precedes, for each sample number t, x M (t)=((1-γ) / 2)×x L (t)+((1+γ) / 2)×x R (t) A downmix signal is obtained. When the preceding channel information indicates that no channel is preceded, for each sample number t, x M (t)=(x L (t)+x R (t)) / 2 to obtain the downmixed signal.
[0286] [Monaural encoding unit 160]
[0287] The mono encoding unit 160 is the same as the mono encoding unit 160 of the second embodiment. The downmix signal output by the downmix unit 112 is input to the mono encoding unit 160. The mono encoding unit 160 encodes the input downmix signal to obtain a mono code CM and outputs it (step S160). The mono encoding unit 160 can also use any encoding method, for example, an encoding method such as the 3GPP EVS standard. The encoding method can be an encoding method that performs encoding processing independently from the stereo encoding unit 174 described later, that is, an encoding method that performs encoding processing without using the stereo code CS' obtained by the stereo encoding unit 174 or information obtained in the encoding processing performed by the stereo encoding unit 174, or an encoding method that performs encoding processing using the stereo code CS' obtained by the stereo encoding unit 174 or information obtained in the encoding processing performed by the stereo encoding unit 174.
[0288] [Stereo encoding unit 174]
[0289] The left channel input audio signal and the right channel input audio signal input into the encoding device 104 are input into the stereo encoding unit 174. The stereo encoding unit 174 encodes the input left channel input audio signal and the input right channel input audio signal to generate and output a stereo code CS' (step S174). The stereo encoding unit 174 may use any encoding method, for example, a stereo encoding method compatible with the stereo decoding method of the MPEG-4 AAC standard, an encoding method that independently encodes the input left channel input audio signal and the input right channel input audio signal, or a combined encoding method of all the encoding methods may be used as the stereo code CS'. The encoding method can be an encoding method that performs encoding processing independently from the mono encoding unit 160, that is, an encoding method that performs encoding processing without using the mono code CM obtained in the mono encoding unit 160 or the information obtained in the encoding processing performed by the mono encoding unit 160, or it can be an encoding method that performs encoding processing using the mono code CM obtained in the mono encoding unit 160 and the information obtained in the encoding processing performed by the mono encoding unit 160.
[0290] <Fourth embodiment>
[0291] As can be seen from the description of the above embodiments, if the encoding device encodes at least a downmixed signal obtained from the input sound signal of the left channel and the input sound signal of the right channel to obtain the encoding, then no matter what type of encoding device is used, a structure can be adopted to obtain the downmixed signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel. In addition, not limited to the encoding device, if the signal processing device performs signal processing on at least a downmixed signal obtained from the input sound signal of the left channel and the input sound signal of the right channel to obtain the signal processing result, then no matter what type of signal processing device is used, a structure can be adopted to obtain the downmixed signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel. Moreover, as a downmixing device used in the upstream stage of these encoding devices and signal processing devices, a structure can also be adopted to obtain the downmixed signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel. These methods will be described as the fourth embodiment.
[0292] <<Sound Signal Coding Device 105>>
[0293] like Figure 18 As shown, the audio signal encoding device 105 of the fourth embodiment includes: a left-right relationship information estimation unit 183, a downmixing unit 112, and an encoding unit 195. The audio signal encoding device 105 of the fourth embodiment performs the following steps for each frame: Figure 19 The processing of step S183, step S112, and step S195 is illustrated. Hereinafter, the sound signal coding apparatus 105 according to the fourth embodiment will be described with reference to the description of the second embodiment as appropriate.
[0294] [Left-Right Relationship Information Estimation Unit 183]
[0295] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment. Based on the input sound signal of the left channel and the input sound signal of the right channel, the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and the preceding channel information, which is information indicating which of the input sound signal of the left channel and the input sound signal of the right channel precede, are obtained and output (step S183).
[0296] [Downmixing Unit 112]
[0297] The downmixing unit 112 is the same as the downmixing unit 112 of the second embodiment, and performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmixed signal and outputs it so that the larger the left-right correlation coefficient γ is, the more the input sound signal of the preceding channel of the left channel and the right channel is included in the downmixed signal (step S112).
[0298] [Coding unit 195]
[0299] At least the downmixed signal output by the downmixing unit 112 is input to the encoding unit 195. The encoding unit 195 encodes at least the input downmixed signal to obtain a sound signal code and outputs it (step S195). The encoding unit 195 may also encode the input sound signal of the left channel and the input sound signal of the right channel, or may include the code obtained by the encoding in the sound signal code and output it. In this case, Figure 18 As indicated by the middle dotted lines, the input audio signal of the left channel and the input audio signal of the right channel are also input to the encoding unit 195 .
[0300] <<Sound Signal Processing Device 305>>
[0301] like Figure 20 As shown, the audio signal processing apparatus 305 of the fourth embodiment includes: a left-right relationship information estimation unit 183, a downmixing unit 112, and a signal processing unit 315. The audio signal processing apparatus 305 of the fourth embodiment performs the following steps on each frame: Figure 21 The processing of step S183, step S112, and step S315 are exemplified. The following describes the differences between the sound signal processing device 305 of the fourth embodiment and the sound signal coding device 105 of the fourth embodiment.
[0302] [Signal processing unit 315]
[0303] At least the downmixed signal output by the downmixing unit 112 is input to the signal processing unit 315. The signal processing unit 315 performs signal processing on at least the input downmixed signal, obtains a signal processing result, and outputs it (step S315). The signal processing unit 315 may also perform signal processing on the input audio signal of the left channel and the input audio signal of the right channel to obtain a signal processing result. In this case, Figure 20As indicated by the middle dashed line, the left and right channel input audio signals are also input to the signal processing unit 315. The signal processing unit 315 may, for example, perform signal processing on the input audio signals of each channel using a downmix signal to obtain output audio signals for each channel. Alternatively, the signal processing may be performed on decoded left and right channel audio signals obtained by decoding the code CS' obtained by the stereo encoder 174 of the third embodiment using a decoding device having a decoding unit corresponding to the stereo encoder 174. In other words, the left and right channel input audio signals input to the audio signal processing unit 305 are not necessarily digital audio signals or audio signals obtained by collecting the code CS' using two microphones and performing A / D conversion. The left and right channel input audio signals input to the audio signal processing unit 305 may be decoded left and right channel audio signals obtained by decoding the code, or may be any audio signals obtained as long as they are two-channel stereo audio signals.
[0304] In the case where the left channel input audio signal and the right channel input audio signal input to the audio signal processing device 305 are decoded left channel audio signals and right channel decoded audio signals obtained by decoding the coded audio signals by other devices, one or both of the left-right correlation coefficient γ and the preceding channel information obtained by the left-right relationship information estimation unit 183 may be obtained by other devices. In the case where one or both of the left-right correlation coefficient γ and the preceding channel information are obtained by other devices, Figure 20 As indicated by the dotted line, either or both of the left-right correlation coefficient γ and the preceding channel information obtained by other means may be input to the sound signal processing device 305. In this case, the left-right relationship information estimation unit 183 only needs to obtain the left-right correlation coefficient γ or the preceding channel information that has not been input to the sound signal processing device 305. If both the left-right correlation coefficient γ and the preceding channel information are input to the sound signal processing device 305, the sound signal processing device 305 may not include the left-right relationship information estimation unit 183 and may not perform step S183. That is, if Figure 20 As indicated by the two-dot chain line, the audio signal processing device 305 includes a left-right relationship information acquisition unit 185. This unit obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the left and right channel input audio signals, and preceding channel information, which indicates which of the left and right channel input audio signals preceded the previous one (step S185). The left-right relationship information estimation unit 183 and step S183 in each of the aforementioned devices can also be considered part of the left-right relationship information acquisition unit 185 and step S185.
[0305] <<Sound Signal Downmixing Device 405>>
[0306] like Figure 22 As shown, the audio signal downmixing device 405 of the fourth embodiment includes: a left-right relationship information acquisition unit 185 and a downmixing unit 112. The audio signal downmixing device 405 performs the following steps for each frame: Figure 23 The processing of steps S185 and S112 is illustrated. The sound signal downmixing device 405 will be described below with reference to the description of the second embodiment as appropriate. Furthermore, similarly to the sound signal processing device 305, the left-channel and right-channel input sound signals input to the sound signal downmixing device 405 may be digital sound signals or sound signals collected by two microphones and obtained by A / D conversion, or may be decoded left-channel and right-channel sound signals obtained by decoding encoded sound. Any sound signals may be used as long as they are stereo two-channel sound signals.
[0307] [Left-right relationship information acquisition unit 185]
[0308] The left-right relationship information acquisition unit 185 obtains and outputs a left-right correlation coefficient γ, which is a correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and preceding channel information, which is information indicating which of the input sound signal of the left channel and the input sound signal of the right channel precedes (step S185).
[0309] When both the left-right correlation coefficient γ and the preceding channel information are obtained by other means, as shown in FIG. Figure 22 As indicated by the dot-dash line, the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ and the preceding channel information input to the audio signal downmixing device 405 from another device, and outputs them to the downmixing unit 112 .
[0310] In the case where both the left and right correlation coefficient γ and the preceding channel information are not obtained by other means, such as Figure 22 As indicated by the middle dashed line, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. Similar to the left-right relationship information estimation unit 183 of the second embodiment, the left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and preceding channel information from the left-channel input audio signal and the right-channel input audio signal, and outputs the obtained information to the downmixing unit 112.
[0311] If neither the left-right correlation coefficient γ nor the preceding channel information is obtained by other means, as in Figure 22As shown by the middle dotted line, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 of the left-right relationship information acquisition unit 185, similar to the left-right relationship information estimation unit 183 of the second embodiment, obtains the left-right correlation coefficient γ or the preceding channel information not obtained by other devices from the left-channel input sound signal and the right-channel input sound signal, and outputs it to the downmixing unit 112. Regarding the left-right correlation coefficient γ obtained by other devices or the preceding channel information obtained by other devices, as shown in FIG. Figure 22 As indicated by the dot-dash line in the figure, the left-right relationship information acquisition unit 185 outputs the left-right correlation coefficient γ or the preceding channel information input from another device to the downmixing unit 112 .
[0312] [Downmixing Unit 112]
[0313] The downmixing unit 112 is the same as the downmixing unit 112 of the second embodiment. Based on the preceding channel information and the left-right correlation coefficient obtained by the left-right relationship information obtaining unit 185, the downmixing unit 112 performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmixed signal and outputs the downmixed signal so that the larger the left-right correlation coefficient γ is, the more the input sound signal of the preceding channel of the input sound signal of the left channel and the input sound signal of the right channel is included in the downmixed signal (step S112).
[0314] For example, when the sampling number is set to t and the input audio signal of the left channel is set to x L (t), set the input sound signal of the right channel to x R (t), set the downmix signal to x M At (t), when the preceding channel information indicates that the left channel precedes, the downmixing unit 112 calculates the value of x for each sample number t. M (t)=((1+γ) / 2)×x L (t)+((1-γ) / 2)×x R (t) A downmix signal is obtained. When the preceding channel information indicates that the right channel precedes, for each sample number t, x M (t)=((1-γ) / 2)×x L (t)+((1+γ) / 2)×x R (t) A downmix signal is obtained. When the preceding channel information indicates that no channel is preceded, for each sample number t, x M (t)=(x L (t)+x R (t)) / 2 to obtain the downmixed signal.
[0315] <Program and Recording Medium>
[0316] The processing of each part of the above-mentioned encoding device, decoding device, audio signal encoding device, audio signal processing device and audio signal downmixing device can also be realized by a computer. In this case, the processing content of the function to be possessed by each device is described in a program. Then, Figure 24 The storage unit 1020 of the computer 1000 shown reads the program and operates the arithmetic processing unit 1010, the input unit 1030, the output unit 1040, etc., thereby realizing various processing functions in the above-mentioned devices on the computer.
[0317] The program describing the processing contents can be recorded on a computer-readable recording medium. The computer-readable recording medium is, for example, a non-transitory recording medium, specifically, a magnetic recording device, an optical disk, or the like.
[0318] The program can be distributed, for example, by selling, transferring, or renting a removable recording medium such as a DVD or CD-ROM that stores the program. Furthermore, the program can be distributed by storing it in a storage device of a server computer and transferring it from the server computer to other computers via a network.
[0319] A computer executing such a program first temporarily stores a program recorded on a transportable recording medium or transferred from a server computer in its own non-transitory storage device, namely, auxiliary recording unit 1050. Then, when executing a process, the computer causes storage unit 1020 to read the program stored in its own non-transitory storage device, namely, auxiliary recording unit 1050, and executes the process according to the read program. Alternatively, as another method of executing the program, the computer may directly read the program from the transportable recording medium into storage unit 1020 and execute the process according to the program. Furthermore, the computer may sequentially execute the process according to the received program each time the program is transferred from the server computer to the computer. Furthermore, a configuration may be adopted in which the above-described process is executed by a so-called ASP (Application Service Provider)-type service, which implements the processing function solely through execution instructions and result acquisition, without transferring the program from the server computer to the computer. Furthermore, the program in this embodiment includes information for the computer's processing, namely information based on the program (data that, while not direct instructions to the computer, specifies the nature of the computer's processing, etc.).
[0320] Furthermore, in this embodiment, the present apparatus is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0321] Furthermore, it is apparent that appropriate changes can be made without departing from the spirit of the present invention.
Claims
1. A program for downmixing a sound signal, wherein the downmixing step obtains a signal obtained by mixing a first channel input sound signal and a second channel input sound signal, namely, a downmixed signal, the program comprising the following processing: obtaining preceding channel information indicating which of the first-channel input audio signal and the second-channel input audio signal precedes, and a coefficient indicating correlation between the first-channel input audio signal and the second-channel input audio signal; and The downmix signal is obtained by performing a weighted average of the first-channel input sound signal and the second-channel input sound signal based on the preceding channel information and the coefficient representing the correlation, so that a larger coefficient representing the correlation includes a greater proportion of the input sound signal of the preceding channel between the first-channel input sound signal and the second-channel input sound signal.
Citation Information
Patent Citations
Sound coding device and sound coding method
WO2006070751A1