Audio signal downmixing method, audio signal downmixing device, program
The audio signal downmixing method enhances coding efficiency by determining the leading channel and applying a correlation coefficient to generate a monaural signal from two-channel audio signals, addressing the limitations of existing methods in signal processing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2025-03-14
- Publication Date
- 2026-07-29
AI Technical Summary
Existing methods for obtaining a monaural signal from a two-channel audio signal are inadequate for signal processing such as encoding, as they do not effectively utilize the left and right channel signals to enhance coding efficiency.
An audio signal downmixing method that determines the leading channel and applies a correlation coefficient to weight the signals, incorporating a time difference to generate a downmix signal, which is then processed to obtain a monaural signal suitable for encoding.
The method enables the generation of a monaural signal suitable for encoding from a two-channel audio signal, improving coding efficiency and signal processing capabilities.
Smart Images

Figure 0007896717000025 
Figure 0007896717000026 
Figure 0007896717000027
Abstract
Description
[Technical Field]
[0001] This invention relates to a technique for obtaining a monaural audio signal from a two-channel audio signal in order to encode an audio signal in monaural, encode an audio signal using a combination of monaural and stereo encoding, process an audio signal in monaural, or perform signal processing using a monaural audio signal with a stereo audio signal. [Background technology]
[0002] Patent Document 1 describes a technique for obtaining a monaural audio signal from a 2-channel audio signal and for embedding and decoding the 2-channel audio signal and the monaural audio signal. Patent Document 1 discloses a technique for obtaining a monaural signal by averaging the input left channel audio signal and the input right channel audio signal for each corresponding sample, encoding the monaural signal (monaural encoding) to obtain a monaural code, decoding the monaural code (monaural decoding) to obtain a monaural local decoded signal, and encoding the difference (predicted residual signal) between the input audio signal and the predicted signal obtained from the monaural local decoded signal for both the left and right channels. In the technology described in Patent Document 1, for each channel, a signal with an amplitude ratio obtained by adding a delay to the monaural local decoded signal is used as a prediction signal. A prediction signal with a delay and amplitude ratio that minimizes the error between the input sound signal and the prediction signal is selected, or a prediction signal with a delay difference and amplitude ratio that maximizes the cross-correlation between the input sound signal and the monaural local decoded signal is used. The prediction signal is then subtracted from the input sound signal to obtain a prediction residual signal, and this prediction residual signal is used as the target for encoding / decoding, thereby suppressing the degradation of the sound quality of the decoded sound signal for each channel. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] WO2006-070751 publication [Overview of the project] [Problems that the invention aims to solve]
[0004] The technology described in Patent Document 1 improves the coding efficiency of each channel by optimizing the delay and amplitude ratio applied to the monaural locally decoded signal when obtaining the prediction signal. However, in the technology described in Patent Document 1, the monaural locally decoded signal is obtained by encoding and decoding a monaural signal obtained by averaging the sound signals of the left channel and the right channel. In other words, the technology described in Patent Document 1 has the problem that it does not employ any method to obtain a monaural signal useful for signal processing such as coding from two-channel sound signals. The present invention aims to provide a technique for obtaining a monaural signal useful for signal processing such as encoding from a two-channel audio signal. [Means for solving the problem]
[0005] One aspect of the present invention is an audio signal downmixing method for obtaining a downmix signal, which is a signal obtained by mixing a first channel input audio signal and a second channel input audio signal.
[0006] This audio signal downmixing method includes the steps of: obtaining a code indicating which of the first channel input audio signal and the second channel input audio signal is leading; and a downmixing step of obtaining a downmix signal from the leading channel and the other channel based on a degree determined based on the code and a correlation coefficient which is a coefficient indicating the magnitude of the correlation between the first channel input audio signal and the second channel input audio signal.
[0007] In the downmix step, the signal that precedes the first channel input audio signal is given a greater weight than the other signal. Furthermore, the process includes a step of obtaining a time difference between the first channel input sound signal and the second channel input sound signal, wherein the sign indicates the positive or negative value of the time difference. . [Effects of the Invention]
[0008] According to the present invention, a monaural signal useful for signal processing such as encoding can be obtained from a two-channel audio signal. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing examples of encoding devices of the first reference embodiment and the second embodiment. [Figure 2] This is a flowchart illustrating an example of the processing of the first reference form of encoding device. [Figure 3] This is a block diagram showing an example of a decoding device in the first reference form. [Figure 4] This is a flowchart illustrating an example of the processing of a decoding device in the first reference configuration. [Figure 5] This flowchart shows an example of the processing of the left channel subtractive gain estimation unit and the right channel subtractive gain estimation unit of the first reference form. [Figure 6] This flowchart shows an example of the processing of the left channel subtractive gain estimation unit and the right channel subtractive gain estimation unit of the first reference form. [Figure 7] This flowchart shows an example of the processing of the left channel subtractive gain decoder and the right channel subtractive gain decoder in the first reference configuration. [Figure 8] This flowchart shows an example of the processing of the left channel subtractive gain estimation unit and the right channel subtractive gain estimation unit of the first reference form. [Figure 9] This flowchart shows an example of the processing of the left channel subtractive gain estimation unit and the right channel subtractive gain estimation unit of the first reference form. [Figure 10] This is a block diagram showing examples of the encoding device of the second reference embodiment and the first embodiment. [Figure 11] This flowchart shows an example of the processing of the second reference form of encoding device. [Figure 12] This is a block diagram showing an example of a second reference form of decoding device. [Figure 13] This flowchart shows an example of the processing of a decoding device in the second reference form. [Figure 14] This is a flowchart showing an example of the processing of the encoding device according to the first embodiment. [Figure 15] This is a flowchart showing an example of the processing of the encoding device according to the second embodiment. [Figure 16] This is a block diagram showing an example of an encoding device according to the third embodiment. [Figure 17]This is a flowchart showing an example of the processing of the encoding device according to the third embodiment. [Figure 18] This is a block diagram showing an example of an audio signal coding device according to the fourth embodiment. [Figure 19] This is a flowchart showing an example of the processing of the sound signal encoding device according to the fourth embodiment. [Figure 20] This is a block diagram showing an example of an audio signal processing device according to the fourth embodiment. [Figure 21] This is a flowchart showing an example of processing in the sound signal processing device of the fourth embodiment. [Figure 22] This is a block diagram showing an example of an audio signal downmixing device according to the fourth embodiment. [Figure 23] This is a flowchart showing an example of the processing of the audio signal downmixing device according to the fourth embodiment. [Figure 24] This figure shows an example of the functional configuration of a computer that implements each device in the embodiment of the present invention. [Modes for carrying out the invention]
[0010] First, let's explain the notation in the specification. The superscript "^" for a character x, such as ^x, should ideally be placed directly above the character x. However, due to the constraints of notation in the specification, it is sometimes written as ^x. <First reference form> Before describing embodiments of the invention, the first and second reference embodiments will be described, which will be the encoding and decoding devices that form the basis for carrying out the invention of the second embodiment and the invention of the first embodiment. In the specification and claims, the encoding device may be referred to as an audio signal encoding device, the encoding method as an audio signal encoding method, the decoding device as an audio signal decoding device, and the decoding method as an audio signal decoding method.
[0011] ≪Encoding device 100≫ The first reference embodiment of the encoding device 100, as shown in Figure 1, includes a downmix unit 110, a left channel subtractive gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtractive gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, and a stereo encoding unit 170. The encoding device 100 encodes the input 2-channel stereo time-domain sound signal in frame units of a predetermined time length, for example, 20 ms, to obtain and output a monaural code CM, a left channel subtractive gain code Cα, a right channel subtractive gain code Cβ, and a stereo code CS, which will be described later. The 2-channel stereo time-domain sound signal input to the encoding device is, for example, a digital audio signal or sound signal obtained by AD conversion after sound such as speech or music is picked up by two microphones, and consists of a left channel input sound signal and a right channel input sound signal. The codes output by the encoding device, namely the monaural code CM, the left channel subtractive gain code Cα, the right channel subtractive gain code Cβ, and the stereo code CS, are input to the decoding device. For each frame, the encoding device 100 performs the processing from steps S110 to S170 as illustrated in Figure 2.
[0012] [Downmix section 110] The downmix unit 110 receives the left channel input sound signal and the right channel input sound signal input to the encoding device 100. The downmix unit 110 obtains a downmix signal, which is a signal obtained by mixing the left channel input sound signal and the right channel input sound signal from the input left channel input sound signal and the right channel input sound signal, and outputs it (step S110).
[0013] For example, if the number of samples per frame is T, the downmix unit 110 receives the input audio signal x of the left channel, which is input to the encoding device 100 on a frame-by-frame basis. L (1), x L (2), ..., x L (T) and the input audio signal x of the right channel R (1), x R (2), ..., x R(T) is input. Here, T is a positive integer. For example, if the frame length is 20 ms and the sampling frequency is 32 kHz, then T is 640. The downmixing unit 110 obtains and outputs a series based on the average value of the sample values for corresponding samples of the input sound signal of the input left channel and the input sound signal of the right channel as the downmixing signal x M (1), x M (2),..., x M (T). That is, when each sample number is t, x M (t) = (x L (t) + x R (t)) / 2.
[0014] [Left-channel subtraction gain estimation unit 120] The left-channel subtraction gain estimation unit 120 is input with the input sound signal x L (1), x L (2),..., x L (T) of the left channel input to the encoding device 100, and the downmixing signal x M (1), x M (2),..., x M (T) output by the downmixing unit 110. The left-channel subtraction gain estimation unit 120 obtains and outputs the left-channel subtraction gain α and the left-channel subtraction gain sign Cα, which is a sign representing the left-channel subtraction gain α, from the input sound signal of the input left channel and the downmixing signal (step S120). The left-channel subtraction gain estimation unit 120 obtains the left-channel subtraction gain α and the left-channel subtraction gain sign Cα by a well-known method exemplified by the method of obtaining the amplitude ratio g in Patent Document 1 or the method of encoding the amplitude ratio g, or by a method based on the newly invented principle of minimizing the quantization error. The principle of minimizing the quantization error and the method based on this principle will be described later.
[0015] [Left-channel signal subtraction unit 130] The left-channel signal subtraction unit 130 is input with the input sound signal x L (1), x L (2),..., x L(T) and the downmix signal x output by the downmix unit 110 M (1), x M (2), ..., x M (T) and the left channel subtractive gain α output by the left channel subtractive gain estimation unit 120 are input. The left channel signal subtraction unit 130 subtracts the sample value x of the downmix signal for each corresponding sample t. M (t) is multiplied by the left channel subtraction gain α, resulting in α × x M (t) The sample value x of the input sound signal for the left channel L The value x subtracted from (t) L (t)-α×x M The sequence by (t) is the left channel difference signal y L (1), y L (2), ..., y L (T) is obtained and output (step S130). That is, y L (t) = x L (t)-α×x M (t) In the encoding device 100, in order to avoid delays and computational processing required to obtain the local decoded signal, the left channel signal subtraction unit 130 does not use the quantized downmix signal, which is the local decoded signal of the monaural encoding, but rather the unquantized downmix signal x obtained by the downmix unit 110. M (t) may be used. However, if the left channel subtraction gain estimation unit 120 obtains the left channel subtraction gain α by a well-known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing quantization error, the encoding device 100 shall be provided with means for obtaining a local decoded signal corresponding to the monaural code CM after the monaural encoding unit 160 or within the monaural encoding unit 160, and the left channel signal subtraction unit 130 shall obtain the downmix signal x M (1), x M (2), ..., x M Instead of (T), a quantized downmix signal ^x, which is the local decoded signal for monaural coding, is used, similar to conventional encoding devices such as those described in Patent Document 1. M (1), ^x M (2), ..., ^x M The left channel difference signal may be obtained using (T).
[0016] [Right channel subtractive gain estimation unit 140] The right channel subtractive gain estimation unit 140 receives the input audio signal x of the right channel input to the encoding device 100. R (1), x R (2), ..., x R (T) and the downmix signal x output by the downmix unit 110 M (1), x M (2), ..., x M (T) and are input. The right channel subtractive gain estimation unit 140 obtains and outputs the right channel subtractive gain β and the right channel subtractive gain code Cβ, which is a code representing the right channel subtractive gain β, from the input sound signal and downmix signal of the input right channel (step S140). The right channel subtractive gain estimation unit 140 obtains the right channel subtractive gain β and the right channel subtractive gain code Cβ by a well-known method, such as the method for determining the amplitude ratio g and the method for encoding the amplitude ratio g as exemplified in Patent Document 1, or by a newly devised method based on a principle for minimizing quantization error. The principle for minimizing quantization error and the method based on this principle will be described later.
[0017] [Right channel signal subtraction unit 150] The right channel signal subtraction unit 150 receives the input sound signal x of the right channel input to the encoding device 100. R (1), x R (2), ..., x R (T) and the downmix signal x output by the downmix unit 110 M (1), x M (2), ..., x M (T) and the right channel subtractive gain β output by the right channel subtractive gain estimation unit 140 are input. The right channel signal subtraction unit 150 subtracts the sample value x of the downmix signal for each corresponding sample t. M (t) is multiplied by the right channel subtraction gain β, resulting in β × x M (t) The sample value x of the input audio signal for the right channel R The value x subtracted from (t) R (t)-β×xM The sequence by (t) is the right channel difference signal y R (1), y R (2), ..., y R (T) is obtained and output (step S150). That is, y R (t) = x R (t)-β×x M (t) The right channel signal subtraction unit 150, similar to the left channel signal subtraction unit 130, does not require delay or computational processing to obtain the local decoded signal in the encoding device 100, but rather the unquantized downmix signal x obtained by the downmix unit 110 is used instead of the quantized downmix signal which is the local decoded signal of the monaural encoding. M (t) may be used. However, if the right channel subtraction gain estimation unit 140 obtains the right channel subtraction gain β by a well-known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing quantization error, the encoding device 100 shall be provided with means for obtaining a local decoded signal corresponding to the monaural code CM after the monaural encoding unit 160 or within the monaural encoding unit 160, and in the right channel signal subtraction unit 150, similar to the left channel signal subtraction unit 130, the downmix signal x M (1), x M (2), ..., x M Instead of (T), a quantized downmix signal ^x, which is the local decoded signal for monaural coding, is used, similar to conventional encoding devices such as those described in Patent Document 1. M (1), ^x M (2), ..., ^x M The right channel difference signal may be obtained using (T).
[0018] [Monaural encoding section 160] The monaural encoding unit 160 receives the downmix signal x output by the downmix unit 110. M (1), x M (2), ..., x M (T) is input. The monaural encoding unit 160 encodes the input downmix signal using a predetermined encoding method b MEncode with b bits to obtain and output a monaural code CM (step S160). That is, from the downmixed signals x M (1), x M (2),..., x M (T) of the input T samples, obtain and output a b-bit monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used. M x(1) M x(2) M ... x(T), obtain and output a b-bit monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used. M Encode with b bits to obtain and output a monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used.
[0019] [Stereo Encoder 170] The left channel difference signals y L (1), y L (2),..., y L (T) output by the left channel signal subtraction unit 130 and the right channel difference signals y R (1), y R (2),..., y R (T) output by the right channel signal subtraction unit 150 are input to the stereo encoder 170. The stereo encoder 170 encodes the input left channel difference signal and right channel difference signal with a total of b bits using a predetermined encoding method to obtain and output a stereo code CS (step S170). That is, from the input left channel difference signals y L (1), y L (2),..., y L (T) of T samples and the input right channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left channel difference signal and right channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. L y(1) L y(2) L ... y(T) and the right channel difference signals y R (1), y R (2),..., y R (T) output by the right channel signal subtraction unit 150 are input to the stereo encoder 170. The stereo encoder 170 encodes the input left channel difference signal and right channel difference signal with a total of b bits using a predetermined encoding method to obtain and output a stereo code CS (step S170). R y(1) R y(2) R ... y(T) and the input left channel difference signals y L (1), y L (2),..., y L (T) of T samples, obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left channel difference signal and right channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. s Encode with b bits to obtain and output a stereo code CS (step S170). That is, from the input left channel difference signals y L (1), y L (2),..., y L (T) of T samples L y(1) L y(2) L ... y(T) and the input right channel difference signals y R (1), y R (2),..., y R (T) of T samples R y(1) R y(2) R ... y(T), obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left channel difference signal and right channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. S Encode with b bits to obtain and output a stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left channel difference signal and right channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding.
[0020] When encoding the input left-channel differential signal and right-channel differential signal independently, the stereo encoding unit 170 encodes the left-channel differential signal with b L bits and encodes the right-channel differential signal with b R bits. That is, the stereo encoding unit 170 obtains the b L (1), y L (2),..., y L (T) of the left-channel differential signal of the input T samples to obtain the b L bit left-channel differential code CL, and obtains the b R (1), y R (2),..., y R (T) of the right-channel differential signal of the input T samples to obtain the b R bit right-channel differential code CR, and outputs the combination of the left-channel differential code CL and the right-channel differential code CR as the stereo code CS. Here, the sum of the b L bits and the b R bits is b S bits.
[0021] When encoding the input left-channel differential signal and right-channel differential signal together in one encoding method, the stereo encoding unit 170 encodes the left-channel differential signal and the right-channel differential signal with a total of b S bits. That is, the stereo encoding unit 170 obtains and outputs the b L (1), y L (2),..., y L (T) of the left-channel differential signal of the input T samples and the b R (1), y R (2),..., y R (T) of the right-channel differential signal of the input T samples to obtain the b S bit stereo code CS.
[0022] <<Decoder device 200>> The decoding device 200 of the first reference embodiment, as shown in Figure 3, includes a monaural decoding unit 210, a stereo decoding unit 220, a left channel subtractive gain decoding unit 230, a left channel signal summing unit 240, a right channel subtractive gain decoding unit 250, and a right channel signal summing unit 260. The decoding device 200 decodes the input monaural code CM, left channel subtractive gain code Cα, right channel subtractive gain code Cβ, and stereo code CS in frame units of the same time length as the corresponding encoding device 100, and outputs a time-domain decoded audio signal (left channel decoded audio signal and right channel decoded audio signal, described later) in a frame unit. The decoding device 200 may also output a time-domain decoded audio signal (a monaural decoded audio signal, described later) in monaural time domain, as shown by the dashed line in Figure 3. The decoded audio signal output by the decoding device 200 is made audible by, for example, DA conversion and playback by a speaker. The decoding device 200 performs the processing from step S210 to step S260, as illustrated in Figure 4, for each frame.
[0023] [Monaural decoding unit 210] The monaural decoding unit 210 receives the monaural code CM input to the decoding device 200. The monaural decoding unit 210 decodes the input monaural code CM using a predetermined decoding method to produce a monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M (T) is obtained and output (step S210). As the predetermined decoding method, the decoding method corresponding to the encoding method used in the monaural encoding unit 160 of the corresponding encoding device 100 is used. The number of bits of the monaural code CM is b M That is the case.
[0024] [Stereo decoding unit 220] The stereo decoding unit 220 receives the stereo code CS input to the decoding device 200. The stereo decoding unit 220 decodes the input stereo code CS using a predetermined decoding method to obtain the left channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L (T) and the right channel decoded difference signal ^y R(1), ^y R (2), ..., ^y R (T) and are obtained and output (step S220). As the predetermined decoding method, the decoding method corresponding to the encoding method used in the stereo encoding unit 170 of the corresponding encoding device 100 is used. The total number of bits of the stereo code CS is b S That is the case.
[0025] [Left channel subtractive gain decoding unit 230] The left channel subtractive gain decoding unit 230 receives the left channel subtractive gain code Cα input to the decoding device 200. The left channel subtractive gain decoding unit 230 decodes the left channel subtractive gain code Cα to obtain and output the left channel subtractive gain α (step S230). The left channel subtractive gain decoding unit 230 decodes the left channel subtractive gain code Cα using a decoding method corresponding to the method used by the left channel subtractive gain estimation unit 120 of the corresponding encoding device 100 to obtain the left channel subtractive gain α. The method by which the left channel subtractive gain decoding unit 230 decodes the left channel subtractive gain code Cα to obtain the left channel subtractive gain α when the left channel subtractive gain estimation unit 120 of the corresponding encoding device 100 obtains the left channel subtractive gain α and the left channel subtractive gain code Cα using a method based on the principle of minimizing quantization error will be described later.
[0026] [Left channel signal summing unit 240] The left channel signal summer 240 receives the monaural decoded sound signal^x output by the monaural decoding unit 210. M (1), ^x M (2), ..., ^x M (T) and the left channel decoded difference signal ^y output by the stereo decoding unit 220. L (1), ^y L (2), ..., ^y L (T) and the left channel subtractive gain α output by the left channel subtractive gain decoding unit 230 are input. The left channel signal summing unit 240 adds the sample value ^y of the left channel decoded difference signal for each corresponding sample t. L (t) and the sample value of the monaural decoded sound signal ^x M α×^x is the value obtained by multiplying (t) by the left channel subtraction gain α.M (t) and the sum of the two values ^y L (t) + α × ^x M The sequence by (t) is the left channel decoded sound signal ^x L (1), ^x L (2), ..., ^x L (T) is obtained and output (step S240). That is, ^x L (t=^y L (t) + α × ^x M (t)
[0027] [Right channel subtractive gain decoding unit 250] The right channel subtractive gain decoding unit 250 receives the right channel subtractive gain code Cβ input to the decoding device 200. The right channel subtractive gain decoding unit 250 decodes the right channel subtractive gain code Cβ to obtain the right channel subtractive gain β and outputs it (step S250). The right channel subtractive gain decoding unit 250 decodes the right channel subtractive gain code Cβ using a decoding method corresponding to the method used by the right channel subtractive gain estimation unit 140 of the corresponding encoding device 100 to obtain the right channel subtractive gain β. The method by which the right channel subtractive gain decoding unit 250 decodes the right channel subtractive gain code Cβ to obtain the right channel subtractive gain β when the right channel subtractive gain estimation unit 140 of the corresponding encoding device 100 obtains the right channel subtractive gain β and the right channel subtractive gain code Cβ using a method based on the principle of minimizing quantization error will be described later.
[0028] [Right channel signal summing unit 260] The right channel signal summer 260 receives the monaural decoded sound signal^x output by the monaural decoding unit 210. M (1), ^x M (2), ..., ^x M (T) and the right channel decoded difference signal ^y output by the stereo decoding unit 220. R (1), ^y R (2), ..., ^y R (T) and the right channel subtractive gain β output by the right channel subtractive gain decoding unit 250 are input. The right channel signal summing unit 260 adds the sample value ^y of the right channel decoded difference signal for each corresponding sample t.R (t) and the sample value of the monaural decoded sound signal ^x M The value obtained by multiplying (t) by the right channel subtraction gain β is β×^x M (t) and the sum of the two values ^y R (t) + β × ^x M The sequence by (t) is the right channel decoded sound signal ^x R (1), ^x R (2), ..., ^x R (T) is obtained and output (step S260). That is, ^x R (t=^y R (t) + β × ^x M (t)
[0029] [Principle for minimizing quantization error] The principle for minimizing quantization error is explained below. When the left channel difference signal and the right channel difference signal input to the stereo coding unit 170 are combined and coded within a single coding scheme, the number of bits b used for coding the left channel difference signal is... L and the number of bits b used to encode the right channel difference signal R Although it may not be explicitly defined, below, the number of bits used to encode the left channel difference signal is b L The number of bits used to encode the right channel difference signal is b R We will explain based on this assumption. Furthermore, while we will primarily discuss the left channel below, the same principles apply to the right channel.
[0030] The above-described encoding device 100 processes the input audio signal x of the left channel. L (1), x L (2), ..., x L From each sample value of (T), the downmix signal x M (1), x M (2), ..., x M The left channel difference signal y consists of the values obtained by subtracting the values obtained by multiplying each sample value of (T) by the left channel subtraction gain α. L (1), y L (2), ..., y L (T) bL Encoded in bits, downmix signal x M (1), x M (2), ..., x M (T) b M Encoded in bits. Also, the decoding device 200 described above is b L The left channel decoded difference signal ^y from the bit sign L (1), ^y L (2), ..., ^y L (T) (hereinafter also referred to as the "quantized left channel difference signal") is decoded, b M Decoded mono sound signal from bit code ^x M (1), ^x M (2), ..., ^x M After decoding (T) (hereinafter also referred to as the "quantized downmix signal"), the quantized downmix signal ^x obtained by decoding is then processed. M (1), ^x M (2), ..., ^x M The quantized left channel difference signal ^y obtained by decoding is the value obtained by multiplying each sample value of (T) by the left channel subtraction gain α. L (1), ^y L (2), ..., ^y L Adding (T) to each sample value results in the left channel decoded sound signal, which is the left channel decoded sound signal^x. L (1), ^x L (2), ..., ^x L (T) is obtained. The encoding device 100 and the decoding device 200 should be designed such that the energy of the quantization error in the decoded sound signal of the left channel obtained in the above process is small.
[0031] The energy of the quantization error in the decoded signal obtained by encoding and decoding the input signal (hereinafter, for convenience, referred to as "quantization error caused by encoding") is, in most cases, roughly proportional to the energy of the input signal and tends to decrease exponentially with respect to the number of bits per sample used for encoding. Therefore, the average energy per sample of the quantization error caused by encoding the left channel difference signal is a positive number σ L2 Using this, the following equation (1-0-1) can be used to estimate the average energy per sample of the quantization error caused by coding the downmix signal, which is a positive number σ M 2 Using this, we can estimate it as shown in equation (1-0-2) below.
number
number
[0032] Let's assume here that the input audio signal x of the left channel L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x M Assume that the individual sample values are close enough that (T) can be considered to belong to the same sequence. For example, the input audio signal x of the left channel. L (1), x L (2), ..., x L (T) and the input signal x of the right channel R (1), x R (2), ..., x R Cases where (T) is obtained by capturing sound emitted from a sound source equidistant from two microphones in an environment with little background noise or reverberation fall under this condition. Under this condition, the left channel difference signal y L (1), y L (2), ..., y L Each sample value of (T) corresponds to the downmix signal x M (1), x M (2), ..., x M This is equivalent to the value obtained by multiplying each sample value of (T) by (1-α). Therefore, the energy of the left channel difference signal is (1-α) of the energy of the downmix signal. 2 Since it can be expressed as a multiple, the above σ L 2 The above σ M2 Using (1-α) 2 ×σ M 2 Since this can be replaced, the average energy per sample of the quantization error caused by encoding the left channel difference signal can be estimated as shown in equation (1-1) below.
number
number
[0033] Assuming that the quantization error resulting from the encoding of the left channel difference signal and the quantization error in the sequence of values obtained by multiplying each sample value of the decoded quantized downmix signal by the left channel subtraction gain α are uncorrelated, the average energy per sample of the quantization error in the decoded left channel sound signal can be estimated by the sum of equations (1-1) and (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error in the decoded left channel sound signal can be found as shown in equation (1-3) below.
number
[0034] In other words, the input audio signal x of the left channel L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x MUnder the condition that (T) is considered to be the same sequence, and each sample value is close enough, in order to minimize the quantization error of the decoded sound signal of the left channel, the left channel subtraction gain estimation unit 120 should calculate the left channel subtraction gain α using equation (1-3). The left channel subtraction gain α obtained by equation (1-3) is a value greater than 0 and less than 1, and is equal to the number of bits used for the two encodings, b L and b M When they are equal, it is 0.5, and the number of bits b for encoding the left channel difference signal. L The number of bits b for encoding the downmix signal M The more than this value, the closer it is to 0 than 0.5, and the number of bits b for encoding the downmix signal. M The number of bits b for encoding the left channel difference signal L The more than 0.5, the closer the value is to 1.
[0035] The same applies to the right channel, and the input audio signal x for the right channel R (1), x R (2), ..., x R (T) and downmix signal x M (1), x M (2), ..., x M In order to minimize the quantization error of the right channel decoded sound signal under the condition that each sample value is close enough that (T) can be considered to be the same sequence, the right channel subtractive gain estimation unit 140 should calculate the right channel subtractive gain β using the following equation (1-3-2).
number
[0036] Next, the input audio signal x of the left channel L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x M This section explains the principle for minimizing the energy of the quantization error in the decoded audio signal of the left channel, including the case where (T) cannot be considered to be the same sequence.
[0037] Left channel input audio signal x L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x M (T) Normalized inner product value r L This can be expressed by the following equation (1-4).
number
[0038] Left channel input audio signal x L (1), x L (2), ..., x L (T) is, for each sample number t, x L (t)=r L ×x M (t)+(x L (t)- r L ×x M (t)) can be decomposed as follows: where x L (t)- r L ×x M The sequence formed by each value of (t) is an orthogonal signal x L (1), x L (2), ..., x L If we denote it as (T), then according to this decomposition, each sample value y of the left channel difference signal L (t) = x L (t)-αx M (t) is the downmix signal x M (1), x M (2), ..., x M (T) Each sample value x M (t) is the normalized inner product r L and using the left channel subtractive gain α (r L The value obtained by multiplying by -α) (r L -α)×x M (t) and each sample value x of the orthogonal signal L '(t) and sum (r L -α)×x M (t)+x L This is equivalent to (t). Orthogonal signal x L (1), x L (2), ..., x L'(T) is the downmix signal x M (1), x M (2), ..., x M Because it exhibits orthogonality with respect to (T), i.e., the inner product is 0, the energy of the left channel difference signal is equal to the energy of the downmix signal (r L -α) 2 It is expressed as the sum of the multiplied value and the energy of the quadrature signal. Therefore, the left channel difference signal is b L The average energy per sample of the quantization error resulting from bit encoding is a positive number σ 2 Using this, we can estimate it as shown in equation (1-5) below.
number
[0039] Assuming that the quantization error resulting from the encoding of the left channel difference signal and the quantization error in the sequence of values obtained by multiplying each sample value of the quantized downmix signal obtained by decoding by the left channel subtraction gain α are uncorrelated, the average energy per sample of the quantization error in the decoded left channel sound signal can be estimated by the sum of equations (1-5) and (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error in the decoded left channel sound signal can be found as shown in equation (1-6) below.
number
[0040] In other words, in order to minimize the quantization error of the decoded sound signal of the left channel, the left channel subtraction gain estimation unit 120 should calculate the left channel subtraction gain α using equation (1-6). That is, considering the principle of minimizing the energy of this quantization error, the left channel subtraction gain α has a normalized inner product value r L b is the number of bits used for encoding. L and b MThe correction coefficient, which is a value determined by [the specified factor], should be used as the product of [the specified factor] and [the specified factor]. The correction coefficient is a value greater than 0 and less than 1, and the number of bits b for encoding the left channel difference signal. L and the number of bits b for encoding the downmix signal M When they are the same, it is 0.5, and the number of bits b for encoding the left channel difference signal. L The number of bits b for encoding the downmix signal M The more than 0.5, the closer it is to 0, and the number of bits b for encoding the left channel difference signal. L The number of bits b for encoding the downmix signal M The less than this value, the closer it is to 1 than to 0.5.
[0041] The same applies to the right channel; in order to minimize the quantization error of the decoded sound signal of the right channel, the right channel subtraction gain estimation unit 140 should calculate the right channel subtraction gain β using the following equation (1-6-2).
number
number
[0042] [Estimation and decoding of subtractive gain based on the principle of minimizing quantization error] This section describes specific examples of subtractive gain estimation and decoding based on the principle of minimizing the quantization error described above. In each example, the left channel subtractive gain estimation unit 120 and the right channel subtractive gain estimation unit 140, which estimate the subtractive gain in the encoding device 100, and the left channel subtractive gain decoding unit 230 and the right channel subtractive gain decoding unit 250, which decode the subtractive gain in the decoding device 200, will be described.
[0043] [Example 1] Example 1 shows the input audio signal x for the left channel. L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x M The principle for minimizing the quantization error energy of the decoded audio signal of the left channel, including the case where (T) cannot be considered to be the same sequence, and the input audio signal x of the right channel R (1), x R (2), ..., x R (T) and downmix signal x M (1), x M (2), ..., x M This is based on a principle that minimizes the energy of the quantization error in the right channel's decoded sound signal, including cases where (T) cannot be considered to be the same sequence.
[0044] [Left channel subtraction gain estimation unit 120] The left channel subtractive gain estimation unit 120 generates a candidate α for the left channel subtractive gain. cand (a) and the code Cα corresponding to the candidate cand Multiple pairs (A pairs, a=1, ..., A) with (a) are stored in advance. The left channel subtractive gain estimation unit 120 performs the following steps S120-11 to S120-14 as shown in Figure 5.
[0045] The left channel subtractive gain estimation unit 120 first calculates the input audio signal x of the left channel that has been input to it. L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x M From (T), equation (1-4) gives the normalized dot product r of the downmix signal with respect to the input sound signal of the left channel. L (Step S120-11) obtains the left channel subtraction gain estimation unit 120 in the stereo encoding unit 170 the left channel difference signal y L (1), y L (2), ..., y L Number of bits b used for encoding (T) L Then, in the mono encoding unit 160, the downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M The left channel correction coefficient c is calculated using the following equation (1-7) with T being the number of samples per frame. L Obtain (Step S120-12).
number
[0046] Furthermore, in the stereo encoding unit 170, the left channel difference signal y L (1), y L (2), ..., y L Number of bits b used for encoding (T) L If the number of bits b of the stereo code CS output by the stereo coding unit 170 is not explicitly determined. s Half of (i.e., b s / 2) to the number of bits b L It can be used as follows. Also, the left channel correction coefficient c L This is not a value obtained from equation (1-7) itself, but a value greater than 0 and less than 1, and is the left channel difference signal y L (1), y L (2), ..., y L Number of bits b used for encoding (T) L and downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M When they are the same, it is 0.5, and the number of bits b L The number of bits is b M The more than 0.5, the closer it is to 0, and the number of bits b L The number of bits is b M The smaller the value, the closer it is to 1 than 0.5. These same principles apply to the examples described later.
[0047] [Right channel subtraction gain estimation unit 140] The right channel subtractive gain estimation unit 140 generates a candidate β for the right channel subtractive gain. cand (b) The code Cβ corresponding to the candidate cand Multiple pairs (B pairs, b=1, ..., B) of (b) are stored in advance. The right channel subtractive gain estimation unit 140 performs the following steps S140-11 to S140-14 as shown in Figure 5.
[0048] The right channel subtractive gain estimation unit 140 first calculates the input sound signal x of the right channel that has been input to it. R (1), x R (2), ..., x R (T) and downmix signal x M (1), x M (2), ..., x M From (T), the normalized dot product value r of the downmix signal with respect to the input sound signal of the right channel is obtained by equation (1-4-2). R (Step S140-11) obtains the right channel subtractive gain estimation unit 140 in the stereo encoding unit 170, the right channel difference signal y R (1), y R (2), ..., y R Number of bits b used for encoding (T) R Then, in the mono encoding unit 160, the downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M The right channel correction coefficient c is calculated using the following equation (1-7-2) with the number of samples per frame T and the number of samples per frame T. R Obtain (Step S140-12).
number
[0049] Furthermore, in the stereo encoding unit 170, the right channel difference signal y R (1), y R (2), ..., y R Number of bits b used for encoding (T) R If the number of bits b of the stereo code CS output by the stereo coding unit 170 is not explicitly determined. s Half of (i.e., b s / 2) to the number of bits b R It can be used as follows. Also, the right channel correction coefficient c R This is not a value obtained directly from equation (1-7-2), but rather a value greater than 0 and less than 1, and is the right channel difference signal y R (1), y R (2), ..., y R Number of bits b used for encoding (T) R and downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M When they are the same, it is 0.5, and the number of bits b R The number of bits is b M The more than 0.5, the closer it is to 0, and the number of bits b R The number of bits is b M The smaller the value, the closer it is to 1 than 0.5. These same principles apply to the examples described later.
[0050] [Left channel subtractive gain decoding unit 230] The left channel subtractive gain decoding unit 230 stores the same candidate left channel subtractive gain α as stored in the left channel subtractive gain estimation unit 120 of the corresponding encoding device 100. cand (a) and the code Cα corresponding to the candidate cand Multiple pairs (A pairs, a=1, ..., A) of (a) are pre-stored. The left channel subtractive gain decoder 230 receives the stored code Cα cand (1), ..., Cα cand (A) The candidate left channel subtractive gain corresponding to the input left channel subtractive gain code Cα is obtained as the left channel subtractive gain α (step S230-11).
[0051] [Right channel subtractive gain decoding unit 250] The right channel subtractive gain decoding unit 250 stores a candidate right channel subtractive gain β, which is the same as the one stored in the right channel subtractive gain estimation unit 140 of the corresponding encoding device 100. cand (b) The code Cβ corresponding to the candidate cand Multiple pairs (B pairs, b=1, ..., B) of (b) are pre-stored. The right channel subtractive gain decoder 250 receives the stored code Cβ cand (1), ..., Cβ cand (B) The candidate right channel subtractive gain corresponding to the input right channel subtractive gain code Cβ is obtained as the right channel subtractive gain β (step S250-11).
[0052] Note that the same subtractive gain candidate and sign can be used for both the left and right channels. Assuming A and B are the same value as described above, the left channel subtractive gain candidate α stored in the left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 is used. cand (a) and the code Cα corresponding to the candidate cand (a) and the candidate right channel subtractive gain β stored in the right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250. cand (b) The code Cβ corresponding to the candidatecand (b) and the other two may be the same.
[0053] [Example 1 variation] Number of bits b used for encoding the left channel difference signal in the encoding device 100 L b is the number of bits used by the decoding device 200 to decode the left channel difference signal, and b is the number of bits used by the encoding device 100 to encode the downmix signal. M The value of is the number of bits used by the decoding device 200 to decode the downmix signal, so the correction coefficient c L The same value can be calculated by both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r L The quantized value of the dot product of the normalized dot product is ^r, which is the target of encoding and decoding. L Correction coefficient c L The left channel subtractive gain α can be obtained by multiplying by . The same applies to the right channel. This form will be explained as a variation of Example 1.
[0054] [Left channel subtraction gain estimation unit 120] The left channel subtractive gain estimation unit 120 generates candidate values for the normalized inner product of the left channel r Lcand (a) and the code Cα corresponding to the candidate cand Multiple pairs (A pairs, a=1, ..., A) with (a) are stored in advance. The left channel subtractive gain estimation unit 120 performs steps S120-11 and S120-12, which were also explained in Example 1, and steps S120-15 and S120-16 below, as shown in Figure 6.
[0055] The left channel subtractive gain estimation unit 120 first processes the input left channel input sound signal x, similar to step S120-11 of the left channel subtractive gain estimation unit 120 in Example 1. L (1), x L (2), ..., x L (T) and downmix signal x M (1), x M (2), ..., x MFrom (T), equation (1-4) gives the normalized dot product r of the downmix signal with respect to the input sound signal of the left channel. L (Step S120-11) The left channel subtraction gain estimation unit 120 then obtains a candidate r of the normalized inner product value of the stored left channel. Lcand (1), ..., r Lcand The normalized dot product value r obtained in step S120-11 of (A) L The closest candidate (normalized inner product value r) L (Quantized value of)^r L Obtaining the stored code Cα cand (1), ..., Cα cand (A) The closest candidate ^r L The code corresponding to this is obtained as the left channel subtractive gain code Cα (step S120-15). Also, the left channel subtractive gain estimation unit 120, similar to step S120-12 of the left channel subtractive gain estimation unit 120 in Example 1, obtains the left channel difference signal y in the stereo coding unit 170. L (1), y L (2), ..., y L Number of bits b used for encoding (T) L Then, in the mono encoding unit 160, the downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M The left channel correction coefficient c is calculated using equation (1-7) with T being the number of samples per frame. L (Step S120-12) The left channel subtraction gain estimation unit 120 then calculates the quantized value ^r of the normalized inner product obtained in step S120-15. L and the left channel correction coefficient c obtained in step S120-12 L The value obtained by multiplying by is the left channel subtraction gain α (step S120-16).
[0056] [Right channel subtraction gain estimation unit 140] The right channel subtractive gain estimation unit 140 generates candidate values for the normalized inner product of the right channel r Rcand(b) The code Cβ corresponding to the candidate cand Multiple pairs (B pairs, b=1, ..., B) of (b) are stored in advance. The right channel subtractive gain estimation unit 140 performs steps S140-11 and S140-12, which were also explained in Example 1, and steps S140-15 and S140-16 below, as shown in Figure 6.
[0057] The right channel subtractive gain estimation unit 140 first processes the input right channel input sound signal x, similar to step S140-11 of the right channel subtractive gain estimation unit 140 in Example 1. R (1), x R (2), ..., x R (T) and downmix signal x M (1), x M (2), ..., x M From (T), the normalized dot product value r of the downmix signal with respect to the input sound signal of the right channel is obtained by equation (1-4-2). R (Step S140-11) The right channel subtraction gain estimation unit 140 then obtains candidate r of the normalized inner product value of the stored right channel. Rcand (1), ..., r Rcand (B) The normalized dot product value r obtained in step S140-11 R The closest candidate (normalized inner product value r) R (Quantized value of)^r R The stored code Cβ is obtained. cand (1), ..., Cβ cand (B) The closest candidate ^r R The code corresponding to is obtained as the right channel subtractive gain code Cβ (step S140-15). Also, the right channel subtractive gain estimation unit 140, similar to step S140-12 of the right channel subtractive gain estimation unit 140 in Example 1, obtains the right channel difference signal y in the stereo coding unit 170. R (1), y R (2), ..., y R Number of bits b used for encoding (T) R Then, in the mono encoding unit 160, the downmix signal x M (1), x M(2), ..., x M Number of bits b used for encoding (T) M The right channel correction coefficient c is calculated using equation (1-7-2) with T being the number of samples per frame. R (Step S140-12) obtains the quantized value ^r of the normalized inner product obtained in step S140-15. The right channel subtraction gain estimation unit 140 then obtains the quantized value ^r of the normalized inner product obtained in step S140-15. R and the right channel correction coefficient c obtained in step S140-12 R The value obtained by multiplying by is the right channel subtraction gain β (step S140-16).
[0058] [Left channel subtractive gain decoding unit 230] The left channel subtractive gain decoding unit 230 stores candidate r of the normalized dot product of the left channel, which is the same as that stored in the left channel subtractive gain estimation unit 120 of the corresponding encoding device 100. Lcand (a) and the code Cα corresponding to the candidate cand Multiple pairs (A pairs, a=1, ..., A) with (a) are stored in advance. The left channel subtractive gain decoder 230 performs the following steps S230-12 to S230-14 as shown in Figure 7.
[0059] The left channel subtractive gain decoding unit 230 uses the stored code Cα cand (1), ..., Cα cand (A) The candidate for the left channel normalized inner product value corresponding to the input left channel subtractive gain code Cα is the decoded value of the left channel normalized inner product ^r L This is obtained as (step S230-12). Also, the left channel subtractive gain decoding unit 230 uses the left channel decoded difference signal ^y in the stereo decoding unit 220. L (1), ^y L (2), ..., ^y L Number of bits b used to decode (T) L Then, in the monaural decoding unit 210, the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M Number of bits b used to decode (T) MThe left channel correction coefficient c is calculated using equation (1-7) with T being the number of samples per frame. L (Step S230-13) The left channel subtractive gain decoder 230 then decodes the normalized inner product value obtained in step S230-12, ^r L And the left channel correction coefficient c obtained in step S230-13 L The value obtained by multiplying by is the left channel subtraction gain α (step S230-14).
[0060] Furthermore, if the stereo code CS is a combination of the left channel difference code CL and the right channel difference code CR, the stereo decoding unit 220 will process the left channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L Number of bits b used to decode (T) L This is the number of bits in the left channel difference code CL. In the stereo decoding unit 220, the left channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L Number of bits b used to decode (T) L If this is not explicitly determined, the number of bits b of the stereo code CS input to the stereo decoding unit 220 is determined. s Half of (i.e., b s / 2) to the number of bits b L It can be used as follows: In the monaural decoding unit 210, the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M Number of bits b used to decode (T) M This is the number of bits in the monaural code CM. Left channel correction coefficient c L This is not a value obtained directly from equation (1-7), but rather a value greater than 0 and less than 1, and is the left channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L Number of bits b used to decode (T) L and mono decoded sound signal^x M (1), ^x M(2), ..., ^x M Number of bits b used to decode (T) M When they are the same, it is 0.5, and the number of bits b L The number of bits is b M The more than 0.5, the closer it is to 0, and the number of bits b L The number of bits is b M The less than 0.5, the closer the value is to 1.
[0061] [Right channel subtractive gain decoding unit 250] The right channel subtractive gain decoding unit 250 stores candidate r of the normalized inner product of the right channel, which is the same as that stored in the right channel subtractive gain estimation unit 140 of the corresponding encoding device 100. Rcand (b) The code Cβ corresponding to the candidate cand Multiple pairs (B pairs, b=1, ..., B) of (b) are stored in advance. The right channel subtractive gain decoder 250 performs the following steps S250-12 to S250-14 as shown in Figure 7.
[0062] The right channel subtractive gain decoder 250 receives the stored code Cβ cand (1), ..., Cβ cand (B) The candidate for the normalized dot product of the right channel corresponding to the input right channel subtractive gain code Cβ is the decoded value of the normalized dot product of the right channel ^r R This is obtained as (step S250-12). Also, the right channel subtractive gain decoding unit 250 processes the right channel decoded difference signal ^y in the stereo decoding unit 220. R (1), ^y R (2), ..., ^y R Number of bits b used to decode (T) R Then, in the monaural decoding unit 210, the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M Number of bits b used to decode (T) M The right channel correction coefficient c is calculated using equation (1-7-2) with T being the number of samples per frame. R(Step S250-13) The right channel subtractive gain decoder 250 then decodes the normalized inner product value obtained in step S250-12 ^r R And the right channel correction coefficient c obtained in step S250-13 R The value obtained by multiplying by is the right channel subtraction gain β (step S250-14).
[0063] Furthermore, if the stereo code CS is a combination of the left channel difference code CL and the right channel difference code CR, the stereo decoding unit 220 will process the right channel decoded difference signal ^y R (1), ^y R (2), ..., ^y R Number of bits b used to decode (T) R This is the number of bits in the right channel difference code CR. In the stereo decoding unit 220, the right channel decoded difference signal ^y R (1), ^y R (2), ..., ^y R Number of bits b used to decode (T) R If this is not explicitly determined, the number of bits b of the stereo code CS input to the stereo decoding unit 220 is determined. s Half of (i.e., b s / 2) to the number of bits b R It can be used as follows: In the monaural decoding unit 210, the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M Number of bits b used to decode (T) M This is the number of bits in the monaural code CM. Right channel correction coefficient c R This is not a value obtained directly from equation (1-7-2), but rather a value greater than 0 and less than 1, and is the right channel decoded difference signal ^y R (1), ^y R (2), ..., ^y R Number of bits b used to decode (T) R and mono decoded sound signal^x M (1), ^x M (2), ..., ^x M Number of bits b used to decode (T) MWhen they are the same, it is 0.5, and the number of bits b R The number of bits is b M The more than 0.5, the closer it is to 0, and the number of bits b R The number of bits is b M The less than 0.5, the closer the value is to 1.
[0064] Furthermore, the same normalized inner product candidate values and signs can be used for the left and right channels. Assuming A and B are the same value as described above, the left channel normalized inner product candidate values r stored in the left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 are used. Lcand (a) and the code Cα corresponding to the candidate cand (a) and the candidate r of the normalized inner product value of the right channel stored in the right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250. Rcand (b) The code Cβ corresponding to the candidate cand (b) and the other two may be the same.
[0065] Note that the code Cα is essentially the code corresponding to the left channel subtractive gain α, and is referred to as the left channel subtractive gain code for the purpose of consistency in the descriptions of the encoding device 100 and the decoding device 200. However, since it represents a normalized inner product value, it could also be called the left channel inner product code. The same applies to the code Cβ, which could also be called the right channel inner product code.
[0066] [Example 2] Example 2 illustrates an example where the normalized dot product value takes into account the input values of past frames. In Example 2, while the optimality within a frame, i.e., minimizing the energy of the quantization error in the decoded audio signal of the left channel and the energy of the quantization error in the decoded audio signal of the right channel, is not strictly guaranteed, it reduces the rapid fluctuations between frames in the left channel subtractive gain α and the right channel subtractive gain β, thereby reducing the noise generated in the decoded audio signal due to these fluctuations. In other words, Example 2 takes into account not only the energy of the quantization error in the decoded audio signal but also the auditory quality of the decoded audio signal.
[0067] In Example 2, the encoding side, i.e., the left channel subtractive gain estimation unit 120 and the right channel subtractive gain estimation unit 140, differs from that of Example 1, but the decoding side, i.e., the left channel subtractive gain decoding unit 230 and the right channel subtractive gain decoding unit 250, is the same as in Example 1. The following explanation will focus on the differences between Example 2 and Example 1.
[0068] [Left channel subtraction gain estimation unit 120] The left channel subtractive gain estimation unit 120 performs the following steps S120-111 to S120-113 and steps S120-12 to S120-14 as described in Example 1, as shown in Figure 8.
[0069] The left channel subtractive gain estimation unit 120 first calculates the input audio signal x of the left channel that has been input to it. L (1), x L (2), ..., x L (T) and the input downmix signal x M (1), x M (2), ..., x M (T) and the dot product E used in the previous frame L Using (-1) and , the dot product value E used in the current frame is obtained by the following equation (1-8). L (0) is obtained (step S120-111).
number
[0070] The left channel subtractive gain estimation unit 120 also takes the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M Using (-1) and , the energy E of the downmix signal used in the current frame is calculated by the following equation (1-9). M (0) is obtained (step S120-112).
number
[0071] The left channel subtractive gain estimation unit 120 then calculates the dot product value E used in the current frame obtained in steps S120-111. L (0) and the energy E of the downmix signal used in the current frame obtained in steps S120-112 M Using (0), the normalized inner product value r L This is obtained using the following formula (1-10) (step S120-113).
number
[0072] The left channel subtractive gain estimation unit 120 also performs step S120-12, and then calculates the normalized inner product value r obtained in step S120-11. L Instead, the normalized inner product value r obtained in step S120-113 described above. L Step S120-13 is performed using [the specified method], and then step S120-14 is performed.
[0073] Note that the above ε L and ε M The closer to 1 the normalized inner product value r is, the better. L The normalized dot product r is more likely to include the influence of the input audio signal and downmix signal from the left channel of past frames. L or the normalized inner product value r L This reduces the frame-to-frame variation in the left channel subtractive gain α obtained by this method.
[0074] [Right channel subtraction gain estimation unit 140] The right channel subtractive gain estimation unit 140 performs the following steps S140-111 to S140-113 and steps S140-12 to S140-14 as described in Example 1, as shown in Figure 8.
[0075] The right channel subtractive gain estimation unit 140 first calculates the input sound signal x of the right channel that has been input to it. R (1), x R (2), ..., x R (T) and the input downmix signal x M (1), x M (2), ..., x M (T) and the dot product E used in the previous frame R Using (-1) and , the dot product value E used in the current frame is obtained by the following equation (1-8-2): R (0) is obtained (step S140-111).
number
[0076] The right channel subtractive gain estimation unit 140 also takes the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M Using (-1) and , according to equation (1-9), the energy E of the downmix signal used in the current frame is M (0) is obtained (step S140-112). The right channel subtractive gain estimation unit 140 calculates the energy E of the obtained downmix signal. M (0) The energy E of the downmix signal used in the previous frame M The energy E of the downmix signal used in the current frame is stored in the right channel subtractive gain estimation unit 140 as (-1) for use in the next frame. The left channel subtractive gain estimation unit 120 also uses equation (1-9) to estimate the energy of the downmix signal used in the current frame. M Since (0) is obtained, either step S120-112 performed by the left channel subtractive gain estimation unit 120 or step S140-112 performed by the right channel subtractive gain estimation unit 140 may be performed by only one of them.
[0077] The right channel subtractive gain estimation unit 140 then calculates the dot product value E obtained in step S140-111 to be used in the current frame. R (0) and the energy E of the downmix signal used in the current frame obtained in steps S140-112. M Using (0), the normalized inner product value r R This is obtained using the following formula (1-10-2) (step S140-113).
number
[0078] The right channel subtractive gain estimation unit 140 also performs step S140-12, and then calculates the normalized inner product value r obtained in step S140-11. R Instead, the normalized inner product value r obtained in step S140-113 described above. R Step S140-13 is performed using [the specified method], and then step S140-14 is performed.
[0079] Note that the above ε R and ε M The closer to 1 the normalized inner product value r is, the better. R The normalized inner product r is more likely to include the influence of the input audio signal and downmix signal of the right channel of past frames. R or the normalized inner product value r R This reduces the frame-to-frame variation in the right channel subtractive gain β obtained.
[0080] [Example 2 modified version] The same modifications can be made to Example 2 as to the modified version of Example 1. This form will be described as the modified version of Example 2. In the modified version of Example 2, the encoding side, i.e., the left channel subtractive gain estimation unit 120 and the right channel subtractive gain estimation unit 140, differs from the modified version of Example 1, but the decoding side, i.e., the left channel subtractive gain decoding unit 230 and the right channel subtractive gain decoding unit 250, is the same as the modified version of Example 1. The differences between the modified version of Example 2 and the modified version of Example 1 are the same as in Example 2, so in the following, the modified version of Example 2 will be explained with appropriate reference to the modified version of Example 1 and Example 2.
[0081] [Left channel subtraction gain estimation unit 120] The left channel subtractive gain estimation unit 120, similar to the left channel subtractive gain estimation unit 120 in the modified example of Example 1, generates candidate r of the normalized inner product of the left channel. Lcand (a) and the code Cα corresponding to the candidate candMultiple pairs (A pairs, a=1, ..., A) with (a) are stored in advance. The left channel subtractive gain estimation unit 120 performs the same steps S120-111 to S120-113 as in Example 2, and the same steps S120-12, S120-15, and S120-16 as in the modified example of Example 1, as shown in Figure 9. Specifically, it is as follows.
[0082] The left channel subtractive gain estimation unit 120 first calculates the input audio signal x of the left channel that has been input to it. L (1), x L (2), ..., x L (T) and the input downmix signal x M (1), x M (2), ..., x M (T) and the dot product E used in the previous frame L Using (-1) and , the dot product value E used in the current frame is obtained by equation (1-8). L (0) is obtained (steps S120-111). The left channel subtractive gain estimation unit 120 also takes the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M Using (-1) and , according to equation (1-9), the energy E of the downmix signal used in the current frame is M (0) is obtained (steps S120-112). The left channel subtractive gain estimation unit 120 then determines the dot product value E to be used in the current frame obtained in step S120-111. L (0) and the energy E of the downmix signal used in the current frame obtained in steps S120-112 M Using (0), the normalized inner product value r is obtained by equation (1-10). L The left channel subtractive gain estimation unit 120 then obtains a candidate r of the normalized inner product value of the stored left channel. Lcand (1), ..., r Lcand The normalized dot product value r obtained in step S120-113 of (A) LThe closest candidate (normalized inner product value r) L (Quantized value of)^r L Obtaining the stored code Cα cand (1), ..., Cα cand (A) The closest candidate ^r L The code corresponding to this is obtained as the left channel subtractive gain code Cα (step S120-15). Also, the left channel subtractive gain estimation unit 120 obtains the left channel difference signal y in the stereo coding unit 170. L (1), y L (2), ..., y L Number of bits b used for encoding (T) L Then, in the mono encoding unit 160, the downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M Using the number of samples per frame T, the left channel correction coefficient c is calculated by equation (1-7). L (Step S120-12) The left channel subtraction gain estimation unit 120 then calculates the quantized value ^r of the normalized inner product obtained in step S120-15. L and the left channel correction coefficient c obtained in step S120-12 L The value obtained by multiplying by is the left channel subtraction gain α (step S120-16).
[0083] [Right channel subtraction gain estimation unit 140] The right channel subtractive gain estimation unit 140, similar to the right channel subtractive gain estimation unit 140 in the modified example of Example 1, generates candidate r of the normalized inner product of the right channel. Rcand (b) The code Cβ corresponding to the candidate cand Multiple pairs (B pairs, b=1, ..., B) with (b) are stored in advance. The right channel subtraction gain estimation unit 140 performs the same steps S140-111 to S140-113 as in Example 2, and the same steps S140-12, S140-15, and S140-16 as in the modified example of Example 1, as shown in Figure 9. Specifically, it is as follows.
[0084] The right channel subtractive gain estimation unit 140 first calculates the input sound signal x of the right channel that has been input to it. R (1), x R (2), ..., x R (T) and the input downmix signal x M (1), x M (2), ..., x M (T) and the dot product E used in the previous frame R Using (-1) and , the dot product value E used in the current frame is obtained by equation (1-8-2). R (0) is obtained (steps S140-111). The right channel subtractive gain estimation unit 140 also takes the input downmix signal x M (1), x M (2), ..., x M (T) and the energy E of the downmix signal used in the previous frame M Using (-1) and , according to equation (1-9), the energy E of the downmix signal used in the current frame is M (0) is obtained (steps S140-112). The right channel subtractive gain estimation unit 140 then determines the dot product value E to be used in the current frame obtained in step S140-111. R (0) and the energy E of the downmix signal used in the current frame obtained in steps S140-112. M Using (0), the normalized inner product value r is obtained by equation (1-10-2). R (Steps S140-113) The right channel subtractive gain estimation unit 140 then obtains candidate r of the normalized inner product value of the stored right channel. Rcand (1), ..., r Rcand (B) The normalized dot product value r obtained in step S140-113 R The closest candidate (normalized inner product value r) R (Quantized value of)^r R The stored code Cβ is obtained. cand (1), ..., Cβ cand (B) The closest candidate ^r RThe code corresponding to this is obtained as the right channel subtractive gain code Cβ (step S140-15). Also, the right channel subtractive gain estimation unit 140 obtains the right channel difference signal y in the stereo coding unit 170. R (1), y R (2), ..., y R Number of bits b used for encoding (T) R Then, in the mono encoding unit 160, the downmix signal x M (1), x M (2), ..., x M Number of bits b used for encoding (T) M Using the number of samples per frame T, the right channel correction coefficient c is calculated using equation (1-7-2). R (Step S140-12) obtains the quantized value ^r of the normalized inner product obtained in step S140-15. The right channel subtraction gain estimation unit 140 then obtains the quantized value ^r of the normalized inner product obtained in step S140-15. R and the right channel correction coefficient c obtained in step S140-12 R The value obtained by multiplying by is the right channel subtraction gain β (step S140-16).
[0085] [Example 3] For example, if the sounds such as speech or music contained in the left channel input audio signal are different from the sounds such as speech or music contained in the right channel input audio signal, the downmix signal may contain components from both the left channel input audio signal and the right channel input audio signal. Therefore, the larger the value used for the left channel subtraction gain α, the more it sounds as if the left channel decoded audio signal contains sounds originating from the right channel input audio signal that should not be audible. Similarly, the larger the value used for the right channel subtraction gain β, the more it sounds as if the right channel decoded audio signal contains sounds originating from the left channel input audio signal that should not be audible. Therefore, although minimizing the energy of the quantization error in the decoded audio signal is not strictly guaranteed, the left channel subtraction gain α and the right channel subtraction gain β may be set to values smaller than those obtained by Example 1, taking auditory quality into consideration. Similarly, the left channel subtraction gain α and the right channel subtraction gain β may be set to values smaller than those obtained by Example 2.
[0086] Specifically, for the left channel, in Examples 1 and 2, the normalized inner product value r L and left channel correction coefficient c L The multiplicative value c L ×r L In Example 3, the quantized value of r was used as the left channel subtraction gain α, but in Example 3, the normalized inner product value r L and left channel correction coefficient c L λ is a predetermined value greater than 0 and less than 1. L The multiplicative value λ L ×c L ×r L Let the quantized value be the left channel subtraction gain α. Therefore, as in Example 1 and Example 2, the multiplication value c L ×r L The left channel subtractive gain code Cα is used as the target for encoding in the left channel subtractive gain estimation unit 120 and decoding in the left channel subtractive gain decoding unit 230, and the multiplied value c is the result of the left channel subtractive gain coding. L ×r L The left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 multiply the quantized value c L ×r L Quantized value and λ L Alternatively, the left channel subtractive gain α may be obtained by multiplying by the normalized inner product value r. L and left channel correction coefficient c L and a predetermined value λ L The multiplicative value λ L ×c L ×r L The left channel subtractive gain code Cα is used as the target for encoding in the left channel subtractive gain estimation unit 120 and decoding in the left channel subtractive gain decoding unit 230, and the multiplier value λ L ×c L ×r L It may also be used to represent the quantized value.
[0087] Similarly, for the right channel, in Examples 1 and 2, the normalized inner product value r R and the right channel correction coefficient c R The multiplicative value c R ×r RIn Example 3, the quantized value of r was used as the right channel subtraction gain β, but in Example 3, the normalized inner product value r R and the right channel correction coefficient c R λ is a predetermined value greater than 0 and less than 1. R The multiplicative value λ R ×c R ×r R Let the quantized value be the right channel subtraction gain β. Therefore, as in Example 1 and Example 2, the multiplication value c R ×r R The right channel subtractive gain code Cβ is used as the target for encoding in the right channel subtractive gain estimation unit 140 and decoding in the right channel subtractive gain decoding unit 250, and the multiplied value c is the right channel subtractive gain code Cβ. R ×r R The right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250 multiply the quantized value c R ×r R Quantized value and λ R Alternatively, the right channel subtractive gain β may be obtained by multiplying by . Or, the normalized inner product value r R and left channel correction coefficient c R and a predetermined value λ R The multiplicative value λ R ×c R ×r R The right channel subtractive gain code Cβ is used for encoding in the right channel subtractive gain estimation unit 140 and decoding in the right channel subtractive gain decoding unit 250, and the multiplier value λ is used for the right channel subtractive gain code Cβ. R ×c R ×r R It may also be used to represent the quantized value of λ. R is λ L It's best to set it to the same value as [this value].
[0088] [Example 3 modified version] As mentioned above, the correction coefficient c L The same value can be calculated by both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r can be calculated in the same way as in the modified examples of Example 1 and Example 2. L The left channel subtractive gain code Cα is the target of encoding in the left channel subtractive gain estimation unit 120 and decoding in the left channel subtractive gain decoding unit 230, and the normalized inner product value r LThe left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 calculate the normalized inner product value r such that it represents the quantized value of r. L Quantized value and left channel correction coefficient c L λ is a predetermined value greater than 0 and less than 1. L Alternatively, the left channel subtractive gain α may be obtained by multiplying by the normalized inner product value r. L λ is a predetermined value greater than 0 and less than 1. L The multiplicative value λ L ×r L The left channel subtractive gain code Cα is used as the target for encoding in the left channel subtractive gain estimation unit 120 and decoding in the left channel subtractive gain decoding unit 230, and the multiplier value λ L ×r L The left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 multiply the quantized value λ L ×r L Quantized value and left channel correction coefficient c L Alternatively, you could multiply by this to obtain the left channel subtractive gain α.
[0089] The same applies to the right channel, and the correction coefficient c R The same value can be calculated by both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r can be calculated in the same way as in the modified examples of Example 1 and Example 2. R The dot product value r, which is the normalized value of the right channel subtractive gain code Cβ, is the target of encoding in the right channel subtractive gain estimation unit 140 and decoding in the right channel subtractive gain decoding unit 250. R The right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250 calculate the normalized inner product value r such that it represents the quantized value of r. R Quantized value and right channel correction coefficient c R λ is a predetermined value greater than 0 and less than 1. R Alternatively, the right channel subtraction gain β may be obtained by multiplying by the normalized inner product value r. R λ is a predetermined value greater than 0 and less than 1. R The multiplicative value λ R ×r RThe right channel subtractive gain code Cβ is used for encoding in the right channel subtractive gain estimation unit 140 and decoding in the right channel subtractive gain decoding unit 250, and the multiplier value λ is used for the right channel subtractive gain code Cβ. R ×r R The right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250 multiply the quantized value λ R ×r R Quantized value and right channel correction coefficient c R Alternatively, the right channel subtraction gain β may be obtained by multiplying by .
[0090] [Example 4] The auditory quality issue described at the beginning of Example 3 occurs when the correlation between the left channel input sound signal and the right channel input sound signal is small, and this issue does not occur as much when the correlation between the left channel input sound signal and the right channel input sound signal is large. Therefore, in Example 4, instead of the predetermined value in Example 3, the left-right correlation coefficient γ, which is the correlation coefficient between the left channel input sound signal and the right channel input sound signal, is used. As the correlation between the left channel input sound signal and the right channel input sound signal is large, priority is given to reducing the energy of the quantization error in the decoded sound signal, and as the correlation between the left channel input sound signal and the right channel input sound signal is small, priority is given to suppressing the deterioration of auditory quality.
[0091] In Example 4, the encoding side differs from that of Examples 1 and 2, but the decoding side, namely the left channel subtractive gain decoding unit 230 and the right channel subtractive gain decoding unit 250, is the same as in Examples 1 and 2. The differences between Example 4 and Examples 1 and 2 will be explained below.
[0092] [Left-right relationship information estimation unit 180] The encoding device 100 in Example 4 also includes a left-right relationship information estimation unit 180, as shown by the dashed line in Figure 1. The left-right relationship information estimation unit 180 receives the input sound signal of the left channel and the input sound signal of the right channel input to the encoding device 100. The left-right relationship information estimation unit 180 obtains a left-right correlation coefficient γ from the input sound signals of the left channel and the right channel and outputs it (step S180).
[0093] The left-right correlation coefficient γ is the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and the sample sequence x of the input audio signal of the left channel L (1), x L (2), ..., x L (T) and the sample sequence x of the input audio signal from the right channel R (1), x R (2), ..., x R The correlation coefficient γ of (T) may be 0, or the correlation coefficient that takes time difference into account, for example, the correlation coefficient γ between the sample sequence of the input sound signal of the left channel and the sample sequence of the input sound signal of the right channel, which is shifted by τ samples later than the said sample sequence. τ That's fine.
[0094] This τ is information corresponding to the difference (so-called arrival time difference) between the arrival time from the sound source primarily emitting sound in a given space to the left channel microphone and the arrival time from that sound source to the right channel microphone, assuming that the sound signal obtained by AD conversion of the sound picked up by the left channel microphone placed in a given space is the input sound signal for the left channel, and the sound signal obtained by AD conversion of the sound picked up by the right channel microphone placed in the same space is the input sound signal for the right channel. Hereafter, this will be called the left-right time difference. The left-right time difference τ can be determined by any well-known method, such as the method described in the left-right relationship information estimation unit 181 of the second reference embodiment. That is, the correlation coefficient γ mentioned above. τ This information corresponds to the correlation coefficient between the sound signal captured by the microphone for the left channel from the sound source and the sound signal captured by the microphone for the right channel from the same sound source.
[0095] [Left channel subtraction gain estimation unit 120] The left channel subtractive gain estimation unit 120, instead of step S120-13, uses the normalized inner product value r obtained in step S120-11 or step S120-113. L And the left channel correction coefficient c obtained in step S120-12. LThen, the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 by the value obtained in step S120-13") is obtained. The left channel subtractive gain estimation unit 120 then, instead of step S120-14, obtains the stored left channel subtractive gain candidate α cand (1), ..., α cand The multiplication value γ×c obtained in step S120-13" of (A) L ×r L The closest candidate (multiplication value γ × c) L ×r L The quantized value of is obtained as the left channel subtraction gain α, and the stored code Cα cand (1), ..., Cα cand The code corresponding to the left channel subtractive gain α in (A) is obtained as the left channel subtractive gain code Cα (step S120-14).
[0096] [Right channel subtraction gain estimation unit 140] The right channel subtractive gain estimation unit 140, instead of step S140-13, uses the normalized inner product value r obtained in step S140-11 or step S140-113. R And the right channel correction coefficient c obtained in step S140-12. R Then, the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 by the value obtained in step S140-13") is obtained. The right channel subtractive gain estimation unit 140 then, instead of step S140-14, stores a candidate β of the right channel subtractive gain. cand (1), ..., β cand The multiplication value γ×c obtained in step S140-13” of (B) R ×r R The closest candidate (multiplication value γ × c) R ×r R The quantized value of is obtained as the right channel subtraction gain β, and the stored code Cβ cand (1), ..., Cβ cand The code corresponding to the right channel subtractive gain β from (B) is obtained as the right channel subtractive gain code Cβ (step S140-14).
[0097] [Example 4 modified version] As mentioned above, the correction coefficient c L The same value can be calculated by both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r L The product of the left-right correlation coefficient γ and γ × r L The left channel subtractive gain code Cα is used as the target for encoding in the left channel subtractive gain estimation unit 120 and decoding in the left channel subtractive gain decoding unit 230, and the left channel subtractive gain code Cα is multiplied by the value γ × r L The left channel subtractive gain estimation unit 120 and the left channel subtractive gain decoding unit 230 multiply the quantized value γ × r L Quantized value and left channel correction coefficient c L Alternatively, you could multiply by this to obtain the left channel subtractive gain α.
[0098] The same applies to the right channel, and the correction coefficient c R The same value can be calculated by both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r R The product of the left-right correlation coefficient γ and γ × r R The right channel subtractive gain code Cβ is used for encoding in the right channel subtractive gain estimation unit 140 and decoding in the right channel subtractive gain decoding unit 250, and the right channel subtractive gain code Cβ is multiplied by the value γ × r R The right channel subtractive gain estimation unit 140 and the right channel subtractive gain decoding unit 250 multiply the quantized value γ × r R Quantized value and right channel correction coefficient c R Alternatively, the right channel subtraction gain β may be obtained by multiplying by .
[0099] <Second reference form> The second reference form of encoding and decoding devices will be described.
[0100] ≪Encoding device 101≫ The second reference embodiment encoding device 101, as shown in Figure 10, includes a downmix unit 110, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, a left / right relationship information estimation unit 181, and a time shift unit 191. The second reference embodiment encoding device 101 differs from the first reference embodiment encoding device 100 in that it includes a left / right relationship information estimation unit 181 and a time shift unit 191, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 use the signal output by the time shift unit 191 instead of the signal output by the downmix unit 110, and in addition to the above-mentioned codes, it also outputs a left / right time difference code Cτ, which will be described later. The other configurations and operations of the encoding device 101 of the second reference embodiment are the same as those of the encoding device 100 of the first reference embodiment. The encoding device 101 of the second reference embodiment performs the processing from steps S110 to S191, as illustrated in Figure 11, for each frame. The differences between the encoding device 101 of the second reference embodiment and the encoding device 100 of the first reference embodiment will be explained below.
[0101] [Left-Right Relationship Information Estimation Unit 181] The left-right relationship information estimation unit 181 receives the input sound signal of the left channel and the input sound signal of the right channel, both input to the encoding device 101. The left-right relationship information estimation unit 181 obtains and outputs a left-right time difference τ and a left-right time difference code Cτ, which is a code representing the left-right time difference τ, from the input sound signal of the left channel and the input sound signal of the right channel (step S181).
[0102] The left-right time difference τ is information that corresponds to the difference (so-called arrival time difference) between the arrival time from the sound source primarily emitting sound in a given space to the left channel microphone and the arrival time from that sound source to the right channel microphone, assuming that the sound signal obtained by AD conversion of the sound picked up by the left channel microphone placed in a given space is the input sound signal for the left channel, and the sound signal obtained by AD conversion of the sound picked up by the right channel microphone placed in the same space is the input sound signal for the right channel. In addition to the arrival time difference, the left-right time difference τ is also intended to include information about which microphone the sound arrives at first, so the left-right time difference τ can take both positive and negative values relative to either input sound signal. In other words, the left-right time difference τ is information that indicates how much earlier the same sound signal is contained in the left channel input sound signal and the right channel input sound signal. In the following, if the same audio signal is present in the left channel's input audio signal before the right channel's input audio signal, we will say that the left channel is leading. Conversely, if the same audio signal is present in the right channel's input audio signal before the left channel's input audio signal, we will say that the right channel is leading.
[0103] The left-right time difference τ may be determined by any known method. For example, the left-right relationship information estimation unit 181 determines a predetermined τ max from τ min up to (for example, τ max is a positive number, τ min The number of candidate samples τ (where τ is a negative number) cand Regarding this, the sample sequence of the input audio signal for the left channel and the number of candidate samples τ cand γ is a value (hereinafter referred to as the correlation value) that represents the magnitude of the correlation between the sample sequence of the right channel input sound signal located at a position shifted by a certain number of minutes from the sample sequence in question and the sample sequence of the right channel input sound signal. cand Calculate the correlation value γ cand The number of candidate samples τ that maximizes this candis obtained as the left-right time difference τ. That is, in this example, when the left channel is ahead, the left-right time difference τ is a positive value, and when the right channel is ahead, the left-right time difference τ is a negative value. The absolute value of the left-right time difference τ represents the value (the number of samples ahead) of how much the ahead channel is ahead of the other channel. For example, when calculating the correlation value γ cand using only the samples within a frame, when τ cand is a positive value, the partial sample sequence x R (1 + τ cand ), x R (2 + τ cand ),..., x R (T) of the input sound signal of the right channel and the partial sample sequence x cand of the input sound signal of the left channel, which is shifted forward by the candidate sample number τ L (1), x L (2),..., x L (T - τ cand ) from the partial sample sequence, are used to calculate the absolute value of the correlation coefficient as the correlation value γ cand . When τ cand is a negative value, the partial sample sequence x L (1 - τ cand ), x L (2 - τ cand ),..., x L (T) of the input sound signal of the left channel and the partial sample sequence x cand of the input sound signal of the right channel, which is shifted forward by the candidate sample number -τ R (1), x R (2),..., x R (T + τ cand ) from the partial sample sequence, are used to calculate the absolute value of the correlation coefficient as the correlation value γ cand . Of course, one or more samples of the input sound signal in the past consecutive to the sample sequence of the input sound signal in the current frame may also be used to calculate the correlation value γ cand . In this case, the sample sequences of the input sound signals in the past frames may be stored in a storage unit (not shown) in the left-right relationship information estimation unit 181 for a predetermined number of frames in advance.
[0104] For example, instead of the absolute value of the correlation coefficient, the correlation value γ can be calculated using the phase information of the signal, as shown below. cand The left-right relationship information estimation unit 181 first calculates the input sound signal x of the left channel. L (1), x L (2), ..., x L (T) and the input audio signal x of the right channel R (1), x R (2), ..., x R By performing a Fourier transform on each of (T) as shown in equations (3-1) and (3-2) below, we obtain the frequency spectrum X at each frequency k from 0 to T-1. L (k) and X R (k) is obtained.
number
number
number
number
number
number
[0105] Furthermore, the left-right relationship information estimation unit 181 may encode the left-right time difference τ using a predetermined encoding scheme to obtain a left-right time difference code Cτ, which is a code that can uniquely identify the left-right time difference τ. As the predetermined encoding scheme, a well-known encoding scheme such as scalar quantization may be used. The predetermined number of each candidate sample is τ max from τ min It may be any integer value up to τ max from τ min It may include fractional or decimal values between the given number and τ. max from τ min It is not necessary to include any integer value between the given value and τ. max =-τ min It may be so, or it may not be so. Also, when dealing with special input audio signals where one channel always precedes the other, τ max τ min We can also take τ as a positive number. max τ min You can also treat it as a negative number.
[0106] Furthermore, when the encoding device 101 performs subtractive gain estimation based on the principle of minimizing the quantization error of Example 4 or a modified example of Example 4 described in the first reference embodiment, the left-right relationship information estimation unit 181 further calculates the correlation value between the sample sequence of the input sound signal of the left channel and the sample sequence of the input sound signal of the right channel, which is located at a position shifted later than the said sample sequence by a left-right time difference τ, i.e., τ max from τ min Number of candidate samples up to τ cand The correlation value γ calculated for cand The maximum value among them is output as the left-right correlation coefficient γ (step S180).
[0107] [Time Shift Section 191] The time shift unit 191 receives the downmix signal x output by the downmix unit 110. M (1), x M (2), ..., x M(T) and the left - right time difference τ output by the left - right relationship information estimation unit 181 are input. When the left - right time difference τ is a positive value (that is, when the left - right time difference τ indicates that the left channel is ahead), the time shift unit 191 outputs the down - mix signals x M (1), x M (2),..., x M (T) as they are to the left - channel subtraction gain estimation unit 120 and the left - channel signal subtraction unit 130 (that is, it is determined to be used by the left - channel subtraction gain estimation unit 120 and the left - channel signal subtraction unit 130), and outputs the signal x M (1 - |τ|), x M (2 - |τ|),..., x M (T - |τ|) which is the delayed down - mix signal x M' (1), x M' (2),..., x M' (T) to the right - channel subtraction gain estimation unit 140 and the right - channel signal subtraction unit 150 (that is, it is determined to be used by the right - channel subtraction gain estimation unit 140 and the right - channel signal subtraction unit 150). When the left - right time difference τ is a negative value (that is, when the left - right time difference τ indicates that the right channel is ahead), the signal x M (1 - |τ|), x M (2 - |τ|),..., x M (T - |τ|) which is the delayed down - mix signal x M' (1), x M' (2),..., x M' (T) to the left - channel subtraction gain estimation unit 120 and the left - channel signal subtraction unit 130 (that is, it is determined to be used by the left - channel subtraction gain estimation unit 120 and the left - channel signal subtraction unit 130), and the down - mix signals x M (1), x M (2),..., x M(T) is output directly to the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150 (i.e., it is decided to use it in the right channel subtraction gain estimation unit 140 and the right channel signal subtraction unit 150), and if the left-right time difference τ is 0 (i.e., the left-right time difference τ indicates that neither channel is leading), the downmix signal x M (1), x M (2), ..., x M (T) is output as is to the left channel subtractive gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtractive gain estimation unit 140, and the right channel signal subtraction unit 150 (i.e., it is decided to use it in the left channel subtractive gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtractive gain estimation unit 140, and the right channel signal subtraction unit 150) (step S191). That is, for the channel with the shorter arrival time of the left and right channels, the input downmix signal is output as is to the subtractive gain estimation unit and the signal subtraction unit of that channel, and for the channel with the longer arrival time of the left and right channels, the input downmix signal is delayed by the absolute value |τ| of the left-right time difference τ and the signal is output to the subtractive gain estimation unit and the signal subtraction unit of that channel. Furthermore, since the time shift unit 191 uses the downmix signals of past frames to obtain a delayed downmix signal, a storage unit (not shown) within the time shift unit 191 stores the downmix signals input in past frames for a predetermined number of frames. Also, if the left channel subtractive gain estimation unit 120 and the right channel subtractive gain estimation unit 140 obtain the left channel subtractive gain α and right channel subtractive gain β by a well-known method as exemplified in Patent Document 1, rather than a method based on the principle of minimizing quantization error, the encoding device 101 is provided with means for obtaining a local decoded signal corresponding to the monaural code CM after the monaural encoding unit 160 or within the monaural encoding unit 160, and the time shift unit 191 receives the downmix signal x M (1), x M (2), ..., x M Instead of (T), use the quantized downmix signal ^x, which is the local decoded signal of the monaural encoding. M(1), ^x M (2), ..., ^x M The above processing may be performed using (T). In this case, the time shift unit 191 uses the downmix signal x M (1), x M (2), ..., x M Replace (T) with the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M Output (T) and delay downmix signal x M' (1), x M' (2), ..., x M' (T) is replaced with the delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' Output (T).
[0108] [Left channel subtraction gain estimation unit 120, left channel signal subtraction unit 130, right channel subtraction gain estimation unit 140, right channel signal subtraction unit 150] The left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 perform the same operations as described in the first reference embodiment, using the downmix signal x output by the downmix unit 110. M (1), x M (2), ..., x M Instead of (T), the downmix signal x input from the time shift unit 191 M (1), x M (2), ..., x M (T) or delayed downmix signal x M' (1), x M' (2), ..., x M' This is done using (T) (steps S120, S130, S140, S150). That is, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 use the downmix signal x determined by the time shift unit 191. M (1), x M(2), ..., x M (T) or delayed downmix signal x M' (1), x M' (2), ..., x M' Using (T), the same operation as described in the first reference form is performed. Note that the time shift unit 191 is downmixed signal x M (1), x M (2), ..., x M Replace (T) with the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M Output (T) and delay downmix signal x M' (1), x M' (2), ..., x M' (T) is replaced with the delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' When (T) is output, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 receive the quantized downmix signal ^x from the time shift unit 191. M (1), ^x M (2), ..., ^x M (T) or delayed quantized downmix signal^x M' (1), ^x M' (2), ..., ^x M' The above process is performed using (T).
[0109] ≪Decoding device 201≫ The decoding device 201 of the second reference embodiment, as shown in Figure 12, includes a monaural decoding unit 210, a stereo decoding unit 220, a left channel subtractive gain decoding unit 230, a left channel signal summing unit 240, a right channel subtractive gain decoding unit 250, a right channel signal summing unit 260, a left-right time difference decoding unit 271, and a time shift unit 281. The decoding device 201 of the second reference embodiment differs from the decoding device 200 of the first reference embodiment in that, in addition to the codes described above, the left-right time difference code Cτ, which will be described later, is also input, it includes a left-right time difference decoding unit 271 and a time shift unit 281, and the left channel signal summing unit 240 and the right channel signal summing unit 260 use the signal output by the time shift unit 281 instead of the signal output by the monaural decoding unit 210. The other configurations and operations of the decoding device 201 of the second reference embodiment are the same as those of the decoding device 200 of the first reference embodiment. The decoding device 201 of the second reference embodiment performs the processing from steps S210 to S281, as illustrated in Figure 13, for each frame. The differences between the decoding device 201 of the second reference embodiment and the decoding device 200 of the first reference embodiment will be explained below.
[0110] [Left-right time difference decoding unit 271] The left-right time difference decoding unit 271 receives the left-right time difference code Cτ input to the decoding device 201. The left-right time difference decoding unit 271 decodes the left-right time difference code Cτ using a predetermined decoding method to obtain and output the left-right time difference τ (step S271). The predetermined decoding method is a decoding method that corresponds to the encoding method used by the left-right relationship information estimation unit 181 of the corresponding encoding device 101. The left-right time difference τ obtained by the left-right time difference decoding unit 271 is the same value as the left-right time difference τ obtained by the left-right relationship information estimation unit 181 of the corresponding encoding device 101, and τ max from τ min It is any value within the range up to [a certain value].
[0111] [Time Shift Section 281] The time shift unit 281 receives the monaural decoded sound signal^x output by the monaural decoding unit 210. M (1), ^x M (2), ..., ^x M(T) and the left-right time difference τ output by the left-right time difference decoding unit 271 are input. The time shift unit 281 receives the monaural decoded sound signal ^x when the left-right time difference τ is a positive value (i.e., when the left-right time difference τ indicates that the left channel is ahead). M (1), ^x M (2), ..., ^x M (T) is output directly to the left channel signal summer 240 (i.e., it is decided to use it in the left channel signal summer 240), and the monaural decoded sound signal is delayed by |τ| samples to ^x M (1-|τ|), ^x M (2-|τ|), ..., ^x M The delayed monaural decoded sound signal ^x is (T-|τ|). M' (1), ^x M' (2), ..., ^x M' (T) is output to the right channel signal summer 260 (i.e., it is decided to use it in the right channel signal summer 260), and if the left-right time difference τ is a negative value (i.e., the left-right time difference τ indicates that the right channel is ahead), the monaural decoded sound signal is |τ|sample delayed to ^x M (1-|τ|), ^x M (2-|τ|), ..., ^x M The delayed monaural decoded sound signal ^x is (T-|τ|). M' (1), ^x M' (2), ..., ^x M' (T) is output to the left channel signal summer 240 (i.e., it is decided to use it in the left channel signal summer 240), and the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M (T) is output directly to the right channel signal summer 260 (i.e., it is decided to use it in the right channel signal summer 260), and if the left-right time difference τ is 0 (i.e., the left-right time difference τ indicates that neither channel is preceding the other), the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M(T) is output as is to the left channel signal summer 240 and the right channel signal summer 260 (i.e., it is decided to use it in the left channel signal summer 240 and the right channel signal summer 260) (step S281). Note that the time shift unit 281 uses the monaural decoded sound signals of past frames to obtain the delayed monaural decoded sound signal, so the storage unit (not shown) in the time shift unit 281 stores the monaural decoded sound signals input in past frames for a predetermined number of frames.
[0112] [Left channel signal summing unit 240, right channel signal summing unit 260] The left channel signal summer 240 and the right channel signal summer 260 perform the same operations as described in the first reference embodiment, processing the monaural decoded sound signal^x output by the monaural decoding unit 210. M (1), ^x M (2), ..., ^x M Instead of (T), the monaural decoded sound signal ^x input from the time shift unit 281 M (1), ^x M (2), ..., ^x M (T) or delayed mono decoded sound signal^x M' (1), ^x M' (2), ..., ^x M' This is done using (T) (steps S240, S260). That is, the left channel signal summing unit 240 and the right channel signal summing unit 260 use the monaural decoded sound signal^x determined by the time shift unit 281. M (1), ^x M (2), ..., ^x M (T) or delayed mono decoded sound signal^x M' (1), ^x M' (2), ..., ^x M' Use (T) to perform the same action as described in the first reference form.
[0113] <First Embodiment> The first embodiment is a modification of the encoding device 101 of the second reference embodiment, which generates a downmix signal considering the relationship between the input audio signal of the left channel and the input audio signal of the right channel. The encoding device of the first embodiment will be described below. Note that the code obtained by the encoding device of the first embodiment can be decoded by the decoding device 201 of the second reference embodiment, so the description of the decoding device will be omitted.
[0114] ≪Encoding device 102≫ As shown in Figure 10, the encoding device 102 of the first embodiment includes a downmix unit 112, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, a left / right relationship information estimation unit 182, and a time shift unit 191. The encoding device 102 of the first embodiment differs from the encoding device 101 of the second reference embodiment in that it includes a left / right relationship information estimation unit 182 instead of a left / right relationship information estimation unit 181, a downmix unit 112 instead of a downmix unit 110, and, as shown by the dashed line in Figure 10, the left / right relationship information estimation unit 182 obtains and outputs a left / right correlation coefficient γ and preceding channel information, and the output left / right correlation coefficient γ and preceding channel information are input to the downmix unit 112 and used. The other configurations and operations of the encoding device 102 of the first embodiment are the same as those of the encoding device 101 of the second reference embodiment. The encoding device 102 of the first embodiment performs the processing from steps S112 to S191, as illustrated in Figure 14, for each frame. The differences between the encoding device 102 of the first embodiment and the encoding device 101 of the second reference embodiment will be explained below.
[0115] [Left-right relationship information estimation unit 182] The left-right relationship information estimation unit 182 receives the input sound signal of the left channel and the input sound signal of the right channel, both input to the encoding device 102. The left-right relationship information estimation unit 182 obtains and outputs the left-right time difference τ, the left-right time difference code Cτ which represents the left-right time difference τ, the left-right correlation coefficient γ, and the preceding channel information from the input sound signal of the left channel and the input sound signal of the right channel (step S182). The process by which the left-right relationship information estimation unit 182 obtains the left-right time difference τ and the left-right time difference code Cτ is the same as that of the left-right relationship information estimation unit 181 in the second reference embodiment.
[0116] The left-right correlation coefficient γ is information corresponding to the correlation coefficient between the sound signal that reaches the left channel microphone and is picked up from the sound source, and the sound signal that reaches the right channel microphone and is picked up from the same sound source, in the assumption described above in the section explaining the left-right relationship information estimation unit 181 of the second reference form. The leading channel information is information corresponding to which microphone the sound emitted by the sound source reaches first, and it is information that indicates whether the same sound signal is included first in the input sound signal of the left channel or the input sound signal of the right channel, and it is information that indicates which channel, the left channel or the right channel, is leading.
[0117] In the example described above in the section explaining the left-right relationship information estimation unit 181 of the second reference form, the left-right relationship information estimation unit 182 calculates the correlation value between the sample sequence of the input sound signal of the left channel and the sample sequence of the input sound signal of the right channel, which is located at a position shifted later than the said sample sequence by a left-right time difference τ, i.e., τ max from τ min Number of candidate samples up to τ cand The correlation value γ calculated for candThe maximum value among these is obtained and output as the left-right correlation coefficient γ. Furthermore, if the left-right time difference τ is a positive value, the left-right relationship information estimation unit 182 obtains and outputs information indicating that the left channel is leading as leading channel information, and if the left-right time difference τ is a negative value, it obtains and outputs information indicating that the right channel is leading as leading channel information. If the left-right time difference τ is 0, the left-right relationship information estimation unit 182 may obtain and output information indicating that the left channel is leading as leading channel information, or it may obtain and output information indicating that the right channel is leading as leading channel information, but it is preferable to obtain and output information indicating that neither channel is leading as leading channel information.
[0118] [Downmix section 112] The downmix unit 112 receives the left channel input sound signal input to the encoding device 102, the right channel input sound signal input to the encoding device 102, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 182, and the leading channel information output by the left-right relationship information estimation unit 182. The downmix unit 112 obtains and outputs a downmix signal by weighting the left channel input sound signal and the right channel input sound signal so that the input sound signal of the leading channel of the right channel input sound signal is included in the downmix signal to a greater extent as the left-right correlation coefficient γ increases (step S112).
[0119] For example, if the absolute value or normalized value of the correlation coefficient is used for the correlation value as described in the explanation section of the left-right relationship information estimation unit 181 of the second reference form, then the obtained left-right correlation coefficient γ will be a value between 0 and 1. Therefore, the downmix unit 112 uses the weight determined by the left-right correlation coefficient γ for each corresponding sample number t to determine the input sound signal x of the left channel. L (t) and the input audio signal x of the right channel R (t) is weighted and added together to obtain the downmix signal x M(t) is sufficient. Specifically, the downmix unit 112 will use x if the preceding channel information indicates that the left channel is preceding, that is, if the left channel is preceding. M (t) = ((1+γ) / 2) × x L (t) + ((1-γ) / 2) × x R (t) If the preceding channel information indicates that the right channel is preceding, i.e., if the right channel is preceding, x M (t) = ((1-γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t), as the downmix signal x M The goal is to obtain (t). The downmix unit 112 obtains the downmix signal in this way, and the smaller the left-right correlation coefficient γ is, that is, the smaller the correlation between the input sound signal of the left channel and the input sound signal of the right channel, the closer the downmix signal is to the signal obtained by averaging the input sound signals of the left channel and the right channel. The larger the left-right correlation coefficient γ is, that is, the larger the correlation between the input sound signal of the left channel and the input sound signal of the right channel, the closer the downmix signal is to the input sound signal of the preceding channel between the input sound signal of the left channel and the input sound signal of the right channel.
[0120] Furthermore, when neither channel is preceding, the downmix unit 112 should average the input sound signals of the left and right channels to obtain and output a downmix signal so that the input sound signals of the left and right channels are included in the downmix signal with equal weight. Therefore, when the preceding channel information indicates that neither channel is preceding, the downmix unit 112 should, for each sample number t, average the input sound signal x of the left channel. L (t) and the input audio signal x of the right channel R (t) averaged x M (t)=(x L (t)+x R (t) / 2 downmix signal x M Let (t) be the case.
[0121] <Second Embodiment> The encoding device 100 of the first reference embodiment may also be modified to generate a downmix signal considering the relationship between the input audio signal of the left channel and the input audio signal of the right channel, and this embodiment will be described as the second embodiment. Note that the code obtained by the encoding device of the second embodiment can be decoded by the decoding device 200 of the first reference embodiment, so the description of the decoding device will be omitted.
[0122] ≪Encoding device 103≫ As shown in Figure 1, the encoding device 103 of the second embodiment includes a downmix unit 112, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, and a left / right relationship information estimation unit 183. The encoding device 103 of the second embodiment differs from the encoding device 100 of the first reference embodiment in that it includes a downmix unit 112 instead of a downmix unit 110, and, as shown by the dashed line in Figure 1, includes a left / right relationship information estimation unit 183. The left / right relationship information estimation unit 183 obtains and outputs a left / right correlation coefficient γ and preceding channel information, and the outputted left / right correlation coefficient γ and preceding channel information are input to the downmix unit 112 and used. The other configurations and operations of the encoding device 103 of the second embodiment are the same as those of the encoding device 100 of the first reference embodiment. Furthermore, the operation of the downmix unit 112 of the encoding device 103 in the second embodiment is the same as the operation of the downmix unit 112 of the encoding device 102 in the first embodiment. The encoding device 103 in the second embodiment performs the processing from steps S112 to S183, as illustrated in Figure 15, for each frame. The differences between the encoding device 103 in the second embodiment and both the encoding device 100 in the first reference embodiment and the encoding device 102 in the first embodiment will be explained below.
[0123] [Left-right relationship information estimation unit 183] The left-right relationship information estimation unit 183 receives the input sound signal of the left channel and the input sound signal of the right channel, both input to the encoding device 103. The left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ and leading channel information from the input sound signals of the left and right channels (step S183).
[0124] The left-right correlation coefficient γ and leading channel information obtained and output by the left-right relationship information estimation unit 183 are the same as those described in the first embodiment. That is, the left-right relationship information estimation unit 183 can be the same as the left-right relationship information estimation unit 182, except that it does not need to obtain and output the left-right time difference τ and the left-right time difference code Cτ.
[0125] For example, the left-right relationship information estimation unit 183, τ max from τ min Number of candidate samples up to τ cand Regarding this, the sample sequence of the input audio signal for the left channel and the number of candidate samples τ for each channel. cand The correlation value γ between the sample sequence of the right channel input sound signal, which is shifted by a minute from the sample sequence in question, and the sample sequence of the right channel input sound signal. cand The maximum value among them is obtained as the left-right correlation coefficient γ and output, and τ when the correlation value is at its maximum is obtained. cand If the value is positive, information indicating that the left channel is leading is obtained and output as leading channel information, and the τ at which the correlation value is maximum is obtained. cand If the value is negative, information indicating that the right channel is leading is obtained and output as leading channel information. The left-right relationship information estimation unit 183 determines the τ when the correlation value is at its maximum. cand If the value is 0, information indicating that the left channel is leading may be obtained and output as leading channel information, or information indicating that the right channel is leading may be obtained and output as leading channel information, but it is preferable to obtain and output information indicating that neither channel is leading as leading channel information.
[0126] <Third Embodiment> Even for an encoding device that stereo encodes the input audio signals of each channel rather than the difference signals of each channel, a configuration may be adopted that considers the relationship between the input audio signals of the left channel and the input audio signals of the right channel to obtain a downmix signal, and this configuration will be described as a third embodiment.
[0127] <<Encoding device 104>> The encoding device 104 of the third embodiment includes a left / right relationship information estimation unit 183, a downmix unit 112, a monaural encoding unit 160, and a stereo encoding unit 174, as shown in Figure 16. For each frame, the encoding device 104 of the third embodiment performs the processing of steps S183, S112, S160, and S174, as illustrated in Figure 17. The encoding device 104 of the third embodiment will be described below with due reference to the description of the second embodiment as appropriate.
[0128] [Left-right relationship information estimation unit 183] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment. The left-right relationship information estimation unit 183 receives the input sound signal of the left channel and the input sound signal of the right channel, which are input to the encoding device 104. The left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and the leading channel information, which is information indicating which of the left channel input sound signal and the right channel input sound signal is leading (step S183).
[0129] [Downmix section 112] The downmix unit 112 is the same as the downmix unit 112 of the second embodiment. The downmix unit 112 receives the input sound signal of the left channel input to the encoding device 104, the input sound signal of the right channel input to the encoding device 104, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 183, and the preceding channel information output by the left-right relationship information estimation unit 183. The downmix unit 112 obtains and outputs a downmix signal by weighting the input sound signals of the left channel and the right channel input sound signals so that the input sound signal of the preceding channel of the left channel input sound signal is included in the downmix signal to a greater extent as the left-right correlation coefficient γ increases (step S112).
[0130] For example, let the sample number be t, and the input audio signal of the left channel be x L (t) Let the input audio signal of the right channel be x R (t) Let the downmix signal be x M (t) If the leading channel information indicates that the left channel is leading, the downmix unit 112 will, for each sample number t, x M (t) = ((1+γ) / 2) × x L (t) + ((1-γ) / 2) × x R (t) is used to obtain a downmix signal, and if the leading channel information indicates that the right channel is leading, then for each sample number t, x M (t) = ((1-γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t) is used to obtain a downmix signal, and if the leading channel information indicates that no channel is leading, then for each sample number t, x M (t)=(x L (t)+x R A downmix signal is obtained by (t) / 2.
[0131] [Monaural encoding section 160] The monaural encoding unit 160 is the same as the monaural encoding unit 160 of the second embodiment. The monaural encoding unit 160 receives the downmix signal output by the downmix unit 112 as input. The monaural encoding unit 160 encodes the input downmix signal to obtain a monaural code CM and outputs it (step S160). The monaural encoding unit 160 may use any encoding scheme, for example, an encoding scheme such as the 3GPP EVS standard may be used. The encoding scheme may be an encoding scheme that performs encoding processing independently of the stereo encoding unit 174, which will be described later, that is, an encoding scheme that performs encoding processing without using the stereo code CS' obtained by the stereo encoding unit 174 or the information obtained in the encoding processing performed by the stereo encoding unit 174, or an encoding scheme that performs encoding processing using the stereo code CS' obtained by the stereo encoding unit 174 or the information obtained in the encoding processing performed by the stereo encoding unit 174.
[0132] [Stereo encoding section 174] The stereo encoding unit 174 receives the left channel input sound signal and the right channel input sound signal input to the encoding device 104. The stereo encoding unit 174 encodes the input left channel sound signal and the input right channel sound signal to obtain and output a stereo code CS' (step S174). The stereo encoding unit 174 may use any encoding scheme; for example, it may use a stereo encoding scheme that corresponds to the stereo decoding scheme of the MPEG-4 AAC standard, or it may use an encoding scheme that encodes the input left channel sound signal and the input right channel sound signal independently, and the stereo code CS' can be obtained by combining all the codes obtained by encoding. The encoding method may be an encoding method that performs encoding processing independently of the monaural encoding unit 160, that is, an encoding method that performs encoding processing without using the monaural code CM obtained by the monaural encoding unit 160 or the information obtained in the encoding processing performed by the monaural encoding unit 160, or it may be an encoding method that performs encoding processing using the monaural code CM obtained by the monaural encoding unit 160 or the information obtained in the encoding processing performed by the monaural encoding unit 160.
[0133] <Fourth Embodiment> As can be seen from the above descriptions of embodiments, as long as the encoding device at least encodes the downmix signal obtained from the left channel input sound signal and the right channel input sound signal to obtain a code, any encoding device may adopt a configuration that considers the relationship between the left channel input sound signal and the right channel input sound signal to obtain the downmix signal. Furthermore, not limited to encoding devices, as long as the signal processing device at least processes the downmix signal obtained from the left channel input sound signal and the right channel input sound signal to obtain a signal processing result, any signal processing device may adopt a configuration that considers the relationship between the left channel input sound signal and the right channel input sound signal to obtain the downmix signal. Moreover, as a downmix device used in front of these encoding devices and signal processing devices, a configuration that considers the relationship between the left channel input sound signal and the right channel input sound signal to obtain the downmix signal may be adopted. These forms will be described as the fourth embodiment.
[0134] ≪Sound signal encoding device 105≫ The audio signal coding device 105 of the fourth embodiment includes a left / right relationship information estimation unit 183, a downmix unit 112, and an coding unit 195, as shown in Figure 18. For each frame, the audio signal coding device 105 of the fourth embodiment performs the processing of steps S183, S112, and S195 illustrated in Figure 19. The audio signal coding device 105 of the fourth embodiment will be described below with due reference to the description of the second embodiment.
[0135] [Left-right relationship information estimation unit 183] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment, and from the input left channel sound signal and the input right channel sound signal, it obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the input left channel sound signal and the input right channel sound signal, and the leading channel information, which is information indicating which of the left channel sound signal and the input right channel sound signal is leading (step S183).
[0136] [Downmix section 112] The downmix unit 112 is the same as the downmix unit 112 of the second embodiment, and obtains and outputs a downmix signal by weighting and averaging the input sound signals of the left channel and the right channel so that the input sound signal of the preceding channel among the input sound signals of the left channel and the input sound signal of the right channel is included in larger proportions as the left-right correlation coefficient γ increases (step S112).
[0137] [Encoding unit 195] The encoding unit 195 receives at least the downmix signal output by the downmix unit 112. The encoding unit 195 encodes the input downmix signal to obtain an audio signal code, which is then output (step S195). The encoding unit 195 may also encode the input audio signals of the left channel and the right channel, and may include the code obtained from this encoding in the audio signal code before outputting it. In this case, as shown by the dashed lines in Figure 18, the encoding unit 195 receives the input audio signals of the left channel and the right channel.
[0138] ≪Sound signal processing device 305≫ The audio signal processing device 305 of the fourth embodiment includes a left / right relationship information estimation unit 183, a downmix unit 112, and a signal processing unit 315, as shown in Figure 20. For each frame, the audio signal processing device 305 of the fourth embodiment performs the processing of steps S183, S112, and S315 as illustrated in Figure 21. The differences between the audio signal processing device 305 of the fourth embodiment and the audio signal encoding device 105 of the fourth embodiment will be explained below.
[0139] [Signal processing unit 315] The signal processing unit 315 receives at least the downmix signal output by the downmix unit 112. The signal processing unit 315 performs at least signal processing on the input downmix signal to obtain and output the signal processing result (step S315). The signal processing unit 315 may also process the input sound signals of the left channel and the right channel to obtain the signal processing result. In this case, as shown by the dashed lines in Figure 20, the signal processing unit 315 also receives the input sound signals of the left channel and the right channel. For example, the signal processing unit 315 may perform signal processing on the input sound signals of each channel using the downmix signal to obtain the output sound signals of each channel as the signal processing result. Alternatively, it may perform this signal processing on the decoded sound signals of the left channel and the decoded sound signals of the right channel obtained by decoding the code CS' obtained by the stereo coding unit 174 in the third embodiment using a decoding device equipped with a decoding unit corresponding to the stereo coding unit 174. In other words, it is not essential that the left channel input sound signal and the right channel input sound signal input to the sound signal processing device 305 are digital audio signals or sound signals obtained by AD conversion after being picked up by two microphones. The left channel input sound signal and the right channel input sound signal input to the sound signal processing device 305 may be decoded sound signals for the left channel and decoded sound signals for the right channel obtained by decoding the codes, or any sound signal obtained in any way as long as it is a stereo two-channel sound signal.
[0140] When the left channel input sound signal and the right channel input sound signal input to the sound signal processing device 305 are decoded left channel sound signal and decoded right channel sound signal obtained by decoding the codes in a separate device, either or both of the left-right correlation coefficient γ and the preceding channel information obtained by the left-right relationship information estimation unit 183 may be obtained by the separate device. If either or both of the left-right correlation coefficient γ and the preceding channel information are obtained by the separate device, as shown by the dashed line in Figure 20, the sound signal processing device 305 should be configured to receive either or both of the left-right correlation coefficient γ and the preceding channel information obtained by the separate device. In this case, the left-right relationship information estimation unit 183 only needs to obtain the left-right correlation coefficient γ or the preceding channel information that was not input to the sound signal processing device 305. If both the left-right correlation coefficient γ and the preceding channel information are input to the sound signal processing device 305, the sound signal processing device 305 does not need to have the left-right relationship information estimation unit 183 and step S183 does not need to be performed. In other words, the sound signal processing device 305, as shown by the dashed line in Figure 20, includes a left-right relationship information acquisition unit 185, and the left-right relationship information acquisition unit 185 should acquire and output the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and the leading channel information, which is information indicating which of the input sound signal of the left channel or the input sound signal of the right channel is preceding (step S185). It should also be said that the left-right relationship information estimation unit 183 and step S183 of each of the above-mentioned devices are also within the scope of the left-right relationship information acquisition unit 185 and step S185.
[0141] ≪Audio Signal Downmixing Device 405≫ The audio signal downmixing device 405 of the fourth embodiment includes a left / right relationship information acquisition unit 185 and a downmixing unit 112, as shown in Figure 22. The audio signal downmixing device 405 performs the processing of steps S185 and S112 illustrated in Figure 23 for each frame. The audio signal downmixing device 405 will be described below with due reference to the description of the second embodiment. Similar to the audio signal processing device 305, the input audio signal of the left channel and the input audio signal of the right channel input to the audio signal downmixing device 405 may be digital audio signals or sound signals obtained by AD conversion after being picked up by two microphones, or they may be decoded audio signals of the left channel and right channel obtained by decoding the codes, or they may be any audio signals obtained as long as they are stereo 2-channel audio signals.
[0142] [Left-right relationship information acquisition unit 185] The left-right relationship information acquisition unit 185 obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and the leading channel information, which is information indicating which of the input sound signal of the left channel or the input sound signal of the right channel is leading (step S185).
[0143] If both the left-right correlation coefficient γ and the preceding channel information are obtained by a separate device, the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ and the preceding channel information input to the sound signal downmixing device 405 from the separate device and outputs them to the downmixing unit 112, as shown by the dashed line in Figure 22.
[0144] If both the left-right correlation coefficient γ and the preceding channel information are not obtained by a separate device, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183, as shown by the dashed line in Figure 22. The left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and the preceding channel information from the input sound signal of the left channel and the input sound signal of the right channel, similar to the left-right relationship information estimation unit 183 of the second embodiment, and outputs them to the downmix unit 112.
[0145] If either the left-right correlation coefficient γ or the preceding channel information is not obtained by another device, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183, as shown by the dashed line in Figure 22. The left-right relationship information estimation unit 183 of the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ or the preceding channel information that is not obtained by another device from the input sound signal of the left channel and the input sound signal of the right channel, similar to the left-right relationship information estimation unit 183 of the second embodiment, and outputs it to the downmix unit 112. For the left-right correlation coefficient γ or the preceding channel information obtained by another device, the left-right relationship information acquisition unit 185 outputs the left-right correlation coefficient γ or the preceding channel information input to the sound signal downmix device 405 from another device to the downmix unit 112, as shown by the dashed line in Figure 22.
[0146] [Downmix section 112] The downmix unit 112 is the same as the downmix unit 112 of the second embodiment. Based on the leading channel information and left-right correlation coefficient acquired by the left-right relationship information acquisition unit 185, the downmix unit obtains and outputs a downmix signal by weighting and averaging the input sound signals of the left channel and the right channel, such that the input sound signal of the leading channel among the left channel input sound signals is included in the downmix signal to a greater extent as the left-right correlation coefficient γ increases (step S112).
[0147] For example, let the sample number be t, and the input audio signal of the left channel be x L (t) Let the input audio signal of the right channel be x R (t) Let the downmix signal be x M (t) If the leading channel information indicates that the left channel is leading, the downmix unit 112 will, for each sample number t, x M (t) = ((1+γ) / 2) × x L (t) + ((1-γ) / 2) × x R (t) is used to obtain a downmix signal, and if the leading channel information indicates that the right channel is leading, then for each sample number t, x M(t) = ((1-γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t) is used to obtain a downmix signal, and if the leading channel information indicates that no channel is leading, then for each sample number t, x M (t)=(x L (t)+x R A downmix signal is obtained by (t) / 2.
[0148] <Program and recording medium> The processing of each part of the above-mentioned encoding device, decoding device, sound signal encoding device, sound signal processing device, and sound signal downmixing device may be implemented by a computer. In this case, the processing content of the functions that each device should have is described by a program. This program is then loaded into the memory unit 1020 of the computer 1000 shown in Figure 24, and the arithmetic processing unit 1010, input unit 1030, output unit 1040, etc. are made to operate, thereby realizing the various processing functions of each of the above-mentioned devices on the computer.
[0149] The program describing this process can be recorded on a computer-readable recording medium. Computer-readable recording media are, for example, non-temporary recording media, specifically magnetic recording devices, optical discs, etc.
[0150] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0151] A computer executing such a program first stores the program recorded on a portable recording medium or transferred from a server computer in its own non-temporary storage device, the auxiliary recording unit 1050. Then, when processing is to be executed, the computer reads the program stored in the auxiliary recording unit 1050 into the storage unit 1020 and executes the processing according to the loaded program. Alternatively, the computer may directly read the program from the portable recording medium into the storage unit 1020 and execute the processing according to that program. Furthermore, each time a program is transferred to this computer from a server computer, it may sequentially execute the processing according to the received program. Alternatively, the above processing may be executed by a so-called ASP (Application Service Provider) type service, where the server computer does not transfer programs to this computer, but the processing function is realized only by execution instructions and result acquisition. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (data that is not a direct instruction to the computer but has the property of defining the processing of the computer).
[0152] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware.
[0153] It goes without saying that the invention may be modified as appropriate without departing from its spirit.
Claims
1. A method for downmixing audio signals to obtain a downmix signal, which is a signal obtained by mixing the first channel input audio signal and the second channel input audio signal, The steps include obtaining a code that indicates which of the first channel input sound signal and the second channel input sound signal is preceding, The system includes a downmix step that obtains the downmix signal from the preceding channel and the other channel, based on a degree determined by the aforementioned symbol and a correlation coefficient which is a coefficient indicating the magnitude of the correlation between the first channel input sound signal and the second channel input sound signal. In the downmix step, the signal that precedes the first channel input audio signal and the second channel input audio signal is given a greater weight than the other signal. Furthermore, the step includes obtaining a time difference from the first channel input sound signal and the second channel input sound signal. The sign indicates the positive or negative sign of the time difference. Audio signal downmixing method.
2. An audio signal downmixing device that obtains a downmix signal, which is a signal obtained by mixing the first channel input audio signal and the second channel input audio signal, A left-right relationship information acquisition unit acquires a code indicating which of the first channel input sound signal and the second channel input sound signal is preceding, The system includes a downmixing unit that obtains the downmix signal from the preceding channel and the other channel, based on a degree determined by the aforementioned symbol and a correlation coefficient which is a coefficient indicating the magnitude of the correlation between the first channel input sound signal and the second channel input sound signal. In the downmix section, the signal that precedes the first channel input audio signal and the second channel input audio signal is given a greater weight than the other signal. A time difference is obtained from the first channel input sound signal and the second channel input sound signal. The sign indicates the positive or negative sign of the time difference. Audio signal downmixing device.
3. A program for causing a computer to execute each step of an audio signal downmixing method, which is a signal obtained by mixing the first channel input audio signal and the second channel input audio signal, The aforementioned audio signal downmixing method is: The steps include obtaining a code that indicates which of the first channel input sound signal and the second channel input sound signal is preceding, Based on the degree determined by the aforementioned symbol and the relationship coefficient which is a coefficient indicating the magnitude of the correlation between the first channel input sound signal and the second channel input sound signal, a downmix step is performed to obtain the downmix signal from the preceding channel and the other channel. It has, In the downmix step, the signal that precedes the first channel input audio signal and the second channel input audio signal is given a greater weight than the other signal. Furthermore, the step includes obtaining a time difference from the first channel input sound signal and the second channel input sound signal. The sign indicates the positive or negative sign of the time difference. program.