Sound signal downmix method, sound signal downmix device, and program

The sound signal downmixing method addresses the challenge of obtaining a useful monaural signal from two-channel audio signals by determining the leading channel and its correlation, enabling efficient encoding and processing while maintaining sound quality.

JP2025087910AActive Publication Date: 2025-06-10NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025041391
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-09
Filing Date
2025-03-14
Publication Date
2025-06-10
Estimated Expiration
2040-11-04

AI Technical Summary

Technical Problem

Existing techniques for obtaining a monaural audio signal from two-channel audio signals lack effective methods for signal processing, such as encoding, as they rely on averaging the left and right channels to obtain the monaural signal.

Method used

A sound signal downmixing method that determines a code indicating the leading channel and its correlation with the other channel, then uses this information to weight the signals appropriately and generate a downmix signal for processing.

Benefits of technology

This method allows for the effective extraction of a monaural signal suitable for signal processing, such as encoding, from two-channel audio signals, improving coding efficiency and maintaining sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087910000001_ABST
    Figure 2025087910000001_ABST
Patent Text Reader

Abstract

To provide a technique for obtaining a monaural signal useful for signal processing such as coding from two-channel sound signals.SOLUTION: A sound signal downmix device 405 is a device for obtaining a downmix signal that is a mixed signal of a first channel input sound signal and a second channel input sound signal. The sound signal downmix device 405 has: a right / left relation information acquiring unit 185 that acquires a sign indicating whether the first channel input sound signal or the second channel input sound signal precedes; and a downmix unit 112 that obtains downmixed signals from the preceding channel and the other channel, based on a degree determined based on a relational coefficient which is a coefficient indicating a magnitude of correlation among the sign, the first channel input sound signal, and the second channel input sound signal. In the downmix section 112, the preceding signal between the first channel input sound signal and the second channel input sound signal is given a weight greater than that of the other signal.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for obtaining a monaural audio signal from a two-channel audio signal in order to encode an audio signal in monaural, or to encode an audio signal by using both monaural encoding and stereo encoding, or to perform signal processing on an audio signal in monaural, or to perform signal processing on a stereo audio signal by using a monaural audio signal.

Background Art

[0002] As a technique for obtaining a monaural audio signal from a two-channel audio signal and performing embedded encoding / decoding on the two-channel audio signal and the monaural audio signal, there is the technique of Patent Document 1. In Patent Document 1, a monaural signal is obtained by averaging the input left-channel audio signal and the input right-channel audio signal for each corresponding sample, the monaural signal is encoded (monaural encoding) to obtain a monaural code, the monaural code is decoded (monaural decoding) to obtain a monaural locally decoded signal, and for each of the left channel and the right channel, a difference (prediction residual signal) between the input audio signal and a prediction signal obtained from the monaural locally decoded signal is encoded. In the technique of Patent Document 1, for each channel, a signal obtained by giving a delay to the monaural locally decoded signal and giving an amplitude ratio is used as a prediction signal, and a prediction signal having a delay and an amplitude ratio that minimize the error between the input audio signal and the prediction signal is selected, or a prediction signal having a delay difference and an amplitude ratio that maximize the cross-correlation between the input audio signal and the monaural locally decoded signal is used, the input audio signal is subtracted from the prediction signal to obtain a prediction residual signal, and the prediction residual signal is made the target of encoding / decoding, thereby suppressing the deterioration of the sound quality of the decoded audio signal of each channel.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technology of Patent Document 1, the coding efficiency of each channel can be improved by optimizing the delay and amplitude ratio given to the monaural local decoded signal when obtaining the prediction signal. However, in the technology of Patent Document 1, the monaural local decoded signal is obtained by encoding and decoding a monaural signal obtained by averaging the sound signals of the left channel and the right channel. That is, the technology of Patent Document 1 has a problem that there is no contrivance for obtaining a monaural signal useful for signal processing such as encoding processing from the two-channel sound signals. An object of the present invention is to provide a technique for obtaining a monaural signal useful for signal processing such as encoding processing from two-channel sound signals.

Means for Solving the Problems

[0005] One aspect of the present invention is a sound signal downmixing method for obtaining a downmix signal which is a signal obtained by mixing a first channel input sound signal and a second channel input sound signal.

[0006] This sound signal downmixing method includes a step of obtaining a code indicating which of the first channel input sound signal and the second channel input sound signal is leading, and a degree determined based on the code and a relational coefficient which is a coefficient indicating the magnitude of the correlation between the first channel input sound signal and the second channel input sound signal, and a downmixing step of obtaining a downmix signal from the leading channel and the other channel.

[0007] In the downmixing step, a greater weight is given to the leading signal between the first channel input sound signal and the second channel input sound signal than to the other signal.

Effects of the Invention

[0008] According to the present invention, a monaural signal useful for signal processing such as encoding processing can be obtained from two-channel sound signals.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Embodiments for Carrying Out the Invention

[0010] First, the notation method in the specification will be explained. For a certain character x, the superscript "^" such as ^x should originally be written directly above "x". However, due to the constraints of the specification notation, it may also be written as ^x. <First Reference Embodiment> Before explaining the embodiments of the invention, as the first reference embodiment and the second reference embodiment, the encoding device and the decoding device in the forms that are the basis for implementing the invention of the second embodiment and the invention of the first embodiment will be explained. In the specification and the claims, the encoding device may also be referred to as an audio signal encoding device, the encoding method may be referred to as an audio signal encoding method, the decoding device may be referred to as an audio signal decoding device, and the decoding method may be referred to as an audio signal decoding method.

[0011] ≪Encoding Device 100≫ As shown in FIG. 1, the encoding device 100 of the first reference form includes a downmixing unit 110, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a monaural encoding unit 160, and a stereo encoding unit 170. The encoding device 100 encodes an input two-channel stereo time-domain audio signal in units of frames having a predetermined time length of, for example, 20 ms, and obtains and outputs a monaural code CM, a left-channel subtraction gain code Cα, a right-channel subtraction gain code Cβ, and a stereo code CS, which will be described later. The two-channel stereo time-domain audio signal input to the encoding device is, for example, a digital audio signal or an acoustic signal obtained by collecting sound such as voice or music with two microphones respectively and performing AD conversion, and consists of a left-channel input audio signal and a right-channel input audio signal. The codes output by the encoding device, that is, the monaural code CM, the left-channel subtraction gain code Cα, the right-channel subtraction gain code Cβ, and the stereo code CS, are input to the decoding device. The encoding device 100 performs the processing from step S110 to step S170 illustrated in FIG. 2 for each frame.

[0012] [Downmixing Unit 110] The downmixing unit 110 receives the left-channel input audio signal input to the encoding device 100 and the right-channel input audio signal input to the encoding device 100. The downmixing unit 110 obtains and outputs a downmixing signal, which is a signal obtained by mixing the input left-channel input audio signal and right-channel input audio signal (step S110).

[0013] For example, assuming the number of samples per frame is T, the downmixing unit 110 receives the left-channel input audio signal x L (1), x L (2),..., x L (T) and the right-channel input audio signal x R (1), x R (2),..., x R(T) is input. Here, T is a positive integer. For example, if the frame length is 20 ms and the sampling frequency is 32 kHz, then T is 640. The downmixing unit 110 obtains and outputs a series of average values of the sample values for corresponding samples of the input sound signal of the input left channel and the input sound signal of the right channel as a downmixing signal x M (1), x M (2), ..., x M (T). That is, when each sample number is t, x M (t) = (x L (t) + x R (t)) / 2.

[0014] [Left-channel subtraction gain estimation unit 120] The left-channel subtraction gain estimation unit 120 is input with the input sound signal x L (1), x L (2), ..., x L (T) of the left channel input to the encoding device 100, and the downmixing signal x M (1), x M (2), ..., x M (T) output by the downmixing unit 110. The left-channel subtraction gain estimation unit 120 obtains and outputs a left-channel subtraction gain α and a left-channel subtraction gain sign Cα that represents the left-channel subtraction gain α from the input sound signal of the input left channel and the downmixing signal (step S120). The left-channel subtraction gain estimation unit 120 obtains the left-channel subtraction gain α and the left-channel subtraction gain sign Cα by a well-known method exemplified by the method of obtaining the amplitude ratio g in Patent Document 1 and the method of encoding the amplitude ratio g, or by a method based on the principle of minimizing the quantization error newly invented. The principle of minimizing the quantization error and the method based on this principle will be described later.

[0015] [Left-channel signal subtraction unit 130] The left-channel signal subtraction unit 130 is input with the input sound signal x L (1), x L (2), ..., x L(T) and the downmix signal x output by the downmix unit 110 M (1), x M (2), ..., x M (T), and the left-channel subtraction gain α output by the left-channel subtraction gain estimation unit 120 are input. For each corresponding sample t, the left-channel signal subtraction unit 130 multiplies the sample value x M (t) of the downmix signal by the left-channel subtraction gain α to obtain a value α×x M (t) and subtracts it from the sample value x L (t) of the input audio signal of the left channel to obtain a value x L (t)-α×x M (t) to form a series of left-channel difference signals y L (1), y L (2), ..., y L (T) and outputs it (step S130). That is, y L (t)=x L (t)-α×x M (t). In the encoding device 100, in order not to require a delay or arithmetic processing amount for obtaining a local decoded signal, in the left-channel signal subtraction unit 130, instead of the quantized downmix signal that is the local decoded signal of the monaural encoding, the non-quantized downmix signal x M (t) obtained by the downmix unit 110 may be used. However, when the left-channel subtraction gain α is obtained by a well-known method such as that exemplified in Patent Document 1 rather than a method based on the principle of minimizing the quantization error, a means for obtaining a local decoded signal corresponding to the monaural code CM is provided after the monaural encoding unit 160 or within the monaural encoding unit 160 of the encoding device 100. In the left-channel signal subtraction unit 130, instead of the downmix signal x M (1), x M (2), ..., x M (T), similar to a conventional encoding device such as Patent Document 1, the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M (T) that is the local decoded signal of the monaural encoding may be used to obtain the left-channel difference signal.

[0016] [Right channel subtraction gain estimation unit 140] The right channel subtraction gain estimation unit 140 receives the input audio signal x of the right channel input to the encoding device 100 R (1), x R (2), ..., x R (T), and the downmix signal x output from the downmix unit 110 M (1), x M (2), ..., x M (T). The right channel subtraction gain estimation unit 140 obtains and outputs the right channel subtraction gain β and the right channel subtraction gain sign Cβ, which is a sign representing the right channel subtraction gain β, from the input right channel input audio signal and the downmix signal (step S140). The right channel subtraction gain estimation unit 140 obtains the right channel subtraction gain β and the right channel subtraction gain sign Cβ by a well-known method exemplified by the method of obtaining the amplitude ratio g in Patent Document 1 and the method of encoding the amplitude ratio g, or by a method based on the principle of minimizing the quantization error newly invented. The principle of minimizing the quantization error and the method based on this principle will be described later.

[0017] [Right channel signal subtraction unit 150] The right channel signal subtraction unit 150 receives the input audio signal x of the right channel input to the encoding device 100 R (1), x R (2), ..., x R (T), the downmix signal x output from the downmix unit 110 M (1), x M (2), ..., x M (T), and the right channel subtraction gain β output from the right channel subtraction gain estimation unit 140. The right channel signal subtraction unit 150 subtracts the value β×x M (t) obtained by multiplying the sample value x of the downmix signal by the right channel subtraction gain β from the sample value x of the input audio signal of the right channel M (t) to obtain the value x R (t)-β×x R (t) for each corresponding sample t.M The series by (t) is the right-channel difference signal y R (1), y R (2), ..., y R (T) and output it (step S150). That is, y R (t) = x R (t) - β × x M (t). In the right-channel signal subtraction unit 150, similar to the left-channel signal subtraction unit 130, in order not to require the delay and arithmetic processing amount for obtaining the local decoded signal in the encoding device 100, instead of the quantized downmix signal which is the local decoded signal of the monaural encoding, the non-quantized downmix signal x M (t) obtained by the downmix unit 110 may be used. However, when the right-channel subtraction gain β is obtained by a well-known method as exemplified in Patent Document 1 rather than a method based on the principle of minimizing the quantization error, a means for obtaining a local decoded signal corresponding to the monaural code CM is provided in the subsequent stage of the monaural encoding unit 160 of the encoding device 100 or within the monaural encoding unit 160. Similar to the left-channel signal subtraction unit 130, in the right-channel signal subtraction unit 150, the downmix signal x M (1), x M (2), ..., x M (T) is replaced with the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M (T) which is the local decoded signal of the monaural encoding, similar to a conventional encoding device such as Patent Document 1, to obtain the right-channel difference signal.

[0018] [Monaural Encoding Unit 160] The downmix signal x M (1), x M (2), ..., x M (T) output by the downmix unit 110 is input to the monaural encoding unit 160. The monaural encoding unit 160 encodes the input downmix signal in a predetermined encoding method as b MEncode with b bits to obtain and output a monaural code CM (step S160). That is, from the downmixed signals x M (1), x M (2),..., x M (T) of the input T samples, obtain and output a b-bit monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used. M x(1) M x(2) M ... to x(T), obtain and output a b-bit monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used. M Encode with b bits to obtain and output a monaural code CM. Any encoding method may be used as the encoding method. For example, an encoding method such as the 3GPP EVS standard may be used.

[0019] [Stereo Encoding Unit 170] The left-channel difference signals y L (1), y L (2),..., y L (T) output by the left-channel signal subtraction unit 130 and the right-channel difference signals y R (1), y R (2),..., y R (T) output by the right-channel signal subtraction unit 150 are input to the stereo encoding unit 170. The stereo encoding unit 170 encodes the input left-channel difference signal and right-channel difference signal with a total of b bits using a predetermined encoding method to obtain and output a stereo code CS (step S170). That is, from the input left-channel difference signals y L (1), y L (2),..., y L (T) of T samples and the input right-channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left-channel difference signal and right-channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. L y(1) L y(2) L ... to y(T) and the right-channel difference signals y R (1), y R (2),..., y R (T) output by the right-channel signal subtraction unit 150 are input to the stereo encoding unit 170. The stereo encoding unit 170 encodes the input left-channel difference signal and right-channel difference signal with a total of b bits using a predetermined encoding method to obtain and output a stereo code CS (step S170). That is, from the input left-channel difference signals y L (1), y L (2),..., y L (T) of T samples and the input right-channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left-channel difference signal and right-channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. R y(1) R y(2) R ... to y(T), the stereo encoding unit 170 encodes the input left-channel difference signal and right-channel difference signal with a total of b bits using a predetermined encoding method to obtain and output a stereo code CS (step S170). That is, from the input left-channel difference signals y L (1), y L (2),..., y L (T) of T samples and the input right-channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left-channel difference signal and right-channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding. s Encode with b bits to obtain and output a stereo code CS (step S170). That is, from the input left-channel difference signals y L (1), y L (2),..., y L (T) of T samples and the input right-channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. L y(1) L y(2) L ... to y(T) and the input right-channel difference signals y R (1), y R (2),..., y R (T) of T samples, obtain and output a total of b bits of stereo code CS. R y(1) R y(2) R ... to y(T), obtain and output a total of b bits of stereo code CS. S Any encoding method may be used as the encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or a method that independently encodes each of the input left-channel difference signal and right-channel difference signal may be used. The stereo code CS may be the combination of all the codes obtained by encoding.

[0020] When the input left-channel differential signal and right-channel differential signal are each independently encoded, the stereo encoder 170 encodes the left-channel differential signal with b L bits and encodes the right-channel differential signal with b R bits. That is, the stereo encoder 170 obtains a b L (1), y L (2),..., y L (T) from the left-channel differential signal of the input T samples, and obtains a b L bit left-channel differential code CL, and obtains a b R (1), y R (2),..., y R (T) from the right-channel differential signal of the input T samples, and obtains a b R bit right-channel differential code CR, and outputs the combination of the left-channel differential code CL and the right-channel differential code CR as the stereo code CS. Here, the sum of the b L bits and the b R bits is b S bits.

[0021] When the input left-channel differential signal and right-channel differential signal are encoded together in one encoding method, the stereo encoder 170 encodes the left-channel differential signal and the right-channel differential signal with a total of b S bits. That is, the stereo encoder 170 obtains a b L (1), y L (2),..., y L (T) of the input T samples of the left-channel differential signal and a b R (1), y R (2),..., y R (T) of the input T samples of the right-channel differential signal, and obtains and outputs a b S bit stereo code CS.

[0022] <<Decoder 200>> As shown in FIG. 3, the decoder 200 of the first reference embodiment includes a monaural decoder 210, a stereo decoder 220, a left-channel subtraction gain decoder 230, a left-channel signal adder 240, a right-channel subtraction gain decoder 250, and a right-channel signal adder 260. The decoder 200 decodes the input monaural code CM, left-channel subtraction gain code Cα, right-channel subtraction gain code Cβ, and stereo code CS in frame units of the same time length as the corresponding encoder 100, and obtains and outputs a decoded audio signal in the time domain of two-channel stereo in frame units (left-channel decoded audio signal and right-channel decoded audio signal described later). As shown by the dashed line in FIG. 3, the decoder 200 may also output a decoded audio signal in the time domain of monaural (monaural decoded audio signal described later). The decoded audio signal output by the decoder 200 is, for example, DA-converted and reproduced by a speaker to be made audible. For each frame, the decoder 200 performs the processes of steps S210 to S260 illustrated in FIG. 4.

[0023] [Monaural Decoder 210] The monaural code CM input to the decoder 200 is input to the monaural decoder 210. The monaural decoder 210 decodes the input monaural code CM by a predetermined decoding method to obtain a monaural decoded audio signal ^x M (1), ^x M (2), ..., ^x M (T) and outputs them (step S210). As the predetermined decoding method, a decoding method corresponding to the encoding method used in the monaural encoder 160 of the corresponding encoder 100 is used. The number of bits of the monaural code CM is b M is.

[0024] [Stereo Decoder 220] The stereo code CS input to the decoder 200 is input to the stereo decoder 220. The stereo decoder 220 decodes the input stereo code CS by a predetermined decoding method to obtain a left-channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L (T) and a right-channel decoded difference signal ^y R(1), ^y R (2), ..., ^y R (T) and, obtain and output (step S220). As a predetermined decoding method, a decoding method corresponding to the encoding method used in the stereo encoding unit 170 of the corresponding encoding device 100 is used. The total number of bits of the stereo code CS is b S is.

[0025] [Left channel subtraction gain decoding unit 230] The left channel subtraction gain decoding unit 230 receives the left channel subtraction gain code Cα input to the decoding device 200. The left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα to obtain and output the left channel subtraction gain α (step S230). The left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα by a decoding method corresponding to the method used in the left channel subtraction gain estimation unit 120 of the corresponding encoding device 100 to obtain the left channel subtraction gain α. Regarding the method by which the left channel subtraction gain decoding unit 230 decodes the left channel subtraction gain code Cα to obtain the left channel subtraction gain α when the left channel subtraction gain estimation unit 120 of the corresponding encoding device 100 obtains the left channel subtraction gain α and the left channel subtraction gain code Cα based on the principle of minimizing quantization error, it will be described later.

[0026] [Left channel signal addition unit 240] The left channel signal addition unit 240 receives the monaural decoded audio signal ^x output by the monaural decoding unit 210 M (1), ^x M (2), ..., ^x M (T) and, the left channel decoded difference signal ^y output by the stereo decoding unit 220 L (1), ^y L (2), ..., ^y L (T) and, the left channel subtraction gain α output by the left channel subtraction gain decoding unit 230, are input. The left channel signal addition unit 240, for each corresponding sample t, the sample value ^y of the left channel decoded difference signal L (t) and, the sample value ^x of the monaural decoded audio signal M (t) and the value α×^x obtained by multiplying by the left channel subtraction gain αM The value ^y obtained by adding (t) and L (t) + α × ^x M The left-channel decoded audio signal ^x of the series by (t) L (1), ^x L (2), ..., ^x L is obtained and output as (T) (step S240). That is, ^x L (t) = ^y L (t) + α × ^x M is (t).

[0027] [Right-channel subtraction gain decoding unit 250] The right-channel subtraction gain decoding unit 250 receives the right-channel subtraction gain symbol Cβ input to the decoding device 200. The right-channel subtraction gain decoding unit 250 decodes the right-channel subtraction gain symbol Cβ to obtain and output the right-channel subtraction gain β (step S250). The right-channel subtraction gain decoding unit 250 decodes the right-channel subtraction gain symbol Cβ by a decoding method corresponding to the method used by the right-channel subtraction gain estimation unit 140 of the corresponding encoding device 100 to obtain the right-channel subtraction gain β. The method by which the right-channel subtraction gain decoding unit 250 decodes the right-channel subtraction gain symbol Cβ to obtain the right-channel subtraction gain β when the right-channel subtraction gain estimation unit 140 of the corresponding encoding device 100 obtains the right-channel subtraction gain β and the right-channel subtraction gain symbol Cβ based on the principle of minimizing quantization error will be described later.

[0028] [Right-channel signal addition unit 260] The right-channel signal addition unit 260 receives the monaural decoded audio signal ^x output by the monaural decoding unit 210 M (1), ^x M (2), ..., ^x M (T), the right-channel decoded difference signal ^y output by the stereo decoding unit 220 R (1), ^y R (2), ..., ^y R (T), and the right-channel subtraction gain β output by the right-channel subtraction gain decoding unit 250. The right-channel signal addition unit 260, for each corresponding sample t, the sample value ^y of the right-channel decoded difference signalR (t) and the sample value ^x of the monaural decoded sound signal M The value β×^x obtained by multiplying (t) and the right-channel subtraction gain β M The value ^y obtained by adding (t) and R (t)+β×^x M The sequence by (t) is the right-channel decoded sound signal ^x R (1), ^x R (2), ..., ^x R Obtained and output as (T) (step S260). That is, ^x R (t)=^y R (t)+β×^x M (t).

[0029] [Principle of minimizing quantization error] Hereinafter, the principle of minimizing quantization error will be described. When encoding the left-channel difference signal and the right-channel difference signal input to the stereo encoding unit 170 together in one encoding method, the number of bits b used for encoding the left-channel difference signal L And the number of bits b used for encoding the right-channel difference signal R May not be positively determined. However, hereinafter, the number of bits used for encoding the left-channel difference signal is b L And the number of bits used for encoding the right-channel difference signal is b R Will be described. Also, hereinafter, mainly the left channel will be described, but the same applies to the right channel.

[0030] The above-described encoding device 100 subtracts the value obtained by multiplying each sample value of the input sound signal x L (1), x L (2), ..., x L (T) of the left channel from the value obtained by multiplying each sample value of the downmix signal x M (1), x M (2), ..., x M (T) by the left-channel subtraction gain α to obtain the left-channel difference signal y L (1), y L (2), ..., y L (T) by bL The downmix signal x is encoded with M (1), x M (2), ..., x M (T) to b M The above-mentioned decoding device 200 encodes the data with b L Bit code to left channel decoded difference signal ^y L (1), ^y L (2), ..., ^y L (T) (hereinafter also referred to as the “quantized left channel difference signal”), and M From the bit code to the mono decoded sound signal ^x M (1), ^x M (2), ..., ^x M (T) (hereinafter also referred to as the “quantized downmix signal”), the quantized downmix signal ^x obtained by the decoding is M (1), ^x M (2), ..., ^x M The quantized left channel difference signal ^y obtained by multiplying each sample value of (T) by the left channel subtraction gain α is L (1), ^y L (2), ..., ^y L By adding it to each sample value of (T), the left channel decoded sound signal ^x L (1), ^x L (2), ..., ^x L Encoding apparatus 100 and decoding apparatus 200 should be designed so that the energy of the quantization error in the decoded sound signal in the left channel obtained by the above process is small.

[0031] The energy of the quantization error (hereinafter, for convenience, referred to as the "quantization error caused by encoding") in the decoded signal obtained by encoding and decoding the input signal is roughly proportional to the energy of the input signal in many cases, and tends to be exponentially smaller with respect to the number of bits per sample used for encoding. Therefore, the average energy per sample of the quantization error caused by encoding the left channel difference signal is a positive number σ L2 can be estimated as in the following formula (1-0-1), and the average energy per sample of the quantization error caused by encoding the downmix signal is a positive number σ M 2 can be estimated as in the following formula (1-0-2). [Number] [Number]

[0032] Here, assume that the input audio signal x of the left channel L (1), x L (2),..., x L (T) and the downmix signal x M (1), x M (2),..., x M (T) are such that each sample value is close enough to be regarded as the same sequence. For example, the input audio signal x of the left channel L (1), x L (2),..., x L (T) and the input signal x of the right channel R (1), x R (2),..., x R (T) are obtained by picking up the sound emitted by a sound source equidistant from two microphones in an environment with little background noise and reverberation. A case like this corresponds to this condition. Under this condition, each sample value of the left channel difference signal y L (1), y L (2),..., y L (T) is equivalent to the value obtained by multiplying each sample value of the downmix signal x M (1), x M (2),..., x M (T) by (1-α). Therefore, since the energy of the left channel difference signal can be expressed as (1-α) times the energy of the downmix signal, the above σ 2 times, so the above σ L 2 is the above σ M2 is used to obtain (1 - α) 2 ×σ M 2 Since it can be replaced, the average energy per sample of the quantization error caused by encoding the left channel difference signal can be estimated as shown in the following equation (1-1). [Number] Also, the average energy per sample of the quantization error of the signal added to the quantized left channel difference signal in the decoding device, that is, the average energy per sample of the quantization error of the series of values obtained by multiplying each sample value of the quantized downmix signal obtained by decoding by the left channel subtraction gain α, can be estimated as shown in the following equation (1-2). [Number]

[0033] Assuming that the quantization error caused by encoding the left channel difference signal and the quantization error of the series of values obtained by multiplying each sample value of the quantized downmix signal obtained by decoding by the left channel subtraction gain α do not have a correlation with each other, the average energy per sample of the quantization error of the decoded sound signal of the left channel is estimated as the sum of Equation (1-1) and Equation (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error of the decoded sound signal of the left channel is obtained as shown in the following equation (1-3). [Number]

[0034] That is, the input sound signal x of the left channel L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x MIn the condition where each sample value is close enough to be regarded as the same sequence (T), in order to minimize the quantization error of the decoded audio signal of the left channel, the left-channel subtraction gain estimator 120 may obtain the left-channel subtraction gain α using Equation (1-3). The left-channel subtraction gain α obtained by Equation (1-3) is a value greater than 0 and less than 1, and is b, the number of bits used for the two encodings. L and b M is 0.5 when they are equal, and the number of bits b L for encoding the left-channel difference signal is M closer to 0 than 0.5 the more it is than the number of bits b M for encoding the downmix signal, and the number of bits b L for encoding the left-channel difference signal is

[0035] The same applies to the right channel. For the input audio signal x R (1), x R (2), ..., x R (T) of the right channel and the downmix signal x M (1), x M (2), ..., x M (T), in the condition where each sample value is close enough to be regarded as the same sequence (T), in order to minimize the quantization error of the decoded audio signal of the right channel, the right-channel subtraction gain estimator 140 may obtain the right-channel subtraction gain β using the following Equation (1-3-2).

Equation

[0036] Next, the input audio signal x of the left channel L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M including the case where they cannot be regarded as the same sequence, the principle of minimizing the energy of the quantization error of the decoded audio signal of the left channel will be described.

[0037] The input audio signal x of the left channel L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T) of the normalized inner product value r L is represented by the following formula (1-4).

Equation

[0038] The input audio signal x of the left channel L (1), x L (2),..., x L (T) is, for each sample number t, x L (t)=r L ×x M (t)+(x L (t)- r L ×x M (t)) can be decomposed. Here, x L (t)- r L ×x M (t) is a series composed of each value of x L ’(1), x L ’(2),..., x L ’(T). According to this decomposition, each sample value y of the left-channel difference signal L (t)=x L (t)-αx M (t) is the downmix signal x M (1), x M (2),..., x M (T) of each sample value x M (t) multiplied by the normalized inner product value r L and (r L -α) obtained by multiplying by the left-channel subtraction gain α (r L -α)×x M (t), and the sum of each sample value x L ’(t) of the orthogonal signal (r L -α)×x M (t)+x L ’(t) is equivalent. The orthogonal signal x L ’(1), x L ’(2),..., x L’(T) is the downmix signal x M (1), x M (2), ..., x M To show the orthogonality property, i.e., the inner product is 0, for the downmix signals x(1), x(2), ..., x(T), the energy of the left channel difference signal is expressed as the sum of the energy of the downmix signal multiplied by (r L -α) 2 and the energy of the orthogonal signal. Therefore, the average energy per sample of the quantization error generated by encoding the left channel difference signal with b L bits can be estimated as shown in Equation (1-5) using a positive number σ 2 .

Equation

[0039] Assuming that the quantization error generated by encoding the left channel difference signal and the quantization error of the series of values obtained by multiplying each sample value of the quantized downmix signal obtained by decoding by the left channel subtraction gain α have no correlation with each other, the average energy per sample of the quantization error of the decoded sound signal of the left channel is estimated as the sum of Equation (1-5) and Equation (1-2). The left channel subtraction gain α that minimizes the energy of the quantization error of the decoded sound signal of the left channel is obtained as shown in Equation (1-6) below.

Equation

[0040] That is, in order to minimize the quantization error of the decoded sound signal of the left channel, the left channel subtraction gain estimator 120 may obtain the left channel subtraction gain α using Equation (1-6). That is, considering the principle of minimizing the energy of this quantization error, the left channel subtraction gain α includes the normalized inner product value r L and the number of bits b L used for encoding and b MA value determined by [[ID=]] should be used after multiplying it by a correction coefficient. The correction coefficient is a value greater than 0 and less than 1, and is the number of bits b for encoding the left channel difference signal L and the number of bits b for encoding the downmix signal M When they are the same, it is 0.5, and the number of bits b for encoding the left channel difference signal L is the number of bits b for encoding the downmix signal M The larger the number of bits b for encoding the left channel difference signal L is the number of bits b for encoding the downmix signal M The smaller the number of bits b for encoding the left channel difference signal

[0041] The same applies to the right channel. In order to minimize the quantization error of the decoded sound signal of the right channel, the right channel subtraction gain estimator 140 may obtain the right channel subtraction gain β by the following formula (1-6-2).

Equation

Equation

[0042] [Estimation and Decoding of Subtraction Gain Based on the Principle of Minimizing Quantization Error] A specific example of the estimation and decoding of the subtraction gain based on the principle of minimizing the quantization error described above will be explained. In each example, the left-channel subtraction gain estimation unit 120 that estimates the subtraction gain in the encoding device 100, the right-channel subtraction gain estimation unit 140, the left-channel subtraction gain decoding unit 230 that decodes the subtraction gain in the decoding device 200, and the right-channel subtraction gain decoding unit 250 will be described.

[0043] [[Example 1]] Example 1 is based on the principle of minimizing the energy of the quantization error of the decoded audio signal of the left channel, including cases where the input audio signals x L (1), x L (2),..., x L (T) and the downmix signals x M (1), x M (2),..., x M (T) cannot be regarded as the same sequence, and the principle of minimizing the energy of the quantization error of the decoded audio signal of the right channel, including cases where the input audio signals x R (1), x R (2),..., x R (T) and the downmix signals x M (1), x M (2),..., x M (T) cannot be regarded as the same sequence.

[0044] [[[Left-Channel Subtraction Gain Estimation Unit 120]]] The left-channel subtraction gain estimation unit 120 has candidates α for the left-channel subtraction gain cand (a) and a code Cα corresponding to the candidate cand (a), and a plurality of sets (A sets, a = 1, ..., A) of the pair are stored in advance. The left-channel subtraction gain estimation unit 120 performs the following steps S120-11 to S120-14 shown in FIG. 5

[0045] The left-channel subtraction gain estimation unit 120 first inputs the input sound signal x of the left channel L (1), x L (2), ..., x L (T) and the downmix signal x M (1), x M (2), ..., x M (T), and obtains a normalized inner product value r of the downmix signal with respect to the input sound signal of the left channel according to Equation (1-4) L (step S120-11). Further, the left-channel subtraction gain estimation unit 120 uses the number of bits b used for encoding the left-channel difference signal y L (1), y L (2), ..., y L (T) in the stereo encoding unit 170, the number of bits b used for encoding the downmix signal x L (1), x M (2), ..., x M (2), ..., x M (T) in the monaural encoding unit 160, the number of samples T per frame, and obtains a left-channel correction coefficient c according to the following Equation (1-7) M (step S120-12). L

Equation

[0046] Note that when the number of bits b L (1), y L (2),..., y L (T) used for encoding the left-channel difference signal y L is not positively determined, then half of the number of bits b s of the stereo code CS output by the stereo encoding unit 170 (that is, b s / 2) may be used as the number of bits b L . Also, the left-channel correction coefficient c L is not the value obtained from the formula (1-7) itself, but a value greater than 0 and less than 1. When the number of bits b L (1), y L (2),..., y L (T) used for encoding the left-channel difference signal y L is the same as the number of bits b M (1), x M (2),..., x M (T) used for encoding the downmix signal x M , it is 0.5. The more the number of bits b L is more than the number of bits b M , the closer it is to 0 than 0.5. The less the number of bits b L is less than the number of bits b M , the closer it is to 1 than 0.5. The same applies to each example described later.

[0047] [[[Right-channel subtraction gain estimation unit 140]]] The right-channel subtraction gain estimation unit 140 stores a plurality of sets (B sets, b = 1, ..., B) of a candidate β of the right-channel subtraction gain cand (b) and a code Cβ corresponding to the candidate. cand (b). The right-channel subtraction gain estimation unit 140 performs the following steps S140-11 to S140-14 shown in FIG. 5.

[0048] First, the right-channel subtraction gain estimation unit 140 calculates the normalized inner product value r of the input sound signal x of the input right channel R (1), x R (2),..., x R (T) and the downmix signal x M (1), x M (2),..., x M (T) with respect to the input sound signal of the right channel of the downmix signal according to Equation (1-4-2) to obtain R r (step S140-11). Also, the right-channel subtraction gain estimation unit 140 calculates the number of bits b used for encoding the right-channel difference signal y R (1), y R (2),..., y R (T) in the stereo encoding unit 170, the number of bits b used for encoding the downmix signal x R (1), x M (2),..., x M (2),..., x M (T) in the monaural encoding unit 160, the number of samples T per frame, and uses them to obtain the right-channel correction coefficient c M by the following Equation (1-7-2) (step S140-12). R

Equation

[0049] Note that when the number of bits b R (1), y R (2),..., y R (T) used for encoding the right-channel difference signal y R is not positively determined, the number of bits b of the stereo code CS output by the stereo encoder 170 s Half of (that is, b s / 2) may be used as the number of bits b R . Also, the right-channel correction coefficient c R is not the value obtained by the formula (1-7-2) itself, but a value greater than 0 and less than 1. The number of bits b R (1), y R (2),..., y R (T) used for encoding the right-channel difference signal y R and the number of bits b M (1), x M (2),..., x M (T) used for encoding the downmix signal x M are the same, it is 0.5. The more the number of bits b R is more than the number of bits b M , the closer it is to 0 than 0.5. The less the number of bits b R is less than the number of bits b M , the closer it is to 1 than 0.5. The same applies to each example described later.

[0050] [[[Left channel subtraction gain decoding unit 230]]] The left channel subtraction gain decoding unit 230 stores the same candidate α for the left channel subtraction gain as that stored in the left channel subtraction gain estimation unit 120 of the corresponding encoding device 100 cand (a) and the code Cα corresponding to the candidate cand (a) and a plurality of sets (set A, a = 1,..., A) are stored in advance. The left channel subtraction gain decoding unit 230 stores the stored code Cα cand (1),..., Cα cand (A), and obtains the candidate for the left channel subtraction gain corresponding to the input left channel subtraction gain code Cα as the left channel subtraction gain α (step S230-11).

[0051] [[[Right channel subtraction gain decoding unit 250]]] The right channel subtraction gain decoding unit 250 stores the same candidate β for the right channel subtraction gain as that stored in the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 cand (b) and the code Cβ corresponding to the candidate cand (b) and a plurality of sets (set B, b = 1,..., B) are stored in advance. The right channel subtraction gain decoding unit 250 stores the stored code Cβ cand (1),..., Cβ cand (B), and obtains the candidate for the right channel subtraction gain corresponding to the input right channel subtraction gain code Cβ as the right channel subtraction gain β (step S250-11).

[0052] Note that the same candidates and codes for the subtraction gain may be used for the left and right channels. By setting the above-mentioned A and B to the same value, the candidate α for the left channel subtraction gain stored in the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 cand (a) and the code Cα corresponding to the candidate cand (a), and the candidate β for the right channel subtraction gain stored in the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 cand (b) and the code Cβ corresponding to the candidatecand (b) may be the same as the set with (b).

[0053] [[Example 1 Variation]] The number of bits b used for encoding the left channel difference signal in the encoding device 100 L is the number of bits used for decoding the left channel difference signal in the decoding device 200, and the number of bits b used for encoding the downmix signal in the encoding device 100 M Since the value of is the number of bits used for decoding the downmix signal in the decoding device 200, the correction coefficient c L can be calculated with the same value in both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r L Taking the quantization value ^r of the normalized inner product value as the object of encoding and decoding, the correction coefficient c L is multiplied by to obtain the left channel subtraction gain α. The same applies to the right channel. This form will be described as a variation of Example 1. L Lcand

[0054] [[[Left Channel Subtraction Gain Estimation Unit 120]]] The left channel subtraction gain estimation unit 120 stores in advance a plurality of sets (A sets, a = 1,..., A) of the candidate r of the normalized inner product value of the left channel and the code Cα Lcand (a) corresponding to the candidate. As shown in FIG. 6, the left channel subtraction gain estimation unit 120 performs step S120-11 and step S120-12 described in Example 1, and the following step S120-15 and step S120-16. cand (a).

[0055] First, the left channel subtraction gain estimation unit 120, in the same manner as step S120-11 of the left channel subtraction gain estimation unit 120 in Example 1, inputs the input sound signal x of the input left channel L (1), x L (2),..., x L (T) and the downmix signal x M (1), x M (2),..., x MFrom (T), the normalized inner product value r of the input audio signal of the left channel of the downmix signal is obtained by Expression (1-4). L (Step S120-11). Next, the left-channel subtraction gain estimation unit 120 determines the candidate r of the normalized inner product value of the stored left channel Lcand (1),..., r Lcand (A). Among them, the candidate closest to the normalized inner product value r obtained in step S120-11 L is obtained (the quantization value of the normalized inner product value r L ) ^r L . Then, the stored code Cα cand (1),..., Cα cand (A). The code corresponding to the closest candidate ^r among them L is obtained as the left-channel subtraction gain code Cα (step S120-15). Also, in the same manner as step S120-12 of the left-channel subtraction gain estimation unit 120 in Example 1, the left-channel subtraction gain estimation unit 120 determines the number of bits b used for encoding the left-channel difference signals y L (1), y L (2),..., y L (T) in the stereo encoding unit 170 L , the number of bits b used for encoding the downmix signals x M (1), x M (2),..., x M (T) in the monaural encoding unit 160 M , and the number of samples T per frame, and uses them to obtain the left-channel correction coefficient c L by Expression (1-7) (step S120-12). Next, the left-channel subtraction gain estimation unit 120 multiplies the quantization value ^r of the normalized inner product value obtained in step S120-15 L by the left-channel correction coefficient c obtained in step S120-12 L to obtain the left-channel subtraction gain α (step S120-16).

[0056] [[[Right-channel subtraction gain estimation unit 140]]] The right-channel subtraction gain estimation unit 140 is provided with candidates r for the normalized inner product values of the right channel Rcand(b) and the code Cβ corresponding to the candidate cand A plurality of sets (B sets, b = 1, ..., B) of the pair of (b) are stored in advance. As shown in FIG. 6, the right-channel subtraction gain estimation unit 140 performs step S140-11 and step S140-12 described in Example 1, and the following step S140-15 and step S140-16.

[0057] First, the right-channel subtraction gain estimation unit 140, similar to step S140-11 of the right-channel subtraction gain estimation unit 140 in Example 1, calculates the normalized inner product value r of the input sound signal x of the right channel in the input right-channel R (1), x R (2), ..., x R (T) and the downmix signal x M (1), x M (2), ..., x M (T) according to Equation (1-4-2) (step S140-11). Next, the right-channel subtraction gain estimation unit 140 selects the candidate r of the normalized inner product value of the stored right channel R The candidate (quantized value) ^r of the normalized inner product value r that is closest to the normalized inner product value r obtained in step S140-11 among (1), ..., r Rcand (1), ..., r Rcand (B) is obtained, and the code corresponding to the closest candidate ^r among the stored codes Cβ R (1), ..., Cβ R (B) is obtained as the right-channel subtraction gain code Cβ (step S140-15). Also, the right-channel subtraction gain estimation unit 140, similar to step S140-12 of the right-channel subtraction gain estimation unit 140 in Example 1, determines the number of bits b used for encoding the right-channel difference signal y R in the stereo encoding unit 170 cand (1), y cand (2), ..., y R (T), and the number of bits b used for encoding the downmix signal x R (1), y R (2), ..., y R (T) in the monaural encoding unit 160 R and the downmix signal x M (1), x M(2),..., x M The number of bits b used for encoding (T) M Using the number of samples T per frame, the right-channel correction coefficient c is obtained by Equation (1-7-2) (step S140-12). The right-channel subtraction gain estimation unit 140 then uses the quantization value ^r of the normalized inner product value obtained in step S140-15 R and the right-channel correction coefficient c obtained in step S140-12 R to obtain the value obtained by multiplying them as the right-channel subtraction gain β (step S140-16). R

[0058] 〔〔〔Left-channel subtraction gain decoding unit 230〕〕〕 The left-channel subtraction gain decoding unit 230 stores the same candidates r for the normalized inner product value of the left channel as those stored in the left-channel subtraction gain estimation unit 120 of the corresponding encoding device 100 Lcand (a) and the code Cα corresponding to the candidate cand (a). A plurality of pairs (Group A, a = 1,..., A) are stored in advance. The left-channel subtraction gain decoding unit 230 performs the following steps S230-12 to S230-14 shown in FIG. 7

[0059] The left-channel subtraction gain decoding unit 230 uses the stored codes Cα cand (1),..., Cα cand (A) to obtain the candidate for the normalized inner product value of the left channel corresponding to the input left-channel subtraction gain code Cα as the decoded value ^r of the normalized inner product value of the left channel (step S230-12). Also, the left-channel subtraction gain decoding unit 230 uses the number of bits b used for decoding the left-channel decoded difference signal ^y L (1), ^y L (2),..., ^y L (T) in the stereo decoding unit 220 L and the number of bits b used for decoding the monaural decoded sound signal ^x L (1), ^x M (2),..., ^x M (T) in the monaural decoding unit 210 M to obtain the decoded value ^r of the normalized inner product value of the left channel (step S230-12). Also, the left-channel subtraction gain decoding unit 230 uses the number of bits b used for decoding the left-channel decoded difference signal ^y Mand the number of samples T per frame, and obtain the left-channel correction coefficient c by Equation (1-7) L (Step S230-13). Next, the left-channel subtraction gain decoder unit 230 multiplies the decoded value ^r of the normalized inner product value obtained in Step S230-12 L by the left-channel correction coefficient c obtained in Step S230-13 L to obtain the left-channel subtraction gain α as the multiplied value (Step S230-14).

[0060] Note that when the stereo code CS is the combination of the left-channel difference code CL and the right-channel difference code CR, the number of bits b L (1), ^y L (2),..., ^y L (T) used for decoding the left-channel decoded difference signal ^y L is the number of bits of the left-channel difference code CL. When the number of bits b L (1), ^y L (2),..., ^y L (T) used for decoding the left-channel decoded difference signal ^y L is not positively determined, then half of the number of bits b s of the stereo code CS input to the stereo decoder unit 220 (that is, b s / 2) may be used as the number of bits b L . The number of bits b M (1), ^x M (2),..., ^x M (T) used for decoding the monaural decoded sound signal ^x M is the number of bits of the monaural code CM. The left-channel correction coefficient c L is not the value obtained by Equation (1-7) itself, but a value greater than 0 and less than 1, and the number of bits b L (1), ^y L (2),..., ^y L (T) used for decoding the left-channel decoded difference signal ^y L and the monaural decoded sound signal ^x M (1), ^x M(2), ..., ^x M The number of bits b used for decoding (T) M is 0.5 when they are the same, and the number of bits b L is the number of bits b M is closer to 0 than 0.5 as it is more than the number of bits b L is the number of bits b M may be a value closer to 1 than 0.5 as it is less than the number of bits b

[0061] 〔〔〔Right channel subtraction gain decoding unit 250〕〕〕 In the right channel subtraction gain decoding unit 250, the candidate r of the normalized inner product value of the right channel, which is the same as that stored in the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 Rcand (b) and the code Cβ corresponding to the candidate cand (b) and a plurality of sets (B sets, b = 1,..., B) are stored in advance. The right channel subtraction gain decoding unit 250 performs the following steps S250-12 to S250-14 shown in FIG. 7

[0062] The right channel subtraction gain decoding unit 250 stores the stored code Cβ cand (1),..., Cβ cand (B), and the candidate of the normalized inner product value of the right channel corresponding to the input right channel subtraction gain code Cβ is used as the decoded value ^r of the normalized inner product value of the right channel R to obtain (step S250-12). Also, the right channel subtraction gain decoding unit 250, in the stereo decoding unit 220, the right channel decoded difference signal ^y R (1), ^y R (2),..., ^y R The number of bits b used for decoding (T) R and, in the monaural decoding unit 210, the monaural decoded sound signal ^x M (1), ^x M (2),..., ^x M (T) The number of bits b used for decoding M and the number of samples T per frame, and using them, the right channel correction coefficient c is calculated by Equation (1-7-2) Rto obtain (step S250-13). The right-channel subtraction gain decoder 250 then decodes the normalized inner product value ^r obtained in step S250-12 R and the right-channel correction coefficient c obtained in step S250-13 R to obtain the value obtained by multiplying them as the right-channel subtraction gain β (step S250-14).

[0063] Note that when the stereo code CS is a combination of the left-channel difference code CL and the right-channel difference code CR, the number of bits b used for decoding the right-channel decoded difference signal ^y R (1), ^y R (2),..., ^y R (T) in the stereo decoder 220 is the number of bits of the right-channel difference code CR. When the number of bits b used for decoding the right-channel decoded difference signal ^y R in the stereo decoder 220 is not positively determined, half of the number of bits b of the stereo code CS input to the stereo decoder 220 R (i.e., b R (2),..., ^y R (T) may be used as the number of bits b R . The number of bits b used for decoding the monaural decoded sound signal ^x s (1), ^x s (2),..., ^x R (T) in the monaural decoder 210 is the number of bits of the monaural code CM. The right-channel correction coefficient c M (1), ^x M (2),..., ^x M (T) is not the value obtained by the formula (1-7-2) itself, but a value greater than 0 and less than 1, and the number of bits b used for decoding the right-channel decoded difference signal ^y M (1), ^y R (2),..., ^y R (1), ^y R (2),..., ^y R (T) and the number of bits b used for decoding the monaural decoded sound signal ^x R (1), ^x M (2),..., ^x M (2),..., ^x M (T) in the monaural decoder 210 Mis 0.5 when they are the same, and the number of bits b R is the number of bits b M the more it is greater than the number of bits b, the closer it is to 0 than 0.5, and the number of bits b R is the number of bits b M the less it is than the number of bits b, it may be a value closer to 1 than 0.5.

[0064] Note that for the left channel and the right channel, the same normalized inner product value candidates and signs may be used. By setting A and B described above to the same value, the candidate r of the normalized inner product value of the left channel stored in the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 Lcand (a) and the sign Cα corresponding to the candidate cand (a), and the candidate r of the normalized inner product value of the right channel stored in the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 Rcand (b) and the sign Cβ corresponding to the candidate cand (b) may be made the same.

[0065] Note that the sign Cα is substantially a sign corresponding to the left channel subtraction gain α. For the purpose of aligning the language in the description of the encoding device 100 and the decoding device 200, etc., it is called the left channel subtraction gain sign, but considering that it represents a normalized inner product value, it may also be called the left channel inner product sign, etc. The same applies to the sign Cβ, and it may also be called the right channel inner product sign, etc.

[0066] [[Example 2]] Example 2 will describe an example in which a value considering the values of the inputs of past frames is used as the normalized inner product value. In Example 2, the optimality within a frame, that is, the minimization of the quantization error energy of the decoded audio signal of the left channel and the minimization of the quantization error energy of the decoded audio signal of the right channel are not strictly guaranteed. However, it reduces the sudden fluctuations between frames of the left channel subtraction gain α and the sudden fluctuations between frames of the right channel subtraction gain β, and reduces the noise generated in the decoded audio signal due to the fluctuations. That is, Example 2 takes into account the auditory quality of the decoded audio signal in addition to reducing the energy of the quantization error of the decoded audio signal.

[0067] In Example 2, the encoding side, that is, the left channel subtraction gain estimation unit 120 and the right channel subtraction gain estimation unit 140 are different from those in Example 1, but the decoding side, that is, the left channel subtraction gain decoding unit 230 and the right channel subtraction gain decoding unit 250 are the same as those in Example 1. Hereinafter, the differences between Example 2 and Example 1 will be mainly described.

[0068] 〔〔〔Left channel subtraction gain estimation unit 120〕〕〕 As shown in FIG. 8, the left channel subtraction gain estimation unit 120 performs the following steps S120-111 to S120-113 and steps S120-12 to S120-14 described in Example 1.

[0069] The left channel subtraction gain estimation unit 120 first uses the input left channel input audio signals x L (1), x L (2),..., x L (T), the input downmix signals x M (1), x M (2),..., x M (T), and the inner product value E L (-1) used in the previous frame to obtain the inner product value E L (0) to be used in the current frame by the following formula (1-8) (step S120-111).

Equation

[0070] The left-channel subtraction gain estimation unit 120 also uses the input downmix signals x M (1), x M (2),..., x M (T), and the energy E M (-1) of the downmix signal used in the previous frame, to obtain the energy E M (0) of the downmix signal used in the current frame by the following formula (1-9) (step S120-112).

Equation

[0071] Next, the left-channel subtraction gain estimation unit 120 uses the inner product value E L (0) used in the current frame obtained in step S120-111 and the energy E M (0) of the downmix signal used in the current frame obtained in step S120-112 to obtain the normalized inner product value r L by the following formula (1-10) (step S120-113).

Equation

[0072] The left-channel subtraction gain estimation unit 120 also performs step S120-12, and then uses the normalized inner product value r obtained in step S120-11 L to perform step S120-13 and further perform step S120-14, using the normalized inner product value r obtained in step S120-113 described above instead of r L

[0073] Note that the above ε L and ε M The closer to 1 the normalized inner product value r L , the more likely it is that the input sound signal and the downmix signal of the left channel of the past frame are included. The variation between frames of the left-channel subtraction gain α obtained by the normalized inner product value r L or the normalized inner product value r L becomes smaller.

[0074] 〔〔〔Right-channel subtraction gain estimation unit 140〕〕〕 As shown in FIG. 8, the right-channel subtraction gain estimation unit 140 performs the following steps S140-111 to S140-113 and steps S140-12 to S140-14 described in Example 1.

[0075] The right-channel subtraction gain estimation unit 140 first uses the input sound signal x of the input right channel R (1), x R (2),..., x R (T), the input downmix signal x M (1), x M (2),..., x M (T), and the inner product value E used in the previous frame R (-1) to obtain the inner product value E R (0) used in the current frame according to the following formula (1-8-2) (step S140-111).

Equation

[0076] The right-channel subtraction gain estimation unit 140 also uses the input downmix signals x M (1), x M (2),..., x M (T), the energy E M (-1) of the downmix signal used in the previous frame, and obtains the energy E M (0) of the downmix signal used in the current frame by the formula (1-9) (step S140-112). The right-channel subtraction gain estimation unit 140 stores the obtained energy E M (0) of the downmix signal in the right-channel subtraction gain estimation unit 140 for use in the next frame as the "energy E M (-1) of the downmix signal used in the previous frame". Note that the left-channel subtraction gain estimation unit 120 also obtains the energy E M (0) of the downmix signal used in the current frame by the formula (1-9). Therefore, only one of step S120-112 performed by the left-channel subtraction gain estimation unit 120 and step S140-112 performed by the right-channel subtraction gain estimation unit 140 may be performed.

[0077] Next, the right-channel subtraction gain estimation unit 140 uses the inner product value E R (0) used in the current frame obtained in step S140-111 and the energy E M (0) of the downmix signal used in the current frame obtained in step S140-112 to obtain the normalized inner product value r R by the following formula (1-10-2) (step S140-113).

Number

[0078] The right channel subtraction gain estimation unit 140 also performs step S140-12, and then uses the normalized inner product value r obtained in step S140-11 R to perform step S140-13 in place of the normalized inner product value r obtained in step S140-113 described above R and further performs step S140-14

[0079] Note that the above ε R and ε M The closer to 1, the more likely the normalized inner product value r R will include the influence of the input sound signal and the downmix signal of the right channel of the past frame, and the variation between frames of the right channel subtraction gain β obtained by the normalized inner product value r R or the normalized inner product value r R will be reduced

[0080] 〔〔Modification Example of Example 2〕〕 Regarding Example 2 as well, the same modifications as those of the modification example of Example 1 with respect to Example 1 can be made. This form will be described as a modification example of Example 2. In the modification example of Example 2, the encoding side, that is, the left channel subtraction gain estimation unit 120 and the right channel subtraction gain estimation unit 140 are different from those of the modification example of Example 1, but the decoding side, that is, the left channel subtraction gain decoding unit 230 and the right channel subtraction gain decoding unit 250 are the same as those of the modification example of Example 1. Since the differences between the modification example of Example 2 and the modification example of Example 1 are the same as those between Example 2 and Example 1, hereinafter, the modification example of Example 2 will be described with appropriate reference to the modification example of Example 1 and Example 2

[0081] 〔〔〔Left Channel Subtraction Gain Estimation Unit 120〕〕〕 The left channel subtraction gain estimation unit 120 has, in the same manner as the left channel subtraction gain estimation unit 120 of the modification example of Example 1, candidates r for the normalized inner product values of the left channel Lcand (a) and the code Cα corresponding to the candidate cand(A set of pairs with (a) is stored in advance in multiple sets (set A, a = 1, ..., A). As shown in FIG. 9, the left-channel subtraction gain estimation unit 120 performs the same steps S120-111 to S120-113 as in Example 2, and the same steps S120-12, S120-15, and S120-16 as in the modification of Example 1. Specifically, it is as follows.)

[0082] (The left-channel subtraction gain estimation unit 120 first uses the input left-channel input sound signals x L (1), x L (2),..., x L (T), the input downmix signals x M (1), x M (2),..., x M (T), and the inner product value E L (-1) used in the previous frame to obtain the inner product value E L (0) for the current frame using Equation (1-8) (step S120-111). The left-channel subtraction gain estimation unit 120 also uses the input downmix signals x M (1), x M (2),..., x M (T) and the energy E M (-1) of the downmix signal used in the previous frame to obtain the energy E M (0) of the downmix signal for the current frame using Equation (1-9) (step S120-112). Next, the left-channel subtraction gain estimation unit 120 uses the inner product value E L (0) for the current frame obtained in step S120-111 and the energy E M (0) of the downmix signal for the current frame obtained in step S120-112 to obtain the normalized inner product value r L using Equation (1-10) (step S120-113). Next, the left-channel subtraction gain estimation unit 120 uses the candidate r Lcand (1),..., r Lcand (A) of the normalized inner product values of the left channel stored and the normalized inner product value r LThe candidate closest to (the quantized value of the normalized inner product value r L )^r L is obtained, and the stored code Cα cand (1),..., Cα cand (A) among the closest candidates ^r L The corresponding code is obtained as the left channel subtraction gain code Cα (step S120-15). Also, the left channel subtraction gain estimation unit 120 uses the number of bits b L used for encoding the left channel difference signals y L (1), y L (2),..., y L in the stereo encoding unit 170, and the number of bits b M used for encoding the downmix signals x M (1), x M (2),..., x M (T) in the monaural encoding unit 160, and the number of samples T per frame, and obtains the left channel correction coefficient c L by Equation (1-7) (step S120-12). The left channel subtraction gain estimation unit 120 then obtains the left channel subtraction gain α as the value obtained by multiplying the quantized value ^r L of the normalized inner product value obtained in step S120-15 and the left channel correction coefficient c L obtained in step S120-12 (step S120-16).

[0083] [[[Right channel subtraction gain estimation unit 140]]] Similar to the right channel subtraction gain estimation unit 140 in the modification of Example 1, the right channel subtraction gain estimation unit 140 stores in advance a plurality of pairs (B pairs, b = 1,..., B) of the candidate r Rcand (b) of the normalized inner product value of the right channel and the code Cβ cand (b) corresponding to the candidate. As shown in FIG. 9, the right channel subtraction gain estimation unit 140 performs the same steps S140-111 to S140-113 as in Example 2, and the same steps S140-12, S140-15, and S140-16 as in the modification of Example 1. Specifically, it is as follows.

[0084] The right-channel subtraction gain estimation unit 140 first uses the input right-channel input audio signal x R (1), x R (2),..., x R (T), the input downmix signal x M (1), x M (2),..., x M (T), and the inner product value E R (-1) used in the previous frame to obtain the inner product value E R (0) to be used in the current frame according to Equation (1-8-2) (step S140-111). The right-channel subtraction gain estimation unit 140 also uses the input downmix signal x M (1), x M (2),..., x M (T) and the energy E M (-1) of the downmix signal used in the previous frame to obtain the energy E M (0) of the downmix signal to be used in the current frame according to Equation (1-9) (step S140-112). Next, the right-channel subtraction gain estimation unit 140 uses the inner product value E R (0) obtained in step S140-111 and the energy E M (0) of the downmix signal obtained in step S140-112 to obtain the normalized inner product value r R according to Equation (1-10-2) (step S140-113). Next, the right-channel subtraction gain estimation unit 140 obtains the candidate r Rcand (1),..., r Rcand (B) of the normalized inner product value of the right channel stored in the memory that is closest to the normalized inner product value r R obtained in step S140-113 (the quantization value of the normalized inner product value r R )^r R and obtains the code Cβ cand (1),..., Cβ cand (B) in the memory corresponding to the closest candidate ^r RObtain the corresponding code as the right-channel subtraction gain code Cβ (step S140-15). Also, the right-channel subtraction gain estimator 140 determines the number of bits b R (1), y R (2), ..., y R (T) used for encoding the right-channel difference signals y R and the number of bits b M (1), x M (2), ..., x M (T) used for encoding the downmix signals x M and the number of samples T per frame, and obtains the right-channel correction coefficient c R by Equation (1-7-2) (step S140-12). Next, the right-channel subtraction gain estimator 140 multiplies the quantized value ^r R of the normalized inner product value obtained in step S140-15 by the right-channel correction coefficient c R obtained in step S140-12 to obtain the right-channel subtraction gain β (step S140-16).

[0085] [[Example 3]] For example, when the sounds such as voices and music included in the input sound signal of the left channel are different from the sounds such as voices and music included in the input sound signal of the right channel, the downmix signal may include components of both the input sound signal of the left channel and the input sound signal of the right channel. Therefore, the larger the value used for the left-channel subtraction gain α, the more likely it is that the sound derived from the input sound signal of the right channel that should not be originally audible is included in the left-channel decoded sound signal, and the larger the value used for the right-channel subtraction gain β, the more likely it is that the sound derived from the input sound signal of the left channel that should not be originally audible is included in the right-channel decoded sound signal. Thus, although the minimization of the energy of the quantization error of the decoded sound signal is not strictly guaranteed, considering the auditory quality, the left-channel subtraction gain α and the right-channel subtraction gain β may be set to values smaller than the values obtained in Example 1. Similarly, the left-channel subtraction gain α and the right-channel subtraction gain β may be set to values smaller than the values obtained in Example 2.

[0086] Specifically, for the left channel, in Examples 1 and 2, the multiplication value c of the normalized inner product value r L and the left channel correction coefficient c L was used as the left channel subtraction gain α, which was the quantization value of c L ×r L However, in Example 3, the quantization value of the multiplication value λ L of the normalized inner product value r L and the left channel correction coefficient c L with a predetermined value λ greater than 0 and less than 1 L ×c L ×r L is used as the left channel subtraction gain α. Therefore, similar to Examples 1 and 2, the multiplication value c L ×r L is used as the target for encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, and the left channel subtraction gain symbol Cα represents the quantization value of the multiplication value c L ×r L so that the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 multiply the quantization value of the multiplication value c L ×r L and λ L to obtain the left channel subtraction gain α. Alternatively, the multiplication value λ L of the normalized inner product value r L and the left channel correction coefficient c L with a predetermined value λ L ×c L ×r L is used as the target for encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230, and the left channel subtraction gain symbol Cα represents the quantization value of the multiplication value λ L ×c L ×r L This may also be the case.

[0087] Similarly, for the right channel, in Examples 1 and 2, the multiplication value c of the normalized inner product value r R and the right channel correction coefficient c R was used as the right channel subtraction gain α, which was the quantization value of c R ×r RThe quantization value was used as the right-channel subtraction gain β. In Example 3, the normalized inner product value r R and the right-channel correction coefficient c R and a predetermined value λ that is greater than 0 and less than 1 R the multiplication value λ R ×c R ×r R The quantization value of is used as the right-channel subtraction gain β. Therefore, similar to Example 1 and Example 2, the multiplication value c R ×r R is used as the object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250, and the right-channel subtraction gain symbol Cβ is the multiplication value c R ×r R The quantization value of represents, and the right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 multiply the quantization value of the multiplication value c R ×r R by λ R to obtain the right-channel subtraction gain β. Alternatively, the normalized inner product value r R and the left-channel correction coefficient c R and a predetermined value λ R the multiplication value λ R ×c R ×r R is used as the object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250, and the right-channel subtraction gain symbol Cβ is the multiplication value λ R ×c R ×r R The quantization value of represents. Note that λ R should be the same value as λ L .

[0088] 〔〔Modified Example of Example 3〕〕 As described above, the correction coefficient c L can calculate the same value in both the encoding device 100 and the decoding device 200. Therefore, similar to the modified example of Example 1 and the modified example of Example 2, the normalized inner product value r L is used as the object of encoding in the left-channel subtraction gain estimation unit 120 and decoding in the left-channel subtraction gain decoding unit 230, and the left-channel subtraction gain symbol Cα is the normalized inner product value r Lsuch that the quantization value of L and the left channel correction coefficient c L and a predetermined value λ greater than 0 and less than 1 L are multiplied to obtain the left channel subtraction gain α. Alternatively, the normalized inner product value r L and a predetermined value λ greater than 0 and less than 1 L of the multiplication value λ L ×r L is used as the object of encoding in the left channel subtraction gain estimator 120 and decoding in the left channel subtraction gain decoder 230, and the left channel subtraction gain symbol Cα is such that the quantization value of the multiplication value λ L ×r L and the left channel correction coefficient c L ×r L are multiplied to obtain the left channel subtraction gain α. L

[0089] The same applies to the right channel, and the correction coefficient c R can calculate the same value in both the encoding device 100 and the decoding device 200. Therefore, similar to the modified example of Example 1 and the modified example of Example 2, the normalized inner product value r R is used as the object of encoding in the right channel subtraction gain estimator 140 and decoding in the right channel subtraction gain decoder 250, and the right channel subtraction gain symbol Cβ is such that the quantization value of the normalized inner product value r R and the right channel correction coefficient c R and a predetermined value λ greater than 0 and less than 1 R are multiplied to obtain the right channel subtraction gain β. Alternatively, the normalized inner product value r R and a predetermined value λ greater than 0 and less than 1 R of the multiplication value λ R ×r R ×r R ​The quantization value of the multiplication value λ R ×r R is used as the target for encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250, and the right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 multiply the quantization value of the multiplication value λ R ×r R by the right-channel correction coefficient c R to obtain the right-channel subtraction gain β.

[0090] 〔〔Example 4〕〕 The problem of auditory quality described at the beginning of Example 3 occurs when the correlation between the input sound signal of the left channel and the input sound signal of the right channel is small, and this problem hardly occurs when the correlation between the input sound signal of the left channel and the input sound signal of the right channel is large. Therefore, in Example 4, instead of the predetermined value in Example 3, the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, is used. As the correlation between the input sound signal of the left channel and the input sound signal of the right channel increases, priority is given to reducing the energy of the quantization error of the decoded sound signal. As the correlation between the input sound signal of the left channel and the input sound signal of the right channel decreases, priority is given to suppressing the deterioration of auditory quality.

[0091] In Example 4, the encoding side is different from Examples 1 and 2, but the decoding side, that is, the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250, are the same as in Examples 1 and 2. Hereinafter, the differences between Example 4 and Examples 1 and 2 will be described.

[0092] 〔〔〔Left-Right Relationship Information Estimation Unit 180〕〕〕 The encoding apparatus 100 of Example 4 also includes a left-right relationship information estimation unit 180 as shown by the dashed line in FIG. 1. The left-right relationship information estimation unit 180 receives the input sound signal of the left channel input to the encoding apparatus 100 and the input sound signal of the right channel input to the encoding apparatus 100. The left-right relationship information estimation unit 180 obtains and outputs the left-right correlation coefficient γ from the input left-channel input sound signal and right-channel input sound signal (step S180).

[0093] The left-right correlation coefficient γ is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel. The sample sequence x of the input sound signal of the left channel L (1), x L (2), ..., x L (T) and the sample sequence x of the input sound signal of the right channel R (1), x R (2), ..., x R (T) of the correlation coefficient γ 0 It may also be the correlation coefficient considering the time difference, for example, the correlation coefficient γ between the sample sequence of the input sound signal of the left channel and the sample sequence of the input sound signal of the right channel that is shifted by τ samples after the sample sequence. τ It may be.

[0094] This τ is the information corresponding to the difference (so-called arrival time difference) between the arrival time from the sound source mainly emitting sound in the space to the microphone for the left channel and the arrival time from the sound source to the microphone for the right channel, assuming that the sound signal obtained by AD-converting the sound picked up by the microphone for the left channel arranged in a certain space is the input sound signal of the left channel, and the sound signal obtained by AD-converting the sound picked up by the microphone for the right channel arranged in the space is the input sound signal of the right channel. Hereinafter, it is called the left-right time difference. The left-right time difference τ may be obtained by any well-known method, and may be obtained by the method described in the left-right relationship information estimation unit 181 of the second reference embodiment. That is, the above-mentioned correlation coefficient γ τ is the information corresponding to the correlation coefficient between the sound signal that reaches and is picked up by the microphone for the left channel from the sound source and the sound signal that reaches and is picked up by the microphone for the right channel from the sound source.

[0095] 〔〔〔Left channel subtraction gain estimation unit 120〕〕〕 Instead of step S120-13, the left channel subtraction gain estimation unit 120 uses the normalized inner product value r obtained in step S120-11 or step S120-113 L and the left channel correction coefficient c obtained in step S120-12 LAnd, a value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 is obtained (step S120-13”). Next, instead of step S120-14, the left-channel subtraction gain estimator 120 uses the candidate α of the stored left-channel subtraction gain cand (1), ..., α cand Among (A), the multiplication value γ×c obtained in step S120-13” L ×r L The candidate closest to (the quantization value of the multiplication value γ×c L ×r L ) is obtained as the left-channel subtraction gain α, and the stored code Cα cand (1), ..., Cα cand Among (A), the code corresponding to the left-channel subtraction gain α is obtained as the left-channel subtraction gain code Cα (step S120-14”).

[0096] 〔〔〔Right-channel subtraction gain estimator 140〕〕〕 Instead of step S140-13, the right-channel subtraction gain estimator 140 uses the normalized inner product value r obtained in step S140-11 or step S140-113 R And the right-channel correction coefficient c obtained in step S140-12 R And the left-right correlation coefficient γ obtained in step S180, and obtains a multiplied value (step S140-13”). Next, instead of step S140-14, the right-channel subtraction gain estimator 140 uses the candidate β of the stored right-channel subtraction gain cand (1), ..., β cand Among (B), the multiplication value γ×c obtained in step S140-13” R ×r R The candidate closest to (the quantization value of the multiplication value γ×c R ×r R ) is obtained as the right-channel subtraction gain β, and the stored code Cβ cand (1), ..., Cβ cand Among (B), the code corresponding to the right-channel subtraction gain β is obtained as the right-channel subtraction gain code Cβ (step S140-14”).

[0097] 〔〔Modified example of Example 4〕〕 As described above, correction coefficient c L can be calculated with the same value in both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r L and the multiplication value γ×r of the left - right correlation coefficient γ L are used as the objects of encoding in the left - channel subtraction gain estimation unit 120 and decoding in the left - channel subtraction gain decoding unit 230, such that the left - channel subtraction gain symbol Cα represents the quantization value of the multiplication value γ×r L and the left - channel subtraction gain estimation unit 120 and the left - channel subtraction gain decoding unit 230 multiply the quantization value of the multiplication value γ×r L by the left - channel correction coefficient c L to obtain the left - channel subtraction gain α.

[0098] The same applies to the right channel. The correction coefficient c R can be calculated with the same value in both the encoding device 100 and the decoding device 200. Therefore, the normalized inner product value r R and the multiplication value γ×r of the left - right correlation coefficient γ R are used as the objects of encoding in the right - channel subtraction gain estimation unit 140 and decoding in the right - channel subtraction gain decoding unit 250, such that the right - channel subtraction gain symbol Cβ represents the quantization value of the multiplication value γ×r R and the right - channel subtraction gain estimation unit 140 and the right - channel subtraction gain decoding unit 250 multiply the quantization value of the multiplication value γ×r R by the right - channel correction coefficient c R to obtain the right - channel subtraction gain β.

[0099] <Second Reference Embodiment> The encoding device and the decoding device of the second reference embodiment will be described.

[0100] ≪Encoding Device 101≫ As shown in FIG. 10, the encoding device 101 of the second reference embodiment includes a downmixing unit 110, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimation unit 181, and a time shift unit 191. The encoding device 101 of the second reference embodiment is different from the encoding device 100 of the first reference embodiment in that it includes a left-right relationship information estimation unit 181 and a time shift unit 191, the left-channel subtraction gain estimation unit 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimation unit 140, and the right-channel signal subtraction unit 150 use the signal output from the time shift unit 191 instead of the signal output from the downmixing unit 110, and in addition to the above-described respective codes, it also outputs a left-right time difference code Cτ, which will be described later. Other configurations and operations of the encoding device 101 of the second reference embodiment are the same as those of the encoding device 100 of the first reference embodiment. The encoding device 101 of the second reference embodiment performs the processing from step S110 to step S191 illustrated in FIG. 11 for each frame. Hereinafter, differences between the encoding device 101 of the second reference embodiment and the encoding device 100 of the first reference embodiment will be described.

[0101] [Left-Right Relationship Information Estimation Unit 181] The left-right relationship information estimation unit 181 receives the input audio signal of the left channel input to the encoding device 101 and the input audio signal of the right channel input to the encoding device 101. The left-right relationship information estimation unit 181 obtains and outputs a left-right time difference τ and a left-right time difference code Cτ, which is a code representing the left-right time difference τ, from the input audio signal of the left channel and the input audio signal of the right channel (step S181).

[0102] The left-right time difference τ is information corresponding to the difference (so-called arrival time difference) between the arrival time from a sound source mainly emitting sound in the space to the microphone for the left channel and the arrival time from the sound source to the microphone for the right channel, assuming that the sound signal obtained by AD-converting the sound picked up by the microphone for the left channel arranged in a certain space is the input sound signal for the left channel, and the sound signal obtained by AD-converting the sound picked up by the microphone for the right channel arranged in the space is the input sound signal for the right channel. Note that in order to include not only the arrival time difference but also information on which microphone the sound reaches earlier in the left-right time difference τ, the left-right time difference τ can take both positive and negative values with respect to one of the input sound signals as a reference. That is, the left-right time difference τ is information indicating how much earlier the same sound signal is included in either the input sound signal for the left channel or the input sound signal for the right channel. Hereinafter, when the same sound signal is included earlier in the input sound signal for the left channel than in the input sound signal for the right channel, it is also said that the left channel is leading, and when the same sound signal is included earlier in the input sound signal for the right channel than in the input sound signal for the left channel, it is also said that the right channel is leading.

[0103] The left-right time difference τ may be obtained by any well-known method. For example, the left-right relationship information estimation unit 181, for each candidate sample number τ max from τ min to τ max (for example, τ min is a positive number, τ cand is a negative number), calculates a value (hereinafter referred to as a correlation value) γ cand representing the magnitude of the correlation between the sample sequence of the input sound signal for the left channel and the sample sequence of the input sound signal for the right channel shifted by the candidate sample number τ cand positions later than the sample sequence, and determines the candidate sample number τ cand at which the correlation value γ candis obtained as the left - right time difference τ. That is, in this example, when the left channel leads, the left - right time difference τ is a positive value, and when the right channel leads, the left - right time difference τ is a negative value. The absolute value of the left - right time difference τ represents the value (the number of leading samples) indicating how much the leading channel leads the other channel. For example, when calculating the correlation value γ cand using only the samples within the frame, when τ cand is a positive value, the partial sample sequence x R (1 + τ cand ), x R (2 + τ cand ),..., x R (T) of the input sound signal of the right channel and the partial sample sequence x cand of the input sound signal of the left channel, which is shifted τ L (1), x L (2),..., x L (T - τ cand ) samples before the partial sample sequence by the number of candidate samples τ cand are used to calculate the absolute value of the correlation coefficient as the correlation value γ cand . When τ L is a negative value, the partial sample sequence x cand (1 - τ L ), x cand (2 - τ L ),..., x cand of the input sound signal of the left channel and the partial sample sequence x R (1), x R (2),..., x R (T+τ cand ) of the input sound signal of the right channel, which is shifted - τ cand samples before the partial sample sequence by the number of candidate samples, are used to calculate the absolute value of the correlation coefficient as the correlation value γ cand . Of course, one or more samples of the input sound signal in the past consecutive to the sample sequence of the input sound signal in the current frame may also be used to calculate the correlation value γ

[0104] Also, for example, instead of the absolute value of the correlation coefficient, the correlation value γ may be calculated using the phase information of the signal as follows. cand In this example, the left - right relationship information estimation unit 181 first Fourier - transforms each of the input sound signals x L (1), x L (2),..., x L (T) of the left channel and the input sound signals x R (1), x R (2),..., x R (T) of the right channel according to the following equations (3 - 1) and (3 - 2) to obtain the frequency spectra X L (k) and X R (k) at each frequency k from 0 to T - 1.

Number

Number

Number

Number

[0105] ​Further, the left - right relationship information estimation unit 181 may encode the left - right time difference τ using a predetermined encoding method to obtain a left - right time difference code Cτ, which is a code that can uniquely identify the left - right time difference τ. As the predetermined encoding method, a well - known encoding method such as scalar quantization may be used. Note that each of the predetermined candidate sample numbers may be each integer value from τ max to τ min or may include fractional values or decimal values between τ max and τ min or may not include any integer value between τ max and τ min . Also, τ max = - τ min may or may not be true. Further, in the case of targeting a special input sound signal in which a certain channel always precedes, τ max and τ min may both be set as positive numbers, or τ max and τ min may both be set as negative numbers.

[0106] Note that when the encoding device 101 performs estimation of the subtraction gain based on the principle of minimizing the quantization error of Example 4 or a modified example of Example 4 described in the first reference form, the left - right relationship information estimation unit 181 further calculates the correlation value between the sample sequence of the input sound signal of the left channel and the sample sequence of the input sound signal of the right channel shifted by the left - right time difference τ from the sample sequence, that is, the maximum value of the correlation values γ max calculated for each candidate sample number τ min from τ cand to τ cand , and outputs it as the left - right correlation coefficient γ (step S180).

[0107] [Time shift unit 191] The time shift unit 191 receives the down - mix signal x M (1), x M (2),..., x M(T) and the left-right time difference τ output by the left-right relationship information estimation unit 181 are input. When the left-right time difference τ is a positive value (that is, when the left-right time difference τ indicates that the left channel is leading), the time shift unit 191 outputs the downmix signal x M (1), x M (2), ..., x M (T) as it is to the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130 (that is, it is determined to be used by the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130), and delays the downmix signal by |τ| samples (the number of samples corresponding to the absolute value of the left-right time difference τ, the number of samples corresponding to the magnitude represented by the left-right time difference τ), and the delayed downmix signal x M (1 - |τ|), x M (2 - |τ|), ..., x M (T - |τ|) is output to the right-channel subtraction gain estimation unit 140 and the right-channel signal subtraction unit 150 (that is, it is determined to be used by the right-channel subtraction gain estimation unit 140 and the right-channel signal subtraction unit 150). When the left-right time difference τ is a negative value (that is, when the left-right time difference τ indicates that the right channel is leading), the downmix signal is delayed by |τ| samples, and the delayed downmix signal x M' (1), x M' (2), ..., x M' (T) is output to the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130 (that is, it is determined to be used by the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130). When the left-right time difference τ is a negative value (that is, when the left-right time difference τ indicates that the right channel is leading), the downmix signal is delayed by |τ| samples, and the delayed downmix signal x M (1 - |τ|), x M (2 - |τ|), ..., x M (T - |τ|) is output to the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130 (that is, it is determined to be used by the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130), and the downmix signal x M' (1), x M' (2), ..., x M' (T) is output to the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130 (that is, it is determined to be used by the left-channel subtraction gain estimation unit 120 and the left-channel signal subtraction unit 130), and the downmix signal x M (1), x M (2), ..., x M(T) is directly output to the right-channel subtraction gain estimator 140 and the right-channel signal subtraction unit 150 (that is, it is determined to be used by the right-channel subtraction gain estimator 140 and the right-channel signal subtraction unit 150). When the left-right time difference τ is 0 (that is, when it means that the left-right time difference τ does not lead any channel), the downmix signal x M (1), x M (2), ..., x M (T) is directly output to the left-channel subtraction gain estimator 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimator 140, and the right-channel signal subtraction unit 150 (that is, it is determined to be used by the left-channel subtraction gain estimator 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimator 140, and the right-channel signal subtraction unit 150) (step S191). That is, for the channel with the shorter arrival time among the left and right channels, the input downmix signal is directly output to the subtraction gain estimator and the signal subtraction unit of the corresponding channel. For the channel with the longer arrival time among the left and right channels, the signal obtained by delaying the input downmix signal by the absolute value |τ| of the left-right time difference τ is output to the subtraction gain estimator and the signal subtraction unit of the corresponding channel. Note that since the time shift unit 191 uses the downmix signals of past frames to obtain the delayed downmix signal, the storage unit (not shown) in the time shift unit 191 stores the downmix signals input in past frames for a predetermined number of frames in advance. Also, when the left-channel subtraction gain estimator 120 and the right-channel subtraction gain estimator 140 obtain the left-channel subtraction gain α and the right-channel subtraction gain β by a well-known method as exemplified in Patent Document 1 instead of the method based on the principle of minimizing quantization error, a means for obtaining a local decoded signal corresponding to the monoral code CM is provided in the subsequent stage of the monoral encoding unit 160 of the encoding device 101 or within the monoral encoding unit 160. In the time shift unit 191, instead of the downmix signal x M (1), x M (2), ..., x M (T), the quantized downmix signal ^x that is the local decoded signal of monoral encoding M(1), ^x M (2), ..., ^x M (T) may be used to perform the above-described processing. In this case, the time shift unit 191 outputs the downmix signal x M (1), x M (2), ..., x M (T) and outputs the quantized downmix signal ^x M (1), ^x M (2), ..., ^x M (T), and outputs the delayed downmix signal x M' (1), x M' (2), ..., x M' (T) and outputs the delayed quantized downmix signal ^x M' (1), ^x M' (2), ..., ^x M' (T).

[0108] [Left channel subtraction gain estimation unit 120, left channel signal subtraction unit 130, right channel subtraction gain estimation unit 140, right channel signal subtraction unit 150] The left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 perform the same operations as described in the first reference embodiment on the downmix signal x M (1), x M (2), ..., x M (T) input from the time shift unit 191 instead of the downmix signal x M (1), x M (2), ..., x M (T) or the delayed downmix signal x M' (1), x M' (2), ..., x M' (T). That is, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 perform operations on the downmix signal x M (1), x M(2),..., x M (T) or the delayed downmix signal x M' (1), x M' (2),..., x M' (T) is used to perform the same operations as described in the first exemplary embodiment. Note that the time shift unit 191 uses the downmix signal x M (1), x M (2),..., x M (T) is replaced by the quantized downmix signal ^x M (1), ^x M (2),..., ^x M (T) to output the delayed downmix signal x M' (1), x M' (2),..., x M' (T) is replaced by the delayed quantized downmix signal ^x M' (1), ^x M' (2),..., ^x M' When (T) is output, the left channel subtraction gain estimation unit 120, the left channel signal subtraction unit 130, the right channel subtraction gain estimation unit 140, and the right channel signal subtraction unit 150 use the quantized downmix signal ^x M (1), ^x M (2),..., ^x M (T) or the delayed quantized downmix signal ^x M' (1), ^x M' (2),..., ^x M' (T) to perform the above-described processing.

[0109] ≪Decoder 201≫ As shown in Fig. 12, the decoding device 201 of the second reference embodiment includes a monaural decoding unit 210, a stereo decoding unit 220, a left-channel subtraction gain decoding unit 230, a left-channel signal addition unit 240, a right-channel subtraction gain decoding unit 250, a right-channel signal addition unit 260, a left-right time difference decoding unit 271, and a time shift unit 281. The decoding device 201 of the second reference embodiment differs from the decoding device 200 of the first reference embodiment in that, in addition to the above-described respective codes, a left-right time difference code Cτ described later is also input, it includes the left-right time difference decoding unit 271 and the time shift unit 281, and the left-channel signal addition unit 240 and the right-channel signal addition unit 260 use the signal output from the time shift unit 281 instead of the signal output from the monaural decoding unit 210. Other configurations and operations of the decoding device 201 of the second reference embodiment are the same as those of the decoding device 200 of the first reference embodiment. The decoding device 201 of the second reference embodiment performs the processing from step S210 to step S281 illustrated in Fig. 13 for each frame. Hereinafter, differences between the decoding device 201 of the second reference embodiment and the decoding device 200 of the first reference embodiment will be described.

[0110] [Left-Right Time Difference Decoding Unit 271] The left-right time difference decoding unit 271 receives the left-right time difference code Cτ input to the decoding device 201. The left-right time difference decoding unit 271 decodes the left-right time difference code Cτ by a predetermined decoding method to obtain and output a left-right time difference τ (step S271). As the predetermined decoding method, a decoding method corresponding to the encoding method used by the left-right relationship information estimation unit 181 of the corresponding encoding device 101 is used. The left-right time difference τ obtained by the left-right time difference decoding unit 271 is the same value as the left-right time difference τ obtained by the left-right relationship information estimation unit 181 of the corresponding encoding device 101, and is any value within the range from τ max to τ min and up to.

[0111] [Time Shift Unit 281] The time shift unit 281 receives the monaural decoded audio signal ^x output from the monaural decoding unit 210 M (1), ^x M (2),..., ^x M(T) and the left - right time difference τ output by the left - right time - difference decoding unit 271 are input. When the left - right time difference τ is a positive value (that is, when the left - right time difference τ indicates that the left channel is leading), the time - shift unit 281 outputs the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M (T) as it is to the left - channel signal addition unit 240 (that is, it is determined to be used by the left - channel signal addition unit 240), and outputs the signal ^x M (1 - |τ|), ^x M (2 - |τ|), ..., ^x M (T - |τ|), which is the delayed monaural decoded sound signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to the right - channel signal addition unit 260 (that is, it is determined to be used by the right - channel signal addition unit 260). When the left - right time difference τ is a negative value (that is, when the left - right time difference τ indicates that the right channel is leading), the time - shift unit 281 outputs the signal ^x M (1 - |τ|), ^x M (2 - |τ|), ..., ^x M (T - |τ|), which is the delayed monaural decoded sound signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to the left - channel signal addition unit 240 (that is, it is determined to be used by the left - channel signal addition unit 240), and outputs the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M (T) as it is to the right - channel signal addition unit 260 (that is, it is determined to be used by the right - channel signal addition unit 260). When the left - right time difference τ is 0 (that is, when the left - right time difference τ indicates that neither channel is leading), the time - shift unit 281 outputs the monaural decoded sound signal ^x M (1), ^x M (2), ..., ^x M(T) is directly output to the left-channel signal adder 240 and the right-channel signal adder 260 (that is, it is determined to be used by the left-channel signal adder 240 and the right-channel signal adder 260) (step S281). Note that since the time shift unit 281 uses the monaural decoded audio signal of the past frame to obtain the delayed monaural decoded audio signal, the storage unit (not shown) in the time shift unit 281 stores the monaural decoded audio signal input in the past frame for a predetermined number of frames in advance.

[0112] [Left-channel signal adder 240, Right-channel signal adder 260] The left-channel signal adder 240 and the right-channel signal adder 260 perform the same operations as those described in the first reference embodiment, using the monaural decoded audio signal ^x M (1), ^x M (2), ..., ^x M (T) Instead, the monaural decoded audio signal ^x input from the time shift unit 281 M (1), ^x M (2), ..., ^x M (T) or the delayed monaural decoded audio signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to perform (steps S240, S260). That is, the left-channel signal adder 240 and the right-channel signal adder 260 use the monaural decoded audio signal ^x M (1), ^x M (2), ..., ^x M (T) or the delayed monaural decoded audio signal ^x M' (1), ^x M' (2), ..., ^x M' (T) to perform the same operations as those described in the first reference embodiment.

[0113] <First Embodiment> The first embodiment is a modification of the encoding device 101 of the second reference form, which generates a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel. Hereinafter, the encoding device of the first embodiment will be described. Note that since the code obtained by the encoding device of the first embodiment can be decoded by the decoding device 201 of the second reference form, the description of the decoding device is omitted.

[0114] <<Encoding Device 102>> As shown in FIG. 10, the encoding device 102 of the first embodiment includes a downmixing unit 112, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimation unit 182, and a time shift unit 191. The encoding device 102 of the first embodiment is different from the encoding device 101 of the second reference form in that it includes a left-right relationship information estimation unit 182 instead of the left-right relationship information estimation unit 181, includes a downmixing unit 112 instead of the downmixing unit 110, and as shown by the dashed line in FIG. 10, the left-right relationship information estimation unit 182 obtains and outputs the left-right correlation coefficient γ and the preceding channel information, and the output left-right correlation coefficient γ and the preceding channel information are input to the downmixing unit 112 and used. Other configurations and operations of the encoding device 102 of the first embodiment are the same as those of the encoding device 101 of the second reference form. The encoding device 102 of the first embodiment performs the processing from step S112 to step S191 illustrated in FIG. 14 for each frame. Hereinafter, the differences between the encoding device 102 of the first embodiment and the encoding device 101 of the second reference form will be described.

[0115] [Left-Right Relationship Information Estimation Unit 182] The left-right relationship information estimation unit 182 receives the input audio signal of the left channel input to the encoding device 102 and the input audio signal of the right channel input to the encoding device 102. The left-right relationship information estimation unit 182 obtains and outputs the left-right time difference τ, the left-right time difference code Cτ which is a code representing the left-right time difference τ, the left-right correlation coefficient γ, and the preceding channel information from the input audio signal of the left channel and the input audio signal of the right channel (step S182). The process by which the left-right relationship information estimation unit 182 obtains the left-right time difference τ and the left-right time difference code Cτ is the same as that of the left-right relationship information estimation unit 181 in the second reference embodiment.

[0116] The left-right correlation coefficient γ is information corresponding to the correlation coefficient between the audio signal that has reached and been picked up by the microphone for the left channel from the sound source and the audio signal that has reached and been picked up by the microphone for the right channel from the sound source in the assumption described in the explanation part of the left-right relationship information estimation unit 181 in the second reference embodiment. The preceding channel information is information corresponding to which microphone the sound emitted by the sound source reaches earlier, and is information indicating which of the input audio signal of the left channel and the input audio signal of the right channel contains the same audio signal earlier, and is information indicating which of the left channel and the right channel is preceding.

[0117] In the example described in the explanation part of the left-right relationship information estimation unit 181 in the second reference embodiment, the left-right relationship information estimation unit 182 calculates the correlation value between the sample sequence of the input audio signal of the left channel and the sample sequence of the input audio signal of the right channel that is shifted later than the sample sequence by the left-right time difference τ, that is, τ max from τ min to each candidate sample number τ cand and obtains the correlation coefficient γ candThe maximum value among them is obtained and output as the left-right correlation coefficient γ. Further, when the left-right time difference τ is a positive value, the left-right relationship information estimation unit 182 obtains and outputs information indicating that the left channel is leading as the leading channel information. When the left-right time difference τ is a negative value, the left-right relationship information estimation unit 182 obtains and outputs information indicating that the right channel is leading as the leading channel information. When the left-right time difference τ is 0, the left-right relationship information estimation unit 182 may obtain and output information indicating that the left channel is leading as the leading channel information, or may obtain and output information indicating that the right channel is leading as the leading channel information. However, it is preferable to obtain and output information indicating that neither channel is leading as the leading channel information.

[0118] [Downmixing unit 112] The downmixing unit 112 receives the input audio signal of the left channel input to the encoding device 102, the input audio signal of the right channel input to the encoding device 102, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 182, and the leading channel information output by the left-right relationship information estimation unit 182. The downmixing unit 112 weights and averages the input audio signal of the left channel and the input audio signal of the right channel so that the input audio signal of the leading channel among the input audio signal of the left channel and the input audio signal of the right channel is included in the downmixing signal to a greater extent as the left-right correlation coefficient γ is larger, and obtains and outputs a downmixing signal (step S112).

[0119] For example, if the absolute value of the correlation coefficient or the normalized value is used for the correlation value as in the example described in the explanation part of the left-right relationship information estimation unit 181 of the second reference embodiment, since the obtained left-right correlation coefficient γ is a value between 0 and 1, the downmixing unit 112 uses the weight determined by the left-right correlation coefficient γ for each corresponding sample number t. L (t) and the input audio signal x of the right channel R (t) are weighted and added to obtain a downmixing signal x MIt may be set as (t). Specifically, when the precedence channel information indicates that the left channel is the leading channel, that is, when the left channel is the leading channel, x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t). When the precedence channel information indicates that the right channel is the leading channel, that is, when the right channel is the leading channel, x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t), and the downmix signal x M (t) may be obtained in this way. By obtaining the downmix signal in this manner in the downmix unit 112, the downmix signal is closer to the signal obtained by the average of the input sound signal of the left channel and the input sound signal of the right channel as the left-right correlation coefficient γ is smaller, that is, as the correlation between the input sound signal of the left channel and the input sound signal of the right channel is smaller. The larger the left-right correlation coefficient γ is, that is, the larger the correlation between the input sound signal of the left channel and the input sound signal of the right channel is, the closer it is to the input sound signal of the leading channel between the input sound signal of the left channel and the input sound signal of the right channel.

[0120] In addition, when neither channel is leading in the downmix unit 112, it is preferable to obtain and output the downmix signal by averaging the input sound signal of the left channel and the input sound signal of the right channel so that the input sound signal of the left channel and the input sound signal of the right channel are included in the downmix signal with the same weight. Therefore, when the precedence channel information indicates that neither channel is leading, for each sample number t, the input sound signal x L (t) of the left channel and the input sound signal x R (t) of the right channel are averaged to obtain x M (t) = (x L (t) + x R (t)) / 2 as the downmix signal x M (t).

[0121] <Second Embodiment> For the encoding device 100 of the first reference embodiment, a modification may be made to generate a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel. This embodiment will be described as the second embodiment. Since the code obtained by the encoding device of the second embodiment can be decoded by the decoding device 200 of the first reference embodiment, the description of the decoding device will be omitted.

[0122] <<Encoding Device 103>> As shown in FIG. 1, the encoding device 103 of the second embodiment includes a downmix unit 112, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a monaural encoding unit 160, a stereo encoding unit 170, and a left-right relationship information estimation unit 183. The encoding device 103 of the second embodiment is different from the encoding device 100 of the first reference embodiment in that it includes a downmix unit 112 instead of the downmix unit 110, and as shown by the dashed line in FIG. 1, it includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ and the previous channel information, and the output left-right correlation coefficient γ and previous channel information are input to the downmix unit 112 and used. The other configurations and operations of the encoding device 103 of the second embodiment are the same as those of the encoding device 100 of the first reference embodiment. Also, the operation of the downmix unit 112 of the encoding device 103 of the second embodiment is the same as the operation of the downmix unit 112 of the encoding device 102 of the first embodiment. The encoding device 103 of the second embodiment performs the processing from step S112 to step S183 illustrated in FIG. 15 for each frame. Hereinafter, the differences between the encoding device 103 of the second embodiment and both the encoding device 100 of the first reference embodiment and the encoding device 102 of the first embodiment will be described.

[0123] [Left-Right Relationship Information Estimation Unit 183] The left-right relationship information estimation unit 183 receives the input audio signal of the left channel input to the encoding device 103 and the input audio signal of the right channel input to the encoding device 103. The left-right relationship information estimation unit 183 obtains and outputs the left-right correlation coefficient γ and the preceding channel information from the input audio signal of the left channel and the input audio signal of the right channel (step S183).

[0124] The left-right correlation coefficient γ and the preceding channel information obtained and output by the left-right relationship information estimation unit 183 are the same as those described in the first embodiment. That is, the left-right relationship information estimation unit 183 may be the same as the left-right relationship information estimation unit 182 except that it does not obtain and output the left-right time difference τ and the left-right time difference code Cτ.

[0125] For example, the left-right relationship information estimation unit 183 max from τ min to each candidate sample number τ cand for the sample sequence of the input audio signal of the left channel and the sample sequence of the input audio signal of the right channel at a position shifted by each candidate sample number τ cand samples after the sample sequence, obtains and outputs the maximum value of the correlation value γ cand as the left-right correlation coefficient γ, and when τ cand at which the correlation value is the maximum value is a positive value, obtains and outputs information indicating that the left channel is leading as the preceding channel information, and when τ cand at which the correlation value is the maximum value is a negative value, obtains and outputs information indicating that the right channel is leading as the preceding channel information. When τ cand at which the correlation value is the maximum value is 0, the left-right relationship information estimation unit 183 may obtain and output information indicating that the left channel is leading as the preceding channel information, or may obtain and output information indicating that the right channel is leading as the preceding channel information, but it is preferable to obtain and output information indicating that neither channel is leading as the preceding channel information.

[0126] <Third Embodiment> Even for an encoding device that stereo-encodes the input audio signal of each channel instead of the differential signal of each channel, a configuration may be adopted in which a downmix signal is obtained in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel, and this mode will be described as the third embodiment.

[0127] ≪Encoding Device 104≫ As shown in FIG. 16, the encoding device 104 of the third embodiment includes a left-right relationship information estimation unit 183, a downmix unit 112, a monaural encoding unit 160, and a stereo encoding unit 174. For each frame, the encoding device 104 of the third embodiment performs the processes of step S183, step S112, step S160, and step S174 illustrated in FIG. 17. Hereinafter, the encoding device 104 of the third embodiment will be described with appropriate reference to the description of the second embodiment.

[0128] [Left-Right Relationship Information Estimation Unit 183] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment. The left-right relationship information estimation unit 183 receives the input audio signal of the left channel input to the encoding device 104 and the input audio signal of the right channel input to the encoding device 104. The left-right relationship information estimation unit 183 obtains and outputs from the input audio signal of the left channel and the input audio signal of the right channel a left-right correlation coefficient γ that is the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and precedence channel information that is information indicating which of the input audio signal of the left channel and the input audio signal of the right channel is ahead (step S183).

[0129] [Downmix Unit 112] The downmixing unit 112 is the same as the downmixing unit 112 of the second embodiment. The input to the downmixing unit 112 includes the input audio signal of the left channel input to the encoding device 104, the input audio signal of the right channel input to the encoding device 104, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 183, and the preceding channel information output by the left-right relationship information estimation unit 183. The downmixing unit 112 weights and averages the input audio signal of the left channel and the input audio signal of the right channel so that the input audio signal of the preceding channel among the input audio signal of the left channel and the input audio signal of the right channel is included in the downmixing signal to a greater extent as the left-right correlation coefficient γ is larger, and obtains and outputs a downmixing signal (step S112).

[0130] For example, assuming the sample number is t, the input audio signal of the left channel is x L (t), the input audio signal of the right channel is x R (t), and the downmixing signal is x M (t), when the preceding channel information indicates that the left channel is preceding, for each sample number t, the downmixing unit 112 obtains the downmixing signal by x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t). When the preceding channel information indicates that the right channel is preceding, for each sample number t, the downmixing unit 112 obtains the downmixing signal by x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t). When the preceding channel information indicates that neither channel is preceding, for each sample number t, the downmixing unit 112 obtains the downmixing signal by x M (t) = (x L (t) + x R (t)) / 2.

[0131] [Mono Encoding Unit 160] The monaural encoding unit 160 is the same as the monaural encoding unit 160 of the second embodiment. The downmix signal output by the downmix unit 112 is input to the monaural encoding unit 160. The monaural encoding unit 160 encodes the input downmix signal to obtain and output a monaural code CM (step S160). The monaural encoding unit 160 may use any encoding method. For example, an encoding method such as the 3GPP EVS standard may be used. The encoding method may be an encoding method that performs encoding processing independently of the stereo encoding unit 174 described later, that is, an encoding method that performs encoding processing without using the stereo code CS' obtained by the stereo encoding unit 174 or the information obtained in the encoding processing performed by the stereo encoding unit 174, or an encoding method that performs encoding processing using the stereo code CS' obtained by the stereo encoding unit 174 or the information obtained in the encoding processing performed by the stereo encoding unit 174.

[0132] [Stereo Encoding Unit 174] The input to the stereo encoding unit 174 is the input audio signal of the left channel input to the encoding device 104 and the input audio signal of the right channel input to the encoding device 104. The stereo encoding unit 174 encodes the input audio signal of the left channel and the input audio signal of the right channel to obtain and output a stereo code CS' (step S174). The stereo encoding unit 174 may use any encoding method. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard may be used, or an encoding method that independently encodes each of the input audio signal of the left channel and the input audio signal of the right channel may be used. The combined codes obtained by encoding may be used as the stereo code CS'. The encoding method may be an encoding method that performs encoding processing independently of the monaural encoding unit 160, that is, an encoding method that performs encoding processing without using the monaural code CM obtained by the monaural encoding unit 160 or the information obtained in the encoding processing performed by the monaural encoding unit 160, or an encoding method that performs encoding processing using the monaural code CM obtained by the monaural encoding unit 160 or the information obtained in the encoding processing performed by the monaural encoding unit 160.

[0133] <Fourth Embodiment> As can be understood from the description of the above embodiments, as long as an encoding device obtains a code by encoding at least a downmix signal obtained from an input audio signal of a left channel and an input audio signal of a right channel, no matter what kind of encoding device it is, a configuration for obtaining a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel may be adopted. Further, not limited to the encoding device, as long as a signal processing device obtains a signal processing result by at least signal-processing a downmix signal obtained from an input audio signal of a left channel and an input audio signal of a right channel, no matter what kind of signal processing device it is, a configuration for obtaining a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel may be adopted. Furthermore, as a downmix device used before these encoding devices and signal processing devices, a configuration for obtaining a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel may be adopted. These modes will be described as the fourth embodiment.

[0134] <<Audio signal encoding device 105>> As shown in FIG. 18, the audio signal encoding device 105 according to the fourth embodiment includes a left-right relationship information estimation unit 183, a downmix unit 112, and an encoding unit 195. The audio signal encoding device 105 according to the fourth embodiment performs the processes of step S183, step S112, and step S195 illustrated in FIG. 19 for each frame. Hereinafter, the audio signal encoding device 105 according to the fourth embodiment will be described with appropriate reference to the description of the second embodiment.

[0135] [Left-right relationship information estimation unit 183] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment, and obtains and outputs a left-right correlation coefficient γ, which is a correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and precedence channel information, which is information indicating which of the input audio signal of the left channel and the input audio signal of the right channel is preceding, from the input audio signal of the left channel and the input audio signal of the right channel that are input (step S183).

[0136] [Downmix section 112] The downmix section 112 is the same as the downmix section 112 of the second embodiment. The downmix section 112 weights and averages the input audio signal of the left channel and the input audio signal of the right channel so that the input audio signal of the leading channel among the input audio signal of the left channel and the input audio signal of the right channel in the downmix signal is included more as the left-right correlation coefficient γ is larger, and obtains and outputs a downmix signal (step S112).

[0137] [Encoding section 195] At least the downmix signal output by the downmix section 112 is input to the encoding section 195. The encoding section 195 encodes at least the input downmix signal to obtain and output an audio signal code (step S195). The encoding section 195 may also encode the input audio signal of the left channel and the input audio signal of the right channel, and may also include the codes obtained by this encoding in the audio signal code and output them. In this case, as shown by the broken line in FIG. 18, the input audio signal of the left channel and the input audio signal of the right channel are also input to the encoding section 195.

[0138] ≪Audio signal processing apparatus 305≫ As shown in FIG. 20, the audio signal processing apparatus 305 of the fourth embodiment includes a left-right relationship information estimation section 183, a downmix section 112, and a signal processing section 315. The audio signal processing apparatus 305 of the fourth embodiment performs the processes of step S183, step S112, and step S315 illustrated in FIG. 21 for each frame. Hereinafter, the differences between the audio signal processing apparatus 305 of the fourth embodiment and the audio signal encoding apparatus 105 of the fourth embodiment will be described.

[0139] [Signal processing section 315] The signal processing unit 315 receives at least the downmix signal output from the downmix unit 112. The signal processing unit 315 performs at least signal processing on the input downmix signal to obtain a signal processing result and outputs it (step S315). The signal processing unit 315 may also perform signal processing on the input sound signal of the left channel and the input sound signal of the right channel to obtain a signal processing result. In this case, as shown by the dashed line in FIG. 20, the input sound signal of the left channel and the input sound signal of the right channel are also input to the signal processing unit 315. The signal processing unit 315 may, for example, perform signal processing using the downmix signal on the input sound signal of each channel to obtain the output sound signal of each channel as the signal processing result, or may perform this signal processing on the decoded sound signal of the left channel and the decoded sound signal of the right channel obtained by decoding the code CS' obtained by the stereo encoding unit 174 of the third embodiment using a decoding device provided with a decoding unit corresponding to the stereo encoding unit 174. That is, it is not essential that the input sound signal of the left channel and the input sound signal of the right channel input to the sound signal processing device 305 are digital audio signals or acoustic signals obtained by collecting and AD-converting the sounds with two microphones respectively. The input sound signal of the left channel and the input sound signal of the right channel input to the sound signal processing device 305 may be the decoded sound signal of the left channel and the decoded sound signal of the right channel obtained by decoding the code, or may be any sound signal obtained in any way as long as it is a stereo two-channel sound signal.

[0140] When the input audio signal of the left channel and the input audio signal of the right channel input to the audio signal processing device 305 are the decoded audio signal of the left channel and the decoded audio signal of the right channel obtained by decoding the code in another device, etc., there are cases where either one or both of the left-right correlation coefficient γ and the preceding channel information obtained by the left-right relationship information estimation unit 183 are obtained in another device. When either one or both of the left-right correlation coefficient γ and the preceding channel information are obtained in another device, as shown by the dashed line in FIG. 20, the audio signal processing device 305 may be configured to input either one or both of the left-right correlation coefficient γ and the preceding channel information obtained in another device. In this case, the left-right relationship information estimation unit 183 may obtain the left-right correlation coefficient γ or the preceding channel information that was not input to the audio signal processing device 305. When both the left-right correlation coefficient γ and the preceding channel information are input to the audio signal processing device 305, the audio signal processing device 305 may not include the left-right relationship information estimation unit 183 and may not perform step S183. That is, as shown by the double-dashed line in FIG. 20, the audio signal processing device 305 includes a left-right relationship information acquisition unit 185, and the left-right relationship information acquisition unit 185 may obtain and output the left-right correlation coefficient γ, which is the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and the preceding channel information, which is information indicating which of the input audio signal of the left channel and the input audio signal of the right channel is preceding (step S185). Note that the left-right relationship information estimation unit 183 and step S183 of each device described above can also be said to be within the scope of the left-right relationship information acquisition unit 185 and step S185.

[0141] ≪Audio Signal Downmixing Device 405≫ As shown in FIG. 22, the audio signal downmixing device 405 according to the fourth embodiment includes a left-right relationship information acquisition unit 185 and a downmixing unit 112. For each frame, the audio signal downmixing device 405 performs the processes of step S185 and step S112 illustrated in FIG. 23. Hereinafter, the audio signal downmixing device 405 will be described with appropriate reference to the description of the second embodiment. Similar to the audio signal processing device 305, the input audio signal of the left channel and the input audio signal of the right channel input to the audio signal downmixing device 405 may be digital audio signals or acoustic signals obtained by collecting sound with two microphones respectively and performing AD conversion, or may be the decoded audio signals of the left channel and the right channel obtained by decoding the codes, or may be audio signals obtained in any manner as long as they are stereo two-channel audio signals.

[0142] [Left-Right Relationship Information Acquisition Unit 185] The left-right relationship information acquisition unit 185 obtains and outputs a left-right correlation coefficient γ, which is the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, and leading channel information, which is information indicating which of the input audio signal of the left channel and the input audio signal of the right channel is leading (step S185).

[0143] When both the left-right correlation coefficient γ and the leading channel information are obtained by another device, as shown by the dashed-dotted line in FIG. 22, the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ and the leading channel information input from the other device to the audio signal downmixing device 405 and outputs them to the downmixing unit 112.

[0144] When both the left-right correlation coefficient γ and the leading channel information are not obtained by another device, as shown by the broken line in FIG. 22, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and the leading channel information from the input audio signal of the left channel and the input audio signal of the right channel in the same manner as the left-right relationship information estimation unit 183 of the second embodiment, and outputs them to the downmixing unit 112.

[0145] When either the left-right correlation coefficient γ or the leading channel information is not obtained by another device, as shown by the dashed line in FIG. 22, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 of the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ that is not obtained by another device or the leading channel information that is not obtained by another device from the input sound signals of the left channel and the right channel in the same manner as the left-right relationship information estimation unit 183 in the second embodiment, and outputs it to the downmixing unit 112. For the left-right correlation coefficient γ obtained by another device or the leading channel information obtained by another device, as shown by the dashed-dotted line in FIG. 22, the left-right relationship information acquisition unit 185 outputs the left-right correlation coefficient γ or the leading channel information input from another device to the sound signal downmixing device 405 to the downmixing unit 112.

[0146] [Downmixing unit 112] The downmixing unit 112 is the same as the downmixing unit 112 in the second embodiment. Based on the leading channel information and the left-right correlation coefficient acquired by the left-right relationship information acquisition unit 185, the input sound signal of the leading channel among the input sound signals of the left channel and the right channel is included in the downmixing signal to a greater extent as the left-right correlation coefficient γ is larger. The input sound signals of the left channel and the right channel are weighted and averaged to obtain and output a downmixing signal (step S112).

[0147] For example, when the sample number is t, the input sound signal of the left channel is x L (t), the input sound signal of the right channel is x R (t), and the downmixing signal is x M (t), then when the leading channel information indicates that the left channel is leading, for each sample number t, x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t) is used to obtain the downmixing signal. When the leading channel information indicates that the right channel is leading, for each sample number t, x M(t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R When obtaining a downmix signal by (t) and indicating that the preceding channel information does not precede any channel, for each sample number t, x M (t) = (x L (t) + x R (t)) / 2 to obtain a downmix signal.

[0148] <Program and Recording Medium> The processing of each part of the above-described encoding device, decoding device, audio signal encoding device, audio signal processing device, and audio signal downmixing device may be realized by a computer. In this case, the processing content of the functions that each device should have is described by a program. Then, by causing this program to be read into the storage unit 1020 of the computer 1000 shown in FIG. 24 and operating it on the arithmetic processing unit 1010, input unit 1030, output unit 1040, etc., various processing functions in the above-described devices are realized on the computer.

[0149] The program describing this processing content can be recorded on a computer-readable recording medium. A computer-readable recording medium is, for example, a non-temporary recording medium, and specifically, a magnetic recording device, an optical disk, etc.

[0150] Also, the distribution of this program is performed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, it is also possible to configure the distribution of this program by storing the program in the storage device of a server computer and transferring the program from the server computer to other computers via a network.

[0151] A computer that executes such a program first stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer once in an auxiliary recording unit 1050 which is its own non-temporary storage device. Then, at the time of executing processing, this computer reads the program stored in the auxiliary recording unit 1050 which is its own non-temporary storage device into a storage unit 1020, and executes processing according to the read program. Also, as another execution form of this program, it is also possible that the computer directly reads the program from the portable recording medium into the storage unit 1020 and executes processing according to the program. Further, every time a program is transferred from the server computer to this computer, it is also possible to sequentially execute processing according to the received program. Also, it is also possible to configure to execute the above-described processing by a so-called ASP (Application Service Provider) type service that does not transfer the program from the server computer to this computer and realizes the processing function only by the execution instruction and result acquisition. Note that the program in this embodiment includes information for use in processing by an electronic computer that conforms to the program (data etc. that are not direct instructions for the computer but have the property of defining the processing of the computer).

[0152] Also, in this embodiment, the present apparatus is configured by causing a computer to execute a predetermined program, but at least a part of these processing contents may be realized hardware-wise.

[0153] Needless to say, appropriate changes can be made within the scope not departing from the gist of the present invention.

Claims

1. 1. A sound signal downmixing method for obtaining a downmix signal which is a signal obtained by mixing a first channel input sound signal and a second channel input sound signal, comprising the steps of: obtaining a code representing whether the first channel input sound signal or the second channel input sound signal is leading; a downmix step of obtaining the downmix signal from the preceding channel and the other channel based on a degree determined based on the code and a correlation coefficient which is a coefficient indicating a magnitude of correlation between the first channel input sound signal and the second channel input sound signal, A sound signal downmixing method, wherein in the downmixing step, a leading signal of the first channel input sound signal and the second channel input sound signal is given a greater weight than the other signal.

2. 1. A sound signal downmixing device for obtaining a downmix signal which is a signal obtained by mixing a first channel input sound signal and a second channel input sound signal, a left-right relationship information acquiring unit that acquires a code indicating which of the first channel input sound signal and the second channel input sound signal is preceding; a downmix unit that obtains the downmix signal from the preceding channel and another channel based on a degree determined based on the code and a correlation coefficient that is a coefficient indicating a magnitude of correlation between the first channel input sound signal and the second channel input sound signal, In the downmixing section, a leading signal of the first channel input sound signal and the second channel input sound signal is given a greater weight than the other signal.

3. A program for causing a computer to execute each step of a sound signal downmixing method for obtaining a downmix signal, which is a signal obtained by mixing a first channel input sound signal and a second channel input sound signal, comprising: The audio signal downmixing method includes: obtaining a code representing whether the first channel input sound signal or the second channel input sound signal is leading; a downmix step of obtaining the downmix signal from the preceding channel and the other channel based on a degree determined based on the code and a correlation coefficient which is a coefficient indicating a magnitude of correlation between the first channel input sound signal and the second channel input sound signal; having In the downmixing step, a leading signal of the first channel input sound signal and the second channel input sound signal is given a greater weight than the other signal. program.

Citation Information

Patent Citations

  • Audio signal downmixing method, audio signal downmixing device and program

    JP7655369B2

  • Stereo signal converter, stereo signal reverse converter, and methods for both

    WO2009122757A1

  • Stereo acoustic signal encoding apparatus, stereo acoustic signal decoding apparatus, and methods for the same

    WO2010084756A1

  • Down-mixing device, encoder, and method therefor

    WO2010140350A1

  • Audio signal coding device and audio signal decoding device

    WO2014068817A1