Sound signal mixing method, sound signal encoding method, sound signal mixing device, sound signal encoding device, and recording medium

Through the sound signal downmix method, the correlation coefficient and advance channel information of the sound signal inputted by the left and right channels are used to generate a downmix signal and perform mono encoding, which solves the problem of difficulty in extracting useful mono signals in the prior art, and realizes efficient encoding processing.

CN115280411BActive Publication Date: 2025-06-20NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080098232.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-09
Filing Date
2020-11-04
Publication Date
2025-06-20
Estimated Expiration
2040-11-04

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract mono signals useful for encoding processing from two-channel sound signals.

Method used

Through a sound signal downmix method, the correlation coefficient and advance channel information of the input sound signals of the left and right channels are obtained, weighted averaged to generate the downmix signal, and the downmix signal is encoded mono.

Benefits of technology

It is realized that mono signals useful for encoding processing are extracted from the two-channel sound signal, which improves encoding efficiency and suppresses the sound quality degradation of the decoded sound signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280411B_ABST
    Figure CN115280411B_ABST
Patent Text Reader

Abstract

An audio signal mixing device is an audio signal mixing device that obtains a mixed signal, i.e., a mixed signal, which is a mixture of a left-channel input audio signal and a right-channel input audio signal. It includes: a left-right relationship information acquisition unit 185 that obtains prior channel information and a left-right correlation coefficient. The prior channel information is information indicating which of the left-channel input audio signal and the right-channel input audio signal is prior, and the left-right correlation coefficient is the correlation coefficient between the left-channel input audio signal and the right-channel input audio signal; and a mixing unit 112 that performs weighted averaging on the left-channel input audio signal and the right-channel input audio signal according to the prior channel information and the left-right correlation coefficient to obtain a mixed signal, so that the larger the left-right correlation coefficient, the more the input audio signal of the prior channel in the left-channel input audio signal and the right-channel input audio signal is included.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for obtaining a monaural audio signal from a stereo audio signal for encoding an audio signal in mono, or encoding an audio signal using both mono encoding and stereo encoding, or signal-processing an audio signal in mono, or signal-processing a stereo audio signal using a monaural audio signal. Background Art

[0002] As a technique for obtaining a monaural audio signal from a stereo audio signal by performing embedded encoding / decoding on the stereo audio signal and the monaural audio signal, there is the technique of Patent Document 1. Patent Document 1 discloses the following technique: for each corresponding sample, the input left-channel audio signal and the input right-channel audio signal are averaged to obtain a monaural signal, the monaural signal is encoded (mono encoding) to obtain a mono encoding, the mono encoding is decoded (mono decoding) to obtain a mono partial decoding signal, and for each of the left and right channels, the difference between the input audio signal and the prediction signal obtained from the mono partial decoding signal (prediction residual signal) is encoded. In the technique of Patent Document 1, for each channel, a signal obtained by delaying the mono partial decoding signal and giving an amplitude ratio is used as the prediction signal, and a prediction signal having a delay and amplitude ratio that minimizes the error between the input audio signal and the prediction signal, or a prediction signal having a delay difference and amplitude ratio that maximizes the cross-correlation between the input audio signal and the mono partial decoding signal is selected, the prediction signal is subtracted from the input audio signal to obtain a prediction residual signal, and the prediction residual signal is made the object of encoding / decoding, thereby suppressing the deterioration of the sound quality of the decoded audio signal for each channel.

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: WO2006-070751 Summary of the Invention

[0006] Problems to be Solved by the Invention

[0007] In the technology of Patent Document 1, by optimizing the delay and amplitude ratio given to the mono local decoded signal when obtaining the prediction signal, the coding efficiency of each channel can be improved. However, in the technology of Patent Document 1, the mono local decoded signal is obtained by encoding / decoding a mono signal obtained by averaging the sound signals of the left channel and the right channel. That is, there is a problem in the technology of Patent Document 1 that there is no study on obtaining a mono signal useful for signal processing such as encoding processing from the two-channel sound signals.

[0008] In the present invention, an object is to provide a technology for obtaining a mono signal useful for signal processing such as encoding processing from two-channel sound signals.

[0009] Means for solving the problem

[0010] A sound signal downmixing method according to an aspect of the present invention obtains a downmixing signal, which is a signal obtained by mixing a left-channel input sound signal and a right-channel input sound signal, and is characterized by including: a left-right relationship information obtaining step of obtaining antecedent channel information and a left-right correlation coefficient, where the antecedent channel information is information indicating which of the left-channel input sound signal and the right-channel input sound signal is antecedent, and the left-right correlation coefficient is the correlation coefficient between the left-channel input sound signal and the right-channel input sound signal; and a downmixing step of performing weighted averaging on the left-channel input sound signal and the right-channel input sound signal according to the antecedent channel information and the left-right correlation coefficient to obtain the downmixing signal, so that the larger the left-right correlation coefficient γ is, the more the input sound signal of the antecedent channel in the left-channel input sound signal and the right-channel input sound signal is included.

[0011] In a sound signal downmixing method according to an aspect of the present invention, it is characterized in that, assuming the sampling number is t, the left-channel input sound signal is x L (t), the right-channel input sound signal is x R (t), the downmixing signal is x M (t), and the left-right correlation coefficient is γ. In the downmixing step, when the antecedent channel information indicates that the left channel is antecedent, for each sampling number t, the downmixing signal is obtained by x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t). When the antecedent channel information indicates that the right channel is antecedent, for each sampling number t, the downmixing signal is obtained by x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t). When the antecedent channel information indicates that neither channel is antecedent, for each sampling number t, through x My(t) = (x L (t) + x R (t)) / 2 to obtain the downmixed signal.

[0012] A method for downmixing audio signals according to one embodiment of the present invention is characterized in that, as the audio signal downmixing step, it includes the above-described audio signal downmixing method, and further includes: a mono encoding step of encoding the downmixed signal obtained in the downmixing step to obtain a mono encoding; and a stereo encoding step of encoding the left-channel input audio signal and the right-channel input audio signal to obtain a stereo encoding.

[0013] Advantages of the Invention

[0014] According to the present invention, a mono signal useful for signal processing such as encoding processing can be obtained from a two-channel audio signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a block diagram showing an example of an encoding apparatus for a first reference mode and a second embodiment.

[0016] Figure 2 It is a flowchart showing an example of the processing of an encoding apparatus for a first reference mode.

[0017] Figure 3 It is a block diagram showing an example of a decoding apparatus for a first reference mode.

[0018] Figure 4 It is a flowchart showing an example of the processing of a decoding apparatus for a first reference mode.

[0019] Figure 5 It is a flowchart showing an example of the processing of a left-channel subtraction gain estimation unit and a right-channel subtraction gain estimation unit for a first reference mode.

[0020] Figure 6 It is a flowchart showing an example of the processing of a left-channel subtraction gain estimation unit and a right-channel subtraction gain estimation unit for a first reference mode.

[0021] Figure 7 It is a flowchart showing an example of the processing of a left-channel subtraction gain decoding unit and a right-channel subtraction gain decoding unit for a first reference mode.

[0022] Figure 8 It is a flowchart showing an example of the processing of a left-channel subtraction gain estimation unit and a right-channel subtraction gain estimation unit for a first reference mode.

[0023] Figure 9 It is a flowchart showing an example of the processing of a left-channel subtraction gain estimation unit and a right-channel subtraction gain estimation unit for a first reference mode.

[0024] Figure 10 It is a block diagram showing an example of an encoding device for a second reference method and a first embodiment.

[0025] Figure 11 It is a flowchart showing an example of the processing of an encoding device for a second reference method.

[0026] Figure 12 It is a block diagram showing an example of a decoding device for a second reference method.

[0027] Figure 13 It is a flowchart showing an example of the processing of a decoding device for a second reference method.

[0028] Figure 14 It is a flowchart showing an example of the processing of an encoding device for a first embodiment.

[0029] Figure 15 It is a flowchart showing an example of the processing of an encoding device for a second embodiment.

[0030] Figure 16 It is a block diagram showing an example of an encoding device for a third embodiment.

[0031] Figure 17 It is a flowchart showing an example of the processing of an encoding device for a third embodiment.

[0032] Figure 18 It is a block diagram showing an example of an audio signal encoding device for a fourth embodiment.

[0033] Figure 19 It is a flowchart showing an example of the processing of an audio signal encoding device for a fourth embodiment.

[0034] Figure 20 It is a block diagram showing an example of an audio signal processing device for a fourth embodiment.

[0035] Figure 21 It is a flowchart showing an example of the processing of an audio signal processing device for a fourth embodiment.

[0036] Figure 22 It is a block diagram showing an example of an audio signal mixing device for a fourth embodiment.

[0037] Figure 23 It is a flowchart showing an example of the processing of an audio signal mixing device for a fourth embodiment.

[0038] Figure 24 It is a diagram showing an example of the functional structure of a computer for each device in the embodiments implementing the present invention. Detailed Embodiments

[0039] <First Embodiment>

[0040] First, the method of recording in the specification will be described. For the "^" in the superscript like "^x" of a certain character x, it should originally be recorded directly above "x". However, due to limitations in the recording expression of the specification, it is sometimes recorded as ^x.

[0041] <First reference mode>

[0042] Before describing the embodiments of the invention, as the first reference mode and the second reference mode, an encoding device and a decoding device that form the basis for implementing the invention of the second embodiment and the invention of the first embodiment will be described. Additionally, in the specification and claims, the encoding device is sometimes referred to as an audio signal encoding device, the encoding method as an audio signal encoding method, the decoding device as an audio signal decoding device, and the decoding method as an audio signal decoding method.

[0043] <<Encoding device 100>>

[0044] As Figure 1 shown, the encoding device 100 of the first reference mode includes: a down mix unit 110, a left channel subtraction gain estimation unit 120, a left channel signal subtraction unit 130, a right channel subtraction gain estimation unit 140, a right channel signal subtraction unit 150, a mono encoding unit 160, and a stereo encoding unit 170. The encoding device 100 encodes the input two-channel stereo time-domain audio signal in units of frames with a specified time length of 20 ms to obtain and output the mono encoding CM, the left channel subtraction gain encoding Cα, the right channel subtraction gain encoding Cβ, and the stereo encoding CS described later. The two-channel stereo time-domain audio signal input to the encoding device is, for example, a digital audio signal or an audio signal obtained by separately picking up sounds such as voices and music with two microphones and performing AD conversion, and is composed of the input audio signal of the left channel and the input audio signal of the right channel. The encodings output by the encoding device, namely the mono encoding CM, the left channel subtraction gain encoding Cα, the right channel subtraction gain encoding Cβ, and the stereo encoding CS, are input to the decoding device. The encoding device 100 performs the processing of steps S110 to S170 as Figure 2 illustrated for each frame.

[0045] [Down mix unit 110]

[0046] The input audio signal of the left channel input to the encoding device 100 and the input audio signal of the right channel input to the encoding device 100 are input to the down mix unit 110. The down mix unit 110 obtains a down mix signal that is a mixture of the input audio signal of the left channel and the input audio signal of the right channel from the input signals and outputs it (step S110).

[0047] For example, when the number of samples per frame is set to T, the input audio signal x of the left channel input to the encoding device 100 in units of frames is input to the mixing unit 110 L (1), x L (2),..., x L (T) and the input audio signal x of the right channel R (1), x R (2),..., x R (T). Here, T is a positive integer. For example, when the frame length is 20 ms and the sampling frequency is 32 kHz, T is 640. The mixing unit 110 obtains and outputs a sequence of the average values of the sampling values of the corresponding each sample based on the input audio signal of the left channel and the input audio signal of the right channel as the mixed signal x M (1), x M (2),..., x M (T). That is, if the sample number is set to t, it is x M (t) = (x L (t) + x R (t)) / 2.

[0048] [Left channel subtraction gain estimation unit 120]

[0049] The input audio signal x of the left channel input to the encoding device 100 is input to the left channel subtraction gain estimation unit 120 L (1), x L (2),..., x L (T) and the mixed signal x output by the mixing unit 110 M (1), x M (2),..., x M (T). The left channel subtraction gain estimation unit 120 obtains the left channel subtraction gain α and the encoding representing the left channel subtraction gain α, that is, the left channel subtraction gain encoding Cα, and outputs them (step S120). The left channel subtraction gain estimation unit 120 obtains the left channel subtraction gain α and the left channel subtraction gain encoding Cα by a known method exemplified by the method of obtaining the amplitude ratio g in Patent Document 1, the method of encoding the amplitude ratio g, or a method based on the principle of minimizing a newly proposed quantization error. The principle of minimizing the quantization error and the method based on this principle will be described later.

[0050] [Left channel signal subtraction unit 130]

[0051] The input audio signal x of the left channel input to the encoding device 100 is input to the left channel signal subtraction unit 130L (1), x L (2),..., x L (T), the downmix signal x output from the downmix unit 110 M (1), x M (2),..., x M (T), the left-channel subtraction gain α output from the left-channel subtraction gain estimation unit 120. For each corresponding sample t, the left-channel signal subtraction unit 130 subtracts the value α×x obtained by multiplying the sample value x L (t) of the downmix signal from the sample value x M (t) of the input sound signal of the left channel M (t) to obtain a value x L (t) - α×x M (t) as the left-channel difference signal y L (1), y L (2),..., y L (T) is obtained and output (step S130). That is, y L (t) = x L (t) - α×x M (t). In the encoding device 100, in order not to require the delay or computational processing amount for obtaining the local decoded signal, in the left-channel signal subtraction unit 130, instead of the locally decoded signal of the mono encoding, which is the quantized downmix signal, the unquantized downmix signal x M (t) obtained by the downmix unit 110 can be used. However, when the left-channel subtraction gain estimation unit 120 uses a known method such as that exemplified in Patent Document 1 instead of a method based on the principle of minimizing the quantization error to obtain the left-channel subtraction gain α, a unit for obtaining the local decoded signal corresponding to the mono encoding CM is provided at the subsequent stage of the mono encoding unit 160 of the encoding device 100 or within the mono encoding unit 160. In the left-channel signal subtraction unit 130, instead of the downmix signal x M (1), x M (2),..., x M (T), similar to existing encoding devices such as Patent Document 1, the quantized downmix signal ^x M (1), ^x M (2),..., ^x M (T) as the locally decoded signal of the mono encoding can also be used to obtain the left-channel difference signal.

[0052] [Right-channel subtraction gain estimation unit 140]

[0053] The input sound signal x of the right channel input to the encoding device 100 is input to the right-channel subtraction gain estimation unit 140R (1), x R (2),..., x R (T) and the downmix signal x output from the downmix unit 110 M (1), x M (2),..., x M (T). The right-channel subtraction gain estimation unit 140 obtains the right-channel subtraction gain β and the encoding representing the right-channel subtraction gain β, i.e., the right-channel subtraction gain encoding Cβ, based on the input right-channel input audio signal and the downmix signal, and outputs them (step S140). The right-channel subtraction gain estimation unit 140 obtains the right-channel subtraction gain β and the right-channel subtraction gain encoding Cβ by a known method exemplified by the method of obtaining the amplitude ratio g in Patent Document 1, the method of encoding the amplitude ratio g, or a method based on the principle of minimizing the newly proposed quantization error. The principle of minimizing the quantization error and the method based on this principle will be described later.

[0054] [Right-channel signal subtraction unit 150]

[0055] The input audio signal x of the right channel input to the encoding device 100 is input to the right-channel signal subtraction unit 150 R (1), x R (2),..., x R (T), the downmix signal x output from the downmix unit 110 M (1), x M (2),..., x M (T) and the right-channel subtraction gain β output from the right-channel subtraction gain estimation unit 140. The right-channel signal subtraction unit 150 subtracts the value obtained by multiplying the sampled value x R (t) of the downmix signal by the right-channel subtraction gain β, i.e., β × x M (t), from the sampled value x M (t) of the input audio signal of the right channel for each corresponding sample t, and obtains and outputs the sequence of values x R (t) - β × x M (t) as the right-channel difference signal y R (1), y R (2),..., y R (T) (step S150). That is, y R (t) = x R (t) - β × x M(t). In the right-channel signal subtraction unit 150, similar to the left-channel signal subtraction unit 130, in order to avoid the delay and computational processing amount of obtaining the partial decoded signal in the encoding device 100, instead of using the quantized downmix signal of the partial decoded signal that is mono-encoded, the unquantized downmix signal x obtained by the downmix unit 110 is used M (t) is sufficient. However, when the right-channel subtraction gain estimation unit 140 obtains the right-channel subtraction gain β by a well-known method such as that exemplified in Patent Document 1, rather than a method based on the principle of minimizing the quantization error, a unit for obtaining a partial decoded signal corresponding to the mono-encoding CM is provided at the subsequent stage of the mono-encoding unit 160 in the encoding device 100 or within the mono-encoding unit 160. Similar to the left-channel signal subtraction unit 130, in the right-channel signal subtraction unit 150, instead of the downmix signal x M (1), x M (2),..., x M (T), similar to existing encoding devices such as Patent Document 1, the quantized downmix signal ^x of the partial decoded signal that is mono-encoded can also be used M (1), ^x M (2),..., ^x M (T) to obtain the right-channel differential signal.

[0056] [Mono-encoding unit 160]

[0057] The downmix signal x output by the downmix unit 110 is input to the mono-encoding unit 160 M (1), x M (2),..., x M (T). The mono-encoding unit 160 encodes the input downmix signal with a specified encoding method by b M bits to obtain the mono-encoding CM and outputs it (step S160). That is, from the input downmix signal x of T samples M (1), x M (2),..., x M (T), a b M -bit mono-encoding CM is obtained and output. As the encoding method, any method can be used, and for example, an encoding method such as the 3GPP EVS standard can also be used.

[0058] [Stereo-encoding unit 170]

[0059] The left-channel differential signal y output by the left-channel signal subtraction unit 130 is input to the stereo-encoding unit 170 L (1), y L (2),..., y L(T), and the right-channel differential signal y output from the right-channel signal subtraction unit 150 R (1), y R (2),..., y R (T). The stereo encoding unit 170 encodes the input left-channel differential signal and right-channel differential signal in a prescribed encoding method to obtain a stereo encoding CS and outputs it (step S170). That is, from the input left-channel differential signal y s (1), y L (1), y L (2),..., y L (T) and the input T-sampled right-channel differential signal y R (1), y R (2),..., y R (T) to obtain a stereo encoding CS of a total of b s bits and outputs it. As the encoding method, any method can be used. For example, a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard can be used, or a method of independently encoding the input left-channel differential signal and right-channel differential signal can be used, or the encoding obtained by combining all the encodings obtained by encoding can be used as the stereo encoding CS.

[0060] When encoding the input left-channel differential signal and right-channel differential signal independently, the stereo encoding unit 170 encodes the left-channel differential signal with b L bits and encodes the right-channel differential signal with b R bits. That is, the stereo encoding unit 170 obtains a left-channel differential encoding CL of b L (1), y L (2),..., y L (T) from the input T-sampled left-channel differential signal y L bits, and obtains a right-channel differential encoding CR of b R (1), y R (2),..., y R (T) from the input T-sampled right-channel differential signal y R bits, and outputs the encoding obtained by combining the left-channel differential encoding CL and the right-channel differential encoding CR as the stereo encoding CS. Here, the sum of b L bits and b R bits is b S bits.

[0061] When encoding the input left-channel differential signal and right-channel differential signal by combining them in one encoding method, the stereo encoding unit 170 uses a total of b SBits encode the left - channel differential signal and the right - channel differential signal. That is, the stereo encoding unit 170 encodes from the input left - channel differential signal y L (1), y L (2),..., y L (T) and the input right - channel differential signal y R (1), y R (2),..., y R (T) to obtain b S - bit stereo encoding CS and outputs it.

[0062] <<Decoding device 200>>

[0063] As Figure 3 shown, the decoding device 200 of the first reference method includes: a mono - decoding unit 210, a stereo - decoding unit 220, a left - channel subtraction gain decoding unit 230, a left - channel signal addition unit 240, a right - channel subtraction gain decoding unit 250, and a right - channel signal addition unit 260. The decoding device 200 decodes the input mono - encoding CM, left - channel subtraction gain encoding Cα, right - channel subtraction gain encoding Cβ, and stereo encoding CS in frame units of the same time length as the corresponding encoding device 100, and obtains the decoded sound signals in the time domain of two - channel stereo in frame units (the left - channel decoded sound signal and the right - channel decoded sound signal described later) and outputs them. The decoding device 200 can also, as Figure 3 shown by the dotted line, output the decoded sound signal in the time domain of mono (the mono - decoded sound signal described later). For example, by performing DA conversion and reproducing the decoded sound signal output by the decoding device 200 with a speaker, the decoded sound signal can be received. The decoding device 200 performs the processing of steps S210 to S260 exemplified in Figure 4 .

[0064] [Mono - decoding unit 210]

[0065] The mono - encoding CM input to the decoding device 200 is input to the mono - decoding unit 210. The mono - decoding unit 210 decodes the input mono - encoding CM in a prescribed decoding method to obtain the mono - decoded sound signal ^x M (1), ^x M (2),..., ^x M (T) and outputs it (step S210). As the prescribed decoding method, a decoding method corresponding to the encoding method used in the mono - encoding unit 160 of the corresponding encoding device 100 is used. The number of bits of the mono - encoding CM is b M .

[0066] [Stereo - decoding unit 220]

[0067] The stereo encoding CS input to the decoding device 200 is input to the stereo decoding unit 220. The stereo decoding unit 220 decodes the input stereo encoding CS in a prescribed decoding method to obtain a left-channel decoded difference signal ^y L (1), ^y L (2),..., ^y L (T), and a right-channel decoded difference signal ^y R (1), ^y R (2),..., ^y R (T) and outputs them (step S220). As the prescribed decoding method, a decoding method corresponding to the encoding method used in the stereo encoding unit 170 of the corresponding encoding device 100 is used. The total number of bits of the stereo encoding CS is b S .

[0068] [Left-channel subtraction gain decoding unit 230]

[0069] The left-channel subtraction gain encoding Cα input to the decoding device 200 is input to the left-channel subtraction gain decoding unit 230. The left-channel subtraction gain decoding unit 230 decodes the left-channel subtraction gain encoding Cα and outputs the left-channel subtraction gain α (step S230). The left-channel subtraction gain decoding unit 230 decodes the left-channel subtraction gain code Cα by a decoding method corresponding to the method used in the left-channel subtraction gain estimation unit 120 of the corresponding encoding device 100 to obtain the left-channel subtraction gain α. Regarding the method by which the left-channel subtraction gain decoding unit 230 decodes the left-channel subtraction gain encoding Cα and obtains the left-channel subtraction gain α when the left-channel subtraction gain α and the left-channel subtraction gain encoding Cα are obtained by the method based on the principle of minimizing the quantization error in the left-channel subtraction gain estimation unit 120 of the corresponding encoding device 100, it will be described later.

[0070] [Left-channel signal addition unit 240]

[0071] The mono decoded sound signal ^x M (1), ^x M (2),..., ^x M (T), the left-channel decoded difference signal ^y L (1), ^y L (2),..., ^y L (T), and the left-channel subtraction gain α output from the left-channel subtraction gain decoding unit 230 are input to the left-channel signal addition unit 240. The left-channel signal addition unit 240 adds the sampled values of the left-channel decoded difference signal ^y L(t), and the sampled values ^x of the mono decoded audio signal M The value α × ^x obtained by multiplying (t) by the left channel subtraction gain α M The value ^y obtained by adding (t) L (t) + α × ^x M The sequence of (t) is used as the left channel decoded audio signal ^x L (1), ^x L (2),..., ^x L (T) is obtained and output (step S240). That is, it is ^x L (t) = ^y L (t) + α × ^x M (t).

[0072] [Right channel subtraction gain decoding unit 250]

[0073] The right channel subtraction gain code Cβ input to the decoding device 200 is input to the right channel subtraction gain decoding unit 250. The right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ and outputs the right channel subtraction gain β (step S250). The right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ by a decoding method corresponding to the method used in the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 to obtain the right channel subtraction gain β. The method by which the right channel subtraction gain decoding unit 250 decodes the right channel subtraction gain code Cβ and obtains the right channel subtraction gain β when the right channel subtraction gain β and the right channel subtraction gain code Cβ are obtained by the method based on the principle of minimizing the quantization error in the right channel subtraction gain estimation unit 140 of the corresponding encoding device 100 will be described later.

[0074] [Right channel signal addition unit 260]

[0075] The mono decoded audio signal ^x output by the mono decoding unit 210 is input to the right channel signal addition unit 260 M (1), ^x M (2),..., ^x M (T), the right channel decoded differential signal ^y output by the stereo decoding unit 220 R (1), ^y R (2),..., ^y R (T), and the right channel subtraction gain β output by the right channel subtraction gain decoding unit 250. The right channel signal addition unit 260 adds, for each corresponding sample t, the sampled value ^y of the right channel decoded differential signal R (t), and the sampled values ^x of the mono decoded audio signal M The value β × ^x obtained by multiplying (t) by the right channel subtraction gain βM The value obtained by adding (t) ^y R (t) + β × ^x M The sequence of (t) as the right-channel decoded sound signal ^x R (1), ^x R (2),..., ^x R (T) is obtained and output (step S260). That is, it is ^x R (t) = ^y R (t) + β × ^x M (t).

[0076] [Principle of minimizing quantization error]

[0077] Hereinafter, the principle of minimizing quantization error will be described. When the left-channel differential signal and the right-channel differential signal input to the stereo encoding unit 170 are combined and encoded in one encoding method, the number of bits b used in the encoding of the left-channel differential signal L and the number of bits b used in the encoding of the right-channel differential signal R may not be clearly determined, but hereinafter, the number of bits used in the encoding of the left-channel differential signal is b L , and the number of bits used in the encoding of the right-channel differential signal is b R . In addition, hereinafter, mainly the left channel will be described, but the same applies to the right channel.

[0078] The above-described encoding device 100 uses b L bits to subtract from each sampled value of the input sound signal x from the left channel L (1), x L (2),..., x L (T) the value obtained by multiplying each sampled value of the downmix signal x M (1), x M (2),..., x M (T) by the left-channel subtraction gain α, and encodes the resulting left-channel differential signal y L (1), y L (2),..., y L (T), and encodes the downmix signal x M using b M (1), x M (2),..., x M (T). In addition, the above-described decoding device 200 decodes the left-channel decoded differential signal ^y according to the encoding of b L bits L (1), ^y L (2),..., ^y L(hereinafter, also referred to as "quantized left channel difference signal") is decoded, and according to the b M bit encoding, the monaural decoded sound signal ^x M (1), ^x M (2),..., ^x M (T) (hereinafter, also referred to as "quantized mixdown signal") is decoded, and then the quantized mixdown signal ^x obtained by decoding M (1), ^x M (2),..., ^x M (T) is multiplied by the left channel gain α for each sample value, and the obtained value is added to the quantized left channel difference signal ^y obtained by decoding L (1), ^y L (2),..., ^y L (T) for each sample value, so as to obtain the left channel decoded sound signal ^x as the decoded sound signal of the left channel L (1), ^x L (2),..., ^x L (T). The encoding device 100 and the decoding device 200 should be designed to reduce the energy of the quantization error of the left channel decoded sound signal obtained in the above processing.

[0079] In most cases, the energy of the quantization error (hereinafter, for convenience, referred to as "quantization error due to encoding") of the decoded signal obtained by encoding / decoding the input signal is approximately proportional to the energy of the input signal, and tends to decrease exponentially with respect to the value of the number of bits per sample used in the encoding. Therefore, the average energy per sample of the quantization error generated by the encoding of the left channel difference signal uses a positive number σ L 2 can be estimated as shown in the following formula (1-0-1), and the average energy per sample of the quantization error generated by the encoding of the mixdown signal uses a positive number σ M 2 can be estimated as shown in the following formula (1-0-2).

[0080] [Mathematical formula 1]

[0081]

[0082] [Mathematical formula 2]

[0083]

[0084] Here, let the input sound signal x of the left channel L (1), x L (2),..., x L (T) and the mixdown signal xM (1), x M (2),..., x M (T) The values of each sampled value that are approximately regarded as the same sequence. For example, the input sound signal x of the left channel L (1), x L (2),..., x L (T) and the input signal x of the right channel R (1), x R (2),..., x R (T) are equivalent to the situation where the sound emitted from a sound source at an equal distance from two microphones is picked up in an environment with a lot of background noise or reverberation. Under this condition, the left-channel differential signal y L (1), y L (2),..., y L (T) Each sampled value is equivalent to the value obtained by multiplying each sampled value of the downmixed signal x M (1), x M (2),..., x M (T) by (1 - α). Therefore, the energy of the left-channel differential signal is represented by (1 - α) times the energy of the downmixed signal. So the above-mentioned σ 2 times, so the above-mentioned σ L 2 can use the above-mentioned σ M 2 can be replaced by (1 - α) 2 × σ M 2 , so the average energy per sample of the quantization error generated by encoding the left-channel differential signal can be estimated as shown in the following formula (1-1).

[0085] [Mathematical formula 3]

[0086]

[0087] In addition, the average energy per sample of the quantization error of the signal obtained by adding the quantized left-channel differential signal in the decoding device, that is, the average energy per sample of the quantization error of the sequence of values obtained by multiplying each sampled value of the quantized downmixed signal obtained by decoding by the left-channel subtraction gain α, can be estimated as shown in the following formula (1-2).

[0088] [Mathematical formula 4]

[0089]

[0090] If it is assumed that the quantization error generated due to the encoding of the left-channel differential signal and the quantization error of the sequence of values obtained by multiplying each sampled value of the quantized downmix signal obtained by decoding by the left-channel subtraction gain α are not correlated with each other, the average energy of each sample of the quantization error of the decoded sound signal of the left channel is estimated by the sum of Equation (1-1) and Equation (1-2). The left-channel subtraction gain α that minimizes the energy of the quantization error of the decoded sound signal of the left channel is obtained as shown in Equation (1-3) below.

[0091] [Mathematical formula 5]

[0092]

[0093] That is to say, under the condition that the input sound signal x L (1), x L (2),..., x L (T) of the left channel and the downmix signal xx M (1), x M (2),..., x M (T) are approximately regarded as the same sequence of values for each sampled value, in order to minimize the quantization error of the decoded sound signal of the left channel, the left-channel subtraction gain estimation unit 120 can obtain the left-channel subtraction gain α through Equation (1-3). The left-channel subtraction gain α obtained from Equation (1-3) is a value greater than 0 and less than 1, and is 0.5 when the number of bits b L used for the two encodings is equal to b M . The greater the number of bits b L used for encoding the left-channel differential signal is than the number of bits b M used for encoding the downmix signal, the closer the value is to 0.5. The greater the number of bits b M used for encoding the downmix signal is than the number of bits b L used for encoding the left-channel differential signal, the closer the value is to 0.5.

[0094] The same applies to the right channel. Under the condition that the input sound signal x R (1), x R (2),..., x R (T) of the right channel and the downmix signal x M (1), x M (2),..., x M (T) are approximately regarded as the same sequence of values for each sampled value, in order to minimize the quantization error of the decoded sound signal of the right channel, the right-channel subtraction gain estimation unit 140 can obtain the right-channel subtraction gain β through the following Equation (1-3-2).

[0095] [Mathematical formula 6]

[0096]

[0097] The right-channel subtraction gain β obtained by Equation (1-3-2) is a value greater than 0 and less than 1. The number of bits used in the two encodings, i.e., b R is equal to b M is 0.5. The number of bits b R used to encode the right-channel differential signal is M more than the number of bits b M used to encode the downmix signal, and the closer it is to 0.5. The number of bits b R used to encode the downmix signal is

[0098] Next, the principle of minimizing the energy of the quantization error of the decoded sound signal of the left channel is described, including the case where the input sound signal x L (1), x L (2),..., x L (T) of the left channel and the downmix signal x M (1), x M (2),..., x M (T) cannot be regarded as the same sequence.

[0099] The input sound signal x L (1), x L (2),..., x L (T) of the left channel and the downmix signal x M (1), x M (2),..., x M (T) have a normalized inner product value r L expressed by the following Equation (1-4).

[0100] [Mathematical formula 7]

[0101]

[0102] The normalized inner product value r L obtained by Equation (1-4) is a real value, which is M the sequence of sampled values obtained by multiplying each sampled value of the downmix signal x M (1), x M (2),..., x L (T) by the real value r L ' to obtain r M '×x L '×x M(2),...,r L '×x M The sequence x obtained by taking the difference between the sequence of the obtained sampled values and each sampled value of the input sound signal of the left channel at (T). L (1)-r L '×x M (1),x L (2)-r L '×x M (2),...,x L (T)-r L '×x M The real value r for which the energy of (T) becomes minimum. L The same value.

[0103] The input sound signal x of the left channel L (1),x L (2),...,x L (T) can be decomposed for each sampling number t as x L (t) = r L ×x M (t)+(x L (t)-r L ×x M (t)). Here, when the sequence composed of each value of x L (t)-r L ×x M (t) is set as the orthogonal signal x L ’(1),x L ’(2),...,x L ’(T), according to this decomposition, each sampled value y of the left channel difference signal L (t) = x L (t)-αx M (t) is equivalent to the value obtained by multiplying each sampled value x of the downmix signal x M (1),x M (2),...,x M (T) by the normalized inner product value r M and (r L - α) using the left channel subtraction gain α, and adding the orthogonal signal's each sampled value x L - α)×x L (t) and x M ’(t), i.e., (r L - α)×x L - α)×x M (t)+x L ’(t). The orthogonal signal x L ’(1),x L ’(2),...,xL ’(T) represents with respect to the downmixed signal x M (1), x M (2),..., x M (T) has the property of orthogonality, i.e., the inner product is 0. Therefore, the energy of the left-channel difference signal can be expressed as the sum of the energy obtained by multiplying the energy of the downmixed signal by (r L -α) 2 times and the energy of the orthogonal signal. Thus, the average energy per sample of the quantization error generated by encoding the left-channel difference signal with b L bits is a positive number σ 2 and can be estimated as shown in Equation (1-5) below.

[0104] [Mathematical formula 8]

[0105]

[0106] If it is assumed that the quantization error generated by encoding the left-channel difference signal and the quantization error of the sequence obtained by multiplying each sampled value of the quantized downmixed signal obtained by decoding by the left-channel subtraction gain α are not correlated with each other, then the average energy per sample of the quantization error of the decoded sound signal of the left channel is estimated by the sum of Equation (1-5) and Equation (1-2). The left-channel subtraction gain α that minimizes the energy of the quantization error of the decoded sound signal of the left channel is obtained as shown in Equation (1-6) below.

[0107] [Mathematical formula 9]

[0108]

[0109] That is to say, in order to minimize the quantization error of the decoded sound signal of the left channel, the left-channel subtraction gain estimation unit 120 only needs to obtain the left-channel subtraction gain α through Equation (1-6). That is to say, if the principle of minimizing the energy of this quantization error is considered, the left-channel subtraction gain α should use the value obtained by multiplying the normalized inner product value r L by a correction coefficient, and this correction coefficient is a value determined by the number of bits used for encoding, i.e., b L and b M The correction coefficient is a value greater than 0 and less than 1. It is 0.5 when the number of bits b L used for encoding the left-channel difference signal and the number of bits b M used for encoding the downmixed signal are the same. The more the number of bits b L used for encoding the left-channel difference signal is than the number of bits b M used for encoding the downmixed signal, the closer it is to 0.5. The number of bits b used for encoding the left-channel difference signalL is smaller than the number of bits b used for encoding the downmixed signal M the closer it is to 0.5.

[0110] The same applies to the right channel. In order to minimize the quantization error of the decoded sound signal of the right channel, the right-channel subtraction gain estimation unit 140 can obtain the right-channel subtraction gain β by the following formula (1-6-2).

[0111] [Mathematical formula 10]

[0112]

[0113] Here, r R is the normalized inner product value of the input sound signal x R (1), x R (2),..., x R (T) of the right channel and the downmixed signal x M (1), x M (2),..., x M (T), and is represented by the following formula (1-4-2).

[0114] [Mathematical formula 11]

[0115]

[0116] That is, if the principle of minimizing the energy of the quantization error is considered, the right-channel subtraction gain β should use the value obtained by multiplying the normalized inner product value r R by a correction coefficient, and this correction coefficient is determined by the number of bits for encoding, that is, b R and b M . This correction coefficient is a value greater than 0 and less than 1. The number of bits b R used for encoding the right-channel difference signal is smaller than the number of bits b M used for encoding the downmixed signal, the closer it is to 0.5. The number of bits for encoding the right-channel difference signal is smaller than the number of bits for encoding the downmixed signal, the closer it is to 0.5.

[0117] 〔Estimation and Decoding of Subtraction Gain Based on the Principle of Minimizing Quantization Error〕

[0118] A specific example of the estimation and decoding of the subtraction gain based on the principle of minimizing the above quantization error will be described. In each example, the left-channel subtraction gain estimation unit 120 and the right-channel subtraction gain estimation unit 140 that estimate the subtraction gain in the encoding device 100, and the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250 that decode the subtraction gain in the decoding device 200 will be described.

[0119] [[Example 1]]

[0120] Example 1 is based on the following principles: including the principle of minimizing the energy of the quantization error of the decoded sound signal of the left channel, excluding the cases where the input sound signals x L (1), x L (2),..., x L (T) and the downmix signal x M (1), x M (2),..., x M (T) are not regarded as the same series; and including the principle of minimizing the energy of the quantization error of the decoded sound signal of the right channel, excluding the cases where the input sound signals x R (1), x R (2),..., x R (T) and the downmix signal x M (1), x M (2),..., x M (T) are not regarded as the same series.

[0121] [[Left Channel Subtraction Gain Estimation Unit 120]]

[0122] In the left channel subtraction gain estimation unit 120, multiple sets of (Group A, a = 1,..., A) are pre-stored for the candidate α cand (a) of the left channel subtraction gain and the code Cα cand (a) corresponding to this candidate. The left channel subtraction gain estimation unit 120 performs the following steps S120-11 to S120-14 as shown in Figure 5 Figure.

[0123] The left channel subtraction gain estimation unit 120 first obtains the normalized inner product value r of the input sound signal of the left channel for the downmix signal according to the input left channel input sound signals x L (1), x L (2),..., x L (T) and the downmix signal x M (1), x M (2),..., x M (T) through Equation (1-4) (step S120-11). Then, the left channel subtraction gain estimation unit 120 uses the number of bits b of the code for the left channel difference signals y L (1), y L (2),..., y L (2),..., y L (T) in the stereo encoding unit 170 L, the number of bits b for encoding the downmixed signals x M (1), x M (2),..., x M (T), and the number of samples T per frame, the left-channel correction coefficient c is obtained by the following formula (1-7) M (Step S120-12). L (Step S120-12).

[0124] [Mathematical formula 12]

[0125]

[0126] Next, the left-channel subtraction gain estimation unit 120 obtains the normalized inner product value r obtained in step S120-11 L and the left-channel correction coefficient c obtained in step S120-12 L multiplied to obtain a value (step S120-13). Then, the left-channel subtraction gain estimation unit 120 obtains the candidate α of the stored left-channel subtraction gain cand (1),..., α cand (A) that is closest to the multiplication value c obtained in step S120-13 L ×r L as the left-channel subtraction gain α, and obtains the encoding Cα corresponding to the left-channel subtraction gain α in the stored encoding Cα L ×r L as the left-channel subtraction gain encoding Cα (step S120-14). cand (1),..., Cα cand (A) as the left-channel subtraction gain encoding Cα (step S120-14).

[0127] In addition, when the number of bits b used in the encoding of the left-channel differential signal y L (1), y L (2),..., y L (T) is not explicitly determined, half of the number of bits b of the stereo encoding CS output by the stereo encoding unit 170 L is used as the number of bits b s (i.e., b s / 2). In addition, the left-channel correction coefficient c L may not be the value obtained by formula (1-7) itself, but a value greater than 0 and less than 1, and is the following value: for the left-channel differential signal y L (1), y L (2),..., y L (2),..., y L (T) in the encoding of the number of bits b LFor the downmixed signal x M (1), x M (2),..., x M (T), the number of bits b M of the encoding is 0.5 when they are the same. When the number of bits b L is more than the number of bits b M , it is closer to 0 than 0.5. When the number of bits b L is less than the number of bits b M , it is closer to 1 than 0.5. The same applies to each of the following examples.

[0128] [[Right channel subtraction gain estimation unit 140]]

[0129] In the right channel subtraction gain estimation unit 140, multiple sets (B sets, b = 1,..., B) of candidates β cand (b) of the right channel subtraction gain and the encoding Cβ cand (b) corresponding to the candidates are pre-stored. The right channel subtraction gain estimation unit 140 performs the following steps S140-11 to S140-14 as Figure 5 shown.

[0130] The right channel subtraction gain estimation unit 140 first obtains the normalized inner product value r R (1), x R (2),..., x R (T) of the input sound signal of the right channel for the downmixed signal and the downmixed signal x M (1), x M (2),..., x M (T) through Equation (1-4-2) (step S140-11). Then, the right channel subtraction gain estimation unit 140 uses the number of bits b R of the encoding for the right channel difference signal y R (1), y R (2),..., y R (T) in the stereo encoding unit 170, the number of bits b R of the encoding for the downmixed signal x M (1), x M (2),..., x M (T) in the mono encoding unit 160, and the number of samples per frame T to obtain the right channel correction coefficient c M through the following Equation (1-7-2) (step S140-12). R (step S140-12).

[0131] [Mathematical formula 13]

[0132]

[0133] Next, the right-channel subtraction gain estimation unit 140 obtains the normalized inner product value r obtained in step S140-11 R multiplied by the right-channel correction coefficient c obtained in step S140-12 R to obtain a value (step S140-13). Next, the right-channel subtraction gain estimation unit 140 obtains the candidate β closest to the stored right-channel subtraction gain cand (1),...,β cand (B) of the multiplication value c obtained in step S140-13 R ×r R as the right-channel subtraction gain β, and obtains the code corresponding to the right-channel subtraction gain β in the stored coding Cβ R ×r R as the right-channel subtraction gain Cβ (step S140-14). cand (1),...,Cβ cand (B).

[0134] In addition, when the number of bits b used in the coding of the right-channel difference signal y R (1),y R (2),...,y R (T) in the stereo coding unit 170 is not explicitly determined, half of the number of bits b of the stereo coding CS output by the stereo coding unit 170 R is used as the number of bits b s (i.e., b s / 2). In addition, the right-channel correction coefficient c R may not be the value obtained from Equation (1-7-2) itself, but a value greater than 0 and less than 1, and is as follows: in the right-channel difference signal y R used for the coding of(1),y R (2),...,y R (T) when the number of bits b R is the same as the number of bits b R used for the coding of the downmix signal x M (1),x M (2),...,x M (T)(T), it is 0.5. The more the number of bits b M is than the number of bits b R , the closer it is to 0 than 0.5. The fewer the number of bits b M is than the number of bits b R , the closer it is to 1 than 0.5. The same applies to each of the examples described later. M ​

[0135] [[Left - channel subtraction gain decoding unit 230]]

[0136] In the left - channel subtraction gain decoding unit 230, similar to the part stored in the left - channel subtraction gain estimation unit 120 of the corresponding encoding device 100, multiple sets (Group A, a = 1, …, A) of left - channel subtraction gain candidates α cand (a) and the corresponding encoding Cα cand (a) are stored in advance. The left - channel subtraction gain decoding unit 230 obtains the left - channel subtraction gain candidate corresponding to the input left - channel subtraction gain encoding Cα cand (1), …, Cα cand (A) as the left - channel subtraction gain α (step S230 - 11).

[0137] [[Right - channel subtraction gain decoding unit 250]]

[0138] In the right - channel subtraction gain decoding unit 250, similar to the part stored in the right - channel subtraction gain estimation unit 140 of the corresponding encoding device 100, multiple sets (Group B, b = 1, …, B) of right - channel subtraction gain candidates β cand (b) and the corresponding encoding Cβ cand (b) are stored in advance. The right - channel subtraction gain decoding unit 250 obtains the right - channel subtraction gain candidate corresponding to the input right - channel subtraction gain encoding Cβ cand (1), …, Cβ cand (B) as the right - channel subtraction gain β (step S250 - 11).

[0139] In addition, the same subtraction gain candidates and encodings can be used for both the left - channel and the right - channel. Let the above - mentioned A and B be the same value. It is also possible to make the left - channel subtraction gain candidates α cand (a) and the corresponding encoding Cα cand (a) stored in the left - channel subtraction gain estimation unit 120 and the left - channel subtraction gain decoding unit 230, and the right - channel subtraction gain candidates β cand (b) and the corresponding encoding Cβ cand (b) stored in the right - channel subtraction gain estimation unit 140 and the right - channel subtraction gain decoding unit 250 the same.

[0140] [[Modification example of Example 1]]

[0141] The number of bits b used for encoding the left - channel differential signal in the encoding device 100 Lis the number of bits used for decoding the left - channel differential signal in the decoding device 200, and the number of bits b used for encoding the down - mixed signal in the encoding device 100 M The value is the number of bits used for decoding the down - mixed signal in the decoding device 200, so the correction coefficient c L Even in the encoding device 100 and the decoding device 200, the same value can be calculated. Therefore, the normalized inner - product value r L can be used as the object of encoding and decoding, and the quantized value ^r of the normalized inner - product value in the encoding device 100 and the decoding device 200 L is multiplied by the correction coefficient c L , to obtain the left - channel subtraction gain α. The same applies to the right channel. This method will be described as a modified example of Example 1.

[0142] 〔〔Left - channel subtraction gain estimation unit 120〕〕

[0143] In the left - channel subtraction gain estimation unit 120, multiple sets (Group A, a = 1, …, A) of candidates r of the normalized inner - product values of the left - channel Lcand (a) and the codes Cα cand (a) corresponding to the candidates are pre - stored. As Figure 6 shown, the left - channel subtraction gain estimation unit 120 performs step S120 - 11 and step S120 - 12 described in Example 1, and the following step S120 - 15 and step S120 - 16.

[0144] The left - channel subtraction gain estimation unit 120 first, in the same way as step S120 - 11 of the left - channel subtraction gain estimation unit 120 in Example 1, according to the input left - channel input audio signal x L (1), x L (2),..., x L (T) and the down - mixed signal x M (1), x M (2),..., x M (T), obtains the normalized inner - product value r of the input audio signal of the left - channel with respect to the down - mixed signal through formula (1 - 4) L (step S120 - 11). Then, the left - channel subtraction gain estimation unit 120 obtains the candidate r of the normalized inner - product value of the left - channel stored Lcand (1),..., r Lcand (A) that is closest to the normalized inner - product value r L obtained through step S120 - 11 (the quantized value of the normalized inner - product value r L ) ^r L , and obtains the code Cα cand (1),..., Cαcand The closest candidate in (A)^r L The corresponding code is used as the left-channel subtraction gain code Cα (step S120-15). In addition, similar to step S120-12 of the left-channel subtraction gain estimation unit 120 in Example 1, the left-channel subtraction gain estimation unit 120 uses the number of bits b of the codes for the left-channel differential signals y L (1), y L (2),..., y L (T) in the stereo encoding unit 170 L , the number of bits b of the codes for the downmix signals x M (1), x M (2),..., x M (T) in the mono encoding unit 160 M , and the number of samples per frame T, and obtains the left-channel correction coefficient c through Equation (1-7) (step S120-12). Next, the left-channel subtraction gain estimation unit 120 obtains the quantization value of the normalized inner product value obtained in step S120-15^r L and multiplies it by the left-channel correction coefficient c obtained in step S120-12 L to obtain the left-channel subtraction gain α (step S120-16). L

[0145] [[Right-channel subtraction gain estimation unit 140]]

[0146] In the right-channel subtraction gain estimation unit 140, multiple sets (Group B, b = 1,..., B) of candidates r of the normalized inner product values of the right channel and the codes Cβ Rcand (b) corresponding to the candidates are pre-stored in a group. As cand shown, the right-channel subtraction gain estimation unit 140 performs steps S140-11 and S140-12 described in Example 1, and the following steps S140-15 and S140-16. Figure 6

[0147] The right-channel subtraction gain estimation unit 140 first, similar to step S140-11 of the right-channel subtraction gain estimation unit 140 in Example 1, according to the input sound signal x of the right channel R (1), x R (2),..., x R (T) and the downmix signal x M (1), x M (2),..., x M (T), obtains the normalized inner product value r of the input sound signal of the right channel with respect to the downmix signal through Equation (1-4-2) R(Step S140-11). Next, the right-channel subtraction gain estimation unit 140 obtains the candidate r of the normalized inner product value of the stored right channel Rcand (1),..., r Rcand The normalized inner product value r obtained through step S140-11 in (B) R The closest candidate (quantized value of the normalized inner product value r R ) ^r R , and obtains the encoding Cβ corresponding to the stored cand (1),..., Cβ cand The closest candidate r in (B) R as the right-channel subtraction gain encoding Cβ (step S140-15). Then, the right-channel subtraction gain estimation unit 140, in the same manner as step S140-12 of the right-channel subtraction gain estimation unit 140 in Example 1, uses the number of bits b of the encoding for the right-channel difference signal y R (1), y R (2),..., y R (T) in the stereo encoding unit 170 R , the number of bits b of the encoding for the downmixed signal x M (1), x M (2),..., x M (T) in the mono encoding unit 160 M , and the number of samples T per frame, and obtains the right-channel correction coefficient c through Equation (1-7-2) R (step S140-12). Next, the right-channel subtraction gain estimation unit 140 obtains the value obtained by multiplying the quantized value ^r of the normalized inner product value obtained in step S140-15 R by the right-channel correction coefficient c obtained in step S140-12 R as the right-channel subtraction gain β (step S140-16).

[0148] 〔〔Left-channel subtraction gain decoding unit 230〕〕

[0149] In the left-channel subtraction gain decoding unit 230, in the same manner as the part stored in the left-channel subtraction gain estimation unit 120 of the corresponding encoding device 100, multiple sets (Group A, a = 1,..., A) of candidates r of the normalized inner product values of the left channel are pre-stored Lcand (a) and the group of encodings Cα cand (a) corresponding to the candidate. The left-channel subtraction gain decoding unit 230 performs the following steps S230-12 to S230-14 as shown in Figure 7 .

[0150] The left-channel subtraction gain decoding unit 230 obtains, as a decoded value ^r of the normalized inner product value of the left channel, the candidate of the normalized inner product value of the left channel corresponding to the input left-channel subtraction gain code Cα cand (1),..., Cα cand (A) (step S230-12). Further, the left-channel subtraction gain decoding unit 230 uses the number of decoded bits b L for decoding the left-channel decoding differential signal ^y L (1), ^y L (2),..., ^y L (T) in the stereo decoding unit 220 L , the number of decoded bits b M for decoding the mono audio signal ^x M (1), ^x M (2),..., ^x M (T) in the mono decoding unit 210 L , and the number of samples T per frame, and obtains the left-channel correction coefficient c through Equation (1-7) L (step S230-13). Next, the left-channel subtraction gain decoding unit 230 obtains the value obtained by multiplying the decoded value ^r of the normalized inner product value obtained in step S230-12 L by the left-channel correction coefficient c obtained in step S230-13 as the left-channel subtraction gain α (step S230-14).

[0151] In addition, when the stereo encoding CS is an encoding obtained by combining the left-channel differential encoding CL and the right-channel differential encoding CR, the number of decoded bits b L for decoding the left-channel decoding differential signal ^y L (1), ^y L (2),..., ^y L (T) in the stereo decoding unit 220 L is the number of bits of the left-channel differential encoding CL. When the number of decoded bits b L for decoding the left-channel decoding differential signal ^y L (1), ^y L (2),..., ^y s (T) in the stereo decoding unit 220 is not explicitly determined s , half of the number of bits b L of the input stereo encoding CS to the stereo decoding unit 220 (i.e., b M / 2) can be used as the number of bits b M . The number of decoded bits bM The number of decoded bits b of (T) M is the number of bits of the monaural encoding CM. The left-channel correction coefficient c L is not a value obtained from Equation (1-7) itself, but a value greater than 0 and less than 1, and is the following value: for the left-channel decoded differential signal ^y L (1), ^y L (2),..., ^y L The number of decoded bits b of (T) L and for the monaural decoded sound signal ^x M (1), ^x M (2),..., ^x M The number of decoded bits b of (T) M is 0.5 when they are the same, and the number of bits b L The more the number of bits b M is, the closer it is to 0 than 0.5, and the number of bits b L The fewer the number of bits b M is, the closer it is to 1 than 0.5.

[0152] [[Right-channel subtraction gain decoding unit 250]]

[0153] In the right-channel subtraction gain decoding unit 250, in the same way as the part stored in the right-channel subtraction gain estimation unit 140 of the corresponding encoding device 100, multiple sets (group B, b = 1,..., B) of candidates r of the normalized inner product values of the right channel are pre-stored Rcand (b) and the encoding Cβ cand (b) of the group corresponding to the candidate. The right-channel subtraction gain decoding unit 250 performs the following steps S250-12 to step S250-14 as Figure 7 shown.

[0154] The right-channel subtraction gain decoding unit 250 obtains the candidate of the normalized inner product value of the right channel corresponding to the input right-channel subtraction gain encoding Cβ among the stored encodings Cβ cand (1),..., Cβ cand (B) as the decoded value ^r of the normalized inner product value of the right channel (step S250-12). In addition, the right-channel subtraction gain decoding unit 250 uses the number of decoded bits b for the right-channel decoded differential signal ^y R (1), ^y R (2),..., ^y R (T) in the stereo decoding unit 220 R and the number of decoded bits b for the monaural decoded sound signal ^x R (1), ^x M (1), ^x M(2),...,^x M The number of decoded bits b of (T) M and the number of samples T per frame, the right-channel correction coefficient c is obtained by Equation (1-7-2) R (Step S250-13). Next, the right-channel subtraction gain decoding unit 250 obtains the decoded value ^r of the normalized inner product value obtained in Step S250-12 R and the right-channel correction coefficient c obtained in Step S250-13 R The value obtained by multiplying R is used as the right-channel subtraction gain β (Step S250-14).

[0155] In addition, when the stereo encoding CS is an encoding obtained by combining the left-channel differential encoding CL and the right-channel differential encoding CR, in the stereo decoding unit 220, for the right-channel decoded differential signal ^y R (1),^y R (2),...,^y R The number of decoded bits b of (T) R is the number of bits of the right-channel differential encoding CR. In the stereo decoding unit 220, for the right-channel decoded differential signal ^y R (1),^y R (2),...,^y R The number of decoded bits b of (T) R When not explicitly determined, half of the number of bits b of the stereo encoding CS input to the stereo decoding unit 220 s (i.e., b s / 2) is used as the number of bits b R That's it. In the mono decoding unit 210, for the mono decoded audio signal x M (1),^x M (2),...,^x M The number of decoded bits b of (T) M is the number of bits of the mono encoding CM. The right-channel correction coefficient c R is not the value obtained by Equation (1-7-2) itself, but a value greater than 0 and less than 1, and is the following value: when used for the right-channel decoded differential signal ^y R (1),^y R (2),...,^y R The number of decoded bits b of (T) R and the number of bits b for mono decoding the sound signal ^x M (1),^x M (2),...,^x M The number of decoded bits b of (T) M are the same, it is 0.5, the number of bits b RBit number b M The larger the bit number b is, the closer it is to 0 compared to 0.5. R Compared to the bit number b M The smaller the bit number b is, the closer it is to 1 compared to 0.5.

[0156] In addition, it is only necessary to use the same normalized inner product value candidates and perform encoding for the left and right channels. Assuming that the above A and B are the same values, it is also possible to make the candidate r of the normalized inner product value of the left channel stored in the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 Lcand (a) And the code Cα corresponding to this candidate cand (a) Group, the candidate r of the normalized inner product value of the right channel stored in the right channel subtraction gain estimation unit 140 and the right channel subtraction gain decoding unit 250 Rcand (b) And the code Cβ corresponding to this candidate cand (b) Group are the same.

[0157] In addition, the code Cα is substantially the code corresponding to the left channel subtraction gain α, and is called the left channel subtraction gain encoding for the purpose of making the statements match in the description of the encoding device 100 and the decoding device 200, etc. However, from the perspective of the code representing the normalized inner product value, it can also be called the left channel inner product encoding, etc. Similarly, the code Cβ can also be called the right channel inner product encoding, etc.

[0158] [[Example 2]]

[0159] An example in which the value obtained by also considering the input value of the past frame is used as the normalized inner product value will be described as Example 2. Example 2 does not strictly guarantee the optimality within the frame, that is, the minimization of the energy of the quantization error of the decoded sound signal of the left channel and the minimization of the energy of the quantization error of the decoded sound signal of the right channel, but reduces the sharp changes between frames of the left channel subtraction gain α and the sharp changes between frames of the right channel subtraction gain β, and reduces the noise generated in the decoded sound signal due to this change. That is, in addition to reducing the energy of the quantization error of the decoded sound signal, Example 2 also considers the auditory quality of the decoded sound signal.

[0160] In Example 2, the encoding side, that is, the left channel subtraction gain estimation unit 120 and the right channel subtraction gain estimation unit 140, is different from Example 1, but the decoding side, that is, the left channel subtraction gain decoding unit 230 and the right channel subtraction gain decoding unit 250, is the same as Example 1. Hereinafter, the differences between Example 2 and Example 1 will be mainly described.

[0161] [[Left Channel Subtraction Gain Estimation Unit 120]]

[0162] As Figure 8As shown, the left-channel subtraction gain estimation unit 120 performs the following steps S120-111 to step SS120-113, and steps S120-12 to step SS120-14 described in Example 1.

[0163] The left-channel subtraction gain estimation unit 120 first uses the input sound signal x of the left channel L (1), x L (2),..., x L (T), the input mixed-down signal x M (1), x M (2),..., x M (T), and the inner product value E L (-1) used in the previous frame, and obtains the inner product value E L (0) for the current frame through the following formula (1-8) (step S120-111).

[0164] [Mathematical formula 14]

[0165]

[0166] Here, ε L is a predetermined value greater than 0 and less than 1, and is pre-stored in the left-channel subtraction gain estimation unit 120. In addition, the left-channel subtraction gain estimation unit 120 stores the obtained inner product value E L (0) as the "inner product value E L (-1) used in the previous frame" for use in the next frame in the left-channel subtraction gain estimation unit 120.

[0167] The left-channel subtraction gain estimation unit 120 also uses the input mixed-down signal x M (1), x M (2),..., x M (T) and the energy E of the mixed-down signal used in the previous frame M (-1), and obtains the energy E of the mixed-down signal for the current frame through the following formula (1-9) (step S120-112). M (0)

[0168] [Mathematical formula 15]

[0169]

[0170] Here, ε M is a predetermined value greater than 0 and less than 1, and is pre-stored in the left-channel subtraction gain estimation unit 120. In addition, the left-channel subtraction gain estimation unit 120 stores the obtained energy E of the mixed-down signal M(0) As the energy E M (-1) of the downmixed signal used in the previous frame, it is used in the next frame and stored in the left channel subtraction gain estimation unit 120.

[0171] Next, the left channel subtraction gain estimation unit 120 uses the inner product value E L (0) used in the current frame obtained in step S120-111 and the energy E M (0) of the downmixed signal used in the current frame obtained in step S120-112, and obtains a normalized inner product value r L (step S120-113) through the following formula (1-10).

[0172] [Mathematical formula 16]

[0173] r L = E L (0) / E M (0)…(1-10)

[0174] The left channel subtraction gain estimation unit 120 also performs step S120-12. Next, instead of the normalized inner product value r L obtained in step S120-11, it uses the normalized inner product value r L obtained in the above step S120-113 to perform step S120-13, and then performs step S120-14.

[0175] In addition, the closer the above ε L and ε M are to 1, the more likely it is that the influence of the input sound signal and the downmixed signal of the left channel of the past frame is included in the normalized inner product value r L , and the variation between frames of the left channel subtraction gain α obtained through the normalized inner product value r L becomes smaller. L [[Right channel subtraction gain estimation unit 140]]

[0176] As

[0177] shown, the right channel subtraction gain estimation unit 140 performs the following steps S140-111 to S140-113, and steps S140-12 to S140-14 described in Example 1. Figure 8 The right channel subtraction gain estimation unit 140 first uses the input right channel input sound signals x

[0178] (1), x R (2),..., x R (2),..., x R (T), the input downmixed signal x M(1), x M (2),..., x M (T), and the inner product value E used in the previous frame R (-1), and obtains the inner product value E used in the current frame through the following formula (1-8-2) R (0) (step S140-111).

[0179] [Mathematical formula 17]

[0180]

[0181] Here, ε R is a predetermined value greater than 0 and less than 1, and is pre-stored in the right-channel subtraction gain estimation unit 140. In addition, the right-channel subtraction gain estimation unit 140 stores the obtained inner product value E R (0) as the "inner product value E used in the previous frame R (-1)" for use in the next frame in the right-channel subtraction gain estimation unit 140.

[0182] The right-channel subtraction gain estimation unit 140 also uses the input downmixed signal x M (1), x M (2),..., x M (T) and the energy E of the downmixed signal used in the previous frame M (-1), and obtains the energy E of the downmixed signal used in the current frame through formula (1-9) M (0) (step S140-112). The right-channel subtraction gain estimation unit 140 stores the obtained energy E of the downmixed signal M (0) as the "energy E of the downmixed signal used in the previous frame M (-1)" for use in the next frame. In addition, the left-channel subtraction gain estimation unit 120 also obtains the energy E of the downmixed signal used in the current frame through formula (1-9) M (0), so only one of the step S120-112 performed by the left-channel subtraction gain estimation unit 120 and the step S140-112 performed by the right-channel subtraction gain estimation unit 140 may be performed.

[0183] Next, the right-channel subtraction gain estimation unit 140 uses the inner product value E R (0) used in the current frame obtained in step S140-111 and the energy E M (0) of the downmixed signal used in the current frame obtained in step S140-112, and obtains the normalized inner product value r through the following formula (1-10-2) R(Step S140-113).

[0184] [Mathematical formula 18]

[0185] r R = E R (0) / E M (0)…(1-10-2)

[0186] The right-channel subtraction gain estimation unit 140 also performs step S140-12. Next, instead of the normalized inner product value r R obtained in step S140-11, the normalized inner product value r R obtained in the above step S140-113 is used to perform step S140-13, and then step S140-14 is performed.

[0187] In addition, the above ε R and ε M The closer to 1, the more likely the normalized inner product value r R is to include the influence of the input sound signal and the downmix signal of the right channel of the past frame. The normalized inner product value r R , and the frame-to-frame variation of the right-channel subtraction gain β obtained through the normalized inner product value r R is smaller.

[0188] [[Example 2 Variation]]

[0189] Regarding Example 2, the same variations as those of the variation of Example 1 with respect to Example 1 can also be made. This method will be described as a variation of Example 2. The variation of Example 2 is different from the variation of Example 1 on the encoding side, that is, the left-channel subtraction gain estimation unit 120 and the right-channel subtraction gain estimation unit 140, but the decoding side, that is, the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250, is the same as the variation of Example 1. The points different from the variation of Example 1 of the variation of Example 2 are the same as those of Example 2. Therefore, hereinafter, the variation of Example 2 will be described with appropriate reference to the variation of Example 1 and Example 2.

[0190] [[Left-Channel Subtraction Gain Estimation Unit 120]]

[0191] In the left-channel subtraction gain estimation unit 120, similar to the left-channel subtraction gain estimation unit 120 of the variation of Example 1, multiple sets (Group A, a = 1,…, A) of candidates r Lcand of the normalized inner product values of the left channel and the codes Cα cand (a) corresponding to the candidates are pre-stored. As Figure 9As shown, the left - channel subtraction gain estimation unit 120 performs the same steps S120 - 111 to S120 - 113 as in Example 2, the same steps S120 - 12, S120 - 15, and S120 - 16 as in the modified example of Example 1. Specifically, as described below.

[0192] The left - channel subtraction gain estimation unit 120 first uses the input sound signal x of the left - channel L (1), x L (2),..., x L (T), the input down - mixed signal x M (1), x M (2),..., x M (T), and the inner - product value E L ( - 1) used in the previous frame, and obtains the inner - product value E L (0) used in the current frame through Equation (1 - 8) (step S120 - 111). The left - channel subtraction gain estimation unit 120 also uses the input down - mixed signal x M (1), x M (2),..., x M (T) and the energy E M ( - 1) of the down - mixed signal used in the previous frame, and obtains the energy E M (0) of the down - mixed signal used in the current frame through Equation (1 - 9) (step S120 - 112). Then, the left - channel subtraction gain estimation unit 120 uses the inner - product value E L (0) used in the current frame obtained in step S120 - 111 and the energy E M (0) of the down - mixed signal used in the current frame obtained in step S120 - 112, and obtains the normalized inner - product value r L (step S120 - 113) through Equation (1 - 10). Then, the left - channel subtraction gain estimation unit 120 obtains the candidate r Lcand (1),..., r Lcand (A) of the normalized inner - product value of the left - channel stored, and the candidate r L closest to the normalized inner - product value r L obtained in step S120 - 113 (quantized value of r L ), and obtains the code corresponding to this closest candidate r cand (1),..., Cα cand (A) of the stored coding Cα L as the left - channel subtraction gain coding Cα (step S120 - 15). In addition, the left - channel subtraction gain estimation unit 120 uses the left - channel differential signal y used in the stereo coding unit 170L (1), y L (2),..., y L Number of bits b of the code for (T) L and, in the mono encoding unit 160, for the downmix signal x M (1), x M (2),..., x M Number of bits b of the code for (T) M and the number of samples T per frame, obtain the left-channel correction coefficient c through Equation (1-7) L (Step S120-12). Next, the left-channel subtraction gain estimation unit 120 obtains the quantization value ^r of the normalized inner product value obtained in Step S120-15 L and the left-channel correction coefficient c obtained in Step S120-12 L and multiplies them to obtain the left-channel subtraction gain α (Step S120-16).

[0193] 〔〔Right-channel subtraction gain estimation unit 140〕〕

[0194] In the right-channel subtraction gain estimation unit 140, similar to the right-channel subtraction gain estimation unit 140 of the modification of Example 1, multiple sets (Group B, b = 1,..., B) of candidates r for the normalized inner product values of the right channel are pre-stored Rcand (b) and the code Cβ cand (b) corresponding to the candidate. As Figure 9 shown, the right-channel subtraction gain estimation unit 140 performs the same Steps S140-111 to S140-113 as in Example 2, and the same Steps S140-12, S140-15, and S140-16 as in the modification of Example 1. Specifically, as follows

[0195] The right-channel subtraction gain estimation unit 140 first uses the input sound signal x of the right channel R (1), x R (2),..., x R (T), the input downmix signal x M (1), x M (2),..., x M (T), and the inner product value E R (-1) used in the previous frame, and obtains the inner product value E R (0) used in the current frame through Equation (1-8-2) (Step S140-111). The right-channel subtraction gain estimation unit 140 also uses the input downmix signal x M (1), x M (2),..., x M(T) and the energy E of the downmixed signal used in the previous frame M (-1), and obtain the energy E of the downmixed signal used in the current frame through Equation (1-9) M (0) (Step S140-112). Then, the right-channel subtraction gain estimation unit 140 uses the inner product value E used in the current frame obtained in Step S140-111 R (0) and the energy E of the downmixed signal used in the current frame obtained in Step S140-112 M (0), and obtain the normalized inner product value r through Equation (1-10-2) R (Step S140-113). Then, the right-channel subtraction gain estimation unit 140 obtains the candidate r of the normalized inner product value of the stored right channel Rcand (1),..., r Rcand (B) and the normalized inner product value r obtained in Step S140-113 R The closest candidate (quantized value of the normalized inner product value r R )^r R , and obtain the stored code Cβ cand (1),..., Cβ cand (B) and the code corresponding to this closest candidate r R as the right-channel subtraction gain code Cβ (Step S140-15). In addition, the right-channel subtraction gain estimation unit 140 uses the number of bits b of the code for the right-channel difference signal y R (1), y R (2),..., y R (T) in the stereo encoding unit 170 R 、the number of bits b of the code for the downmixed signal x M (1), x M (2),..., x M (T) in the mono encoding unit 160 M 、and the number of samples T per frame, and obtain the right-channel correction coefficient c through Equation (1-7-2) R (Step S140-12). Then, the right-channel subtraction gain estimation unit 140 obtains the value obtained by multiplying the quantized value ^r of the normalized inner product value obtained in Step S140-15 R by the right-channel correction coefficient c obtained in Step S140-12 R as the right-channel subtraction gain β (Step S140-16).

[0196] [[Example 3]]

[0197] For example, when the sounds such as voices and music included in the input sound signal of the left channel are different from the sounds such as voices and music included in the input sound signal of the right channel, since the components of the input sound signal of the left channel can also be included in the mix-down signal, there is the following problem: the larger the value of the left-channel subtraction gain α is used, the more the sound included in the left-channel decoded sound signal from the input sound signal of the right channel that should not be heard originally is heard, and the larger the value of the right-channel subtraction gain β is used, the more the sound included in the right-channel decoded sound signal from the input sound signal of the left channel that should not be heard originally is heard. Therefore, although the minimization of the energy of the quantization error of the decoded sound signal is not strictly guaranteed, considering the auditory quality, the left-channel subtraction gain α and the right-channel subtraction gain β can also be set to values smaller than the values obtained by Example 1. Similarly, the left-channel subtraction gain α and the right-channel subtraction gain β can also be set to values smaller than the values obtained by Example 2.

[0198] Specifically, regarding the left channel, in Example 1 and Example 2, the quantization value of the multiplication value c L of the normalized inner product value r L and the left-channel correction coefficient c L ×r L is set as the left-channel subtraction gain α. In Example 3, the quantization value of the multiplication value λ L of the normalized inner product value r L , the left-channel correction coefficient c L , and λ L ×c L ×r L which is a predetermined value greater than 0 and less than 1 is set as the left-channel subtraction gain α. Therefore, similar to Example 1 and Example 2, the multiplication value c L ×r L can also be used as the object of encoding in the left-channel subtraction gain estimation unit 120 and decoding in the left-channel subtraction gain decoding unit 230. As the left-channel subtraction gain encoding Cα represents the quantization value of the multiplication value c L ×r L , the left-channel subtraction gain estimation unit 120 and the left-channel subtraction gain decoding unit 230 multiply the quantization value of the multiplication value c L ×r L by λ L to obtain the left-channel subtraction gain α. Alternatively, the multiplication value λ L of the normalized inner product value r L , the left-channel correction coefficient c L , and the preset value λ L ×c L ×r LAs an object of encoding in the left-channel subtraction gain estimation unit 120 and decoding in the left-channel subtraction gain decoding unit 230, the left-channel subtraction gain encoding Cα represents the quantization value of the multiplication value λ L ×c L ×r L Similarly, for the right channel, in Example 1 and Example 2, the quantization value of the multiplication value c

[0199] ×r R of the normalized inner product value r R and the right-channel correction coefficient c R is set as the right-channel subtraction gain β. In contrast, in Example 3, the quantization value of the multiplication value λ R ×c R ×r R of the normalized inner product value r R the right-channel correction coefficient c R and λ which is a predetermined value greater than 0 and less than 1 R is set as the right-channel subtraction gain β. Therefore, similar to Example 1 and Example 2, the multiplication value c R ×r R can also be used as an object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250. Just as the right-channel subtraction gain encoding Cβ represents the quantization value of the multiplication value c R ×r R the right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 multiply the quantization value of the multiplication value c R ×r R by λ R to obtain the right-channel subtraction gain β. Alternatively, the multiplication value λ R ×c R ×r R of the normalized inner product value r R the left-channel correction coefficient c R and the predetermined value λ R can be used as an object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250. The right-channel subtraction gain encoding Cβ represents the quantization value of the multiplication value λ R ×c R ×r R ×r R In addition, it suffices to set λ R to the same value as λ L .

[0200] 〔〔Variant of Example 3〕〕

[0201] As described above, both the encoding device 100 and the decoding device 200 can calculate the same value. Therefore, similar to the modified examples of Example 1 and Example 2, the normalized inner product value r L represents the object of encoding in the left-channel subtraction gain estimation unit 120 and decoding in the left-channel subtraction gain decoding unit 230. The left-channel subtraction gain encoding Cα represents the normalized inner product value r L of the quantization value. The left-channel subtraction gain estimation unit 120 and the left-channel subtraction gain decoding unit 230 multiply the quantization value of the normalized inner product value r L by the left-channel correction coefficient c L and λ, which is a predetermined value greater than 0 and less than 1, L to obtain the left-channel subtraction gain α. Alternatively, the normalized inner product value r L and λ, which is a predetermined value greater than 0 and less than 1, L of the multiplication value λ L ×r L can be used as the object of encoding in the left-channel subtraction gain estimation unit 120 and decoding in the left-channel subtraction gain decoding unit 230. Similar to the left-channel subtraction gain encoding Cα representing the quantization value of the multiplication value λ L ×r L the left-channel subtraction gain estimation unit 120 and the left-channel subtraction gain decoding unit 230 multiply the quantization value of the multiplication value λ L ×r L by the left-channel correction coefficient c L to obtain the left-channel subtraction gain α.

[0202] Similarly for the right channel, the correction coefficient c R can be calculated to have the same value in both the encoding device 100 and the decoding device 200. Therefore, similar to the modified examples of Example 1 and Example 2, the normalized inner product value r R can be used as the object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250. Similar to the right-channel subtraction gain encoding Cβ representing the quantization value of the normalized inner product value r R the right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 multiply the quantization value of the normalized inner product value r R by the right-channel correction coefficient c R and λ, which is a predetermined value greater than 0 and less than 1, R to obtain the right-channel subtraction gain β. Alternatively, the normalized inner product value r R can be multiplied by λ, which is a predetermined value greater than 0 and less than 1, R to obtain the multiplication value λ R ×r RAs an object of encoding in the right-channel subtraction gain estimation unit 140 and decoding in the right-channel subtraction gain decoding unit 250, the right-channel subtraction gain encoding Cβ represents the quantization value of the multiplication value λ R ×r R The right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 multiply the quantization value of the multiplication value λ R ×r R by the right-channel correction coefficient c R to obtain the right-channel subtraction gain β.

[0203] [[Example 4]]

[0204] When the correlation between the input sound signal of the left channel and the input sound signal of the right channel is small, the problem of the auditory quality described at the beginning of Example 3 occurs. This problem hardly occurs when the correlation between the input sound signal of the left channel and the input sound signal of the right channel is large. Therefore, in Example 4, instead of the pre-determined value in Example 3, the correlation coefficient of the input sound signal of the left channel and the input sound signal of the right channel, that is, the left-right correlation coefficient γ, is used. Thus, the greater the correlation between the input sound signal of the left channel and the input sound signal of the right channel, the more the energy of the quantization error of the decoded sound signal is preferentially reduced, and the smaller the correlation between the input sound signal of the left channel and the input sound signal of the right channel, the more the deterioration of the auditory quality is preferentially suppressed.

[0205] The encoding side of Example 4 is different from those of Example 1 and Example 2, but the decoding sides, that is, the left-channel subtraction gain decoding unit 230 and the right-channel subtraction gain decoding unit 250, are the same as those of Example 1 and Example 2. Hereinafter, the differences between Example 4 and Example 1 and Example 2 will be described.

[0206] [[Left-Right Relationship Information Estimation Unit 180]]

[0207] As shown by the dashed line in Figure 1 , the encoding apparatus 100 of Example 4 further includes a left-right relationship information estimation unit 180. The input sound signal of the left channel input to the encoding apparatus 100 and the input sound signal of the right channel input to the encoding apparatus 100 are input to the left-right relationship information estimation unit 180. The left-right relationship information estimation unit 180 obtains the left-right correlation coefficient γ based on the input sound signal of the left channel and the input sound signal of the right channel and outputs it (step S180).

[0208] The left-right correlation coefficient γ is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and can also be the sampling sequence x L (1), x L (2),..., x L (T) of the input sound signal of the left channel and the sampling sequence x R (1), xR (2),...,x R The correlation coefficient γ0 of (T), or it can also be a correlation coefficient considering the time difference, for example, the correlation coefficient γ between the sampling column of the input sound signal of the left channel and the sampling column of the input sound signal of the right channel located at a position deviated by τ samplings behind this sampling column. τ 。

[0209] Let τ be the information equivalent to the difference in arrival times (so-called arrival time difference) from the sound source that mainly emits sound in this space to the microphone for the left channel and from this sound source to the microphone for the right channel when the sound signal obtained by AD-converting the sound collected by the microphone for the left channel arranged in a certain space is the input sound signal of the left channel and the sound signal obtained by AD-converting the sound collected by the microphone for the right channel arranged in this space is the input sound signal of the right channel. Hereinafter, it is called the left-right time difference. The left-right time difference τ can be obtained by any known method, or can be obtained by the method described by the left-right relationship information estimation unit 181 of the second reference method, etc. That is, the above-mentioned correlation coefficient γ τ is information equivalent to the correlation coefficient between the sound signal collected from the sound source to the microphone for the left channel and the sound signal collected from this sound source to the microphone for the right channel.

[0210] 〔〔Left channel subtraction gain estimation unit 120〕〕

[0211] The left channel subtraction gain estimation unit 120 replaces step S120-13, and obtains the normalized inner product value r obtained in step S120-11 or step S120-113 L , the left channel correction coefficient c obtained in step S120-12 L , and the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 (step S120-13”). Then, the left channel subtraction gain estimation unit 120 replaces step S120-14, and obtains the candidate α of the left channel subtraction gain stored cand (1),...,α cand The multiplication value γ×c obtained in step S120-13” in (A) L ×r L The candidate closest to (the quantization value of the multiplication value γ×c L ×r L ) as the left channel subtraction gain α, and obtains the stored encoding Cα cand (1),...,Cα cand The encoding corresponding to the left channel subtraction gain α in (A) as the left channel subtraction gain Cα (step S120-14”).

[0212] [[Right channel subtraction gain estimation unit 140]]

[0213] The right channel subtraction gain estimation unit 140 replaces step S140-13 to obtain the normalized inner product value r obtained in step S140-11 or step S140-113 R , the right channel correction coefficient c obtained in step S140-12 R , and the value obtained by multiplying the left-right correlation coefficient γ obtained in step S180 (step S140-13”). Then, the right channel subtraction gain estimation unit 140 replaces step S140-14, and the candidate β of the right channel subtraction gain stored cand (1),..., β cand (B) The multiplication value γ×c obtained in step S140-13” R ×r R The closest candidate (the quantization value of the multiplication value γ×c R ×r R ) is obtained as the right channel subtraction gain β, and the encoding Cβ corresponding to the right channel subtraction gain β stored cand (1),..., Cβ cand (B) The encoding corresponding to the right channel subtraction gain β is obtained as the right channel subtraction gain Cβ (step S140-14”).

[0214] [[Modification example of Example 4]]

[0215] As described above, both the encoding device 100 and the decoding device 200 can calculate the same value. Therefore, the multiplication value γ×r of the normalized inner product value r L and the left-right correlation coefficient γ L can be used as the object of encoding in the left channel subtraction gain estimation unit 120 and decoding in the left channel subtraction gain decoding unit 230. Just as the left channel subtraction gain encoding Cα represents the quantization value of the multiplication value γ×r L , the left channel subtraction gain estimation unit 120 and the left channel subtraction gain decoding unit 230 multiply the quantization value of the multiplication value γ×r L by the left channel correction coefficient c L to obtain the left channel subtraction gain α.

[0216] The same applies to the right channel. The correction coefficient c R can also calculate the same value in the encoding device 100 and the decoding device 200. Therefore, the multiplication value γ×r of the normalized inner product value r R and the left-right correlation coefficient γ RAs an object of encoding by the right-channel subtraction gain estimation unit 140 and decoding by the right-channel subtraction gain decoding unit 250, the right-channel subtraction gain encoding Cβ represents a multiplication value γ×r R As with the quantization value of, the right-channel subtraction gain estimation unit 140 and the right-channel subtraction gain decoding unit 250 use the multiplication value γ×r R of the quantization value and the right-channel correction coefficient c R and multiply them to obtain the right-channel subtraction gain β.

[0217] <Second reference method>

[0218] The encoding device and decoding device of the second reference method will be described.

[0219] <<Encoding device 101>>

[0220] As Figure 10 shown, the encoding device 101 of the second reference method includes: a downmixing unit 110, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a mono encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimation unit 181, and a time shift unit 191. The encoding device 101 of the second reference method is different from the encoding device 100 of the first reference method in that: it includes a left-right relationship information estimation unit 181 and a time shift unit 191; the left-channel subtraction gain estimation unit 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimation unit 140, and the right-channel signal subtraction unit 150 use the signal output by the time shift unit 191 instead of the signal output by the downmixing unit 110; and in addition to the above-mentioned respective codes, it also outputs the left-right time difference encoding Cτ described later. The other structures and operations of the encoding device 101 of the second reference method are the same as those of the encoding device 100 of the first reference method. The encoding device 101 of the second reference method performs the processing of steps S110 to S191 exemplified as Figure 11 shown. Hereinafter, the differences between the encoding device 101 of the second reference method and the encoding device 100 of the first reference method will be described.

[0221] [Left-right relationship information estimation unit 181]

[0222] The input sound signal of the left channel input to the encoding device 101 and the input sound signal of the right channel input to the encoding device 101 are input to the left-right relationship information estimation unit 181. The left-right relationship information estimation unit 181 obtains the left-right time difference τ and the left-right time difference encoding Cτ as the encoding representing the left-right time difference τ from the input left-channel input sound signal and right-channel input sound signal and outputs them (step S181).

[0223] The left-right time difference τ is the following information: When it is assumed that the sound signal obtained by AD-converting the sound collected by the microphone for the left channel disposed in a certain space is the input sound signal of the left channel, and the sound signal obtained by AD-converting the sound collected by the microphone for the right channel disposed in this space is the input sound signal of the right channel, it is the information equivalent to the difference in arrival times (so-called arrival time difference) from the sound source that mainly emits sound in this space to the microphone for the left channel and from this sound source to the microphone for the right channel. In addition, not only the arrival time difference, but also the information on which microphone arrives earlier is included in the left-right time difference τ, and the left-right time difference τ can take positive and negative values based on either of the input sound signals. That is, the left-right time difference τ is the information indicating which of the input sound signals of the left channel and the input sound signal of the right channel contains the same sound signal. Hereinafter, when the same sound signal is included in the input sound signal of the left channel earlier than the input sound signal of the right channel, it is also referred to as the left channel leading, and when the same sound signal is included in the input sound signal of the right channel earlier than the input sound signal of the left channel, it is also referred to as the right channel leading.

[0224] The left-right time difference τ can also be obtained by any known method. For example, the left-right relationship information estimation unit 181 calculates, for each candidate sampling number τ max from τ min (for example, τ max is a positive number, τ min is a negative number) cand , a value (hereinafter referred to as the correlation value) γ cand representing the correlation magnitude between the sampling sequence of the input sound signal of the left channel and the sampling sequence of the input sound signal of the right channel at a position that is deviated backward from this sampling sequence by the candidate sampling number τ cand , and obtains the candidate sampling number τ cand for which the correlation value γ cand becomes the maximum as the left-right time difference τ. That is, in this example, when the left channel leads, the left-right time difference τ is a positive value, and when the right channel leads, the left-right time difference τ is a negative value. The absolute value of the left-right time difference τ is a value (the number of leading samples) indicating approximately how much the leading channel leads with respect to the other channel. For example, when calculating the correlation value γ cand using only the samples within a frame, when τ cand is a positive value, calculate a partial sampling sequence x R (1 + τ cand ) of the input sound signal of the right channel, x R (2 + τ cand ),..., x R (T), and a partial sampling sequence that is deviated forward from this partial sampling sequence by the candidate sampling number τ candPartial sampling column x of the input sound signal of the left channel at the position of L (1), x L (2),..., x L (T - τ cand ) as the correlation value γ by taking the absolute value of the correlation coefficient cand , when τ cand is negative, calculate the partial sampling column x of the input sound signal of the left channel L (1 - τ cand ), x L (2 - τ cand ),..., x L (T), and the partial sampling column x of the input sound signal of the right channel at the position that is deviated forward from this partial sampling column by the candidate sampling number -τ cand R (1), x R (2),..., x R (T + τ cand ) as the correlation value γ by taking the absolute value of the correlation coefficient cand That's all right. Of course, in order to calculate the correlation value γ cand , it is also possible to use one or more samplings of the past input sound signal that is continuous with the sampling column of the input sound signal of the current frame. In this case, it is only necessary to store the sampling column of the input sound signal of the past frame in a storage section (not shown) in the left - right relationship information estimation section 181 by an amount of a predetermined number of frames.

[0225] In addition, for example, instead of taking the absolute value of the correlation coefficient, the information on the phase of the signal can be used to calculate the correlation value γ as follows cand . In this example, the left - right relationship information estimation section 181 first performs Fourier transforms on the input sound signal x of the left channel L (1), x L (2),..., x L (T) and the input sound signal x of the right channel R (1), x R (2),..., x R (T) according to the following equations (3 - 1) and (3 - 2), thereby obtaining the spectra X L (k) and X R (k) for each frequency k from 0 to T - 1.

[0226] [Mathematical formula 19]

[0227]

[0228] [Mathematical formula 20]

[0229] ​

[0230] The left - right relationship information estimation unit 181 uses the obtained spectra X L (k) and X R (k), and obtains the spectrum of the phase difference at each frequency k through the following formula (3 - 3)

[0231] [Mathematical formula 21]

[0232]

[0233] By performing an inverse Fourier transform on the obtained spectrum of the phase difference, as in the following formula (3 - 4), for each candidate sampling number τ max to τ min , the phase difference signal ψ(τ cand ) is obtained cand .

[0234] [Mathematical formula 22]

[0235]

[0236] The absolute value of the obtained phase difference signal ψ(τ cand ) represents a certain correlation corresponding to the rationality of the time difference between the input sound signal x L (1), x L (2),..., x L (T) of the left channel and the input sound signal x R (1), x R (2),..., x R (T) of the right channel. Therefore, the absolute value of this phase difference signal ψ(τ cand ) with respect to each candidate sampling number τ cand is used as the correlation value γ cand . The left - right relationship information estimation unit 181 obtains the candidate sampling number τ cand for which the absolute value of this phase difference signal ψ(τ cand ), that is, the correlation value γ cand becomes the maximum, as the left - right time difference τ. Additionally, instead of directly using the absolute value of the phase difference signal ψ(τ cand ) as the correlation value γ cand , a normalized value such as the relative difference of the average of the absolute values of the phase difference signals obtained for each of a plurality of candidate sampling numbers located before and after τ cand with respect to the absolute value of the phase difference signal ψ(τ cand ) can be used. That is, for each τ cand , a previously determined positive number τ cand can also be used.range , the average value is obtained through the following formula (3-5), and the obtained average value ψ is used c (τ cand ) and the phase difference signal ψ(τ cand ), and the normalized correlation value obtained through the following formula (3-6) is used as γ cand for use.

[0237] [Mathematical formula 23]

[0238]

[0239] [Mathematical formula 24]

[0240]

[0241] In addition, the normalized correlation value obtained through formula (3-6) is a value between 0 and 1, which represents that the more reasonable τ is as the left and right time difference, the closer it is to 1, and the more unreasonable τ is as the left and right time difference, the closer it is to 0. cand As the left and right time difference, the closer it is to 1, and τ cand As the left and right time difference, the closer it is to 0, which is a value with this property.

[0242] In addition, the left and right relationship information estimation unit 181 encodes the left and right time difference τ using a specified encoding method to obtain the left and right time difference encoding Cτ, which is an encoding that can uniquely determine the left and right time difference τ. As the specified encoding method, a well-known encoding method such as scalar quantization can be used. In addition, each of the pre-determined candidate sampling numbers can be each integer value from τ max to τ min , or can include fractional values and decimal values between τ max and τ min , or can not include any integer value between τ max and τ min . In addition, it can be τ max =-τ min , or not. In addition, in the case of a special input sound signal where a certain sound channel must be prior, it can also be that both τ max and τ min are positive numbers, or both τ max and τ min are negative numbers.

[0243] In addition, when the encoding device 101 estimates the subtraction gain based on the principle of minimizing the quantization error described in Example 4 or a modified example of Example 4 in the first reference method, the left and right relationship information estimation unit 181 also outputs the correlation value between the sampling sequence of the input sound signal of the left channel and the sampling sequence of the input sound signal of the right channel at a position that is deviated backward from this sampling sequence by the left and right time difference τ, that is, for τmax to τ min each candidate sampling number τ cand the calculated correlation value γ cand The maximum value in is used as the left - right correlation coefficient γ (step S180).

[0244] [Time - shift section 191]

[0245] Input the down - mixed signal x output from the down - mixing section 110 to the time - shift section 191 M (1), x M (2),..., x M (T) and the left - right time difference τ output from the left - right relationship information estimation section 181. When the left - right time difference τ is positive (i.e., the left - right time difference τ indicates that the left channel leads), the down - mixed signal x M (1), x M (2),..., x M (T) is directly output to the left - channel subtraction gain estimation section 120 and the left - channel signal subtraction section 130 (i.e., it is determined to be used in the left - channel subtraction gain estimation section 120 and the left - channel signal subtraction section 130), and the signal x obtained by delaying the down - mixed signal by |τ| samples (the number of samples equal to the absolute value of the left - right time difference τ, the number of samples of the magnitude indicated by the left - right time difference τ) M (1 - |τ|), x M (2 - |τ|),..., x M (T - |τ|), that is, the delayed down - mixed signal x M' (1), x M' (2),..., x M' (T) is output to the right - channel subtraction gain estimation section 140 and the right - channel signal subtraction section 150 (i.e., it is determined to be used in the right - channel subtraction gain estimation section 140 and the right - channel signal subtraction section 150). When the left - right time difference τ is negative (i.e., the left - right time difference τ indicates that the right channel leads), the signal x obtained by delaying the down - mixed signal by |τ| samples M (1 - |τ|), x M (2 - |τ|),..., x M (T - |τ|), that is, the delayed down - mixed signal x M' (1), x M' (2),..., x M' (T) is output to the left - channel subtraction operation section 120 and the left - channel signal subtraction section 130 (i.e., it is determined to be used in the left - channel subtraction operation section 120 and the left - channel signal subtraction section 130), and the down - mixed signal x M (1), x M (2),..., x M(T) is directly output to the right-channel subtraction gain estimation unit 140 and the right-channel signal subtraction unit 150 (that is, it is determined to be used in the right-channel subtraction gain estimation unit 140 and the right-channel signal subtraction unit 150). In the case where the left-right time difference τ is 0 (that is, the left-right time difference τ indicates the case where neither channel leads), the downmixed signal x M (1), x M (2),..., x M (T) is directly output to the left-channel subtraction gain estimation unit 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimation unit 140, and the right-channel signal subtraction unit 150 (that is, it is determined to be used in the left-channel subtraction gain estimation unit 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimation unit 140, and the right-channel signal subtraction unit 150) (step S191). That is, for the channel with the shorter arrival time among the left channel and the right channel, the input downmixed signal is directly output to the subtraction gain estimation unit of that channel and the signal subtraction unit of that channel. For the channel with the longer arrival time among the left channel and the right channel, the signal obtained by delaying the input downmixed signal by the absolute value |τ| of the left-right time difference τ is output to the subtraction gain estimation unit of that channel and the signal subtraction unit of that channel. In addition, since the time shift unit 191 uses the downmixed signals of past frames to obtain the delayed downmixed signal, in a storage unit (not shown) within the time shift unit 191, the downmixed signals input in past frames are stored in a predetermined number of frames. In addition, when the left-channel subtraction gain estimation unit 120 and the right-channel subtraction gain estimation unit 140 do not use a method based on the principle of minimizing quantization error, but use a well-known method as exemplified in Patent Document 1 to obtain the left-channel subtraction gain α and the right-channel subtraction gain β, a unit for obtaining a local decoded signal corresponding to the mono coding CM is provided at the subsequent stage of the mono coding unit 160 of the encoding device 101 or within the mono coding unit 160. In the time shift unit 191, instead of the downmixed signal x M (1), x M (2),..., x M (T), the quantized downmixed signal ^x M (1), ^x M (2),..., ^x M (T) which is the local decoded signal of the mono coding can be used to perform the above processing. In this case, the time shift unit 191 outputs the quantized downmixed signal ^x M (1), ^x M (2),..., ^x M (T) to replace the downmixed signal x M (1), x M (2),..., x M (T), and outputs the delayed quantized downmixed signal ^xM' (1), ^x M' (2),..., ^x M' Replace the delayed mix-down signal x with (T) M' (1), x M' (2),..., x M' (T).

[0246] [Left-channel subtraction gain estimation unit 120, left-channel signal subtraction unit 130, right-channel subtraction gain estimation unit 140, right-channel signal subtraction unit 150]

[0247] The left-channel subtraction gain estimation unit 120, left-channel signal subtraction unit 130, right-channel subtraction gain estimation unit 140, and right-channel signal subtraction unit 150 use the mix-down signal x input from the time shift unit 191 M (1), x M (2),..., x M (T) or the delayed mix-down signal x M' (1), x M' (2),..., x M' To perform the same operations as described in the first reference mode, instead of the mix-down signal x output by the mix-down unit 110 M (1), x M (2),..., x M (T) (Steps S120, S130, S140, S150, S150). That is, the left-channel subtraction gain estimation unit 120, left-channel signal subtraction unit 130, right-channel subtraction gain estimation unit 140, and right-channel signal subtraction unit 150 use the mix-down signal x determined by the time shift unit 191 M (1), x M (2),..., x M (T) or the delayed mix-down signal x M' (1), x M' (2),..., x M' (T) and perform the same operations as described in the first reference mode. In addition, the time shift unit 191 outputs the quantized mix-down signal ^x M (1), ^x M (2),..., ^x M Instead of the mix-down signal x M (1), x M (2),..., x M (T), it outputs the delayed quantized mix-down signal ^x M' (1), ^x M' (2),..., ^x M' Instead of the delayed mix-down signal x M' (1), xM' (2),...,x M' In the case of (T), the left-channel subtraction gain estimation unit 120, the left-channel signal subtraction unit 130, the right-channel subtraction gain estimation unit 140, and the right-channel signal subtraction unit 150 use the quantized downmixed signal ^x M (1),^x M (2),...,^x M (T) or the delayed quantized downmixed signal ^x M' (1),^x M' (2),...,^x M' (T) to perform the above processing.

[0248] [[Decoding device 201]]

[0249] As Figure 12 shown, the decoding device 201 of the second reference method includes: a mono decoding unit 210, a stereo decoding unit 220, a left-channel subtraction gain decoding unit 230, a left-channel signal addition unit 240, a right-channel signal subtraction gain decoding unit 250, a right-channel signal addition unit 260, a left-right time difference decoding unit 271, and a time shift unit 281. The difference between the decoding device 201 of the second reference method and the decoding device 200 of the first reference method is that: in addition to inputting the above-mentioned respective codes, it is also input with the left-right time difference code Cτ described later; it includes a left-right time difference decoding unit 271 and a time shift unit 281; and the left-channel signal addition unit 240 and the right-channel signal addition unit 260 use the signal output by the time shift unit 281 instead of the signal output by the mono decoding unit 210. The other structures and operations of the decoding device 201 of the second reference method are the same as those of the decoding device 200 of the first reference method. The decoding device 201 of the second reference method performs the processing from step S210 to step S281 as Figure 13 illustrated for each frame. Hereinafter, the differences between the decoding device 201 of the second reference method and the decoding device 200 of the first reference method will be described.

[0250] [Left-right time difference decoding unit 271]

[0251] The left-right time difference code Cτ input to the decoding device 201 is input to the left-right time difference decoding unit 271. The left-right time difference decoding unit 271 decodes the left-right time difference code Cτ in a prescribed decoding method to obtain the left-right time difference τ and outputs it (step S271). As the prescribed decoding method, the decoding method corresponding to the encoding method used in the left-right relationship information estimation unit 181 of the corresponding encoding device 101 is used. The left-right time difference τ obtained by the left-right time difference decoding unit 271 is the same value as the left-right time difference τ obtained by the left-right relationship information estimation unit 181 of the corresponding encoding device 101, which is from τmax to τ min any value within the range of

[0252] [Time-shifting unit 281]

[0253] Input the mono decoded sound signal ^x output by the mono decoding unit 210 to the time-shifting unit 281 M (1), ^x M (2),..., ^x M (T) and the left-right time difference τ output by the left-right time difference decoding unit 271. When the left-right time difference τ is positive (i.e., the left-right time difference τ indicates that the left channel leads), the mono decoded sound signal ^x M (1), ^x M (2),..., ^x M (T) is directly output to the left channel signal adder 240 (i.e., determined to be used in the left channel signal adder 240), and the signal ^x obtained by delaying the mono decoded sound signal by |τ| samples M (1 - |τ|), ^x M (2 - |τ|),..., ^x M (T - |τ|), that is, the delayed mono decoded sound signal ^x M' (1), ^x M' (2),..., ^x M' (T) is output to the right channel signal adder 260 (i.e., determined to be used in the right channel signal adder 260). When the left-right time difference τ is negative (i.e., the left-right time difference τ indicates that the right channel leads), the signal ^x obtained by delaying the mono decoded sound signal by |τ| samples M (1 - |τ|), ^x M (2 - |τ|),..., ^x M (T - |τ|), that is, the delayed mono decoded sound signal ^x M' (1), ^x M' (2),..., ^x M' (T) is output to the left channel signal adder 240 (i.e., determined to be used in the left channel signal adder 240), and the mono decoded sound signal ^x M (1), ^x M (2),..., ^x M (T) is directly output to the right channel signal adder 260 (i.e., determined to be used in the right channel signal adder 260). When the left-right time difference τ is 0 (i.e., the left-right time difference τ indicates that neither channel leads), the mono decoded sound signal ^x M (1), ^x M (2),..., ^xM (T) is directly output to the left-channel signal adder 240 and the right-channel signal adder 260 (that is, it is determined to be used in the left-channel signal adder 240 and the right-channel signal adder 260) (step S281). In addition, since the time-shift unit 281 uses the mono decoded audio signal of the past frame to obtain the delayed mono decoded audio signal, in a storage unit (not shown) within the time-shift unit 281, the mono decoded audio signal input in the past frame is stored in an amount of a pre-determined number of frames.

[0254] [Left-channel signal adder 240, right-channel signal adder 260]

[0255] The left-channel signal adder 240 and the right-channel signal adder 260 use the mono decoded audio signal ^x input from the time-shift unit 281 M (1), ^x M (2),..., ^x M (T) or the delayed mono decoded audio signal ^x M' (1), ^x M' (2),..., ^x M' to perform the same operations as those described in the first reference mode, instead of the mono decoded audio signal ^x output by the mono decoder 210 M (1), ^x M (2),..., ^x M (T) (steps S240, S260). That is, the left-channel signal adder 240 and the right-channel signal adder 260 use the mono decoded audio signal ^x determined by the time-shift unit 281 M (1), ^x M (2),..., ^x M (T) or the delayed mono decoded audio signal ^x M' (1), ^x M' (2),..., ^x M' (T) to perform the same operations as those described in the first reference mode.

[0256] <First Embodiment>

[0257] A modification of the encoding device 101 of the second reference mode, which generates a downmix signal in consideration of the relationship between the input audio signal of the left channel and the input audio signal of the right channel, is the first embodiment. Hereinafter, the encoding device of the first embodiment will be described. In addition, since the encoding obtained by the encoding device of the first embodiment can be decoded by the decoding device 201 of the second reference mode, the description of the decoding device is omitted.

[0258] <<Encoding Device 102>>

[0259] As Figure 10 shown, the encoding device 102 of the first embodiment includes: a downmixing unit 112, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a mono encoding unit 160, a stereo encoding unit 170, a left-right relationship information estimation unit 182, and a time shift unit 191. The encoding device 102 of the third embodiment is different from the encoding device 101 of the second reference embodiment in that: it includes a left-right relationship information estimation unit 182 instead of the left-right relationship information estimation unit 181, and includes a downmixing unit 112 instead of the downmixing unit 110. As Figure 10 shown by the dashed line in the figure, the left-right relationship information estimation unit 182 obtains the left-right correlation coefficient γ and the leading channel information and outputs them. The output left-right correlation coefficient γ and leading channel information are input to the downmixing unit 112 and used. The other structures and operations of the encoding device 102 of the first embodiment are the same as those of the encoding device 101 of the second reference embodiment. The encoding device 102 of the first embodiment performs the processing of steps S112 to S191 as Figure 14 illustrated for each frame. Hereinafter, the differences between the encoding device 102 of the first embodiment and the encoding device 101 of the second reference embodiment will be described.

[0260] [Left-Right Relationship Information Estimation Unit 182]

[0261] The input sound signal of the left channel input to the encoding device 102 and the input sound signal of the right channel input to the encoding device 102 are input to the left-right relationship information estimation unit 182. The left-right relationship information estimation unit 182 obtains the left-right time difference τ, the left-right time difference code Cτ as the code representing the left-right time difference τ, the left-right correlation coefficient γ, and the leading channel information based on the input left-channel input sound signal and right-channel input sound signal, and outputs them (step S182). The process by which the left-right relationship information estimation unit 182 obtains the left-right time difference τ and the left-right time difference code Cτ is the same as that of the left-right relationship information estimation unit 181 of the second reference embodiment.

[0262] The left-right correlation coefficient γ is information equivalent to the correlation coefficient of the sound signal collected by the microphone for the left channel from the sound source and the sound signal collected by the microphone for the right channel from the sound source as conceived in the description part of the left-right relationship information estimation unit 181 of the second reference embodiment. The leading channel information is information equivalent to which microphone the sound from the sound source reaches earlier, is information indicating which of the input sound signals of the left channel and the input sound signals of the right channel contains the same sound signal first, and is information indicating which of the left channel and the right channel is the leading channel.

[0263] Based on the example in the description part of the left-right relationship information estimation unit 181 in the second reference mode, the left-right relationship information estimation unit 182 calculates the correlation value between the sampling column of the input sound signal of the left channel and the sampling column of the input sound signal of the right channel at a position that is deviated backward from the sampling column by the left-right time difference τ, that is, for each candidate sampling number τ max from τ min to τ cand and outputs the maximum value of the calculated correlation values γ cand as the left-right correlation coefficient γ. In addition, when the left-right time difference τ is positive, the left-right relationship information estimation unit 182 obtains and outputs the information indicating that the left channel is ahead as the ahead channel information. When the left-right time difference τ is negative, the left-right relationship information estimation unit 182 obtains and outputs the information indicating that the right channel is ahead as the ahead channel information. When the left-right time difference τ is 0, the left-right relationship information estimation unit 182 can obtain and output the information indicating that the left channel is ahead as the ahead channel information, or can obtain and output the information indicating that the right channel is ahead as the ahead channel information, or can obtain and output the information indicating that neither channel is ahead as the ahead channel information.

[0264] [Mixing unit 112]

[0265] The input sound signal of the left channel input to the encoding device 102, the input sound signal of the right channel input to the encoding device 102, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 182, and the ahead channel information output by the left-right relationship information estimation unit 182 are input to the mixing unit 112. The mixing unit 112 performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a mixed signal and outputs it so that the larger the left-right correlation coefficient γ in the mixed signal, the larger the input sound signal of the ahead channel in the input sound signals of the left channel and the right channel (step S112).

[0266] For example, if the absolute value or normalized value of the correlation coefficient is used for the correlation value as in the example in the description part of the left-right relationship information estimation unit 181 in the second reference mode, the obtained left-right correlation coefficient γ is a value between 0 and 1. Therefore, the mixing unit 112 uses the weight determined by the left-right correlation coefficient γ for each corresponding sampling number t and weights and adds the input sound signal x L (t) of the left channel and the input sound signal x R (t) of the right channel to obtain the mixed signal x M (t). Specifically, when the ahead channel information is the information indicating that the left channel is ahead, that is, when the left channel is ahead, it is set as x M (t) = ((1 + γ) / 2) × x L(t)+((1 - γ) / 2)×x R (t), when the precedence channel information indicates information indicating that the right channel takes precedence, that is, when the right channel takes precedence, it is set to x M (t) = ((1 - γ) / 2)×x L (t)+((1 + γ) / 2)×x R (t), to obtain the downmixed signal x M (t) is sufficient. The downmixed signal is obtained in this way by the downmixing unit 112. Thus, the smaller the left - right correlation coefficient γ, that is, the smaller the correlation between the input audio signal of the left channel and the input audio signal of the right channel, the closer the downmixed signal is to the signal obtained by averaging the input audio signal of the left channel and the input audio signal of the right channel. If the left - right correlation coefficient γ is larger, that is, the correlation between the input audio signal of the left channel and the input audio signal of the right channel is larger, the closer the downmixed signal is to the input audio signal of the precedence channel among the input audio signals of the left channel and the right channel.

[0267] In addition, when neither channel takes precedence, the downmixing unit 112 can average the input audio signal of the left channel and the input audio signal of the right channel and output the downmixed signal in such a way that the input audio signal of the left channel and the input audio signal of the right channel are included in the downmixed signal with the same weight. Therefore, when the precedence channel information indicates that neither channel takes precedence, for each sampling number t, the downmixed signal x L (t) and the input audio signal x R (t) of the right channel are averaged, and the resulting x M (t) = (x L (t)+x R (t)) / 2 is set as the downmixed signal x M (t).

[0268] <Second Embodiment>

[0269] For the encoding device 100 of the first reference mode, a modification can also be made to generate the downmixed signal by considering the relationship between the input audio signal of the left channel and the input audio signal of the right channel. This mode will be described as the second embodiment. In addition, since the encoding obtained by the encoding device of the second embodiment can be decoded by the decoding device 200 of the first reference mode, the description of the decoding device is omitted.

[0270] <<Encoding Device 103>>

[0271] As Figure 1As shown, the encoding device 103 of the second embodiment includes: a downmixing unit 112, a left-channel subtraction gain estimation unit 120, a left-channel signal subtraction unit 130, a right-channel subtraction gain estimation unit 140, a right-channel signal subtraction unit 150, a mono encoding unit 160, a stereo encoding unit 170, and a left-right relationship information estimation unit 183. The encoding device 103 of the second embodiment is different from the encoding device 100 of the first reference embodiment in that: it includes a downmixing unit 112 instead of the downmixing unit 110, and as shown by the dashed line in Figure 1 it includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and the leading channel information and outputs them. The output left-right correlation coefficient γ and leading channel information are input to the downmixing unit 112 and used. The other structures and operations of the encoding device 103 of the second embodiment are the same as those of the encoding device 100 of the first reference embodiment. In addition, the operation of the downmixing unit 112 of the encoding device 103 of the second embodiment is the same as the operation of the downmixing unit 112 of the encoding device 102 of the first embodiment. The encoding device 103 of the second embodiment performs the processing of Figure 15 the steps S112 to S183 illustrated. Hereinafter, the points where the encoding device 103 of the second embodiment is different from both the encoding device 100 of the first reference embodiment and the encoding device 102 of the first embodiment will be described.

[0272] [Left-Right Relationship Information Estimation Unit 183]

[0273] The input sound signal of the left channel input to the encoding device 103 and the input sound signal of the right channel input to the encoding device 103 are input to the left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and the leading channel information based on the input left-channel input sound signal and right-channel input sound signal and outputs them (step S183).

[0274] The left-right correlation coefficient γ and the leading channel information obtained and output by the left-right relationship information estimation unit 183 are the same as the information described in the first embodiment. That is, except that the left-right relationship information estimation unit 183 may not obtain the left-right time difference τ and the left-right time difference encoding Cτ and not output them, it may be the same as the left-right relationship information estimation unit 182.

[0275] For example, regarding each candidate sampling number τ max from τ min to τ cand , the left-right relationship information estimation unit 183 compares the sampling sequence of the input sound signal of the left channel with the sampling sequence of the input sound signal of the right channel at a position shifted backward by each candidate sampling number τ cand from this sampling sequence, and the correlation value γ candThe maximum value in [the above] is obtained as the left - right correlation coefficient γ and output, and τ when the correlation value is the maximum value cand In the case where τ cand is positive, information indicating that the left channel leads is obtained as the leading - channel information and output, and τ when the correlation value is the maximum value cand In the case where τ cand is negative, information indicating that the right channel leads is obtained and output as the leading - channel information. When τ is the maximum value of the correlation value cand When τ cand is 0, the left - right relationship information estimation unit 183 can obtain and output information indicating that the left channel leads as the leading - channel information, can also obtain and output information indicating that the right channel leads as the leading - channel information, and can also obtain and output information indicating that neither channel leads as the leading - channel information.

[0276] <Third Embodiment>

[0277] For an encoding device that performs stereo encoding on the input audio signals of each channel instead of the differential signals of each channel, a structure that obtains a down - mix signal by considering the relationship between the input audio signal of the left channel and the input audio signal of the right channel can also be adopted, and this method will be described as the third embodiment.

[0278] <<Encoding Device 104>>

[0279] As Figure 16 shown, the encoding device 104 of the third embodiment includes: a left - right relationship information estimation unit 183, a down - mix unit 112, a mono - encoding unit 160, and a stereo - encoding unit 174. The encoding device 104 of the third embodiment performs the processing of step S183, step S112, step S160, and step S174 as exemplified for each frame. Hereinafter, the encoding device 104 of the third embodiment will be described with appropriate reference to the description of the second embodiment. Figure 17

[0280] [Left - Right Relationship Information Estimation Unit 183]

[0281] The left - right relationship information estimation unit 183 is the same as the left - right relationship information estimation unit 183 of the second embodiment. The input audio signal of the left channel input to the encoding device 104 and the input audio signal of the right channel input to the encoding device 104 are input to the left - right relationship information estimation unit 183. The left - right relationship information estimation unit 183 obtains the correlation coefficient of the input audio signal of the left channel and the input audio signal of the right channel, that is, the left - right correlation coefficient γ, and the leading - channel information indicating which one of the input audio signal of the left channel and the input audio signal of the right channel leads, and outputs them (step S183).

[0282] [Down - Mix Unit 112] ​

[0283] The downmixing unit 112 is the same as the downmixing unit 112 of the second embodiment. The input sound signal of the left channel input to the encoding device 104, the input sound signal of the right channel input to the encoding device 104, the left-right correlation coefficient γ output by the left-right relationship information estimation unit 183, and the leading channel information output by the left-right relationship information estimation unit 183 are input to the downmixing unit 112. The downmixing unit 112 performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmixed signal and output it so that the larger the left-right correlation coefficient γ in the downmixed signal, the larger the input sound signal of the leading channel in the input sound signals of the left channel and the right channel (step S112).

[0284] For example, when the sampling number is set to t, the input sound signal of the left channel is set to x L (t), the input sound signal of the right channel is set to x R (t), and the downmixed signal is set to x M (t), when the leading channel information indicates that the left channel leads, for each sampling number t, the downmixed signal is obtained by x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t). When the leading channel information indicates that the right channel leads, for each sampling number t, the downmixed signal is obtained by x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t). When the leading channel information indicates that neither channel leads, for each sampling number t, the downmixed signal is obtained by x M (t) = (x L (t) + x R (t)) / 2.

[0285] [Mono Encoding Unit 160]

[0286] The mono encoder 160 is the same as the mono encoder 160 in the second embodiment. The downmix signal output from the downmix unit 112 is input to the mono encoder 160. The mono encoder 160 encodes the input downmix signal to obtain a mono code CM and outputs it (step S160). The mono encoder 160 may use any encoding method. For example, it may use an encoding method such as the 3GPP EVS standard. The encoding method may be an encoding method that performs encoding processing independently of the stereo encoder 174 described later, that is, an encoding method that performs encoding processing without using the stereo code CS' obtained by the stereo encoder 174 or the information obtained in the encoding processing performed by the stereo encoder 174, or may be an encoding method that performs encoding processing using the stereo code CS' obtained by the stereo encoder 174 or the information obtained in the encoding processing performed by the stereo encoder 174.

[0287] [Stereo Encoder 174]

[0288] The input audio signal of the left channel input to the encoding device 104 and the input audio signal of the right channel input to the encoding device 104 are input to the stereo encoder 174. The stereo encoder 174 encodes the input audio signal of the left channel and the input audio signal of the right channel to obtain a stereo code CS' and outputs it (step S174). The stereo encoder 174 may use any encoding method. For example, it may use a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard, or may use an encoding method that encodes the input audio signal of the left channel and the input audio signal of the right channel independently, or may use the combined encoding of all the encodings obtained by encoding as the stereo code CS'. The encoding method may be an encoding method that performs encoding processing independently of the mono encoder 160, that is, an encoding method that performs encoding processing without using the mono code CM obtained in the mono encoder 160 or the information obtained in the encoding processing performed by the mono encoder 160, or may be an encoding method that performs encoding processing using the mono code CM obtained in the mono encoder 160 or the information obtained in the encoding processing performed by the mono encoder 160.

[0289] <Fourth Embodiment>

[0290] As can be seen from the description in the above embodiments, if an encoding device encodes at least a downmix signal obtained from an input sound signal of a left channel and an input sound signal of a right channel to obtain an encoding, then regardless of what kind of encoding device, a structure that obtains a downmix signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel can be adopted. In addition, not limited to an encoding device, if a signal processing device performs signal processing on at least a downmix signal obtained from an input sound signal of a left channel and an input sound signal of a right channel to obtain a signal processing result, then regardless of what kind of signal processing device it is, a structure that obtains a downmix signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel can be adopted. Moreover, as a downmix device used in the front stage of these encoding devices and signal processing devices, a structure that obtains a downmix signal by considering the relationship between the input sound signal of the left channel and the input sound signal of the right channel can also be adopted. These methods will be described as the fourth embodiment.

[0291] <<Audio signal encoding device 105>>

[0292] As Figure 18 shown, the audio signal encoding device 105 of the fourth embodiment includes: a left-right relationship information estimation unit 183, a downmix unit 112, and an encoding unit 195. The audio signal encoding device 105 of the fourth embodiment performs the processing of step S183, step S112, and step S195 as Figure 19 illustrated for each frame. Hereinafter, the audio signal encoding device 105 of the fourth embodiment will be described with appropriate reference to the description of the second embodiment.

[0293] [Left-right relationship information estimation unit 183]

[0294] The left-right relationship information estimation unit 183 is the same as the left-right relationship information estimation unit 183 of the second embodiment. Based on the input sound signal of the left channel and the input sound signal of the right channel that are input, it obtains the correlation coefficient of the input sound signal of the left channel and the input sound signal of the right channel, that is, the left-right correlation coefficient γ, and the information indicating which one of the input sound signal of the left channel and the input sound signal of the right channel leads, that is, the leading channel information, and outputs them (step S183).

[0295] [Downmix unit 112]

[0296] The downmix unit 112 is the same as the downmix unit 112 of the second embodiment. It performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmix signal and outputs it, so that the larger the left-right correlation coefficient γ is, the more the input sound signal of the leading channel among the input sound signal of the left channel and the input sound signal of the right channel is included in the downmix signal (step S112).

[0297] [Coding unit 195]

[0298] At least the downmixed signal output by the downmixing unit 112 is input to the encoding unit 195. The encoding unit 195 encodes at least the input downmixed signal to obtain a sound signal code and outputs it (step S195). The encoding unit 195 may also encode the input sound signal of the left channel and the input sound signal of the right channel, or may also include the code obtained by the encoding in the sound signal code and output it. In this case, Figure 18 As indicated by the middle dotted line, the input audio signal of the left channel and the input audio signal of the right channel are also input to the encoding unit 195 .

[0299] <<Sound signal processing device 305>>

[0300] like Figure 20 As shown, the audio signal processing device 305 of the fourth embodiment includes: a left-right relationship information estimation unit 183, a downmixing unit 112, and a signal processing unit 315. The audio signal processing device 305 of the fourth embodiment performs the following steps on each frame: Figure 21 The processing of step S183, step S112, and step S315 are exemplified. The following describes the differences between the sound signal processing device 305 of the fourth embodiment and the sound signal coding device 105 of the fourth embodiment.

[0301] [Signal processing unit 315]

[0302] At least the downmixed signal output by the downmixing unit 112 is input to the signal processing unit 315. The signal processing unit 315 performs signal processing on at least the input downmixed signal, obtains a signal processing result, and outputs it (step S315). The signal processing unit 315 may also perform signal processing on the input audio signal of the left channel and the input audio signal of the right channel to obtain a signal processing result. In this case, Figure 20As shown by the dashed line in the figure, the input audio signal of the left channel and the input audio signal of the right channel are also input to the signal processing unit 315. For example, the signal processing unit 315 can perform signal processing on the input audio signals of each channel using the downmix signal, and obtain the output audio signals of each channel as the signal processing result. It can also perform this signal processing on the decoded audio signal of the left channel and the decoded audio signal of the right channel obtained by decoding the encoded CS’ obtained by the stereo encoding unit 174 of the third embodiment by a decoding device having a decoding unit corresponding to the stereo encoding unit 174. That is, it is not necessary that the input audio signal of the left channel and the input audio signal of the right channel input to the audio signal processing device 305 are digital audio signals or audio signals obtained by collecting and performing AD conversion on the input audio signals of the left and right channels using two microphones respectively. The input audio signal of the left channel and the input audio signal of the right channel input to the audio signal processing device 305 can be the decoded audio signal of the left channel and the decoded audio signal of the right channel obtained by decoding the encoding, and as long as they are stereo two-channel audio signals, they can be any obtained audio signals.

[0303] In the case where the input audio signal of the left channel and the input audio signal of the right channel input to the audio signal processing device 305 are the decoded audio signal of the left channel and the decoded audio signal of the right channel obtained by decoding the encoding by other devices, etc., sometimes the same left-right correlation coefficient γ and any one or both of the preceding channel information obtained by the left-right relationship information estimation unit 183 are obtained by other devices. In the case where any one or both of the left-right correlation coefficient γ and the preceding channel information are obtained by other devices, as Figure 20 shown by the dotted line in the figure, it is only necessary to input to the audio signal processing device 305 any one or both of the left-right correlation coefficient γ and the preceding channel information obtained by other devices. In this case, the left-right relationship information estimation unit 183 only needs to obtain the left-right correlation coefficient γ or the preceding channel information that is not input to the audio signal processing device 305. In the case where both the left-right correlation coefficient γ and the preceding channel information are input to the audio signal processing device 305, the audio signal processing device 305 can not have the left-right relationship information estimation unit 183 and not perform step S183. That is, as Figure 20 shown by the double dotted line in the figure, the audio signal processing device 305 is provided with a left-right relationship information acquisition unit 185, and the left-right relationship information acquisition unit 185 only needs to obtain the correlation coefficient between the input audio signal of the left channel and the input audio signal of the right channel, that is, the left-right correlation coefficient γ, and the information indicating which of the input audio signal of the left channel and the input audio signal of the right channel is preceding, that is, the preceding channel information, and output it (step S185). In addition, the left-right relationship information estimation unit 183 and step S183 of the above-mentioned respective devices can also be said to be within the scope of the left-right relationship information acquisition unit 185 and step S185.

[0304] <<Sound signal downmixing device 405>>

[0305] like Figure 22 As shown, the audio signal downmixing device 405 of the fourth embodiment includes: a left-right relationship information acquisition unit 185 and a downmixing unit 112. The audio signal downmixing device 405 performs the following steps on each frame: Figure 23 The processing of step S185 and step S112 is illustrated. The sound signal downmixing device 405 is described below with reference to the description of the second embodiment as appropriate. In addition, similarly to the sound signal processing device 305, the input sound signal of the left channel and the input sound signal of the right channel input to the sound signal downmixing device 405 can be a digital sound signal or sound signal collected by two microphones and AD-converted, or a decoded sound signal of the left channel and a decoded sound signal of the right channel obtained by decoding the encoded sound. As long as it is a stereo two-channel sound signal, it can be any sound signal.

[0306] [Left-right relationship information acquisition unit 185]

[0307] The left-right relationship information acquisition unit 185 obtains and outputs the left-right correlation coefficient γ, which is the correlation coefficient between the input sound signal of the left channel and the input sound signal of the right channel, and the preceding channel information, which is information indicating which of the input sound signal of the left channel and the input sound signal of the right channel precedes (step S185).

[0308] When both the left-right correlation coefficient γ and the preceding channel information are obtained by other means, as Figure 22 As indicated by the dot-dash line, the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ and the preceding channel information input to the audio signal downmixing device 405 from other devices, and outputs them to the downmixing unit 112 .

[0309] In the case where both the left-right correlation coefficient γ and the preceding channel information are not obtained by other means, such as Figure 22 As shown by the middle dotted line, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. The left-right relationship information estimation unit 183 obtains the left-right correlation coefficient γ and the preceding channel information from the input audio signal of the left channel and the input audio signal of the right channel, and outputs them to the downmixing unit 112, similarly to the left-right relationship information estimation unit 183 of the second embodiment.

[0310] In the case where neither the left-right correlation coefficient γ nor the preceding channel information is obtained by other means, such as Figure 22As shown by the dashed line in the figure, the left-right relationship information acquisition unit 185 includes a left-right relationship information estimation unit 183. Similar to the left-right relationship information estimation unit 183 in the second embodiment, the left-right relationship information estimation unit 183 of the left-right relationship information acquisition unit 185 obtains the left-right correlation coefficient γ that has not been obtained by other devices or the leading channel information that has not been obtained by other devices from the input sound signal of the left channel and the input sound signal of the right channel, and outputs it to the downmixing unit 112. Regarding the left-right correlation coefficient γ obtained by other devices or the leading channel information obtained by other devices, as Figure 22 As shown by the dotted line in the figure, the left-right relationship information acquisition unit 185 outputs the left-right correlation coefficient γ or the leading channel information of the input sound signal downmixing device 405 from other devices to the downmixing unit 112.

[0311] [Downmixing unit 112]

[0312] The downmixing unit 112 is the same as the downmixing unit 112 in the second embodiment. Based on the leading channel information and the left-right correlation coefficient obtained by the left-right relationship information acquisition unit 185, the downmixing unit 112 performs weighted averaging on the input sound signal of the left channel and the input sound signal of the right channel to obtain a downmixing signal and outputs it, so that the larger the left-right correlation coefficient γ, the more the input sound signal of the leading channel in the input sound signals of the left channel and the right channel is included in the downmixing signal (step S112).

[0313] For example, when the sampling number is set to t, the input sound signal of the left channel is set to x L (t), the input sound signal of the right channel is set to x R (t), and the downmixing signal is set to x M (t), when the leading channel information indicates that the left channel leads, for each sampling number t, the downmixing unit 112 passes x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t) to obtain the downmixing signal. When the leading channel information indicates that the right channel leads, for each sampling number t, the downmixing unit 112 passes x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t) to obtain the downmixing signal. When the leading channel information indicates that neither channel leads, for each sampling number t, the downmixing unit 112 passes x M (t) = (x L (t) + x R (t)) / 2 to obtain the downmixing signal.

[0314] [Program and recording medium]

[0315] The processing of each part of the above-described encoding devices, decoding devices, voice signal encoding device, voice signal processing device, and voice signal mixing device can also be implemented by a computer. In this case, the processing content of the functions that each device should have is described by a program. Then, the storage unit 1020 of the computer 1000 as shown in Figure 24 reads in this program, and the arithmetic processing unit 1010, input unit 1030, output unit 1040, etc. are operated, so as to implement various processing functions in the above-described devices on the computer.

[0316] The program describing this processing content can be recorded on a recording medium readable by a computer. A computer-readable recording medium is, for example, a non-transitory recording medium. Specifically, it is a magnetic recording device, an optical disc, etc.

[0317] Moreover, the distribution of this program is carried out, for example, by removable recording media such as DVDs and CD-ROMs that record this program through sales, transfers, rentals, etc. Furthermore, it can also be configured such that this program is stored in the storage device of a server computer, and via a network, by forwarding this program from the server computer to other computers, the program is distributed.

[0318] A computer that executes such a program, for example, first temporarily stores the program recorded in the transportable recording medium or the program forwarded from the server computer in its own non-temporary storage device, namely the auxiliary recording unit 1050. Then, when executing the processing, the storage unit 1020 reads in the program stored in its own non-temporary storage device, namely the auxiliary recording unit 1050, and executes the processing according to the read-in program. Moreover, as another execution mode of this program, the computer can also directly read the program from the transportable recording medium into the storage unit 1020 and execute the processing according to this program. Furthermore, it can also be configured such that each time the program is forwarded from the server computer to this computer, the processing according to the received program is executed successively. Moreover, it can also be configured as a so-called ASP (Application Service Provider) type of service that realizes the processing function only through the execution instruction and result acquisition without forwarding the program from the server computer to this computer, and executes the above-described processing. Moreover, in the program in this mode, it includes information for the processing of an electronic computer, namely information based on the program (although not a direct instruction for the computer, but data having the nature of prescribing the processing of the computer, etc.).

[0319] Moreover, in this mode, it is assumed that this device is constituted by executing a prescribed program on a computer, but at least a part of these processing contents can also be implemented hardware-wise.

[0320] In addition, of course, appropriate changes can be made without departing from the gist of the present invention.

Claims

1. A method for mixing sound signals, which obtains a mixed signal, i.e., a mixed-down signal, formed by mixing a left-channel input sound signal and a right-channel input sound signal, characterized in that, Comprising: A step of obtaining left-right relationship information to obtain leading channel information and a left-right correlation coefficient, where the leading channel information is information indicating which one of the left-channel input sound signal and the right-channel input sound signal leads, and the left-right correlation coefficient is the correlation coefficient between the left-channel input sound signal and the right-channel input sound signal; And A downmixing step of performing weighted averaging on the left-channel input sound signal and the right-channel input sound signal according to the leading channel information and the left-right correlation coefficient to obtain a downmixed signal, such that the larger the left-right correlation coefficient, the greater the inclusion of the input sound signal of the leading channel in the left-channel input sound signal and the right-channel input sound signal.

2. The method for mixing sound signals according to claim 1, characterized in that, Let the sampling number be t, and the input sound signal of the left channel be x L (t), and the input sound signal of the right channel be x R (t), and the downmix signal be x M (t), and the left-right correlation coefficient be γ In the downmixing step, When the prior channel information indicates that the left channel is prior, for each sampling number t, through x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t), the downmix signal is obtained. When the preceding channel information indicates that the right channel is the preceding channel, for each sampling number t, through x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t), the downmix signal is obtained. When the preceding channel information indicates that no channel precedes, for each sampling number t, the downmixed signal is obtained by x M (t)=(x L (t)+x R (t)) / 2 3. A method for encoding sound signals, characterized in that, Including the sound signal downmixing method according to claim 1 or 2 as a sound signal downmixing step, Further comprising: A mono encoding step of encoding the downmixed signal obtained in the downmixing step to obtain a mono encoding; and A stereo encoding step of encoding the left-channel input sound signal and the right-channel input sound signal to obtain a stereo encoding.

4. A device for mixing sound signals, which obtains a mixed signal, i.e., a mixed-down signal, formed by mixing a left-channel input sound signal and a right-channel input sound signal, characterized in that, Comprising: A left-right relationship information obtaining unit that obtains leading channel information and a left-right correlation coefficient, where the leading channel information is information indicating which one of the left-channel input sound signal and the right-channel input sound signal leads, and the left-right correlation coefficient is the correlation coefficient between the left-channel input sound signal and the right-channel input sound signal; And A downmixing unit that performs weighted averaging on the left-channel input sound signal and the right-channel input sound signal according to the leading channel information and the left-right correlation coefficient to obtain a downmixed signal, such that the larger the left-right correlation coefficient, the greater the inclusion of the input sound signal of the leading channel in the left-channel input sound signal and the right-channel input sound signal.

5. The device for mixing sound signals according to claim 4, characterized in that, Let the sampling number be t, the left-channel input audio signal be x L (t), the right-channel input audio signal be x R (t), the mixed signal be x M (t), and the left-right correlation coefficient be γ. The downmixing unit executes: When the preceding channel information indicates that the left channel is the preceding channel, for each sampling number t, the downmix signal is obtained by x M (t) = ((1 + γ) / 2) × x L (t) + ((1 - γ) / 2) × x R (t). When the prior channel information indicates that the right channel is prior, for each sampling number t, through x M (t) = ((1 - γ) / 2) × x L (t) + ((1 + γ) / 2) × x R (t) the downmix signal is obtained, In the case where the said preceding channel information indicates that no channel precedes, for each sampling number t, through x M (t) = (x L (t) + x R (t)) / 2, the said downmixed signal is obtained.

6. A device for encoding sound signals, characterized in that, Including the sound signal downmixing device according to claim 4 or 5 as a sound signal downmixing unit, Further comprising: A mono encoding unit that encodes the downmixed signal obtained by the downmixing unit to obtain a mono encoding; and A stereo encoding unit that encodes the left-channel input sound signal and the right-channel input sound signal to obtain a stereo encoding.

7. A computer-readable recording medium, which records a program for causing a computer to execute the processes of the steps of the method for mixing sound signals according to claim 1 or 2.

8. A computer-readable recording medium that records a program for causing a computer to execute the processes of the steps of the sound signal encoding method according to claim 3.

Citation Information

Patent Citations

  • Sound coding device and sound coding method

    WO2006070751A1

  • Stereo encoding method and device as well as encoder

    CN101826326A

  • Stereo signal converter, stereo signal reverse converter, and methods for both

    CN101981616A