Audio signal downmixing method, audio signal encoding method, audio signal downmixing device, audio signal encoding device, and program
By employing delayed crosstalk addition and downmixing techniques, the method effectively generates a monaural signal from two-channel sound signals, enhancing encoding efficiency by utilizing left-right channel correlations and timing differences.
Patent Information
- Application Number
- JP2023544861
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2041-09-01
AI Technical Summary
Existing technologies for obtaining a monaural signal from two-channel sound signals are inadequate for signal processing such as encoding, as they do not effectively utilize the relationship between the left and right channels to produce a monaural signal that is useful for encoding.
The proposed method involves a delayed crosstalk addition step, left-right relationship information acquisition, and a downmix step to generate a downmix signal that incorporates the input sound signal of the preceding channel to a greater extent based on the left-right correlation value, followed by mono and stereo encoding steps.
This approach enables the generation of a monaural signal suitable for encoding from two-channel sound signals, improving encoding efficiency by leveraging the correlation and timing differences between the left and right channels.
Smart Images

Figure 0007803346000021 
Figure 0007803346000022 
Figure 0007803346000023
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for obtaining a monaural sound signal from two-channel sound signals in order to encode the sound signal in monaural, encode the sound signal using a combination of monaural and stereo encoding, process the sound signal in monaural, or process a stereo sound signal using a monaural sound signal. [Background technology]
[0002] A technology for obtaining a monaural sound signal from two-channel sound signals and for embedded encoding / decoding the two-channel sound signal and the monaural sound signal is disclosed in Patent Document 1. Patent Document 1 discloses a technology for obtaining a monaural signal by averaging an input left-channel sound signal and an input right-channel sound signal for each corresponding sample, encoding the monaural signal (monaural encoding) to obtain a monaural code, decoding the monaural code (monaural decoding) to obtain a monaural locally decoded signal, and encoding, for each of the left and right channels, a difference (prediction residual signal) between the input sound signal and a prediction signal obtained from the monaural locally decoded signal. In the technology of Patent Document 1, for each channel, a signal obtained by delaying and assigning an amplitude ratio to a monaural locally decoded signal is used as a predicted signal, and a predicted signal having a delay and amplitude ratio that minimizes the error between the input sound signal and the predicted signal is selected, or a predicted signal having a delay and amplitude ratio that maximizes the cross-correlation between the input sound signal and the monaural locally decoded signal is used, and a predicted signal is subtracted from the input sound signal to obtain a predicted residual signal, which is then subjected to encoding / decoding, thereby suppressing deterioration in sound quality of the decoded sound signal of each channel. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2006 / 070751 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology of Patent Document 1 can improve the coding efficiency of each channel by optimizing the delay and amplitude ratio given to the monaural locally decoded signal when obtaining a predicted signal. However, in the technology of Patent Document 1, the monaural locally decoded signal is obtained by encoding and decoding a monaural signal obtained by averaging the left channel sound signal and the right channel sound signal. In other words, the technology of Patent Document 1 has a problem in that it does not incorporate any measures to obtain a monaural signal useful for signal processing such as encoding from two channel sound signals. An object of the present invention is to provide a technique for obtaining a monaural signal useful for signal processing such as encoding from a two-channel sound signal. [Means for solving the problem]
[0005] One aspect of the present invention is a delayed crosstalk addition step for obtaining, for each of the two channels, a signal obtained by adding the input sound signal of the channel and a signal obtained by delaying the input sound signal of the other channel and multiplying the delayed sound signal by a weighting value that is a predetermined positive real value and has an absolute value smaller than 1, as a delayed crosstalk-added signal of the channel; a left-right relationship information acquisition step for obtaining preceding channel information that indicates which of the delayed crosstalk-added signals of the two channels is preceding, and a left-right correlation value that indicates the magnitude of correlation between the delayed crosstalk-added signals of the two channels; and a downmix step for obtaining, based on the left-right correlation value and the preceding channel information, a downmix signal in which the input sound signal of the preceding channel out of the input sound signals of the two channels is included to a greater extent the larger the left-right correlation value is; The present invention is characterized by having the following.
[0006] One aspect of the present invention is an audio signal encoding method that includes the above-described audio signal downmixing method as an audio signal downmixing step, and is characterized by including a mono encoding step of encoding the downmixed signal obtained in the downmixing step to obtain a mono code, and a stereo encoding step of encoding two-channel input audio signals to obtain a stereo code. [Effects of the Invention]
[0007] According to the present invention, a monaural signal useful for signal processing such as encoding can be obtained from a two-channel sound signal. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing a sound signal downmixing device according to a first embodiment. [Figure 2] 3 is a flowchart showing the processing of the sound signal downmixing device of the first embodiment. [Figure 3]FIG. 10 is a block diagram illustrating an example of a sound signal downmixing device according to a second embodiment. [Figure 4] 10 is a flowchart illustrating an example of processing performed by the sound signal downmixing device of the second embodiment. [Figure 5] FIG. 10 is a block diagram illustrating an example of a sound signal encoding device according to a third embodiment. [Figure 6] 10 is a flowchart illustrating an example of processing performed by a sound signal encoding device according to a third embodiment. [Figure 7] FIG. 10 is a block diagram showing an example of a sound signal processing device according to a fourth embodiment. [Figure 8] 10 is a flowchart showing an example of processing by the sound signal processing device of the fourth embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of the functional configuration of a computer that realizes each device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] First Embodiment The two-channel sound signals that are the subject of signal processing such as encoding are often digital sound signals obtained by AD converting sounds picked up by a left-channel microphone and a right-channel microphone placed in a certain space. In this case, what is input to a device that performs signal processing such as encoding are a left-channel input sound signal, which is a digital sound signal obtained by AD converting the sound picked up by the left-channel microphone placed in the space, and a right-channel input sound signal, which is a digital sound signal obtained by AD converting the sound picked up by the right-channel microphone placed in the space. These left-channel input sound signal and right-channel input sound signal often contain sounds emitted from each sound source present in the space, with a given difference between the arrival time from the sound source to the left-channel microphone and the arrival time from the sound source to the right-channel microphone (so-called arrival time difference).
[0010] In the technology of Patent Document 1 described above, a signal obtained by delaying and giving an amplitude ratio to a monaural locally decoded signal is used as a prediction signal, and the prediction signal is subtracted from an input sound signal to obtain a prediction residual signal, which is then encoded / decoded. That is, for each channel, the more similar the input sound signal and the monaural locally decoded signal are, the more efficient the encoding. However, for example, if only sound emitted by a single sound source present in a certain space is included in a left channel input sound signal and a right channel input sound signal with an arrival time difference, and the monaural locally decoded signal is obtained by encoding and decoding a monaural signal obtained by averaging the left channel input sound signal and the right channel input sound signal, even though the left channel input sound signal, the right channel input sound signal, and the monaural locally decoded signal all contain only sound emitted by the same sound source, the degree of similarity between the left channel input sound signal and the monaural locally decoded signal is not very high, and the degree of similarity between the right channel input sound signal and the monaural locally decoded signal is also not very high. Thus, simply averaging the left channel input sound signal and the right channel input sound signal to obtain a monaural signal may not provide a monaural signal that is useful for signal processing such as encoding.
[0011] Therefore, the sound signal downmixing device of the first embodiment performs downmixing processing that takes into account the relationship between the left channel input sound signal and the right channel input sound signal so as to obtain a monaural signal that is useful for signal processing such as encoding. The sound signal downmixing device of the first embodiment will be described below.
[0012] As shown in FIG. 1 , the sound signal downmixing device 100 of the first embodiment includes a left-right relation information estimation unit 120 and a downmixing unit 130. The sound signal downmixing device 100 obtains and outputs a downmix signal (described later) from an input two-channel stereo time-domain sound signal in frame units of a predetermined time length, for example, 20 ms. The sound signal downmixing device 100 receives a two-channel stereo time-domain sound signal, which may be, for example, a digital sound signal obtained by collecting sound such as speech or music with two microphones and performing AD conversion, a digital decoded sound signal obtained by encoding and decoding the digital sound signal, or a digitally processed sound signal obtained by signal processing the digital sound signal, and which is composed of a left channel input sound signal and a right channel input sound signal. The downmix signal, which is a time-domain monaural sound signal obtained by the sound signal downmixing device 100, is input to at least a sound signal encoding device that encodes the downmix signal or a sound signal processing device that processes the downmix signal. If the number of samples per frame is T, the sound signal downmixing device 100 receives the left channel input sound signal x L (1), x L (2), ..., x L (T) and right channel input sound signal x R (1), x R (2), ..., x R (T) is input, and the audio signal downmixing device 100 generates a downmix signal x in units of frames. M (1), x M (2), ..., x M (T) is obtained and output, where T is a positive integer, and for example, if the frame length is 20 ms and the sampling frequency is 32 kHz, T is 640. The sound signal downmixing apparatus 100 performs the processes of steps S120 and S130 illustrated in FIG. 2 for each frame.
[0013] [Left-right relationship information estimation unit 120] The left-right relation information estimation unit 120 receives the left channel input sound signal input to the sound signal downmixing device 100 and the right channel input sound signal input to the sound signal downmixing device 100. The left-right relation information estimation unit 120 obtains and outputs a left-right correlation value γ and preceding channel information from the left channel input sound signal and the right channel input sound signal (step S120).
[0014] The preceding channel information is information corresponding to whether a sound emitted from a main sound source in a space reaches the left channel microphone or the right channel microphone arranged in the space earlier. In other words, the preceding channel information is information indicating whether the same sound signal is contained first in the left channel input sound signal or the right channel input sound signal. If the same sound signal is contained first in the left channel input sound signal, it is said that the left channel is leading or the right channel is trailing. If the same sound signal is contained first in the right channel input sound signal, it is said that the right channel is leading or the left channel is trailing. The preceding channel information is information indicating which channel, the left channel or the right channel, is leading. The left-right correlation value γ is a correlation value that takes into account the time difference between the left channel input sound signal and the right channel input sound signal. In other words, the left-right correlation value γ is a value that represents the magnitude of the correlation between the sample sequence of the input sound signal of the leading channel and the sample sequence of the input sound signal of the trailing channel, which is shifted τ samples later than the sample sequence. This τ will hereinafter also be referred to as the left-right time difference. The preceding channel information and the left-right correlation value γ are information that indicates the relationship between the left channel input sound signal and the right channel input sound signal, and therefore can also be said to be left-right relationship information.
[0015] For example, if the absolute value of the correlation coefficient is used as the value representing the magnitude of the correlation, the left-right relationship information estimation unit 120 may use a predetermined τ max From τ min up to (e.g., τ max is a positive number, τ min is a negative number) cand , the sample sequence of the left channel input sound signal and the number of candidate samples τcand the absolute value γ of the correlation coefficient with a sample sequence of the right channel input sound signal that is shifted after the sample sequence by cand The maximum value of the left-right correlation value γ is obtained and output, and τ when the absolute value of the correlation coefficient is the maximum value cand When the value of is positive, the information indicating that the left channel is leading is obtained and output as leading channel information, and the absolute value of the correlation coefficient is the maximum value. cand When τ is a negative value, the left-right relationship information estimation unit 120 obtains and outputs information indicating that the right channel is leading as leading channel information. cand When is 0, information indicating that the left channel is leading may be obtained and output as leading channel information, or information indicating that the right channel is leading may be obtained and output as leading channel information, but it is preferable to obtain and output information indicating that neither channel is leading as leading channel information.
[0016] The number of predetermined candidate samples is τ max From τ min It may be any integer value up to τ max From τ min It may contain fractional or decimal values between τ max From τ min It is also possible that the integer values between τ and τ are not included. max =-τ min Assuming that an input sound signal is targeted in which it is not known which channel is leading, τ max Let be a positive number, and τ min It is better to set the absolute value of the correlation coefficient γ cand In calculating the above, one or more samples of a past input sound signal that are consecutive to the sample sequence of the input sound signal of the current frame may also be used. In this case, the sample sequence of the input sound signal of the past frames may be stored in a memory unit (not shown) in the left-right relationship information estimation unit 120 for a predetermined number of frames.
[0017] Alternatively, for example, instead of the absolute value of the correlation coefficient, a correlation value using the phase information of the signal may be expressed as γ cand In this example, the left-right relation information estimation unit 120 first calculates the left channel input sound signal x L (1), x L (2), ..., x L (T) and right channel input sound signal x R (1), x R (2), ..., x R By Fourier transforming each of (T) as in the following equations (1-1) and (1-2), the frequency spectrum X at each frequency k from 0 to T-1 is obtained. L (k) and X R Obtain (k).
number
number
[0018] Next, the left-right relationship information estimation unit 120 calculates the frequency spectrum X at each frequency k obtained by the formulas (1-1) and (1-2). L (k) and X R Using (k), the spectrum φ(k) of the phase difference at each frequency k is obtained by the following equation (1-3).
number
[0019] Next, the left-right relationship information estimation unit 120 performs an inverse Fourier transform on the spectrum of the phase difference obtained by equation (1-3) to obtain τ as shown in the following equation (1-4): max From τ min The number of candidate samples up to τ cand Regarding the phase difference signal ψ(τ cand ) is obtained.
number
[0020] The phase difference signal ψ(τ cand ) is the absolute value of the left channel input sound signal x L (1), x L (2), ..., x L (T) and right channel input sound signal x R (1), x R (2), ..., x R Since the left-right relationship information estimation unit 120 estimates the number of candidate samples τ cand phase difference signal ψ(τ cand ) is the absolute value of the correlation value γ cand That is, the left-right relationship information estimation unit 120 uses the phase difference signal ψ(τ cand ) is the absolute value of the correlation value γ cand The maximum value of γ is obtained as the left-right correlation value γ and output. cand When the correlation value is a positive value, information indicating that the left channel is leading is obtained and output as leading channel information, and τ cand When τ is a negative value, the left-right relationship information estimation unit 120 obtains and outputs information indicating that the right channel is leading as leading channel information. cand When the correlation value γ is 0, the left-right relation information estimation unit 120 may obtain and output information indicating that the left channel is leading as the leading channel information, or may obtain and output information indicating that the right channel is leading as the leading channel information, but it is preferable to obtain and output information indicating that neither channel is leading as the leading channel information. cand as the phase difference signal ψ(τ cand Instead of using the absolute value of each τ cand Regarding the phase difference signal ψ(τ cand τ for the absolute value of cand Alternatively, a normalized value such as a relative difference between the average of the absolute values of the phase difference signals obtained for each of a plurality of candidate samples before and after the phase difference signal may be used. cand For a predetermined positive number τ rangeUsing the above, the average value is calculated using the following formula (1-5), and the obtained average value ψ c (τ cand ) and the phase difference signal ψ(τ cand ) and normalized the correlation value obtained by the following equation (1-6) as γ cand It may also be used as.
number
number
[0021] The normalized correlation value obtained by equation (1-6) is a value between 0 and 1, and τ cand is plausibly close to 1 as the left-right time difference, and τ cand is a value that indicates that the time difference between the left and right sides is so close to 0 that it is unlikely to occur.
[0022] [Downmix section 130] The downmixing unit 130 receives as input the left channel input sound signal input to the sound signal downmixing device 100, the right channel input sound signal input to the sound signal downmixing device 100, the left-right correlation value γ output by the left-right relationship information estimation unit 120, and the preceding channel information output by the left-right relationship information estimation unit 120. The downmixing unit 130 obtains and outputs a downmix signal by weighting and adding the left channel input sound signal and the right channel input sound signal so that the input sound signal of the preceding channel, either the left channel input sound signal or the right channel input sound signal, is included in the downmix signal to a greater extent the larger the left-right correlation value γ (step S130).
[0023] For example, if the absolute value or normalized value of the correlation coefficient is used as the correlation value as in the example described above in the description of the left-right relationship information estimation unit 120, the left-right correlation value γ input from the left-right relationship information estimation unit 120 is a value between 0 and 1, and therefore the downmixing unit 130 uses a weight determined by the left-right correlation value γ for each corresponding sample number t to calculate the left channel input sound signal x L(t) and the right channel input sound signal x R (t) is weighted and added to the downmix signal x M For example, if the preceding channel information indicates that the left channel is preceding, that is, if the left channel is preceding, the downmixing unit 130 may set x M (t)=((1+γ) / 2)×x L (t)+((1-γ) / 2)×x R (t), if the leading channel information is information indicating that the right channel is leading, that is, if the right channel is leading, then x M (t)=((1-γ) / 2)×x L (t)+((1+γ) / 2)×x R (t), as the downmix signal x M When the downmix unit 130 obtains a downmix signal in this manner, the smaller the left-right correlation value γ, i.e., the smaller the correlation between the left channel input sound signal and the right channel input sound signal, the closer the downmix signal is to a signal obtained by averaging the left channel input sound signal and the right channel input sound signal, and the larger the left-right correlation value γ, i.e., the greater the correlation between the left channel input sound signal and the right channel input sound signal, the closer the downmix signal is to the input sound signal of the preceding channel, either the left channel input sound signal or the right channel input sound signal.
[0024] If none of the channels are preceding, the downmixing unit 130 preferably performs weighted addition of the left channel input sound signal and the right channel input sound signal so that the left channel input sound signal and the right channel input sound signal are included in the downmixed signal with the same weight, and outputs the downmixed signal. In other words, if the preceding channel information indicates that none of the channels are preceding, the downmixing unit 130 preferably performs weighted addition of the left channel input sound signal and the right channel input sound signal to obtain the downmixed signal. Specifically, for each sample number t, the downmixing unit 130 performs weighted addition of the left channel input sound signal and the right channel input sound signal to obtain the downmixed signal. L (t) and the right channel input sound signal x R (t) averaged x M (t)=(x L (t)+xR (t) / 2 is the downmix signal x M It is recommended to use (t).
[0025] Second Embodiment When the left channel microphone and the right channel microphone are arranged at positions apart in space, and for example, when a sound source emitting a sound is close to the left channel microphone, the sound emitted by the sound source may be hardly contained in the input sound signal picked up by the right channel microphone. In such a case, it would be better for the sound signal downmixing device to convert the left channel input sound signal into a downmix signal useful for signal processing such as encoding. However, in such a case, since the right channel input sound signal hardly contains the sound emitted from the sound source, the sound signal downmixing device 100 of the first embodiment downmixes the left channel input sound signal to a downmix signal useful for signal processing such as encoding. cand If the preceding channel information indicates that the right channel is preceding, a downmix signal containing a larger amount of right channel input sound signals than left channel input sound signals is obtained. In such a case, the sound signal downmixing device 100 of the first embodiment may obtain a small value as the left-right correlation value γ, and may obtain a downmix signal containing a signal close to the average of the left channel input sound signal and the right channel input sound signal. Furthermore, in such a case, τ candThe values of the left-right correlation value γ and the left-right correlation value γ may vary significantly from frame to frame, and the downmix signal obtained by the sound signal downmixing device 100 of the first embodiment may vary significantly from frame to frame. That is, the sound signal downmixing device 100 of the first embodiment has a problem in that it does not necessarily obtain a downmix signal useful for signal processing such as encoding when either the left channel input sound signal or the right channel input sound signal contains a significant amount of sound emitted by the sound source, but the other of the left channel input sound signal and the right channel input sound signal does not contain a significant amount of sound emitted by the sound source. The sound signal downmixing device of the second embodiment is able to obtain a downmix signal useful for signal processing such as encoding, even when either the left channel input sound signal or the right channel input sound signal contains a significant amount of sound emitted by the sound source, but the other of the left channel input sound signal and the right channel input sound signal does not contain a significant amount of sound emitted by the sound source. The sound signal downmixing device of the second embodiment will be described below, focusing on the differences from the sound signal downmixing device of the first embodiment.
[0026] As shown in Fig. 3, the sound signal downmixing device 200 includes a delayed crosstalk adding unit 210, a left-right relationship information estimating unit 220, and a downmixing unit 230. The sound signal downmixing device 200 obtains and outputs a downmix signal (described later) from a left channel input sound signal and a right channel input sound signal, which are input two-channel stereo time domain sound signals, in units of frames each having a predetermined time length of, for example, 20 ms. The sound signal downmixing device 200 performs the processes of steps S210, S220, and S230 shown in Fig. 4 for each frame.
[0027] [Outline of delay crosstalk addition unit 210] The delayed crosstalk adder 210 receives the left channel input sound signal input to the sound signal downmixing device 200 and the right channel input sound signal input to the sound signal downmixing device 200. The delayed crosstalk adder 210 obtains and outputs a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal from the left channel input sound signal and the right channel input sound signal (step S210). The process by which the delayed crosstalk adder 210 obtains the left channel delayed crosstalk-added signal and the right channel delayed crosstalk-added signal will be described after the left-right relationship information estimation unit 220 and the downmixing unit 230 have been described.
[0028] [Left-right relationship information estimation unit 220] The left-right relation information estimation unit 220 receives as input the left channel crosstalk-added signal output by the delayed crosstalk adder 210 and the right channel crosstalk-added signal output by the delayed crosstalk adder 210. The left-right relation information estimation unit 220 obtains and outputs a left-right correlation value γ and preceding channel information from the left channel crosstalk-added signal and the right channel crosstalk-added signal (step S220). The left-right relation information estimation unit 220 performs the same processing as the left-right relation information estimation unit 120 of the sound signal downmixing device 100 of the first embodiment, but uses the left channel crosstalk-added signal instead of the left channel input sound signal and the right channel crosstalk-added signal instead of the right channel input sound signal.
[0029] That is, the left-right relationship information estimation unit 220 obtains leading channel information, which is information indicating which of the delayed crosstalk-added signals of the two channels is leading, and a left-right correlation value γ, which is a value indicating the magnitude of correlation between the delayed crosstalk-added signals of the two channels.
[0030] [Downmix section 230] The downmixing unit 230 receives as input the left channel input sound signal input to the sound signal downmixing device 200, the right channel input sound signal input to the sound signal downmixing device 200, the left-right correlation value γ output by the left-right relationship information estimation unit 220, and the preceding channel information output by the left-right relationship information estimation unit 220. The downmixing unit 230 obtains and outputs a downmix signal by weighting and adding the left channel input sound signal and the right channel input sound signal so that the input sound signal of the preceding channel, either the left channel input sound signal or the right channel input sound signal, is included more significantly in the downmix signal the larger the left-right correlation value γ (step S230). In other words, the downmixing unit 230 is the same as the downmixing unit 130 of the sound signal downmixing device 100 of the first embodiment, except that the downmixing unit 230 uses the left-right correlation value γ and the preceding channel information obtained by the left-right relationship information estimation unit 220 instead of the left-right relationship information estimation unit 120.
[0031] That is, based on the left-right correlation value γ and the preceding channel information, the downmix unit 230 obtains a downmix signal by weighting and adding the input sound signals of the two channels so that the input sound signal of the preceding channel out of the input sound signals of the two channels is included to a greater extent the larger the left-right correlation value.
[0032] [Details of the delay crosstalk addition unit 210] In a case where a sound emitted by a sound source is significantly included in the left channel input sound signal but not significantly included in the right channel input sound signal (hereinafter also referred to as the "first case"), the downmixing unit 230 can obtain a downmixed signal useful for signal processing such as encoding by allowing the downmixing unit 230 to obtain a signal that mainly includes the left channel input sound signal as the downmixed signal. In order for the downmixing unit 230 to obtain a signal that mainly includes the left channel input sound signal as the downmixed signal, the left channel input sound signal must be preceding and the left-right correlation value must be a large value. In order for the left-right relation information estimation unit 220 to obtain such preceding channel information and left-right correlation value, when the sound emitted by the sound source is significantly included in the left channel input sound signal but not significantly included in the right channel input sound signal, the left-right relation information estimation unit 220 can obtain preceding channel information and left-right correlation value by regarding a signal that has been processed so that a signal identical to the left channel input sound signal is included in the right channel input sound signal later than the left channel input sound signal as the right channel input sound signal.
[0033] In a case where sound emitted by a sound source is significantly included in the right channel input sound signal but not significantly included in the left channel input sound signal (hereinafter also referred to as the "second case"), the downmixing unit 230 can obtain a downmixed signal useful for signal processing such as encoding by configuring the downmixing unit 230 to obtain a signal that mainly includes the right channel input sound signal as the downmixed signal. In order for the downmixing unit 230 to obtain a signal that mainly includes the right channel input sound signal as the downmixed signal, the right channel input sound signal must be preceding and the left-right correlation value must be a large value. In order for the left-right relation information estimation unit 220 to obtain such preceding channel information and left-right correlation value, when the sound emitted by the sound source is significantly included in the right channel input sound signal but not significantly included in the left channel input sound signal, the left-right relation information estimation unit 220 can obtain preceding channel information and left-right correlation value by regarding a signal that has been processed so that a signal identical to the right channel input sound signal is included in the left channel input sound signal later than the right channel input sound signal as the left channel input sound signal.
[0034] In cases other than these (i.e., when neither the first nor the second case applies), the left-right relationship information estimation unit 220 is preferably configured to obtain preceding channel information and a left-right correlation value, similar to the left-right relationship information estimation unit 120 of the first embodiment. That is, the signal processing described above is required to be such that, when both the left channel input sound signal and the right channel input sound signal contain a significant amount of sound emitted from the sound source, the left-right correlation value and the preceding channel information are not affected, and when either the left channel input sound signal or the right channel input sound signal contains a significant amount of sound emitted from the sound source, a large left-right correlation value is obtained. Experiments conducted by the inventors have shown that this processing is preferably performed by adding a delayed signal of the input sound signal of each channel to the input sound signal of the other channel, with the amplitude reduced to about 1 / 100. However, reducing the amplitude to about 1 / 100 is not essential; it is sufficient to at least reduce the amplitude, and the extent to which the amplitude is reduced can be determined by taking into consideration the types of signals the left channel input sound signal and the right channel input sound signal are.
[0035] Therefore, for each channel, the delayed crosstalk adder 210 obtains, as the delayed crosstalk-added signal of that channel, a signal obtained by adding together the input sound signal of that channel and a signal obtained by delaying the input sound signal of the other channel and multiplying it by a weighting value whose absolute value is a predetermined value less than 1. Specifically, the delayed crosstalk adder 210 obtains, as the left-channel delayed crosstalk-added signal, a signal obtained by adding together the left-channel input sound signal and a signal obtained by delaying the right-channel input sound signal and multiplying it by a weighting value whose absolute value is a predetermined value less than 1, and obtains, as the right-channel delayed crosstalk-added signal, a signal obtained by adding together the right-channel input sound signal and a signal obtained by delaying the left-channel input sound signal and multiplying it by a weighting value whose absolute value is a predetermined value less than 1. It is essential that the weighting value has an absolute value less than 1, and experiments by the inventors have shown that a value of around 0.01 is good, but the weighting value may be set to a predetermined value taking into consideration the types of signals the left-channel input sound signal and the right-channel input sound signal are. Therefore, it is not essential that the weight given to the delayed right channel input sound signal and the weight given to the delayed left channel input sound signal have the same value.
[0036] The delay amount of the input sound signal of the other channel may be any amount of delay as long as it allows the left-right relation information estimation unit 220 to obtain the preceding channel information described above in the first and second cases. When the sound emitted by the sound source is significantly included in the left channel input sound signal but not significantly included in the right channel input sound signal, the delayed crosstalk addition unit 210 calculates τ cand To ensure that τ is a positive value, we set the number of candidate samples τ candand the left channel input sound signal delayed by the delay amount a is included in the right channel delayed crosstalk added signal. In addition, when the sound emitted by the sound source is significantly included in the right channel input sound signal but not significantly included in the left channel input sound signal, the delayed crosstalk adder 210 calculates τ cand To ensure that τ is a negative value, we set the number of candidate samples τ cand The absolute value of any one of the negative values of the above is set as the delay amount a, and the right channel input sound signal delayed by the delay amount a is included in the left channel delayed crosstalk added signal. From the above, the delay amount of the left channel input sound signal in the right channel delayed crosstalk added signal is determined by the number of candidate samples τ cand The delay amount of the right channel input sound signal in the left channel delayed crosstalk added signal may be any one of the positive values of the number of candidate samples τ cand It is preferable that the absolute value of any of the negative values of
[0037] [First Example of Delay Crosstalk Adder 210] As a first example of the delayed crosstalk adder 210, processing in the time domain will be described. In the first example, in order to minimize a decrease in the accuracy with which the left-right relationship information estimation unit 220 obtains the left-right correlation value γ and preceding channel information without increasing the memory amount for processing by the delayed crosstalk adder 210 or the algorithm delay due to the processing by the delayed crosstalk adder 210, it is preferable to set both the delay amount of the right channel input sound signal in the left channel delayed crosstalk added signal and the delay amount of the left channel input sound signal in the right channel delayed crosstalk added signal to about one sample. Therefore, in the first example, an example in which the delay amount is one sample will be described first. If the number of samples per frame is T, the sample number is t, and the sample numbers in the frame range from 1 to T, then the left channel input sound signal sample with sample number t is denoted by x L(t), and the right channel input sound signal sample at sample number t is x R (t), and the left channel delayed crosstalk added signal sample of sample number t is y L (t), and the delayed crosstalk-added signal sample of the right channel at sample number t is y R (t) and the weight value is w, the delayed crosstalk adder 210 calculates the left channel delayed crosstalk added signal y L (1), y L (2), ..., y L (T) is obtained by the following equation (2-1), and the right channel delayed crosstalk added signal y R (1), y R (2), ..., y R (T) can be obtained from the following formula (2-2).
number
number
[0038] The delayed crosstalk adder 210 includes a storage unit (not shown) for storing the last sample of the left channel input sound signal of the immediately preceding frame and the last sample of the right channel input sound signal of the immediately preceding frame, and converts the last sample of the left channel input sound signal of the immediately preceding frame into x L (0) is used in equation (2-2) for the first sample of the left channel input sound signal of the frame to be processed, and the last sample of the right channel input sound signal of the immediately preceding frame is used as x R (0) can be used in the equation (2-1) of the frame to be processed. L (0) = 0 and the process equivalent to equation (2-2) R Alternatively, the delayed crosstalk adder 210 may perform processing equivalent to equation (2-1) where (0) = 0. That is, for the first sample of a frame, the delayed crosstalk adder 210 may treat the input sound signal as a delayed crosstalk-added signal as is for each channel.
[0039] When the delayed crosstalk adder 210 performs processing in the time domain corresponding to a delay amount a other than 1 (where a>0), the above processing can be performed using equations in which t-1 in equations (2-1) and (2-2) is replaced with ta. However, the delay amounts in equations (2-1) and (2-2) do not need to be the same, and the weighting values in equations (2-1) and (2-2) do not need to be the same. For these reasons, the delayed crosstalk adder 210 sets a1 and a2 to predetermined positive values and w1 and w2 to predetermined values whose absolute values are smaller than 1, and calculates the left channel delayed crosstalk added signal y L (1), y L (2), ..., y L (T) is obtained by the following equation (2-1'), and the right channel delayed crosstalk added signal y R (1), y R (2), ..., y R (T) can be obtained by the following formula (2-2').
number
number
[0040] [Second Example of Delay Crosstalk Adder 210] As a second example of the delayed crosstalk adder 210, processing in the frequency domain will be described. First, an example of processing in the frequency domain corresponding to the first example in which the delay amount of the right channel input sound signal in the left channel delayed crosstalk added signal and the delay amount of the left channel input sound signal in the right channel delayed crosstalk added signal are both set to one sample will be described. If the frequency number is k and the frequency numbers in the frequency spectrum frame range from 0 to T-1, the frequency spectrum sample of the left channel input sound signal with frequency number k is expressed as X L (k), and the frequency spectrum sample of the right channel input sound signal with frequency number k is X R(k), and the frequency spectrum sample of the left channel delayed crosstalk added signal of frequency number k is Y L (k), and the frequency spectrum sample of the delayed crosstalk-added signal of the right channel of frequency number k is Y R (k) and the weight value is w, the delayed crosstalk adder 210 calculates the frequency spectrum X L (0), X L (1), ..., X L (T-1) is obtained by equation (1-1), and the frequency spectrum X of the right channel input sound signal is R (0), X R (1), ..., X R (T-1) is obtained by equation (1-2), and the frequency spectrum Y of the left channel delayed crosstalk added signal is L (0), Y L (1), ..., Y L (T-1) is obtained by the following equation (2-3), and the frequency spectrum Y R (0), Y R (1), ..., Y R (T-1) can be obtained by the following formula (2-4).
number
number
[0041] When the delay crosstalk adder 210 performs processing in the frequency domain corresponding to a delay amount a that is not 1 (where a>0), the following equations (2-3) and (2-4) are satisfied:
number
number
number
number
[0042] The frequency spectrum Y obtained by the delay crosstalk adder 210 using equations (2-3) and (2-4), or equations (2-3') and (2-4') L (0), Y L (1), ..., Y L (T-1) and Y R (0), Y R (1), ..., Y R (T-1) is the time domain left channel delayed crosstalk added signal y L (1), yL (2), ..., y L (T) and right channel delayed crosstalk added signal y R (1), y R (2), ..., y R (T) is a frequency spectrum obtained by Fourier transforming (T). Therefore, the delayed crosstalk adder 210 may output the frequency spectrum obtained by equations (2-3) and (2-4), or equations (2-3') and (2-4'), as a delayed crosstalk-added signal in the frequency domain, and the delayed crosstalk-added signal in the frequency domain output by the delayed crosstalk adder 210 may be input to the left-right relationship information estimation unit 220. The left-right relationship information estimation unit 220 may use the input frequency-domain delayed crosstalk-added signal as the frequency spectrum, without performing a process of Fourier transforming the delayed crosstalk-added signal in the time domain to obtain a frequency spectrum.
[0043] Third Embodiment A coding device that codes an audio signal may include the audio signal downmixing device of the second embodiment as an audio signal downmixing unit, and this configuration will be described as a third embodiment.
[0044] ≪Sound signal encoding device 300≫ As shown in FIG. 5 , the sound signal encoding device 300 of the third embodiment includes an sound signal downmixing unit 200 and an encoding unit 340. The sound signal encoding device 300 of the third embodiment encodes an input two-channel stereo time-domain sound signal in frame units with a predetermined time length of, for example, 20 ms, to obtain and output sound signal codes. The two-channel stereo time-domain sound signal input to the sound signal encoding device 300 is, for example, a digital audio signal or acoustic signal obtained by capturing sound such as speech or music with two microphones and performing AD conversion, and is composed of a left channel input sound signal and a right channel input sound signal. The sound signal code output by the sound signal encoding device 300 is input to a sound signal decoding device. The sound signal encoding device 300 of the third embodiment performs the processes of steps S200 and S340 shown in FIG. 6 for each frame. The sound signal encoding device 300 of the third embodiment will be described below with appropriate reference to the description of the second embodiment.
[0045] [Sound signal downmix unit 200] The sound signal downmixing unit 200 obtains and outputs a downmix signal from the left channel input sound signal and the right channel input sound signal input to the sound signal encoding device 300 (step S200). The sound signal downmixing unit 200 is similar to the sound signal downmixing device 200 of the second embodiment, and includes a delayed crosstalk addition unit 210, a left-right relationship information estimation unit 220, and a downmixing unit 230. The delayed crosstalk addition unit 210 performs the above-mentioned step S210, the left-right relationship information estimation unit 220 performs the above-mentioned step S220, and the downmixing unit 230 performs the above-mentioned step S230. In other words, the sound signal encoding device 300 includes the sound signal downmixing device 200 of the second embodiment as the sound signal downmixing unit 200, and performs the processing of the sound signal downmixing device 200 of the second embodiment as step S200.
[0046] [Encoding unit 340] The encoding unit 340 receives at least the downmix signal output by the sound signal downmixing unit 200. The encoding unit 340 encodes at least the input downmix signal to obtain and output a sound signal code (step S340). The encoding unit 340 may also encode the left channel input sound signal and the right channel input sound signal, and may output the code obtained by this encoding as part of the sound signal code. In this case, as indicated by the dashed lines in FIG. 5, the left channel input sound signal and the right channel input sound signal are also input to the encoding unit 340.
[0047] The encoding process performed by the encoding unit 340 may be any encoding process. For example, when the input downmix signal x M (1), x M (2), ..., x M (T) may be encoded using a monaural encoding method such as the 3GPP EVS standard to obtain an audio signal code. For example, in addition to encoding the downmix signal to obtain a monaural code, the left channel input audio signal and the right channel input audio signal may be encoded using a stereo encoding method corresponding to the stereo decoding method of the MPEG-4 AAC standard to obtain a stereo code, and the monaural code and the stereo code may be combined and output as an audio signal code. For example, in addition to encoding the downmix signal to obtain a monaural code, the left channel input audio signal and the right channel input audio signal may be encoded using a difference or weighted difference between the downmix signal and the left channel input audio signal for each channel to obtain a stereo code, and the monaural code and the stereo code may be combined and output as an audio signal code.
[0048] <Fourth embodiment> A signal processing device that processes an audio signal may include the audio signal downmixing device of the second embodiment as an audio signal downmixing section, and this configuration will be described as a fourth embodiment.
[0049] <Sound signal processing device 400> As shown in FIG. 7 , the sound signal processing device 400 of the fourth embodiment includes a sound signal downmixing unit 200 and a signal processing unit 450. The sound signal processing device 400 of the fourth embodiment processes input two-channel stereo time-domain sound signals in frame units of a predetermined time length, for example, 20 ms, to obtain and output the signal processing results. The two-channel stereo time-domain sound signals input to the sound signal processing device 400 are, for example, digital sound signals or audio signals obtained by capturing sounds such as speech or music with two microphones and performing AD conversion, or digital sound signals or audio signals obtained by processing the digital sound signals or audio signals, or digital decoded sound signals or decoded audio signals obtained by decoding stereo codes by a stereo decoding device, and are composed of a left channel input sound signal and a right channel input sound signal. The sound signal processing device 400 of the fourth embodiment performs the processes of steps S200 and S450 illustrated in FIG. 8 for each frame. The sound signal processing device 400 of the fourth embodiment will be described below with appropriate reference to the description of the second embodiment.
[0050] [Sound signal downmix unit 200] The sound signal downmixing unit 200 obtains and outputs a downmix signal from the left channel input sound signal and the right channel input sound signal input to the sound signal processing device 400 (step S200). The sound signal downmixing unit 200 is similar to the sound signal downmixing device 200 of the second embodiment, and includes a delayed crosstalk adding unit 210, a left-right relationship information estimating unit 220, and a downmixing unit 230. The delayed crosstalk adding unit 210 performs the above-mentioned step S210, the left-right relationship information estimating unit 220 performs the above-mentioned step S220, and the downmixing unit 230 performs the above-mentioned step S230. In other words, the sound signal processing device 400 includes the sound signal downmixing device 200 of the second embodiment as the sound signal downmixing unit 200, and performs the processing of the sound signal downmixing device 200 of the second embodiment as step S200.
[0051] [Signal processing unit 450] At least the downmix signal output by the sound signal downmix unit 200 is input to the signal processing unit 450. The signal processing unit 450 performs at least signal processing on the input downmix signal to obtain and output the signal processing result (step S450). The signal processing unit 450 may also perform signal processing on a left channel input sound signal and a right channel input sound signal to obtain the signal processing result. In this case, as indicated by the dashed lines in Fig. 7, the left channel input sound signal and the right channel input sound signal are also input to the signal processing unit 450, and the signal processing unit 450 performs signal processing on the input sound signal of each channel using the downmix signal, for example, to obtain the output sound signal of each channel as the signal processing result.
[0052] <Programs and recording media> The processing of each unit of the above-mentioned sound signal downmixing device, sound signal encoding device, and sound signal processing device may be implemented by a computer, in which case the processing content of the functions to be possessed by each device is described by a program. Then, by loading this program into the storage unit 1020 of the computer 1000 shown in Fig. 9 and running the arithmetic processing unit 1010, input unit 1030, output unit 1040, etc., the various processing functions of each of the above-mentioned devices are implemented on the computer.
[0053] The program describing the processing contents can be recorded on a computer-readable recording medium, such as a non-transitory recording medium, specifically a magnetic recording device, an optical disk, or the like.
[0054] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.
[0055] A computer that executes such a program, for example, first stores the program recorded on a portable recording medium or transferred from a server computer in its own non-transitory storage device, auxiliary storage unit 1050. Then, when executing a process, the computer loads the program stored in auxiliary storage unit 1050, its own non-transitory storage device, into storage unit 1020 and executes processing in accordance with the loaded program. Alternatively, as another form of execution of this program, the computer may load the program directly from a portable recording medium into storage unit 1020 and execute processing in accordance with the program. Furthermore, each time a program is transferred from a server computer to this computer, the computer may execute processing in accordance with the received program. Alternatively, the server computer may not transfer the program to this computer, but may instead execute the processing function by issuing an execution instruction and obtaining the results, thereby executing the above-described processing through a so-called ASP (Application Service Provider) type service. Note that the program in this embodiment includes information used for processing by a computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that define computer processing).
[0056] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0057] It goes without saying that other modifications are possible without departing from the spirit of the present invention.
Claims
1. 1. A sound signal downmixing method for obtaining a downmix signal that is a monaural sound signal from two-channel input sound signals, comprising: a delayed crosstalk addition step for obtaining, for each of the two channels, a signal obtained by adding an input sound signal of the channel and a signal obtained by delaying the input sound signal of the other channel and multiplying the delayed input sound signal by a weighting value that is a predetermined positive real value smaller than 1, as a delayed crosstalk-added signal of the channel; a left-right correlation information acquisition step for acquiring leading channel information, which is information indicating which of the delayed crosstalk-added signals of the two channels is leading, and a left-right correlation value, which is a value indicating the magnitude of correlation between the delayed crosstalk-added signals of the two channels; a downmixing step of obtaining a downmix signal in which the input sound signal of the preceding channel out of the input sound signals of the two channels is included to a greater extent as the left-right correlation value is larger, based on the left-right correlation value and the preceding channel information; A method for downmixing an audio signal, comprising:
2. 2. A sound signal downmixing method according to claim 1, comprising: The delay crosstalk adding step includes: The input sound signals of the two channels are defined as a left channel input sound signal and a right channel input sound signal, respectively, and the delayed crosstalk-added signals of the two channels are defined as a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, respectively. The sample number is defined as t, and each sample of the left channel input sound signal is defined as x L (t), and each sample of the right channel input sound signal is x R (t), and each sample of the left channel delayed crosstalk added signal is y L (t), and each sample of the right channel delayed crosstalk added signal is y R (t), and a predetermined positive value is a 1 , a 2 Let w be a predetermined positive real number whose absolute value is less than 1. 1 , w 2 As, Each sample y of the left channel delayed crosstalk added signal L (t) [Equation 17] obtained by Each sample y of the right channel delayed crosstalk added signal R (t) [Equation 18] obtained by A method for downmixing audio signals.
3. The audio signal downmixing method according to claim 1 or 2 includes an audio signal downmixing step, a mono encoding step of encoding the downmix signal obtained in the downmixing step to obtain a mono code; a stereo encoding step of encoding the two-channel input sound signals to obtain stereo codes; The sound signal encoding method further comprises:
4. 1. A sound signal downmixing device for obtaining a downmix signal that is a monaural sound signal from two-channel input sound signals, comprising: a delay crosstalk addition unit that, for each of the two channels, obtains a signal obtained by adding an input sound signal of the channel and a signal obtained by delaying the input sound signal of the other channel and multiplying the delayed input sound signal by a weighting value that is a predetermined positive real value smaller than 1, as a delay crosstalk-added signal of the channel; a left-right correlation information acquisition unit that acquires leading channel information, which is information indicating which of the delayed crosstalk-added signals of the two channels is leading, and a left-right correlation value, which is a value indicating the magnitude of correlation between the delayed crosstalk-added signals of the two channels; a downmix unit that obtains a downmix signal in which the input sound signal of the preceding channel of the input sound signals of the two channels is included to a greater extent as the left-right correlation value is larger, based on the left-right correlation value and the preceding channel information; a sound signal downmixing device including:
5. 5. A sound signal downmixing apparatus according to claim 4, The delay crosstalk addition unit The input sound signals of the two channels are defined as a left channel input sound signal and a right channel input sound signal, respectively, and the delayed crosstalk-added signals of the two channels are defined as a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, respectively. The sample number is defined as t, and each sample of the left channel input sound signal is defined as x L (t), and each sample of the right channel input sound signal is x R (t), and each sample of the left channel delayed crosstalk added signal is y L (t), and each sample of the right channel delayed crosstalk added signal is y R (t), and a predetermined positive value is a 1 , a 2 Let w be a predetermined positive real number whose absolute value is less than 1. 1 , w 2 As, Each sample y of the left channel delayed crosstalk added signal L (t) [Equation 21] obtained by Each sample y of the right channel delayed crosstalk added signal R (t) [Equation 22] obtained by Sound signal downmixer.
6. The audio signal downmixing device according to claim 4 or 5 is included as an audio signal downmixing unit, a monaural encoding unit that encodes the downmix signal obtained by the downmix unit to obtain a monaural code; a stereo encoding unit that encodes the two-channel input sound signals to obtain stereo codes; The sound signal encoding device further comprises:
7. A program for causing a computer to execute the processing of each step of the audio signal downmixing method according to claim 1 or 2.
8. A program for causing a computer to execute the processing of each step of the sound signal encoding method according to claim 3.
Citation Information
Patent Citations
Stereo signal encoding method, stereo signal encoding device, and program
JP2013003330A
Acoustic reproduction device and acoustic reproduction method
JP2015170926A
Sound coding device and sound coding method
WO2006070751A1
Down-mixing device, encoder, and method therefor
WO2010140350A1
Sound signal downmixing method, sound signal coding method, sound signal downmixing device, sound signal coding device, program, and recording medium
WO2021181746A1