Target mid-side signal for audio applications
Patent Information
- Application Number
- JP2024553721
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-03-03
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2043-03-03
Smart Images

Figure 0007927078000025 
Figure 0007927078000026 
Figure 0007927078000027
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims priority to the following priority applications: U.S. Provisional Patent Application No. 63 / 318,226 filed on 9 March 2022, U.S. Provisional Patent Application No. 63 / 423,786 filed on 8 November 2022, and European Patent Application No. 22183794.1 filed on 8 July 2022.
[0002] This invention relates to source separation, signal enhancement, and signal processing. Furthermore, this invention relates to a method for extracting a target mid-range audio signal and a target side-range audio signal from a stereo audio signal. The invention further relates to a processing device using the aforementioned method. [Background technology]
[0003] In the field of audio processing, audio source separation is important in several applications. In one simple application, the separated speech audio signal is retained or given additional gain, while the background audio signal is excluded or attenuated relative to the speech audio signal. This can improve the intelligibility of the speech audio signal.
[0004] In audio source separation, the signal that is the target of estimation or extraction is called the target signal. The target signal is not limited to speech; it can be any general audio signal (such as an instrument), or even multiple audio signals that must be separated from an audio mix that contains noise and / or other additional audio signals that are not the target.
[0005] For a stereo audio signal containing two audio signals associated with the left and right audio channels, one processing method that implicitly performs source separation is to extract a mid audio signal and a side audio signal from the left and right audio signals, respectively, where the mid and side audio signals are proportional to the sum and difference of the left and right audio signals, respectively. The mid signal emphasizes audio components that are equal in magnitude and in phase between the left and right audio signals, while the side audio signal attenuates or removes such signals. Thus, calculating the mid and side signals from the left and right audio signals constitutes an efficient method of enhancing or attenuating in-phase, center-panned sources, respectively. The mid and side signals can then be further inversely converted back into the left and right (conventional stereo) audio signals.
[0006] It should be noted that, insofar as the target is a center-panned source, source separation is implicitly performed. That is, the mid signal includes this source, while the side signals do not. Side audio signals mainly consist of potentially uninteresting background audio signals, and attenuating or excluding these background audio signals can improve the intelligibility of the center-panned in-phase audio source.
[0007] A drawback of conventional mid and side signal calculations is that implicit source separation based on mid / side audio signal separation fails when the desired audio source is not center-panned. To address this, specific mid and side signal calculation techniques have been developed for stereo audio signals with specific panning characteristics. These techniques can effectively target stationary, non-center-panned audio sources. See, for example, "Stereo Music Source Separation via Bayesian Modeling" by Master, Aaron (Ph.D. dissertation, Stanford University, 2006).
[0008] While these techniques mitigate some of the shortcomings of basic mid- and side-signal extraction, such pan-specific extraction techniques still cannot isolate more common audio sources present in a stereo audio signal, such as moving audio sources, sources with reverberation, or sources that are spatially dominant at various points in the time-frequency space. [Overview of the Initiative] [Problems that the invention aims to solve]
[0009] Therefore, an object of this disclosure is to provide such an improved method and an audio processing system that performs enhanced audio source separation on a stereo audio signal. [Means for solving the problem]
[0010] According to a first aspect of the present invention, a method is provided for extracting a target mid-range audio signal from a stereo audio signal, wherein the stereo audio signal includes a left audio signal and a right audio signal. The method includes the steps of: obtaining a plurality of consecutive time segments of the stereo audio signal, each time segment comprising a representation of a portion of the stereo audio signal; and obtaining at least one of a target pan parameter and a target phase difference parameter for each frequency band of a plurality of frequency bands of each time segment of the stereo audio signal. The target pan parameter represents the distribution of the magnitude ratio between the left audio signal and the right audio signal over a time segment in a frequency band, and the target phase difference parameter represents the distribution of the phase difference between the left audio signal and the right audio signal of the stereo audio signal over a time segment.
[0011] The method further comprises extracting a partial mid-range signal representation for each time segment and each frequency band, the partial mid-range signal representation being based on a weighted sum of a left audio signal and a right audio signal, the respective weights of the left audio signal and the right audio signal being based on at least one of the target pan parameter Θ and the target phase difference parameter for each frequency band and each time segment, and forming a target mid-range audio signal by combining the partial mid-range signal representations for each frequency band and each time segment.
[0012] Obtaining at least one of the target pan parameter and / or target phase difference parameter may include receiving the target pan parameter and / or target phase difference parameter, accessing those parameters, or requesting those parameters. At least one of the target pan parameter and / or target phase difference parameter may be replaced with a default value for at least one time segment and frequency band.
[0013] Consecutive time segments refer to segments that represent parts of an audio signal in time, where later time segments represent later parts of the audio signal, and earlier time segments represent earlier parts of the audio signal. Consecutive time segments may or may not overlap in time.
[0014] A portion of a stereo audio signal can be represented in any time-domain or frequency-domain format. The frequency-domain representation can be any linear time-frequency domain representation, such as a Short-Time Fourier Transform (STFT) representation or a Quadrature Mirror Filter (QMF) representation.
[0015] This invention is at least partially based on the understanding that, by obtaining target pan parameters and / or target phase difference parameters for each time segment and each frequency band, it is possible to extract a target mid-audio signal that always targets a time-varying and / or frequency-varying audio source in a stereo audio signal. Furthermore, since the target mid-audio signal is generated using individual target pan parameters for each time segment and each frequency band, this method enables the extraction of a target mid-audio signal that simultaneously targets two or more frequency-separated audio sources in a stereo audio signal.
[0016] In some embodiments, the weights of the left and right audio signals are based on target pan parameters such that the left or right audio signal with the larger magnitude is given a greater weight.
[0017] In other words, the left or right audio signal associated with a greater magnitude or power (indicated by the target pan parameter) is given a greater weight and contributes more to the formation of the target mid-audio signal.
[0018] In some embodiments, the method further includes extracting partial side signal representations for each time segment and each frequency band, the partial side signal representations being based on a weighted difference between a left audio signal and a right audio signal, the respective weights of the left audio signal and the right audio signal being based on at least one of the target pan parameter and the target phase difference parameter for each frequency band and each time segment, and forming a target side audio signal by combining each partial side signal representation for each frequency band and each time segment.
[0019] In other words, the target side audio signal can be formed in parallel with the target mid audio signal, and the target mid audio signal and the target side audio signal form a complete representation of the stereo audio signal. It is understood that any embodiment relating to the formation or processing of the target mid audio signal can be performed in the same way as the formation and processing of the target side audio signal as described below.
[0020] If a stereo audio signal includes, for example, an ideally center-panned audio source, or a single audio source panned under the constant power law, the target mid audio signal ideally captures all the signal energy of the audio source, while the target side audio signal does not capture such energy. However, in the case of a typical audio source(s), the target mid audio signal captures the target audio source(s)(s) and possibly some other non-target sounds, while the target side audio signal captures the non-target sounds and likely removes or nearly removes the target source(s). By including a target side audio signal in addition to the target mid audio signal, it is ensured that all audio signals of the left and right audio signal pair are present in the target mid and target side audio signal pair. For example, this allows for lossless reconstruction of the left and right audio signals.
[0021] Any function described in relation to a method may have a corresponding feature in a system or device, and vice versa.
[0022] The present invention will be described in more detail with reference to the accompanying drawings illustrating preferred embodiments of the present invention. [Brief explanation of the drawing]
[0023] [Figure 1] This is a block diagram showing an analysis device according to several embodiments.
[0024] [Figure 2] This is a flowchart showing a method according to one embodiment.
[0025] [Figure 3]This is a block diagram showing alternative analytical devices according to several embodiments.
[0026] [Figure 4a] This figure shows a set of target pan parameters and target phase difference parameters distributed across multiple time segments and frequency bands, according to several embodiments.
[0027] [Figure 4b] This figure shows a set of partial mid-range signal representations that form a target mid-range audio signal, according to several embodiments.
[0028] [Figure 4c] This figure shows a set of partial side signal representations that form a target side audio signal, according to several embodiments.
[0029] [Figure 5a] This is a block diagram of an analysis device according to several embodiments.
[0030] [Figure 5b] This is a block diagram of an analysis system using alternative pan parameters and alternative phase difference parameters according to several embodiments.
[0031] [Figure 6a] This is a block diagram of an audio processing system having a mono-source separator and a mid / side processor, according to several embodiments.
[0032] [Figure 6b] This is a block diagram of another audio processing system having a stereo processor that operates on alternate left audio signals and alternate right audio signals, according to several embodiments.
[0033] [Figure 6c]This is a block diagram of yet another audio processing system that applies an estimated stereo soft mask to a target mid-audio signal and reconstructs the processed left and right audio signals, according to several embodiments. [Modes for carrying out the invention]
[0034] The systems and methods disclosed in this application can be implemented as software, firmware, hardware, or a combination thereof. In the hardware embodiments, task partitioning does not necessarily correspond to partitioning into physical units. Conversely, a single physical component may have multiple functions, and a single task may be performed collaboratively by several physical components.
[0035] Computer hardware can be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify the actions to be performed by such computer hardware. Furthermore, this disclosure relates to any set of computer hardware that individually or collectively execute instructions that perform any one or more of the concepts discussed herein.
[0036] Certain or all components can be executed by one or more processors, which receive computer-readable (also called machine-readable) code containing a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. This includes any processor capable of executing a set of instructions (sequential or otherwise) that specify the action to be performed. Thus, one example is a typical processing system (i.e., computer hardware) comprising one or more processors. Each processor may include one or more of the following: CPU, graphics processing unit, and programmable DSP unit. The processing system may further include a memory subsystem, which may include a hard drive, SSD, RAM, and / or ROM. A bus subsystem for communication between components may also be included. Software may reside in the memory subsystem and / or may reside in the processors while the computer system is executing the software.
[0037] One or more processors can operate as standalone devices, or they can connect to other processors (one or more), and can, for example, be connected to a network. Such networks can be built on a variety of different network protocols and can be the internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0038] Software can be distributed across computer-readable media, which may include computer storage media (i.e., non-temporary media) and communication media (i.e., temporary media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile media, removable and non-removable media, implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, various forms of physical (non-temporary) storage media, such as EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store desired information and that can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media (temporary) typically include any information distribution media, which embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transport mechanisms.
[0039] The following description includes an introductory section on the extraction of basic mid- and side-audio signals, along with a detailed description of a preferred embodiment at present.
[0040] Extraction of basic mid-range and side-range audio signals A stereo audio signal, including the left audio signal L and the right audio signal R, can also be represented as a mid audio signal M and a side audio signal S. The basic mid audio signal M and side audio signal S are constructed from the left audio signal L and the right audio signal R using the following analytical formulas: M=0.5L+0.5R Equation 1 S=0.5L-0.5R Formula 2 The original left audio signal L and right audio signal R can be reconstructed from the mid audio signal M and side audio signal S using the following synthesis formula. L=M+S formula 3 R=MS Equation 4
[0041] The mid audio signal M enhances the audio signal features that are in phase and located in the center of the stereo mix, while the side audio signal S attenuates the audio signal features that are in phase. For example, if a stereo audio signal contains a pair of center-panned left audio signals L and right audio signals R (equal magnitude and in-phase components in each of the left and right audio signals), the mid audio signal M will include the stereo audio signal, while the side audio signal S will remove this stereo audio signal. The basic mid / side audio signals then target the center-panned audio signal that is included in the mid audio signal M but not in the side audio signal S.
[0042] Stereo signals in the time domain and time-frequency domain The left audio signal L, right audio signal R, mid audio signal M, and side audio signal S can be represented in the time domain or the frequency domain. The frequency domain representation can be, for example, a short-time Fourier transform (STFT) representation or an orthogonal modulation filter (QMF) representation. For example, the left audio signal L and the right audio signal R can be represented in the frequency domain by the following equations.
number
number
number
[0043] The detected magnitude parameter U and the detected pan parameter θ form a representation of the stereo audio signal called the stereo-polar magnitude (SPM) representation. While the SPM representation can be replaced with any equivalent representation, the SPM representation of the stereo audio signal is used in the following explanation.
[0044] The detected pan parameter θ is in the range of 0 to π / 2, where 0 indicates a stereo audio signal where only the left audio signal L is non-zero, and π / 2 indicates a stereo audio signal where only the right audio signal R is non-zero. A detected pan parameter value of θ = π / 4 indicates a centered stereo audio signal where the magnitudes of the left audio signal L and the right audio signal R are equal. The detected pan parameter is obtained from the stereo input signal, but mathematically it also coincides with the source mixing model. Magnitude │S x │, phase ψ x , and known pan parameters Θ x Audio source S x Regarding this, the left audio signal L and the right audio signal R can be expressed by the following equations.
number
number
[0045] Mid-audio and side-audio signals with fixed panning The basic mid audio signal M and side audio signal S are extracted by equally weighting the left audio signal L and the right audio signal R, respectively, to target an audio source (a centrally located audio source) that is assumed to be present in equal proportions to the left audio signal L and the right audio signal R. For audio sources that are not centrally located, the weighting of the left audio signal L and the right audio signal R in equations 1 and 2 should be adjusted.
[0046] For example, the coefficients of the left audio signal L and the right audio signal R in equations 1 and 2 can be changed from 0.5 under the constraint that the sum of the coefficients must equal 1. This allows, for an audio source that appears predominantly in the left audio signal, the coefficients of the left audio signal L and the right audio signal R in equation 1 to be set to, for example, 0.7 and 0.3 (the same applies to the configuration of the side signal in equation 2). However, this results in the undesirable situation that when the audio source moves between the left audio signal L and the right audio signal R (assuming constant power law mixing), the magnitudes of the mid audio signal M and the side audio signal S change considerably.
[0047] To avoid this problem, a known or estimated "target" panning parameter Θ x for an audio source S x , the mid audio signal M and the side audio signal S can be extracted from the left audio signal L and the right audio signal R according to the following equations. M=cos(Θ x )L+sin(Θ x )R Equation 14 S=sin(Θ x )L-cos(Θ x )R Equation 15 In this case, the left audio signal L and the right audio signal R can be reconstructed as the following equations. L=cos(Θ x )M+sin(Θ x )S Equation 16 R=sin(Θ x )M-cos(Θ x )S Equation 17
[0048] For example, when Θ x = 0, which indicates that the audio source appears only in the left audio signal, the mid audio signal M is equal to the left audio signal L, the side audio signal S is equal to the right audio signal R, and Θ xThe opposite is true when =π / 2. Similarly, when the audio source appears with equal magnitude to the left audio signal L and the right audio signal R (i.e., Θ x The mid-audio signal M (=π / 4) is the left audio signal L and the right audio signal R, weighted equally. Furthermore, when the audio source is located to the left of the center, the left audio signal L and the right audio signal R are in the range of Θ 0 to π / 4. x Weighted using Θ, when the audio source is located to the right of the center, the left audio signal L and the right audio signal R are in the range of π / 4 to π / 2. x The coefficients used to weight the left audio signal L and the right audio signal R to construct the mid audio signal M and the side audio signal S, and the coefficients used to perform the reverse construction, are thus fixed target pan parameter Θ. x It is based on.
[0049] Extraction of target mid-range audio signal and target side-range audio signal One aspect of the present invention relates to the generation of a target mid audio signal M and / or a target side audio signal S based on at least one of a target pan parameter Θ and a target phase difference parameter Φ obtained for each time segment and each frequency band of a stereo audio signal. The target pan parameter Θ represents an estimated or indicated value of the magnitude ratio between the left audio signal L and the right audio signal R in each time segment and each frequency band of sound, corresponding to one or more target sources. The target phase difference parameter Φ represents an estimated or indicated value of the phase difference between the left audio signal L and the right audio signal R in each time segment and each frequency band of sound, corresponding to one or more target sources.
[0050] In some embodiments, the pan parameter θ and / or phase difference parameter Φ for each time segment and each frequency band 111, 112 are the median, mean, mode, numbered percentile, maximum or minimum pan parameter θ and / or phase difference parameter Φ of the time segment. Generally, the detected pan parameter θ and / or detected phase difference parameter φ exist for each sample of the stereo audio signal. On the other hand, the target pan parameter θ and / or target phase difference parameter Φ used to generate the target mid audio signal M and the target side audio signal S are not necessarily of such fine granularity and may represent statistical features (e.g., mean) of multiple samples, such as the mean of all samples within a given time segment. For example, the detected pan parameter θ and / or detected phase difference parameter φ may assume many different values more than 1000 times per second, while the target pan parameter θ and / or phase difference parameter Φ may change only a few times per second.
[0051] Figure 1 shows an analysis device 10 that extracts a target mid-audio signal M and a target side audio signal S from a stereo signal containing a left audio signal L and a right audio signal R. The extraction of the mid-audio signal M and the side audio signal S is performed in the extraction unit 12 of the analysis device 10, which acquires the left audio signal L and the right audio signal R along with target pan parameters Θ and / or target phase difference parameters Φ for multiple time segments 100 and frequency bands 111, 112.
[0052] The extraction unit 12 acquires the left audio signal L and the right audio signal R (stereo audio signal) as multiple consecutive time segments. Each time segment contains a representation of a portion of the left audio signal L and the right audio signal R. Alternatively, the extraction unit 12 is configured to divide the left audio signal L and the right audio signal R into multiple consecutive time segments. Each time segment contains two or more frequency bands of the stereo audio signal, and each frequency band in each time segment is associated with at least one of the acquired target pan parameter Θ and the acquired target phase difference parameter Φ.
[0053] In the embodiment shown in Figure 1, the analysis device 10 receives the left audio signal L and the right audio signal R, and the target pan parameter Θ and / or target phase difference parameter Φ for each frequency band 111, 112 of each time segment. The frequency band numbers range from band 1 to band B. Θ (t,1) and Φ (t,1) This represents the target pan parameter Θ and target phase difference parameter Φ of the first frequency band 111 of time segment t, while Θ (t,2) and Φ (t,2) This represents the target pan parameter and target phase difference parameter for the second frequency band 112 of time segment t.
[0054] Using the left audio signal L and the right audio signal R, along with the target pan parameter Θ and / or the target phase difference parameter Φ, the extraction unit 12 extracts the target mid audio signal M and the target side audio signal S. In some embodiments, the extraction unit 12 extracts only one of the target mid audio signal M and the target side audio signal S, such as only the mid audio signal M.
[0055] To enable the modeling of the inter-channel phase difference between the left audio signal L and the right audio signal R, the true source phase of the target audio source is defined as being proportionally closer to the left or right audio signal that exhibits greater magnitude. That is, in the case of a stereo audio signal where the left audio signal L is dominant in power or magnitude, the phase of the left audio signal L is closer to the true source phase of the audio source compared to the phase of the right audio signal R, and the opposite is true in the case of a stereo audio signal where the right audio signal R is dominant. Audio source S x However, it is modeled to appear in the left audio signal L and the right audio signal R according to the following mixing formula.
number
[0056] Based on the mixing models described in equations 17 and 18 above, the extraction device 12 can target a source having a target pan parameter Θ and a target phase difference parameter Φ using the following analytical formulas.
number
[0057] In Equation 20, the target mid-audio signal M is based on the weighted sum of the left audio signal L and the right audio signal R, and the weight of the left audio signal is The file is JPEG0007927078000013.jpg527, and the weight of the right audio signal is The image is JPEG0007927078000014.jpg525, and the pan parameter Θ and / or phase difference parameter Φ are obtained for each time segment and frequency band 111, 112.
[0058] Similarly, in Equation 20, the target side audio signal S is based on the weight difference between the left audio signal L and the right audio signal R, and the weight of the left audio signal is The file is JPEG0007927078000015.jpg526, and the weight of the right audio signal is The image is JPEG0007927078000016.jpg525, and the pan parameter Θ and / or phase difference parameter Φ are obtained for each time segment and frequency band 111, 112.
[0059] The weights of the left audio signal L and the right audio signal R in the weighted sum of Equation 20 and the weighted difference of Equation 21, used to extract the mid audio signal M and the side audio signal S, are complex values, with a real-valued magnitude coefficient, e.g., cos(Θ), and a complex-valued phase coefficient, e.g., Includes JPEG0007927078000017.jpg517. The acquired target pan parameter Θ and / or target phase difference parameter Φ influence the weights so that the target mid audio signal M and / or target side audio signal S can be extracted.
[0060] For example, the real-valued magnitude coefficient can be based on Θ such that the left or right audio signal with the larger magnitude is given a greater weight in the weighted sum. The complex-valued coefficient can also be based on Φ such that a larger phase difference Φ means a larger difference exists between the phases of each weight in the weighted sum. In some embodiments, the complex-valued coefficient of each weight in the weighted sum is based on both Φ and Θ such that the weight of the left audio signal L or right audio signal R with the larger magnitude is given a smaller phase, for example, one of the weights The filename is JPEG0007927078000018.jpg517.
[0061] Similarly, the real-valued magnitude coefficients in the weighted difference (for target-side signal extraction) can be based on a target pan parameter Θ such that the left or right audio signal with the larger magnitude is given a smaller weight in the weighted difference. Alternatively, the complex-valued coefficients can be based on Φ such that a larger phase difference Φ means a larger difference exists between the phases of each weight in the weighted difference. In some embodiments, the complex-valued coefficients of each weight in the weighted difference are based on both Φ and Θ such that the weight of the left or right audio signal with the larger magnitude is given a smaller phase, for example, one of the weights is The filename is JPEG0007927078000019.jpg526.
[0062] In some embodiments, only one of the target pan parameter Θ and the target phase difference parameter Φ is obtained for at least one time segment and frequency band 111, 112, while the other is assigned a default value. For example, if only the target pan parameter Θ is obtained, the target phase difference parameter Φ is assumed to be Φ=0, and if only the target phase difference parameter Φ is obtained, the target pan parameter Θ is assumed to be Θ=π / 4. These default values assume an audio source in the center of a stereo audio signal with no inter-channel phase difference, but other default values are also possible.
[0063] The extraction device 12 extracts the partial mid-signal component or representation M of each frequency band in each time segment. (t,n) By combining these, the target mid-audio signal M is obtained. Similarly, the extraction device 12 obtains a partial side signal representation S of each frequency band in each time segment. (t,n) By combining these, the target side audio signal S is obtained. Partial mid signal component or representation M (t,n) This is generated for each frequency band 1...B in time segment t, and represents all partial mid-signal representations of time segment t M (t,n) By combining these, the target mid-audio signal time segment M (t) A target mid-audio signal time segment M is generated. (t) The sequence forms the target mid-range audio signal M. The target side-range audio signal S is also formed in a similar manner. Partial mid-range signal representation M in frequency band 1...B (t,n) Combining these results in a partial target mid-audio signal M using a composite filter bank. (t,n) This may include processing.
[0064] The method performed by the analysis device 10 will be described in detail with further reference to the flowchart in Figure 2. A stereo audio signal is acquired by the analysis device 10 in step S1, and the target pan parameter Θ and / or target phase difference parameter Φ are acquired by the analysis device 10 for multiple time segments 100 and frequency bands 111, 112 in step S2.
[0065] In step S3, partial mid-signal representation M (t,n) However, the mid / side extractor 12 extracts the signal (for example, according to equation 19 above), and this partial mid signal representation is based on a weighted sum of the left audio signal L and the right audio signal R. The respective weights of the left audio signal L and the right audio signal R are based on at least one of the target pan parameter Θ and the target phase difference parameter Φ for each frequency band 111, 112 and each time segment. Similarly, the partial side signal representation S (t,n) However, extracted in step S31 (for example, according to equation 20 above), this partial side signal representation is based on the weighted difference between the left audio signal L and the right audio signal R.
[0066] In some embodiments, the target side audio signal S is not calculated, and the extractor unit 12 is configured, for example, to extract only the target mid audio signal M. Alternatively, the target side audio signal S is calculated but reduced to a small value or 0 in step S51 before the reconstruction of the left audio signal L and right audio signal R in step S6 (for example, using reconstruction formulas 21 and 22 below). In some cases, the target mid audio signal M is expected to capture substantially all of the energy of the target audio source, so the target side audio signal S may mainly consist of undesirable noise or background audio that can be excluded or attenuated.
[0067] Following steps S3 and S31, the method proceeds to steps S4 and S41. These steps involve partial mid-signal representation M (t,n) and partial side signal representation S (t,n) The method includes forming a target mid-audio signal M and a target side audio signal S by combining these into a target mid-audio signal M and a target side audio signal S, respectively. The method then proceeds with steps S5 and S51, which include processing the target mid-audio signal M and / or the target side audio signal S, examples of which include attenuating the target side audio signal S or providing the target mid-audio signal M to a mono-source separator, as will be described below with reference to Figures 6a, 6b and 6c.
[0068] Finally, the method proceeds to step S6, which includes reconstructing the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S. The reconstruction of the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S is performed by a synthesizer, as will be described in more detail below with reference to Figures 5a and 5b.
[0069] As an example of the operation of the analysis device 10, consider an audio source with a left audio signal L and a right audio signal R having time-varying pan. This audio source is panned at a constant speed from completely right to completely left under the constant power law as time t progresses from 0 seconds to 10 seconds. For example, in the case of the basic mid-audio signal and side audio signal extracted for this audio source according to equations 1 and 2 above, the mid-audio signal M captures substantially all of the energy of the audio source at t=5 and substantially does not capture energy at t=0 and t=10, while the side audio signal S captures substantially all of the energy of the audio source at t=0 and t=10 and substantially does not capture energy at t=5. The target mid-audio signal M, extracted using the varying pan parameter Θ and / or phase difference parameter Φ for each time segment and each frequency band, contains substantially all of the audio source energy for any t∈[0,10], while the target side audio signal S contains substantially none of the time-varying energy of the audio source.
[0070] The constant power law ensures that the audio signal power (proportional to the sum of the squares of the amplitudes) of the left audio signal L and the right audio signal R remains constant for all pan angles. For example, the audio signal amplitudes of the left and right audio signals are scaled using a scaling factor dependent on the pan angle to ensure that the square of the sum of the left and right audio signals is equal for all pan angles. Unlike the linear panning law, which ensures that the sum of the signal amplitudes is constant and allows the total signal power to change with the pan angle, the constant power panning law ensures that the perceived audio signal power is constant for all pan angles. The constant power law is described, for example, in "Loudness Concepts & Pan Laws" (Introduction to Computer Music, Carnegie Mellon University) by Anders Oeland and Roger Dannenberg.
[0071] This demonstrates how the target mid-audio signal M and target side audio signal S can target and isolate one or more audio sources in a stereo audio signal whose time and frequency are changing. Furthermore, if a second audio source of a different frequency is moved at a different speed than the first audio source between the left audio signal L and the right audio signal R (or, for example, stationary in the center pan), the target mid-audio signal M targets both the first and second audio sources in their respective frequency bands. This means that, despite both audio sources being of different frequencies and shifting at different speeds between the left audio signal L and the right audio signal R, virtually all of the energy from both audio sources is present in the target mid-audio signal M.
[0072] Figure 3 shows a combinatorial analyzer 10' configured to extract the target mid-audio signal M and target side audio signal S according to the above, and also to extract or determine the target pan parameter Θ and / or target phase difference parameter Φ for each frequency band 111, 112 and each time segment of the left audio signal L and right audio signal R using a parameter extractor unit 14. Thus, the target pan parameter Θ and / or target phase difference parameter Φ are obtained together with the left audio signal L and right audio signal R, or determined from the left audio signal L and right audio signal R. An example method of doing this is described below.
[0073] The target phase difference parameter Φ can be calculated, for example, by a parameter extractor 14, as the typical phase difference or dominant phase difference between the left audio signal L and the right audio signal R in each time segment and each frequency band 111, 112. For example, in the time-frequency domain (e.g., the STFT domain), the detected phase difference can be calculated for each STFT tile as φ = Arg(R / L), and by analyzing the distribution with respect to φ, the value of Φ can be estimated to be the dominant value or typical value for that time segment and frequency band. Similarly, the detected pan parameter θ can be calculated for each STFT tile as θ = arctan(│R│ / │L│), and by analyzing the distribution with respect to θ, the value of θ can be estimated to be the dominant value or typical value for that time segment and frequency band. Here, L and R represent the left audio signal L and the right audio signal R in the time segment and frequency bands 111, 112. The standard or dominant values estimated from the distribution can be the median, mean, mode, numbered percentiles, minimum, or maximum.
[0074] In "Dialog Enhancement via Spatio-Level Filtering and Classification" (Convention Paper 10427 of the 149th Convention of the Audio Engineering Society) by Master, A et al., a method is proposed for obtaining indicators of the average value and spread of pan and inter-channel phase difference labeled thetaMid, thetaWidth, phiMid, and phiWidth. The pan parameter Θ used by the analysis device 10 can be based on thetaMiddle proposed in this conference paper, and similarly, the target phase difference parameter Φ used by the analysis device 10 can be based on the phiMiddle of the so-called "Shift and Squeeze" parameter or "S&S" parameter proposed in this paper.
[0075] For each time segment and each frequency band within that time segment, the audio processing system calculates the magnitude U squared, i.e., U 2A 51-bin histogram can be created for θ weighted by . The system does the same for the detected pan parameter φ, and a version of φ in the range of 0 to 2*pi called φ2, however, each of these histograms uses 102 bins. Each histogram is smoothed over the time segment over their given dimensions. For the smoothed θ histogram, the system detects the target pan parameter as the highest peak called thetaMiddle, as well as the width around this peak called thetaWidth, which is needed to capture 40% of the energy in the histogram. The system does the same for φ and φ2, recording phiMiddle, phi2Middle, phiWidth, and phi2Width, but for width, which requires 80% energy capture. The system records the final values of phiMiddle and phiWidth based on which has a higher concentration in phi space, as indicated by the smaller phiWidth value.
[0076] For example, the parameter extractor 14 in Figure 3 extracts the detected pan parameter θ for each time segment and each frequency band within that time segment from the total signal power U from Equation 7. 2A histogram is created using the weighted values, and the same is done for the detected phase difference parameter φ. These histograms can have any number of bins, such as 10 or more bins, or 50 or more bins. In one embodiment, each histogram has 51 bins. Each histogram is smoothed using a smoothing function suitable for at least the time dimension and / or time and the θ or φ dimension. For the smoothed θ histogram, the system detects the θ value of the highest peak, called thetaMiddle, and also the width around this peak, called thetaWidth, which is necessary to capture 40% of the energy in the histogram. The same is done for φ, and phiMiddle and phiWidth are recorded, but the width requires 80% energy capture. It is understood that the parameter extractor 14 can be configured to either obtain the detected pan parameter θ and / or the detected phase difference parameter φ and / or the detected magnitude U, or to determine the detected pan parameter θ and / or the detected phase difference parameter φ and / or the detected magnitude U from the left audio signal L and the right audio signal R (for example, using Equation 7).
[0077] Details of time segments and frequency bands Referring to Figure 4a, multiple time segments 110, 120 are shown, extending from the first time segment to time segment t. Each time segment 110, 120 contains multiple frequency bands 111, 112, 121, 122, extending from the first frequency band to frequency band B. The time segments 110, 120 and the frequency bands 111, 112, 121, 122 form a time-frequency tile representation, where each frequency band of each time segment forms its own time / frequency tile.
[0078] As shown in Figure 4a, each frequency band 111, 112, 121, and 122 of each time segment 110 and 120 is associated with its respective target pan parameter Θ and / or target phase difference parameter Φ. At least two frequency bands 111, 112, 121, and 122 (or subbands) are defined for each time segment 110 and 120, such as a high-frequency band containing frequencies above the threshold frequency and a low-frequency band containing frequencies below the threshold frequency. That is, each time segment 110 and 120 includes frequency band 1...B, where B is at least 2. Experiments have shown that octave frequency bands or quasi-octave frequency bands are suitable for targeting dialogue audio sources such as frequency bands with edges at 0Hz, 400Hz, 800Hz, 1600Hz, 3200Hz, 6400Hz, 13200Hz, and 24000Hz. However, different bandwidth distributions are also possible.
[0079] The duration of each time segment can correspond to 10ms to 1s of a stereo audio signal. In some embodiments, each time segment corresponds to multiple frames of a stereo audio signal, such as 10 frames, with each frame representing a portion of the stereo audio signal, such as 10ms to 200ms, or 50ms to 100ms. Experiments have shown that representing one time segment using 10 overlapping frames, each 50ms to 100ms long with 75% overlap, represents a good trade-off between fast response time, parameter stability, parameter reliability, and computational cost for typical dialogue sources. However, for non-dialogue audio sources, such as music, longer or shorter duration time segments are also possible.
[0080] Figure 4b shows the partial mid-audio signal representation M extracted for each time segment 210, 220 and each frequency band 211, 212, 221, 222. (t,b)This shows that, as seen in this figure, the first time segment 210 is a partial mid-audio signal representation M of the first frequency band 211. (1,1) , second partial mid-audio signal representation M of the second frequency band 212 (1,1) The last mid-range audio signal representation M in frequency band B, etc. (1,B) This includes up to the first time segment 210, a partial mid-audio signal representation M. (1,1) ...M (1,B) By combining these, the first time segment M of the target mid-audio signal M (1) This is generated. Similarly, the subsequent time segment M of the target mid-audio signal M is generated. (2) ...M (t) However, the subsequent time segment 220's partial mid-audio signal representation M (1,1) ...M (1,B) It is generated by combining these. Finally, each time segment M (1) ...M (t) By combining these elements in the correct order, the target mid-audio signal M is generated.
[0081] Further reference to Figure 4c, the partial side audio signal representation S extracted for each time segment 310, 320 and each frequency band 311, 312, 321, 322. (t,b) This is shown. The partial mid-audio signal representation M described above (t,b) In a similar manner to the combination, a partial side audio signal representation S (t,b) These are combined to form the target side audio signal S. Partial mid audio signal representation M (t,b) and / or partial side audio signal representation S (t,b) It is understood that the extraction and combination can be performed in a single step by a single processing unit, or as two or more steps performed by two or more processing units.
[0082] Reconstruction of left and right audio signals from target mid audio signal and target side audio signal. Figure 5a shows a block chart illustrating a synthesis unit 20 that performs the reconstruction of the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S. The synthesis unit 20 comprises a reconstructor unit 22 that reconstructs the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S using the pan parameter Θ and / or the phase difference parameter Φ. For example, the reconstructor unit 22 implements the following synthesis formulas for the target mid audio signal M and the target side audio signal S extracted using equations 20 and 21.
number
[0083] As seen in equations 22 and 23, the reconstruction of the left audio signal L and the right audio signal R is based on the weighted sum and weighted difference of the target mid audio signal M and the target side audio signal S. However, these equations can be modified to downweight or remove the contribution from the side signals. For example, the left audio signal L and the right audio signal R are based on the left signal... Mid weight coefficient and right signal of JPEG0007927078000021.jpg526 The mid-weighting coefficients in JPEG0007927078000022.jpg526 can be used to reconstruct the signal from the target mid-audio signal M alone. Here, these weighting coefficients are based on at least one of the target pan parameter Θ and the target phase difference parameter Φ. In this example, the contribution from the side signal is 0. Other contributions between 0 and the values shown in equations 22 and 23 are also possible.
[0084] Similarly, the left audio signal L and the right audio signal R can also be reconstructed by taking into account the target side audio signal S. In such a case, the left audio signal L is based on the weighted sum of the target mid audio signal M and the target side audio signal S, and the weights JPEG0007927078000023.jpg562 is based on at least one of the pan parameter and the phase difference parameter. In contrast, the right audio signal R is based on the weighted difference between the target mid audio signal M and the target side audio signal S, and the weight JPEG0007927078000024.jpg562 is based on at least one of the pan parameter Θ and the phase difference parameter Φ. The contribution from the mid-signal can also be down-weighted or reduced to 0.
[0085] Figure 5b shows the alternative pan parameter Θ. ALT and / or alternative phase difference parameter Φ ALT Reconstruction is performed using equations 22 and 23, which have the following characteristics: As a result, the alternative left audio signal L ALT and alternative right audio signal R ALT This is a block chart showing a synthesis apparatus 20 that provides the signal. In the illustrated embodiment, the left audio signal L located in the center has no inter-channel phase difference. CEN and right audio signal R CEN To generate a center-panned stereo audio signal including Θ for all time segments and frequency bands, ALT =Θ CEN=π / 4 and Φ ALT =Φ CEN =0.
[0086] Alternative mid weighting factors and alternative side weighting factors are defined by replacing Θ and Φ of the mid weighting factors and side weighting factors described in Equations 22 and 23 with the corresponding alternative parameters Θ ALT , Φ ALT obtained for each time segment and each frequency band. In some embodiments, the alternative left audio signal L ALT and the alternative right audio signal R ALT are generated using only the target mid audio signal M and the alternative mid weighting factors (e.g., the coefficients for the target mid audio signal M in Equations 22 and 23 having Θ ALT , Φ ALT in place of Θ, Φ). Alternatively, the alternative left audio signal L ALT and the alternative right audio signal R ALT are generated using both the target mid audio signal M and the target side audio signal S, and the alternative mid weighting factors and alternative side weighting factors (e.g., the coefficients for the target mid audio signal M and the target side audio signal S in Equations 22 and 23 having Θ ALT , Φ ALT in place of Θ, Φ, respectively).
[0087] Processing unit that implements target mid-audio signals and target side-audio signals. Hereinafter, FIGS. 6a, 6b and 6c are block diagrams showing an audio processing system comprising an analysis apparatus 10 and a synthesis apparatus 20, in combination with additional audio processing units such as a mono source separator 30, a mid / side processor 40, a stereo processor 50 and a stereo soft mask estimator 60 that perform audio processing on the target mid audio signal M and the target side audio signal S or the alternative left audio signal L CEN and the alternative right audio signal R CEN .
[0088] Figure 6a shows an audio processing system that processes a stereo audio signal via a target mid audio signal M and a target side audio signal S extracted by the analysis device 10. As can be seen in this figure, the target mid audio signal M is used as an input to a mono source separator 30. In addition, the target mid audio signal M and the target side audio signal S are provided to a mid / side processor 40 that applies enhancement and / or attenuation to the audio signals.
[0089] The target mid audio signal M separates audio sources in the target mid audio signal M, and the source-separated mid audio signal M sep is provided to a mono source separation system 30 configured to generate . For example, the mono source separation system 30 outputs a source-separated mid audio signal M with improved intelligibility of at least one audio source present in the target mid audio signal M sep is generated.
[0090] The source-separated mid audio signal M sep is provided to the mid / side processor 40 together with the target side audio signal S, and a processed mid audio signal M' and a side audio signal S' are generated. The mid / side processor 40 processes the source-separated mid audio signal M sepGain can be applied to and / or the target side signal S can be attenuated. In some cases, the mid / side processor 40 sets the target side audio signal S to 0. The processed mid audio signal M' and side audio signal S' are then provided to a synthesizer 20, which includes a left / right audio signal reconstruction unit 22. The left / right audio signal reconstruction unit 22 reconstructs the processed left audio signal L' and right audio signal R' based on the processed mid audio signal M' and side audio signal S' and the pan parameter Θ and / or phase difference parameter Φ obtained by the analysis unit 10.
[0091] In some embodiments, source-separated mid-audio signal M sep The signal is directly supplied to the synthesizer 20, which then processes at least the source-separated mid-audio signal M. sep Based on (and optionally the target side audio signal S) and the target pan parameter Θ and / or target phase difference parameter Φ, the processed left audio signal L' and right audio signal R' are reconstructed.
[0092] Alternatively, the mono source separation system 30 may be omitted, and the target mid-audio signal M and target side audio signal S are provided directly to the mid-side audio processing unit 40, which extracts the processed mid-audio signal M' and side audio signal S' from the target mid-audio signal M and target side audio signal S.
[0093] Figure 6b shows the left audio signal L located in the center of the alternative. CEN and right audio signal R CEN The diagram shows how the target mid audio signal M and target side audio signal S are used to extract and enable the stereo processing system 50 to be used for a center-panned stereo audio signal.
[0094] The analysis device 10a acquires arbitrary left audio signal L and right audio signal R, and target pan parameter Θ and / or target phase difference parameter Φ for each time segment and each frequency band, and extracts target mid audio signal M and target side audio signal S. The target mid audio signal M and target side audio signal S are provided to the synthesizer 20a, which uses an alternative centered pan parameter Θ to represent a center-panned stereo audio signal. CEN and / or phase difference parameter Φ CEN The set also acquires the left audio signal L located in the center. Therefore, the analysis device 10a and the synthesis device 20a acquire the left audio signal L located in the center. CEN and right audio signal R CEN Generates a center-panned stereo signal containing the original stereo signal from any given original stereo signal.
[0095] Center-panned alternative left audio signal L CEN and alternative right audio signal R CEN This is provided to the stereo processing system 50. The stereo processing system 50 performs stereo source separation on the center-panned stereo source and processes the center-located left audio signal L' CEN and right audio signal R' CEN It is configured to output the processed left audio signal L' located in the center. CEN and right audio signal R' CEN For example, the left audio signal L located in the center. CEN and right audio signal R CEN It features improved intelligibility of at least one audio source present in the system.
[0096] The processed left audio signal L' located in the center. CEN and right audio signal R' CEN This is then supplied to the second analysis device 10b, which processes the left audio signal L' located in the center. CENand right audio signal R' CEN And the pan parameter Θ located in the center. CEN and / or phase difference parameter Φ CEN The processed mid-audio signal M' and side audio signal S' are extracted using these parameters. Finally, the second synthesizer 20b reconstructs the processed left audio signal L' and right audio signal R' using the original pan parameter Θ and / or phase difference parameter Φ obtained or determined by the first analyzer 10a.
[0097] The centrally located pan parameter Θ indicates a central pan with no inter-channel phase difference. CEN and / or phase difference parameter Φ CEN This is the alternative pan parameter Θ ALT and / or alternative phase difference parameter Φ ALT It should be noted that this is one of many possible values. For example, the parameter Θ is the middle of the alternatives. CEN , Φ CEN For example, Θ is any alternative parameter that strictly indicates the right or left stereo audio signal, with or without non-zero phase difference. ALT , Φ ALT It can be replaced with this.
[0098] It should be further noted that the target mid-audio signal M and target side audio signal S supplied between the first analyzer 10a and the first synthesizer 20a can undergo mono signal separation and / or mid-side processing as described with respect to Figure 6 above. The same applies to the processed target mid-audio signal M' and target side audio signal S' supplied between the second analyzer 10b and the second synthesizer 20b. That is, the mono source separator 30 and / or mid / side processor 40 from Figure 6a can be used to process the mid-audio signal and side audio signal passed between the first analyzer 10a and the first synthesizer 20a and between the second analyzer 10b and the second synthesizer 20b. In addition, although the first analyzer 10a and the first synthesizer 20a are shown as separate units, these devices can be combined into a single unit, and referring to equations 20, 21, 22, and 23, the alternative left audio signal L CEN and alternative right audio signal R CEN It is clear that the operation to generate can be performed as a single step. The same applies to the second analysis apparatus 10b and the second synthesis apparatus 20b.
[0099] In some embodiments, an alternative (e.g., centrally located) left audio signal L CEN and alternative right audio signal R CEN This is generated by the first synthesizer 20a using only the target mid-audio signal M. Similarly, the second analyzer 10b can be configured to extract only the processed target mid-audio signal M'.
[0100] Figure 6c shows an analysis device 10 operating with a filter estimator configured to receive a left audio signal L and a right audio signal R and estimate a filter for source separation. This filter estimator is, for example, a stereo soft mask estimator 60 configured to receive a left audio signal L and a right audio signal R and estimate a soft mask F. The stereo soft mask estimator 60 can be any soft mask estimator and can be configured to output, for example, a soft mask F for stereo source separation. Conventionally, the soft mask F is applied to the left audio signal L and the right audio signal R of a stereo audio signal.
[0101] In the embodiment shown in Figure 6c, the estimated stereo soft mask F is applied to the target mid-audio signal M, and the source-separated mid-audio signal M sep Forms a source-separated mid-audio signal M sep The mid / side processor 40 is provided with the target side audio signal S, and the mid / side processor 40 outputs the processed target mid audio signal M' and the processed target side audio signal S'. The processed side audio signal S' can be down-weighted, including down-weighting up to 0, compared to S. The processed target mid audio signal M' and the processed target side audio signal S' are provided to the synthesis unit 20, which reconstructs the processed left audio signal L' and right audio signal R'. Optionally, the mid / side processor is omitted, thereby resulting in the source-separated mid audio signal M'. sep , and optionally, the target side audio signal S is directly supplied to the synthesis unit 20 for reconstruction.
[0102] The stereo soft mask estimator 60 is configured to determine or acquire the pan parameter Θ and / or the phase difference parameter Φ and provide these parameters to the analysis device 10. Alternatively, the pan parameter Θ and / or the phase difference parameter Φ may be acquired by the analysis device 10 from another source or determined by the analysis device 10 based on the left audio signal L and the right audio signal R.
[0103] Unless otherwise specified, as will be apparent from the following discussion, throughout this disclosure, any discussion using terms such as “processing,” “computing,” “calculating,” “requesting,” and “analyzing” is understood to refer to the operation and / or processes of computer hardware or computing systems, or similar electronic computing devices, that manipulate data expressed as physical quantities, such as electronic quantities, and / or convert such data into other data expressed as similar physical quantities.
[0104] In the above description of exemplary embodiments of the present invention, it should be understood that, for the purpose of streamlining the disclosure and aiding in the understanding of one or more aspects of the invention, various features of the invention may, in some cases, be grouped together in a single embodiment, in the figures of this embodiment, or in the description. However, this method of disclosure should not be interpreted as reflecting the intention that the claimed invention requires more features than those explicitly enumerated in each claim. On the contrary, as reflected in the following claims, aspects of the invention exist in fewer features than all the features of the single disclosed embodiment described above. Thus, the claims following this detailed description are thus explicitly incorporated into this detailed description, with each claim independently existing as an individual embodiment of the invention. Furthermore, while some embodiments described herein include some features that are not included in other embodiments, it is intended that combinations of features from different embodiments fall within the scope of the invention and form different embodiments, as will be understood by those skilled in the art. For example, in the following claims, any of the embodiments described in the claims may be used in any combination.
[0105] Furthermore, some of the embodiments described herein are methods or combinations of elements of methods that can be carried out by a processor of a computer system or by other means that perform the function. Thus, a processor having the instructions necessary to carry out such a method or element of a method forms a means that carries out the method or element of a method. When a method includes several elements, for example, several steps, it should be noted that the ordering of such elements is not implied unless specifically specified. Furthermore, the elements of the embodiments of the apparatus described herein are examples of means that carry out the function carried out by this element for the purpose of carrying out the present invention. The description provided herein contains a great many specific details. However, it will be understood that embodiments of the present invention can be carried out without these specific details. In other cases, known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description. Those skilled in the art will recognize that the present invention is by no means limited to the preferred embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, the stereo processing system in Figure 6c can handle a mid-audio signal M or a source-separated mid-audio signal M sepA mono-source separator 30 operating on the target mid-audio signal may be further provided. As a further example, it is assumed that all components or processes related to the formation of the target mid-audio signal and the target side-audio signal, and the processing of the mid-audio signal and the side-audio signal, can be configured to form or process the target mid-audio signal only, or the target mid-audio signal and the target side-audio signal. Similarly, reconstructing the left-audio signal and the right-audio signal (e.g., alternative or centrally located left-audio signal and right-audio signal) based on the target mid-audio signal and the target side-audio signal may include reconstructing the left-audio signal and the right-audio signal using only the target mid-audio signal, or using both the target mid-audio signal and the target side-audio signal.
[0106] Various features and aspects will be understood from the following enumerated exemplary embodiments ("EEE").
[0107] EEE1. A method for extracting a target mid audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), and the method is: (S1) obtaining a plurality of consecutive time segments of the stereo audio signal, wherein each time segment includes a representation of a portion of the stereo audio signal. For each frequency band of the multiple frequency bands of each time segment of the stereo audio signal, A target pan parameter (Θ) representing the distribution over the time segment of the magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band, A target phase difference parameter (Φ) represents the distribution of the phase difference between the left audio signal (L) and the right audio signal (R) of the stereo audio signal over that time segment, To obtain at least one of the following (S2), For each time segment and each frequency band, extract a partial mid-range signal representation (211, 212) (S3), wherein the partial mid-range signal representation (211, 212) is based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band. The target mid-audio signal (M) is formed by combining the partial mid-signal representations (211, 212) of each frequency band and each time segment (S4), Methods that include...
[0108] EEE2. The method according to EEE1, wherein the weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight.
[0109] EEE3. The method according to EEE1 or 2, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).
[0110] EEE4. The method according to EEE3, wherein the complex-valued weights are based on the target phase difference parameter (Φ) and the target pan parameter (Θ) such that the phase difference generated by the application of the complex-valued weights is smaller with respect to the left audio signal or the right audio signal having a larger magnitude.
[0111] EEE5. Providing the stereo audio signal to a stereo source separation system (60) configured to output a filter for stereo source separation, The filter is applied to the target mid-audio signal to process the target mid-audio signal (M sep ) forming, The method described in any one of EEE1 to EEE4, further including the above.
[0112] EEE6. Providing the target mid-range audio signal (M) to a mono source separation system (30) configured to perform audio source separation on a mono audio signal and output the processed mono audio signal, The processed mono audio signal is processed into the target mid-audio signal (M sep ) to be used as, The method described in any one of EEE1 to 5, further including the above.
[0113] EEE7. Reconstructing the processed left audio signal (L') and processed right audio signal (R') to form the processed stereo audio signal. It further includes, Each of the processed left audio signal (L') and the processed right audio signal (R') is weighted using the respective left and right weight coefficients to obtain the processed target mid-audio signal (M'). sep The method according to EEE5 or 6, wherein the left weighting coefficient and the right weighting coefficient are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band.
[0114] EEE8. Alternative left audio signal (L) that forms an alternative stereo audio signal. cen ) and alternative right audio signal (R cen ) to reconstruct, It further includes, The aforementioned alternative left audio signal (L cen ) and the alternative right audio signal (R cen Each of the ) is based on at least the weighted contribution of the target mid-audio signal, each contribution is based on the respective alternative mid-weighting coefficient, and each mid-weighting coefficient is based on the alternative pan parameter (Θ) for each time segment and each frequency band. cen ) and alternative phase difference parameter (Φ cen A method according to any one of EEE1-7, which is based on at least one of the following.
[0115] EEE9. Processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen Performing stereo signal processing on the alternative stereo audio signal which forms a processed alternative stereo audio signal including ) Reconstructing the processed target mid-audio signal (M''), It further includes, The processed target mid audio signal (M') is weighted using a weighting coefficient to obtain the processed alternative left audio signal (L'). cen ) and the processed alternative right audio signal (R' cen It is based on the sum of ) The weighting coefficient is the alternative pan parameter (Θ) for each frequency band and each time segment. cen ) and the alternative phase difference parameter (Φ cen The method described in EEE8, which is based on at least one of the following.
[0116] EEE10. The alternative pan parameter (Θ cen ) and / or the alternative phase difference parameter (Φ cen ) is the method described in EEE8 or 9, which indicates the central pan.
[0117] EEE11. Extracting partial side signal representations (311, 312) for each time segment and each frequency band (S31), wherein the partial side signal representations (311, 312) are based on the weight difference between the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment. The target side audio signal is formed by combining partial side signal representations for each frequency band and each time segment (S41), The method described in any one of EEE1 to 10, further including the above.
[0118] EEE12. Performing processing (S5, S51) on at least one of the target mid audio signal (M) and the target side audio signal (S) to form a processed target mid audio signal (M') and a processed target side audio signal (S'), Reconstructing the processed left audio signal (L') and processed right audio signal (R') that form the processed stereo audio signal (S6), It further includes, The processed left audio signal (L') is based on a weighted sum of the processed target mid audio signal (M') and the processed target side audio signal (S'), wherein the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The method according to EEE11, wherein the processed right audio signal (R') is based on a weighted difference between the processed target mid audio signal (M') and the processed target side audio signal (S'), the weight being based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ).
[0119] EEE13. Processing at least one of the target mid audio signal (M) and the target side audio signal (S) is: The target side audio signal (S) is attenuated, Applying gain to the target mid-audio signal (M), Perform mono signal source separation on the target mid-audio signal (M), Applying a stereo source separation filter to the target mid-audio signal (M), A method of EEE12, which includes at least one of the following.
[0120] EEE14. Obtaining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) is: To obtain at least one of the multiple detected pan parameters (θ) and the multiple detected phase difference parameters (φ), For each detected pan parameter (θ) and each detected phase difference parameter (φ), the detected magnitude parameter (U) is obtained, The method involves determining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) by calculating the average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), wherein the average is a weighted average weighted using the detected magnitude parameter (U). A method according to any one of EEE1-13, including the method described in EEE1-13.
[0121] EEE15. Sort each of the multiple detected pan parameters (θ) and the multiple detected phase difference parameters (φ) into a predetermined number of bins. It further includes, The method according to EEE14, wherein each bin is associated with at least one of a detected pan parameter value and a detected phase difference parameter value, and the mean of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter value and the detected phase difference parameter value.
[0122] EEE16. An audio processing system configured to extract a target mid audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), and the audio processing system is configured to extract a target mid audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), Obtaining a plurality of consecutive time segments of the stereo audio signal, wherein each time segment includes a representation of a portion of the stereo audio signal, For each frequency band of the multiple frequency bands of each time segment of the stereo audio signal, A target pan parameter (Θ) representing the distribution over the time segment of the magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band, A target phase difference parameter (Φ) representing the distribution of the phase difference between the left audio signal (L) and the right audio signal (R) in the frequency band over the time segment, To obtain at least one of the following, Extracting a partial mid-range signal representation (211, 212) for each time segment and each frequency band, wherein the partial mid-range signal representation (211, 212) is based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band. The target mid-audio signal (M) is formed by combining the partial mid-signal representations of each frequency band and each time segment, An audio processing system comprising an extractor unit (12) configured to perform the following.
[0123] EEE17. The audio processing system according to EEE16, wherein the weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight.
[0124] EEE18. The audio processing system according to EEE16 or 17, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).
[0125] EEE19. The audio processing system according to EEE18, wherein the complex-valued weights are based on a target phase difference parameter (Φ) and a target pan parameter (Θ) such that the phase difference generated by the application of the complex-valued weights is smaller with respect to the left audio signal or the right audio signal having a larger magnitude.
[0126] EEE20. Stereo source separation system (60) configured to output a filter for stereo source separation, Furthermore, The stereo audio signal is provided to the stereo source separation system (60), and the filter for stereo source separation is applied to the target mid-audio signal to process the target mid-audio signal (M sep An audio processing system described in any one of EEE16-19, which forms an )
[0127] EEE21. A mono source separation system (30) configured to perform audio source separation on a mono audio signal and output the processed mono audio signal. Furthermore, The target mid-range audio signal (M) is provided to the mono source separation system (30), and the processed mono audio signal is the processed target mid-range audio signal (M) sep An audio processing system described in any one of the EEE16-20 standards, used as an audio processing system.
[0128] EEE22. Reconstructor unit (20) configured to reconstruct the processed left audio signal (L') and right audio signal (R') that form the processed stereo audio signal. Furthermore, Each of the processed left audio signal (L') and right audio signal (R') is weighted using the respective left and right weight coefficients to obtain the processed target mid-range audio signal (M'). sep The audio processing system according to EEE20 or 21, wherein the left weight coefficient and the right weight coefficient are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band.
[0129] EEE23. Alternative left audio signal (L) that forms an alternative stereo audio signal. cen ) and alternative right audio signal (R cen A reconstructor unit (20) configured to rebuild ) Furthermore, The aforementioned alternative left audio signal (L cen ) and the alternative right audio signal (R cenEach of the ) is based on at least the weighted contribution of the target mid-audio signal, each contribution is based on the respective alternative mid-weighting coefficient, and each mid-weighting coefficient is based on the alternative pan parameter (Θ) for each time segment and each frequency band. cen ) and alternative phase difference parameter (Φ cen An audio processing system described in any one of EEE16-22, based on at least one of the following.
[0130] EEE24. Processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen A stereo processor (50) is configured to perform stereo signal processing on the alternative stereo audio signal which forms a processed alternative stereo audio signal including ), A second extractor unit configured to extract the processed target mid-range audio signal (M''), Furthermore, The processed target mid audio signal (M') is weighted using a weighting coefficient to obtain the processed alternative left audio signal (L'). cen ) and the processed alternative right audio signal (R' cen It is based on the sum of ) The weighting coefficient is the alternative pan parameter (Θ) for each frequency band and each time segment. cen ) and the alternative phase difference parameter (Φ cen An audio processing system described in any one of EEE16-23, based on at least one of the following.
[0131] EEE25. The alternative pan parameter (Θ cen ) and / or the alternative phase difference parameter (Φ cen ) indicates a central pan, as described in EEE23 or 24 of the audio processing system.
[0132] EEE26. The extraction unit 12 is, Extracting partial side signal representations (311, 312) for each time segment and each frequency band, wherein the partial side signal representations (311, 312) are based on a weighted difference between the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment. The target side audio signal is formed by combining partial side signal representations for each frequency band and each time segment. An audio processing system described in any one of the EEE16-25, further configured to perform the following.
[0133] EEE27. A mid / side processor (40) configured to process at least one of the target mid audio signal (M) and the target side audio signal (S) to form a processed target mid audio signal (M') and a processed target side audio signal (S'), A reconstructor unit (20) configured to reconstruct the processed left audio signal (L') and the processed right audio signal (R') that form the processed stereo audio signal, Furthermore, The processed left audio signal (L') is based on a weighted sum of the processed target mid audio signal (M') and the processed target side audio signal (S'), wherein the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The audio processing system according to any one of EEE16 to 26, wherein the processed right audio signal (R') is based on a weight difference between the processed target mid audio signal (M') and the processed target side audio signal (S'), and the weight is based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ).
[0134] EEE28. The processing performed by the mid / side processor (40) is: The attenuation of the target side audio signal (S), The application of gain to the target mid-audio signal (M), Mono signal source separation performed on the target mid-audio signal (M), The application of a stereo source separation filter to the target mid-audio signal (M), An audio processing system as described in EEE27, comprising at least one of the following.
[0135] EEE29. Obtain at least one of the multiple detected pan parameters (θ) and the multiple detected phase difference parameters (φ), For each detected pan parameter (θ) and each detected phase difference parameter (φ), the detected magnitude parameter (U) is obtained, The method involves determining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) by calculating the average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), wherein the average is a weighted average weighted using the detected magnitude parameter (U). An audio processing system according to any one of EEE16 to 28, further comprising a parameter extractor unit (14) configured to perform the following.
[0136] EEE30. The parameter extractor unit (14) is further configured to sort each of the multiple detected pan parameters (θ) and the multiple detected phase difference parameters (φ) into a predetermined number of bins. The audio processing system according to EEE29, wherein each bin is associated with at least one of a detected pan parameter value and a detected phase difference parameter value, and the average of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter value and the detected phase difference parameter value.
[0137] EEE31. A computer program product which, when executed by a computer, includes instructions that cause the computer to perform any one of EEE1 to EEE15.
[0138] EEE32. The magnitudes of the left audio signal and the right audio signal are determined. The method or the audio processing device described above is In response to the determination that the magnitude of the left audio signal is greater than the magnitude of the right audio signal, the weight of the left audio signal is determined to be greater than the weight of the right audio signal. In response to the determination that the magnitude of the right audio signal is greater than the magnitude of the left audio signal, the weight of the right audio signal is determined to be greater than the weight of the left audio signal. An audio processing device described in EEE2, or any one of EEE3 to EEE15 when dependent on EEE2, including the method described in EEE17, or any one of EEE18 to EEE30 when dependent on EEE17.
[0139] EEE33. The magnitudes of the left audio signal and the right audio signal are determined. The method or the audio processing device described above is In response to the determination that the magnitude of the left audio signal is greater than the magnitude of the right audio signal, the weight of the left audio signal is determined such that it causes a first phase shift when applied to the left audio signal, and the weight of the right audio signal is determined such that it causes a second phase shift when applied to the right audio signal. In response to the determination that the magnitude of the right audio signal is greater than the magnitude of the left audio signal, the weight of the right audio signal is determined such that it causes a first phase shift when applied to the right audio signal, and the weight of the left audio signal is determined such that it causes a second phase shift when applied to the left audio signal. Includes, The first phase shift is smaller than the second phase shift, the method according to EEE4 or the method according to any one of EEE5 to EEE15 when dependent on EEE4, or the audio processing device according to EEE19 or the audio processing device according to any one of EEE20 to EEE30 when dependent on EEE19.
[0140] EEE34. The method according to EEE8 or the method according to any one of EEE9-15 when dependent on EEE8, or the audio processing device according to EEE23 or the audio processing device according to any one of EEE24-30 when dependent on EEE23, wherein the alternative left audio signal and the alternative right audio signal are a second left audio signal and a second right audio signal, and / or the alternative stereo signal is a second stereo signal.
[0141] EEE35. The alternative stereo audio signal (L cen , R cen Performing stereo signal processing on the alternative stereo audio signal (L cen , R cenApply source separation to the processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen A method according to EEE9 or 10, or a method according to any one of EEE11 to 15 when dependent on EEE9 or 10, including outputting ).
Claims
1. A computer implementation method for extracting a target mid-range audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), and the method is: (S1) Obtaining a plurality of consecutive time segments of the stereo audio signal, wherein each time segment includes a representation of a portion of the stereo audio signal. For each frequency band of the multiple frequency bands of each time segment of the stereo audio signal, A target pan parameter (Θ) representing the distribution over a time segment of the magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band, wherein the target pan parameter (Θ) represents statistical features of a plurality of samples within the time segment, and the statistical features are the median, mean, mode, numbered percentile, or maximum or minimum pan parameter Θ, and A target phase difference parameter (Φ) representing the distribution of the phase difference between the left audio signal (L) and the right audio signal (R) of the stereo audio signal over the time segment, wherein the target phase difference parameter (Φ) represents statistical features of a plurality of samples within the time segment, and the statistical features are the median, mean, mode, numbered percentile, or maximum or minimum phase difference parameter Φ. To obtain at least one of the following (S2), (S3) extracting partial mid-range signal representations (211, 212) for each time segment and each frequency band, wherein the partial mid-range signal representations (211, 212) are based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band. The target mid-audio signal (M) is formed by combining the partial mid-signal representations (211, 212) of each frequency band and each time segment (S4), Includes, The weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight, in the method.
2. The method according to claim 1, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).
3. The method according to claim 2, wherein the complex-valued weights are based on the target phase difference parameter (Φ) and the target pan parameter (Θ) such that the phase generated by the application of the complex-valued weights is smaller with respect to the left audio signal or the right audio signal having a larger magnitude.
4. The stereo audio signal is provided to a stereo source separation system (60) configured to output a filter for stereo source separation, The filter is applied to the target mid-audio signal to process the target mid-audio signal (M sep ) forming, The method according to claim 1, further comprising:
5. The objective is to provide the target mid-range audio signal (M) to a mono source separation system (30) configured to perform audio source separation on a mono audio signal and output the processed mono audio signal, The processed mono audio signal is processed into the target mid-audio signal (M sep ) to be used as, The method according to claim 1, further comprising:
6. Reconstructing the processed left audio signal (L') and processed right audio signal (R') that form the processed stereo audio signal, It further includes, Each of the processed left audio signal (L') and the processed right audio signal (R') is weighted using the respective left and right weight coefficients to obtain the processed target mid-range audio signal (M'). sep The method according to claim 4, wherein the left weighting coefficient and the right weighting coefficient are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band.
7. Alternate left audio signal (L) that forms an alternate stereo audio signal cen ) and alternative right audio signal (R cen ) to reconstruct, It further includes, The alternative left audio signal (L cen ) and the alternative right audio signal (R cen ) are each based on at least the weighted contribution of said target mid audio signal, each contribution being based on a respective alternative mid weighting factor, each mid weighting factor being based on at least one of an alternative panning parameter (Θ cen ) and an alternative phase difference parameter (Φ cen ) for each time segment and each frequency band, the method according to claim 1.
8. Processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen Performing stereo signal processing on the alternative stereo audio signal which forms a processed alternative stereo audio signal including ) Reconstructing the processed target mid-audio signal (M'), It further includes, The processed target mid-range audio signal (M') is weighted using a weighting coefficient to obtain the processed alternative left-range audio signal (L'). cen ) and the processed alternative right audio signal (R' cen It is based on the sum of ) The weighting coefficient is the alternative pan parameter (Θ) for each frequency band and each time segment. cen ) and the alternative phase difference parameter (Φ cen The method according to claim 7, which is based on at least one of the following.
9. The alternative pan parameter (Θ cen ) and / or the alternative phase difference parameter (Φ cen The method according to claim 7, wherein ) represents the central pan.
10. The aforementioned alternative stereo audio signal (L cen , R cen Performing stereo signal processing on the alternative stereo audio signal (L cen , R cen Apply source separation to the processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen The method according to claim 8, which includes outputting ).
11. S31 extracts partial side signal representations (311, 312) for each time segment and each frequency band, wherein the partial side signal representations (311, 312) are based on the weight difference between the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment. The target side audio signal is formed by combining partial side signal representations for each frequency band and each time segment (S41), The method according to claim 1, further comprising:
12. Performing processing (S5, S51) on at least one of the target mid-audio signal (M) and the target side audio signal (S) to form a processed target mid-audio signal (M') and a processed target side audio signal (S'), Reconstructing the processed left audio signal (L') and processed right audio signal (R') that form the processed stereo audio signal (S6), It further includes, The processed left audio signal (L') is based on a weighted sum of the processed target mid audio signal (M') and the processed target side audio signal (S'), wherein the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The method according to claim 11, wherein the processed right audio signal (R') is based on a weighted difference between the processed target mid audio signal (M') and the processed target side audio signal (S'), the weight being based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ).
13. Processing at least one of the target mid audio signal (M) and the target side audio signal (S) is: The target side audio signal (S) is attenuated, Applying gain to the target mid-audio signal (M), Performing mono signal source separation on the target mid-audio signal (M), Applying a stereo source separation filter to the target mid-audio signal (M), The method according to claim 12, comprising at least one of the following.
14. Obtaining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) is: Obtaining at least one of multiple detected pan parameters (θ) and multiple detected phase difference parameters (φ), For each detected pan parameter (θ) and each detected phase difference parameter (φ), the detected magnitude parameter (U) is obtained, The method for determining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) is to calculate the average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), wherein the average is a weighted average weighted using the detected magnitude parameter (U). The method according to claim 1, including the method described in claim 1.
15. Sort each of the multiple detected pan parameters (θ) and the multiple detected phase difference parameters (φ) into a predetermined number of bins. It further includes, The method according to claim 14, wherein each bin is associated with at least one of a detected pan parameter value and a detected phase difference parameter value, and the mean of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter value and the detected phase difference parameter value.
Citation Information
Patent Citations
Signal processor, signal processing method and computer-readable recording medium recording signal processing program
JP1999289599A
Sound signal processor and sound signal processing method
JP2006080708A
Sound separating device, sound separating method, sound separating program, and computer-readable recording medium
WO2006090589A1
Perceptual optimization of magnitude and phase for time-frequency and softmask source separation systems
WO2021252795A2
Separation of panned sources from generalized stereo backgrounds using minimal training
WO2021252912A1