Target mid-side signal for audio applications

The method improves audio source separation in stereo audio signals by using time and frequency-specific pan and phase parameters to weight left and right audio signals, effectively isolating non-centered and dynamic audio sources.

JP2025517590AActive Publication Date: 2025-06-10DOLBY LABORATORIES LICENSING CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024553721
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2023-03-03
Publication Date
2025-06-10
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Conventional methods for extracting mid and side audio signals from stereo audio signals fail to effectively separate audio sources that are not centered, particularly moving or reverberant sources.

Method used

A method that extracts a target mid-audio signal and a target side-audio signal from a stereo audio signal by obtaining target pan parameters and phase difference parameters for each time segment and frequency band, allowing for the weighting of left and right audio signals to isolate the target audio source.

Benefits of technology

This method enables enhanced audio source separation, effectively capturing target audio sources that change over time and frequency, while minimizing the capture of non-target sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517590000001_ABST
    Figure 2025517590000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and an audio processing apparatus for extracting a target mid (and optionally target side) audio signal from a stereo audio signal. The method includes obtaining a plurality of consecutive time segments of the stereo audio signal (S1), and obtaining at least one of a target pan parameter Θ and a target phase difference parameter Φ for each of a plurality of frequency bands of each time segment of the stereo audio signal (S2). The method further includes extracting partial mid signal representations 211, 212 based on at least one of the target pan parameter Θ and the target phase difference parameter Φ of each frequency band for each time segment and each frequency band (S3), and forming a target mid audio signal M by combining the partial mid signal representations 211, 212 of each frequency band and each time segment (S4).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority to the following prior applications: namely, U.S. Provisional Patent Application No. 63 / 318,226, filed on March 9, 2022; U.S. Provisional Patent Application No. 63 / 423,786, filed on November 8, 2022; and European Patent Application No. 22183794.1, filed on July 8, 2022.

[0002] The present invention relates to source separation, signal enhancement, and signal processing. The present invention further relates to a method for extracting a target mid - audio signal and a target side - audio signal from a stereo audio signal. The present invention further relates to a processing device using the aforementioned method.

Background Art

[0003] In the field of audio processing, audio source separation is important in several applications. In one simple application, the separated speech audio signal is retained or given additional gain, while the background audio signal is excluded or attenuated relative to the speech audio signal. This can improve the intelligibility of the speech audio signal.

[0004] In audio source separation, the signal to be estimated or extracted as the target is called the target signal. The target signal is not limited to speech and can be any general audio signal (such as musical instruments), or even a plurality of audio signals that must be separated from an audio mix including noise and / or additional non - target audio signals.

[0005] For a stereo audio signal including two audio signals respectively associated with left and right audio channels, one processing method that implicitly performs source separation is to extract a mid audio signal and a side audio signal from the left and right audio signals, where the mid audio signal and the side audio signal are respectively proportional to the sum and difference of the left and right audio signals. The mid signal emphasizes audio components that are equal in magnitude and in phase between the left and right audio signals, while the side audio signal attenuates or removes such signals. Thus, the calculation of the mid and side signals from the left and right audio signals constitutes an efficient way to enhance or attenuate in-phase center-panned sources respectively. The mid and side signals can be further inverse-transformed into the left and right (conventional stereo) audio signals.

[0006] Note that source separation is implicitly performed as long as the center-panned source is of interest. That is, the mid signal includes this source while the side signal does not. The side audio signal mainly includes potentially uninteresting background audio signals, and by attenuating or excluding these background audio signals, the intelligibility of the center-panned in-phase audio source can be improved.

[0007] The drawback of conventional mid and side signal calculations is that implicit source separation based on mid / side audio signal separation fails when the desired audio source is not centered. To address this, specific calculation techniques for mid and side signals have been developed for stereo audio signals with specific panning characteristics. These techniques can successfully target stationary non-centered audio sources. See, for example, "Stereo Music Source Separation via Bayesian Modeling" by Master, Aaron (Ph.D. dissertation, Stanford University, 2006).

[0008] These techniques mitigate some of the shortcomings of basic mid and side signal extraction, but such pan-specific extraction techniques still cannot separate more general audio sources present in stereo audio signals, such as moving audio sources, reverberant sources, or sources that are spatially dominant at various points in the time-frequency space. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION

[0009] Accordingly, an object of the present disclosure is to provide such an improved method and an audio processing system that performs enhanced audio source separation on stereo audio signals. MEANS FOR SOLVING THE PROBLEM

[0010] According to a first aspect of the present invention, there is provided a method for extracting a target mid-audio signal from a stereo audio signal, the stereo audio signal including a left audio signal and a right audio signal. The method includes obtaining a plurality of consecutive time segments of the stereo audio signal, each time segment including a representation of a portion of the stereo audio signal, and for each frequency band of a plurality of frequency bands of each time segment of the stereo audio signal, obtaining at least one of a target pan parameter and a target phase difference parameter. The target pan parameter represents the distribution over time segments of the magnitude ratio between the left audio signal and the right audio signal in the frequency band, and the target phase difference parameter represents the distribution over time segments of the phase difference between the left audio signal and the right audio signal of the stereo audio signal.

[0011] The method further includes, for each time segment and each frequency band, extracting a partial mid-signal representation, the partial mid-signal representation being based on a weighted sum of the left audio signal and the right audio signal, the respective weights of the left audio signal and the right audio signal being based on at least one of the target pan parameter Θ and the target phase difference parameter for each frequency band and each time segment, and forming a target mid-audio signal by combining the partial mid-signal representations for each frequency band and each time segment.

[0012] Obtaining at least one of the target pan parameter and the target phase difference parameter can include receiving the target pan parameter and / or the target phase difference parameter, accessing those parameters, or determining those parameters. At least one of the target pan parameter and / or the target phase difference parameter can be replaced with a default value for at least one time segment and frequency band.

[0013] A continuous time segment means a segment that represents a later time portion of an audio signal for a later time segment and an earlier time portion of the audio signal for an earlier time segment. The continuous time segments may or may not have overlapping times.

[0014] The representation of a portion of a stereo audio signal can be any time-domain representation or any frequency-domain representation. The frequency-domain representation can be any linear time-frequency domain such as a short-time Fourier transform (STFT) representation or a quadrature mirror filter (QMF) representation.

[0015] The present invention is at least partially based on the understanding that a target mid-audio signal can be extracted that always targets an audio source that changes over time and / or frequency in a stereo audio signal by obtaining target pan parameters and / or target phase difference parameters for each time segment and each frequency band. Also, since the target mid-audio signal is generated using individual target pan parameters for each time segment and each frequency band, the method enables the extraction of a target mid-audio signal that targets two or more audio sources separated in frequency in a stereo audio signal.

[0016] In some embodiments, the weights of the left and right audio signals are based on target pan parameters such that the left or right audio signal with the greater magnitude is given the greater weight.

[0017] That is, the left or right audio signal associated with the greater magnitude or power (as indicated by the target parameter) is given a greater weight and contributes more to the formation of the target mid audio signal.

[0018] In some embodiments, the method further comprises, for each time segment and each frequency band, extracting a partial side signal representation, the partial side signal representation being based on a weighted difference between the left audio signal and the right audio signal, wherein the respective weights of the left audio signal and the right audio signal are based on at least one of the target parameter and the target phase difference parameter for each frequency band and each time segment, and forming a target side audio signal by combining each partial side signal representation for each frequency band and each time segment.

[0019] In other words, the target side audio signal can be formed in parallel with the target mid audio signal, and the target mid audio signal and the target side audio signal together form a complete representation of the stereo audio signal. It is understood that any embodiments related to the formation or processing of the target mid audio signal can be carried out in the same manner as the formation and processing of the target side audio signal as described below.

[0020] When a stereo audio signal includes, for example, an ideal center-panned audio source or a single audio source panned under the constant power law, the target mid-audio signal ideally captures all the signal energy of the audio source, while the target side-audio signal does not capture such energy. However, in the case of a general audio source(s), the target mid-audio signal captures the target audio source(s) and probably some other non-target sounds, while the target side-audio signal captures non-target sounds and probably removes or almost removes the target source(s). By including the target side-audio signal in addition to the target mid-audio signal, it is ensured that all the audio signals of the pair of left and right audio signals are present in the pair of target mid-audio signal and target side-audio signal. For example, this enables lossless reconstruction of the left and right audio signals.

[0021] Any function described with respect to a method can have a corresponding feature in a system or device, and vice versa.

[0022] The present invention will be described in more detail with reference to the accompanying drawings showing presently preferred embodiments of the invention.

Brief Description of the Drawings

[0023]

Figure 1

[0024]

Figure 2

[0025]

Figure 3

[0026]

Figure 4a

[0027]

Figure 4b

[0028]

Figure 4c

[0029]

Figure 5a

[0030]

Figure 5b

[0031]

Figure 6a

[0032]

Figure 6b

[0033]

Figure 6c

[0034] The systems and methods disclosed in this application can be implemented as software, firmware, hardware, or combinations thereof. In hardware embodiments, the partitioning of tasks does not necessarily correspond to the partitioning into physical units. On the contrary, one physical component may have multiple functions, and one task may be executed collaboratively by several physical components.

[0035] Computer hardware can be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify operations performed by that computer hardware. Further, the present disclosure relates to any collection of computer hardware that individually or jointly executes instructions for implementing any one or more of the concepts discussed herein.

[0036] One or more processors can implement certain or all components by receiving computer-readable (also called machine-readable) code that includes a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify the actions to be taken is included. Thus, one example is a conventional processing system (i.e., computer hardware) that includes one or more processors. Each processor can include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system can further include a memory subsystem that includes a hard drive, an SSD, RAM, and / or ROM. A bus subsystem for communicating between components may be included. Software can reside in the memory subsystem and / or can reside within the processor during execution of the software by the computer system.

[0037] One or more processors can operate as a stand-alone device or can be connected to other processors (s), e.g., network connected. Such networks can be built on a variety of different network protocols and can be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0038] Software can be distributed on computer-readable media that can include computer storage media (i.e., non-transitory media) and communication media (i.e., transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile media, removable and non-removable media, implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, various forms of physical (non-transitory) storage media such as EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Further, communication media (transitory) typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media, as is well known to those skilled in the art.

[0039] Including an introductory section on the extraction of basic mid-audio signals and side-audio signals, the following description includes details of the presently preferred embodiments.

[0040] Extraction of basic mid-audio signal and side-audio signal A stereo audio signal including a left audio signal L and a right audio signal R is represented as a mid-audio signal M and a side-audio signal S as another representation. The basic mid-audio signal M and side-audio signal S are constructed from the left audio signal L and the right audio signal R using the following analytical expressions: M = 0.5L + 0.5R Equation 1 S = 0.5L - 0.5R Equation 2 The original left audio signal L and right audio signal R can be reconstructed from the mid audio signal M and side audio signal S using the following synthesis equations. L = M + S Equation 3 R = M - S Equation 4

[0041] The mid audio signal M enhances the audio signal features that are in-phase and at the center of the stereo mix, while the side audio signal S attenuates the in-phase audio signal features. For example, if the stereo audio signal includes a pair of left audio signal L and right audio signal R that are centered panned (equal magnitude and in-phase components in each of the left audio signal and right audio signal), the mid audio signal M includes the stereo audio signal, while the side audio signal S removes this stereo audio signal. And the basic mid / side audio signal targets the centered panned audio signal that is included in the mid audio signal M but not in the side audio signal S.

[0042] Stereo signal in time domain and time-frequency domain The left audio signal L, right audio signal R, mid audio signal M, and side audio signal S can be represented in the time domain or the frequency domain. The frequency domain representation can be, for example, a short-time Fourier transform (STFT) representation or an orthogonal modulation filter (QMF) representation. For example, the left audio signal L and right audio signal R can be represented in the frequency domain by the following equations.

Equation

Equation

Equation

[0043] The detected magnitude parameter U and the detected pan parameter θ form a representation of the stereo audio signal called the Stereo-Polar Magnitude (SPM) representation. The SPM representation can be replaced by any equivalent representation, but the SPM representation of the stereo audio signal is used in the following explanation.

[0044] The detected pan parameter θ is in the range of 0 to π / 2, where 0 indicates a stereo audio signal in which only the left audio signal L is non-zero, and π / 2 indicates a stereo audio signal in which only the right audio signal R is non-zero. The detected pan parameter value of θ = π / 4 indicates a centered stereo audio signal where the magnitudes of the left audio signal L and the right audio signal R are equal. The detected pan parameter is obtained from the stereo input signal and also mathematically corresponds to the model of source mixing. Magnitude │S x │, phase ψ x , and an audio source S x having a known pan parameter Θ x , the left audio signal L and the right audio signal R can be expressed as the following equations.

Equation

Equation

[0045] Mid-audio signal and side-audio signal with fixed pan The basic mid-audio signal M and side-audio signal S are extracted by equally weighting each of the left audio signal L and the right audio signal R in order to target an audio source (a centered audio source) assumed to be present in equal proportions in the left audio signal L and the right audio signal R. In the case of an audio source not located in the center, the weighting of the left audio signal L and the right audio signal R in Equations 1 and 2 should be adjusted.

[0046] For example, the coefficients of the left audio signal L and the right audio signal R in Equations 1 and 2 can be changed from 0.5 under the constraint that the sum of the coefficients must be equal to 1. Thus, for an audio source that mainly appears in the left audio signal, the coefficients of the left audio signal L and the right audio signal R in Equation 1 can be, for example, 0.7 and 0.3 (similarly for the configuration of the side signal in Equation 2). However, this results in an undesirable situation where the magnitudes of the mid audio signal M and the side audio signal S change significantly when the audio source moves between the left audio signal L and the right audio signal R (assuming constant-power mixing).

[0047] To avoid this problem, for an audio source S x having a known or estimated "target" pan parameter Θ x the mid audio signal M and the side audio signal S can be extracted from the left audio signal L and the right audio signal R according to the following equations. M = cos(Θ x )L + sin(Θ x )R Equation 14 S = sin(Θ x )L - cos(Θ x )R Equation 15 In this case, the left audio signal L and the right audio signal R can be reconstructed as the following equations. L = cos(Θ x )M + sin(Θ x )S Equation 16 R = sin(Θ x )M - cos(Θ x )S Equation 17

[0048] For example, when Θ x = 0, indicating that the audio source appears only in the left audio signal, the mid audio signal M is equal to the left audio signal L, the side audio signal S is equal to the right audio signal R, and Θ xWhen Θ = π / 2, it is the opposite. Similarly, when the audio source appears with equal magnitude in the left audio signal L and the right audio signal R (i.e., Θ x = π / 4), the mid audio signal M is obtained by equally weighting the left audio signal L and the right audio signal R. Further, when the audio source is located to the left of the center, the left audio signal L and the right audio signal R are weighted using Θ x in the range of 0 to π / 4, and when the audio source is located to the right of the center, the left audio signal L and the right audio signal R are weighted using Θ x in the range of π / 4 to π / 2. The coefficients used to weight the left audio signal L and the right audio signal R to construct the mid audio signal M and the side audio signal S, and the coefficients used to construct the inverse thereof, are based on this fixed target pan parameter Θ x .

[0049] Extraction of target mid-audio signal and target side-audio signal One aspect of the present invention relates to the generation of a target mid audio signal M and / or a target side audio signal S based on at least one of a target pan parameter Θ and a target phase difference parameter Φ obtained for each time segment and each frequency band of a stereo audio signal. The target pan parameter Θ represents an estimated value or a display value of the magnitude ratio between the left audio signal L and the right audio signal R in each time segment and each frequency band of the sound, corresponding to one or more target sources. The target phase difference parameter Φ represents an estimated value or a display value of the phase difference between the left audio signal L and the right audio signal R in each time segment and each frequency band of the sound, corresponding to one or more target sources.

[0050] In some embodiments, the pan parameter Θ and / or the phase difference parameter Φ for each time segment and each frequency band 111, 112 are the median, average, mode, numbered percentile, maximum or minimum pan parameter Θ and / or phase difference parameter Φ of the time segment. Generally, the detected pan parameter θ and / or the detected phase difference parameter φ exist for each sample of the stereo audio signal. On the other hand, the target pan parameter Θ and / or the target phase difference parameter Φ used to generate the target mid audio signal M and the target side audio signal S are not necessarily of such fine granularity, and may represent statistical features (such as the average) of a plurality of samples such as the average of all samples within a predetermined time segment. For example, the detected pan parameter θ and / or the detected phase difference parameter φ can be assumed to have different values more than 1000 times per second, whereas the target pan parameter Θ and / or the phase difference parameter Φ change only a few times per second.

[0051] FIG. 1 shows an analysis device 10 that extracts a target mid audio signal M and a target side audio signal S from a stereo signal including a left audio signal L and a right audio signal R. The extraction of the mid audio signal M and the side audio signal S is performed in an extraction unit 12 of the analysis device 10, and the extraction unit 12 acquires the left audio signal L and the right audio signal R together with the target pan parameter Θ and / or the target phase difference parameter Φ of a plurality of time segments 100 and frequency bands 111, 112.

[0052] The extraction unit 12 acquires the left audio signal L and the right audio signal R (stereo audio signal) as a plurality of consecutive time segments. Each time segment includes an expression of a part of the left audio signal L and the right audio signal R. Alternatively, the extraction unit 12 is configured to divide the left audio signal L and the right audio signal R into a plurality of consecutive time segments. Each time segment includes two or more frequency bands of the stereo audio signal, and each frequency band in each time segment is associated with at least one of the acquired target pan parameter Θ and the acquired target phase difference parameter Φ.

[0053] In the embodiment shown in FIG. 1, the analysis device 10 receives the left audio signal L and the right audio signal R, and the target pan parameter Θ and / or the target phase difference parameter Φ of each frequency band 111, 112 of each time segment. The numbers of the frequency bands range from band 1 to band B. Θ (t,1) and Φ (t,1) represent the target pan parameter Θ and the target phase difference parameter Φ of the first frequency band 111 of the time segment t, while Θ (t,2) and Φ (t,2) represent the target pan parameter and the target phase difference parameter of the second frequency band 112 of the time segment t.

[0054] Using the left audio signal L and the right audio signal R, and the target pan parameter Θ and / or the target phase difference parameter Φ, the extraction unit 12 extracts the target mid audio signal M and the target side audio signal S. In some embodiments, the extraction unit 12 extracts only one of the target mid audio signal M and the target side audio signal S, such as only the mid audio signal M.

[0055] To enable modeling of the inter-channel phase difference between the left audio signal L and the right audio signal R, the true source phase of the target audio source is defined as being proportionally closer to the left audio signal or the right audio signal that exhibits a greater magnitude. That is, in the case of a stereo audio signal where the left audio signal L is dominant in terms of power or magnitude, the phase of the left audio signal L is closer to the true source phase of the audio source compared to the phase of the right audio signal R, and vice versa in the case of a stereo audio signal where the right audio signal R is dominant. The audio source S x is modeled to appear in the left audio signal L and the right audio signal R according to the following mixing formula.

Equation

[0056] Based on the mixing models described in the above equations 17 and 18, the extraction device 12 can target a source having the target pan parameter Θ and the target phase difference parameter Φ using the following analytical expressions.

Equation

[0057] In equation 20, the target mid audio signal M is based on the weighted sum of the left audio signal L and the right audio signal R, and the weight of the left audio signal is JPEG2025517590000014.jpg527, and the weight of the right audio signal is JPEG2025517590000015.jpg525, and the pan parameter Θ and / or the phase difference parameter Φ are obtained for each time segment and frequency band 111, 112.

[0058] Similarly, in equation 20, the target side audio signal S is based on the weighted difference between the left audio signal L and the right audio signal R, and the weight of the left audio signal is JPEG2025517590000016.jpg526, and the weight of the right audio signal is JPEG2025517590000017.jpg525, and the pan parameter Θ and / or the phase difference parameter Φ are obtained for each time segment and frequency band 111, 112.

[0059] The weights of the left audio signal L and the right audio signal R in the weighted sum of Equation 20 and the weighted difference of Equation 21 used to extract the mid-audio signal M and the side audio signal S are complex numbers, including a real-valued magnitude coefficient, e.g., cos(Θ), and a complex-valued phase coefficient, e.g., including and JPEG2025517590000018.jpg517. The obtained target pan parameter Θ and / or target phase difference parameter Φ affect the weights so that the target mid-audio signal M and / or target side audio signal S can be extracted.

[0060] For example, the real-valued magnitude coefficient can be based on Θ such that the left or right audio signal with the greater magnitude is given a greater weight in the weighted sum. Also, the complex-valued coefficient can be based on Φ such that the greater the phase difference Φ, the greater the difference between the phases of the respective weights in the weighted sum. In some embodiments, the complex-valued coefficient of each weight in the weighted sum is based on both Φ and Θ such that the weight of the left or right audio signal with the greater magnitude is given a smaller phase, e.g., one of the weights is JPEG2025517590000019.jpg517.

[0061] Similarly, the real-valued magnitude coefficient in the weighted difference (for target side signal extraction) can be based on the target pan parameter Θ such that the left or right audio signal with the greater magnitude is given a smaller weight in the weighted difference. Also, the complex-valued coefficient can be based on Φ such that the greater the phase difference Φ, the greater the difference between the phases of the respective weights in the weighted difference. In some embodiments, the complex-valued coefficient of each weight in the weighted difference is based on both Φ and Θ such that the weight of the left or right audio signal with the greater magnitude is given a smaller phase, e.g., one of the weights is, It is JPEG2025517590000020.jpg526.

[0062] In some embodiments, only one of the target pan parameter Θ and the target phase difference parameter Φ is obtained for at least one time segment and frequency bands 111, 112, and a default value is assigned to the other. For example, if only the target pan parameter Θ is obtained, the target phase difference parameter Φ is assumed to be Φ = 0, and if only the target phase difference parameter Φ is obtained, the target pan parameter Θ is assumed to be Θ = π / 4. These default values are based on an audio source at the center of a stereo audio signal having no inter-channel phase difference, but other default values are possible.

[0063] The extraction device 12 obtains the target mid-audio signal M by combining the partial mid-signal components or representations M of each frequency band in each time segment. Similarly, the extraction device 12 obtains the target side-audio signal S by combining the partial side-signal representations S of each frequency band in each time segment. The partial mid-signal components or representations M (t,n) are generated for each frequency band 1...B of the time segment t, and by combining all the partial mid-signal representations M of the time segment t (t,n) the target mid-audio signal time segment M (t,n) is generated, and the sequence of the target mid-audio signal time segments M (t,n) forms the target mid-audio signal M. The target side-audio signal S is also formed in a similar manner. Combining the partial mid-signal representations M (t) of the frequency bands 1...B can include processing the partial target mid-audio signal M (t) using a synthesis filter bank. (t,n) (t,n) (t,n) (t,n)

[0064] Referring further to the flowchart in FIG. 2, the method executed by the analysis device 10 will be described in detail. A stereo audio signal is acquired by the analysis device 10 at S1, and a target pan parameter Θ and / or a target phase difference parameter Φ are acquired by the analysis device 10 for a plurality of time segments 100 and frequency bands 111, 112 at step S2.

[0065] At step S3, a partial mid-signal representation M (t,n) is extracted by the mid / side extractor 12 (e.g., according to the above formula 19), and this partial mid-signal representation is based on the weighted sum of the left audio signal L and the right audio signal R. The respective weights of the left audio signal L and the right audio signal R are based on at least one of the target pan parameter Θ and the target phase difference parameter Φ for each frequency band 111, 112 and each time segment. Similarly, a partial side-signal representation S (t,n) is extracted at step S31 (e.g., according to the above formula 20), and this partial side-signal representation is based on the weighted difference between the left audio signal L and the right audio signal R.

[0066] In some embodiments, the target side audio signal S is not calculated, and the extractor unit 12 is an extractor unit 12 configured to extract only the target mid audio signal M, for example. Alternatively, the target side audio signal S is calculated, but is reduced to a small value or 0 at step S51 before the reconstruction of the left audio signal L and the right audio signal R at step S6 (e.g., using the following reconstruction formulas 21 and 22). Since the target mid audio signal M in some cases is expected to capture substantially all of the energy of the target audio source, the target side audio signal S may mainly contain unwanted noise or background audio that can be excluded or attenuated.

[0067] Following steps S3 and S31, the method proceeds to steps S4 and S41. These steps involve forming target mid-audio signal M and target side-audio signal S by combining partial mid-signal representation M (t,n) and partial side-signal representation S (t,n) with target mid-audio signal M and target side-audio signal S respectively. The method then continues with steps S5 and S51 which involve performing processing on target mid-audio signal M and / or target side-audio signal S. Examples of processing include attenuating target side-audio signal S or providing target mid-audio signal M to a mono-source separator as will be described with respect to FIGS. 6a, 6b, and 6c below.

[0068] Finally, the method proceeds to step S6 which involves reconstructing left-audio signal L and right-audio signal R from target mid-audio signal M and target side-audio signal S. The reconstruction of left-audio signal L and right-audio signal R from target mid-audio signal M and target side-audio signal S is performed by a synthesizing device as will be described in more detail with respect to FIGS. 5a and 5b below.

[0069] As an example of the operation of the analysis device 10, consider an audio source of a left audio signal L and a right audio signal R having a panned time-varying pan. This audio source pans at a constant speed from completely right to completely left under the constant power law as time t progresses from 0 seconds to 10 seconds. For example, in the case of the basic mid-audio signal and side-audio signal extracted for this audio source according to the above equations 1 and 2, the mid-audio signal M captures substantially all of the energy of the audio source at t = 5 and substantially no energy at t = 0 and t = 10, while the side-audio signal S captures substantially all of the energy of the audio source at t = 0 and t = 10 and substantially no energy at t = 5. The target mid-audio signal M extracted using the changing pan parameter Θ and / or phase difference parameter Φ for each time segment and each frequency band contains substantially all of the audio source energy for any t ∈ [0, 10], while the target side-audio signal S substantially does not contain the energy of the time-varying audio source.

[0070] The constant power law ensures that the audio signal power (proportional to the sum of the squares of the amplitudes) of the left audio signal L and the right audio signal R remains constant for all panning angles. For example, the audio signal amplitudes of the left audio signal and the right audio signal are scaled using a scaling factor that depends on the panning angle to ensure that the sum of the squares of the left audio signal and the right audio signal is equal for all panning angles. Different from the linear panning law that ensures the sum of the signal amplitudes is constant and allows the total signal power to vary with the panning angle, the constant power panning law ensures that the perceived audio signal power is constant for all panning angles. The constant power law is described, for example, in "Loudness Concepts & Pan Laws" by Anders Oeland and Roger Dannenberg (Introduction to Computer Music, Carnegie Mellon University).

[0071] This shows how the target mid audio signal M and the target side audio signal S can target and separate audio source(s) whose time and frequency vary in a stereo audio signal. Additionally, if a second audio source of a different frequency is moved at a different speed between the left audio signal L and the right audio signal R (or is stationary, for example, at the center pan), the target mid audio signal M targets the first audio source and the second audio source in individual frequency bands. This means that substantially all of the energy from both audio sources is present in the target mid audio signal M, even though both audio sources are of different frequencies and are shifted at different speeds between the left audio signal L and the right audio signal R.

[0072] FIG. 3 shows a combined analysis device 10' configured to extract the target mid-audio signal M and the target side-audio signal S according to the above, and also extract or obtain the target pan parameter Θ and / or the target phase difference parameter Φ of each frequency band 111, 112 and each time segment of the left audio signal L and the right audio signal R using the parameter extractor unit 14. Therefore, the target pan parameter Θ and / or the target phase difference parameter Φ are obtained together with the left audio signal L and the right audio signal R, or obtained from the left audio signal L and the right audio signal R. An exemplary method for doing this will be described next.

[0073] The target phase difference parameter Φ can be calculated, for example, by the parameter extractor 14 as the typical phase difference or the dominant phase difference between the left audio signal L and the right audio signal R in each time segment and each frequency band 111, 112. For example, in the time-frequency domain (e.g., the STFT domain), the detected phase difference can be calculated as φ = Arg(R / L) for each STFT tile, and by analyzing the distribution of φ, the value of Φ can be estimated to be the dominant value or the typical value of that time segment and frequency band. Similarly, the detected pan parameter θ can be calculated as θ = arctan(|R| / |L|) for each STFT tile, and by analyzing the distribution of θ, the value of Θ can be estimated to be the dominant value or the typical value of that time segment and frequency band. Here, L and R represent the left audio signal L and the right audio signal R in the time segment and the frequency bands 111, 112. The typical value or the dominant value estimated from the distribution can be the median, the mean, the mode, the numbered percentile, the minimum value, or the maximum value.

[0074] Master, in "Dialog Enhancement via Spatio-Level Filtering and Classification" (Convention Paper 10427 of the 149th Convention of the Audio Engineering Society) by A, a method for obtaining indicators of the average value and spread of panning and inter-channel phase differences labeled with thetaMid, thetaWidth, phiMid, and phiWidth is proposed. The panning parameter Θ used by the analysis device 10 can be based on thetaMiddle proposed in this conference paper. Similarly, the target phase difference parameter Φ used by the analysis device 10 can be based on the so-called "Shift and Squeeze" parameter or the phiMiddle of the "S&S" parameter proposed in this paper.

[0075] For each time segment and each frequency band within this time segment, the audio processing system squares the magnitude U, that is, U 2A 51-bin histogram regarding θ weighted by [weighting factor] can be created. The system performs the same for the detected pan parameter φ and a version of φ in the range 0 to 2*pi called φ2. However, these histograms each use 102 bins. The histograms are each smoothed over time segments across their given dimension. For the smoothed θ histogram, the system detects the target pan parameter as the highest peak called thetaMiddle and also detects the width around this peak called thetaWidth that is necessary to capture 40% of the energy in the histogram. The system performs the same for φ and φ2, recording phiMiddle, phi2Middle, phiWidth, and phi2Width, but requires 80% energy capture for the width. The system records the final values of phiMiddle and phiWidth based on which has a higher concentration in the phi space as indicated by the smaller phiWidth value.

[0076] For example, parameter extractor 14 in FIG. 3, for each time segment and each frequency band within this time segment, divides the detected pan parameter θ by the total signal power U from Equation 7 2Create a histogram of the weighted ones using [the given method], and do the same for the detected phase difference parameter φ. These histograms can have any number of bins, such as 10 or more bins, or 50 or more bins. In one embodiment, each histogram has 51 bins. Each histogram is smoothed using a smoothing function suitable for at least the time dimension and / or the time and θ or φ dimensions. For the smoothed θ histogram, the system detects the θ value of the highest peak called thetaMiddle, and also detects the width around this peak called thetaWidth, which is necessary to capture 40% of the energy in the histogram. The same is done for φ, and phiMiddle and phiWidth are recorded, but 80% energy capture is required for the width. The parameter extractor 14 can be configured to either obtain the detected pan parameter θ and / or the detected phase difference parameter φ and / or the detected magnitude U, or to determine the detected pan parameter θ and / or the detected phase difference parameter φ and / or the detected magnitude U from the left audio signal L and the right audio signal R (e.g., using Equation 7).

[0077] Details of time segment and frequency band Referring to FIG. 4a, a plurality of time segments 110, 120 from a first time segment to time segment t are shown. Each time segment 110, 120 includes a plurality of frequency bands 111, 112, 121, 122 from a first frequency band to frequency band B. The time segments 110, 120 and the frequency bands 111, 112, 121, 122 form a time-frequency tile representation, and each frequency band of each time segment forms a respective time / frequency tile.

[0078] As shown in FIG. 4a, each frequency band 111, 112, 121, 122 of each time segment 110, 120 is associated with a respective target pan parameter Θ and / or target phase difference parameter Φ. At least two frequency bands 111, 112, 121, 122 (or sub-bands), such as a high frequency band including frequencies above a threshold frequency and a low frequency band including frequencies below the threshold frequency, are defined for each time segment 110, 120. That is, each time segment 110, 120 includes frequency bands 1...B, where B is at least 2. Experiments have shown that octave frequency bands or quasi-octave frequency bands are suitable for targeting dialog audio sources such as frequency bands having edges at 0 Hz, 400 Hz, 800 Hz, 1600 Hz, 3200 Hz, 6400 Hz, 13200 Hz, and 24000 Hz. However, different band distributions are also possible.

[0079] The duration of each time segment can correspond to 10 ms to 1 s of the stereo audio signal. In some embodiments, each time segment corresponds to a plurality of frames of the stereo audio signal, such as 10 frames, and each frame represents, for example, 10 ms to 200 ms of the stereo audio signal, such as 50 ms to 100 ms of the stereo audio signal. Experiments have shown that representing one time segment using 10 overlapping frames with a length of 50 ms to 100 ms with a 75% overlap for each frame is a good trade-off between fast responsiveness, parameter stability, parameter reliability, and computational cost for normal dialog sources. However, for audio sources other than dialog, such as music, time segments with longer or shorter durations may also be considered.

[0080] FIG. 4b shows the partial mid-audio signal representation M extracted for each time segment 210, 220 and each frequency band 211, 212, 221, 222 (t,b)is shown. As can be seen in this figure, the first time segment 210 includes the partial mid-audio signal representation M of the first frequency band 211 (1,1) , the second partial mid-audio signal representation M of the second frequency band 212 (1,1) , etc., up to the last partial mid-audio signal representation M of the frequency band B (1,B) . By combining the partial mid-audio signal representations M (1,1) ...M (1,B) of the first time segment 210, the first time segment M of the target mid-audio signal M (1) is generated. Similarly, subsequent time segments M (2) ...M (t) of the target mid-audio signal M are generated by combining the partial mid-audio signal representations M (1,1) ...M (1,B) of the subsequent time segment 220. Finally, by combining each time segment M (1) ...M (t) in order, the target mid-audio signal M is generated.

[0081] Referring further to FIG. 4c, the partial side-audio signal representations S (t,b) extracted for each time segment 310, 320 and each frequency band 311, 312, 321, 322 are shown. In a manner similar to the combination of the partial mid-audio signal representations M (t,b) described above, the partial side-audio signal representations S (t,b) are combined to form the target side-audio signal S. It is understood that the extraction and combination of the partial mid-audio signal representations M (t,b) and / or the partial side-audio signal representations S (t,b) can be performed in a single step by a single processing unit or as two or more steps performed by two or more processing units.

[0082] Reconstruction of left audio signal and right audio signal from target mid-audio signal and target side-audio signal Figure 5a shows a block diagram of a synthesizer 20 that performs the reconstruction of the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S. The synthesizer unit 20 includes a reconstructor unit 22 that reconstructs the left audio signal L and the right audio signal R from the target mid audio signal M and the target side audio signal S using the pan parameter Θ and / or the phase difference parameter Φ. For example, the reconstructor unit 22 implements the following synthesis equations of the target mid audio signal M and the target side audio signal S extracted using equations 20 and 21.

Number

[0083] As seen in equations 22 and 23, the reconstruction of the left audio signal L and the right audio signal R is based on the weighted sum and weighted difference of the target mid audio signal M and the target side audio signal S. However, these equations can be modified to downweight or remove the contribution from the side signal. For example, the left audio signal L and the right audio signal R are the midweight coefficient of JPEG2025517590000022.jpg526 for the left signal and It is possible to reconstruct from only the target mid-audio signal M using each of the mid-weighting coefficients of JPEG2025517590000023.jpg526. Here, these weighting coefficients are based on at least one of the target pan parameter Θ and the target phase difference parameter Φ. In this example, the contribution from the side signal is 0. Other contributions between 0 and the values shown in equations 22 and 23 are also possible.

[0084] Similarly, the left audio signal L and the right audio signal R can also be reconstructed by taking into account the target side audio signal S. In such a case, the left audio signal L is based on a weighted sum of the target mid-audio signal M and the target side audio signal S, and the weights JPEG2025517590000024.jpg562 are based on at least one of the pan parameter and the phase difference parameter. In contrast, the right audio signal R is based on a weighted difference between the target mid-audio signal M and the target side audio signal S, and the weights JPEG2025517590000025.jpg562 are based on at least one of the pan parameter Θ and the phase difference parameter Φ. The contribution from the mid signal can also be reduced to down-weighting or 0.

[0085] FIG. 5b is a block diagram showing a synthesizer 20 that performs reconstruction using equations 22 and 23 having an alternative pan parameter Θ ALT and / or an alternative phase difference parameter Φ ALT and as a result gives an alternative left audio signal L ALT and an alternative right audio signal R ALT In the illustrated embodiment, for all time segments and frequency bands, to generate a centered panned stereo audio signal including a left audio signal L CEN and a right audio signal R CEN having no inter-channel phase difference, Θ ALT = Θ CEN= π / 4, and Φ ALT = Φ CEN = 0.

[0086] The alternative mid-weighting coefficient and the alternative side-weighting coefficient are obtained by replacing Θ and Φ of the mid-weighting coefficient and the side-weighting coefficient described in Equations 22 and 23 with the corresponding alternative parameters Θ ALT , Φ ALT obtained for each time segment and each frequency band. In some embodiments, the alternative left audio signal L ALT and the alternative right audio signal R ALT are generated using only the target mid-audio signal M and the alternative mid-weighting coefficient (e.g., the coefficients of the target mid-audio signal M in Equations 22 and 23 having Θ ALT , Φ ALT instead of Θ, Φ). Alternatively, the alternative left audio signal L ALT and the alternative right audio signal R ALT are generated using both the target mid-audio signal M and the target side-audio signal S, and the alternative mid-weighting coefficient and the alternative side-weighting coefficient (e.g., the coefficients of the target mid-audio signal M and the target side-audio signal S in Equations 22 and 23 having Θ ALT , Φ ALT instead of Θ, Φ respectively).

[0087] Processing device for implementing target mid-audio signal and target side-audio signal In the following, FIGS. 6a, 6b, and 6c show a block diagram of an audio processing system including an analysis device 10 and a synthesis device 20 in combination with additional audio processing units such as a monosource separator 30, a mid / side processor 40, a stereo processor 50, and a stereo soft mask estimator 60 that perform audio processing on the target mid-audio signal M and the target side-audio signal S or the alternative left audio signal L CEN and the alternative right audio signal R CEN .

[0088] Figure 6a shows an audio processing system that processes a stereo audio signal via a target mid audio signal M and a target side audio signal S extracted by an analysis device 10. As can be seen in this figure, the target mid audio signal M is used as an input to a mono source separator 30. In addition, the target mid audio signal M and the target side audio signal S are provided to a mid / side processor 40 that applies enhancement and / or attenuation to the audio signal.

[0089] The target mid audio signal M is provided to a mono source separation system 30 configured to separate the audio source in the target mid audio signal M and generate a source-separated mid audio signal M sep For example, the mono source separation system 30 generates a source-separated mid audio signal M that improves the intelligibility of at least one audio source present in the target mid audio signal M sep

[0090] The source-separated mid audio signal M sep is provided to the mid / side processor 40 together with the target side audio signal S, and a processed mid audio signal M’ and side audio signal S’ are generated. The mid / side processor 40 processes the source-separated mid audio signal M sep ​can be applied to the gain and / or the target side signal S can be attenuated. In some cases, the mid / side processor 40 sets the target side audio signal S to 0. The processed mid audio signal M’ and side audio signal S’ are then provided to a synthesizing device 20 including a left and right audio signal reconstruction unit 22. The left and right audio signal reconstruction unit 22 reconstructs the processed left audio signal L’ and right audio signal R’ based on the processed mid audio signal M’ and side audio signal S’, and the pan parameter Θ and / or phase difference parameter Φ obtained by the analysis device 10.

[0091] In some embodiments, the source-separated mid audio signal M sep is provided directly to the synthesizing device 20, and the synthesizing device 20 reconstructs the processed left audio signal L’ and right audio signal R’ based on at least the source-separated mid audio signal M sep (and optionally the target side audio signal S) and the target pan parameter Θ and / or target phase difference parameter Φ.

[0092] Alternatively, the monaural source separation system 30 is omitted, and the target mid audio signal M and target side audio signal S are provided directly to the mid-side audio processing unit 40, and the processing unit 40 extracts the processed mid audio signal M’ and side audio signal S’ from the target mid audio signal M and target side audio signal S.

[0093] FIG. 6b shows how the target mid audio signal M and target side audio signal S are used to extract the left audio signal L CEN and right audio signal R CEN located at the center alternately and enable the stereo processing system 50 to be used for the center-panned stereo audio signal.

[0094] Analysis device 10a acquires an arbitrary left audio signal L and right audio signal R, and target pan parameter Θ and / or target phase difference parameter Φ for each time segment and each frequency band, and extracts a target mid audio signal M and a target side audio signal S. The target mid audio signal M and the target side audio signal S are provided to a synthesizing device 20a, and the synthesizing device 20a acquires a set of an alternative center-positioned pan parameter Θ CEN and / or phase difference parameter Φ CEN as well. Therefore, the analysis device 10a and the synthesizing device 20a generate a center-panned stereo signal including the center-positioned left audio signal L CEN and right audio signal R CEN from an arbitrary original stereo signal.

[0095] The center-panned alternative left audio signal L CEN and alternative right audio signal R CEN are provided to a stereo processing system 50. The stereo processing system 50 is configured to perform stereo source separation on the center-panned stereo source and output a processed center-positioned left audio signal L’ CEN and right audio signal R’ CEN . The processed center-positioned left audio signal L’ CEN and right audio signal R’ CEN are characterized by, for example, improved intelligibility of at least one audio source existing in the center-positioned left audio signal L CEN and right audio signal R CEN .

[0096] The processed center-positioned left audio signal L’ CEN and right audio signal R’ CEN are then provided to a second analysis device 10b, and the second analysis device 10b processes the center-positioned left audio signal L’ CENand the right audio signal R' CEN and the pan parameter Θ located at the center CEN and / or the phase difference parameter Φ CEN are used to extract the processed mid audio signal M' and side audio signal S'. Finally, the second synthesizer 20b reconstructs the processed left audio signal L' and right audio signal R' using the original pan parameter Θ and / or phase difference parameter Φ obtained or determined by the first analyzer 10a.

[0097] The pan parameter Θ located at the center indicating a center pan with no inter-channel phase difference CEN and / or the phase difference parameter Φ CEN is one option among many possible values of the alternative pan parameter Θ ALT and / or the alternative phase difference parameter Φ ALT Note that it is. For example, the alternative pan parameter Θ located at the center CEN , Φ CEN can be replaced by any alternative pan parameter Θ that indicates a strictly right or left stereo audio signal, regardless of the presence or absence of a non-zero phase difference, for example ALT , Φ ALT .

[0098] It should be further noted that the target mid-audio signal M and the target side-audio signal S supplied between the first analysis device 10a and the first synthesis device 20a can undergo mono signal separation processing and / or mid-side processing as described with respect to FIG. 6 above. The same applies to the processed target mid-audio signal M' and target side-audio signal S' supplied between the second analysis device 10b and the second synthesis device 20b. That is, the mono source separator 30 and / or the mid / side processor 40 from FIG. 6a can be used to process the mid-audio signal and the side-audio signal passed between at least one of between the first analysis device 10a and the first synthesis device 20a and between the second analysis device 10b and the second synthesis device 20b. Additionally, although the first analysis device 10a and the first synthesis device 20a are shown as separate units, these devices can be combined into a single unit, and referring to equations 20, 21, 22, 23, the alternative left audio signal L CEN and the alternative right audio signal R CEN It is clear that the operations to generate can be performed as a single step. The same applies to the second analysis device 10b and the second synthesis device 20b.

[0099] In some embodiments, the alternative (e.g., centrally located) left audio signal L CEN and the alternative right audio signal R CEN are generated by the first synthesis device 20a using only the target mid-audio signal M. Similarly, the second analysis device 10b can be configured to extract only the processed target mid-audio signal M'.

[0100] FIG. 6c shows an analysis device 10 that receives a left audio signal L and a right audio signal R and operates with a filter estimator configured to estimate a filter for source separation. This filter estimator is, for example, a stereo soft mask estimator 60 configured to receive the left audio signal L and the right audio signal R and estimate a soft mask F. The stereo soft mask estimator 60 can be any soft mask estimator, and can be configured to output, for example, a soft mask F for stereo source separation. Conventionally, the soft mask F has been applied to the left audio signal L and the right audio signal R of the stereo audio signal.

[0101] In the embodiment shown in FIG. 6c, the estimated stereo soft mask F is applied to the target mid audio signal M to form a source-separated mid audio signal M sep The source-separated mid audio signal M sep is provided to a mid / side processor 40 together with the target side audio signal S, and the mid / side processor 40 outputs a processed target mid audio signal M' and a processed target side audio signal S'. The processed side audio signal S' can be downweighted, including downweighting to 0, compared to S. The processed target mid audio signal M' and the processed target side audio signal S' are provided to a synthesis unit 20, and the synthesis unit 20 reconstructs a processed left audio signal L' and a right audio signal R'. Optionally, the mid / side processor is omitted, whereby the source-separated mid audio signal M sep , and optionally the target side audio signal S, are provided directly to the synthesis unit 20 for reconstruction.

[0102] The stereo soft mask estimator 60 is configured to obtain or acquire the binaural parameter Θ and / or the phase difference parameter Φ and provide these parameters to the analysis device 10. Alternatively, the binaural parameter Θ and / or the phase difference parameter Φ may be obtained from other locations by the analysis device 10, or may be obtained by the analysis device 10 based on the left audio signal L and the right audio signal R.

[0103] Unless otherwise specified, as will be apparent from the following discussion, throughout this disclosure, discussions using terms such as "processing", "computing", "calculating", "obtaining", "analyzing", etc. refer to the operation and / or process of a computer hardware or computing system, or similar electronic computing device, that manipulates data represented as physical quantities such as electronic quantities and / or converts these data into other data also represented as physical quantities.

[0104] In the foregoing description of exemplary embodiments of the present invention, for the purposes of streamlining the disclosure and aiding in the understanding of one or more aspects of the various inventions, it is to be understood that the various features of the present invention may, in some cases, be grouped together in a single embodiment, a figure of this embodiment, or the description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. To the contrary, as the following claims reflect, aspects of the present invention lie in less than all of the features of the single disclosed embodiment described above. Thus, the claims following this detailed description hereby expressly incorporate each claim, in its independent capacity, as an individual embodiment of the present invention. Further, although some embodiments described herein include some features that are not other features included in other embodiments, as will be understood by those of ordinary skill in the art, combinations of features of different embodiments are within the scope of the present invention and are intended to form different embodiments. For example, in the following claims, any of the embodiments recited in the claims can be used in any combination.

[0105] Furthermore, some of the embodiments are described herein as a method or a combination of method elements that can be implemented by a processor of a computer system or by other means that execute a function. Thus, a processor having instructions necessary to execute such a method or method element forms means for executing the method or method element. Note that when a method includes several elements, for example several steps, the ordering of such elements is not implied unless otherwise specified. Further, the elements described herein of the apparatus embodiments are an example of means for performing the functions performed by this element for the purpose of carrying out the present invention. The description provided herein sets forth numerous specific details. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description. Those skilled in the art will recognize that the present invention is in no way limited to the preferred embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, the stereo processing system in FIG. 6c may be a mid-audio signal M or a source-separated mid-audio signal M sepIt can further include a monosource separator 30 that operates on . As a further example, it is assumed that all components or processes related to the formation of the target mid-audio signal and the target side-audio signal, and the processing of the mid-audio signal and the side-audio signal, can be configured to form or process only the target mid-audio signal, or the target mid-audio signal and the target side-audio signal. Similarly, reconstructing the left audio signal and the right audio signal (e.g., alternative or centrally located left and right audio signals) based on the target mid-audio signal and the target side-audio signal can include reconstructing the left and right audio signals using only the target mid-audio signal, or using both the target mid-audio signal and the target side-audio signal.

[0106] Various features and aspects will be understood from the following enumerated exemplary embodiments (「EEE」:enumerated exemplary embodiment).

[0107] EEE1. A method for extracting a target mid-audio signal (M) from a stereo audio signal, the stereo audio signal including a left audio signal (L) and a right audio signal (R), the method comprising: obtaining (S1) a plurality of consecutive time segments of the stereo audio signal, each time segment including a representation of a portion of the stereo audio signal; for each frequency band of a plurality of frequency bands of each time segment of the stereo audio signal, a target pan parameter (Θ) representing the distribution over the time segment of the magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band; A target phase difference parameter (Φ) representing the distribution of the phase difference between the left audio signal (L) and the right audio signal (R) of the stereo audio signal over the time segment, acquiring at least one of (S2); extracting partial mid-signal representations (211, 212) for each time segment and each frequency band, wherein the partial mid-signal representations (211, 212) are based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) of each frequency band (S3); forming the target mid-audio signal (M) by combining the partial mid-signal representations (211, 212) of each frequency band and each time segment (S4); A method comprising.

[0108] EEE2. The method according to EEE1, wherein the weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight.

[0109] EEE3. The method according to EEE1 or 2, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).

[0110] EEE4. The method according to EEE3, wherein the complex-valued weights are based on the target phase difference parameter (Φ) and the target pan parameter (Θ) such that the phase difference generated by the application of the complex weights is smaller with respect to the left audio signal or the right audio signal having a larger magnitude.

[0111] providing the stereo audio signal to a stereo source separation system (60) configured to output a filter for stereo source separation; applying the filter to the target mid audio signal to form a processed target mid audio signal (M sep ); The method according to any one of EEE1 to 4, further comprising.

[0112] providing the target mid audio signal (M) to a mono source separation system (30) configured to perform audio source separation on a mono audio signal and output a processed mono audio signal; using the processed mono audio signal as the processed target mid audio signal (M sep ); The method according to any one of EEE1 to 5, further comprising.

[0113] reconstructing a processed left audio signal (L') and a processed right audio signal (R') that form a processed stereo audio signal; further comprising, each of the processed left audio signal (L') and the processed right audio signal (R') is based on the processed target mid audio signal (M sep ) weighted using respective left and right weighting factors, the left and right weighting factors being based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band. The method according to EEE5 or 6.

[0114] reconstructing an alternative left audio signal (L cen ) and an alternative right audio signal (R cen ) that form an alternative stereo audio signal; further comprising, the substitute left audio signal (L cen ) and the substitute right audio signal (R cen ) are each based at least on the weighted contribution of the target mid audio signal, each contribution is based on a respective substitute midweighting coefficient, and each midweighting coefficient is based on at least one of the substitute pan parameter (Θ cen ) and the substitute phase difference parameter (Φ cen ) for each time segment and each frequency band, the method according to any one of EEE1 to 7.

[0115] EEE9. Performing stereo signal processing on the substitute stereo audio signal including the processed substitute left audio signal (L’ cen ) and the processed substitute right audio signal (R’ cen ) to form a processed substitute stereo audio signal, Reconstructing the processed target mid audio signal (M’’), further comprising, the processed target mid audio signal (M’) is based on the sum of the processed substitute left audio signal (L’ cen ) and the processed substitute right audio signal (R’ cen ) weighted using a weighting coefficient, the weighting coefficient is based on at least one of the substitute pan parameter (Θ cen ) and the substitute phase difference parameter (Φ cen ) for each frequency band and each time segment, the method according to EEE8.

[0116] EEE10. The substitute pan parameter (Θ cen ) and / or the substitute phase difference parameter (Φ cen ) indicates center pan, the method according to EEE8 or 9.

[0117] EEE11. Extracting partial side signal representations (311, 312) for each time segment and each frequency band (S31), where the partial side signal representations (311, 312) are based on a weighted difference between the left audio signal (L) and the right audio signal (R), and the weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment. Forming a target side audio signal by combining the partial side signal representations for each frequency band and each time segment (S41). The method according to any one of EEE1 to EEE10, further comprising.

[0118] EEE12. Executing processing (S5, S51) on at least one of the target mid audio signal (M) and the target side audio signal (S) to form a processed target mid audio signal (M') and a processed target side audio signal (S'). Reconstructing a processed left audio signal (L') and a processed right audio signal (R') that form a processed stereo audio signal (S6). Further comprising The processed left audio signal (L') is based on a weighted sum of the processed target mid audio signal (M') and the processed target side audio signal (S'), and the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The processed right audio signal (R') is based on a weighted difference between the processed target mid audio signal (M') and the processed target side audio signal (S'), and the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The method according to EEE11.

[0119] EEE13. Performing the processing of at least one of the target mid-audio signal (M) and the target side-audio signal (S) includes attenuating the target side-audio signal (S), applying a gain to the target mid-audio signal (M), performing mono signal source separation on the target mid-audio signal (M), applying a stereo source separation filter to the target mid-audio signal (M), and including at least one of the above, the method according to EEE12.

[0120] EEE14. Obtaining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) includes obtaining at least one of a plurality of detected pan parameters (θ) and a plurality of detected phase difference parameters (φ), for each detected pan parameter (θ) and each detected phase difference parameter (φ), obtaining a detected magnitude parameter (U), and obtaining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) by calculating an average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), where the average is a weighted average weighted by the detected magnitude parameter (U), and including the above, the method according to any one of EEE1 to EEE13.

[0121] EEE15. Further including sorting each of the at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ) into a predetermined number of bins Each bin is associated with at least one of the detected pan parameter values and the detected phase difference parameter values, and the average of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter values and the detected phase difference parameter values, according to the method described in EEE14.

[0122] EEE16. An audio processing system configured to extract a target mid audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), and the audio processing system obtains a plurality of consecutive time segments of the stereo audio signal, each time segment including a representation of a portion of the stereo audio signal, and for each frequency band of a plurality of frequency bands of each time segment of the stereo audio signal, obtains at least one of a target pan parameter (Θ) representing the distribution over the time segment of the magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band, and a target phase difference parameter (Φ) representing the distribution over the time segment of the phase difference between the left audio signal (L) and the right audio signal (R) in the frequency band, and extracts a partial mid-signal representation (211, 212) for each time segment and each frequency band, the partial mid-signal representation (211, 212) being based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) being based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) of each frequency band. ​Forming the target mid audio signal (M) by combining the partial mid signal representations for each frequency band and each time segment; An audio processing system comprising an extractor unit (12) configured to perform the above.

[0123] EEE17. The audio processing system according to EEE16, wherein the weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight.

[0124] EEE18. The audio processing system according to EEE16 or 17, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).

[0125] EEE19. The audio processing system according to EEE18, wherein the complex-valued weights are based on the target phase difference parameter (Φ) and the target pan parameter (Θ) such that the phase difference generated by the application of the complex weights becomes smaller with respect to the left audio signal or the right audio signal having a larger magnitude.

[0126] EEE20. A stereo source separation system (60) configured to output a filter for stereo source separation, further comprising, the stereo audio signal is provided to the stereo source separation system (60), and the filter for stereo source separation is applied to the target mid audio signal to form a processed target mid audio signal (M sep ), the audio processing system according to any one of EEE16 to 19.

[0127] A mono-source separation system (30) configured to perform audio source separation on a mono audio signal and output a processed mono audio signal further comprising wherein the target mid audio signal (M) is provided to the mono-source separation system (30), and the processed mono audio signal is used as a processed target mid audio signal (M sep ), the audio processing system according to any one of EEE16 to 20

[0128] A reconstructor unit (20) configured to reconstruct a processed left audio signal (L') and a processed right audio signal (R') that form a processed stereo audio signal further comprising wherein each of the processed left audio signal (L') and the processed right audio signal (R') is based on the processed target mid audio signal (M sep ) weighted using respective left and right weighting factors, and the left and right weighting factors are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band, the audio processing system according to EEE20 or 21

[0129] A reconstructor unit (20) configured to reconstruct an alternative left audio signal (L cen ) and an alternative right audio signal (R cen ) that form an alternative stereo audio signal further comprising wherein the alternative left audio signal (L cen ) and the alternative right audio signal (R cenEach of them is based at least on the weighted contribution of the target mid-audio signal, each contribution is based on respective alternative mid-weighting coefficients, and each mid-weighting coefficient is based on at least one of alternative pan parameters (Θ cen ) and alternative phase difference parameters (Φ cen ) in an audio processing system according to any one of EEE16 - 22.

[0130] EEE24. A stereo processor (50) configured to perform stereo signal processing on the processed alternative stereo audio signal including the processed alternative left audio signal (L’ cen ) and the processed alternative right audio signal (R’ cen ) before forming the processed alternative stereo audio signal; A second extractor unit configured to extract the processed target mid-audio signal (M’’); further comprising The processed target mid-audio signal (M’) is based on the sum of the processed alternative left audio signal (L’ cen ) and the processed alternative right audio signal (R’ cen ) weighted using a weighting coefficient, The weighting coefficient is based on at least one of the alternative pan parameters (Θ cen ) and the alternative phase difference parameters (Φ cen ) in each frequency band and each time segment in an audio processing system according to any one of EEE16 - 23.

[0131] EEE25. The alternative pan parameter (Θ cen ) and / or the alternative phase difference parameter (Φ cen ) indicates center pan in an audio processing system according to EEE23 or 24.

[0132] EEE26. The extractor unit 12 For each time segment and each frequency band, extracting a partial side signal representation (311, 312), wherein the partial side signal representation (311, 312) is based on a weighted difference between the left audio signal (L) and the right audio signal (R), and the weights of the left audio signal (L) and the right audio signal (R) respectively are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment, forming a target side audio signal by combining each partial side signal representation for each frequency band and each time segment, The audio processing system according to any one of EEE16 - 25, further configured to perform the above.

[0133] EEE27. A mid / side processor (40) configured to process at least one of the target mid audio signal (M) and the target side audio signal (S) to form a processed target mid audio signal (M') and a processed target side audio signal (S'), A reconstructor unit (20) configured to reconstruct a processed left audio signal (L') and a processed right audio signal (R') that form a processed stereo audio signal, further comprising, The processed left audio signal (L') is based on a weighted sum of the processed target mid audio signal (M') and the processed target side audio signal (S'), and the weights are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ), The processed right audio signal (R’) is based on a weighted difference between the processed target mid audio signal (M’) and the processed target side audio signal (S’), and the weight is based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ), the audio processing system according to any one of EEE16 to 26.

[0134] EEE28. The processing executed by the mid / side processor (40) is attenuation of the target side audio signal (S), application of a gain to the target mid audio signal (M), mono signal source separation executed on the target mid audio signal (M), application of a stereo source separation filter to the target mid audio signal (M), the audio processing system according to EEE27, including at least one of

[0135] EEE29. Obtaining at least one of a plurality of detected pan parameters (θ) and a plurality of detected phase difference parameters (φ), for each detected pan parameter (θ) and each detected phase difference parameter (φ), obtaining a detected magnitude parameter (U), determining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) by calculating an average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), the average being a weighted average weighted by the detected magnitude parameter (U), the audio processing system according to any one of EEE16 to 28, further comprising a parameter extraction unit (14) configured to perform

[0136] EEE30. The parameter extractor unit (14) is further configured to sort each of the at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ) into a predetermined number of bins, Each bin is associated with at least one of the detected pan parameter value and the detected phase difference parameter value, and the average of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter value and the detected phase difference parameter value, the audio processing system according to EEE29.

[0137] EEE31. A computer program product comprising instructions that, when the program is executed by a computer, cause the computer to execute the method according to any one of EEE1-15.

[0138] EEE32. The magnitude of each of the left audio signal and the right audio signal is obtained, The method or the audio processing device is, in response to a determination that the magnitude of the left audio signal is greater than the magnitude of the right audio signal, determining that the weight of the left audio signal is greater than the weight of the right audio signal; in response to a determination that the magnitude of the right audio signal is greater than the magnitude of the left audio signal, determining that the weight of the right audio signal is greater than the weight of the left audio signal; comprising the method according to EEE2 or any one of EEE3-15 when dependent on EEE2, or the audio processing device according to EEE17 or any one of EEE18-30 when dependent on EEE17.

[0139] EEE33. The magnitude of each of the left audio signal and the right audio signal is obtained, The method or the audio processing device determines the weight of the left audio signal to cause a first phase shift when applied to the left audio signal, and determines the weight of the right audio signal to cause a second phase shift when applied to the right audio signal, in response to a determination that the magnitude of the left audio signal is greater than the magnitude of the right audio signal; determines the weight of the right audio signal to cause a first phase shift when applied to the right audio signal, and determines the weight of the left audio signal to cause a second phase shift when applied to the left audio signal, in response to a determination that the magnitude of the right audio signal is greater than the magnitude of the left audio signal; and includes wherein the first phase shift is smaller than the second phase shift, the method according to EEE4 or any one of EEE5 to 15 when dependent on EEE4, or the audio processing device according to EEE19 or any one of EEE20 to 30 when dependent on EEE19.

[0140] EEE34. The alternative left audio signal and the alternative right audio signal are a second left audio signal and a second right audio signal, respectively, and / or the alternative stereo signal is a second stereo signal, the method according to EEE8 or any one of EEE9 to 15 when dependent on EEE8, or the audio processing device according to EEE23 or any one of EEE24 to 30 when dependent on EEE23.

[0141] EEE35. Performing stereo signal processing on the alternative stereo audio signal (L cen , R cen ) means that the alternative stereo audio signal (L cen , R cen) Apply source separation to obtain the processed alternative left audio signal (L’ cen ) and the processed alternative right audio signal (R’ cen ) and output them. The method according to EEE9 or 10, or the method according to any one of EEE11 - 15 when dependent on EEE9 or 10.

Claims

1. A computer-implemented method for extracting a target mid-audio signal (M) from a stereo audio signal, wherein the stereo audio signal includes a left audio signal (L) and a right audio signal (R), and the method comprises: obtaining (S1) a plurality of consecutive time segments of the stereo audio signal, each time segment including a representation of a portion of the stereo audio signal; for each frequency band of a plurality of frequency bands of each time segment of the stereo audio signal, obtaining (S2) at least one of: a target pan parameter (Θ) representing a distribution over the time segment of a magnitude ratio between the left audio signal (L) and the right audio signal (R) in the frequency band; and a target phase difference parameter (Φ) representing a distribution over the time segment of the phase difference between the left audio signal (L) and the right audio signal (R) of the stereo audio signal; extracting (S3) partial mid-signal representations (211, 212) for each time segment and each frequency band, the partial mid-signal representations (211, 212) being based on a weighted sum of the left audio signal (L) and the right audio signal (R), and the respective weights of the left audio signal (L) and the right audio signal (R) being based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) of each frequency band; forming (S4) the target mid-audio signal (M) by combining the partial mid-signal representations (211, 212) of each frequency band and each time segment. A method as described above.

2. The method according to claim 1, wherein the weights of the left audio signal (L) and the right audio signal (R) are based on the target pan parameter (Θ) such that the left audio signal or the right audio signal having a larger magnitude is given a larger weight.

3. The method according to claim 1 or 2, wherein the weights of the left audio signal (L) and the right audio signal (R) are complex-valued weights, and the phase of the complex-valued weights is based on the target phase difference parameter (Φ).

4. ​ ​ The method according to claim 3, wherein the complex-valued weight is based on the target phase difference parameter (Φ) and the target pan parameter (Θ) such that the phase difference generated by the application of the complex weight becomes smaller with respect to the left audio signal or the right audio signal having a larger magnitude.

5. Providing the stereo audio signal to a stereo source separation system (60) configured to output a filter for stereo source separation; Apply the filter to the target mid-audio signal to form a processed target mid-audio signal (M sep ). The method according to any one of claims 1 to 4, further comprising:

6. Providing the target mid audio signal (M) to a mono source separation system (30) configured to perform audio source separation on a mono audio signal and output a processed mono audio signal; Use the processed mono audio signal as the processed target mid audio signal (M sep ), and The method according to any one of claims 1 to 5, further comprising:

7. Reconstructing a processed left audio signal (L') and a processed right audio signal (R') that form a processed stereo audio signal; further comprising Each of the processed left audio signal (L') and the processed right audio signal (R') is based on the processed target mid audio signal (M) weighted using respective left and right weighting factors sep ) and the left and right weighting factors are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each time segment and each frequency band, the method according to claim 5 or 6.

8. Reconstructing an alternative left audio signal (L cen ), and an alternative right audio signal (R cen ) that form an alternative stereo audio signal. further comprising The said substitute left audio signal (L cen ), and the said substitute right audio signal (R cen ), each of which is based on at least the weighted contribution of the said target mid audio signal, each contribution is based on respective substitute mid weight coefficients, and each mid weight coefficient is based on at least one of the substitute pan parameters (Θ cen ) and substitute phase difference parameters (Φ cen ) of each time segment and each frequency band. The method according to any one of claims 1 to 7.

9. Processed alternative left audio signal (L’ cen ), and processed alternative right audio signal (R’ cen ), performing stereo signal processing on the alternative stereo audio signal to form a processed alternative stereo audio signal including the Reconstructing a processed target mid audio signal (M''); further comprising The processed target mid-audio signal (M') is based on the sum of the processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen ) weighted using a weight coefficient, The weight coefficient is based on at least one of the alternative pan parameters (Θ cen ) and the alternative phase difference parameter (Φ cen ) for each frequency band and each time segment. The method according to claim 8.

10. The substitution parameter (Θ cen ) and / or the substitution phase difference parameter (Φ cen ) indicates a center pan, the method according to claim 8 or 9.

11. Performing stereo signal processing on the alternative stereo audio signal (L cen , R cen ) includes applying source separation to the alternative stereo audio signal (L cen , R cen ) to output the processed alternative left audio signal (L' cen ) and the processed alternative right audio signal (R' cen ), the method according to claim 9 or 10.

12. For each time segment and each frequency band, extracting a partial side signal representation (311, 312) (S31), wherein the partial side signal representation (311, 312) is based on a weighted difference between the left audio signal (L) and the right audio signal (R), and the weights of the left audio signal (L) and the right audio signal (R) are based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) for each frequency band and each time segment; Forming a target side audio signal by combining each partial side signal representation for each frequency band and each time segment (S41); The method according to any one of claims 1 to 11, further comprising:

13. Performing at least one of the processing (S5, S51) on the target mid-audio signal (M) and the target side-audio signal (S) to form a processed target mid-audio signal (M') and a processed target side-audio signal (S'); Reconstructing the processed left audio signal (L') and the processed right audio signal (R') that form the processed stereo audio signal (S6); Further comprising; The processed left audio signal (L') is based on a weighted sum of the processed target mid-audio signal (M') and the processed target side-audio signal (S'), and the weight is based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ); The processed right audio signal (R') is based on a weighted difference between the processed target mid-audio signal (M') and the processed target side-audio signal (S'), and the weight is based on at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ). The method according to claim 12.

14. Performing at least one of the processing on the target mid-audio signal (M) and the target side-audio signal (S) includes: Attenuating the target side-audio signal (S); Applying a gain to the target mid-audio signal (M); Performing mono signal source separation on the target mid-audio signal (M); Applying a stereo source separation filter to the target mid-audio signal (M); The method according to claim 13, including at least one of the above.

15. Obtaining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) includes: Obtaining at least one of a plurality of detected pan parameters (θ) and a plurality of detected phase difference parameters (φ); For each detected pan parameter (θ) and each detected phase difference parameter (φ), obtaining a detected magnitude parameter (U); Determining at least one of the target pan parameter (Θ) and the target phase difference parameter (Φ) by calculating an average of at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ), wherein the average is a weighted average weighted by the detected magnitude parameter (U); The method according to any one of claims 1 to 14, comprising: **Claim 16** Sorting each of the at least one of the plurality of detected pan parameters (θ) and the plurality of detected phase difference parameters (φ) into a predetermined number of bins; further comprising each bin is associated with at least one of the detected pan parameter value and the detected phase difference parameter value, and the average of at least one of the plurality of detected pan parameters and the plurality of detected phase difference parameters (φ) is based on the detected pan parameter value and the detected phase difference parameter value. The method according to claim 15.

Citation Information

Patent Citations

  • Signal processor, signal processing method and computer-readable recording medium recording signal processing program

    JP1999289599A

  • Sound signal processor and sound signal processing method

    JP2006080708A

  • Sound separating device, sound separating method, sound separating program, and computer-readable recording medium

    WO2006090589A1

  • Perceptual optimization of magnitude and phase for time-frequency and softmask source separation systems

    WO2021252795A2

  • Separation of panned sources from generalized stereo backgrounds using minimal training

    WO2021252912A1