Method, apparatus, and system for detecting and extracting spatially recognizable subband audio sources
Through the time-frequency domain representation and spatial parameter processing methods, the problem of extracting space-recognizable subband audio sources from dual-channel audio mixing is solved, and efficient and robust audio source extraction is achieved, suitable for amplitude translation, delay mixing and band changes audio sources.
Patent Information
- Application Number
- CN202180041824.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-11
- Filing Date
- 2021-06-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-06-11
AI Technical Summary
The prior art is difficult to efficiently detect and extract spatially identifiable subband audio sources from two-channel audio mixing, especially in the presence of amplitude translation, delay mixing, reverb or band changes.
Using the time frequency domain representation and spatial parameter processing methods, the spatial parameters and levels of the time frequency chip are calculated, these parameters are modified using shift and extrusion parameters, soft mask values are generated, and applied to the time frequency chip to generate an estimated audio source, and finally the time frequency chip is assembled into chunks for further processing.
Efficient and robust extraction of spatially identifiable subband audio sources from dual channel mixing, capable of handling amplitude translation, delayed mixing, and band changes, requiring little training data or delay.
Smart Images

Figure CN115715413B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 038,048, filed on June 11, 2020, and European Patent Application No. 20179447.6, filed on June 11, 2020, each of which is incorporated herein by reference in its entirety. Technical field
[0003] The present disclosure generally relates to audio signal processing, and more particularly to audio source separation techniques. Background art
[0004] Two - channel audio mixing (e.g., stereo mixing) is created by mixing multiple audio sources together. There are several examples where it is desirable to detect and extract individual audio sources from a two - channel mix, including but not limited to: remixing applications, where audio sources are re - positioned in a two - channel mix; up - mixing applications, where audio sources are positioned or re - positioned in a surround - sound mix; and audio source enhancement applications, where certain audio sources (e.g., speech / dialogue) are boosted and added back to a two - channel or surround - sound mix. Summary of the invention
[0005] Details of the disclosed embodiments are set forth in the following drawings and description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims.
[0006] In an embodiment, a method includes: transforming, using one or more processors, one or more frames of a two - channel time - domain audio signal into a time - frequency domain representation including a plurality of time - frequency tiles, wherein a frequency domain of the time - frequency domain representation includes a plurality of frequency bins that are grouped into a plurality of sub - bands; for each time - frequency tile: calculating, using one or more processors, a spatial parameter and a level of the time - frequency tile; modifying, using one or more processors, the spatial parameter using shift and squeeze parameters; obtaining, using one or more processors, a softmask value for each frequency bin using the modified spatial parameter, the level, and sub - band information; and applying, using one or more processors, the softmask value to the time - frequency tile to generate a modified time - frequency tile of the estimated audio source.
[0007] In an embodiment, a plurality of frames of time-frequency slices are assembled into a plurality of chunks, each chunk including a plurality of subbands, and the method includes: for each subband in each chunk: using one or more processors to calculate spatial parameters and levels for each time-frequency slice in the chunk; using one or more processors to modify the spatial parameters using shift parameters and squash parameters; using one or more processors to obtain soft mask values for each frequency bin using the modified spatial parameters, levels, and subband information; and using one or more processors to apply the soft mask values to the time-frequency slices to generate modified time-frequency slices of the estimated audio source.
[0008] In an embodiment, the method further includes: using one or more processors to transform the modified time-frequency slices into a plurality of time-domain audio source signals.
[0009] In an embodiment, the spatial parameters include panning and phase difference for each of the time-frequency slices.
[0010] In an embodiment, the method includes: for each subband, determining the statistical distribution of the panning parameter and the statistical distribution of the phase difference parameter; determining the shift parameters as the panning parameter and the phase difference parameter corresponding to the peaks of the respective statistical distributions of the panning parameter and the phase difference parameter; and determining the squash parameter as the width around the peaks of the respective distributions of the panning parameter and the phase difference parameter to capture a predetermined amount of audio energy.
[0011] In an embodiment, the predetermined amount of audio energy is at least forty percent of the total energy in the statistical distribution of the panning parameter and at least eighty percent of the total energy in the statistical distribution of the phase difference parameter.
[0012] In an embodiment, the soft mask values are obtained from a look-up table or function of a spatial level filtering (SLF) system trained for a center-panned target source.
[0013] In an embodiment, transforming one or more frames of a stereo time-domain audio signal into a frequency-domain signal includes applying a short-time Fourier transform (STFT) to the stereo time-domain audio signal.
[0014] In an embodiment, the plurality of frequency bins are grouped into octave subbands or approximate octave subbands.
[0015] In an embodiment, the spatial parameters include a translation parameter and a phase difference parameter for each time-frequency slice, and calculating the shift parameter and the squeeze parameter further includes: optionally assembling consecutive frames of the time-frequency slice into chunks, each chunk including a plurality of subbands; for each subband in each chunk: creating a smoothed level-parameter-weighted histogram on the translation parameter; creating a smoothed level-parameter-weighted first phase difference histogram on a first phase difference parameter, wherein the first phase difference parameter has a first range; creating a smoothed level-parameter-weighted second phase difference histogram on a second phase difference parameter, wherein the second phase difference parameter has a second range different from the first range; detecting a translation peak in the smoothed translation histogram; determining the translation peak width; determining the translation median; detecting a first phase difference peak in the smoothed first phase difference histogram; determining the first phase difference peak width; determining the first phase difference median; detecting a second phase difference peak in the smoothed second phase difference histogram; determining the second phase difference peak width; and determining the second phase difference median, wherein the shift parameter includes the translation median and either the first phase difference median or the second phase difference median, and the squeeze parameter includes the translation peak width and either the first phase difference peak width or the second phase difference peak width. The statistical distribution of the translation parameter in the above embodiment may include a smoothed level-parameter-weighted histogram on the translation parameter. The statistical distribution of the phase difference parameter may include a first phase histogram and a second phase histogram. Determining the translation parameter corresponding to the peak of the statistical distribution of the translation parameter and the width around the peak of the statistical distribution of the translation parameter may include: detecting the translation peak, determining the translation peak width, and determining the translation median. Determining the phase difference parameter corresponding to the peak of the statistical distribution of the phase difference parameter and the width around the peak of the statistical distribution of the phase difference parameter may include: detecting the first phase difference peak and the second phase difference peak, determining the first phase difference peak width and the second phase difference peak width, and determining the first phase difference median and the second phase difference median.
[0016] In an embodiment, the method further includes: determining which of the first phase difference peak width and the second phase difference peak width is narrower (after adjustment), wherein the shift parameter includes the translation median and the one having the narrower peak among the first phase difference median or the second phase difference median, and the squeeze parameter includes the translation peak width and the narrower one of the first phase difference peak width or the second phase difference peak width. It should be understood that "(after adjustment) narrower" indicates that the second phase difference value is used only when the second phase difference value is significantly narrower than the first phase difference value; this helps to ensure Stability of the value. In an embodiment, the value is twice as narrow. The term "narrower (after adjustment)" also means that for the same amount of captured audio energy, more energy is concentrated around the peak.
[0017] In an embodiment, the spatial parameters include a translation parameter and a phase difference parameter for each time-frequency bin, and calculating the shift parameter and the squeeze parameter further includes: for each subband in each chunk: creating a smoothed level parameter weighted histogram on the translation parameter; creating a smoothed level parameter weighted first phase difference histogram on a first phase difference parameter, where the first phase difference parameter has a first range; creating a smoothed level parameter weighted second phase difference histogram on a second phase difference parameter, where the second phase difference parameter has a second range different from the first range; detecting a translation peak in the smoothed translation histogram; determining the translation peak width; determining the translation median; detecting a first phase difference peak in the smoothed first phase difference histogram; determining the first phase difference peak width; determining the first phase difference median; detecting a second phase difference peak in the smoothed second phase difference histogram; determining the second phase difference peak width; and determining the second phase difference median, where the shift parameter includes the translation median and the first phase difference median or the second phase difference median, and the squeeze parameter includes the translation peak width and the first phase difference peak width or the second phase difference peak width.
[0018] In an embodiment, the method further includes: determining which of the first phase difference peak width and the second phase difference peak width is narrower (after adjustment), where the shift parameter includes the translation median and the one of the first phase difference median or the second phase difference median that has the narrower peak, and the squeeze parameter includes the translation peak width and the narrower of the first phase difference peak width or the second phase difference peak width.
[0019] In an embodiment, the first phase difference range is from -π to π radians, and the second phase difference range is from 0 to 2π radians.
[0020] In an embodiment, the translation histogram and the first and second phase histograms are smoothed in time using the translation histograms and the phase difference histograms created for previous and subsequent chunks; or weighted data in previous and subsequent chunks is collected and then directly used to form the histograms.
[0021] In an embodiment, the translation peak width captures at least forty percent of the total energy in the translation histogram, and the first phase difference peak width and the second phase difference peak width each capture at least eighty percent of the total energy in their respective histograms.
[0022] In an embodiment, the shift parameters and squeeze parameters for each subband in each chunk are converted to be present for each of the one or more frames.
[0023] In an embodiment, the translation shift parameter and the squeeze parameter are converted to exist for each frame using linear interpolation, and the first phase difference shift parameter or the second phase difference shift parameter is converted to exist for each frame using zero-order hold.
[0024] In an embodiment, the method further comprises determining a single shifted median value and a single shifted peak width value per unit time for the one or more subbands in the one or more chunks.
[0025] In an embodiment, the soft mask values are smoothed in time and frequency.
[0026] In an embodiment, an apparatus includes one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the aforementioned methods.
[0027] In an embodiment, a non-transitory computer-readable storage medium has instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform any of the aforementioned methods.
[0028] Certain embodiments disclosed herein provide one or more of the following advantages. Efficiently and robustly extract spatially identifiable sub-band audio sources from a two-channel mix. The system is robust because it can extract any spatially identifiable sub-band audio source, including amplitude-shifted and non-amplitude-shifted audio sources, such as audio sources mixed or recorded with a delay between channels, audio sources mixed or recorded with reverberation, and audio sources with spatial characteristics that vary across frequency subbands. The system is also efficient and requires little to no training data or latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In the accompanying drawings referred to below, various embodiments are illustrated in the form of block diagrams, flowcharts, and other diagrams. Each block in the flowchart or block diagram may represent a module, program, or code portion that includes one or more executable instructions for performing the specified logical function. Although these blocks are illustrated in a particular order for performing these method steps, they may not necessarily be executed strictly in the order illustrated. For example, depending on the nature of the corresponding operations, these blocks may be executed in reverse order or simultaneously. It should also be noted that each block in the block diagram and / or flowchart, and combinations thereof, can be implemented by a dedicated system based on software or hardware for performing the specified function / operation, or by a combination of dedicated hardware and computer instructions.
[0030] Figure 1 is a block diagram of a system for detecting and extracting spatially identifiable subband audio sources from a stereo mix according to an embodiment.
[0031] Figures 2A - 2B is a visual depiction of the input and output of a spatial level filter (SLF) trained to extract translational sources according to an embodiment.
[0032] Figure 3 is a flowchart of a process for detecting and extracting spatially identifiable subband audio sources from a stereo mix according to an embodiment.
[0033] Figure 4 illustrates an apparatus architecture for implementing the system and process described in reference Figures 1 to 3 described, according to an embodiment.
[0034] Like reference numerals used in the various drawings indicate like elements. DETAILED DESCRIPTION
[0035] The disclosed embodiments allow for the detection and extraction (audio source separation) of spatially identifiable subband audio sources from a stereo audio mix. As used herein, a "spatially identifiable" subband audio source is a subband audio source whose energy is spatially concentrated within an octave subband or an approximate octave subband.
[0036] The disclosed embodiments are mainly used in the context of a sound source separation system that takes two-channel (stereo) signals as input and operates in a frequency domain such as the short-time Fourier transform (STFT) domain. In a typical sound source separation system, four basic steps are used.
[0037] First, an application front end transforms the two-channel time-domain audio signal into the frequency domain. In an embodiment, the STFT is typically used, which produces a spectrogram (e.g., magnitude and phase) of the input signal in the frequency domain. The elements of the STFT output can be referenced by indicating their indices in time and frequency; each such element can be referred to as a time-frequency bin. Each time point corresponds to a frame number, which includes a plurality of frequency bins that can be subdivided or grouped into subbands. The STFT parameters (e.g., window type, hop size) are selected by those of ordinary skill in the art to be relatively optimized for the source separation problem. According to the STFT representation, the described system calculates the spatial parameter θ(Θ) and and the level parameter U (both defined below) and notes the associated quasi-octave subband b.
[0038] Second, detect the presence of audio sources and the parameters that describe the spatial identity of these audio sources.
[0039] Third, use the spatial parameter θ(Θ) and and the level parameter U to perform the extraction of the estimated (multiple) audio sources by applying an amplitude soft mask (e.g., values in the continuous range [0,1]) to each bin of the STFT representation of each channel (e.g., each bin of each time-frequency slice of the left and right channels).
[0040] Fourth, convert the STFT-domain estimate of the (multiple) audio sources into a two-channel time-domain estimate by performing an inverse short-time Fourier transform (ISTFT) on the STFT representation of each channel. It should be noted that although this step is described as "fourth" in sequence in this context, there may be other optional processing that occurs in the STFT domain before this fourth step. In an embodiment, the ISTFT is performed after other STFT-domain processing is completed.
[0041] The parameters of each bin in the STFT representation include two spatial parameters θ(Θ) and and the parameter U, which are defined and calculated as follows.
[0042] θ(Θ) is the detected translation of each time-frequency slice (ω, t), defined as:
[0043]
[0044] Where "full left" is 0 radians, "full right" is π / 2 radians, and "dead center" is π / 4 radians. It should be noted that the "detected translation" can also be considered as the interchannel difference represented as a continuous value from 0 to π / 2.
[0045] is the detected phase difference for each time-frequency bin and is defined as:
[0046]
[0047] where, ranges from -π to π radians, where 0 means the phases detected in the two channels are the same. For some content, may be concentrated around + / -π, i.e., at the opposite ends of the range defined here as Therefore, define whose data is the same as but rotated on the unit circle so that the range is from 0 to 2π. Mathematically, this just means that any value below 0 is set to its previous value plus 2π. It should be noted that is useful in a specific part of the system.
[0048] U is the detected level for each time-frequency bin and is defined as:
[0049] U(ω,t) = 10*log 10 (|X R (ω, t)| 2 +|X L (ω, t)| 2 , [3]
[0050] This is the decibel (dB) version of the "Pythagorean" magnitude of the two channels. It can be considered as a mono magnitude spectrogram. The version of U in equation [3] is on a dB scale and can also be referred to as U dB. Various scales of U can also be used at various points in the system. For example, U-power is U-power(ω, t) = (|XR(ω, t)|2 + |XL(ω, t)|2). Additional versions of U can be generated by raising U to various exponents (powers). This is particularly relevant to all mentions of "level weighted histograms" in this article. It should be understood that such mentions imply that various powers can be used when applying level weighting; powers between 1 and 2 are recommended, and U-power (power of 2) is recommended in the specific steps mentioned earlier.
[0051] Each frequency bin ω is understood to represent a specific frequency. However, the data can also be grouped within subbands, which are collections of contiguous bins, where each frequency bin ω belongs to a subband. Grouping the data within subbands is particularly useful for certain estimation tasks performed in the system. In an embodiment, octave subbands or approximate octave subbands are used, but other subband definitions can also be used. Some examples of banding include defining the band edges as follows, where the values are listed in Hz:
[0052] [0, 400, 800, 1600, 3200, 6400, 13200, 24000],
[0053] [0, 375, 750, 1500, 3000, 6000, 12000, 24000], and
[0054] [0, 375, 750, 1500, 2625, 4125, 6375, 10125, 15375, 24000].
[0055] It should be noted that if the "octave" definition is strictly followed, there could be an infinite number of such bands, where the lowest band approaches an infinitesimal width, so some selection is needed to allow a finite number of subbands. In an embodiment, the lowest band is selected to be equal in size to the second band, but other conventions can be used in other embodiments.
[0056] In an embodiment, the system processes groups of consecutive frames, also referred to hereinafter as "chunks". This allows the use of data from multiple frames to achieve more stable spatial property estimation. By using chunks, rather than just a longer frame length, the advantages of a specific frame length (e.g., between 50 - 100 ms) (e.g., quasistationarity, optimality of source separation) are retained. The chunks can be overlapped by selecting a chunk hop size that is less than the number of frames in the chunk. In an embodiment, the system uses chunks of 10 frames and a chunk hop size of 5 frames. Since these frames themselves will be hopping with a frame hop size of 1024 samples (assuming a sampling rate of 48 kHz) and are 4096 samples in length, these chunks will require approximately 277 milliseconds of data. Depending on the computational, latency, and data stability implementation requirements, smaller or larger chunks or hop sizes can be used, where the amount of lookahead and lookback used is also determined by the implementation requirements. In an embodiment, the chunks have 5 lookahead frames and 5 lookback frames.
[0057] In an embodiment, the robust, efficient sound source separation system described herein uses a spatial level filtering (SLF) system. A spatial level filter (SLF) is a system that has been trained to achieve the following purpose: extract a target source with a given level distribution and specified spatial parameters from a mixture including a background with a given level distribution and spatial parameters. For illustrative and practical purposes, the following description of SLF will assume that the target spatial parameters consist only of a translation parameter Θ1, and further assume that Θ1 corresponds to a centered translated source. The techniques described herein can also be used in combination with an SLF trained to extract a target source whose spatial parameters are not so constrained; such techniques are described below in the context of shift parameters and squeeze parameters.
[0058] The translation parameter Θ1 exists in the context of the following signal model, in which the target source s1 and the background b are mixed into two channels depending on the context, hereinafter referred to as the "left channel" (x1 or XL) and the "right channel" (x2 or XR).
[0059] It is assumed that the target source s1 is amplitude translated using a constant power law. Since other translation laws can be converted to the constant power law, the use of the constant power law in the signal model 100 is non-restrictive. Under the constant power law translation, the source s1 mixed into the left / right (L / R) channels is described as follows:
[0060] x1 = cos(Θ1)s1, [1]
[0061] x2 = sin(Θ1)s1, [2]
[0062] where Θ1 ranges from 0 (translated to the leftmost source) to π / 2 (translated to the rightmost source). We can represent it in the short-time Fourier transform (STFT) domain as
[0063] X L = cos(Θ i ),S i , [3]
[0064] X R = sin(Θ1)S1. [4]
[0065] Then recall that the "target source" is assumed to be translated, which means that the source can be characterized by Θ1. By inspection, it should be clear that if the signal contains only the target source at a given point in the time-frequency space, the detected translation parameter θ(Θ) above will produce a perfect estimate of the target source translation parameter Θ1.
[0066] Returning to the concept of how to use SLF, recall Θ(ω, t) above, And the above definitions of U(ω, t), which can also be denoted as and are understood to exist for each time-frequency bin (ω, t). θ(Θ) and are the detected "spatial parameters", and U is the detected "level parameter". Further, it should be noted that the frequency value ω of the time-frequency bin under discussion is a member of the approximate octave sub-band b for which the SLF is trained. In one embodiment, for each time-frequency bin (ω, t) in the time-frequency representation, the SLF takes an input of four values and outputs a single STFT soft mask value. Thus, the STFT soft mask value is determined by any trained SLF which, for each time-frequency bin, takes four inputs and produces one output. The soft mask value is multiplied by the input mixed representation value to produce the estimated target source value.
[0067] It should be noted that the SLF which takes four input values and produces one output value can exist in the form of a function (four inputs, one output) or a table (four-dimensional, where the values stored in the table represent the output value). In an embodiment, the SLF used takes the form of a table. Table lookup 106 is a technique for accessing the values in the table using any method familiar to those skilled in the art.
[0068] A visual description of the input and output of a typical trained SLF lookup table is as Figures 2A - 2B shown. Figures 2A - 2B The illustrated non-limiting exemplary SLF system is one example SLF system that can be used in the disclosed embodiments. Other SLF systems can also be used, which: 1) are trained to extract a centered translated source; 2) have at least four inputs, which include: Θ as defined above, U, and the sub-band b; 3) have at least one output, which is a floating-point value from 0 to 1 (inclusive of the end values); 4) perform input / output operations for each STFT bin; 5) have an output of STFT size, which is composed of floating-point values (referred to as soft masks) for each STFT bin; and 6) have an input STFT representation which is multiplied by the soft mask value to obtain the estimated source output STFT representation, which is then transformed into the estimated source signal in the two-channel time domain.
[0069] The spatial Θ and parameters detected for the training data will have a distribution in each sub-band. When there is a centered translated source, these values give some notion of the "spread" or "width" of such data. In an embodiment, during training, a histogram analysis of the data in each sub-band is performed, which tracks the width to capture 40% of the energy relative to Θ or relative to Capture 80% of the data. These widths are respectively recorded as the "reference θ width" and "reference width" for each sub-band. For Figures 2A - 2B the exemplary SLF system depicted in , the reference Θ widths (on 7 sub-bands) are [0.1 0.07 0.04 0.10 0.12 0.2 0.12], and the reference
[0070] widths are [0.6 0.5 0.4 0.6 0.8 1.0 1.0].
[0071] The exemplary audio source separation system described herein is designed based on an investigation of typical audio source mixture examples including conversations. The system utilizes the information discovered during the investigation. The next section will briefly summarize the results of the investigation, related assumptions, and related system objectives.
[0072] · The sub - band spatial concentration is related to the intelligible conversation source. When the U-power weighted 2-D histogram is plotted on the Θ and sub-band distribution of the data, if there is a concentrated peak (e.g., most of the energy is concentrated in less than 10% of the space), then the bandpass signal will also be intelligible - or as intelligible as an octave bandpass speech signal. Therefore, the system will attempt to identify, parameterize, and capture such energy.
[0073] · The octave sub - band accuracy can be good enough for the identification and extraction of "delayed sources". Inter-channel delay estimation is more than calculating in the STFT domain More challenging problems, especially when there is a large amount of interference. However, for many or most typical mixed or delayed recordings, there is still sufficient concentration within the octave subbands relative to such that the source can be identified and extracted based on . This is an important observation because it allows for source separation without explicitly estimating the delay. The values of Θ and around which the energy is concentrated will vary with the subband. Given these observations, the system will estimate the concentration in each subband per unit time.
[0074] · For some examples, extracting one source per sub - frequency band is effective and efficient. In source separation, the task is to extract one or more sources per unit time depending on the goal or context. When the goal is to efficiently extract spatially distinguishable sources (e.g., dialogue) from typical entertainment content, experiments have shown that extracting one source per approximately octave subband may be sufficient in terms of the resulting output audio quality. This is because it is rare for two sources to simultaneously dominate the same subband. This is a version of "W-disjoint orthogonality" that makes a similar observation for each STFT (higher frequency resolution) bin. It should be emphasized that audio source separation still occurs in the individual STFT bins; only source identification and spatial parameter estimation are performed, for which octave subband processing has been found to be sufficient. Based on the observations, the system will attempt to parameterize only one source per subband per unit time.
[0075] · For voice sources, avoid certain frequencies when identifying spatial parameters or performing extraction. Some speech energy exists at very low frequencies, depending on the fundamental frequency of the speaker. In the best-case scenario, this energy can be used to identify spatial parameters and perform extraction. In practice, due to the presence of special effects and other backgrounds, this scenario rarely exists in typical entertainment content. For this reason, when detecting dialogue, data below approximately 175 Hz is excluded, and when extracting dialogue, no attempt is made to extract data below approximately 117 Hz. For similar reasons and due to computational costs, frequencies above approximately 13200 Hz are not considered for detection or extraction.
[0076] · If the assumptions are violated, further attention is required. The above observations lead to the design of the source separation system described below, which identifies and extracts sources based on detectable subband spatial concentration. It is assumed that the target source is at least as spatially distinguishable in the subband as any interfering source. This generally also requires that the target source be at least at the same level as the interfering source in the subband.
[0077] Figure 1is a block diagram of an exemplary system 100 for detecting and extracting spatially identifiable sub-band audio sources from a stereo mix according to an embodiment. System 100 includes a transform module 101, a parameter extraction module 102, a detection module 103, a parameter modification module 104, a table lookup module 105, a lookup table 106, a soft mask application module 107, and an inverse transform module 108. Each of these modules may be implemented in hardware or software, or a combination of hardware and software. In an embodiment, system 100 may be implemented using a device architecture as shown in reference Figure 4 as shown. Each module will now be described in turn with reference to Figure 1 .
[0078] Referring to Figure 1 on the left side, transform module 101 transforms a stereo time-domain mixed audio signal (e.g., a stereo signal) into a frequency-domain representation, such as an STFT-domain representation (e.g., a spectrogram / time-frequency slice), using windows and parameters familiar to those skilled in the art. In an embodiment, the window is a 4096-point square root of a Hann window with a 1024-frame hop, and the STFT is a 4096-point FFT for a 48 kHz sampled input. Other windows, such as a Gaussian window, may also be used. Within limits, scales that maintain the hop size and frame length in milliseconds may be used to obtain lower or higher sampling rates.
[0079] Extraction module 102 calculates the above parameters for each time-frequency slice (bin and frame) in the STFT representation i.e., if an example has 1000 frames and uses 2049 unique STFT bins (assuming a 4096-point STFT), then each parameter will have 2,049,000 values.
[0080] In an embodiment, the U parameter is adjusted based on the measured input data level. For each frame, a buffer of data is assembled for the current frame and some reasonable number of previous frames. This is intended to be a long-term measurement. For practical purposes, the buffer length is typically several seconds (e.g., 5 seconds). For the data in the buffer, the level of the frame is calculated using the loudness relative to full scale (LKFS) method. Other methods may also be used. However, whatever method is used, it should match the method used to calculate the training data level. It should be noted that it is assumed that similar but longer-term measurements have been performed on the training data previously to produce the measured training data level.
[0081] In an embodiment, the level parameter U is then adjusted as follows: Udb = Udb - (measured training data level - measured input data level + additional level shift), where the measured training data level is the total level value in dB (such as the LKFS of the training data as described above). The measured input data level is the input data level value in dB (such as in LKFS), which, as described above, is measured in real time for each frame.
[0082] The additional level shift is an optional user-selectable value. This value is used in subsequent parts of system 100 described below but is addressed here. By selecting a positive value, the user can specify that the input data is at a higher level than it actually is, which drives the system to use more selective values of the SLF system. The system operator can select this parameter via an interface, examples of which include parameter selection in an API call or editing the text of a configuration file.
[0083] Figures 2A - 2B is a sampled representation of the input and output of the SLF system, providing an example of a relevant SLF system, but any SLF system can be utilized. Figures 2A - 2B The schematic in is a four-dimensional graph. The four input variables are represented by the left and right axes and the in and out axes of each subplot, as well as the vertical subplot index and the horizontal subplot index. These variables correspond respectively to the input variables: (1) modified θ, (2) modified (3) sub-band b, (4) level U. It should be noted that for practical reasons, the horizontal subplot dimension (level U) does not depict all levels stored in the SLF lookup table; doing so would require approximately 128 subplots because 1 dB increments are used within a 128 dB range in the table. In practice, finer or coarser increments can be used for higher precision or higher lookup efficiency, respectively. The output variable is represented by the vertical value of each subplot; this corresponds to a soft mask value between 0 and 1.
[0084] When viewing Figures 2A - 2B , it should be noted that there are many "not shown" subplots from left to right. Using a positive value for the additional level shift corresponds to moving from a given subplot corresponding to the input level to a more right subplot (or corresponding not shown data) corresponding to a higher input level. A negative value corresponds to moving to a more left subplot (or corresponding data). It can generally be observed that moving to a more right subplot (or moving to the data in the table corresponding to such a subplot, whether or not it is included in Figures 2A - 2B results in a more selective (less "flat") filtering. This is associated with less background capture but more artifacts in the source estimate. Conversely, using a lower value has the opposite effect, such as more background capture but fewer artifacts.
[0085] The detection module 103 detects a spatially identifiable audio source for each sub-band. The recommended method for doing this involves histograms and is described in detail below. However, any method that meets the following conditions (such as distribution estimation from a Parzen window) meets the design requirements of the system: (1) estimating the peak of the relevant distribution on θ and above, (2) estimating the range of the said distribution with respect to θ and capturing a large amount of energy, such as a predetermined amount of audio energy (recommended to be 40% for θ and 80% for (ranging from -π to π) and (ranging from 0 to 2π). It should be noted that for a conversational audio source with little energy above 13 kHz, the cost of detecting the highest octave may not be worth it. Therefore, this procedure may only be applicable to sub-bands with a minimum frequency equal to or below 13 kHz. The detection module 103 assembles consecutive frame data into chunks (e.g., 10-frame chunks). For each sub-band in each chunk (if in the first sub-band, data below 175 Hz is excluded as recommended above), the detection module 103 creates a U-power weighted histogram on Θ, which is smooth on Θ. Additionally, the same process is applied to (ranging from -π to π) and The U-power weighted histogram can use any number of bins (e.g., 51 bins with respect to Θ, bins with respect to Figures 2A - 2B ). Since the lower sub-bands have fewer data points, they will require more smoothing. In another embodiment, fewer histogram bins can be used for the lower sub-bands, while more histogram bins can be used for the higher sub-bands. Smoothing can be performed using techniques familiar to those skilled in the art. However, in the preferred embodiment, it is recommended to use smoothing kernels on each of Θ and
[0086] These smoothing kernels correspond to the following fractional values of the data range of Θ or and For the histogram, the situation is similar. The recommended weights are as follows: the current chunk is 1.0, the previous chunk is 0.4, the chunk before the previous chunk is 0.2, and the next chunk is 0.1. Depending on the application, the smoothing method can (1) share the weighted data across time and then create a histogram from the smoothed data, or (2) first create a histogram and then share the weighted histogram across time to smooth the histogram. When memory and computational power are limited, method (2) can be used.
[0087] Referring again to Figure 1 , the detection module 103 picks up and detects the peak width as follows. For the Θ histogram, the Θ value of the detected peak, called "thetaMiddle", is detected, and also the width around this peak required to capture 40% of the energy in the capture histogram, called "thetaWidth". For and , the same process is applied, and middle (phiMiddle), middle (phi2Middle), width (phiWidth), and width (phi2Width) are recorded, but when recording the width, 80% rather than 40% of the energy is required to be captured. Recall that the range of Θ is from 0 (leftmost) to π / 2 (rightmost), so the maximum thetaWidth value is always less than π / 2. Recall that ranges from -π to π radians (representing all phase values on the unit circle), so the maximum width value will always be less than 2π. It should also be noted that the 80% and 40% energy captures are recommended values, and other percentages can also be selected.
[0088] Now, given the widths of and , based on the smaller width value, it indicates which parameter has a higher concentration in the space to record the middle and width final values. However, width is only selected if it is at least less than half of the width. This allows reducing the rapid alternation between and when there is a very wide distribution of quasi-random data with respect to
[0089] Now, for each subband and chunk, thetaMiddle, thetaWidth, middle, and The width parameters are all known. (Recall that subbands and bins are different: there are only about 7 subbands, but there are likely 2049 unique bins. Frames and chunks are also different; there are multiple frames in each chunk.). By using first-order linear interpolation, the θ mid, θ width, and the width parameters are converted to per-frame presence, but other techniques familiar to those skilled in the art can also be used. By using zero-order hold, the mid parameter is converted to per-frame presence to avoid rapid phase changes in cases where some chunks are close to or equal to +π and some chunks are close to or equal to -π. The parameters θ mid and θ width are also referred to hereinafter as "θ shift and squeeze" parameters, and the parameters mid and width are also referred to hereinafter as " shift and squeeze" parameters. These four parameters are collectively referred to hereinafter as "shift and squeeze" or "S&S" parameters.
[0090] The S&S parameters can conceptually be understood as representing the difference between the detected Θ and the concentration of the data, as well as the concentration for an ideal center-shifted source with limited or no background. This concept will later allow the system to use the S&S parameters to modify the detected data in such a way that the SLF designed for a center-shifted source can be used to extract target sources with arbitrary concentrations in Θ and . In most cases, such applications should be understood to be optimal and recommended. However, the SLF used does not need to be trained only for center-shifted sources, the S&S parameters do not need to be calculated only relative to center-shifted sources, and the system does not need to limit itself to using only a single trained SLF model to perform target source extraction. By calculating the S&S parameters relative to the trained SLF target source parameters, any SLF model can be used, including a greater number of models. For efficiency considerations, the system uses a single center-shifted source SLF.
[0091] The above steps are performed for each Θ and Generate values corresponding to "mid" and "width". In some embodiments, it may also be desirable to have a single overall "mid" value for Θ per unit time that takes into account the data in all subbands. To achieve this, the weighted sum of the Θ histograms of most subbands is calculated as follows for a given chunk before peak picking. Due to the special effect of spatial ambiguity at low frequencies - which may particularly challenge the detection of the speech source, subband 1 is optionally completely ignored. The weight of subband 2 is reduced by scaling its histogram by a certain factor (e.g., 0.1). The other subband histograms are weighted equally (e.g., by scaling each by 1.0). It should be noted that while the higher octave subbands tend to have lower energy per bin, these subbands have more bins, thus offsetting this effect and ensuring that all subbands have a perceptually relevant chance to influence the single Θ estimate. Once the combined Θ histogram for a given chunk has been created as described above, the histogram is smoothed as described above relative to other time chunks to obtain θ mid, etc. Next, simple peak picking is performed. The picked peak is the single Θ value for each chunk. In an embodiment, linear interpolation is applied between chunks to obtain these values for each frame. The single Θ value for each frame obtained in this way is also referred to as "single theta (singleTheta)" hereinafter.
[0092] Referring again to Figure 1 , the parameter modification module 104 uses shift and squash (S&S) parameters to modify the parameters input to the SLF system values. The steps for this part are as follows. Processing is done frame-by-frame and subband-by-subband. That is, the following steps assume processing within one frame and one subband. As previously mentioned, any subbands whose frequencies are mostly or completely outside the considered range (e.g., frequencies above 13 kHz) can optionally be skipped; of course, if the S&S parameter detection skips a corresponding subband, those subbands should be skipped as they will have no data to act on. Unless otherwise stated, the data described in the variables herein is specific to the considered frame and subband. For example, "θ mid" is understood to have a value for each frame and subband, so a reference to θ mid implies consideration of the current frame and subband.
[0093] As suggested above, when considering the SLF system output values for the first subband, frequencies below approximately 117 Hz (no input given) can be ignored, or equivalently, the corresponding soft mask values can be set to zero after they are calculated. It should be noted the key difference between bins and subbands. For each bin in a single subband, the "raw data" of Θ, as well as U, are single. For example, subband 4 can contain 136 bins. All 136 bins for a particular frame have a single value, but corresponding to θ mid, θ width, for that frame, in the middle and a single value for "sub-band 4" of the width.
[0094] In an embodiment, the Θ value is modified according to its S&S parameters as follows.
[0095] Calculate: squeezeFactor = θ width / (reference θ width value corresponding to the trained SLF to be applied). If the squeeze factor is outside the range [1.0, 1.5], bring it back into this range. It should be noted that values higher than 1.5 can be used to allow for more fully capturing more diffuse sources. A squeeze factor of 1.5 provides a good balance for extracting spatially distinguishable sources. To make the system more selective, they can be narrowed by multiplying the reference θ width (and reference width) value by 0.5 or other suitable factor.
[0096] Calculate: shiftFactor = θ middle - π / 4 (for this frame and sub-band). It should be noted that π / 4 is used here because it represents a centered shifted source. The trained SLF system to be used should be for a centered shifted source.
[0097] Calculate: distsFromMiddle = θ middle - (original θ data for each bin in this frame and this sub-band).
[0098] Calculate: newDistsFromMiddle = distsFromMiddle / squeezeFactor.
[0099] Calculate: thetaModified = θ middle + newDistsFromMiddle - shiftFactor;
[0100] If thetaModified exceeds the range [0, 2*π], limit it to this range.
[0101] Use a similar method to modify the value according to the S&S parameters. It should be noted that there are some key differences from the θ case.
[0102] Calculate:
[0103] This may cause some data to go outside the range [-π, π], so, use a cyclic processing of the phase to bring all values back into this range. That is, add 2*π to any value below -π and subtract 2*π from any value above π.
[0104] Calculate:
[0105] At this point, the squeezing factor value should be limited as was θ above. However, additional realities are considered here. By definition, sources with "extreme" Θ values near 0 (left most) or π / 2 (right most) are expected to always have a broad distribution on . Thus, when extreme values are taken in the middle of θ, it is not optimal to impose strict limits on the "squeezing" of the dimension. To ensure reasonable limits are imposed, the following procedure is carried out. First, calculate the "theoretical maximum squeezing" (tpms) based on the corresponding reference width value as follows: tpms = 2 * π / (the reference width of the subband). This value is only relevant for values reasonably close to those outside the center Θ values, that is, those roughly outside the range 0.231 to 1.3398. Recall that the entire range of Θ is 0 to π / 2. For values in the middle range from 0.231 to 1.3398, use the conventional maximum squeezing factor, which is 1.5 for the conventional maximum squeezing factor. For values very close to 0 or π / 2 (values within 5% of these), use the theoretical maximum. For those values in the remaining range between the values of interest, perform a simple linear interpolation based on how far the middle value of θ is within the range to obtain the maximum squeezing factor.
[0106] Next, limit the squeezing factor calculated previously to the value calculated in the previous step.
[0107] Finally, calculate: Modified At this point, there should be no values outside the range -π to π.
[0108] At this point, the modified θ, the modified and U have been modified. It should be noted that U has been scaled previously to address the level difference between the detected input signal level and the training data level, as well as any additional level shift specified by the user.
[0109] Referring again to Figure 1 , the table lookup module 105 obtains the soft mask value from the SLF lookup table 106 and the soft mask application module 107 applies the soft mask value to the STFT time-frequency bins. The input values (modified θ, modified and (U) for obtaining soft mask values for each frame and bin from the lookup table 106. Although the lookup table 106 is provided as an example embodiment, the SLF itself can be implemented in different ways, including but not limited to lookup tables, functions, nested tables and / or functions, (multiple) neural networks, etc. with four input values and one output value. Since the SLF to be used corresponds to the center translation source, any of these methods can take advantage of the fact that for a typical general background in the training data, the center SLF should be symmetric about π / 4. This can be achieved by averaging the data on both sides of θ == π / 4 during smoothing, effectively halving the training data with respect to θ. The required effective memory can also be reduced by treating any modified θ values above π / 4 as the same as the modified θ values below that value in the opposite way. This also increases the consistency of the system output.
[0110] As previously mentioned, in a non-limiting example, Figures 2A - 2B a sampled representation of n SLFs is shown. The output is shown on the vertical axis of each subplot. The four input variables are the left-right (Θ) axes and the input-output axes of each subplot, as well as the vertical (subband b) subplot index and the horizontal (level U) subplot index. The output variable is between 0 and 1 (inclusive of the end values) and represents the fraction of the corresponding input STFT that should be passed to the output. Since there is one (four-dimensional) input per STFT bin, there is also one output per STFT bin. The result of applying the SLF function is a representation of the STFT magnitude, which is composed of values between 0 and 1 (also called the soft mask). This soft mask representation is called "source mask 1".
[0111] The U value will be needed in subsequent steps. Therefore, the U value is returned to the previously described unscaled original value (not required for SLF input).
[0112] In an embodiment, the soft mask values and / or signal values are smoothed in time and frequency using techniques familiar to those skilled in the art. Assuming a 4096-point FFT, smoothing with respect to frequency can be used, which uses a smoother [0.17 0.33 1.0 0.33 0.17] / sum([0.17 0.33 1.0 0.33 0.17]). For higher or lower FFT sizes, some reasonable scaling should be performed on the smoothing range and coefficients. Assuming a hop size of 1024 samples, a smoother of approximately [0.1 0.55 1.0 0.55 0.1] / sum([0.1 0.55 1.0 0.55 0.1]) can be used with respect to time. If the hop size or frame length changes, the smoothing should be adjusted appropriately.
[0113] Referring again to Figure 1, the inverse transform module 108 performs an inverse STFT on the STFT representation of the estimated audio source. In an embodiment, the inverse STFT is performed using the same synthesis window (post window) as the analysis window, such as the square root of a Hann window. Since there are two STFT representations, there are now two time-domain signals.
[0114] The output of the inverse transform module 108 is a two-channel time-domain audio signal that combines the audio sources extracted from six (or seven) of the seven subbands. In some examples, this is all that is needed, and this single time-domain signal can then be processed or utilized. In other examples, it may be desirable to have each subband signal separately. This is especially meaningful when the subband signals may have very different θ and / or values. For example, if subbands 1-4 have the leftmost θ sources and subbands 5 and 6 have sources to the right of center, the system can be configured to produce a bandpass output by processing in the STFT domain before the inverse transform module 108 or by bandpass filtering the estimated extracted audio source signals.
[0115] Figures 2A - 2B is a visual depiction of the input and output of an SLF system trained to extract translation sources according to an embodiment. More specifically, Figures 2A - 2B is Figure 1 an example of a trained SLF lookup table described in
[0116] Figure 3 is a flowchart of a process 300 for detecting and extracting spatially distinguishable subband audio sources from a two-channel mix according to an embodiment. Process 300 can be implemented using, for example, the device architecture 400 described in Figure 4 .
[0117] Process 300 can begin by transforming a two-channel time-domain audio signal (e.g., a stereo signal) into a frequency-domain representation (301) that includes time-frequency tiles having a plurality of frequency bins. For example, a stereo audio signal can be transformed into an STFT representation of time-frequency tiles, as described in reference Figure 1 .
[0118] Process 300 continues by calculating spatial parameters and level parameters (302) for each time-frequency tile. For example, process 300 calculates Θ, and the U parameter for each time-frequency tile, as described in reference Figure 1 .
[0119] Process 300 continues by calculating shift parameters and squeeze parameters (303) using the spatial parameters and level parameters (Θ, and U), and modifying the spatial parameters using the shift parameters and squeeze parameters (304). For example, it can be as described in referenceFigure 1 Calculate the shift parameter and the squeeze parameter as described.
[0120] Process 300 continues with the modified spatial parameters Obtain a soft mask value (305). For example, the modified spatial parameters can be used Select a soft mask value from a trained SLF lookup table, such as Figures 2A - 2B The exemplary SLF lookup table shown.
[0121] Process 300 continues to apply the soft mask value to the time-frequency bins to generate the time-frequency bins of the estimated audio source (306). For example, the soft mask value is a continuous value (fraction) between 0 and 1, and these continuous values are multiplied by their corresponding magnitudes in the dimension in the bins of the STFT bins. Since the soft mask value is a fraction, applying the soft mask value to the STFT bins will effectively reduce the magnitudes in all frequency bins that do not contain audio source data.
[0122] Process 300 continues to inverse-transform the time-frequency bins of the estimated audio source into a two-channel time-domain estimate of the audio source (307).
[0123] Figure 4 is according to an embodiment Figure 1 A block diagram of the device architecture 400 of the system 100 shown. The device architecture 400 can be used in any computer or electronic device capable of performing the above mathematical calculations. The features and processes described herein can be implemented in one or more of an encoder, a decoder, or an intermediate device. These features and processes can be implemented in hardware or software, or a combination of hardware and software.
[0124] In the example shown, the device architecture 400 includes one or more processors (401) (e.g., CPU, DSP chip, ASIC), one or more input devices (402) (e.g., keyboard, mouse, touch surface), one or more output devices (e.g., LED / LCD display), a memory 404 (e.g., RAM, ROM, flash memory), and an audio subsystem 406 (e.g., media player, audio amplifier, and support circuitry) coupled to a speaker 406. Each of these components is coupled to one or more buses 407 (e.g., system, power, peripheral, etc.). In an embodiment, the features and processes described herein can be implemented as software instructions stored in the memory 404 or any other computer-readable medium and executed by one or more processors 401. Other architectures with more or fewer components are also possible, such as architectures that use a hybrid of software and hardware to implement the functions and processes described herein.
[0125] Although this document contains many details of specific embodiments, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially so claimed, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may relate to a sub-combination or a variant of a sub-combination. The logical flow depicted in the figures does not require the particular order shown or an ordered sequence to achieve the desired result. Additionally, other steps may be provided or steps may be removed from the described flow, and other components may be added to or removed from the described system. Accordingly, other embodiments are within the scope of the following claims.
[0126] Aspects of the present invention can be understood from the following enumerated example embodiments (EEEs):
[0127] EEE1. A method, comprising:
[0128] Using one or more processors to transform one or more frames of a stereo time-domain audio signal into a time-frequency domain representation including a plurality of time-frequency bins, wherein a frequency domain of the time-frequency domain representation includes a plurality of frequency bins, and the plurality of frequency bins are grouped into a plurality of subbands;
[0129] For each time-frequency bin:
[0130] Using the one or more processors to calculate a spatial parameter and a level of the time-frequency bin;
[0131] Using the one or more processors to modify the spatial parameter using a shift parameter and a squash parameter;
[0132] Using the one or more processors to obtain a soft mask value for each frequency bin using the modified spatial parameter, the level, and subband information; and
[0133] Using the one or more processors to apply the soft mask value to the time-frequency bin to generate a modified time-frequency bin of an estimated audio source.
[0134] EEE2. The method according to EEE 1, wherein a plurality of frames of the time-frequency bins are assembled into a plurality of chunks, each chunk including a plurality of subbands, and the method comprises:
[0135] For each subband in each chunk:
[0136] Using the one or more processors, compute spatial parameters and levels for each time-frequency bin in the group of bins; using the one or more processors, modify the spatial parameters using shift parameters and squash parameters;
[0137] Using the one or more processors, obtain soft mask values for each frequency bin using the modified spatial parameters, the levels, and subband information; and
[0138] Using the one or more processors, apply the soft mask values to the time-frequency bins to generate modified time-frequency bins of the estimated audio source.
[0139] EEE3. The method according to EEE 2, wherein the spatial parameters include translation parameters and phase difference parameters for each time-frequency bin, and computing the shift parameters and squash parameters further includes:
[0140] For each subband in each group of bins:
[0141] Create a smoothed level parameter weighted histogram on the translation parameters;
[0142] Create a smoothed level parameter weighted first phase difference histogram on a first phase difference parameter, wherein the first phase difference parameter has a first range;
[0143] Create a smoothed level parameter weighted second phase difference histogram on a second phase difference parameter, wherein the second phase difference parameter has a second range different from the first range;
[0144] Detect a translation peak in the smoothed translation histogram;
[0145] Determine the translation peak width;
[0146] Determine the translation median;
[0147] Detect a first phase difference peak in the smoothed first phase difference histogram;
[0148] Determine the first phase difference peak width;
[0149] Determine the first phase difference median;
[0150] Detect a second phase difference peak in the smoothed second phase difference histogram;
[0151] Determine the second phase difference peak width; and
[0152] Determine the second phase difference median,
[0153] Wherein, the shift parameter includes the translation intermediate value and the first phase difference intermediate value or the second phase difference intermediate value, and the squeeze parameter includes the translation peak width and the first phase difference peak width or the second phase difference peak width.
[0154] EEE4. The method according to EEE 3, further comprising: determining which one of the first phase difference peak width and the second phase difference peak width is narrower, wherein the shift parameter includes the translation intermediate value and the one with the narrower peak among the first phase difference intermediate value or the second phase difference intermediate value, and the squeeze parameter includes the translation peak width and the narrower one among the first phase difference peak width or the second phase difference peak width.
[0155] EEE5. The method according to any one of EEE 1 to 4, further comprising:
[0156] Using the one or more processors to transform the modified time-frequency bins into a plurality of time-domain audio source signals.
[0157] EEE6. The method according to any one of EEE 1 to 5, wherein the spatial parameter includes the translation and phase difference of each of the time-frequency bins.
[0158] EEE7. The method according to any one of EEE 1 to 6, wherein the soft mask value is obtained from a look-up table or function of a spatial level filtering (SLF) system trained for a central translation target source.
[0159] EEE8. The method according to any one of EEE 1 to 7, wherein transforming one or more frames of a stereo time-domain audio signal into a frequency-domain signal includes applying a short-time Fourier transform (STFT) to the stereo time-domain audio signal.
[0160] EEE9. The method according to any one of EEE 1 to 8, wherein grouping a plurality of frequency bins into octave subbands or approximate octave subbands.
[0161] EEE10. The method according to any one of EEE 1 to 9, wherein the spatial parameter includes a translation parameter and a phase difference parameter for each time-frequency bin, and calculating the shift parameter and the squeeze parameter further includes:
[0162] Assembling consecutive frames of the time-frequency bins into chunks, each chunk including a plurality of subbands;
[0163] For each subband in each chunk:
[0164] Creating a smoothed level parameter weighted histogram on the translation parameter;
[0165] Create a smoothed level parameter weighted first phase difference histogram on a first phase difference parameter, where the first phase difference parameter has a first range;
[0166] Create a smoothed level parameter weighted second phase difference histogram on a second phase difference parameter, where the second phase difference parameter has a second range different from the first range;
[0167] Detect a translation peak in the smoothed translation histogram;
[0168] Determine the translation peak width;
[0169] Determine the translation median;
[0170] Detect a first phase difference peak in the smoothed first phase difference histogram;
[0171] Determine the first phase difference peak width;
[0172] Determine the first phase difference median;
[0173] Detect a second phase difference peak in the smoothed second phase difference histogram;
[0174] Determine the second phase difference peak width; and
[0175] Determine the second phase difference median,
[0176] where the shift parameter includes the translation median and the first phase difference median or the second phase difference median, and the squeeze parameter includes the translation peak width and the first phase difference peak width or the second phase difference peak width.
[0177] EEE11. The method according to EEE 10, further comprising: determining which of the first phase difference peak width and the second phase difference peak width is narrower, where the shift parameter includes the translation median and the first phase difference median or the second phase difference median having the narrower peak, and the squeeze parameter includes the translation peak width and the first phase difference peak width or the second phase difference peak width that is narrower.
[0178] EEE12. The method according to EEE 10 or 11, where the first range is from -π to π radians, and the second range is from 0 to 2π radians.
[0179] EEE13. A method as described in any one of EEEs 10 to 12, wherein the translation histogram and the phase difference histogram created for the previous block and the subsequent block are used to temporally smooth the translation histogram and the first phase histogram and the second phase histogram; or weighted data in the previous block and the subsequent block are collected and then directly used to form the histogram.
[0180] EEE14. A method as described in any one of EEEs 10 to 13, wherein the translation peak width captures at least forty percent of the total energy in the translation histogram, and the first phase difference peak width and the second phase difference peak width each capture at least eighty percent of the total energy in their respective histograms.
[0181] EEE15. The method of any one of EEEs 10 to 14, wherein the shift parameter and the squeeze parameter of each subband in each chunk are converted to be present for each frame of the one or more frames. EEE16. The method of any one of EEEs 10 to 14, wherein the shift parameter and the squeeze parameter of each subband in each chunk are converted to be present for each frame of the one or more frames.
[0182] EEE16. A method as described in any one of EEEs 10 to 15, wherein the translation shift parameter and the squeeze parameter are converted to exist for each frame using linear interpolation, and the first phase difference shift parameter or the second phase difference shift parameter is converted to exist for each frame using zero-order hold.
[0183] EEE17. The method as described in any one of EEEs 10 to 16 further includes: determining a single shifted median value and a single shifted peak width value per unit time for the one or more subbands in the one or more chunks. EEE17.
[0184] EEE18. The method as described in any one of EEEs 10 to 17, wherein the soft mask value is smoothed in time and frequency.
[0185] EEE19. A device comprising:
[0186] one or more processors;
[0187] A memory storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform any one of the aforementioned methods described in EEE 1 to 18.
[0188] EEE20. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform any of the aforementioned methods as described in EEE 1 to 18.
Claims
1. A method for audio signal processing, comprising: Using one or more processors to transform one or more frames of a stereo time-domain audio signal into a time-frequency domain representation including a plurality of time-frequency bins, wherein a frequency domain of the time-frequency domain representation includes a plurality of frequency bins, and the plurality of frequency bins are grouped into a plurality of subbands; For each of the plurality of time-frequency bins: Using the one or more processors to calculate a spatial parameter and a level of the time-frequency bin; Using the one or more processors to modify the spatial parameter using a shift parameter and a squeeze parameter; Using the one or more processors to obtain a soft mask value for each frequency bin using the modified spatial parameter, the level, and subband information; and Using the one or more processors to apply the soft mask value to the time-frequency bin to generate a modified time-frequency bin of the estimated audio source, Wherein the spatial parameter includes a translation parameter and a phase difference parameter for each of the time-frequency bins, and wherein the method further comprises, for each of the plurality of subbands: Determining a statistical distribution of the translation parameter and a statistical distribution of the phase difference parameter; Determining the shift parameter as the translation parameter and the phase difference parameter corresponding to peaks of the respective statistical distributions of the translation parameter and the phase difference parameter; and Determining the squeeze parameter as a width around the peaks of the respective statistical distributions of the translation parameter and the phase difference parameter to capture a predetermined amount of audio energy.
2. The method according to claim 1, wherein, The predetermined amount of audio energy is at least forty percent of the total energy in the statistical distribution of the translation parameter and at least eighty percent of the total energy in the statistical distribution of the phase difference parameter.
3. The method according to claim 1 or 2, wherein, Determining the statistical distribution of the translation parameter further comprises: Creating a smoothed level parameter weighted translation histogram on the translation parameter; Wherein determining the statistical distribution of the phase difference parameter further comprises: Creating a smoothed level parameter weighted first phase difference histogram on a first phase difference parameter, wherein the first phase difference parameter has a first range; Creating a smoothed level parameter weighted second phase difference histogram on a second phase difference parameter, wherein the second phase difference parameter has a second range different from the first range; Wherein determining the translation parameter corresponding to the peak of the statistical distribution of the translation parameter and the width around the peak of the statistical distribution of the translation parameter further comprises: Detecting a translation peak in the smoothed level parameter weighted translation histogram; Determining a translation peak width; Determining a translation median; and Wherein determining the phase difference parameter corresponding to the peak of the statistical distribution of the phase difference parameter and the width around the peak of the statistical distribution of the phase difference parameter further comprises: Detecting a first phase difference peak in the smoothed level parameter weighted first phase difference histogram; Determining a first phase difference peak width; Determining a first phase difference median; Detecting a second phase difference peak in the smoothed level parameter weighted second phase difference histogram; Determining a second phase difference peak width; and Determine the second phase difference median value, wherein the shift parameter includes the translation median value and the first phase difference median value or the second phase difference median value, and the squeeze parameter includes the translation peak width and the first phase difference peak width or the second phase difference peak width.
4. The method according to claim 3, further comprising: Determine which of the first phase difference peak width and the second phase difference peak width is narrower, wherein the shift parameter includes the translation median value and the phase difference median value with the narrower peak among the first phase difference median value or the second phase difference median value, and the squeeze parameter includes the translation peak width and the phase difference peak width that is narrower among the first phase difference peak width or the second phase difference peak width.
5. The method according to claim 1 or 2, further comprising: Using the one or more processors to transform the modified time-frequency slice into a plurality of time-domain audio source signals.
6. The method according to claim 1 or 2, wherein The soft mask value is obtained from a lookup table or function of a spatial level filtering (SLF) system trained for a centered translation target source.
7. The method according to claim 1 or 2, wherein, Transforming one or more frames of a stereo time-domain audio signal into a frequency-domain signal includes applying a short-time Fourier transform (STFT) to the stereo time-domain audio signal.
8. The method according to claim 1 or 2, wherein A plurality of frequency bins are grouped into octave subbands or approximate octave subbands.
9. The method according to claim 3, wherein, The first range is from - to radians, and the second range is from 0 to radians.
10. The method according to claim 1 or 2, wherein The plurality of frames of the time-frequency slice are assembled into a plurality of chunks, each chunk including a plurality of subbands, and wherein the method is performed for each subband in each chunk.
11. The method according to claim 3, wherein, The smoothed level parameter weighted translation histogram and the smoothed level parameter weighted first phase histogram and the smoothed level parameter weighted second phase histogram are smoothed in time using the translation histogram and the phase difference histograms created for the previous chunk and the subsequent chunk; or the weighted data in the previous chunk and the subsequent chunk are collected and then directly used to form the smoothed level parameter weighted translation histogram and the smoothed level parameter weighted first phase histogram and the smoothed level parameter weighted second phase histogram.
12. The method according to claim 3, wherein, The translation peak width captures at least forty percent of the total energy in the translation histogram, and the first phase difference peak width and the second phase difference peak width each capture at least eighty percent of the total energy in their respective histograms.
13. The method according to claim 10, wherein, The shift parameter and the squeeze parameter for each subband in each chunk are converted to exist for each frame in the one or more frames.
14. The method according to claim 3, wherein, The shift parameter and the squeeze parameter are converted to exist for each frame using linear interpolation, and the first phase difference parameter or the second phase difference parameter is converted to exist for each frame using zero-order hold.
15. The method according to claim 3, further comprising determining a single translation median value and a single translation peak width value per unit time for the one or more subbands in the one or more chunks.
16. The method according to claim 1 or 2, wherein The soft mask value is smoothed in time and frequency.
17. An apparatus for audio signal processing, comprising: One or more processors; A memory that stores instructions which, when executed by the one or more processors, cause the one or more processors to perform any one of the preceding method claims 1 to 16.
18. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform any one of the methods described in the preceding claims 1 to 16.
19. A computer program product comprising a program that, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Acoustic source separation systems
CN111133511A
Source separation using a circular model
WO2014047025A1