Separate the generalized stereo background and translation source using minimal training.
By using Bayesian models and spatial level filter lookup tables, the problem of efficiently extracting and enhancing audio sources from two-channel mixing was solved, especially for dialogue sources in amplitude-shifted mixing, achieving high-quality source separation and enhancement effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-11
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies struggle to efficiently extract and enhance audio sources from two-channel mixing, especially dialogue sources using amplitude-shift mixing, and often require large amounts of training data or latency.
Using a Bayesian model and a spatial level filter (SLF) lookup table, the signal-to-noise ratio (SNR) distribution is generated by detecting the level and spatial parameters in the frequency domain representation of the audio signal and applying the power law to perform amplitude shift. The lookup table is then generated using training data to achieve source separation.
It can extract high-quality audio sources, especially amplitude-shifted dialogue sources, with almost no training data or latency, to enhance dialogue and reduce artifacts.
Smart Images

Figure CN115699171B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 038,046, filed June 11, 2020, and European Patent Application No. 20179449.2, filed June 11, 2020, which are incorporated herein by reference. Technical Field
[0003] This disclosure relates generally to audio signal processing, and more specifically to audio source separation techniques. Background Technology
[0004] Two-channel audio mixing (e.g., stereo mixing) is created by mixing multiple audio sources together. Several examples exist that aim to detect and extract individual audio sources from two-channel mixes, including, but not limited to: remixing applications that reposition audio sources within a two-channel mix; upmixing applications that position or reposition audio sources within a surround sound mix; and audio source enhancement applications that enhance certain audio sources (e.g., speech / dialogue) and add them back into a two-channel or surround sound mix. Summary of the Invention
[0005] Details of the disclosed embodiments are set forth in the accompanying drawings and description below. Other features, objects, and advantages will be apparent from this specification, the drawings, and the claims.
[0006] In one embodiment, a method includes: using one or more processors to obtain a frequency domain representation of a first sample set from multiple target source level distributions and spatial distributions in multiple sub-frequency bands; using the one or more processors to obtain a frequency domain representation of a second sample set from multiple background level distributions and spatial distributions in the multiple sub-frequency bands; using the one or more processors to add the first sample set and the second sample set to create a combined sample set; using the one or more processors, for each sub-band in the multiple sub-frequency bands, detecting level parameters and spatial parameters of each sample in the combined sample set; and within each sub-band in the multiple sub-frequency bands, transmitting the detected level parameters and spatial parameters relative to the target source. The level and spatial distributions corresponding to the background are weighted; using the one or more processors, the weighted level parameters, spatial parameters, and signal-to-noise ratio (SNR) of each sample in the combined sample set within the plurality of sub-bands are stored in a table; and using the one or more processors, the table is re-indexed by the sub-bands and the weighted level and spatial parameters, such that the table includes the target percentile SNR of the sub-bands and the weighted level and spatial parameters, and for a given input of the sub-bands and the quantized detected spatial and level parameters, the estimated SNR associated with the sub-bands and the quantized detected spatial and level parameters is obtained from the table.
[0007] In an embodiment, the method further includes smoothing data based on one or more of the detected levels, one or more of the spatial parameters, or the sub-band index.
[0008] In this embodiment, the frequency domain representation is a short-time Fourier transform (STFT) domain representation.
[0009] In an embodiment, the spatial parameters include the translation and phase difference between the two channels of the mixed audio signal.
[0010] In this embodiment, the target source is amplitude-shifted using the power law.
[0011] In this embodiment, the target percentile SNR is the 25th percentile.
[0012] In one embodiment, a method includes: using one or more processors to transform one or more frames of a two-channel time-domain audio signal into a time-frequency domain representation comprising a plurality of time-frequency slices, wherein the frequency domain of the time-frequency domain representation comprises a plurality of frequency bins, the plurality of frequency bins being grouped into a plurality of sub-bands; for each time-frequency slice: using the one or more processors to calculate spatial parameters and levels of the time-frequency slice; using the one or more processors to generate a percentile signal-to-noise ratio (SNR) for each frequency bin in the time-frequency slice; using the one or more processors to generate a fractional value for the bin based on the SNR of the bin; and using the one or more processors to apply the fractional value of the bin of the time-frequency slice to generate a modified time-frequency slice of an estimated audio source.
[0013] In one embodiment, multiple frames of a time-frequency slice are grouped into multiple blocks, each block comprising multiple sub-bands. The method includes: for each sub-band in each block: using the one or more processors to calculate the spatial parameters and levels of each time-frequency slice in the block; using the one or more processors to generate a percentile signal-to-noise ratio (SNR) for each frequency cell in the time-frequency slice; using the one or more processors to generate a fractional value for the cell based on the SNR of the cell; and using the one or more processors to apply the fractional value of the cell in the time-frequency slice to generate a modified time-frequency slice of the estimated audio source.
[0014] In one embodiment, the method includes using the one or more processors to transform the modified time-frequency slice into a plurality of time-domain audio source signals.
[0015] In an embodiment, the spatial parameters include translation and phase difference between channels for each of the time-frequency slices.
[0016] In this embodiment, the score is obtained from a lookup table or function of a spatial level filtering (SLF) system trained for a translation target source.
[0017] In an embodiment, transforming one or more frames of a two-channel time-domain audio signal into a frequency-domain signal includes applying a short-time frequency transform (STFT) to the two-channel time-domain audio signal.
[0018] In this embodiment, multiple frequency modules are grouped into octave subbands or approximately octave subbands.
[0019] The specific embodiments disclosed herein offer one or more of the following advantages. The disclosed embodiments allow for the extraction of a target source (source separation) from a recording consisting of a source plus some background. More specifically, the disclosed embodiments allow for the extraction of sources that use (fully or mostly) amplitude panning blending, the most common way to blend dialogue in TV and film. The ability to extract such sources enables dialogue enhancement (extracting and then enhancing the dialogue in the blend) or upmixing. Additionally, the ability to extract high-quality source estimates with minimal training data or latency distinguishes it from most other source separation methods. Attached Figure Description
[0020] Various embodiments are illustrated in the accompanying drawings, which are referenced below, in the form of block diagrams, flowcharts, and other figures. Each block in a flowchart or block diagram may represent a module, program, or code section containing one or more executable instructions for performing a specified logical function. Although these blocks are illustrated in a specific order of performing these method steps, they may not necessarily be performed in a strictly illustrated order. For example, depending on the nature of the respective operations, these blocks may be performed in reverse order or simultaneously. It should also be noted that each block and combination thereof in the block diagrams and / or flowcharts may be implemented by a software-based or hardware-based dedicated system for performing the specified function / operation, or by a combination of dedicated hardware and computer instructions.
[0021] Figure 1 The illustration depicts a signal model for source separation that illustrates temporal mixing according to an embodiment.
[0022] Figure 2 This is a block diagram of a system according to an embodiment for generating a trained spatial level filter (SLF) lookup table to extract translation sources.
[0023] Figure 3 It is a visual depiction of the inputs and outputs of an SLF lookup table trained to extract translation sources according to an embodiment.
[0024] Figure 4 This is a block diagram of a system according to an embodiment for detecting and extracting spatially identifiable sub-band audio sources from a two-channel mix using a trained SLF to extract translation sources.
[0025] Figure 5 This is a flowchart of a process for generating an SLF lookup table trained to extract translation sources, according to an embodiment.
[0026] Figure 6 This is a flowchart illustrating the process of using an SLF trained to extract translation sources from a two-channel mix to detect and extract spatially identifiable sub-band audio sources, according to an embodiment.
[0027] Figure 7 An implementation reference according to an embodiment is shown. Figures 1 to 6 A block diagram of the device architecture describing the system and process.
[0028] The same reference numerals used in all the figures indicate the same elements. Detailed Implementation
[0029] Signal Models and Assumptions
[0030] Figure 1 The illustration depicts a signal model 100 for source separation using temporal mixing according to an embodiment. Signal model 100 assumes that the target source s1 and background b are mixed into two channels in the basic temporal domain according to the context, hereinafter referred to as the "left channel" (x1 or X). L ) and "right channel" (x2 or X R These two channels are input to the source separation system 101, which estimates...
[0031] Assume that the target source s1 is amplitude-shifted using the power law. Since other translation laws can be converted to the power law, using the power law in signal model 100 is not restrictive. Under the power law shift, the source s1 mixed into the left / right (L / R) channels is described as follows:
[0032] x1 = cos(Θ1)s1, [1]
[0034] x2 = sin(Θ1)s1, [2]
[0036] Here, Θ1 ranges from 0 (shifted to the leftmost source) to π / 2 (shifted to the rightmost source). This can be represented in the Short-Time Fourier Transform (STFT) domain as...
[0037] X L =cos(Θ1)S1, [3]
[0039] X R =sin(Θ1)S1. [4]
[0041] Continuing in the STFT domain, adding background B to each channel is represented as follows:
[0042] X L =cos(Θ1)S1+cos(Θ) B )|B|e j∠B , [5]
[0044] [6]
[0046] Background B includes the additional parameter ∠B and These parameters describe the phase difference between the left channel phases of S1 and B in the STFT space, and the inter-channel phase difference between the left and right channel phases of B. It should be noted that equations [5] and [6] do not need to include... The parameters are assumed to be zero because the inter-channel phase difference of the translation source is defined as zero. It is assumed that the target S1 and background B do not share a specific phase relationship in the STFT space; therefore, the distribution of ∠B is modeled as a uniform distribution.
[0047] There is a key spatial difference between the target source and the background. Spatially, Θ1 is considered a specific single value (the "translation parameter" of the target source S1), but Θ B and Φ B Each has a certain statistical distribution, which allows the use of statistical models (e.g., Bayesian models) to perform source separation.
[0048] Then, to recap, the "target source" is assumed to have been translated, meaning it can be characterized by Θ1. We assume the inter-channel phase difference of the target source is zero. Its audio level also exhibits a distribution L. S =|S1|, which is assumed to be known at least in the approximate octave subband. It is assumed that the spatial information is entirely specified by the translation parameters of the source.
[0049] Background B is characterized in Θ B Furthermore, there is also a phase difference between the audio channels. The distribution exists on the background level L. B =|B|, which is assumed to be known at least in the approximate octave subband.
[0050] For the purposes of this model, the source and background can only be modeled at points in time when both are assumed to be "active". In this sense, for this purpose, it is assumed that the source and background are always "on" or "off", while separation should assume that both the target source and background are "on". It can be seen that if the target source is active but the background is inactive, the extraction is still nearly perfect. If the target source and translation parameters are unknown, they can be estimated using techniques known to those skilled in the art. In some cases, such as most music, there may be harmonic relationships between the target source and the background. In signal model 100, such relationships are not modeled separately; it is assumed that the distribution includes a certain degree of harmonic overlap, which is suitable for the given application.
[0051] Training process
[0052] Figure 2This is a block diagram of a system 200 for generating an SLF lookup table trained to extract translation sources, according to an embodiment. An SLF is a system trained to extract a target source having a given level distribution and specified spatial parameters from a mixture comprising a background having a given level distribution and spatial parameters.
[0053] System 200 includes a target source parameter database 201, a target source distribution sampler 202, a transform 203, a parameter detector 204, a reindexer 205, a target SNR selector 206, a trained SLF lookup table 207, a background parameter database 208, a background distribution sampler 209, and a transform 210. Distribution samplers 202 and 209 and transforms 203 and 210 are... Figure 2 While they are shown as separate blocks, samplers 202, 209 and transforms 203, 210 can actually be combined into a single module (e.g., a software module) that operates on the target source database and background databases 201, 208.
[0054] The training procedure implemented by System 200 aims to create a Bayesian model that predicts the relative fraction of energy belonging to the target source for each STFT domain bin or tile, given a two-channel input (e.g., L / R stereo input). To aid in achieving this goal, four parameters are used that are detectable for two-channel inputs in the STFT domain.
[0055] The first parameter is b, which represents the approximate octave subband. This parameter is obtained through a trivial mapping from a given frequency band ω to its corresponding subband b. An example of a subband boundary is given below.
[0056] The second parameter is the detected "translation" for each (ω,t) slice, defined as:
[0057] Θ(ω, t) = arctan(|X) R (ω,t)| / |X L (ω,t)|), [7]
[0059] "Full left" is 0, while "full right" is π / 2.
[0060] The third parameter is the detected "phase difference" for each slice. This is defined as:
[0061] [8]
[0063] Its range is from -π to π, where 0 means that the phase detected in both channels is the same.
[0064] The fourth parameter is the detected "level" of each chip, defined as:
[0065] U(ω, t) = 10 * log 10 (|X R (ω,t)| 2 +|X L (ω,t)| 2 ), [9]
[0067] This is the Pythagorean amplitude for two channels. It can be considered a single-amplitude spectrogram.
[0068] Each frequency bin ω is understood to represent a specific frequency. However, data can also be grouped into subbands, which are collections of consecutive bins, where each frequency bin ω belongs to a subband. Grouping data into subbands is particularly useful for certain estimation tasks performed in the system. In the embodiments, octave subbands or approximately octave subbands are used, but other subband definitions may also be used. Some examples of banding include defining band edges as follows, where values are listed in Hz:
[0069] [0,400,800,1600,3200,6400,13200,24000],
[0070] [0,375,750,1500,3000,6000,12000,24000], and
[0071] [0,375,750,1500,2625,4125,6375,10125,15375,24000].
[0072] It should be noted that if the definition of "octave" is strictly followed, there could be an infinite number of such frequency bands, where the lowest frequency band width is close to infinitesimal. Therefore, some selection is required to allow for a finite number of subbands. In this embodiment, the lowest frequency band is chosen to be equal in size to the second frequency band, but other conventions may be used in other embodiments. Throughout this document, the terms "subband" and "frequency band" are used interchangeably.
[0073] To understand how to build a Bayesian system based on these four parameters, let's first review the Bayesian rules:
[0074] p(A|B)=p(B|A)p(A) / p(B).
[10]
[0076] In this context, the goal of the training process is to allow estimation of the SNR distribution for each spectrogram image, given some observations. The observations b, Θ, and... are described above. U. Bayes' rule is given by the following formula:
[0077]
[11]
[0079] The goal now is to train a Bayesian system that can produce all quantities on the right side of equation
[11] , so that quantities on the left side of equation
[11] can be estimated. To do this, p(SNR) is estimated by considering the level distribution over the target source in the background.
[0080] When mixing targets and backgrounds with various SNR values, the parameters in each frequency band b are considered. To estimate the distribution The conditional probability. The procedure for generating this data involves: using distributed samplers 202 and 209 to generate numerous data samples from the target source and background databases 201 and 208, respectively, by sampling from their known or assumed spatial and level distributions. Transformers 203 and 210 create STFT thresholds using the properties of the samples.
[0081] To recap, the target source is assumed to have specific translation parameters; therefore, the training procedure described herein explicitly specifies the translation parameters of the target source that are desired to be extracted later. The example embodiment described herein assumes that the target source has Θ1 = π / 4, which corresponds to a center-translated source. When generating training data, it is assumed that a random phase relationship exists between the target and the background as described above. In practice, this can be implemented by setting one phase value to zero and another phase value to various samples on the unit circle.
[0082] To create training data, the frequency domain representations output by transformation modules 203 and 210 are summed together (e.g., ...). Figure 1 (As shown in signal model 100) to create a frequency domain representation of the combination. It should be noted that during Bayesian training, the combinations of target and background data items will be very large, and such a large number of combinations will have a relatively small target-to-background ratio of equal quantization.
[0083] To efficiently utilize this reality, the training process creates uniformly sampled datasets separately for the following: target-background SNR (0 to 37 dB, although a larger range can be selected), phase difference between target and background (0 to 2π), background Θ (0 to π / 2), and background The amplitude (from 0 to π). For all possible combinations of this data, the training process calculates the detected amplitude. The values are then stored in storeThetaHat, storePhiHat, and storeUdBHat, respectively. It should be noted that this calculation still does not consider the specific spatial and level distributions on each of the target and background. They are simply mapped from all potential combinations of relevant input attributes to the detected Θ, And a lookup table for U. Using these tables will improve the efficiency of the later training process.
[0084] Next, we combine specific spatial and level data of the target and background. To recap, the goal is to obtain... In fact, The distribution of each variable in the model can be represented by a quantized probability density function (pdf), and the SNR can also be quantized. In the embodiment, for Amplitude (0 to π) uses 51 quantization levels, Θ (0 to π / 2) uses 51 quantization levels, U (example range 0 to 127 dB) uses 1 dB increments, and DNR (example range -40 dB to +60 dB) uses 1 dB increments. Given this quantization, information... It can be stored in a multidimensional array "storePopularity" of the following size: 7 frequency bands × 101 trained SNRs (-40 to 60) × 51 Θ bins × 51 The array is divided into 128 dB levels (e.g., 0 to 127). Then, for each item, the values stored in the array represent the probability (or similarly, "popularity") of a particular combination relative to other combinations in the array. For example, array element (4,49,26,26,90) represents the Θ value for the DNR (49th value) at band 4 and +8 dB, the detected π / 4 (26th value), and the π / 2 (26th value). The amplitude value and the level U of 89dB (the 90th value) indicate its "popularity".
[0085] In order to obtain The training process exhaustively (or by sampling) iterates through all possible combinations of spatial and level data from the target and source. At this point, when specific SNR, phase difference, background Θ, and background parameters are observed in the training data... At that time, Θ is retrieved using the data previously stored in storeThetaHat, storePhiHat, and storeUdBHat respectively. And U, to reduce training computation. This lookup can also be called "parameter detection" and is by Figure 2Block 204 is executed. Importantly, the popularity of each such combination is also used, indicated by spatial and level distribution values for the target and background; these popularity values are weighted onto the `storePopularity` array, and in doing so, p(SNR) is combined according to expectation. The aforementioned `storePopularity` array is created by looping through all such combinations and recording their popularity. This array may be sparse or noisy, and therefore should be smoothed using techniques familiar to those skilled in the art. An example technique is smoothing along one or more dimensions of the table.
[0086] At this stage, the data required for Bayesian analysis is obtained, but not in the desired lookup table or function format. The final step in the training process is to extract data from a storePopularity of the following size. Get available 7 frequency bands, 101 trained SNRs (-40 to 60), 51 Θ bins, 51 The chamber is 128 dB level (e.g., 0 to 128). To understand this, you need to... Let's review the correspondence. It can be equivalently represented as or equivalent These five metrics are the same as those in storePopularity.
[0087] This reindexing or remapping is caused by Figure 2 Blocks 205 and 206 in the code were completed. It should be remembered that the expected... Not a single set of values, but some detected values for each frequency band b. Given a set of SNR distributions, decisions are made on how to concisely describe these distributions to maintain control over the size of the representation; typical ways to do this include using the mean, median, or other parameters. Considering the practical application requirements for designing this system, the 25th and 50th percentiles of each SNR distribution are used in this embodiment.
[0088] In order to obtain The training process works to perform reindexing (block 205) and target SNR selection (block 206). The fundamental goal is to assemble and characterize the detected frequencies from storePopularity with respect to a given frequency band b. All SNR data corresponding to the triplet. Since the frequency bands are treated as independent, this is equivalent to considering the goal as finding each of N individual exercises for each of N frequency bands. Block 205 performs this task. This block iterates through each frequency band and checks the distribution level of each sample of the following variables: detected Θ, detected Θ, and detected Θ. The detected level. For each such value, it is derived from storePopularity (which consists of all SNRs and their values at a given detected Θ). A buffer is created based on the popularity of specific combinations of U values. More specifically, the buffer is a subset of storePopularity, as follows: storePopularitySmoothed(band index, (all data), Θ index, (Index, U index). The next block 206 is a buffer of analyzed values, and in this embodiment, the 25th and 50th percentile values in the trained SLF lookup table (207) are detected and recorded. Specifically, these values are recorded in new arrays, percentile25SNRvalues and percentile50SNRvalues, respectively, with each array indexed as (band index, detected Θ index, detected Θ index). Index, detected U index), this is for The desired representation.
[0089] Due to the inherent sparsity of the training data, some buffers from which percentile SNRs are calculated may have too few data points to produce reliable percentile SNR values. To address this issue, two example techniques can be used, but other techniques are also possible. One technique involves sharing data from adjacent frequency bands, Θ values, etc., before calculating the percentile SNR. Data on the U value or U value (prioritizing bandwidth and U level sharing). Another technique is to first calculate the percentile SNR even from sparse data, and then replace or smooth the percentile SNR values with SNR values from adjacent U values (or bandwidth if needed) if they appear unstable.
[0090] At this stage, the reindexing was completed and the application of the training system was described. The system has reindexed tables such that the indexes of these tables represent Θ, The quantization value of Θ and U, and the index b of the frequency band in question. To obtain the soft mask value using such a table, the function takes Θ, U, and the index b of the frequency band in question as input. The Θ and U values were quantized into 51, 51, and 128 levels, respectively. From the detected Θ, The conversion from U values to their indices is trivial and follows the same quantization used when performing the quantization distribution described above. The function accesses the values in the table corresponding to these quantization index levels (and the indices of the frequency band b corresponding to the frequency bin ω in question).
[0091] It should be noted that although percentile25SNRvalues and percentile50SNRvalues are obtained from a table with a specific index in this case, the SNR values can actually be obtained by using any (not necessarily quantized) Θ. A more general function of the values of U and b is given. In fact, seeking values from Θ, The functions for obtaining softmask values using U and b do not require access to a table to output the softmask values. They can directly compute the softmask values using curves or general functions (including trained neural networks) that approximate and / or interpolate values from the table. This is achieved by checking... Figure 3 (Representation of the 25th percentile SNR system) It is readily apparent that the curve can fit the data represented in the table. In embodiments using the table, the table is understood not as a restrictive method for obtaining soft mask values, but rather as a recommended and efficient method for doing so. Neural networks that approximate or interpolate the table can be constructed using techniques familiar to those skilled in the art.
[0092] Figure 3 This is a visual depiction of the input and output of an SLF lookup table trained to extract translation sources, according to an embodiment. More specifically, Figure 3 The reference shows Figure 2 The description describes a visual representation of a trained 25th percentile four-dimensional (4D) SLF lookup table of a center-translated target source. The SLF lookup table is large but also repetitive. Techniques familiar to those skilled in the art can be used to reduce the lookup time and memory required to store information in the table (e.g., entropy encoding) or to convert the information in the table into a continuous function as described above.
[0093] As mentioned above, Figure 3 The visual representation is 4D. The four input variables are the left and right Θ axes and the in and out axes of each subplot. The axis, as well as the vertical (subband b) subgraph index and the horizontal (level U) subgraph index. It should be noted that, for practical reasons, the horizontal subgraph dimension (level U) does not depict all levels stored in the SLF lookup table; doing so would require approximately 128 subgraphs because 1dB increments are used within the 128dB range of the table. In practice, finer or coarser increments could be used for higher precision or higher lookup efficiency, respectively. When viewing... Figure 3 It should be noted that there are many "not shown" subplots from left to right.
[0094] The output variable of the SLF lookup table is a soft mask value between 0 and 1 (inclusive), displayed on the vertical axis of each subgraph. The soft mask value represents the fraction of the corresponding input STFT that should be passed to the output. Since each STFT has one (four-dimensional) input, each STFT also has one output. The result of applying the SLF table / function is a representation of the STFT size, composed of values between 0 and 1.
[0095] As mentioned above, soft mask values generated by percentile25SNRvalues or percentile50SNRvalues can be used, although other percentiles can also be used. Generally, using percentile25SNRvalues yields a source separation solution that strikes a balance between including some background and introducing some artifacts in source estimation. Using percentile50SNRvalues provides a solution with fewer artifacts but more background. The application of soft mask parameters in… Figure 4 It is shown in block 404.
[0096] In this embodiment, techniques familiar to those skilled in the art are used to smooth the soft mask values and / or signal values in terms of time and frequency. Assuming a 4096-point FFT, a smoothing relative to frequency can be used, employing a smoother of [0.17 0.33 1.0 0.33 0.17] / total ([0.17 0.33 1.0 0.33 0.17]). For higher or lower FFT sizes, some reasonable scaling should be performed on the smoothing range and coefficients. Assuming a jump size of 1024 samples, a smoother of approximately [0.1 0.55 1.0 0.55 0.1] / total ([0.1 0.55 1.0 0.55 0.1]) relative to time can be used. If the jump size or frame length changes, the smoothing can be adjusted accordingly.
[0097] Example Application
[0098] Figure 4 This is a block diagram of a system 400 for detecting and extracting spatially identifiable sub-band audio sources from a two-channel mix using an SLF, according to an embodiment. System 400 includes a transform 401, a parameter calculator 402, a table lookup 403, a soft mask applicator 404, and an inverse transform 405. The table lookup 403 operates on a database 406, which stores information as referenced. Figure 2 The described SLF lookup table is trained to detect translation sources. For this example application, it is assumed that the target source to be extracted has a known translation parameter, or that the detection of such a parameter is performed using any number of techniques known to those skilled in the art. One example technique for detecting the translation parameter is peak picking from a level-weighted histogram of the Θ values.
[0099] refer to Figure 4 Transformer 401 is applied to the two-channel input signal (e.g., a stereo mix signal). In this embodiment, system 400 uses STFT parameters, including window type and jump size, which are known to those skilled in the art to be relatively optimal for source separation problems. However, other STFT parameters may also be used. Based on the STFT representation, parameter calculator 402 calculates the parameters for each octave subband b. The values are used by table lookup 403 to perform a table lookup against the SLF lookup table stored in database 406. The table lookup generates a percentile SNR (e.g., the 25th percentile) for each STFT chip or cell. Based on the SNR, system 400 calculates the fraction of the STFT to be used as the input to the Bayesian estimation output. For example, if the estimated percentile SNR is 0 dB, the fraction passing through the input will be 0.5 or 50%, since the estimated target source and background have the same level U. The general formula follows the assumptions of the Wiener filter and is: input fraction = 10^(SNR / 20) / (10^(SNR / 20)+1). Next, soft mask applicator 404 multiplies the input STFT for each channel by the fractional value between 0 and 1 for each STFT chip. Inverse transform 405 then inverse transforms the STFT representation to obtain a two-channel time-domain signal representing the estimated target source.
[0100] Although the foregoing example embodiments use an STFT time-frequency representation (e.g., a chip), any suitable time-frequency representation may be used.
[0101] Although the example source separation application described above uses an SLF lookup table, other embodiments may use SLF functions instead of lookup tables.
[0102] Example process
[0103] Figure 5 This is a flowchart of a process 500 for generating a trained SLF lookup table to extract translation sources, according to an embodiment. Process 500 can be derived from, for example, references... Figure 7 The described device architecture is implemented using 700.
[0104] Process 500 begins by obtaining frequency domain representations of samples from the target source level distribution and spatial distribution in the subband (501), obtaining frequency domain representations of samples from (multiple) background level distributions and spatial distributions (502), and combining the first sample set with the second sample set to create a combined sample set (503), as referenced. Figure 2 As described.
[0105] Process 500 continues: For each sub-band, the level and spatial parameters of each sample in the combined sample set are detected (504), and within each sub-band, the detected level and spatial parameters are weighted by their corresponding level and spatial distributions relative to the target source and (multiple) backgrounds (505), as referenced. Figure 2 As described.
[0106] Process 500 continues: For each sample in the combined sample set, the weighted level parameters and spatial parameters, along with the SNR and subband, are stored in table (506), as referenced. Figure 2 and Figure 3 As described.
[0107] Process 500 continues: the stored parameters and SNR are reindexed so that the table includes the target percentile SNR of the subband and the weighted level and spatial parameters, and for a given input of the subband and the quantized detected spatial and level parameters, the estimated SNR of the subband and associated with the quantized detected spatial and level parameters can be obtained from the table (507), as referenced. Figure 2 and Figure 3 As described. Then, the SLF lookup table is stored (in a database) for use by source separation applications, such as [reference needed]. Figure 4 and 6 (As described).
[0108] Figure 6 This is a flowchart of process 600, according to an embodiment, using a trained SLF to detect translation sources to detect and extract spatially identifiable sub-band audio sources from a two-channel mix. Process 600 can be derived from, for example, references... Figure 7 The described device architecture is implemented using 700.
[0109] Process 600 may begin by transforming a two-channel time-domain audio signal into a frequency-domain representation that includes time-frequency slices having multiple frequency bins grouped into subbands (601). For example, an STFT representation of each channel of the two-channel time-domain audio signal can be created using STFTs.
[0110] Process 600 continues: calculate the spatial and level parameters for each frequency cell (602). For example, the parameters can be calculated using equations [7]-[9].
[0111] Process 600 continues: for each time-frequency slice, a percentile SNR (603) for each frequency block in the slice is generated; a fractional value for that frequency block is generated based on the SNR of the frequency block (604); and the fractional values are applied to the corresponding frequency blocks in the time-frequency slice to generate a modified time-frequency slice of the estimated audio source (605), as referenced. Figure 4 As described. The SLF lookup table / function is trained to detect translation sources, as referenced. Figure 2 and Figure 5 As described herein, the aforementioned scores are also referred to herein as soft mask values, and are real numbers between 0 and 1 (inclusive), representing the fraction of the corresponding input STFT passed to the output. The result of applying the SLF table / function is a representation of the STFT size, which consists of values between 0 and 1. In embodiments, techniques familiar to those skilled in the art are used to smooth the soft mask values and / or SNR values in time and frequency.
[0112] The process 600 may optionally continue by: inversely transforming the estimated time-frequency slice of the target audio source into a two-channel time-domain estimate of the target audio source (606), as referenced. Figure 4 As described. It should be noted that some embodiments may utilize the time-frequency slice of the estimated audio source in the frequency domain, while other embodiments may utilize a two-channel time-domain estimate of the estimated audio source.
[0113] Example device architecture
[0114] Figure 7 An implementation reference according to an embodiment is shown. Figures 1 to 6 A block diagram of the device architecture 700 describing the system and process.
[0115] Device architecture 700 can be used in any computer or electronic device capable of performing the above mathematical calculations.
[0116] In the example shown, device architecture 700 includes one or more processors 701 (e.g., CPU, DSP chip, ASIC), one or more input devices 702 (e.g., keyboard, mouse, touch surface), one or more output devices (e.g., LED / LCD display), memory 704 (e.g., RAM, ROM, flash memory), and an audio subsystem 706 coupled to a speaker 706 (e.g., media player, audio amplifier, and support circuitry). Each of these components is coupled to one or more buses 707 (e.g., system, power supply, peripherals, etc.). In embodiments, the features and processes described herein can be implemented as software instructions stored in memory 704 or any other computer-readable medium and executed by one or more processors 701. Other architectures with more or fewer components are also possible, such as those using a hybrid of software and hardware to implement the functions and processes described herein.
[0117] While this document contains numerous details of specific implementations, these details should not be construed as limiting the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even initially stated so, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof. The logical flow depicted in the drawings does not require the specific order or ordered sequence shown to achieve the desired result. Additionally, other steps may be provided from the described flow, or steps may be deleted, and other components may be added to or removed from the described system. Therefore, other implementations are within the scope of the following claims.
Claims
1. A method for signal processing, comprising: One or more processors are used to sample the first spatial and level distribution of the sound source to obtain a first sample set representing the target source; The one or more processors are used to sample the second spatial and level distribution of the sound source to obtain a second sample set representing the background source; The frequency domain representation of the first sample set in multiple sub-bands is obtained using the one or more processors; The frequency domain representation of the second sample set in the plurality of sub-frequency bands is obtained using the one or more processors; The frequency domain representation of the first sample set and the frequency domain representation of the second sample set are paired and combined using the one or more processors to create multiple training datasets; Using the one or more processors, for each of the plurality of sub-frequency bands, detect the corresponding parameter values of each training dataset in the training dataset; For each of the plurality of sub-frequency bands, the one or more processors are used to weight the influence of the popularity array of each combination of the corresponding parameter values through the first spatial and level distribution and the second spatial and level distribution; Using the one or more processors, the popularity array and the signal-to-noise ratio (SNR) values corresponding to the plurality of sub-bands are stored in a table, wherein each corresponding SNR value represents the level ratio of the target source to the background source in the corresponding training dataset; as well as Using the one or more processors, the table is reindexed based on the Bayesian relationship between the distribution of the corresponding parameter values and the corresponding signal-to-noise ratio (SNR) values, such that for a given input set of parameter values, the estimated SNR value associated with the given input set of parameter values for the plurality of sub-bands is obtained from the reindexed table via a lookup operation.
2. The method of claim 1, further comprising: Smooth the data in the popularity array across one or more dimensions of the table.
3. The method as described in any one of claims 1 or 2, wherein, The frequency domain representation is a short-time Fourier transform domain representation.
4. The method as described in any one of claims 1 or 2, wherein, The spatial parameters represented by the corresponding parameter values include the translation and phase difference between the two channels of the mixed audio signal corresponding to the plurality of training datasets.
5. The method as described in any one of claims 1 or 2, wherein, The target source is amplitude-shifted using the power law.
6. A method for signal processing, comprising: One or more frames of a two-channel time-domain audio signal are transformed into a time-frequency domain representation comprising multiple time-frequency slices using one or more processors, wherein the time-frequency domain representation comprises multiple frequency bins and the multiple frequency bins are grouped into multiple sub-bands; For each slice represented in the time-frequency domain: The one or more processors are used to calculate the set of corresponding spatial parameter values and the corresponding levels; Using the one or more processors and a lookup table, a corresponding percentile signal-to-noise ratio (SNR) value is generated for each of the plurality of sub-bands, wherein the lookup table is trained based on a Bayesian relationship between the distribution of parameter values and the corresponding SNR values, such that in response to a given input set of parameter values, the lookup table returns the estimated SNR value associated with the given input set of parameter values for the plurality of sub-bands, and wherein each of the corresponding SNR values represents the level ratio of the target source to the background source in the corresponding sub-band of the plurality of sub-bands; Using the one or more processors, a soft mask of fractional values is generated based on the corresponding percentile SNR values, wherein each fractional value represents the audio portion of a corresponding sub-band belonging to the target source in the corresponding chip represented in the time-frequency domain; and Using the one or more processors, a soft mask of the fractional value is applied to the plurality of sub-bands in the respective slice to generate a modified time-frequency slice corresponding to the target source.
7. The method of claim 6, further comprising: The modified time-frequency slice is transformed into a corresponding time-domain audio source signal using one or more processors.
8. The method as described in any one of claims 6 or 7, wherein, The transformation includes applying a short-time frequency transformation to the two-channel time-domain audio signal.
9. The method as described in any one of claims 6 or 7, wherein, The multiple frequency compartments are grouped to form octave subbands or near-octave subbands.
10. An apparatus for signal processing, comprising: One or more processors; as well as The memory stores instructions that, when executed by the one or more processors, cause the device to perform any one of the methods described in claims 1 to 9.
Citation Information
Patent Citations
Method, device and system for detecting and extracting spatially recognizable sub-band audio source
CN115715413A
Binaural noise reduction
US20120128164A1
Method of operating a hearing aid system and a hearing aid system
US20160261961A1